# Context Engineering — Chapter-22: Context validation - **Source**: /library/Context Engineering/source-file.pdf - **PDF pages**: 177–189 - **Pages without text**: none --- Context validation The complaint reached the Vila Nova Clinic front desk before nine. At 2 p.m., on one provider’s schedule, the same interval had two patients at the door: one with an appointment, and one there for a work-in, one of the extra appointments squeezed into a full schedule. One of the two waited forty minutes to be seen. The task that lands in your hands is small and clear. Scheduling an appointment in VilaSchedule cannot clash with an appointment that already exists at that time for that provider, and the check that prevents it has to go in today. You assemble the packet the way chapter 17 taught you. It opens with the convention the task might violate: business rules live in the domain and never in the controller. In the middle comes the task material: the line from the standing rules table in the living doc, verified in continuous integration (CI) three days ago, and the two files of the scheduling slice where the diff is going to happen. The request closes the packet. A little over a thousand tokens, no old conversation, no whole file without a reason. The answer comes back good, and that is what makes the day bad. The code is in the domain, the function name follows the convention, and the test comes with it. In the middle of the text, dropped in passing the way you would give someone context, this sentence appears: since the schedule accepts two appointments at the same time when one of them is a work-in, the check considers regular appointments only. You are reading fast, and the sentence sounds like domain knowledge. Clinics that take a work-in on top of a full hour are real, and VilaSchedule worked that way until June. You accept the diff. Two days later the front desk schedules another appointment on top of a work-in, and that is when you find out. The clinical coordinator removed the overlap last month: the work-in started taking its own interval at the end of the block, and no time on the clinic’s schedule has taken two appointments since. The agent’s sentence described with precision a system that stopped existing five weeks ago, and your code implemented it. The error survives every explanation the earlier chapters gave you. Context was not missing: the line that contradicts the sentence was in the packet, in the standing rules table, with the date of the check on top of it. It was not chapter 9’s dead documentation, because the document was right and a test in CI keeps it right. It was not a bad restart from chapter 18, because the session was new and had one task only. What happened is that a correct packet does not force the model to believe it. Your window held the standing rule and the contrary belief at the same time, and the contrary belief had going for it everything the model read about clinics before it met yours. Checking a belief is not validating input Context validation is checking whether what the AI states about the state of the project matches what the project says today, before you act on the answer. The word validation is already taken in architecture vocabulary, and confusing the two senses would be expensive. In the previous book in this trilogy, FOCUS Architecture (https://books.kodel.com.br/en/books/focus/), validating input is the use case’s job, the piece that book defines as the only place for business rules; a controller that validates, applies rules and picks a path is exactly the habit it takes apart, and the VilaSchedule convention that opened the packet earlier in this chapter comes from there. That kind of validation protects the system from invalid data, and it runs in production, on every request, forever. This one happens earlier, at your desk, once per statement, and what you are examining is not the patient’s data: it is the sentence the AI wrote about your project. One takes care of what the user sends; the other takes care of what the model believes. What goes into the check is a statement of state, that is, what the system does today: a standing rule, a limit, a number, the name of a field, a file or a table, the current behavior of a flow. Everything that is preference and suggestion stays out. If the agent proposes another name for the function or argues for splitting the file in two, that is project conversation, and code review settles it. Mixing the two is the shortest path to checking nothing, because anyone who tries to check every sentence of every answer gives up by Tuesday. Two ways to state what is not so Ji and colleagues, in “Survey of Hallucination in Natural Language Generation” (DOI 10.1145/3571730), published in ACM Computing Surveys in 2023, treat hallucination as generated content that does not hold up against the source offered to the model, and they separate two types that bear directly on your day. Intrinsic hallucination contradicts the material that was in the input. Extrinsic hallucination states something the input can neither confirm nor deny. The sentence about the overlap is of the first type, and it is the type that hurts more, because it contradicts the intuition you brought from Part II. You put the right information in the window and got the opposite of it back, with the same confident prose as always. The right information was one table line in the middle of the packet; the wrong information is the pattern of a whole industry, seen thousands of times during training. Where the two contradict each other, the table line has to win a dispute, and chapter 5 already showed with numbers that this dispute is not won on merit. The second type shows up all the time and goes unnoticed because it contradicts nothing. If the packet did not bring the schema of the work-ins table, the sentence “the field that marks cancellation is canceled_at ” neither contradicts nor confirms anything inside the window: it simply cannot be checked there. You read it, it looks reasonable, and from then on the field name circulates in the session as though it were a fact read in the code, which chapter 4 already described as the output of the cycle turning into the input of the cycle. The distinction is worth the small effort of memorizing it, because it changes where you look. Against intrinsic hallucination, the check is to reread your own packet; it costs seconds and it also hands you a diagnosis: if your line was there and lost, it is in the wrong position, and chapter 17 says where it goes. Against the extrinsic kind there is no check inside the window; it can only be settled outside, in the project. And when the failure comes from an incomplete packet, the first place to look is the line chapter 17 told you to leave in the state note, the one that records what opened the packet on the way back: it says in seconds whether the information never went in or went in and lost the dispute. The confidence of the answer is identical in both cases, and that is the warning I would keep: you cannot detect either one by the tone of the sentence, and that is why you need a procedure instead of a nose for it. The statement, the check and the repair The excerpt below shows the session that opened this chapter, on the path it should have taken, with each [...] marking what did not fit on this page. Reconstruction: I condensed a VilaSchedule session about the time conflict check into one representative excerpt, with the lines rewritten; the outdated statement, the check and the repair are the mechanics that repeat. ## The statement, in the middle of an answer that is right elsewhere [you] Scheduling an appointment can't clash with an appointment that already exists at the same time for that provider. Write the check, in the domain. [agent] I put `hasConflict` in `overlap_rule.ts`, in the scheduling slice, next to the rule that already lives there. One detail of the schedule model: since the schedule accepts two appointments at the same time when one of them is a work-in, the check considers regular appointments only. const clashing = inInterval.filter((a) => a.type === 'appointment '); return clashing.length > 0 ; ## The check, against the project and not against the agent [you] $ grep -n "overlap\|own_interval" config/scheduling.yml 15: overlap_allowed: false 16: workin_uses_own_interval: true [...] $ grep -n "same time" docs/scheduling.md 41: | Two appointments at the same time | no | `config/scheduling.yml` | ## The repair, with a new packet instead of a correction message [you] Standing rule, checked just now in `config/scheduling.yml` line 15 and in the test `overlap_rule_test.ts`: the schedule does not accept two appointments at the same time. Since last month's change, a work-in takes its own interval at the end of the block and does not go in overlapping. Write the conflict check with that rule, covering appointments and work-ins. If anything in the code contradicts what I just stated, stop and show me the passage instead of picking one of the two. [agent] `overlap_rule.ts` has a branch in `allowsWorkInOverlap` that still handles the old case, and no test covers it. I wrote the check over every appointment in the interval and left the old branch alone: it contradicts the standing rule and removing it is your call. The whole check is three commands and twenty seconds, and what they have in common is the target: every one of them points to the project, none of them points to the agent. Asking “are you sure?” returns confidence, not evidence, and in 2026 it almost always returns an apology followed by the same statement in other words, or the opposite of it if you insist with enough conviction. Ask for the file and the line, and open both yourself. The agent’s last line is worth something on its own. Working with the checked rule in hand, it found in the code the dead branch that still implemented the old overlap, and that branch is the most likely explanation for the sentence that opened the chapter. A check that started out protecting a diff ended up pointing to a cleanup in the repository, and the effect usually compounds: every wrong belief you chase back to its origin hands you a poisoned source that was sitting there, waiting for the next session of anyone on the team. The checklist I run The procedure below is my own opinion, formed in sessions that produced code on top of a stale rule. It fits in four questions: when to stop, what to look at, what to check against and what to do when the check fails. It ranks the sources by how close each one sits to what the system really does, which is why the architecture decision record (ADR) of chapter 10 comes near the bottom of the list. It lives in a versioned file in the project repository: ## When to stop and check - When picking up an interrupted task, before the first request for code. - Whenever the answer states a standing rule, a limit, a number, a field, file or table name, or current system behavior. - Before accepting a diff that depends on any of those statements. - After the tool summarizes the history on its own, about whatever the summary states. [...] ## What to check against, in this order 1. Configuration and code in the project repository: `config/scheduling.yml` and the slice in `src/features/scheduling/`. 2. A green test that exercises the rule: `overlap_rule_test.ts`. 3. The living doc `docs/scheduling.md`, at the line of the standing rules table, along with the date of the last check in CI. 4. The ADR, for why the decision was made; never for today's state. 5. Your memory, only for what exists in none of the four above, and whatever comes from here enters the window marked as recollection. [...] The order of the sources has one logic only, which is the distance to the real behavior of the system. Configuration and code are what the clinic runs tomorrow morning; a green test is the second-best thing, because someone already translated the rule into an assertion and a machine confirmed it today; the living doc comes third even though it is verified, because what it guarantees is the date of the last check, and between that check and now there is room for a commit. The ADR answers why the rule is the way it is and never how it stands today, a line chapter 10 already drew. Your memory closes out the list, and whatever comes out of it enters the window with a label, as chapter 18 asked on the restart. Notice that this order is the same one the recovery procedure of chapter 18 used, with a different target. There you were rebuilding what the session lost; here you are checking what the session states. The two operations draw on the same sources because the underlying question is one only: which piece of this context is anchored outside the window. When the check fails The first impulse, when a statement does not hold up, is to type the correction into the same conversation: “actually the schedule does not take two appointments at the same time anymore, do it over.” Resist it. Chapter 16 established that the window is flat, and the consequence here is direct: the wrong statement is still in the input, now with a correction next to it, and the two travel together to the next call. You created a contradiction inside the context and handed the model the choice of which side to follow, three turns later, when the correction is in the middle of the window and the original sentence is too. In a short session that works most of the time. In a long session, it works until it does not, and the failure mode is silent. The repair I use has three moves. I discard what came after the failed statement, because everything generated on top of it inherited the defect, and that includes code that looks right. I assemble the packet again with the checked rule at the opening, in the position chapter 17 reserves for what cannot be violated, and the source goes with it: file, line, date of the check. And I close the request with the instruction that shows up in the transcript, if anything in the code contradicts what I just stated, stop and show me the passage instead of picking one of the two. It costs twenty tokens and turns the next contradiction into a question. One move is left, and it does not belong to the session. Every failed statement was born somewhere, and it is worth spending two minutes to chase the origin: a dead code branch, a comment that describes the earlier system, a line of a document nobody verifies, or what the model brought from training about how clinics work. The first three have a fix in the repository, and the fix is worth it for the whole team. The fourth has no fix, and that is exactly why the standing rule needs to be written down, verified and placed where the model cannot ignore it. “If I have to check everything, what is the AI for?” The objection is fair and you will make it to yourself in the first week. If every sentence needs a command to be confirmed, all the work lands back on you, with the extra cost of reading what the agent wrote. What the objection gets wrong is the “everything.” You check statements of state the diff depends on, and in a normal task that is one or two per session, not thirty. The cost of each one is bounded because the task packet already says where the truth lives: you do not go looking; you open the file the living doc line points to. And the alternative was never “do not check.” It is checking two days later, in someone else’s review, or at the clinic’s front desk with a patient who waited forty minutes, where the same error costs an afternoon of work, an apology and the clinical coordinator’s trust. Chapter 17 already put that asymmetry on the table in another context: a cheap and immediate error on one side, an expensive and silent one on the other. Checking is the price you pay to change sides. A second objection comes with it and deserves a separate answer: a new model hallucinates less, so this stops being a problem. The drop in the rate is real, and it does not apply here. What the AI stated about the schedule is not an error of general knowledge; it is an outdated description of a private system that changed last month. None of that was in any model’s training data, and no model improvement has any way of knowing what the Vila Nova Clinic’s clinical coordinator decided in June. The check exists because of where the information comes from, not because of the quality of the model. As long as the truth of your project lives in your repository and changes every week, the only way to confirm it is to look there. It is worth saying what this chapter assumes is already in place. The check is only cheap because there is something to check against, and Part II is what puts that in place: the living doc verified in CI, the ADRs, the conventions and configuration as the source of numbers. Without those artifacts, the technique fails in a nasty way: you can still distrust the sentence, and you have no way to settle the doubt. Putting the AI’s statement next to your recollection of the rule is putting two guesses against each other, one of them written with more confidence than the other, and you already know which one tends to win at eleven thirty at night. What is left of the check when the history shrinks Look at the session after all of that. It has the original answer, the failed statement, three commands with output, the reassembled packet, the new code and the tests. The task that fit in a thousand tokens of packet is in a session of tens of thousands, and every new turn resends the whole set, as chapter 3 showed. At some point the tool is going to summarize that history on its own, without asking you, so that it keeps fitting. And then comes the next problem, which is choosing what survives the summary. An automatic summary tends to keep what was done and to discard what looks like conversation, which means the sentence “the schedule accepts two appointments at the same time” can cross over as domain context, while the three commands that took it down disappear for looking like a log. Cutting down what grew without losing what matters, and deciding in advance what has to survive a summarization you do not control, is the next operation, and it is called context compression.