No puede seleccionar más de 25 temas Los temas deben comenzar con una letra o número, pueden incluir guiones ('-') y pueden tener hasta 35 caracteres de largo.

12KB

Context Engineering — Chapter-12: Living documentation

  • Source: /library/Context Engineering/source-file.pdf
  • PDF pages: 74–81
  • Pages without text: none

Living documentation The previous chapter ended by pointing to an artifact that promises to describe the system as it is today. VilaSchedule has one. It is called docs/scheduling.md . It was written with care by a dev who left the team last year, and it is one task away from doing damage. The task arrives: Vila Nova Clinic’s clinical coordinator wants a schedule utilization report: the percentage of time each provider spends with patients each day. You hand the request to the agent, and it does what agents do well in 2026: it combs through the repository for context before writing any code. It finds the docs/ folder, finds the file with the perfect name and reads this:

How the schedule works

Each provider's schedule is divided into 20-minute intervals, generated from the weekly schedule configured in the system. Every appointment takes exactly one open interval; no appointments are booked outside the regular intervals. Patients with urgent needs are referred by the front desk to an urgent care center.

If you followed chapter 8, you have already spotted the two problems. The clinic’s intervals run 30 minutes, not 20; they changed more than a year ago, and nobody went back to the document. And “no appointments are booked outside the regular intervals” was true when the text was written, but the work-in from the previous chapter, that extra appointment squeezed into a full schedule, was implemented, was shipped, and is used every day. The agent saw no problem at all. It computed utilization by dividing the day into 20-minute blocks and ignored work-ins entirely, because the document stated that they do not exist. The report came back with absurd numbers: providers at 140% utilization on ordinary days, days packed with work-ins showing up as idle. The code was clean, the report’s tests passed, and all of it was wrong. Notice the mechanism, because it is the same one from chapter 1: the model does not verify the input; it answers it. A document in the window does not arrive with a stamp that says “trust this 60%.” It arrives as text, and confident prose sitting in the file with the most official name in the repository weighs heavily on the model’s attention. The outdated doc is not neutral context the model can ignore; it is poisoned context that competes with the code for the truth and wins often, because prose is easier to “understand” than a thousand lines of validation. Hence this chapter’s thesis, which I will state as my own opinion, backed by an argument from Cyrille Martraire that I will get to in a moment: outdated documentation is worse than absent documentation. Without the file, the agent would have read the scheduling code, found the 30-minute constant and the work-in entity, and the report would have been right, slower and more expensive, but right. With the file, it received a ready answer, and a wrong one, and ready answers are exactly what models prefer. Absence forces you to read the source of truth; a lie lets you skip it.

The document that describes the present The docs/scheduling.md you just read has a name: dead documentation, a document that promises to describe the system as it stands, but that was only true on the day somebody wrote it. Notice that it is not a badly written spec. The work-in spec from chapter 8 is a dated record of the intent behind a change, and aging is part of its job, the way it is part of a commit’s job. The dead document fails because it took on a different commitment: to say what is true now. An artifact with that commitment and with no mechanism that forces it to keep it is a scheduled lie; all that is missing is the date. The way out has a book behind it: in Living Documentation (Addison-Wesley, 2019), Cyrille Martraire set out the discipline of treating documentation as something generated or verified from the source of truth, instead of written alongside it and kept in sync by goodwill. The central idea is simple to state: the knowledge already exists in the system, in the code, in the tests, in the configuration; to document is to extract and present that knowledge, not to duplicate it by hand. Whatever does get duplicated by hand needs a mechanical checker that screams when the copy diverges from the original. See how that looks in VilaSchedule. The replacement for the dead document opens like this:

Scheduling: how it works today

System: VilaSchedule (Vila Nova Clinic)

Owner: the scheduling team Last verified: 2026-07-24, by the doc_scheduling_test test (the build fails in continuous integration if the table of current rules diverges from the configuration) Three header lines, and every one of them works. “How it works today” in the title declares the artifact’s commitment, the same one the dead document took on and broke. The named owner is accountable for keeping it current, and “last verified” replaces the usual “last updated”: the date does not say when somebody last touched the text; it says when a machine last checked that the text is still true, and it says which machine. For a model reading that file, the header calibrates the trust the dead document demanded in the dark. The heart of the document is the part that lied most in the dead version, the rules, and it is where the verification is anchored:

Current rules

Rule Value Where it is defined
Regular interval duration 30 min config/scheduling.yml

| Work-in duration | 15 min | config/scheduling.yml | | Work-ins per provider per day | 2 | config/scheduling.yml | | Maximum lead time for an appointment | 60 days | config/scheduling.yml | | Blocked schedule accepts a work-in | no | BlockRule (tests) | The third column is what separates this document from the previous one. Every value points to the place in the code it comes from, and that link is not decorative: it is the contract the test named in the header executes. The last section of the file explains the mechanism:

How to maintain this

This document is verified in continuous integration: the doc_scheduling_test test reads the table of current rules and compares each value against config/scheduling.yml. Anyone who changes the configuration without updating the table breaks the build, and the build points to the row that diverged.

The test is twenty lines long: a parser for the markdown table and five comparisons against the configuration file. That is not much code for what it buys. On the day the clinical coordinator raises the work-in limit to 3, somebody will edit config/scheduling.yml , the build will break and point to the offending row, and that person will fix the document in the same commit as the change, not “later.” The dead doc depended on memory; the living one depends on a test, and tests do not forget. That swap of failure modes is what Martraire proposed: do not promise discipline; install a mechanism. Not everything in the file is verifiable that way, and that is fine. The prose overview, which describes intervals, appointments, work-ins and blocks in a single paragraph, has no test to check it; what protects it is being short, stable and made of concepts that change rarely, not of values that change all the time. The rule of thumb I use: numbers, limits and behaviors that fit in configuration or in a test go into the verified part; prose is reserved for what the code does not say on its own, the vocabulary of the domain and the general shape of the flow. The smaller the unverified part, the smaller the surface where rot can start. “All docs rot, so why write them?” The objection you will hear when you propose this to the team is honest, and whoever raises it usually has scars: all documentation rots, so writing it is just scheduling a lie. I agree with the diagnosis and disagree with the conclusion, in two steps. First: documentation rots when maintenance depends on somebody remembering. The dead document at the start of the chapter did not rot by bad luck; it rotted because nothing

happened when it diverged from the system: no build broke, no test failed, no owner was held accountable. The living document does not promise that nobody will forget; it promises that forgetting has an immediate and cheap consequence (a red build today) instead of a late and expensive one (a wrong report a year from now). The objection is exactly right about an artifact with no mechanism, and it does not apply to one that has a mechanism. Second: the objection smuggles in the idea that the alternative to a doc that rots is no doc at all, and the opening scenario shows the real cost of that alternative when there is a model in the cycle. With no document, every session pays again to read the code and rebuild what the document would have said; chapter 6 showed you what that kind of repeated rebuilding costs, in tokens and in the chance of error on every round. The living document is the external memory chapters 3 and 4 showed the model does not have: instead of rebuilding the scheduling flow on every session, the agent reads one page verified three days ago. The choice was never between a doc that lies and pure code; it is between paying to extract the knowledge once, with verification, or paying on every session, with no guarantee. One version of the objection deserves a separate answer: “then let’s document only the bare minimum.” Yes. That is the corollary, not a refutation. VilaSchedule’s living document is one page long, and its “what this document does not cover” section hands the history of the rules off to other artifacts instead of absorbing that history itself. Minimal living documentation is the only kind the mechanism can protect end to end; the 80-page manual has no test that saves it, and at this point in the book you know it could not fit in a context window, and everyone in the room knows it too.

What each artifact answers Now that you have read two chapters of Part II, you can tell where each piece of information belongs. “The work-in limit becomes 2 per day” is the intent behind a change: a spec, a dated record, chapter 8. “The work-in limit is 2 per day” is the present state: a living doc, verified against the configuration, this chapter. When the agent goes to implement something new about the schedule, the living doc enters the window as a trustworthy portrait of the terrain, and the spec for the change enters as the target; the two artifacts complement each other without competing for the same role, and neither of them needs to be long, because each answers a single question. One question is missing, and it shows up on the first day somebody uses these artifacts for real. The living doc says the interval runs 30 minutes; the spec says the work-in started out at 15. Neither says why. Why fixed intervals instead of a free-form schedule? Who decided, when, against which alternatives? In VilaSchedule, that answer today sits where it sits on most teams: in a Slack thread from two years ago, in the memory of a dev who has left, nowhere at all. And an agent that does not know the why behind a decision is an agent one refactor away from undoing it with the best of intentions. Recording decisions, with their context and their consequences, is a third kind of artifact, and it is the subject of the next chapter.

Powered by TurnKey Linux.