Du kannst nicht mehr als 25 Themen auswählen Themen müssen entweder mit einem Buchstaben oder einer Ziffer beginnen. Sie können Bindestriche („-“) enthalten und bis zu 35 Zeichen lang sein.

19KB

Context Engineering — Chapter-21: Context recovery

  • Source: /library/Context Engineering/source-file.pdf
  • PDF pages: 163–176
  • Pages without text: none

Context recovery The task is the same one as in the previous chapter, and this time the packet was right. A canceled work-in, one of the extra appointments squeezed into a full schedule, stops counting toward the limit of two per provider per day. The window holds twelve hundred tokens of the minimum packet from chapter 17, and on top of that a whole morning of work. At 10:25 a.m. you decided that the new check would live in day_limits.ts , in the domain. At 10:40 a.m. you dropped the idea of filtering the canceled ones in the controller, because the house convention does not allow it. At 10:50 a.m. you found out, by reading the code, that the status field is what decides what counts. At 10:55 a.m. you dropped the idea of adding a new column to the work- in, which would solve today and leave the rule written in two places. At 11:20 a.m. the machine rebooted on its own for an update, and the session died in the middle of the task. You reopen the tool and start again. It is no big deal: you reassemble the packet in a minute, because every piece of it lives in a file. So you type the opening from memory, three sentences that sum up where you were: clinic scheduling system, a canceled work-in does not count toward the limit, already worked on this earlier today, pick up from there. The answer comes back worse than the first one. The agent suggests filtering the canceled ones in the controller, in the endpoint listing, with a clean line and a test alongside it. That is exactly the path you dropped at 10:40 a.m. And this time you accept it, because the suggestion looks reasonable, because it is eleven thirty, and above all because the reason you had for

rejecting it died with the session. On the first morning, you spent twenty minutes and two passes through the code to reach the conclusion that this path would not do. On the second, you took the path as if you had never seen it before. Rule out the usual suspects before you blame the tool, because none of them was at the scene. No project material was missing: the conventions, the living documentation and the files in the diff all came back into the window whole, because every one of them has an address. What vanished was the morning: four decisions, two rejected paths and one discovery that came up inside the session and never left it. I call it losing the thread because what you lose is not information about the project; it is the path you had already walked. Recovery is reassembling what had no address Context recovery is rebuilding the context that was lost, out of artifacts that live outside the session. The word doing the work in that definition is “outside”: recovery does not happen inside the dead window; it happens from what survived the window. The cheapest outside material to bring back is the transcript of the session itself. Most tools in 2026 keep the conversation on disk and know how to reopen an interrupted session, like Claude Code’s claude --resume , and when the resume goes back far enough to cover the whole morning it is the first thing to run. The scene in the opening goes wrong at exactly that point: starting over from memory with a resume command at hand is throwing away the only faithful copy of the morning that still existed. Resuming the session, though, is not the same operation as recovering the context, and the difference shows up in three situations. The first is having nothing to resume: the session ran

on another machine, in an ephemeral environment spun up from scratch, in a tool that does not persist the conversation, or past the retention window. The second is the transcript coming back without what you need, because the tool summarized the history while you worked and the summary kept what was done without the reasons, which is the topic of chapter 20. The third shows up when everything works: the resume hands the window back as it was, with the decisions mixed in with the drafts, with file excerpts already out of date and with the test output from two hours ago. It hands back the bloated packet of chapter 17, and it hands it back flat. Recovery is the opposite: reassemble the minimum, the task packet plus the few lines that decide the morning. What no command brings back is what has no copy anywhere outside the window: not in the transcript, not in a project file, not in a note of your own. That part is not recovered; it is reinvented. And reinvention that passes for a restart, for picking the task back up where it stopped, is what makes you pay for a morning of work twice. The target of recovery is smaller than the feeling of loss suggests, and it is worth mapping it against the layers of chapter 16 before you start anything. Layer 0, the tool’s own standing load, comes back on its own in the new session. Layer 1 lives in a versioned project file and comes back by reading. Layer 2 lives in the feature’s artifacts and comes back the same way. You reassemble the two together with the operation of chapter 17, packing again, and the cost of that is the cost of any new session. Layer 4, the turn layer, died and does not need to come back: you reproduce the test error by running the test. What is left is layer 3, the session layer, and it is the whole of recovery.

What makes layer 3 expensive is not its size. The four lines you lost would take up fifty tokens. It is where they came from: each one is the residue of work already done once. You read the code, you tested a hypothesis, you rejected a path, and to rebuild the line means doing that work again or remembering it. Memory gives that part back worst of all, because it gives back conclusions without the reasons. You remember that the controller would not do; you do not remember why, and without the why the conclusion does not hold up against a well-written suggestion at eleven thirty. The idea of writing down, outside the window, what the window does not keep shows up as a recommendation in the tooling literature. In “Effective context engineering for AI agents” (Anthropic, 2025, anthropic.com/engineering), the same text that treats context as a finite resource describes structured note- taking outside the window, notes the agent writes to a file and rereads later, as one of the ways to sustain tasks that last longer than one window. What holds for the agent holds for you, only the pen is in your hand. The routine that follows is my own opinion, formed over restarts that went badly. Recovery is not prevention Before the routine, a boundary, because two techniques in this book treat the same problem at different moments and confusing them costs a lot. Context compression, the topic of chapter 20, acts before the loss: it reduces what grew in the window and preserves what matters, and it also takes care of what survives when the tool summarizes the history on its own. Context recovery acts after the loss: it rebuilds the thread from what exists outside the session. One operates on a context that is still there; the other on a context that is gone.

The practical consequence is that this chapter is not going to teach you to avoid the loss. Prevention has a chapter of its own, and reaching it before you come through here would get the order of the lessons backward: nobody who has never lost a morning writes a single note. The only prevention this chapter asks for is the state note, and the state note is not compression: it does not summarize the history; it records state. The routine I use to restart a task The first step is to reassemble the packet before you talk. It sounds obvious and almost nobody does it, because the new session invites you to explain instead of to load. Typing “clinic scheduling system, TypeScript” means rewriting from memory a layer 1 line that already exists, finished, in a file. Redo the packet of chapter 17, with the same four questions, and most of what you lost comes back at no cost to memory at all. The second is to rebuild layer 3 from evidence before memory, and I order the evidence by reliability. I start with the uncommitted diff, which is the hardest evidence there is: what is on disk happened, and it shows where the work stopped without depending on anyone to remember. Then I run the tests, because their result is the second most reliable piece of evidence, and it costs almost nothing to get, and a red test with a descriptive name tells you the intention along with the state. Then I read the state note, if there is one. Only then do I fall back on the transcript of the dead session, through the tool’s resume or through the file on disk, when it exists and still reaches back to the morning, which in 2026 depends entirely on which tool you use. That is the topic of Part IV. It comes fourth not because it is unreliable, but because it is bloated: it comes back whole, with

decisions and drafts tangled together, and mining the reasons out of it costs more reading than checking the diff. Last, and only for what is left over, I use my own memory. The third step is to separate what you checked from what you remember, and to say so in the window. The window is flat, as chapter 16 established: a sentence verified in the code and a reconstructed guess come in with the same weight, and the model has no way to tell them apart if you do not tell them apart. One line settles it: “this I just checked in the diff; this part is my recollection of the morning, check it before you use it.” A marked recollection turns into a question, which costs one turn; an unmarked one turns into a premise, which costs the rest of the task. The fourth is to test the restart before you ask for code. The test has three questions, and they are the same ones that define what layer 3 held: what is already done; which decision is closed; and what has been dropped and why. Ask the agent to answer all three before it writes a single line, with the explicit instruction to say it does not know instead of guessing. If it answers all three with what is in the packet, the thread is back. If it reopens the path you had dropped, a piece is missing, and it is much better to find that out in a paragraph of the reply than in an accepted diff. The state note The state note is the artifact that makes this routine cheap. It is a working file, one you write during the task, alongside it and outside version control, that answers the three questions of the restart test plus a fourth. Below is an excerpt from the note for the canceled work-in task, with each [...] marking what did not fit on this page:

Where the diff stopped

  • src/features/workins/day_limits.ts: the limit check already ignores canceled work-ins. Still missing the case of a work-in canceled and rescheduled on the same day. [...]

Closed decisions (do not reopen)

  • 10:25 a.m.: the new check stays in day_limits.ts, next to the count already there. Reason: the convention puts business rules in the domain. [...]

Dropped (and why)

  • 10:40 a.m.: filtering canceled ones in the controller. Reason: it violates the convention of validation in the domain.
  • 10:55 a.m.: adding a counts_toward_limit column to the work-in. Reason: it would solve today and leave the rule written in two places.

Open (where the next session starts)

  • Does a work-in canceled and rescheduled on the same day count once or not at all? The clinic coordinator has not answered yet. Until the answer comes, the code treats it as not at all and the test records the open question in its name. [...] The fourth section is the one I took longest to adopt and the one that saves the most time. “Open” records the next question, and without it the restart starts in the wrong place: you come back, understand where you stopped, and spend ten minutes rediscovering what you were about to do. Notice also the shape of those entries, which always carry the reason right next to them. A rejected path with no reason does not survive the first suggestion to the contrary, because the reason is the only thing you have left to push back with.

Each section of the note keeps the minimum that no other artifact in the project keeps. The standing work-in rule does not go in: it is in the living documentation verified in continuous integration (CI), and chapter 9 already settled that address. A decision that outlives the task does not go in: it becomes an architecture decision record (ADR), in the shape chapter 10 gave it. A message for the team does not go in, and here it is worth recalling the anti-pattern of chapter 12, the “ATTENTION, Friday’s deploy is on hold” that lives forever in the persistent file. The state note is the address that was missing for layer 3 material: it starts with the task, it dies with the task, and it pollutes no file the whole team carries into every session. Its cost is one line per decision, written at the moment of the decision. I write a line when I drop something, which is the moment the information is most expensive and most volatile, and then again when I close off a piece of the diff. All told, a few seconds at a time. The two restarts, side by side The excerpt below compares the restart that opened the chapter with the restart guided by the note, on the same task and with the same loss. Reconstruction: I condensed two restarts of the canceled work-in task into one representative excerpt, with the lines rewritten; the mechanics are those of any session that dies in the middle of the work.

Restart 1: starting over from memory

[you] Context: clinic scheduling system, TypeScript. A canceled work-in can't count toward the limit of 2 per provider per day. I already worked on this earlier today. Pick up from there. [agent] Sure. The most direct path is to filter canceled ones out of the listing in the work-ins endpoint, in workin_controller.ts: const dayWorkIns = workIns.filter((w) => w.status !== ‘canceled ‘); [...]

Restart 2: with the state note

[you] Picking up an interrupted task. Below are the task packet and the

state note I wrote during this morning's work. Before writing a single line of code, answer me in three sentences: what is already done, which decision is closed and what has been dropped. Whatever is not in the note, say you don't know instead of assuming. [task packet and state note pasted here] [agent] Done: the check in day_limits.ts already ignores canceled work-ins, and the new test in day_limits_test.ts is red because the day query, in workin_repository.ts, still counts the canceled ones. Closed: the check lives in the domain, in day_limits.ts, and the field that decides what counts is status. Dropped: filtering in the controller, for violating the convention of validation in the domain, and adding a counts_toward_limit column, for duplicating the rule.

I don't know what to do with a work-in canceled and rescheduled on the same day; the note says the clinical coordinator has not answered yet. [...] The difference between the two is not in the size of what you typed; the second restart has one more sentence of instruction. It is in the order and in the license. The order is state before request: until the agent declares where the work stopped, you do not know whether you are talking to someone who read the morning or to someone who is guessing at it. The license is the “say you don’t know,” and its value shows up in the agent’s last line, when it hands the open question back instead of picking an answer for it. Without that line, the open question becomes a silent decision inside the diff. “In 2026 the agent handles it on its own” The objection is a live one: the agent reads the project on its own, runs the diff on its own and rebuilds the context without you narrating anything. Why keep a note by hand? That part is largely right, and it became the second step of the routine. The agent reads the diff faster than you do, runs the whole suite without complaining and assembles layers 1 and 2 better than your eleven-thirty memory. Delegate that. What it does not do is remember what nobody wrote down, and layer 3 was never on disk. Looking at the code, it sees the check in the domain and concludes, reasonably enough, that this was the

choice; it has no way of knowing that the controller was dropped over a convention, and the new column over a duplicated rule. Where evidence is missing, it fills the gap with the most plausible alternative, and it hands the invention back with the same confidence with which it hands back what it read. Automatic reconstruction is excellent for what is written down, and it is also exactly what produces restart 1. A second objection arrives with it: why not leave the session alive forever, and never need a restart? Because the eternal session is the bloated packet of chapter 17 growing on its own with every turn, and because chapter 3 already showed the ending: the history is resent whole, it hits the ceiling, and the tool cuts the beginning with no warning. A long session does not avoid the loss; it postpones the loss and chooses for you what gets lost. It is worth saying what this chapter assumes. Recovery is cheap in proportion to what has an address outside the session: step one costs a minute because Part II gave an address to layers 1 and 2, in the living documentation verified in CI, the ADRs, the conventions, the spec and the persistent file. Without those artifacts, the whole restart falls back on the last place in that order of reliability: your own memory. What degrades is not the time, which is five minutes of typing. It is the fidelity: every restart reintroduces a slightly different version of the same work-in rule, and a task interrupted three times ends up with code stitched together from three versions of it, none of them checked. Not everything that came back is still true Look at the window right after a good restart and notice what you just assembled. There is material you checked in the diff two minutes ago, material the note recorded at 10:40 a.m., material

you remember from the morning and material the agent filled in by inference while it read the code. All four are in the same window, and all four look equally like fact, and only you know which is which, for now. Add the interval to that. While the session was dead, the clinical coordinator may have answered the open question, somebody may have changed the work-in limit in the project, and the rule you rebuilt from memory may have changed last month with you none the wiser. Recovery gives the thread back; it does not guarantee that the thread is still tied at the other end. Checking whether what the AI believes matches what the project says today, and doing that before the code goes out, is the next operation, and it is called context validation.

Powered by TurnKey Linux.