25'ten fazla konu seçemezsiniz Konular bir harf veya rakamla başlamalı, kısa çizgiler ('-') içerebilir ve en fazla 35 karakter uzunluğunda olabilir.

18KB

Context Engineering — Chapter-23: Context compression

  • Source: /library/Context Engineering/source-file.pdf
  • PDF pages: 190–203
  • Pages without text: none

Context compression Start with the failure, because it is small and fits in one line of code: at 4:50 p.m. the agent hands back the last test case with an example inside it: a morning block that ends at noon and a work- in positioned at 12:00. No alarm goes off, because the example is consistent with everything in the window: the work-in, one of the extra appointments squeezed into a full schedule, takes its own interval, that interval sits at the end of the block, and the morning block ends at noon. Given what the window holds, the conclusion is airtight. And it is exactly what you forbade an hour and forty minutes earlier, at 3:10 p.m., and forbade with a reason. Rewind the afternoon to find the point where the prohibition evaporated. The task was an afternoon’s worth of work: now that a work-in takes its own interval at the end of the provider’s block, what was left to decide was where that interval falls when the morning block ends right at the lunch break, without the front desk calling the clinical coordinator for every patient at the door. At 3:10 p.m. the decision closed, and it closed with a reason: a work-in does not run into the lunch break, because the clinical coordinator uses the break for test follow-ups and the front desk has no way to turn away someone who is already there; a block that ends at noon sends the work-in to the end of the afternoon block. The conversation that produced that sentence took some fifteen minutes, went through two alternatives and one phone call, and stayed entirely inside the window. The afternoon went on. You created the rule file in the domain, wrote the test next to it with four cases, adjusted the repository query to bring the end of the block along with the day’s work-

ins, ran the suite three times, pasted test output, dropped two approaches, knocked down an expired statement about time overlap. The window filled up the normal way: nothing wrong went into it, there was simply a lot of it. At 4:40 p.m. the tool announced, in one discreet line, that it had summarized the conversation history. You barely looked. The session went on, the agent went on answering, and ten minutes later out came the 12:00 example. Notice what does not explain this error. Durable context was not missing: the break rule was not in the living doc because it had been born that afternoon, and it is fair that it was not there. It is not hallucination in the sense of chapter 19, because the agent contradicted no source that was in the window; the decision simply was not there anymore to be contradicted. It is not a bad restart from chapter 18, because the session never died. What happened was a mechanism that compressed the session while you worked, chose what to preserve without asking you, and chose what was done over why. Compressing is choosing what is left Context compression means reducing what has already grown inside the window while preserving what the task cannot lose. The word doing the work there is “preserving.” Shrinking history is easy and any blind cut shrinks it; the operation only has value because it defines, before the cut, what gets through. That draws a clean boundary between compression and the context packing from chapter 17. There you select at the door what will come in, with the window empty and all the power to say no. Compression works on a window that is already full: the material is in, it has already been used, it has already produced a

decision, and now it needs to fit in less space than it takes up. The packing question is what deserves to come in. The compression question is what deserves to stay. And it acts on one layer only. Chapter 16 showed that the session layer, the one that lasts hours, is the one that swells over the day, and it is the one that gets summarized. The project layer and the task layer do not go through the summary: they live in files, with lifetimes of months and of days, and you reassemble them in the next window whenever you want. It exists because the ceiling of chapter 3 is a real ceiling. A long session reaches it, and reaching the ceiling leaves three ways out: stop the session, let the beginning fall off, or summarize. Only the third one decides what stays, and that is why the tools chose the third. The general mechanism is automatic session summarization: when the history gets close to the limit, the tool asks the model itself for a summary of the conversation, replaces the history with that summary and carries the session on from there. In 2026, Claude Code calls that moment compaction, offers the /compact command to trigger it, runs the same mechanism on its own when the window gets tight and accepts instructions about what to preserve (Claude Code documentation, Anthropic, 2025-2026). The name of the command belongs to a moment in time and will change; the observable effect is what matters here, and the mechanics of configuring this on your machine are left for Part IV. If tomorrow the tool calls it something else, the problem of this chapter stays the same, because it is born out of the existence of the ceiling, not out of the name of the command. Compression and recovery solve the same loss at opposite moments. Compression acts before the loss: the information is still alive in the window and you decide what survives the cut. Context recovery, chapter 18, acts after the loss: the session has already died or the thread has already vanished, and the job is to

reassemble from the outside in. Good compression reduces how often you will need chapter 18, and it does not replace the restart routine from there, which is still the right answer once the damage is done. What a summary optimizes for The summarizing machine is good at what it sets out to do, so start by knowing what that is. A session summary is abstractive, not extractive. An extractive summary cuts sentences out of the original and stacks them. An abstractive one reads the history and writes a new text. Maynez and colleagues measured the price of that rewriting in “On Faithfulness and Factuality in Abstractive Summarization” (arXiv:2005.00661), at the Association for Computational Linguistics (ACL) conference in 2020: over news summaries produced by neural systems, they found content not supported by the source document in most of the summaries, and most of that content was extrinsic, the kind the source can neither confirm nor deny. It is the same intrinsic and extrinsic pair from chapter 19, now applied to the summary instead of the answer. Two practical consequences come out of that. The first is that the summary can state something the session never said, and the check from chapter 19 applies to it as it applies to any answer. The second is quieter and it is the one that ruined your afternoon: rewriting is choosing, and what is not chosen disappears without leaving a hole. A passage lost from an extractive text leaves a visible gap. A well-written abstractive summary has no gap at all. It is coherent, fluent and complete in itself, and nothing in it warns you that a decision about the lunch break ever existed.

And the selection is not random. The summary is written to carry the task on, so it keeps what looks like the state of the work: files created, functions written, tests that passed, suggested next step. The why looks like conversation. The 3:10 p.m. decision arrived wrapped in fifteen minutes of discussion, one rejected alternative and a phone call, and all of that sounds, to the model writing the summary, like a preamble to the real work. This is how the summary came out that day:

Summary of the conversation so far

The user is implementing work-in positioning in VilaSchedule, a clinic scheduling system written in TypeScript. A work-in takes its own interval at the end of the provider's block. Work completed:

  • Created src/features/workins/position_in_block.ts with the function workInPosition, which returns the work-in time counting back from the end of the provider's block. [...]
  • The suite is green, except for the full block case, still pending. Suggested next step: handle the full block case and review the function names. That summary is good at what it set out to do. It tells you where the work stopped and lets you carry on from there. It just does not know that a rule about the lunch break exists, and nothing in the text suggests it should know. Complaining about it is complaining that a tool does what it does. The useful question is this: whose job was it to make sure the 3:10 p.m. decision got through? The anchors you write beforehand The opening scene already answers the most common criticism of this technique: a summary loses the important decision. It does lose it. It will keep losing it, because no model can guess which of the afternoon’s forty sentences are the three that cannot disappear. The way out is not to trust the summary more. It is to write down, before compression runs, what needs to survive it. I call those pre-compaction anchors: a handful of lines that live outside the window and that compression cannot erase, because they are not inside it. They are born at the moment the decision is born, not at summary time. Wait for summary time and you are writing from memory, and memory at that point has already gone through the same filter the summary is about to use.

The admission criterion is a single question: if this session disappears right now, does this come back for free? The code comes back, because it is on disk. The standing rule comes back, because it is in the chapter 9 living doc. The reason behind an architecture decision comes back, because it is in the chapter 10 architecture decision record (ADR). None of that is an anchor; all of that is a pointer. An anchor is what exists only inside this session and took work to be born: today’s decision that has not become a record yet, the rejected path with its reason, the statement that already proved false. In the work-in session, the sheet looked like this:

Decisions closed in this task (with the reason)

  • 3:10 p.m.: a work-in does not run into the lunch break. A morning block that ends at noon sends the work-in to the end of the afternoon block. Reason: the clinical coordinator uses the break for test follow-ups and the front desk has no way to turn away a patient already at the door. [...]

Dropped (and why)

  • 3:20 p.m.: pushing the work-in to the first open slot after lunch. Reason: it breaks the rule that a work-in takes its own interval at the end of the block and brings overlap back. [...]

Statements that already failed (do not reintroduce)

  • 3:55 p.m.: “the schedule accepts two appointments at the same time when one of them is a work-in”. The rule expired in June; checked in config/scheduling.yml, line 15. [...]

The instruction that goes with the summary

When summarizing this session, preserve the closed decisions with their reason, the rejected paths with their reason and the failed statements, literally. You may discard test output, pasted file excerpts and the narrative of the attempts. If you cut anything beyond

that, say what was cut. Notice three things. The reason travels with the decision, because a decision with no reason is an orphan rule and the next session will want to renegotiate it. The rejected path goes in with the same weight as the decision, because without it the rejected alternative comes back ten minutes later, dressed up as a new idea. And the failed statement from chapter 19 lives here: knocking down an expired statement costs one check, and paying for that check twice in the same day is a bad deal, which is what happens when the summary takes the sentence and leaves the check behind. The last block of the sheet is what turns an anchor into an instruction. You do not depend on compression guessing. You name what to preserve literally, you name what can be thrown away, and you ask for the cut to be declared. That last request is cheap, and it gives back what was missing in the opening scene: knowing that something was dropped. With the anchors on the table, the same session, compressed at the same moment and down to the same size, produces a different summary:

Closed decisions (do not reopen)

  • A work-in does not run into the lunch break: a morning block that ends at noon sends the work-in to the end of the afternoon block. Reason: the clinical coordinator uses the break for test follow-ups.

[...]

Statement that already failed in this session (do not reintroduce)

  • “The schedule accepts two appointments at the same time when one of them is a work-in.” The rule expired in June; checked in config/scheduling.yml, line 15.

Where the diff stopped

  • position_in_block.ts and the test next to it: four cases, three green. [...]

Cut from this summary on purpose

Test output, pasted file excerpts and the narrative of the attempts. All of that reproduces by running the suite or reading the disk.

The two summaries are about the same size. The difference is not in how much was preserved; it is in what. The first kept the trail of the work, which the disk already kept. The second kept what existed only in the conversation and sent the trail of the work away, because the trail reproduces by running the suite. The last block of that summary is the badge of honor: it says what it threw away, and with that you know where to look if you miss something. “Then turn automatic summarization off” There are people who draw the opposite conclusion: if the summary loses things, turn the summary off and always work with the full history. The objection sounds prudent, and acting on it is a bad deal. Turning it off does not make the window grow. It still has the ceiling of chapter 3, and what changes is what happens when you touch it: instead of a silent loss, you get a hard stop in the middle of the afternoon, or worse, a tool that starts letting the beginning of the history fall off in silence, which is the same loss with no choice behind it. Compression does not invent the problem; it answers it. A long session will be cut one way or another; the real choice is between deciding for yourself what stays and handing that decision to a mechanism that does not know which sentence of your afternoon was the important one. If the sheet reminded you of the state note from chapter 18, it did so for a good reason: the two hold the same material, a closed decision and a rejected path, always with the reason attached. What changes is who it is written for. You write the note when you stop, and it speaks to tomorrow’s session; you write the sheet while you work, and it speaks to the summary that will run in a little while. If you already keep the note, write the anchors inside

it, in the same file, without duplicating a single line. What does not work is putting both off until the end of the day, because by then the summary has already run. It is worth recording the honest discomfort that is left. Writing an anchor costs time while you are in the middle of the reasoning, and that is exactly the moment when stopping is least appealing. I write them anyway, and the calculation that convinces me is the one from chapter 6: every anchor line costs once, and the lost decision costs the whole discussion again, plus the wrong code that came out in the meantime, plus the complaint that arrives through the front desk.

What this chapter assumes is in place Compression is the operation in Part III that leans hardest on what came before it, and it is only safe because most of what it discards has an address outside the session. All of Part II is holding this operation up from below. The living documentation from chapter 9, verified in continuous integration (CI), holds the standing rule, so the summary can forget it with no damage. The ADR from chapter 10 holds the reason behind architecture decisions, so they do not need to become anchors. The conventions from chapter 11 hold what cannot be violated, and the spec holds what was agreed for the task. With those four in place, the anchor sheet stays short, and a short sheet is a sheet you keep. Without them, the arithmetic flips. If the only copy of the standing rule is the sentence the agent said at 3:10 p.m., and the only copy of the reason is the conversation that produced it, then the session history has become the project’s knowledge repository. Compressing a knowledge repository is not compressing; it is destroying. And no anchor saves a project that would need to anchor everything. One window, one task Packing, checking, compressing and recovering all manage the same thing: one window that carries one task to the end. That was the premise of the five operations from chapter 16 on, and it works for as long as it is true.

It stops being true early. On an ordinary Thursday you are on the morning work-in, the clinical coordinator asks for a no-show report for today and the front desk integration test breaks for an unrelated reason. Three tasks, one session. Now there is no single thread to compress: what is essential for the work-in is noise for the report, and the anchor sheet of one task has nothing to do with the other’s. Any summary of that session will mix three topics and serve all three badly, however well written it may be. The question moves. It stops being what fits in this window and becomes how many windows the work needs, who assembles each one and what one hands to the next when it finishes. That is where Part III leaves the day-to-day operations and enters context architecture decisions, and the first name in that conversation is isolation.

Powered by TurnKey Linux.