Nevar pievienot vairāk kā 25 tēmas Tēmai ir jāsākas ar burtu vai ciparu, tā var saturēt domu zīmes ('-') un var būt līdz 35 simboliem gara.

17KB

Context Engineering — Chapter-20: Context packing

  • Source: /library/Context Engineering/source-file.pdf
  • PDF pages: 151–162
  • Pages without text: none

Context packing Thirty-four thousand tokens. That is the size of the input that just left your machine, and the inventory shows where it came from: the project conventions, pasted in full; the living documentation for scheduling; the architecture decision record (ADR) that explains the schedule of fixed intervals; the work-in spec; the whole work-ins folder, read file by file; the scheduling and appointments folders, because they touch on the subject; the complete output of the last npm test . At the end of that pile sits the request, which fits in one sentence: a canceled work-in stops counting toward the limit of two work-ins per provider per day. The matching change fits in a few lines of code, and the hurry is real: the front desk canceled two of the morning’s work-ins, VilaSchedule kept refusing the third, and the coordinator wants this settled today. The giant packet was not born of carelessness. It was born of the diligence of someone who learned the previous chapter’s lesson: no discarded conversation, everything that came in is current project material. You just never decided how much of each thing came in, or in what order. Back comes a diff in workin_controller.ts , filtering canceled ones out of the application programming interface (API) response, and a suggestion to touch config/scheduling.yml . The count that produces the limit lives in src/features/workins/day_limits.ts , a file that was in the window, pasted in full, fifteen minutes earlier. It does not appear in the diff. The convention that puts business rule validation in the domain, and never in the controller, was also in the window, and the first line of the answer broke it.

Notice that the usual complaint does not work here. Context was not missing: the right file was there, the right rule was there, the spec was there. It was not the poisoned context of chapter 12 either, because not one pasted line was wrong. What you have is a two-line task buried in a packet of thirty-four thousand tokens, assembled as a precaution, with nobody deciding how much of each thing was needed or in what order it should arrive. Packing is choosing the minimum, and choosing means saying no Context packing is selecting and packing, for one specific task, what goes into the window: what information, in what form and in what position. It is the operation that sits between the layers of chapter 16 and the first token of work. The layers answer what exists and how long each thing lasts; they do not answer how much of each fits in this task. A complete layer 1 is the whole conventions document, with all of its items, when today’s task may violate two of them. A complete layer 2 is the whole living documentation and the whole spec, when the change touches one line of the standing rules table. Packing is the act of trimming each layer down to the size of the task, and the hard part of it is the subtraction: you have to say no to material that is true, and useful on another day, about the system you are working on right now. The practice comes recommended where you would expect. In the guide “Claude Code: Best practices for agentic coding,” published by Anthropic in 2025 (anthropic.com/engineering), two recommendations run in that direction: be specific in the request, naming the files that matter instead of describing the task in vague terms, and clear the context between tasks, so the previous task’s material does not ride along into the next one’s

window. Both say the same thing from different angles: the packet is per task, and what is left over from the last task does not belong in it. The inventory of the bloated packet Before you assemble a lean packet, it is worth looking at the bloated one item by item, with the size of each piece beside it. Below is the inventory of the session that opened this chapter, reconstructed after the mistake, with each [...] marking the items that did not fit on this page:

Bloated packet: the whole project to change one work-in rule

[...] 2. All of docs/conventions.md, pasted because the code is new: ~1,200 tokens. The task may violate two of its lines. 3. All of docs/scheduling.md, the living doc: ~1,500 tokens. The task depends on one line of the standing rules table. 4. All of docs/adr/adr-001-fixed-intervals.md: ~900 tokens. No line of this task touches the interval model. [...]

  1. The five files in src/features/workins/, in full: ~3,500 tokens. The diff happens in two of them.
  2. The six files in src/features/scheduling/, in full, “so it understands the schedule”: ~4,800 tokens. [...]
  3. The full output of the last npm test: 340 green lines and one red, ~5,000 tokens.
  4. Yesterday's conversation about the utilization report, which came along because the session was never closed: ~9,000 tokens.
  5. The request, in the last message: “a canceled work-in can't count toward the day limit; fix it”. ~20 tokens. Estimated total: ~34,000 tokens before the first answer. The request takes up 20 of them.

The inventory shows three things the session never made visible. The first is the proportion: the description of the task is 0.06% of the packet, and the rest is scenery. The second is where the waste comes from, which is not the wrong items but the right ones entering whole: the conventions are in force, the living documentation is current, the files in the work-ins slice are the ones the task lives in, and not one of them needed to enter in full. The convention the answer broke was inside item 2, diluted in twelve hundred tokens of other rules this task runs no risk of breaking. The third is position. The request closes the packet, which is the only right place for it, and the space just before it, which shares with it the end of the window that gets the most attention, went to fourteen thousand tokens of test log and conversation about another task. No item in the packet says where the diff should happen, and the file that produces the wrong number sits in the middle of item 7, among four other files the task does not touch. The window is not uniform That third observation is the point where packing stops being a way to save tokens and becomes a design decision. Chapter 5 introduced the finding of Liu and colleagues in “Lost in the Middle: How Language Models Use Long Contexts,” published in the Transactions of the Association for Computational Linguistics (TACL) in 2024 (arXiv:2307.03172): when they held the relevant information and the question fixed and varied only the position of the document that contains the answer, performance traced a U-shaped curve, high at the ends and worst in the middle. In that chapter the finding explained why your session degrades. Here it becomes an assembly instruction, because the position of each piece of the packet is your choice.

Translated into the layers of chapter 16, that curve gives the packet three zones. The top is where what cannot be violated under any circumstances goes, and that is layer 1 material: the few house rules this task is likely to break. The middle is the shadow zone, and that is why it is where the volume goes, the layer 2 material the agent will consult but does not have to recite: the files where the diff happens, the standing rules the change does not alter. The end is the position closest to generation, and that is where the request goes, with the layer 4 material it carries. Two practical consequences come out of that. The first is that the request never opens the packet; it closes it. Writing the task first and then pasting six files buries the instruction in the middle of your own window. The second is that the critical rule should not be left for the middle in the hope that the model will find it: if it holds for everything the session produces, it opens the packet, whatever the three lines cost. The rest competes for room in the middle, and the middle is where you pay for every token twice, in money and in diluted attention. Four questions that assemble the packet The criterion I use to assemble a packet fits in four questions, in this order, and I think the order matters more than the questions. The first is: what diff does this task produce? Start from the output, not the input. When you answer “two lines in day_limits.ts and one query in workin_repository.ts , plus the tests next to both,” the core of the packet is assembled, because what goes in is what surrounds that diff. The question also works as an alarm: if you do not know which diff the task produces, the problem is not one of context but one of spec, and no packing fixes that.

The second is: if I drop this, does the answer change? Apply it item by item, and accept the honest answer. ADR-001 explains why the schedule runs on thirty-minute intervals, and the answer to this task is identical with or without it in the window: out. The acceptance criteria of the work-in spec describe the behavior of the whole feature, and the change touches the count: out. The convention about validation in the domain does change the answer, because the wrong answer broke exactly that one: in. The third is: what is the cheapest form that does the job? There is a ladder of granularity between citing and pasting, and almost everyone jumps straight to the last rung. The cheapest rung is the pointer, the file path, which costs one line and lets the agent go look if it needs to. The middle one is the excerpt, the twelve lines of the count instead of the three-hundred-line file. The most expensive is the whole file, which is justified when the diff happens inside it. The living documentation of chapter 9 enters as one table row; the file where the diff happens enters in full. The fourth is: where does each thing go, and what is the ceiling? Position you already know how to decide. The ceiling is a number you declare before you assemble, and it exists to make the subtraction mandatory: with no ceiling, every item passes the second question by a wide margin, because when in doubt anything can change the answer. It is the same discipline as a performance budget, and it serves the same end, which is to force the choice while it is still cheap. The same request, packed The packet that comes out of those four questions, for the same task as the opening, fits in a little over a thousand tokens. An excerpt appears below, in the order it enters the window:

Opening the packet: what cannot be violated (layer 1)

  • Business rule validation lives in the domain, never in the controller (project conventions).
  • Domain terms match what the clinic says: Appointment, WorkIn, Provider. No synonyms (Booking, Visit, Slot) and no generics (Item, Entity, Record). [...]

Middle of the packet: the task material (layer 2)

  • What changes: the day's work-in count starts ignoring canceled work-ins. The limit stays at 2 per provider per day.
  • Standing rules this change does not touch (living doc, verified in CI): a work-in is for the current day only; 15 minutes long; a blocked schedule takes no work-in.
  • Where the diff happens: src/features/workins/day_limits.ts, the

limit check, and src/features/workins/workin_repository.ts, the query that counts the day's work-ins, plus the tests next to both. [...]

Closing the packet: the request (layer 4)

Change the day's work-in count to ignore canceled ones, keeping the limit of 2 per provider. Start with the test that describes the new rule, next to day_limits.ts. If you need any file that is not in this packet, ask before assuming. [...] Two lines of that packet do work that is not obvious. “The limit stays at 2 per provider per day” is a constraint disguised as context: it blocks the most likely creative reading, which is to touch the number while the file is already open. And “ask before assuming” is the line that makes the minimum packet safe, because it turns a selection mistake into a question instead of turning it into invention. Together they cost thirty tokens. The packet file also records what was left out and why, and that section is not bureaucracy: it is what you reread when the answer comes back bad. If the agent gets it wrong for lack of information you excluded on purpose, that exclusion line becomes the

correction for the next assembly. With no record, the temptation is to go back to pasting everything, which is what produced the opening of the chapter. For the same reason, when you close the loop, leave one line in the state note of chapter 18 saying what opened its packet: which sources came in and at what length. It is one line, not a system, and it is what makes a future mistake diagnosable without archaeology: the question “what was the model looking at when it got this wrong?” finally has a written answer. “If I forget the right file, it will make something up” The objection comes in two parts, and both are fair. The first: I do not know in advance what the model is going to need, and missing material is worse than extra material, because when something is missing it makes something up. The second, more current: in 2026 the agent reads files on its own, runs searches on its own and assembles whatever context it wants, so packing by hand has become wasted work. The first part assumes a symmetry that does not exist. Missing material produces, at worst, one question or one extra read, and the clinic’s packet asks for exactly that in its last line; the cost is one turn, and the mistake is visible right away. Extra material produces a fluent, confident, wrong answer that gets past your tired eye and shows up in someone else’s review two days later. Erring on the side of too little is a cheap, immediate, self- correcting mistake; erring on the side of too much is expensive, silent and hard to attribute. When the two mistakes cost different amounts, the default goes to the cheap side. And you have a way

to check this without arguing: the clean-session A/B test of chapter 5, run again with the minimum packet on one side and the dump on the other, on the same task. The second part confuses who does the work with whether the work exists. When the agent decides on its own which files to read, it is doing packing, only with no admission criterion, no ceiling and no control over position; the result enters the window in the most expensive form there is, the whole file, and it drags the tool log along with it. That is how the loop in chapter 4 filled the window by itself. What you pack, in an agent that searches, is different: instead of pasting the files, you hand over the map (where the truth lives, which files the diff touches), the ceiling and the license to ask. Which tool offers which control over that search is the subject of Part IV; the admission criterion is yours in any of them. It is worth saying what this chapter assumes is already done. Packing is selection, not writing: every piece of the minimum packet is an excerpt of an artifact Part II told you to build, the verified living documentation, the ADRs, the conventions, the task spec. If those artifacts do not exist, the technique degrades in a specific way: you can still keep the packet small, but the little you choose becomes your own memory of the rule, typed on the spot, with nothing to verify it. A small packet with a from- memory paraphrase is worse than a large packet with a source, because it concentrates all of the model’s attention on a version nobody checked. What the packet cannot carry Look at the lean packet one last time and notice which of the layers of chapter 16 does not appear in it. Layer 3, the session layer, is missing, and it is missing for a structural reason: at

minute zero it is empty. Everything it will contain (that the count comes out of a single query, that the check fits in the file that already exists, that solving it in the controller was dropped and why) is born during the work, inside the window, and has no copy anywhere in the repository. That makes layer 3 the only part of the packet you cannot reassemble from scratch. When the session blows past the window, when the tool compacts the history or when you close the laptop and come back on Thursday, the layer 1 and layer 2 material comes back with one command, because it has an address, and layer 3 comes back from memory, badly and in pieces. That is why the next operation assembles nothing: it rebuilds the thread of a task in progress after the session lost it, and it is called context recovery.

Powered by TurnKey Linux.