Você não pode selecionar mais de 25 tópicos Os tópicos devem começar com uma letra ou um número, podem incluir traços ('-') e podem ter até 35 caracteres.

16KB

Context Engineering — Chapter-19: Context layers

  • Source: /library/Context Engineering/source-file.pdf
  • PDF pages: 140–150
  • Pages without text: 144

Context layers Wednesday morning, a small task: change a work-in rule in VilaSchedule, Vila Nova Clinic’s scheduling system, which has carried the examples in this book since Part II. A work-in is the patient the front desk squeezes into a schedule that is already full. You open the session and put the context together in the way that looks careful. You paste the whole conventions document, you paste the whole scheduling schema and, to bring the agent up to speed, you paste yesterday’s conversation too, the one where you and a colleague spent half an hour on work-ins. Sixty thousand tokens before the first question, and not one line of it is false: the conventions are in force, the schema is the one running in production, and the conversation happened. The code comes back wrong in a disconcerting way. It respects the naming convention, it gets the tables right, and it implements a parameter called allowsWorkInDuringPartialBlock , which exists nowhere in the system. You hunt for the origin and find it: it is in yesterday’s conversation, in the message where your colleague asked whether a blocked schedule could take a work-in during a partial block. Ten minutes later, in that same conversation, you answered no, that a blocked schedule takes no work-in under any circumstances, and that was the end of it. To you, that was a discarded guess. To the model, it was text in the input, carrying exactly the same weight as the rule the project’s living documentation has verified in continuous integration (CI) since 2024.

Notice what does not explain the error. It was not lack of context: the correct rule was in the window, pasted twice, in the living documentation and in the conventions. It was not excess in the simple sense of size: sixty thousand tokens fit comfortably in any 2026 window. What happened is that the window is flat. It has no column saying “this has always been true” and another saying “this was thrown out yesterday at 3:40 p.m..” Everything arrives as text, and chapter 3 already gave the reason: every call reassembles the whole input and hands it to a model that was never in your Wednesday. It has no way of knowing that one sentence is a decision and the other is a draft, because the two arrive identical. A layer is a lifetime, not a folder Working in context layers means organizing what enters the window by the lifetime of the information: how long that piece of text stays true and useful before it turns into dead weight. A context layer is a set of information that shares a lifetime, enters together and leaves together. One note on vocabulary before going on, because the word layer already has an owner on the shelf. In the previous book in this trilogy, FOCUS Architecture (https://books.kodel.com.br/en/books/focus/), a layer is the organization of folders by technical type, all the views in one folder, all the services in another, and that is precisely the shape that book refuses. The criterion it adopts instead is the axis of change, things that change together live together, and the cut that comes out of it is the vertical slice. None of that is in play here. A context layer is not a folder, it does not describe architecture and it does not say where the code lives; it is a band of lifetime inside a single call’s window. If you want a bridge

between the two ideas, the bridge is mine and not the previous book’s: both group by what changes together, one on disk, the other in time. Why lifetime, and not subject, size, or source? Because lifetime is what predicts the moment a piece of information goes from help to hindrance. Yesterday’s conversation was useful for thirty minutes and turned into poison the next day; the block rule has been useful since 2024 and will still be useful in the next session anyone on the team opens. Sorting by subject would put both in the same drawer, “work-ins,” which is exactly the mistake Wednesday made. Anthropic stated the general principle in “Effective context engineering for AI agents” (2025, anthropic.com/engineering): context is a finite resource, to be curated, and not a warehouse where everything that might help gets dumped as a precaution. Curating demands a bar for entry, and the bar in this chapter is validity over time. The layers of a session The first layer lasts as long as the project. It is the identity of the system, the conventions that hold for all new code, the decisions nobody reopens on each task. In VilaSchedule, it is knowing that the schedule is built from fixed intervals, that Appointment and WorkIn are the names the clinic uses and never a synonym, and that a test lives next to the file it tests. That information changes on a scale of months, when it changes, and Part II already gave its address: the persistent context file from chapter 12, which points to the conventions, the architecture decision records (ADRs) and the living documentation instead of copying them. The second lasts as long as the task. It is the spec for what you are about to build, the rules in force for the flow you are about to touch and the list of files where the diff will happen. It is born

when the task starts, it dies when the task closes, and it is the layer most sessions assemble without noticing, from memory and badly. Nothing in it is news to the project: the artifacts from chapters 8 to 11 already contain every piece, and the work here is choosing which ones enter, not writing them again. The third lasts as long as the session. It is what the work found out after it started: that the day’s work-in count comes out of a single query, that the new check fits in the file that already exists, that the alternative of solving it in the controller was dropped and for what reason. None of those sentences existed when you opened the editor, none survives the end of the day on its own and all of them are expensive to rebuild. It is the most fragile layer that matters, and that is why chapter 18 exists. The fourth lasts one turn. The error the test just spat out, the twelve lines you pasted for the edit, the request you are typing right now. Useful life of one or two message exchanges. After that, it does not stay neutral: it becomes a distractor, in the exact sense chapter 5 measured, an old version of the code competing with the current one for the model’s attention. There is also a layer you do not assemble, and ignoring it is what keeps the token count from ever adding up. Call it layer 0: the tool’s system instruction, the definitions of every available tool, the user-level global instruction files. In chapter 2 you measured that baggage in your own agent, with a two-sentence prompt, and saw tens of thousands of tokens riding along. In 2026, each tool gives you a different degree of control over that material, and which knob exists in which tool is the subject of Part IV. What matters here is the accounting: you start every session with the window partly occupied by a layer you did not choose.

The diagram is the stack of one session, and the arrows point in the direction the cut travels when the window gets tight: from the bottom up. At the bottom sit the volatile, fat layers, because test output and file snippets weigh far more than a line of convention; at the top, the stable, small ones. The solid arrows mark what leaves early and without mercy; the dashed one, what leaves only when the whole job is done. Layer 1 stays out of the queue while the project is the same, and layer 0 always stays out of it, because it is not yours to cut. The Wednesday in the opening was a placement error in the stack: a piece of layer 4 information, born the day before and expired the same day, entered as though it were layer 1. The stack has one more silent dividend, and it comes from the caching in chapter 6. The discount providers give in July 2026 applies to the prefix of the input that repeats byte for byte between calls, and assembly by layers produces exactly that prefix: layers 0 and 1 at the top of the payload, unchanged during the session, with whatever changes each turn entering after them. The same order that protects the stable from the cut makes every call in the session cheaper; an edit at the top, mid-session, invalidates the cache from there down and the next call pays full price. Stable first was already the discipline of discarding; the provider’s meter charges for the same order. The same session, annotated by layer Naming the layers is only worth it if you can point, in a real session, to which one each piece belongs. Below is the next VilaSchedule task, stopping the same patient from getting two work-ins on the same day, with the session context annotated

item by item. It is a maintenance session reconstructed for teaching, and each [...] marks a part of the annotation that did not fit on this page:

Layer 2: task (lives for days; leaves when the task closes)

  • What to build: a patient can have only 1 work-in per day across the whole clinic, whatever the provider (change spec).
  • Standing work-in rules this change does not touch: current day only; at most 2 per provider per day; 15 minutes long; a blocked schedule takes no work-in (living doc, verified in CI). [...]

Layer 3: session (lives for hours; dies when you close the session)

  • Decided at 10:20 a.m.: the new check goes into day_limits.ts, next to the count already there; no new file.
  • Dropped at 10:35 a.m.: doing the check in the controller. Reason: the convention puts the business rule in the domain.

Layer 4: turn (lives one turn; leaves after use)

  • Output of the last npm test: one red test, “rejects a second work-in for the same patient on the same day”.
  • The 12 lines of day_limits.ts pasted in for the edit. [...] One line of that annotation deserves attention because it looks like another one you have already seen in this chapter. “Dropped at 10:35 a.m.: doing the check in the controller” is layer 3, and it has to survive, because it is the only thing keeping you and the agent from reopening the same discussion at 3 p.m., spending the same time again and running the risk of deciding differently. Yesterday’s conversation, at the top of the chapter, was also a rejected path, and there the right answer was to keep it from entering. The difference is in the lifetime of the scope that produced it: the path rejected at 10:35 a.m. belongs to the task in progress and holds while the task lasts; yesterday’s belonged to a conversation that had closed. When a rejected path deserves to last longer than the session, it stops being an annotation and becomes an ADR, with the address chapter 10 gave it. What the layers let you decide The first decision is one of address. Every piece of information has a layer, and the layer says where it lives when it is not in the window. Layer 1 lives in a versioned file in the repository, read in

every session. Layer 2 lives in the feature’s artifacts, loaded when the task starts. Layer 3 lives in the session and, if it needs to last longer, it has to be written down somewhere before the session dies. Layer 4 lives nowhere: used, done. With that map, the announcement anti-pattern from chapter 12, that “HEADS UP: Friday deploys are suspended” living forever in the persistent file, gets a one-sentence diagnosis: it is layer 4 content written at the address of layer 1. You no longer have to judge line by line whether it deserves to be there; you ask how long it holds and the address settles itself. The second is the order of the cut. Every long session reaches the point where something has to go, and with no criterion the tool cuts by age, oldest first, which throws out exactly what you settled on at the start of the session, as chapter 3 showed. With layers, the cut has a direction: the turn goes first, then the session, and the session layer leaves summarized, never dropped in silence. Cutting by layer instead of cutting by age also has empirical support. The report “Context Rot: How Increasing Input Tokens Impacts LLM Performance,” published by Chroma in 2025 (research.trychroma.com) and detailed in chapter 5, measured replications of long conversations where the models did better receiving only the relevant excerpt of the history than receiving the complete history, both carrying the same information. Discarding the volatile layer is not controlled loss; in a large context, it is a gain in quality. The third is diagnosis. When the answer comes back wrong, you have a new question to ask before cursing the model: which layer did the information that produced this error come from? If it came from layer 1, you have a wrong line in a file that enters every session on the team, and the fix is worth weeks. If it came from layer 4, as on Wednesday, the fix is one of admission, deciding that this material does not enter again. If the right information was there and got lost anyway, you are facing the

position-and-volume problem chapter 5 measured, and the remedy is a different one. Three causes, three different fixes, and without the layers all three turn into the same generic complaint that the AI is no good. “This is bureaucracy for a twenty-minute session” The objection is fair and you will hear it from anyone on a deadline: nobody is going to stop and classify context by lifetime before asking for a ten-line adjustment. Nobody is, and the chapter does not ask for that. The classification is not one more step in your day; it is the name of what you already do by default, and its cost shows up once, when you decide where each kind of information lives. After that, the twenty-minute session inherits the finished work: layer 1 is already in the file the tool loads on its own, layer 2 is already in the feature’s artifacts, and what is left for you to assemble is the smallest part. In the session where everything fits with room to spare and nothing goes wrong, the layers charge nothing and are not missed. They charge in the session that went wrong, and there the alternative to having a vocabulary is rereading sixty thousand tokens looking for where an invented parameter came from. It is worth saying what this chapter assumes is already done. Layers 1 and 2 are selection, not writing: they choose among the artifacts Part II told you to build, the verified living documentation, the ADRs, the conventions, the task spec and the persistent file. If those artifacts do not exist, the technique still works as a mental model, and it degrades in a predictable way: you start filling the two layers from your head, in every session, paying for the same work again and introducing variation each round, because the version of the work-in rule you remember

today is not the one from last Thursday. With no durable source, the project layer turns into folklore, and folklore in the position of highest attention in the window is what chapter 12 called poisoned context. With the layers named, you know what exists and how long each thing lasts. What you still do not know is how much of each layer fits in this task. A complete layer 1 is the whole conventions document, which has twenty items when you need three; a complete layer 2 is the whole living documentation, when the task touches a single flow. A full layer is still a full window, and the decision of which subset enters, in what order and in what position, is not settled by lifetime: it depends on the task, and there is research showing that the position of what you pasted changes the odds of the model finding it. Choosing the minimum and putting it where it works is the next operation, and it is called context packing.

Powered by TurnKey Linux.