# Context Engineering — Chapter-24: Context isolation - **Source**: /library/Context Engineering/source-file.pdf - **PDF pages**: 204–216 - **Pages without text**: none --- Context isolation Thursday, and the day arrives with three things. The migration that renames the schedule’s time field has to run before the end of the week. The Vila Nova Clinic’s clinical coordinator asked for a utilization report by provider, which the front desk wants by Monday. And the work-in, one of the extra appointments squeezed into a full schedule, still needs its position inside the block settled: that job stalled yesterday and is still waiting for its last test case. You handle all three in the same session, because all three are VilaSchedule and the window is already warm. You start with the migration, paste the schema, discuss the field’s new name, settle on scheduled_start . You move to the report, sketch the utilization query and find out that work-ins and appointments have to be added up with different weights, because one lasts fifteen minutes and the other lasts thirty. You go back to the work-in and ask for the last test case, the one for the full block. The test comes back with two things that do not exist. It reads scheduled_start from a work-in whose field is still called time , because the migration has not run anywhere, and it calls dayUtilization , a function that exists only in the report sketch, three turns earlier, inside another slice. You fix both and ask for the report broken down by provider. Back comes a query that leaves the lunch break intervals out of the calculation, with a polite note explaining that a work-in does not run into the break. The sentence is true: it is yesterday’s decision, and it belongs to the work-in task. In a utilization report, it hides from the clinical coordinator exactly the part of the day the coordinator wants to look at. Notice what does not explain these errors. Context was not missing: the three tasks were in the window with current, checked material. It was not a badly assembled packet in the sense of chapter 17, because each of the three packets, looked at on its own, is right. It was not a bad summary from chapter 20, because nothing has been compressed yet, and it was not an outdated statement from chapter 19, because everything the agent said is true somewhere in VilaSchedule. What you have is a flat window with three tasks in it. Every piece of context is good for one of them and a distractor for the other two, and the model has no way of knowing which of the three each sentence belongs to, because all three arrive looking the same. One context per task Context isolation means giving each task a window of its own, with a packet of its own, that does not see the other tasks’ history and hands back a small result to the session that opened it. The principle fits in four words: one context per task. What makes an isolated context work is not the tool that opens it; it is the text that draws its boundary, and that text has a name in this book: the subtask contract, the sheet that says what that context receives, what it returns and what it does not need to know. Before going on, a note on vocabulary, because the word contract is already spoken for elsewhere on this shelf. In the previous book in this trilogy, FOCUS Architecture (https://books.kodel.com.br/en/books/focus/), a slice’s public contract is the surface it lets a neighbor call, and the rule of thumb there is that a slice talks to a slice through the front door, never by importing a neighbor’s internal file. That one lives in the repository, holds for everybody and is verified by a tool, as chapter 14 showed. A subtask contract is another thing and runs on another clock: it is session text, written for one specific task, born when you split the work and dead when the subtask delivers. One governs the code; the other governs a window. The mechanism that instantiates the principle today has a name of its own. In 2026, agent tools call it a subagent: the main session opens a child context, hands it a request, lets it work on its own and closes it, taking back only the result. The name belongs to this moment and the button moves with every version; the observable effect is what matters here, and it does not depend on the tool. Two windows open side by side, each with its own packet, are already one context per task, and you have been doing that by hand since before subagents existed. Which tool offers which control over child contexts is Part IV’s business. Anthropic described the arrangement in “How we built our multi-agent research system” (2025, anthropic.com/engineering): a lead agent breaks the question down, opens subagents that search in parallel, each with its own window, and gets back the condensed finding instead of the path it took. They report a clear gain over the single-agent baseline on their internal research evaluation, and they report the price too, on the order of fifteen times more tokens than an ordinary conversation. And they add the caveat that matters most to you: the gain shows up on a task that splits into independent searches, and not on a task whose parts depend on one another, which describes a good deal of the work of writing code. When splitting is worth the coordination cost Splitting exacts a fixed price, and the price has three parts: writing the input contract, reading the return and reconciling what came back with the main session. None of them goes away with a better tool. The criterion below is my own opinion, formed by splits that went wrong, and it has four conditions that have to hold at the same time. Fail one, do not split. The first is the disjoint diff. List, before you start, the files each task is going to write. If the lists overlap on so much as one file, the two stay in the same window: two contexts editing the same file produce a write conflict in the best case and a silent overwrite in the worst. That condition is only answerable because Part II gave you boundary lines: with chapter 14’s front door in place, you know where a task is allowed to write. The second is the small return. What comes back to the main session has to fit on one page and has to be the result, not the path it took. If the only way to use the subtask’s work is to load its whole session back in, you have split nothing; you have only postponed the bloat. A small return is what keeps coordination cheap, and it is the condition most people forget. The third is the closed shared decision. No project decision that holds for both tasks can be open at the moment of the split. If the field’s name, the format of the return or the rule both of them consult is still under discussion, close it first and write it where the decision can be read, or do not split. The fourth is the contract you can write today. Can you say, right now, before the subtask starts, what it receives and what it returns? If you cannot, the problem is not one of context; it is one of spec, as chapter 8 already pointed out, and no split fixes that. A subtask that only defines itself while it runs turns into a round trip, and every round trip pays the coordination cost again. With all four conditions met, one piece of arithmetic is left, and it decides the borderline case. The split is worth it when the subtask is several times bigger than its contract. Sweeping the whole repository to answer one question fits in three lines of contract and eats dozens of files: split. Renaming a function fits in one line of request and the contract would be the size of the task: do not split, because you would write the work twice. When in doubt, the default is the single window; isolating is a justified exception, not a default stance. Thursday, split again Run the criterion on the three tasks from the opening and it decides on its own. The schedule migration and the work-in adjustment fail the first condition before the second question: the work-in test reads the field the migration renames, and the repository mapping shows up in both lists. They fail the third as well, because when you started the day the field’s name was still open. The right split between those two is not in space; it is in time: close the name, run the migration, and only then go back to the work-in, in the same window, with the field already existing. Sequencing is not defeat; it is the recognition that one task is the input to the other. The utilization report passes all four. Its diff lives entirely in src/features/reports/ , a slice that already exists and that talks to scheduling, appointments and work-ins through each one’s front door. The return fits on one page, because what the main session needs to know is which files were created and what was missing at the doors it queried. Two decisions hold for both sides: a canceled work-in does not count toward utilization, and the report covers the closed day. You close both in thirty seconds, before splitting. And the contract can be written today, because the clinical coordinator said what the report has to show. A fourth task shows up that nobody asked for, and it is the cleanest case of all. Chapter 19 left one question open: where else in VilaSchedule is there code that assumes two appointments at the same time? Answering that means reading dozens of files, following imports, opening old tests, and handing back six lines: file, line and the suspect passage. Empty diff, minimal return, no shared decision, a three-line contract. It is the shape of work where splitting pays best, and not by accident: it is exactly the breadth-first search Anthropic describes as the success case of the arrangement. Doing that sweep in the work-in window would fill the session with forty files the work-in task does not use, and chapter 5 already measured what that does to the next answer. The subtask contract The utilization report’s contract, excerpted, with each [...] marking what did not fit on this page: # Subtask contract: utilization report by provider [...] ## What it receives (input packet, assembled before it starts) Opening the packet, what cannot be violated: - Business rules live in the domain, never in the controller (project conventions). - A slice talks to a slice through the index: `reports` queries `scheduling`, `appointments` and `workins` through each one's `index.ts`, never through an internal file. [...] - Decisions already closed in the main session that hold here: a canceled work-in does not count toward utilization; the report covers the closed day, never the current one. [...] ## What it returns (fixed format, fits on one page) 1. The files created or changed, one line per file. 2. The questions it asked at each front door and what was missing in the answers. 3. The decisions it had to make on its own, with the reason for each. 4. What it assumed for lack of information, marked as an assumption. [...] ## What it does not need to know - The discussion about the work-in's position inside the provider's block, which is running in the main session. - The schedule's schema migration under way: the report asks the slices' index and knows no table. - The main session's history, its test output and the dead ends it has already abandoned there. ## Write boundary It creates and edits files only inside `src/features/reports/` and the tests next to them. If it needs any change in `scheduling/`, `appointments/` or `workins/`, it stops and hands the request back instead of editing. [...] Three sections of that text do the heavy lifting. The one about what it does not need to know is the strangest to write and the most valuable: it is the list of true things you are barring from coming in, and every line of it matches one of the morning’s errors. The write boundary is the criterion’s first condition turned into an instruction, and what makes it verifiable is not the agent’s goodwill; it is chapter 14’s lint waiting on the other side. And the fixed format of the return is the second condition: by asking for files, questions, decisions and assumptions, you get one page instead of a transcript, and the marked assumptions become your checklist when the result arrives. Notice what the contract inherits instead of repeating. The standing rules come in as an excerpt from the living doc, as chapter 17 taught you, and not as a paraphrase of your own. The main session’s closed decisions are copied in from chapter 20’s anchor sheet, which already existed. The contract is chapter 17’s context packing applied to a smaller task, and that is why it costs less than it looks: you are not writing new material; you are cutting from what already has an address. Two subtasks at once, each on its own ground So far isolation has been treated as a split, and the split as a sequence: you open the subtask, it works, the return comes back. But the four-condition criterion has a consequence that deserves to be said out loud, because it is where the arrangement pays the coordination cost with the most room to spare: two subtasks that each pass the criterion with respect to the other can run at the same time. The diff is disjoint, the shared decisions are closed, each one has its own contract and its own return; nothing in the arrangement demands that the second wait for the first. The Anthropic piece cited earlier in this chapter described subagents searching in parallel; your version, as someone who writes code, is the utilization report and the sweep for expired rules running the same afternoon, each in its own window, while your main session goes on with the work-in discussion. Running in parallel exacts a price that running in sequence never charged. In sequence, two subtasks with disjoint diffs can share the same working directory, because one finishes before the other touches disk. Running at once, they cannot: even with disjoint target files, two contexts in the same working tree fight over the branch, the git index and the build state, and the first git checkout from one pulls the rug out from under the other. The answer that became the standard in 2026 is to give each context a working copy of its own with git worktree , git’s native mechanism for materializing more than one working directory from the same repository, one branch in each, without cloning history. Each subtask edits in its own worktree; what goes back to the main repository goes back by the road all code travels, the merge. The naive alternative, cloning the whole repository per subtask, works but duplicates history and configuration at every split; the worktree exists exactly so you do not pay that. The tools absorbed the pattern: in July 2026, Claude Code creates a worktree per parallel terminal session and per isolated subagent, and Cursor gives each agent in multi-agent mode a workspace of its own via worktree; other tools do the same. The principle came before all of them: isolated context with isolated writing, and the contract’s write boundary now standing on physically separate ground. “The subagent loses sight of the whole” The strongest objection to this chapter does not come from people who never split. It comes from people who split and got burned. Walden Yan, of Cognition, published the most direct argument against the arrangement in 2025, in “Don’t Build Multi-Agents” (cognition.ai/blog): every action carries an implicit decision, and two contexts working apart make different implicit decisions, so the pieces come back correct but do not fit together. The example is building a clone of a game with two subagents: one hands back a background in one visual style, the other hands back the character in another style, and putting the two together gives you two well-executed pieces and one incoherent result. The recommendation drawn from that is to work with a single thread and to share the whole trace of what happened, not individual messages, even if that costs window. The diagnosis is right; the general conclusion drawn from it is where I part ways. The game clone scene fails the criterion’s third condition before it starts: the visual style is a shared decision nobody closed, and that is why each context invented one. It fails the first as well, because the background and the character meet on the same screen and often in the same file. Splitting there was a mistake, and the criterion rules it out. What the argument does not show is the report case, where the shared decision was closed and written down before the split, nor the sweep case, where there is no shared decision at all because nothing is written. What is left of the objection still stands, and it is worth recording instead of hiding it: the isolated context really does not see the whole, and that is what it is for. You are the one who needs to see the whole, and the contract is the instrument for it. It carries in the decisions that already hold, it draws the boundary and it forces the return to declare assumptions. Outside those conditions, this book’s position is the same as Yan’s: single window, single thread, and the discomfort of carrying too much context instead of the damage of pieces that do not fit. “Re-explaining the context to each one is expensive” The second objection is one of arithmetic, and the number behind it is real: a multi-agent system eats far more tokens than a conversation, and it is Anthropic itself that publishes the order of magnitude. If every isolated context has to receive conventions, standing rules and decisions before it starts, you pay for the whole packet several times instead of once. The arithmetic is wrong in two places. The first is what it compares. The subtask contract is not the session rewritten for another reader; it is chapter 17’s minimum packet for a smaller task, and it would be paid either way, because the task would exist inside the single window too. What the split adds to the cost is the return and the reconciliation, and the criterion’s second condition exists to keep both small. The second place is what it ignores on the other side of the scale. In a window with three tasks, chapter 3 already explained what happens: the whole history travels again on every turn, so the report sketch is resent on every question about the work-in, and the migration schema travels along with the test. You were already paying for the three tasks on every turn. The difference is that you were also paying in wrong answers. It is worth saying what this chapter assumes is in place. The criterion’s first condition depends entirely on chapter 14’s boundary lines: with no declared front door and no lint rule holding it up, you have no way to state that two diffs are disjoint, and the isolated context writes wherever it can reach. And the contract depends on Part II’s durable sources, the living doc that is verified, the architecture decision records (ADRs) and the conventions, because every line of it is an excerpt from an existing artifact. Without them, the technique degrades in a specific way: every split turns into typing from memory, and three isolated contexts receive three slightly different versions of the same work-in rule, none of them checked. That is where the criticism about sight of the whole lands squarely, because the whole was written down nowhere. The packet that fits in no window at all With the criterion in hand, Thursday turns into four contexts and the main session stops mixing subjects. Except that one of the contracts does not close, and its problem is not one of splitting. The utilization report has to classify each interval according to the clinic’s care policies: what counts as a no-show, what counts as a schedule block, how much grace time each insurance plan accepts. That lives in a document of hundreds of pages that the clinical coordinator updates every month. It does not fit in the input packet, it does not fit in the isolated context’s window and it does not fit in the main window, and splitting the work into more contexts does not shrink the document by a single line. The question changes axis: instead of how many windows the work uses, it becomes what goes into the window whole and what stays outside to be fetched when the question comes up. It is the choice between embedding and retrieving, and it is the next chapter’s subject.