Du kan inte välja fler än 25 ämnen Ämnen måste starta med en bokstav eller siffra, kan innehålla bindestreck ('-') och vara max 35 tecken långa.

8.7KB

Context Engineering — Chapter-06: The context cycle

  • Source: /library/Context Engineering/source-file.pdf
  • PDF pages: 37–42
  • Pages without text: none

The context cycle Nine on Monday morning, and you open a session to chase down a bug. The first hour is great: direct answers, code on target, the assistant seems to read your mind. Around eleven, something sours. The answers turn wordy, the assistant revisits decisions already made, goes back to an approach you and the assistant dropped at nine thirty, and mixes the current bug with a refactor you mentioned in passing. You changed nothing about the way you ask. The session just… aged. The first three chapters gave you the pieces to explain that: the answer is a function of the input, the input is measured in tokens inside a finite window, and the “conversation” is the history sent again on every call. This chapter assembles the pieces into the picture that was missing: an AI session is a feedback loop, and feedback loops have a property every engineer respects: what goes into them does not come out on its own. The shape of the cycle Follow one turn of conversation, from click to click. You write a message. The tool assembles the input: system instruction, plus the accumulated history, plus your new message. The model processes that input and generates an output. The tool shows the output to you and appends it to the history. On the next turn, everything starts over, with one difference: the history now contains the output of the previous turn. In diagram form:

The arrow that says “becomes history” is this whole chapter. The output of every turn becomes input for every turn that follows. The model at eleven on Monday is not answering your question; it is answering your question added to two hours of everything already said, by you and by itself. The conversation feeds on its own production, and the application programming interface (API) contract the previous chapters cited (the complete list of messages sent again on every call, per the public docs of the chat APIs in 2026) guarantees that nothing escapes the circuit on its own. Notice a consequence that usually escapes attention: the model reads itself. Each of its answers comes back as input, with the same standing as any other text in the context, competing for attention with your instructions. If an answer came out verbose, the verbosity is now part of the context and pulls the next

answers toward the same tone. If an answer came out with an error nobody corrected, the error circulates as if it were an established fact of the conversation. The cycle does not tell signal from noise; it only accumulates and resends. Where the cycle swells If the cycle only accumulated your questions and the useful answers, the growth would be slow and nearly harmless. A real session swells much faster than that, and at predictable points. They are worth mapping, because they are the same in every tool. The first point is the diagnostic paste. Stack trace, deploy log, dump of an API response: you paste it to illustrate one specific problem, you solve the problem in two turns, and the paste stays. A 300-line log from a problem solved at nine oh five is still circulating at noon, reprocessed and competing for attention on every turn, hours after it lost all usefulness. By chapter 2’s yardstick, that is tens of thousands of tokens of dead weight per paste. The second is tool output. Modern coding assistants run commands, read files and run tests, and every result goes into the history: the complete directory listing, the 800-line file read to answer a 10-line question, the verbose output of the test runner with its 40 progress bars. You typed none of it, but it is all in the cycle, and you pay for it on every turn that follows. The third is the model’s own prose. Assistant answers tend to run long: they recap the request, explain the obvious, offer alternatives nobody asked for. Every decorative paragraph of every answer becomes permanent input. In long sessions, a meaningful share of the history is the model quoting, summarizing and repeating the model, layer upon layer.

The fourth is topic drift. The refactor mentioned in passing, the side question about a library, the old bug that came back into the conversation by association: every detour deposits the context of one subject into the cycle of another. The eleven o’clock session mixes the bug with the refactor because, in the input, the two subjects really are mixed, side by side, with similar weights, and the model has no way to know which of them is alive and which is residue. The arithmetic of accumulation The cycle has an arithmetic property worth seeing in round numbers, because it surprises even people who have understood the picture: the cost of a session does not grow with its length; it grows with its square. Suppose a well-behaved session, with no monstrous paste in it: each turn adds, between your message and the model’s answer, some 2,000 tokens to the history. On turn 1, the input holds 2,000 tokens. On turn 10, it holds 20,000, because it carries the nine previous turns. On turn 50, 100,000. Now add up what the model processed over the whole session: it is not the final size of the history but the sum of the inputs of every turn, 2,000 plus 4,000 plus 6,000, and so on. For 50 turns, that sum passes 2.5 million tokens processed, for a conversation whose text, read end to end, runs to 100,000. Every token you deposit in the cycle is not read just once; it is reread on every turn still to come. Redo the math with your own numbers, because it is grade- school arithmetic: if each turn adds T tokens and the session has N turns, the total processed is roughly T times N squared, divided by two. Doubling the length of the session quadruples the processing; the cycle reprocesses the 30,000-token paste you made on turn 5 of a 50-turn session 45 times. That is why the

difference between pasting a whole log and pasting the 10 relevant lines is not an aesthetic one: in the cycle, every excess is multiplied by the number of turns left. Keep that multiplication in mind. Chapter 5 shows what it does to quality; chapter 6 converts it into money. Reading a session as a cycle With the picture in hand, reread the Monday at the start of this chapter as an engineer, not as a frustrated user. At nine, the cycle was clean: system instruction, your description of the bug, little else. Small input, concentrated attention, sharp answers. At nine thirty, in came the stack trace and the discarded approach; the approach was discarded in the conversation but not in the input, where it is still present, with the same weight as any valid decision. At ten, the assistant read three whole files and ran the tests twice; four fat outputs went into the cycle. At eleven, the input of each turn is dozens of times larger than it was at nine, and your current question is a tiny slice of it. The assistant “revisiting” the discarded approach is not a regression of the model: it is the discarded approach, alive in the input, winning a contest for attention that got more tangled with every turn. The symptom you feel as the session souring is the sum of two effects the next chapters measure: quality drops because attention gets diluted in swollen context, and cost rises because every turn reprocesses the whole pile. Neither one is an accident; both are the physics of the cycle. And notice that none of it required bad faith or a glaring mistake from anybody: you used the tool exactly as it presents itself, and the swelling came along. The cycle degrades by default; keeping a session healthy is active work, and the next parts of the book exist for that work.

One personal opinion before I close: of the dozens of degraded sessions I have debugged, the cause was almost never an exotic one. It was ordinary accumulation, from the four categories above, that nobody looked at. The habit of asking “what is circulating in my cycle right now?” solves more bad sessions than any change of model. This chapter closes the mechanism of Part I: you know what the model sees, in what unit, with what limit, why memory is a resend and how the session feeds itself. What is missing is the evidence that accumulating costs you. The next chapter presents the research that measured the drop in quality in large contexts, including the effect with a diagnosis for a name, “lost in the middle,” and the phenomenon the 2025 literature named context rot. An uncomfortable spoiler: the window may well hold all your tokens without complaining; the model’s attention, as you are about to see in the data, does not, and it degrades long before any error shows up on your screen.

Powered by TurnKey Linux.