No puede seleccionar más de 25 temas Los temas deben comenzar con una letra o número, pueden incluir guiones ('-') y pueden tener hasta 35 caracteres de largo.

9.8KB

Context Engineering — Chapter-07: Context rot: why large contexts degrade quality

  • Source: /library/Context Engineering/source-file.pdf
  • PDF pages: 43–48
  • Pages without text: none

Context rot: why large contexts degrade quality The logic looks flawless: the model’s window, the ceiling chapter 2 measured, holds a million tokens, your entire project comes to three hundred thousand, so you paste the entire project and never again field a question about a missing file. You do that, and the first answers are impressive. Then you ask for a change that depends on a rule defined in a file in the middle of the paste, and the assistant reinvents the rule, gets it wrong and says so with full confidence. The rule was there. You check: it is there, literally, in the context. The model simply answered as if it were not. Chapter 2 warned you that fitting and working are different things; chapter 4 showed the cycle filling the window on its own. This chapter brings the part that was missing: the evidence. Quality degradation in large contexts is not just your impression, nor forum folklore. It is a measured phenomenon, replicated and published, with a name, a curve and evaluation methods of its own. Knowing the literature changes your diagnosis: you stop asking “why is the model dumb?” and start asking “where in my context is the information dying?” The U-shaped curve: “lost in the middle” The most cited result in the literature came out of Nelson Liu’s group at Stanford, in the paper “Lost in the Middle: How Language Models Use Long Contexts,” published in the

Transactions of the Association for Computational Linguistics (TACL) in 2024 (arXiv:2307.03172). The experiment is elegant in its simplicity: they give the model a set of documents and a question whose answer is in exactly one of them, and they vary only the position of the relevant document inside the context. If the model used the context uniformly, position would not matter. It matters enormously. Performance traces a U-shaped curve: high when the relevant information is at the start of the context, high when it is at the end, and visibly worse when it is in the middle. In some configurations of the study, the model with the answer in the middle of the context did worse than the same model with no document at all, answering from training memory. Hold on to that picture: the middle of your context is a shadow zone. The rule the assistant reinvented at the top of this chapter did not vanish; it was buried in the trough of the curve. The finding does not depend on one specific model: in the paper itself, the curve shows up in models from different vendors and at varying context sizes, and the Chroma report you will meet later in this chapter finds positional degradation again while measuring 18 models from a later generation. That makes the phenomenon structural, not a defect the next version will fix. Translate the curve into your session. The start of the context is the system instruction and the very beginning of the conversation; the end is your last message. The middle is everything else, and chapter 4 showed the cycle pushing everything there: each new turn displaces the previous one further from the ends. That architecture decision made 30 messages ago now lives in the worst neighborhood in the context. Needles, haystacks and the test that became a

standard Before academia formalized the curve, practitioners had already been measuring the problem with a homegrown test that became an industry standard: the needle in a haystack, published by Greg Kamradt in 2023 as an open repository (github.com/gkamradt/LLMTest_NeedleInAHaystack). The recipe: hide a random sentence (the needle) at a controlled position in any long text (the haystack), ask the model for the needle, repeat while varying the position and the size of the haystack, and chart the map of hits. Kamradt’s maps showed the same pattern as the literature: retrieval degrading as the context grows and as the needle sinks into certain regions. The test matters for two reasons. First, because it is the number vendors started displaying (“99% on needle in a haystack”) when they announce giant windows, and now you know how to read that number for what it is: the grade on one specific exam, not a general guarantee about long context. Second, and more important, because the exam is far too easy for what you do for a living. Finding an out-of-place sentence planted in a text that never mentions it is search, almost a grep; your real work requires the model to connect, synthesize and reason about what it found. A model can ace the synthetic haystack and keep stumbling in your Monday session. Which brings us to the study that measured exactly that. Context rot: degradation in tasks that ought to be trivial In 2025, Kelly Hong, Anton Troynikov and Jeff Huber at Chroma published a technical report on how a growing input degrades the performance of large language models (LLMs). The report is

“Context Rot: How Increasing Input Tokens Impacts LLM Performance,” available at research.trychroma.com, and it evaluates those 18 models, the largest each vendor offered at the time, from Anthropic, OpenAI and Google. Their question: holding the task fixed and trivial, what happens when only the size of the input grows? The answer: performance drops, consistently and measurably, even in tasks an intern would solve before their first coffee. Replications of the needle in a haystack with needles that require a minimal inferential step (the needle says “I wrote about that in chemistry class” and the question asks about “high school”) degrade much faster than literal search. Distractors, wrong answers planted to resemble the needle, make everything worse as the context grows. And the most counterintuitive finding: in replications of a long conversation, the models did better when they received only the relevant portion of the history than when they received the complete history, even though that history contained the same information. More context, with the answer unchanged, produced a worse result. The name the authors gave the phenomenon, context rot, stuck, and I use it here. The underlying explanation is the one you have been carrying since chapter 1, now with engineering vocabulary: attention is a finite budget. The article “Effective context engineering for AI agents,” published by Anthropic in 2025 (anthropic.com/engineering), puts it this way: every new token dilutes the attention budget available to all the others, and context should be treated as a resource to curate, not as a warehouse. The window is how much you can store; attention is how much the model can actually use. The former has doubled in size several times in recent years; the latter is still the bottleneck. Diagnosing rot in your session

The literature gives you three objective symptoms to look for in a degraded transcript, and they are worth looking for on your next bad afternoon. The instruction is present and ignored: the rule is in the context, you check, and the answer violates it. A classic symptom of the middle of the curve, like the session in chapter 3, where what you agreed on in message 7 died in the shadow long before any window truncation, the cut the tool makes when the history no longer fits. A distractor wins: the answer uses the wrong version of a piece of information that exists in two versions in the context (the discarded approach, the old code before the refactor). It is the effect Chroma measured, and chapter 4’s cycle manufactures distractors all day, because nothing that goes in comes out. Quality drops with the age of the session, with no change in the kind of request: rot in its pure form, performance as a decreasing function of the size of the input, a small-scale replica of the report’s chart. And there is a cheap test that turns suspicion into evidence, with no tooling whatsoever: the clean-session A/B test. When an answer is bad in a long session, copy only the essentials (the question, the code that matters, the rule that matters) into a fresh session and repeat the request. If the answer from the clean session is visibly better, you have just reproduced the Chroma experiment at your own desk: same relevant information, less haystack around it, better result. Do that three or four times and you will never again need a paper to convince you that swollen context degrades; you will have seen it in your own code. The test also works as a yardstick for deciding when a session should be closed: if the clean A/B wins by a wide margin, the old session has rotted beyond repair.

Notice what the three symptoms have in common: none of them produces an error, a warning or a log. The call returns success, the text reads as fluent and confident, and the degradation only shows up if you are measuring quality on your own. Context rot is a silent failure, the worst kind of failure to debug. A personal opinion: after I learned about the U-shaped curve, I stopped fighting with degraded sessions and started closing them guilt-free, the same way I restart a process with a memory leak instead of arguing with it. The session is not a relationship; it is a buffer. You do not fix a rotted context with one more instruction at the end, which only pushes more material into the middle; you fix it by starting over smaller. One more thing makes this worse, and it closes this part of the book. You pay for everything that rots in your context: every token in the shadow of the middle, every distractor, every dead log from the cycle shows up on the bill, per call, at list price. Quality falling and the bill rising are the same phenomenon seen from two angles, and the next chapter does the math on the second angle, in dollars, with a formula you can redo with your own numbers.

Powered by TurnKey Linux.