Du kannst nicht mehr als 25 Themen auswählen Themen müssen entweder mit einem Buchstaben oder einer Ziffer beginnen. Sie können Bindestriche („-“) enthalten und bis zu 35 Zeichen lang sein.

9.4KB

Context Engineering — Chapter-05: Memory and limits

  • Source: /library/Context Engineering/source-file.pdf
  • PDF pages: 31–36
  • Pages without text: none

Memory and limits Every long session with AI has its moment of betrayal. You set the rules at the start, the assistant confirmed them, the work flowed for an hour, and then, on a trivial request, it handed you code that violated everything you had agreed on, as confidently as ever. The instinctive reaction is personal: “it ignored me.” The real explanation is more clinical and more useful: nobody ignored you, because there was nobody there to remember. The excerpt below comes from a working session with an AI assistant. It is a reconstruction: I condensed some 60 messages from actual sessions in 2026 into one representative excerpt, and the full file says so. The mechanics it illustrates happen, to some degree, in every long session you have ever had: [message 45, user] Now create the order cancellation handler. [message 46, assistant] Here is the handler: async function cancelOrder(req, res) { const ord = await db.query(‘select * from orders where id = $1’, [req.param s.id]); if (!ord) return res.status(404).send(‘not found’); await db.query(‘update orders set status = $1 where id = $2’, [‘canceled’, req.params.id]);

res.send(‘ok’); } [message 47, user] You just broke the three rules we agreed on: single quotes, the name shortened to “ord” and an inline query in the handler, with no repository. Thirty-nine messages earlier, in message 7, the assistant had confirmed the three rules in writing. This chapter explains what happened between message 7 and message 46, because in that gap lives the difference between using AI and being used by it. The chat’s memory is a replay The previous chapters left the pieces on the table: the model is stateless, nothing persists between calls, and each call carries at most one window’s worth of tokens. Put the pieces together and the question is unavoidable: if nothing persists, how does the chat answer message 46 knowing what happened in message 45? The answer sits in the contract of the chat application programming interface (API), publicly documented by the vendors (the docs for Anthropic’s Messages API and for OpenAI’s chat API, as of 2026): every request sends the complete list of the conversation’s messages. When you type message 46, the tool does not send message 46; it sends messages 1 through 45, plus 46, all in a single input. The model reads that whole input from scratch, as if for the first time, because for it this is the first time. It always is. The chat’s “memory,” then, is the tool sending the history again: a text file that grows with every turn, one message from you plus one answer from the model, and is reprocessed in full each time.

No model is following your session. What exists is a session retold in full, hundreds of times, to a model with no memory. The illusion works because the replay is faithful. Until the day it is not. Where the illusion breaks If the tool resent the history intact forever, the illusion of memory would be perfect and this chapter would end here. It does not end because chapter 2 imposed a ceiling: the context window is finite. A real working session produces tokens at a rate you now know how to estimate: each answer with code, a few thousand; each stack trace you paste, a few thousand more; that 200-line log table you dumped into the conversation to diagnose a bug, tens of thousands. The session at the start of this chapter had all of that between message 8 and message 44. When the sum hits the ceiling, the tool has to decide what to do, and none of the options preserves the illusion. Tools generally do one of three things: cut the oldest messages, summarize the start of the conversation into a paragraph and discard the original, or some combination of the two. In all of them, something that was in the conversation drops out of the input. And chapter 1 already delivered the verdict on what is not in the input: for the model, it does not exist. Now you can reconstruct the betrayal of message 46 without a single metaphor. The three rules you agreed on lived in messages 6 and 7, the oldest point in the conversation. The session grew until it hit the limit. The tool cut or summarized the oldest stretch to make room, and the rules went with it, without warning, because no tool tells you what it discarded. On the next call, the model received a conversation that, as far as it could see, had never contained a rule about quotes, names or the repository

layer that keeps database access out of the handler. It did not break the agreement; the agreement never reached it. The confidence in the answer stayed the same because, from the model’s point of view, nothing was missing. There is a second, subtler failure mode, which does not even require overflowing the window: what you agreed on can sit in the input and still lose, in the competition for attention, to tens of thousands of more recent tokens. The rule is there, but it does not carry much weight. There is published research measuring where and how much that happens, and chapter 5 is about exactly that. What matters here is that both modes produce the same symptom on your screen: the assistant “forgets,” and the cause never surfaces. What about the tools that claim to have memory? You may object, and rightly so: plenty of tools advertise persistent memory. The chat that remembers your name between sessions, the coding assistant that keeps project preferences, the “memories” feature that summarizes old conversations. Does that contradict what this chapter claims? It does not contradict it; it confirms it. Open the documentation for any of those features and you will find the same architecture: the tool writes facts to its own storage (a file, a database) and, on every new call, injects the relevant facts into the context, along with the rest of the input. The memory lives outside the model and reaches it through the only way in: the input of the call. The model still has no memory. The note-taking happens in the tool, which hands the notes back before every call.

The distinction sounds pedantic, but it changes what you do. If memory is injected context, it obeys everything you have already learned about context: it takes up tokens in the window, competes for attention with the rest of the input and only works if the tool decides to inject the right fact at the right moment. When your tool’s memory feature “fails,” the investigation is the usual one: was the fact stored? Was it injected on this call? Did it arrive with enough weight to win that competition? Three questions, three possible points of failure, none of them mystical. Keep the general rule in mind, because it applies to every promise of memory you will run into: there is no model that remembers; there is context somebody assembled. The useful question is never “does this tool have memory?” but “what does this tool inject into the context, when, and how much of that do I control?” Work with the memory that exists, not the one you imagine The corrected mental model has practical consequences right away. First: stop treating the start of the session as a vault. Everything you establish in message 6 has an expiration date, because it is the first thing truncation takes. If an instruction has to survive the whole session, it has to live somewhere that gets resent every time, like the permanent instruction files the tools offer, or it has to be repeated when it matters. Repeating an instruction looks inelegant to anyone thinking about the don’t repeat yourself (DRY) principle; it is ordinary engineering to anyone who knows that the tool reassembles the input on every call.

Second: stop stretching sessions out of convenience. Every turn reprocesses the whole history, so a 300-message session carries the dead weight of the first 250 in every new question, paying in attention and, as chapter 6 will show, in money. When the subject changes, a fresh session with a short summary of what matters almost always beats the old session with everything in it. Third: when the assistant “forgets,” diagnose instead of swearing at it, because the symptom tells you the cause if you know how to read it. The right question is the one from chapter 1: was the agreement still in the input of this call? If your tool shows how much of the window is consumed, look. If the session stayed far from the ceiling, the problem is attention and not truncation, and the treatment is different. Telling the two cases apart is half of diagnosing any degraded session. One honest warning, and this one is mine: no technique in this book gives the model real memory, because there is nowhere to keep it, and I distrust anyone who promises otherwise without showing where the context is assembled. All context engineering does is decide, deliberately, what goes into the next call, instead of leaving that decision to a truncation algorithm that does not know your project. Resending the history explains the basic mechanism, and it opens a bigger question: if every output of the model feeds back into the input of the next call, the session is a cycle that feeds on itself, and everything that enters it (good code, a garbage log, idle chatter) goes around forever. The next chapter maps that cycle out in full and marks the exact points where it balloons, one by one.

Powered by TurnKey Linux.