Nie możesz wybrać więcej, niż 25 tematów Tematy muszą się zaczynać od litery lub cyfry, mogą zawierać myślniki ('-') i mogą mieć do 35 znaków.

11KB

Context Engineering — Chapter-03: How LLMs use context

  • Source: /library/Context Engineering/source-file.pdf
  • PDF pages: 16–22
  • Pages without text: none

How LLMs use context On Tuesday you asked the AI assistant for a spreadsheet import function and got back clean code, in your project’s style, with the error cases handled. On Thursday you asked for practically the same thing and got a loose script, with generic names, ignoring the conventions the assistant itself had followed two days earlier. The request was the same. The model was the same. The result, the opposite. The most common explanation you will hear is some variation of “that is just how AI is, a lottery.” That explanation is comfortable and wrong. There is a concrete variable that changed between Tuesday and Thursday, and it has a name: the context. On Tuesday the conversation already held pieces of your code, the discussion about the error pattern and two of your own examples. On Thursday you opened a new session and sent the request cold. The model did not get worse. The information it received got worse. This chapter sets the mental model that holds up the whole book: what exactly the model sees when it answers, and what it does not see at all. Without that, every technique in the parts ahead turns into a memorized recipe. The model only sees the input Start with the term that gives the book its name. Context is everything the model receives as input in one call: your question, the conversation history, the system instructions, the pasted

files, the tool results. All of it, concatenated, forms a single block of text that the model reads at once. In Tuesday’s example, the context was your request plus the code excerpts and the earlier discussion; on Thursday it was only the request. Every time the model processes a context and produces an answer, an inference happens. An inference is one call to the model: text goes in, text comes out, and nothing else takes part. There is no side channel through which the model consults your repository, your intentions or yesterday’s conversation. If the information is not in that call’s input, for the model it does not exist. That sounds obvious written this way, but almost nobody works as if it were true. When you complain that “the AI should know” your project uses a certain pattern, you are crediting the model with knowledge that never came through the only door there is: the input. The model should not know. You should have told it. The practical consequence flips the usual question. Instead of “why did the model get it wrong?,” ask “what was in the input that would justify the right answer?” In most of the frustrating sessions you have had, the true answer is: nothing. Thursday’s cold request did not carry the project’s error pattern, so the model picked some pattern. From its own point of view, Thursday’s answer was as good as Tuesday’s: coherent with the input it received. Attention: how the model weighs what you sent Inside one inference, the model does not treat the context as a uniform bag of words. The architecture behind today’s large language models (LLMs), described by Vaswani and coauthors in the 2017 paper “Attention Is All You Need” (arXiv:1706.03762),

turns on a mechanism called attention. For the purposes of this book, the mental model is enough: when it generates each word of the answer, the model assigns weights to every piece of the input, deciding how much each piece influences the next word. Pieces with high weight pull the answer; pieces with low weight barely take part. A minimal example. Suppose the input: “Our backend is in Go. Always answer with examples in the backend’s language. How do I open a file?” When the answer is generated, attention links “the backend’s language” to “Go,” and the example comes out in Go. If the sentence about the backend were absent, the weight would spread to the rest and the model would pick the most likely language given everything else in the conversation, maybe Python. The answer changes without the final question having changed a single letter. Two consequences of that mechanism matter to you. The first: everything in the context takes part in the contest for attention, including what you pasted without thinking. That 300-line log you dropped into the conversation to illustrate an error is still there: it still receives weights and still competes with your instructions at every word generated. The second: attention is a statistical mechanism, not an exact search. The model does not “find” your instruction the way grep finds a string; it weighs your instruction against all the rest. Instructions can lose the contest. Chapters 4 and 5 show when and why that happens more and more often. If you want to go down to the real mechanism, with the matrices and the attention heads, the 2017 paper by Vaswani and coauthors (arXiv:1706.03762) is the primary source. To use AI well, the mental model above is enough, and this book does not go past it.

Nothing survives between calls One piece is missing, and it is the one that knocks down the most expensive illusion: that the model remembers. An LLM is stateless between calls: it keeps no state. Once an inference ends, the model retains nothing of what it processed. The public documentation of the chat application programming interfaces (APIs) of the major providers (the docs for Anthropic’s Messages API and for OpenAI’s API, in 2026) describes the same contract: every request sends the full list of messages in the conversation, and the server answers that list. There is no live session on the other side, no “brain” that follows you from one question to the next. There is a function: context in, answer out, done. If it helps, picture a peculiar call center. Every time you call this company, whoever answers is a completely different person, with no access at all to what you dealt with on earlier calls: no customer database, no history, no “as we discussed yesterday.” Everything that agent knows about your case is what you say on this call. In return, this is the best-prepared agent on the planet: every language, every framework, every pattern ever published. And it is exactly that breadth that creates the problem. Faced with a vague request, the agent has no way to know which of the thousand correct answers on hand is the right one for your case, so it picks the answer that is most likely in general, which is rarely yours. All the knowledge in the world, with no focus, produces a generic answer; the focus is what you bring, in what you say during the call. Every call to the model is that phone call: it starts over from zero, with someone on the line who knows everything and remembers nothing.

“But the chat does remember the conversation,” you will say, “it answers my second question knowing about the first.” It answers because the chat interface resends the whole conversation with every message you send. The memory you notice does not live in the model; it lives in the text the tool piles up and resends. It is a legitimate stage trick, and chapter 3 takes it apart in detail, with a real transcript of the point where it breaks. For now, hold on to the contract: one call, one context, one answer, no residue. That contract explains the Thursday at the start of the chapter in full. The new session had no access to Tuesday’s session, because there is no place where Tuesday could have been kept. You did not lose model quality from one day to the next; you lost the context, and the quality went with it. Where the context hides Before closing the mental model, a second illusion deserves to be undone: that the context is only what you type. In practice, the text you write tends to be the smallest part of what the model receives. When you use a coding assistant, the tool assembles the input on its own before calling the model. It usually includes a system instruction (the text that defines the assistant’s behavior), project configuration files you may not even remember exist, excerpts from the files open in the editor, results of searches the tool itself ran and the output of every command it executed. You type one line; the model receives tens of thousands of words. Try it in your own tool: look for the option that shows the request it sends or how much the session has consumed. The first time you see the whole package tends to be uncomfortable, like

opening the payload of a request you thought was lean and finding megabytes of extras. And each of those extras, you now know, competes for attention with your instruction. That is not a flaw in the tools; it is their job. Assembling context automatically is what makes a coding assistant more useful than a bare chat. But it hands you a new responsibility: knowing what is being assembled on your behalf. If you have never looked at the real input, you have no way to diagnose why the output came out wrong. Throughout the book, “look at the context” will show up as the first step of almost every diagnosis, the same way “look at the log” is the first step of almost every production investigation. What changes in your practice Put the three pieces together. The model only sees the input. Inside the input, attention weighs each piece, and everything competes. Between calls, nothing persists. Out of those three sentences comes a working definition that the rest of the book only refines: the quality of the answer is a function of the quality of the information present in the context of that call. Notice what that definition does to your room to maneuver. You do not control the model’s weights, you do not control the training, and you do not control the architecture. You control one single thing: what goes into the context. That single thing determines, more than any other variable within your reach, whether you get Tuesday’s code or Thursday’s script. Here is my own position: after years of using AI in production every day, I have not seen any adjustment of tool, model or phrasing return as much as treating the input with the same care I give a public interface. That practice is what this book systematizes.

One dimension is still missing, and this chapter treated it as abstract: the input has a size, and that size has a limit and a price. The model does not read “text” but tokens, and the context window that receives them is a finite resource. The next chapter defines those two measures, because without them you cannot reason about what fits, what costs and what stays out.

Powered by TurnKey Linux.