Du kan inte välja fler än 25 ämnen Ämnen måste starta med en bokstav eller siffra, kan innehålla bindestreck ('-') och vara max 35 tecken långa.

13KB

Context Engineering — Chapter-04: Tokens and context windows

  • Source: /library/Context Engineering/source-file.pdf
  • PDF pages: 23–30
  • Pages without text: none

Tokens and context windows You ask the agent to analyze your project. It reads file after file, runs searches and dumps command output, and the session moves along fine until, with no warning, the tool announces it is going to “compact the conversation” or simply starts answering while ignoring instructions you gave twenty minutes ago. Nobody showed you what piled up, how much fit or how heavy each read was. You are negotiating with an invisible limit, in a unit you cannot see. The previous chapter established that the answer is a function of what is in the input. This chapter gives the two measures of that input: the unit it is counted in, the token, and the container that limits it, the context window. With both in hand you start to estimate what fits and to predict when it will overflow; chapters 5 and 6 turn the same measures into the cost of wasted room, in quality and money. The model reads tokens, not words A token is the smallest unit of text the model processes: a piece of a word, a whole short word, a punctuation mark or a space. Before the model processes anything, the context goes through a tokenizer, which slices it into those units and turns them into numbers. The model never sees letters; it sees the sequence of numbers the tokenizer produced.

The dominant slicing algorithm derives from byte pair encoding (BPE), described for use in language models by Sennrich, Haddow and Birch in 2016 (arXiv:1508.07909). The principle can be stated in one sentence: character sequences that show up often in the training corpus become a single token; rare sequences get broken into smaller pieces. “The” is one token. An identifier like calculateTotalWithDiscount becomes several. A minimal example, with the open-source tokenizer tiktoken, which OpenAI publishes on GitHub (github.com/openai/tiktoken): “the quick brown fox” comes to 4 tokens, one per word, because all four words are common. The same sentence in Portuguese, “a raposa marrom veloz,” comes to 7, because English dominates the training corpus of these tokenizers and words in other languages get sliced more often. Both numbers come from tiktoken 0.13.0, encoding o200k_base , measured in July 2026, and the older encoding cl100k_base returns the same pair. That asymmetry gives you the two rules of thumb you will use every day: in English, one token is roughly 4 characters, or about three quarters of a word; the same content in another language costs more tokens, on the order of 20 to 40% more in Romance languages and well above that in Japanese, Chinese or any other non-Latin script. They are approximations, and the only exact measure is running your own tool’s tokenizer, but to estimate orders of magnitude they are enough. Now apply that yardstick to your own day. A page of running text lands in the hundreds of tokens. A 300-line code file, a few thousand. The output of that build command the agent ran and captured in full, tens of thousands. And here is what the agent era changed: you are no longer the one pasting text into the conversation; the agent is the one piling it up, one Read at a time, one search at a time, one log at a time, and all of it goes into the same count. You do not need precision; you need to stop treating those reads as weightless. Do the exercise once, to calibrate your

instinct: take a file the agent read in your last session, count the characters and divide by four. Compare that number with the usage your tool reports for the session. After the first measurement, you will never again send an agent off to “read the whole project” without a second of hesitation, and that hesitation is exactly the habit this chapter wants to build. Every read has a size in tokens, and from the next chapter on that size takes center stage. What tokenization explains as a bonus Knowing that the model sees tokens, and not letters, undoes a few mysteries you have probably watched happen and chalked up to the model being dumb. Before the example, a reminder that defuses some frustration: a large language model (LLM) is not an intelligent entity in the sense the fluent conversation suggests. The whole mechanism is one thing: predicting the most likely text in the output given the text in the input. There is no understanding, no intention and no “somebody” on the other side who knows what they are saying; the intelligence you perceive is something you project onto it, an impression fluency creates. Whether that prediction amounts to some form of intelligence is an open debate among researchers, but what matters here is the mechanism. If you expect to be talking to something genuinely intelligent, you will expect the model to have abilities the mechanism simply does not have. The classic case: you ask how many letter “r”s there are in a word and the model botches a count a child gets right. The error is only shocking because of the expectation above; you assumed intelligence where there is text prediction. It stays frustrating until you remember that the model never saw the letters. The word arrived as one or two tokens, whole numbers in a sequence,

and what characters make up each token is not part of what the model processes directly. Asking it to count letters is like asking you to count the bytes of an image by looking at the photo: the information exists at some level of the representation, but not at the level where you operate. The same reasoning explains why swapping a word for a synonym sometimes changes the answer more than it should: different words slice into different tokens, with different statistical neighborhoods in the training data. And it explains why long code identifiers full of abbreviations use more tokens than clean names, because the tokenizer slices what it has never seen. None of those effects requires you to memorize the tokenizer’s vocabulary. What they require is that you remember there is a slicing layer between your text and the model, and that this layer has a countable cost. It is that countable cost that matters from here on. If every piece of text has a price in tokens, the next question is unavoidable: how many tokens fit in one call? The context window is the container The second measure is the limit. The context window is the largest number of tokens one call to the model holds, everything included: system instructions, conversation history, every file the agent read, every command output it captured and the answer the model is going to generate. The answer counts too, because the model generates token by token inside the same window it read the input in. And in the reasoning models that are the default at the major providers as of July 2026, the answer you read is not everything the model generated: before it come the thinking tokens, the internal draft the tool hides or summarizes, which takes up room in the window like any other generated

token. A ten-line answer may have cost a few thousand tokens of draft, and a budget that ignores that invisible portion runs out of window sooner than the math predicted. When the total gets close to the ceiling, something has to give: the call fails, the answer comes out truncated, or the tool compacts or discards part of the history to make room, as in the scene that opened this chapter. Of the three, the third is the most treacherous, and chapter 3 shows the damage it does. How big is the window? Here I refuse to print a table, on purpose. Window numbers age in months, and a book that pinned them down would be lying to you before its second printing. Take the order of magnitude, anchored in time: in 2026, the major providers’ frontier models (Anthropic, OpenAI, Google), the largest each one offers, have windows between hundreds of thousands and a few million tokens, and the exact numbers are on each model’s public page, one click away. By the time you read this paragraph, the values will have grown. The mechanics described here will not have. Do the math that matters: a window of hundreds of thousands of tokens holds roughly a few hundred pages of text or a small code project in its entirety. That sounds like plenty. And that is where the trap is, the one that separates the people who have read this book from the people who have read the marketing page. A big window is no license to fill it The natural reaction to windows growing is “great, now I can send the agent to read the whole repository and let the model sort it out.” That reaction assumes the window works like a disk, where taking up 10% or 90% amounts to the same thing as long as it fits. It does not.

Remember chapter 1: inside the window, every token competes for attention every time a word is generated. Filling the window changes that contest. Your instruction, which dominated the attention in a lean input, now competes with tens of thousands of tokens of log, dead code and old conversation. There is published research measuring how much quality drops as the context grows and where the drop is worst, and chapter 5 is entirely about it. For now, note the asymmetry: the window grew because it is an easy number to sell, but the capacity to hold tokens and the capacity to use those tokens well are different things, and the second did not keep up with the first. There is also a part no model page advertises: every token in the window is billed. Providers price per million input tokens and per million output tokens, so a window full of garbage costs real money on every call, even when quality survives. Chapter 6 does that math with you. Measure it yourself: what travels with a one-line question You do not have to take abstract numbers on faith; the experiment fits in one prompt. While writing this chapter, I opened a fresh session of my coding agent, cleared the history and asked it for one thing only: Repeat back exactly what you received in this request. I want to see everything that is in the context besides my prompt. Two sentences. A few dozen tokens. The answer listed what else was in the window on that call, and the list is long: the tool’s system instruction, with rules about how it should behave and

the complete state of the repository; the schemas of every tool the agent can call, the machine-readable description of what each one takes, plus a list of another hundred tools available on demand; my global instruction file, which dragged in a whole Flutter style guide, useless for a project that was not Flutter; a behavior mode injected by a session hook, a script the tool runs on its own at startup; the descriptions of some thirty-five installed skills, packaged instruction sets the agent loads on demand; and instructions from three Model Context Protocol (MCP) servers. Added up, the material that traveled with my question measured tens of thousands of tokens, three orders of magnitude larger than the prompt itself. None of that is a flaw in the tool; it is the price of a well-equipped agent. That is not the point, though: every subsequent call in the session reloads that baggage, and I had put a good part of it there myself and forgotten. Run the same prompt in your own tool before you read on. Knowing what your one-line question drags along with it is the first act of context engineering this book asks of you. The yardstick you take from this chapter Recall the session that opened the chapter. In that session, the agent read, say, a dozen files of a few hundred lines each, plus two build outputs. At a few thousand tokens per file and tens of thousands per log, the sum passes a hundred thousand tokens before you notice. The tool’s system instruction, the conversation history and the room needed for the answers pushed the total against the ceiling of the window, and the tool started to discard history to survive and took your instructions with it. None of that was invisible; it was only unmeasured. Now you measure: you

estimate the tokens of each piece, you know the sum competes for a finite window, and you know that filling the window has a double cost, in attention and money. One piece is still missing before the mechanism closes. If every call is isolated, as chapter 1 showed, and every call carries at most one window of tokens, as this chapter measured, then why does it feel as though the chat remembers what you said ten messages ago? The answer is that the tools send everything again for you, and understanding that trick explains why long sessions forget what was agreed on. That is the next chapter.

Powered by TurnKey Linux.