No puede seleccionar más de 25 temas Los temas deben comenzar con una letra o número, pueden incluir guiones ('-') y pueden tener hasta 35 caracteres de largo.

6.2KB

Context Engineering — Chapter-08: Token economics: the real cost of bad context

  • Source: /library/Context Engineering/source-file.pdf
  • PDF pages: 49–52
  • Pages without text: none

Token economics: the real cost of bad context The bill from your application programming interface (API) provider arrives and it is 40% higher than last month’s. You pull up the dashboards: same team, same projects, no new AI feature shipped. Perceived usage did not change; consumption did. Nobody can point to where the money went, because the money went nowhere visible: it was burned token by token, in context no human asked for and no model used. The meeting ends with the worst possible conclusion, “that’s just what AI costs,” which is the claim “AI is a lottery” from chapter 1 restated in dollars, and equally false. And if you are on a flat-rate plan, the problem is the same: the plan covers a fixed amount of usage, and the more you use, the sooner you hit the limit and your work gets interrupted. Have you ever paid for a month of an AI tool and had it last one or two weeks? The previous chapter showed what swollen context does to quality. This chapter shows what it does to cash, and the second problem tends to convince management where the first one does not. The good news: unlike quality, cost is arithmetic simple enough to do on the back of a napkin. By the end, you will have a one-line formula to estimate your own waste, with your own numbers. How the meter runs

Model providers charge per token processed, with public prices per million tokens, split into input (what the model reads) and output (what the model generates). Anthropic’s, OpenAI’s and Google’s price pages list the current values; this book does not print them as a table, because token prices change faster than a book can be printed and shipped. For the math here, what counts is two stable properties of the prices, not the numbers: input tends to be several times cheaper than output, and what you pay is proportional to volume, regardless of merit. Like a taxi meter in traffic, it runs the same whether it is metering the snippet of code that solved the bug or the dead log that has been circulating since turn 5. Reasoning models added a third property to the bill, and as of July 2026 it holds true across the major providers: the thinking tokens from chapter 2, the draft the model generates before the answer, are billed as output, the more expensive of the two columns on any provider’s price page, even when the tool does not show them or shows only a summary. That is the easiest share of the budget to underestimate, because it is invisible in the answer and right there on the bill: you look at ten generated lines and the usage page records a few thousand output tokens. A task that triggers long reasoning pays for that draft on every call, and cutting irrelevant context also cuts how much the model drafts about it. If you pay for a fixed monthly subscription instead of paying per token, do not skip this section thinking it is a problem for finance. The meter is there all the same, just hidden: providers cap the usage of those plans with quotas that reset on a rolling window (per session, per day or per week, depending on the provider; check its usage limits page), and what counts against the quota is the same volume of processed tokens that would show up on an API bill. Your currency is not the dollar but the quota, and swollen context does not show up as red ink on a

spreadsheet: it turns into the “limit reached” notice in the middle of a task, on Wednesday morning. Read everything this chapter says about dollar costs this way as well: every useless token in the cycle moves up the moment when the tool stops answering and you sit waiting for the quota to renew. That last sentence is the key to the chapter. Put it together with the cycle from chapter 4: your tool resends and reprocesses the whole history on every turn, so every useless token is not billed once but on all the remaining turns of the session. The 30,000- token paste on turn 5 of a 50-turn session shows up on the bill 45 times. In a chat session with a human in the loop, that adds up to dollars. The trouble is that the industry stopped keeping a human in the loop. Agent scale: the multiplier nobody budgets for An agent is a model in a loop: it receives a task, decides on an action, reads the result, decides the next one, dozens of times, without you clicking anything. The coding assistant that runs tests, reads files and iterates until the test passes is an agent. And each of those iterations is a full call, with the whole accumulated context in the input, at list price, the undiscounted rate on the provider’s price page. That is where the multiplier lives. In chat, what limits the number of calls is your patience; in an agent, it is the task. A routine coding task easily fires off 30 to 50 chained calls, and each one carries the whole cycle: the files read, the test outputs, the logs. The article “Effective context engineering for AI agents,” by Anthropic (2025), uses exactly that scenario to argue that context is a finite, critical resource: when an agent is running, it reprocesses and pays for every irrelevant token dozens of times per task, hundreds of times per day, thousands of

times per month. The waste that was pocket change in chat becomes a meaningful line on the bill, and the 40% jump in the bill at the top of this chapter stops being a mystery: all it took was the team adopting agents without adopting context hygiene, the discipline of deciding what gets into the input and taking out what no longer earns its place. On a flat-rate plan, the same multiplier applies to the quota. The task’s 30 to 50 calls eat into the limit exactly as they would eat into a budget, and that is why the agent subscriber hits the ceiling far more often than the chat user ever did: the provider did not shrink the plan; the agent’s loop multiplied the volume processed per task by dozens, dragging the dead weight along on every iteration. Do the math yourself Enough qualitative talk. Here is the calculation, and you can redo it by hand, swapping in the values of your own operation:

Powered by TurnKey Linux.