您最多选择25个主题 主题必须以字母或数字开头,可以包含连字符 (-),并且长度不得超过35个字符

7.0KB

Context Engineering — Chapter-09: Parametric calculation: cost of irrelevant context

  • Source: /library/Context Engineering/source-file.pdf
  • PDF pages: 53–57
  • Pages without text: none

Parametric calculation: cost of irrelevant context Variables (plug in your own tool’s and provider’s numbers): D = tokens of irrelevant context per call P = price per million input tokens, in dollars C = model calls per task T = tasks per day daily waste = (D / 1,000,000) × P × C × T Worked example (2026 values, dated on purpose; the formula outlives the prices): D = 30,000 (a log pasted and never removed from the cycle) P = 3 (dollars per million input tokens, the order of magnitude of a mid-tier model in 2026) C = 40 (a coding agent iterates dozens of times per task) T = 50 (a small team, a few tasks per dev per day) daily waste = (30,000 / 1,000,000) × 3 × 40 × 50 = 0.03 × 3 × 40 × 50 = $180 per day annual waste ≈ 180 × 250 working days = $45,000

Forty-five thousand dollars a year, on a small team, because of a single log forgotten in the cycle. And notice how conservative the example was: 30,000 tokens is one paste, not a whole project; $3 per million is the order of magnitude of a mid-tier model in 2026, and frontier models cost multiples of that; 40 calls per task is a disciplined agent. Redo it with your real numbers, which are in your provider’s billing dashboard and in your tool’s token counter. In most operations I have seen, the honest math is scarier than the example. If you are on a flat-rate plan, redo the math without P: D × C × T gives the waste in tokens per day, not in dollars. In the example above, 30,000 × 40 × 50 is 60 million daily tokens of irrelevant context that the team’s quota absorbs; for a single dev with 3 tasks a day, it is still 3.6 million. You do not need to know the exact size of your quota (most providers do not publish it in tokens) to draw the conclusion that matters: every token in that pile shortens the plan, and cutting D is the difference between a subscription that lasts the month and one that lasts two weeks, exactly the question at the top of this chapter. Two fair objections deserve an answer before I close out the math. The first: “prices per token only fall; the problem solves itself.” Prices do fall, and consumption per task rises faster, because agents multiply calls and larger windows invite larger contexts; the industry’s aggregate bill keeps growing, and yours probably does too. The second: “my provider has prompt caching.” It does, and you should use it: providers offer a large discount for spans of input repeated between calls, which targets exactly this reprocessing. In July 2026, the numbers are large enough to change a decision: Anthropic charges 10% of the input price for a token read from cache, with a 25% surcharge on the write; OpenAI takes about 50% off automatically, once the repeated prefix reaches 1,024 tokens; Google takes 90% off on Gemini 2.5 and later. All three publish those values on their

caching documentation pages (references in the appendix), and all three impose the same condition: the discount applies to the prefix that matches byte for byte from the start of the input. That has an engineering consequence chapter 16 picks up: what is stable in your session (instructions, conventions, tool definitions) lives at the top of the payload and does not change mid-session, because editing one line at the top invalidates the cache from there on and the next call reprocesses everything at full price. But caching discounts the price of the irrelevant token; it does not make it free, it does not make it fit better in the window and it does not take it out of the competition for attention from chapter 5. Caching is a painkiller, not a cure: context that should not be there still should not be there, at a discount. What to measure tomorrow morning The formula only works if you feed it your own numbers, and you can get all four in minutes, with no new tool. D, the irrelevant tokens per call, is the most laborious to measure and the most revealing: open the usage breakdown for a recent session in your tool, look at what makes up the input and ask, item by item, “did this contribute to any answer after the turn it entered on?” Add up whatever fails the test: resolved logs, file reads that mattered once, the side conversation. The first audit tends to find more dead weight than live context, and you do not need precision; the formula is linear in D, so getting D wrong by half only gets the result wrong by half. P is on your provider’s public price page, in the row for the model you actually use, input column; if you are on a fixed plan, drop P and keep the result in tokens, which is the currency of your

quota. C, the calls per task, shows up in your agent tool’s log, and if it does not expose that, count one typical task by hand, once. T you know by heart: how many AI tasks the team runs per day. Did the math? Now take the step that turns a number into a decision: compare the annual waste with the cost of avoiding it. On a flat-rate plan the comparison is even sharper, because you already feel it in your week: write down what day of the week (or of the month) you hit the limit today, apply the hygiene techniques for two weeks and write it down again. Every extra day before you hit the ceiling is the same saving, paid in uninterrupted working time instead of dollars. The techniques in the next parts of the book (context assembled from a specification, short sessions, curation of what enters the cycle) cost discipline, not a software license. When the calculated waste exceeds the cost of the hours spent on hygiene, and it crosses that line early, the practice justifies itself, in whatever spreadsheet your management uses. Quality and cost are the same bug Put the two problems side by side, because they are the same defect with two bills. The dead log in your cycle degrades the answer (chapter 5, through dilution of attention) and costs money or days of quota (this chapter, through billed reprocessing). There is no trade-off between quality and cost here, and that is rare in engineering: removing irrelevant context improves both at once. It is the kind of alignment that turns a technical practice into a business argument. And now an opinion: it was this math, not the U-shaped curve, that gave me cover to invest project time in context hygiene without having to ask permission.

Part I built the mechanism (chapters 1 to 4) and priced the damage (chapters 5 and 6). What is missing is the conclusion that gives the book its name: if quality and cost depend on what is in the context, then the variable the industry spent years optimizing, the wording of the prompt, was never the main lever. The next chapter closes the part by arguing exactly that, giving prompt engineering its due, and naming the discipline that takes its place at the center of the practice.

Powered by TurnKey Linux.