Vous ne pouvez pas sélectionner plus de 25 sujets Les noms de sujets doivent commencer par une lettre ou un nombre, peuvent contenir des tirets ('-') et peuvent comporter jusqu'à 35 caractères.

22KB

Context Engineering — Chapter-30: Measuring context: how to evaluate whether your context improves results

  • Source: /library/Context Engineering/source-file.pdf
  • PDF pages: 270–284
  • Pages without text: none

Measuring context: how to evaluate whether your context improves results The colleague who asked how much it improved was not being ironic. They watched you change the way you work for three weeks, write the packet before asking for code, stop in the middle of a task to check a rule against the repository, write down a decision before the tool summarized the session. They want to know whether that is worth three weeks of their own. You open your mouth to answer and what comes out is: “it seems a lot better.” And it does seem that way. Thursdays started looking less like that Thursday, the review handed back less diff, the afternoon that used to vanish down an approach you had already abandoned became the exception. Except that you look at what holds the sentence up and find nothing beyond your memory of the last few weeks. It is the same kind of source you refused in chapter 19, when the agent confidently asserted a scheduling rule that had changed: confident prose, with nothing backing it up outside the head of the person saying it. Across the table somebody remembers the Tuesday before last, the one where the VilaSchedule schedule export ate the whole afternoon, and concludes out loud that nothing changed. Two impressions, no measurement, and the tie goes to whoever speaks with more conviction. You are not sure yourself, for that

matter. Maybe the tasks of those three weeks were smaller, or you knew that part of the system better, and what feels like improvement is luck of the calendar. The price of not knowing shows up in the next decision. The clinical coordinator wants the reports module by the end of the month and you have to choose where to spend the half day left in the week: writing the living doc for the no-show flow, putting together the subtask contract that does not exist yet, or writing code. With no number, the choice is a guess, and a month from now you will defend the guess with the same sentence you used today. Notice that no technique is missing here. What is missing is evidence, and evidence starts with counting something, which is exactly what you never did. “Evaluating that is work for a machine learning team” The first thing you do is look up how this gets measured, and the search hands back a world that is not yours. In 2026, evaluating a system built on large language models (LLMs) means a set of labeled cases, one model judging the output of another (what the field calls LLM-as-judge), execution tracing, a dashboard and a regression pipeline. None of that fits between two maintenance tasks at the clinic, and the natural conclusion is that measuring is for people with a dedicated team. You close the tab and go back to “it seems better.” Before accepting the barrier, it is worth reading the people who built that world. Hamel Husain, in “Your AI Product Needs Evals,” published in 2024 (hamel.dev), argues for something close to the opposite of what the ecosystem suggests: evaluation starts simple, over the real cases you have already seen fail, and the step nobody skips without paying for it is looking at your own data

one item at a time, before any dashboard or generic metric. The infrastructure comes later, pulled in by what the inspection showed, and not as a condition for starting. His text speaks to teams building a product on top of an LLM; carrying it over to your case is my doing, and it simplifies the arithmetic even further, because you are not evaluating a product for thousands of users. You are evaluating your own way of working, and the set of cases you need to look at is the work you already did this week. So an eval, in this chapter, means something modest: a count over the work you already do, written down in a text file checked into the repo alongside the project. No metric here requires a service, a database or instrumentation of your flow, and the restriction is deliberate, not a poor version of the right way. An instrument that has to be built before it produces the first number dies in the third week, and you end up with no infrastructure and no measurement. Count by hand first. If the manual count ever starts to strain, the problem will be well defined and automating it becomes an easy decision, with data on the table. What counts as right the first time The metric I count is called the first-pass rate: the share of turns whose first result was accepted with no course correction. The name is not my invention, and it is worth knowing where it comes from, because the family resemblance is right there in the names. In manufacturing, first-pass yield is an old Lean Six Sigma metric, the share of units that come off the line with no rework and no scrap. Software quality calls the same idea the first-time pass rate, counting the task that cleared review without coming back. And model evaluation has a close relative in pass@1, defined by Mark Chen and coauthors in “Evaluating Large Language Models Trained on Code” (arXiv:2107.03374),

from 2021. In that notation, pass@k is the probability that at least one of k answers generated for the same problem passes the automated tests that come with the problem, and the number after the at sign says how many attempts the model had. With k equal to 1 the model answers once, and the metric becomes the chance that it solves the problem on the first attempt, which is the kinship with what I count here. I borrowed the name from all three, and all three are older and better established than anything in this book. What is mine is the framing, and only that: the unit that enters the count and the line between what counts as a hit and what does not, which are the subject of the next two sections. The framing is also what makes any published number useless to you. The 90% a consultancy announces, the 85% to 95% a diagnostic platform reports and the pass@1 of a benchmark came out of another unit, another process and another acceptance criterion, so none of them is a target or a floor for you. The only legitimate reference point is yourself, two weeks ago. The unit is the turn of chapter 25, the full pass through pack, run, validate and distill. A task that needed three turns enters the count as three lines, not as one. If you are not running the loop yet, use the task as the unit: the number gets coarser and still works, as long as you do not switch units halfway through. The numerator is where the metric earns or loses its value, because “right the first time” is elastic and gets looser along with your mood at 6 p.m. My line is the course correction. If getting to the result you accepted took reassembling the packet, contradicting a statement, pointing out a file that was missing or switching approach, the turn does not count as right the first time. You repaired the context along the way, and that is exactly what the metric is trying to see.

The other side of the line matters just as much. If the agent wrote the test, saw red and worked on its own until it went green, against the target you declared before it started, the turn counts. That is the run step working the way chapter 25 asked for, not the context failing. Name adjustments, formatting and style preferences do not cost the turn either, because none of them came from information missing in the window. The edge cases will show up on the second day, and for them the rule that matters more than any definition of mine is this one: decide the borderline case once, write the decision at the top of the file and do not touch it inside the batch. A batch is my name for one closed block of counting: fifteen turns or two weeks, whichever comes first. Consistency matters more than accuracy, because your number is not going to be compared with anybody else’s. It is going to be compared with your own, from two weeks ago, and a criterion that swings turns any difference into noise. Thirty seconds per turn Collection has a set time, and the time already exists in your day: the distill step. When you close the turn, while you write the anchors and decide what is promoted to the durable sources, add a line to the counting file, with the turn still open on the screen. Writing it down at the end of the week, from memory, produces a record with exactly as much backing as the “it seems better” of the opening, with the added problem that it looks like data. The line has four fields: the number of the turn, what it did in half a dozen words, whether it came out right the first time and, when it did not, the reason. The last one does the work. Write the reason in free text and by the end of the month you will have fourteen distinct reasons, none of them countable. Choose from a closed list, written before the first batch, and each reason already

points to a technique from this part. Four weeks of VilaSchedule maintenance look like this, abridged, with each [...] marking what did not fit on this page:

Counting rules, fixed before the first batch

  • Unit: one turn of the pack-run-validate-distill loop. A task that needed three turns enters as three lines.
  • Batch: closes at 15 turns or two weeks, whichever comes first.
  • A turn counts as right the first time when its first result passed validation and was accepted with no course correction: no reassembling the packet, no contradicting a statement, no pointing out a missing file, no switching approach. [...]
  • Reason for failure: chosen from this closed list. With two causes in the same turn, I write down the first one that showed up.
    • incomplete packet: the packet was missing a file or a rule the

task depended on.

  • unchecked statement: I accepted a statement of state the project contradicted.
  • lost decision: the session summary carried away a closed decision.
  • restart from memory: I came back to the task with no state note.
  • mixed topics: the window carried more than one task.
  • stale data: I pasted schedule state that changed after the pasting.
  • ill-defined task: the problem was in the spec, not in the context.
  • A turn that failed for incomplete packet gets a fifth field: what opened that turn's packet, copied from the state note line.
  • I write the line at the distill step, with the turn still on

screen. Never at the end of the week, from memory.

Batch 1: 2026-06-01 to 2026-06-12 (first two weeks of the loop)

# Turn 1st? Reason
1 canceled work-in does not count yes
2 test for the per-day limit yes
3 time conflict no unchecked statement
4 work-in position in the block no lost decision
5 weekly utilization report no mixed topics
[...]
13 restart of the migration no restart from memory
14 batch cancellation no unchecked statement
[...]

Closing the batches

  • Batch 1: 5 of 14 right the first time (36%).
  • Batch 2: 11 of 15 right the first time (73%).
  • Dominant reason in batch 1: unchecked statement, 4 of the 9 failures. Validating, in batch 1, meant running the suite; no statement of state was checked against the project.
  • The single deliberate change between the batches: the validate step started running the validation checklist, with its order of sources, before I accepted the diff. [...]
  • One failure beyond the reach of context: ill-defined task in same-day rescheduling. The fix is in the spec, not in the packet.
  • Turns per task, median: 2 in batch 1, 2 in batch 2. The rate went up without my slicing the turns thinner to make the count easier. Three things in that record are worth more than the rate. The first is the reason column, and it is the only reason the file exists: a number on its own tells you something got worse, the reason tells you what. The second is the same-day rescheduling line, the

one that failed for ill-defined task . Not every bad turn is a context problem, and a context metric that does not admit this becomes an excuse: with no such category on the list, the spec failure would be counted as a packet failure and you would go fix the wrong thing. The third is the last line of the closing, the turns per task. Without it, the rate has an easy and unintended loophole in it, which is slicing the turn until each one is trivial; the number goes up and nothing improves. The two counts together shut that door. The number on its own decides nothing What do you compare against? Yourself, in the previous batch, and nothing else. A first-pass rate depends on the kind of task, the system, the model, your acceptance criterion and the day of the week, so putting it next to somebody else’s, another team’s or a number somebody posted means nothing at all. Comparing your 36% with your 73% does mean something, because both measurements came out of the same imperfect instrument. Compare apples to apples. Over what window of time? A batch closes at fifteen turns or two weeks, whichever comes first, and the two limits exist for different reasons. Below fifteen turns, one bad task moves the number ten points and you end up reacting to nothing. Above two weeks, you get a more reliable number about a decision that has already cost six weeks of work done the wrong way. Between precision and reaction time, prefer reaction time: the person measuring here is the person doing the work. How do you read the result? A few points of variation between batches is noise and asks nothing of you. A large move, up or down, asks for an explanation, and the explanation is never in the rate: it is in the column beside it. The dominant reason of the

batch picks your next technique, and the map is direct. incomplete packet sends you back to the four questions of chapter 17, mostly to the first one: which diff does this task produce? unchecked statement is chapter 19’s checklist coming into the validate step. lost decision is chapter 20’s anchor sheet written before the tool summarizes. restart from memory is chapter 18’s state note. mixed topics is chapter 21’s criterion for splitting. stale data is chapter 23 warning you that the piece of data belonged in a tool the agent could call, not in text pasted into the window. And then comes the one rule I follow strictly: one change per batch. If you apply three new techniques at the same time, the next batch will tell you it improved and will not tell you which change did it, and you end up with a routine full of rituals nobody knows the use of. That is what gave the record above its value: between batch 1 and batch 2 only the validate step changed, and that is why the 37-point difference has a single cause you can point to, one you can defend in a conversation. Measuring without deciding produces a vanity metric, a number that looks like management and changes nothing, and the difference between the two is the sentence that comes after the number. If your rate went up and you had not deliberately changed anything, you got lucky, not methodical, and the next batch may take the luck back. If it fell for two batches running and the reason column does not change, the bottleneck may not be context: it may be the spec, the size of the tasks or the model you picked. Writing that conclusion down is worth as much as writing down the rest, because it is what keeps you from spending three months optimizing what was already fine. Two counts that fit in the same file

The first is the size of the turn’s opening packet, in tokens. In 2026, every agent tool shows that number in some corner of the screen, and it costs one more column in your line. At the end of the batch you take the median, and the median speaks directly to chapter 6: the opening packet is what the cycle resends on every turn of that pass, charged in money or in quota. When the first- pass rate goes up while the median packet goes down, you have the argument that was missing in the conversation at the top of this chapter, and it fits in two columns: more hits with fewer tokens. The second comes from outside, and that is what makes it worth twice as much: how many turns came back from somebody else’s review. The data already exists in your flow, somebody already produced it for you, and it is the only number in this chapter that does not pass through your own judgment. A first-pass rate going up while post-review rework goes up with it is a sign that you loosened the acceptance criterion without noticing. Everything that requires assembly stays outside the file: a quality dashboard, a database of recorded runs, a model judging a model, any instrumentation of your workflow. Those things exist, they solve real problems and they are not the problem of this chapter. Also outside, as a continuous count, is chapter 5’s clean-session A/B test: running that on every turn costs more than the benefit, and it remains the best one-off diagnostic tool you have, for the day the rate falls and you suspect the session simply rotted. “Fifteen turns prove nothing” The objection is fair, and the answer is not a statistical one. You are not publishing a result; you are choosing between carrying on and changing course, and for that choice the yardstick is the size of the effect. A three-point difference between batches does not

move you; the 37-point difference of the record above does, and no significance test would change what you are going to do on Monday. Add to that the fact that the most actionable part of the file depends on no sample at all: four failures for the same reason already tell you what to fix, even if the rate itself means nothing. The second objection is more serious: you are grading your own homework. True, and the bias has a known direction, upward, especially on the turn you badly want to call finished at 6:30 p.m. Three things hold that bias within acceptable limits: the criterion written before the first batch, the note taken at the moment of the turn and the comparison always against yourself, which carries the same bias on both sides of the account. And the cheap external check is already in the previous paragraph, in the turns handed back by review, which do not pass through you. The third one stings because it is right: the tasks of one batch are not the tasks of the other, and the improvement may be in them, not in your context. You do not eliminate that without a laboratory you do not have. You can reduce it: write down the size of the task in two coarse categories, fits in one turn and does not fit, and compare inside the category when the difference between batches is too large to swallow. After that, accept the coarse measurement for what it is. The alternative in play was never a perfect measurement; it was the “it seems better” of the opening, which has all of these biases and the bias of memory on top. It is worth naming what this chapter assumes is already in place. Collection assumes nothing beyond a text file, and that is why it opens the chapter. It is the fixing that assumes things. The reason column only turns into action because each reason has an address in the repository: unchecked statement only has something to check against if the verified living documentation of chapter 9, the architecture decision record (ADR) of chapter 10, the conventions of chapter 11 and the configuration as a source of

numbers exist; incomplete packet only has a cheap fix if the rule that was missing is written somewhere the next packet knows how to cite. With none of those artifacts, the rate stays perfectly collectable and degrades where it matters: each failure becomes a fix that dies with the session, the same reason comes back in the next batch, and the number does not go up. You would have measured with precision a problem with no address. Your rate and the team’s Suppose it works. Two batches later you have 73%, a reason column pointing to the next fix and a sentence that replaces “it seems better.” The colleague from the opening accepts the number, adopts the practice and asks the next question, which is worse: how do I do this here? The first half of the question is about tooling, and this book has been putting the answer off on purpose since chapter 16. Which file your agent loads on its own before the first token, where it keeps persistent context, how it decides what to compact when the window gets tight, what changes when it runs in the terminal, inside the integrated development environment (IDE) or in a continuous integration (CI) run with nobody watching. The principles are the same in all of them; the controls are not, and every chapter in this part left a piece of that bill for Part IV to pay. That is where the principles become configuration, with the care not to become the manual of a tool that changes its name next year. The second half is more interesting. Your rate went up because part of the context is in the repository and part of it is in you, in your way of choosing what enters the packet. The part that is in the repository the colleague inherits on the first clone. The part that is in you enters nobody’s onboarding, human or agent, and it

is what makes the new developer take three months to get where you got in three weeks. Turning context into an asset of the team, with a repository standard, governance over what goes in and a way in for whoever arrives tomorrow, is the second subject of Part IV.

Powered by TurnKey Linux.