25'ten fazla konu seçemezsiniz Konular bir harf veya rakamla başlamalı, kısa çizgiler ('-') içerebilir ve en fazla 35 karakter uzunluğunda olabilir.

18KB

Context Engineering — Chapter-29: Development loops with AI

  • Source: /library/Context Engineering/source-file.pdf
  • PDF pages: 257–269
  • Pages without text: none

Development loops with AI Tuesday, 10:20 a.m. The task is to record who authorized the work-in when the front desk at Vila Nova Clinic goes past the limit of two per provider and the clinical coordinator releases the third. You open the session, paste the three work-in rules the change might break, point to the two files where the diff happens and ask for the test before the code. The agent writes the test, you run it, it goes red for the right reason. It writes the rule, the test goes green. You open config/scheduling.yml to confirm the limit is still two, update the line in the living doc and push the commit. 11:10 a.m., review with no comments. Thursday, 10:20 a.m. The task is the same size and on the same subject: stop the same patient from getting two work-ins on the same day, with any provider at the clinic. You open the session with the window still warm from Tuesday’s conversation, ask for the change directly, and the agent hands back a filter in the controller, an approach VilaSchedule dropped in two earlier tasks because it broke the house convention. You correct it in the same conversation, it redoes the work in the domain, and in the middle of the answer it states that the day limit is always per provider, so counting per patient is redundant. The sentence sounds reasonable and you move on. 5:30 p.m., the review hands the diff back: the new count counts per provider, and the front desk can still give the same patient two work-ins in two different exam rooms. Two tasks of the same size, in the same week, on the same system, with the same person and the same model. Forty minutes on one side, an afternoon on the other. And the question

that lingers is not why Thursday went wrong; it is the question before that, the one you cannot answer: what exactly did you do on Tuesday that you did not do on Thursday? Notice that no technique is missing here. The eight previous chapters are in your head, and you used pieces of them on both days. On Tuesday you packed the minimum, and on Thursday you packed something too. On Tuesday you checked the limit against the project configuration, and on Thursday you checked nothing. On Tuesday you started with the test, and on Thursday you started with the request. None of that was decided: it happened. What got written down from the two days was the diff, and a diff records what you produced, never the route you took. With no record of the route there is no way to compare Tuesday and Thursday, and with no comparison you are left with the only explanation there is, that on some days the AI is good and on others it is not. Technique is not cadence Every chapter in this part delivered a criterion and none of them delivered a moment. Chapter 17 taught you to assemble the minimum packet and did not say how many times per task you reassemble it. Chapter 19 gave you the yardstick for checking what the AI asserts and did not say at what point in the task the checking pays off most. Chapter 20 taught you to write the anchors before the summary and did not say when you stop working to write them. Each technique on its own is a right answer to a question you have to remember to ask, and remembering is exactly what fails at 5 p.m. A cadence is the fixed order in which those questions come up, without depending on your memory. It gives you no new capability: it gives you repetition, and repetition is what turns

eight occasional right moves into a predictable result. It is also what makes the error diagnosable, and that is the larger gain. When Thursday goes wrong inside a declared order, you do not ask what happened to the AI; you ask which step was skipped, and the answer fits in one word. Before I give the thing a name, one caveat, because this book has already used the word for something else. Chapter 4 called the involuntary mechanism of the session the context cycle: the output of each turn comes back as the input of the next, the history grows on its own and the window degrades by default. That cycle runs with you or without you, and it never stops. The loop of this chapter is the opposite in intent: it is voluntary, you impose it on top of the other one, and it exists precisely to manage what the context cycle does by itself. One is the physics of the session; the other is your work discipline inside it. Pack, run, validate, distill The reference loop of this book is called the pack-run-validate- distill loop, and the four steps bring no new technique at all: each one is a chapter you have already read, placed in a position. Call each complete pass through the four a turn. The word is chapter 4’s, stretched one notch: there it named a single round with the model, here it names a full pass through the loop, and in both it is the turn, not the task, that is the unit that repeats: a small task fits in one, and an afternoon’s task takes five. Pack is the context packing of chapter 17, choosing item by item what enters the window, applied on top of the lifetime layers of chapter 16. You open the turn by deciding what goes in: the few lines of layer 1 the task might violate, the layer 2 material it consumes, the request closing the packet. This is also where the three context architecture decisions of chapters 21 to 23 come in,

and they are made once, before the first token, not in the middle of the work: what to embed, what to leave outside to be retrieved by a search, what to expose as a tool because it changes faster than the session, and whether this task deserves a context of its own whether the four conditions are met. Run is the task itself, in small steps, with a check at every step, and that way of working is not my invention. In the guide “Claude Code: Best practices for agentic coding,” published by Anthropic in 2025 (anthropic.com/engineering), the recommendation is to set the target before the implementation, by writing the test or describing the expected result, and to check against that target at every step, instead of asking for the whole change and reviewing at the end. What I add is the boundary of chapter 21: when the subtask passes the four conditions, it runs in a context of its own, with a written contract, and what comes back to the main turn is a one-page return, never the whole session. Validate is the context validation of chapter 19, checking each statement the AI asserts against the project, in the position where it costs least: before accepting the diff, and not two days later, in somebody else’s review. You do not check everything. You check the statements of state the diff depends on, one by one, in the order of sources from that chapter. On the Thursday this chapter opened with, the turn would have stopped exactly at the sentence about the limit always being per provider, and the missing check would have cost one command. Distill is the step almost everybody skips, and it is the one that closes the loop. It has two halves. The first is the context compression of chapter 20, writing down what has to survive the summarizer: you write the anchors, the decision with its reason and the drop with its reason, so they live through the moment the tool summarizes the session. The second looks outside the

session: what this turn produced that is worth more than the turn is promoted to the durable sources of Part II, the new line in the living doc, the architecture decision record (ADR) when the decision is architectural, the convention when the pattern starts to hold for all new code. Distilling is what keeps the next turn from starting over from memory, and that is why it is not optional: a turn that does not distill delivers code and delivers nothing to the next one. Where recovery comes in You may have noticed that the context recovery of chapter 18, the routine for rebuilding a lost thread, does not appear among the four steps, and the omission is deliberate. Recovery is not a stage of the turn; it is what you do when the turn breaks in the middle, because the machine restarted, the window blew up or the task slept until Thursday. The loop runs with a live session. When there is no live session, you come in through the routine of chapter 18, assemble the packet again, rebuild the thread from the state note the previous distillation left behind and go back to the loop at pack, never at the step where the work stopped. The better your distilling, the less recovery you need; and the worse it is, the more recovery becomes your main job. One turn on Thursday Run Thursday’s task inside the loop and see where it would have changed course. What follows is the turn written step by step, abridged, with each [...] marking what did not fit on this page:

1. Pack

Context architecture decisions, made before the first token:

  • Embed: the standing work-in rules and the two conventions this task may violate. They are small, stable and every work-in task uses the same ones.
  • Retrieve: nothing. No clinic policy decides this rule.
  • Expose as a tool: nothing. The task changes a counting rule and does not need the schedule's current state.
  • Isolate in its own context: no. The diff covers two files and the contract would be the size of the task. [...]

2. Run

  • Step 1: the test that describes the new rule, red, next to day_limits.ts. [...]
  • Check at every step: run the step's test before asking for the next one. No new step with the suite red for a reason you do not understand yet.
  • What the run leaves in layer 3: the day count comes out of a single query and the change needs a second one; the new check stays in day_limits.ts; doing the check in the controller was dropped, because the convention keeps the rule in the domain.

3. Validate

Statements of state this turn produced, and what each was checked against:

  • “The standing limit is 2 work-ins per provider per day”: checked against the standing-rules table in docs/scheduling.md, and against the green test that exercises the limit.

[...] When one of them fails: discard what the session generated after the statement, assemble the packet again with the verified rule at the top, citing file and line, and redo the turn from the step that depended on it.

4. Distill

  • To the anchor sheet, which survives this session's summarization: the decision to check in the domain, with the reason; the drop of the controller, with the reason.
  • To the state note, which survives the end of the session: where the diff stopped, the closed decision, the drop and the open question (does a work-in canceled and rescheduled on the same day count once or not at all).
  • To the project's durable sources, which survive the task: the new

line in the living doc's standing rules table, “1 work-in per patient per day across the whole clinic”, verified in CI by this turn's test. [...]

When the turn breaks

Recovery is not a step of this loop. It comes in when the session loses the thread in the middle of a turn, from a machine restart, a blown window or a day's gap: pick up from the state note of the previous distillation, check what came back before asking for code and restart the turn at pack, never at the step where the work stopped. Three lines of that sheet would have saved the lost afternoon on their own. The first is the drop of the controller, which on Thursday you had to correct in conversation and which here comes in already decided, because the distillation of an earlier turn recorded it. The second is the statement about the limit, which comes out of the agent’s head and becomes a line checked against the standing rules table, with the green test beside it. The

third is the last one in the distill step: the new rule is promoted to the living doc, and the next person to touch work-ins gets that rule in the packet instead of finding it in review. Notice what that turn does not have. It has no technique you did not know before this chapter, no tool, no new file beyond the three Part II was already asking for. What it has is order, and order is what makes Thursday comparable with Tuesday. Calibrate without breaking it The part you have to adapt is the cadence, meaning the size of the turn and how often it repeats. The reference loop does not say that every task fits in one turn or that every turn lasts an hour, and there is a single rule I use for sizing: the turn ends where validation is possible. If you can verify the result after two lines, the turn is two lines. If the only verification available is the whole suite running in twelve minutes, the turn grows until it holds one suite run, because a turn smaller than your verification cycle is ceremony with no payoff. On an exploratory task, where you do not know the target yet, the first turn delivers an answer and not a diff: the run step becomes reading, and the validate step checks the statements the reading produced. Granularity is the second knob, and it changes who does each step. On a small task, the four steps are yours and happen in the same window. On a large task, pack and distill stay yours, run can live in an isolated context with the contract of chapter 21, and validate can be partly automated, because a test that exercises the rule is better validation than a command you type. When you are paired with somebody else, the distill step usually becomes the closing conversation of the day, and the anchor sheet becomes its agenda.

Three things I do not touch, and I say that as an opinion formed in turns that cost me dearly. The order of the four steps, because validating before running has nothing to validate and packing after running is self-deception. The obligation to distill, because it is the only step whose benefit shows up tomorrow and is therefore the first to be sacrificed today. And declaring the target before running, because with no declared target the validate step turns into a read of the diff through the tired eyes of somebody who already wants to go home. “That is ceremony for a ten-minute task” The objection comes up in the first week and you will make it yourself. Four named steps, to change one constant? The answer is the same one chapter 16 gave about layers: the loop adds no work to your day; it only names the order of what you already do when things go well. On the ten-minute task, packing is one sentence, running is one request, validating is one command and distilling is deciding that none of it deserves to survive, which is a legitimate decision and takes two seconds. The loop charges you on the turn that goes wrong, and there it is the only thing that answers the question this chapter opened with. A second objection is more up to date: in 2026 the agent plans, writes the code, runs the test and summarizes the session on its own, so the cadence is already built into the tool. There is a lot of truth in that, and you should delegate everything it covers. What the agent does not do is the beginning and the end of the loop. It does not choose the admission criterion for the packet, because what it knows about your project is whatever fits in the window and it has no way to know that the fixed-interval ADR is irrelevant to this task. And it does not decide what from this session deserves to become a durable record for the team, because that decision depends on what somebody else will need

three weeks from now, information that is in no window at all. Pack and distill stay yours even when run and validate go by themselves, and which tool automates which step is the subject of Part IV. It is worth naming what this chapter assumes is already in place. Distilling is only cheap because there is somewhere to distill to: the verified living doc of chapter 9, the ADR of chapter 10, the conventions of chapter 11 and the spec of chapter 8 are the destination of what the turn produced and the origin of what the next turn packs. With none of those artifacts, the loop degrades in a specific and cruel way: distillation has no address, everything the turn learned stops at the task’s state note, dies with it, and the next turn starts over packing from memory. You would be back to the swings of the opening, now with process on top, which is the worst combination available. Two weeks later, the same feeling Run the loop for two weeks and something changes. Thursdays start to look more like Tuesdays, less of the diff comes back from review, and the afternoon that used to disappear down an already dropped route becomes the exception. You tell a colleague about it and they ask how much it improved. You answer that it seems a lot better, and you notice, as you say the sentence, that it is exactly the same kind of sentence you refused in chapter 19 when the agent asserted something with no source. The problem now is one of evidence. You have a repeatable cadence, and repeatability is the precondition for any measurement: the turns of the loop are comparable with one another because they follow the same order, which the Tuesday and the Thursday of the opening were not. What is missing is counting something about them. How many turns came out

right on the first try, how many came back from review, how many stopped at the validate step and why. None of that requires infrastructure, a dashboard or an evaluation tool: it requires a text file, one column and the habit of writing things down. How to put together simple evaluations of your own work, and how to compute the first-pass rate, the share of turns that come out right on the first try, by hand, is the subject of the next chapter.

Powered by TurnKey Linux.