25개 이상의 토픽을 선택하실 수 없습니다. Topics must start with a letter or number, can include dashes ('-') and can be up to 35 characters long.

27KB

Context Engineering — Chapter-34: A complete AI-guided implementation

  • Source: /library/Context Engineering/source-file.pdf
  • PDF pages: 334–353
  • Pages without text: none

A complete AI-guided implementation Four files and three empty folders sit on disk. By the end of this chapter they will have grown into the whole Vila Nova Clinic system, built across five sessions. Each session leans on a different technique from Part III, because each one ran into a different problem. The first assembles the packet and then stalls on a dependency. The second audits what the session claimed. The third overflows the window. The fourth loses a morning of work and gets resumed twice. The fifth pushes an entire slice out of the main window. Here is what the five cost, added up from the tables at the bottom of each transcript: Session Iterations Time Output tokens Cost 1, bootstrap 32 153.8 s 11,353 $0.86 2, providers 49 276.2 s 22,218 $1.61 3, scheduling 64 551.7 s 44,113 $2.70 4, monthly report 66 500.3 s 43,837 $2.07 5, workins 41 602.2 s 53,155 $2.20 Total 252 2,084.2 s 174,676 $9.44

Thirty-five minutes of waiting on a machine, and under ten dollars. Keep the row for session 5 in mind, the priciest one, because it is the counterexample to what intuition says. Every session ran on July 29, 2026, on Claude Code 2.1.220 (Anthropic) with the Sonnet 5 model, in the non-interactive -- print mode with --permission-mode acceptEdits , which authorizes writing a file but not running git commit . The complete transcripts are at github.com/jckodel/context-engineering-companion-en, tag v1.0-part-v ; the excerpts the prose uses are printed here. Only the names belong to that particular tool: --print , --continue , - -resume , /compact , the Explore subagent type. As of July 2026, Cursor, the Codex command-line interface (CLI) from OpenAI and the Gemini CLI from Google each resume a session, compress a history and farm out a subtask in their own way, with their own syntax. The cycle underneath is what carries over: pack in layers, check claims against the project, write the anchor before you compress, leave a state note for whoever resumes, and put a contract around anything you push out of the window. Session 1: the skeleton, and three sentences that paid for the session Techniques in play: context layers (chapter 16) and packing (chapter 17). The request arrives split into zones, and the zone at the top is the one that does the work.

Opening the packet: what cannot be violated (layer 1)

  • No new dependency. The stack is closed in CLAUDE.md: TypeScript,

Hono, node:sqlite, Vitest. If you think you need something else, ask instead of installing.

  • Organization by feature. No file in src/ outside a feature folder, except the entry point of the server.
  • An interface with a single implementation is forbidden. The request asks for a minimal hypertext transfer protocol (HTTP) server with a health route, plus database access ready for the repositories still to come. The session reads the conventions, the spec, package.json and tsconfig.json , writes nothing, and answers this: Before writing any file, I found a point that blocks the skeleton: package.json already declares “dev”: “node src/server.ts”, but the Hono package alone does not run on Node: it exposes only app.fetch (Web Fetch API), with no built-in HTTP server. The standard way to serve that on Node is the @hono/node-server package (same maintainers as Hono), which today is neither in package.json nor installed. CLAUDE.md is explicit: “No dependency beyond those... Before installing any package, ask.” So I am asking before deciding on my own: [...] Nine iterations, 72.5 seconds, $0.26, and not one line of code. Zero is the right answer here, and the contrast with the previous chapter could not be sharper: the same tool, on the same day,

installed five dependencies without asking when the directory was empty. This time it stopped dead on a single one, because one line of layer 1 reached the window ahead of the request. I approved the package and asked for one more thing: record the decision in CLAUDE.md alongside the stack, so no future session has to ask again. The session rewrote the stack line before it did anything else:

  • HTTP server with Hono, served by @hono/node-server (the official adapter from the Hono team for running on Node.js). Decision made on 2026-07-29; do not ask again. That is layer 1 growing by a single conversational turn, the move chapter 16 calls promotion: a decision made in one session that binds every session after it moves up a layer and lands in a file. The skeleton itself fits in ten lines, and no commit happened, because acceptEdits authorizes an edit, never a git command. Half the request went unanswered. The session delivered the server and said nothing about database access. I called that out and told it to answer before writing anything, and out came the sharpest architecture call of the session: That access should not exist now: with no second feature needing it, creating shared/ would be abstracting before the second real implementation, so the node:sqlite connection is born inside the first *.repository.ts that the scheduling or the providers feature requires, and it only moves to shared/ the day another feature proves it needs the same code.

Three iterations, 12.9 seconds, $0.21, no new file. The rule behind that refusal was already sitting in the conventions; what turned it into a decision was the demand for one sentence of justification before any file got written. Drop that demand and the path of least resistance is to create the folder. The session as a whole: 32 iterations, 153.8 seconds, $0.86. The packet-free session in the previous chapter burned $1.40 to build a system I had to throw away. Same model on both sides. Session 2: five statements, three checks and a defect that was not a statement Technique in play: context validation (chapter 19). This slice registers a provider along with a weekly schedule, and it is the first one to pull material from the old system into the window, since that is where the schedule format comes from. Legacy material brings a risk: the session may present a 2019 rule as if it still held today. The packet guards against that by asking the session to say what it understood, before any code: Before writing any file, tell me in up to five lines what you understood the system does today in this area: what the schedule is, what rules it enforces and what already exists in the repository for you to reuse. Five statements came back. The second one was worth checking:

  1. The inherited format (docs/legacy/code-map.md) keeps this in the weekly_schedule table (provider_id, weekday, start, end), with the time as an integer of minutes since midnight: I need to keep that format in the new table. It also found something nobody had asked about. The Providers section of the spec demands two refusal messages and prints neither one, while CLAUDE.md requires copying the exact sentence from the spec. So the session asked what those sentences said. I could not just type the answer into the chat. A sentence that shows up in a conversation lives only there, and the next session would ask the same question. Instead, both messages went into the spec and got committed before the following turn, and my prompt pointed to the file and the line rather than repeating the text. The session did not open the spec first. It went and verified that the commit I claimed to have made actually existed. The full slice came out clean: the rule in the use case, the error translated in the HTTP file. Ten green tests. Now comes the part chapter 19 calls checking against the project: three commands, none of them addressed to the session. The first one asks whether the two messages in the code are truly the ones from the spec and not a paraphrase: matches the spec: Provider needs at least one working day matches the spec: Invalid schedule range The second command checks the inherited format against the code running at the clinic, not against the document describing that code, because a document was the session’s only source. A

grep through the old src/schedule.js turns up provider_id , weekday and the loop stepping forward in 30-minute increments. Two statements verified in twenty seconds. The third check covers nothing the session ever claimed, and it is the one that matters: $ grep -rn “:memory:|DatabaseSync” src/ | grep -v test src/providers/providers.repository.ts:1:import { DatabaseSync } from “node:sq lite”; src/providers/providers.repository.ts:23: db: DatabaseSync = new DatabaseSyn c(":memory:"), That parameter carries a default, src/server.ts calls the function with no argument, and the server the clinic starts with npm run dev drops every provider on restart. All ten tests stay green, because an in-memory database is precisely what a test wants. Notice what happened, because this is where chapter 19 meets its limit in the field. All five statements were correct. The defect rode in on a decision the session never announced, tucked into a parameter default. Auditing what the session said will never catch that. A different question does, and it belongs in every slice with an external effect: when the system runs for real, where does that thing live? Fixing it meant dropping the default on both sides and letting the entry point decide. The fix uncovered the next problem, which the session spotted and did not follow through on: with the database opened at the top of the module, any import triggers it, and npm test began writing the production file at the project root. The remedy was to move that call inside the block that only fires when the file runs directly, and one command confirmed it:

$ rm vilaschedule.db && npm test && ls vilaschedule.db Test Files 4 passed (4) Tests 10 passed (10) ls: vilaschedule.db: No such file or directory Forty-nine iterations, 276.2 seconds, $1.61. Half of that was repair work, and both repairs traced back to checks that cost one command apiece. Session 3: the window fills up, and what is left is not what you think Technique in play: context compression (chapter 20). This slice books and cancels an appointment, and the packet ends with a question I planted deliberately: the appointment’s weekday has to come out of a date string, and the clinic’s repository already does that math the wrong way. The session caught the trap by itself, in the right place: Before writing any code: the inherited format (docs/legacy/CLAUDE.md, section “Known traps”) documents the wrong way that already exists in the old system: today() uses toISOString(), which returns the date in UTC, and after 9 p.m. in Brasilia time the day rolls over to the next one. It is the same risk new Date(“YYYY-MM-DD”).getDay() would carry here: the Date constructor reads a YYYY-MM-DD string as midnight UTC, but .getDay() reads it in the local time zone of the process. The inherited trap chapter 29 wrote into the legacy packet has just paid for itself. The diagnosis is right; the proposal that follows it is expensive, hand-rolled calendar arithmetic to avoid depending on Date at all. I threw that out with a reason, and the

decision that stuck is Date.UTC followed by getUTCDay , so the code that writes the date and the code that reads it both sit in coordinated universal time (UTC). Nineteen green tests. Before the next turn, the decision and the discard went onto the anchor sheet from chapter 20, each with its reason. One question decides what gets in: if this session vanished right now, would this come back for free? tsconfig.json comes back. The conventions come back. The spec messages come back. The reason for reading dates in UTC does not come back from anywhere, because nobody ever wrote it down. With the slice done and 28 green tests, it was time to compress on purpose. Here are the compaction numbers almost nobody looks at: {“type”:“system”,“subtype”:“compact_boundary”,“compact_metadata”:{ “trigger”:“manual”,“pre_tokens”:71665,“post_tokens”:7783, “cumulative_dropped_tokens”:63882,“duration_ms”:72550}} The compact_boundary field marks the cut: 71,665 tokens in, 7,783 out, 63,882 dropped. Eleven percent survived. The operation ran for 72.6 seconds and cost $0.31 while producing not one token of useful work. That is the figure to remember: compression is an expense, never a saving. The summary kept the cancellation decision word for word, reason included, exactly as the instruction demanded. It carried nothing about Date.UTC , nothing about discarding the hand-rolled arithmetic, nothing about the 9 p.m. bug. That is no failure of the summary: none of it was there to keep. Chaining turns with -- continue spun up a new session file partway through, so the cancellation turns ran without any conversational memory of the

booking turns. They still worked, because everything they needed sat in CLAUDE.md , in the conventions, in the spec and in the code already on disk. That sentence deserves a pause, because it sets up the next technique: compression is the second thing that erases context. The first is the tool deciding on its own where a session ends. For the survival test I made a new request that depended on the lost decision: is there an earlier decision about how to derive the weekday, and what was the reason behind it? The compressed session nailed the what, since that lives in code. On the why, it said this: The reason, confirmed by the test book-appointment.usecase.test.ts:141 (“fixes the day of the week from the date”), is to avoid depending on the local time zone of the process. The sentence is nearly right, and it is not the decision anybody made. The real one named a specific bug in the old system, with a file and a line, and banned a class of code as a consequence. Worse than the vagueness is “confirmed by the test”: a test confirms behavior, never a reason, and the source it cites does not support the claim it is attached to. This is chapter 19’s extrinsic hallucination, aimed at the project’s own history, and it cost 12.0 seconds and $0.20. The remedy is not a better summary. It is getting the decision out of the conversation entirely. The reason became the first entry in docs/decisions.md , together with what it forbids and what it discards, and CLAUDE.md had pointed there since chapter 29. Here is the same question again, this time in a fresh session that inherited nothing:

Yes, there is: D1 in docs/decisions.md. [...] Reason: in the legacy system (src/utils.js:3), the date goes through toISOString(), and an appointment booked after 9 p.m. in the Brasilia time zone ends up stored with the next day's date, a bug the front desk has lived with since 2019. Forbidden as a consequence: new Date(string).getDay() and toISOString() to get a date in any file of the project. Thrown out: Zeller/Sakamoto. One file read, 8.6 seconds, $0.10, the reason quoted verbatim and the discard along with it. Compare that with 12.0 seconds, $0.20 and an invented reason propped up by a false source. The cheapest anchor is the one you never have to carry, because it turned into a file in the project. Session 4: the same task resumed twice Technique in play: context recovery (chapter 18). The task is the monthly report the front desk asked for, and I cut the session off deliberately at the halfway mark, leaving behind two settled decisions, one discard with its reason, one open question and an extraction that had been agreed on and never done. None of it written to a file. I saved the half-finished diff as a patch, ran the first resumption, restored the disk to that exact point with git apply , and started the second resumption from there. The only variable between the two runs is what reaches the window. Resumption 1 got the sentence everybody types: I started yesterday, I stopped halfway, pick up where I left off. It performed better than that method’s reputation suggests, and the reason is this book’s whole argument: it ran git diff , ran npm test , read the packet still sitting on disk, read docs/decisions.md , read the spec and the conventions. Anything stored in an artifact came back by itself. It also spotted something I had missed:

Conflict with the spec. docs/scheduling-spec.md lists “reports and billing” under Out of scope (line 85). packet-04 asks for exactly one report. Was this agreed with the coordination and the spec is merely out of date, or is it a one-off exception? That kind of finding surfaces only for a newcomer who arrives with no packet and has to orient themselves by reading the entire spec. I wrote the packet with the spec one line away and never noticed. What it failed to recover is exactly what chapter 18 predicts. Both decisions settled in the previous turn came back as open questions, offering the same two alternatives I had already chosen between. Twelve iterations, 88.6 seconds and $0.34 just to arrive at those questions. Pushed to keep moving, it settled them in its own way: // src/scheduling/generate-monthly-report.usecase.ts import { dayOfWeek } from “./book-appointment.usecase.ts”; Now the report use case depends on the booking use case just to compute a date. Thirty-five green tests. That is not what I decided, and nobody reviewing the pull request later would have any way to know the question had already been answered the other way. Resumption 2 began from the same disk with two extra files in the window: the task packet and the state note from chapter 18, recording where the diff stopped, which decisions were settled, what got dropped and why, and what stayed open. My prompt demanded three answers before any code, and told it to say “I

don’t know” rather than assume. Three commands later: all three answers, each reason attached to its decision, the discard quoted with both of its reasons, and the “I don’t know” in precisely the right spot: One open point the note records explicitly: I do not know whether a canceled appointment enters the count of the month. The coordination of the clinic has not answered yet, and for that reason it should not count until there is an answer. Four iterations, 20.0 seconds, $0.13. The two resumptions side by side: | | Resumption 1 | Resumption 2 | |---|---|---| | What reached the window | one sentence | packet and note | | Iterations to know where it was | 12 | 4 | | Time to know where it was | 88.6 s | 20.0 s | | Cost to know where it was | $0.34 | $0.13 | | Settled decisions recovered | 0 of 3 | 3 of 3 | | Discard recovered | no | yes, with both reasons | | Open question | lost | handed back as “I don't know” |

| Result of the code | diverges from the decision | follows the decision | The $0.21 gap is the most misleading number in that table. The last two rows are what matter. It is also worth recording what the two resumptions shared, because it turns this book’s argument into evidence: both recovered the database format, the conventions, the weekday decision and the state of the tests entirely on their own. None of that required memory, because none of it lived in memory alone. A state note adds only what has no other address. Session 5: the slice that left the main window Technique in play: context isolation (chapter 21). This slice is the waitlist with an automatic work-in on cancellation, the clinic’s term for the 15-minute appointment squeezed into a day that is already booked, and it cuts across all three features of the system. I ran the four-condition test before splitting anything, and the slice failed two conditions. It failed the disjoint-diff condition because the work-in fires on cancellation, and cancellation lives in src/scheduling/ . It failed the settled-shared-decision condition because the dependency between the two slices still pointed both ways: one of them has to import the other, and that choice reshapes the design on both sides. Chapter 21’s answer in that situation is not to split more carefully. It is to settle first. The direction became a decision entry, with the reason and the discard attached:

D3: workins knows scheduling, scheduling does not know workins

[...]

The practical consequence is that the route the front desk uses to cancel changes owner: it is now served by workins.http.ts, which composes the cancellation with the work-in attempt, and it leaves scheduling.http.ts. Reason: the alternative would make the two import each other, and a cycle between features is the beginning of the lump that organizing by feature exists to avoid. Of the two possible directions, the one we keep i s the one that follows the domain's real dependency: a work-in does not exist without an appointment, an appointment exists without a work-in. Once that was settled, the diff turned disjoint again and all four conditions passed. The subtask contract runs to four sections, and the costliest one to write is the list of true things you are keeping out:

What it does not need to know

  • The history of the four earlier sessions, their test output and the paths that were already thrown out over there.
  • How the monthly report was built, and the discussion about where dayOfWeek should live, which is already over.
  • That this repository is the companion to a book, and what the transcripts in transcripts/ are doing here.
  • The legacy system of the clinic and its map in docs/legacy/. Then the writing boundary, which turns the first condition of the test into an instruction: It creates and edits files only inside src/workins/ and the tests next to them. Two named exceptions, and only those: registering the new route in src/server.ts and removing the cancellation route from src/scheduling/scheduling.http.ts, which changes owner because of D3. If you need any other change in src/scheduling/ or in src/providers/, stop and hand the request back instead of editing.

The request that opened the isolated session runs to one line, because the contract is already the task. Back came eight new files, the two named exceptions, forty-seven green tests and a clean compile. The code is not the point; the shape of the delivery is. Every decision it made alone carried a reason pointing to a line of the contract, and every assumption arrived labeled. That label is your checklist when the result lands:

  • [ASSUMPTION] Any BookAppointmentError while attempting the work-in (not just the schedule) moves on to the next one on the list; another kind of error propagates.
  • [ASSUMPTION] POST /appointments/:id/cancellation still returns the canceled appointment (200); the work-in is a side effect, it does not show up in the body. The first of those two overreaches. The spec covers exactly one refusal, a time that does not fit the schedule; the session treated any booking refusal whatsoever as grounds for calling the next patient. That is safe today, and it stops being true the day the spec grows a third refusal. The label exists for exactly that reason. A fourth item arrived that nobody asked for:
  1. I edited src/scheduling/scheduling.http.test.ts to remove the tests of the route that changed owner, even though it is not one of the two named exceptions: I treated it as inseparable from “removing the route from scheduling.http.ts”, since leaving them would break npm test. If that reading is too broad, say so.

That hole is the contract’s fault. Authorizing a file for editing and forgetting the test sitting next to it is a sloppy boundary. Still, the instruction said stop and hand the request back, and the session chose to deliver instead. What redeems the episode is that the deviation showed up in the delivery, rather than hiding behind forty-seven green tests. The price of the subtask is the number that defies intuition: 39 iterations, 539.2 seconds, $1.89, the single most expensive step in the entire case study. Isolation is not cheap. What the money bought was an entire slice built without dragging one thing from the four earlier sessions back into the window. The second kind of isolation appeared in that same session, and it is the cleaner one: a read-only sweep handed to the tool’s own subagent, hunting for date math outside the settled decision and for front desk messages that did not come verbatim from the spec. My prompt told the session to delegate rather than sweep inside its own window. The subagent ran fifteen search commands, opened half a dozen files and reported back two sentences, each with a file and a line. The main session spent 2 iterations and 1,026 output tokens; all the file reading happened on the far side, and the work-in window stayed clean. One detail about the subagent’s inheritance beats the savings: it received the three-line contract, not the conversation. That is why its report fits in two sentences, and why it had to name where it found each item, with a file and a line. Nothing else would have made the answer checkable from outside. The technique’s limit showed up through a mistake of my own. I fired the isolated session from the wrong directory, so it loaded another project’s context file. It burned four commands hunting for the contract, found it, navigated to the right repository and

did the right work anyway, because the contract named files by path. A contract that said “follow the project conventions” would have followed some other project’s conventions without a word. What the five sessions add up to The system runs: provider registration with a schedule, booking and cancellation with a conflict check, the monthly report, the waitlist and the automatic work-in. Forty-seven green tests when the fifth session closed, $9.44 spent, 252 machine iterations. Where the money went is the reading that matters. The two priciest sessions are the one that overflowed the window ($2.70) and the one that isolated a slice ($2.20), and they ran up the bill for opposite reasons: the first carried everything, the second carried nothing. The cheapest was the first ($0.86), the session that wrote the least code and said no the most. Three of the five sessions moved something out of the conversation and into a file: the dependency decision went into CLAUDE.md , the two error messages went into the spec, the reason behind the date calculation went into the decision record. Not one of those writes produced a line of code, and all three killed a question the next session would otherwise have asked again. One fair objection before I close. Look at these five sessions and you could argue I steered too much, and that a better agent would handle all of it alone. So look at which four interventions moved the outcome most: fixing the message in the spec rather than in the chat, stripping a default off a parameter, rejecting hand- rolled calendar arithmetic with a reason, and settling which slice

depends on which. Not one of them is information that sat on disk while the agent failed to read it. Every one is a project decision, made by a person who answers for it. Five sessions produced eight failures, and one of them slipped through all five unnoticed. The next chapter opens that record.

Powered by TurnKey Linux.