Du kannst nicht mehr als 25 Themen auswählen Themen müssen entweder mit einem Buchstaben oder einer Ziffer beginnen. Sie können Bindestriche („-“) enthalten und bis zu 35 Zeichen lang sein.

19KB

Context Engineering — Chapter-35: Post-mortem: where the context failed and how it was recovered

  • Source: /library/Context Engineering/source-file.pdf
  • PDF pages: 354–366
  • Pages without text: none

Post-mortem: where the context failed and how it was recovered The rule was written in three places. In docs/conventions.md , in the section about the boundary between features. In CLAUDE.md , which opens every session. And, spelled out, in the fifth session’s subtask contract. One feature talks to another through the front door: the public use case of the other feature, never one of its internal files. During the audit before publishing the code, the first command I ran turned up this: src/scheduling/book-appointment.usecase.ts:1 import type { ProvidersRepository } from “../providers/providers.repository.ts”; src/scheduling/scheduling.http.ts:2 same thing That violation appeared in session 3 and survived sessions 4 and 5 without anyone tripping over it. The workins slice, where a work-in is the clinic’s 15-minute appointment squeezed into a booked day, got the rule last, spelled out in its contract, and obeyed it. The scheduling slice had the same rule in the same context file and broke it. Forty-seven green tests, a clean compile, no warning anywhere.

If writing the rule in three places was not enough, what would be? To answer that, this chapter opens the record of all eight failures from the case study, and the answer keeps the same shape every time: a context failure has an address, and the fix goes into a file, a contract or a check command. Not one of the eight fixes amounted to asking the session to pay closer attention. The eight are in transcripts/failures.md , written on the spot, during the sessions, with the symptom, the diagnosis and the cost. Seven trace back to a numbered session from the previous chapter. The eighth, the one with the import above, has no session of origin: it was found in the final audit, after all five, and it is the only one of the set that nobody watched happen. The delivery that came back incomplete and said nothing Failure F01 hit in session 1. My request had two parts, the server and the database access, and only one came back. The session skipped the second part and never said it was skipping it. When I called that out on the next turn, it answered correctly: do not create the shared folder before the second use. The diagnosis lies in how I shaped the request, not in the session: a compound instruction disappears when the turn is interrupted midway. Turn 1 ended in a question about a dependency, turn 2 answered the question, and the second half of the original request stayed two turns behind a conversation that had changed subject. A request in two parts comes back in two parts, or it becomes two requests.

F03, in session 2, is the same failure wearing a different hat. After we fixed the database hiding in a parameter default, the call that opens the database moved to the top of src/server.ts . The server’s test imports that module, so npm test began creating the production database at the project root. The session watched it happen, noted that the file had appeared, and stopped there. What both of them teach is a question for the end of every repair: what did this fix start doing that it did not do before? It costs one command, and here it saved us from a production database opened by the test suite on the machine that serves the clinic. The decision nobody made out loud F02 is the most instructive of the set, because it is the only one that went through one of this book’s techniques working the way it should and escaped anyway. In session 2, the providers repository started out with the in-memory database as the parameter’s default, and the server called the chain with no argument. The system the clinic brings up with npm run dev lost every provider on each restart. The session claimed nothing false. That is the problem: it claimed nothing at all. Chapter 19’s validation technique checks a statement about state against the project, and all five statements from that session held up. The decision slipped in through a parameter default, which is where an infrastructure choice goes unnoticed, and the ten tests stayed green because an in-memory database is precisely what a test wants. The fix adds one question to the checklist, and it is mandatory in every slice with an external effect: once a database, a file or the network enters the slice, where does that thing live when the

system runs for real? Go find the answer in the code, not in what the session tells you. The reason that died with the session F05, in session 3, is the failure the whole book had been predicting. When asked, after the compaction, whether a previous decision existed about how to get the day of the week and what the reason for it was, the session found the practice in the code and explained the reason as “confirmed by the test book- appointment.usecase.test.ts:141 .” The practice was right. The real reason was something else: a specific bug in the old system, with file and line, and a class of code forbidden as a consequence. Two wrong things in a single sentence. The first is a reason invented because it sounded plausible, which is what is left when the reason was never written down anywhere: the code preserves the choice and loses the why. The second is worse and easier to let through: a test confirms behavior, never a reason, and the source it cited did not support the claim. An answer with a file and line reference looks verified, and this one was not. The fix was not writing a better summary, and here is where chapter 20’s compression stops helping: an anchor written before compressing only preserves what was already written somewhere, and a reason that never left the conversation has nothing to anchor. It was taking the decision out of the conversation and putting it in chapter 10’s record, one entry per closed decision, with the reason, what was discarded and what is forbidden as a consequence. The same question, after that, in a session with nothing inherited, was answered with one file read in 8.6 seconds and $0.10, compared with the $0.20 the invented answer cost.

F04 is its twin, in the same session 3, and it is the easiest one to repeat without noticing. As it wrapped up the cancellation, the session announced: I also saved to memory the decision about errors outside the spec, to keep that pattern in future use cases. The decision is right and the place is wrong. The file landed in the tool’s memory directory, on my machine, outside the VilaSchedule repository. git status does not show it; the commit does not carry it; the next dev clones the project and gets nothing. The rule was worth having for the team, but it was filed away on one machine. The fix was to promote the rule to a section of docs/conventions.md , which is versioned project material in the sense chapter 12 gives the term, and to delete the memory file, so that two copies do not sit there diverging over time. That is the only one of the eight failures that depends on a detail of the tool, because the directory where Claude Code keeps memory is its own. In July 2026, Cursor, the Codex command- line interface (CLI) and the Gemini CLI each keep state of their own outside the repository, and the question that catches the failure is the same in all four: when the session announces that it saved something, where did it save it, and does git see that place? The packet that was wrong F06 happened in session 4 and the fault is mine, and it is a writing mistake. docs/scheduling-spec.md listed reports under Out of scope, and the session’s packet asked for a monthly report. The

contradiction was one line away and survived two whole turns. The detail that redeems the record is the most uncomfortable one: in the first turn the session ran a grep that matched the word report in the spec, read the line that says “out of scope” and moved on to the report’s design without mentioning the subject. The contradiction was caught by a later session that had no packet in the window and was reading the spec to find out on its own what was going on. The likeliest explanation is that the packet asserted the scope with authority, so the spec entered the window to confirm something already decided rather than as a source allowed to disagree. The resumption had no packet at all, so it read the entire section to get its bearings. The fix was in the spec, which brought the report into scope with the rule about the canceled appointment. The error at the source is the packet’s, and the lesson is about whoever writes it: the packet is the only piece of the flow that nobody checks. Checking layer 1 and layer 2 against the spec is a two-minute read, and it is worth doing before sending, not after two sessions have worked on top of it. The boundary drawn halfway F07 is from session 5, inside the isolated subtask. The contract authorized writing in src/workins/ and two more named exceptions, and ordered the subtask to stop and hand the request back for any other change in the other two slices. The subtask also edited the route’s test file when that route changed owners, which was not among the exceptions. Item 4 of the delivery, already printed in the previous chapter, declared the deviation

and gave the justification in one line: it treated those tests as inseparable from removing the route, because leaving them would break npm test . The episode survives review because the subtask declared the deviation instead of burying it under forty-seven green tests. The worrying part is that my instruction told it to stop and hand the request back, and it delivered anyway, which is precisely the choice chapter 21’s isolation removes when the contract is drawn well. The hole is in the contract: without that edit the suite would break, so the deviation was necessary. A writing boundary is drawn by unit of change, not by file, and authorizing a file is authorizing the test next to it. A contract that separates the two is going to be disobeyed for a good reason, which is the worst kind of disobedience to catch afterwards. The rule nobody enforced Back to the failure I opened with, F08. Its cause is dull and it is the most important one in the chapter: the violation crept in because the providers slice had no reading door at all. There was only the use case for registering a provider. The session needed the provider’s schedule, the only way to reach it was the repository, and importing the repository worked. No test went red, no type complained, and a context file has no way of refusing an import . The fix was to create the door that was missing, a use case to check the schedule, and to pass the function in place of the whole repository. A good side effect: scheduling’s test doubles shrank

from a repository of four methods to a one-line function, and the suite went from forty-seven tests to forty-nine, counting the two that came with the new use case. Two conclusions come out of that. The first is that a rule in the packet is not an enforced rule: what keeps the violation from happening is the other slice having the door ready, and when it does not, the rule loses to the only thing that works. The second is that the check that catches this is not reading, it is a command: $ grep -rn ‘from “../’ src --include=’*.ts’ | grep -v ‘.test.ts’ src/scheduling/scheduling.http.ts:2:import type { CheckSchedule } from “../pr oviders/check-schedule.usecase.ts”; src/scheduling/book-appointment.usecase.ts:1:import type { CheckSchedule } fr om “../providers/check-schedule.usecase.ts”; src/workins/workins.http.ts:2:import { AppointmentNotFoundError } from “../sc heduling/cancel-appointment.usecase.ts”; src/workins/workins.http.ts:3:import type { createCancelAppointment } from “. ./scheduling/cancel-appointment.usecase.ts”; src/workins/workins.http.ts:4:import type { createBookAppointment } from “../ scheduling/book-appointment.usecase.ts”; src/workins/cancel-appointment-with-workin.usecase.ts:4:} from “../scheduling /cancel-appointment.usecase.ts”; src/workins/cancel-appointment-with-workin.usecase.ts:5:import { BookAppointm entError } from “../scheduling/book-appointment.usecase.ts”; src/workins/cancel-appointment-with-workin.usecase.ts:6:import type { createB ookAppointment } from “../scheduling/book-appointment.usecase.ts”; Eight lines, every one of them ending in .usecase.ts . Reading that output takes ten seconds and answers the question none of the five sessions answered. The grep that must come back empty is the same one with a filter at the end, looking for repository in the list: before the fix it matched two lines, now it matches none. The fix came with an amendment to the convention, because the raw rule would also forbid what the tests legitimately do:

The rule applies to production code. A test sets the scenario up as an entry point, and for that reason it may build the repository of the other feature to write the data it needs, the same way src/server.ts does. Without that sentence written down, the audit is not reproducible: the next person to run the command would find the tests in the list and would not know whether that is a violation or an exception. What each failure cost Failure Session Cost of the fix Reason from ch. 26 F01, partial delivery 1 12.9 s, $0.21 none F02, decision by default 2 75.6 s, $0.60 none F03, trail of the repair 2 48.3 s, $0.23 none F04, decision outside the repo 3 two edits by hand lost decision F05, reconstructed reason 3 8.6 s, $0.10 lost decision

reason F06, packet against the spec 4 amendment to the spec incomplete packet F07, boundary without the tests 5 none none F08, rule not enforced final audit 6 files, by hand none The four fixes that consumed session turns add up to $1.14, compared with the $9.44 the build cost. Twelve percent, and that is the easy reading. The hard reading is the cost column in the other four lines: they cost zero in dollars because they were fixed by hand, which means the cost was human attention, which shows up on no invoice and is the project’s scarcest resource. F08 is the extreme case: it cost an audit that only happened because I had decided to publish the code. The last column is the bridge to chapter 26 and the result is uncomfortable. The closed list of failure reasons I use in the first- pass count has seven entries, and all seven are printed here so you can run the test yourself: incomplete packet , unchecked statement , lost decision , restart from memory , mixed topics , stale data and ill-defined task . Of the eight failures, only one falls cleanly into one of them, F05 into lost decision . F04 is a lost decision through a mechanism the category did not foresee, a tool writing outside the repository. F06 is not an incomplete packet: the packet was complete and contradicted the spec. And five failures have no category at all.

That is what a post-mortem is for. The list gained three entries, with the matching technique next to each one: silent partial delivery , when the request comes back halfway with no warning, which points to the four questions of chapter 17 and the instruction to answer before writing; decision by default , when the choice comes in through a default value instead of a statement, which points to the question about external effects in the checklist of chapter 19; and rule with no check , when the rule is written and nothing verifies it, which points to a command at the end of the slice. The list is mine and yours is going to look different, because it is made of the failures your project produced, and not of the ones mine produced. One caveat from chapter 26 worth repeating here: one change per batch. Applying the three new checks all at once, next week, makes the number go up without saying which of them was responsible. The legacy system as counterpoint The legacy packet cut both ways, and both are worth stating. It headed off what would have been the most expensive failure of the set. The toISOString trap was written in the legacy packet with file and line, and session 3 found the risk before writing the first date function, and cited the section on known traps. Without that paragraph, the natural path was to use the Date constructor with the date string, which is the same bug as the 2019 system’s, reproduced in new code, in a project that decided to ignore time zones. And it made one risk worse. The code map of the legacy system is a document, and a document is an authoritative source that ages without warning. In session 2, it stated the format of the weekly

schedule table and cited the map, and the statement was right. But what checked it was not the map: it was a grep in src/schedule.js as it runs at the clinic. A legacy packet is evidence, so check it against the source. The day the map drifts from the code, anyone reading the map alone will confidently assert something that has stopped being true. Post-mortem script Five questions, none of them about the clinic, and none of them asking anyone for more attention. Every answer is a command or a file. First: did the request come back whole? Compare the delivery with the request item by item, and treat a part not delivered and not mentioned as a failure, even when the decision not to deliver it was right. Second: did the slice pick up an external effect, and where does that effect live when the system runs? The question applies to a database, a file, the network and a queue alike. Go find the answer in the code, with a grep , not in what the session tells you. Third: was any decision made today that stayed only in the conversation or only in the tool’s memory? If so, it has to become a project file before the session closes, with the reason and what was discarded along with it. If the tool announced that it saved something, check whether git sees the place. Fourth: did the packet contradict any of the project’s sources? Read layer 1 and layer 2 against the spec before sending. It is the piece nobody checks, because it is the piece that authorizes all the others.

Fifth: which written rule has no mechanical check? Pick one per batch, write the command that verifies it, and run it at the end of every slice. If the rule does not fit into any command, it is going to depend on somebody remembering, and F08 shows how long a rule like that survives without being enforced. Three objections to close. The first is that eight failures in five sessions is a bad number. It is the number a project of five sessions has when somebody writes them down; the honest comparison is not with zero, it is with the same project with no record, where the eight would have happened and none would have a name. The second is that a post-mortem with no production incident is ceremony. VilaSchedule never went into production, and even so two of the eight failures end in a production database opened by the test suite, on the machine that serves the clinic. Here the exercise cost one table and produced three checks; in a project that is already live, it costs the same and the incident costs more. The third is that half of this is my mistake, not the AI’s. True, and the record says so: F06’s packet is mine, F07’s contract is mine, and the cut that left the providers slice without a reading door, the root of F08, is mine as well. That is not a concession tacked onto the end of a chapter. It is the conclusion of the entire part. Context is an engineering artifact, and an engineering artifact fails where somebody designed the failure in.

Powered by TurnKey Linux.