選択できるのは25トピックまでです。 トピックは、先頭が英数字で、英数字とダッシュ('-')を使用した35文字以内のものにしてください。

13KB

Context Engineering — Chapter-18: Context for brownfield projects

  • Source: /library/Context Engineering/source-file.pdf
  • PDF pages: 130–139
  • Pages without text: none

Context for brownfield projects This chapter is for the developer at the end of the previous chapter: the one who on Monday inherits a system that has been in production for years, with no spec, no doc, no architecture decision record (ADR), with decisions living in the memory of people who have already left. That is brownfield work, the opposite of the greenfield project the chapters before this one assumed, where the tree and the artifacts are born with the first commit. If all your projects were born with the artifacts of Part II, you can skip to Part III with a clear conscience; if you have ever opened a repository and felt you were reading an excavation, this chapter is yours. And a disclosure before we start: the legacy repository used here is a teaching reconstruction. I built a real git repo, with fourteen commits, dates from 2019 to 2024 and fictional authors, modeled on the scheduling system Vila Nova Clinic used before VilaSchedule. The commands and the git output are real, and every one of them can be rerun in that repo; the story behind them is invented to fit the book’s domain. Monday, then. The clinic wants to evolve the old system for as long as the new one does not cover everything, and the repository you receive has nine JavaScript files, zero tests, zero documentation and a message history like “fix” and “tweaks.” Anyone who read chapters 8 through 12 will instinctively spot the whole quartet missing at once, the spec, the living doc, the ADRs and the conventions, and the tempting shortcut is to ask the AI to produce all of it in one shot: dump the repo into the window and type “document this system.” The result looks great. Out comes a fluent summary, with sections, and it says the

system schedules appointments with a configurable duration. It says the work-in limit, the cap on the extra appointments squeezed into a full schedule, is set per profile, and that reminders go out over SMS. Three plausible claims, three lies: the duration is fixed at 30 minutes, the limit is a hardcoded 2, and the reminder goes out over WhatsApp. Chapter 1 explained the mechanism: where the input does not support the answer, the model fills the gap with the most likely pattern from training, and generic scheduling systems have configurable duration and SMS. Documentation extracted this way is chapter 9’s lying doc, produced at industrial scale. The criticism you will hear, and that I partly share, is this: AI hallucinates when it summarizes legacy code, and extraction with no verification produces false context, worse than no context at all, because it enters the window of every future session wearing the face of a fact. My answer is not to avoid AI; it is a routine in which each step extracts one kind of signal, produces a named artifact from Part II and demands proof: every extracted claim points to the evidence that supports it, a file and a line, a commit, the output of a command, or it is explicitly marked as a hypothesis. The steps are ordered from the cheapest signal to the most expensive: first what the repository hands over for free, last the reading that costs tokens and attention. Step 1: structure and names The cheapest signal you already know from chapter 13: the file listing. Before reading any file, ask the repo what it holds: $ git ls-files db.js src/appointment.js

src/block.js src/reminder.js src/report.js src/schedule.js src/utils.js src/whatsapp.js src/workin.js Legacy code rarely screams the domain the way the good tree of chapter 13 does, but it is almost never entirely mute: here the names hand you schedule, appointment, work-in, block, reminder and WhatsApp before you spend a token on reading. This step’s artifact is the start of the persistent file from chapter 12, an honest draft about your own ignorance, with every item tagged by how sure you are:

  • Old Node.js clinic scheduling system (callbacks, var), raw MySQL, no framework in sight. [seen]
  • Topics the names scream: schedule, appointment, work-in, block, reminder, report, WhatsApp. [seen]
  • reminder.js + whatsapp.js: appointment confirmation or reminder over WhatsApp. [?] The [?] marker is the first defense against the objection raised at the start: what is a deduction from a name is stated as a deduction, and the steps that follow either promote each hypothesis or rule it out.

Step 2: git archaeology The second signal is also free and almost always ignored: the history. Adam Tornhill built a whole book, Your Code as a Crime Scene (2015), on treating version history as behavioral evidence: the code says what the system does; the history says where the system hurts. Start with the wide shot: $ git log --oneline 9fb6029 urgent prod fix 23af5af tweaks acb4ef8 whatsapp d1ca331 workin cant happen when schedule is blocked e43a4b8 fix 64ef4a5 schedule block 98a005c change workin limit to 2 (dr cecilia) e309a69 wip e15e1e8 fix workin e9aff58 workin 4efadb2 tweaks c237feb insurance report c5e903a fix 84386f2 first version Bad messages, but information all the same: the system was born in 2019, the work-in arrived in 2020, and one message names a person. The authors tell you who to look for, and the change count per file, Tornhill’s hotspot, tells you where maintenance piles up: $ git shortlog -sn HEAD 5 Paulo Tanaka 5 Renato Alves 4 Marcia Lima $ git log --format= --name-only | sort | uniq -c | sort -rn | head -5 5 src/workin.js

3 src/reminder.js 3 src/appointment.js 1 src/whatsapp.js 1 src/utils.js The most edited file in the repository is the work-in file, which already tells the expensive reading of step 3 where to begin. And the history of one specific file tells the story of a rule: $ git log --date=short --format=’%h %ad %an %s’ -- src/workin.js 9fb6029 2024-05-29 Paulo Tanaka urgent prod fix d1ca331 2022-08-04 Paulo Tanaka workin cant happen when schedule is blocked 98a005c 2021-01-15 Marcia Lima change workin limit to 2 (dr cecilia) e15e1e8 2020-04-10 Renato Alves fix workin e9aff58 2020-04-02 Renato Alves workin There is a fossilized why. The work-in limit did not start at 2; it started at 3 and a physician named Cecilia had it cut. The git blame command (documented, like every command in this section, in the official git reference at git-scm.com/docs) confirms that the line carrying the current limit came from exactly that commit: $ git blame -L 13,13 --date=short src/workin.js 98a005ca (Marcia Lima 2021-01-15 13) if (rows[0].n >= 2) return cb(new Er ror(‘workin limit’)); This step’s artifact is chapter 10’s ADR, in the variant only brownfield work needs: the reconstructed ADR, which records the decision found in the dig and says in its status line how confident it is, instead of pretending it was there from the start:

Status: reconstructed by git archaeology on 2026-08-01; not confirmed with whoever decided it.

Context

The work-in was born on 2020-04-02 (commit e9aff58) accepting up to 3 per day. On 2021-01-15, commit 98a005c, by Marcia Lima, cut the limit to 2 with the message “change workin limit to 2 (dr cecilia)". There is no record of the reason beyond the message. Notice that everything up to here came out of commands, not out of a model’s opinion. The first two steps cost minutes, fit any repo and produce context no hallucination can contaminate, because there was no generation at all: only a transcript of evidence. Step 3: AI-guided reading Now the expensive signal: the code itself. Michael Feathers, in Working Effectively with Legacy Code (2004), defines legacy code as code with no tests, with no safety net to say what it actually does; his central recommendation is to characterize the existing behavior before changing anything. Guided reading is that

characterization done with an agent, and the word that governs it is guided: instead of dumping the repo and asking for a summary, you open one session per topic, and you start with the hotspot that step 2 pointed out. The questions have to be the kind whose answer forces the model to cite the exact place. Not “what does this system do?,” but “which conditions make create in src/workin.js reject a work-in, and what line is each one on?.” A question with an address has a verifiable answer; a panoramic question lets the model’s training answer in the repo’s place. For every claim the agent makes, the rule is the one behind step 1’s markers: either it comes with a file and a line you check in seconds, or it is demoted to a hypothesis, or it is thrown out. The summary from the start of the chapter dies in that funnel: “limit as a parameter per profile” does not survive “show me the line.” This step produces two artifacts. The first is chapter 11’s conventions document, in the observed variant: not what the team agreed on, because there is no team to agree, but what the code repeats with enough consistency for the next session’s agent to imitate:

  • Times are whole minutes from midnight: start and end in src/schedule.js and src/appointment.js; conversion to text only at the edge, in minutesToTime (src/utils.js).
  • Dates travel as YYYY-MM-DD strings (today() in src/utils.js); never as a Date object between modules.
  • Error-first callbacks everywhere; no use of Promise or async/await

in the repository. The second is chapter 9’s living doc for the flow you read, with every claim anchored in the code that supports it:

  1. A work-in is always for the current day: the date comes from utils.today() and is not a parameter of the create function.
  2. A blocked schedule rejects a work-in before any other check (block.isBlocked, error ‘blocked’).
  3. The system counts the provider's work-ins for the day and rejects the request once the provider already has two (error ‘workin limit’). Compare it with the hallucinated summary from the start of the chapter: same model, same repo, and the difference is all in the protocol. The doc with addresses costs more per paragraph, and that is why it comes after the free signals and starts with the hotspot, not with the whole repo. Step 4: generating the artifacts incrementally

The final temptation is the heroic three-month push: repeat step 3 until the whole legacy code base has doc, conventions and ADRs, and only then touch the code. Nobody is going to give you those three months, and they would be badly spent: a good chunk of that code will never be touched again, and context for code nobody touches is a cost with no reader. The last step’s rule is to extract on demand: each real task pays only for the extraction it needs, and the collection of artifacts grows in the order you change the system, which is exactly the order of usefulness. This step’s artifact is the one from chapter 8: the spec of the first real change, written on ground the earlier steps have firmed up. In the clinic’s system, the first task to arrive is to stop a duplicate work-in for the same patient, and the spec opens by citing the extracted behavior instead of restating it from memory:

Business rules

  • A patient can have only 1 work-in per day in the whole clinic, regardless of the provider.
  • The current work-in rules stay as they are; this change only adds the duplicate check. The session that implements that spec gets the persistent file started in step 1, the reconstructed ADR from step 2 and the flow doc from step 3 in its window, and each of those artifacts was cheap because it came in the right order. At the end of the task, whatever the session learned goes back into the artifacts, in the

maintenance cycle that chapters 9 and 12 already described. Six months of tasks later, the legacy code has the quartet in the parts that matter, and nobody ever had to ask for those three months. And Cecilia, if she still sees patients, deserves a visit: the reconstructed ADR becomes a confirmed ADR with one conversation, and the status line records the promotion. That closes Part II: you know how to build context in the project you control from the first commit and how to extract context from the project you inherited with none. What the two situations have in common is the result, a shelf of artifacts: specs, living doc, ADRs, conventions, persistent file, a structure that informs. What neither of them settles is the question every session reopens: out of all those artifacts, what enters the window of this task, in what order, in what form, and what stays out? A full shelf with a finite window is an operational problem, and operating context has techniques of its own: layers, packing, retrieval, validation, compression, isolation. They are Part III, which starts in the next chapter.

Powered by TurnKey Linux.