Du kannst nicht mehr als 25 Themen auswählen Themen müssen entweder mit einem Buchstaben oder einer Ziffer beginnen. Sie können Bindestriche („-“) enthalten und bis zu 35 Zeichen lang sein.

20KB

Context Engineering — Chapter-33: Preparing a project (from scratch and from a legacy system)

  • Source: /library/Context Engineering/source-file.pdf
  • PDF pages: 317–333
  • Pages without text: none

Preparing a project (from scratch and from a legacy system) The directory was empty, and the request was the one everybody makes the first time: “Implement scheduling for the clinic.” The session looked around, found nothing, and what came back was not code. The working directory is empty. There is no existing project for me to base myself on. Since the instruction is quite generic, I need more details before I start implementing. A few questions so we can line up the scope:

  1. Project context: is this a new project from scratch, or is there a repository/code somewhere else that I should use?
  2. Technology stack: which language/framework do you want to use? (e.g. Node.js + React, Flutter, .NET, etc.) Database (Postgres, MySQL, SQLite, Supabase...)?
  3. Scope of “scheduling”: what should this feature cover? [...]
  4. Authentication/access profiles: receptionist, doctor, patient, admin, do they have different permissions? Six questions, with the fourth and fifth trimmed from the block above. Two iterations, 14.5 seconds, $0.06 and zero lines of code. And then came the second turn, which is the one that matters, because it is the turn you have already taken: “I don’t know, do it whichever way you think is best.”

From there the session worked alone for 54 turns. It picked Node.js with TypeScript, Express and Prisma on top of SQLite, modeled patient, provider, service type, weekly availability and appointment, implemented the overlap check, wrote six tests with Jest and supertest, and all of them passed. Cost of the turn: 330.5 seconds and $1.34. Cost of the whole session: $1.40. What it delivered works. And none of it is Vila Nova Clinic’s system.

  • Project and domain named with synonyms and generics (clinic-scheduling, ServiceType), against the convention of using the term the clinic says.
  • ServiceType.durationMinutes leaves the duration configurable. At the clinic, the slot is fixed at 30 minutes.
  • There is no work-in, no waitlist and no schedule block. The three rules the front desk uses every day.
  • Prisma, Express, Jest, zod and supertest went in with no question: five dependencies the real project does not want.
  • src/routes/ and src/services/ organize by technical layer, not by feature. Look at the shape of that list. There is not a single bug in it. The session got right everything that could be gotten right alone, and got wrong everything that depended on knowing where it was. The initial context packet is exactly the list of what cannot be gotten right alone, written to disk before the first session opens. And the list was already there in turn 1. The six questions the session asked are the table of contents of the packet: project context, stack, scope, entities, conventions, access. Preparing a project is answering that in a file, once, instead of answering it in the conversation, every time, with a slightly different answer in each session.

This chapter builds the packet in the two scenarios you will find yourself in: the repository that does not exist yet and the repository that already runs in production. They are different packets, with the same job and with a difference in kind that shows up halfway through. A caveat about the tool The case study runs on Claude Code, and it shows up by name from here to the end of Part V, because a transcript of it is text and fits on a printed page. The choice is not part of the method. In July 2026, Claude Code (Anthropic) reads CLAUDE.md at the root of the project; Cursor reads rules in .cursor/rules/ ; the Codex command-line interface (CLI, OpenAI) and the Gemini CLI (Google) read AGENTS.md . The name of the file changes per tool and will change again with time. The three questions the file answers, which are what to build, how it is done here and where the truth lives, have not changed in any of them. Every artifact this chapter uses is printed here in full, and the companion repository is there for checking, not for required reading: github.com/jckodel/context-engineering-companion- en, tag v1.0-part-v . The path from scratch: three files and a tree VilaSchedule starts as three context files and an empty folder structure. The order you write them in matters, because each one answers a different question and one of them only makes sense after the other two. The spec, which says what to build

First what to build. The versioned spec as the source of scope is the central idea of Spec Driven Development (https://books.kodel.com.br/en/books/sdd/), and Part II of this book has already used it as session input. Here it is the first file of the project, written before a line of code exists, and not documentation produced afterwards to justify what was done.

Context

Vila Nova Clinic books appointments in fixed 30-minute intervals per provider. Today the front desk tracks that in a spreadsheet, and the daily pa in is twofold: an appointment booked on top of another one, and an interval that opens up when someone cancels with nobody telling whoever wa s waiting.

Providers

Every provider has a name, a specialty and a weekly schedule: per day of the week, the time work starts and ends. A day with no hours listed is a d ay on which the provider does not work.

  • A provider is created with at least one working day in the schedule; an empty schedule is rejected with “Provider needs at least one working day”.
  • The range of a day has its start before its end, and both fall on the hour or on the half hour, to match the 30-minute intervals. A range that violates either one of the two is rejected with “Invalid schedule range”. Both rejection messages are quoted word for word, and that is a practical choice: the receptionist reads those exact strings off the screen, so the project convention tells the session to copy the sentence from the spec rather than write an equivalent one. Leave the sentence out of the spec and the session writes its own, a different one every time. The section that looks least necessary is the one that saves the most session time:

Out of scope

Authentication and access control, a graphical interface, notifying the patient through any channel, billing, multi-site operation,

time zones (everything in the clinic's local time) and versioned database migration. Compare that with what session 00 delivered without this paragraph, and with the sixth question it asked in turn 1. An explicit scope does not stop the session from working; it stops the session from working on something else. The conventions, which say how it is done here Second, how it is done here. This is the file session 00 did not have when it chose src/routes/ and src/services/ . Rules that apply to all new code. What a tool checks on its own does not live here.

Organization

  • One folder per feature inside src/: scheduling/, providers/, workins/. The criterion for the folder is the axis of change, not the technical layer.
  • Inside the feature, the files sit directly in the folder. A technical subfo lder

(domain/, infra/, usecases/) is forbidden; whatever grows too large becomes a new feature, never a layer.

  • shared/ only appears when the same code proves necessary in two features. Before that, duplicating is cheaper than abstracting early.

Boundary between features

One feature talks to another through the front door: the public use case of the other feature, imported through its folder path. Importing an internal file of another feature is forbidden, and it is the first sign that the boundary is in the wrong place. The second sentence of the file is what keeps it from growing until nobody reads it. Indentation, semicolons and import order are checked by a formatter; writing that here spends context window to say what a tool already guarantees. What lives in this file is what no linter has any way of knowing. The organization by feature and the rule of the late shared folder come from FOCUS Architecture (https://books.kodel.com.br/en/books/focus/), where shared/ starts out absent, and code only moves up into it when reuse has

proven itself in at least two real slices, written and working. Banning a technical subfolder inside the slice is stricter than FOCUS asks for, and that one is mine: in a small slice, domain/ and infra/ are layering sneaking back in, and skipping them costs nothing while the slice still fits on one screen. The context file, which says where the truth lives Third, the file the tool loads in every session. It comes last because its job is to point to the first two. Chapter 12 laid out the anatomy, and what is worth noticing here is the size.

Where the truth lives

  • What to build: docs/scheduling-spec.md. With no spec in the window, ask before implementing.
  • How we do things here: docs/conventions.md, mandatory for every new file.
  • Why we did it this way: docs/decisions.md, one entry per closed decision, with the reason. Before changing anything that has an entry there, read the entry.
  • What has been built already: src/ itself, one folder per feature.

Rules for every session

  • Domain terms match what the clinic says: Appointment, WorkIn, Provider, Waitlist. No synonyms (Booking, Visit, Slot) and no generics (Item, Entity, Record).
  • An error message shown to the front desk comes from the spec, copied literally.
  • A business rule lives in the use case; the HTTP file translates the error into a status and decides nothing. Four pointers and three rules. The fourth pointer is the most interesting of the set: docs/decisions.md does not exist yet at this point, and it shows up in the middle of the implementation, in the next chapter, when the first decision is closed. Leaving the pointer ready before the file is what makes the session ask where to record something instead of recording it in the conversation. The rules for every session are the three that session 00 broke because it did not know them: the domain vocabulary, where the error message comes from and where the business rule lives. None of them can be deduced from an empty repository. The empty tree

What is still missing is the structure. Three folders with an empty file inside each one, so git includes them in the commit: $ git ls-tree -r --name-only 2de75c4 | grep ‘^src/’ src/placeholder.test.ts src/providers/.gitkeep src/scheduling/.gitkeep src/workins/.gitkeep There is no shared/ . The convention says that folder shows up when the reuse proves itself in two features, and creating it now would offer the session a convenient home for anything it could not place. The names of the three folders are not decoration: they are the spec translated into axes of change, and the session that opens here gets the organization by feature as a done deal instead of a recommendation. That closes the packet: three files, none longer than two pages, and writing all three costs less than the morning session 00 burned building a system for the wrong domain. Somebody always objects at this point that a packet is just waterfall sneaking back in. It is not, because nothing here freezes a decision. I amended the VilaSchedule spec in the middle of the implementation, during the fourth session, when the monthly report came into scope and the clinical coordinators ruled that a canceled appointment does not count. The file is versioned exactly so it can change and leave a trail. The other objection is stronger: modern AI infers all of this from the code. It does, and it infers a lot. Session 00 picked a plausible stack, implemented the conflict check nobody asked for explicitly and wrote tests on its own. What it got wrong was the domain vocabulary, the fixed interval, the work-in and the organization by feature. A work-in is the clinic’s name for the

15-minute appointment squeezed into a day that is already booked, which is not the walk-in an American front desk would picture. Those four things were not on disk anywhere, and no amount of intelligence pulls them out of thin air. The legacy path: a packet that shows its evidence The other scenario is the common one. The clinic’s system has been running since 2019, has nine files, 160 lines, zero tests and zero documentation. There is no spec to version; there is code to read. The old system’s repository used here is a teaching reconstruction: the files, the commits and the command outputs are real and reexecutable, and the story behind them is made up for the book. Chapter 15 has the technique, and this chapter applies the four steps without teaching any of them again: structure and names, git archaeology, AI-guided reading and incremental generation of the artifacts. Compared with the path from scratch, which the rest of this chapter calls greenfield, what changes is the nature of what comes out. In greenfield, the packet declares intent, and you are the authority. In legacy, the packet reports what exists, and the authority is the code. That has a consequence for the format: every claim carries a file and a line, and whatever was not confirmed gets marked. The old system has no spec, no living documentation and no architecture decision record (ADR), so there is nothing to inherit and everything to extract. This file is the persistent context of the system that has been

running at the clinic since 2019, extracted from the repository itself on 2026-07-29. The system has no spec, no doc and no ADR: everything here was pulled from the code and from the git history, and every claim carries the evidence that holds it up. An item marked with [?] is an unconfirmed hypothesis; treat it as a question, never as a fact. The section that does the heavy lifting is the one with the rules the code enforces today. They are written nowhere in the old system, and breaking any one of them produces a bug no test catches, because there is no test.

  • A time is an integer number of minutes since midnight: the intervals are built from start and end of the weekly_schedule table and advance 30 at a time (src/schedule.js:10). The conversion to text happens only at the edge, in minutesToTime (src/utils.js:6).
  • A work-in is always for the current day: the date comes from utils.today() and is not a parameter of create (src/workin.js:6).
  • The limit of 2 work-ins per provider per day is hard-coded, not

configurable (src/workin.js:13).

  • The limit is checked before the reason and the user (src/workin.js:13 to :15), so a work-in with no reason on a full provider returns “workin limit”, never “no reason”. The last one is the kind of rule that only shows up in a line-by- line reading: the order of the checks is observable from the outside, because it decides which error message reaches the front desk. A session that rewrites that function in the “natural” order changes the message without changing the apparent behavior, and nobody notices until the receptionist calls in complaining about an error that makes no sense. Then comes the section greenfield does not have, and the most valuable one in the two pages:

Known traps

  • today() uses toISOString, which returns the date in UTC (src/utils.js:3). After 9 p.m. in the Brasília time zone, the current day of the work-in becomes the next day. No test covers that.
  • src/workin.js is the most touched file in the repository, with 5 of the 14 commits. A change in there has a history of breaking production (commit 9fb6029, “urgent prod fix”).
  • The commit c237feb is called “insurance report”, but the current code of src/report.js has nothing about insurance. [?] An inherited trap is a known bug that is not going to be fixed now. The first one on that list is the date arriving in coordinated universal time (UTC) while the clinic reads it in Brasília time, three hours behind. Writing the trap into the packet does two things at once, pulling in opposite directions: the session does not reproduce the pattern in new code, and the session does not “fix” the pattern along the way in a task that was about something else. The toISOString bug comes back in the next two chapters, and in both ways. The git archaeology goes into the code map, and it is the map that answers the third predictable objection, which is that the legacy system is too big to map: $ git log --format= --name-only | sort | uniq -c | sort -rn | head -3 5 src/workin.js 3 src/reminder.js 3 src/appointment.js

The map does not need to cover the system; it needs to cover the next change. The one at the clinic has nine lines because the system has nine files; in a system of nine hundred, you map the slice you are going to touch, and the command above says which one that is. Nobody at the clinic had to be interviewed to find out that work-ins are the hot spot of the repository. | File | Subject | Talks to | |---|---|---| | db.js | MySQL pool and query | everybody | | src/schedule.js | generates the intervals of the day from the weekly sche dule | db | | src/appointment.js | books an appointment | db, schedule, block | | src/workin.js | creates the work-in of the day | db, utils, block | | src/block.js | says whether the schedule is blocked | db | | src/reminder.js | builds and fires the reminder | db, whatsapp, util s | | src/whatsapp.js | POST to the messaging API | https | | src/report.js | counts the appointments of the month | db | | src/utils.js | today's date and a readable time | nothing |

The third legacy artifact is the one that takes the place of the greenfield conventions, with a difference that sits in its first sentence: Nobody agreed on these rules: they were read from the code on 2026-07-29 and they stand as a description to imitate, so that what the sessi on writes does not clash with what is already there. Each one cites where it was observed.

  • Asynchrony via error-first callbacks, always. No Promise in the repository (src/appointment.js:5, src/workin.js:5).
  • A business rule error becomes a new Error with a short lowercase string: 'taken', 'workin limit', 'no reason' (src/appointment.js:19, src/workin.js:13).
  • Table and column names in snake_case (provider_id, weekday, created_by).

An observed convention is not a desired convention, and the file says so to your face. Nobody on this project wants var and callbacks; what I want is for the new code not to clash with the file it enters, and the request to imitate the local style is what avoids the partial modernization that leaves half the file in one paradigm and half in the other. The decision to modernize exists, it is big and it belongs to another conversation, not of today’s task. How to know the packet is ready There is no checklist that closes this, but there is a cheap test: open a session with the packet loaded and give it the first real request. If it asks something that is already written, the packet is not pointing where it should. If it asks something that is written nowhere, you have just found the next line of the packet, and it cost you one question instead of a morning. What the prepared session does not do is ask the six questions this chapter opened with. It can keep asking, and it will: about the order of two checks, about an edge case the spec did not foresee, about where to record a decision it has just made. Those are the good questions, the ones that need a person. The six from turn 1 did not. VilaSchedule now has a spec, conventions, a context file and three empty folders. The next chapter opens the first session on top of that and goes all the way to the system working, in five sessions, with the cost in dollars of each one. Two things this chapter planted come back there: the pointer to the docs/decisions.md that does not exist yet, and the toISOString bug in the old system.

Powered by TurnKey Linux.