Ви не можете вибрати більше 25 тем Теми мають розпочинатися з літери або цифри, можуть містити дефіси (-) і не повинні перевищувати 35 символів.

12KB

Context Engineering — Chapter-11: Specifications

  • Source: /library/Context Engineering/source-file.pdf
  • PDF pages: 65–73
  • Pages without text: none

Specifications From this chapter on, the book’s examples live in a single place: VilaSchedule, the scheduling system of the fictional Vila Nova Clinic, with its appointments, providers, schedule and work-ins. You are the dev responsible for it, and the story starts on a Friday, with the clinical coordinator asking for the feature the front desk has spent months begging for: the work-in, an extra appointment squeezed into a provider’s schedule for the patient who cannot wait. You open the coding agent and type the request the way it reached you: “Add support for appointment work-ins in the providers’ schedule.” The agent nails it. It reads the right files, follows the project’s conventions, creates the migration, the endpoint and the screen, and delivers all of it with tests passing. You do the review on Monday and the code is clean. Except that the work-in accepts any future date, though the clinic’s rule allows same-day only. There is no limit per provider, though the clinical coordinator had set two per day. The duration came out at 30 minutes, the same as a regular appointment, though the agreed figure was 15. And a schedule blocked for vacation accepts a work-in without complaint. Four business decisions, four plausible guesses, four errors. Notice that nothing the agent got wrong was knowable from the input. The clinical coordinator set the limit of two work-ins per day in a meeting; the manager gave the 15-minute duration in a voice message; the same-day rule lived in your head. Chapter 1 summed up the mechanism that dooms that request: what is not in the window does not exist for the model. The model did not

guess the rules because there was no rule at all in the input, only a one-line request, and it did what models do with gaps: it filled them with the most likely pattern from training, fluently and confidently. The naive fix is the one you probably already reach for: typing a bigger request. It works once and evaporates. The paragraph you wrote in the chat dies with the session, leaves the window when chapter 4’s cycle tightens, and tomorrow another dev, or another agent, redoes the request from memory, with other words and other gaps. Business information typed into a prompt is short- lived context for a long-lived decision. The durable form of that information has a name, and it is older than large language models (LLMs): the specification, or spec. What a spec carries A specification is the document that fixes, before implementation, what to build: the problem, the business rules, the examples that prove the expected behavior and what stays out. I wrote a whole book about that method, Spec Driven Development (https://books.kodel.com.br/en/books/sdd/): executable specs as the unit of work when you build with AI, the flow that goes from intent to implementation and the discipline around it. This chapter does not reteach the method and does not assume you have read the book: the angle here is deliberately narrower. What matters about the spec in this book is one thing only: it is a source of context, probably the densest one your project produces. How to write it and use it to drive development is the subject of that book; what it is worth inside a context window, and why, is the subject of this one.

Dense in what sense? In the sense chapters 5 and 6 gave the word: every token in the window competes for attention and shows up on the bill, so the right question for any artifact is how much decision information it delivers per token it takes up. A well-written spec is almost nothing but decision: rule, limit, example, exclusion. Compare it with the alternatives the cycle usually drags into the window: the Slack thread with twenty messages of social context for one useful sentence, the build log, the 2,000-line file read in full to find one function. The spec is the extract of all those conversations, already filtered by a human who knew what mattered. When you put it in the input, you are handing the model the result of that curation, not the raw material. Here is what that looks like in VilaSchedule. The work-in spec, the one that Friday at the start of the chapter deserved, opens by setting out the problem and the goal:

Context

Vila Nova Clinic works with a schedule of fixed 30-minute intervals per provider. Patients with low-complexity urgency ask to be seen the same day, and the front desk solves that with work-ins: extra appointments accommodated in a provider's schedule without taking a regular open interval.

Goal

Let the front desk record a work-in in a provider's schedule for the same day, respecting the limits the clinical coordinator sets. Notice what those ten lines already settle for a model. “Fixed 30- minute intervals” anchors the vocabulary of the domain; “without taking a regular open interval” kills the most obvious and wrong reading, the one where a work-in is just an appointment in a free interval; “for the same day” already shows up in the goal, ahead of any rule. A model that read that passage does not have to guess what “work-in” means at this particular clinic, and domain knowledge is exactly the kind of information no training contains, because it was born in a meeting only your team attended. Then come the rules, and here is the direct answer to Monday’s four errors:

Business rules

  • A work-in can only be created for the same day; a work-in for a future date is forbidden (that is what the regular appointment is for).
  • Each provider accepts at most 2 work-ins per day; the limit belongs to the clinical coordinator and the front desk cannot change it.
  • The work-in goes into the gap between two consecutive taken intervals and has a fixed duration of 15 minutes.
  • A provider with a blocked schedule (vacation, conference, sick leave) receives no work-in under any circumstance.
  • The work-in records who created it (the front desk user) and the reason the patient gave, both required. Each of those lines is a guess the model no longer makes. Those are 120-odd tokens, and back at the opening of the chapter they would have saved four rounds of rework: the implementation, the review that caught the errors, the meeting to reconfirm the rules and the reimplementation. Chapter 6’s math rarely works out this cleanly.

Examples are the part the model understands best Rules stated in prose still leave room for interpretation. How does “at most 2 work-ins per day” refuse the third one? Silently? With what message? The next section of the spec closes that gap the only way that leaves no room for a second reading, with concrete examples:

Acceptance criteria

  1. Given a provider with 1 work-in today, when the front desk creates the second work-in, then the system accepts it and the day's schedule shows both work-ins between the regular intervals.
  2. Given a provider with 2 work-ins today, when the front desk tries to create the third one, then the system refuses with the message “Work-in limit for the day reached for this provider.”
  3. Given a provider with a blocked schedule today, when the front desk tries to create a work-in, then the system refuses

and states the reason for the block. 4. Given the work-in form with no reason filled in, when the front desk tries to save, then the system refuses and points at the required field. The three-part shape, given, when, then, is a convention borrowed from behavior-driven development, and what makes it worth the ceremony is that it forces each criterion into a checkable form: starting state, action, expected result. Specifying by concrete examples, instead of by abstract rules alone, is established practice from long before generative AI: Gojko Adzic documented it in Specification by Example (Manning, 2011), describing teams that traded ambiguous requirements for key examples they validated with the people who understood the business. The original argument was about humans: examples expose misunderstandings the abstract rule hides. With LLMs the argument picks up another layer: chapter 7 showed that examples inside the input, few-shot, are among the prompt techniques with the most measurable effect. Acceptance criteria are few-shot for behavior: each “given, when, then” is a solved case the model uses as an answer key, from the exact text of the error message to the handling of the empty field. And they pay off twice, because the same criterion that guided the implementation becomes, later, the yardstick for verification: you ask the agent to check the implementation against the four criteria, one by one, and the spec that was input becomes a test. There is one more section, the one almost everybody skips, and for context it is worth as much as the rules:

Out of scope

  • Work-in for a future date (that is a regular appointment).
  • Notifying the patient by text message or WhatsApp (its own spec).
  • Automatic reordering of the schedule after the work-in. Call that negative context: the list of what the model should not build. Models are generous by default; ask for a work-in and a notification system may well be thrown in for free, because in training those features travel together. Each line of the out-of- scope section prunes one of those unasked-for extras before it costs tokens to generate, review and undo. Since chapter 4 taught you that every token circulates in the cycle, saying what not to do stops being bureaucracy and becomes hygiene. The complete file also has a status header and an open-questions section, empty because the clinical coordinator answered the questions before approval. Now run the experiment that closes the argument, in the spirit of chapter 7: the same Friday, the same agent, the same one-line request, but with the spec in the window ahead of it. The four guesses disappear, because they stopped being gaps. The request did not improve; the input did. It is chapter 7’s scenario A, the prompt on its own, turning into scenario B, the same prompt with the context in front of it, and now the context comes from a versioned artifact instead of a heroic piece of typing. The waterfall objection

The resistance you will meet, in the team or in yourself, comes in two classic forms. “Writing a document before coding is going back to waterfall” is the first, and it aims at the wrong target: what made waterfall a problem was the batch size, months of frozen specification before the first line of code, not the act of writing down intent. The work-in spec runs one page and covers one feature; writing it cost less than the meeting it avoided. The second objection is more serious: “the spec goes stale, six months from now it lies.” I grant the fact and reject the conclusion. The spec fixes the intent of a change at the moment it was decided; it is a dated record, like a commit, and a dated record does not lie; it ages. A document that promises to describe the system’s present does lie when it goes stale, and it is a different artifact, with a different maintenance discipline. That distinction matters for your context window. A year from now, VilaSchedule will have work-ins with rules that evolved: maybe three per day, maybe a work-in by telemedicine. Today’s spec will still be useful for answering “why does the limit exist and where did it come from?” but it will be dangerous input for an agent to implement on top of, because it describes the system that was, not the one that is. What is true now has to live in an artifact that follows the code, generated or verified from it, and building that artifact without falling into documentation that rots in silence is the subject of the next chapter. Before you turn the page, hold on to the takeaway from this one: next time a one- line request is about to become an AI session, ask which rules the model would otherwise have to guess at. If the answer is “in a meeting” or “in my head,” you already know which artifact to write first.

Powered by TurnKey Linux.