No puede seleccionar más de 25 temas Los temas deben comenzar con una letra o número, pueden incluir guiones ('-') y pueden tener hasta 35 caracteres de largo.

44KB

FOCUS Architecture — Chapter-21: FOCUS + AI: The Duo That Scales

  • Source: /library/FOCUS Architecture/source-file.pdf
  • PDF pages: 640–676
  • Pages without text: none

FOCUS + AI: The Duo That Scales In this chapter, you’ll: write an architectural prompt for a feature of your own, with folder structure, pasted contracts, and explicit bans; review a generated diff with the slice checklist and name the chapter 19 card each violation breaks; say which FOCUS piece neutralizes which GitClear number, and by what mechanism. Chapter 3 opened a debt. It measured code degrading in the AI era, pointed at the three guard-rails that would catch it, and promised that FOCUS as a whole would close the argument here. Time’s up. You’ll watch the same Rosie’s Coffee Shop feature come out of a model four different ways, count the differences instead of reaching for adjectives, and leave with a prompt and a checklist ready to use tomorrow. The idea holding up this chapter is the narrow search space, and it’s already been yours since chapter 3: the fewer correct ways there are to complete the code, the fewer chances the generator has of picking the wrong one. What was missing was the other half. Narrowing the search space has two dials, not one. One tells the generator WHAT to build, and it’s called the specification; the other tells it WHERE each decision lives, and it’s this book’s architecture. Turn only the first and you get the right thing in a

tangle. Turn only the second and you get the wrong thing, neatly arranged. And turn neither and you’re doing vibe coding (Karpathy, February 2025), with the results you already know. The one-sentence request Before any technique, the disaster. I asked Claude Opus 4.8, in July 2026, to implement Rosie’s Coffee Shop’s tab split in Dart. The model generated the code; I asked and read what came back. Each run went in a clean context, with a disposable configuration profile, no file from this project nearby, and the word FOCUS never appearing anywhere. One caveat before the numbers, and it holds for the whole chapter. One run per condition isn’t a controlled study. A model’s output isn’t deterministic, and if you repeat the experiment you’ll get a different result, probably better in one spot and worse in another. The studies that back the argument are chapter 3’s: GitClear, DORA, and arXiv. What follows is my own observation, with the model and the date on its badge. The vague request, in full: “People at Rosie’s are always asking to split the table’s bill. Can you build that in Dart for me?” One sentence. It’s the request Rosie herself would make at the counter, and it’s also the request plenty of people paste into the chat at eleven at night. What came back was 530 lines of code, not counting comments and blank lines. An interactive terminal menu, three split modes (even, by consumption, by item), a configurable service charge, a discount field, and a settle-up algorithm for when one person pays for everyone and the rest square up later. None of it was asked for. The request had one idea; the output has seven.

A quick vocabulary note, the same diff chapter 11 already defined: from here on, the diff is always whatever the model handed back. Three excerpts from the output, copied as they came back, with the location comments added by me: Dart // splitter.dart: business refusal thrown as an exception void addPerson(String name) { final n = name.trim(); if (n.isEmpty) throw ArgumentError(‘Empty name.'); if (people.any((p) => p.toLowerCase() == n.toLowerCase())) { throw ArgumentError('“$n” is already at the table.'); } people.add(n); } // main.dart: the funnel that catches all three exceptions and moves on void _attempt(void Function() action) {

try { action(); } on FormatException catch (e) { print(’ x ${e.message}'); } on ArgumentError catch (e) { print(’ x ${e.message}'); } on StateError catch (e) { print(’ x ${e.message}'); } } // test.dart: the same five-line block, repeated three times check(‘rejects negative price’, () { try { Item(‘x’, -1); return false;

} on ArgumentError { return true; } }()); Read the first excerpt by what kind of refusal it is. “This person is already at the table” is a business decision, exactly as legitimate as “this tab is already paid,” and the code turns it into an ArgumentError , the type Dart reserves for a programmer’s malformed argument. Empty name, negative price, a count under one, a tab with nobody on it: all of it becomes an exception. The second excerpt shows where those exceptions go to die. _attempt catches FormatException , ArgumentError , and StateError in the same funnel, prints the message, and hands control back to the menu. A typo and a business refusal leave through the same pipe, with the same face. That’s chapter 19’s Card 5, domain try/catch, and it’s the damage that chapter described as the one the tools commit straight out of the box. The third excerpt is a different animal. That five-line block shows up three times in a row in the test file, body swapped and frame identical, and a second five-line block shows up twice more further down. GitClear, on the same definition chapter 3 already used, calls any five-or-more-line stretch that reappears after normalizing whitespace and stripping comments a duplicate. That’s two duplicated blocks in a single output. None of chapter 19’s six cards covers duplication, and I’d rather say so than force the fit: duplication isn’t a slice anti-pattern, it’s the metric the vertical slice neutralizes, and it comes back next section with the number attached.

One pain I expected never showed up, and it’s worth recording rather than hiding. I expected files touched outside the feature, some utility being born in a shared directory, some global config changed. None of that happened, in any run. The model stayed inside the feature the whole time, and the measure of files touched outside the slice came out zero in every condition. Here, it didn’t separate anything. Of the three error cases the feature needs to handle, the vague request’s output handles one. A people count under one has a path, via the exception you just saw. Tab not found and tab already paid don’t exist in the code: with no word for “tab” in the request, the model never even invented the concept of a tab stored somewhere. You can’t handle the error of a concept that never got born. And then the experiment proved me wrong I ran the same condition, still with no structure and no architecture, with a complete statement this time: the three refusal cases named, the integer-cents requirement, and the requirement that the parts sum to exactly the total. Still no folder structure, no prescribed return type, no ban of any kind, and no word FOCUS. The model got it right. Exhaustive types for the three refusal cases, integer cents from start to finish, zero try/catch in the entire code, zero duplicated blocks, a passing test file, 200 lines. It’s not the result I expected, and it’s the result this chapter publishes. Hiding this, or rerunning the round until a disaster showed up, would be fabricating the same anti-solution chapter 3 accuses AI of fabricating.

What the comparison shows, then, is a single variable: the statement. With a one-sentence request, the generator invented six features nobody wanted and handled one error case out of three. With the request spelled out, it got there without receiving any architecture at all. That’s exactly the “spec” half of this chapter’s argument, demonstrated by accident and in full: the specification says what, and without it the generator decides what on its own. That leaves the question that matters. If a good statement already produces good code, what does architecture add? The answer isn’t in the first generation, and the round 2 section goes looking for it where it lives, which is the same feature’s second change. One more observation, from a run I threw out for a method flaw. This experiment’s first attempt ran in an environment I thought was isolated and wasn’t: the generator inherited my own coding instructions left over in the context, and handed back code with a sealed Result, comments in the format I use, and even a final report in my own working style, none of it asked for by the five- line statement. That run was voided for the count. But it’s accidental evidence for the thesis: a model with architecture contracts in its context produces structured code without the request asking for it. I found that out by getting the isolation wrong. Each piece against a number The three GitClear numbers chapter 3 asked you to hold on to aren’t a portrait of tragedy. Each one has a FOCUS piece that neutralizes it, and the mechanism is concrete in every case, not a hope for discipline.

Code duplication rose 81% relative to the pre-AI era (GitClear, 2024-2026). The piece is chapter 11’s vertical slice, and the mechanism is the generator’s reach. A model recreates the utility it can’t see; that’s how chapter 3’s loyalty rule got its second copy. When the whole feature fits in one directory, the context you paste is the slice, and whatever already exists inside it is in plain sight, so recreating it stops being the path of least resistance. Duplication between different slices still exists, and chapter 19 already signed that lease with its eyes open. Error masking rose 47% (GitClear, 2026). The piece is the exhaustive Result from chapters 8 and 15, and the mechanism is the compiler. A switch over a sealed family doesn’t compile with a case missing, and a generator wanting to swallow the failure would have to write, in visible code, the line that swallows it. Ignoring it is still possible; ignoring it silently isn’t. That’s the difference between a reviewable decision and an invisible omission. Refactoring dropped from about 25% of changed lines in 2021 to under 10% in 2024 (GitClear, 2024). The piece is chapter 14’s pure use case, and the mechanism is the test with no infrastructure. Nobody refactors what they can’t tell is broken. A rule that lives in a pure function has a test that runs in milliseconds, no database, no network, and no screen, and the cost of touching it drops to the point where touching it is worth it. A rule scattered across four files with I/O in the middle has the opposite cost, and what happens to it is new code piling on top, which is precisely what the number measures.

Outside GitClear the direction repeats, and it’s what convinces me the numbers aren’t an artifact of one methodology alone: the 2024 DORA research and the arXiv 2409.19182 study measured delivery speed climbing while stability and maintainability fall, when nothing constrains the shape of what the generator produces. Three independent sources, same direction. Their target is never the tool; it’s the terrain. The spec says what, the architecture says where

The flow above is Spec Kit’s, and it fits on half a page. /speckit- specify takes the request in plain language and produces a specification: what the feature does, for whom, with which error cases and which acceptance criteria. No line of code shows up at this stage, and that’s on purpose: the specification is where questions are still cheap to answer. /speckit-plan takes the finished specification and decides how to implement it in this project, with this architecture, in these languages. Only then comes generation. And the arrow that matters most is the one going back: when review flags something, what changes first is the specification, not the generated code. Fixing only the code lets the next generation repeat the same mistake, because its source is still sitting there.

What you just read is the skeleton. Each of those stages has its own rules, its own pitfalls, and a way of going wrong that only shows up after the third feature; to follow this chapter, the skeleton is enough. That’s what Spec Driven Development (2026, https://books.kodel.com.br/en/books/sdd) covers, where I teach you to run the flow I only use here, and you do not need to have read it to go on from here. Now the argument itself, with both sides failing on their own. Spec without architecture you’ve already watched work, and it’s the complete-statement case from the previous section: the right thing, with all three cases handled: everything inside one file that runs the calculation, holds the data, and prints the output. The program is correct. The question “where do I go when the rule changes” has exactly one answer, and it’s “somewhere in that file.” Architecture without spec is the reverse picture, and I’ve seen it happen more times than I’d like: the generator gets the four folders, the contracts, the bans, and hands back a textbook slice that implements a feature nobody asked for, with split-by- consumption and a service charge each neatly tucked into its own file. Every piece in the right place, and the problem it solves is the wrong one. The pair works because the two dials constrain different things. The specification cuts down the space of WHAT can be built; the architecture cuts down the space of WHERE each decision can live. Turning only one leaves the other axis free, and the free axis is where the generator improvises. Anatomy of the architectural prompt An architectural prompt is the prompt that names your architecture’s layers, types, and contracts, pastes those contracts into the request itself as code, and lists what’s off-limits. It

doesn’t describe the implementation. It describes the box the implementation has to fit inside. This is the prompt that generated round 2, in full, unedited and with no extra context. It went to the model exactly as it appears here: Implement Rosie's Coffee Shop's “split tab” feature, in Dart. What the feature does: the cashier gives the table number for an open tab and the number of people; the system returns how much each person pays. Values are integers in cents, never floating point, and the parts must sum to exactly the tab's total. Refusal cases the business already knows about: a people count under 1, tab not found, tab already paid.

Where each thing lives

The feature is a vertical slice. Create exactly these folders and nothing outside them: features/tab/ view/ what the person sees; decides nothing orchestrator/ receives the event, calls the use case and repository, publishes state usecases/ the business rule, as a pure function data/ data access; knows about the outside world

Contracts that already exist, use these, don't invent others

// infrastructure failure, never a business refusal enum Failure { noConnection, unavailable, notAuthorized } // what the repository returns sealed class LookupResult {} class TabFound extends LookupResult { TabFound(this.tab); final Tab tab; } class TabNotFound extends LookupResult {} class InfraFailure extends LookupResult { InfraFailure(this.failure); final Failure failure;

} // the event the view fires class SplitTab { const SplitTab({required this.table, required this.people}); final int table; final int people; } // the state the view receives sealed class TabState {} class Loading extends TabState {} class SplitReady extends TabState { SplitReady(this.data); final SplitData data; } class SplitRejected extends TabState { SplitRejected(this.reason); final String reason; }

Bans

  • Don't create or change any file outside features/tab/.
  • Don't use try/catch for business refusal. Business refusal is a return value, with its own type, and the caller decides with an exhaustive switch.
  • The use case doesn't receive a repository, doesn't do I/O, and doesn't import anything from data/. Data comes in as an argument, the result goes out as a return value.
  • The orchestrator doesn't calculate any business logic. It calls the use case.
  • The view doesn't decide anything. It fires the event and renders the state.
  • Don't create an interface with a single implementation just for ceremony.

Definition of done

The code compiles with Dart 3 and runs. A demo main exercises all four paths: a successful split with a remainder ($100.00 among 3 people should give 3334, 3333, 3333 cents), an invalid people count, tab not found, and tab already paid.

One caveat before walking through the parts, because it jumps out at anyone who read ch. 11. The prompt above is from July 2026 and asks for four technical-role folders inside the slice. FOCUS doesn’t ask for that today: the slice is flat, and what grows inside it is a sub-feature, not a drawer. The prompt is printed as it was handed over that day because it’s what produced the measurements in this section, and rewriting it here would mean showing an input nobody ran. If you’re going to use this prompt, swap the second part for the flat tree from ch. 11; the rest stands as is, and what the section measures still holds, because what was being measured is the effect of constraining structure, not the specific folder layout. Five parts, and it’s worth walking through them one by one, because each one cuts a different axis of the search space. The first part is the business goal, and it’s the specification in miniature: what the feature does, with what input, with what result, and the integer-cents requirement chapter 20 taught me to never leave implicit. The three refusal cases come named. It’s the same information that made the difference between the two no-architecture runs, and it’s still just as necessary here: no structural ban fixes a statement that doesn’t say what the business refuses. The second part is the folder structure, written as a tree, with one sentence per folder stating its job. Notice the sentence describes responsibility, not content. “what the person sees; decides nothing” is more restrictive than a list of files, because it holds for files that don’t exist yet. The third part is the one most people forget: the contracts pasted as code, not described in prose. Describing a type in prose leaves the generator free to reinvent it under another name, another shape, and another semantics, and you get back a generic Result<T, E> where there should be a LookupResult with three variants. Pasted,

the type is a hard constraint: the model continues the text you started. It’s the same reason SplitData shows up in the state contract without being defined in the prompt; the generator has to produce it under that name for the rest to fit. The fourth part is the bans, and they’re the direct translation of chapter 19’s cards into the language of the request. No file outside the slice closes Card 6. No try/catch for business refusal closes Card 5. A use case that doesn’t receive a repository closes Card 2, an orchestrator that doesn’t calculate closes Card 1, and the ban on a single-implementation interface closes Card 4. The bans are negative on purpose. They say where you can’t go, and leave the how up to whoever generates. The fifth part is the definition of done, with the canonical output spelled out: $100.00 among 3 people gives 3334, 3333, 3333. A verifiable criterion in the prompt is worth more than three paragraphs of desired quality, because the generator can check it on its own before handing you the result. Try it: copy the prompt above, swap in your own feature and contracts, and send it to the model you use. Then count how many of the five parts you’d have written without this list. My bet is on the first and the fourth; the third is the one that usually gets left out, and it’s the one that holds the result together the most. The same feature under the prompt The previous section’s prompt went to a clean context byte for byte, in the same isolation as the other runs. It needed no revision: the first output already brought the slice’s files in the

four folders and compiled, which was the agreed criterion. What came back was 198 lines of code across six files, and the two listings below are what the model handed back, uncurated by me. The listings come out in Dart only. The prompt named Dart, so that’s the run that existed; printing the same slice with the TypeScript symbol stacked next to it would claim the model produced this in both languages, and nobody ran the second one. The equivalent slice in all ten official languages is in the Try It at the end of the section. The use case, in full: Dart import ‘tab.dart’; /// Business refusal is a return value, not an exception. Whoever calls /// decides with an exhaustive switch. sealed class SplitResult {} class SplitCalculated extends SplitResult { SplitCalculated(this.data); final SplitData data; }

class InvalidPeopleCount extends SplitResult { InvalidPeopleCount(this.people); final int people; } class TabAlreadyPaid extends SplitResult { TabAlreadyPaid(this.table); final int table; } /// Pure function: data comes in as an argument, the result goes out as /// a return value. No repository, no I/O, no import from data/. /// /// The remainder is handed out one unit at a time to the first people /// in line, so the parts sum to exactly the total. SplitResult splitTab({

required Tab tab, required int people, }) { if (people < 1) { return InvalidPeopleCount(people); } if (tab.paid) { return TabAlreadyPaid(tab.table); } final base = tab.totalInCents ~/ people; final remainder = tab.totalInCents % people; final parts = List.generate( people,

(index) => index < remainder ? base + 1 : base, ); return SplitCalculated( SplitData( table: tab.table, totalInCents: tab.totalInCents, partsInCents: parts, ), ); } The file imports exactly one thing, tab.dart , the folder neighbor holding the data types. No import from data/ , no repository in the signature, no I/O call. Both business refusals come out as variants of the return; each one carries the data the screen will need for its message: InvalidPeopleCount carries the number given, TabAlreadyPaid carries the table number. The doc comments came from the model too, and they repeat the prompt’s bans almost

word for word, which is a cheap and useful signal: the generator recorded the constraint in the artifact, and whoever reads the file next gets the constraint along with it. The orchestrator, without the event and state types the prompt already pasted: Dart /// Receives the event, calls the repository and the use case, publishes /// state. Calculates no business logic: translates results into state. class TabOrchestrator { TabOrchestrator({ required TabRepository repository, required void Function(TabState) publish, }) : _repository = repository, _publish = publish; final TabRepository _repository; final void Function(TabState) _publish;

Future onSplitTab(SplitTab event) async { _publish(Loading()); final lookup = await _repository.findTab(event.table); switch (lookup) { case TabNotFound(): _publish(SplitRejected(‘Table ${event.table}: no tab found.')); case InfraFailure(:final failure): _publish(SplitRejected(_describeFailure(failure))); case TabFound(:final tab): _emitSplit(tab, event.people); } }

void _emitSplit(Tab tab, int people) { final result = splitTab(tab: tab, people: people); switch (result) { case SplitCalculated(:final data): _publish(SplitReady(data)); case InvalidPeopleCount(people: final given): _publish(SplitRejected( ‘Invalid people count: $given. Must be at least 1.', )); case TabAlreadyPaid(:final table): _publish(SplitRejected(‘Table $table: tab already paid.')); } }

} Two exhaustive switches, one over the repository’s result and one over the use case’s result, and no arithmetic in between. The three error cases show up in two different places, and that’s design, not carelessness: invalid people count and tab already paid are refusals the business rule knows about, so they come out of the use case; tab not found is missing data, so it comes out of the repository as TabNotFound . The orchestrator is where the two families turn into the same thing for the screen, which is a message. The generator also reported, unprompted, two decisions it had to make on its own. It put the demo main in view/ , because the ban on creating files outside the four folders left no room for the program’s composition point; that’s my prompt’s flaw, not the generator’s. And it noted that the InfraFailure arm is handled but never exercised, because the demo repository is an in-memory map with no way to go down. It chose to say so rather than invent an artificial failure just to make the case count look tidy. Now the three measures, counted with the same definition across all three outputs, before any prose: Measure Vague statement Complete statement Architectural prompt files touched outside the slice 0 0 0 duplicated blocks of 5+ lines 2 0 0 error cases 1 3 3

handled, out of 3 domain try/catch 1 0 0 lines of code 530 200 198 The file measure is files outside the slice, not the total file count, on purpose. The output under the architectural prompt has six files against three for the others, because the slice has four folders; counting the total would rank size and call the expected result a defect. Outside the slice, every file touched is coupling the boundary should have blocked, and there the number ranks quality in the same direction across all three outputs. This round it came out zero across the board, so it didn’t separate anything. Look at the table without playing favorites. Columns two and three match on the first four rows. Over a good statement, architecture didn’t improve any of the three measures, for the simplest reason there is: they were already on target. Anyone trying to sell architecture with this table is selling what it doesn’t show. The fourth measure: the second change This book’s yardstick has always been a different one, and it’s time to use it. Specification and architecture don’t pay off on the first draft; they pay off on the same feature’s second change, when someone needs to touch what already exists. So I handed all three outputs the same new business request, in the same isolation, with no mention of architecture in any of them: each person pays their own part separately, the cashier marks who’s

paid, and the tab only closes once every part is paid. Each base became a repository with an initial commit, and the diff was measured against it. Measure Vague Complete Architectural files touched 3 of 3 3 of 3 7 of 7, 1 new lines added 396 370 330 lines removed 1 46 51 where the new rule lives +147 in the usual file +248 in the same file new file, 61 previous behavior broken preserved preserved The row that separates the three outputs is the fourth. Under architecture, the new rule was born as its own 61-line file, one pure function next to the others, and the rest of the diff is wiring: the orchestrator gained an arm, the view gained a button. Without architecture, the same rule went in as 248 lines inside the file that already held everything, and that file went on to do one more thing. Both programs work. The difference isn’t in working; it’s in the answer to “where do I look for this next time,” which is the question you’ll ask six months from now, probably with Rosie waiting on the phone. Now the row I won’t use. The vague-statement base broke previous behavior, and the other two didn’t. It’s tempting to say architecture prevented the regression, and it would be false: the complete-statement base, which has no architecture at all, survived exactly as well as the one that does. I chased the hypothesis that the other two had the same defect hidden by a

missing test, gave all three the same isolated edge case (split 3 cents among 4 people and pay part by part), and the defect only exists in the vague base. The defect is real, and it belongs to the vague statement alone. But the variable that produced it is still the statement, not the architecture. This data supports the claim that architecture changes where the change lands; it doesn’t support the claim that it prevents regression, and I’d rather publish the smaller, true claim. One last thing the fourth measure showed, and one I hadn’t predicted. The second change’s request repeated no contract: it said nothing about Result, about refusal as a value, about pure functions. The output under architecture handed back the new rule with business refusal as a return value, created explicit states for an open tab and a closed tab, and refused to re-split a tab with a part already paid. The contracts pasted into the first prompt kept governing the second generation without anyone repeating them, because they were sitting in the code the generator read before it wrote anything. A contract that lives in the repository doesn’t need to be pasted twice. Try it: open https://focus.kodel.com.br/en/dart/21-01 (or swap dart for kotlin , ts , java , csharp , go , php , python , swift , or rust ) and run the full slice, with all four paths. Then delete one variant from the orchestrator’s switch and watch your language’s compiler demand the missing case. In languages with no exhaustiveness check, the same route shows what stands in for the compiler instead, which is chapter 18’s discipline. Slice-guided review

You received a whole slice at once. Six files, 198 lines, all of it plausible. Reading in the order the model wrote it is the worst option available: that order is generation’s order, not the system’s, and it leads you to judge each file by what it looks like instead of by the place it occupies. Slice-guided review is walking the diff in FOCUS’s order, view, event, orchestrator, repository, and use case, and asking the question that fits at each stop. You don’t read files; you follow a piece of data’s path from the customer’s finger to the database and back. A file can get visited twice, and sometimes it does. Architectural review checklist is the list of questions you ask at those stops. It isn’t new: it’s chapter 19’s six cards, same names and same order, rewritten as a question you can answer by looking at the diff. Keep the difference between the two lists straight, because it’s confusing on first read. The numbers are still the cards’ numbers; the order you ask the questions in is the slice’s order, and the two don’t line up. The first question you ask before opening a single file, just looking at the diff’s list of paths: did a new file show up outside the slice? That’s Card 6, premature shared/ , and it’s cheap because the answer is in the file names. Then start walking. At the orchestrator stop, Card 1: is there business calculation inside the event handler? At the use case stop, Card 2: did the signature pick up a repository dependency? At the repository stop, Card 3: does some type serve the whole coffee shop and speak the database’s language? Card 5, domain try/catch, gets three stops instead of one. At the orchestrator and the use case the question is whether there’s a try/catch around a legitimate business refusal, and the right

answer is none. At the repository boundary the try/catch is legitimate, and the question changes: does it translate the exception into a value, or does it swallow it and move on? That leaves Card 4, layer for ceremony, which has no stop of its own because it has all of them. A file that only passes things through, an interface with one implementation, a DTO identical to the model, a field-by-field mapper: it’s the question you repeat every time you open a new file in the diff, wherever it sits along the path. Apply it to the vague statement’s diff and see what happens. Card 6 answers before you read a single line: no file outside the slice, and that’s the answer for every run in this chapter. Card 1 finds no orchestrator to flag, because that output doesn’t have one; the rule and the menu live in the same place, which is worse than Card 1’s damage and isn’t Card 1’s damage. Card 5 flags it, and flags it twice: ArgumentError for “this person is already at the table” is business refusal turning into an exception, and _attempt catching three exception types in the same funnel is Card 5’s generic catch in the flesh. Card 3 is a clean pass, and it matters as much as the flags do. There’s no generic repository in that output because there’s no repository at all; the data lives in lists inside the tab object. The checklist passes clean, and passing clean is the right answer. A checklist that flags every single item isn’t reviewing, it’s complaining, and you stop trusting it by the third time. The case Card 5 catches in any language Card 5’s symptom shows up with a different accent in every language. In Python it takes an almost idiomatic shape, and it’s the one that slips past review the most: Python

class TabRepository: def init(self, service: TabService) -> None: self._service = service def find_tab(self, table: int) -> Tab | None: tab = None try: tab = self._service.read(table) except Exception: pass return tab Point this repository at a service that’s up and ask for a tab that doesn’t exist. Then point it at a service that’s down and ask for a tab that does exist. Both calls return None . The cashier’s screen will say the same thing in both cases, and they’re opposite cases:

in the first, the tab doesn’t exist and the cashier needs to double- check the number; in the second, the tab exists and it’s the system that failed to read it. Why does the generator prefer this shape? Because the apparent goal of whoever’s asking is that the program doesn’t crash, and swallowing the exception meets that goal in one line, with no need for the generator to know what the business wants when the service goes down. Handling it for real costs a decision the prompt never gave. And Python has no exhaustive switch demanding the missing variant, so nothing in the environment complains; the program runs, the tests pass, the rushed reviewer sees three harmless lines. It’s the same mechanism as chapter 18: where the language doesn’t stop you, convention has to. The fix is the one from chapters 8 and 15. The try/catch stays where it’s legitimate, at the repository boundary, and translates the library’s exception exactly once, into a value the rest of the system understands: Python LookupResult = TabFound | TabNotFound | InfraFailure class TabRepository: def init(self, service: TabService) -> None: self._service = service

def find_tab(self, table: int) -> LookupResult: try: tab = self._service.read(table) except ConnectionError: return InfraFailure(Failure.NO_CONNECTION) if tab is None: return TabNotFound(table) return TabFound(tab) Three changes, all small. except Exception became except ConnectionError , because catching everything also catches the AttributeError from your own typo. pass became a named return, InfraFailure , the type from chapters 8 and 15. And missing data stopped being the same thing as a read failure: now they’re two distinct variants, and whoever calls has to choose what to do with each one. The two calls from the previous paragraph now print different things, which is the least you’d expect from two different cases. Two criticisms I take seriously

The first criticism is the strongest one this chapter faces, and it deserves the unvarnished version: models are going to keep improving, generation after generation; two years from now the generator will produce better-structured code than the average team produces today, and this whole apparatus of pasted contracts, bans, and checklists will be dead weight nobody maintains. It’s happened before, to other defensive disciplines. My answer is a dated fact. Between 2024 and 2026 the models improved a great deal, and GitClear’s numbers got worse over the same period: the 47% rise in error masking was measured in 2026, over the most capable generation to date, not over the 2022 models. If generator quality solved the problem, the curve would have turned. It didn’t turn because the bottleneck was never generation. It’s review and maintenance, and both are still done by people, at the same old pace, over a volume of code that keeps growing. A better model produces more plausible code per hour, and plausible code is exactly what eats human review. I’d be glad to be wrong about this, and the test is public: when a GitClear report shows duplication and masking falling with nothing having changed in how repositories are structured, this section goes obsolete, and I’ll say so. The second criticism is more practical and almost always comes from someone who’s already tried it: writing the specification costs more than writing the code. For the tab split, this section’s architectural prompt has more lines than the use case it produced. For a two-screen feature, the time it takes to write the statement, paste the contracts, and list the bans outruns the time it takes to just write the thing. The criticism is right, and that’s exactly why it doesn’t get answered with “but it looks nicer.” It gets answered with the fourth measure instead. Specification and architecture don’t pay off the first time; they pay off on the same feature’s second change, which is when someone needs to

figure out where the rule lives. On the first draft you pay for the prompt and get the same result you’d have gotten writing it straight. On the second, you get a new 61-line file instead of 248 lines stacked into a file that already did something else, and you get the contracts governing the new generation without anyone repeating a thing. My position, spelled out in full: for code that’s getting thrown away next week, don’t write any specification, and skip this entire chapter with my blessing. For code Rosie is going to run her register on for the next ten years, the second change always shows up. Pitfalls Trusting the checklist and giving up on reading the code. The checklist narrows the search; it doesn’t replace reading what the diff does. None of the six items asks whether the cents split is right, whether the remainder was distributed, or whether the total adds up. A slice can pass all six items and still charge table four the wrong amount. Use the checklist to clear the known error classes in two minutes, and spend what’s left reading the business rule, the one part only you know how to check. Pasting in too much context and blowing the window. Once chapter 11’s context window is blown, something gets dropped, and what gets dropped first is usually the beginning, which is exactly where your contracts were sitting. The temptation is to paste the whole repository so the generator can “understand the system.” The result is a prompt where the important ban ends up buried under thirty irrelevant files. The slice is the cut that makes the context fit: the feature’s folder, the contracts it uses, and nothing more. If your prompt doesn’t fit in a slice, the problem probably isn’t the window’s size.

The architectural prompt that turns into hand-written code in prose. There’s a point where detailing the request stops constraining and starts dictating: when the prompt says which loop to use, how to name the index variable, and in what order to run the checks, you wrote the program in English and asked for a translation. At that point the earlier criticism about cost is dead right, and by a wide margin. The prompt names the contract and the boundary; the implementation is what you’re delegating. If the generator solves the problem in a way you wouldn’t have chosen, but it respects the contracts and passes the checklist, its way is fine. Q&A I ran the prompt here and the model handed back something else. Did I get it wrong? No. A model’s output isn’t deterministic, and mine wouldn’t repeat identically if I ran it again today. The success criterion isn’t matching this chapter’s listings; it’s whether what came back passes the previous section’s checklist: files only inside the slice, business refusal as a value, a use case with no repository, an orchestrator with no arithmetic, all three error cases with a path. If it passes, it’s good, even under different names. I don’t use AI to write code. Does this chapter apply to me? It does, and outside the experiment sections it barely mentions AI at all. The checklist asks about the code, not about where it came from: a pull request from a human teammate at six on a Friday evening gets reviewed by the same six questions, in the same order, with the same outcome. The architectural prompt becomes the issue’s text, which is also a statement written before the implementation.

If the complete statement already produced good code, is architecture optional? For a small feature’s first generation, this chapter’s numbers say yes, and I’m not going to pretend otherwise. The math changes on the second change, and changes again on the same system’s fifth feature, once “where does this live” stops having an obvious answer. Architecture is what keeps the answer obvious after the system grows. Does the prompt need to paste the contracts every time? On the slice’s first generation, yes. Once the slice exists in the repository, the contracts are in the code the generator reads before it writes, and this chapter’s second change showed they keep holding without being repeated. Paste them again when the slice is new or when the contract has changed. Quick tip Keep your architectural prompt in a versioned file next to the feature, not in the chat history. It’s the one artifact that goes stale in silence when the contracts change, and a file in the repository shows up in the diff when someone touches the type it pastes. Bonus: whoever joins the team reads the prompt and understands the slice faster than reading the code. Quick reference Symptom in the generated diff FOCUS piece that blocks it upfront Business refusal turning into an exception Sealed Result, refusal is a variant (chapter 8)

Generic catch that logs and moves on Single translation at the boundary (chapter 15) Rule recreated, the generator never saw it Vertical slice: the rule fits in the context Business arithmetic in the handler Pure use case, called by the orchestrator (chapter 14) Use case with a repository in the signature Data goes in, Result comes out (chapter 14) New file outside the slice Banned in the prompt, Card 6 checks it (chapter 11) Inflated scope, feature nobody asked for Statement first, with a definition of done Exercises

  1. Write the architectural prompt for the coffee shop’s inventory feature (deduct items when the tab closes, warn when an item hits its minimum). Use this section’s five parts: business goal with the refusal cases named, folder structure with one sentence per folder, the contracts pasted as code, the bans, and a definition of done with a checkable numeric result. Then count the refusal cases you named: if there are fewer than two, you probably still don’t know what the feature does when something goes wrong.
  2. Run the tab split on the model you use, with the first section’s one-sentence request, and apply the checklist to the output. How many of the six items flagged something? Can you predict, before rerunning it with the complete statement,

which items will stop flagging just because of the statement, and which only stop once you paste the contracts? Tip 21 Narrow the search space before you ask for the search. The statement cuts down what can be built; the architecture cuts down where each decision can live. Next chapter: no more isolated slices, and no more one-feature examples. You’ll build Rosie’s Coffee Shop’s entire app, one specification at a time, with everything the previous twenty-one chapters left on the table.

Powered by TurnKey Linux.