您最多选择25个主题 主题必须以字母或数字开头,可以包含连字符 (-),并且长度不得超过35个字符

18KB

FOCUS Architecture — Chapter-03: AI Writes Fast. So What?

  • Source: /library/FOCUS Architecture/source-file.pdf
  • PDF pages: 42–56
  • Pages without text: none

AI Writes Fast. So What? In this chapter, you’ll: quote from memory the three GitClear numbers that measure code degradation in the AI era: duplication +81%, error masking +47%, and refactoring dropping from ~25% to under 10% of changed lines; define vibe coding with author and year, and separate it from “using AI”; point out, in Rosie’s coupon bug, where each of FOCUS’s three guardrails would have caught the defect. You’re going to ask an AI for a feature and get it back, done, in twelve minutes. You’ll test it on screen, watch it work, and ship before lunch. Three weeks later Rosie’s register will charge the wrong amount, and no log will explain why. This chapter shows where the defect hid, how much of it is already measured at scale, and which three structures would have stopped it at the door. Chapter 2 ended with a promise: the fastest intern in the world was about to find the menu screen. She did. That tangle of business rule, formatting, and persistence crammed into one file cost a human team six months of shortcuts; I’ll guess, with the bluntness of someone exaggerating on purpose, that an AI delivers the same tangle 40 times faster. The 40 is my hyperbole,

not a measurement. The question it carries is serious: what happens to the cost of change when tangled code stops taking months to exist and starts existing in minutes? The twelve-minute coupon Wednesday, late afternoon. Rosie saw the competitor’s coffee shop on her phone and showed up with the request ready: “I want a discount coupon button, the kind where you type WELCOME10 and get 10% off.” You already have four tasks in the queue. So you paste the request into an AI, along with the order- flow file, and twelve minutes later there’s a new handler, applyCoupon, with code validation, total calculation, and even a friendly message for an expired coupon. You type WELCOME10, the price drops 10%, and Rosie applauds from the counter. Deploy done, queue resumed. Three weeks later, the customer at table four complains. She has seven orders in the loyalty program and used the new campaign’s coupon; the app charged only the coupon discount, without adding the loyalty discount. The cashier checks, and the complaint holds up. You open the day’s log: nothing. No exception, no warning, no line out of place. The app recorded no defect because, as far as it was concerned, no defect happened. The archaeology takes half an hour and turns up two discoveries. First: the handler that came out of the model never calls the loyalty calculation that already existed in the order flow; it recomputes the total from scratch and carries its own copy of the rule, rewritten from whatever the model saw in the file. Two handlers now hold the same business rule, each on its own, and the new copy was born out of date: it uses the old threshold of ten orders, not the five Rosie adopted months earlier. With seven orders, the customer clears today’s threshold and misses the one

the copy kept. This has a name. Knowledge duplication is the same business decision written in two places that don’t know about each other; when the decision changes, someone has to remember every address, and chapter 2 already showed how that recall fails. The second discovery is worse. At the end of the new handler sat a consistency check that compared the coupon total against the order-flow total and threw an exception on any mismatch. The AI wrapped that check in an empty try/catch, commented “avoids blocking checkout.” The alarm existed. Someone switched it off during installation. Error masking is exactly this: code that catches or swallows a failure without handling it, and hands the user the appearance of success in place of the problem. The coupon bug wasn’t silent by accident; it was silenced by design. The GitClear yardstick One coffee shop case doesn’t prove a trend. Numbers do, and someone counted them: GitClear, a company that analyzes code quality, examined hundreds of millions of real changes across repositories for its 2024, 2025, and 2026 reports. Hold on to three of those numbers; they’re the spine of this entire book. Code duplication rose 81% relative to the pre-AI era (GitClear, 2024- 2026). Error masking rose 47% (GitClear, 2026): the generator favors a silent catch, safe navigation, and stubs (facade implementations that return some value without doing the work) that hide defects over ones that handle them. And refactoring dropped from about 25% of changed lines in 2021 to under 10% in 2024 (GitClear, 2024): new code piles on top of new code, and almost nobody tidies up. The rest of the numbers in the same set of reports back up those three. Duplicated blocks of five or more lines grew roughly 8x in 2024 (GitClear, 2024). The same report notes that 2024 was the

first year copy-paste outpaced moved code: more copying happened than reuse (GitClear, 2024). Cross-file calls, reuse between files, fell 35%, and legacy code maintenance dropped 74% (GitClear, 2024-2026). Outside GitClear the direction repeats: the arXiv 2409.19182 study (2024) and the 2024 DORA (DevOps Research and Assessment) research found delivery speed climbing while stability and maintainability fall when nothing constrains the generator. Notice the framing. The measured problem is never “used AI”; it’s what gets generated when nothing limits the shape of the output. This way of generating has its own name. Vibe coding is the term Andrej Karpathy coined in February 2025 for the practice of accepting code from an LLM (Large Language Model) without reading it, guided only by the surface result: it ran, it looked fine on screen, move on. Notice the gap between that term and “using AI.” Someone who uses AI with review and a boundary stays in command of the code; someone who does vibe coding delegated the reading too. The coupon handler was classic vibe coding, and I was the first to do exactly the same: the code looked right, the screen worked, and twelve minutes is too tempting to resist. Why the generator fails this way An LLM doesn’t optimize for your system to last; it optimizes for the next answer to look correct to whoever reads it. Those are different goals. Inside a single file, “looking correct” and “being correct in the system” nearly line up, which is why the coupon handler worked so well in the demo. Across the whole system the two goals split apart: the loyalty rule that already existed sat outside the context the model could see, so recreating it inside the new handler was the path of least resistance. The same logic applies to failure. Handling an exception means deciding what the business wants in every bad case; swallowing the exception

makes today’s demo pass. Without a boundary that forces the handling, the empty catch is the exit the generator has learned to prefer, and GitClear’s +47% (2026) shows that preference at industrial scale. The contrast fits in one listing. Rosie’s inventory lookup, first as the AI delivered it, then as it looks once the failure becomes a return value; chapter 8 builds that Result piece by piece, so don’t worry about the sealed syntax yet: Dart // what the AI delivered: looks like it works Future availableUnits(String item) async { var units = 0; try { units = await inventory.check(item); } catch (e) {} return units; }

// the same lookup with Result: failure becomes a value the caller handles sealed class InventoryQuery {} class Available extends InventoryQuery { Available(this.units); final int units; } class InventoryDown extends InventoryQuery {} Future checkInventory(String item) async { try { return Available(await inventory.check(item)); } on Exception { return InventoryDown();

} } TypeScript // what the AI delivered: looks like it works async function availableUnits(item: string): Promise { let units = 0; try { units = await inventory.check(item); } catch (e) {} return units; } // the same lookup with Result: failure becomes a value the caller handles type InventoryQuery =

| { kind: “available”; units: number } | { kind: “inventoryDown” }; async function checkInventory(item: string): Promise { try { return { kind: “available”, units: await inventory.check(item) }; } catch (e) { return { kind: “inventoryDown” }; } } Read the two halves as two answers to the same question: “what happens when the inventory service goes down?” The top half answers with a polite lie. The offline inventory becomes 0, the menu shows “out of stock” for an item that exists, and no log reports anything; it’s the twin of the coupon’s empty catch. The bottom half changes the return type: whoever calls checkInventory gets Available or InventoryDown back and has to decide what the screen does in each case. The lie is no longer an option on the table.

Who enforces that decision depends on the language, and the difference matters. In Dart, a switch over a sealed class is exhaustive by construction: miss a case and the program doesn’t compile. In TypeScript, exhaustiveness is optional; the compiler only complains if you ask it to, by assigning the unhandled case to a variable of type never in the default branch. Without that explicit request, an incomplete switch slips right through. Try it: open https://focus.kodel.com.br/en/dart/03-01 and run the file a few times; the example inventory fails at random. The naive version prints 0 as if the item had run out, and the Result version prints the inventory-down warning. Then delete the InventoryDown case from the switch and watch the compiler refuse the program. The same comparison in TypeScript is at https://focus.kodel.com.br/en/ts/03-01. Put the chapter’s pieces together and the coupon case stops looking like an isolated accident. It’s one full turn of a cycle that feeds itself:

Every generation accepted without reading adds a copy; the next rule change forgets one of them; the catch masks the mismatch; the silence in the logs turns into confidence that everything is fine; that confidence authorizes generating even more. The cycle doesn’t stop on its own. It stops when some structure breaks one of the links, and that’s what the rest of this book is about. Three guardrails against the same bug Call the central idea a narrow search space: give the generator (and the reviewer) contracts so tight that only one correct way exists to complete the code. A model that can return anything will, sooner or later, return the wrong thing wearing the face of the right one; a model squeezed by types, tests, and a boundary

errs less and, when it does err, errs loud. FOCUS narrows that space with three guardrails, the same barriers that keep a car from leaving the road without taking the wheel out of your hands. Each one would have caught the coupon bug at a different point. The first is the compiler armed with exhaustive types. If the coupon’s consistency check returned a Result like the one in the listing, instead of throwing an exception, the empty try/catch wouldn’t even be possible: the handler would be forced to declare, in visible code, what to do with a total mismatch. “Ignore the failure” would still exist as a decision, but a written, reviewable one, never an invisible omission. Chapter 8 builds this guardrail. The second is the pure-function test. If the loyalty rule lived in a single function with no screen, no database, and no network nearby, it would have one address and a full-name test suite; the outdated copy inside the coupon handler would have nowhere to come from, and a threshold change from ten orders to five would break a test in the same second. Chapters 14 and 17 build this guardrail. The third is the slice. With the coupon flow isolated in its own slice, with an explicit contract for talking to the rest of the system, the possible damage from a bad handler stays confined to the size of the slice: the blast radius of a twelve-minute generation becomes a directory, not the whole system. Chapter 11 builds this guardrail. Notice that none of the three demands a better AI or a superhuman reviewer; all three change the ground where any generator, human or not, is capable of erring. So is AI the problem?

No, and I wouldn’t have written this book if I thought so. I use AI every day (I used it to write code for this book too), and the speed gain from shipping that coupon is real; the same evidence that shows the degradation shows the gain. This book covers the code structure that makes generation fit inside a boundary. Steering that power with specifications is the subject of Spec Driven Development (2026, https://books.kodel.com.br/en/books/sdd), and you do not need to have read it to follow from here. The problem was never the fastest intern in the world. The problem is handing her a system where the business rule lives on four screens, no type forces failure handling, and any file can reach any other; on that ground, speed only amplifies the tangle from chapter 2. Structure first, generation second, and the pair scales; chapter 21 closes this argument with all of FOCUS on the table. Pitfalls “The tests pass, so it’s correct.” Watch out for who wrote the tests. When the AI generates the code and the tests in the same breath, the test tends to lock in the generated behavior, not the business rule: the coupon handler’s test asserted that a total mismatch returns the coupon total without complaint: the error masking had a test guaranteeing the lie stayed alive. A test is worth what it demands, not whether it passes; the rule for who demands what arrives in chapters 14 and 17. “I review every diff myself, this won’t happen to me.” Rereading chapter 2 helps here: the human reviewer approves the fourth copy at 6pm on a Friday. AI multiplies diff volume by a factor no review discipline keeps up with; trusting your system’s safety to a tired human’s infinite attention is betting against the odds. Structure that makes the error impossible to compile doesn’t get tired. “To err is human” is real: sooner or later, someone errs, always.

“So I’ll ban AI on the team.” A ban throws away the real speed gain and doesn’t remove the cause: the ground without a boundary is still there, and hurried humans produce the same tangle in slow motion, as three reasonable pull requests already proved in chapter 2. The right target is the ground, not the tool. Q&A Won’t these GitClear numbers age badly? They will, which is why each one carries its report year right next to it. What this chapter asks you to keep is the mechanism, ownerless copy plus swallowed failure, which stays explainable even after the percentages change. Don’t newer models write better code and retire this whole discussion? They write better code, faster, and that cuts both ways: they also fail faster. GitClear’s 2026 report measured error masking rising in exactly the most capable generation to date. As long as the generator’s goal is to look correct to whoever reads it, the shape of the output stays the responsibility of whoever sets the boundary: you. Is error masking an invention of the AI era? No; the empty catch has existed as long as exceptions have, written by people. The new part is scale: what used to be an occasional slip by a rushed developer became the statistical preference of a tool that writes a large share of the world’s new code, with a 47% rise measured by GitClear in 2026. Quick tip Before accepting any AI-generated diff, hunt it for masking patterns: git diff | grep -nE “catch\s*((.*))?\s*{\s*$” catches an empty-body catch on its first line, and it’s worth repeating the search for ?. and suspicious default values like ?? 0 .

Better still: turn on your language’s lint rule ( empty_catches in Dart, no-empty in ESLint) as a CI (Continuous Integration) failure, the battery that runs on every push before code merges, and the most common masking dies before the merge, at zero cost in human attention. Quick reference Situation What to do Citing the numbers +81% duplication, +47% error masking (GitClear 2024-2026) Citing the refactoring drop ~25% to under 10% of lines (GitClear, 2024) AI feature passed the demo Hunt for the recreated rule and the silenced catch Log too clean after a bug Suspect error masking at the source Team debates “AI, yes or no” Debate the boundary instead: chs. 8, 14/17, and 11 Telling vibe coding from using AI Accepting without reading (Karpathy, Feb 2025) is vibe Exercises

  1. Open https://focus.kodel.com.br/en/dart/03-01 and add a third case to the sealed class: ItemOutOfStock, for when the

service reports zero units. Watch the order of events: the compiler flags the incomplete switch before you run anything at all. That’s the feeling of working inside a narrow search space, and it’s the one you’ll build in chapter 8. 2. A coworker asked an AI for Rosie’s coupon-application screen and got back the code at https://focus.kodel.com.br/en/dart/03-02 (React version at https://focus.kodel.com.br/en/ts/03-02). The code compiles and the demo passes. Can you find the single spot of error masking hidden in it, name the information the user loses there, and propose what the function should return instead? Tip 3 Never accept an error from an AI that the compiler wouldn’t have caught. Next chapter: before you learn to structure what you build, you’ll learn the cheapest trick of the trade: deciding what not to build. With those scissors in hand, Part II begins: the foundations that turn the diagnosis of these three chapters into daily practice.

Powered by TurnKey Linux.