DRY Isn’t About Code In this chapter, you’ll: decide, faced with two duplicated snippets, whether they carry the same knowledge or just the same text, and unify or keep them apart with a business justification; apply the rule of three when the honest answer is “I don’t know”; explain why undoing the wrong abstraction costs more than deleting a duplication. Two identical functions sit side by side, staring back at you. Every trained instinct you own says delete one and extract the other into a shared spot, because duplication is a sin and you learned that before you learned to code properly. This chapter tells the story of the day that instinct cost Rosie money, and hands you the question that separates cleanup from a trap: do these two snippets change for the same reason? Chapter 4 ended on a taunt: when the same code shows up twice, is deleting it always the answer? That chapter’s scissors decided what not to build. This chapter’s decide what not to unify, and both cut along the same edge: real need, never reflex. There’s even a name for the family resemblance: AHA (Avoid Hasty Abstractions), coined by Kent C. Dodds, is YAGNI applied to abstractions. Build only once the need actually arrives, never because you predict it’s coming. That holds for an entire pricing engine, and it holds for a five-line extracted function.
Extraction by reflex Rosie runs two promotions in the app. The loyalty program gives 10% off to whoever stamps their tenth purchase; it exists to reward the customer who comes back every week. The daily combo gives 10% off the cappuccino-and-cheese-bread pair; it exists to move stock that’s at risk of sitting unsold in the display case. Two different business decisions, made on different days, for different reasons. In the code, they were born like this: Dart // Loyalty program: rewards customers who come back every week. int loyaltyDiscount(int totalInCents) { return (totalInCents * 10) ~/ 100; } // Daily combo: moves stock that's at risk of sitting unsold. int dailyComboDiscount(int totalInCents) { return (totalInCents * 10) ~/ 100; }
The bodies are identical, character for character. A coworker opens the file for a different task, spots the repetition, and feels the itch you know well. Two identical snippets, one obvious refactor, thirty seconds of work. He extracts the shared function, deletes both copies, and pushes the commit with the message “remove duplication.” Nobody objects in review; the diff even shrank the file. Dart // The function that unified the two “identical” rules. // Three months later loyalty jumped to 15%, and the fix landed here. int calculateDiscount(int totalInCents) { return (totalInCents * 15) ~/ 100; } int loyaltyDiscount(int totalInCents) => calculateDiscount(totalInCents); int dailyComboDiscount(int totalInCents) => calculateDiscount(totalInCents);
TypeScript function calculateDiscount(totalInCents: number): number { return Math.floor((totalInCents * 15) / 100); } function loyaltyDiscount(totalInCents: number): number { return calculateDiscount(totalInCents); } function dailyComboDiscount(totalInCents: number): number { return calculateDiscount(totalInCents); } What actually changes is the math. Dart has ~/ , an operator that divides and truncates in one step, so the result already comes out whole. TypeScript only knows floating-point division, and / returns a fraction the moment the total isn’t round: Math.floor sits there to cut that fraction off before half a cent sneaks onto Rosie’s tab. Go is the cheap counterpoint, where neither gesture
is needed, because / between two int truncates by the language’s own definition. Same business rule, three spellings. That’s why the listings from here on run in Dart alone. Three months later Rosie decides to boost loyalty: 15% starting Monday. The task lands on someone who’s never seen this file. That person looks for the rule, finds calculateDiscount , swaps 10 for 15, tests the loyalty flow, it works. On Monday the daily combo wakes up giving 15% too. On a $20 tab, the combo’s discount used to be $2 and becomes $3; the difference comes straight out of Rosie’s margin on every cappuccino-and-cheese-bread pair sold, and nobody notices until the month closes out. The loyalty test passed. There was no test saying the combo should stay at 10%, because in everyone’s head it was “the” discount. Try it: run the bug’s two stages at https://focus.kodel.com.br/en/dart/05-01 or https://focus.kodel.com.br/en/ts/05-01. Stage 1 runs the unified code with loyalty at 15% and shows the combo discount jumping right along with it; stage 2 undoes the extraction and prints the correct values. Notice the size of the fix before you keep reading. Where was the bug born? The reflex answer is “the combo was missing a test,” and it’s true and insufficient. The bug was born in the extraction. To see that precisely, you need the original definition of DRY, the one almost nobody quotes in full. What Hunt and Thomas actually wrote DRY (Don’t Repeat Yourself) appeared in The Pragmatic Programmer (Andy Hunt and Dave Thomas, 1999), and the original wording doesn’t mention code at all:
“Every piece of knowledge must have a single, unambiguous, authoritative representation within a system.” The word carrying the whole sentence is knowledge. Knowledge, here, is a decision about the business or about the system: loyalty pays 10%, the card reader fee is 3.49%, a canceled order never reaches the kitchen. Text is the shape that decision takes in a file. DRY forbids duplicating knowledge. About text, it says nothing at all. In the twentieth-anniversary edition (2019), Hunt and Thomas spent a whole section undoing the misunderstanding this chapter is attacking. Two clarifications matter here. First: DRY is wider than code and covers database schemas, documentation, and build scripts; if the same decision lives in the schema and in a validation, it’s duplicated, even without a single repeated line of code. Second, and this is the sentence that takes the coworker’s commit apart: two pieces of code that are textually identical but represent different decisions do NOT violate DRY. Identical code with distinct meanings even has a name: accidental duplication, the coincidence of two independent decisions producing, for now, the same text. And notice accidental: it isn’t something bad, catastrophic. Accidental just means something that happens by chance, without intent. Now the bug’s rediagnosis is short. loyaltyDiscount and dailyComboDiscount were accidental duplication: two pieces of knowledge, the reward for repeat visits and the push on inventory, that happened to be worth 10% in the same quarter. The text was one; the reasons to change it were two. Whoever extracted calculateDiscount didn’t remove a duplication, because there never was one. They removed the boundary between two of
Rosie’s decisions, and from then on any change to one would drag the other along. The DRY violation was the extraction, not the duplication. The commit message had it backwards. Try it: before you turn the page, judge another case from the same app. The number 0.0349, the fee the card processor charges per card sale, shows up in three files: recording the sale, closing out the day, and simulating a price. Same knowledge, or coincidence? Unify, or keep separate? Decide now; the answer comes in the next section. The inverse case: the card reader fee If you answered “unify,” you got it right, and the reason matters more than the verdict. Look at the current state: Dart // payment.dart: deducts the fee when recording a card sale int netSaleAmount(int amountInCents) => amountInCents - (amountInCents * 0.0349).round(); // closing.dart: projects what the card reader pays out at day's end int dailyPayout(int cardTotalInCents) => cardTotalInCents - (cardTotalInCents * 0.0349).round();
// pricing.dart: shows how much of an item's margin the fee eats int feeOnPrice(int priceInCents) => (priceInCents * 0.0349).round(); The three snippets aren’t even alike; each function carries a different name, parameter, and slightly different math. And yet this is exactly the duplication DRY forbids. The three 0.0349s are the same knowledge: the fee written into Rosie’s contract with the card processor. There’s a single document in the real world that defines that number. When the processor adjusts it to 3.79%, all three spots need to change together, at the same moment, for the same reason; whoever forgets the third one creates a day’s closing that never matches the statement. The test was never “is the text the same?” The test is “is there a single decision behind it?” Here there is, so the representation has to be single too: Dart // fees.dart: the one line that changes when the processor adjusts its rate const cardReaderFee = 0.0349; int feeOn(int amountInCents) => (amountInCents * cardReaderFee).round();
// payment.dart int netSaleAmount(int amountInCents) => amountInCents - feeOn(amountInCents); // closing.dart int dailyPayout(int cardTotalInCents) => cardTotalInCents - feeOn(cardTotalInCents); // pricing.dart int feeOnPrice(int priceInCents) => feeOn(priceInCents); Put the two verdicts side by side, because together they’re the chapter’s lesson. The discounts had identical text and different knowledge: keep them separate. The fee had different text and a single piece of knowledge: unify right away, without waiting for a third occurrence. The eye that compares characters gets both cases wrong. The question that gets both right doesn’t look at the code; it looks at Rosie’s business. Timing tools
Knowing the right question doesn’t erase the hard case: what about when you can’t answer it? A small toolbox has built up around exactly that impasse, and each piece has an owner and an address. The first is from Sandi Metz, in “The Wrong Abstraction” (2016): “prefer duplication over the wrong abstraction.” The wrong abstraction is the function or class that unifies snippets that never carried the same knowledge, exactly what calculateDiscount became. Metz’s argument is about interest. The wrong abstraction doesn’t sit still waiting for you to undo it: the next almost-matching case shows up, someone adds a parameter to accommodate it, then a conditional, and every patch raises the price of taking the whole thing apart. Duplication just sits there instead, repeated and harmless, until someone understands it. The second you already met in the opening: Kent C. Dodds’s AHA (kentcdodds.com, “AHA Programming,” 2019). Dodds doesn’t ask you to duplicate forever; he asks you to wait for the abstraction to reveal itself, instead of forcing it at the first resemblance. It’s the same muscle from chapter 4: you don’t build capacity on a forecast, and you don’t abstract on one either. Abstracting at the first coincidence is betting that two snippets will evolve together before you have any evidence of it. The third tool answers “wait until when?” The rule of three, which Martin Fowler records in Refactoring (1999), says: the first time you write it, the second time you duplicate with your eyes open, the third time you extract. The third occurrence is the missing evidence; with three uses in hand, the abstraction’s real shape shows up, and you build it knowing exactly what it needs to cover. Two occurrences are still too small a sample to guess the right boundary.
The fourth is less a rule and more a mnemonic reminder. Conlin Durbin coined WET (Write Everything Twice) in “What is WET code?” (dev.to, 2018): tolerate the second copy and only abstract on the third. It’s the rule of three dressed up as a pun on DRY, and it works as a short answer for the coworker who flags any second occurrence as debt. None of these four pieces contradicts Hunt and Thomas. Metz, Dodds, Fowler, and Durbin regulate the timing of abstracting similar-looking text; DRY demands a single representation for a single piece of knowledge. The whole flow fits in one diagram: Run the chapter’s two cases through it. The discounts enter the question node and exit through “no”: loyalty changes when Rosie wants to reward more, the combo changes when inventory gets tight, independent reasons. The fee enters and exits through
“yes”: one contract, one number, three points of use. You’ll exercise the “not sure” branch in the exercises, with a pair that has no obvious answer on purpose. The same knowledge outside the code The 1999 sentence talks about a system, not a file, and that’s why it aged well. Rosie’s decision about the loyalty discount has more places to settle into today than it had back then, and three of them aren’t code. Semantic duplication is the same decision written twice in different words. The use case requires the tenth purchase to unlock the 10%; the customer screen works out how many stamps are missing and prints “two to go”, with the arithmetic redone right there. No text search finds that pair, because there’s no repeated text: what repeats is the rule. When Rosie starts requiring twelve purchases, the use case changes and the screen keeps counting to ten. The criterion is the one it always was: both change for the same reason, so the decision needs a single representation, and the screen asks instead of recomputing. Prompt duplication is the rule that comes to live in the request you write to ask for code as well. You paste into the request that the loyalty discount is 10% from the tenth purchase on, you get the function back and, from then on, the rule lives in two places: in the file and in the text of the request, which usually sits saved in some project instructions file. When the discount goes up to 15%, the file changes and the request doesn’t, and the next answer comes back with the old version, now carrying the authority of something fresh off the machine. The way out is the card reader fee’s: the request cites the file instead of repeating the rule, and what it carries is the address, not the number.
Context duplication is the copy someone makes to spare the reader from leaving the file. The comment that re-explains the discount rule at the top of the repository, the README that reproduces the use case’s signature, the snippet pasted into the team’s documentation. Each copy is a photograph of the day it was taken, and none of them breaks when the original changes: they just go quietly wrong, which is the worst way to be wrong. All three go through the same question, and that’s what keeps the extension from turning into another crusade against repeated text. The example code that shows up three times in this chapter isn’t duplication: it exists to teach, and it ages along with the page. The discount rule copied into the request is: it exists to decide, and it decides wrong the moment it ages. The critique: DRY as a coupling factory Every popular principle collects criticism, and the most serious one against DRY is this: DRY breeds premature abstraction that couples what evolves separately. Whole teams, trained to hunt duplication, produce layers of generic helpers that nobody can change without breaking three screens. The critique describes real damage; you saw a miniature of it in Rosie’s combo. Except its target is the misunderstanding, not the principle. Whoever extracted calculateDiscount was violating DRY, not applying it: they unified two pieces of knowledge into one representation, the literal opposite of the 1999 sentence. DRY correctly read and AHA don’t compete for territory. DRY governs knowledge: a single representation for a single decision. AHA governs timing: without evidence it’s the same decision, wait. The dispute between the two only exists once DRY turns into “delete all repeated text,” and that version isn’t on a single page Hunt and Thomas wrote.
Here’s my position, so you can calibrate your own: between duplicating and risking the wrong abstraction, I duplicate and sleep fine. Undoing a duplication that turned out to be a single piece of knowledge is search and replace, ten minutes with the editor and the tests. Undoing the wrong abstraction is surgery: every caller depends on it in its own way, the accommodation parameters already created combinations nobody tested, and removing it means understanding every use case at once. Both mistakes are possible; the prices aren’t in the same league. Pitfalls The shared/ utility born on the second occurrence. You write a function, notice another feature has something similar, and create shared/utils.dart for both right away. It’s reflex extraction with a fancier address: with two occurrences you rarely know whether there’s one piece of knowledge or two, and the shared directory invites the rest of the team to hang parameters off it. The way out is the diagram from the previous section: ask the knowledge question; when unsure, rule of three, and the utility only gets born on the third occurrence, shaped by whatever the three uses actually need. “So I’m never extracting anything again.” That’s the mirror- image conclusion, and it costs just as much as the original. When the knowledge is genuinely one thing, like the card reader fee, unifying isn’t optional and doesn’t need a third occurrence: leaving it scattered is betting that three spots will change together by hand, forever, without a single slip. The rule of three is the way out of “I don’t know,” never a veto over “yes.” Unifying the text to “stay ready.” The coworker argues that if the two rules ever truly converge, the code will already be prepared. You know this argument from chapter 4: it’s presumed
capability, now wearing a function’s shape. If the rules do converge, that day’s extraction will be cheap and informed. Today’s is a guess with the power to spread bugs. Q&A How do I find out if two snippets are the same knowledge? Look for the decision’s source outside the code. The fee has a signed contract; the discounts have two business motivations with different owners. If the question “who’s in charge of this number?” points to two places, it’s two pieces of knowledge, whether the text matches or not. Wouldn’t the extraction be defensible if calculateDiscount had a good test? A test on the combo would have turned the silent bug into a visible failure, and that alone would be worth a lot. But the coupling would still be there: every change to loyalty would still run into the combo, now with a red test in the way. A good test exposes the wrong abstraction; only undoing it fixes it. Does DRY apply outside code? It does, and that’s the part the 2019 edition makes a point of underlining: database schema, documentation, and build also carry knowledge. If the menu says the cappuccino costs $11.95 and a constant in the app says 1195 cents, those are two representations of the same decision, and one of them is going to rot. Why 1195 cents instead of 11.95? Because money in floating point is a risk, as chapter 4 already flagged: float represents 0.10 as a binary approximation, and cents vanish in long sums. An integer is always worth the same thing, no matter how the language represents it (float, double, decimal) or how it travels (JSON, ProtoBuf); JSON parsers, in particular, decode a broken number into a float, so what leaves one side as an integer arrives whole on the other. The 0.0349 fee in
the listings can stay a float because it’s a multiplier, not stored money; the result lands back in whole cents inside that same line’s round() . Quick tip Before you extract a function to kill a duplication, run git log -p on the snippets involved. If they changed in separate commits, for separate reasons, that’s strong evidence of two pieces of knowledge; if every change to one always came bundled with the other, unifying is probably overdue. Tip 5 Before you unify two identical snippets, ask whether they change for the same reason. Quick reference Situation Fix Same text, same knowledge Unify now, don’t wait for a third time Same text, different knowledge Keep separate: it’s accidental duplication Not sure if it’s the same knowledge Rule of three: wait for the third time Wrong abstraction already in place Undo it and duplicate back; then reassess
Exercises
measure how many spots you’d have to edit today if it changed; that number is your risk of a closing that never matches. Next chapter: you learned to sniff out duplicated knowledge in functions and constants, but what about when the duplication lives in the shape of your classes, and “do they change for the same reason” becomes the question that decides an entire system’s design?
Powered by TurnKey Linux.