Du kannst nicht mehr als 25 Themen auswählen Themen müssen entweder mit einem Buchstaben oder einer Ziffer beginnen. Sie können Bindestriche („-“) enthalten und bis zu 35 Zeichen lang sein.

37KB

FOCUS Architecture — Chapter-17: Test Each Piece the Way It Asks to Be Tested

  • Source: /library/FOCUS Architecture/source-file.pdf
  • PDF pages: 498–539
  • Pages without text: none

Test Each Piece the Way It Asks to Be Tested In this chapter, you’ll: write the loyalty feature’s suite for Rosie’s Coffee Shop: chapter 14’s use case test with no double at all, and chapter 13’s orchestrator flow test with chapter 15’s repository fake; justify in writing where integration testing pays for itself and where it doesn’t, operation by operation, across the coffee shop; explain why a test that checks call order breaks on a refactor that doesn’t change behavior. Since chapter 10, FOCUS has promised that every layer is easy to test. That promise comes due today. You’re going to watch a green suite turn red without the program changing behavior, and you’re going to see why the other suite, written against that same refactor, stays green. The difference between the two isn’t a matter of style: it’s what each one chose to assert. You arrive here with four pieces built and none of them tested. Chapter 12’s dumb View fires an event and renders state. Chapter 13’s orchestrator converts an event into state. Chapter 14’s loyalty discount is a pure function. Chapter 15’s repository translates an infrastructure exception into a Result, and it brought along an in-memory fake. Chapter 16 split the two tracks, command and

query. Each of these pieces asks for a different kind of proof, and this chapter’s thesis is that you don’t choose which: the architecture already chose for you. Whoever separated rule from IO earned a cheap testing base. Whoever didn’t pays in doubles. Before the technique, the pain. The anti-solution: the suite that asserts the how The operation that crosses all four layers is paying the tab. The orchestrator receives the event, looks up the tab in the repository, tells it to mark the tab paid, and publishes the state. A developer sits down to test this and does the thing that looks like the most rigorous move in the world: puts a double in place of the repository and checks whether the orchestrator called the right methods, in the right order, the right number of times. Test double (Gerard Meszaros’s term) is the name he gave, in xUnit Test Patterns (2007), to any object that stands in for a real collaborator during a test. Meszaros cataloged five kinds; this chapter uses two, and the difference between them is this whole chapter’s axis. A mock is an interaction checker: it asserts which methods got called and in what order. Keep that definition in mind. In the “Repository: the fake you already have” section, it gets contrasted with the definition of a fake. Notice it takes just ONE double for the pain to show up. The orchestrator only receives one injected collaborator, the repository, and the whole suite leans on it. Dart class TabRepositoryMock extends Mock implements TabRepository {}

// Builds the double already taught to respond, fires the event, and // returns the double for the assertions. Future payTableFour() async { final repository = TabRepositoryMock(); final tab = Tab( 4, const [(name: “Espresso”, priceInCents: 700)], 700, true, ); when(() => repository.findTab(any())) .thenReturn(TabFound(tab)); when(() => repository.markAsPaid(any())) .thenReturn(TabSaved(tab));

final orchestrator = TabOrchestratorWithPayment( repository, refactored: refactored, ); orchestrator.add(PayTab(4)); await Future.delayed(const Duration(milliseconds: 50)); await orchestrator.close(); return repository; } test(“looks up the tab before marking it paid”, () async { // Arrange + Act final repository = await payTableFour();

// Assert: the ORDER of the calls. Not one line looks at the state // that went out the door. verifyInOrder([ () => repository.findTab(4), () => repository.markAsPaid(any()), ]); }); test(“looks up the tab twice”, () async { // Arrange + Act final repository = await payTableFour(); // Assert: the CALL COUNT. This is the line the refactor knocks down, // without a single comma of observable behavior changing. verify(() => repository.findTab(4)).called(2); });

TypeScript it(“looks up the tab before marking it paid”, () => { // Arrange + Act const repository = payTableFour(); // Assert: the ORDER of the calls. const find = vi.mocked(repository.findTab); const mark = vi.mocked(repository.markAsPaid); expect(find.mock.invocationCallOrder[0]).toBeLessThan( mark.mock.invocationCallOrder[0]!, ); }); it(“looks up the tab twice”, () => {

// Arrange + Act const repository = payTableFour(); // Assert: the COUNT. expect(vi.mocked(repository.findTab)).toHaveBeenCalledTimes(2); }); Read the two assertions and ask what they know about paying a tab. The answer is nothing. They know findTab got called before markAsPaid , and that the first one got called twice. Not one line looks at the state the orchestrator published, which is the only thing the waiter’s screen ever sees. The suite asserts the HOW, and the how is exactly the part you have the right to change. Why called(2) and not called(1) ? Because the handler is written in a silly way, on purpose: it looks up the tab once to validate, then looks it up again to pay. It’s the kind of duplication nobody reread. The mockist suite, green, is photographing exactly that defect. Try it: open https://focus.kodel.com.br/en/dart/17-01 (or your language’s route) and run it. Before you look at the output, answer this: if someone fixes the duplicated lookup, which of the two assertions falls?

Now someone comes along and fixes it. The refactor extracts a private method _pay and reuses the result of the first lookup instead of querying the repository a second time. One lookup, not two. Dart // AFTER version: _pay extracted, the result of the first lookup // reused, a single call. Same states, same payloads. Future _onPayTabAfter( PayTab event, Emitter emit, ) async { emit(Loading()); final lookup = _repository.findTab(event.table); switch (lookup) { case InfraFailure(): emit(Failed(“no connection”));

case TabFound(:final tab): await _pay(tab, emit); } } Before showing the break, prove the refactor changed nothing. This order isn’t ceremony: if the behavior had changed, the mockist suite would be right to complain, and this chapter’s whole argument would collapse. The proof is running the flow suite against both versions and comparing the output character by character. $ diff <(grep -v ‘^[0-9:.]* ' /tmp/flow-before.txt)
<(grep -v ‘^[0-9:.]* ' /tmp/flow-after.txt) $ diff /tmp/output-before.txt /tmp/output-after.txt Both diffs come out empty. Same sequence of states, same payloads, same output. The program does exactly what it did before. Now run the mockist suite against the new version: 00:00 +0: loading test/mockist_test.dart 00:00 +0: (setUpAll) 00:00 +0: looks up the tab before marking it paid 00:00 +1: looks up the tab twice

00:00 +1 -1: looks up the tab twice [E] Expected: <2> Actual: <1> Unexpected number of calls package:matcher expect package:mocktail/src/mocktail.dart 595:5 VerificationResult.called test/mockist_test.dart 73:48 main. 00:00 +1 -1: (tearDownAll) 00:00 +1 -1: Some tests failed. Failing tests: test/mockist_test.dart: looks up the tab twice And the flow suite, against that same new version: 00:00 +0: loading test/flow_test.dart 00:00 +0: paying a tab that's found publishes Loading then Ready 00:00 +1: network outage on the LOOKUP publishes Loading then Failed 00:00 +2: failure on the SAVE publishes Loading then Failed 00:00 +3: All tests passed! Red on one side, green on the other. Expected: <2> / Actual: <1> is the test saying the program stopped doing something it never promised to do, against code whose observable behavior hasn’t changed by a comma since the previous version. The developer who did the refactor now has two bad options: undo the improvement, or open the test and adjust the number. Everyone picks the second one, at three in the afternoon on a Friday. That’s when the test turns into a stamp: it starts recording what the code does, and a test that records what the code does catches no defect at all.

Use case: a pure test, no double at all Move up a layer and look at the loyalty discount from chapter 14. The table from chapter 10 says the Use Case forbids “IO, framework, and domain exceptions.” Read that ban as a testing promise: if the layer can’t touch IO or a framework, there’s no dependency to fake. Zero doubles, because there’s nothing to double. This is where the investment in purity from chapters 7 and 14 gets paid back, with interest. applyLoyaltyDiscount takes a tab and a number of points, and returns one of the three variants of DiscountResult , without querying a database, without asking any framework’s permission, and without depending on anything you’d need to set up first. You build the input by hand, call the function, and compare the output. That’s it. A function is the easiest thing there is to test. There are three cases, one per variant. The third is the most interesting, because it sends spare points alongside an out-of- stock item: it proves the ORDER of the rule, not just the result. Dart test(“applies 10% when there are enough points”, () { // Arrange: the tab is a value. Nobody needs a database to build one. const tab = Tab(4, [ Item(“Espresso”, 700), Item(“Cappuccino”, 1195),

], 4000); // Act: the function under test is pure. Calling it is the whole test. final result = applyLoyaltyDiscount(tab, 120); // Assert: the variant that came out, and the value it carries. check(result).isA().has( (r) => r.tab.totalInCents, “totalInCents”, ).equals(3600); }); test(“rejects for insufficient points and states how many are needed”, () { // Arrange const tab = Tab(4, [Item(“Espresso”, 700)], 700);

// Act final result = applyLoyaltyDiscount(tab, 40); // Assert check(result) .isA() .has((r) => r.pointsNeeded, “pointsNeeded”) .equals(100); }); test(“rejects for an out-of-stock item, before checking points”, () { // Arrange: spare points, but one item out of stock. The rule's order // is what decides the outcome. const tab = Tab(4, [ Item(“Espresso”, 700), Item(“cheese bread”, 600, outOfStock: true),

], 1300); // Act final result = applyLoyaltyDiscount(tab, 500); // Assert check(result) .isA() .has((r) => r.name, “name”) .equals(“cheese bread”); }); Line by line. The Arrange block builds the tab with const : no database, no factory, no builder. The Act block is one line, because the function needs nothing beyond its arguments. The Assert block uses package:checks , which chains isA() (is it?) to assert the variant and has(...) (does it have?) to drill down into the field. The // Arrange , // Act , and // Assert comments are subgoal labels: each one names the goal of the block that follows, and they exist because the listing runs past fifteen lines.

Two things deserve an explanation. The 4000 isn’t the sum of the two items, which comes to 1895, and the difference is on purpose: applyLoyaltyDiscount doesn’t add up a single item; it takes 10% off the totalInCents that already arrived calculated, and if it ever starts recalculating from the items, this is the test that breaks. The second is the Tab . Here it carries items , because the out-of-stock rule scans the list. In the flow tests, which never look at a single item, the published snippet carries a Tab reduced to table, total, and canPay , which is why this chapter’s listings build tabs of different shapes. Now the other nine. What changes from one language to the next is how each one expresses “this is the variant that came out,” and how much the compiler helps. TypeScript it(“applies 10% when there are enough points”, () => { // Arrange const tab: Tab = { table: 4, items: [item(“Espresso”, 700), item(“Cappuccino”, 1195)], totalInCents: 4000, };

// Act const result = applyLoyaltyDiscount(tab, 120); // Assert expect(result.kind).toBe(“discountApplied”); if (result.kind === “discountApplied”) { expect(result.tab.totalInCents).toBe(3600); } }); The if after expect is uncomfortable, and it’s telling you something true. The discriminated union only narrows the type inside a block that tests the discriminant, and expect doesn’t narrow anything as far as the compiler is concerned. Where Dart writes isA().has(...) , TypeScript has to assert twice: once for the runner, once for the type checker. Kotlin · Swift From here on, the listings call confirm , and there’s no point searching for that function in any library: it’s a three-line checker defined right inside the snippet, which prints “ok” or “FAILED” next to the case name and makes the program exit

with an error if any case failed. It exists because these languages’ playgrounds run a main , not a test runner, and three lines are enough to do here what a runner would. confirm( “rejects for insufficient points and states how many are needed”, declined == DiscountResult.NotEligibleForDiscount(40, 100), ) confirm( “rejects for an out-of-stock item before checking points”, outOfStock == DiscountResult.ItemOutOfStock(“cheese bread”), ) Kotlin and Swift ship stacked because the assertion is the same sentence in both: compare the whole result against the expected variant, by value equality. data class in Kotlin and enum with associated values in Swift give structural equality for free, so the test doesn’t need to drill down field by field. In both, a when or switch that forgot a variant wouldn’t even compile. Java · C#

confirm(“applies 10% with enough points”, applied instanceof DiscountApplied d && d.tab().totalInCents() == 3600); confirm(“rejects for insufficient points and states how many are needed”, declined.equals(new NotEligibleForDiscount(40, 100))); Java and C# also ship together, for the same reason: record in both languages generates structural equality, and the pattern matching of instanceof (Java) and is (C#) ties the type test and the field extraction into a single expression. The difference between the two is which compiler complains about an incomplete switch , and it’s Java: in C#, exhaustiveness over a sealed hierarchy earns a warning, not an error. Rust #[test] fn applies_ten_percent_with_enough_points() { // Arrange let tab = Tab {

table: 4, items: vec![Item::new(“Espresso”, 700), Item::new(“Cappuccino”, 1195) ], total_in_cents: 4000, }; // Act let result = apply_loyalty_discount(&tab, 120); // Assert match result { DiscountResult::DiscountApplied(t) => { assert_eq!(t.total_in_cents, 3600); } other => panic!(“expected DiscountApplied, got {other:?}"), }

} Rust leaves the Kotlin-and-Swift group for a practical reason: it has a test runner built in. cargo test finds any function tagged with #[test] in the same file as the code, so the listing shows the language’s own testing idiom instead of the hand-rolled checker. The match in the Assert block is exhaustive by the compiler’s own demand, and the other arm exists to give a readable error message, not to paper over a typing gap. PHP confirm( ‘applies 10% with enough points’, $applied instanceof DiscountApplied && $applied->tab->totalInCents === 3600, ); PHP has no sealed union and no automatic structural equality, so the test combines instanceof with a field-by-field comparison and uses === to avoid the type coercion of == . It’s more verbose than Java for the same reason Java is more verbose than Kotlin: every feature the language lacks reappears as a line of test. Go // Act

result := ApplyLoyaltyDiscount(tab, 120) // Assert applied, ok := result.(DiscountApplied) if !ok { t.Fatalf(“expected DiscountApplied, got %T”, result) } if got, want := applied.Tab.TotalInCents, 3600; got != want { t.Errorf(“total = %d, want %d”, got, want) } Go is the counterpoint in form. There’s no assertion library here, and not by oversight: if got != want { t.Errorf } is the language’s idiom, and the test gets longer in exchange for having nothing to learn beyond if . Notice the t.Fatalf in the first block against the t.Errorf in the second. The first one aborts, because continuing without the tab makes no sense; the second one records and moves on. Python

Act

applied = apply_loyalty_discount(with_points, 120) declined = apply_loyalty_discount(without_points, 40) out_of_stock = apply_loyalty_discount(with_shortage, 500)

Assert

confirm( “applies 10% with enough points”, isinstance(applied, DiscountApplied) and applied.tab.total_in_cents == 3600, ) confirm( “no variant went unhandled”, {type(applied), type(declined), type(out_of_stock)} == set(DiscountResult.args),

) Python is the counterpoint in content. Its suite has four cases, not three. In Dart, Kotlin, Swift, Java, and Rust, a switch that forgot a variant of DiscountResult doesn’t compile, and forgetting a case turns into a compile error instead of a missing test. In Python, exhaustiveness only exists if an external type checker runs with assert_never , and that checker doesn’t run in the CI (Continuous Integration) of anyone who just runs pytest . The work the compiler did for free in the other five languages comes back to your own suite. The price is the fourth case: it gathers the types of the three already-computed results into a set and compares that set against what the union declares. Try it: open https://focus.kodel.com.br/en/dart/17-02 (or your language’s route) and delete the out-of-stock item case. Which test breaks: the out-of-stock one, or the discount-applied one? Repository: the fake you already have Go back to chapter 15 and look at what it left ready. Alongside the real repository and the boundary’s try/catch , that chapter published a second implementation of the same contract, in memory, called TabRepositoryFake . It’s been there since snippet 15- 02. This chapter doesn’t invent the fake: it harvests what chapter 15 planted. Dart class TabRepositoryFake implements TabRepository {

TabRepositoryFake({this.simulateNetworkOutage = false}); // The network outage becomes a flag, not a socket: the test controls // the outcome without touching any IO at all. final bool simulateNetworkOutage; @override LookupResult findTab(int table) { if (simulateNetworkOutage) { return InfraFailure(Failure.noConnection); } return TabFound(Tab(table, 1895, true)); } @override

SaveResult save(Tab tab) { if (simulateNetworkOutage) { return InfraFailure(Failure.noConnection); } return TabSaved(tab); } @override SaveResult markAsPaid(Tab tab) => save(tab); } Three things about this class matter. Its size isn’t one of them. First: it implements the three methods of the contract, the same three chapter 15 declared, not one more. Second: none of them touch network, disk, or a database; the network outage is a boolean flag, not a socket. Third: it fits on one screen, and it fits because chapter 15’s contract was designed small on purpose, with verbs from the feature instead of save(T) and getAll() .

A fake is exactly that: a simplified, working implementation of a contract, which you use to check STATE. Meszaros (2007) draws the line right here. The mock in the anti-solution checked interaction: which methods got called, and in what order. The whole suite depended on the orchestrator continuing to call the repository the same way it called it on the day the test was written. The fake checks nothing. It works, and the checking is done by your test’s assertion, which looks at the result. A mock asserts the path; a fake lets you look at the destination. Meszaros cataloged five kinds of test double, and this chapter uses two on purpose. Dummy, stub, and spy have their place, and I’m not going to teach them here: the distinction that changes an architecture decision is fake versus mock, and carrying the whole taxonomy would only make you memorize names. Orchestrator: flow test The orchestrator from chapter 13 is the middle piece, and the table from chapter 10 says it forbids “deciding rules and persisting.” Read that again as a testing promise: if it doesn’t decide rules and doesn’t persist, its only collaborator is the repository, and the repository already has a fake. What’s left to test? The flow. An event goes in, a sequence of states comes out. Before the test, the target. Chapter 13 declared the PayTab event and never registered a handler for it, so this chapter builds one, in a subclass called TabOrchestratorWithPayment . There’s no new business rule in it: it’s chapter 13’s cycle stitched together with chapter 15’s lookup and save, and the operation itself is a command, in the exact sense chapter 16 gave that word. Dart

class TabOrchestratorWithPayment extends TabOrchestrator { TabOrchestratorWithPayment(this._repository, {bool refactored = false}) : super(_repository) { on( refactored ? _onPayTabAfter : _onPayTabBefore, ); } final TabRepository _repository; } Now the test. There are three cases, and the fun part is the distance between them. Dart class FakeThatDoesNotSave extends TabRepositoryFake { @override SaveResult markAsPaid(Tab tab) =>

InfraFailure(Failure.noConnection); } blocTest<TabOrchestratorWithPayment, TabState>( “paying a tab that's found publishes Loading then Ready”, // Arrange: chapter 15's fake in its default setting, which finds // and saves. build: () => TabOrchestratorWithPayment( TabRepositoryFake(), refactored: refactored, ), // Act: an event goes in. act: (orchestrator) => orchestrator.add(PayTab(4)), // Assert: the sequence of states that comes out, not a word about // calls. expect: () => [isA(), isA()],

); blocTest<TabOrchestratorWithPayment, TabState>( “network outage on the LOOKUP publishes Loading then Failed”, // Arrange: the same fake, a different argument. build: () => TabOrchestratorWithPayment( TabRepositoryFake(simulateNetworkOutage: true), refactored: refactored, ), act: (orchestrator) => orchestrator.add(PayTab(4)), expect: () => [isA(), isA()], ); blocTest<TabOrchestratorWithPayment, TabState>( “failure on the SAVE publishes Loading then Failed”, // Arrange: a fake that finds the tab and doesn't save it.

build: () => TabOrchestratorWithPayment( FakeThatDoesNotSave(), refactored: refactored, ), act: (orchestrator) => orchestrator.add(PayTab(4)), expect: () => [isA(), isA()], ); TypeScript class FakeThatDoesNotSave extends TabRepositoryFake { override markAsPaid(_tab: Tab): SaveResult { return { kind: “infraFailure”, failure: “noConnection” }; } } // Fires the event and returns just the names of the published states.

function statesFrom(repository: TabRepositoryFake): string[] { const orchestrator = new TabOrchestratorWithPayment(repository, refactored) ; orchestrator.payTab(4); return orchestrator.states.map((s) => s.kind); } it(“publishes loading then ready when it finds the tab”, () => { // Arrange: chapter 15's fake in its default setting. const repository = new TabRepositoryFake(); // Act + Assert expect(statesFrom(repository)).toEqual([“loading”, “ready”]); });

it(“publishes loading then failed on a network outage during LOOKUP”, () => { // Arrange: the same fake, a different argument. const repository = new TabRepositoryFake(true); // Act + Assert expect(statesFrom(repository)).toEqual([“loading”, “failed”]); }); it(“publishes loading then failed on a failure during SAVE”, () => { // Arrange: a fake that finds the tab and doesn't save it. const repository = new FakeThatDoesNotSave(); // Act + Assert expect(statesFrom(repository)).toEqual([“loading”, “failed”]); });

Compare the first case with the second. The only difference between them is one construction argument on the fake, simulateNetworkOutage: true . You didn’t set up an expectation, didn’t teach the double how to respond, didn’t configure anything: it’s a flag chapter 15 had already left ready. The third case is this chapter’s argument in miniature. The default fake fails at the lookup and never gets to save, so the SAVE error branch would go untested. To cover it, extend the fake and override one method. FakeThatDoesNotSave has a single line of body. With a mock, the same scenario would cost one more chained expectation, and one more expectation is one more assertion about the orchestrator’s insides: one more spot that can break on the next refactor. About blocTest : it’s a convenience of the Dart ecosystem, not a requirement of the architecture. What matters is that the assertion is about the OBSERVABLE SEQUENCE OF STATES, and the TypeScript version right above does the same thing with an array and a toEqual . If you switch libraries tomorrow, these three tests stay valid, because what they assert is what comes out the door. Try it: open https://focus.kodel.com.br/en/dart/17-03 (or your language’s route) and comment out the markAsPaid line in the handler. Which of the three tests turns red? View and integration: where each one pays its own cost

Two ends are still missing. Chapter 12’s View forbids “business rules and data access,” and that ban answers the question on its own: there’s nothing to fake in a dumb View, because it has no collaborator. What’s left is rendering. If the screen just draws the state it received, field by field, a widget test would only confirm that Flutter knows how to draw text, and that isn’t your problem. Don’t write a single test. I’m explicitly authorizing the blank page here, because the alternative is people writing widget tests out of guilt or a sense of completeness that brings no value at all. The trigger is a rendering conditional. The moment the screen chooses between two drawings, it earned a behavior of its own, and a behavior of its own deserves proof. Rosie’s Coffee Shop’s pay button is the case: it shows up enabled or disabled depending on canPay , and a widget test that mounts the screen with canPay: false and looks for the disabled button catches the inverted boolean, the most common defect there is. On the other end, the real repository. It’s the only layer that talks to a database, and that’s why it’s the only one where the fake isn’t enough. The justification is concrete and singular: no fake ever catches a broken migration. You can have a hundred green tests against TabRepositoryFake and production still falls over because the paid_at column changed type in the database. The integration test runs against a real database, one per driver, and it exists to catch exactly what the fake can’t know: whether the SQL is right, whether the migration ran, whether the mapping matches. There are few of them, and they’re expensive. Few, because each one boots infrastructure; expensive, because each one takes time. One per driver is enough, because what you’re testing is the translation, and it’s the same for every query on that driver. This test can also be replaced by a SQL test, run with a tool or with a script executed by hand, to confirm the database really is what

the application expects. On PostgreSQL, the typical tool is pgTAP, a unit-testing framework written in SQL that runs inside the database itself. Rosie’s Coffee Shop, operation by operation: Operation Strategy Why Apply discount pure test pure function: input by hand, zero doubles Flag item out of stock pure test the rule decides, and never touches IO Pay the tab flow with a fake the sequence of states matters Add item flow with a fake the orchestrator only connects the ends Pay button widget test canPay chooses between two drawings List the menu no test no conditional, nothing to fail Save a paid tab integration a fake catches no migration or bad SQL Look up a tab integration checks the column-to-field mapping

The dashed line is the fake’s boundary. Everything above it runs in memory, in milliseconds, without booting anything. Below it lives the cost, and it stays confined to a single layer because chapter 15 put the try/catch in one place, and one place only.

The two critiques The first critique is the most serious, and the anti-solution already proved it: mocks couple the test to the implementation. You watched a suite turn red without the program changing behavior. Shai Yallin, in “Fake, Don’t Mock” (2023), argues that a double checking interaction turns the test into a copy of the code, and Martin Fowler, in “Mocks Aren’t Stubs” (2007), named the two schools behind that split. The classicist tests by state: real objects or fakes stand in, and the result gets inspected. The mockist tests by interaction: collaborators get replaced, and calls get checked. FOCUS sides with the first, and not out of taste: with rules living in pure functions and IO sitting behind a small contract, the classical school comes cheap, and the mockist one gets expensive for nothing. That doesn’t ban mocks. I use a mock when the dependency has no way of getting a fake that matches the real thing, typically a third-party SDK whose behavior I don’t control and can’t reproduce without guessing. Outside that, I write the fake and sleep better. If you disagree, the test is empirical: refactor the inside of one of your own services without changing behavior, and count how many tests break. The second critique is the dispute between Mike Cohn’s testing pyramid (Succeeding with Agile, 2009), which calls for many unit tests at the base, integration in the middle, and few end-to-end tests at the top, and Kent C. Dodds’s Testing Trophy (2018), whose motto is “Write tests. Not too many. Mostly integration.” and which shifts the weight to the middle. Which one is right? Neither, and the question is broken. Dodds is right about the base he saw: in an architecture where the business rule lives scattered across controllers and services stuffed with dependencies, a unit test only exists behind a wall of mocks, and a wall of mocks is

fragile and proves nothing. His answer was to move up a level. FOCUS’s answer was a different one: move the rule into a pure function, which is chapter 14 in full. When the rule lives in a pure function, the base of the pyramid gets cheap again, genuinely cheap, because applyLoyaltyDiscount needs no double at all. The shape of your suite is a consequence of your architecture, not a choice you make before you start coding. If your base is expensive, the pyramid isn’t the problem. Pitfalls The fake that lies. It’s this chapter’s main pitfall, and it’s silent. Your fake returns TabFound for any table; the real repository returns InfraFailure when the table doesn’t exist. The tests stay green and production fails, and the worst part is the suite stays green the whole time the bug is happening. The way out is a three-step discipline, and not one of the steps is writing more tests. First: the fake and the real one implement the SAME interface, and you never add a method to the fake that the contract doesn’t have. If chapter 15 declared three methods, the fake has three. Second: whenever the real one gains a new observable behavior, such as a new failure or a new Result variant, the fake gains the matching one in the same commit. The sealed Failure family helps here: adding a variant breaks every non-exhaustive switch that’s missing a default , and the compiler points you straight at the fake. Third: whenever doubt about divergence shows up, it’s a question about the real implementation, and the integration test is what answers it. A fake that lies is a fake that aged alone.

Chasing 100% coverage. Coverage measures lines executed, not claims made. A suite that runs every line and asserts nothing scores 100% and catches no defect at all. This chapter doesn’t promise full coverage, and it names what it deliberately skips: the View without a conditional, the formatter that only formats cents, and the real repository, left for the integration test. Testing the View without a conditional. If you write a widget test for a screen that only draws the state it received, you’ll be testing the framework, and you’ll pay for it every time you change a padding. Q&A What if the refactor had changed behavior? Then the mockist suite would be right to break, and I’d have no argument at all. That’s exactly why the preservation gets demonstrated first, with two empty diffs, before any mention of the break. A test that breaks when behavior changes is a good test. The mockist’s problem is breaking when behavior does NOT change. What does this chapter deliberately not test? The View without a rendering conditional, for having no behavior of its own. The real repository, which calls for integration and stays out of scope here. And the end-to-end path, the one that boots the whole app: it exists, it’s expensive, and one per critical flow is enough. If mocks are so bad, why does the library exist? Because it solves the case where you don’t control the dependency and can’t build a fake that matches the real thing. Mock is a tool of last resort, not first choice. The question I ask before using one is: could I write a fake faithful to this? When the answer is yes, the fake wins.

Do I need a fake per repository implementation? No. The fake belongs to the CONTRACT, not the implementation. One contract, one fake, and as many real repositories as the app needs. Quick tip Open your current project’s suite and search for verify , toHaveBeenCalled , assert_called_with , or your tool’s equivalent. Each hit is a claim about the INSIDE of something. Don’t delete anything yet: just count, and compare that count against the number of assertions on return values. The ratio between the two numbers is how much your suite is going to hurt on the next refactor. Quick reference Layer Test strategy Chapter View widget test only where there’s a conditional 12 Orchestrator flow test with a fake repository 13 and 15 Use Case pure test, no double at all 14 Repository fake for consumers, integration for the real one 15 Each row is born from a ban in chapter 10’s table. The View forbids “business rules and data access,” so there’s nothing to fake in it. The Orchestrator forbids “deciding rules and

persisting,” so its only collaborator is the repository, which already has a fake. The Use Case forbids “IO, framework, and domain exceptions,” so there’s no dependency to fake. The Repository does “CRUD (fetch and save), the only place an infra exception exists and becomes a Result,” and forbids “business rules”: it’s the only boundary that needs a fake, and the only one that pays for integration. Exercises

  1. Rosie decided to give a discount to customers celebrating a birthday. Add the case to the use case test at https://focus.kodel.com.br/en/dart/17-02 (or your language’s route): build a tab, pass the date, and assert the variant that comes out. The completion criterion is the diff: it must contain only the use case’s test file, and you can’t create, touch, or configure a single test double. If you needed one, the rule leaked outside the pure function.
  2. Open https://focus.kodel.com.br/en/dart/17-03 (or your language’s route) and simulate a network outage in the happy path’s flow test: assert Loading followed by Failed . The criterion is the count: exactly one construction argument on the fake changes. Can you write a third case, the one where the tab is found and doesn’t save, without touching the TabRepositoryFake chapter 15 published? Tip 17 Test what the piece promises, not how it delivers. A mock checks the how, and the how changes.

Next chapter: the architecture is complete and tested in one language, and chapter 18 opens Part IV by rebuilding the same slice of Rosie’s Coffee Shop in all ten, so you can find out what in FOCUS is an idea and what was just Dart’s accent.

Powered by TurnKey Linux.