Du kannst nicht mehr als 25 Themen auswählen Themen müssen entweder mit einem Buchstaben oder einer Ziffer beginnen. Sie können Bindestriche („-“) enthalten und bis zu 35 Zeichen lang sein.

22KB

Context Engineering — Chapter-25: RAG vs direct context

  • Source: /library/Context Engineering/source-file.pdf
  • PDF pages: 217–232
  • Pages without text: none

RAG vs direct context The utilization report’s subtask stalls on its first column. To say whether the open 2 p.m. interval on Tuesday counts as provider idle time, the isolated context has to know what Vila Nova Clinic treats as a no-show, what it treats as a schedule block and what grace period each insurance plan allows before the appointment turns into a work-in, one of the extra appointments squeezed into a full schedule. None of that is in VilaSchedule. It is in the care policies, which the clinical coordinator keeps in a shared folder outside the repository and revises every month. You ask for the material and get back the list of what is in the folder: Reconstructed for teaching: a sample of the policy base that Vila Nova Clinic's clinical coordinator keeps outside the VilaSchedule repository. The real base has 14 documents and around 320 pages; here are three of them, shortened, with the sections the examples in this book cite. [...] Review cycle: monthly. The clinical coordinator publishes the new

edition on the first business day of the month, and the previous one stops being in force that same day. [...] | cancellation.md | Cancellation, no-show, late arrival | 18 | 2026-07-01 | | insurance.md | Rules by insurance plan | 96 | 2026-07-01 | | workins.md | Work-ins and schedule blocks | 11 | 2026-06-01 | The 320 pages run past 200,000 tokens. You do what looks reasonable and paste in only the three documents that look relevant, some 60,000 tokens, and then the arithmetic of chapter 6 kicks in: the cycle resends the whole input on every turn, and an agent task with forty calls pays for those 60,000 forty times. That is 2.4 million tokens of policy per task, to answer a question that fits in two lines, and most of that text is about insurance plans this week’s report never mentions. It is chapter 5’s pain and chapter 6’s bill in the same session: you burned the budget before the first question. The second blow arrives on the first of the month. The clinical coordinator publishes the new edition, the free cancellation window goes from 24 to 48 hours, and your pasted copy keeps answering 24 with the same confidence as before. By copying, you have just re-created chapter 9’s dead document, except that this one lives in your window and has no owner and no test to cover it.

Notice what does not explain the problem. It is not a badly assembled packet in chapter 17’s sense: trimming requires knowing beforehand which passage the task will use, and here you only find out when the question shows up, interval by interval. It is not bad isolation from chapter 21: splitting the work into more contexts does not shrink the document by a single line. What you have is information with three properties at once, large, mutable and used in pieces, and information like that has no place inside the window. Fetching the passage when the question comes up Retrieval-augmented generation (RAG) is the arrangement in which the knowledge base stays outside the window and a search brings in, at the moment the question appears, the passage that answers it. Only that passage goes in. The 320 pages stay where they were, and what travels in the cycle is the handful of paragraphs today’s task actually consulted. This is the second time in the book that a session pulls text in from outside, and the two operations are worth keeping apart. Chapter 18 rebuilt the thread of a session that got lost, and the source it rebuilt from was the trail of your own work: notes, commits, the state of the repository. What this chapter describes starts somewhere else. The source is a base nobody lost, the text was never in the session, and the operation runs while the work is going well rather than after it has broken. One repairs the window; the other feeds it. The name comes from a 2020 paper. Lewis and colleagues presented “Retrieval-Augmented Generation for Knowledge- Intensive NLP Tasks” (arXiv:2005.11401) at NeurIPS, the Conference on Neural Information Processing Systems,

proposing a model that combines the parametric memory of trained weights with a non-parametric memory, an index of passages that a retriever queries and that the generator conditions on. Notice the gap between that and what the industry calls RAG today: in the paper, retriever and generator were trained together, pieces of one model; in 2026, RAG is the name for practically any arrangement in which retrieved text is pasted into the prompt of an off-the-shelf model. The name stuck and the design changed, which is worth remembering the next time somebody cites the paper to defend an implementation the paper does not describe. The 2026 mechanism has named parts. You cut the base into passages, turn each one into an embedding, a vector that stands for the meaning of the text, store those vectors in a vector database and, at question time, search by proximity, almost always mixed with keyword search. I record the names as a dated instance, the way chapter 21 treated the subagent: which database and how to keep the index current are Part IV’s business. The principle that survives the replacement of all those parts fits in one sentence: fetch the passage when the question comes up, instead of carrying the base along just in case. Notice, by the way, that you already do this with no infrastructure at all. When the agent runs a search in the repository and reads only the two files that matched, it is retrieving; the index is the file system itself, and the retriever is grep . And between grep and the vector database there is a step almost nobody counts as retrieval, though it has the best signal-to- noise ratio for code: structural search. Code is not running prose. It has symbols, definitions, references and a syntax tree, and the tools that understand that structure answer questions grep can only approximate: where this function is defined, who calls it, what this module exports. In 2026 that reaches the agent by more than one route: the language servers of the Language

Server Protocol (LSP), the same ones that feed your editor’s go- to-definition, and the syntax tree parsers, with tree-sitter as the instance that became standard across the Part IV tools. The answer to a search like that comes back exact, small and with no spurious matches: “who calls dayLimits ” gives back the three callers, not the forty lines that contain the word “limit.” For the code base, that step postpones the vector index for a long time. It is for the coordinator’s prose policies, where there is no syntax tree to consult, that search by meaning earns its place. Size, mutability and how each task uses it The useful question is not whether RAG works. It is which information deserves to leave the window, and you make that decision per piece of information, never for a whole project. Three axes are enough to decide. The first is size, and you measure it after chapter 17’s ladder, not before. Do not ask whether the base is large; ask whether what is left of it fits in the packet after you apply pointer, excerpt and whole file. VilaSchedule’s living doc is one page and goes in as three table lines. The policies are 320 pages and they do not shrink, because the task does not know beforehand which paragraph it will need. The second is mutability, measured against your own work cycle. A project convention changes over months, and a copy of it in the window ages slowly. The policy base gets a new edition on the first of every month, with a declared owner and a declared effective range, and any copy you keep turns into a lie on a known date. There is a third degree of mutability, the data that changes between your question and the model’s answer, and it fits in neither of this chapter’s two destinations.

The third is how each task uses the information, and it is the axis most people forget. What matters is not how many times the base is consulted; it is whether every task uses the same piece or each task uses a different one. Information that nearly every task consults, always the same, tends to be the longest-lived, layer 1 of chapter 16, and it belongs in the packet whatever it costs. Information from which each task consumes one unpredictable paragraph is a natural candidate for search. The three axes point to three destinations, and the third one only gets its name here because the next chapter is entirely about it: embed, retrieve or expose as a tool. To embed, here, is to put the text in the packet by hand; it has nothing to do with the embedding of two sections back, which is a vector. The two words are neighbors in spelling and nothing else. Here is the cheat sheet I use to decide:

  • Embed: the text goes into the task packet, chosen by you before the session starts.
  • Retrieve: the text stays outside the window and a search brings the passage in at the moment the question comes up.
  • Expose as a tool: no text goes in; the model asks the question and the system answers with the value as of now. [...] | Axis | Embed | Retrieve | Expose |

|---|---|---|---| | Size | Fits whole | Does not fit even trimmed | Not applicable | | Mutability | Months | Weeks or months | Between question and answer | | Use | Nearly every task | One passage per task | Always, a fresh value | | Choice of passage | Yours, beforehand | The search's, on the spot | There i s no passage | | Typical failure | Large and visible | Wrong and silent | Down | | Cost to maintain | None | One index per edit | One integration | [...]

  1. Does it fit embedded? If the whole piece fits in the task packet along with the rest, embed it and stop here. Do not index what fits.
  2. Is it born stale in the window? If the value changes between the moment of pasting and the moment of answering, neither embedding nor retrieving works: expose it as a tool.
  3. What is left large and mutable on a slow cycle? Retrieve that,

and only after meeting the three conditions below. 4. When torn between embedding and retrieving, embed. The mistake of embedding is expensive and visible; the mistake of retrieving is cheap and invisible. [...] The fourth question is this book’s position, and it deserves a defense, not just a restatement. Embedding is the default until it hurts The defense has three parts, and the first is the asymmetry between the two errors. The oversized packet fails in a way you see: the bill goes up, the tool’s token counter says so, the window gets tight and quality drops the way chapter 5 measured. The search fails in a way you do not see: it gives back three plausible paragraphs, the model answers fluently about them and nothing on screen says that the paragraph that settled the question stayed in the base. Too much context is an expensive, loud mistake; a search that misses is a cheap, quiet one. Between a failure that screams and a failure that smiles, the default goes to the one that screams. The second part is who does the choosing. Packing is your own admission criterion, applied beforehand, with the whole task in view, and chapter 17 showed that the hard part of it is deliberate subtraction. Retrieving hands that admission over to a ranker that does not know the task, only the wording of the question,

and that decides by textual similarity. When similarity gets it wrong, it gets it wrong with no warning and no record of what was left out. The third is the cost of maintenance, which nobody adds up while the two are being compared. An index is one more artifact in your project, and artifacts age. The policy base gets a new edition every month, and an index built from the June edition will keep answering from June long after July is out, without a word of complaint. A stale index is chapter 9’s dead document with a search on top, which makes it worse: easier to consult and just as false. Hence the rule I use, and I state it as an opinion: embedding is the default until it hurts. Hurting has three symptoms, and I want all three before indexing anything. The first is that the information does not fit even after chapter 17’s ladder, already trimmed to the minimum, and still takes up tens of thousands of tokens in every task. The second is that each task consumes a different piece and you cannot predict which one; if you can, the predictable piece goes back into the packet and the problem is over. The third is that the source changes on a cycle that is not yours, on a set date, and your copy ages between one task and the next. One symptom on its own is not enough: a huge base whose passage you know beforehand is an excerpt, not a search. The best test of that rule is in the coordinator’s own folder. The work-in document is eleven pages, and two of its sections are exactly the kind of thing that should never leave the window:

2. Who authorizes it

The front desk grants up to 2 work-ins per provider per day. Beyond

that, only the clinical coordinator authorizes it, case by case, and records the reason in the day's report. [...]

4. Blocked schedule

A schedule blocked for vacation, a conference or a long procedure takes no work-in under any circumstances. There is no partial block at this clinic: a block is either blocked or open. Those two rules fit in three lines, hold in every task that touches work-ins and have changed once in two years. They show none of the three symptoms, so they stay embedded, and they already were: they are the same lines that chapter 9’s living doc verifies in continuous integration (CI) and that chapter 17’s packet carries at the top. Notice what that does to the decision: the same folder, from the same owner, in the same month, has a document that goes to search and a document that goes to the packet. If you index the whole folder because the folder is large, you have handed the ranker the most consulted rule in the system, and the day it does not rank high enough is the day the agent reinvents chapter 16’s allowsWorkInDuringPartialBlock . “RAG retrieves the wrong passage”

The most serious criticism of retrieval does not come from people who have never used it. It comes from people who have put it into production and cataloged the damage. Barnett and colleagues published “Seven Failure Points When Engineering a Retrieval Augmented Generation System” (arXiv:2401.05856) in 2024, drawn from real systems in three domains, and four of the seven points live in retrieval: the content simply is not in the base; it is there, but it does not rank high enough; it comes up, but it does not enter the window because of the cut; it enters, but with the wrong specificity, answering in general terms what the question wanted in particular. The last one is what bites here, and it is treacherous because the retrieved passage is true. Ask the base what a patient’s grace period is, and the text that most resembles the question is section 4 of the cancellation document, which answers in prose, with the same words, that a patient up to 10 minutes late is seen inside their own interval. The answer the report needs is in another document, in a table with none of those words:

1. Contracted grace period

Each contract sets its own grace period, and it prevails over the general rule in section 4 of cancellation.md: | Plan | Grace period | After that | |---|---|---|

| Southline Health | 20 min | Work-in at the end of the block | | UniHealth | 10 min | Rescheduling | The report comes out with the general rule applied to everybody, and the Southline Health patient who arrived 15 minutes late shows up as a no-show, which turns into a charge on the month’s bill. Nobody suspects anything, because the answer is plausible, coherent and traceable to an official document that does say that. Three things answer that criticism, and none of them is a better ranker. The first is this chapter’s criterion, which shrinks the surface at risk: only what shows all three symptoms goes to search, so conventions, standing rules, the task spec and the work-in rules never pass through a ranker. Index everything and every question becomes a lottery; index only what does not fit and you draw a few times a day. The second is to require an address on whatever comes back. A retrieved passage that arrives on its own is impossible to check; a passage that arrives with document, section and effective date takes five seconds to read, and those five seconds are what tell you it came from cancellation.md when the question was about an insurance plan. That condition depends on the source having citable units, which leads to the next criticism. The third is to treat what comes back as a statement, not as truth. Chapter 19 already gave you the yardstick for what the AI asserts, and a retrieved passage falls under the same rule: it is a claim about the clinic, with an address and a date, waiting to be checked. How much rigor you apply depends on what is at stake. To pick the label of a column in an internal report, a retrieved

passage is enough. For a line that turns into a charge on a patient’s bill, no retrieved passage goes to production without the clinical coordinator having looked at it. “Chunking fragments meaning” The second criticism attacks the step before the search. Chunking is cutting the base into units small enough to fit in the window and specific enough to be found. Every cut is a bet about where meaning ends, and the cancellation document shows the bet being lost:

2. Late cancellation

A cancellation made less than 24 hours ahead is recorded as a late cancellation and carries a charge of 50% of the self-pay rate for the appointment. [...]

6. Exceptions by insurance plan

The charges in sections 2 and 3 do not apply to the plans listed in appendix B of insurance.md, which prohibit charging the patient for

cancellation and no-show by contract. Between the two sections there are four others, and no reasonable chunker keeps the two in the same passage. The question about how much a late cancellation costs retrieves section 2, which answers 50% with no sign that section 6 exists. The answer is confident, it is citable and it is wrong for two of the clinic’s plans. The reply I often hear is that the chunker needs to improve, with overlap between passages, cutting by heading, a hierarchy of sections. That improves things at the margins and does not solve this, because meaning was not fragmented by the chunker: it was fragmented by the person who wrote the document, when the rule was separated from its own exception by four sections. The fix is upstream and you already know it from chapter 9: the base has to be written in units that survive the cut, each rule next to the exception that limits it, each unit with a title, an owner and an effective date. Living documentation is not a privilege reserved for code artifacts. A policy base written that way becomes searchable, and the same base written as running prose keeps producing passages that are true and misleading. What is left of the criticism still stands, and I would rather record it than paper over it. When the base belongs to somebody else, the clinical coordinator, legal, a vendor, you cannot rewrite it, and then the upstream fix is not available. Two ways out remain, and both come at a cost. You retrieve larger units, the whole document instead of the passage, paying in tokens what you cannot pay in editing, which in 2026 is workable for documents a few dozen pages long and remains unworkable for the whole base. Or you take that part out of the automation and send the

question to a person. The third way out, index it however you can and trust what comes back, is the one that produces the wrong charge on the bill. It is worth saying what this chapter assumes is in place. The decision to embed depends on chapter 17’s packet, and without it the alternative to search is the dump, which makes any retrieval look great by comparison. The decision to retrieve depends on the source having chapter 9’s properties, an owner, an effective range and citable units, because without them the passage comes back with no address and you have no way to know which edition it came from. And checking what came back depends on chapter 19’s validation. Without those three, the technique degrades in a specific and known way: the base enters the window through the search door instead of the copy door, with the same lack of provenance as before, and now with a layer of infrastructure between you and the error. The column the search does not answer With the criterion applied, almost all of the utilization report comes together. The work-in rules go into the packet, the cancellation and insurance policies stay in the base and come up passage by passage, with address and effective date, and chapter 21’s isolated context fits in one window again. One column is left over, and it fits neither destination. The clinical coordinator wants to see, next to yesterday’s utilization, which of tomorrow’s intervals are still open. That number is in no policy: it is in VilaSchedule’s database and it changes with every appointment the front desk makes while the report runs. Pasted into the window, it is born stale. Indexed, it goes stale at the first appointment after indexing, and rebuilding the index every minute for a value read once is work thrown away. It is the

case of the third degree of mutability, the data that changes between your question and the answer, and this chapter’s table already gave its destination without explaining how: expose it as a tool. Information like that is not read; it is asked for, and the next chapter is about what a question like that costs the window before it is answered.

Powered by TurnKey Linux.