# Context Engineering — Chapter-26: MCP and tools as dynamic context - **Source**: /library/Context Engineering/source-file.pdf - **PDF pages**: 233–245 - **Pages without text**: none --- MCP and tools as dynamic context The last column of the report is the one the clinical coordinator most wants to see, and it is the only one that does not close. Along with utilization for the day that just closed, they want to know which open slots are left on tomorrow’s schedule, so the front desk can start calling the waitlist today. You do what anybody would do. You export tomorrow’s schedule, four providers, blocks from 8 a.m. to 7 p.m., with patient and insurance plan in every taken block, and you paste nearly three hundred lines into the window along with the request. None of that is large in the sense of chapter 22: the export covers a single day, it fits with room to spare, and this time you knew in advance exactly which passage the task was going to use. The session goes well for half an hour. The agent reads the pasted schedule, builds the column and throws in a suggestion: Dr. Alves’s Wednesday morning is nearly empty, and the patient at the top of the waitlist could take 8:30 a.m. You pass that on to the front desk and hear the one line that ruins your afternoon: the 8:30 was taken at 2:10 p.m., and so was the 9:00. The file sitting in your window was exported at 2:07 p.m. Notice what does not explain this error. Context was neither missing nor in excess: the packet had the day, the four providers and nothing else, as chapter 17 asks. It was not chapter 22’s wrong retrieval, because there was no search anywhere along the way. It was not chapter 19’s stale claim, because the agent repeated faithfully what was in the window, and what was in the window was true at the moment you copied it. Every piece of text you paste carries an invisible date, the instant of the copy, and you never had to think about that date because conventions, architecture decision records (ADRs) and the living doc age in months. A schedule ages in minutes. Pasted, it is stale on arrival; indexed, it goes stale at the first appointment after the indexing, and rebuilding the index every minute for a value read once does not hold up. Information you do not read but ask for A tool, in this book, is a question your system lets the model ask, with the answer produced the moment you ask. Exposing a piece of information as a tool means deciding that it never enters the window as text: what enters is the right to ask, and the value only shows up after somebody pulls the trigger. What separates this from chapter 22’s two destinations is time, not size. Embedding and retrieving both work on text that already exists. You choose beforehand which passage goes in, or you delegate that choice to a search at the moment of the question, but in both cases the text was written somewhere before the session began. A tool’s answer was not written anywhere: it is computed when you ask, and that is why it reaches the window with an age of zero. Call that dynamic context and hold on to the principle, because it is what survives the replacement of every name in this chapter: information whose value changes faster than your session is fetched at the moment of the question instead of being loaded beforehand. It is chapter 22’s sentence taken one step further. There, what the search brought in was a passage from a document that stayed the same while you worked; here, what the question brings back is a value that did not even exist when the session opened. You have been doing this since chapter 4 without calling it that. When the agent runs the test suite and reads the output, it is not reading an old file: it is asking the system what the state of the suite is right now, and the answer comes into existence right there. When it runs a search in the repository, same thing. Tools predate any protocol, and the habit this chapter asks for is recognizing which of your pastes are, underneath, questions you are answering by hand. The name this has in 2026 In 2026, the mechanism that standardizes these questions is called Model Context Protocol (MCP). It is an open protocol, published by Anthropic at the end of 2024 in “Introducing the Model Context Protocol” (anthropic.com/news, with the specification at modelcontextprotocol.io), and the problem it attacks is plumbing. Before it, every AI tool talked to every data source through an integration somebody wrote by hand, and the bill was the length of one list times the length of the other. With a protocol in the middle, the side that holds the data publishes its questions once, and any client that speaks the protocol can start asking them. The observable effect in VilaSchedule is this: the availability query is written once, on the system side, and turns up in anybody’s session, in any tool that speaks the protocol, with nobody pasting a thing. Which server to use, how to install it, where the configuration file lives and what to do when it does not come up is Part IV’s business, which translates this book’s principles tool by tool. What matters here is only the idea the mechanism is an instance of. I am recording the date because it will matter. The protocol is from 2024, it took over the market during 2025 and it is what exists as I write. Nothing guarantees it will be the standard of the next decade, in the same way that chapter 21’s subagent and chapter 22’s vector database are dated answers to questions older than they are. What does not change is the property that forces the arrangement: there is information whose value changes between the moment you assemble the packet and the moment the model answers, and for that information no copy works, pasted or indexed. The definition is what the model reads Publishing the question is the easy part. What decides whether the tool gets used correctly is its definition, the text that describes it to the model and that travels in the window before any call. Here is VilaSchedule’s, excerpted, with each [...] marking what did not fit on this page: ## Name scheduling_open_slots ## Description (this is the text the model reads to decide to call it) Returns the open slots on a provider's schedule, on one day, as they stand ri ght now. Use it whenever the answer depends on what is open or taken right now: proposing a time to the patient, checking whether the day still fits a work-in, confirming that a time mentioned in the conversation is still free. Do not use it for: the work-in rules (limit per day, duration, blocked schedule), which are in the project's living doc and already came in the packet; closed days in the past, which come from the utilization report; insurance, cancellation or no-show rules, which are in the clinic's policies. Read only. This tool does not schedule, does not cancel and does not move any appointment. [...] | Name | Type | Required | Accepted values | |---|---|---|---| | provider_id | text | yes | provider active at the clinic | | date | date YYYY-MM-DD | yes | from today up to 60 days out | | duration_min | integer | no | 15 or 30; defaults to 30 | Errors it returns instead of guessing: UNKNOWN_PROVIDER, DATE_OUT_OF_WINDOW and SCHEDULE_BLOCKED (the day exists and takes no appointment at all). [...] { "queried_at": "2026-07-28T14:31:07-03:00", "provider": "Marina Alves", "date": "2026-07-29", "duration_min": 30, "open": ["08:00", "10:30", "11:00", "16:30"], "blocks": [ {"start": "12:00", "end": "13:00", "reason": "break"} ], "workins_today": {"used": 1, "remaining": 1} } [...] Four things in that text do the heavy lifting, and three of them talk about what the tool does not do. The first is the “do not use it for.” The model picks which tool to call by reading the description, and that pick is a guess of the same kind chapter 22’s ranker makes, except that here you write the text the guess is made from. A description that only says what the tool does invites the model to use it for everything that sounds close: asked whether a work-in, one of the extra appointments squeezed into a full schedule, fits at 3 p.m., it queries the open slots, sees that the block is free and answers yes, ignoring the limit of two per day that was in the packet. Saying where the rule question should go costs three lines and heads off the detour. It is the “what it does not need to know” section of chapter 21’s subtask contract, turned inside out. The second is the stamp. The queried_at field is the most important thing in the return and the easiest to forget, because at the instant the answer reaches the window it becomes pasted text like any other, and the opening problem starts over: ten turns later, that list of open slots has the age of the conversation. With the stamp, the model has a way to know the answer has aged and you have something to check against. Without it, the tool has merely pushed the aging from hours to turns and hidden the clock. The third is the size of the return. A tool that hands back the day’s whole schedule has solved nothing: it moved the dump from the opening to a later turn, with the ceremony of an integration along the way. The second condition of chapter 21’s criterion, the small return, applies in full here, and for the same reason: what comes back has to be the answer, not the base. The fourth is the boundary, and it is borrowed from the previous book. In FOCUS Architecture (https://books.kodel.com.br/en/books/focus/), the second book in this trilogy, the rule of thumb is that a slice talks to a slice through the front door, never by importing a neighbor’s internal file. The tool is that same door, opened to a caller that is not code: the model asks the public surface of the scheduling slice, the same one reports and workins have queried since chapter 14, and never touches a table. The extension is mine and not the previous book’s, which deals with a slice calling a slice; what I take from there is the rule about where you come in. The gain is the usual one: the limit of two work-ins per day comes out of the function the system already runs in production, so on the day the clinical coordinator changes that number, the return changes with it and nobody has to remember to edit the tool. Every tool is context paid for before the question Now the arithmetic. The definition you just read travels in the window on every call of the session, including the ones that have nothing to do with scheduling. It is layer 0 in the sense of chapter 16, the standing load you do not assemble per task and that is already there when the session opens. In chapter 2, when you asked your agent for the list of what had traveled along with a two-sentence prompt, the tool schemas showed up in that list, next to the instructions of the connected MCP servers, and the total measured tens of thousands of tokens before any work. That changes how you look at a tool catalog. Each one you connect is a bet that its question will come up often enough to justify the space its definition takes in every session, including the weeks when it is not called once. Twenty tools turned on just in case are chapter 17’s bloated packet again, with the added problem that they are invisible: they do not show up in what you typed, and the item-by-item inventory you learned to make there is almost never made here. The cost does not stop at the token. Anthropic takes this up in “Effective context engineering for AI agents” (2025, anthropic.com/engineering), the same text that supported the idea of context as a curated resource in chapter 16. The recommendation is that each tool have a clear purpose and not overlap with the others, because a bloated set produces an ambiguous decision point, and the yardstick they propose is direct: if a human engineer cannot say with certainty which of two tools to use in a situation, there is no reason to expect the agent to choose better. Two schedule queries with similar names cost more than the sum of their definitions. They cost you wrong calls. The discipline, then, is chapter 17’s, applied to the catalog. Declare the ceiling before connecting, measure what your session’s standing load already consumes and put every new tool through the packet’s second question: if you take this out, does the answer change? For VilaSchedule’s open slots query, it changes in every task that touches scheduling. For a tool that handed back the year’s holidays, it would change nothing: that is a twelve-row table that ages once a year, and their destination is the packet. When the data calls for a tool Chapter 22’s table gave the destination without saying how to recognize it. Three conditions have to hold at the same time, and the criterion is my own opinion, formed by integrations that should never have existed. The first is a shelf life shorter than the session. Ask how long the value stays true after being copied. If the answer is months, embed it. If it is weeks and the text does not fit even trimmed, retrieve it. If it is minutes, no copy works, and that is where exposing comes in. Tomorrow’s open slots change with every appointment the front desk schedules, a patient’s no-show history changes when they miss one, and the count of work-ins already used today changes while you read this sentence. The second is a question you can state, with a small answer. You need to be able to write, right now, the name of the question, the parameters it takes and the format of what it returns, the same way chapter 21 required the subtask contract to be writable today. “Which times on this provider’s schedule are open on this day” passes. “What is going on at the clinic” does not, and the temptation to expose a tool like that ends in the predictable place: you have reinvented the dump, now with latency. The third is that a source exists that can answer right now. A tool presupposes a system on the other side, with the answer computable at the instant of the question. If the data lives in a spreadsheet somebody updates every Monday, it does not change on every query: it changes every Monday, and its destination is to be embedded with the date attached, or retrieved. Fail one of the three and you do not expose. And there is a case that passes all three and still does not pay off: the value the task looks up once, at the start, and whose later change does not alter the result. Yesterday’s utilization is like that, because the day has closed. Pasting the number with the time of the query beside it costs one line and settles it. Exposing is for the data the task asks about several times over the course of the work, or that has to be right at the instant of the answer because somebody is going to act on it, which is the case of the patient on the waitlist. “That is a whole integration to read four times” The objection is fair and comes from people who have paid for an integration. Writing, publishing and maintaining a tool costs more than copy and paste. The arithmetic goes wrong in two places. The first is that the query already exists. VilaSchedule’s scheduling slice has published the day’s availability since chapter 14, because reports and workins need it, and chapter 21’s subtask contract already listed it among the available front doors. What the tool adds is the definition, the text you just read, and publishing it to a caller outside the repository. I am not proposing a new piece in the architecture; I am proposing that the door the neighboring slices already use be opened to the model as well. The second is what the comparison ignores on the other side. The alternative does not cost zero. Either you export the schedule every turn, and then the most expensive person in the process has become the tool, or you export it once, and then the wrong call to the patient on the waitlist all over again. Add the one page of definition to the handful of lines of query that already exist, then compare that total with maintaining one index per edit, which was the price chapter 22 charged for retrieval. “And when the tool is down?” It does go down, and chapter 22’s table already named the typical failure of exposing: it goes down. It is worth comparing with the neighbors before treating that as a grave defect. Embedding fails large and visible, retrieving fails wrong and silent, and exposing fails in a way nobody mistakes for success, because the call does not come back and the work stops at that point. Of the three, it is the one that produces the fewest wrong decisions. The real risk is the model filling the gap on its own. That is why the definition declares the expected errors, by name, so the return says “unknown provider” instead of handing back an empty list that the model reads as “no open slots.” A named error is what separates the tool that stops from the tool that misleads. And chapter 19’s yardstick still holds: what comes back from a call is a claim with an address and a time, not permanent truth. The difference is that here the address is your own system and the time comes stamped, which makes checking cheap rather than something you skip. It is worth naming what this chapter assumes is already in place. The tool depends on chapter 14’s boundary: with no front door published, it ends up written straight against the database, and the first thing anybody does after writing a raw query is reimplement the work-in limit rule inside it, because the return needs that rule. Then the clinic has two truths about the same limit, and the one that answers the model is not the one that runs in production. The description depends on chapter 9’s living doc, because the “do not use it for” has to point to a place where the rule is written down and verified, and not to your memory. And the catalog depends on chapter 17’s packet discipline, because with no ceiling it grows by addition and layer 0 eats the window before the first request. What the three decisions still do not say With the open slots column settled, the utilization report closes, and with it closes the block that began on chapter 21’s Thursday. You have four destinations for any information that shows up in a task: embed what is small and stable, retrieve what is large, mutable on a slow cycle and consumed in pieces, expose what changes faster than your session and give the work that is large, divisible and disjoint in its diff a context of its own. These are architecture decisions, and you make them once per piece of information and per task, not on every question. There is, however, one question none of the four decisions answered, and it was inside this chapter the whole time without anybody asking it. The schedule you exported carried patient and insurance plan in every block, and it left the clinic the moment it entered the window. The document the front desk forwards you tomorrow will come in the same way, and everything this book has taught so far treats what comes in as possibly wrong, too large or stale, never as possibly confidential, and never as possibly ill-intentioned. How much authority each piece of text gains when it enters the packet, and what happens when one of them arrives carrying instructions of its own, is the subject of the next chapter.