Você não pode selecionar mais de 25 tópicos Os tópicos devem começar com uma letra ou um número, podem incluir traços ('-') e podem ter até 35 caracteres.

12KB

Context Engineering — Chapter-10: Prompt engineering vs context engineering: why the prompt became a second-order variable

  • Source: /library/Context Engineering/source-file.pdf
  • PDF pages: 58–64
  • Pages without text: none

Prompt engineering vs context engineering: why the prompt became a second-order variable You probably keep a collection of prompts that “work” in some notes file: the one that asks for code step by step, the one that tells the model to act as a senior engineer, the one that promises the model a cash tip. A coworker sends another on Slack, “this one changed my life,” and the collection grows. And your results stay inconsistent, like Tuesday and Thursday in chapter 1. You have swapped magic prompts three times this year. If phrasing were the main lever, the third swap would have solved it. This chapter closes Part I with the thesis of the book. To defend it, I first have to do justice to the discipline it demotes: prompt engineering worked, still works and solves real problems. The argument is not that optimizing prompts is silly; it is that the prompt is the second-order variable, and you spent years optimizing the wrong term of the equation. What prompt engineering really solves Prompt engineering is the practice of improving a model’s result by adjusting the instruction: phrasing, structure, examples, assigned role. Rejecting that practice as superstition would be a straw man, and straw men do not survive contact with attentive readers, so let me set down what it demonstrably delivers.

Format instructions work: asking for the answer in JavaScript Object Notation (JSON) against a given schema, capping the length, requiring code with no comments. Role assignment works as a shortcut to a register: “answer as a security reviewer” shifts the vocabulary and the focus of the answer. Examples inside the prompt (the technique known as few-shot prompting) measurably improve performance on classification and extraction tasks. Breaking the request into explicit steps reduces error in reasoning tasks. None of that is folklore; it is the bread and butter of anyone building a product on top of a large language model (LLM), it is documented in every major provider’s official guides, and this book assumes you will go on using all of it. I use it every day. Now let me mark out the limits. Look again at the list: format, register, examples, decomposition. All of it operates on how the model should process and present the available information. No item creates information that is not there. And all of Part I has just shown that the bottleneck in your real sessions is almost never the processing; it is the available information. The perfect prompt does not contain your project’s error pattern (chapter 1), does not keep what the session settled on from falling out of the context window (chapter 3), does not dig the rule out of the dead zone in the middle of the input, where chapter 5’s U-shaped curve bottoms out, and does not take the dead log off the bill (chapter 6). Prompt engineering stops working exactly where your problems start. The same prompt, opposite results The thesis fits into an experiment you can run today, and it answers the third question this book asked you at the start of this part.

Take a good prompt, honestly good: “Fix the discount calculation bug described below. Follow the project’s conventions, handle the error cases in the existing pattern and write the matching regression test.” Run it in two scenarios, both in an ordinary chat: the kind with no access to your files. Scenario A: fresh session, the prompt on its own, plus the file with the bug pasted into the conversation. The model does not know the conventions, so it invents a plausible pattern; it does not know the error pattern, so it picks exceptions where the project uses a typed result, an error returned as a value instead of thrown; it does not know the test framework, so it guesses the most popular one. The answer is fluent, it is confident, and every line of it is rework. Scenario B: the same fresh session, the same prompt, word for word, but ahead of it you put the project’s convention guide, an example of a handler in the current error pattern and an existing test as a reference. The same model, with the same instruction, now produces code that passes your review. The prompt did not change by a single word; the result went from unacceptable to ready. The variable that decided the outcome was the context, and that is what “second order” means, with no hyperbole: with the right context, a mediocre prompt gets the job done; with the wrong context, the best prompt in your collection hallucinates elegantly. Phrasing adjusts at the margin; information decides the result. You may object that scenario A is dated, and the objection is fair: if you use a modern coding agent, it almost never starts from scratch. Before touching the bug, it lists files, searches the repository, reads the handler next to it and finds the test framework on its own. Just watch what the agent does in those first seconds: it is not reasoning better about the same prompt; it is assembling scenario B by itself. Automatic exploration is

context engineering carried out by the tool, and the fact that agent vendors have built that step into everything they ship is the strongest piece of public evidence for this chapter’s thesis: they found out, by measuring, that the result was decided there. The proof that the variable is still the context shows up when that assembly fails, and it fails often: the agent reads the wrong file in a monorepo, a single repository holding many projects; the convention lives in a document the search does not reach; and what the session agreed on has already left the window, as chapter 3 showed. The prompt is the same, the agent is the same, and the result degrades all the way to scenario A. Delegating the assembly of the window does not remove the discipline; it only changes who carries it out, and the rest of the book is about you taking that control instead of hoping the automatic step gets it right. The discipline that takes its place The industry noticed that inversion and named it in public. In June 2025, Tobi Lütke, chief executive officer (CEO) of Shopify, wrote on X that he liked the term “context engineering” better than “prompt engineering,” because “it describes the core skill better: the art of providing all the context for the task to be plausibly solvable by the LLM.” Andrej Karpathy, formerly of OpenAI and Tesla, endorsed the term that same month, in the same place, and defined it as “the delicate art and science of filling the context window with just the right information for the next step” in any industrial-strength LLM application. That pair of posts became the turning point in the vocabulary shift, and the term caught on fast because it did not invent a new practice: it only put a name on what agent practitioners had already learned the hard way in production.

The working definition of this book: context engineering is the discipline of deciding deliberately what goes into the model’s window on each call, with what structure and at what cost. Those three terms hold everything Part I established. Deliberately, because chapter 4 showed that, with no decision, the cycle decides for you, and decides badly. Each call, because chapter 3 showed there is no memory, only reassembly. At what cost, because chapters 5 and 6 showed that context has a double price, in attention and in dollars. The definition shifts the question itself. Prompt engineering asks “how do I ask better?”; context engineering asks “what does the model need to know, and how do I guarantee that this, and only this, is in the window?” You answer the first question once and it turns into a note in your prompt file. You answer the second one again on every task, because the information needed changes on every task, and that is why one is a trick and the other is engineering. And let me spell out the limits of the thesis, because a thesis with no declared domain turns into a slogan. “The prompt is a second- order variable” holds on this book’s terrain: long, situated tasks, with a repository, where the right answer depends on what the model knows about your system, and that knowledge is not in the weights; it is in your files. That is the terrain of the developers this book serves. Outside it, the hierarchy inverts, and I should say where: in a short, self-contained task, with no external knowledge (classifying tickets, extracting fields against a schema, routing messages, locking the output format), there is no context to engineer beyond half a page, and the instruction with good examples is the biggest lever available, as the list at the start of the chapter showed. Anyone building that kind of pipeline is right to spend a week on the prompt. The bounded thesis comes out stronger, not weaker: it says when each

discipline rules, instead of demoting either one across the board. On your terrain, phrasing is still worth a percentage point or two at the margin. Just do not confuse the margin with the center. The objections that deserve an answer Two criticisms of that vocabulary shift have circulated since 2025, and both deserve an answer instead of silence. The first: “context engineering is just prompt engineering with a new name, consultant marketing.” The answer sits on the technical boundary Part I drew. Prompt engineering operates inside one message: phrasing, structure, examples. Context engineering operates on the system that assembles the window: what the tool injects, what the cycle accumulates, what survives truncation, what each token costs. You solve one by editing text; the other demands understanding the mechanism of chapters 1 to 4 and measuring the effects of chapters 5 and 6. Calling both by the same name is like calling both the query and the schema design “writing structured query language (SQL)”: the name covers the notation, not the job, and the confusion is only possible from a distance. The second criticism is more serious: “all of this is transitory; better models will do away with curation.” I will concede part of that: windows grow, attention improves, and part of today’s hygiene will be unnecessary tomorrow. But the underlying limit is not one of engineering; it is one of logic: no model, however good, guesses information it never received. Your project’s error pattern, the decision made in yesterday’s meeting, the client’s constraint: either that goes into the window, or it does not exist for the model. Better models reduce the cost of imperfect context;

they do not remove the need for the right context. The discipline survives the next generation of models because the problem it solves does not live in the model. Where the right context comes from Part I ends here, and it ends on an open question on purpose. If quality is a function of the context, and the context has to be assembled deliberately on every call, the question that defines the rest of the book is: where does the right context come from? Part II’s answer has a name and you already know it from another context, if you have read my previous book: specification. In Spec Driven Development (https://books.kodel.com.br/en/books/sdd/), I argue that executable specifications replace loose prompts as the unit of work when you build with AI; you do not need to have read that book to follow this one, but the bridge between the two is exactly the chapter that comes next. A well-written spec is, among other things, perfectly packaged context: what to build, the constraints, the examples, the acceptance criteria, everything scenario B had and scenario A did not, in auditable and reusable form. Part II shows how specifications, architecture documents and recorded decisions become the raw material that fills the context window, and what changes in your routine when the context stops being improvised copy-paste and becomes an engineering artifact. You close this part knowing why the AI “forgets,” why the session degrades, how much that costs and which variable actually changes the outcome. Keep the prompt collection, which still has its uses; just demote it from strategy to tactic. What is left is learning to assemble the variable that rules, and that is where we are going.

Powered by TurnKey Linux.