Treat the context window like a budget, not a bucket
Most agent failures we debug are not reasoning failures. They are context failures. Here is how we budget the window at Replace Works.
By Replace Works
When an agent misbehaves, the first instinct is to blame the model or rewrite the prompt. In our experience most of these failures are simpler and more boring than that. The agent did the wrong thing because it was handed the wrong context, too much of it, in the wrong order. The model was fine. The plumbing was not.
The mental shift that fixed this for us was to stop treating the context window like a bucket you pour everything into, and start treating it like a budget you spend deliberately. Every token you add competes with every other token for the model's attention. Dumping an entire document, a long history, and ten tool definitions into one call does not make the agent smarter. It makes it distracted.
Start by deciding what the agent actually needs to make the next decision. Not what might be useful, not what is convenient to pass along, but what is required for this specific step. Most of the time that is a small slice: the user's current request, a short summary of what has happened so far, and the one or two facts that bear on the choice in front of it. Everything else is noise that raises cost and lowers reliability.
Retrieval is where budgets get blown the fastest. It is tempting to fetch the top twenty chunks and let the model sort it out. Fewer, better chunks almost always beat more chunks. We would rather return three highly relevant passages than twenty mediocre ones, because the irrelevant seventeen actively pull the model toward wrong answers. If your retrieval cannot be confident, it is better to say so and let the agent ask a clarifying question than to flood the window and hope.
History is the other silent budget killer. A long conversation does not need to be replayed verbatim on every turn. Summarize older turns into a compact running state and keep only the most recent exchanges in full. The agent does not need the transcript. It needs to know where things stand. This single change tends to cut token usage dramatically while improving answers, because the important facts are no longer buried under small talk.
Tools deserve the same discipline. Every tool definition you expose is context the model has to read and reason about on every call. An agent with thirty tools spends a meaningful share of its attention just deciding which tool not to use. Give it the smallest set that covers the job. If you have many tools, route to a relevant subset before the call rather than presenting all of them at once.
Order matters more than people expect. Models pay more attention to the start and end of the context than the middle. Put the instruction and the most decision-relevant facts where they will be seen, not in the soft middle where they get lost. When something must be obeyed, repeat it at the end. This is not elegant, but it is reliable, and reliability is the whole game in production.
The practical test we use is simple. For any agent step, can you point at every chunk of context and say why it is there and what decision it supports? If you cannot justify a token, cut it. We have never regretted removing context. We have regretted adding it many times.
None of this is glamorous. There is no clever prompt that substitutes for knowing exactly what the agent needs and giving it that and nothing more. But budgeting the window is the highest leverage work we do on agents. It is cheaper than a bigger model, faster than more retrieval, and it is usually the difference between a demo that impresses and a system that holds up.