Context Engineering for Agents
Context engineering means choosing the instructions, files, and recent results an agent needs for its next action, while maintaining durable records rather than relying on the prompt alone. The context window limits how much information the model can receive at once; it does not guarantee that every included fact will be used well. Lost in the Middle, for example, found that retrieval quality can depend on where information appears in a long input. The aim is to supply relevant, current evidence, not simply to fill the window.
This page concerns the working set for one task: what the model must see now and what can remain outside its window. Candidate discovery is covered by the retrieval pipeline; retaining and revising information between tasks is covered by agent memory.
Open full-size imageMove along the horizontal axis: the answer-bearing passage changes position within 20 documents, about 4,000 tokens in total. The vertical axis is question-answering accuracy, not retriever recall. In this experiment with GPT-3.5-Turbo-0613, middle positions perform worse than either end; the dashed line shows the closed-book baseline without documents. The result motivates testing evidence placement on the model you use, rather than treating this historical curve as a benchmark for every model.
What belongs in context
The first four layers may enter the prompt. Durable state stays outside it and is selected back in when needed. A transcript is a record of interaction, not the system of record for the project.
Select before compressing
Search for symbols, call sites, tests, definitions, and exact evidence before asking a model to summarize anything. A summary cannot recover material that was never retrieved, and repeated summaries can quietly turn a guess into a fact.
A practical input budget is:
There is no universal percentage for rules, evidence, and output. The next action should determine the packet. Local models may also hit KV-cache or memory limits before the advertised context limit becomes useful.
Example task packet
For a parser that fails on a split UTF-8 sequence, a useful packet might be:
contract:
write_scope: [src/parser.py, tests/test_parser.py]
success: targeted test and parser suite pass
state:
attempted: boundary check after byte slicing
result: still fails on a split multibyte code point
working_set:
- parser function and two callers
- three relevant tests
- exact traceback
open_question: should truncation operate on bytes or Unicode code points?
The entire repository, full test log, and every earlier hypothesis would be larger but not necessarily more helpful.
Ways to keep the working set small
Prompt caching works best when genuinely stable instructions and schemas stay unchanged. Exact savings depend on the provider, model, request shape, and cache policy; measure them rather than copying a headline percentage.
Large outputs should be stored in full and represented in context by a useful slice plus an artifact reference. Head-and-tail truncation is one option, not a guarantee that the important line survives.
Compression without laundering uncertainty
A summary should keep:
- the source or artifact behind a claim;
- whether the statement was observed, inferred, or merely proposed;
- failed attempts that prevent repetition;
- unresolved contradictions;
- the next decision, not every conversational turn.
Structured state is often easier to inspect than a prose recap, but it is not automatically correct. Update it atomically where possible and check it against the current files when resuming.
How to test a context strategy
Move the same evidence between the beginning, middle, and end. Add irrelevant and conflicting documents. Include a case where the right answer is “insufficient evidence.” Measure task success, missed constraints, attribution, latency, and token use.
Useful failure labels include context stuffing, recency bias, summary drift, retrieval bias, and hidden-state dependence. These names are diagnostic shortcuts, not reasons to add another framework. The remedy is usually to remove irrelevant material, retrieve better evidence, or move durable state out of the prompt.