The first problem was not prompt size. It was that a documentation-first system had made discovery part of almost every meaningful task.
The documentation held more than instructions. It distinguished intended design from current behaviour, recorded decisions and proposals, and told an agent which sources carried authority. Before changing the system, an agent needed to find the right parts and understand how they related.
That made the work more traceable, but it also made poor discovery expensive. Search too narrowly and the agent missed a qualification or authoritative source. Search broadly and the routes, candidate files, rejected material and search history remained in the same context as the work itself. The cost did not end when discovery ended. It was carried through every later reasoning and implementation step.
A conversation about the cost of agentic workflows made the problem more concrete. Some individual runs were costing more than $2. That number depended on the model, tools and length of the task; it was not a benchmark, and it did not come from the retrieval experiments in this note. What mattered was how quickly a modest per-run cost could compound when discovery triggered more calls, larger contexts and retries.
I had also been experimenting more generally with when to use sub-agents. They can be valuable when a bounded piece of work can be isolated and returned as a useful hand-off. They can be wasteful when they repeat the same reading, discard context the next agent needs or insert another serial step without reducing the work around it.
Documentation discovery sat directly on that boundary. Should the main agent search and continue with everything it had accumulated, or should a separate agent search widely and return only the evidence needed for the mission?
The goal was fast, accurate and affordable discovery across local files and documentation. I was not trying to make the smallest prompt. I was trying to stop discovery from consuming the context needed for the actual task without making the complete system slower, less reliable or more expensive.
Separate discovery from the mission
My first change was therefore architectural rather than a cleverer search. A focused sub-agent could inspect the wider source collection, then hand the main agent a small packet identifying the authoritative sections it should read.
The original discovery prompt and sub-agent implementation were not preserved well enough to reproduce. The first comparison was captured, however. Both arms received the same read-only planning task in fresh contexts.
| Measure | Raw discovery | Separate discovery |
|---|---|---|
| Documentation read | 25 files | 3 sections plus instructions |
| Lines retained | 1,735 | 125 |
| Words retained | 17,954 | 1,984 |
| Rough token proxy | ~33,400 | ~4,200 |
The 88–89% reduction was large enough to justify a proper experiment. It was not evidence that the complete system was cheaper or better. It measured the main context, not the discovery work that produced the hand-off. It also did not measure evidence recall or comparable end-to-end latency.
The controlled replay exposed the difference. Separating discovery reduced the median retrieval material passed into the main task by 59.8%, but the complete system used more estimated tokens, took longer and found fewer of the required sources. Protecting the mission context was useful. The first way I did it simply moved too much cost and risk into another part of the system.
The result that first gave this note its name came from a later candidate. It used 30% fewer median aggregate estimated tokens, but follow-up searches rose from six to thirteen, recall fell slightly and the candidate treated incomplete evidence as sufficient.
It looked like a clean turning point: compress the context beyond some minimum, and the system has to recover what was removed. I started thinking of that boundary as the context compression frontier—the point between removing noise and removing information the system still needs.
The later experiments made that explanation less tidy. Another candidate reduced both tokens and follow-up searches, yet lost far more of the required evidence. A third improved recall while making discovery slower. A fourth lowered the median token estimate while increasing the total, the median latency and the slowest runs.
I still find the frontier useful as a working idea, but not as a line I have located. The more durable result was that token use, retrieval depth, latency and evidence quality could move independently. The boundary appears to depend on the task, model, retrieval method and measure being observed.
What the later tests changed
The main experiment suites used eleven or twelve documentation questions. Some had one clear source. Others crossed several areas, used different terminology from the documentation, contained conflicting material or required a claim to be checked against the implementation. Required sources were defined before execution, and the later candidates ran three times in fresh contexts.
The suites and measurement boundaries evolved between phases, so these are not points on one controlled curve. I use them as repeated examples of the measures separating, not as a pooled benchmark.
| Candidate change | What improved | What it cost elsewhere |
|---|---|---|
| Compact document map | Median aggregate tokens fell 30% | Follow-ups rose from 6 to 13; recall fell; unsupported and false-complete claims appeared |
| Structured section index | Median aggregate tokens fell 25.2%; follow-ups fell from 6 to 2 | Recall fell from 96.4% to 77.4–79.8%; unsupported claims and false-complete packets increased |
| Relationship expansion | Recall rose to 96.4%; false-complete packets fell; median tokens fell 5.1% | Final packets still missed needed relationships; total tokens rose 0.9%, follow-ups 60% and discovery latency 17.5% |
| Reviewed fast-path order | Median aggregate tokens fell 9.2% | Total tokens rose 0.3%; median end-to-end latency rose 3%, p95 rose 49.8%, and quality was unstable |
The structured section index changed my interpretation most. It reduced the initial text and produced fewer follow-up searches, but it often isolated the best-matching section from a required sibling or surrounding qualification. The system did not always recover through another search. Sometimes it simply stopped with an incomplete evidence set.
Relationship expansion exposed a different gap. It made broad recall better and reduced false-complete hand-offs, but visibility in discovery did not guarantee that the final packet admitted the relationship the question needed. The mechanism genuinely improved what the system could see; the improvement stopped at the hand-off.
The fast-path test exposed a measurement problem rather than a compression mechanism. Its median token estimate improved while the total increased slightly. Median latency moved in the wrong direction and p95 moved much further. A single median would have made the candidate look cheaper while hiding the cost in the distribution and the unstable quality between repetitions.
Caching provided the cleanest local speed improvement. Warm candidate discovery was about 70% faster and the tested invalidation rules returned no stale evidence. Yet aggregate measured text rose, while whole-run p95 latency also rose. A component had become faster; the complete path had not produced a clean win.
No single change improved every measure at once. Each moved real cost or risk into a different part of the system — which is the pattern this note exists to describe, not a verdict on any one mechanism.
Sometimes compression creates recovery work. Sometimes the system fails to notice what is missing. The common failure is allowing one favourable measure to stand in for the complete system.
More context is not the same as more understanding
The original instinct to remove context was not irrational. Long context windows tell us how much a model can accept, not how reliably it will use every part of what it receives.
In controlled multi-document question-answering and key-value retrieval experiments, Liu and colleagues found that model performance often fell when relevant information appeared in the middle of a long context, including for models designed to accept long inputs ( Liu and colleagues, 2024). Those tasks are narrower than an agent searching technical documentation, but they support the practical concern: adding more material can make evidence available without making it equally usable.
Compression can also genuinely help. LongLLMLingua compressed prompts across several long-context benchmarks and reported improvements in task performance as well as reductions in cost and latency at the tested compression ratios ( Jiang and colleagues, 2024). That is useful counterevidence to any simple claim that smaller context must be worse. The method, models and benchmark tasks differ from mine, which is precisely the point: the frontier is likely to depend on what is removed, what the task needs and how well the remaining context exposes it.
My results sit between those findings. Too much context can bury the evidence. Too little can remove the route or relationship needed to find it. The aim is not minimal context. It is sufficient, usable context.
More searches are not automatically waste
I also need to be careful about treating iteration count as a quality measure.
For genuinely multi-step questions, later searches may depend on facts discovered earlier. In experiments across four multi-hop question-answering datasets, Trivedi and colleagues found that interleaving retrieval with reasoning improved both retrieval and downstream answers compared with a single retrieve-and-read step ( Trivedi and colleagues, 2023). Those are benchmark questions rather than software work, but they establish an important boundary: iterative retrieval can be the right design when the question itself unfolds in stages.
The distinction I care about is therefore not simply one search versus many. It is whether each additional search resolves a necessary dependency or recovers information that an earlier compression step unnecessarily removed.
The experiments do not establish a clean causal latency curve between serial loops and elapsed time. Several early phases had incompatible or replay-contaminated timing boundaries. The later fast-path test did record comparable end-to-end timings and found a 3% median regression alongside a 49.8% p95 regression, but that does not show that search count caused the delay. Model and host variance remained, and the candidate changed route order rather than iteration depth alone. Time only became a first-class concern after these differences were impossible to ignore.
The practical conclusion is narrower: measure end-to-end latency directly. Do not infer it from prompt size, aggregate tokens or the number of tool calls. They describe different things.
The frontier belongs to the whole system
The Theory of Constraints gave me a useful way to think about what had happened. Improving one stage does not necessarily improve the flow of the whole system. It may simply move the constraint somewhere less visible.
I had optimised the hand-off to the main agent. The cost moved into discovery and recovery. Then I optimised aggregate token use. The pressure appeared in serial iteration and evidence quality.
There are at least two context boundaries here. The mission context contains what the main agent needs to complete the task. The discovery context contains what the retrieval layer needs to find, relate and qualify the evidence. Separating them can prevent broad search from overwhelming the mission. Compressing either one too far can still damage the complete result.
Model capability may move those boundaries. A stronger model might reconstruct a missing relationship or choose a better recovery search; a smaller one may fail earlier. I did not control model tier, so that remains a hypothesis. Task shape and consequence matter too. A reversible lookup can tolerate a different balance from a decision that depends on complete source coverage.
This is also why I did not assume that a vector or graph database was the answer. They may improve semantic or relationship discovery as the source collection grows, but they add indexing, invalidation, latency and operating costs of their own. I have since tested all three — lexical, vector and graph — and written up what each showed in Testing databases for AI documentation retrieval.
What I would measure now
For retrieval where the answer needs to be traceable, evidence quality is a gate. The system must find the required sources, distinguish current material from proposals, expose conflicts and avoid claiming completeness when important gaps remain.
After that, I would keep four measures separate:
- aggregate tokens across discovery and the mission
- dependent retrieval depth, not just the number of calls
- end-to-end latency
- the model capability and operating cost required to sustain the result
A subsequent experiment changed part of this hierarchy. I had treated aggregate tokens as though temporary discovery context and the context retained by a continuing mission imposed the same burden. In I measured the wrong context—and the system carried the cost forward, I test that assumption directly and explain why aggregate usage remains important without deciding the architecture on its own.
Their priority depends on the work. An interactive tool may care most about latency. A high-volume background process may care more about cost. A complex investigation may accept both for stronger coverage. The mistake is letting the easiest number to obtain stand in for the system.
I began by trying to stop discovery from consuming the context needed for the actual task. I ended up with a more useful question: what is the smallest hand-off that preserves enough information for the next part of the system to work well?
I do not yet know where that frontier sits, or whether it is useful to describe it as a single point. My working hypothesis is that compression helps while it removes noise and fails when it removes information the next stage still needs. The results show that failure does not have one symptom: it may appear as recovery work, silent incompleteness, weaker claim support or worse latency at the tail.
Token savings can continue while any of those things happen. That is what makes the result easy to misread.
Further reading
These books offer related lenses rather than direct proof of the experiment. I am including them as a reading list to examine, not as endorsements of every claim they contain.
- Thinking in Systems, Donella Meadows—feedback loops, delays and choosing a system boundary
- The Goal, Eliyahu M. Goldratt and Jeff Cox—constraints, flow and why local optimisation can fail to improve the whole
- Information Retrieval: Implementing and Evaluating Search Engines, Stefan Büttcher, Charles L. A. Clarke and Gordon V. Cormack—the foundations of search, indexing and retrieval evaluation