I thought I had already corrected this mistake.

I had stopped treating the size of one prompt as the cost of a documentation retrieval system. Search could make the main task look smaller while moving tokens, latency and missed evidence into another context. So I began measuring the whole route: discovery, hand-off, recovery and the mission that followed.

That was the argument behind The context compression frontier. Evidence quality came first, then aggregate tokens across discovery and the mission, retrieval depth, end-to-end latency and model cost. The improvement was real. A local saving could no longer hide work moved elsewhere.

But I had replaced one incomplete measure with a total that flattened an important difference. Some context belonged to a temporary evidence job. Some belonged to the mission that would continue through planning, implementation and review. I added their tokens together as though their location and lifetime did not matter.

The metric contained an architecture

The original retrieval design used a separate agent. It could search widely, identify the authoritative evidence and return a compact hand-off. The main task would keep its context for the work it had actually been asked to do.

An early controlled replay challenged that design. Isolation reduced the retrieval material retained by the main task, but the complete run used more estimated tokens, took longer and found fewer required sources. The isolated design was undone. That was a defensible response to the evidence available at the time.

The evaluation rule still carried two assumptions. First, the isolated agent owned only a bounded search pass, not the complete evidence job the downstream task needed. Second, each enabling component had to beat the complete incumbent on its own. Isolation, a complete-job contract, a capable worker and a validated hand-off could each be rejected before their interaction was tested.

Aggregate tokens then became more than a diagnostic. They became an architectural preference. A raw prompt used one context. Any orchestrated alternative paid for another context, so the additional setup and reasoning counted against it before the boundary could demonstrate why it existed.

I had designed a benchmark that could measure orchestration, but was predisposed to reward its absence.

The context that ends is different

At the same time, I was developing a broader theory in What if the agent is not the unit of work?. The goals, dependencies, evidence and accepted state should persist. The agents performing bounded parts of that work can be temporary.

The retrieval experiment exposed the measurement consequence. If the agent is temporary, its working context is not automatically continuing system state. It can investigate, compare, verify and then end. What should survive is the result the next stage needs: the requested answer, exact sources, material conflicts, checks performed and unresolved gaps.

Two contexts can use tokens without imposing the same continuing burden
BoundaryTemporary evidence workerContinuing mission
ResponsibilityDiscover, challenge and verify evidenceDecide and perform the requested work
LifetimeEnds when the evidence job is completeContinues across later iterations
Useful outputMinimum-sufficient evidence hand-offAccepted decisions and changed system state
Material that should not persistSearch trace, rejected candidates and process historyNothing the mission still needs

Context retained by the mission remains present as that work continues. It occupies capacity before the next decision even begins. In this workflow, later requests continue from that accumulated state until material is removed or compressed. A cached token may be cheaper to process, but caching does not make irrelevant material useful or restore the context-window headroom it occupies.

The worker's context has a different shape. It may be large—sometimes it should be—but it terminates at the boundary. Only the hand-off becomes part of the mission's future. The burden is not invisible; it still appears in aggregate tokens, elapsed time and provider spend. It simply does not need to be carried through every later stage of the work.

A case that could not fit inside the limit

The next experiment gave the isolated worker a complete evidence responsibility. It received the downstream task boundary, built a ledger of what the task needed to know, searched for contrary evidence and verified implementation claims before declaring coverage complete.

Numeric limits became page sizes rather than definitions of completeness. This mattered because the test case deliberately required nine independently owned parts of the system. Earlier retrieval allowed six evidence sections in a final packet and exposed eight candidates by default. A required ninth source could therefore disappear before the agent had a chance to decide whether it mattered.

Bounded transfer is useful. Silent terminal truncation is not. If nine sources are genuinely necessary, nine is the irreducible evidence floor. The boundary should remove duplicate explanation and discarded process history, not facts required by the next decision.

What the composed experiment found

The final controlled comparison used the same broad planning case and the same top-capability model—gpt-5.6-sol—for both architectures. Each arm ran three times, with the order alternated. The decision rule was fixed before those runs.

The evidence worker had to cover all nine required source groups in every repetition. Its returned result and the downstream task both had to pass separate semantic checks. Its median task-session peak also had to improve beyond the measured variation within the two arms.

Controlled comparison: the evidence remained complete while the continuing task carried substantially less context
ArchitectureRequired-source coverageTask peak contextSemantic checks
Single context · run 18 of 969,495 tokensPassed
Single context · run 29 of 962,592 tokensPassed
Single context · run 38 of 960,024 tokensPassed
Evidence worker · run 19 of 917,736 tokensPassed
Evidence worker · run 29 of 918,409 tokensPassed
Evidence worker · run 39 of 918,290 tokensPassed

Median continuing-task peak fell from 62,592 to 18,290 tokens: a 70.78% reduction. The reduction was larger than the measured within-arm variation. Every isolated repetition covered all nine source groups, and both the hand-off and the downstream task passed their fixed semantic checks.

The isolated workers were not small. Their own peak contexts ranged from 68,654 to 72,022 tokens, and the composed system used more aggregate billed tokens and wall time than the single-context baseline. Those are real costs and the next optimisation surface. They did not change the result the architecture was introduced to produce: complete usable evidence with a materially leaner continuing mission.

What the result supports—and what it does not

The experiment supports one part of the unit-of-work theory. A temporary agent can own a coherent, independently verifiable job, return durable evidence and disappear without forcing the mission to inherit its entire working history. Persistence can belong to the work rather than the worker.

It also challenges the easy version of that theory. Temporary agents are not free. A weak or incomplete brief can isolate the wrong work. A lossy summary can make the boundary look lean by hiding what was lost. The worker may need a more capable model, deeper investigation and executable validation before its result is safe to carry forward.

This does not prove that smaller retained context makes every model reason better. It does not quantify hallucination, drift or what a later compaction would have removed. Those are not required for the narrower system finding. The experiment asked whether complete evidence could reach the mission while the mission retained substantially less context. In this case, it could.

Preserved headroom is still a meaningful architectural result. The system does not have to wait for accumulated search history to become a visible failure before avoiding it. Less irrelevant context at the start means more capacity remains for the work and iterations that follow. The exact downstream benefit will depend on the model, mission length and what the system later needs to remember.

What I would measure now

I have not stopped measuring the whole system. I have stopped asking one total to make every decision.

The revised hierarchy is:

  1. Accuracy is a constraint. The evidence must be right, complete, correctly labelled and honest about gaps.
  2. Effectiveness is a gate. The result and sources must let the mission verify, correct, review or continue the work.
  3. Mission context measures leanness. Peak retained context in the continuing task and the size of the minimum-sufficient hand-off show what the work must carry forward.
  4. Worker and aggregate costs remain visible.Tokens, latency, retries and spend show where the temporary component should improve. A genuine affordability ceiling can still block adoption when it is declared in advance.
  5. Composition must be tested as composition.Components establish their intended effects separately; the combined architecture must show that those effects coexist.

This is not permission to hide an expensive worker behind a clean hand-off. It is a decision about where each measure has authority. Aggregate tokens describe resource consumption. They do not, on their own, describe how much context the continuing work has been forced to remember.

The system carried the cost forward

The mistake was not measuring tokens. It was assuming every token had the same architectural consequence.

When discovery happened inside the mission, the mission kept the routes, rejected candidates, verification steps and source material after discovery had finished. The search was over, but its context continued into the work. I had made the system carry a completed process forward.

The better boundary was not the smallest packet a numeric limit would allow. It was the smallest hand-off that remained sufficient for the next responsible decision. Finding that boundary required a worker capable of doing more, so the mission could safely retain less.

I had begun with a theory that the agent might not be the unit of work. The experiment did not prove that theory in general. It gave me one concrete system in which the theory held: the agent stopped, the useful evidence persisted and the mission continued without carrying the agent's whole journey.

Further reading