I had built document retrieval more times than I could easily count. Each version belonged to its project, and each felt a little better than the last. Eventually I realised the useful part was not another search technique. It was the boundary that kept the search out of the work it was meant to support.

The pattern usually began innocently. A project accumulated enough documentation that asking an AI to read it all became wasteful. I added routes to the likely sources, a way to find candidates and rules for distinguishing current guidance from proposals, implementation records and historical evidence.

Then the retrieval grew responsibilities. It had to look for contrary evidence. It had to verify claims about the current code. It had to admit uncertainty instead of filling a gap with a plausible answer. Eventually it needed to return a compact, traceable result that another piece of work could safely use.

By then I no longer had a clever search prompt. I had a small information system hiding inside a repository.

The observation that kept returning

One change made a larger practical difference than I expected: I moved the complete retrieval job into a fresh context.

The worker could search, reject candidates, follow links, read whole sections and check the implementation as deeply as the question required. The continuing mission received the usable finding and its provenance. Most candidate search, rejected routes and intermediate analysis ended with the worker.

The measured reason for doing this was context. Retrieval could consume a large amount of working material without leaving all of it in the conversation that still had to plan, change and review the software.

An earlier controlled comparison in DeployOptix supported that narrow claim: isolation reduced the context retained by the continuing task without losing the evidence required by that case. It was enough to adopt the boundary there. It was not yet the baseline I wanted for a reusable system.

The more interesting effect was something I had not yet measured cleanly. The work felt better.

I appeared to need fewer rounds of refinement. Bugs were easier to notice. Drift between the request, the evidence and the implementation seemed to stand out earlier. With less working material competing for attention, the decisions still to be made appeared more distinct.

Those are observations from my work, not a controlled result. I had not recorded every iteration, correction or defect under comparable conditions. A stronger model, a clearer task or my own growing familiarity could also explain some of the change. The controlled comparison measured retained context and evidence sufficiency; it did not measure refinement count, bug detection or drift.

The material left behind was not just noise

There is an important cost hidden inside the cleaner result. Raw documentation does more than answer explicit questions.

Across many files, an AI can encounter the same decision made repeatedly, see which exceptions attract warnings, notice the terms a project uses consistently and infer what its authors treat as normal practice. No single paragraph has to declare "this is our common knowledge" for that pattern to influence the work.

A compact evidence result can preserve every fact named in a benchmark and still lose that distributed signal. The mission may know which rule applies without acquiring the broader feel for why the repository regards that rule as ordinary. On an unfamiliar task, those weak signals may help it infer a sensible default or recognise that a technically valid change is out of character.

Calling all of that material "retrieval debris" would therefore be a straw man. Search traces, duplicate reasoning and rejected candidates may have little continuing value. Source bodies, recurring choices and adjacent explanations can carry useful knowledge that the retrieval question did not know to request.

The controlled comparison did not test whether isolation preserved that kind of inferred practice. Its oracle described explicit evidence needs and judged the returned result and downstream answer against them. Passing those gates established sufficiency for that task, not that the isolated mission learnt everything useful the single-context mission could have inferred from the surrounding corpus.

The real design problem is not how to remove as much context as possible. It is how to stop incidental retrieval work from crowding the mission while retaining the explicit evidence and the ambient patterns the mission will need to exercise good judgement. I had evidence for the first half of that boundary. The second remained an open question.

Cleaner context did not merely look cheaper. It seemed to make the remaining work easier to see—but anything useful removed with it had become harder to see at all.

That distinction began to bother me. If the effect mattered, "this feels better" was not a durable way to improve it.

The local system worked

DeployOptix became the fullest version of the project-specific approach. Repository Markdown remained authoritative. Small routing documents narrowed the search. A helper found likely candidates without loading a complete index into the mission. One isolated evidence worker owned discovery, source reading, contrary-evidence search and implementation checks.

It returned a bounded evidence packet. The packet had to name the needs it covered, the exact sources that supported them, the checks it performed, the conflicts it found and the gaps it could not close. A deterministic validator rejected malformed or incomplete hand-offs before the mission could use them.

This was not the crude old version waiting for a proper rewrite. It was useful, tested and increasingly careful. It also belonged completely to DeployOptix.

Its instructions assumed that repository's documentation structure. Its terminology described that product. Its runtime, contracts and tests evolved beside unrelated application code. When another project needed the same capability, I could copy the pattern, but the new copy immediately began a separate life.

Copying was not reuse

Copying felt efficient because I was no longer starting from a blank page. It was still a weak form of reuse.

A correction made in one project did not reach the others. A benchmark improvement belonged to the repository that ran it. Contract changes had to be rediscovered or transferred by hand. Similar capabilities slowly acquired different meanings while continuing to look alike.

The repeated code was only the visible duplication. The deeper duplication was learning. Each project could become better at retrieving its own documents, but there was no single thing being maintained and no stable history against which to judge improvement.

I wanted to be able to install the capability, keep the project's knowledge with the project and improve the shared machinery without quietly changing what that knowledge meant.

Finding the boundary worth extracting

My first instinct was to extract the retrieval skill. That was too small.

The behaviour I valued depended on several things moving together: authoritative knowledge, compact discovery, isolated analysis, a strict hand-off and an executable gate between the result and the continuing mission. Packaging only the prompt would have preserved the visible instruction while leaving its guarantees behind.

AI Library became the place for that larger boundary.

It stores accepted project knowledge as revisioned Books made from independently addressable Entries. A deterministic reference map gives retrieval a compact surface, but it remains derived: the worker must inspect the selected Entries before using them as evidence. Plans and receipts make changes to canonical knowledge explicit. Documentation can be projected elsewhere without the projection silently becoming the source of truth.

The project installing the Library still owns its knowledge. The installer owns the runtime, contracts and provider adapters. That separation matters because maintaining the retrieval machinery should not grant it authority to rewrite a project's intent.

What changed when the boundary moved
Project-specific systemReusable Library
Repository Markdown carries authority directlyRevisioned Books preserve accepted knowledge
Local routers and candidate searchVerified reference map and lazy Entry loading
Repository-owned evidence contractInstallable runtime and provider-neutral contracts
Improvements remain localOne maintained capability can serve several projects

The benchmark became part of the product

Extracting a system creates a dangerous moment. The new repository is cleaner. The concepts have names. Installation works. Contracts validate. The architecture feels more deliberate. None of that proves the retrieval is better.

This time I did not want the next improvement to rest mainly on remembered frustration or a satisfying run. AI Library carries a benchmark laboratory beside the capability under development. Cases define the required sources, points and conflicts before a run. Results keep quality, context, tokens, latency, tool output and hand-off size separate rather than allowing one attractive number to stand for the whole system.

The canonical Books and reference-map branch established the starting line. Its job was not to announce that the new system had won. It made the knowledge representation and discovery path concrete, pinned and comparable so later work could show whether an apparent improvement survived contact with the same task.

The starting line in numbers

Before changing discovery, I ran the existing isolated boundary and a single-context counterpart through a Sol benchmark at medium reasoning. The matrix produced 48 valid run records: four repetitions of each arm across six cases, including the same release-readiness problem with clearly irrelevant, adjacent and near-topic distractors.

This was a baseline, not a verdict. Nineteen runs failed a deterministic quality gate. The other 29 still needed blinded semantic review. None was decision-ready. Those limitations are part of the result rather than a footnote to remove later.

Mission input stayed almost flat while the single-context route grew with harder distractors
Distractor bandSingle-context missionIsolated missionReduction
Small · clearly off-topic104,614 tokens13,918 tokens86.7%
Medium · adjacent engineering148,222 tokens13,880 tokens90.6%
Large · near-topic247,098 tokens13,885 tokens94.4%

Values are medians across four runs per arm. The isolated ranges were 13,830–14,304, 13,847–13,915 and 13,818–14,153 tokens from small to large. The single-context large band ranged from 152,331 to 357,594.

Codex did not expose maximum live context for these runs. "Mission input" is the provider-reported input consumed by the final mission context, not a snapshot of its context window. It is the continuing-mission burden this benchmark can compare consistently, and it is the measure I will carry into later parts of the series unless better telemetry replaces it.

Protecting the mission moved work; it did not remove it
Distractor bandSingle total inputIsolated total inputSingle timeIsolated time
Small104,614218,40156.3 s128.7 s
Medium148,222352,57351.7 s147.4 s
Large247,098423,66470.0 s118.0 s

Total input includes every model context in the arm. In every band, the isolated route protected the final mission by using more input overall and taking longer.

Across small, medium and large bands, the hand-off remained near 1,500 estimated tokens. The isolated mission made no tool calls and received no tool-result material; the retrieval work ended on the other side of the boundary.
The resource pattern was clearer than the quality verdict
Distractor bandIsolated outcomesSingle-context outcomes
Small2 failed · 2 awaiting review2 failed · 2 awaiting review
Medium3 failed · 1 awaiting review2 failed · 2 awaiting review
Large1 failed · 3 awaiting review4 failed

"Failed" means a deterministic gate failed. "Awaiting review" means deterministic checks did not settle semantic quality. Neither label means the run passed end to end.

This is the comparison shape I want to keep: mission protection, total-system cost, boundary size, quality and evidential status. A later improvement should have to move through all five views. It should not get to choose the one number that makes it look successful.

That is as far as this part of the story goes. The attempts to improve that starting point belong to the next one.

A shared library raises the cost of being wrong

There is a reasonable argument for leaving this machinery inside each project. Local systems can fit local terminology. Their changes affect one repository. A small project may need little more than careful search and a few clear documents.

A reusable library can turn one mistaken assumption into a distributed defect. It can also create ceremony around a problem that was previously solved well enough with a file and a command. Reuse is not automatically simplification.

That is why the extracted boundary includes explicit non-goals. The Library does not get to originate human intent. Its reference map is not truth. Its evidence result is not truth. Installing it does not make every project's knowledge model identical.

The shared part should be the integrity of the route: how knowledge is identified, revised, discovered, inspected and handed back, including a way to challenge whether a compact hand-off preserved enough of the surrounding practice. The meaning travelling through that route remains the project's.

Somewhere the learning can accumulate

I began because I was tired of rebuilding document retrieval. That explanation is true, but incomplete.

I could have copied the latest version again. The stronger reason to extract it was that I had observed something I cared about and could not yet explain confidently. Keeping retrieval work out of the continuing mission appeared to reduce more than context. It seemed to reduce the friction of seeing what was wrong.

AI Library does not prove that observation. It gives me a stable capability, an inspectable boundary and a benchmark from which to investigate both what cleaner context removes and what it must retain.

Reusable code means another project can inherit what already works. A maintained benchmark means it may also inherit what I learn next—and the evidence that the change helped rather than merely feeling better.

I did not extract the system because I had finished improving it. I extracted it so the improvement would have somewhere to accumulate.