When I extracted document retrieval into AI Library, I did not yet have a better retrieval system. I had an isolated evidence boundary, a benchmark and somewhere the next experiment could be compared. The first experiment was to put a map inside that boundary.
Isolation decides where discovery happens and what crosses back into the continuing work. It does not decide how the isolated worker finds evidence. The worker could still search the repository conventionally, open broad groups of documents and spend much of its own context discovering how the knowledge was organised.
The previous note, I kept rebuilding document retrieval, ended with that boundary measured and reusable. This one is only about the first retrieval layer placed inside it and the Luna and Sol comparisons used to decide whether that layer was ready to merge.
Isolation was the boundary, not the search method
I initially described the comparison too loosely, as though Isolation and reference-map retrieval were competing architectures. They were not. Both conditions used the same fresh evidence worker, the same admitted hand-off and the same downstream mission boundary.
The only intended difference was discovery:
| Condition | Discovery inside the worker | Mission hand-off |
|---|---|---|
| Isolation | Ordinary repository discovery | Validated evidence result |
| Isolation + reference map | Map-first routing, then selected Entries | Validated evidence result |
Preserving that distinction kept the earlier Isolation benchmark meaningful. A discovery layer could be added, removed or changed without quietly redefining what the isolation result had measured.
The map was allowed to route, not answer
AI Library stored accepted project knowledge as canonical Books. Each Book contained independently addressable Entries with stable identity, status, relationships and an exact content revision. The Markdown inside an Entry remained the authoritative human-readable source.
The reference map was deliberately weaker. It was generated deterministically from those Books and contained titles, summaries, headings, bounded search terms, relationships and canonical locators. It could help a worker decide which Entries to inspect, but it could not become evidence by itself.
Before a query, the runtime rebuilt the expected map from the current Books and compared it byte for byte with the stored one. A missing, stale or altered map blocked retrieval. After selection, the worker still had to read the canonical Entries, look for contrary material and return exact source revisions. Admission checked those claims again before the result crossed the isolation boundary.
The map selected where to look. It never acquired the authority to say what the project meant.
The first benchmark kept finding ways to cheat
A clean comparison required more than giving one worker a map. Early attempts produced numbers that looked valid while the worker had reached the answer by another route.
In one run, the map worker queried the new layer and then read the generated documentation projection directly. In another, it inferred the source repository from the benchmark artefact path and enumerated the Book Entries before searching. Later, it opened runtime source to discover the command shape. None of those runs measured map-first retrieval, even when the answer itself was correct.
The benchmark changed with each failure. Model executions moved into disposable fixtures outside the source checkout. The Docs projection was withheld during retrieval and restored only for scoring. The worker received a prepared, revision-bound query request and its exact invocation. A protocol audit rejected repository enumeration, raw-map inspection, direct Docs reads, runtime-source inspection and escape back to the development checkout.
The audit itself also failed once by treating two commands separated by a semicolon as one prohibited source read. That false rejection became a tested command-boundary rule. Without these controls, I could not tell whether I was measuring the retrieval layer or the model's ability to route around it.
Luna made the result look like a regression
The first accepted pair used GPT-5.6 Luna at medium reasoning on the small release-readiness case. Both arms ran from the same repository revision. Both produced admitted evidence and complete deterministic coverage.
| Metric | Isolation | + Reference map | Change |
|---|---|---|---|
| Elapsed time | 83.3 s | 79.0 s | 5.2% lower |
| Total input | 110,145 tokens | 129,081 tokens | 17.2% higher |
| Fresh input | 26,689 tokens | 47,417 tokens | 77.7% higher |
| Deterministic coverage | Complete | Complete | No measured loss |
One pair is descriptive evidence, not a stable estimate. The map condition queried 10 of the 12 available Entries, so its intended lazy loading was only weakly selective in this run.
If I had reduced the experiment to answer quality and elapsed time, the map would have looked successful. If I had reduced it to token use, it would have looked plainly worse. Keeping both views showed the actual state: the capability worked, but this run did not support the context-reduction goal.
Sol gave the same layer a different answer
I then ran the same small case with GPT-5.6 Sol at medium reasoning, pinned to the final branch revision. This was again one matched pair, with ordinary Isolation first and the reference-map condition second.
| Metric | Isolation | + Reference map | Change |
|---|---|---|---|
| Elapsed time | 109.9 s | 89.6 s | 18.5% lower |
| Total input | 226,664 tokens | 163,499 tokens | 27.9% lower |
| Fresh input | 51,816 tokens | 48,555 tokens | 6.3% lower |
| Retrieval input | 212,452 tokens | 149,302 tokens | 29.7% lower |
| Tool-result material | 64,935 characters | 29,520 characters | 54.5% lower |
| Required-point recall | 0.75 | 1.00 | One missed point recovered |
Both results were admitted. Ordinary Isolation missed the unresolved-authority point and failed a deterministic quality gate. The map condition achieved complete source, point and conflict recall with no failed deterministic gate. Its semantic quality remained unassessed until a blinded review could be applied.
On Sol, the map condition was faster, used less input, returned less tool material and recovered a required point that ordinary discovery missed. It was the kind of result I had hoped the Luna run would produce.
It was still one observation. Execution order was not balanced across repetitions, and two models producing different cost patterns did not explain the cause. The reasonable conclusion was narrower: the reference-map path could improve retrieval, but its effect was model-sensitive in the evidence available.
The benchmark did not need to choose a winner
The branch had implemented canonical Books, a verified derived map, map-first candidate discovery, lazy Entry inspection and a benchmark capable of refusing invalid routes. The Sol run showed a meaningful benefit. The Luna run prevented that benefit from becoming a general claim.
That was enough to merge the capability, but not enough to make it the default or call it optimised. Isolation remained the baseline. Reference-map discovery remained an explicit layer inside it. The accepted position was that performance was still under evaluation and any broader claim needed balanced repetitions across more cases.
This is where this part of the work ends: AI Library had its first real retrieval layer and two qualified measurements of it. What happened when I tried to make that layer more selective belongs to the next experiment, not this one.
The first benchmark did not tell me that the map was better. It told me the map was real enough to test—and that the model was part of the retrieval system I was testing.