I tested three database-backed designs for AI documentation retrieval: a lexical full-text index, a vector store and a typed relationship graph. Each brought real strengths and real costs, and together the evidence says more about where retrieval actually breaks than about any single technology winning or losing.
The system they plugged into is the one described in The context compression frontier: compact maps, section indexes, relationship expansion and caching, with repository Markdown as the authoritative source. Each design changed only how candidate sources were discovered. A database could propose a source, but it could not decide that the source governed the question, label its status, resolve a conflict or make a claim supported.
How the three were tested
These were three SQLite-backed retrieval designs rather than a comparison of database products. One indexed words, one indexed semantic vectors and one stored explicit relationships between documents.
The cases asked documentation questions in fresh contexts, with the required evidence defined before execution. I tracked recall, input and output tokens, latency and answer quality as separate measures, so a gain on one axis could not quietly absorb a loss on another. What follows is what each design showed on those axes.
| Candidate | Strengths | Limits |
|---|---|---|
| SQLite FTS5 and BM25 | The best recall of any design tested, with roughly 10% less retrieval input | Ranked snippets arrived without surrounding context; answers built on them misread status and overclaimed |
| sqlite-vec and FastEmbed | Safe build and invalidation, aimed at a genuine vocabulary-mismatch failure | Its conservative trigger never fired, and a direct query still missed the terminology source |
| Persisted typed graph | Typed, provenance-complete relationships with safe lifecycle handling | At this corpus size, read-time link following already surfaced the same evidence |
The lexical index: wider discovery, thinner context
The lexical design replaced per-query Markdown scanning with a persisted section-level full-text index ranked by BM25. Its case is efficiency: discovery becomes cheaper and wider at once.
The evidence supports that case. It recovered 50 of 54 required evidence groups — the widest recall of any design I tested, and it surfaced sources that direct scanning missed. Median main retrieval input fell 10.4%, and median aggregate input fell 9.4%.
The cost showed up in what it handed downstream. Ranked snippets arrive stripped of the surrounding sections that carry qualifications and status. The agent working from them classified a proposal as current guidance, made six unsupported consequential claims and declared coverage complete when it was not. Median aggregate output rose 17.7% — more words resting on worse evidence.
The index did not write those claims, and blaming BM25 for them would be lazy. But the trade-off is real and it is structural: a lexical index widens what an agent can see while thinning the context each source arrives with. Adopting one means pairing it with a judgement step strong enough to absorb that wider, and thinner, candidate pool.
The vector store: right target, wrong trigger
Of the three, the vector store was aimed at the most genuine gap. One test case turned on a vocabulary mismatch: the question used different words from the source that answered it, and exact retrieval could not rank that source in its top eight. That is precisely the failure embeddings exist to fix.
The engineering held up. The index built and invalidated safely, and the design was deliberately conservative: exact retrieval ran first, and semantic search was allowed only when exact retrieval returned no candidates.
That trigger turned out to describe an event that does not happen. Exact search always returned something — eight candidates, just not always the right ones — so semantic search activated zero times in 27 observations. "No results" and "wrong results" are different failures, and I had built a detector for the wrong one. That is a finding about fallback design as much as about vectors: a fallback only earns its keep if its activation rule matches the failure it exists to catch.
A separate diagnostic sharpened the picture. Running the vector query directly, outside the trigger, it also failed to rank the required terminology source in its top eight. So the mechanism itself, not just its gating, was not yet solving the target case. The costs were measurable meanwhile: about 32 MB of index and 155.6 seconds of build time over the scale fixture.
The typed graph: durable structure ahead of need
The graph's case is that relationships between documents are real structure, and re-deriving them on every query is wasted work. Persist them once, typed and with provenance, and traversal becomes cheap and repeatable.
It kept every promise. Edges were extracted only from explicit phrases in the Markdown, every edge retained its provenance, and adds, edits, moves, deletes, cycles and stale state were all handled safely. Of the three designs, it was the most complete piece of engineering, and it made relationships visible earlier and more repeatably than read-time traversal does.
What the evidence could not show, at roughly one hundred documents, was that earlier visibility changing any outcome. The graph surfaced 42 documents across its runs; nine were admitted to the evidence packet, read and used; all 72 required evidence groups were recovered with or without it, because an agent could still follow the same explicit links at read time. The graph's value is a function of corpus size and relationship density, and this corpus had not yet crossed the threshold where that value becomes measurable.
The hand-off was the constraint
Retrieval sounds like one action. In this system it was a sequence:
- discover candidate documents
- decide which candidates are authoritative and relevant
- admit a bounded set to the evidence packet
- read the cited sections in their proper context
- use them without hiding conflicts or gaps
All three databases changed the first step. The hard failures lived later: a proposal mistaken for current guidance, an adjacent qualification missed, a visible source never admitted, an answer claiming more certainty than its evidence supported.
Candidate recall is not evidence recall. Evidence recall is not correct use.
This is why the graph result matters more than its query time. Forty-two surfaced documents became nine used documents. The binding constraint sat in the hand-off between visibility and judgement, and that constraint is indifferent to which database sits in front of it.
What would make each worth adopting
The decision cases came from a real corpus of roughly one hundred documents. Synthetic fixtures pushed the operational tests past eleven thousand files, but those tested build time, storage and invalidation — not the terminology, authority conflicts and relationship density of a real corpus at that size. At one hundred documents, an agent can still search Markdown, follow explicit links and read surrounding sections without a specialised data layer. At ten or one hundred times that, the picture may change. These experiments do not tell me where.
So each design has a condition under which its strengths start to pay. The lexical index pays when scanning cost is the measured pain and the judgement steps downstream have been hardened enough to absorb a wider candidate pool. The vector store pays when a repeatable vocabulary miss exists and a direct vector query demonstrably recovers it — the mechanism has to solve the case before it earns a place in the pipeline. The graph pays when relationships become implicit, dense or expensive enough that following them at read time starts to fail.
Start with the miss
I began these experiments with an architectural question: which database should sit behind documentation discovery?
The better question was earlier: what is the retrieval failure I am trying to change?
If exact search cannot bridge a vocabulary mismatch, semantic retrieval may help. If relationships are expensive to rebuild, a graph may help. If scanning is too slow, a lexical index may help. But a new data layer cannot repair a source that is found and then misclassified, excluded or used without its qualification.
All three mechanisms worked. The evidence just points somewhere else: this system needed to preserve the meaning of evidence after discovery more than it needed a better way to find candidates.