AI found 08:42. The answer had been right until the timetable changed.

This is a small thought experiment. Imagine that you are planning a train journey. The operator publishes an 08:42 departure, and you save a screenshot in a shared folder. A family member adds the time to a shared calendar and asks an assistant to prepare the itinerary. At that moment, the live timetable, screenshot, calendar and itinerary agree. Nothing is stale.

The screenshot still creates a question. Is it a convenient reference, a dated record of what was shown, or the place everyone should now maintain? If nothing on it answers that question, it looks more authoritative than it is.

Later, engineering work moves the train to 09:12. The operator updates its timetable. The screenshot, calendar and itinerary do not change. When a person or assistant looks for the departure, the old answer may arrive first as a complete and plausible set of results. The three artefacts are not independent confirmations. They are descendants of one old representation.

The practical problem is easy to miss because it did not begin when the times diverged. It began when the copied time lost its declared relationship to the source that governed it.

The copy problem starts before the values disagree

  1. 01 · CopiedSame content
    Governing source08:42
    Unlabelled copy08:42

    The values match; the copy's role is unclear.

  2. 02 · ChangedDifferent content
    Governing source09:12
    Unlabelled copy08:42

    The ambiguity has become a visible divergence.

  3. 03 · ReusedOne lineage
    ScreenshotCalendarItinerary

    Three artefacts agree because they share one source.

Accuracy and authority are different properties. The copy first loses a clear governing relationship; only later does its value become stale and seed derived work.
The content may not decay immediately. Certainty about its authority can.

The problem begins before the copy is stale

I have found “single source of truth” too easy to hear as “put everything in one database”. That is neither realistic nor especially useful. Different parts of a system can govern different things, and the same information often needs to appear in several places.

The distinction I care about is narrower. For consequential information, there should be a declared governing source for the current value within a defined scope. There should also be an accountable owner or process for settling changes and conflicts.

That authority might be a database record, an approved policy, a signed decision or a live service. It is not authoritative merely because of where it is stored. It is authoritative because the surrounding system says what it governs, for whom and during which period.

This leaves room for more than one valid account. A past policy can remain true about the past. Two prices can be correct for different regions. A contextual summary can deliberately explain the same evidence for a different audience. Scope and time are part of the claim.

An unmanaged copy is different. It independently restates the information without making its relationship to the governing source visible. Even while the values match, a reader may not know which one should change, which one should settle a disagreement or whether either is still current.

That is why the first loss is not necessarily accuracy. It is authority information.

Finding an answer is also a stopping decision

Most searches cannot continue indefinitely. We stop when the expected value of looking again feels lower than the cost, effort or delay.

In an experiment on sequential choice, Andrew Caplin, Mark Dean and Daniel Martin found that most participants searched until they reached a satisfactory threshold rather than inspecting every available option (Caplin, Dean and Martin, 2011). In a separate experiment involving 115 people completing three online search tasks, Glen Browne, Mitzi Pitts and James Wetherbe found that people used several different stopping rules depending on the task (Browne, Pitts and Wetherbe, 2007).

Neither study examined stale workplace documents, and neither says that people always accept the first plausible answer. They support a more modest point: stopping is part of information seeking, and the stopping rule changes with the task.

That matters when an old copy looks complete. A visible contradiction can prompt another search. A credible answer with no visible conflict can remove the reason to keep looking.

Outdated information does not need to defeat the governing source. In a consequential workflow, it may only need to be found and accepted first.

An AI system inherits what retrieval gives it

The AI case has a different mechanism. A model does not independently decide that it has searched enough in the same way a person does. Its retriever, tools, prompt, orchestration and limits decide which material reaches it and when the search ends.

If retrieval returns the old screenshot without its date or relationship to the live timetable, the model can produce a perfectly coherent itinerary from the wrong premise. Fluency cannot recover information the workflow failed to provide.

An EMNLP 2025 paper tested retrieval-augmented question answering across sources with different reliability and found that standard relevance-based retrieval remained vulnerable when source reliability was not represented. The authors' reliability-aware method improved results in their tested settings (Hwang and colleagues, 2025).

Their experiments concern multi-source question answering, not organisational document versions, so they do not prove the rule proposed here. They do show why relevance and authority should not be treated as synonyms. A document can be an excellent semantic match and still be the wrong source to govern the answer.

Agreement can share one ancestor

Once the screenshot has produced a calendar entry and itinerary, later searches may find several artefacts that all say 08:42. The volume looks reassuring.

It should not be counted as independent corroboration. Each artefact confirms consistency with its parent, not the truth of the original premise.

This is where provenance becomes more than administrative metadata. The W3C PROV model describes provenance as information about the entities, activities and people involved in producing an object. Its primer notes that provenance can help people judge whether information should be trusted and understand how something was generated (W3C PROV primer).

The standard does not nominate a governing source for every domain. What it gives us is a useful language for the missing relationship: origin, derivation, revision, time and responsibility.

The same kind of missing path appears in two nearby notes. In “Sometimes the breadcrumbs are the story”, we lose the path from a problem and its evidence to a polished decision. Here, the missing path belongs to the record itself: can a copied value still lead us back to the source that governs it now?

In “When AI velocity outruns feedback”, I looked at what happens next, after a weak decision becomes context for fast downstream work. If a representation does not reveal what it came from or what governs it now, neither a person nor an AI system has a sound basis for deciding how much authority to give it before that loop begins.

Copies are often useful

Removing every duplicate would be both impossible and undesirable.

Systems need replicas for availability, caches for speed, exports for exchange, backups for recovery and snapshots for audit or history. People need summaries, translations and views that make complex information usable. A copy becomes a problem because its governing relationship is absent or misleading, not simply because it exists.

Two questions help me separate useful representations from competing ones.

First, how does this information relate to the governing source?

  • Synchronised: retrieved from or updated under a declared contract with the governing source
  • Referenced: linked back without independently redefining the value
  • Snapshotted: preserved as an immutable record of a stated point in time

Second, what has happened to its meaning?

  • Represented: restated without intentionally changing the scope or interpretation
  • Contextualised: interpreted for a purpose whose audience, scope and limitations are visible

These are not exclusive boxes. A contextual view can be synchronised. A historical explanation can be based on a snapshot. The useful test is whether a reader can tell what the artefact is and what it is not.

Make the relationship easier to see

For low-consequence, short-lived information, an informal copy may be proportionate. The cost of managing every note can exceed the cost of occasionally checking it again.

The stronger discipline belongs where stale information could change a decision, affect other people or create difficult-to-reverse work. Before relying on such information, a person or system should be able to determine:

  • where it came from
  • what currently governs the value
  • the scope and period for which it applies
  • whether it is live, derived or historical
  • whether a newer or conflicting source exists

That does not require one universal registry or an elaborate governance platform. Sometimes the right design is a link instead of copied text. Sometimes it is a visible “as at” date, a version identifier, an owner, or a retrieval rule that prefers the governing source and surfaces conflicts.

The goal is not to make information harder to share. It is to let information travel without quietly changing its authority.

The old screenshot is harmless when it clearly says what it was: the timetable as seen at a particular moment. It becomes risky when it presents itself as the answer to a current question.

Outdated information may only need to be found and accepted first. Good provenance gives the reader a reason to look one step further.

Further reading

These books offer related perspectives rather than direct proof of the argument. I am including them as a reading list to examine, not as endorsements of every claim they contain.