I introduced a rule in the interest of quality: work should always be tested against the latest main branch before it was accepted.

The reasoning seemed difficult to dispute. A pull request had been developed against an earlier version of the system. If the shared system had changed, the responsible thing was to update the branch and establish that its code, tests and review still applied.

At low volume, this looked like caution. At scale, it became a race.

Several pull requests would approach completion in parallel. Each had been reviewed and had accumulated the evidence appropriate to its change. One would merge first. main had now moved, so the others were updated and their evidence was treated as stale. Their checks started again.

While those checks ran, another pull request would finish and merge. main would move again. The remaining work would restart again.

The race was not fair. A small change with a short validation path was more likely to finish first. The largest changes, and the ones requiring the widest or most consequential assurance, spent longer exposed to another merge. The work that needed the most evidence therefore had more opportunities for that evidence to be invalidated.

I had built a safeguard that repeatedly charged the highest assurance cost to the work already paying the most for assurance.

The safeguard was enforcing the wrong fact

The system could prove that main had moved. It could not prove that the movement mattered to the claim under review.

Those are different facts.

Suppose a pull request changes the presentation of a static note. While it is being reviewed, another pull request changes an unrelated Worker diagnostic. The base revision is now different, but the assumptions supporting the note change may be untouched. Repeating every check does not necessarily add assurance. It may only reproduce evidence we already had at considerable cost.

The opposite mistake is easy too. Git may report that two changes merge cleanly because they edit different lines or different files. They can still interact through an API contract, dependency, configuration value, database migration, build tool or runtime assumption. Absence of a textual conflict is not evidence of absence of an interaction.

So neither of the simple rules is sufficient:

  • retest everything because main moved
  • accept everything because Git found no conflict

The question is not whether the repository changed. It is whether the change can reach the boundary that the existing evidence was meant to support.

I could see the churn more clearly than its cost

I examined the recent pull request history in the repository where this was happening. Of the latest 100 pull requests, nineteen had force-push events. There were twenty such events in total. Ten preserved the rewritten tip commit's stable patch identity, and four preserved the complete pull request patch against main.

That does not mean every one of those events was waste. Some rebases may have resolved meaningful interactions. Some checks may have been cheap or may have found something important.

One case made the mechanism concrete. A comparison using range-diff and patch-id established that the change itself was identical after a rebase. Exact-head assurance and multiple validators were nevertheless run again. The repository revision had changed, so evidence about an unchanged patch lost its authority.

The available history has an important limit. Local validation runs were not recorded centrally with their cause, duration and token use. I can establish that branch and evidence churn occurred. I cannot honestly turn the frustration I experienced into a precise total cost after the fact.

That missing measurement is now part of the correction. A system that makes repeated work invisible cannot tell us whether its caution is proportionate.

I had made the commit the unit of evidence

This exposed a contradiction in another idea I have been developing.

In What if the agent is not the unit of work?, I argued that goals, dependencies, evidence and accepted state should persist. The agents performing the next bounded operation can be temporary. Continuity should belong to the work rather than to a worker and its conversation.

But my validation process did not allow evidence to persist with the work. It attached authority to an exact repository state, then discarded that authority whenever the state moved—even when the relevant work did not.

There are good reasons to record exact revisions. Evidence without provenance is difficult to audit. If the pull request's patch changes, checks against the previous patch cannot silently stand in for it. If conflict resolution alters the code, that is new authoring and deserves fresh applicable evidence.

The mistake was treating provenance as the complete validity rule.

Evidence has assumptions as well as an origin. A unit test may support a narrow behaviour. An integration test may support an interaction between two boundaries. A review may support a claim about a particular patch. A deployment check may support the behaviour of a complete candidate in a specific environment.

When the system changes, those claims should not all become stale in the same way.

If the work is the persistent unit, its evidence needs a persistent account of what it supports, what it depended on and what would invalidate it.

From freshness to relevance

I am now thinking about freshness as an impact assessment rather than a binary comparison with the tip of main.

When the base branch moves but the pull request does not, the system can inspect the intervening change. It can compare the affected paths, but paths alone are not enough. It also needs to consider dependencies, contracts, configuration, test infrastructure and runtime boundaries.

That produces three broad outcomes.

If the intervening change cannot reach the boundary supported by the existing evidence, retain that evidence. The base moved; the relevant assumptions did not.

If the changes may interact, run the checks that exercise that interaction. This may be possible against a temporary merge result without rewriting the reviewed branch or repeating its complete authoring cycle.

If the interaction is consequential, unknown or difficult to reverse, require fresh integration evidence against the current system. Uncertainty is not a reason to wave the work through.

If conflict resolution or another edit changes the pull request itself, obtain fresh evidence for the changed patch. Even then, that does not automatically mean running every validator the repository possesses. The evidence should remain proportionate to the affected boundary and consequence.

This is not a relaxation of quality. It is an attempt to make each check answer a real question.

Watching the system does not mean watching every run

There is a popular aspiration for AI processes: arrange them so they no longer need you.

I understand the appeal. An always-on system that requires an always-on human has automated activity without automating responsibility for the flow of work. I still want systems that can continue safely when I am not present.

But independence of execution is not independence from examination.

The validation loop worked as designed. The agents were not disobedient. The checks did not fail to run. The system reliably enforced the rule I had given it. Only by watching the complete behaviour—rather than celebrating each successful validation—did I notice that the rule was producing an expensive and uneven race.

Watching, in this sense, does not mean supervising every agent or approving every command. It means preserving enough observability to question the system's proxies:

  • Is repository movement a good proxy for invalid evidence?
  • Is the number of completed checks a good proxy for assurance?
  • Is a clean merge a good proxy for compatibility?
  • Is unattended activity a good proxy for useful progress?

Human judgement belongs at the boundary where those proxies are chosen, challenged and corrected. A mature system should then encode the better rule, retain the supporting evidence and bring back the cases it cannot resolve.

The aim is less continuous intervention, not less accountability.

The harder rule needs better state

“Retest against the latest main” was attractive partly because it required very little understanding. Compare two revisions. If they differ, begin again.

An impact-aware freshness rule needs more from the system. It needs to know:

  • the pull request patch and the base against which it was assessed
  • the claims each piece of evidence supports
  • the boundaries and dependencies that could affect those claims
  • the changes introduced to main while the work waited
  • the consequence of accepting a mistaken reuse decision
  • which uncertainties require human judgement

That state will not always be complete. Dependency maps can be wrong. Configuration and runtime coupling can hide outside obvious paths. A confident classification can create false assurance more efficiently than a blunt retest rule.

The conservative boundary matters. Related or unknown changes should not be declared irrelevant merely to make the queue move. Consequential work may justify fresh end-to-end evidence even when a narrower analysis looks safe. Low-risk, reversible and demonstrably independent changes can justify reuse that the old rule forbade.

The right level of concern is not a universal setting. It is a decision made from interaction risk, consequence, reversibility and the quality of the available evidence.

What I am changing

I am replacing the automatic freshness reset with a narrower, auditable gate. The system will record the previously assessed base and head, inspect the intervening base delta and make the interaction assessment explicit.

It will not be allowed to call a change safe merely because paths do not overlap. Dependency, contract, configuration, test-infrastructure and runtime interactions must be considered. If any relevant category is related or unknown, existing evidence cannot simply be reused.

The first version will probably be conservative. That is appropriate while I learn where the classifications fail. I want to measure evidence retained, checks rerun, elapsed time, tokens consumed, human interventions and defects or interactions found after reuse. Avoided work is only an improvement if the accepted result remains dependable.

This may reveal another constraint. The assessment itself may become expensive. The metadata needed to support evidence reuse may cost more to maintain than retesting some small changes. A useful gate must be able to choose the cheap check when analysis would cost more than running it.

That would not make the correction wrong. It would locate the next decision: when is understanding the change cheaper and safer than repeating the proof?

What actually became stale?

The original safeguard was not foolish. Testing against the shared system is how integration problems become visible before users find them. The failure was allowing a fact that was easy to detect—main moved—to decide a question it could not answer—the evidence no longer applies.

At scale, that shortcut produced more than wasted computation. It changed the shape of the queue. The fastest work repeatedly won, while the work carrying the greatest assurance burden was given the most chances to start again.

I do not want to solve that by ignoring change or waving through anything that merges cleanly. I want the system to preserve valid evidence, challenge changed assumptions and spend its assurance effort where the work can actually interact.

The repository will keep moving. The harder and more useful question is what that movement changed.