I built the validator myself. I trusted what it checked. I also made every commit wait between one and three hours for it.

The assurance process was detailed and the checks were valuable, but they had to run serially. Once I put all of them on the path of every commit, the pipeline stopped feeling like a delivery system and started feeling like a waiting system.

The mental picture I keep is a road narrowing into one inspection lane. The inspection may be necessary. Its capacity still becomes part of the road.

The obvious conclusion would have been that the assurance was too thorough. I do not think that was the problem. I had put all of the checking at the same point, at the same frequency and behind a serial execution model. A useful safeguard had become a poor checkpoint.

This was one experience, not a controlled study. I did not measure a team gradually bypassing the process, and I did not establish a universal limit for how long a validator should take. What the experience gave me was a more useful question.

Can the assurance system process work at the rate and level of risk the production system creates?

A slow checkpoint is not automatically the bottleneck

I first described the lesson too absolutely: a system can move only as quickly as its slowest mandatory validator.

That is memorable, but incomplete. A long-running check does not necessarily limit delivery. It may run in parallel, reuse a previous result, inspect only affected work, or sit outside the critical path while a later boundary remains protected.

The narrower claim is this: when every change must pass through a validator, and that validator cannot sustainably process changes at the rate they arrive, it constrains the rate of completed work. If the mismatch continues, the input has to slow, validation capacity has to increase, some work has to wait, or the control has to change.

Queueing theory gives this relationship firmer language. For a stable process, Little's law relates the average number of items in a system to their arrival rate and average time in the system (Little, 1961). It does not tell us which software checks to run. It does help explain why work in progress and waiting time cannot be treated as unrelated measures.

There is also evidence that build latency changes how developers work. Jaspan and Green describe experiments at Google in which moderate improvements to build latency were associated with more builds and faster completion of small and medium changes (Jaspan and Green, 2023). Their setting does not establish a universal threshold for every codebase, but it supports treating feedback time as part of the development system rather than a harmless delay around it.

In my case, one to three hours was plainly wrong for a per-commit gate. That does not make the same suite worthless. I had designed the checks, but not the way work would flow through them.

Minimum sufficient assurance is not minimum effort

I use minimum sufficient assurance as a working phrase, not an established industry threshold.

The order matters. First decide what evidence is sufficient for the consequence of being wrong. Then find the smallest, fastest and most reliable set of controls that can provide that evidence.

“Minimum” applies to redundant friction, not to the required level of safety.

A formatting error and a destructive database change should not pass through the same assurance path. A reversible internal change may be safe to test more deeply after merge. A release, deletion or irreversible decision may need strong evidence before the boundary can be crossed. In some systems, broad integration or security checks must remain pre-merge. The shape follows the risk, not a universal sequence.

A layered alternative to one deep gate

  1. 01 · While changingFast local evidenceformatting · types · schemas · contracts
  2. 02 · Change boundaryTarget what can be affectedaffected tests · review · governing rules
  3. 03 · System evidenceRun broad checks with capacityintegration · regression · consistency · security
04 · Consequential boundaryDo not cross while required evidence is failingrelease · deletion · irreversible action
Depth follows consequence and reversibility, not simply elapsed time. The layers can overlap; the protected boundary determines which evidence must be complete.

Moving a check later does not make failure optional. It requires an owner, a response expectation and a boundary that cannot be crossed while the necessary evidence is failing.

The better design may combine several approaches: run the relevant parts immediately, parallelise independent checks, reuse evidence that has not been invalidated, and reserve the comprehensive suite for the point where its breadth justifies its cost.

The validator needs evidence too

A validator is another operated system. It consumes time and money, depends on assumptions and sources of truth, produces failures that people have to understand, and changes as the system around it changes.

That makes a few questions worth asking regularly:

  • Which meaningful failures does this check detect?
  • How quickly and clearly does it return a result?
  • How often is it false, flaky or ignored?
  • Which work really invalidates its previous evidence?
  • What does it cost to execute and maintain?
  • Which consequential boundary does it protect?

A five-second check that prevents irreversible loss may be extraordinarily valuable. A slow check that rarely detects relevant risk may need to be redesigned, moved or removed. Duration alone does not decide; detection value and consequence do.

I did not observe people learning to bypass the assurance process in this experience. I think the pressure is still worth noticing. A control that makes normal work repeatedly painful creates an incentive to work around it, but that is a design risk, not a result this incident proved.

The same caution applies to correction cost. Delayed feedback can leave more context to reconstruct and more adjacent changes to consider. Whether a particular correction actually becomes more expensive depends on the work, its coupling and what happened while it waited.

Production can scale faster than assurance

The mechanism is not limited to a test suite.

If several people or AI agents produce work in parallel while one person must personally review every consequential decision, production capacity and assurance capacity have grown at different rates. That person may become a gate, but validation is not inherently serial. Predictable checks can be automated, relevant review can be distributed closer to the work, and scarce human attention can be kept for ambiguity, exceptions and high-impact judgement.

AI makes this mismatch easier to create because starting another stream of production is cheap. I have written separately about the point at which AI output exceeds accountable human capacity. Here the narrower lesson is that automation does not only increase the work being checked. It can also help decompose, target and operate the checks—provided people retain ownership of what those checks actually establish.

This is where “validate the validators” earns its place. It is not a call for another giant layer of process. It is a reminder to inspect whether the assurance system remains accurate, useful and proportionate to the work it governs.

What I took from the experience

The lesson was not that slow checks are bad, or that every pipeline should prefer speed. Some assurance is necessarily expensive. Some risks deserve deliberate waiting.

What changed for me was the unit of design. I stopped thinking only about whether each individual check was good and started asking whether the complete assurance path could work at the pace of the system around it.

For any important validator, I now want to know:

  • what risk it is reducing
  • where the earliest useful evidence appears
  • how much work can arrive while it runs
  • whether independent work can proceed safely
  • what boundary must remain closed when it fails
  • who owns the response and the validator itself

The one-to-three-hour suite may still have been valuable. It was making every commit stop at its deepest checkpoint that exposed the mismatch.

I still want the evidence. I no longer assume every change should wait in the same queue to produce it.

Further reading

These books offer related perspectives rather than direct proof of this reflection. I am including them as a reading list to examine, not as endorsements of every claim they contain.