I left orchestrated agents working while I was away and returned to a queue of pull requests. At first, it looked like leverage. Code, tests and documentation had appeared while I was doing something else. Then I opened the first change.
Every pull request carried questions that generation had not answered. Was this the right change? Which assumptions had it made? What else could it affect? Which evidence would justify keeping it? The agents had produced asynchronously. In the workflow I had built, those decisions had waited for me.
The queue consumed the day I had intended to spend on planning and deeper work. Each delayed review also required me to rebuild context that had been live when the agent began. The agents had preserved their outputs. They had not preserved my state of understanding.
It cost one full day. It cost only one because, after that, I stopped reviewing the changes properly.
The queue was still there. I had not become faster at clearing it. I carried on producing and deferred assurance, intending to test the combined result rather than understand and validate each change as it arrived.
I tried routing review through AI too. It found useful issues, but it did not remove the delay. A general reviewer sat in the same serial path. The agent inspected the code and produced one combined report; I waited for it, then judged whether its findings matched the intended change and were supported by evidence. The problem was not AI review itself. It was the shape of the review—broad, synchronous and unable to separate work that could run independently from decisions that still required human judgement.
The escape route worked in one narrow sense: I was moving again. Once the changes were combined, however, useful coverage depended on me remembering what each one could affect and selecting the right tests. I did not. Validation became closer to clicking through the obvious paths, seeing that they appeared to work and treating that feeling as stability.
I had recovered the appearance of velocity. Assurance was the price.
By assurance, I mean the proportionate evidence and accountable judgement needed to decide whether a change can be kept, repaired, combined or stopped.
The arithmetic did not add up
Several agents had worked in parallel. They produced useful material that would have taken me much longer to write alone. That gain was real. So was the lost day, and so was the weaker evidence after I escaped the queue.
Those facts were the first clue. I had measured work produced while the system was charging me for work completed.
Generated code remains candidate work until it has been understood, verified, integrated and accepted. AI had increased the arrival rate of candidate work without proportionally increasing the system's capacity to establish evidence, reconcile findings and make accountable decisions. Candidate code had become cheap. Completed change had not.
That distinction matters beyond my desk. The review bottleneck is increasingly visible wherever AI-assisted authoring increases the number and size of changes reaching a smaller group of accountable reviewers. A practical guide from Cortex describes the same pressure and responds with risk-based review, specialist AI reviewers, earlier local feedback and retained human sign-off (Buenahora, 2026). It is one practitioner account, not proof that every team has the same constraint. It does show that this is not an eccentric concern produced by one unusual day.
The first suspect
My first suspect was pull-request size. AI can produce a large, internally coherent change with alarming ease. The result may be technically organised and still ask too much of the person who must reconstruct its intent, challenge its assumptions and judge how it fits the rest of the system.
I made the units smaller. I split work into changes that a person or agent could read, explain and validate without holding an entire feature in one review. Where one change depended on another, I used stacked pull requests to keep the sequence visible.
This helped. Each review became less intimidating. A finding invalidated less work. The stack preserved decisions that a larger batch would have hidden.
Then the queue became longer.
Reducing batch size reduced the cost of understanding one change, but it did not reduce the number of arrivals. Smaller pull requests were useful. They were not an answer for the complete system.
Little's law gives the clue a firmer boundary. In a stable process, the average number of items in a system relates to their arrival rate and average time in the system (Little, 1961). It does not prescribe a pull-request workflow or prove what happened in mine. It does tell us that if candidate changes arrive faster than they can be completed, waiting will grow somewhere.
The queue was not inert
I had treated the pull requests as passive inventory: an inconvenient list that would remain where I left it. That assumption was wrong.
A pending pull request contains a possible future that does not yet exist on main. Until it merges, mainremains canonical, but it becomes a poorer forecast of the system that new work may eventually meet.
New changes can repeat a decision already waiting in review, depend on an interface about to move, update documentation that another branch has rewritten or pass tests against assumptions that will not survive integration. Stacks help with dependencies I already know. Independent stacks can still meet at a shared file, contract, environment or architectural decision.
Review delay therefore creates a reinforcing loop. More work remains pending. Later branches become more speculative. Speculation produces conflicts, repeated validation and context reconstruction. That integration work consumes the same assurance capacity whose delay created it.
How apparent velocity weakens assurance
- ProductionCandidate changes arrive fasterparallel agents increase the arrival rate
- QueueAssurance waitsreview, evidence and acceptance stay scarce
- AgeingContext and baselines divergelater work becomes more speculative
- ReworkThe constraint receives more workconflicts · retesting · reconstructed intent
Delay creates integration work that returns to the same constrained assurance path. Bypass makes the queue look smaller while justified confidence becomes weaker.
The other suspects
Pull-request size was not the only plausible explanation. A single slow validator could have been holding everything up. The problem could have been poor sequencing, an overused test environment, vague intent or simply too much work in progress. Formation—the work of deciding what was valuable and coherent enough to begin—could have been the real constraint all along.
Each suspect explained part of the scene. Smaller changes helped comprehension. Better stacks exposed known order. Limiting arrivals would reduce immediate pressure. A faster validator would shorten one feedback loop. Clearer intent would prevent agents from making plausible but divergent decisions.
None explained the whole pattern. The congestion crossed tests, AI and human review, specialist judgement, integration and acceptance. The general reviewer added useful scrutiny, but as one broad, synchronous pass it remained inside the same constrained path. Smaller units made the queue more legible without completing it. When I bypassed several kinds of assurance at once, apparent velocity returned.
The Theory of Constraints offered the best-fitting current hypothesis: accelerating one stage does not accelerate the complete system once another stage governs its throughput. AI had accelerated production and exposed assurance as the next candidate constraint.
This remains a hypothesis about my current system, not a new law of software delivery. Formation may become binding again. Integration environments may prove to be the real scarce resource. Better assurance may simply move the queue elsewhere. The point of the next intervention is to find out.
A theory is not a plan
Saying “assurance is the constraint” does not tell me what to build. A collection of extra reviewers and checks could make the bottleneck more expensive from its first day. A rule that every change must pass every control would buy confidence by recreating the queue. A rule that people should simply review faster would rename the problem as personal failure.
The proposed response is a composable assurance system. It is a plan, not implemented behaviour. The current workflow already contains useful foundations—risk-selected review, deterministic checks, bounded review lenses and a separate repair path. Those components still operate inside a flow capable of reproducing the incident described here. None of the phases below has yet proved the wider system.
Composability is the important part. Each assurance component should declare:
- what change or affected boundary selects it
- what evidence it consumes and produces
- whether it reads, executes against mutable state or writes
- where in the lifecycle its evidence can be established faithfully
- what a failure obliges
- what completion does not prove
- what it costs and which other components it depends on
That contract lets the system select only the components a change owes. It also lets one component move, narrow, improve or disappear without redesigning the entire assurance path. A monolithic review process decays as its assumptions change. A composable one makes those assumptions visible and testable.
How one change selects its assurance
Selection uses two independent questions. First: how consequential is the change, and how difficult would it be to reverse? That determines the minimum depth of evidence. Second: which boundaries does this exact change affect? That selects the relevant deterministic workers and specialist judgement within that depth.
Paths are useful clues, but they do not decide consequence. Two changes to the same source file may owe different assurance if one adjusts internal formatting while the other changes authorisation or stored data.
A deterministic worker answers a predefined question with repeatable evidence. An Agent Persona applies bounded, contextual judgement to one concern, remains read-only and returns findings rather than acceptance. Neither substitutes for the person accountable for the change's intent or release.
The operating system for one exact change
Consider two changes. A contained internal refactor may select a compiler, focused tests and no specialist Persona. A change to authentication that also alters stored account state may select security, contract and persistence judgement, migration and provider evidence, plus an explicit accountable acceptance decision. Running the deeper set for the refactor would be waste. Running the lighter set for the authentication change would be missing assurance.
Deterministic does not mean harmless to run concurrently. A test may mutate a database or consume a shared environment. Read-only Personas can usually share an exact-head workspace. Workers that compete for mutable state must be isolated or ordered. The point is not maximum parallelism. It is parallelism with a reason.
Writing changes the evidence
Review should not alter the branch it claims to have reviewed. Repair does. That is why the proposed system separates a read-only Reviewer from one controlled Fixer boundary.
The Fixer does not treat a finding as established merely because another agent published it. It traces the finding, decides whether it is supported and records what happened to it. Repair workers may be composed behind the Fixer, but they do not independently push, publish or declare success. Writing authority stays narrow.
The proposed resolution modes are hypotheses to test rather than permanent policy. One mode might escalate everything except the minimum safety floor. Another might repair bounded, mechanical work and escalate ambiguity. A later mode might attempt a best-supported interpretation while making the remaining judgement explicit. The experiment must discover which allocation is accurate, economical and responsible.
Repair, escalation and independent re-review
Seven phases, one uncertainty at a time
The plan does not begin by launching more agents. It begins by measuring the existing system, then introduces one new source of uncertainty at a time. Selection is tested before execution. Read-only execution is tested before writing. One pull request is understood before several are allowed to move together.
1. Observe the current system
Measure the path from intended change to delivered change: arrivals, completion, waiting, critical-path time, stale work, repeated work, assurance performed or bypassed, defects, incomplete outcomes and recovery. Missing historical data remains missing; it should not be reconstructed into false precision.
Why first: without a baseline, later automation can look successful merely because it produces more activity.
2. Decompose assurance into components
Define the fixed vocabulary, component contracts, evidence envelopes and ownership boundaries. No new specialist review or automated repair is introduced yet.
Why second: a system cannot select, measure or replace assurance while review remains one undifferentiated box.
3. Shadow the selection
For real changes, produce an assurance plan without changing the live workflow. Compare each plan with human judgement. Did it identify the affected boundaries? Did it miss consequential evidence? Did it select work with no positive reason?
Why third: prove that the system can choose the right work before trusting it to execute—or omit—anything.
4. Prove the system on one change
Run selected independent Personas and deterministic workers concurrently inside one pull request. Exercise unclear intent, stale heads, failed isolation, incomplete work and reconciliation as deliberately as the happy path.
Why fourth: establish the inner operating system and its critical path without adding cross-pull-request scheduling.
5. Introduce controlled resolution
Add escalation first, then bounded repair, then more interpretive repair if the evidence justifies it. Reclassify every produced head and keep independent re-review separate from fixing.
Why fifth: writing changes the object being assured. It belongs after read-only selection and reconciliation have been shown to work.
6. Prove coordination without concurrency
Process a chosen set serially. Then introduce an immutable wave: a frozen set of candidate changes, explicit dependency order, restartable fresh-context takes, a terminal disposition for every member and assurance over the consolidated result.
Why sixth: prove membership, recovery, consolidation and failure accounting before concurrency makes a bad outcome harder to attribute.
7. Add bounded concurrency and keep auditing
Only now allow several pull requests to progress asynchronously. Keep the ceiling configurable and separate from the worker concurrency inside each pull request. Continue auditing omitted assurance, unnecessary work, displaced queues, stale results and components whose cost has outgrown their value.
Why last: concurrency should optimise a system already shown to behave safely. It cannot be the mechanism that makes the system correct.
How the plan earns permission to scale
- 1Measureestablish the baseline without changing the flow
- 2Define componentscontracts, authority, evidence and limits
- 3Shadow selectioncompare proposed assurance with human judgement
- 4Prove one changeselected fan-out, reconciliation and stale paths
- 5Control repairescalate, write, reclassify and re-review
- 6Prove wavesserial coordination, recovery and consolidation
- 7Add concurrencybounded cross-change scale, followed by audit
Why one pull request comes before many
There are two distinct kinds of concurrency. Inside one pull request, independent assurance components can overlap and shorten the critical path. Across pull requests, several complete assurance systems compete for repositories, environments, model capacity and integration order.
Introducing both at once would make success difficult to explain and failure difficult to locate. If the first experiment is slow, I need to know whether selection was poor, workers were serial, reconciliation dominated or cross-change contention caused it. Proving the inner system first removes one variable.
The same logic explains serial batches and waves. A wave freezes membership before it starts. Later arrivals wait for another wave. Every candidate ends as eligible for consolidation, returned to authoring or carried forward because a dependency failed. The successful subset is consolidated, then boundary assurance determines whether the next wave may begin.
This is not a permanent queue or a claim that one batch shape will suit every system. It is an accounting boundary. It makes missing work, partial failure and a changing baseline visible before concurrency is permitted to hide them.
How another team could begin
The first useful experiment requires no Agent Persona, workflow engine or new merge gate. Take a small set of recent changes and reconstruct the assurance each actually needed.
- Write one sentence describing the intended change.
- Identify its consequence, reversibility and affected boundaries.
- List the deterministic evidence that was required.
- List the questions that required contextual judgement.
- Record what ran unnecessarily and what important evidence was absent.
- Record where feedback first became available and where it was finally used.
- Compare the result with a person accountable for accepting the change.
Repeat that as a shadow exercise before automating selection. The aim is not to invent a universal risk score. It is to discover the boundaries, evidence and decisions your own system repeatedly owes.
What would change my mind
The assurance hypothesis becomes weaker if improving assurance capacity does not improve completed flow, if the queue simply moves to formation or integration, or if human decision load remains unchanged because reconciliation still requires reading every underlying report.
The proposed architecture should be revised or undone if it misses consequential boundaries, produces stale confidence, makes recovery slower, increases escaped defects, or costs more than the blanket work it replaces. Faster review is not a success if quality or assurance degrades. Perfect-looking assurance is not a success if the system becomes too expensive to use.
A change in allocation is not automatically a failure. A question initially assigned to a person may become deterministic once its rule is understood and a reliable check exists. A specialist Persona may be narrowed after repeated false positives. A slow check may move to a later faithful boundary. A component that no longer protects enough value should be removed.
That evolution is why the components need versioned contracts rather than live self-modification. The system learns through measured, reviewed changes to its parts—not by allowing a worker to rewrite its own authority while it runs.
Assurance, speed and confidence at scale
The goal is not a shorter pull-request list at any price. It is a delivery system in which quality, assurance and speed improve together.
- Quality asks whether the delivered change met its purpose without unacceptable regressions or consequences.
- Assurance asks whether the system selected and completed the required fresh evidence, exposed omissions and preserved responsibility.
- Speed asks whether end-to-end feedback, completion and recovery improved—not merely generation.
Confidence is not another green indicator. It is justified when the evidence remains proportionate, current and capable of revealing what the system does not know.
The next level of scale is not reached when agents can write code faster. It is reached when purposeful change can be produced, assured and delivered without outrunning the evidence required to keep it.
If those loops can operate at least as quickly as valuable work can be formed and responsibly accepted, scarce human attention moves towards a healthier constraint: deciding what is worth building, resolving the difficult trade-offs and judging whether the result created real value.
Scale is not more code. It is the ability to turn purposeful decisions into verified, deliverable change without production, assurance or deployment becoming the next avoidable queue.
The case remains open
This post establishes the incident, the working theory and the proposed investigation. It does not claim that the plan has solved the bottleneck. Future posts will follow the phases: what I built, what the evidence showed, what degraded, which components changed and whether each phase was kept, revised or undone.
The investigation continues ideas from When AI works from an unfinished map, AI multiplies output, not human capacity, When AI velocity outruns feedback and The safeguard that made every change wait. Those posts describe the pressure. This one supplies the plan whose assumptions the series must now expose and test.
Further reading
These sources offer related perspectives rather than proof that the proposed assurance system will work. That result must come from operating, measuring and revising it.
- A practical guide to risk-based code review, Cristina Buenahora—risk routing, specialist AI review, local feedback and human attention in a code-review context
- The Goal, Eliyahu M. Goldratt and Jeff Cox—constraints, flow and why improving one stage does not necessarily improve the complete system
- The Principles of Product Development Flow, Donald G. Reinertsen—queues, batch size, work in progress and the economics of product-development flow