I have repeatedly ended up with between thirty and 150 open pull requests.
At first, the queue looked like leverage. While I focused on work that needed more of my judgement, context or presence, agents continued with the bounded work I had delegated. Code, tests and documentation existed that had not existed before. Measured at the point of generation, an extraordinary amount of work had been completed.
Then I began trying to accept it.
Each pull request needed to be understood, reviewed, verified and integrated. Some needed repair. Some had been created before a related decision settled. Some were correct against the version of the system they had started from, but the system had moved while they waited.
The list became a death march. Eventually I reached the point I can only describe as “f&%k it, push them through”. The queue cleared, but not because I had found a responsible way to absorb the work. Exhaustion had changed the standard.
I have already written about parts of this experience. AI multiplies output, not human capacity is about production growing faster than accountable understanding. When AI works from an unfinished map is about parallel work filling unresolved decisions with different assumptions. The case of the review bottleneck follows the constraint into assurance and sets out the verification and repair system I am now beginning to build.
That system decomposes the evidence a change needs, assigns the right checks and kinds of judgement, runs independent verification in parallel and keeps writing authority narrow. I still think that work matters.
But a thought stopped me while I was considering it:
What if I am improving the part of the system that receives the queue when I should also be asking why the system manufactured that queue?
One accepted reality, many possible ones
Asynchronous software work usually begins from a shared point. A branch takes the current source, changes it somewhere else and later attempts to return.
That branch is not yet the system. It is one possible future of the system. Start thirty branches and there are thirty possible futures developing at once.
The first one accepted into the main branch changes the shared reality. The remaining twenty-nine do not necessarily become wrong. Independent changes may still merge cleanly, and some evidence may remain relevant. But they were produced against a state that is no longer authoritative in exactly the same way.
Their continued validity may now need to be established again. A branch may need rebasing. Tests may need rerunning against the new combination. A review may need to consider an interaction that did not exist before. A repair can change the branch sufficiently to make earlier evidence stale. A clean textual merge can still combine two changes whose assumptions disagree.
This resembles an established idea in computing. Optimistic concurrency allows work to proceed without holding a lock, then validates it before commitment. It is valuable when conflicts are uncommon. When contention rises, the cost of validation, retry and rollback rises with it. The analogy is imperfect—software changes are not database transactions—but it gives me a useful way to see a branch: it is speculative work until the current system can accept it. H. T. Kung and John Robinson described the underlying concurrency model in their 1981 paper, On Optimistic Methods for Concurrency Control.
GitHub's merge queue makes part of this cost visible in ordinary delivery. It tests queued changes against the latest target branch and the changes ahead of them. A conflicting or failing change can be removed and later combinations rebuilt. Even moving a pull request to the front can require every in-progress entry to be rebuilt because the commit graph has changed (GitHub Docs).
The important point is not that every merge ejects every other change. It is that concurrent changes are not fixed pieces waiting in an inert list. They remain related to an authoritative state that continues to move.
I later turned that qualification into a concrete assurance question: when the shared state moves, which evidence actually became stale? The validation race that crippled my PR workflow follows the safeguard, the repeated revalidation race and the impact-aware rule that replaced it.
The cost of change may compound
I had been thinking about the queue as thirty or 150 separate units of work. Finish one review, reduce the list by one. Finish another, reduce it again.
That assumes the cost of each item remains roughly stable while it waits.
My experience suggests that assumption fails when the work is coupled. The first accepted change can alter what the second needs. Repairing the second can alter what the third needs. Later changes may contain code, tests, documentation or conclusions inherited from an earlier branch that is no longer going to survive.
The cost is not only resolving a merge conflict. It can include discovering which reality an artefact describes, deciding whether its evidence still applies, reconstructing the context in which it was produced and finding its descendants.
I do not yet know the shape of that cost well enough to call it exponential or attach a formula to it. Coupling, reversibility, batch size and the quality of the changes will all matter. The narrower claim is enough for now: the cost of a waiting change is not necessarily constant when other changes continue to alter the system around it.
That creates a difficult possibility. Starting more work can increase visible production while reducing the rate at which the complete system can accept it.
The numbers may be true and still mislead
AI productivity is often made visible through numbers that are easy to count:
- code generated
- tasks completed
- pull requests opened
- agents running concurrently
- time to first output
- the duration of an individual coding task
These measurements can be accurate. The problem begins when a measure of one stage is used to make a claim about the whole system.
Generated code is not accepted change. An opened pull request is not delivered value. An agent reporting completion does not establish that the result remains correct, compatible, understood or worth keeping.
If the intended outcome is valuable, maintainable software, the measurement has to include the path from decision to accepted change: formation, generation, review, verification, integration, rework, recovery and the human attention each requires.
Using a partial measure to support a system-wide productivity claim can produce a false conclusion even when the number itself is correct.
This is one reason current productivity evidence remains difficult to interpret. In a 2025 randomised study, METR found that experienced open-source developers using early-2025 AI tools took longer on the selected tasks despite believing they were faster. A later attempt produced estimates pointing in the other direction, but METR considered them unreliable partly because adoption changed who was willing to participate and because developers running several agents found active time difficult to report (METR, 2026). Neither result measures the system I am describing. Together they are a useful warning that “AI makes software faster” is too coarse a sentence.
Eight visible hours and sixteen missing ones
There is another mismatch in the way I have been thinking about productivity.
A person has a limited period of sustained attention in a day. It therefore makes intuitive sense to maximise what can be produced during that period. If four agents can finish four tasks while I finish one, the parallel version looks obviously better.
But an AI system does not share the same working day. It can, in principle, continue after I stop. If it needs me to review every result, select every next task and repeatedly tell it to continue, that extra operating time is mostly theoretical. I remain the clock governing the workflow.
The comparison may therefore be between two different shapes of progress:
- a large burst of speculative work produced during the hours a person can supervise it
- a steadier flow of accepted work that can continue whenever valid work is available
The second may initially look slower because fewer agents are visibly active and less candidate work arrives at once. Over a full day or week, it may complete more useful work by avoiding the queue that suffocates the first approach.
The missing sixteen hours are not free. Compute has a cost. Unattended failure can also run for longer. The system still needs permissions, limits, evidence and safe stopping conditions. The question is not how to keep every machine busy for twenty-four hours. It is whether the system can continue making credible progress without requiring a human to remain continuously present.
I may be treating agents too much like people
The word agent gives us a useful mental model. It also quietly brings an organisation with it.
We give an agent a role, a ticket, a workspace and a branch. We ask it to own a piece of work, retain its context and report when it is finished. When we want more capacity, we create more agents. When an agent stops, we return to its conversation and encourage it to continue.
These are recognisably human patterns of work. They make a complicated technology easier to understand, but that does not mean they make the best use of it.
An AI agent can be created for a bounded operation and disappear afterwards. Another can resume later. It does not require an enduring identity, a career inside the project or permanent ownership of a ticket. A human can participate in the same work where human context, authority or judgement is required, but the entire system does not have to be built around a durable worker holding it together.
This is the inversion I want to examine:
The goals, dependencies, evidence and accepted state should persist. The agents performing the next valid part of the work can be temporary.
Temporary does not mean interchangeable in every situation. Different models, tools and people have different capabilities and authorities. It means the continuity of the work should not depend on any one of them remaining alive, available or able to remember the whole journey.
A documentation retrieval experiment later gave me one concrete version of that boundary. A temporary evidence worker investigated and verified the wider source collection, then disappeared. Only the accurate, minimum-sufficient evidence entered the continuing task. The result supports part of this theory, while also exposing the cost and capability the temporary worker may require: I measured the wrong context—and the system carried the cost forward.
Stateless agents may still be useful. But stateless executors require stateful work. If the state disappears with the conversation, the next agent can only repeat the discovery or guess what its predecessor understood.
The system should decide what is ready
The alternative begins with a persistent representation of the work required to move towards a goal. It would hold the tasks, dependencies, requirements, current state and evidence needed to accept each transition.
The system could then expose the work that is genuinely ready. Independent items could run in parallel. Related work could proceed in order. A blocked item could wait without an agent consuming context or improvising around the blocker. Completion of one accepted step could trigger the next.
An available executor—an AI agent, a person or a deterministic tool—could take a bounded piece of eligible work. When it finished, its useful result would enter the persistent state and the executor could disappear.
This is not a case for making all work synchronous. Exploration may benefit from several independent perspectives. Read-only checks against an immutable change may run together. Changes in genuinely separate parts of a system may advance independently. Anthropic reported that its multi-agent research system performed particularly well on breadth-first work with independent directions, while coding tasks and other domains with shared context or many dependencies were a poorer fit. It also reported much higher token use for multi-agent work (Anthropic, 2025).
The distinction I want the system to make is not synchronous or asynchronous. It is ready or not ready, independent or coupled, speculative or accepted.
Parallelism should follow from the state of the work, not from the number of agents available.
Relentlessness belongs to the system
The appeal of agents is often framed as autonomy: give an agent a goal and let it continue until it succeeds.
I am beginning to prefer a different kind of autonomy. The system continues to advance while any valid work remains. Individual attempts can stay narrow. Agents can finish, fail or stop at defined boundaries. The persistence comes from the system's ability to record the outcome and choose the next valid transition, not from one agent behaving heroically inside an ever-growing context.
That system may sometimes choose to wait. If a task depends on an unsettled contract or a human decision, starting it anyway does not demonstrate autonomy. It creates another possible reality that may later need to be unwound.
Waiting can be correct behaviour. The system can release that executor, move other ready work forward and resume the blocked path when its dependency is resolved.
The system is relentless because progress can survive interruption. It is not unstoppable. It should stop when the work is exhausted, when only human decisions remain, when evidence is insufficient, when a cost or safety boundary is reached or when the goal is no longer worth pursuing.
An always-on system should need less of an always-on human
Externalising the work would also make it visible in a way that a collection of private agent conversations is not.
The current state could be represented as a graph or command hub: what is ready, what is running, what is blocked, what has become stale, which goals are moving and where a person is genuinely required. Events such as a merge, test result, webhook or decision could update that state and release further work.
That changes what an agent interface or mobile application is for. Instead of helping a person supervise more conversations and repeatedly press “continue”, it could present the exceptions that require their authority:
- a consequential decision
- missing context
- incompatible goals
- a cost or risk boundary
- an evidence gap the system cannot resolve
The person answers the bounded question. The system records the decision and continues.
An always-on system should reduce the need for an always-on human, not make constant supervision more convenient.
The dashboard could still mislead. A beautiful graph full of activity can become another vanity metric. Its primary measures should concern accepted progress: completion towards the goal, time and cost per accepted outcome, repeated or invalidated work, human attention, failure and recovery. Agent count, token use and generated artefacts remain useful diagnostic measures, but they are not the outcome.
The gravitational pull of what already exists
There is a final complication in thinking about this with AI.
The models helping us design agentic systems have learned from the systems, organisations and explanations people have already produced. Ask for an agent architecture and familiar shapes readily appear: workers, managers, teams, roles, tickets and hand-offs.
Those patterns may be appropriate. Existing knowledge deserves more than novelty for novelty's sake. But a model's ability to name an established pattern does not establish that the pattern is right for a new kind of executor.
I noticed that pull while discussing this idea. Each time I questioned the worker model, the conversation returned to improving the workflow around the workers: a better queue, a clearer graph, a stronger scheduler, another plan to test. Those are useful components, but they can domesticate the question before it has been allowed to challenge the premise.
Are human-shaped agent systems common because they suit the properties of the technology? Or because human organisations gave us the easiest language with which to imagine it?
I do not know. That uncertainty is part of this note.
A theory to build from
I am not claiming that a persistent work system with temporary executors will solve the queues I created.
Maintaining its dependencies may become another bottleneck. It may serialise work that could safely run in parallel. A wrong goal or accepted decision could move relentlessly through it. Human judgement, ambiguity and responsibility do not disappear because a graph makes them visible.
This is a starting theory: some of the problems attributed to AI output, review, context and insufficient agent coordination may begin earlier, in the decision to organise cheap temporary executors around a human model of parallel work.
I plan to start designing and testing from there. The current verification and repair work gives me somewhere practical to begin. I want to compare rapid candidate production with sustained accepted progress, and measure the costs that disappear from generation metrics: revalidation, repair, context, integration, human intervention and recovery.
The result may show that the familiar model is right, that the alternative only works under narrow conditions, or that I have simply moved the constraint again.
But the question has changed for me.
AI has already given us extraordinary sprint speed. I have seen what happens when I use it for repeated bursts: the work races ahead, stops at the next constraint, waits for me to catch up and then sprints again. It can look fast at every half-mile while making the whole course slower and more expensive.
The hare's speed is not the problem. The problem is failing to sustain useful progress across the complete course. I want to build a system that can use that speed when the path is clear, wait when it is blocked and keep moving until the goal is reached.
The measure is not how impressive each sprint looks. It is whether the system completes the marathon efficiently.