At the start of a recent day, I thought I was about to begin onboarding data. By the end, I had built much of the system around that work and onboarded none of it.

The day was not wasted. With AI's help, I drafted structures, compared alternatives, found gaps, refined language and tested the result. The emerging system was better for that work.

What surprised me was where the time had gone.

It had gone into deciding which boundaries mattered, how information would move, what later work could rely on and how I would know the pieces agreed. AI made candidate answers easier to produce. It did not remove the work of making those answers fit together into one system I could depend on.

An unfinished map does not always stop a journey. If a shared landmark is missing, people can still set off. Each traveller simply has to decide for themselves what the blank space means. Their routes may each look sensible and still fail to agree.

That was close to the problem I had encountered. The tasks were available to start, but some were only available because I—or the AI—could fill in decisions that the system had not yet made.

The gap quietly becomes a decision

I have been using two working labels:

  • A hard dependency stops later work because a required input, decision or capability does not exist yet
  • A context dependency allows later work to begin, but only by making assumptions that the unfinished work would otherwise resolve

These are my labels for a planning problem, not established technical categories.

A hard dependency is visible. A migration cannot run without a schema. A deployment cannot happen without an artefact. The task reports that it is blocked.

A context dependency is quieter. Documentation can be written while the schema is changing, but it has to choose names and meanings. Tests can be drafted before a contract settles, but they have to imagine the behaviour they will validate. A technical design can proceed before a boundary is agreed, but it will draw that boundary somewhere.

Humans do this too. We use judgement to bridge gaps because waiting for perfect information would stop nearly everything. The difference with AI is the rate at which a plausible local interpretation can become several polished artefacts.

The day showed me the dependency between unfinished context and later work. Parallel AI work can deepen it.

A blank on a plan remains visible. In generated work, the blank is often replaced by a plausible answer. Missing context becomes an assumption, and the assumption can become code, documentation or a test before anyone has agreed to it.

A task is independent when it does not need the unresolved decision. If it does, proceed provisionally only when the assumption is explicit, reversible and tracked; otherwise the honest state is blocked.

Reasonable work can still pull apart

One task assumes a boundary. Another uses a slightly different definition. A test makes one version executable. Documentation makes another version look settled. Each output may be reasonable on its own while the collection becomes harder to reconcile.

Parallel work changes the shape of the risk. A person drafting tests, another writing documentation and another shaping an interface may notice a conflict through conversation, shared review or simply by watching the same work evolve. Human teams miss these signals too; coordination is not automatic.

Separate AI tasks do not automatically share the same conversation or see when their interpretations conflict. Unless the workflow gives them a common reference and a way to reconcile assumptions, each can make a different local choice at the same time.

At volume, those choices can become code, tests and documentation before one reviewer sees the fork.

The AI tasks have not independently chosen to make the system drift. Each has responded to the context and instructions available in its task. The problem is that several locally reasonable answers do not add up to agreement across a system.

Software research gives this mechanism some support without proving the AI claim. Cataldo and Herbsleb studied two large software projects and compared the coordination that technical dependencies required with the coordination developers actually performed. Gaps between the two were associated with more software failures, while closer alignment was associated with better productivity (Cataldo and Herbsleb, 2013). The projects were large organisational settings, not AI-assisted solo work, so the result should not be treated as a direct measurement of this experience. It does support the narrower point that dependencies create real coordination work, and missing that work has consequences.

This is also where this reflection differs from two earlier notes. AI multiplies output, not human capacity is about the review and comprehension that arrive after production. When AI velocity outruns feedback is about what happens after a weak decision starts producing descendants. This problem sits one step earlier: should the dependent work have begun before the shared decision had settled?

Making more is easier than agreeing on one direction

By convergence, I mean the work of making boundaries, decisions and outputs agree well enough that later work can rely on them.

AI can help with almost every part of that process. It can generate alternatives, compare them, challenge weak reasoning and test a chosen structure.

Parallelism can fail in two directions. Tasks can fill the same blank differently and pull apart, or inherit the same incomplete premise and agree too easily. Five alternatives are not a decision, and five agents repeating one premise are not independent confirmation.

During this particular day of tightly coupled formation work, AI scaled production more easily than convergence. That is a field observation, not a universal limit on what the tools can do.

Recent productivity research is a useful warning against universal claims here. In an early-2025 randomised trial, METR studied 16 experienced open-source developers completing 246 tasks in mature repositories they knew well. With the tools available at the time, the developers took 19% longer when AI was allowed, despite expecting to be faster (Becker and colleagues, 2025). A late-2025 follow-up produced raw estimates pointing towards gains, but METR judged selection effects and time measurement problems too serious for a reliable current estimate (Becker and colleagues, 2026).

Neither result proves anything about my scaffolding day. Together, they show why “AI makes software work faster” is too coarse a sentence. The answer changes with the task, tool, prior context and work needed to integrate the output. Generation time is only one part of the loop.

The slow part is not always the waste

I have started to distinguish between two bottlenecks.

Like the dependency terms above, these are working labels for what I observed and what may follow—not established categories.

The visible formation bottleneck is the time spent making something coherent enough to be depended upon. It feels slow because the downstream work is waiting and the uncertainty is still visible.

The hidden correction bottleneck is the possible later cost of letting dependent work proceed on incompatible assumptions. Output continues, so the system can look fast until someone has to identify what inherited which premise, decide what remains valid and bring the pieces back into agreement.

The cost is not only correcting the first artefacts. It is finding every later output that treated one of their assumptions as settled context.

I did not observe that complete correction loop during this particular day. It is a risk inferred from the dependency structure and from other experiences of unpicking AI-assisted work, not a measured outcome of this example.

That distinction matters because a visible bottleneck is not automatically waste. Some formation work is the work: learning what the system needs to be, resolving coupled decisions and testing whether the result holds together.

But waiting is not automatically wise either. A slow process may still contain duplicated work, vague ownership, unnecessary review or poor tools. The useful question is whether the time is resolving a real dependency or merely preserving friction.

Three honest states for work that could start

When shared context is still forming, I now find it useful to place downstream work into one of three temporary states:

  • Independent — the task neither consumes nor silently pre-decides the unfinished context
  • Blocked — the task cannot proceed without inventing or committing an unresolved decision
  • Provisional — the task can proceed if its assumption is explicit, reversible and linked to a clear review trigger

The labels do not make the judgement for me. They make the judgement visible.

Provisional work is the category that needs the most care. An assumption that is merely written down can still become permanent through neglect. I need to know where it was used, what would invalidate it and when it will be revisited. If I cannot answer those questions, “provisional” may simply be a more comfortable word for hidden rework.

A practical test is:

If the upstream decision changes tomorrow, what would I need to find and undo?

If the answer is small, visible and reversible, proceeding may be sensible. If the answer is “I do not know”, the work is not as independent as it looks.

This is not an argument for turning every decision into a gate. Low-consequence, reversible work should move. Useful parallelism comes from finding work that is genuinely independent and from making provisional assumptions easy to see—not from treating all available production capacity as evidence that everything should start.

What I am keeping from the day

The scaffolding day was not evidence that AI had failed. It made many individual steps faster and improved the material I could consider.

What it did not do was turn coupled decisions into independent ones.

I had underestimated how much shared context the later work would need. The day felt slower than expected because I had measured it against generation, not against the job of creating a dependable starting point.

The question I am keeping is not only, “Can this task start?”

It is, “What will it have to assume if it does?”

Further reading

These books offer related lenses rather than direct proof of the argument. I am including them as a reading list to examine, not as endorsements of every claim they contain.