I was trying to understand why an AI repair workflow could look complete while leaving consequential defects behind.

The investigation began with code review. I gave the same historical pull request to three independent AI reviewers. The strongest reviewer found four valid defects. The other two found two each. There was only one overlap, so the three reviewers exposed seven distinct problems between them.

That result became The best AI reviewer still missed three defects. Its lesson was not that three opinions are automatically better than one. It was that the models followed different evidence paths, and that independently adjudicated differences revealed boundaries any single reviewer had missed.

Finding defects was only the first half of the workflow. Someone—or something—still had to repair them.

I gave the same seven validated findings to three fresh agents: Luna, Terra and Sol. Each worked from an identical copy of the affected code. Each changed the implementation, ran its checks and published a structured resolution record. Every record said Resolved. Every one marked all seven findings Fixed.

When I independently evaluated the resulting code, only one repair was complete.

That became When 'resolved' only meant well formed. The workflow could prove that the right fields were present, that every finding had received a disposition and that the agents' chosen tests passed. It could not prove that the original behavioural requirements now held.

The obvious explanation was that the repair agents had made mistakes. That was true, but it did not explain why all three could produce such coherent accounts of success.

I began to wonder whether the problem had entered earlier.

Perhaps we had captured the findings badly

A code-review finding has to travel between two kinds of work. A reviewer discovers a possible failure by moving through code, documentation, state and consequences. A resolver receives a short record of that investigation and has to reconstruct enough of it to change the right behaviour.

The native findings were written as review comments. They were accurate, but their prose could combine the observed failure, location, consequence, boundary and expected direction in one narrative. A person can often recover those relationships from context. An agent may attend to one part, produce a plausible local repair and quietly leave the larger obligation open.

Perhaps the findings contained the right facts in the wrong form.

The theory also fitted an earlier question I had explored in What changes when behaviour comes before the code?. In that experiment, agents performed better when the required decisions were present. A behaviour model and fact-matched prose achieved the same correctness, although the structured condition used fewer median total implementation tokens.

That experiment suggested that detail mattered more than notation, but it used small generation tasks that both detailed conditions completed perfectly. A real repair task might expose a difference hidden by that ceiling. Review findings contain causal relationships, partial evidence and existing code. A clearer behavioural structure might help the resolver preserve those relationships through the repair.

It was a plausible theory. I wanted it to survive comparison, not persuasion.

I changed the representation, not the evidence

I returned to the same historical code and the same seven validated findings. This time, each of the three model families received two fresh repair runs.

The control received the findings in their native form. The treatment received a normalised record that separated the same material into named fields. Across 56 compared fields, the semantic payloads were identical. The structured condition did not receive an extra example, a hidden proof hint or knowledge of what the earlier agents had missed.

The comparison was deliberately narrow: would reorganising the same facts make the required behaviour easier to recover and repair?

Each of the six sessions worked in its own disposable branch and pull request. The agents could inspect the repository, change the code, run checks and publish a resolution. They could not read one another's work.

Two condition-blind evaluators then assessed the anonymous candidate commits against the seven original behavioural requirements. A repair received credit only when both evaluators accepted it.

I expected the structured findings either to help or to make little difference. They made every matched result worse.

Conservative agreement between two condition-blind evaluators
Model familyNative findingsStructured findings
Luna3 of 7 repaired2 of 7 repaired
Terra3 of 7 repaired2 of 7 repaired
Sol6 of 7 repaired4 of 7 repaired
Total12 of 21 repaired8 of 21 repaired

The structured condition produced thirteen false Fixed dispositions, against nine in the native condition. It also introduced two critical correctness defects; the native repairs introduced none.

This was not one model reacting badly to an unfamiliar template. Luna, Terra and Sol all repaired fewer findings correctly when the same information was normalised.

The treatment decision was straightforward: undo it.

The cheaper-looking route was worse

The structured runs looked better on two measures of effort. Their median run took 39 turns rather than 52, and their median same-model billed-token ratio was 0.911. Elapsed time did not materially improve, and the model families varied enough that the aggregate should not be mistaken for a stable performance law.

Quality had already decided the experiment. A workflow that consumes fewer turns while leaving more required behaviour unresolved has not necessarily become more efficient. It may simply have reached its visible completion boundary sooner.

That is what made the result more useful than a straightforward failure. All six agents published the same confident outcome: Resolved, with all seven findings marked Fixed. Under the conservative reconciliation, neither blinded evaluator judged any of the six candidates complete.

The clearer records made some of the work easier to process. They did not make the success claim more trustworthy.

A cleaner route to the wrong completion boundary is not an improvement.

The models did not ignore the structure

It would be tempting to say that models are less affected by behavioural input than I had thought. That is not what happened.

The models were affected. Every structured condition produced a different and worse result than its native counterpart. What failed was the assumption that clearer organisation would affect the work beneficially.

That is evidence against a simple theory of better prompting: identify the right facts, organise them clearly and better behaviour will follow. It is not evidence that structure never helps, or that native prose contains some universal advantage.

The benchmark used one schema on one historical case. It tells me to reject this intervention, not every structured finding format.

Description was not the missing authority

The structured record could name the source, location, observed problem, consequence and expected direction. It still could not establish that the candidate repair satisfied the original behaviour.

Each resolver interpreted the finding, chose a change and selected or authored checks that supported its interpretation. Its own tests passed. The repository validator accepted its resolution record. The behavioural invariant could still remain open.

Improving the description at the entrance did not repair the authority gap at the exit.

The findings did not carry an executable, finding-specific proof obligation. Nor did the publication boundary require independent evidence before accepting Fixed. More descriptive fields could not supply a capability the workflow did not possess.

The format may still have contributed to the regression. Normalisation might have removed useful relationships held in the native prose. Named fields might have encouraged the agents to complete each local part without reconstructing the whole behavioural obligation. The structured form may have made a narrower interpretation feel more complete.

Those are possible mechanisms, not findings. The experiment changed the representation; it was not designed to explain why that representation performed worse.

A failed theory can improve the next question

Adding more descriptive fields is now a weak next move. The stronger theory targets the closure boundary instead.

A finding could carry, or resolve to, a reviewer-owned behavioural probe or an explicit proof obligation. The workflow could demonstrate that the probe fails against the reviewed source and passes against the candidate. The resolution record could retain the identity of that evidence, and Fixed could remain unavailable until an independent route confirms the transition.

That theory may also fail. A probe can encode the wrong requirement. An evaluator can share the reviewer's blind spot. Stronger assurance can add cost until the workflow becomes unusable. Proof-shaped evidence can become another well-formed record that reality does not support.

The point is not that proof obligations are already the answer. The format experiment removed one branch from the search and gave us a more precise next question.

What this result can carry

This was one historical case, seven findings and one observation per model-condition cell. It cannot estimate run-to-run variance or establish that structured findings generally reduce repair quality. Seven of the 42 finding-level judgements differed between the two evaluators before conservative reconciliation. Their agreement is useful evidence, not perfect ground truth.

A different schema, task, agent workflow or finding corpus may behave differently. Models and tools will also change underneath any conclusion.

The narrower result is still worth keeping. On this controlled case, reformatting the same facts did not make three model families more faithful to the required behaviour. It made every matched result worse.

I could protect the original theory by treating this as the wrong template, an unlucky run or a task too complicated for the comparison. Any of those may eventually prove relevant. None is established by this experiment.

We are not trying to prove that behavioural modelling works, or that it does not. We are trying to discover which parts of the system change outcomes we care about.

The structured findings were a plausible answer. They were not this answer.