The three resolution records told the same reassuring story.
Each agent had received an identical codebase and the same seven validated review findings. Each traced the findings, changed the code, ran the required local verification, pushed a repair and published a structured resolution comment. Every comment said Resolved. Every finding was marked fixed.
The records were not malformed. Their source heads, finding identifiers, outcomes and produced commits were present. The repository's helper accepted them. From the workflow's own visible signals, all three agents had completed the task.
Then I gave the three anonymous patches to an independent evaluator. It knew the seven behavioural requirements but not which model had produced each repair. Only one patch satisfied all seven.
The benchmark had to change when the agents could write
The experiment followed an earlier benchmark of three independent code reviewers. Review is read-only, so three agents can inspect the same pinned snapshot without changing one another's input.
Resolution mutates its target. The first agent to push would otherwise change the code, head identity and review lineage seen by the next. A fair comparison therefore required three disposable pull requests. Their trees were identical, their review findings were identical and each agent was confined to its own branch.
Sol, Terra and Luna ran at the same medium reasoning setting. They were prohibited from reading the original later repair, sibling branches or one another's results. The evaluator saw anonymous source-and-produced commit pairs rather than model names or rankings.
| Resolver | Self-reported result | Blind quality gate | Outcome |
|---|---|---|---|
| Sol | Resolved, 7/7 fixed | Passed all seven behavioural requirements | Complete repair |
| Terra | Resolved, 7/7 fixed | Several partial or incomplete repairs | Incomplete repair |
| Luna | Resolved, 7/7 fixed | Multiple contract and fail-open gaps | Incomplete repair |
The records were valid
It would be easy to describe the two failures as dishonest claims. I do not think that is what happened. The agents produced repairs consistent with their interpretation of the findings, then selected evidence that reinforced that interpretation.
Terra strengthened several important paths but left cleanup retries insufficiently bound to their intended target, handled Windows path identity incompletely and consulted a narrower source of required-check rules than the finding required. Luna's version could fail open when required-check discovery failed, made a cleanup binding optional and did not preserve the complete resolution lineage.
Their tests passed because those tests did not expose every missing behaviour. Their resolution records passed because the grammar could establish that each finding had a disposition, not whether the stated disposition was correct.
The system had validated the shape of the evidence, then treated that as evidence of the behaviour.
Three different questions had collapsed into one
A trustworthy repair workflow has at least three distinct questions:
- Record validity
- Does the resolution name the right source, produced head, findings and dispositions?
- Verification validity
- Did the required checks run successfully against the produced code?
- Claim validity
- Does the produced behaviour actually satisfy what each finding required?
The first two questions are necessary. Without provenance and current verification, a repair cannot responsibly travel through the system. But neither question establishes the third. A record can refer to the correct commit. A test can pass on that commit. The relevant behaviour can still be absent.
This is different from the problem in What the green light left out. There, the underlying result is correct but its representation omits the context needed for the next decision. Here, the representation is complete enough to read and the underlying success claim is what remains unproved.
The actor and the judge shared a blind spot
The agent was not literally the only judge. Deterministic validators and tests stood outside it. But the agent still interpreted the finding, chose the repair and often supplied the targeted evidence meant to establish success. The deterministic layer then checked only the behaviours it knew how to reach.
That creates a closed confidence loop. A mistaken interpretation shapes the patch. The same interpretation shapes the new tests or the selection of existing tests. The passing checks strengthen the original belief. Finally, a grammar validator confirms that the resulting story contains all the required fields.
Independent evaluation interrupted that loop by returning to the original behavioural obligations. It did not ask whether the patch looked plausible or whether the author could produce a coherent account. It asked whether each required scenario now held.
Independence should be proportionate
I do not think every generated change needs another frontier model, a hidden test suite and a blind tribunal. A reversible local edit may justify self-verification followed by an ordinary review. Independence has a cost, and unnecessary assurance can become the queue that stops useful work.
The case becomes stronger when an agent changes shared state, deletes resources, handles credentials or decides whether another assurance step may be skipped. Those repairs are hard to judge from a green test run alone because the important failure often lives in the state the test did not create.
A practical system can add independence in layers: contract tests authored before the repair, deterministic probes that the resolver cannot redefine, a separate reviewer for consequential patches, and occasional hidden or adversarial evaluation to audit the workflow itself. The aim is not ritual separation. It is to create at least one route to reality that does not inherit the repair's central assumption.
What this experiment does not prove
This was one seven-finding fixture in one repository. The blinding was procedural rather than enforced by separate repositories. I did not execute a live merge, a concurrent push race or the destructive cleanup journeys described by some findings. The evaluator itself may have blind spots.
The result does not establish that Sol will generally repair more accurately than Terra or Luna. It establishes something narrower about this workflow: all three resolvers could cross its visible completion boundary while two still carried consequential defects.
In The case of the review bottleneck, I proposed that a Fixer can prove what it changed but should not independently accept its own repair. At the time that was a design principle. This experiment gave me a concrete reason to keep it.
A resolution record still matters. It preserves provenance, forces every finding to receive a disposition and makes the agent's reasoning available for challenge. But the record is an accountable claim, not the final authority on whether the world now matches it.