I expected a code-review benchmark to tell me which model was best. It did, but that was not the most useful answer.
I ran three Codex model families—Sol, Terra and Luna—against the same historical pull-request snapshot at the same medium reasoning setting. Each reviewer worked independently. They could read the repository and run local checks, but they could not read one another's work, the existing review comments or the later commits that repaired the change.
Sol was the strongest individual reviewer. It found four valid defects. Terra and Luna found two each. If the purpose had simply been to rank them, the result would have been clear.
But the reports overlapped only once. Between them, the three reviewers exposed seven distinct defects. Terra found one consequential omission Sol missed. Luna found two more. The best reviewer had still left almost half of the combined finding set unseen.
A snapshot with known consequences
The benchmark used a pre-fix checkpoint from a substantial pull request: fourteen changed files and roughly 1,365 added lines around a branch-mutation workflow. Later repair commits gave me a useful, if incomplete, source of adjudication. I could ask whether a reported defect corresponded to a boundary that the subsequent repair actually changed.
That does not turn the later implementation into perfect ground truth. A later change can be unnecessary, incomplete or mistaken. I therefore traced each finding through the code and governing contract as well. All seven reported defect families were supported by that combined evidence.
| Reviewer | Valid reports | Unique contribution | What stood out |
|---|---|---|---|
| Sol | 4 | 3 | Strongest reproduction and failure scenarios |
| Terra | 2 | 1 | Resolution-lineage reasoning |
| Luna | 2 | 2 | Cleanup and required-check boundaries |
| Combined | 8 reports | 7 distinct defects | Materially broader than any single review |
They were not making the same mistakes
The distinction was not simply depth. Sol did go deeper in several places. It reproduced a Windows path-comparison failure and gave precise scenarios for losing unpushed work and deleting a remote branch non-atomically.
Terra's most useful contribution came from a different reading of the system. It traced the implementation back to the documented review-and-repair lineage and noticed that the code did not enforce it. Luna concentrated elsewhere again: a cleanup retry was not bound tightly enough to its intended target, and the required-check validation admitted a weaker result than the contract required.
These were not stylistic preferences. They were different ways for a consequential mutation boundary to act on the wrong repository state, discard work or proceed with insufficient evidence.
This suggests that model diversity can operate like a change in viewpoint. The same code and instructions do not guarantee the same search path. One reviewer follows operating-system behaviour, another follows state lineage, and another follows cleanup and gate semantics. Their differences become useful when the work remains independent long enough for those paths to develop.
Agreement is not the decision rule
An ensemble can sound like a vote: if two reviewers agree, the finding wins. That would have weakened this experiment. Most of its value came from findings reported by only one model.
A lone report still had to be reproduced or traced to reachable behaviour. A repeated report still needed the same treatment. Agreement could increase the reason to investigate, but it did not make the finding true. Nor did disagreement make it false.
Review diversity expands the search. Evidence still decides what survives it.
This also changes the role of reconciliation. Combining reports is not a request for a fourth agent to smooth them into one confident summary. The reconciler has to preserve provenance, recognise when two identifiers describe one defect, and keep unsupported or disputed findings visible until they are adjudicated.
The ensemble was still incomplete
Seven valid defects sounds impressive. It was not exhaustive. Comparison with the later production review showed several consequential boundaries that all three medium-effort runs missed, including freshness, base-tip and durable pre-mutation-record requirements.
I also ran the same reviewers against two smaller changes. On one, all three returned no findings. On the other, Sol reported one credible defect while Terra and Luna remained clean. That restraint matters: an ensemble that produces more findings by rewarding suspicion would only move work into adjudication.
The sample is much too small to establish a stable hierarchy between model families. It is one repository, one main fixture and a handful of additional reviews. Models, tools and prompting will also change. The result supports a workflow choice for consequential work, not a universal model ranking.
Where I would use this
Three reviews are not automatically better than one. They cost more, create more findings to adjudicate and can repeat the same blind spots. For a small, reversible change, one capable reviewer plus deterministic checks may be proportionate.
The ensemble becomes more plausible when a change crosses several failure boundaries, mutates shared state or is hard to reverse. Even there, I would keep it explicit rather than make every pull request pay the cost. The useful measures are not raw finding count or majority agreement, but adjudicated unique findings, missed consequential boundaries, elapsed time and the human effort required to reconcile the result.
This extends the provider portability explored in Free your skills. Let the models compete., but changes the question. Portable skills let several models attempt the same work. Independent evidence tells us whether their differences protect anything that matters.
The benchmark gave me a winner. The more durable lesson was that selecting the winner would have thrown away three real defects.