The first test told me to specify the work. It did not tell me which model I needed to do it.

When Sol received every implementation decision, ordinary prose and a behaviour model both reached full acceptance. That made the value of complete requirements difficult to miss. It did not tell me whether behavioural structure made it safer or cheaper to step down the model ladder.

That was the stronger promise I wanted to test next. Could the same behaviour model help Terra or Luna match Sol while using fewer tokens, less time, fewer repairs or less money? Or would equally complete prose work just as well?

I rebuilt the benchmark around that question: harder tasks, exact fact parity, three GPT-5.6 tiers and measured repair turns. The answer was untidy. Luna matched Sol under both formats. The only final failure came from Terra working from prose. Behaviour changed the cost profile, but it did not produce one consistent advantage.

What the first benchmark could not tell me

In the first experiment, short implementation requests failed while complete prose specifications and behaviour models both passed every task. I found strong evidence for specificity and only a qualified signal for behavioural structure.

It also left the most interesting claim untested: could a structured behaviour model let a cheaper model produce work comparable to a frontier model?

Repeating the original tasks would not answer that. Prose and behaviour had already reached a ceiling. The direct condition also contained fewer implementation facts, so it mixed prompt structure with missing information.

The second version had to make the work harder and the comparison fairer.

A smaller, harder comparison

I created two new JavaScript tasks. One handled subscription renewals, retry scheduling, grace periods, cancellation, duplicate commands and time-dependent behaviour. The other handled atomic stock reservations, optimistic versions, expiry, authorisation, idempotency and deterministic output.

Each task had a manifest of implementation facts. Every fact appeared in both detailed conditions:

  • Prose: complete conventional requirements
  • Behaviour: the same requirements organised as states, transitions, guards, transformations, invariants and boundaries

Exact fact parity passed across all 64 facts. The models saw the same starter code, medium reasoning effort and visible smoke tests. Hidden acceptance tests were added only after a coding turn. A failing implementation could receive one repair turn with normalised failure feedback.

The primary matrix contained twelve cells: two tasks, two representations and three GPT-5.6 tiers—Sol, Terra and Luna— with one run per cell.

One run per cell is calibration, not confirmation. It can show whether a proposed comparison has enough variation to deserve a larger study. It cannot estimate how often any individual result would repeat.

The study had to become smaller before it could continue

I had originally registered 24 cells, including a direct condition and a previous-generation reference model. The first harder cell used 259,449 tokens after repair. Even an initial-turn-only projection put the full study above my four-million-token limit.

Continuing until the limit stopped the runner would have left an incomplete, order-biased matrix. I stopped after that first cell, recorded the failure of the resource estimate and reduced the design to its primary question before continuing.

The amended twelve-cell study retained prose versus behaviour across the three current model tiers and allowed at most one repair. The direct and previous-generation arms were deferred. The final collection used 2.19 million tokens and 43.8 agent minutes including isolation checks, safely inside the original limits.

The model ranking never arrived

Final and first-pass acceptance, two tasks per cell group
ModelBehaviour finalProse finalBehaviour first-passProse first-pass
Sol2/22/21/21/2
Terra2/21/21/21/2
Luna2/22/21/22/2

Behaviour finished six out of six cells at full acceptance. Prose finished five out of six. That looks like a behaviour advantage until the individual runs are restored to the table.

Sol tied under both formats. Luna tied under both formats and passed both prose tasks first-time. The whole final-quality difference came from one Terra prose run on stock reservation, which remained at 13/14 after repair.

Mean final prose acceptance was therefore 100% for Sol, 96.4% for Terra and 100% for Luna. Capability did not degrade in model order. Behaviour removed the observed Sol–Terra gap, but there was no Sol–Luna gap to narrow.

A one-cell anomaly is a reason to investigate, not permission to draw a trend line through it.

Behaviour won one comparison, not the experiment

The resource measures were just as uneven. Behaviour was clearly better on Terra: higher final quality, 20.6% fewer total tokens, 19.7% fewer non-cached tokens, 10.9% less time and 15.6% lower estimated API cost.

On Luna, prose was better on most implementation-efficiency measures. Behaviour used 5.7% more total tokens, 40.5% more non-cached tokens, 7.4% more time, one additional repair and 18.1% more estimated cost. Quality tied.

Sol split the difference. Behaviour used 13.2% fewer total tokens, but 30.9% more non-cached tokens and 7.8% more estimated cost. Time and repairs were effectively equal.

Aggregate implementation resources, six cells per condition
MeasureBehaviourProseBehaviour relative to prose
Total tokens1,008,8721,134,75411.1% fewer
Non-cached tokens305,640263,84215.8% more
Elapsed time21.54 minutes21.70 minutes0.8% less
Repair turns321 more
Estimated API cost$1.4371$1.42400.9% more

The aggregate keeps the same ambiguity. Behaviour reduced total processed tokens and tool calls, but used more non-cached tokens, required more repairs and cost slightly more at the frozen rates. It saved about ten seconds across all six runs.

“More efficient” needs an object. A format can reduce context processed while increasing uncached work. It can reduce tool calls while creating another repair. It can be faster without becoming cheaper by an amount that matters.

Repairs were part of the result

Five of the twelve cells needed a repair. Those turns consumed 407,871 tokens—19% of all valid implementation tokens. A study that stopped at first-pass acceptance would capture quality but miss a material part of the cost of reaching usable software.

The lone final failure is instructive. Terra working from prose first missed a command-boundary requirement. The repair fixed that failure but violated the separately specified rule that duplicate identifiers must be checked before every other guard. The score stayed at 13/14, but the failing behaviour moved.

That was not a missing fact. Both requirements were present in both prompt formats. Nor was it a harness failure. It was a local repair that corrected one boundary while disturbing another.

This is closer to real maintenance than a first-pass-only benchmark. Completion includes the route through feedback, not merely the quality of the first generated patch.

The specification still mattered most

The second experiment did not repeat the direct-request arm, so it cannot independently reproduce the first experiment's specificity result. But it did test what happened after the implementation facts were held constant.

Under that constraint, the lowest-priced model matched the frontier model's final quality. The behaviour representation did not reliably make that possible: Luna's prose results were at least as strong and cheaper to obtain.

One possible mechanism is that the requirements removed enough uncertainty for model capability to stop being the binding constraint. The agents still had to implement interacting behaviour, but they did not have to invent the contract while doing it.

That interpretation remains bounded by two tasks. Different repositories, longer horizons or requirements that demand more architectural judgement may restore a clear tier gradient. This calibration cannot locate that boundary.

A behaviour model can still be the better artefact

Prompt efficiency is not the only reason to model behaviour. States, guards and invariants can make a system easier to inspect before implementation. A durable model can be reviewed, versioned and reused when the code changes.

A behaviour model may be a better software artefact without being a cheaper implementation prompt.

That distinction matters because the estimated construction effort was higher for behaviour: roughly 30–40 additional minutes across the two task specifications. The observed implementation-time difference was nowhere near enough to recover it on one use. Reuse may change that calculation, but this experiment did not measure reuse.

Why I am not confirming this result yet

The calibration has enough variation to be interesting and not enough to justify multiplying the same matrix. Eleven of twelve cells ultimately passed. The expected model-tier gradient did not appear. The apparent behaviour advantage is concentrated in one non-monotonic Terra result.

A useful next experiment would add independently authored tasks that produce repeatable first-pass variation across tiers. It would keep exact fact parity, treat first-pass quality and repair burden as co-primary outcomes, and add repetitions only after the tasks demonstrate that they can distinguish the models.

For now, I would make two practical choices. Write requirements completely enough that the implementation contract is not left to inference. Then test the least expensive suitable model on representative work instead of assuming its place in a product name predicts its place in the result.

Use a behaviour model when its structure helps people inspect, change and preserve the design. Do not promise that the format itself will make implementation cheaper. In this experiment, it sometimes did. Just not reliably.