I read a LinkedIn post and could not tell whether I had found a new development method or a name for something familiar.

I came across the idea in a LinkedIn post about behaviour-first agentic coding. It described writing the expected behaviour first—states, transitions, data transformations and boundaries—then letting a coding agent derive the implementation from that model. The model could be inspected, changed, versioned and reused.

My first reaction was not disbelief. It was recognition without clarity. This sounded close to the systematic agentic design I was already trying to do, but perhaps with the behaviour promoted into a more durable artefact.

Was it a development system, a specification template, a software model, or simply the familiar work of resolving requirements before coding? I left a comment asking for a concrete before and after—and whether the claim of fewer tokens had been measured.

I did not want to invent what the author meant or pretend my interpretation described their practice. But the question was testable. So I built a small pilot around the part I could examine: what changes when the same coding agent receives an ordinary brief, a complete prose specification or a structured behaviour model?

What I thought might be changing

Several different ideas were bundled together. The behaviour model might improve the representation of decisions. It might help mainly because it contains more decisions than the brief it replaces. It might reduce implementation work by moving that work into specification. Or it might create a reusable model whose value appears across later changes rather than the first implementation.

Those mechanisms have different consequences. A template can make a prompt easier to write. A behavioural model can become a versioned source of intent. A development system needs to connect that intent to implementation, verification and future change. Calling all three “better prompting” would hide the interesting part.

Three ways to ask for the same software

I created three bounded JavaScript tasks: an inventory-image manifest, an approval workflow and a webhook-delivery queue. Each began with a minimal repository, a public interface, one visible smoke test and a hidden acceptance suite.

For each task, the coding agent received one of three prompts:

  • Direct: a concise feature brief resembling an ordinary ticket
  • Prose: every registered requirement fact in ordinary prose
  • Behaviour: the same facts as prose, organised as states, transitions, transformations and boundaries

The last two prompts had mechanical fact-ID parity. That distinction matters. Behaviour versus direct tests the whole behaviour-first workflow, but mixes two changes: more decisions and a different representation. Prose versus direct isolates specificity. Behaviour versus prose is the closer test of the representation itself.

The pilot used one frontier coding model at the same medium reasoning setting for every run. Two repetitions of each task and condition produced eighteen valid comparisons.

The obvious result was specificity

The short briefs performed badly. Their median hidden-test score was 11.8%, and none of the six implementations passed a complete suite.

Both detailed conditions passed everything: six out of six for prose and six out of six for the behaviour model.

Corrected pilot results, six valid runs per condition
ConditionMedian scoreAll-pass runsMedian total tokensMedian time
Direct11.8%0/6126,289216.4 seconds
Prose100%6/6130,890173.3 seconds
Behaviour100%6/6109,211156.0 seconds

That is a striking difference, but it is not evidence that state-and-transition notation caused the reliability gain. Ordinary prose containing the same decisions performed identically.

The safer conclusion is simpler: an agent cannot reliably implement decisions it was never given. When a brief leaves behaviour open, the model has to invent it. Capable models can produce plausible software, but plausible is not the same as specified.

The token result depends on which tokens we mean

Behaviour and matched prose achieved identical correctness, while the behaviour condition used 16.6% fewer median total implementation tokens and finished about 10% faster.

Against the short brief, behaviour-first used 13.5% fewer median total tokens and finished 27.9% faster. The token reduction missed my pre-registered 15% threshold, while the time reduction was substantial.

There is an important wrinkle. The provider reports cached input separately. When I subtracted that cached subset, behaviour used 4.4% more median non-cached tokens than matched prose, not fewer. Against the short brief it still used 22.2% fewer.

“Fewer tokens” is incomplete without saying whether cached context is included and whether the concern is model work, latency or price.

Total processed tokens and non-cached tokens are both useful views. Neither is a direct measure of human reasoning, and a token count alone is not a monetary price.

The benchmark failed before it helped

My first detailed runs repeatedly scored 7/8 and 9/10. At first glance, that looked like a small but stable model failure. On inspection, the problem was in my specification.

One inventory fact described trimming a CSV header as a whole, while the hidden test accepted whitespace around each header cell. An approval fact required non-empty fields but did not say that invalid values must throw a TypeError, although the evaluator required it.

The agents had implemented what I wrote. The tests measured what I meant.

I corrected the fact inventory, recorded and excluded the affected results, and reran only those cells. All eight replacements passed completely. The discarded results remain in the experiment rather than disappearing from the story.

That feels central to the behaviour-first idea. A versionable software model is valuable because it makes decisions inspectable. It can also make an incorrect decision look unusually authoritative. Moving reasoning earlier does not remove reasoning. It moves both leverage and risk upstream.

What the experiment cost

I can report implementation execution cost in tokens and elapsed agent time. I cannot honestly report the whole route: I did not instrument the human and model effort used to design, review and correct the fact inventories and prompts.

Recorded collection cost; this is usage, not a monetary bill
ScopeTotal tokensNon-cached tokensAgent time
18 valid pilot runs2,196,731601,59556.4 minutes
All result-bearing attempts3,478,149981,38183.4 minutes

The second row includes harness failures, a contaminated run and results invalidated by the specification corrections. It is the more honest account of what collection consumed, even though those attempts do not belong in the comparison.

At the valid pilot's average, the planned forty-five-run confirmation would require roughly another 5.5 million tokens and 141 sequential agent-minutes before model differences, exclusions or human work. That estimate is useful for deciding whether to continue; it is not a price quote.

Why I am not adding more repetitions yet

This pilot does not test the smaller-model claim. It also does not establish whole-route savings. More importantly, prose and behaviour have both reached a correctness ceiling on these tasks. Repeating the same experiment would make that ceiling more precise without revealing whether behavioural structure can improve correctness over equally complete prose.

Before a second version, I would change four things:

  1. use independently authored ordinary briefs that better represent strong everyday practice
  2. use harder tasks that leave room between detailed conditions
  3. run direct, prose and behaviour conditions on both model sizes
  4. record specification-production time and tokens, then automate correction turns and tighten execution isolation

The pilot has done its job: it validated the fixtures, exposed defects and estimated the cost of a larger study. It should not be promoted into a general benchmark merely because its first result is interesting.

What I think it is now

I no longer think the useful distinction is the diagram or template. Nor do I think behaviour-first coding is wholly separate from careful software design.

The meaningful shift is treating expected behaviour as a first-class artefact: something that can be inspected, changed, versioned, reused and checked before implementation. If you already make states, transitions, transformations and boundaries explicit before asking an agent to code, you may already be doing much of it.

What makes it more than a prompt template is the role the artefact plays afterwards. If behaviour is changed before the code, versioned with the system, reused by later agents and checked against the implementation, it begins to function as a model of the software. If it is written once and discarded after generation, it is closer to a particularly good brief.

The pilot supports the value of doing the specification work. It offers a promising but qualified signal that behavioural structure can encode the same decisions more efficiently than prose. It does not show that the structure itself improved correctness, and it does not yet support the broader claims about smaller models or lower end-to-end cost.

And it reinforced a less glamorous lesson: when code is generated from a model of the software, the model belongs inside the system we test.