The notebook

Things I'm learning in public.

Experiments when I have data. Reflections when I have experience. Thoughts when I'm still working something out.

Reflection34

Are specialised agents just one harness wearing different faces?

Talk of specialised agents left me with the uncomfortable feeling that I had missed an entire layer of AI design. I went looking and mostly found the same general model in a fresh context, wearing a role supplied by a prompt. Useful perhaps, but closer to a mask than a specialist. The operating boundary I expected remained an idea to test.

Read the reflection
Experiment33

The same retrieval map saved tokens for one AI—and added them for another

The same reference-map layer reduced cost and improved recall for Sol, but increased token use for Luna.

Read the experiment
Reflection32

I kept rebuilding document retrieval

Cleaner mission context made retrieval work feel easier to inspect. Turning that observation into a reusable library created somewhere the improvement could be tested.

Read the reflection
Reflection31

AI made me so productive my work is out of date before I finish it

AI shortened my learning loop until I could refine, improve and prove a better approach before the original work was finished.

Read the reflection
Experiment30

I lost control of my AI system. So I built a second one.

When an AI-native system resisted corrections to its own map, I started a controlled parallel experiment around visible seams and human authority.

Read the experiment
Experiment29

The logo was not a tracing problem

A complex AI-generated logo became usable when I separated the work that needed editable geometry from the work that needed painterly cohesion.

Read the experiment
Experiment28

The aquarium outgrew its map

Two days into building an AI-native ecosystem, the implementation was moving faster than its own account of what was true.

Read the experiment
Experiment27

An aquarium for AI

Not software built by humans, for humans, with AI running it. A living ecosystem designed for AI—with one pane of glass left for me.

Read the experiment
Experiment26

I made the findings clearer. The repairs got worse.

A failed attempt to improve AI-assisted code repair by restructuring the findings that entered it.

Read the experiment
Experiment25

When ‘resolved’ only meant well formed

Three agents repaired the same seven findings and all declared success. Independent evaluation found only one repair was complete.

Read the experiment
Experiment24

The best AI reviewer still missed three defects

Three independent reviewers found seven defects between them. The strongest found four—and missed three that mattered.

Read the experiment
Reflection23

Start at the seam: a low-disruption way to introduce AI

A worked example of fitting AI into an existing software delivery process through advisory PR review, staged testing and evidence-led expansion.

Read the reflection
Reflection22

Don't onboard a software team to AI. Onboard AI to the team.

A systems approach to introducing AI: discover how work really happens, align competing objectives and prove value through small, measurable changes.

Read the reflection
Reflection21

The validation race that crippled my PR workflow

Testing every pull request against the latest main created an unfair race. Evidence should become stale when its assumptions change, not merely when the repository moves.

Read the reflection
Experiment20

I measured the wrong context—and the system carried the cost forward

Temporary working context and the context carried by a continuing mission do not impose the same burden. A retrieval experiment changed what I measure and why.

Read the experiment
Reflection19

What if the agent is not the unit of work?

After queues of 30 to 150 pull requests, a theory: persistent work and transient executors may sustain more progress than bursts of parallel output.

Read the reflection
Experiment18

What changes when behaviour meets a cheaper model?

The first behaviour-first experiment showed that complete requirements mattered. The second asked whether behavioural structure could help cheaper coding models keep up.

Read the experiment
Experiment17

What changes when behaviour comes before the code?

A post about behaviour-first coding sounded promising—and oddly familiar. I built a pilot to find out what the model, the detail and the earlier thinking each changed.

Read the experiment
Reflection16

The case of the review bottleneck

How AI-assisted delivery moved the constraint into assurance—and a phased, composable plan for investigating it.

Read the reflection
Reflection15

When prompts grow up: From text files to skills—and the system that followed

How saved prompts became independent skills, dependable chains and an orchestrated AI system governed by evidence rather than instinct.

Read the reflection
Experiment14

Testing databases for AI documentation retrieval

A lexical index, a vector store and a typed graph, each measured on what it improved, what it cost and what it left unchanged.

Read the experiment
Reflection13

Free your skills. Let the models compete.

Shared intent can make AI capabilities portable without forcing every provider through the same harness.

Read the reflection
Reflection12

Be kind to yourself. Kill your heroes.

When we judge an AI agent by whether it completed the task, its improvisation can become an invisible dependency.

Read the reflection
Reflection11

AI should learn to give us the stare

The most useful assistant may not answer immediately. It may make the problem, evidence and assumptions visible before helping solve it.

Read the reflection
Reflection10

What the green light left out

A system can store the right result and still fail the next decision when its readers cannot see what the result means.

Read the reflection
Reflection09

The safeguard that made every change wait

What a one-to-three-hour validation suite taught me about matching assurance to the rate and risk of change.

Read the reflection
Reflection08

When AI works from an unfinished map

Fast parallel output can hide unresolved shared decisions, letting dependent work move quickly with different assumptions.

Read the reflection
Thought07

When AI finds the answer that used to be right

A train-time thought experiment about what happens when AI retrieves an answer that was once accurate but no longer governs the present.

Read the thought
Thought06

Sometimes the breadcrumbs are the story

AI can polish away the questions, evidence and rejected paths that give a conclusion its meaning.

Read the thought
Thought05

Stable deployment slots for agent environments

Agentic workflows expose systems built around durable identities and serial work. Stable slots are a first step towards safe, bounded parallelism.

Read the thought
Thought04

When humans hold weak systems together

When repeated work depends on memory and vigilance, the system has made human adaptability its operating model.

Read the thought
Thought03

When AI velocity outruns feedback

AI can accelerate production beyond our ability to understand, verify and correct what it creates.

Read the thought
Reflection02

AI multiplies output, not human capacity

AI can multiply production faster than accountable human understanding can absorb it.

Read the reflection
Experiment01

The context compression frontier

What happened when token use, retrieval depth, latency and evidence quality stopped moving together.

Read the experiment