The aquarium moved. Then I noticed the label on the glass was wrong.

Two days ago, I wrote about building an aquarium for AI: a persistent world where temporary agents could do useful work without borrowing continuity from my head.

I expected the early lesson to be about autonomy. Could the system continue without me? Could I stop reading every line of code and still remain responsible for what it did?

The more immediate lesson was less dramatic. The work could move faster than the system's own account of the work.

What changed in two days

Jörmungandr began as a name, a set of boundaries and a small collection of proofs. It now has 45 commits, 255 tracked files and roughly 22,000 lines across documentation, contracts, fixtures, tests and implementation.

Those numbers describe output, not value. The more useful account is narrower.

Day threeMore has been proved; much has not been built
There is evidence for
  • durable, provider-neutral work contracts
  • the same typed exchange across two communication paths
  • restart reconstruction and idempotent replay
  • bounded retry, dead letters and authority escalation
  • independently verified recovery in two fixture classes
  • a separate assurance facility and richer control view
There is still no
  • production persistence technology
  • scheduler or live dispatch path
  • production lease renewal or recovery policy
  • external authoritative mutation
  • verified delivery through a real provider
  • live projection behind the pane of glass

A bounded proof can be complete while the phase remains in progress and the product still does not exist.

The first two bounded phases are complete. The third has moved through a series of small recovery proofs: claiming work, surviving restart, waiting, retrying, recording dead letters, escalating, verifying a repair and reconstructing why it happened.

This is going better than I expected as an experiment. It is not yet going at all as a production system.

The source of truth was true in several places

The repository is deliberately its own memory. Conversations are disposable. Accepted intent belongs in canonical documents. Code and executable checks establish what has actually been implemented.

That worked. A new agent could enter, query the project and recover enough context to continue. I no longer had to explain the whole idea each time.

But the memory was not one object. It was a collection of documents with different jobs: detailed assessments, a component plan, current-position summaries, a repository introduction and routes telling agents what to read.

The detailed record said the transition runtime had advanced through seven bounded slices. The summary above it still said two. Both sentences lived in the source of truth.

The first visible driftThe record grew faster than its summaries
01ImplementA bounded slice earns executable evidence
02RecordThe owning document gains the assessment
03PropagateSummaries and routes still describe an earlier state
04ResumeThe next participant receives two current positions

The next executor does not need the old conversation. It does need the durable sources to agree about their scope.

This was not a catastrophic failure. The underlying evidence remained available and the disagreement was easy to find. But it exposed a mechanism I had underestimated.

A source of truth does not stay singular because I call it one. Every summary is a cache. Every index is a projection. Every sentence that says “current” creates a maintenance obligation.

AI makes creating those projections cheap. It does not make agreement between them automatic.

“Complete” needed a boundary

The other recurring problem was the word “complete”. Agents like completion. So do project plans. A passing test and a cleanly named commit make a satisfying ending.

But Jörmungandr contains several different endings nested inside one another.

One word, four scopesCompletion only means something with its boundary attached
  1. ProofCompleteOne fixture demonstrates one declared behaviour
  2. Recovery classBoundedThe behaviour holds for the two cases exercised
  3. P3 runtimeIn progressGeneral recovery and production choices remain open
  4. EcosystemNot operationalNo live scheduler, dispatch or external mutation exists

Remove the scope and a true claim becomes a misleading one.

“The recovery proof is complete” can be accurate. “Recovery is complete” is not. The difference is only a few words, but it separates evidence from theatre.

I had already designed evidence, authority and independent verification into the runtime. I had not applied the same discipline consistently to the project's own status.

That may be the most useful test of an AI-native system: do its principles survive contact with its own construction?

Control did not disappear. It changed shape

I started this experiment because I was worried that wanting to understand the complete codebase was keeping me inside the loop.

I have now stopped trying to understand every implementation movement. There is too much of it, and much of it is exactly the kind of mechanical work I wanted temporary executors to absorb.

I am not doing less thinking. I am thinking at a different boundary.

My attention has moved towards the meaning of the system. What is authoritative? Which claim outranks another? What evidence is sufficient? When is a result stale? Who may change external state? What does “complete” mean here?

These decisions are smaller in volume than the code, but they have a larger blast radius. A weak implementation can fail a test. A weak definition of success can make every test pass.

So far, letting go has not meant surrendering control. It has meant moving control out of constant participation and into explicit boundaries that another participant can inspect and challenge.

What I have learnt so far

The persistent world matters more than the intelligence of any individual agent. Changing agents was easy. Preserving exact meaning across their arrival and departure was the real work.

Documentation is not durable memory merely because it is committed. It needs ownership, reconciliation and a way to notice when one current claim has outlived another.

Boundaries have to travel with results. “Passed”, “safe” and “complete” are dangerous when separated from the fixture, revision, authority and limitations that made them true.

And acceleration does not remove coordination work. It changes its form. Less time is spent producing each artefact; more care is needed to keep the relationships between artefacts honest.

How is it going?

Better than I expected, if I describe it as a way of finding the real problems.

Worse than the screenshots suggest, if I describe it as an autonomous software ecosystem.

The project can now demonstrate several things I had only described on day one. It can reconstruct bounded work after a restart, reject stale or duplicated meaning, distinguish a failed delivery from blocked work and require independent evidence before one kind of repair resumes.

It cannot yet take a real piece of work through a production scheduler, dispatch it to a replaceable executor, change an external system and show me the verified result through a live control surface.

The gap between those two paragraphs is the project.

The next thing through the glass

My first instinct was to keep adding movement: another recovery class, another adapter, another route through the world.

I think the more important next step is to make the map harder to leave behind. Current-position claims need to be treated as projections of evidence, not prose that somebody remembers to update after the interesting work is done.

I wanted a system I did not have to hold in my head. Two days in, it has shown me the difference between not holding every implementation detail and not holding responsibility for what the system means.

The first may be freedom. The second would be abdication.