The first warning was a skill that stopped being dependable even though I had not changed it.
It had worked repeatedly. Then a later model followed the same instructions differently, chose a less reliable route through the tools or stopped at a point the earlier model had pushed through. The skill still looked correct. Its behaviour had changed underneath me.
This has happened often enough that I no longer think of a skill as a prompt with some supporting files. It is a relationship between an instruction, a model, a provider's tools and the host that joins them together. Any one of those can change while the words in the skill remain still.
At first I repaired each failure where I found it. A stronger instruction here. A new reference file there. Another check to stop a route that had become unreliable. The fixes worked, but only for the provider and model in front of me.
Then I began using more than one provider. Claude and Codex are my current examples. Both are capable. They are not interchangeable.
One may understand a difficult change more quickly. The other may notice the edge case during review. Planning, code, documentation and investigation do not always favour the same model. Even within one category, the result depends on the repository, the tools available, the context supplied and the particular shape of the task.
Using both increased my leverage. It also multiplied the maintenance. I wanted some skills to exist everywhere and others to remain provider-specific. The common ones lived in more than one place. I copied improvements manually, usually when I remembered, and tried to work out whether a difference was deliberate or simply drift.
The drift rarely announced itself. It waited until the next real task, when a familiar capability behaved differently depending on which agent happened to run it. The more useful the collection of skills became, the more effort it took to trust that collection.
Yet I did not want several providers merely for availability. I wanted several AI perspectives on the same problem: one to propose, another to challenge, perhaps a third to find the assumption the first two had shared. At its best, the process felt less like asking a tool for another answer and more like reviewing a difficult problem with a council of capable peers.
That council was only useful if its members were allowed to think and work differently. Giving every provider a harness designed by the same agent made maintenance easier, but it also made the perspectives less independent.
The temptation is to turn those observations into an exclusive routing table: use this model for planning, that model for reviews, another one for documentation. That throws away part of the reason for having several perspectives. A model that is weaker on average can still notice an edge case the current leader misses. Different training, context handling and tools can expose different parts of the problem.
A benchmark can help decide who leads and how much overlapping coverage a task deserves. It should not decide who is allowed to contribute. I did not need a permanent winner or one opinion per task. I needed a system in which the lead was allowed to change while independent perspectives remained available.
The route from useful skill to durable capability
- 01Make today's model reliableBuild a skill around what works now
- 02Watch the relationship changeA later model or provider follows a different route
- 03Duplicate and synchroniseManual copies drift as the skill collection grows
- 04Separate truth from executionShare the contract; preserve routes, overlap and substitution
The leverage created its own maintenance
The obvious way to control that maintenance was to let one provider write and maintain the skills for every other provider. One agent would understand the canonical version and project it everywhere else.
The maintenance became simpler. The council became less diverse.
A skill written for one agent carries more than its stated objective. It carries assumptions about how tools are invoked, how context is discovered, how subagents return work, how browser sessions are isolated, which metadata triggers the skill and what the host considers a successful completion.
Asking that agent to design the other providers' harnesses can produce several implementations. It can also reproduce the first agent's assumptions several times.
I found this particularly unhelpful in code review. The value of asking Claude and Codex to examine the same change was not that they could repeat one checklist twice. It was that they could bring different strengths and failure modes to the same standard of evidence. When one provider defined the other's review harness, some of that independence disappeared before the review began.
A second opinion is less useful when the first opinion wrote its instructions.
Complete separation failed too
The opposite answer is complete separation: independent instructions, skills, references and tests. That preserves difference, but it creates another problem. The separate harnesses drift.
One agent learns that a destructive action must pause. Another does not. One review requires the remote head to be checked before publishing a result. Another reviews a branch that has already moved. A definition changes in the documentation but survives only in an old provider-specific memory.
So the choice is not between one universal harness and several unrelated ones. The useful boundary sits inside the harness.
Some things describe what the work means. Others describe how a particular provider performs it. I had been trying to synchronise both.
Share the contract, not the route
I began treating each skill as a small shared contract with a provider-native adapter for every integration.
The contract holds the durable intent: when the capability applies, what outcome it must produce, which actions are forbidden, what evidence must be preserved and what counts as complete. Shared reference files hold repository and product knowledge. Shared scripts, schemas and test fixtures execute checks that should not acquire a Claude or Codex personality.
Each provider then owns its adapter: the skill description, invocation syntax, tool choreography, context strategy, subagent mechanics and the way it presents results. A third or fourth provider can add another adapter without asking either of the first two to become its designer.
Portability cannot mean identical files
A provider adapter is not a cosmetic translation. As I write this, Codex assembles repository instructions through an AGENTS.md instruction chain, discovers skills from locations such as .agents/skills, and can attach OpenAI-specific metadata. Claude Code uses CLAUDE.md and CLAUDE.local.md, supports imported and nested context, and gives skill frontmatter control over invocation, tools and isolated execution. Those mechanisms will continue to evolve.
These differences decide when a reference enters the context, which instruction wins, whether the model or only the user may invoke a capability, what tools it can use and how another agent is brought into the work. Flattening them into one file format would create the appearance of portability while discarding useful provider behaviour.
I found a small example of this drift while revisiting my own system. My Claude adapter carried an invocation rule in prose because Claude did not expose an equivalent control when I wrote it. Current Claude skill frontmatter provides explicit invocation controls. The capability had not changed, but its best adapter had. That is the boundary I now try to protect: a shared capability contract, a provider-native adapter and model-specific tuning and evaluation.
One contract, many actors, several routes
Claude and Codex are the examples from my work, not a boundary around the design. Any provider that exposes enough of an API, command-line tool or integration surface to run the capability can participate. Two providers may be enough. Three or more can add coverage, capacity or a specialist route. The contract does not need to know which one will be best next month.
The ledger carries only the difference
I use a small shared ledger—a neutral manifest—to record facts that cannot live reliably inside one provider's skill file.
A provider may encode invocation policy as metadata while another needs prose. One may support interface labels and default prompts that do not exist elsewhere. A skill may be deliberately available to one provider and absent from another. These are facts about the relationship between the adapters, not facts either adapter can express alone.
The ledger does not contain another copy of the complete skill. That would create a third implementation to maintain. It holds only the neutral intent and the mapping information required to tell whether the adapters remain in step.
A read-only comparison checks the adapters against that intent. It detects missing skills, unexpected provider scope, conflicting descriptions, forked shared assets and policy that has disappeared during translation. Detection is shared; correction belongs to the provider whose adapter is wrong.
Let each provider write its own side
The rule that made the largest practical difference was surprisingly simple: one writer for each provider adapter.
Claude may improve the Claude adapter. Codex may improve the Codex adapter. Either may reveal that the common contract is missing something, but neither silently rewrites the other agent's working method.
This prevents a synchronisation loop in which each agent sees unfamiliar syntax, “fixes” it according to its own mechanics and leaves the other side looking wrong again. It also changes how capability develops. The agent using a skill can refine it, benchmark it and test its failure modes with the tools that will actually execute it.
Shared changes still cross the boundary. If a code review must verify the exact commit, preserve stable finding identities or publish only one assurance result, that belongs in the common contract. Whether an agent uses a particular subagent type, browser surface or prompt structure to satisfy it does not.
Hidden memory is an undeclared dependency
I will be blunt: I hate MEMORY.md and any hidden AI instruction or system prompt as a foundation for software development. Providers will have internal instructions I cannot control. My own process should not depend on them. What feels like helpful adaptation on one task is invisible configuration on the next machine. It cannot be reviewed with the change, and its absence is rarely reported as an error.
I work across multiple machines, and many of my agents run in fresh sandboxes. A machine-local or provider-local memory can make the same repository, skill and task behave differently depending on where it runs. That hidden difference can cause havoc precisely because it is difficult to reproduce, inspect, test or remove.
Hidden memory is an undeclared dependency. If the work needs it, put it where the team, the next machine and the test harness can see it.
Same repository and task, two context models
- Same repository and task
- Machine or provider-local memory changes
- Behaviour diverges without a declared error
- Repository-owned agent file
- Lazy-loaded, versioned references
- Provider-native route proves the shared contract
My alternative is a small, healthy, provider-native agent file kept with the repository. It tells the agent what to load and when instead of trying to contain everything. Durable decisions live in versioned reference files that can be lazily loaded for the task. The shared ledger holds portable capability intent. Provider-specific files adapt the route without hiding the rules.
At most, I treat automatic memory as a disposable cache for non-governing preferences. Architecture, safety rules, release policy and the definition of done belong in source control. If an agent must know something to produce a correct result, it is not memory. It is part of the system.
| Shared across providers | Allowed to vary by route | Measured continuously |
|---|---|---|
| Capability intent and boundaries | Instruction files, discovery and precedence | Evidence coverage and missed findings |
| Versioned authoritative reference files | Reference discovery and lazy-loading mechanics | Correction and recovery work |
| Evidence schemas and reusable checks | Skill packaging, metadata and invocation | Latency, tokens and monetary cost |
| Safety constraints and stopping rules | Tools, subagents and provider optimisation | Performance on representative tasks |
| Representative tasks and acceptable quality bar | Usage limits, cost and available capacity | Substitute quality, later verification and continuity |
| Collaboration standard and evidence format | Access, preferences and environment capabilities | Cross-provider completion and review |
Coverage is not consensus
The architecture describes where differences belong. It does not establish that several routes will improve the result. Multiple routes can improve coverage, but agreement between them is not proof.
Research on language-model ensembles gives reasons for both interest and caution. Wang and colleagues found that sampling several reasoning paths and selecting the most consistent answer improved performance across the arithmetic and reasoning benchmarks they studied ( Wang and colleagues, 2023). Those were benchmark questions using repeated samples from a model, not autonomous software agents, but the result supports a modest idea: one plausible path need not exhaust the useful search space.
Results about mixing different models are less tidy. An early Mixture-of-Agents preprint reported gains from layering outputs from several models ( Wang and colleagues, 2024). A later preprint found that repeatedly sampling the strongest model often outperformed mixing models on its tested benchmarks, while also identifying cases where mixture helped ( Li and colleagues, 2025). Neither establishes how a multi-provider coding workflow will behave. Together they challenge the easy assumption that adding models automatically adds value. Diversity and capability both need to be measured.
Software engineering has an older warning. Knight and Leveson asked 27 teams to develop versions of the same program from a common specification, then ran one million tests. More than one version failed on the same input substantially more often than an independence model predicted ( Knight and Leveson, 1986). The study involved student-written programs for one defined problem, not language models. Its enduring warning is narrower: separate implementations can still make correlated errors.
Claude and Codex can agree because the requirement is clear. They can also agree because the same incomplete reference file led both in the same direction. Provider diversity cannot recover evidence the system never supplied. A shared contract needs independent challenge, observable evidence and human judgement—not a majority vote dressed as assurance.
Benchmark the capability, not the brand
Once the capability is separated from the provider, model selection becomes a revisable operational decision.
I can run representative reviews, documentation changes, planning problems and code tasks through each route. I can compare whether required evidence was found, whether important risks were missed, how much correction followed and what the complete task cost in time and money. The benchmark belongs to the capability. A model is one current candidate for leading or contributing to it.
I also want to measure marginal coverage, not only rank each route in isolation. What useful finding did a second model add? Which shared assumption did it challenge? Where did it create noise rather than insight? A lower individual score does not imply zero value in combination, just as a collection of high individual scores does not guarantee useful diversity.
This avoids two forms of lock-in: the reputation a model earned from work it no longer performs best, and the routing table that turns a temporary lead into exclusive ownership. A provider should lead or support a role because current evidence supports it, not because one good month hardened into policy.
The tests should remain proportionate. Not every task needs two providers, an ensemble and a human panel. Low-risk, reversible work can use one route that clears the evidence gate. Ambiguous or consequential work may justify another perspective, an independent review or a second implementation. The important part is that overlap remains easy to add, and that adding or removing a provider does not require redefining what good work means.
A second provider is also spare capacity
Different perspectives are not the only benefit. If I hit my token or usage limit with Codex, I can give the next suitable task to Claude, or the other way around, rather than waiting for that provider to become available again.
I do not expect identical output. Because both adapters work from the same capability contract, references and evidence gates—and have been tested on representative tasks—I can have moderate confidence that either will reach a similar acceptable level of quality. They may take different routes or expose different edge cases. That difference is part of the value.
Usage limits, cost, latency and available capacity therefore belong beside quality when work is routed. The preferred model can lead while it is available. Another proven provider can take the next suitable task when it is not. Consequential work may still justify additional review, but work that only needs moderate confidence does not have to stop because one provider's allowance has run out.
When the allowance resets, I can bring the original model back to test or verify the work completed in its absence. That closes the loop: the substitute keeps new tasks moving, the returning provider supplies an independent check, and the findings improve the benchmarks used next time. I can keep moving at greater scale without requiring one model to be permanently available or treating the substitute's quality as an article of faith.
Keep the next suitable task moving
Availability is not evenly distributed
Usage limits are only one kind of interruption. I have had Codex become unavailable and a Claude Code update behave badly in one of my environments. Because each provider has its own adapter, I can identify the degraded route, pivot and continue working without redesigning the capability around the outage or regression.
The same separation matters inside a team. People have favourite tools. Some have access only to Claude, some only to Codex, some to both, and others use another integration entirely. Requiring everyone to use the same provider can make or break collaboration for a reason that has little to do with the quality of their contribution. They need the same visible intent, references and evidence standard, not necessarily the same agent.
Collaboration should require a shared standard, not a shared favourite provider.
Not every actor is human, and not every sandbox is equal. Environments differ in network access, installed tools, permissions, browser state and available integrations. A provider-native adapter can work with the environment available to that actor while still proving the same result. Forcing one exact route across every environment makes the system look consistent while making it brittle.
This is what I think of as fake lock-in: dependence created by scaffolding that assumes the same provider, file layout, tools or hidden state everywhere. Where the capability does not genuinely require those assumptions, liberating the actors—people and agents—to use the route available to them creates practical flexibility without weakening the shared standard.
Portability is a testable property
It is easy to call a harness provider-agnostic because its top-level prompt contains no model name. The stronger test is substitution.
Can another provider discover the authoritative context? Can it execute or replace the shared checks? Can it produce evidence that the next part of the system understands? Can the original provider disappear without taking the definition of success with it?
The answer does not need to be yes for every capability. Some skills may legitimately depend on a provider-specific tool. Lock-in is not automatically wrong. Unexamined lock-in is the problem: a dependency hidden inside local state, syntax or one agent's undocumented workaround.
The shared ledger makes the intended scope visible. The comparison makes drift visible. Independent benchmarks make current capability visible. Substitution tests make continuity and quality trade-offs visible.
Free the system, not the prompt
I began with a familiar question: should I use Claude or Codex?
The more useful question was: which parts of this capability must survive a change of model or provider?
The answer was not a perfectly neutral prompt. It was a system boundary. Intent, constraints, references and evidence are durable. Each adapter can follow its provider's native skill, tool and orchestration conventions. Tests keep those different routes attached to the same meaning without requiring them to become identical.
Today's strongest model can lead the work. Tomorrow's can take the lead. Other models can still challenge, review or expose an edge case. A specialist provider can lead one capability without excluding other perspectives or becoming the architecture for all the others.
The same freedom applies to people and agents. They do not need identical provider access or identical sandboxes to contribute. They need a visible common standard and a route capable of satisfying it.
Free your skills from the model that made them. Then let every provider prove what it can do.
Further reading
These books offer related lenses rather than direct proof of this particular multi-provider design. I am including them as a reading list to examine, not as endorsements of every claim they contain.
- Design Rules, Volume 1: The Power of Modularity, Carliss Y. Baldwin and Kim B. Clark—how stable design rules allow independently designed modules to evolve and compete
- Building Evolutionary Architectures, Neal Ford, Rebecca Parsons, Patrick Kua and Pramod Sadalage—fitness functions and guided change when today's design must remain able to evolve
- AI Engineering, Chip Huyen—evaluation, model choice and the wider application stack around foundation models