Not every ecosystem begins with a grand design. Mine began with a text file and a blank space for a repository URL.

I used to keep my most useful prompts in files. The stable instructions sat at the top; variables such as the repository, branch or task waited at the bottom. When I needed one, I copied it into a session, filled in the blanks and set the AI moving.

This was more reliable than remembering the wording each time. It also gave me the small satisfaction of having automated a process without having automated anything at all.

Then skills arrived.

I moved roughly the same prompts into skill folders and congratulated myself on entering the future. The files had metadata. The agent could discover them. I no longer had to perform the ancient ritual of copy, paste and replace all.

In truth, I had discovered packaging.

A prompt the agent can find

A skill is easy to describe as a reusable prompt, and that is a useful place to begin. It preserves instructions for a class of work instead of asking someone to reconstruct them in every conversation.

The current OpenAI documentation describes a skill as a package of instructions, resources and optional scripts. An agent initially sees the skill's name and description, loads its full instructions when the work matches, and can also be told explicitly to use it ( OpenAI, Build skills).

That discovery changes more than convenience. My old prompt library depended on me knowing the right file existed. A skill can announce what it is for and enter the work when its description matches the request.

But discovery does not make the instructions good. A prompt placed inside a skill folder can still be vague, brittle and dependent on the author standing nearby to interpret its results. Mine were.

The skills became useful enough to expose their own weakness. They worked repeatedly, so their inconsistencies stopped looking like isolated AI moments and started looking like design problems.

Every frustration added a responsibility

The first problem was output. Two runs could complete the same task and return results with different structure, evidence and levels of detail. Both might look plausible. Only one might be usable by whatever came next.

So I added output standards tailored to the work. A review skill should distinguish evidence from inference. A research skill should expose sources and uncertainty. A skill producing an artefact should say where it saved it and what it verified. “Done” was no longer an adequate interface.

Then input became the problem. A carefully written workflow still produces poor work when the task arrives without the repository, source material, decision boundary or required outcome. I added input checks and guards: not bureaucracy for its own sake, but a way to stop confident work beginning from a malformed request.

A reusable prompt had acquired an interface. It knew something about the inputs it could accept, the outputs it owed and the conditions under which it should pause rather than improvise.

This was the first meaningful change. The skill was no longer only storing what I wanted the agent to do. It was making some of my previously hidden expectations visible.

Some instructions were badly disguised programs

As the skills repeated more work, another pattern appeared. I was asking the model to rediscover procedures I already understood: walk these files, extract these fields, compare these structures, validate this format, produce the same report shape.

Natural-language instructions were useful while the process was still being discovered. Once a step became stable, mechanical and testable, asking a model to improvise it on every run added variation without adding judgement. Some of my prompts were programs waiting for me to admit it.

I began scripting those parts and having the skill execute the scripts. This produced two benefits. The first was context. The agent needed to understand when the operation should run, what it should receive and how to interpret the result; it did not need to carry or reconstruct every implementation step in its working context.

The second was consistency. The same operation could take the same path, report failures in the same shape and produce an output the next stage knew how to consume. I was moving stable mechanics out of probabilistic conversation and into something I could test directly.

The scripts also exposed a mistake I had made while removing the blanks from my old prompt files. I had not always removed the variables; I had hidden them in code. Repository names, paths, URLs, branch conventions and output locations had been baked into scripts because they were true in the project where the script was born. The assumptions remained invisible until I tried to use the same skill somewhere else.

Cross-project use forced me to sanitise those scripts. Genuine inputs became explicit parameters or travelled with the project. A value could be discovered from an authoritative local source, but if it was missing or ambiguous the script had to validate and stop rather than guess. I had removed the blanks from the bottom of the prompt only to bury them somewhere much harder to see.

This was not an argument against every default. A safe default is visible, genuinely shared and cheap to reverse. A baked-in project assumption silently changes the target. The difference matters when a consistent script can now be consistently wrong in several repositories at once.

A third benefit appeared in terminal-facing skills: a trusted script could create a narrow boundary around sensitive data. The agent could request an operation such as configure, authenticate or verify. The script could get or set the value through a credential store, use it internally and return only success, failure or non-sensitive metadata. The model did not need to receive the secret to get the job done.

That boundary had to be real. The secret could not be printed to standard output, placed in command-line arguments, written to model-readable files or copied into logs. A script the agent could freely rewrite was not much of a guard either. The useful pattern was a small, reviewed capability with restricted behaviour: the AI could invoke the operation, but not inspect or extract the value it operated with.

This did not mean scripting everything. Ambiguous research, interpretation and decisions still benefited from the model. The current OpenAI authoring guidance makes a similar distinction: prefer instructions until deterministic behaviour or external tooling justifies a script. Nor did determinism guarantee correctness. A script can perform the wrong operation with admirable consistency. It simply made that operation explicit enough to test, version and hold accountable.

Four stagesIndependent value became system leverage at the hand-offs
  1. 01Saved promptReusable instructions; I supplied the context
  2. 02Independent skillBounded work with an explicit contract
  3. 03Chained skillsOne dependable output became the next input
  4. 04OrchestratorSelection, ordering and recovery became adaptive
Governed byBaselines · representative work · whole-route metrics · system review

One skill to rule them all

My next instinct was to keep adding instructions to the skill that already worked. Each failure produced another paragraph. Each exception produced another branch. The prompt grew references, scripts, validation, tools and opinions about which agent should do the work.

Before long, my reusable prompt had acquired staff. I had not improved a text file. I had accidentally given it a department.

The useful evolution was not prompt, then ever-larger skill. It was prompt → skill → chained skills → orchestrator. Each skill first had to earn value independently. Only then did joining them together multiply rather than merely combine their problems.

Consistent, expected inputs and outputs made the hand-offs possible. Statelessness made them safer. By stateless, I do not mean that the work had no state. I mean that a unit should not quietly depend on the residue of a previous conversation, a machine-local memory or an assumption only its author knows. Required state should enter through a declared input and leave in an inspectable output.

Once that was true, one skill's result could become the next skill's starting point without me standing between them to translate. The value of each unit remained testable on its own, while the chain created a capability none of the units provided alone. That was the point at which the collection became unexpectedly powerful: not because any individual skill had become cleverer, but because the hand-offs had become dependable.

A fixed chain could run a known sequence. An orchestrator added a different job: select the right units, order them, pass only the state each needed, handle failure and decide whether the combined result was sufficient. This was where two different kinds of skill became visible. Some were units of work; others coordinated the route through them.

Responsible compositionThe chain passes declared state, not hidden context
OrchestratorSelect · order · recover · judge
Declared inputSkill ABounded unit of work
Expected outputInspectable hand-off
Declared inputSkill BIndependently testable
Not allowed as undeclared dependencies
  • Previous chat
  • Machine-local memory
  • Author-only assumptions

I use those as working labels rather than a universal taxonomy. The distinction matters because the two do not grow or fail in the same way. A unit can often be improved by making its responsibility narrower and its result clearer. An orchestrator must also be judged by what it delegates, what it retains and whether the pieces still add up to the mission.

Chaining did not remove risk; it gave risk a faster route through the system. A malformed but plausible output could now become a valid-looking input downstream. The contracts, guards and measures mattered more after composition, not less.

Treating an orchestrator as a large prompt hid its real architecture. Treating every instruction as a separate unit produced the opposite problem: a beautifully divided system spending much of its time explaining itself to itself.

The seam is part of the design

Breaking work apart can protect the main agent's context, isolate specialist reasoning and make a failure easier to locate. It can also create duplicated reading, lossy summaries, more tool calls and another boundary at which responsibility becomes ambiguous.

I learnt not to treat decomposition as a virtue in itself. A useful boundary contains a coherent responsibility that can be completed and tested independently. Split it further and the cost of coordination can exceed the benefit of isolation. Leave it larger and unrelated work begins competing for the same context and instructions.

Like database normalisation, both too little and too much can leave the system harder to operate. Unfortunately, skills do not arrive with a helpful line saying “normalise to here”. The seam has to be found by watching where context, quality and responsibility are lost.

I began using lazy-loaded references so detailed guidance entered only the work that needed it. I gave bounded research or implementation tasks to subagents so their discovery did not automatically consume the mission context. This was one of the experiments behind The context compression frontier.

Subagents were useful, but not free. Every delegated task needed a sufficiently complete brief, its own context and a result the main agent could evaluate. For a while, every problem acquired a subagent. Then every subagent needed instructions, supervision and someone to read its report. I had recreated management with a token counter.

The end of “this feels better”

I was still changing skills by taste. I would shorten an instruction, move material into a reference, try another model or split work across agents. If the next result looked cleaner, the change felt successful.

That was tolerable while the prompt was a personal shortcut. It became irresponsible once other skills, agents and decisions depended on its behaviour.

The problem with “better” is that it conceals the dimension being improved. A shorter main prompt may trigger more searches later. A subagent may protect one context while increasing the total tokens, time and hand-offs required to finish. A cheaper model may match the average score while failing more often on the cases where mistakes matter most.

So I began establishing a baseline before making material changes. I assembled representative tasks, defined the evidence or characteristics a good result should contain, ran the existing version and compared the proposed version against it. Where model variation mattered, I repeated the work rather than building a conclusion around one fortunate run.

The measures depended on the skill. I commonly looked at task success, recurring failure types, consistency, input and output tokens, elapsed time, tool calls, retries and the amount of human intervention needed to trust the result. Accuracy was not one universal number; it was a judgement tied to the work and its consequences.

The benchmark did not tell me what to value. It stopped me pretending that a preference was evidence.

Measure where the work moved

Local metrics can make almost any architectural change look successful. Move discovery into another context and the main agent becomes smaller. Remove validation and the skill becomes faster. Route work to a cheaper model and the visible call costs less.

The whole system may have become slower, more expensive or less dependable. The cost has merely moved out of the box being measured.

I therefore started measuring the complete path through the work: what entered, which actors became involved, what they consumed, what they returned, which checks ran and how much correction remained for me. The aim was not to minimise every number. It was to make the trade visible.

That changed model choice too. Not every unit of work needs the agent currently leading the session or the most capable model available. If a smaller or cheaper model produces the required outcome with comparable reliability, using it can preserve expensive capacity for decisions that benefit from it.

But that is a property of the measured capability, not the model's reputation. Another skill may need a different leader. A later model release may move the boundary. This is why I now try to benchmark the capability rather than bind it to the brand.

The ecosystem needs declared dependencies

As skills spread across agents and environments, hidden inputs became more dangerous. A machine-local memory file, a personal instruction or an old conversation could make the same skill behave differently for reasons absent from the work itself.

I came to treat that kind of memory as toxic to a distributed skill system—not because memory is inherently bad, but because required context hidden outside the capability is an undeclared dependency. It can make a weak design look reliable on the machine where it grew up.

If correct work depends on context, that context should travel with the work in an inspectable form: repository instructions, a skill reference, a shared contract, a test case or another source the participating agents can discover.

Models and providers change underneath these systems too. A skill can appear flaky when it describes the route one model happened to follow rather than the outcome that must remain true. Clear intent, boundaries and evidence do not prevent change. They make it possible to tell whether the capability survived it.

The harness is part of the system

There was another maintenance loop I had initially missed. I was using the harness to review changes to skills, but the harness itself was now made of contracts, checks, fixtures, routing rules and assumptions. It was no longer standing outside the system. It was part of it—and therefore capable of being wrong with excellent formatting.

I began reviewing the whole arrangement for three recurring problems: contradictions, where two rules expected incompatible behaviour; gaps, where a responsibility fell between skills; and duplication, where the same policy or procedure had been copied into several places and could drift independently. Passing the current tests did not prove those tests still described a coherent system.

That review needs a stable account of intent without freezing every implementation. I use a shared skills ledger for that separation. The ledger records the neutral capability intent and the differences that must remain visible. Skills and scripts stay free to use the implementation best suited to their provider or project, as long as they continue to satisfy the contract.

This made the review bidirectional. Results could expose a weak implementation, but implementations could also expose missing, duplicated or contradictory intent. The ledger, the skills and the harness each became evidence about the others; none was allowed to become unquestionable simply because it was used to judge the rest.

Accountability grew with the leverage

Metrics are not governance on their own. A fixed evaluation set can be overfitted. A pass rate can hide an important class of failure. Cost and speed are easy to count precisely even when they are not the qualities that matter most.

Human judgement therefore remains in the loop, but it has a better job. Instead of approving a change because the result feels nicer, I can state what I expected to improve, inspect what changed, accept the trade-off deliberately and keep the failures that matter in the next comparison.

This is the accountability I did not know my text files would need. The more work a skill can influence, the less reasonable it is to maintain that skill as private prompt craft. Its boundaries, dependencies and changes become part of the system other work relies on.

My prompts did not become skills when I moved them into new folders. They became skills when their responsibilities became explicit. They became a system when one skill's output became another's input, when agents began coordinating around them and when a quiet edit could alter the whole route through the work.

Prompts grow up gradually. Responsibility should not arrive late.

The text file survived. It just acquired dependants.

Further reading