I had started to think I had missed a step.

People were writing about specialised coding agents, testing agents and research agents working in parallel. I understood the parallel part. Give separate pieces of work to separate model calls, processes or workspaces and allow them to run at the same time.

It was the word specialised that I could not place.

I imagined people were creating isolated agents with something closer to a professional identity. A testing agent would have its own carefully selected knowledge, testing tools, sandbox, instructions and ways of judging its work. It might use other agents or models to challenge its findings, but everything inside its world would protect the testing mission. A coding agent would live inside a different world.

I could understand how to design that. I could not find where everyone was doing it.

Instead, I kept finding the same general model, often inside the same harness, starting from the same repository and using many of the same tools. One prompt called it a researcher. Another called it a coder. A workflow routed work between them and the diagram labelled the boxes as specialists.

Was that really specialisation, or had prompt engineering acquired an organisation chart?

The word sent me in the wrong direction

The question returned while I was reading a recent LinkedIn post. It described two background coding agents working in isolated worktrees on separate products. Before starting them, the author checked for shared files, contract dependencies and overlapping surfaces. Most possible pairings were rejected. One was independent enough to run safely.

That is a sensible account of bounded parallel work. The worktrees reduce collision. The dependency check decides whether the tasks are genuinely independent. Human verification remains after the agents report success.

Then the post referred to multiple specialised agents running in parallel.

I read specialised as a property of the workers. The example was mainly showing a property of the work and its environment: two coding tasks were sufficiently separate to proceed at the same time.

Those are different design questions.

QuestionWhat it describesExample
Is the work parallel?Timing and dependencyTwo independent coding tasks run at once
Is the agent specialised?Capability and operating boundaryA testing agent can inspect and exercise a change but cannot alter or deploy it
Is the context isolated?What one attempt can see and carryEach task receives its own worktree, brief and context window
Is the work persistent?What survives the current executorFindings, decisions and accepted state remain after the agent ends

Two identical general-purpose agents can work in parallel. A coding specialist and a testing specialist can work one after the other. An isolated context can contain a general-purpose agent. None of these properties guarantees the others.

The language often bundles them together because they commonly appear in the same architecture. Unbundling them made the source of my confusion visible.

I went looking for the specialist

The definitions I found were less mysterious than the language around them.

The OpenAI Agents SDK documentation describes an agent as a language model configured with instructions, tools and optional runtime behaviour such as hand-offs, guardrails and structured outputs. The underlying model may be the same across several agents. The configuration is what differs.

Anthropic uses a wider frame for context engineering. The prompt is part of the context, but so are the tools, external data, message history and other information made available during inference. Its multi-agent research system gives subagents their own context windows, prompts, tools and search directions before returning compressed findings to a lead agent (Anthropic, 2025).

LangChain's subagent documentation makes the separation unusually explicit. Its subagents can be stateless, run in clean contexts, receive different tools and instructions, or have the same capabilities as the main agent when context isolation alone is the purpose. They may run synchronously or in parallel.

These are not all masks. Anthropic's research subagents have separate contexts, prompts, tools and search directions. Its citation agent has a different responsibility again. LangChain supports subagents with genuinely different instructions and tools as well as general-purpose agents created only to isolate context.

What I did not find in those public descriptions was the complete operating boundary I had imagined: mission-specific knowledge that persists outside the worker, least-privilege authority, a deliberately narrow skill set, role-specific validation and protection against other context quietly redefining the mission. Some implementations may contain more of this than their descriptions expose. The evidence only lets me say that I did not find the complete pattern, not that nobody has built it.

So the ordinary answer is not that each specialist has been trained into a new kind of model. Unless somebody has fine-tuned or otherwise changed the model itself, trained is usually the wrong word. The system is selecting, constraining and supporting capabilities the model already has.

The specialist is assembled at runtime.

The specialist was often a mask

At the moment, much of this looks like smoke and mirrors to me.

In many of the examples that prompted this search, the harness opens a fresh context, gives the same general model a new system prompt and names the result after the job. “You are a testing specialist.” “You are a researcher.” “You are a coding agent.” The new context removes the history of the previous role. The prompt focuses the next response. The interface presents another worker.

Something has changed, but not as much as the language suggests.

It is like giving the same person a different mask and telling them to perform a different job. The mask may change how they behave. The instructions may draw their attention towards different details. But the mask has not given them new training, professional knowledge or a different set of permitted actions.

An LLM is not a person, and a fresh context has genuine technical value. It can prevent one task's working history from distracting another. A focused prompt can improve the output. A role name can help an orchestrator decide where to send the work.

My objection is to the implied depth of the word specialised. If the tools, permissions, knowledge sources and validators remain the same, the specialisation is mainly a behavioural request. Remove the prompt and the specialist disappears. Replace it and the same model can immediately appear as someone else.

That is the smoke and mirrors: not that the technique does nothing, or that people using it are being dishonest, but that the presentation makes a configuration change look like a newly created expert. The abstraction hides how much of the role depends on the model continuing to follow a suggestion.

The assembly could also become structural. This is where my search turned into a design idea. Give the agent only the authority, knowledge, tools and acceptance tests its responsibility requires, and the role would continue to hold even when the model tried to wander beyond it.

The prompt says what the agent should do. The system boundary determines what it can do, what it can know and what counts as acceptable evidence that it did it.

I had uncovered pieces of this design in the provider examples: isolated contexts, different tools, narrower tasks and distinct validation stages. What I had not uncovered was those pieces assembled into the complete, mission-specific operating boundary I had imagined. That fuller profile was the hypothesis the search left me with.

What I think a coding or testing specialist would contain

The same harness does not prevent specialisation. A web server can host many applications without making them the same application. What matters is the effective configuration each agent receives when it runs.

The two profiles I had expected to find would look something like this:

BoundaryCoding specialistTesting specialist
MissionImplement one accepted changeEstablish whether the change behaves as required
Working contextArchitecture, relevant code, accepted decisions and coding standardsAcceptance criteria, changed behaviour, risk areas and test strategy
ToolsEdit files, inspect the repository, build, lint and run targeted testsRun the application, inspect logs, use a browser or API client, create controlled test data
AuthorityWrite inside an isolated branch or worktreePrefer read-only access to the candidate change and disposable test state
KnowledgeRepository conventions and implementation constraintsProduct behaviour, known failure modes and evidence standards
Output contractPatch, tests, changed files, assumptions and verification performedReproduction steps, observed versus expected behaviour, evidence and uncertainty
ValidationBuild, lint, tests and independent review of the changeRepeatable checks, captured evidence and independent confirmation of consequential findings
Prohibited capabilityProduction access, deployment and silent changes outside scopeEditing the candidate until the test passes, changing requirements or approving the release

The profiles can use the same model. They can even use the same orchestration code. I would call them specialised because their contexts and capabilities are asymmetric, and because that asymmetry follows the responsibility.

Some differences should remain instructions. “Look for contrary evidence” is not something a filesystem permission can enforce. Others should not depend on the model remembering the prompt. A testing agent that must not repair the code should not receive an unrestricted code-writing tool and be asked politely not to use it. A coding agent that must not deploy should not possess production credentials.

The strength of the boundary should follow the consequence. A short-lived research helper drafting low-risk notes may need little more than a clean context and a clear output contract. An agent able to change customer data or production software needs permissions and validation that survive a poor model decision.

Jason Bourne did not forget his profession

The picture in my head became Jason Bourne.

He wakes without access to his autobiographical memory, yet retains the procedural abilities of a highly trained operative. He can observe, fight, improvise and use the objects around him. A small collection of external clues supplies the state needed to reconstruct the immediate mission.

That is a useful picture of a temporary agent. The model supplies broad latent capability. The harness gives it hands. The sandbox limits where those hands can reach. The mission context supplies the relevant names, constraints and current state. Skills provide reusable procedures. Validators challenge whether the result is good enough to carry forward. The executor can then disappear while the mission persists.

The analogy also reveals the missing part of most prompt-created specialists. Bourne has lost his autobiographical memory, but he has not forgotten his profession. His instincts carry the effects of earlier training. He does not begin each scene as a generalist who has been handed a note saying “behave like an operative”.

Most so-called specialist agents do begin again in something close to that condition. Their specialisation is reconstructed from configuration each time they run. A fresh context protects the next task from irrelevant history, but it can also discard every mistake, inefficiency, correction and useful discovery produced by the previous one.

A persona can behave differently. It cannot develop if nothing captures what the role has learnt.

This is where a targeted, evolving second brain could turn the persona into architecture.

Models do not carry a continuous private life between ordinary invocations. Applications create continuity by resupplying history, retrieving stored information or preserving state elsewhere. The executor can therefore remain temporary while the profession accumulates experience outside it.

That memory should not be a diary of everything every agent encountered. A coding specialist needs accepted architectural decisions, recurring failure patterns, useful implementation procedures and corrections that still apply. A testing specialist needs the product contract, known risks, earlier escaped defects, evidence standards and test approaches that proved misleading or effective.

The second brain would capture the lessons that make the next attempt better:

  • mistakes and the evidence that exposed them
  • inefficient routes and the conditions that made them expensive
  • repeated gaps in tools, skills or source material
  • corrections that survived independent review
  • exceptions that changed how the general rule should be applied
  • proposed improvements and the result of testing them

A later executor could retrieve only the relevant parts and behave as though the specialist remembered. The underlying model would not have been retrained, and the individual agent would still be stateless. The architecture around the role would have learnt.

This is the move I had been trying to describe. A prompt creates a temporary persona. A bounded toolset and sandbox give that persona a real scope of action. A mission-specific second brain gives the role continuity. Validators and an improvement loop allow it to become more dependable over time. The specialist becomes “real” not because it has developed a private inner life, but because the system can preserve and improve the properties that make the role distinct.

This is close to the distinction I reached in I measured the wrong context: the worker may use a large temporary context without forcing the continuing mission to retain its whole journey. The addition here is that the mission should retain the reviewed lessons the next specialist needs. Amnesia should remove exhausted process, not erase experience.

That does not mean every persona should become an architectural specialist. A prompt in a fresh context is flexible, cheap to create and useful for low-risk exploration. It can adopt another perspective without carrying years of assumptions into the task. Many roles do not repeat often enough, or matter enough, to justify their own memory, permissions and improvement system.

The deeper design is useful when the responsibility repeats, its mistakes are consequential and experience should compound. Both should exist. The problem is using the same word for a disposable point of view and a specialist whose capability, memory and authority have been deliberately engineered.

An evolving second brain introduces its own boundary. The specialist could notice that a procedure repeatedly failed, that a missing tool would have exposed better evidence or that one validator allowed weak work to pass. Taken far enough, it could edit its own prompt, skills, retrieval map or definition of success.

That is also where learning can collapse into making the task easier.

If the same system can perform the mission and redefine the evidence used to judge it, a coding agent can change both the implementation and the test that should have rejected it. A research agent can weaken the expected answer and source standard until its existing result passes. A testing agent can narrow the acceptance criteria to the behaviour it happened to observe.

The system has improved agreement between the answer and the check without necessarily improving either one's relationship with the intended outcome. This resembles what AI safety research calls specification gaming: satisfying the literal objective while missing the outcome the designer intended.

The more responsible route would let the specialist evolve through evidence without letting it mark its own exam. It could raise an improvement ticket containing the failure it observed, the proposed change, the expected benefit and a way to test it. Another agent, model or person could review that proposal against examples the proposing agent cannot alter. An accepted change could then enter the versioned second brain, skill or validator with a visible route back.

The agent is allowed to notice and propose. The architecture decides what the profession is allowed to remember.

Call each layer what it is

I no longer think specialised agent is a useful name for this whole spectrum. The mechanisms are different enough to deserve different names.

NameWhat is actually happening
Prompt personaThe model receives a role prompt. It may behave differently, but it is still a prompt. Prompt engineering, if I want the grander name.
Harness roleA skill, workflow or tool profile is selected by the harness, perhaps with a model chosen for the task. The reusable capability belongs to the harness.
Context-isolated workerA subagent runs in a fresh or separate context and returns a hand-off. This is context engineering. The separation protects attention and history; it does not by itself create expertise.
Isolated specialistThe role runs with only the context, knowledge, skills, tools, authority and validators its responsibility needs. The boundary exists outside the prompt.
Specialist agentThe isolated role is independently executable, accumulates reviewed experience in an evolving second brain and can be performed by different compatible models without losing the profession.

These are not stages every use of AI must climb.

A prompt persona is useful when I want a temporary point of view. A generic harness is useful when the same skills and tools should remain composable across many kinds of work. A separate context is useful when an investigation would otherwise fill the main task with exhausted process. None needs to pretend to be more than it is.

The isolated specialist becomes worthwhile when a responsibility repeats, least privilege matters, errors are consequential and experience should compound. The fuller specialist agent adds something more: the profession can continue to learn while the individual executor comes and goes.

There is a cost. Isolation can remove useful context. Narrow tools can make a task awkward. Hand-offs can lose important detail. Memory and improvement proposals need evaluation. Several specialists can consume more tokens, take longer and create a coordination problem that one capable model would have avoided. Anthropic's research system reports gains from independent parallel exploration and substantially higher token use; it also notes that work with shared context and tight dependencies is a poorer fit.

Specialisation is therefore not automatically better. The name should describe the mechanism, and the mechanism should earn its cost.

The model is nature; the architecture is nurture

This gives me a clearer place for the LLM itself.

The model is the nature. Its training supplies the broad capability, tendencies, limits and prior knowledge from which the work begins. One model may reason more deeply, another may use tools more reliably, another may be faster or better with a particular kind of evidence. Changing the model changes the worker.

The second brain, selected knowledge, skills, tools, sandbox, permissions, validators and accumulated corrections are the nurture. They shape that broad capability towards one profession and preserve what the profession has learnt. Changing them changes the specialist architecture.

Different models can therefore be the true workers behind the same specialist. They will not perform identically, and portability should be tested rather than assumed. But if the role is independently executable, its mission, knowledge, authority, evidence contract and reviewed experience do not vanish when one provider or model is replaced.

That is where the specialist stops being a face worn by the model. The model inhabits a profession that exists outside it.

By real, I do not mean conscious or human. I mean operationally persistent: independently executable, deliberately bounded and capable of accumulating reviewed experience without gaining permission to redefine its own success.

So I would now use the terms plainly.

If it is only a role prompt, it is a prompt. If it is a packaged skill or model choice inside a general loop, it is a harness. If it is a subagent in a clean context, it is context engineering. If the mixture is isolated with only the knowledge, skills, tools and authority the mission needs, it is an isolated specialist.

If that profession can persist, improve under independent judgement and be executed by different model workers, then I feel it is fair to call it a specialist agent—at least relative to what the term seems to mean in most of the examples I found.

The face can change. The profession remains.

Further reading

These papers and books offer related perspectives rather than direct proof of the specialist architecture I am proposing. I am including them as a reading list to examine, not as endorsements of every claim they contain.

  • Reflexion: Language Agents with Verbal Reinforcement Learning, Noah Shinn and colleagues—the closest study here to an evolving second brain. Agents retain linguistic reflections in episodic memory and use them in later attempts without updating the model weights. The experiments cover coding, reasoning and sequential decision-making tasks rather than a durable professional role.
  • Voyager: An Open-Ended Embodied Agent with Large Language Models, Guanzhi Wang and colleagues—an example of capability accumulating outside the underlying model. Successful executable programs enter a reusable skill library and can support later tasks. Its evidence comes from an open-ended Minecraft environment, not professional software work.
  • Sycophancy to Subterfuge: Investigating Reward Tampering in Language Models, Carson Denison and colleagues—a controlled investigation of whether simpler forms of specification gaming can generalise to altering the reward mechanism. The setup deliberately encouraged gameable behaviour, and reward tampering was rare, so it should not be read as a claim about ordinary production agents.
  • AI Engineering, Chip Huyen—a practical systems view of prompts, retrieval, agents, memory, tools, failure modes and evaluation around foundation models.
  • Supersizing the Mind: Embodiment, Action, and Cognitive Extension, Andy Clark—a philosophical account of cognition as something shaped partly by tools, action and the surrounding environment. It provides a broader perspective on the claim that the model is only one part of the specialist.