I recently wrote that we should not begin by onboarding a software team to AI. We should onboard AI to the team.

That is easier to say than to make concrete.

What does a small first intervention actually look like? How can it produce useful evidence without requiring a new operating model at the same time? Where should it sit, what should it be allowed to do and what would justify letting it do more?

Having thought about this for a long time, and been asked the question many times, this is how I would approach the challenge for a small software team.

The details are deliberately anonymised. The substance is the design, the sequence of decisions behind it and a bounded way to prove its value through real work.

The approach begins with two apparent opportunities: an early pull-request review and a set of focused testing agents. The more useful idea emerged while trying to fit them into the existing process.

The starting point should not be the agents. It should be the seams where work already changes hands.

The existing process is part of the design

The existing development flow is already recognisable.

A work-management system holds work items, acceptance criteria, code and pull requests. When a pull request is ready for another developer, it is posted in an existing review channel. People review it and retain the final merge decision. Testing happens later, with existing application and data facilities that can support test access and reproducible environments.

There are frustrations in that flow. Review time is valuable. A pull request can wait before another person is available. Testing is a larger constraint, and feedback may return after the change has moved through several stages. By then, the developer has had to reload the context and the process has already paid for work that might have been avoided by an earlier check.

Those frustrations do not make the surrounding process disposable.

As I worked through the proposal, it was easy to drift into redesign. Perhaps the system needed a new ticket state. Perhaps automated UI testing needed a new kind of test account. Perhaps the stable database should behave differently. Each suggestion sounded defensible in isolation. Each also added another change the organisation would have to understand, approve and maintain before the proposed service had produced any value.

The corrective constraint was simple: they already have a system. Work with it.

The review-channel post is an existing readiness signal. The linked work item is the source of stated intent. The repository contains the candidate implementation. The data-copy process provides a known starting point. The existing test-access facility provides the required session. Human review already owns the decision.

These are not obstacles around the AI system. They are its integration boundary.

Work crosses three visible boundaries

Existing workflow
Review channelPR ready
Work systemIntent and criteria
RepositoryCandidate change
Test facilitiesKnown data and access
Bounded AI service
  1. TriggerUse the existing hand-off
  2. Early sanity checkReview before deeper cost
  3. Selected testingOnly when justified
  4. Evidence returnedPR and review channel
Human decision
  1. Developer respondsIssues arrive earlier
  2. Human reviewContext and judgement remain
  3. Merge decisionAuthority stays visible

A pull request readiness signal from an existing review channel and intent from a work-management system trigger a bounded AI service. The service reviews the code and, where justified, builds a disposable environment using existing data and test-access facilities. Findings return to the work system, review channel and developer. A person still reviews the work and makes the merge decision.

Illustrative system boundary, not measured performance. The existing workflow remains the source of readiness, intent, code, test access and final decisions. The service consumes bounded inputs and returns evidence to the same places.

Begin at a hand-off that already exists

The first proposed intervention is deliberately small.

When a developer posts a pull request to the existing review channel, an AI reviewer can pick it up asynchronously and perform an initial sanity check. It can look for plausible defects, regressions, missing tests and consequential questions, then place detailed evidence on the pull request and return a short summary to the channel.

It does not merge the change, reject it or move the work item. It does not replace another developer's review. It attempts to find avoidable problems while the change is still fresh in its author's mind, before another person spends time rediscovering them.

This is a useful first seam because the trigger, destination and human owner already exist. People do not have to remember a new command or watch a new dashboard. If the service is unavailable, the original process can continue. If its findings are poor, its access and output can be removed without reconstructing the surrounding way of working.

That makes the intervention low-disruption and reversible. It does not make it risk-free. A reviewer with repository access can still see sensitive material. Low-value comments can create noise. Confident but incorrect findings can waste developer time. A fast response can feel intrusive if it arrives as a wall of criticism before a colleague has even opened the pull request.

The design therefore has to govern not only what the service can access, but also how it appears and behaves.

Give the service an identity, not a disguise

I would give the service its own identity in the systems it uses.

That identity provides a technical boundary. It can be limited to the agreed work area, repositories and review channel. Its activity is distinguishable from human activity. Permissions can be reviewed, changed or revoked without using a developer's personal account.

These are common platform capabilities rather than assumptions about one particular system. Azure DevOps, for example, supports application identities for background automation and allows permissions to be limited within the platform (Microsoft, 2026). Slack apps similarly use bot identities and granular scopes, with no access by default beyond the permissions and conversations granted to them (Slack, 2026). The precise identity and authentication mechanism would still need to match the host organisation's infrastructure and policies.

The identity also provides a social boundary.

A recognisable name, an obviously non-human profile image and a short description can tell people what has entered the conversation. A consistent voice can make its output easier to scan and less abrasive. That voice can be refined using the same evidence used to refine the technical checks.

I would not try to make the service pass as a colleague. Familiarity should not be used to make people lower their guard. The aim is a recognisable service identity whose behaviour can earn trust.

Its interaction contract might include:

  • use calm, concise and non-judgemental language
  • prioritise possible defects and consequential uncertainty over style
  • distinguish a likely defect from a question or optional suggestion
  • attach enough evidence for a developer to inspect the claim
  • state when the available evidence is insufficient
  • consolidate related findings rather than producing a stream of comments
  • put detail on the pull request and only a short summary in the review channel
  • provide a simple way to mark a finding useful, incorrect or unnecessarily intrusive

Several specialist agents may operate behind that identity, but I would begin with one visible assistant. Three automated personalities independently commenting on acceptance criteria, code and UI behaviour could feel less like support and more like being surrounded. One interface can reconcile their findings, remove repetition and make the next decision clearer.

Tone can reduce friction, but it is not evidence of reliability. A friendly reviewer can still be wrong. Trust should grow from calibrated claims, useful findings, visible limits, correction and dependable behaviour—not from a name or profile picture alone.

Least privilege should apply to the service's social presence as well as its technical access.

Let each stage earn the next

The larger opportunity is to move selected testing closer to the pull request. That should not mean running every possible check against every change from the first day.

I separated the proposed work into three testing roles:

  • Acceptance criteria: compare the implementation with the stated intent and criteria of the linked work item
  • Cold code: independently inspect the change for defects, regressions and missing tests
  • UI testing: exercise affected browser journeys when the change and available environment justify it

The order matters.

Acceptance criteria should come before the more expensive work. If the implementation has missed the intended outcome, or the intent is too unclear to assess, there is little value in building an environment and clicking through it. A code review can then look for implementation problems. Only after those stages produce enough reason to continue should the service create an isolated application and database copy for selected UI testing.

Capability expands only after evidence

  1. Existing seamPR-ready hand-offNo new workflow to remember
  2. First interventionAdvisory sanity checkUseful signal before authority
  3. Decision gateIs the signal worth its cost?No: refine, narrow or stop · Yes: continue
  4. Deeper reviewIntent and codeCriteria before expensive testing
  5. Decision gateIs UI testing justified?No: human review · Yes: create an environment
  6. Selected testingDisposable environmentTargeted UI journeys only
  7. Existing authorityHuman final sayReview, merge or return

An existing pull-request hand-off triggers an advisory sanity check. If its signal is not useful enough, the service is narrowed, refined or stopped. If it is useful, acceptance criteria are checked before code review. Only justified changes receive an isolated environment and UI testing. People retain final review, and returned work feeds improvements back into the service.

Capability is added only when the previous stage produces useful enough evidence to justify more cost, access and complexity.

This is more than a cost optimisation. The gates make failure easier to attribute. If the first review produces noise, there is no need to wonder whether a browser environment or test-data problem caused it. If acceptance criteria cannot be assessed, the system can expose that uncertainty instead of generating a polished test result against invented intent.

A proposed two-month proof of value

I would test this through a short proof of value rather than propose the final system in advance.

During the first month, the aim would be to produce the first advisory pull-request reviews, observe how people respond and refine the supporting triggers, evidence format and interaction style. Work on the environment harness and configuration could begin in parallel, but useful review output should arrive before the infrastructure becomes the visible product.

During the second month, one testing role could be introduced at a time: acceptance criteria, then cold-code review, then selected UI testing. The final period would be used to examine findings, false positives, returned work and whether a more capable system was actually justified.

The measures should cover the complete effect rather than celebrate agent activity. I would want to know:

  • how many findings reviewers and authors considered valid and useful
  • how many were incorrect, repetitive or too minor to justify attention
  • when useful feedback first became available
  • whether human review time or avoidable back-and-forth changed
  • whether earlier findings added work that prevented later rework
  • how often the service could not assess unclear intent
  • whether participants found the timing, volume and tone helpful or intrusive
  • what the service cost to operate, monitor and refine
  • whether a second model found enough important, non-overlapping issues to justify its additional cost

A two-month observation will not establish long-term defect reduction or prove that the approach generalises to another environment. It can establish something more modest and immediately useful: whether this intervention produces enough trusted signal to continue, what it costs the complete workflow and which part, if any, has earned expansion.

The acceptable conclusion may be to keep a useful reviewer at low authority. It may be to change the task boundary, repair part of the underlying process, bring a proven capability in-house or stop. A pilot that can only conclude "adopt more AI" has not been designed as an experiment.

Working with a process is not preserving it forever

There is an obvious objection to this approach. What if the existing process is the problem?

Attaching automation to a poor hand-off can make the poor hand-off faster and more permanent. An undocumented workaround may become an API contract. A queue may receive work sooner without gaining capacity to resolve it. An AI reviewer may reproduce existing blind spots while giving them a more confident voice.

Working with the existing process does not mean treating it as correct or untouchable. It means using it as the starting evidence rather than assuming a new tool has earned the right to replace it.

The proposed service may expose missing acceptance criteria, duplicated handoffs or a testing constraint that earlier review cannot relieve. Those are reasons to revisit the process with the people who own it. They are not permission for an external service to redesign it silently in advance.

Compatibility is the starting constraint. Learning may justify change.

Start with somewhere useful to stand

The first AI intervention does not need to display the full capability of the technology. It needs to give people somewhere useful to stand while deciding what should happen next.

In this example, that means beginning at an existing pull-request hand-off, remaining advisory, returning evidence through familiar systems and keeping a person responsible for the decision. A dedicated identity makes access and authorship visible. A restrained interaction style reduces avoidable social friction. Staged testing gives each increase in cost and authority something to prove.

None of that guarantees success. The proposed reviewer may produce too much noise. The environment harness may cost more than the testing delay it is meant to reduce. Earlier findings may move work without improving accepted flow. Those are questions for the proof of value, not gaps to cover with a more confident proposal.

Onboarding AI into existing work begins by respecting that there is already a system, a language and a distribution of responsibility. Find the seam where a small intervention can return useful evidence. Make its access, identity and authority visible. Then let its behaviour decide whether it is invited any further in.