I keep being asked how a software team should begin adopting AI.
The organisations differ. So do the products, pressures, tools and levels of experience. The question often arrives with a familiar mixture of commercial urgency and understandable caution. Customers may begin to expect faster delivery. Competitors may be using AI already. Development and testing are asked to keep pace without sacrificing quality. Some people are enthusiastic; others are uncertain about the technology, its reliability or what it might mean for their work.
There is no responsible answer that begins with a preferred model and applies unchanged to every team.
Giving everyone access to a chat tool is easy. Building a way of working that produces accepted value without quietly increasing risk, review burden or exhaustion is a systems problem.
There is an older lineage behind that claim. In 1951, Eric Trist and Ken Bamforth examined the social and psychological consequences of changing the technology and organisation of coal mining. Their study helped establish the sociotechnical view that the technical and social parts of work have to be understood together (Trist and Bamforth, 1951). A coal face is not a software team, and the study does not validate a modern AI-adoption method. It does give this approach an important starting principle: changing the tool changes the surrounding work, relationships and responsibilities too.
My high-level approach has four phases: discover the system that already exists, align the objectives around it, build shared readiness and introduce AI through bounded experiments. Evidence from each phase determines whether the next step is justified. Implementation then feeds new evidence back into discovery.
This is not a maturity model in which every team is expected to move towards maximum automation. Some work should remain human. Some work is better served by conventional automation. Some AI assistance may prove more expensive to verify and maintain than the task it was meant to improve.
The purpose of the approach is to find out which is which.
Discovery begins with the real system
Before deciding what AI should do, I want to understand how the team currently produces software.
I would begin with the visible material: architecture and product documentation, wikis, runbooks, release notes, work-item templates, coding standards, test plans and standard operating procedures. Their existence is useful, but it does not establish that they are current, complete or followed. A document describes an intended system. Discovery has to find the operating one.
That means speaking with the people doing the work and observing a selection of real tasks. Where do they look for information? Which source do they trust when two sources disagree? What has to be copied between systems? Which steps are routinely skipped, repeated or reconstructed? Where does work wait for a decision, environment or approval?
I would also look for manual, repetitive work that consumes time without making meaningful use of the team's judgement. It is tempting to call every such task an automation opportunity. First, I want to understand why the task exists, what makes an outcome acceptable and what happens when it is wrong. Automating a poorly understood activity can preserve its weakness while making it operate faster and at greater volume.
The less visible part of discovery concerns knowledge and process dependency:
- information that lives mainly in one person's memory
- work that only moves when a particular person is available
- unwritten checks performed by experienced developers or testers
- release knowledge reconstructed from commits, tickets and conversation
- quality activities that are valuable but often skimmed under pressure
- systems whose outputs need interpretation before anyone can act on them
I have written separately about what happens when humans hold weak systems together. In discovery, the point is not to label knowledgeable people as bottlenecks or extract everything they know into a model. Their adaptations may be the reason the system works. The useful question is which routine burdens the system should carry so those people can spend more of their attention on judgement, exceptions and improvement.
The first milestone is therefore not an AI deployment. It is a credible map of the current system: its information, tools, repeated work, dependencies, controls, frustrations and points of human judgement. It should distinguish what is documented from what is observed and identify uncertainty rather than filling every blank with an assumption.
If that map cannot be reviewed and recognised by the team, the discovery phase is not complete.
The sequence has some independent support in current AI governance guidance. The US National Institute of Standards and Technology's voluntary AI Risk Management Framework organises work around governing, mapping, measuring and managing risk. Its mapping outcomes include documenting intended benefits, monetary and non-monetary costs, scope, practitioner proficiency and human oversight (NIST AI RMF 1.0, 2023). The framework is broader than software-team onboarding, its functions are not a prescribed sequence and the 1.0 framework is being revised. The useful agreement is narrower: context, benefit, cost and responsibility need to be made explicit rather than inferred from the tool's capability.
Objectives need to become visible before they become metrics
Once the current system is understood, I would ask what the business, management and team each want to improve.
Those answers may overlap without being identical. The business may be responding to competitive pressure or customer expectations. Management may want more predictable delivery and clearer visibility. Developers may want less repetitive work, better context and room for deeper technical thinking. Testers may want earlier involvement, stronger acceptance criteria and fewer late surprises. Individuals may be wondering whether AI will help them or become a new way to compare their output.
None of these perspectives is automatically wrong. The danger is allowing one to become the unstated definition of success.
If the business says "faster", the next question is: which part of the system, measured from where to where? Faster generation can create slower review. More code can create more testing and integration work. Shorter time to first output can coexist with longer time to an accepted outcome. I explored one version of this in The case of the review bottleneck: accelerating candidate production moved the constraint into assurance.
Alignment does not require pretending the objectives are the same. It requires recording the tensions and making an explicit decision about the outcome a pilot is meant to improve. A useful objective is narrow enough to test and broad enough to include the work displaced elsewhere.
"Use AI to complete more tickets" is not enough. "Reduce the time and rework needed to turn an unclear work item into testable, accepted scope, without reducing team confidence or moving clarification downstream" gives the team something it can investigate.
This is the second milestone: the sponsor, management and participating team can describe the intended outcome, important constraints and unacceptable costs in shared language. Any material disagreement is visible. The team also knows how a decision will be made when the objectives conflict.
Without that agreement, a metric will not resolve the conflict. It will merely give one side of it a number.
Education is part of the operating system
A team does not need to become a group of AI specialists before it can begin. It does need enough shared understanding to use the technology consistently, challenge its outputs and recognise when it is no longer making useful progress.
I would first establish a simple baseline. Which tools have people tried? What worked? Where did they encounter poor or unreliable results? How confident do they feel? What are they concerned about? What do they currently believe the tools can and cannot do?
The answers shape the education. A practical foundation would usually include:
- what current AI tools are useful and unreliable at
- how to give a task an appropriate boundary and enough relevant context
- why confident language is not evidence of a correct answer
- how to verify results against the system and the intended behaviour
- how long conversations can lose, compress or pollute context
- how to recognise repeated retries, rabbit holes and AI wheel-spinning
- privacy, security, intellectual-property and access boundaries
- when to stop, restart, change approach or ask for human input
- when a deterministic script or conventional automation is the better tool
The language matters. Generation is not completion. Output is not accepted value. Delegation is not abdication. A tool that can perform an action has not automatically earned the authority to perform it in the team's system.
Education should also make room for fear without turning every concern into resistance to be overcome. People may be worried about job security, becoming accountable for output they do not understand, losing hard-won craft or being measured against colleagues who use different tools and tasks. Those concerns affect whether people will report failures honestly. A pilot that makes doubt unsafe will receive flattering data until its weaknesses surface somewhere more expensive.
Amy Edmondson's 1999 multimethod field study examined 51 work teams in one manufacturing company. Team psychological safety was associated with learning behaviour, which in turn was associated with performance (Edmondson, 1999). One company and one study cannot tell us how every software team will respond to AI. It does support the practical reason for making uncertainty and failure discussable: a pilot depends on people being willing to ask for help, challenge output and report what did not work.
How people prefer to work matters too. Some developers may want AI close to the terminal, editor and repository. Other people may work more effectively in a desktop application, browser or mobile interface. Testers, product staff and technical writers may need different integrations from developers. Tool selection should follow the team's tasks, access boundaries and working practices rather than asking everyone to imitate the most enthusiastic user.
Company-managed accounts with an agreed provider can create a common starting point and clearer data boundaries. Multiple providers may eventually offer useful specialisation and resilience, but they also multiply administration, cost, interfaces and the knowledge needed to use them well. I would add that complexity when there is an evidenced reason, not as the price of beginning.
Finally, education needs a humane boundary. AI can increase the amount of candidate work a person can generate or supervise. It can also increase context switching, verification demand and the feeling that work should happen at maximum speed continuously. People still need time to think, understand, learn and recover. A sustainable system should not turn a tool's availability into a permanent output expectation.
Lisanne Bainbridge described a related problem in industrial process control: automation can expand rather than eliminate the operator's difficulties, particularly when people retain responsibility for abnormal conditions after routine participation has been removed (Bainbridge, 198390046-8)). The paper predates generative AI and concerns a different environment. Its enduring warning is that automating activity redistributes human work; it does not guarantee that the remaining work becomes easier or safer.
Agree the measurement before the result exists
The last part of readiness is agreeing what will be measured, how it will be captured and where the results will be visible.
This belongs before implementation. Choosing a measure after seeing the result makes it too easy to celebrate whatever moved and ignore the cost that moved somewhere else.
For a bounded workflow, I would consider:
- the current time and effort required for the complete task
- elapsed and active time with the proposed workflow
- how often the output is accepted, corrected or rejected
- human review and verification time
- defects, rework or quality changes visible within the observation period
- time spent designing, teaching, maintaining and supporting the workflow
- licence, model and infrastructure costs
- repeated attempts, context use and other signs of inefficient task design
- confidence, cognitive load, frustration and perceived control
Not every measure belongs in every pilot. Collecting data has a cost, and a measurement people cannot apply consistently may create false precision. For example, asking someone to estimate how long a task would have taken without AI after completing it with AI produces a convenient comparison, but not necessarily a dependable one. Observation of a small baseline may be more useful than a large collection of guesses.
Some valuable workflows will not save time immediately. A read-only review may surface missing acceptance criteria, a risky assumption or an untested edge case. Its immediate effect could be to add work. The possible benefit arrives later as clearer implementation, fewer defects or less rework. If the pilot is judged only by minutes removed from the first stage, the quality control will look like failure precisely when it is doing useful work.
Qualitative feedback has to be treated as evidence too, with appropriate limits. Did the workflow make the task easier to understand? Did it help the person notice something consequential? Did it create more confidence for a good reason, or only because the output looked polished? Would they choose to use it again? What made them stop trusting it?
The measures should evaluate the intervention and the system, not rank individual developers against one another. Individual comparison would distort the pilot, discourage honest reporting and confuse differences in work with differences in capability.
A poor metric can do more harm than good. A hidden one is worthless. The team should be able to see how success is being judged, how the data is interpreted and what the result will be allowed to change.
Donald Campbell gave the measurement problem a sharper boundary while writing about evaluation of planned social change. He warned that the more a quantitative indicator is used for social decision-making, the more it becomes subject to corruption pressure and can distort the process it is meant to monitor (Campbell, 1976). An AI pilot inside a software team is much narrower than the public programmes Campbell examined. The mechanism is still relevant: once adoption, ticket count or time saved becomes a target, people and processes can adapt to improve the number without improving the underlying outcome.
That is the readiness milestone: shared definitions, proportionate access and security boundaries, an education plan, a supported tool choice, a baseline and visible measures agreed before the first pilot is assessed.
Begin with less authority than the technology can hold
Implementation should begin with a small number of reversible workflows whose purpose and boundaries can be explained.
I often describe these as skills: composed prompts or workflows that can be shared and improved from a common baseline. The label is less important than the contract. Each one should make clear:
- its purpose and intended user
- the inputs it requires and information it may access
- the output it should produce
- actions it may and must not perform
- points that require human approval
- completion and verification criteria
- known limitations and examples
The first versions should be deliberately less powerful than they could be. Authority is easier to add after a workflow has demonstrated value than to recover after it has changed something consequential or damaged trust.
For a development and testing team, plausible early candidates include:
- reviewing a work item for ambiguity, missing information, risks and testable acceptance criteria
- investigating a problem read-only while separating observed facts from assumptions and suggested next checks
- reviewing a code change for possible defects, regressions, missing tests and documentation without editing the code
- suggesting test cases, edge cases and areas of risk from agreed scope or a candidate change
- generating bounded tests that remain subject to human review and execution
These are not automatically the right pilots. Discovery may reveal a simpler and more valuable opportunity elsewhere. They are useful examples because their initial authority can remain narrow and their output can be inspected before it changes the system.
I would avoid beginning with automated migrations, security-sensitive changes or production actions. Merge assistance should initially remain advisory when database, security or architectural decisions may be involved. The potential time saving does not justify treating consequence as a later concern.
The operational structure can stay simple. A visible project should contain the pilot work, decisions, feedback and current status. It needs an accountable sponsor, a short regular review, one route for consolidated feedback and a reasonable expectation for decisions or sign-off. When required feedback is unavailable, the honest state is paused. The workflow should not silently expand its authority to keep itself busy.
Documentation for a reusable skill should be generated from the same source that defines the skill where practical. Purpose, inputs, permissions, limitations and examples should not drift into a second manually maintained version. Feedback can return to the source before the documentation is republished.
Expansion should be earned at visible milestones
Small steps are not only a way to reduce fear. They make causality and cost easier to inspect.
A foundation phase might end by producing one demonstrator and a recommendation for what should happen next. A pass at that stage means the team has found a bounded use worth testing under agreed conditions. It does not mean the organisation has proved AI adoption successful.
A pilot milestone asks whether people can use the workflow as intended and whether its outputs meet the agreed acceptance criteria. It should include corrections, rejections, verification effort and the team's experience—not only successful examples.
A value milestone asks whether the workflow improves the complete outcome once its design, training, maintenance and displaced work are included. Where the benefit is delayed, the review should state what leading evidence exists, what has not yet been observed and when the claim will be revisited.
An authority milestone asks whether the evidence justifies allowing the workflow to do more. A read-only reviewer might eventually be allowed to add a comment. A test generator might be allowed to open a candidate change. Each increase should have its own consequence, rollback path, approval boundary and evidence requirement.
Passing a milestone does not oblige the team to expand. A workflow can remain useful at low authority. Failing one does not prove that the people resisted AI or that the technology has no value. It may reveal a poor task boundary, missing context, weak documentation, unsuitable tooling or a process that should be repaired before assistance is added.
This is why implementation loops back into discovery. The pilot will expose things the initial map missed. It may reveal that the task depends on a hidden decision, that maintaining context costs more than expected or that a person was performing a valuable check nobody had named. Review the system, refine the intervention and repeat only when the next experiment is justified.
The objective is net improvement, not maximum adoption
There is no silver-bullet AI onboarding plan.
Every team has a different product, history, risk profile, set of tools and distribution of knowledge. A method that helps one team may add ceremony to another. A task that is safe and reversible in one system may be consequential in the next. The same tool can create relief for one person and cognitive overload for another.
The repeatable part is not the choice of model or workflow. It is the sequence of questions:
- How does the work happen now?
- What outcome are we trying to improve, and whose objective is it?
- What must the team understand and agree before beginning?
- What is the smallest intervention that can produce useful evidence?
- What does accepted value include, and where might the cost move?
- What has the workflow earned the authority to do next?
AI should be an enabler, not an enforcer. That principle applies to people and to systems. Adoption should help a team exercise more useful judgement and remove avoidable burden. It should not force everyone into the same interface, turn activity into surveillance or use commercial urgency to hide unresolved risk.
The goal is not to get as much AI into the organisation as possible. It is to obtain as much worthwhile benefit as the team can sustain, with as little negative cost as the complete system can reasonably achieve.
Start by understanding the team. Let the evidence decide how far the AI should go.
Further reading
These books offer related perspectives rather than direct proof of this approach. I am including them as a reading list to examine, not as endorsements of every claim they contain.
- Thinking in Systems, Donella H. Meadows —a clear introduction to feedback, delays, system boundaries, resilience and why an intervention's effects may appear somewhere other than the point of change
- Accelerate, Nicole Forsgren, Jez Humble and Gene Kim —research-led work on software-delivery performance, measurement and the organisational capabilities associated with stronger outcomes
- Team Topologies, Matthew Skelton and Manuel Pais —a practical lens on value flow, team boundaries, interaction modes and cognitive load
- Human-Centered AI, Ben Shneiderman —an argument for combining high levels of automation with human control, reliable software engineering and organisational governance
- The New Economics for Industry, Government, Education, W. Edwards Deming —a systems view of management that connects variation, knowledge, psychology and cooperation rather than reducing performance to isolated individual output