Picture a pull request that any agentic rollout produces by the dozen. An agent picks up a ticket with a one-line description, writes the change, writes the tests and opens the PR. An automated review gate passes it. The tests are green. Coverage goes up. Everyone is pleased.
Now ask one question: "Who decided what the right answer was?"
The agent read the code, inferred the behaviour and wrote tests that agree with itself. A second agent checked that the tests were tidy. Nobody was wrong, exactly. Nobody was accountable either.
That question is the subject of this post. I've spent 22 years in quality engineering, from my first QA engineer job to building a 50-engineer QA, Release and Reliability function that acts as the enabling team for more than 25 product squads. I've moved organisations from manual-dominant to automation-dominant testing more than once, in regulated estates where an escaped defect is a regulatory event, not a Jira ticket. Now I own the quality strategy for the next transition: agents as the default way of working across an entire engineering organisation.
The question I keep getting from CTOs and QA directors is simple: how should a modern AI organisation look at quality? Here is my answer, in eight parts.
Key takeaways
- The bottleneck moved from producing code to deciding what correct means. Agents make generation cheap. Verification, judgement and accountability are now the scarce resources.
- QA doesn't disappear. It stops being a phase and becomes the owner of the oracle: specs, expected answers, evals, risk tiers and release evidence.
- Humans stay in six specific places: intent, risk acceptance, oracle ownership, eval design, the release decision and production judgement. Everywhere else, let the agents work.
- Keep familiar titles, rewrite the accountabilities. Change what people sign their name to, not what they're called. FDEs earn their cost only when the truth lives in the customer's workflow.
- XP matters more now, not less. Small batches, spec-first, CI, pairing and collective ownership are the control system that stops agent speed turning into instability.
Agents made code cheap. They did not make "correct" cheap. That's still a human job, and it's the job your quality function should be built around.
1. Traditional SDLC vs the agentic SDLC: the work moved, it didn't vanish
In the traditional lifecycle, quality was a place near the end: a test phase, then UAT, then a release weekend. Much of my early career in London banking looked like that: big test plans, long regression cycles and a sign-off meeting.
Agile and XP pulled testing left. On one Tier 1 bank programme I helped take an automated regression cycle from three days to three hours by wiring it into CI. At my current organisation we moved release cycles from 12 weeks to under four hours, and deployment frequency from monthly to hourly, through trunk-based development and CI/CD maturity. That same delivery overhaul across 25+ squads saved roughly $545K a year across the portfolio. None of that needed AI. It needed discipline.
The agentic SDLC changes something different: who does the work, and where humans have to stand.
In the agentic loop, humans own intent, evidence and decisions. Agents and pipelines do the generating and the running.
| Dimension | Traditional SDLC | Agentic SDLC |
|---|---|---|
| Scarce resource | Engineering hours to write code and tests | Clear intent, trusted oracles and reviewer attention |
| Where defects are found | Test phase and UAT, after build | Spec review, automated gates, evals and production signals |
| What a reviewer checks | Syntax, style and logic | Intent, assertions, risk and blast radius |
| Release evidence | Test pass rate and sign-off | Gate results, eval scores against a baseline, risk tier, named owner |
| Main failure mode | Late discovery, slow delivery | Fast, confident, plausible changes nobody fully understood |
| What quality leaders manage | Test coverage and headcount | The decision system: who decides what correct means, and on what evidence |
The last two rows are the ones that matter. The 2025 DORA report, drawn from nearly 5,000 technology professionals, found that 90% of respondents use AI at work and more than 80% believe it has increased their productivity. AI adoption now has a positive relationship with delivery throughput, but it still has a negative relationship with delivery stability. DORA's own explanation: without robust control systems, like strong automated testing, mature version control practices and fast feedback loops, more change volume leads to instability.
Agents increase change volume. Quality is the control system. I covered the throughput side in The AI Productivity Lie; this post is about the org you build around it.
2. If agents write the code and the tests, what does QA do?
This is the question that makes QA teams nervous, and I understand why. Here is my stance: if your QA function's value is executing tests, agents will replace it. If its value is deciding what correct means and proving the product meets it, agents make it more valuable than it has ever been.
The problem with agent-written tests isn't that they're sloppy. It's that they're agreeable: tests written from the implementation encode whatever the code does, bugs included. I went through the research in Your AI Wrote the Tests From the Code. So It Tested the Bug.: generated suites often reach faulty behaviour and then assert that it's correct. Coverage goes up. Confidence shouldn't.
Developers already sense this. In the Stack Overflow 2025 Developer Survey, 84% of respondents said they use or plan to use AI tools, yet more developers actively distrust the accuracy of AI output (46%) than trust it (33%). Only about 3% highly trust it. The most common frustration, cited by 66%, is "AI solutions that are almost right, but not quite."
"Almost right" is a quality problem, not a productivity problem. It is exactly the class of defect a good tester is trained to find: the plausible answer that's wrong at the edge, under load, for the customer nobody modelled.
Developers have already told us where humans stay in the loop: when the answer has to be trusted.
Asked when they'd still want a person's help if AI did most coding, 75% chose "when I don't trust AI's answers." That's the job description for quality engineering, written by developers themselves.
So what does QA actually do? Here's where I'm pointing my own function, and what I'd build anywhere:
- Own the oracle. Expected answers come from the spec, the contract, the domain expert or a curated golden dataset, never only from the code under test.
- Build the control points. Agentic PR review gates that enforce standards before a human sees the change. Contract tests at team boundaries. Eval suites for anything with a model in it.
- Tier the risk. Not every change needs the same scrutiny. A copy change and a change to how money moves are different classes of risk, and the gate should know that.
- Turn production into tests. We built an agent that reads production traces and raises pull requests directly. It now raises six to ten PRs a day that reach production. It removed the triage step. It did not remove the human review step, and it did not remove the need for someone to decide whether the trace showed a bug or a feature.
- Build the skills the agents run on. We've built more than 40 QA skills and agents, covering backend, contract, test strategy, mobile and data/PII testing, so squads adopt agentic testing with guardrails already in place instead of reinventing them squad by squad.
Missing from that list: running regression packs, writing step-by-step test cases and being the last stop before production. Agents are welcome to them.
3. Where does the human in the loop actually sit?
"Human in the loop" is in every vendor deck. Few can say which human, at which point, signing off on which artefact.
My rule is simple: humans own decisions that carry accountability and expected answers. Agents own generation and execution. If you can't name the person who would be in the incident review, you haven't put a human in the loop. You've put a human near the loop.
Six decisions stay with a named person. Between them, agents generate, run and report.
| Decision point | Human owner | What they sign off | What agents do around it |
|---|---|---|---|
| Intent and spec | Product owner with a quality engineer | Acceptance examples, contracts and non-goals, before any code | Draft specs, propose edge cases, flag ambiguity |
| Risk acceptance | Engineering lead with the business owner | The risk tier of the change and the blast radius they're willing to accept | Classify changes, map dependencies, estimate impact |
| Oracle ownership | Quality engineer with a domain expert | The expected answers: golden datasets, reference outputs, invariants | Generate inputs and fixtures; never the expected value alone for critical paths |
| Eval design | Quality engineer | What "good" means for model-driven behaviour, thresholds, and how the LLM judge is calibrated | Run evals at scale, report variance, propose new cases |
| Release decision | A named release owner | Promote, hold or roll back, recorded with the evidence | Assemble the evidence pack, run progressive rollout |
| Production judgement | SRE with a quality engineer | Whether a signal is a defect, a gap in the spec or acceptable behaviour | Watch traces, cluster incidents, draft fixes and new evals |
Two principles make this map work in practice.
First, never let the same agent write the code and the tests that judge it, at least on critical paths. Gergely Orosz's write-up of his conversation with Kent Beck describes TDD as a "superpower" with agents because agents introduce regressions, and notes that Beck had trouble stopping agents from deleting tests to make them "pass." If the agent can edit the oracle, you don't have an oracle.
Second, the human checkpoint has to be cheap enough that people actually do it. A 600-line generated diff at 5pm gets skimmed. The fix is smaller changes and automated gates that catch the mechanical problems first. That's why our autonomous PR review gate enforces standards before a human reaches the change: human attention goes on the questions only a human can answer.
METR is the cautionary tale for anyone tempted to replace this map with self-reported confidence. In METR's early-2025 randomised trial, 16 experienced open-source developers working on 246 real issues in their own repositories took 19% longer with AI tools allowed. Before the tasks they expected a 24% speedup; afterwards they still believed they'd been sped up by 20%.
Feelings are not evidence. METR's later estimates point to speedups, but with intervals that cross zero and heavy selection effects.
To be fair to the tools, METR now calls those results out of date. Its February 2026 update (57 developers, 800+ tasks, late-2025 tools) estimated an 18% speedup for returning developers and 4% for new ones, with confidence intervals crossing zero, and METR is redesigning the study because 30% to 50% of developers withheld tasks they didn't want to do without AI. The tools are probably improving. The lesson holds: people are bad at judging their own AI-assisted output. The loop needs named humans and real evidence, not a confidence survey.
4. How do you reskill the workforce?
I've led this kind of transition before, without AI. On a greenfield credit and lending platform at a Tier 1 bank, I was the first QA hire. I went on to build and run 40 QAs across 15 component teams in three regions, and to define the development pathway from manual testing into automation engineering as the function moved to an automation-first model. That model was backed by a cost-of-quality case that released more than $1M over three quarters.
What that programme taught me is that a reskilling pathway is a career framework, not a training catalogue. People need to see the next role, what it's accountable for and how they'll be judged in it. The same is true now, with one difference: agents are very good at the part of automation engineering that took people months to learn, and poor at the part that takes years: domain knowledge and risk judgement. Knowing that a duplicate settlement is a regulatory incident while a misaligned button is a backlog item is exactly the knowledge an agent doesn't have. Reskill towards that.
The other principle: don't make tooling the barrier. For our in-house conversational AI product, we built a custom LLM evaluation framework on DeepEval with a simple Flask front end, so non-technical QA could run AI validation without writing code. Domain testers shouldn't need Python to be eval designers.
Here's the phased plan I'd use. Run it on real work, one squad at a time, and end each phase with evidence, not a certificate.
Each phase ends with something you can inspect: reviewed PRs, a working eval suite, a gate in the merge path.
| Phase | Focus | What people do | Evidence at the end |
|---|---|---|---|
| Days 1–30 | Read and judge | Pair with a coding agent daily on real tickets. Review agent PRs for assertions and intent: would this test fail if the behaviour were wrong? Label every generated test's oracle source: implementation, spec or human. | Each person can explain an assertion they rejected and why |
| Days 31–60 | Specify and measure | Write acceptance examples before the agent starts. Build one eval suite per squad with a small human-labelled golden set. Calibrate any LLM judge against human labels before it can block anything. | Every squad has a spec-first story in production and an eval suite in CI |
| Days 61–90 | Own the gates | Put risk tiers in the merge path. Turn production incidents and traces into specs or evals weekly. Replace test counts in the quality report with evidence. | A release decision that changed because of evidence the quality team produced |
One non-negotiable for leaders: protect the time. Across four waterfall-to-agile conversions inside Tier 1 banks and more than 55 squads coached through agile and XP, my firm view is that reskilling squeezed into "spare capacity" doesn't happen. Fund the learning, or expect theatre.
5. Do you keep traditional titles like dev and QA?
My answer is mostly yes, and that's deliberately unfashionable.
Our industry answers every shift with a new title: QA, then QE, then SDET, now "AI Quality Engineer". Renaming is cheap and changes nothing that matters.
What matters is accountability. In an agentic org the question isn't "what do we call you?" It's "what do you sign your name to?" Keep a small number of titles that people and recruiters understand, and rewrite what each one owns.
The title can stay. The accountability has to change.
| Today's title | New focus | What they now sign off |
|---|---|---|
| Manual tester | Exploratory and domain risk specialist | Risk scenarios, exploratory findings, the "this is wrong for the customer" call |
| Automation engineer / SDET | Eval and verification engineer | Eval suites, contract tests, oracle quality, gate reliability |
| QA lead | Quality engineering lead | Risk tiers, control points and the squad's release evidence |
| Developer | Engineer who specifies, steers and owns agent output | Every line merged under their name, whoever typed it |
| Business analyst | Spec author | Acceptance examples precise enough to be executable intent |
| Release manager | Release owner | Promote, hold or roll back, with the evidence recorded |
| Head of QA | Head of quality | The organisation's decision system: who decides what correct means, on what evidence |
The developer row is the biggest change. "The agent wrote it" is not an acceptable answer in an incident review. Kent Beck's distinction in Augmented Coding: Beyond the Vibes is the right standard: "In augmented coding you care about the code, its complexity, the tests, & their coverage." The value system, he writes, is similar to hand coding, "tidy code that works. It's just that I don't type much of that code." That's the bar for every engineer in an agentic org.
Structurally, I run quality as an enabling team: a central practice carrying standards, tooling and the hardest testing problems, with capability embedded in 25+ squads, the model from QA is a Mindset, Not a Role. Agents make it more important. If each squad builds its own skills, gates and evals, you get 25 definitions of "tested."
6. Is a forward deployed engineer overkill?
The model comes from Palantir. In its own description of the roles, Palantir separates "Devs", software engineers in product development who build its platforms, from "Deltas", forward deployed software engineers who sit in business development, work as part of a team that directly supports one customer, and measure success by impact on that customer's goal. Andreessen Horowitz's June 2025 piece, Trading Margin for Moat, called the FDE "the hottest job in startups" and noted that, at the time of writing, 22 of the 311 open roles on OpenAI's careers page were forward deployed or solutions engineering roles. OpenAI's own FDE job description says FDEs measure success through "production adoption, measurable workflow impact, and eval-driven feedback that changes product and model roadmaps."
That last phrase is the one quality leaders should notice. A good FDE stands closest to the oracle: they see what "correct" means inside the customer's workflow, where AI products tend to fail.
An FDE is worth it when:
- You ship an AI product into customer environments that differ significantly in data, workflow and policy, and the real definition of "correct" only exists inside each customer.
- The cost of a wrong answer is high and domain-specific: finance, healthcare, legal, regulated operations.
- The FDE has an explicit duty to feed what they learn back into the product as eval cases, specs and reusable components, not just as custom code for one account.
An FDE is overkill when:
- Your product is largely uniform across customers, and a solutions engineer plus good documentation would do.
- It's an internal platform. Inside your own org, the equivalent is an enablement engineer embedded with squads, which a good quality enabling team already provides.
- "Forward deployed" is a new name for unpriced professional services. If every deployment produces bespoke code nobody can reuse, you're a consultancy with a software company's cost base.
My test: after six months, has the FDE made the next deployment faster and the product's eval suite better? If they've only made one customer happy, it's a services contract.
I've sat on the embedded side of this. I spent over a decade as a consultant inside Tier 1 banks, and on one engagement I founded an internal SDET community of practice precisely so the capability would outlive the engagement. FDEs face the same choice: leave a capability behind, or leave a dependency.
7. Do you still need SDLC basics and XP principles?
Yes, more than ever. I learned XP across two ThoughtWorks tenures, in India and London, in teams practising TDD, pairing, continuous integration, trunk-based development and collective code ownership. Every transformation I've led since was built on that foundation, and the agentic one is no exception.
XP practices are a control system for change. Agents increase the volume and speed of change. A bigger engine needs better brakes.
DORA's AI Capabilities Model names seven capabilities that amplify AI's benefits, and two would be familiar to any XP team from 2005: strong version control practices and working in small batches. Frequent commits amplify AI's positive influence on individual effectiveness; frequent use of rollback boosts AI-assisted team performance; small batches amplify AI's positive influence on product performance and reduce friction. As the 2025 report puts it: "AI doesn't fix a team; it amplifies what's already there."
Small batches: more important. Agents can generate a thousand-line change as easily as a ten-line one. The human reviewer can't review both with the same care. Enforce PR size limits on agent output. If the agent produces 800 lines, the agent splits it, not the reviewer.
Test-first or spec-first: more important, in a slightly different form. For humans, TDD is a design forcing function. For agents, the value is that the expected behaviour exists before the code that has to meet it. When I rebuilt a TDD-in-the-agent-loop experiment, strict red-green-refactor on top of a complete spec mostly bought extra turns and tokens. A good spec and TDD are substitutes for the same underlying good: knowing what you're building before you build it. So the non-negotiable is intent first: a spec, acceptance examples or a test, written or approved by a human, before the agent writes code. That's the move I described in From Red-Green-Refactor to Living Documentation.
Continuous integration: more important. CI is where an agent's confident claims meet reality. Flaky CI was always expensive; with agents it's dangerous, because agents happily retry until green. The same goes for agent evals, which is why I treat them as statistical tests rather than pass/fail checks.
Pairing: changed, not gone. Pairing becomes human-and-agent pairing, with the human as navigator. The agent drives; the human holds the design, asks "why", and stops the agent from coding ahead. Beck describes doing exactly this in his augmented coding experiments, intruding more on the design to keep complexity from piling up. Keep some human-human pairing for high-risk changes and for juniors, who otherwise learn only from a machine that's never been on call.
Collective ownership: more important, and harder. XP's collective code ownership said anyone can change any code, and everyone is responsible for it. With agents, the risk is the opposite: nobody owns code because nobody wrote it. Make ownership explicit. Whoever merges it owns it. The squad owns the service. "The AI wrote it" is not a rollback strategy at 3am.
Trunk-based development and continuous delivery: the multiplier. Small, frequent, reversible changes, the kind our move to hourly deployments was built on, are survivable at agent volume. Big, rare, irreversible ones are not, whoever writes them.
What I'd drop is ceremony: sign-off gates without evidence, test case documents nobody reads and metrics that count activity. Agents will generate those faster than ever, which is the best reason to stop asking for them.
8. The pivot: from quality as a phase to quality as the decision system
Stop running quality as a phase, a headcount or a department. Start running it as the organisation's decision system.
In practice that means five shifts:
- From test execution to oracle ownership. Your team's core asset is no longer the regression suite. It's the trusted set of expected answers, specs and evals that tell you whether anything an agent produces is right.
- From coverage to evidence. Report what you can prove, not how much you ran: gate results, eval scores against a baseline, escaped defects, change failure rate and time to restore.
- From gatekeeper to gate builder. Build the control points every squad runs through, and keep humans on the six decisions that need them.
- From headcount to leverage. Quality impact should compound through tooling and standards, not scale with people. Agents make that leverage much larger, if you build the skills once and share them.
- From opinion to accountability. Every release decision has a named owner and recorded evidence. That's not bureaucracy. It's the only thing that survives an incident review.
For a QA director or VP of Engineering, here's what I'd do this quarter:
- Pick one squad and draw the six human decision points for its next three releases. Name the people.
- Add an oracle-source label to generated tests and report the split.
- Put one risk-tiered gate in the merge path, so high-risk changes can't merge on agent review alone.
- Start the 90-day reskilling plan with your quality team, on real work, with time protected.
- Change one line in your quality report from a count of tests to a piece of evidence.
In an agentic org, code is cheap and tests are cheap. The scarce thing is someone willing to say what correct means, and sign their name to it.
Continue reading
- Your AI Wrote the Tests From the Code. So It Tested the Bug. covers why oracles must come from the spec, not the implementation.
- The AI Productivity Lie: Your Engineers Merge Twice the PRs and Ship Nothing Faster covers where the delivery bottleneck moves when AI speeds up coding.
- QA is a Mindset, Not a Role covers the enabling-team model that agentic quality builds on.
Sources
- DORA / Google Cloud, "Announcing the 2025 DORA Report: State of AI-assisted Software Development," September 23, 2025. Nearly 5,000 respondents; 90% use AI at work; more than 80% believe it increased productivity; 30% report little or no trust in AI-generated code; positive relationship with throughput, negative with stability; "AI doesn't fix a team; it amplifies what's already there."
- DORA / Google Cloud, "Introducing DORA's inaugural AI Capabilities Model." Seven capabilities, including strong version control practices and working in small batches. See also the AI Capabilities Model PDF.
- Stack Overflow, "2025 Developer Survey: AI." 84% use or plan to use AI tools; 46% distrust vs 33% trust AI accuracy; 3% highly trust; 66% cite "almost right, but not quite"; 75% would still ask a person "when I don't trust AI's answers."
- Becker, Rush, Barnes and Rein, METR, "Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity," July 10, 2025. 16 developers, 246 issues; 19% longer with AI; expected 24% speedup; believed 20% afterwards.
- METR, "We are Changing our Developer Productivity Experiment Design," February 24, 2026. 57 developers, 800+ tasks; estimated speedups of 18% (returning developers) and 4% (new developers), with confidence intervals crossing zero; 30% to 50% of developers withheld tasks they didn't want to do without AI.
- Kent Beck, "Augmented Coding: Beyond the Vibes," Tidy First?, June 25, 2025.
- Gergely Orosz, "TDD, AI agents and coding with Kent Beck," The Pragmatic Engineer, June 11, 2025.
- Palantir, "Dev versus Delta: Demystifying engineering roles at Palantir," Palantir Blog, April 2019.
- Joe Schmidt, Andreessen Horowitz, "Trading Margin for Moat: Why the Forward Deployed Engineer Is the Hottest Job in Startups," June 4, 2025.
- OpenAI, "Forward Deployed Engineer - London," OpenAI Careers, accessed October 6, 2026.