Picture a pull request that any agentic rollout produces by the dozen. An agent picks up a ticket with a one-line description, writes the change, writes the tests and opens the PR. An automated review gate passes it. The tests are green. Coverage goes up. Everyone is pleased.

Now ask one question: "Who decided what the right answer was?"

The agent read the code, inferred the behaviour and wrote tests that agree with itself. A second agent checked that the tests were tidy. Nobody was wrong, exactly. Nobody was accountable either.

That question is the subject of this post. I've spent 22 years in quality engineering, from my first QA engineer job to building a 50-engineer QA, Release and Reliability function that acts as the enabling team for more than 25 product squads. I've moved organisations from manual-dominant to automation-dominant testing more than once, in regulated estates where an escaped defect is a regulatory event, not a Jira ticket. Now I own the quality strategy for the next transition: agents as the default way of working across an entire engineering organisation.

The question I keep getting from CTOs and QA directors is simple: how should a modern AI organisation look at quality? Here is my answer, in eight parts.

Key takeaways

Agents made code cheap. They did not make "correct" cheap. That's still a human job, and it's the job your quality function should be built around.

1. Traditional SDLC vs the agentic SDLC: the work moved, it didn't vanish

In the traditional lifecycle, quality was a place near the end: a test phase, then UAT, then a release weekend. Much of my early career in London banking looked like that: big test plans, long regression cycles and a sign-off meeting.

Agile and XP pulled testing left. On one Tier 1 bank programme I helped take an automated regression cycle from three days to three hours by wiring it into CI. At my current organisation we moved release cycles from 12 weeks to under four hours, and deployment frequency from monthly to hourly, through trunk-based development and CI/CD maturity. That same delivery overhaul across 25+ squads saved roughly $545K a year across the portfolio. None of that needed AI. It needed discipline.

The agentic SDLC changes something different: who does the work, and where humans have to stand.

Two flows compared. Traditional: requirements, design, build, test phase, UAT and release in sequence, with quality near the end. Agentic: intent and executable spec, agents build code and tests, automated gates, risk review and release decision, progressive rollout and production judgement, with production signals looping back into new specs and evals In the agentic loop, humans own intent, evidence and decisions. Agents and pipelines do the generating and the running.

Dimension Traditional SDLC Agentic SDLC
Scarce resource Engineering hours to write code and tests Clear intent, trusted oracles and reviewer attention
Where defects are found Test phase and UAT, after build Spec review, automated gates, evals and production signals
What a reviewer checks Syntax, style and logic Intent, assertions, risk and blast radius
Release evidence Test pass rate and sign-off Gate results, eval scores against a baseline, risk tier, named owner
Main failure mode Late discovery, slow delivery Fast, confident, plausible changes nobody fully understood
What quality leaders manage Test coverage and headcount The decision system: who decides what correct means, and on what evidence

The last two rows are the ones that matter. The 2025 DORA report, drawn from nearly 5,000 technology professionals, found that 90% of respondents use AI at work and more than 80% believe it has increased their productivity. AI adoption now has a positive relationship with delivery throughput, but it still has a negative relationship with delivery stability. DORA's own explanation: without robust control systems, like strong automated testing, mature version control practices and fast feedback loops, more change volume leads to instability.

Agents increase change volume. Quality is the control system. I covered the throughput side in The AI Productivity Lie; this post is about the org you build around it.

2. If agents write the code and the tests, what does QA do?

This is the question that makes QA teams nervous, and I understand why. Here is my stance: if your QA function's value is executing tests, agents will replace it. If its value is deciding what correct means and proving the product meets it, agents make it more valuable than it has ever been.

The problem with agent-written tests isn't that they're sloppy. It's that they're agreeable: tests written from the implementation encode whatever the code does, bugs included. I went through the research in Your AI Wrote the Tests From the Code. So It Tested the Bug.: generated suites often reach faulty behaviour and then assert that it's correct. Coverage goes up. Confidence shouldn't.

Developers already sense this. In the Stack Overflow 2025 Developer Survey, 84% of respondents said they use or plan to use AI tools, yet more developers actively distrust the accuracy of AI output (46%) than trust it (33%). Only about 3% highly trust it. The most common frustration, cited by 66%, is "AI solutions that are almost right, but not quite."

"Almost right" is a quality problem, not a productivity problem. It is exactly the class of defect a good tester is trained to find: the plausible answer that's wrong at the edge, under load, for the customer nobody modelled.

Two bar charts from the Stack Overflow 2025 Developer Survey. Left: trust in the accuracy of AI output, with 3.1% highly trust, 29.6% somewhat trust, 26.1% somewhat distrust and 19.6% highly distrust. Right: reasons developers would still ask a person if AI does most coding, led by "when I don't trust AI's answers" at 75.3%, then ethical or security concerns at 61.7% Developers have already told us where humans stay in the loop: when the answer has to be trusted.

Asked when they'd still want a person's help if AI did most coding, 75% chose "when I don't trust AI's answers." That's the job description for quality engineering, written by developers themselves.

So what does QA actually do? Here's where I'm pointing my own function, and what I'd build anywhere:

  1. Own the oracle. Expected answers come from the spec, the contract, the domain expert or a curated golden dataset, never only from the code under test.
  2. Build the control points. Agentic PR review gates that enforce standards before a human sees the change. Contract tests at team boundaries. Eval suites for anything with a model in it.
  3. Tier the risk. Not every change needs the same scrutiny. A copy change and a change to how money moves are different classes of risk, and the gate should know that.
  4. Turn production into tests. We built an agent that reads production traces and raises pull requests directly. It now raises six to ten PRs a day that reach production. It removed the triage step. It did not remove the human review step, and it did not remove the need for someone to decide whether the trace showed a bug or a feature.
  5. Build the skills the agents run on. We've built more than 40 QA skills and agents, covering backend, contract, test strategy, mobile and data/PII testing, so squads adopt agentic testing with guardrails already in place instead of reinventing them squad by squad.

Missing from that list: running regression packs, writing step-by-step test cases and being the last stop before production. Agents are welcome to them.

3. Where does the human in the loop actually sit?

"Human in the loop" is in every vendor deck. Few can say which human, at which point, signing off on which artefact.

My rule is simple: humans own decisions that carry accountability and expected answers. Agents own generation and execution. If you can't name the person who would be in the incident review, you haven't put a human in the loop. You've put a human near the loop.

Six cards showing where humans sit in the loop. 1 Intent and spec, owned by product owner and quality engineer. 2 Risk acceptance, owned by engineering lead and business owner. 3 Oracle ownership, owned by quality engineer and domain expert. 4 Eval design, owned by quality engineer. 5 Release decision, owned by a named release owner. 6 Production judgement, owned by SRE and quality engineer Six decisions stay with a named person. Between them, agents generate, run and report.

Decision point Human owner What they sign off What agents do around it
Intent and spec Product owner with a quality engineer Acceptance examples, contracts and non-goals, before any code Draft specs, propose edge cases, flag ambiguity
Risk acceptance Engineering lead with the business owner The risk tier of the change and the blast radius they're willing to accept Classify changes, map dependencies, estimate impact
Oracle ownership Quality engineer with a domain expert The expected answers: golden datasets, reference outputs, invariants Generate inputs and fixtures; never the expected value alone for critical paths
Eval design Quality engineer What "good" means for model-driven behaviour, thresholds, and how the LLM judge is calibrated Run evals at scale, report variance, propose new cases
Release decision A named release owner Promote, hold or roll back, recorded with the evidence Assemble the evidence pack, run progressive rollout
Production judgement SRE with a quality engineer Whether a signal is a defect, a gap in the spec or acceptable behaviour Watch traces, cluster incidents, draft fixes and new evals

Two principles make this map work in practice.

First, never let the same agent write the code and the tests that judge it, at least on critical paths. Gergely Orosz's write-up of his conversation with Kent Beck describes TDD as a "superpower" with agents because agents introduce regressions, and notes that Beck had trouble stopping agents from deleting tests to make them "pass." If the agent can edit the oracle, you don't have an oracle.

Second, the human checkpoint has to be cheap enough that people actually do it. A 600-line generated diff at 5pm gets skimmed. The fix is smaller changes and automated gates that catch the mechanical problems first. That's why our autonomous PR review gate enforces standards before a human reaches the change: human attention goes on the questions only a human can answer.

METR is the cautionary tale for anyone tempted to replace this map with self-reported confidence. In METR's early-2025 randomised trial, 16 experienced open-source developers working on 246 real issues in their own repositories took 19% longer with AI tools allowed. Before the tasks they expected a 24% speedup; afterwards they still believed they'd been sped up by 20%.

Two charts from METR. Left: in the early-2025 study, developers forecast a 24% speedup and believed afterwards they had been 20% faster, but measured task time was 19% longer. Right: follow-up estimates with 95% confidence intervals, showing plus 19% for the original study, minus 18% for returning developers and minus 4% for new developers in late 2025, both later intervals crossing zero Feelings are not evidence. METR's later estimates point to speedups, but with intervals that cross zero and heavy selection effects.

To be fair to the tools, METR now calls those results out of date. Its February 2026 update (57 developers, 800+ tasks, late-2025 tools) estimated an 18% speedup for returning developers and 4% for new ones, with confidence intervals crossing zero, and METR is redesigning the study because 30% to 50% of developers withheld tasks they didn't want to do without AI. The tools are probably improving. The lesson holds: people are bad at judging their own AI-assisted output. The loop needs named humans and real evidence, not a confidence survey.

4. How do you reskill the workforce?

I've led this kind of transition before, without AI. On a greenfield credit and lending platform at a Tier 1 bank, I was the first QA hire. I went on to build and run 40 QAs across 15 component teams in three regions, and to define the development pathway from manual testing into automation engineering as the function moved to an automation-first model. That model was backed by a cost-of-quality case that released more than $1M over three quarters.

What that programme taught me is that a reskilling pathway is a career framework, not a training catalogue. People need to see the next role, what it's accountable for and how they'll be judged in it. The same is true now, with one difference: agents are very good at the part of automation engineering that took people months to learn, and poor at the part that takes years: domain knowledge and risk judgement. Knowing that a duplicate settlement is a regulatory incident while a misaligned button is a backlog item is exactly the knowledge an agent doesn't have. Reskill towards that.

The other principle: don't make tooling the barrier. For our in-house conversational AI product, we built a custom LLM evaluation framework on DeepEval with a simple Flask front end, so non-technical QA could run AI validation without writing code. Domain testers shouldn't need Python to be eval designers.

Here's the phased plan I'd use. Run it on real work, one squad at a time, and end each phase with evidence, not a certificate.

A three-phase 90-day reskilling roadmap. Days 1 to 30, read and judge: everyone pairs with an agent daily, review agent PRs for assertions not syntax, label every test's oracle source. Days 31 to 60, specify and measure: write acceptance examples first, build one eval suite per squad, calibrate an LLM judge against human labels. Days 61 to 90, own the gates: run risk tiers in the merge path, turn production incidents into evals, report evidence not test counts Each phase ends with something you can inspect: reviewed PRs, a working eval suite, a gate in the merge path.

Phase Focus What people do Evidence at the end
Days 1–30 Read and judge Pair with a coding agent daily on real tickets. Review agent PRs for assertions and intent: would this test fail if the behaviour were wrong? Label every generated test's oracle source: implementation, spec or human. Each person can explain an assertion they rejected and why
Days 31–60 Specify and measure Write acceptance examples before the agent starts. Build one eval suite per squad with a small human-labelled golden set. Calibrate any LLM judge against human labels before it can block anything. Every squad has a spec-first story in production and an eval suite in CI
Days 61–90 Own the gates Put risk tiers in the merge path. Turn production incidents and traces into specs or evals weekly. Replace test counts in the quality report with evidence. A release decision that changed because of evidence the quality team produced

One non-negotiable for leaders: protect the time. Across four waterfall-to-agile conversions inside Tier 1 banks and more than 55 squads coached through agile and XP, my firm view is that reskilling squeezed into "spare capacity" doesn't happen. Fund the learning, or expect theatre.

5. Do you keep traditional titles like dev and QA?

My answer is mostly yes, and that's deliberately unfashionable.

Our industry answers every shift with a new title: QA, then QE, then SDET, now "AI Quality Engineer". Renaming is cheap and changes nothing that matters.

What matters is accountability. In an agentic org the question isn't "what do we call you?" It's "what do you sign your name to?" Keep a small number of titles that people and recruiters understand, and rewrite what each one owns.

Seven role mappings from today's title to agentic focus: manual tester to exploratory and domain risk specialist; automation engineer or SDET to eval and verification engineer; QA lead to quality engineering lead owning oracles, risk tiers and gates; developer to engineer who specifies, steers and owns agent output; business analyst to spec author; release manager to release owner; head of QA to head of quality owning the organisation's decision system The title can stay. The accountability has to change.

Today's title New focus What they now sign off
Manual tester Exploratory and domain risk specialist Risk scenarios, exploratory findings, the "this is wrong for the customer" call
Automation engineer / SDET Eval and verification engineer Eval suites, contract tests, oracle quality, gate reliability
QA lead Quality engineering lead Risk tiers, control points and the squad's release evidence
Developer Engineer who specifies, steers and owns agent output Every line merged under their name, whoever typed it
Business analyst Spec author Acceptance examples precise enough to be executable intent
Release manager Release owner Promote, hold or roll back, with the evidence recorded
Head of QA Head of quality The organisation's decision system: who decides what correct means, on what evidence

The developer row is the biggest change. "The agent wrote it" is not an acceptable answer in an incident review. Kent Beck's distinction in Augmented Coding: Beyond the Vibes is the right standard: "In augmented coding you care about the code, its complexity, the tests, & their coverage." The value system, he writes, is similar to hand coding, "tidy code that works. It's just that I don't type much of that code." That's the bar for every engineer in an agentic org.

Structurally, I run quality as an enabling team: a central practice carrying standards, tooling and the hardest testing problems, with capability embedded in 25+ squads, the model from QA is a Mindset, Not a Role. Agents make it more important. If each squad builds its own skills, gates and evals, you get 25 definitions of "tested."

6. Is a forward deployed engineer overkill?

The model comes from Palantir. In its own description of the roles, Palantir separates "Devs", software engineers in product development who build its platforms, from "Deltas", forward deployed software engineers who sit in business development, work as part of a team that directly supports one customer, and measure success by impact on that customer's goal. Andreessen Horowitz's June 2025 piece, Trading Margin for Moat, called the FDE "the hottest job in startups" and noted that, at the time of writing, 22 of the 311 open roles on OpenAI's careers page were forward deployed or solutions engineering roles. OpenAI's own FDE job description says FDEs measure success through "production adoption, measurable workflow impact, and eval-driven feedback that changes product and model roadmaps."

That last phrase is the one quality leaders should notice. A good FDE stands closest to the oracle: they see what "correct" means inside the customer's workflow, where AI products tend to fail.

An FDE is worth it when:

An FDE is overkill when:

My test: after six months, has the FDE made the next deployment faster and the product's eval suite better? If they've only made one customer happy, it's a services contract.

I've sat on the embedded side of this. I spent over a decade as a consultant inside Tier 1 banks, and on one engagement I founded an internal SDET community of practice precisely so the capability would outlive the engagement. FDEs face the same choice: leave a capability behind, or leave a dependency.

7. Do you still need SDLC basics and XP principles?

Yes, more than ever. I learned XP across two ThoughtWorks tenures, in India and London, in teams practising TDD, pairing, continuous integration, trunk-based development and collective code ownership. Every transformation I've led since was built on that foundation, and the agentic one is no exception.

XP practices are a control system for change. Agents increase the volume and speed of change. A bigger engine needs better brakes.

DORA's AI Capabilities Model names seven capabilities that amplify AI's benefits, and two would be familiar to any XP team from 2005: strong version control practices and working in small batches. Frequent commits amplify AI's positive influence on individual effectiveness; frequent use of rollback boosts AI-assisted team performance; small batches amplify AI's positive influence on product performance and reduce friction. As the 2025 report puts it: "AI doesn't fix a team; it amplifies what's already there."

Small batches: more important. Agents can generate a thousand-line change as easily as a ten-line one. The human reviewer can't review both with the same care. Enforce PR size limits on agent output. If the agent produces 800 lines, the agent splits it, not the reviewer.

Test-first or spec-first: more important, in a slightly different form. For humans, TDD is a design forcing function. For agents, the value is that the expected behaviour exists before the code that has to meet it. When I rebuilt a TDD-in-the-agent-loop experiment, strict red-green-refactor on top of a complete spec mostly bought extra turns and tokens. A good spec and TDD are substitutes for the same underlying good: knowing what you're building before you build it. So the non-negotiable is intent first: a spec, acceptance examples or a test, written or approved by a human, before the agent writes code. That's the move I described in From Red-Green-Refactor to Living Documentation.

Continuous integration: more important. CI is where an agent's confident claims meet reality. Flaky CI was always expensive; with agents it's dangerous, because agents happily retry until green. The same goes for agent evals, which is why I treat them as statistical tests rather than pass/fail checks.

Pairing: changed, not gone. Pairing becomes human-and-agent pairing, with the human as navigator. The agent drives; the human holds the design, asks "why", and stops the agent from coding ahead. Beck describes doing exactly this in his augmented coding experiments, intruding more on the design to keep complexity from piling up. Keep some human-human pairing for high-risk changes and for juniors, who otherwise learn only from a machine that's never been on call.

Collective ownership: more important, and harder. XP's collective code ownership said anyone can change any code, and everyone is responsible for it. With agents, the risk is the opposite: nobody owns code because nobody wrote it. Make ownership explicit. Whoever merges it owns it. The squad owns the service. "The AI wrote it" is not a rollback strategy at 3am.

Trunk-based development and continuous delivery: the multiplier. Small, frequent, reversible changes, the kind our move to hourly deployments was built on, are survivable at agent volume. Big, rare, irreversible ones are not, whoever writes them.

What I'd drop is ceremony: sign-off gates without evidence, test case documents nobody reads and metrics that count activity. Agents will generate those faster than ever, which is the best reason to stop asking for them.

8. The pivot: from quality as a phase to quality as the decision system

Stop running quality as a phase, a headcount or a department. Start running it as the organisation's decision system.

In practice that means five shifts:

  1. From test execution to oracle ownership. Your team's core asset is no longer the regression suite. It's the trusted set of expected answers, specs and evals that tell you whether anything an agent produces is right.
  2. From coverage to evidence. Report what you can prove, not how much you ran: gate results, eval scores against a baseline, escaped defects, change failure rate and time to restore.
  3. From gatekeeper to gate builder. Build the control points every squad runs through, and keep humans on the six decisions that need them.
  4. From headcount to leverage. Quality impact should compound through tooling and standards, not scale with people. Agents make that leverage much larger, if you build the skills once and share them.
  5. From opinion to accountability. Every release decision has a named owner and recorded evidence. That's not bureaucracy. It's the only thing that survives an incident review.

For a QA director or VP of Engineering, here's what I'd do this quarter:

In an agentic org, code is cheap and tests are cheap. The scarce thing is someone willing to say what correct means, and sign their name to it.

Continue reading

Sources