Every engineering leader I speak to is being asked the same question by their board. "We bought AI coding tools for everyone. Where is the speed?"
Most of them do not have a good answer. Here is mine, and it will not be popular.
For most organisations, AI has not made software delivery faster. It has made typing faster. Those are not the same thing, and confusing the two is the most expensive mistake our industry is making right now.
Key takeaways
- AI raises individual output and moves the bottleneck downstream onto review, testing and operations.
- "Percentage of code written by AI" is a vanity metric. It rewards volume and punishes deletion.
- Measure the system, not the keyboard: lead time, review wait, change failure rate, rework and escaped defects.
- AI amplifies existing discipline. Small batches, tests first and fast verification are what turn generated code into delivered value.
AI did not remove the bottleneck in software delivery. It moved it.
The numbers nobody puts on the vendor slide
I have spent 22 years watching tools promise speed, so I went looking for data instead of demos. Four studies, read together, tell a very uncomfortable story.
- METR ran a randomised controlled trial with experienced open source developers. With AI tools they were 19% slower. They believed they were 20% faster. That perception gap is the whole story in one number.
- Faros AI tracked over 10,000 developers. Those with high AI adoption completed 21% more tasks and merged 98% more pull requests. PR review time went up 91%. Average PR size went up 154%. Bugs per developer went up 9%. At company level, there was no significant link between AI adoption and better delivery.
- The 2025 DORA report surveyed around 5,000 professionals. 90% use AI at work. AI adoption now correlates with higher throughput, but still correlates with lower delivery stability. 30% report little or no trust in the code it produces.
- GitClear analysed 211 million changed lines of code. Duplicated code rose from 8.3% to 12.3%. Refactoring fell from 25% of changed lines to under 10%. For the first time, copy and paste overtook moved code.

Put those side by side. More code. Bigger PRs. Slower reviews. More bugs. Less refactoring. Less stability.
That is not a productivity revolution. That is a factory that sped up one station and flooded the next one.
Coding was never the bottleneck
In every team I have led, writing code was a minority of the week. The rest went on understanding the problem, waiting for review, fighting environments, clarifying requirements, testing, deploying and fixing what broke.
Anyone who has read The Goal knows what happens when you speed up a step that is not the constraint. You do not get more output. You get more inventory.
In software, inventory is unreviewed pull requests, half tested features and code nobody fully understands sitting in a branch. We have handed every engineer a machine that produces inventory at industrial scale, and then we are surprised the warehouse is full.
The review queue is the new release weekend
I have called a 2,000 line PR a hostage situation. AI has made those hostage situations free to produce.
The reviewer is now the bottleneck, and the reviewer is usually your most senior engineer. You have turned your most expensive people into full time proofreaders for a machine that never gets tired and never says "I am not sure about this bit."
Reviewers are human. Faced with a 600 line generated diff at 5pm, they do what humans do. They skim, approve and hope. The 9% rise in bugs is not a mystery. It is fatigue with a commit hash.
Every discipline is feeling this, not just developers
| Discipline | What changed | What it feels like |
|---|---|---|
| Developers | More code they did not write but now own | "The AI wrote it" is not a rollback strategy at 3am |
| QA and test engineers | Test suites generated by the model that wrote the code | The model marking its own homework: coverage up, confidence not |
| SRE and operations | More, larger, less understood changes | The stability dip DORA measured, felt on the on call rota |
| Product owners and BAs | Vague requirements turn into code in minutes | The clarifying question never gets asked |
| Architects and platform teams | Duplication across services | Every service grows its own copy of the same helper |
| Engineering managers | Output metrics are back | Lines of code, rebranded as "AI acceptance rate" |
If you work in quality engineering, the second row should worry you most. I covered why in AI-Generated Tests Need a Test of Their Own.
The controversial part: stop measuring AI adoption
"Percentage of code written by AI" is the worst engineering metric of this decade.
It rewards volume. It punishes deletion. It tells you nothing about whether a customer got value, whether production is stable, or whether your codebase will still be maintainable in two years. It is Goodhart's Law with a licence fee.
If you want to know whether AI is working, measure the system, not the keyboard. Here is the scorecard I would start with. It is a proposed set, not an industry standard.
| Measure | What it tells you | Warning sign |
|---|---|---|
| Lead time for changes | First commit to running in production | PRs rise, lead time flat |
| Review wait time | How long a PR sits before a human looks at it | Queue grows every sprint |
| Change failure rate and time to restore | Whether speed is costing stability | Failure rate creeping up |
| Rework | Code rewritten within a few weeks of shipping | Churn up, refactoring down |
| Escaped defects per release | What customers actually experience | More incidents per deploy |
If those numbers are not moving, your AI investment is not moving either. It is generating activity that looks like progress on a dashboard.
What the genuinely faster teams do differently
The line from the DORA report worth pinning to every wall: "AI doesn't fix a team; it amplifies what's already there."
The teams I have seen get real gains did not buy better tools. They already had the discipline, and AI multiplied it.
- They made batches smaller, not bigger. PR size limits apply to generated code too. If the agent produces 800 lines, the agent splits it, not the reviewer.
- They write the test before the prompt. The test is the intent. The human defines what correct looks like and the AI writes code to meet it. Never let the same agent write both the code and the tests that judge it. I tested this in Does TDD Survive Inside an AI Agent's Loop?
- They invest in verification, not generation. Contract tests, static analysis and automated quality gates catch mechanical problems before a human opens the diff. The human reviews intent and design, not syntax.
- They budget for deletion. Refactoring time is protected. Duplication is treated as a defect. The goal is less code, not more.
- They measure flow, not output. Leadership asks "did lead time drop" rather than "how many PRs did we merge."
None of this is new. It is XP and continuous delivery, the same practices we have known work for two decades. The difference is that AI has removed the option of ignoring them.
The shift, in one line
AI did not remove the bottleneck in software delivery. It moved it downstream, onto review, testing and operations, and exposed every team that was getting by on slow typing and good luck.
The winners of this era will not be the teams that generate the most code. They will be the teams that need to write the least.
Is AI making your team faster, or just busier? And what is the one metric you would use to prove it?
Continue reading
- The Confident Wrong Answer covers how to verify the code an agent writes before it merges.
- When Efficiency Becomes the Enemy covers why optimising one step can slow the whole system.