Every engineering leader I speak to is being asked the same question by their board. "We bought AI coding tools for everyone. Where is the speed?"

Most of them do not have a good answer. Here is mine, and it will not be popular.

For most organisations, AI has not made software delivery faster. It has made typing faster. Those are not the same thing, and confusing the two is the most expensive mistake our industry is making right now.

Key takeaways

AI did not remove the bottleneck in software delivery. It moved it.

The numbers nobody puts on the vendor slide

I have spent 22 years watching tools promise speed, so I went looking for data instead of demos. Four studies, read together, tell a very uncomfortable story.

More code in, same delivery out: developers with high AI adoption merged 98% more pull requests and completed 21% more tasks, while average PR size rose 154%, review time 91% and bugs per developer 9%, with no significant change in company-level delivery

Put those side by side. More code. Bigger PRs. Slower reviews. More bugs. Less refactoring. Less stability.

That is not a productivity revolution. That is a factory that sped up one station and flooded the next one.

Coding was never the bottleneck

In every team I have led, writing code was a minority of the week. The rest went on understanding the problem, waiting for review, fighting environments, clarifying requirements, testing, deploying and fixing what broke.

Anyone who has read The Goal knows what happens when you speed up a step that is not the constraint. You do not get more output. You get more inventory.

In software, inventory is unreviewed pull requests, half tested features and code nobody fully understands sitting in a branch. We have handed every engineer a machine that produces inventory at industrial scale, and then we are surprised the warehouse is full.

The review queue is the new release weekend

I have called a 2,000 line PR a hostage situation. AI has made those hostage situations free to produce.

The reviewer is now the bottleneck, and the reviewer is usually your most senior engineer. You have turned your most expensive people into full time proofreaders for a machine that never gets tired and never says "I am not sure about this bit."

Reviewers are human. Faced with a 600 line generated diff at 5pm, they do what humans do. They skim, approve and hope. The 9% rise in bugs is not a mystery. It is fatigue with a commit hash.

Every discipline is feeling this, not just developers

Discipline What changed What it feels like
Developers More code they did not write but now own "The AI wrote it" is not a rollback strategy at 3am
QA and test engineers Test suites generated by the model that wrote the code The model marking its own homework: coverage up, confidence not
SRE and operations More, larger, less understood changes The stability dip DORA measured, felt on the on call rota
Product owners and BAs Vague requirements turn into code in minutes The clarifying question never gets asked
Architects and platform teams Duplication across services Every service grows its own copy of the same helper
Engineering managers Output metrics are back Lines of code, rebranded as "AI acceptance rate"

If you work in quality engineering, the second row should worry you most. I covered why in AI-Generated Tests Need a Test of Their Own.

The controversial part: stop measuring AI adoption

"Percentage of code written by AI" is the worst engineering metric of this decade.

It rewards volume. It punishes deletion. It tells you nothing about whether a customer got value, whether production is stable, or whether your codebase will still be maintainable in two years. It is Goodhart's Law with a licence fee.

If you want to know whether AI is working, measure the system, not the keyboard. Here is the scorecard I would start with. It is a proposed set, not an industry standard.

Measure What it tells you Warning sign
Lead time for changes First commit to running in production PRs rise, lead time flat
Review wait time How long a PR sits before a human looks at it Queue grows every sprint
Change failure rate and time to restore Whether speed is costing stability Failure rate creeping up
Rework Code rewritten within a few weeks of shipping Churn up, refactoring down
Escaped defects per release What customers actually experience More incidents per deploy

If those numbers are not moving, your AI investment is not moving either. It is generating activity that looks like progress on a dashboard.

What the genuinely faster teams do differently

The line from the DORA report worth pinning to every wall: "AI doesn't fix a team; it amplifies what's already there."

The teams I have seen get real gains did not buy better tools. They already had the discipline, and AI multiplied it.

  1. They made batches smaller, not bigger. PR size limits apply to generated code too. If the agent produces 800 lines, the agent splits it, not the reviewer.
  2. They write the test before the prompt. The test is the intent. The human defines what correct looks like and the AI writes code to meet it. Never let the same agent write both the code and the tests that judge it. I tested this in Does TDD Survive Inside an AI Agent's Loop?
  3. They invest in verification, not generation. Contract tests, static analysis and automated quality gates catch mechanical problems before a human opens the diff. The human reviews intent and design, not syntax.
  4. They budget for deletion. Refactoring time is protected. Duplication is treated as a defect. The goal is less code, not more.
  5. They measure flow, not output. Leadership asks "did lead time drop" rather than "how many PRs did we merge."

None of this is new. It is XP and continuous delivery, the same practices we have known work for two decades. The difference is that AI has removed the option of ignoring them.

The shift, in one line

AI did not remove the bottleneck in software delivery. It moved it downstream, onto review, testing and operations, and exposed every team that was getting by on slow typing and good luck.

The winners of this era will not be the teams that generate the most code. They will be the teams that need to write the least.

Is AI making your team faster, or just busier? And what is the one metric you would use to prove it?

Continue reading