The pull request looked great. An agent refactored a string utility, then generated twelve unit tests for it. Branch coverage went up. The mutation report was green. Reviewers saw a tidy test file and approved.
One of those tests was assertFalse(StringUtils.equals("abc", new StringBuilder("abc"))).
That assertion is wrong. Two character sequences with the same content should be equal. The test passed because the code under test was buggy, and the model read the buggy code and wrote down what it saw. The test didn't check the behaviour. It photographed it.
That example isn't hypothetical. It's lifted almost verbatim from a 2026 ISSTA study of LLM-generated tests (more on it below). And it's the clearest picture I've seen of the problem most teams are walking into: when the same input that might contain the bug is also the source of the expected answer, your tests can only ever agree with it.
I wrote last month that AI-generated tests need a test of their own, and I still believe mutation testing is the right challenge for weak assertions. Two new papers sharpen where that advice holds, and where it quietly stops working.
Key takeaways
- LLM tests are good at reaching bugs and bad at noticing them. In one study, generated suites triggered roughly a third of hard faults but detected only about 1 to 2% of them, because the assertions were wrong.
- Coverage and mutation score measure regression protection, not correctness. They're useful signals when you can trust the code you generated tests from. When you can't, they tell you very little.
- Separate the oracle from the implementation. Write or generate expected values from the spec, ticket, or contract with the code hidden, and make assertion review a human-owned step in the merge path.
A test generated from the implementation is a regression snapshot. It is not evidence the implementation is right.
The two words that matter: trigger and detect
Every useful test does two things. It triggers a fault, meaning the input drives the code down a path where behaviour is wrong. Then it detects the fault, meaning the assertion notices the wrong output and fails.
Coverage tools only see the first half. A line is "covered" whether the assertion after it checks the right value, the wrong value, or nothing at all. Mutation testing gets closer, because a mutant only dies if some assertion notices the change. But mutation testing has a hidden assumption: it starts from a suite that passes on the current code. If the current code is wrong, and the suite was written to agree with it, mutation testing grades the suite on how well it defends the bug.
For years that assumption was mostly fine, because humans wrote many tests from requirements and memory of what the feature was meant to do. Now the default workflow is "here's the function, write tests for it." The implementation is the prompt. That changes what our metrics mean.
What the evidence says
Study 1: coverage and mutation only work when the code is already right
Junda Zhao, Shurui Zhou and Eldan Cohen at the University of Toronto replicated two classic coverage and mutation studies using LLM-generated tests (arXiv:2607.22880, published at ISSTA 2026). The scale is serious: 11 models in 13 settings, 318 buggy methods from all 17 Defects4J projects, 8,268 generated suites and 101,123 test cases. Each method was given to the models twice, once as the fixed version and once as the buggy one.
The findings split cleanly in two.
When tests were generated from bug-free code, coverage and mutation score were meaningful signals for comparing models. Models with higher branch coverage and raw mutation scores tended to catch more real bugs. That's the regression-testing case: protect today's correct behaviour from tomorrow's change.
When tests were generated from buggy code, coverage correlations with real bug detection were weak in every view the authors analysed. Mutation analysis wasn't even applicable, because it needs a green suite, and the tests that would expose the bug are exactly the ones that fail. The StringUtils.equals example above (Defects4J Lang-14) is the paper's own illustration: from the fixed code the model wrote assertTrue; from the buggy code it wrote assertFalse.
One more result is worth a slide in your next quality review. Suite size barely correlated with mutation score or bug detection. More generated tests did not mean stronger tests.
Study 2: the inputs are fine, the assertions are the problem
Asma Hamidi, Michael Konstantinou, Renzo Degiovanni and Mike Papadakis went further and simulated a fully generated workflow: LLM-written code tested by LLM-written tests (arXiv:2609.09315). They used five models across four Python benchmarks and kept 6,066 of the hardest faults.
Test suites were sampled to maximise statement coverage, branch coverage, or mutation score. On average those suites triggered 32.9%, 38.9% and 32.2% of the faults respectively. They detected only about 1 to 2%. The tests reached the broken behaviour and then asserted that the broken behaviour was correct.
Mutation testing only marginally outperformed plain coverage here, and the authors openly question whether its much higher cost is justified in this setting. They also report that the artificial faults mutation tools inject don't seem to couple with the kinds of mistakes LLMs make.
Then they tried the obvious fix. They took the tests that triggered faults, hid the implementation, and asked a model to write new assertions from the natural-language task description alone. Detection improved, by around 6.6% on average and in the best case from about 1% to 28.6%. Even so, the corrected suites caught at most roughly 30% of faults. The authors' conclusion is blunt: humans still need to reason about the assertions.
Both papers point back to earlier work showing the same bias: LLM oracles tend to capture the program's actual behaviour rather than its expected behaviour (Konstantinou et al., arXiv:2410.21136), and incorrect code measurably degrades generated tests (Huang et al., arXiv:2409.09464).
What I'd change in the pipeline
None of this means "stop generating tests." The test inputs these models produce are genuinely useful; they reach the faulty path a third of the time or more. The fix is to stop letting the same source decide both the input and the expected answer.
1. Label every generated test by where its oracle came from
Add one line of metadata to generated tests, a tag or a docstring field: oracle: implementation, oracle: spec, or oracle: human. It costs nothing and changes the conversation. A suite that is 90% oracle: implementation is a regression net, and you can report it as one. Don't let it count toward "verified behaviour" on any dashboard.
2. Generate expected values with the code hidden
Split generation into two calls. The first may see the code and produce inputs and setup, which is where code access genuinely helps. The second sees only the ticket, acceptance criteria, API contract, or docstring plus those inputs, and writes the assertion. That's the same separation Study 2 tested. It isn't a cure, but it moves the oracle off the thing under suspicion.
If you're working spec-first, this is where spec-driven development pays for itself: the spec is the one artifact the model didn't derive from the code.
3. Make "would this assertion fail on the bug?" a review question
When an agent fixes a bug, require at least one test that fails on the parent commit and passes on the fix. That's the exact definition Study 1 used for real bug detection, and CI can check it mechanically: check out the base, run the new tests, expect a red. A bug-fix PR whose new tests all pass on the old code didn't test the fix.
For new features there's no parent to fail against, so the gate is human: a reviewer reads the assertions, not the test names, and signs off on expected values for the critical paths.
4. Keep mutation testing, scoped to what it's good at
Mutation testing remains the right tool for the question "would this suite notice if someone broke today's behaviour later?" Run it on code that has already been accepted, on the diff, and use surviving mutants to find lazy assertions. Just don't present a mutation score on freshly generated code as evidence that code is correct. Study 1 shows why: on buggy input there's no meaningful score to compute.
5. Report two numbers, not one
Stop reporting a single "test quality" metric for AI-generated suites. Report regression strength (coverage, mutation score on accepted code) and correctness evidence (share of tests with spec or human oracles, and fail-on-parent checks for bug fixes) separately. A CTO can follow that. More importantly, it stops a green dashboard from implying something it never measured.
A gate you can start with this week
A minimal version that fits most CI setups:
ai_generated_tests_gate:
require:
- every_generated_test_has: oracle_source # implementation | spec | human
- bug_fix_prs: at_least_one_new_test_fails_on_base_commit
- critical_paths: oracle_source in [spec, human]
report_separately:
regression_strength: [branch_coverage, mutation_score_on_diff]
correctness_evidence: [pct_spec_or_human_oracles, fail_on_base_count]
never:
- count implementation-oracle tests as behaviour verification
Start with the fail-on-base check for bug fixes. It's the cheapest, it's mechanical, and it catches exactly the assertFalse case from the top of this post.
The leadership point
For two decades the hard part of automation was getting enough tests written. That constraint is gone. The hard part now is the oldest problem in testing: knowing what the right answer is. Models are very good at producing inputs and very willing to agree with whatever code they're shown. That makes the oracle the scarce asset on your team, and it's where your senior testers' judgement should go.
If your AI-generated suite has never disagreed with your code, it hasn't proved the code works. It has only proved the code is consistent with itself.
Continue reading
- Your LLM Judge Is Untested Software. Gate It Like the Rest of the Suite. covers calibrating LLM graders with human agreement, bias probes, and pinned versions before they can block a merge.
- Your Agent Eval Suite Is a Coin Flip. Treat It Like One. covers pass rates, baselines, flake budgets, and failing the build on variance.
- Prompt Regression Testing: Treat Prompts Like Code in CI/CD covers pinning prompts, golden datasets, and failing the build when quality slips.
Sources
- Zhao, Zhou and Cohen, "Do Coverage and Mutation Scores of LLM-Generated Test Suites Correlate with Their Effectiveness? (Replicability Study)," ISSTA 2026 / PACMSE, arXiv:2607.22880. Replication package on GitHub.
- Hamidi, Konstantinou, Degiovanni and Papadakis, "How effective are traditional test criteria at detecting bugs in large language models generated code?", arXiv:2609.09315.
- Konstantinou, Degiovanni and Papadakis, "Do LLMs generate test oracles that capture the actual or the expected program behaviour?", arXiv:2410.21136.
- Huang, Zhang, Harman, Du and Cui, "Measuring the Influence of Incorrect Code on Test Generation," arXiv:2409.09464.