The pull request looked great. An agent refactored a string utility, then generated twelve unit tests for it. Branch coverage went up. The mutation report was green. Reviewers saw a tidy test file and approved.

One of those tests was assertFalse(StringUtils.equals("abc", new StringBuilder("abc"))).

That assertion is wrong. Two character sequences with the same content should be equal. The test passed because the code under test was buggy, and the model read the buggy code and wrote down what it saw. The test didn't check the behaviour. It photographed it.

That example isn't hypothetical. It's lifted almost verbatim from a 2026 ISSTA study of LLM-generated tests (more on it below). And it's the clearest picture I've seen of the problem most teams are walking into: when the same input that might contain the bug is also the source of the expected answer, your tests can only ever agree with it.

I wrote last month that AI-generated tests need a test of their own, and I still believe mutation testing is the right challenge for weak assertions. Two new papers sharpen where that advice holds, and where it quietly stops working.

Key takeaways

A test generated from the implementation is a regression snapshot. It is not evidence the implementation is right.

The two words that matter: trigger and detect

Every useful test does two things. It triggers a fault, meaning the input drives the code down a path where behaviour is wrong. Then it detects the fault, meaning the assertion notices the wrong output and fails.

Coverage tools only see the first half. A line is "covered" whether the assertion after it checks the right value, the wrong value, or nothing at all. Mutation testing gets closer, because a mutant only dies if some assertion notices the change. But mutation testing has a hidden assumption: it starts from a suite that passes on the current code. If the current code is wrong, and the suite was written to agree with it, mutation testing grades the suite on how well it defends the bug.

For years that assumption was mostly fine, because humans wrote many tests from requirements and memory of what the feature was meant to do. Now the default workflow is "here's the function, write tests for it." The implementation is the prompt. That changes what our metrics mean.

What the evidence says

Study 1: coverage and mutation only work when the code is already right

Junda Zhao, Shurui Zhou and Eldan Cohen at the University of Toronto replicated two classic coverage and mutation studies using LLM-generated tests (arXiv:2607.22880, published at ISSTA 2026). The scale is serious: 11 models in 13 settings, 318 buggy methods from all 17 Defects4J projects, 8,268 generated suites and 101,123 test cases. Each method was given to the models twice, once as the fixed version and once as the buggy one.

The findings split cleanly in two.

When tests were generated from bug-free code, coverage and mutation score were meaningful signals for comparing models. Models with higher branch coverage and raw mutation scores tended to catch more real bugs. That's the regression-testing case: protect today's correct behaviour from tomorrow's change.

When tests were generated from buggy code, coverage correlations with real bug detection were weak in every view the authors analysed. Mutation analysis wasn't even applicable, because it needs a green suite, and the tests that would expose the bug are exactly the ones that fail. The StringUtils.equals example above (Defects4J Lang-14) is the paper's own illustration: from the fixed code the model wrote assertTrue; from the buggy code it wrote assertFalse.

One more result is worth a slide in your next quality review. Suite size barely correlated with mutation score or bug detection. More generated tests did not mean stronger tests.

Study 2: the inputs are fine, the assertions are the problem

Asma Hamidi, Michael Konstantinou, Renzo Degiovanni and Mike Papadakis went further and simulated a fully generated workflow: LLM-written code tested by LLM-written tests (arXiv:2609.09315). They used five models across four Python benchmarks and kept 6,066 of the hardest faults.

Test suites were sampled to maximise statement coverage, branch coverage, or mutation score. On average those suites triggered 32.9%, 38.9% and 32.2% of the faults respectively. They detected only about 1 to 2%. The tests reached the broken behaviour and then asserted that the broken behaviour was correct.

Mutation testing only marginally outperformed plain coverage here, and the authors openly question whether its much higher cost is justified in this setting. They also report that the artificial faults mutation tools inject don't seem to couple with the kinds of mistakes LLMs make.

Then they tried the obvious fix. They took the tests that triggered faults, hid the implementation, and asked a model to write new assertions from the natural-language task description alone. Detection improved, by around 6.6% on average and in the best case from about 1% to 28.6%. Even so, the corrected suites caught at most roughly 30% of faults. The authors' conclusion is blunt: humans still need to reason about the assertions.

Both papers point back to earlier work showing the same bias: LLM oracles tend to capture the program's actual behaviour rather than its expected behaviour (Konstantinou et al., arXiv:2410.21136), and incorrect code measurably degrades generated tests (Huang et al., arXiv:2409.09464).

What I'd change in the pipeline

None of this means "stop generating tests." The test inputs these models produce are genuinely useful; they reach the faulty path a third of the time or more. The fix is to stop letting the same source decide both the input and the expected answer.

1. Label every generated test by where its oracle came from

Add one line of metadata to generated tests, a tag or a docstring field: oracle: implementation, oracle: spec, or oracle: human. It costs nothing and changes the conversation. A suite that is 90% oracle: implementation is a regression net, and you can report it as one. Don't let it count toward "verified behaviour" on any dashboard.

2. Generate expected values with the code hidden

Split generation into two calls. The first may see the code and produce inputs and setup, which is where code access genuinely helps. The second sees only the ticket, acceptance criteria, API contract, or docstring plus those inputs, and writes the assertion. That's the same separation Study 2 tested. It isn't a cure, but it moves the oracle off the thing under suspicion.

If you're working spec-first, this is where spec-driven development pays for itself: the spec is the one artifact the model didn't derive from the code.

3. Make "would this assertion fail on the bug?" a review question

When an agent fixes a bug, require at least one test that fails on the parent commit and passes on the fix. That's the exact definition Study 1 used for real bug detection, and CI can check it mechanically: check out the base, run the new tests, expect a red. A bug-fix PR whose new tests all pass on the old code didn't test the fix.

For new features there's no parent to fail against, so the gate is human: a reviewer reads the assertions, not the test names, and signs off on expected values for the critical paths.

4. Keep mutation testing, scoped to what it's good at

Mutation testing remains the right tool for the question "would this suite notice if someone broke today's behaviour later?" Run it on code that has already been accepted, on the diff, and use surviving mutants to find lazy assertions. Just don't present a mutation score on freshly generated code as evidence that code is correct. Study 1 shows why: on buggy input there's no meaningful score to compute.

5. Report two numbers, not one

Stop reporting a single "test quality" metric for AI-generated suites. Report regression strength (coverage, mutation score on accepted code) and correctness evidence (share of tests with spec or human oracles, and fail-on-parent checks for bug fixes) separately. A CTO can follow that. More importantly, it stops a green dashboard from implying something it never measured.

A gate you can start with this week

A minimal version that fits most CI setups:

ai_generated_tests_gate:
  require:
    - every_generated_test_has: oracle_source   # implementation | spec | human
    - bug_fix_prs: at_least_one_new_test_fails_on_base_commit
    - critical_paths: oracle_source in [spec, human]
  report_separately:
    regression_strength: [branch_coverage, mutation_score_on_diff]
    correctness_evidence: [pct_spec_or_human_oracles, fail_on_base_count]
  never:
    - count implementation-oracle tests as behaviour verification

Start with the fail-on-base check for bug fixes. It's the cheapest, it's mechanical, and it catches exactly the assertFalse case from the top of this post.

The leadership point

For two decades the hard part of automation was getting enough tests written. That constraint is gone. The hard part now is the oldest problem in testing: knowing what the right answer is. Models are very good at producing inputs and very willing to agree with whatever code they're shown. That makes the oracle the scarce asset on your team, and it's where your senior testers' judgement should go.

If your AI-generated suite has never disagreed with your code, it hasn't proved the code works. It has only proved the code is consistent with itself.

Continue reading

Sources