Your eval suite went green. Re-run it and it goes red. Nothing in the PR changed — only the model's dice.

I have watched teams build careful prompt regression suites, then quietly destroy the signal with one habit: treating a single stochastic run like a unit test. The job fails at 0.79. Someone hits Re-run. It passes at 0.82. The PR merges. Everyone feels safer. They are not. They have trained the pipeline to manufacture confidence on demand.

This post is about the failure mode sitting between "we have goldens" and "we have a release gate": flaky, non-deterministic agent evals in CI. Complementary to pinning prompts and failing on threshold breach. Complementary to scorecards that separate invariants from averages. Different thesis: when the green check itself is a coin flip, you need pass rates, baselines, flake budgets, and the discipline to fail the build on variance — not only on a bad mean.

Key takeaways

A suite you re-run until green is not a quality gate. It is a slot machine with a merge button.

Why agent evals flake when Playwright does not

Classical flaky tests usually have a root cause you can name: race, clock, shared state, network. Agent and LLM evals flake because the system under test samples. Temperature, decoding, tool-choice jitter, and judge-model variance all move the same case between pass and fail with no code change. Even temperature zero is not a promise of bit-identical output across providers or days.

That does not mean "evals are useless in CI." It means a single pass/fail bit is the wrong contract for a stochastic check. A binary result on one trial is an anecdote. A pass rate over N trials, compared to a pinned baseline, is a measurement.

For a fictional support agent that looks up orders, creates replacements, and issues refunds, the same golden can swing for boring reasons:

Source of swing What it looks like in CI
Model sampling Tool args or wording change; graded score moves a few points.
Judge variance Same transcript, different LLM-as-judge roll → 0.81 then 0.77.
Tool / env noise Latency-shaped recovery fixtures tip into "unresolved" vs "recovered."
Unpinned evaluator Judge or embedding model upgrades under the same API name.
Fixture drift Golden expected state no longer matches a quiet product rule change.

The last two are process bugs. The first three are physics. Your gate has to admit the physics without letting process bugs hide inside "the model is non-deterministic."

Single-run pass/fail is the wrong gate

I still see pipelines that assert faithfulness >= 0.80 on one app call and one judge call. That threshold might be fine as a product bar. As a merge decision on one sample, it is a coin flip with API costs.

If the score's ordinary noise band on a stable commit is roughly ±0.04, a hard cut at 0.80 will fail clean PRs and pass dirty ones depending on the roll. Engineers respond rationally: they re-run. After a week of that, red means "try again," not "investigate." At that point the gate is worse than absent — it burns tokens and teaches the team to ignore red.

Split the suite the way you already split agent release evidence:

Check class Example CI contract
Critical invariants Exactly one replacement; no cross-customer write; schema-valid tool args; refund ≤ paid Deterministic, single-run, merge-blocking
Stochastic quality Task completion under temperature, graded faithfulness, policy adherence via pinned judge Pass rate / baseline delta, merge-blocking only with margin
Trend / research signals Exploratory rubrics, new judge prompts, experimental slices Dashboard and PR comment — not a hard gate until stable

Invariants stay binary. They are cheap and they catch the failures that must never average away. Everything that needs an LLM to grade belongs in the rate-and-delta world.

Pass rates, not folklore retries

When a case is legitimately stochastic, stop asking "did it pass?" Ask "how often does it pass under a fixed protocol?"

Illustrative defaults until your traffic and risk replace them:

Retries belong to the product under test, not the evaluator. If the agent is allowed three tool retries before marking a refund unresolved, that policy is part of the scenario: measure final state, elapsed time, and cost across the whole attempt. If the CI job retries a failed eval until green, you are no longer measuring the agent — you are laundering variance.

A practical rule I put on teams: CI may retry infrastructure failures (runner died, provider 503 before the first token). CI must not retry assertion failures on the same commit without incrementing an explicit flake counter and failing when the counter exceeds budget.

Compare versions like an experiment

Absolute thresholds on noisy scores create flaky gates. Baseline deltas create answerable questions.

On a PR that touches prompts, tools, or agent policy:

  1. Score the candidate against the golden set (with N repeats where needed).
  2. Score the pinned baseline (main, or a frozen release manifest) against the same fixtures, same judge, same seeds/protocol.
  3. Fail only if the candidate is worse by more than a tolerance informed by measured noise — not by vibes in standup.

That converts "is 0.84 good?" (unanswerable) into "is 0.84 worse than main's 0.87 by more than noise?" (answerable). You do not need a research paper in the pipeline. You do need:

For long-tail products, watch the tail, not only the mean. A change that leaves average task success intact while crushing the worst 5–10% of refunds is exactly the regression a mean gate will miss. If your harness can assert on percentiles or on per-slice rates (VIP refunds, partial shipments, lost tool responses), use them for the dimensions where pain lives in the tail.

Fail the build on variance itself

Most teams only fail when the mean drops. The more interesting bug is a case that used to be stable and is now a coin flip.

Track per-case agreement across repeats and across builds on the same commit:

Put a variance gate next to the mean gate for merge-critical slices:

Signal Merge response
Invariant violation in any trial Block
Pass rate below floor vs baseline Block
Case flip-rate above flake budget on unchanged inputs Quarantine or block — do not "retry until green"
Mean inside noise band, variance unchanged Pass (soft metrics)
Mean fine, tail or high-blast-radius slice down Block on that slice

Quarantine is not surrender. It is how you keep the merge gate honest. Unstable cases move to a nightly soak until the rubric, fixture, or product rule is fixed. A quarantined eval that still runs and reports is useful. A flaky eval that blocks merges at random trains people to skip the check.

A flake budget for LLM tests

Web UI tests learned flake budgets the hard way. LLM evals need the same operating model, with clearer ownership.

Define the budget in writing (illustrative starting point — replace with your risk tolerance):

Who owns what

Human review before you:

Wire it without theatre

Keep the mechanics boring. Path-filtered job on prompts/**, evals/**, agent policy, and tool code. Pin versions in the release manifest. Publish the full report: per-case successes/attempts, baseline deltas, quarantined list, token cost.

A minimal decision record on the PR should read like this:

Illustrative: hold. Invariants green. Refund-policy pass rate 3/5 vs baseline 5/5 (delta outside noise). Case partial-ship-refund flip-rate 0.4 — quarantined from merge gate, assigned for rubric fix. No CI assertion retries.

That is actionable. "Eval: failed" with a re-run button is not.

Cost control still matters — N repeats are expensive. Use the same controls as prompt regression: small critical set on PRs, deeper corpus nightly; cheap deterministic layer always on; escalate only borderline stochastic cases to heavier sampling or a stronger judge. A flake-aware suite that finance kills for cost returns you to vibes. Measure spend beside signal.

The through-line

Prompt regression answers: did this instruction change move quality? Release gates answer: is this agent version safe to promote? Trace-to-eval answers: did we freeze the failures users already hit? Flake-aware CI answers a blunter question: can we trust the green check at all?

Stop gating stochastic behaviour on one roll. Measure rates. Compare to a baseline with eyes open about noise. Keep invariants binary and sacred. Refuse evaluator retries that erase disagreement. Budget flakes like you budget incidents. Quarantine what is not yet a measurement. Fail the build when variance itself says the product or the eval got worse.

If your agent eval can only pass by being re-run, you do not have a release signal. You have a lucky draw you are calling quality.

Treat the suite like the probabilistic system it is — and make the merge button earn a rate, not a vibe.

Continue reading