Your eval suite went green. Re-run it and it goes red. Nothing in the PR changed — only the model's dice.
I have watched teams build careful prompt regression suites, then quietly destroy the signal with one habit: treating a single stochastic run like a unit test. The job fails at 0.79. Someone hits Re-run. It passes at 0.82. The PR merges. Everyone feels safer. They are not. They have trained the pipeline to manufacture confidence on demand.
This post is about the failure mode sitting between "we have goldens" and "we have a release gate": flaky, non-deterministic agent evals in CI. Complementary to pinning prompts and failing on threshold breach. Complementary to scorecards that separate invariants from averages. Different thesis: when the green check itself is a coin flip, you need pass rates, baselines, flake budgets, and the discipline to fail the build on variance — not only on a bad mean.
Key takeaways
- Gate stochastic evals on pass rates and baseline deltas, never on a single roll.
- Retries that hide disagreement are not resilience — they are how regressions ship.
- Give LLM tests a flake budget; quarantine unstable cases until the rubric or fixture earns a place in the merge gate.
A suite you re-run until green is not a quality gate. It is a slot machine with a merge button.
Why agent evals flake when Playwright does not
Classical flaky tests usually have a root cause you can name: race, clock, shared state, network. Agent and LLM evals flake because the system under test samples. Temperature, decoding, tool-choice jitter, and judge-model variance all move the same case between pass and fail with no code change. Even temperature zero is not a promise of bit-identical output across providers or days.
That does not mean "evals are useless in CI." It means a single pass/fail bit is the wrong contract for a stochastic check. A binary result on one trial is an anecdote. A pass rate over N trials, compared to a pinned baseline, is a measurement.
For a fictional support agent that looks up orders, creates replacements, and issues refunds, the same golden can swing for boring reasons:
| Source of swing | What it looks like in CI |
|---|---|
| Model sampling | Tool args or wording change; graded score moves a few points. |
| Judge variance | Same transcript, different LLM-as-judge roll → 0.81 then 0.77. |
| Tool / env noise | Latency-shaped recovery fixtures tip into "unresolved" vs "recovered." |
| Unpinned evaluator | Judge or embedding model upgrades under the same API name. |
| Fixture drift | Golden expected state no longer matches a quiet product rule change. |
The last two are process bugs. The first three are physics. Your gate has to admit the physics without letting process bugs hide inside "the model is non-deterministic."
Single-run pass/fail is the wrong gate
I still see pipelines that assert faithfulness >= 0.80 on one app call and one judge call. That threshold might be fine as a product bar. As a merge decision on one sample, it is a coin flip with API costs.
If the score's ordinary noise band on a stable commit is roughly ±0.04, a hard cut at 0.80 will fail clean PRs and pass dirty ones depending on the roll. Engineers respond rationally: they re-run. After a week of that, red means "try again," not "investigate." At that point the gate is worse than absent — it burns tokens and teaches the team to ignore red.
Split the suite the way you already split agent release evidence:
| Check class | Example | CI contract |
|---|---|---|
| Critical invariants | Exactly one replacement; no cross-customer write; schema-valid tool args; refund ≤ paid | Deterministic, single-run, merge-blocking |
| Stochastic quality | Task completion under temperature, graded faithfulness, policy adherence via pinned judge | Pass rate / baseline delta, merge-blocking only with margin |
| Trend / research signals | Exploratory rubrics, new judge prompts, experimental slices | Dashboard and PR comment — not a hard gate until stable |
Invariants stay binary. They are cheap and they catch the failures that must never average away. Everything that needs an LLM to grade belongs in the rate-and-delta world.
Pass rates, not folklore retries
When a case is legitimately stochastic, stop asking "did it pass?" Ask "how often does it pass under a fixed protocol?"
Illustrative defaults until your traffic and risk replace them:
- Run each gated stochastic case N times in one CI job (odd N helps avoid ties — 5 or 7 is a common starting point for consequential workflows; 3 is a floor for cheaper slices).
- Require a minimum pass rate with margin — e.g. at least 4 of 5, or 9 of 10 on a smaller high-value set — not "any one green run."
- Record successes / attempts, not a boolean after silent retries.
- Prefer majority or rate over "best of N." Best-of-N is how a 40% agent looks like a 100% agent in the PR check.
Retries belong to the product under test, not the evaluator. If the agent is allowed three tool retries before marking a refund unresolved, that policy is part of the scenario: measure final state, elapsed time, and cost across the whole attempt. If the CI job retries a failed eval until green, you are no longer measuring the agent — you are laundering variance.
A practical rule I put on teams: CI may retry infrastructure failures (runner died, provider 503 before the first token). CI must not retry assertion failures on the same commit without incrementing an explicit flake counter and failing when the counter exceeds budget.
Compare versions like an experiment
Absolute thresholds on noisy scores create flaky gates. Baseline deltas create answerable questions.
On a PR that touches prompts, tools, or agent policy:
- Score the candidate against the golden set (with N repeats where needed).
- Score the pinned baseline (
main, or a frozen release manifest) against the same fixtures, same judge, same seeds/protocol. - Fail only if the candidate is worse by more than a tolerance informed by measured noise — not by vibes in standup.
That converts "is 0.84 good?" (unanswerable) into "is 0.84 worse than main's 0.87 by more than noise?" (answerable). You do not need a research paper in the pipeline. You do need:
- A noise band measured on a stable commit before you set tolerances (re-measure when model, judge, or corpus changes).
- A smallest regression you care about — e.g. a clear drop in refund-policy adherence — and enough samples that a drop that size is distinguishable from ordinary wobble.
- Honesty when the difference is inside the band: inconclusive is a valid CI outcome for a soft metric. Ship on invariants + no significant delta, or demand more samples — do not force a coin flip into a boolean.
For long-tail products, watch the tail, not only the mean. A change that leaves average task success intact while crushing the worst 5–10% of refunds is exactly the regression a mean gate will miss. If your harness can assert on percentiles or on per-slice rates (VIP refunds, partial shipments, lost tool responses), use them for the dimensions where pain lives in the tail.
Fail the build on variance itself
Most teams only fail when the mean drops. The more interesting bug is a case that used to be stable and is now a coin flip.
Track per-case agreement across repeats and across builds on the same commit:
- A case that passes 5/5 on
mainand 2/5 on the PR is a regression even if a single lucky run was green. - A case that flip-flops 50/50 on an unchanged commit is not guarding anything — it is a tax.
- A sudden widening of score spread after a prompt edit can mean the new instructions are ambiguous. That is a product defect showing up as entropy.
Put a variance gate next to the mean gate for merge-critical slices:
| Signal | Merge response |
|---|---|
| Invariant violation in any trial | Block |
| Pass rate below floor vs baseline | Block |
| Case flip-rate above flake budget on unchanged inputs | Quarantine or block — do not "retry until green" |
| Mean inside noise band, variance unchanged | Pass (soft metrics) |
| Mean fine, tail or high-blast-radius slice down | Block on that slice |
Quarantine is not surrender. It is how you keep the merge gate honest. Unstable cases move to a nightly soak until the rubric, fixture, or product rule is fixed. A quarantined eval that still runs and reports is useful. A flaky eval that blocks merges at random trains people to skip the check.
A flake budget for LLM tests
Web UI tests learned flake budgets the hard way. LLM evals need the same operating model, with clearer ownership.
Define the budget in writing (illustrative starting point — replace with your risk tolerance):
- Merge gate: stochastic suite flake rate below an agreed ceiling (e.g. cases with >20% disagreement across N repeats on a stable commit must leave the gate).
- Nightly: wider corpus may tolerate more research noise, but trend charts must show per-case flip-flops, not only suite green %.
- Cost ceiling: N-repeats multiply token spend — budget repeats where blast radius is high; do not 7× the entire nightly corpus by default.
Who owns what
- QA / Quality Engineering owns flake detection, quarantine, pass-rate floors, and the promotion path back into the merge gate.
- Engineering owns pinning (model, prompt, judge, tools), idempotent tools, and deterministic fixtures for invariants.
- Product owns ambiguous expected behaviour that shows up as "the judge cannot decide" — that is often underspecification, not model weather.
Human review before you:
- Lower a pass-rate floor to make CI green.
- Remove a production-scar case because it flakes.
- Widen a noise tolerance after a high-severity miss.
- Unpin or silently upgrade the judge.
Wire it without theatre
Keep the mechanics boring. Path-filtered job on prompts/**, evals/**, agent policy, and tool code. Pin versions in the release manifest. Publish the full report: per-case successes/attempts, baseline deltas, quarantined list, token cost.
A minimal decision record on the PR should read like this:
Illustrative: hold. Invariants green. Refund-policy pass rate 3/5 vs baseline 5/5 (delta outside noise). Case
partial-ship-refundflip-rate 0.4 — quarantined from merge gate, assigned for rubric fix. No CI assertion retries.
That is actionable. "Eval: failed" with a re-run button is not.
Cost control still matters — N repeats are expensive. Use the same controls as prompt regression: small critical set on PRs, deeper corpus nightly; cheap deterministic layer always on; escalate only borderline stochastic cases to heavier sampling or a stronger judge. A flake-aware suite that finance kills for cost returns you to vibes. Measure spend beside signal.
The through-line
Prompt regression answers: did this instruction change move quality? Release gates answer: is this agent version safe to promote? Trace-to-eval answers: did we freeze the failures users already hit? Flake-aware CI answers a blunter question: can we trust the green check at all?
Stop gating stochastic behaviour on one roll. Measure rates. Compare to a baseline with eyes open about noise. Keep invariants binary and sacred. Refuse evaluator retries that erase disagreement. Budget flakes like you budget incidents. Quarantine what is not yet a measurement. Fail the build when variance itself says the product or the eval got worse.
If your agent eval can only pass by being re-run, you do not have a release signal. You have a lucky draw you are calling quality.
Treat the suite like the probabilistic system it is — and make the merge button earn a rate, not a vibe.
Continue reading
- Prompt Regression Testing: Treat Prompts Like Code in CI/CD covers pinning prompts, golden datasets, and failing the build when quality slips.
- Your AI Agent Passed the Demo. Can It Pass a Release Gate? covers grading final service state and keeping critical invariants off the weighted average.
- Production Failures Belong in Your Eval Suite: The Trace-to-Eval Loop covers turning scored production misses into permanent CI goldens.