Your eval suite has a second model in it. You call it a judge. CI treats it like an oracle.
I have watched teams pin prompts, freeze goldens, and build pass-rate gates — then hand the merge decision to an uncalibrated LLM that prefers the longer answer, the first option, or the model family that wrote the transcript. The agent under test gets a scorecard. The scorer gets a free pass. That is not quality engineering. That is trusting the referee because the jersey says "Evaluator."
This post is about the failure mode sitting under every graded eval: an untested LLM judge used as a merge-blocking signal. Complementary to flake-aware pass rates. Complementary to prompt regression and release gates. Different thesis: the judge is software under test. Until you calibrate it, its score cannot veto a PR.
Key takeaways
- An LLM judge is a stochastic grader with documented biases — not a trusted oracle.
- Calibrate before you gate: human agreement, bias probes, pinned versions, and a promotion path into the merge gate.
- An uncalibrated judge may comment on a PR. It must not block the merge.
If the component that scores your agent has never been tested, you do not have an eval suite. You have vibes with an API bill.
The judge is in the suite — so test it
Most teams already admit the agent is non-deterministic. Fewer admit the grader is too. Same temperature knobs. Same provider drift. Same prompt fragility. Plus a set of failure modes that classical assertions do not have: the scorer can be systematically wrong in a direction that looks like quality.
Well-established patterns show up in production suites every week:
| Failure mode | What it looks like on a support / refund agent |
|---|---|
| Position bias | Pairwise "which refund reply is better?" flips when you swap A and B. |
| Verbosity bias | A padded, polite wrong policy explanation outscores a short correct one. |
| Self-preference | Judge from model family X scores X-family agent transcripts higher on the same rubric. |
| Format sensitivity | Markdown bullets, JSON wrappers, or a slightly rephrased rubric move the grade with no content change. |
Research in 2026 made this harder to hand-wave. JudgeBiasBench (Zhou et al.) documents a taxonomy of judgment biases — 12 bias types across four dimensions — and finds that bias is still prevalent even in strong judges, with length, position, and beauty biases persisting across both generative and discriminative judges. RAND's Judge Reliability Harness (Dev, Sloan, Kavner, Kong, and Sandler) stress-tests judges under formatting, paraphrasing, verbosity, and label flips: no judge they evaluated was uniformly reliable across benchmarks, and simple formatting changes, paraphrasing, and changes in verbosity were enough to surface consistency issues.
You do not need to reproduce those papers in CI. You do need to stop pretending your internal "faithfulness ≥ 0.80" call is exempt from the same physics.
Graded quality without a calibrated judge is cargo cult
Deterministic invariants stay binary: exactly one replacement, no cross-customer write, refund ≤ paid amount, schema-valid tool args. Those checks do not need a judge. Everything that needs an LLM to grade — policy adherence tone, partial-shipment empathy, "did we explain the denial correctly?" — inherits the judge's bugs.
I still see pipelines that:
- Run the agent once.
- Call
gpt-judge-latestwith a rubric pasted from Notion. - Fail the PR if the score dips below a round number chosen in standup.
That stack has three untested layers: the agent sample, the judge sample, and the rubric. Flake-aware CI addresses the first. This post addresses the second. The third is product ownership — if humans disagree on the expected refund explanation, the judge cannot invent consensus.
Split evidence the way you already should for agent release gates:
| Check class | Example | Role of the LLM judge | CI contract |
|---|---|---|---|
| Critical invariants | Refund ≤ paid; one replacement SKU; no PII leak in tool args | None — assert on state | Merge-blocking, single-run |
| Calibrated graded quality | Policy adherence, resolution quality on pinned goldens | Pinned judge + passed calibration gate | Pass rate / baseline delta, merge-blocking only after calibration |
| Exploratory rubrics | New judge prompt, new dimension, experimental slice | Candidate judge | Dashboard / PR comment — not a hard gate |
The middle row is earned. The bottom row is research. Confusing them is how you ship regressions that look like "the eval got pickier" when really the judge got noisier.
What a judge-calibration gate looks like
Treat the judge like a service you are promoting. Same discipline as a new test framework: prove it, pin it, watch it, retire it when it drifts.
1. Human agreement baseline
Before a judge can block merges, measure agreement against a frozen human-labeled set for your domain — not a generic chat leaderboard.
Illustrative starting protocol (replace with your risk):
- Label a calibration set of agent transcripts (support refunds, partials, denials, tool failures) with the same rubric CI will use.
- Score with the candidate judge under a fixed protocol (model id, prompt hash, temperature, pairwise order policy).
- Require agreement above a floor you chose with eyes open — e.g. substantial agreement with humans on binary pass/fail dimensions; report disagreement cases, do not bury them.
- Re-run the baseline when the rubric, product policy, or judge model changes. Stale human labels against a new refund rule are how you calibrate to the past.
Disagreement is a feature. If humans and the judge split on "was this denial explanation acceptable?", that case is not ready to be merge-blocking until product writes a clearer expected behaviour.
2. Bias probes on the paths that matter
You do not need twelve academic bias types on day one. You need probes that match how your judge is actually invoked.
For a fictional support agent graded on refund-policy adherence:
| Probe | Construction | Fail if |
|---|---|---|
| Position | Same two replies, swapped order in a pairwise compare | Preference flips with no content change |
| Verbosity | Correct short reply vs incorrect long, padded reply | Long incorrect wins systematically |
| Paraphrase / format | Same substance, Markdown vs plain, light paraphrase of the agent output | Grade moves outside an agreed band |
| Self-preference (if relevant) | Same transcript scored by judge A vs judge B from the agent’s family | Systematic uplift for the home team |
Keep the probe set small and owned. A weekly bias soak that nobody reads is theatre. A dozen probes that fail the promotion checklist are a gate.
3. Pin the judge like you pin the prompt
"Use the best model" is not a version. Pin:
- Model identifier (and refuse silent
latestaliases in merge CI) - Judge system prompt / rubric file hash
- Decoding settings
- Pairwise order policy (e.g. score both orders, require consistency)
- Optional: embedding or secondary scorer versions if they feed the same gate
When the provider ships a new snapshot under the same marketing name, that is a judge change. Re-run calibration. Do not discover it because refund scores mysteriously rose on Friday.
4. Document how the judge is used
Optional but useful: practice-oriented docs in the spirit of LLJ Cards (Chehbouni et al., arXiv:2609.24516) — context, evaluation criteria, prototyping, pipeline, and evaluation of the judge itself. A one-page "judge card" in the repo beats tribal knowledge in Slack. Who may change the rubric? What blast radius does this score have? When was human agreement last measured?
Promotion rule: uncalibrated cannot block
This is the operating rule I put on teams:
A judge may be merge-blocking only after it passes the calibration gate for this rubric and corpus. Until then it is advisory.
Concretely:
| Judge state | Allowed in CI | Merge-blocking? |
|---|---|---|
| Candidate — new rubric or model, no fresh human baseline | PR comment, nightly dashboard | No |
| Calibrated — agreement + bias probes green; versions pinned | Graded suite with pass rates | Yes, with flake-aware rates |
| Drifted — provider change, rubric edit, or probe regression | Auto-demote to advisory; alert owners | No until re-calibrated |
| Quarantined — chronic disagreement or bias fail | Nightly only | No |
Human review before you:
- Promote a judge into the merge gate.
- Lower the human-agreement floor to make promotion easier.
- Unpin or silently upgrade the judge model.
- Delete a bias probe because it failed after a "harmless" prompt tweak.
- Let a graded score override a failing invariant ("the judge said the refund was fine").
Invariants still win. A calibrated judge that likes a transcript cannot excuse a double refund.
Wire it without theatre
Keep the mechanics boring and path-filtered.
On PRs that touch prompts/**, evals/**, judge configs, or agent policy:
- Run deterministic invariants (always merge-blocking).
- Run the judge calibration suite when judge config or rubric changes — agreement spot-check against the frozen human set is enough if the full set is heavy; full suite nightly.
- Run graded agent evals only with the currently calibrated pinned judge.
- Publish a short decision record.
Illustrative PR note:
hold (advisory judge demoted). Invariants green. Judge
refund-policy-v3failed verbosity probe 4/4 (long incorrect preferred). Demoted from merge gate; graded scores reported as comment only. Agent pass rates vs baseline inconclusive under advisory judge. Owner: QA — recalibrate or revert rubric change.
That is actionable. "Eval score: 0.78" with no judge provenance is not.
Cost still matters. Calibration sets and bias probes are cheaper than 7× sampling the whole agent corpus with a random judge. Spend tokens where blast radius is high: promotion events, rubric edits, provider upgrades. Do not 7× exploratory rubrics in the merge job.
Ownership
| Role | Owns |
|---|---|
| QA / Quality Engineering | Calibration gate, bias probes, promotion/demotion, quarantine, reporting |
| Engineering | Pinning judge versions, deterministic fixtures, keeping invariants off the LLM |
| Product | Rubric meaning, expected behaviour when humans disagree, accepting "inconclusive" over fake precision |
If Product will not label a calibration set, you do not have a graded merge gate. You have a wish.
The through-line
Prompt regression answers: did this instruction change move quality? Release gates answer: is this agent version safe to promote? Trace-to-eval answers: did we freeze the failures users already hit? Flake-aware CI answers: can we trust a green check on a stochastic agent? Judge calibration answers the question underneath the graded half of that suite: can we trust the scorer?
Stop treating LLM-as-judge as a free oracle. Measure human agreement. Probe for position, verbosity, format, and self-preference on the workflows you actually ship. Pin versions. Document the contract. Demote on drift. Keep invariants sacred and binary. Let uncalibrated judges talk. Do not let them veto.
An untested judge blocking merges is not rigorous. It is automation of someone else's bias with your release schedule attached.
Gate the grader like the rest of the suite — or admit the score was never a gate.
Continue reading
- Your Agent Eval Suite Is a Coin Flip. Treat It Like One. covers pass rates, baselines, flake budgets, and failing the build on variance.
- Prompt Regression Testing: Treat Prompts Like Code in CI/CD covers pinning prompts, golden datasets, and failing the build when quality slips.
- Your AI Agent Passed the Demo. Can It Pass a Release Gate? covers grading final service state and keeping critical invariants off the weighted average.
- Production Failures Belong in Your Eval Suite: The Trace-to-Eval Loop covers turning scored production misses into permanent CI goldens.