Your eval suite has a second model in it. You call it a judge. CI treats it like an oracle.

I have watched teams pin prompts, freeze goldens, and build pass-rate gates — then hand the merge decision to an uncalibrated LLM that prefers the longer answer, the first option, or the model family that wrote the transcript. The agent under test gets a scorecard. The scorer gets a free pass. That is not quality engineering. That is trusting the referee because the jersey says "Evaluator."

This post is about the failure mode sitting under every graded eval: an untested LLM judge used as a merge-blocking signal. Complementary to flake-aware pass rates. Complementary to prompt regression and release gates. Different thesis: the judge is software under test. Until you calibrate it, its score cannot veto a PR.

Key takeaways

If the component that scores your agent has never been tested, you do not have an eval suite. You have vibes with an API bill.

The judge is in the suite — so test it

Most teams already admit the agent is non-deterministic. Fewer admit the grader is too. Same temperature knobs. Same provider drift. Same prompt fragility. Plus a set of failure modes that classical assertions do not have: the scorer can be systematically wrong in a direction that looks like quality.

Well-established patterns show up in production suites every week:

Failure mode What it looks like on a support / refund agent
Position bias Pairwise "which refund reply is better?" flips when you swap A and B.
Verbosity bias A padded, polite wrong policy explanation outscores a short correct one.
Self-preference Judge from model family X scores X-family agent transcripts higher on the same rubric.
Format sensitivity Markdown bullets, JSON wrappers, or a slightly rephrased rubric move the grade with no content change.

Research in 2026 made this harder to hand-wave. JudgeBiasBench (Zhou et al.) documents a taxonomy of judgment biases — 12 bias types across four dimensions — and finds that bias is still prevalent even in strong judges, with length, position, and beauty biases persisting across both generative and discriminative judges. RAND's Judge Reliability Harness (Dev, Sloan, Kavner, Kong, and Sandler) stress-tests judges under formatting, paraphrasing, verbosity, and label flips: no judge they evaluated was uniformly reliable across benchmarks, and simple formatting changes, paraphrasing, and changes in verbosity were enough to surface consistency issues.

You do not need to reproduce those papers in CI. You do need to stop pretending your internal "faithfulness ≥ 0.80" call is exempt from the same physics.

Graded quality without a calibrated judge is cargo cult

Deterministic invariants stay binary: exactly one replacement, no cross-customer write, refund ≤ paid amount, schema-valid tool args. Those checks do not need a judge. Everything that needs an LLM to grade — policy adherence tone, partial-shipment empathy, "did we explain the denial correctly?" — inherits the judge's bugs.

I still see pipelines that:

  1. Run the agent once.
  2. Call gpt-judge-latest with a rubric pasted from Notion.
  3. Fail the PR if the score dips below a round number chosen in standup.

That stack has three untested layers: the agent sample, the judge sample, and the rubric. Flake-aware CI addresses the first. This post addresses the second. The third is product ownership — if humans disagree on the expected refund explanation, the judge cannot invent consensus.

Split evidence the way you already should for agent release gates:

Check class Example Role of the LLM judge CI contract
Critical invariants Refund ≤ paid; one replacement SKU; no PII leak in tool args None — assert on state Merge-blocking, single-run
Calibrated graded quality Policy adherence, resolution quality on pinned goldens Pinned judge + passed calibration gate Pass rate / baseline delta, merge-blocking only after calibration
Exploratory rubrics New judge prompt, new dimension, experimental slice Candidate judge Dashboard / PR comment — not a hard gate

The middle row is earned. The bottom row is research. Confusing them is how you ship regressions that look like "the eval got pickier" when really the judge got noisier.

What a judge-calibration gate looks like

Treat the judge like a service you are promoting. Same discipline as a new test framework: prove it, pin it, watch it, retire it when it drifts.

1. Human agreement baseline

Before a judge can block merges, measure agreement against a frozen human-labeled set for your domain — not a generic chat leaderboard.

Illustrative starting protocol (replace with your risk):

Disagreement is a feature. If humans and the judge split on "was this denial explanation acceptable?", that case is not ready to be merge-blocking until product writes a clearer expected behaviour.

2. Bias probes on the paths that matter

You do not need twelve academic bias types on day one. You need probes that match how your judge is actually invoked.

For a fictional support agent graded on refund-policy adherence:

Probe Construction Fail if
Position Same two replies, swapped order in a pairwise compare Preference flips with no content change
Verbosity Correct short reply vs incorrect long, padded reply Long incorrect wins systematically
Paraphrase / format Same substance, Markdown vs plain, light paraphrase of the agent output Grade moves outside an agreed band
Self-preference (if relevant) Same transcript scored by judge A vs judge B from the agent’s family Systematic uplift for the home team

Keep the probe set small and owned. A weekly bias soak that nobody reads is theatre. A dozen probes that fail the promotion checklist are a gate.

3. Pin the judge like you pin the prompt

"Use the best model" is not a version. Pin:

When the provider ships a new snapshot under the same marketing name, that is a judge change. Re-run calibration. Do not discover it because refund scores mysteriously rose on Friday.

4. Document how the judge is used

Optional but useful: practice-oriented docs in the spirit of LLJ Cards (Chehbouni et al., arXiv:2609.24516) — context, evaluation criteria, prototyping, pipeline, and evaluation of the judge itself. A one-page "judge card" in the repo beats tribal knowledge in Slack. Who may change the rubric? What blast radius does this score have? When was human agreement last measured?

Promotion rule: uncalibrated cannot block

This is the operating rule I put on teams:

A judge may be merge-blocking only after it passes the calibration gate for this rubric and corpus. Until then it is advisory.

Concretely:

Judge state Allowed in CI Merge-blocking?
Candidate — new rubric or model, no fresh human baseline PR comment, nightly dashboard No
Calibrated — agreement + bias probes green; versions pinned Graded suite with pass rates Yes, with flake-aware rates
Drifted — provider change, rubric edit, or probe regression Auto-demote to advisory; alert owners No until re-calibrated
Quarantined — chronic disagreement or bias fail Nightly only No

Human review before you:

Invariants still win. A calibrated judge that likes a transcript cannot excuse a double refund.

Wire it without theatre

Keep the mechanics boring and path-filtered.

On PRs that touch prompts/**, evals/**, judge configs, or agent policy:

  1. Run deterministic invariants (always merge-blocking).
  2. Run the judge calibration suite when judge config or rubric changes — agreement spot-check against the frozen human set is enough if the full set is heavy; full suite nightly.
  3. Run graded agent evals only with the currently calibrated pinned judge.
  4. Publish a short decision record.

Illustrative PR note:

hold (advisory judge demoted). Invariants green. Judge refund-policy-v3 failed verbosity probe 4/4 (long incorrect preferred). Demoted from merge gate; graded scores reported as comment only. Agent pass rates vs baseline inconclusive under advisory judge. Owner: QA — recalibrate or revert rubric change.

That is actionable. "Eval score: 0.78" with no judge provenance is not.

Cost still matters. Calibration sets and bias probes are cheaper than 7× sampling the whole agent corpus with a random judge. Spend tokens where blast radius is high: promotion events, rubric edits, provider upgrades. Do not 7× exploratory rubrics in the merge job.

Ownership

Role Owns
QA / Quality Engineering Calibration gate, bias probes, promotion/demotion, quarantine, reporting
Engineering Pinning judge versions, deterministic fixtures, keeping invariants off the LLM
Product Rubric meaning, expected behaviour when humans disagree, accepting "inconclusive" over fake precision

If Product will not label a calibration set, you do not have a graded merge gate. You have a wish.

The through-line

Prompt regression answers: did this instruction change move quality? Release gates answer: is this agent version safe to promote? Trace-to-eval answers: did we freeze the failures users already hit? Flake-aware CI answers: can we trust a green check on a stochastic agent? Judge calibration answers the question underneath the graded half of that suite: can we trust the scorer?

Stop treating LLM-as-judge as a free oracle. Measure human agreement. Probe for position, verbosity, format, and self-preference on the workflows you actually ship. Pin versions. Document the contract. Demote on drift. Keep invariants sacred and binary. Let uncalibrated judges talk. Do not let them veto.

An untested judge blocking merges is not rigorous. It is automation of someone else's bias with your release schedule attached.

Gate the grader like the rest of the suite — or admit the score was never a gate.

Continue reading