Your release gate was green. Production still lied to a customer. That miss belongs in CI — permanently.

I have watched this pattern enough times to stop treating it as bad luck. The prompt regression suite passes. The agent release scorecard passes. Staging fixtures look healthy. Then a real user hits a path no golden ever covered: a refund on a partially shipped order, a tool timeout after the write already landed, a policy edge the product team never wrote down. Someone files a ticket. Someone apologises. Someone tweaks a prompt. The incident channel closes. The next release ships without that failure as a permanent case.

That is the gap this post is about. Pre-release goldens and release gates are necessary. They are also incomplete. The missing discipline is a loop that turns failing production agent traces into regression fixtures — observed, scored, triaged, redacted, and frozen so the same miss fails a future build.

Key takeaways

A production failure that never becomes a golden is an incident you have agreed to rediscover.

What trace-to-eval actually means

Trace-to-eval is not a dashboard of interesting sessions. It is a closed loop with a build consequence.

  1. Observe — capture structured traces for agent runs in production (or a high-fidelity shadow path): inputs, tool calls, tool results, final environment state, cost, and latency.
  2. Score — run online checks on a sample: cheap deterministic invariants on every eligible trace, graded quality on a smaller slice with a pinned judge and rubric.
  3. Triage — decide what the miss is. Agent regression? Underspecified product rule? Infra flake? User noise? Only some of these belong in the eval suite.
  4. Freeze — redact, normalise, and commit the case as a golden with expected state and assertions.
  5. Gate — run that golden in CI alongside the rest of the prompt and agent suite so the next prompt, model, or tool change that reintroduces the failure turns the build red.

If step five never happens, you have observability theatre. Pretty spans. No memory.

This complements the work already covered elsewhere on this site. Prompt regression pins the instruction surface and fails the build when goldens slip. Agent release gates compare candidates on task completion, invariants, and recovery before promotion. Hallucination checks catch confident wrong answers in coding agents. Trace-to-eval is what happens after those gates: the failures users actually hit become the next fixtures those gates run.

What to capture on a trace

If you only store the chat transcript, you will grade the wrong thing. The agent's message can be polite and wrong while the order service is correct — or polite and correct while the backend has two replacements. I care about the state the user paid for.

For a fictional support agent that can look up orders, create replacements, and issue refunds, a useful trace records at least:

Field Why it matters
User input + session context The ask that started the run, including prior turns the agent saw.
Prompt / model / tool versions Without pins you cannot replay or attribute the miss.
Tool call sequence Names, arguments, ordering — including retries and abandoned calls.
Tool results What the environment returned, including errors and timeouts.
Final environment state Order count, refund amount, ticket status — inspected independently of the transcript.
Agent final message Needed for communication checks; never the sole oracle.
Cost and latency Tokens, tool time, end-to-end duration — operational regressions are still regressions.
Outcome labels Success, unresolved, invariant violation, infra failure — set by scorers, not by the agent's self-report.

OpenTelemetry-style spans work fine as the carrier. The point is not the vendor. The point is that every consequential side effect leaves a durable footprint you can assert against later.

Write expected state before you write a grading prompt. If the team cannot say "exactly one replacement on order X, original line preserved, refund ≤ paid amount," the feature is still underspecified — and no amount of LLM-as-judge will fix that.

Online scoring without lighting money on fire

You cannot score every production trace with a frontier judge on every dimension. You also cannot score nothing and call the absence of alerts a quality programme.

The scorecard I would start with for online scoring:

Layer What runs Cadence Release / ops decision
Critical invariants Deterministic checks on state and tool contracts (no duplicate order, no cross-customer write, schema-valid tool args) High sample rate or all eligible writes Any violation pages and enters triage immediately
Graded quality Pinned LLM-as-judge rubric on faithfulness, policy adherence, communication accuracy vs state Lower sample rate (illustrative starting point: 1–5% of sessions, higher on new intents) Trend and threshold breach → triage queue, not instant rollback by default
Operational envelope Latency and cost per completed task vs agreed budget Continuous aggregates Sustained breach is a release and capacity conversation

A few disciplines keep this honest:

Treat sample rates and thresholds as illustrative defaults until your traffic and risk replace them. The operating model matters more than my numbers.

Triage: not every red span is a golden

Promotion without triage fills the suite with noise and trains the team to ignore it. I use four buckets.

1. Agent / prompt / tool regression → promote to golden.
The environment state is wrong (or the message falsely describes a correct state) under a policy the product already agreed. Example: user asked for one replacement; two were created after a lost tool response. Redact, freeze inputs and expected state, add deterministic assertions plus a graded communication check. Link the case to the incident.

2. Product bug or missing rule → fix the product, then decide.
The agent followed an underspecified or wrong policy. Shipping a golden that encodes the wrong rule locks the bug in. Fix the rule or the backend constraint first; then add a golden for the corrected expected state.

3. Infrastructure / dependency failure → track separately.
Provider outage, auth expiry, queue backlog. Do not teach the eval suite that "timeout" is the correct product outcome unless unresolved handling is the behaviour under test. Keep recovery fixtures in the release suite; do not spam CI with raw outage traces.

4. Noise / abuse / out-of-distribution chatter → discard or park.
Jailbreak attempts you already cover with adversarial probes, incomplete sessions, or one-off gibberish. Log for safety analytics if needed; do not dilute the regression corpus.

A simple rule of thumb: if replaying the redacted inputs against the current agent should fail the build when the bug returns, promote. If the right fix is a backend invariant, a product ticket, or an ops runbook, do not pretend a golden is the fix.

Redaction before the scar becomes a fixture

Production traces contain things that must never land in git: names, addresses, payment references, internal IDs that map to real people, free-text that quotes a customer complaint verbatim.

Before promotion:

If you cannot redact a class of traces safely, you do not promote that class. Privacy is a release gate for the eval suite itself.

Close the loop into CI

A promoted case should look like any other golden in the prompt-regression and agent-release suites — same two assertion layers, same pinning habits.

Layer 1 — Deterministic (pass/fail, cheap):

Layer 2 — Graded with a pinned rubric (thresholded):

Run the small "production scars" slice on every PR that touches prompts, tools, or agent policies. Keep a larger scar corpus on nightly or pre-release. Version the dataset. When you add a case, the PR says why — link to the incident or triage ticket.

The mechanics stay boring on purpose: path-filtered CI job, pinned model and judge versions in the manifest, fail on threshold breach, publish the report as an artefact. The only new ingredient is the provenance tag: source: production-trace, incident: INC-1234, promoted_by: qa, redaction: v2.

Who owns the loop

Observability without an owner becomes a museum of screenshots. Eval without an owner becomes a folder nobody trusts.

I put QA / Quality Engineering on the pipeline: sampling config, scorers, triage queue health, redaction standards, and promotion into the suite. Engineering owns instrumentation quality (spans that actually contain state), tool idempotency, and backend invariants that should never rely on the agent alone. Product owns policy clarifications when triage exposes a missing rule.

Humans must review promotion when the case:

Automation should file the candidate and fail the build later. It should not silently grow a golden set that encodes yesterday's confused workaround.

The through-line

Staging goldens test the failures you already imagined. Production is where the failures you did not imagine show up. The teams that get agent quality under control are not the ones with the prettiest trace UI. They are the ones who treat a scored, redacted production miss as unfinished work until it can fail CI.

Observe the run. Score state and invariants, not vibes. Triage with a cold eye. Redact like the fixture will be public someday — because in a repo, it is. Promote only what should break the build forever. Own the loop.

Pre-release gates prove the agent you intended to ship. The trace-to-eval loop proves you intend to remember what broke.

That is how production failures stop being stories you retell in stand-up and start being cases your pipeline refuses to forget.

Continue reading