Judge Model Selection for Automated Replay Evaluation
Treating judge selection as an engineering decision unlocks regression signals worth trusting.

Choosing a judge model for automated replay evaluation is an engineering decision. It carries real tradeoffs, capability, cost, bias, and how well the judge lines up with the failure modes an agent actually produces in production, and teams that treat the choice with that level of rigor end up with regression signals they can trust.
The pressure behind this decision is structural. LangChain's State of Agent Engineering report finds that 57% of organisations now have agents in production, with quality now ranking as the top barrier to deployment EvalAgent / AWS AI Labs. At that scale, manual review of live traffic simply doesn't work as an operating model EvalAgent / AWS AI Labs. Manual evaluation of 100,000 responses takes 50+ days, a number that makes the case for automation concrete rather than theoretical LLM as a Judge: A 2026 Guide to Automated Model Assessment | Label Your Data. That's the case for automated judging generally: it's credible, not just cheap LLM as a Judge: A 2026 Guide to Automated Model Assessment | Label Your Data.
The instinct that follows is predictable: if a model is strong enough, put it in the judge seat and move on. That instinct sounds reasonable until someone asks why, and then it gets harder to defend. Replay evaluation has a narrower job than general benchmarking. It exists to answer one question: does the agent still work on inputs it has already seen and already handled correctly? That's regression protection, not product experimentation, and it's a different task from the open-ended comparison work most judge benchmarks are built around. The narrowness of that job shapes what the judge actually needs to be good at, and that list is not the same list that makes a model look impressive on a leaderboard. What follows is an argument for treating judge selection as an engineering problem with measurable tradeoffs, because that's the only way to get regression signals worth trusting.
What a judge model is doing during replay evaluation
Agent evaluation is not output scoring. It requires assessing execution traces that contain sequences of reasoning steps, tool invocations, error recovery actions, and state transitions, a far more complicated object than a single response to grade. A judge working on replay data has to move through three layers to do this properly. Outcome metrics ask whether the task succeeded, a blunt, binary signal that tells you something broke without telling you where. Trajectory metrics assess whether the agent took the right path, including the tool choices, the order of decisions, and the quality of the reasoning along the way, and most of the useful signal lives here. System metrics track efficiency, cost, and reliability, which matter once an agent is running at production volume.
Skipping straight to outcome scoring hides real problems. A correct final answer can mask flawed reasoning along the way, and a failed output doesn't necessarily mean the agent handled errors poorly, it might have handled them well and still lost on a downstream step. Output-only evaluation misses both cases. A judge that only checks endings is blind to the drift that regression testing exists to catch.
Replay makes this harder still, because the judge isn't just scoring a trace in isolation, it's scoring a trace against a known-good historical one. It has to check whether a changed prompt, tool, or workflow produces a comparably sound trajectory, not merely a comparable ending. This is harder than generic LLM evaluation because the judge must reason about multi-step causal chains, not isolated responses. Traditional metrics like BLEU and ROUGE miss coherence, helpfulness, and factual accuracy even in simple cases, and they have nothing to say about causal chains across a trace. LLM judging is the practical alternative, but that makes the judge's own reliability, its properties, its biases, its blind spots, the bottleneck for the whole evaluation system.
The bias problem: how single-model judges introduce systematic errors into replay signals
Single-model judges carry the fingerprints of their own training. Model biases stemming from training data and optimization choices produce systematic evaluation errors and inconsistent judgments across domains, and no amount of raw capability erases that. Three biases recur constantly in the literature. Verbosity bias is a preference for longer answers regardless of whether they're right. Confidence bias is a preference for answers stated with certainty even when the certainty is misplaced. Position bias is a tendency in pairwise comparisons to favor whichever response occupies a particular slot.
The scale of the inconsistency this produces is visible in benchmark data. Well-known 70B judge-specialized models score in the 63 to 67% range on JudgeBench while scoring 90 to 91% on RewardBench, the same model looking reliable in one domain and shaky in another. That's not noise; it's a pattern, and it means a judge that handles tool-call trajectories well on one kind of task can misread the same kind of trajectory in a different domain.
This matters more in replay eval than almost anywhere else, because the errors don't announce themselves. A postmortem covering two years of incidents at a major retailer found a persistent attribution error rate of roughly 10%, where the model blamed a technology simply because it was mentioned somewhere in the incident thread, pattern-matching standing in for actual causal reasoning augmentcode.com. A biased judge doesn't produce noisy, obviously wrong scores that get caught in review. It produces confidently wrong scores in a consistent direction, which corrupts a regression signal quietly, without tripping any alarm augmentcode.com.
The fix that comes to mind first is an ensemble: run the trace past several judges with different bias profiles and let disagreement surface the risk. Multi-agent evaluation protocols do mitigate individual model bias through this kind of diverse perspective, but they're prohibitively expensive at inference time once you're running replay at production scale. That expense is the hinge the rest of this argument turns on.
The four tradeoffs that should drive judge selection: capability, cost, bias, and domain alignment
Capability against cost is the first and most obvious tension. A stronger model tends to produce better judgment quality, but the inference cost per trace climbs with it, and replay eval runs against far more traces than a benchmark run ever does. Purpose-built judge models can compress that tradeoff rather than eliminate it: Galileo's Luna-2 models, for instance, support full-traffic evaluation at sub-200ms latency and roughly 97% lower cost than a standard LLM-as-judge setup EvalAgent / AWS AI Labs. At replay scale, cost per trace compounds in ways a small pilot does not reveal: what looks affordable across 50 traces becomes a real budget line against full production traffic LLM as a Judge: A 2026 Guide to Automated Model Assessment | Label Your Data.
The second tension is general capability against domain alignment. A model that tops the leaderboards isn't automatically the best judge for a specific failure mode your agent produces. ICSE 2025 research found that incorporating code-specific knowledge improved root cause localization by 28.3% over the previous leading method, a substantial jump that came from domain knowledge, not from a bigger base model EvalAgent / AWS AI Labs. What matters is not which model is strongest in the abstract, but which model is strongest on the failure modes this particular agent actually generates.
Third is bias profile against scoring task. Different judges carry different bias signatures, and some rubrics are far more forgiving of those biases than others: verbosity bias barely matters when scoring binary tool-call correctness, but it can quietly wreck a rubric built around prose quality. Auto-Arena's committee approach, which reaches 92.14% correlation with human preferences, shows that pooling judges with different bias profiles can beat any single judge on its own, though that gain comes at the price of ensemble inference cost JudgePanel.
Fourth is consistency against sensitivity. Replay eval needs a judge with a low false-negative rate on broken trajectories, not simply a high average agreement score, because average accuracy can hide exactly the failures replay exists to catch. Consider an agent that succeeds 75% of the time per trial: pass@3, the chance that at least one of three attempts succeeds, comes out to 98.4%, while pass^3, the chance that all three succeed, is only 42%. A judge that scores individual runs without surfacing that gap will tell a team its regression suite is healthy when the underlying reliability is nowhere close.
How "premature commitment" corrupts root-cause judgments in long traces
As execution logs grow longer and more distributed across sub-agents and tool calls, judge models tend to settle on a plausible-looking failure as the root cause before they've actually explored the rest of the evidence. This is premature commitment, a specific and well-documented failure mode, not a vague worry about model reliability.
In replay eval, this is structurally dangerous. The judge is reading a full multi-step trace and has to attribute a regression to the right layer, the prompt, the tool, the workflow, the memory, rather than just flag that something somewhere went wrong. Get the attribution layer wrong, and the fix that follows targets the wrong component.
The Continual Search framework addresses this directly, by nudging the judge across successive turns to keep exploring evidence it hasn't yet examined rather than settling on the first plausible story. The gap this closes is measurable. On TRAIL, one of the two benchmarks with the longest execution logs, Continual Search raises GPT-5.5's Weighted F1 from 0.426 to 0.500, compared with 0.451 under Passive Continuation Continual Search framework. A judge left to reach its own conclusion and one prompted to keep looking produce different scores, and that difference is not small given how these scores are typically distributed Continual Search framework. Who&When Pro, a benchmark covering more than 12,000 labeled trajectories across agent frameworks, domains, and modalities, shows that premature commitment is not an occasional glitch but a systematic property of judge behavior Who&When Pro (Liu et al., 2026).
The engineering implication is that judge selection can't be separated from judge prompting strategy. The same model, given a passive single-pass prompt versus a search-nudging one, produces meaningfully different attribution accuracy on long traces. Teams that pick a capable model and then leave it running a static prompt are leaving accuracy on the table that a better prompting structure would have recovered for free.
Compact specialized judges vs. large general models: what the evidence shows
The evidence doesn't favor size on its own. JudgePanel, built on a 14B backbone, outperforms judge-specialized models as large as 70B across four evaluation benchmarks, and it posts best-in-class results on both JudgeBench (76.8%) and RewardBench (91.5%) github.com. A compact model, trained the right way, beats larger specialists on cross-domain consistency github.com JudgePanel.
The mechanism behind that result reveals something specific. JudgePanel trains on deliberation traces, records of structured discussion, disagreement, and resolution pulled from an ensemble of strong evaluators, so the model internalizes multi-agent reasoning while still running at single-model inference cost. It gets the benefit of a committee without paying the committee's bill. That stands in sharp contrast to the current default: existing 70B judge-specialized models perform inconsistently across datasets, strong on one benchmark and noticeably weaker on the next, which makes any single one of them a shaky sole source of truth for regression signals across varied agent behavior github.com. JudgePanel also ships a lightweight domain specialization module that adapts to a new evaluation domain with a few hundred labeled samples, a practical option for teams running agent-specific rubrics who can't afford a full retraining cycle.
A separate line of evidence makes the same point from the other direction. Frontier coding assistants, when asked to design and run evaluations without domain-specific evaluation knowledge, manage only a 30% execution success rate while generating more than 12 metrics per agent, a sign that raw capability doesn't automatically translate into good evaluation design Who&When Pro (Liu et al., 2026). EvalAgent, which encodes evaluation domain expertise as reusable, structured skills, improves Eval@1 from 17.5% to 65% and reaches 79.5% preference among human experts over baseline approaches. Both results point at the same conclusion: task-specific knowledge and prompting structure carry as much weight as raw model size, sometimes more.
Practical selection criteria: matching judge properties to your replay eval requirements
Start from the failure modes your production traces actually show. The judge should be evaluated on the error categories an agent produces, not on how it performs on unrelated general benchmarks.
Trace length and complexity set the floor. Short, bounded traces with a handful of tool calls and a clear success or failure condition are usually well served by a capable general model with a carefully structured single-pass prompt. Long, distributed traces involving multiple agents carry real premature commitment risk, and that calls for a judge with explicit search-continuation prompting, or one trained directly on deliberation data.
Attribution granularity comes next. Outcome-only regression checks, simply confirming that a task still passes, tolerate a lower judge capability bar, and cost optimization is a reasonable priority there. Layer-level attribution, pinpointing whether a regression came from the prompt, the tool, the workflow, or memory, demands a higher standard of attribution accuracy, and that standard needs to be validated against labeled traces before the judge is trusted with production decisions.
Throughput and cost budget shape what's even feasible. Replay run against full production volume turns per-trace cost into a real constraint rather than an afterthought. At that point, purpose-built evaluation models, Galileo's Luna-2 at sub-200ms latency and roughly 97% lower cost being one example, become a serious option rather than a curiosity. Ensemble or committee judging helps with bias but multiplies inference cost, so it's better reserved for high-stakes or genuinely ambiguous cases than deployed as the default for routine regression checks.
Domain specificity deserves its own scrutiny. An agent operating in a specialized domain, code, finance, medicine, customer support, needs a judge validated in that domain specifically, since cross-dataset inconsistency in judge models is a documented, not hypothetical, problem. A few hundred labeled traces pulled from real production traffic is often enough to fine-tune or calibrate a compact judge against domain-specific rubrics.
Bias profile against scoring rubric is the last check, and it's easy to skip. Audit a candidate judge's known biases against the rubric it will actually run: verbosity bias barely registers on binary tool-call correctness but does real damage on prose quality scoring. Position bias matters most when the judge runs pairwise comparisons between an old and a new agent version, so randomize ordering and watch for a systematic directional skew before trusting the result.
None of this substitutes for direct validation. Run any candidate judge against a set of traces where the ground truth is already known. A judge's own calibration is the prerequisite for trusting whatever regression signal it produces, not an optional extra step.
How replay validation with a well-selected judge closes the loop from broken trace to verified fix
Shipping a change without replaying it against real historical traces is shipping blind, and a well-selected judge is what turns that replay step into something meaningful rather than a box-checking exercise. Observability platforms that capture and replay detailed agent traces solve half the problem: they get the data in front of a developer, but they leave the harder question unanswered, which step was responsible, why, and how to repair it. The judge is what fills that gap.
A well-chosen judge supports two distinct checks before anything ships augmentcode.com. The regression check asks whether a proposed change reproduces the passing trajectories from the historical trace set, framed around pass^k rather than the more forgiving pass@k, since it's consistency across repeated runs that regression protection actually depends on. The attribution check confirms that the layer targeted by a fix, the prompt, the tool schema, the workflow logic, actually changed in the resulting trajectory, while the other layers held steady.
Putting those two checks together makes the judge the mechanism that connects a broken trace to a verified fix, no longer a mere scoring convenience augmentcode.com. Get the judge selection wrong, mismatched to trace length, blind to the relevant domain, carrying a bias the rubric can't absorb, and that connection breaks quietly, producing scores that look fine right up until the regression it missed reaches production.
Sources
- Auto-Arena: Automating LLM Evaluations with Agent Peer Battles and Committee Discussions
- JudgePanel: A Compact Judge with Panel Deliberation via Adaptive Multi-Reward Reinforcement Learning
- An Empirical Study of Automating Agent Evaluation
- LLM as a Judge: A 2026 Guide to Automated Model Assessment | Label Your Data
- Root-Cause Attribution Is a Search Problem: Continual Search for Long-Horizon Agent Failures
- LLM-as-a-Judge vs Human Evaluation


