Judge Model Selection for Automated Replay Evaluation
Replay evaluation needs judges that can detect multi-step failures, not just score final outputs.
Staff Writer
Darnell spent several years embedded with developer-tooling teams at two venture-backed startups, where he developed a specialty in evaluation infrastructure and replay-based testing workflows. His reporting focuses on the systems and rituals teams use to assess changes before and after they ship.
7 stories
Replay evaluation needs judges that can detect multi-step failures, not just score final outputs.
Replay eval prevents silent agent failures by measuring test coverage across harness layers.
Recording every interaction is the only reliable way to make agent behavior reproducible.
Production agents fail silently in ways staging tests cannot detect or diagnose.
Unattributed fixes mask root causes and compound system liability in production deployments.
Enterprise AI failures overwhelmingly stem from harness defects, not model reasoning itself.
Prompt changes need the same testing rigor as code.