Est.

Regression Test Prioritization Based on Production Failure Frequency

Prioritize regression tests by actual production failure patterns, not imagined ones.

Contributing Editor · · 9 min read
Cover illustration for “Regression Test Prioritization Based on Production Failure Frequency”
Regression Testing · October 7, 2026 · 9 min read · 2,123 words

Regression testing for LLM agents has been built backward. Teams write test cases before deployment, based on failures they imagine, but when they get to production, the failures that actually recur look nothing like the ones they expected. The fix is to rank regression tests by how often a given failure pattern has actually surfaced in live traffic, because only live traffic tells you which harness failures repeat and which were one-offs.

Production traces as a signal for regression testing

A static golden set, built before an agent ever runs in production, reflects the failure modes a developer imagined while writing it. It cannot reflect the failure modes that will actually recur once the agent is handling real traffic, because those failures have not happened yet. That gap matters more for agents than for any prior generation of software, for a reason specific to how agents execute.

Agent runs do not reproduce bit for bit. A model's inference varies even at low temperature, the tools an agent calls can read state that changes between one call and the next, and retry or routing logic can change how many steps a run takes from one execution to another. A hand-written test case that was never seeded from an actual failure is, in effect, testing a trajectory that may never occur in the wild. It checks a path the agent might take, without confirmation that the agent has ever been observed taking it.

The evaluation surface itself is wider than anything a plain LLM application presents. A single agent run moves through planning, tool selection, argument construction, intermediate reasoning, handoffs between steps or sub-agents, and finally an output. A failure introduced at any one of those stages propagates forward through every stage after it. Checking only the final output, which is what most hand-crafted test suites do, catches the symptom and misses the step where the error actually began.

Production traffic shows which failure patterns recur often enough to matter and which are isolated incidents that happened once and won't happen again. Without that distinction, every team runs regression tests in the dark, unable to say which of its test cases guard against something that actually threatens the product and which guard against something that happened once, to one user, under conditions nobody will see again.

How agent failures cluster into attributable, recurring categories

Agent failures are not noise.

Tool errors and schema drift make up the most visible cluster. A tool's contract changes somewhere upstream, the agent keeps following yesterday's schema, and the mismatch triggers a retry loop. Nothing crashes. The agent simply becomes less reliable over time, failing in a way that looks like random flakiness until someone traces it back to the schema change that caused it.

Workflow loops and prompt ambiguity form a second cluster, rooted in goal alignment. An agent with no defined exit condition loops indefinitely. An agent with no fallback handling halts the moment its first tool call fails. An agent given no format specification produces an answer that is correct in substance but unparseable by whatever system expects to consume it. None of these failures are visible by inspecting a single output in isolation; each only becomes visible across the full trajectory of steps the agent took to get there.

A third pattern is harder to catch than either of the first two: fabrication after a failed tool call. When a tool call fails, an agent will sometimes fill the gap with an asserted value that the tool never actually returned. So the original failure compounds silently, and the output that results reads as perfectly normal to anyone who isn't looking for it.

Outcome metrics (did the task succeed or not) can tell you that something broke. They cannot tell you which tool, which argument, or which step was responsible. Only trajectory-level evaluation, which examines the full sequence of actions an agent took rather than just what it produced at the end, can answer that question.

Root cause attribution: assigning each failure to the layer that caused it

A team can know a run failed and still know almost nothing useful. Useful prioritization depends on knowing which layer caused the failure: prompt, tool schema, workflow logic, memory, model, or product logic. Without that attribution, engineering effort gets spent wherever intuition points, leaving the urgency of each failure unclear.

Root cause attribution identifies the first point where a run's execution diverges from a correct one and assigns responsibility to the component that caused the divergence. That single distinction determines where intervention should go: a harness failure should direct engineers to harness engineering, an environment or grader issue should trigger infrastructure repair, and only a failure attributable to the model itself justifies touching the model. Conflating these categories wastes the most expensive kind of engineering time there is, the kind spent retraining or fine-tuning a model to fix something that was never the model's fault.

The practical difficulty is that as execution logs grow longer and more distributed across tools, agents, and retries, automated diagnosis tends to settle on whichever failure looks most plausible rather than tracing the actual causal chain back to its origin. That tendency gets worse specifically in the runs that matter most: the ones with loops, retries, or multi-agent routing, which are exactly the runs where a shallow diagnosis is most likely to stop short of the real cause.

AgentDebugX (arXiv:2607.18754, July 2026) was built to close that gap. It organizes debugging as a closed loop of Detect, Attribute, Recover, and Rerun, and when a case proves too difficult for the standard loop, it escalates to DeepDebug, a multi-turn diagnosis agent that produces an auditable root-cause report complete with supporting evidence and a recommended fix. The output of this kind of attribution is a failure record carrying a layer label: prompt ambiguity, schema drift, a missing exit condition, a retrieval gap. That labeled record is the actual unit a frequency-ranked regression dataset needs in order to grow in a way that reflects reality.

Replay-based record keeping: turning a live failure into a rerunnable regression case

If a production failure cannot be reproduced on demand, it cannot become a regression test. That constraint is the whole reason frequency-ranked prioritization depends on a recording step built to preserve exactly the parts of a run that are non-deterministic: the model's sampled output, the state a tool read at the moment it was called, the path a retry or routing decision happened to take.

Chronicle (arXiv:2609.20625, Chawla and Koul, September 2026) was built to solve that problem directly. It records an agent run at its non-deterministic boundaries as immutable envelopes, then replays a run from that record. That distinction turns a recorded production incident into a regression test that can run inside continuous integration, checked on every commit the same way a unit test would be.

Chronicle's central mechanism is what its authors call cut-point replay. A chosen subset of recorded boundaries gets served from the record, while the rest executes live against new code. That design means a candidate fix can be validated against the exact trajectory that surfaced the original failure, rather than against some synthetic approximation of what the failure might have looked like.

Chronicle's current benchmark supports loops, retries, and multi-agent routing, and these happen to be the failure modes you see most in high-complexity production agents. That coverage spans the failure modes most common in high-complexity production agents, but it does not cover every case, so teams running agents with failure patterns outside that scope should treat Chronicle as a foundation worth extending.

Building a frequency-ranked regression dataset from production traces

The most valuable regression suite is not the one with the largest number of test cases; it is the one where every single case corresponds to a failure pattern that has actually recurred in production, ordered by how often that pattern occurs.

Building that suite follows a loop with four steps. Production traces first get instrumented to surface failures with layer attribution, using the kind of root cause attribution output described above: a labeled record pointing to prompt, tool, workflow, or memory as the source of the failure. Each surfaced failure then gets recorded at its non-deterministic boundaries, using a capture mechanism in the style of Chronicle, so the exact conditions that produced the failure can be replayed later. Each captured failure then gets tagged with its attributed layer and added to a versioned dataset, and a recurrence counter attached to that tag increments every time the same attributed pattern occurs again in a new production trace. Finally, before each regression run, the dataset gets ranked by that recurrence count: the cases with the highest counts run first and serve as release gates, while low-recurrence or single-occurrence cases run later or sit in a review queue.

Ayyad et al. Frequency ranking applies that same logic to production failures rather than to benchmark instances: the subset evaluated on every commit becomes the subset most likely to catch the regressions that actually matter, instead of a sample chosen at random or by guesswork.

Tracking failures per cohort reveals regressions that an aggregate pass rate hides, because a release can raise the overall pass rate while quietly degrading tool-selection accuracy on one specific cohort, a refund workflow or a cancellation path, for instance, and an aggregate number will never show that degradation. The dataset that results from this loop grows on its own. Every new production failure that gets attributed, captured, and tagged adds a real case drawn from an actual incident, not a synthetic one invented at a desk, and patterns that keep recurring push their own position up in the ranking without anyone having to curate the list by hand.

Harness iteration as the right response to most ranked failures

The failure categories that dominate a frequency-ranked production dataset are tool errors, schema drift, prompt ambiguity, and gaps in workflow logic. None of those sit in the model layer, so the correct fix for most ranked cases is a change to the harness, because the model was never the layer the attribution pointed to.

Harness Continual Learning (arXiv:2608.19013, August 2026) formalizes this separation by keeping a frozen foundation model apart from a harness state that evolves rapidly and independently of it. The TaoLive Digital Avatar Agent (TaoLive AIGC LLM Team, arXiv:2608.15763, August 2026) takes a related but distinct approach, called Harness-Aware Training, training compact models to adapt to independently versioned harness components without retraining each time one of those components changes. So operators can change business behavior at high frequency without paying the cost of retraining for every edit. Retraining re-pays its own cost every time the harness changes, and that cost becomes unsustainable once failure patterns start recurring often, which is precisely the situation a frequency-ranked dataset is built to surface.

There's a concrete case for what harness-only fixes can achieve on their own: a coding-agent system improved measurably on Terminal Bench 2.0 purely through harness engineering, adding self-verification, better tracing, improved retry logic, structured termination rules, and more disciplined tool usage, without any change to the base model underneath it.

The practical rule that follows is simple to state. Run attribution on the ranked failures before deciding anything. If most of them point to the prompt, tool, workflow, or memory layer, ship harness fixes first and validate them against the captured traces before any model change is even on the table.

Validating fixes against production traces before shipping

Shipping a fix without replaying it against the traces that originally surfaced the ranked failure is shipping blind. A fix can pass every synthetic test in a suite and still leave the actual production failure pattern untouched.

Replay-based validation closes that gap by running the candidate fix against the captured trace records, the Chronicle-style immutable envelopes described earlier, so the fix gets tested against the exact execution context that produced the original failure rather than a reconstructed or approximated stand-in for it.

The regression gate that matters here operates per cohort, not on an aggregate score. If a candidate raises the overall pass rate but degrades tool-selection accuracy on the specific cohort tied to the ranked failure, it should be blocked outright, because the aggregate improvement hides the one regression the frequency ranking flagged as the highest priority to catch. If the same candidate swings materially across reruns on a fixed control slice, the gate itself is noisy, and that noise needs to be tightened before the gate can be trusted as a release signal.

A validated fix produces more than a passing regression run. It produces an updated dataset in which the fixed pattern's recurrence counter resets to zero, so the next round of production monitoring gets measured against a new baseline.

Sources

  1. Chronicle: Cut-Point Replay for Regression Testing of LLM Agents
  2. AgentDebugX: An Open-Source Toolkit for Failure Observability, Attribution, and Recovery in LLM Agents
  3. [2609.20625] Chronicle: Cut-Point Replay for Regression Testing of LLM Agents

More in Regression Testing