Flaky Agent Regression Tests and How to Diagnose Them
Distinguish real agent regressions from flaky tests by tracing failures through specific layers.

An agent regression test fails intermittently, and the instinct is to call it flaky and move on. That instinct is usually right in traditional software, where flaky tests cluster around known causes like timing and concurrency, and it is exactly the wrong instinct for LLM agents, where the source of non-determinism is not the environment but inference itself. A failed run may not reproduce on the next attempt even when the defect that caused it is completely real and still present in the code. Agents add multi-step trajectories, live tool calls, memory retrieval, and branching workflows on top of this, so the same test can fail for three structurally different reasons that share no common fix. Treating all of them as one undifferentiated category of noise is how a genuine regression gets dismissed and then ships to production.
LLM agent non-determinism versus ordinary test flakiness
Traditional flaky tests come from async timing, concurrency, resource contention, or test-order dependency sitting on top of a codebase that is otherwise deterministic. If you stabilize the environment, the flakiness usually goes away, because the underlying system being tested does not change its behavior from run to run. An LLM agent breaks that assumption at its foundation. The model's own sampling process is not bitwise-reproducible, so if you run the identical prompt against the identical model version, you get no guarantee of the same token sequence, let alone the same tool calls or the same branch through a workflow. That variance does not stay contained to a single step. In a multi-step trajectory, each step's output feeds the next step as input, so a small deviation at step 2 propagates into steps 3, 4, and 5, growing as it goes. The failure a developer actually observes may sit one or more steps downstream of where the instability started. Retrying the test or quarantining it, the standard playbook for ordinary flakiness, fixes nothing because it never touches the actual source. Diagnosing an agent failure correctly means tracing it back through the layers it passed through, not treating the final symptom as the whole story.
Three failure categories that look like flakiness but are traceable regressions
Most of what gets labeled flaky in agent testing actually sorts into three categories, and each one points at a specific layer with a specific fix once it's identified.
Tool schema drift is the first. It has two distinct failure paths. A schema mismatch, where a required field or type no longer matches what the tool expects, throws a runtime error that's visible and easy to catch. A description mismatch is more dangerous: the tool's interface still works mechanically, but its description now nudges the agent to call it in situations it shouldn't, and there's no error signal anywhere in the stack. The silent version of this plays out in a specific sequence: the agent sends a stale or partial payload, the tool coerces or quietly drops the mismatched fields, the response comes back looking valid, the agent narrates a confident result to the user, and the mismatch becomes visible only later when someone notices the outcome was wrong. Throughout that whole sequence, a test watching for HTTP 200 would have passed every time. If a tool schema changes, every status code can stay green even while the agent picks the wrong function for a cancellation, a refund, or a deletion request, so no status-code check will ever catch it. This looks like flakiness because it's intermittent in a specific way: the wrong function only fires when a user's phrasing happens to land on the part of the description that drifted, not on every run.
Prompt structural gaps are the second category. A prompt edit aimed at making responses sound friendlier can quietly strip out a requirement for citations, producing a regression in groundedness that a surface read of the output would never catch. These failures vary in how often they show up, but the same structural gap sits in the prompt regardless of which input triggers it, so once the right evaluator is checking for it, the failure becomes fully traceable.
Memory-driven and workflow loop failures make up the third category. Most harnesses handle long-running agents through manually built heuristics: trajectory summarization, retrieval-based memory, context compaction, retry rules, tool-call validation. None of this gets easier just because context windows get larger. A bigger window does not stop a trajectory from accumulating stale, redundant, or low-signal information over time, and a long-running agent can start looping or degrading in quality for reasons that have nothing to do with the model changing. Its memory state drifted, and that drift is the thing to inspect, not the model.
SRE-level signals and the regressions that look like flakiness
Teams running agents in production usually watch latency, retry rate, token cost, and escalation rate, and all four are legitimate signals. None of them tell you whether answer quality actually got worse. A tool selection regression on refund requests, the kind of failure described above, can run with zero change in latency, leaving every one of those dashboards flat while the agent quietly routes refund requests to the wrong function.
The structural reason this happens is that a single user request typically passes through planning, retrieval, tool selection, function calling, JSON formatting, and response generation, and a regression at any one of those steps can hide behind a final response that still looks fine. If the agent partially recovers from a failure at step 3, the run is slow rather than wrong, and slow is not the alarm anyone is watching for. Research into automated root cause analysis for microservices has identified the same structural gap: existing evaluation methods judge whether a diagnostic approach correctly localizes the responsible service, but that endpoint correctness says nothing about the evidentiary basis for the diagnosis or the actual propagation path connecting the fault to the symptom a user sees. Agent monitoring has the equivalent blind spot. When a retriever index gets rebuilt, latency numbers can stay normal even as the rate of unsupported claims in the agent's output quietly rises, so you need a groundedness or faithfulness check built for that purpose to see it. Product teams only find out when support tickets start contradicting the release notes. Compliance teams find out when they go looking for evidence that a reviewed policy case was actually rechecked and discover it wasn't. Both of these discoveries land days or weeks after the regression itself happened, because none of the signals available in the moment were built to catch it. The existing evaluation frameworks in wide use score whether a given output is acceptable, but they stop there. They don't isolate which layer of a recorded run actually produced the failure, and without that attribution, there's no way to tell a flaky run apart from a regressed one.
Layer-by-layer root cause attribution as the diagnostic method
Classifying a failing agent test correctly means attributing the failure to a specific layer, prompt, tool, workflow, memory, or model, rather than scoring whether the final output looks acceptable. Only attribution at that level of specificity can separate a genuine regression from ordinary variance. The goal of attribution is to answer three questions: which layer is at fault, at which step the critical error occurred, and why it occurred there. The hard part is reasoning backwards from an observed failure, because the interaction data mixes LLM reasoning, tool calls, and environmental feedback into a single tangled trace.
If you ask an LLM-as-judge to explain why a long trace failed, you run into a specific and well-documented failure mode. Ask the question once and the judge gives a confident answer. The cause is premature commitment: the judge settles on a plausible root cause after reading only part of the trace, and a long trace never runs short on plausible-looking causes to settle on. The Continual Search method was built to address this directly. Instead of asking a judge to recheck its own conclusion and confirm it, Continual Search prompts the judge to challenge its current diagnosis: it actively searches for evidence it hasn't yet examined, and it repeats that search over successive turns until the unresolved diagnostic evidence runs out. That reframes root cause attribution as a search problem rather than a scoring problem, which is a meaningfully different exercise: scoring asks whether an output looks right, search asks what evidence would prove the current theory wrong.
If you put this into practice, you build a regression gate around a specific set of per-evaluator metrics. Track Groundedness, HallucinationScore, ToolSelectionAccuracy, and JSONValidation as deltas measured against the last passing baseline, not as standalone scores. Slice the eval-fail-rate by product area, by tool route, and by prompt version, so a regression concentrated in one slice doesn't get diluted into an average that looks fine. Keep trace-linked regression rows that connect a failed eval directly to the specific trajectory step responsible for it, so the question of which layer failed has a concrete answer.
The release decision that falls out of this is a cohort-level one, not an aggregate one. A candidate release can raise the overall pass rate and still be a real regression if any release-critical evaluator or cohort crosses its threshold. Aggregate scores can hide exactly the kind of cohort-level failure that matters most, so the gate has to check cohorts individually rather than trusting the average to surface problems on its own.
Replay-based testing as the mechanism for turning recorded failures into stable regression gates
Knowing which layer failed is only half the problem. You also have to prove that a fix actually resolves that failure without reintroducing the same variance that made the original test unreliable. Most agent testing infrastructure today can observe a run but can't control it: tracing tools record what happened, and evaluation frameworks score whether the output passes, but neither one lets a developer change a single component of a recorded run and check whether that change fixes the failure while holding everything else fixed.
Chronicle, a method for cut-point replay of LLM agent runs accepted as a paper at the REALM workshop at EMNLP 2026 (authored by Tisha Chawla and Susheem Koul), addresses this gap directly. Chronicle records an agent run at its non-deterministic boundaries, and it saves them as fixed, replayable records. Its central operation, cut-point replay, feeds a chosen subset of those recorded boundaries back exactly as they happened, and it executes the remaining subset live against new code. A recorded incident turns into a regression test that runs in continuous integration, with full control over which part of the system is being tested and which part is held constant from the original failure.
You can see what makes this useful in what happens with mutants, bugs deliberately introduced to test whether a test suite actually catches regressions. Chronicle's cut-point tests use a fixed assertion, and they catch every mutant that would let the recorded unsafe action through. A baseline approach that stubs out every boundary and checks the same assertion catches none of them. So cut-point replay catches real regressions that full mocking simply misses, because full mocking strips out the exact conditions under which the regression happens.
Full replay under Chronicle is bit-stable across repeated runs, and the paper's own results confirm this across 20 repetitions. A cut-point test that passes is stable because it was built to be stable, and a cut-point test that fails points to a genuine regression in whatever component changed. Recording overhead is negligible, and full replay issues zero model calls, so you can run this in CI without burning any inference budget, and that matters if your team runs these checks on every pull request.
Teams without infrastructure built specifically for cut-point replay can still apply the underlying principle. Fix a recorded non-deterministic boundary, whether that's a tool response, a retrieved document, or a model output, and replay the rest of the trace around that fixed point. That isolates whether a failure lives in the layer that changed or in ordinary model variance, so it works as the agent-testing equivalent of dependency injection in conventional software. Shipping a fix without validating it against a replayed version of the actual failing trace is shipping on faith: the fix might look correct in isolation, but there's no evidence it resolves the specific failure mode that caused the regression.
Harness iteration as the primary fix path before reaching for model-level solutions
Once attribution points to the layer responsible, prompt, tool interface, workflow logic, or memory handling, the fastest and safest fix is almost always a change to the harness around the model, not a change to the model itself. LLM tool agents can improve substantially without any retraining at all, through changes to prompts, tool interfaces, middleware, state handling, and recovery logic. The harness, not the model's weights, is where most of this optimization work should happen, because harness changes are faster to make, easier to review, and directly testable against the same recorded traces used to diagnose the regression.
A 2026 paper on harness optimization for deterministic agents states this goal precisely: improve task performance by adapting the runtime interface sitting between the frozen model and its environment, without touching model weights or evaluation environments. Attribution identifies the layer. Replay confirms the fix. The harness is where the fix usually lives, and reaching for a model swap or a retraining run before exhausting harness-level options skips past the layer most regressions actually come from.


