Snapshot Testing for Agent Tool Call Sequences
Catch silent agent failures by testing tool call sequences, not just final answers.

An agent can read the right files, reason through a problem out loud, call tools that look sensible in the logs, and still hand back an answer that is wrong in a way no one notices until a customer does. The gap this article is about is simple to state: engineering teams that only grade the final answer are grading the one part of an agent run that tells them the least about what actually happened. A plan, a tool call, an argument binding, a routing decision: each of these is a separate place where a run can go sideways, and none of them show up in a string match against an expected output. The industry's benchmark culture has trained teams to treat the last message as the test, but a model's leaderboard score says nothing about whether it calls your tools in the order your workflow assumes, skips a lint step it used to run, or reads twice as many files as it needs to before making an edit. Tool-call sequences are a steadier signal than the final answer, because they record what the agent actually did rather than how plausible its summary of that sounds.
The specific failure categories that tool-call sequence testing is designed to catch
Schema drift is the clearest case. A tool argument gets renamed somewhere upstream, customer_id becomes account_id, and the agent keeps calling the tool anyway. Because language models are built to be helpful rather than cautious, the model often fills the missing field with a guess instead of failing loudly, so the call succeeds and the downstream data is wrong. Wrong tool order is a related but separate failure: the agent edits a file before reading it, or runs a lint check after a deploy instead of before one, and each individual call can succeed while the sequence as a whole produces an outcome nobody validated. Skipped steps are quieter still, a tool that is supposed to appear in every run simply drops out after a prompt edit or a model swap, with no exception thrown and no CI check tripped. Multi-step workflows carry a fourth problem: binding drift, where an entity the model attaches to a reference in an early step is no longer the thing a later tool call actually receives by the time it reaches step four, a failure mode that research on tool-augmented agents treats as its own distinct category. Agents built with no exit condition, no fallback behavior for a tool that errors, and no format specification for whatever consumes their output downstream carry a structural weakness rather than a one-off bug, and these three gaps alone account for most runs that quietly fail. What ties the whole list together is that none of these failures throws an exception, none of them fails a happy-path eval, and none of them necessarily produces output that looks wrong on its face. They are invisible to anything that only scores the answer, and catching them requires looking at the trajectory itself.
How snapshot testing of tool-call sequences works
Snapshot testing of tool-call sequences takes a known-good run, records its full trajectory, and commits that record as a golden file; every later run gets compared against it, and CI fails the moment a run takes a structurally different path. The golden run itself captures more than just the final answer: the names and order of every tool call, the arguments passed to each one, and which tool the model itself asked for, as distinct from whatever the harness actually ended up executing, with LLM responses and the final output recorded as optional extras. Comparison against that golden file runs across a few distinct dimensions rather than a single pass-fail check. Tool call names and order get compared structurally, using Levenshtein edit distance on the sequence of calls. Arguments get compared through a deep diff, with certain fields configurable to be ignored so that legitimately variable values don't trip a false alarm. The tools the model itself requested get compared separately from what the harness executed, which matters because a model can quietly ask for a different tool even when the harness's own dispatch logic is unchanged. Final responses and output get compared on a semantic basis, either through cosine similarity or an LLM acting as judge. When any one of these dimensions drifts past its configured threshold, the test doesn't just fail silently. It raises a structured diff that a human can read and reason about, which is the point: the output is a reviewable report, not an opaque score. AgentSnap is one working example of this pattern in practice. It records golden runs, writes them out as committed snapshot files, and on every later run replays the same inputs and diffs the new trace against the old one; a representative diff it surfaces looks like [MODEL TOOLS] Model-requested tool sequence changed (edit distance 1): ['search'] -> ['delete_file'], which is exactly the kind of change a prompt injection produces when it silently redirects what the model asks a tool to do. Any version of this tool follows the same harness pattern: wrap the tool executor with a recorder, run it against a fixed list of tasks, write the recorded sequence out as a JSON golden file, and commit it. Running the same harness against a different model and comparing the new trace to the committed golden turns "did the model change behavior" into a concrete, line-by-line diff instead of a guess. The fact that this golden file lives in the repository, checked in alongside the code it tests, is what turns the whole exercise into a gate rather than a one-time sanity check someone ran once and forgot about.
Replay mode and live mode as two complementary CI jobs, not one
Snapshot testing of tool sequences only works as a discipline if it runs as two separate CI jobs, because the two jobs catch two failure classes that don't overlap. Replay mode runs on every pull request. It replays recorded LLM responses instead of calling a real API, which makes it deterministic, free to run, and immune to the flakiness that comes from live model calls; because no live call is made, the comparison shifts to what the code itself sent, and the test fails if the code sends a different prompt, makes a different number of calls, or produces a different tool sequence than the golden run did. That makes replay mode the right place to catch prompt edits that changed behavior by accident, broken tool wiring, a call count that crept up, or a loop that no longer exits the way it used to. Live mode runs on a schedule, typically nightly, and makes real calls against whatever model is currently in production. Live mode catches the drift that replay mode structurally cannot: a provider pushes an update to the model, a temperature setting shifts the distribution of outputs, a new default behavior appears on the vendor's side, none of which show up until the agent is actually talking to the live API, so this is where that kind of regression gets caught before it reaches customers. Within replay mode, tool calls still execute for real by default, but a configuration flag can stub them from the recording entirely so that the run has no side effects at all, which is the right setting for a pure regression check that shouldn't touch any real system. The cadence this produces, cheap and fast on every commit, slower and more expensive on a schedule, is the same split that ordinary software testing has used for decades between unit tests and nightly integration suites. Research on a related technique called cut-point replay, from Chawla and Koul (arXiv 2609.20625, EMNLP REALM 2026), gives the clearest evidence for why selective replay beats stubbing everything. Cut-point replay records an agent run at the points where it behaves non-deterministically, storing each of those points as an immutable envelope; given a proposed fix, it can then serve a chosen subset of those recorded points back to the agent while running the rest of the system live with the new code, turning a past incident into a regression test that isolates the changed code from everything recorded around it. Recording that non-determinism costs almost nothing against the cost of a normal model call, and a full replay makes zero model calls while staying bit-stable across 20 repeated runs. Across all six benchmark incidents tested, cut-point tests built this way failed correctly on faulty code and passed correctly on changes that were either guarded or benign, catching every mutant that let an unsafe action through, while a baseline approach that stubbed every single boundary caught none of them.
Snapshot diffing versus prompt diffing and output scoring
Diffing the tool-call sequence catches a category of regression that neither diffing the prompt text nor scoring the final output can reach, because the failure is in what the model decides to do, not in the words the harness sends it or the answer it eventually prints. Prompt diffing tells you when the text going into the model changed, but it says nothing about a model that reinterprets a stable, unchanged prompt differently after a provider-side update, and that regression appears with no code change anywhere in sight. Output scoring tells you when the final answer got worse, but an agent can produce an answer that reads just as plausibly as before while getting there through the wrong tool path, especially on any task with enough ambiguity that more than one route looks defensible; the scoring layer sits downstream of where the actual mistake happened, so it has nothing to say about the mistake itself. The dimension that closes this gap is the one comparing which tool the model itself asked for, independent of what the harness executed, and it fails the moment that request changes even if the harness's own logic is untouched. The AgentSnap diff cited earlier, [MODEL TOOLS] Model-requested tool sequence changed (edit distance 1): ['search'] -> ['delete_file'], is a real instance of this: a prompt injection rerouted the tool sequence without altering a single line of harness code, and neither an output score nor a prompt diff would have seen it coming. Not every diff raised this way is a regression, and treating every diff as an automatic failure will burn out a team's trust in the gate fast, so the diff has to be read rather than just obeyed. A tool that moves from before an edit to after it is a real process regression and should block the merge or prompt a tightened system-prompt constraint before the gate reopens. Extra search or read calls that still land on the same outcome may just reflect more cautious exploration, worth checking against token count and wall time before deciding whether it's a problem. Fewer read calls happening before an edit than the golden run had is a warning sign for edits made against stale context, and it deserves to be treated as a regression until someone proves otherwise. The same tools appearing in a different order within one phase of the workflow is usually harmless nondeterminism, and the fix there is often to loosen the assertion to compare the set of tools used within a phase rather than their exact order. A tool appearing where none was used before is a capability change worth auditing against sandbox permissions immediately, since it may mean the agent gained access to something it shouldn't have. None of this triage works without configurable thresholds and the ability to exclude fields like timestamps or request IDs from the argument diff, since those fields change on every run for reasons that have nothing to do with agent behavior, and a gate that flags them constantly will train engineers to ignore it. What makes this approach defensible as a long-term practice is that the output an engineer gets back is a diff they can read line by line and reason about.
The harness as the right unit of intervention when a snapshot fails
When a snapshot test catches a regression, the fix almost always belongs in the harness, not in the model itself, because production research on agent systems increasingly treats the harness as the layer doing most of the work. An LLM agent running in production is really a model wrapped in a harness, and the harness is everything around the model: the prompt template it's fed, the set of tools it's permitted to call, the memory or retrieved context placed in front of it, how it plans and revises a multi-step task, and how it checks its own output before returning it. That harness, more than the model sitting underneath it, is increasingly understood to be the main thing driving how an agent actually behaves. That reframing changes where a team should look first when a snapshot diff comes back red. Root-cause attribution as a discipline exists to trace a failed run back to the first point where it diverged from a correct execution and assign the fault to the model, the harness, the environment, or the grader, so that whatever fix follows lands on the right layer instead of the convenient one. A model-side failure should feed back into training objectives, a harness-side failure should drive a change to the prompt template, the tool schema, or the workflow logic, and a failure traced to the environment or the grader should trigger a fix to the benchmark or the test setup rather than to anything upstream of it. Cut-point replay makes this kind of targeted iteration concrete rather than aspirational: it lets a developer change one single component of a recorded run and check directly whether that one change fixes the failure while every other variable in the run stays fixed exactly as recorded. The most useful raw material for this kind of harness work isn't a benchmark at all but production traces, because they show what real users actually did with the system rather than what an engineer guessed they might do when the eval set was written. A workable discipline here keeps a locked golden set reserved strictly for final regression checks, alongside a separate development eval set where a fraction of the cases get swapped out monthly for fresh traces pulled straight from production. That rotation keeps the harness tested against the behavior customers are actually producing, rather than a frozen guess about it from months earlier.


