Skip to main content
A snapshot stores every step of a known-good run, in order, with the input your agent sent to each step and the output it got back. Comparing a new run against it shows exactly what your agent now does differently.

How snapshots fit in

You save a snapshot under a name, from a trace ID or from a run you just executed. The snapshot is stored self-contained, and its source trace is pinned, so neither is removed by trace retention. Saving again under the same name replaces the baseline. You then compare runs against it:
  • In tests, with the MatchesSnapshot assertion. A missing baseline fails the assertion.
  • In production, with CompareSnapshotAsync. Each result is recorded in the dashboard, with expected and actual steps side by side.

What a comparison reports

Every change comes with its position and, for input and output changes, a line diff.

Snapshots and replay

In a replay, model outputs are recorded, so step inputs are where a code or prompt change shows up. A snapshot catches that change with no LLM call. Whether the change made outputs better or worse is a question for a live Judge run.

Deep dives

Catch regressions with snapshots

Save a baseline, check runs in tests, and monitor production.

Replay

Run your agent against recorded traces.