Skip to main content
Veval records what your agent does and tests it against those recordings. You work with a small set of objects: traces record a run, steps record the calls inside it, and scenarios, assertions, replay, and snapshots turn recordings into tests.

Core concepts

Traces

A recorded agent run: input, output, steps, cost, and timing.

Steps

A single LLM call or sub-operation within a trace.

Scenarios

A named set of test items run against your agent.

Assertions

Checks that pass or fail a completed run.

Replay

Run your agent against a recorded trace with mocked LLM responses.

Snapshots

A stored known-good run that new runs are compared against.