Architecture
Your application callsRunAsync with an agent name and a callback. The SDK passes your agent a context. Each LLM call or sub-operation your agent wraps in TrackStepAsync becomes a step. When the run finishes, the SDK ships the trace to Veval and returns your agent’s output. If your agent throws, the SDK ships the trace with an error status and rethrows the exception.
Traces are ingested asynchronously. A trace can take a moment to show up in the dashboard and the API.
Recording
1
Wrap the run
RunAsync starts a trace, runs your agent, and records the input and output.2
Track steps
TrackStepAsync records a step’s input, output, and duration. Use the step handle to attach the model, token counts, and cost.3
Ship the trace
The SDK sends the trace to Veval. It appears in your dashboard with full step detail.
Testing
A recorded trace holds the output of every step. That’s what makes it testable.VevalTestSdk is a test double for the SDK. Load a trace with WithReplay and your agent runs normally, except that each TrackStepAsync call returns the recorded output instead of calling the LLM. See Replay.
After a run, assertions check the result: no errors, cost under a limit, a step or tool call present, an output graded by an LLM judge. A snapshot compares the run against a stored known-good baseline and reports what changed.
Scenarios group these into a named set of test cases. Each scenario run posts pass/fail results to the dashboard so you can track quality over time.
Replay is strict
If your agent callsTrackStepAsync with a step name that isn’t in the recorded trace, VevalTestSdk throws. It never falls through to the live LLM. New LLM calls fail the test instead of quietly costing money.
Next steps
Quickstart
Install the SDK and send your first trace.
Concepts
Traces, steps, scenarios, assertions, replay, and snapshots.
Instrument your agent
Record steps, metadata, and tool calls.
Test with replay
Run your agent against recorded traces in tests.