Skip to main content
A scenario runs your agent against a list of items and checks every result with assertions. Each run posts pass/fail results to the dashboard, so you can track quality over time.

How scenarios fit in

You run a scenario with RunScenarioAsync, passing a scenario name, your agent, the assertions to apply to every item, and the items. Each item runs your agent once. Items can be defined in code or managed in the dashboard. If you don’t pass items, Veval fetches them for the scenario by name, so you can change test cases without redeploying code.

Item types

An item has either an input (synthetic) or a trace_id (trace-backed), not both. A trace-backed item’s trace must have recorded steps. If it has none, Veval throws rather than silently calling the live LLM.

Assertion scope

Results

A scenario run returns whether every item passed, the pass and fail counts, and a result per item with its failure messages.

Deep dives

Run a scenario

Define items, apply assertions, and read results.

Assertions

The checks that pass or fail each item.