Skip to main content
An assertion inspects a completed run and returns null to pass or a failure message to fail. Veval ships built-in assertions for common checks, and you can write your own by implementing ITraceAssertion.

How assertions fit in

You attach assertions to a scenario, to a single scenario item, or to a replay through ReplayOptions. Failure messages are collected on the result, so a failing test tells you which check failed and why.

Built-in assertions

Deterministic checks and the judge

Most assertions are deterministic. They read the recorded steps and need no LLM call. Judge grades the output against a rubric you write, using an LLM. The score and reasoning are recorded on the run whether it passes or fails. A judge call that fails outright counts as an assertion failure, never a silent pass. In replay, model outputs are recorded, so a code or prompt change shows up in step inputs, not outputs. To check whether a change made outputs better or worse, run the judge against live runs.

Deep dives

Write assertions

Use the built-in assertions, configure the judge, and write your own.

Snapshots

Compare runs against a known-good baseline.