Operate ยท 7 min read
Testing and evaluation
Measure quality with datasets and graded runs before you ship.
Prompt changes behave like code changes with no type checker. An evaluation set is how you find regressions before your users do.
Build a dataset
- Collect twenty to fifty real questions with reviewed answers.
- Include the awkward cases: ambiguous phrasing, missing context, hostile input.
- Version the dataset alongside the prompts it grades.
Grade the runs
Combine deterministic checks, such as whether a citation was returned, with model graded checks for tone and faithfulness. Deterministic checks catch the failures that matter most.
Note
Track cost and latency in the same run as quality. A prompt that scores two points higher and costs four times more is rarely the right trade.