Operate ยท 7 min read

Testing and evaluation

Measure quality with datasets and graded runs before you ship.


Prompt changes behave like code changes with no type checker. An evaluation set is how you find regressions before your users do.

Build a dataset

  • Collect twenty to fifty real questions with reviewed answers.
  • Include the awkward cases: ambiguous phrasing, missing context, hostile input.
  • Version the dataset alongside the prompts it grades.

Grade the runs

Combine deterministic checks, such as whether a citation was returned, with model graded checks for tone and faithfulness. Deterministic checks catch the failures that matter most.

Note

Track cost and latency in the same run as quality. A prompt that scores two points higher and costs four times more is rarely the right trade.