← All postsEvaluation

A small evaluation set beats a big opinion

Dana Oyelaran · 2 June 2026 · 5 min read


Teams delay building an evaluation set because it feels like infrastructure work. It is not. Thirty examples in a file is enough to turn an argument about wording into a measurement.

Start from complaints

The best first dataset is the list of answers someone already told you were wrong. Those cases are real, specific, and they are the regressions you most need to catch.

Grade what you can grade cheaply

  • Did it cite a source? Deterministic.
  • Did it refuse when the context was empty? Deterministic.
  • Was the tone right? Model graded, and worth less than you think.

Run it on every change

An evaluation set that runs occasionally is documentation. One that runs on every prompt edit is a test suite, and it is the only thing standing between a one word tweak and a quiet quality drop.