Track · 2:30 · Liner note
Measure Twice, Answer Once
Evals are repeatable tests that score an AI system's answers against known good results, so changes to a prompt or model are judged by evidence rather than impression.
Carpenters measure twice because wood cannot be uncut. A wrong AI answer can be retracted, but only if someone noticed it was wrong. Measure Twice, Answer Once is the case for evals: a fixed set of questions with known good answers, scored every time the prompt, model or data changes, so improvement is a measurement and not a mood.
Start small with real questions from real users, include the awkward ones and score more than one dimension, such as correctness and refusal behavior. A regression caught in the eval is one fewer a user reports.