50–300 real cases (tickets, documents, queries) with agreed correct outputs, weighted toward the hard and costly ones. Synthetic-only sets flatter the system.
Exact-match for structured outputs; rubric scoring for prose; model-graded evaluation with human spot-audit for scale. Per-field accuracy for extraction, per-intent for conversation.
Model swaps, prompt edits, retrieval tweaks — regression-tested like code, because 'it seemed fine' is how quality regressions ship.
Launch thresholds you approve, dashboards after, drift alerts when accuracy moves — evaluation is a living system, not a report.
Skipping the discipline this article describes until an incident, audit, or stalled project forces it — every practice above is cheaper adopted early than retrofitted under pressure.
Let's discuss how we can help you with how to evaluate llm outputs.