Real inputs with graded expected outputs, scored on every model or prompt change — the core practice, wired into CI like any other gate.
Where exact output varies, invariants don't: valid JSON schema, citations present, no forbidden content, length bounds, required fields extracted — mechanical checks on non-deterministic prose.
Accuracy thresholds over the eval set (not per-case perfection), regression alarms on score drops, and human spot-audit sampling — quality as a measured distribution.
Deterministic components (retrieval, tools, gates, fallbacks) get normal tests; injection payloads and failure-mode drills cover the AI-specific attack and error surface.
Skipping the discipline this article describes until an incident, audit, or stalled project forces it — every practice above is cheaper adopted early than retrofitted under pressure.
Let's discuss how we can help you with testing ai features non deterministic.