FitWhen to use it — and when not
Use it when
- Any team changing prompts, models or retrieval regularly
- Selling to technical buyers who ask "how do you know it works?"
- Regulated or high-stakes AI
Skip it when
- Before you have a test set — build 50 real cases first
AnatomyThe parts of the pattern
- Score over versionsOne weighted score per prompt/model version.
- ThresholdThe line a release must stay above.
- Regression alertWhich slice dropped, by how much.
- Failing casesReal examples, before vs after.
- GateBlock or ship, with a record.
GuidelinesDo & don’t
Do
- Break scores down by slice (topic, language, customer).
- Show real failing examples next to the number.
- Make the release gate explicit and logged.
Don’t
- Trust one aggregate number.
- Run evals only when someone remembers.
- Hide judge disagreement — it tells you the metric is shaky.
In productionHow it looks in a shipped product

In the wildReal-world examples
BraintrustLangSmithArize PhoenixOpenAI Evals
Products named for reference only — no affiliation, and the demo above is an original illustration, not a copy of their UI.
For engineersImplementation notes
- Run evals in CI on every prompt/model change; store results keyed by version.
- Mix deterministic checks, LLM-as-judge and human labels; track judge-human agreement.
- Version datasets too — a score change can be a dataset change.