FitWhen to use it — and when not
Use it when
- Forecasts, estimates, extracted facts, classifications
- High-stakes decisions where the user must decide how much to check
- Mixed answers where some parts are solid and some are not
Skip it when
- When you have no real signal — a made-up score is worse than none
- Casual, low-stakes chat
AnatomyThe parts of the pattern
- Overall labelPlain words: high, medium, low — with why.
- Claim markersWeak claims visibly marked inline.
- ReasonWhy it is unsure: one source, old data, conflicting sources.
- Next stepVerify, ask a follow-up, or route to a person.
GuidelinesDo & don’t
Do
- Explain the reason, not just the level.
- Mark the weak part, not the whole answer.
- Pair uncertainty with an action: verify, check source, ask an expert.
Don’t
- Show "87.3% confident" from an uncalibrated model.
- Bury uncertainty in a footnote nobody reads.
- Make everything amber — if everything is flagged, nothing is.
In the wildReal-world examples
Consensus (agreement meter)Elicit (supporting quotes)GitHub Copilot code review
Products named for reference only — no affiliation, and the demo above is an original illustration, not a copy of their UI.
For engineersImplementation notes
- Derive confidence from signals you control: retrieval scores, source agreement, self-consistency across samples, validator checks.
- Calibrate thresholds on a labelled set before you show any level to users.
- Return claims as structured spans with a level and reason so the UI can mark them inline.