FitWhen to use it — and when not
Use it when
- Every AI answer in a product you intend to improve
- Early launches where you are still learning failure modes
- Building eval datasets from real usage
Skip it when
- When nobody reads the feedback — fix the pipeline first
AnatomyThe parts of the pattern
- ThumbsLow-effort signal on every answer.
- Reason chipsAfter thumbs-down: incorrect, too long, ignored instructions, unsafe.
- CommentOptional free text.
- ConsentExplicit choice to share the conversation.
- Follow-throughOffer a retry that uses the feedback.
GuidelinesDo & don’t
Do
- Ask for reasons only after a negative signal.
- Keep reasons specific to your product's failure modes.
- Turn feedback into an immediate improvement (retry with it).
Don’t
- Pop a survey after every answer.
- Collect conversations without saying so.
- Promise "we'll use this to improve" if nobody looks.
In the wildReal-world examples
ChatGPTClaudeGeminiGitHub Copilot
Products named for reference only — no affiliation, and the demo above is an original illustration, not a copy of their UI.
For engineersImplementation notes
- Store feedback with the full request context (prompt version, model, retrieved chunk ids) — without it the signal is unusable.
- Pipe thumbs-down with reasons into an eval queue for triage and labelling.
- Track feedback rate, not just ratio — a falling rate means the UI is hidden.