FitWhen to use it — and when not
Use it when
- Mobile-first and hands-busy contexts (field work, driving, clinics)
- Long, rambling requests that are easier said than typed
- Accessibility — people who cannot type comfortably
Skip it when
- Open-plan offices and public spaces as the only input
- Precise input like code, IDs or numbers without a review step
AnatomyThe parts of the pattern
- Mic buttonOne obvious control, with a clear on/off state.
- Listening stateLive level meter so people know they are being heard.
- Interim transcriptWords appear as spoken; corrections settle in place.
- ReviewEditable text before sending — or a clear auto-send setting.
GuidelinesDo & don’t
Do
- Show interim words immediately; latency is what makes voice feel broken.
- Let people edit the transcript before it goes to the model.
- Make the mic state unmistakable — colour, motion and text.
Don’t
- Send automatically on silence without a way to opt out.
- Hide what was heard — mis-hearings become wrong actions.
- Start listening without an explicit tap.
In the wildReal-world examples
ChatGPT voiceGemini LiveWispr FlowiOS dictation
Products named for reference only — no affiliation, and the demo above is an original illustration, not a copy of their UI.
For engineersImplementation notes
- Stream audio to a streaming STT endpoint and render interim vs final results differently.
- Keep a short client-side silence detector (VAD) for auto-stop, with a user-tunable timeout.
- Request microphone permission only on the first tap, with a one-line reason.