FitWhen to use it — and when not
Use it when
- Any assistant that works with files, pages, repos or threads
- Products with large or paid context windows
- Copilots embedded next to a document or codebase
Skip it when
- Pure chit-chat with no external data
- When context is fully automatic and the user never needs to steer it — still show it on demand
AnatomyThe parts of the pattern
- ChipsOne chip per source with type, name and size; removable.
- Add menuFiles, the current page, folders, threads — with search.
- Context meterHow much of the window is used; warns before the limit.
- Over-limit stateWhat gets cut and how to fix it (remove or summarise).
GuidelinesDo & don’t
Do
- Show the current page or selection as an automatic, removable chip.
- Translate tokens into something human: pages, % of the window.
- Warn before truncation, not after a worse answer.
Don’t
- Silently drop the end of a long document.
- Show raw token counts as the only signal.
- Make users re-attach the same files every message.
In the wildReal-world examples
Cursor (@-mentions)GitHub Copilot ChatClaude ProjectsChatGPT attachments
Products named for reference only — no affiliation, and the demo above is an original illustration, not a copy of their UI.
For engineersImplementation notes
- Count tokens with the provider's tokenizer on the client or a cheap endpoint; estimates drift for code and non-English text.
- Represent context as a list of typed items (file, url, selection) so you can dedupe, cache embeddings and show provenance later.
- When over budget, prefer retrieval or summarisation over truncation, and tell the user which one happened.