Streaming UX that actually feels fast
Priya Raghunathan · 14 July 2026 · 6 min read
Total latency is the number on your dashboard. Time to first visible change is the number your users experience, and the two are rarely improved by the same work.
Show the pipeline, not a spinner
Stream events rather than tokens alone, and surface each stage as it completes: searching, found six sources, drafting. The wait is the same length and feels half as long.
Render sources before the answer
Retrieval finishes long before generation does. Painting the citation list first gives the reader something useful to do during the slowest part of the request.
Budget your preamble
Every reranking pass and guard rail call delays the first token. Measure each one against the perceived improvement it buys, and move whatever can be deferred to after the stream opens.