← All postsEngineering

Streaming UX that actually feels fast

Priya Raghunathan · 14 July 2026 · 6 min read


Total latency is the number on your dashboard. Time to first visible change is the number your users experience, and the two are rarely improved by the same work.

Show the pipeline, not a spinner

Stream events rather than tokens alone, and surface each stage as it completes: searching, found six sources, drafting. The wait is the same length and feels half as long.

Render sources before the answer

Retrieval finishes long before generation does. Painting the citation list first gives the reader something useful to do during the slowest part of the request.

Budget your preamble

Every reranking pass and guard rail call delays the first token. Measure each one against the perceived improvement it buys, and move whatever can be deferred to after the stream opens.