Track · 2:30 · Liner note

Waiting for the First Token

Time to first token is the delay between sending a prompt and seeing the first piece of a model's streamed answer, and it shapes how responsive an AI product feels.

Between pressing Enter and seeing the first word there is a gap. That gap is time to first token: the model has to read the prompt before it can begin streaming its answer. Waiting for the First Token is that short suspense, and it decides whether an AI product feels alive or stuck.

Streaming helps because people read as the answer arrives. Shorter prompts, cached context and smaller models shave the wait, but the trade against quality is real. Measure first-token latency separately from total time, and show progress while the model thinks.