Latency
In one sentence Latency is how long one request takes from send to answer — the waiting a single user actually feels.
Updated
Latency is the time between sending one request and receiving its response — the delay a single user experiences.
It is the answer to "how long until my food arrives?" — nothing else. Not how busy the kitchen is, not how many orders it clears per hour (that is throughput); the minutes between your order and your plate. The two measures famously part ways: a hostel mess feeding 500 people an hour can still make you personally wait forty minutes.
For ML serving, latency decomposes, and each piece is attacked differently: network travel, queueing (waiting for a batching slot or a free GPU), and compute itself. LLMs add a twist worth memorising — two numbers, not one. Time-to-first-token (dominated by processing your prompt) decides how long the screen stays blank; time-per-token decides how fast the answer types out. Streaming exists because the first number matters most to perception.
Measure it honestly: report percentiles, never averages. p50 is the typical case; p99 is what your unluckiest one-in-a-hundred user gets, and averages hide exactly the outliers that generate complaints.
p50 = 300 ms p95 = 900 ms p99 = 2.5 s ← this line is the user experienceStandard levers, cheapest first: cache repeated work (prompt-caching), shrink the model (quantization, distillation), serve from hardware nearer the user, and tune the latency-throughput trade — bigger batches raise GPU efficiency and individual waiting. Real-time products (voice, autocomplete) buy latency at the cost of throughput; offline pipelines do the reverse.
Where to go next
- Full lesson: vLLM
- Related terms: throughput, streaming, inference, batching