AI glossary

Latency

In one sentence Latency is how long one request takes from send to answer — the waiting a single user actually feels.

By Updated

Latency is the time between sending one request and receiving its response — the delay a single user experiences.

It is the answer to "how long until my food arrives?" — nothing else. Not how busy the kitchen is, not how many orders it clears per hour (that is throughput); the minutes between your order and your plate. The two measures famously part ways: a hostel mess feeding 500 people an hour can still make you personally wait forty minutes.

For ML serving, latency decomposes, and each piece is attacked differently: network travel, queueing (waiting for a batching slot or a free GPU), and compute itself. LLMs add a twist worth memorising — two numbers, not one. Time-to-first-token (dominated by processing your prompt) decides how long the screen stays blank; time-per-token decides how fast the answer types out. Streaming exists because the first number matters most to perception.

Measure it honestly: report percentiles, never averages. p50 is the typical case; p99 is what your unluckiest one-in-a-hundred user gets, and averages hide exactly the outliers that generate complaints.

p50 = 300 ms   p95 = 900 ms   p99 = 2.5 s   ← this line is the user experience

Standard levers, cheapest first: cache repeated work (prompt-caching), shrink the model (quantization, distillation), serve from hardware nearer the user, and tune the latency-throughput trade — bigger batches raise GPU efficiency and individual waiting. Real-time products (voice, autocomplete) buy latency at the cost of throughput; offline pipelines do the reverse.

Where to go next