Production and MLOps Section 091

Latency, Load Testing and Capacity

Measuring how fast your service really is, and working out how many machines you need before the traffic arrives.

10 of 10 lessons published Three reading levels on every lesson

Start with “Latency budgets”

Lessons in order

Work top to bottom. Each lesson assumes the one above it.

  1. Latency budgets
  2. Tail latency and percentiles
  3. Load testing a model endpoint
  4. Coordinated omission
  5. Profiling inference code
  6. Tuning CPU inference
  7. Time to first token vs tokens per second
  8. Capacity planning with Little's Law
  9. SLOs and error budgets for model services
  10. Soak testing and slow memory leaks