Production and MLOps Section 093

GPUs: Memory, Scheduling and Cost

The most expensive machine in your stack — how to size it, share it, and stop wasting it.

10 of 10 lessons published Three reading levels on every lesson

Start with “Tensor parallel serving”

Lessons in order

Work top to bottom. Each lesson assumes the one above it.

  1. Tensor parallel serving
  2. Choosing a GPU for inference
  3. Reading GPU utilisation honestly
  4. Sharing one GPU between models
  5. Scheduling GPUs on Kubernetes
  6. Running on spot and preemptible GPUs
  7. CUDA, drivers and container images
  8. Serverless GPU platforms
  9. GPU cost per request
  10. When a CPU is enough