All sections

The whole curriculum, grouped into 14 areas. Start at the top if you are new. If you came here for one thing, filter to its area or search — everything is on this page or one click from it.

14 areas 117 sections 1,162 of 1,167 lessons published 10 new lessons every day

Showing Production and MLOps — 16 sections, 157 lessons. Show all 14 areas

Production and MLOps

Getting a model off your laptop and keeping it alive: data pipelines, serving, scaling, monitoring, cost and on-device.

16 sections 157 of 157 lessons published

MLOps

Getting models out of your laptop and into the real world, reliably.

7 published

  1. What is MLOps?
  2. Experiment tracking
  3. MLflow
  4. Docker for ML
  5. Model serving
  6. Monitoring and model drift
  7. CI/CD for machine learning

Data Engineering for AI

The unglamorous work that decides whether a model succeeds — getting good data and keeping it good.

13 published

  1. Why data quality decides everything
  2. Collecting data
  3. Labelling data
  4. Cleaning data
  5. Handling missing data
  6. Imbalanced data
  7. Data augmentation
  8. Synthetic data
  9. Data versioning
  10. Feature stores
  11. ETL for machine learning
  12. Streaming data
  13. Building a data pipeline

Feature and Data Pipelines in Production

Keeping the numbers a live model reads fresh, correct and fast — the part that breaks most often.

10 published

  1. Point-in-time correctness
  2. Feature freshness
  3. Backfilling a new feature
  4. Streaming feature aggregation
  5. Low-latency feature lookup
  6. Missing features at serve time
  7. Versioning and retiring features
  8. Re-embedding and reindexing
  9. Rebuilding an index with no downtime
  10. Document ingestion pipelines for RAG

Serving Models in Production

The server that sits between your trained model and a real user — how it loads, answers and stays up.

10 published

  1. Serving a model with FastAPI
  2. NVIDIA Triton inference server
  3. BentoML
  4. Ray Serve
  5. Text Generation Inference (TGI)
  6. gRPC vs REST for inference
  7. Warm-up and cold starts
  8. Serving many models on one machine
  9. Streaming token responses
  10. Health checks and readiness probes

Batching and Concurrency

How to serve many users at once on the same hardware, instead of one at a time.

6 published

  1. Dynamic batching
  2. Continuous batching for LLMs
  3. Concurrency in a Python inference server
  4. Admission control and load shedding
  5. Deadlines, timeouts and retry storms
  6. Offline batch scoring

Latency, Load Testing and Capacity

Measuring how fast your service really is, and working out how many machines you need before the traffic arrives.

10 published

  1. Latency budgets
  2. Tail latency and percentiles
  3. Load testing a model endpoint
  4. Coordinated omission
  5. Profiling inference code
  6. Tuning CPU inference
  7. Time to first token vs tokens per second
  8. Capacity planning with Little's Law
  9. SLOs and error budgets for model services
  10. Soak testing and slow memory leaks

Caching and Cost Control

The cheapest inference is the one you never run. How to reuse answers and keep the bill small.

13 published

  1. Estimating how long and how much a run will cost
  2. Caching model predictions
  3. Semantic caching
  4. Prompt prefix caching
  5. Caching embeddings
  6. Cache stampedes and single-flight
  7. Precomputing predictions
  8. Cache invalidation for RAG
  9. Unit economics of an AI feature
  10. Cutting LLM API costs
  11. Model cascades
  12. Self-hosting vs API: the break-even point
  13. The one-box production stack

GPUs: Memory, Scheduling and Cost

The most expensive machine in your stack — how to size it, share it, and stop wasting it.

10 published

  1. Tensor parallel serving
  2. Choosing a GPU for inference
  3. Reading GPU utilisation honestly
  4. Sharing one GPU between models
  5. Scheduling GPUs on Kubernetes
  6. Running on spot and preemptible GPUs
  7. CUDA, drivers and container images
  8. Serverless GPU platforms
  9. GPU cost per request
  10. When a CPU is enough

Scaling and Traffic Management

Adding and removing machines as traffic moves, and sending each request to the right one.

10 published

  1. Autoscaling on the right metric
  2. Autoscaling lag and headroom
  3. Scale to zero, and what it costs you
  4. Load balancing inference traffic
  5. Prefix-aware and sticky routing
  6. Rate limiting and per-user quotas
  7. Multi-tenancy and noisy neighbours
  8. Graceful shutdown and connection draining
  9. Circuit breakers and fallbacks
  10. Serving from more than one region

Registries, Artifacts and Environments

Knowing exactly which model is running, where it came from, and how to build it again next year.

6 published

  1. Model registries
  2. What goes inside a model artifact
  3. Model lineage and traceability
  4. Pinning ML dependencies
  5. Shrinking ML container images
  6. Training and serving environment parity

Testing ML Code and CI

Automatic checks that catch a broken model before your users do.

10 published

  1. Unit testing machine learning code
  2. Testing a training loop
  3. Behavioural tests for models
  4. Golden outputs and regression tests
  5. Testing output you cannot predict
  6. Recording and replaying model API calls
  7. Integration testing an inference server
  8. Evaluation gates in CI
  9. GitHub Actions for ML projects
  10. Automated retraining pipelines

Releasing Models Safely

Putting a new model in front of users a little at a time, and taking it back quickly when it is wrong.

10 published

  1. Shadow deployment
  2. Canary releases for models
  3. Blue-green model deployments
  4. Rolling back a model
  5. Feature flags for model rollouts
  6. Champion and challenger models
  7. Swapping weights with no downtime
  8. Versioning a model API
  9. KServe and Seldon
  10. Switching between model providers

Monitoring Models in Production

Your service is up and answering. This is how you find out whether the answers are still any good.

11 published

  1. Measuring data drift
  2. Concept drift
  3. Training-serving skew
  4. Working with delayed labels
  5. Monitoring when you have no labels
  6. Logging every prediction
  7. Monitoring by segment
  8. Alerts people do not ignore
  9. When model metrics and business metrics disagree
  10. Monitoring a vector index
  11. Feedback loops in production

Observability for LLM Applications

Seeing inside a chain of prompts, retrievals and tool calls when a user says the answer was wrong.

9 published

  1. Tracing an LLM application
  2. Token and cost telemetry
  3. Versioning prompts
  4. Running evaluations on live traffic
  5. LLM-as-judge in production
  6. Capturing user feedback
  7. Redacting personal data from LLM logs
  8. Replaying production traffic
  9. Observability for agent runs

Incident Response for ML Systems

What to do at 2am when the model is wrong, slow or quietly broken — and how to be ready for it.

9 published

  1. On-call for ML systems
  2. Runbooks for model failures
  3. Triaging an ML incident
  4. Silent failures
  5. Debugging a latency spike
  6. When upstream data breaks
  7. Surviving a provider outage
  8. Postmortems for ML incidents
  9. Chaos testing an ML service

Edge and On-device AI

Running models on phones, browsers and tiny boards, where there is no GPU and no cloud.

13 published

  1. What is edge AI?
  2. Why run a model on the device?
  3. Model compression
  4. Quantisation in practice
  5. Pruning
  6. Knowledge distillation
  7. ONNX
  8. TensorFlow Lite
  9. Core ML
  10. Running models in the browser
  11. AI on a Raspberry Pi
  12. Latency and throughput
  13. Power and thermal limits

Other ways to use this site

Learning paths

An ordered route through the lessons for one job: ML developer, AI engineer, computer vision engineer.

Projects

Build real things with full code, a dataset and a deployment guide.

AI glossary

Every term you keep seeing, defined in one plain sentence first.

Error database

Paste the error you got. Find out what it means and how to fix it.