All sections

The whole curriculum, grouped into 14 areas. Start at the top if you are new. If you came here for one thing, filter to its area or search — everything is on this page or one click from it.

14 areas 117 sections 1,162 of 1,167 lessons published 10 new lessons every day

Showing Computer Vision — 13 sections, 143 lessons. Show all 14 areas

Computer Vision

Teaching a computer to see: classification, detection, segmentation, faces, video, documents and 3D.

13 sections 143 of 143 lessons published

Computer Vision

Teaching computers to see and understand images and video.

8 published

  1. What is computer vision?
  2. How images are stored in a computer
  3. OpenCV
  4. Image classification
  5. Object detection
  6. YOLO
  7. Image segmentation
  8. Vision transformers

CNN Backbones and Pretraining

The famous image networks, taken one at a time, plus the modern trick of training them before you have a single label.

13 published

  1. Padding, stride and output shapes
  2. Receptive fields
  3. Residual and skip connections
  4. Inception and GoogLeNet
  5. ResNet
  6. The MobileNet family
  7. EfficientNet and compound scaling
  8. ConvNeXt
  9. Swin transformer
  10. Contrastive learning for images
  11. Masked image modelling
  12. DINO and self-distillation
  13. Choosing a backbone

Training Vision Models

How to take a model somebody else trained and make it work on your few thousand photos.

12 published

  1. Fine-tuning a vision model
  2. Freezing and unfreezing layers
  3. Replacing the classifier head
  4. Linear probing vs fine-tuning
  5. Where ImageNet pretraining stops helping
  6. Few-shot image classification
  7. Domain adaptation for vision
  8. Geometric augmentations
  9. Mixup and CutMix
  10. RandAugment and AutoAugment
  11. Building an Albumentations pipeline
  12. Test-time augmentation

Object Detection in Depth

Every moving part of a detector, from box formats to transformers that skip anchors entirely.

13 published

  1. Bounding box formats
  2. Intersection over union
  3. Anchor boxes
  4. Non-maximum suppression
  5. Feature pyramid networks
  6. RoI pooling and RoI align
  7. R-CNN, Fast R-CNN and Faster R-CNN
  8. YOLO versions compared
  9. Focal loss and RetinaNet
  10. Anchor-free detection
  11. DETR and set prediction
  12. Open-vocabulary detection
  13. Detecting very small objects

Segmentation in Depth

Labelling every single pixel: the three kinds of segmentation, the architectures behind each, and the messy masks they return.

11 published

  1. Semantic, instance and panoptic segmentation
  2. U-Net
  3. DeepLab and atrous convolution
  4. Mask R-CNN
  5. Panoptic segmentation
  6. Mask2Former and universal segmentation
  7. Segment Anything
  8. Dice and other segmentation losses
  9. Weakly supervised segmentation
  10. Cleaning up predicted masks
  11. Image matting and background removal

Faces, People and Pose

Finding people in pictures and describing them, plus the honest limits of doing that to human faces.

12 published

  1. Face detection models
  2. Face alignment and landmarks
  3. Face embeddings and ArcFace
  4. Verification vs identification
  5. Face anti-spoofing and liveness
  6. Deepfake detection
  7. Bias and consent in face recognition
  8. Anonymising faces in a dataset
  9. Heatmaps for keypoints
  10. Top-down vs bottom-up pose estimation
  11. MediaPipe face, hand and body landmarks
  12. 3D human pose and SMPL

Video Understanding and Tracking

What changes when the pictures start moving, and how to keep the same object labelled across a thousand frames.

12 published

  1. Reading and writing video frames
  2. Frame sampling strategies
  3. Optical flow
  4. 3D convolutions for video
  5. SlowFast networks
  6. Video transformers
  7. Action recognition in practice
  8. Video object segmentation
  9. Video anomaly detection
  10. Tracking by detection
  11. SORT, DeepSORT and ByteTrack
  12. Re-identification embeddings

OCR and Document Vision

Turning pictures of text back into text, and pictures of forms back into fields you can query.

11 published

  1. How an OCR system is put together
  2. Preprocessing scans for OCR
  3. Text detection models
  4. CRNN text recognition
  5. CTC loss
  6. Tesseract in practice
  7. TrOCR and transformer OCR
  8. Handwriting recognition
  9. Document layout analysis
  10. Table extraction from documents
  11. LayoutLM and form understanding

3D Vision and Depth

Getting shape and distance back out of flat pictures, from a chessboard calibration to a photorealistic 3D scene.

11 published

  1. Camera models and intrinsics
  2. Calibrating a camera with a chessboard
  3. Homographies and perspective warp
  4. Feature matching with SIFT and ORB
  5. RANSAC for robust fitting
  6. Stereo depth from two cameras
  7. Monocular depth estimation
  8. Structure from motion
  9. Visual SLAM
  10. Neural radiance fields
  11. Gaussian splatting

Image Generation and Restoration

The machinery inside image generators, and the closely related job of repairing photos that already exist.

11 published

  1. Why GANs are hard to train
  2. StyleGAN and latent space
  3. Image-to-image translation
  4. The denoising U-Net
  5. Latent diffusion
  6. Classifier-free guidance
  7. Diffusion samplers and step counts
  8. ControlNet
  9. Inpainting and outpainting
  10. Super-resolution
  11. Watermarking and provenance

Vision Datasets and Annotation

The unglamorous work that decides your accuracy: labelling images well, and finding the errors already in your data.

11 published

  1. Choosing an annotation tool
  2. Writing annotation guidelines
  3. The COCO dataset format
  4. Converting between dataset formats
  5. Building an image dataset from the web
  6. Finding label errors in an image dataset
  7. Active learning for labelling
  8. Model-assisted labelling
  9. Deduplicating an image dataset
  10. Data leakage in image splits
  11. Dataset bias and shortcut learning

Evaluating Vision Models

The numbers that tell you whether a vision model is any good, and how to read them without fooling yourself.

9 published

  1. Mean average precision
  2. Detection error analysis
  3. Segmentation metrics
  4. Keypoint metrics: OKS and PCK
  5. Tracking metrics: MOTA and HOTA
  6. Measuring OCR accuracy
  7. FID and generative image metrics
  8. Testing robustness to corruptions
  9. Grad-CAM and saliency maps

Shipping Vision Models

Getting a model off your laptop and onto a camera feed, without the accuracy quietly falling apart on the way.

9 published

  1. Train and serve preprocessing mismatch
  2. Image formats and compression artefacts
  3. TensorRT for vision models
  4. INT8 calibration for vision models
  5. Preprocessing on the GPU
  6. OpenVINO and CPU inference
  7. RTSP and live camera pipelines
  8. Vision on Jetson and DeepStream
  9. Detecting drift in image data

Other ways to use this site

Learning paths

An ordered route through the lessons for one job: ML developer, AI engineer, computer vision engineer.

Projects

Build real things with full code, a dataset and a deployment guide.

AI glossary

Every term you keep seeing, defined in one plain sentence first.

Error database

Paste the error you got. Find out what it means and how to fix it.