Skip to main content
Learn AI Create the Future — with Pranay
  • Learn
  • Paths
  • Projects
  • Practice
  • Glossary
  • Errors
  • LeetCode
  • Main site

Main site →
Select language

Translation is unavailable right now. Your browser’s own “Translate page” usually works.

Machine translation by Google

Topics

  • 🧱 Foundations 2 sections
  • 🤖 Machine Learning 11 sections
  • 📊 Statistics and Experiments 4 sections
  • 🧭 Doing the Work 9 sections
  • 🧠 Deep Learning 11 sections
  • 🧰 Libraries and Frameworks 4 sections
  • 💬 Language and NLP 15 sections
  • 👁️ Computer Vision 13 sections
    • Computer Vision 8 lessons
    • CNN Backbones and Pretraining 13 lessons
    • Training Vision Models 12 lessons
    • Object Detection in Depth 13 lessons
      • Overview
      • Bounding box formats
      • Intersection over union
      • Anchor boxes
      • Non-maximum suppression
      • Feature pyramid networks
      • RoI pooling and RoI align
      • R-CNN, Fast R-CNN and Faster R-CNN
      • YOLO versions compared
      • Focal loss and RetinaNet
      • Anchor-free detection
      • DETR and set prediction
      • Open-vocabulary detection
      • Detecting very small objects
    • Segmentation in Depth 11 lessons
    • Faces, People and Pose 12 lessons
    • Video Understanding and Tracking 12 lessons
    • OCR and Document Vision 11 lessons
    • 3D Vision and Depth 11 lessons
    • Image Generation and Restoration 11 lessons
    • Vision Datasets and Annotation 11 lessons
    • Evaluating Vision Models 9 lessons
    • Shipping Vision Models 9 lessons
  • 🎙️ Speech, Forecasting and Recommenders 3 sections
  • ✨ Generative AI and LLMs 4 sections
  • 🔬 Inside a Transformer 11 sections
  • ⚙️ Production and MLOps 16 sections
  • 🌍 AI in the Real World 12 sections
  • ⚖️ Safety, Ethics and Law 2 sections

See all topics

  1. Home
  2. Learn
  3. Object Detection in Depth

👁️ Computer Vision · Section 059

🎯 Object Detection in Depth

Every moving part of a detector, from box formats to transformers that skip anchors entirely.

Every lesson in this section is written by Pranay Mahendrakar.

13 of 13 lessons published · Three reading levels on every lesson

Start with “Bounding box formats”

Lessons in order

Work top to bottom. Each lesson assumes the one above it.

  1. 01 Bounding box formats
  2. 02 Intersection over union
  3. 03 Anchor boxes
  4. 04 Non-maximum suppression
  5. 05 Feature pyramid networks
  6. 06 RoI pooling and RoI align
  7. 07 R-CNN, Fast R-CNN and Faster R-CNN
  8. 08 YOLO versions compared
  9. 09 Focal loss and RetinaNet
  10. 10 Anchor-free detection
  11. 11 DETR and set prediction
  12. 12 Open-vocabulary detection
  13. 13 Detecting very small objects
Previous Training Vision Models Next Segmentation in Depth

Learn AI

Learn Artificial Intelligence the easy way.

Learn. Build. Ask. Share.

Free forever. No sign-up, no ads, no tracking.

Learn

  • All topics
  • Learning paths
  • Projects
  • Practice playground
  • AI glossary
  • Error database
  • LeetCode solutions

Start here

  • Python for AI
  • Mathematics for AI
  • Machine Learning
  • Classic Algorithms in Depth
  • Linear Models and Regularisation
  • Ensembles and Gradient Boosting

This site

  • Search
  • New lessons feed
  • Sitemap
  • About Pranay Mahendrakar
  • pranaymahendrakar.com

© 2026 Pranay Mahendrakar. Written for people who are starting from zero.

10 new lessons every day. If something here is wrong or confusing, that is worth fixing — say so.