Skip to main content
Learn AI Create the Future — with Pranay
  • Learn
  • Paths
  • Projects
  • Practice
  • Glossary
  • Errors
  • LeetCode
  • Main site

Main site →
Select language

Translation is unavailable right now. Your browser’s own “Translate page” usually works.

Machine translation by Google

Topics

  • 🧱 Foundations 2 sections
  • 🤖 Machine Learning 11 sections
  • 📊 Statistics and Experiments 4 sections
  • 🧭 Doing the Work 9 sections
  • 🧠 Deep Learning 11 sections
  • 🧰 Libraries and Frameworks 4 sections
    • scikit-learn Properly 7 lessons
    • The HuggingFace Stack 13 lessons
      • Overview
      • Which AutoModel class to use
      • Fast tokenizers, offsets and word_ids
      • Padding, truncation and attention masks
      • Controlling generate()
      • Loading models bigger than your GPU
      • The Hub cache and pushing your own model
      • The Datasets library
      • map, batched and the cache
      • Data collators
      • The Trainer API
      • Accelerate
      • Loading a model in 4-bit
      • QLoRA: fine-tuning a large model on one GPU
    • TensorFlow and Keras 10 lessons
    • JAX and Flax 10 lessons
  • 💬 Language and NLP 15 sections
  • 👁️ Computer Vision 13 sections
  • 🎙️ Speech, Forecasting and Recommenders 3 sections
  • ✨ Generative AI and LLMs 4 sections
  • 🔬 Inside a Transformer 11 sections
  • ⚙️ Production and MLOps 16 sections
  • 🌍 AI in the Real World 12 sections
  • ⚖️ Safety, Ethics and Law 2 sections

See all topics

  1. Home
  2. Learn
  3. The HuggingFace Stack

🧰 Libraries and Frameworks · Section 038

🤗 The HuggingFace Stack

Beyond pipeline(): the model classes, tokeniser details, datasets and trainers you need to build something real.

Every lesson in this section is written by Pranay Mahendrakar.

13 of 13 lessons published · Three reading levels on every lesson

Start with “Which AutoModel class to use”

Lessons in order

Work top to bottom. Each lesson assumes the one above it.

  1. 01 Which AutoModel class to use
  2. 02 Fast tokenizers, offsets and word_ids
  3. 03 Padding, truncation and attention masks
  4. 04 Controlling generate()
  5. 05 Loading models bigger than your GPU
  6. 06 The Hub cache and pushing your own model
  7. 07 The Datasets library
  8. 08 map, batched and the cache
  9. 09 Data collators
  10. 10 The Trainer API
  11. 11 Accelerate
  12. 12 Loading a model in 4-bit
  13. 13 QLoRA: fine-tuning a large model on one GPU
Previous scikit-learn Properly Next TensorFlow and Keras

Learn AI

Learn Artificial Intelligence the easy way.

Learn. Build. Ask. Share.

Free forever. No sign-up, no ads, no tracking.

Learn

  • All topics
  • Learning paths
  • Projects
  • Practice playground
  • AI glossary
  • Error database
  • LeetCode solutions

Start here

  • Python for AI
  • Mathematics for AI
  • Machine Learning
  • Classic Algorithms in Depth
  • Linear Models and Regularisation
  • Ensembles and Gradient Boosting

This site

  • Search
  • New lessons feed
  • Sitemap
  • About Pranay Mahendrakar
  • pranaymahendrakar.com

© 2026 Pranay Mahendrakar. Written for people who are starting from zero.

10 new lessons every day. If something here is wrong or confusing, that is worth fixing — say so.