Libraries and Frameworks Section 038
The HuggingFace Stack
Beyond pipeline(): the model classes, tokeniser details, datasets and trainers you need to build something real.
13 of 13 lessons published Three reading levels on every lesson
Lessons in order
Work top to bottom. Each lesson assumes the one above it.
- Which AutoModel class to use
- Fast tokenizers, offsets and word_ids
- Padding, truncation and attention masks
- Controlling generate()
- Loading models bigger than your GPU
- The Hub cache and pushing your own model
- The Datasets library
- map, batched and the cache
- Data collators
- The Trainer API
- Accelerate
- Loading a model in 4-bit
- QLoRA: fine-tuning a large model on one GPU