Language and NLP Section 045

The BERT Family

The encoder models that read a whole sentence at once — still the cheapest way to classify, tag and search text.

10 of 10 lessons published Three reading levels on every lesson

Start with “Masked language modelling”

Lessons in order

Work top to bottom. Each lesson assumes the one above it.

  1. Masked language modelling
  2. RoBERTa
  3. DeBERTa
  4. ModernBERT
  5. Encoder, decoder and encoder-decoder models
  6. T5 and the text-to-text idea
  7. Fine-tuning BERT for classification
  8. Continued pretraining on your own domain
  9. The 512-token wall
  10. Longformer and long-context encoders