All sections

The whole curriculum, grouped into 14 areas. Start at the top if you are new. If you came here for one thing, filter to its area or search — everything is on this page or one click from it.

14 areas 117 sections 1,162 of 1,167 lessons published 10 new lessons every day

Showing Language and NLP — 15 sections, 140 lessons. Show all 14 areas

Language and NLP

Everything to do with text — cleaning it, embedding it, searching it, and answering questions from it.

15 sections 140 of 140 lessons published

Natural Language Processing

How computers read, understand and write human language.

8 published

  1. What is NLP?
  2. Tokenization
  3. Embeddings
  4. Word2Vec
  5. Attention
  6. BERT
  7. Text classification
  8. Named entity recognition

Text Preprocessing

Getting raw text into a shape a model can use — the quiet step that decides how well everything after it works.

9 published

  1. Text normalisation: case, punctuation and whitespace
  2. Unicode, UTF-8 and mojibake
  3. Stopwords: when removing them helps and when it breaks things
  4. Stemming vs lemmatisation
  5. Splitting text into sentences
  6. Regular expressions for text work
  7. spaCy pipelines
  8. Getting clean text out of PDFs
  9. Detecting the language of a text

Classical NLP That Still Works

The methods from before deep learning that still win on small data, small budgets and tight deadlines.

7 published

  1. Bag of words
  2. TF-IDF
  3. N-gram language models and smoothing
  4. The baseline you must beat before reaching for BERT
  5. BM25
  6. Matching thousands of keywords at once
  7. Keyword and keyphrase extraction

The Embedding Family

The models that turn a word, a sentence or a whole document into a list of numbers you can compare.

10 published

  1. GloVe
  2. FastText and subword vectors
  3. King minus man plus woman, and why it works less often than you think
  4. Static vs contextual embeddings
  5. Sentence-transformers
  6. CLS, mean and max pooling
  7. Fine-tuning an embedding model on your own data
  8. Hard negative mining
  9. Embedding size and Matryoshka truncation
  10. Storing vectors as int8 and binary

The BERT Family

The encoder models that read a whole sentence at once — still the cheapest way to classify, tag and search text.

10 published

  1. Masked language modelling
  2. RoBERTa
  3. DeBERTa
  4. ModernBERT
  5. Encoder, decoder and encoder-decoder models
  6. T5 and the text-to-text idea
  7. Fine-tuning BERT for classification
  8. Continued pretraining on your own domain
  9. The 512-token wall
  10. Longformer and long-context encoders

Sequence Labelling and Structure

Putting a label on every single word instead of the whole sentence — the machinery behind entity extraction.

10 published

  1. Part-of-speech tagging
  2. BIO and BILOU tagging schemes
  3. Conditional random fields
  4. Training NER on your own entities
  5. Nested and overlapping entities
  6. Why token accuracy lies for NER
  7. Entity linking
  8. Coreference resolution
  9. Relation extraction
  10. Dependency parsing

Topic Modelling and Text Clustering

Finding the themes hiding inside a pile of documents nobody has time to read.

9 published

  1. What is topic modelling?
  2. Latent Dirichlet Allocation
  3. NMF for topics
  4. BERTopic
  5. Choosing how many topics
  6. Naming your topics automatically
  7. Topic modelling on tweets and short text
  8. Clustering documents with embeddings
  9. Tracking topics over time

Chunking and Long Documents

How to cut a 200-page document into pieces without cutting the answer in half.

8 published

  1. Chunking strategies compared
  2. Semantic chunking
  3. Chunking by document structure
  4. Chunking source code and tables
  5. Giving each chunk its context back
  6. Retrieve small, return big
  7. Keeping the trail back to the source
  8. Long context or retrieval?

Question Answering

Systems that hand back the answer instead of a list of links — and the many ways they get it wrong.

10 published

  1. Extractive question answering
  2. Teaching a model to say 'not in the document'
  3. Multi-hop questions
  4. Answering questions about tables
  5. Handling 'what about last year?'
  6. Matching a user question to an existing FAQ
  7. Checking that an answer really came from the source
  8. When two documents disagree
  9. Diagnosing a wrong answer
  10. What SQuAD and Natural Questions actually test

Summarisation

Making long text short, either by picking the best sentences or by writing brand-new ones.

10 published

  1. Extractive summarisation
  2. Abstractive summarisation
  3. Summarising a document longer than the context window
  4. Query-focused summarisation
  5. Summarising many documents at once
  6. Summarising meetings and conversations
  7. Fine-tuning your own summariser
  8. Controlling summary length
  9. When a summary says something the source never said
  10. Prompt an LLM or fine-tune a small model?

Sentiment, Opinion and Text Mining

Reading what people actually feel, at a finer grain than thumbs up or thumbs down.

9 published

  1. Sentiment beyond positive and negative
  2. Aspect-based sentiment analysis
  3. Emotion classification
  4. Sarcasm and irony detection
  5. Stance detection
  6. Toxicity and abuse detection
  7. Classifying into a 300-label taxonomy
  8. Zero-shot text classification
  9. Few-shot text classification with SetFit

Multilingual and Indic NLP

Making NLP work in Hindi, Marathi, Tamil and Hinglish — not only in English.

12 published

  1. How neural machine translation works
  2. Running translation models yourself
  3. Translating a low-resource language
  4. XLM-RoBERTa and multilingual encoders
  5. Train in English, run in Hindi
  6. Multilingual sentence embeddings
  7. Question in Hindi, documents in English
  8. Hinglish and code-mixed text
  9. Transliteration between scripts
  10. Normalising Indic scripts
  11. The AI4Bharat and IndicNLP stack
  12. Why Hindi costs three times more tokens than English

Messy Real-World Text

Real user text is misspelled, half in Hindi, full of emoji and pasted out of a scanner. Here is how to survive it.

10 published

  1. Building a spell checker
  2. Fuzzy matching names and addresses
  3. Normalising numbers, dates and currency
  4. Cleaning social media text
  5. Fixing OCR errors
  6. Punctuating a speech transcript
  7. Finding and masking personal data in text
  8. Augmenting text data
  9. When users write 'f r e e m0ney' to dodge your filter
  10. Labelling data with an LLM and a human reviewer

Evaluating Text Systems

How to tell whether your text system got better, when there is no single right answer to compare against.

7 published

  1. BLEU, chrF and COMET
  2. ROUGE and what it misses
  3. BERTScore and embedding-based metrics
  4. Exact match and F1 for question answering
  5. Natural language inference
  6. Using an LLM to grade text
  7. Do your labels even agree?

Other ways to use this site

Learning paths

An ordered route through the lessons for one job: ML developer, AI engineer, computer vision engineer.

Projects

Build real things with full code, a dataset and a deployment guide.

AI glossary

Every term you keep seeing, defined in one plain sentence first.

Error database

Paste the error you got. Find out what it means and how to fix it.