Natural Language Processing

BERT

BERT reads a sentence from both directions at once and gives every word a meaning that depends on the whole sentence, which is why one pretrained BERT can be adapted to dozens of tasks cheaply.

Read these first

On this page 10
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. The training game
  5. Then the cheap part
  6. The thing people get wrong
  7. Where you have already seen it
  8. What is honestly hard
  9. Remember this
  10. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

BERT is a model that reads a whole sentence in both directions before deciding what each word means.

The analogy you have already lived

You pick up a newspaper and one word is hidden under an ink smudge. "I went to the ▓▓▓▓ to deposit my salary."

You fill it in without effort. And notice how you did it. You did not stop at the smudge. You read past it, saw "deposit my salary", and only then decided the word was "bank".

Now try: "I sat on the ▓▓▓▓ of the river and watched the boats." Same smudge, same length, different answer. The words that decided it came after the gap both times.

BERT is a model built entirely around that move.

Why it exists

Before 2018, most language models read strictly left to right. Word by word, never peeking forward. That is a sensible rule when your job is to write the next word. The next word has not been written yet.

It is a bad rule when your job is to understand a sentence that already exists. Half the evidence sits to the right of the word you are looking at, and a left-to-right reader throws it away.

There was a second, more expensive problem. Every task needed its own model, trained from scratch, on its own hand-labelled data. Want to detect spam? Label fifty thousand emails. Want to sort support tickets? Label fifty thousand tickets again. Nothing carried over.

BERT changed both at once. Its name stands for Bidirectional Encoder Representations from Transformers. Bidirectional means it reads both ways. Encoder means it turns text into numbers instead of writing new text.

The training game

BERT learns by playing a fill-in-the-blank game on ordinary text with no labels at all.

Take any sentence from Wikipedia. Hide about one word in seven. Ask the model to guess each hidden word using everything around it, on both sides.

   input:   the chai was [HIDDEN] and sweet
                             |
   the model sees:  "the chai was ____ and sweet"
   the model guesses:  hot
   the real answer:    hot        -> small correction

Nobody labelled anything. The text itself provides the answer, because the answer is the word that was covered up. This is called masked language modelling — masking meaning covering a word so the model has to reconstruct it.

Run that game across billions of sentences and the model has to learn grammar, common sense and word meaning. There is no other way to win.

Then the cheap part

That expensive training happens once. Afterwards you adapt the finished model to your task, using a little of your own labelled data.

   [ BERT, trained once on billions of sentences ]
                       |
        +--------------+--------------+
        |              |              |
   spam or not?   which department?  is this a name?
   (2,000 labels) (5,000 labels)    (3,000 labels)

This is fine-tuning — taking a finished model and continuing its training briefly on your own examples. The heavy lifting was already paid for. You are renting it.

That split is the reason BERT mattered so much. It moved language work from "collect a huge labelled dataset" to "collect a small one".

The thing people get wrong

BERT cannot write text. Ask it to continue a story and it has no mechanism to do so.

It fills gaps in text that already exists. That is a different job from producing text one word at a time, which is what ChatGPT and similar tools do. Those are built the other way around — left to right, on purpose, so they can keep going.

So BERT is not an older, weaker ChatGPT. It is a different tool. It reads. The other writes.

If that distinction feels slippery, read the paragraph again. Almost everyone mixes up these two model families at first. Keeping them apart makes the rest of NLP much easier.

Where you have already seen it

  • Google Search. In 2019 Google announced BERT was being used to interpret search queries, affecting around one in ten English searches. Small words like "to" and "for" started changing results, because BERT reads them in context.
  • Support ticket routing. Your complaint reaches the billing team without a human reading it first.
  • Resume screening tools. Pulling names, companies and skills out of free text.
  • Content moderation. Deciding whether a comment is abusive, where the meaning depends on the whole sentence.

What is honestly hard

BERT reads at most 512 pieces of text at a time — roughly a long paragraph. Feed it a full document and it silently cuts off the rest. Working around this is real engineering, not a setting you flip.

It is also large. The standard version has 110 million numbers inside it, around 440 megabytes to download. That is fine on a server and painful on a phone or a metered mobile connection.

And fine-tuning still needs labelled examples. Fewer than before, but not zero. Anyone who tells you BERT removed the need for labelled data is selling something.

Remember this

  • BERT reads a sentence in both directions, so a word's meaning can depend on what comes after it.
  • It learns by filling in hidden words in ordinary text, which needs no labels.
  • It reads and classifies. It does not write. That is a different family of model.

What to learn next

Developer — Code and libraries.

Three short programs. The first shows what BERT's tokenizer does to your text. The second proves the bidirectional claim. The third does the thing you will actually ship: BERT as a feature extractor with a small classifier on top.

Setup

bash
pip install transformers torch scikit-learn

Everything below runs on CPU. No GPU is needed for inference on a handful of sentences.

Part 1 — what BERT sees

bert_tokens.py
from transformers import AutoTokenizer, logging

logging.set_verbosity_error()          # hide the model-loading chatter

tok = AutoTokenizer.from_pretrained("bert-base-uncased")
for s in ["Bengaluru traffic is unbelievable.", "I sat on the [MASK] of the river."]:
    print(s)
    print("  ", tok.tokenize(s))
    print("  with special tokens:", tok.convert_ids_to_tokens(tok(s)["input_ids"])[:6], "...")
    print()
Output
Bengaluru traffic is unbelievable.
   ['bengal', '##uru', 'traffic', 'is', 'unbelievable', '.']
  with special tokens: ['[CLS]', 'bengal', '##uru', 'traffic', 'is', 'unbelievable'] ...

I sat on the [MASK] of the river.
   ['i', 'sat', 'on', 'the', '[MASK]', 'of', 'the', 'river', '.']
  with special tokens: ['[CLS]', 'i', 'sat', 'on', 'the', '[MASK]'] ...

Three things to notice.

bengal and ##uru. BERT's vocabulary holds 30,522 pieces. "Bengaluru" is not one of them, so it gets split. The ## prefix means "this continues the previous piece". English city names usually survive whole; Indian ones frequently do not, which costs you tokens and a little accuracy. See tokenization for why.

[CLS] at the front. A slot with no meaning of its own, added to every input. After the model runs, the vector sitting in that slot is used as a summary of the whole sentence. It is where a classification head gets attached.

[MASK] survives as one piece. It is a real entry in the vocabulary, not text to be split up.

Part 2 — proving it reads forwards

Both sentences below hide the same word. In both, the deciding evidence sits after the gap. A left-to-right model has no access to it.

bert_fill.py
from transformers import pipeline, logging

logging.set_verbosity_error()
fill = pipeline("fill-mask", model="bert-base-uncased")

sentences = [
    "I went to the [MASK] to deposit my salary.",
    "I sat on the [MASK] of the river and watched the boats.",
]

for s in sentences:
    print(s)
    for r in fill(s, top_k=3):
        print(f"    {r['token_str']:>8}  {r['score']:.3f}")
    print()
Output
I went to the [MASK] to deposit my salary.
        bank  0.895
      office  0.020
       banks  0.007

I sat on the [MASK] of the river and watched the boats.
        bank  0.558
        edge  0.289
       banks  0.059

Same masked position. Same top word. Two completely different meanings, and both times the disambiguating words came later in the sentence. That is the entire argument for bidirectional encoding, in one output block.

Notice the second case is less confident, 0.558 against 0.895, with edge a serious rival. That is correct behaviour. "The edge of the river" is a real phrase. A model that reported 0.99 there would be miscalibrated.

Part 3 — what you will actually build

The usual job is not filling blanks. It is classification. Here BERT converts sentences into vectors, and an ordinary logistic regression makes the decision.

bert_classify.py
import torch
from transformers import AutoTokenizer, AutoModel, logging
from sklearn.linear_model import LogisticRegression

logging.set_verbosity_error()
tok  = AutoTokenizer.from_pretrained("bert-base-uncased")
bert = AutoModel.from_pretrained("bert-base-uncased").eval()

def embed(texts):
    batch = tok(texts, padding=True, truncation=True, return_tensors="pt")
    with torch.no_grad():                       # no training here, so no gradients needed
        out = bert(**batch).last_hidden_state   # one vector per token
    mask = batch["attention_mask"].unsqueeze(-1)
    return ((out * mask).sum(1) / mask.sum(1)).numpy()   # average the real tokens only

train = ["the delivery was fast and the packing was neat",
         "excellent product, works exactly as described",
         "arrived a week late and the box was crushed",
         "stopped working after two days, waste of money"]
labels = ["good", "good", "bad", "bad"]

clf = LogisticRegression(max_iter=1000).fit(embed(train), labels)

tests = ["the courier turned up on time and everything was intact",
         "it broke almost immediately"]
for text, guess in zip(tests, clf.predict(embed(tests))):
    print(f"{guess:4} <- {text}")
Output
good <- the courier turned up on time and everything was intact
bad  <- it broke almost immediately

Read the test sentences against the training sentences carefully. They share almost no words.

Training said "delivery", "fast", "packing". The test says "courier", "on time", "intact". Training said "stopped working", "waste of money". The test says "broke almost immediately".

A word-counting model has nothing to work with here. The what is NLP lesson built exactly such a model and showed where it breaks. This is the gap BERT closes, on four training examples.

Line by line, for the parts that trip people up

.eval() switches off dropout. Forget it and your embeddings change between runs for no visible reason. This is a genuinely common and confusing bug.

torch.no_grad() stops PyTorch from building a graph for backpropagation. You are not training BERT here, so that graph is pure wasted memory. On a long batch, omitting it is the difference between running and an out-of-memory error.

The masked average. padding=True pads short sentences with [PAD] tokens so the batch is rectangular. Averaging over those pad slots dilutes short sentences with meaningless vectors. Multiplying by attention_mask and dividing by its sum averages only real tokens.

Mean pooling, not [CLS]. The [CLS] vector of a pretrained, not fine-tuned BERT is a poor sentence summary. It was trained for next-sentence prediction, a task later shown to contribute little. Mean pooling is the better default when you are not fine-tuning. If sentence similarity is your goal, use a sentence-transformer model instead — see embeddings.

truncation=True. Without it, any input over 512 tokens raises an error. With it, everything past token 512 is discarded silently. Neither is what you want for long documents. You want a chunking strategy you chose on purpose.

Feature extraction or full fine-tuning?

Part 3 froze BERT and trained a small head. The alternative is fine-tuning, where BERT's own weights move too.

Frozen features + headFull fine-tuning
Labelled examples neededhundredsa few thousand
HardwareCPU is fineGPU strongly preferred
Timesecondsminutes to hours
Typical accuracygoodbetter, often by several points
Riskunderfits your domainoverfits small datasets

Start frozen. It gives you a number in five minutes. Fine-tune only when you can show that number is not good enough.

Common mistakes

Using BERT to generate text. It has no mechanism for it. Every token attends to every other token, so there is no "next token" to produce. If you need text out, you need a decoder model.

Judging BERT by its raw [CLS] similarity scores. Cosine similarity between two pretrained BERT sentence vectors is famously weak — worse than averaging older static vectors. Reimers and Gurevych documented this in 2019. Use a model trained for similarity.

Feeding long documents and ignoring the cut. truncation=True throws away everything past 512 tokens without a warning. Your accuracy on long documents will be quietly bad and the loss curve will look fine.

Forgetting that bert-base-uncased lowercases everything. It destroys case, which is a major signal for named entity recognition. Use a cased checkpoint for entity work.

Reaching for BERT before trying TF-IDF. On a clean, keyword-driven classification task, TF-IDF plus logistic regression is sometimes within two points of a fine-tuned BERT, at a thousandth of the cost. Measure the cheap thing first so you know what you bought.

Try it yourself

Take Part 2 and move the evidence before the gap: "My salary needs depositing so I went to the [MASK] ." Compare the confidence with the original. Then delete the evidence entirely — "I went to the [MASK] ." — and watch the top predictions spread out. The model's uncertainty is visible in those numbers, and reading it is a skill worth building.

Then swap bert-base-uncased for bert-base-cased in Part 1 and re-tokenize "Bengaluru traffic is unbelievable." Note how the split changes.

What to learn next

Researcher — Mathematics and papers.

Architecture

BERT (Devlin et al., 2018) is the encoder stack of the original transformer, with no decoder and no causal mask. Every token attends to every other token in both directions at every layer.

BERT-baseBERT-large
Layers L1224
Hidden width H7681024
Attention heads A1216
Parameters110M340M
Feed-forward width30724096

Vocabulary is 30,522 WordPiece tokens for the uncased models. Maximum sequence length is 512, set by the size of the learned absolute position embedding table — this is a hard architectural limit, not a configuration value.

Input representation sums three learned embeddings: token, absolute position, and segment (A or B, for sentence-pair tasks).

Masked language modelling

Sample 15% of token positions. For each selected position i, apply:

text
with probability 0.80  ->  replace token with [MASK]
with probability 0.10  ->  replace token with a uniformly random vocabulary token
with probability 0.10  ->  leave the token unchanged

The loss is cross-entropy over the original tokens at the selected positions only:

text
L_MLM = - sum over i in M of  log p(x_i | x_masked)
  • M is the set of selected positions.
  • x_i is the original token at position i.
  • x_masked is the corrupted input sequence.

Why the 80/10/10 split exists. [MASK] never appears at fine-tuning time, so a model trained only on [MASK] inputs faces a train/test mismatch in its input distribution. The 10% random and 10% unchanged cases force the model to build a useful representation at every position, not only where it sees the mask symbol. The 10% unchanged case in particular means the model cannot assume an unmasked token is correct.

The sample-efficiency cost. The loss is computed on 15% of positions. A causal language model gets a training signal at 100% of positions. Clark et al. (2020) identify this as MLM's central inefficiency and it is the motivation for ELECTRA.

Next sentence prediction, and its refutation

BERT added a binary objective: given segments A and B, predict whether B followed A in the corpus (50% true pairs, 50% random). The [CLS] representation feeds this classifier.

Liu et al. (2019), RoBERTa, ran the ablation and found removing NSP matched or improved downstream performance on every task tested, provided training used full-length contiguous sequences. Lan et al. (2019), ALBERT, offered the explanation: NSP conflates topic prediction with coherence prediction, and topic prediction is learnable from lexical overlap alone. ALBERT's sentence-order prediction — same two segments, swapped or not — removes the topic shortcut and does help.

The practical consequence is that the [CLS] vector of an off-the-shelf BERT is a weak sentence representation. It was optimised for an objective that turned out to be close to degenerate.

Pretraining cost

Original setup: BooksCorpus (800M words) plus English Wikipedia (2,500M words), 1M steps at batch size 256 sequences of 512 tokens, roughly 40 epochs over 3.3B words. Four days on 4 Cloud TPU pods (16 chips) for base; 16 pods (64 chips) for large.

RoBERTa demonstrated that BERT was substantially undertrained. Training on 160GB of text with larger batches, dynamic masking regenerated each epoch, and no NSP produced large gains from the same architecture. This is a recurring pattern: several claimed architectural improvements over BERT are recoverable by training the original longer on more data.

Fine-tuning heads

Task shapeHeadLoss
Sequence classificationlinear on [CLS]cross-entropy over classes
Token classificationlinear on every tokencross-entropy per token
Span extraction (SQuAD)two linear layers giving start and end logitscross-entropy on both
Sentence pairsegment embeddings, linear on [CLS]cross-entropy

Devlin et al. recommend 2 to 4 epochs, batch size 16 or 32, learning rate in {2e-5, 3e-5, 5e-5} with linear warmup and decay. Fine-tuning BERT-large on small datasets is unstable across random seeds; Mosbach et al. (2021) attribute this to optimisation difficulty rather than catastrophic forgetting, and show that longer training with bias correction in Adam largely fixes it. The commonly cited "try more seeds" folklore is treating a symptom.

The encoder family since 2018

  • RoBERTa (Liu et al., 2019) — no NSP, dynamic masking, more data, larger batches. Still a strong default.
  • ALBERT (Lan et al., 2019) — factorised embedding parameterisation and cross-layer parameter sharing. Fewer parameters, not proportionally faster at inference.
  • ELECTRA (Clark et al., 2020) — a small generator corrupts tokens, the discriminator predicts replaced-or-not at every position. Substantially better compute efficiency at small scale.
  • DeBERTa (He et al., 2020) — disentangled attention computing content-to-content, content-to-position and position-to-content terms separately, plus an enhanced mask decoder. DeBERTa-v3 combines this with ELECTRA-style pretraining and remains near the top for encoder classification.
  • DistilBERT (Sanh et al., 2019) — knowledge distillation to 6 layers: about 40% smaller, 60% faster, retaining roughly 97% of GLUE performance.
  • ModernBERT (Warner et al., 2024, arXiv:2412.13663) — the meaningful refresh. RoPE instead of learned absolute positions, giving an 8192-token context; GeGLU feed-forward; alternating local and global attention; unpadding; trained on 2 trillion tokens including code. If you are starting an encoder project today, start here rather than at BERT.

Indic encoders. MuRIL (Khanuja et al., 2021) covers 17 Indian languages with transliterated data included in pretraining, which matters because Hinglish and Romanised Indic text are common in the wild. IndicBERT from AI4Bharat (Kakwani et al., 2020) is the other main option. Both beat multilingual BERT on Indian-language benchmarks by clear margins, largely through better tokenizer coverage.

Why encoders are not obsolete

Decoder-only LLMs dominate the discussion, and for classification at volume they are frequently the wrong tool.

A fine-tuned DeBERTa-v3-base runs classification in single-digit milliseconds on a CPU, costs nothing per call, is deterministic, and produces a calibrated probability you can threshold. A frontier LLM doing the same task costs money per call, adds network latency, returns text you must parse, and gives you no reliable score.

The honest split: use an LLM when you have no labels, when the label set changes weekly, or when the task needs reasoning. Use a fine-tuned encoder when you have a few thousand labels, a stable label set and volume. A common production pattern is to use an LLM to create the labels, then distil into an encoder for serving.

What BERT actually learned

The "BERTology" literature is worth knowing before you over-interpret anything.

Rogers, Kovaleva & Rumshisky (2020), A Primer in BERTology (arXiv:2002.12327), surveys it. Tenney et al. (2019) found the classical NLP pipeline appears roughly in layer order — surface features early, syntax in the middle, semantics late — but with substantial overlap and no clean boundaries. Hewitt & Manning (2019) showed a linear transformation of BERT's space recovers dependency parse tree distances, which is a real structural finding.

Against that: Kovaleva et al. (2019) found heavy redundancy across heads, with a limited number of distinct attention patterns repeated many times. And attention weights are not faithful explanations — Jain & Wallace (2019) is the standard reference, noted also in attention.

Key references

  • Devlin, J., Chang, M.-W., Lee, K. & Toutanova, K. (2018). BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. arXiv:1810.04805
  • Liu, Y. et al. (2019). RoBERTa: A Robustly Optimized BERT Pretraining Approach. arXiv:1907.11692
  • Lan, Z. et al. (2019). ALBERT. arXiv:1909.11942
  • Sanh, V. et al. (2019). DistilBERT. arXiv:1910.01108
  • Clark, K. et al. (2020). ELECTRA. arXiv:2003.10555
  • He, P. et al. (2020). DeBERTa. arXiv:2006.03654
  • Rogers, A., Kovaleva, O. & Rumshisky, A. (2020). A Primer in BERTology. arXiv:2002.12327
  • Mosbach, M., Andriushchenko, M. & Klakow, D. (2021). On the Stability of Fine-tuning BERT. arXiv:2006.04884
  • Khanuja, S. et al. (2021). MuRIL: Multilingual Representations for Indian Languages. arXiv:2103.10730
  • Warner, B. et al. (2024). Smarter, Better, Faster, Longer: ModernBERT. arXiv:2412.13663

Open questions

Why masked prediction transfers at all. There is no satisfying theory explaining why reconstructing 15% of tokens yields representations that transfer to parsing, entailment and coreference. Empirical scaling results exist; a mechanism does not.

Whether bidirectionality still buys anything at scale. Large decoder-only models match or beat encoders on most understanding benchmarks despite the causal constraint. Whether the encoder's remaining advantage is representational or purely a matter of inference cost is unsettled, and it is the question that decides whether the encoder line continues.

What to learn next