The BERT Family

Masked language modelling

Masked language modelling trains a model by hiding random words and asking it to guess them from context on both sides, which is how BERT learned language without any human-written labels.

Read these first

On this page 5
  1. Why it exists
  2. How it works
  3. Where you have already seen it
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Masked language modelling trains a model by blanking out random words in a sentence. It asks the model to guess what was covered up, using the words before and after the gap.

Think about a fill-in-the-blank worksheet from school: "The capital of France is ____." You do not answer from a fact memorised in isolation. You use the whole sentence around the blank to work it out. Cover a different word in a different sentence, and you use the same skill again.

BERT was trained on millions of sentences this way. Words were blanked out, over and over, until it got good at guessing what belongs in the gap. That single repeated exercise is where almost everything BERT knows about language came from.

Why it exists

Training a language model needs enormous amounts of text, and a way to check whether the model's guesses are right. Hand-labelling billions of sentences is not possible. No team could write "correct answers" for that much text.

Masked language modelling solves this with a trick. The text itself already contains the answer. Take any sentence, hide one word, and the original word you hid is automatically the correct answer. No human labeller is required. This is called self-supervised learning, because the supervision comes from the data itself, not from a person labelling it.

This is also why masked language modelling looks both directions at once, unlike a model that predicts the next word left to right. To guess a hidden word from both sides, the model has to actually understand the sentence around the gap. It cannot only continue a pattern forward.

How it works

  Original sentence:   The chai was so [MASK] that it burned my tongue.

  The model reads the whole sentence, both sides of the gap, and
  outputs a probability for every word in its vocabulary:

  potent      17.1%
  sharp        7.7%
  delicious    6.0%
  powerful     3.7%
  ...

  During training, the real answer ("hot") is compared against these
  guesses, and the model is nudged to raise the probability it gives
  to the true word next time.

Do this billions of times, across billions of sentences with different words hidden. The model gradually builds an internal sense of which words fit where. That is another way of saying it learns grammar, common facts, and word meaning, all from the same fill-in-the-blank exercise.

Where you have already seen it

  • Smartphone keyboard predictions, though those usually predict the next word rather than a masked middle word, are trained with a closely related self-supervised idea.
  • Every BERT-based tool — search ranking, spam filters, sentiment analysis — starts from a model pretrained this way. It was fine-tuned for that specific job only afterwards.
  • Grammar and spell checkers flag "the bus have arrived" as wrong. They lean on this same sense of which word is expected in a given slot.

Remember this

  • Masked language modelling hides random words and trains a model to guess them from the surrounding context.
  • It needs no human-labelled data — the original word, before it was hidden, is automatically the correct answer.
  • Because it looks at context on both sides of the gap, it is what makes BERT a bidirectional model.

What to learn next

Developer — Code and libraries.

Hugging Face's pipeline API runs masked language modelling in a couple of lines. distilbert-base-uncased is a distilled, smaller version of BERT — about 260 MB to download the first time, then cached locally.

Setup

bash
pip install transformers torch

Minimal runnable code

fill_mask.py
from transformers import pipeline

fill = pipeline("fill-mask", model="distilbert-base-uncased")
results = fill("Chennai is a [MASK] city in India.")
for r in results[:5]:
    print(f"{r['sequence']}   (score={r['score']:.3f})")
Output
chennai is a metropolitan city in india.   (score=0.236)
chennai is a major city in india.   (score=0.226)
chennai is a coastal city in india.   (score=0.056)
chennai is a port city in india.   (score=0.052)
chennai is a suburban city in india.   (score=0.033)

The first run downloads the model (about 260 MB) and caches it; later runs are fast and need no network access.

Line by line

pipeline("fill-mask", ...) loads a model with a masked-language-modelling head already attached — the layer that turns the model's internal representation of the [MASK] position into a probability over every word in the vocabulary.

[MASK] is a literal token distilbert-base-uncased was trained to expect. Every masked-language model has its own exact mask token spelling — BERT and DistilBERT use [MASK], RoBERTa uses <mask> (see the next lesson) — and using the wrong one silently produces nonsense, since the model has never seen that string in the masked position during training.

score is the model's predicted probability for that exact word filling the gap, out of its entire vocabulary — the five shown are the five highest-scoring words, not the only words the model considered.

Looking under the hood

The pipeline hides one useful detail: how the mask position's output actually becomes a ranked word list. This version does it by hand.

fill_mask_manual.py
import torch
from transformers import AutoTokenizer, AutoModelForMaskedLM

tokenizer = AutoTokenizer.from_pretrained("distilbert-base-uncased")
model = AutoModelForMaskedLM.from_pretrained("distilbert-base-uncased")
model.eval()

text = "The chai was so [MASK] that it burned my tongue."
inputs = tokenizer(text, return_tensors="pt")
mask_position = (inputs["input_ids"][0] == tokenizer.mask_token_id).nonzero(as_tuple=True)[0]

with torch.no_grad():
    logits = model(**inputs).logits

mask_logits = logits[0, mask_position, :].squeeze(0)
top5 = torch.topk(mask_logits, 5)

for score, token_id in zip(top5.values, top5.indices):
    word = tokenizer.decode([token_id])
    prob = torch.softmax(mask_logits, dim=-1)[token_id]
    print(f"{word!r:12} probability={prob.item():.3f}")
Output
'potent'     probability=0.171
'sharp'      probability=0.077
'delicious'  probability=0.060
'powerful'   probability=0.037
'poisonous'  probability=0.032

(inputs["input_ids"][0] == tokenizer.mask_token_id).nonzero(...) finds exactly which token position holds [MASK], since the model's output has one row of logits per token position, and only the masked position's row is the one worth reading.

torch.softmax(mask_logits, dim=-1) converts raw model scores (logits) into probabilities that sum to 1.0 across the entire vocabulary — this is the same softmax operation from Attention, applied here to a vocabulary-sized output instead of an attention score row.

Notice "hot" — the word actually in the original sentence, "burned my tongue" — is not in the top five here. Distilled models trade some accuracy for speed and size; a full-sized bert-base-uncased may rank it differently. This is worth remembering whenever a small model's output looks slightly off: it is a real, expected trade-off, not a bug in your code.

Common mistakes

Using the wrong mask token string for the model. [MASK] works for BERT and DistilBERT; it does nothing useful for RoBERTa, which expects <mask>. Always check tokenizer.mask_token for the model you loaded, rather than assuming.

Masking more than one word and expecting joint reasoning. A masked model predicts each masked position somewhat independently — it does not reliably produce a coherent multi-word phrase across two simultaneous masks the way a proper text generator would.

Treating the top prediction as ground truth. As shown above, a real correct word can rank outside even the top five. The model gives a distribution over plausible words, not a verified fact.

Try it yourself

Change the sentence to something with a more constrained answer, such as "The capital of Tamil Nadu is [MASK]." and check whether "chennai" appears near the top. Factual questions like this are a genuinely useful, informal way to probe what a masked language model picked up during pretraining.

What to learn next

Researcher — Mathematics and papers.

The objective

For a sequence x = (x_1, ..., x_n), a random subset of positions M subset of {1, ..., n} is selected (BERT: 15% of positions), and the model is trained to maximise:

text
L_MLM = sum over i in M of log P(x_i | x_{\M})
  • x_{\M} denotes the sequence with every masked position replaced according to the masking recipe below — the model conditions on the corrupted sequence, not on the true x_i values at masked positions.
  • P(x_i | x_{\M}) is computed by a softmax over the full vocabulary from the encoder's final hidden state at position i.

This is a bidirectional objective: unlike next-token prediction (see The next-token objective), which conditions only on x_{<i}, P(x_i | x_{\M}) conditions on tokens on both sides of position i. This is the direct cause of BERT needing no causal attention mask, unlike a decoder-only model — see Encoder, decoder and encoder-decoder models.

The 80/10/10 masking recipe

Devlin et al. (2019) use a more careful recipe than replacing every masked position outright with [MASK]. Of the selected 15% of positions:

  • 80% are replaced with the literal [MASK] token.
  • 10% are replaced with a random token from the vocabulary.
  • 10% are left unchanged.

This recipe exists to close a train/inference mismatch: [MASK] never appears in real text the model sees at fine-tuning or inference time, so a model trained purely against [MASK] could learn to represent every token as if a mask token were always present, degrading representations of ordinary, unmasked tokens. Mixing in random and unchanged tokens forces the model to build good representations for every input position, since it cannot know in advance which positions the loss will actually be computed on.

Complexity and pretraining cost

Per forward pass, MLM pretraining costs the same as any encoder forward pass — O(n^2 * d) for the attention operation, O(n * d^2) for the feedforward and projection layers, following the same accounting as Attention's researcher block. The loss itself is computed only at masked positions, an O(|M| * |V|) softmax cost that is small relative to the encoder's forward cost for realistic vocabulary sizes |V|.

Because only 15% of positions contribute to the loss per training example, MLM is comparatively sample-inefficient per token processed relative to next-token prediction, which supervises every position in a sequence. This is one contributing reason later encoder work (ELECTRA, see below) targeted denser supervision signals.

Replaced token detection: a denser alternative

Clark et al. (2020), ELECTRA, replace MLM with replaced token detection: a small generator model proposes plausible replacement tokens at masked positions (much like MLM), then a separate discriminator model is trained to classify every token in the sequence as "original" or "replaced" — a binary decision at every position, not only the 15% that were masked.

text
L_disc = sum over i=1..n of  [ -log D(x_i^corrupt) if original,  -log(1 - D(x_i^corrupt)) if replaced ]
  • D(x_i^corrupt) is the discriminator's predicted probability that position i holds the original, unreplaced token.

Because every position contributes to the loss, ELECTRA reaches a given downstream accuracy with substantially less pretraining compute than BERT-style MLM — Clark et al. report ELECTRA-Small matching GPT-level performance with roughly 1/20th the pretraining compute of comparable models, an efficiency argument rather than a capability ceiling difference.

Key references

  • Devlin, J., Chang, M., Lee, K. & Toutanova, K. (2019). BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. arXiv:1810.04805
  • Taylor, W. (1953). "Cloze Procedure": A New Tool for Measuring Readability. Journalism Quarterly. The fill-in-the-blank reading-comprehension test MLM is directly modelled on, predating neural networks entirely.
  • Clark, K., Luong, M., Le, Q. & Manning, C. (2020). ELECTRA: Pre-training Text Encoders as Discriminators Rather Than Generators. arXiv:2003.10555

Current state and open problems

Masked language modelling remains the dominant pretraining objective for encoder-only models, including current strong encoders such as ModernBERT, though the specific masking recipe has been revisited repeatedly — see RoBERTa for what happens when masking is made dynamic rather than fixed once at data-preprocessing time. ELECTRA's denser signal has not fully displaced plain MLM in practice, in part because ELECTRA's two-model generator-discriminator setup adds real training complexity, and because the compute savings matter most at smaller model and data scales than current frontier encoder training operates at. Whether encoder pretraining objectives will keep evolving independently, or gradually converge toward decoder-style next-token training as multi-task and instruction-tuned models blur the encoder/decoder distinction, remains an open direction rather than a settled one.

What to learn next

What to learn next

These follow on from what you just read.

  • The BERT Family

    RoBERTa

    RoBERTa is BERT trained more carefully rather than redesigned — more data, longer training, dynamic masking, and one dropped training objective — and it beat BERT on almost every benchmark from doing that alone.

  • The BERT Family

    DeBERTa

    DeBERTa keeps a word's meaning and its position as two separate pieces of information all the way through the network, instead of merging them at the start, and that one change measurably improved on RoBERTa.

  • The BERT Family

    ModernBERT

    ModernBERT rebuilds BERT's architecture with everything the field learned from six years of large language model research, and reaches 8,192 tokens of context instead of BERT's 512, at real-world speeds.