BERT
In one sentence BERT is a language model that reads a whole sentence in both directions at once to understand it, rather than writing new text one word at a time.
Updated
BERT is a language model built to understand text — it reads a whole sentence in both directions at once, instead of generating new text left to right.
There are two ways to work with a sentence, and they need different tools. A proofreader reads the full line, sees the word after the gap as well as the word before it, and knows what belongs in the middle. A storyteller only has what they have said so far. BERT is the proofreader; GPT-style models are the storyteller. BERT stands for Bidirectional Encoder Representations from Transformers, and "bidirectional" is the whole point.
It was trained by fill-in-the-blank. Google's team took ordinary text, hid about 15% of the tokens, and trained the model to recover them from the words on both sides. To fill "The ___ was so crowded we missed our train", the model has to learn something real about stations, crowds and trains.
What it is good for
"The battery life on this phone is terrible" → BERT → sentiment: negative
"Ravi joined Infosys in Pune last March" → BERT → PERSON, ORG, LOCATION, DATEBERT does not write essays. It produces a rich representation of text that you attach a small task head to, then fine-tune on a few thousand of your own labelled examples. For classification, search ranking, entity extraction and duplicate detection, a fine-tuned BERT-family model often beats prompting a large generative model. It is more accurate, far cheaper and far faster. A base model is around 110 million parameters and runs comfortably on a laptop CPU. Its descendants, RoBERTa, DeBERTa and DistilBERT, are usually the practical picks today, and the original paper is Devlin et al., 2018.
Where to go next
- Full lesson: BERT
- Related terms: transformer, token, embedding, fine-tuning