Static vs contextual embeddings
A static embedding gives a word one fixed vector forever, while a contextual embedding reads the whole sentence first and builds a different vector for the same word each time.
- 11 min read
- 3 reading levels
- Published
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
A static embedding gives a word one vector, forever. A contextual embedding builds a fresh vector for that word, each time, from the sentence around it.
Think about the word "bank" on a form. Filled in near "loan" and "interest rate", you read it as a place that holds money. Filled in near "river" and "fishing", you read it as the ground beside water. Same five letters, completely different meaning. You worked that out instantly, from the surrounding words alone.
GloVe and Word2Vec cannot do that. They give "bank" one vector, decided once during training, used everywhere forever. Contextual embeddings — the kind BERT produces — work differently. They read the whole sentence first, then give "bank" a vector shaped by what is actually around it.
Why it exists
A single fixed vector per word is a real limitation. It is not a rare edge case, either. English is full of words with multiple, unrelated meanings — "bank", "bat", "spring", "book". Even more words shift meaning subtly by context, without being textbook homonyms at all.
Static embeddings handle this by collapsing every use of a word into one average vector. That vector gets dominated by whichever meaning shows up most in training. Rarer meanings do not get blended in fairly. They get outvoted, and effectively lost.
Contextual embeddings fixed this. Models like ELMo, in 2018, and BERT, later that same year, made them mainstream. Both compute a word's vector fresh, every time. They use attention to look at every other word in the sentence first.
How it works
Static (GloVe / Word2Vec):
"bank" -> always the SAME vector, every sentence, forever
Contextual (BERT):
"I sat on the river bank" -> "bank" gets ONE vector, shaped by "river"
"I deposited money at the bank" -> "bank" gets a DIFFERENT vector, shaped by "money"A contextual model reads the entire sentence first. Through attention, it lets every word's meaning be shaped by the words around it. Only then does it produce a vector for each word. The same word, in a different sentence, gets a different vector.
Where you have already seen it
- Google Translate, choosing the right translation for "bank" based on the rest of the sentence.
- Google Search, telling "bank of a river hike" and "bank near me open now" apart.
- Any modern chatbot, correctly handling a word with several meanings, without you ever noticing it was hard.
Remember this
- A static embedding gives a word one vector for life. A contextual embedding builds a new vector each time, based on the sentence.
- This matters most for words with multiple meanings, which static embeddings quietly collapse into one dominant sense.
- Contextual embeddings, powered by attention, are what made models like BERT a real step forward over GloVe and Word2Vec.
What to learn next
- Attention — the mechanism that makes contextual embeddings possible.
- BERT — the model that made contextual embeddings the default choice.
- Sentence-transformers — turning contextual word vectors into a single vector for a whole sentence.
Developer — Code and libraries.
Comparing a static GloVe vector against real contextual vectors from a small transformer model, on the exact same ambiguous word.
Setup
pip install torch transformersThe model used below, sentence-transformers/all-MiniLM-L6-v2, downloads once (about 90MB) and is small enough to run on CPU in under a second per sentence.
The static side: one vector, whatever the sentence
GloVe gives "bank" exactly one vector, so its nearest neighbours are fixed no matter what sentence you had in mind:
import gensim.downloader as api
model = api.load("glove-wiki-gigaword-50")
print("GloVe's nearest neighbours to 'bank':")
for word, score in model.most_similar("bank", topn=6):
print(f" {score:.3f} {word}")GloVe's nearest neighbours to 'bank': 0.870 banks 0.800 securities 0.797 banking 0.785 investment 0.781 exchange 0.767 financial
Every neighbour is about finance. The riverbank sense is not blended in at all — it has been entirely outweighed by the more common financial sense in the training text.
The contextual side: a different vector each time
import torch
import torch.nn.functional as F
from transformers import AutoTokenizer, AutoModel
tok = AutoTokenizer.from_pretrained("sentence-transformers/all-MiniLM-L6-v2")
model = AutoModel.from_pretrained("sentence-transformers/all-MiniLM-L6-v2")
model.eval()
sentences = [
"i sat on the river bank and watched the boats",
"i deposited the cheque at the bank",
"the money is safe in my bank account",
]
labels = ["river bank", "bank (deposit)", "bank (account)"]
def bank_vector(sentence):
tokens = tok.tokenize(sentence)
idx = tokens.index("bank") # find "bank" among this sentence's tokens
enc = tok(sentence, return_tensors="pt")
with torch.no_grad():
hidden = model(**enc).last_hidden_state[0]
return hidden[idx + 1] # +1 skips the [CLS] token at position 0
vecs = [bank_vector(s) for s in sentences]
print("cosine similarity between 'bank' vectors, sentence to sentence:\n")
for i in range(len(sentences)):
for j in range(i + 1, len(sentences)):
sim = F.cosine_similarity(vecs[i], vecs[j], dim=0).item()
print(f" {labels[i]:15} vs {labels[j]:15} cosine = {sim:.3f}")cosine similarity between 'bank' vectors, sentence to sentence: river bank vs bank (deposit) cosine = 0.733 river bank vs bank (account) cosine = 0.690 bank (deposit) vs bank (account) cosine = 0.812
Line by line
The two money-sense vectors are more similar to each other (0.812) than either is to the river sense (0.733, 0.690). The model has, without being told explicitly what "bank" means, separated the two senses of the word by how it is used in each sentence.
The gap is real, but it is not huge. These numbers stay in the 0.69–0.81 range, not something like 0.95 versus 0.20. Transformer hidden states carry a lot of shared structure regardless of meaning — position, sentence topic, grammatical role — and word sense is only one factor pulling the vectors apart. This is a genuine effect, honestly measured, not an idealised textbook result.
tok.tokenize(sentence).index("bank") finds where "bank" landed after tokenization, so the correct row of last_hidden_state can be pulled out. This only works cleanly because "bank" happens to be a single token in this tokenizer's vocabulary — a longer or rarer word could split into several subword pieces, needing a small amount of extra bookkeeping.
Common mistakes
Comparing static and contextual embeddings on the same numeric scale. GloVe's cosine similarities and a transformer's hidden-state cosine similarities are not calibrated the same way — a 0.7 in one space does not mean the same thing as a 0.7 in the other.
Assuming contextual embeddings always separate word senses cleanly. As seen above, the separation is real but partial. For a word with genuinely dozens of subtly different uses, the vectors will not neatly cluster into distinct groups — see pooling strategies for how this gets handled when building a whole-sentence vector instead of a single word's.
Using raw token hidden states as if they were sentence embeddings. The vectors extracted above describe one word's meaning in context — they are not built or trained to summarise a whole sentence on their own. Sentence-transformers covers the pooling step needed to get a proper sentence-level vector.
Try it yourself
Add a fourth sentence about a different ambiguous word, such as "the bat flew out of the cave at dusk" versus "the batsman raised his bat after the century", and adapt the code to compare their vectors for "bat". Expect a similar pattern: two genuinely different meanings, pulled apart, but not by an enormous margin.
What to learn next
- Sentence-transformers — turning these per-word contextual vectors into one vector per sentence.
- Pooling strategies — the exact methods used to combine token vectors like these.
- BERT — the model architecture producing the contextual vectors used above.
Researcher — Mathematics and papers.
The formal distinction
A static embedding is a function f: V -> R^d mapping each vocabulary word to a fixed vector, independent of any sentence. A contextual embedding is a function g: V^n -> R^{n x d} mapping an entire token sequence to a sequence of vectors, where the vector at position i is a function of the entire input sequence, not w_i in isolation. In transformer encoders, this dependency is realised through self-attention: each layer's output at position i is a learned weighted combination of every position's representation from the previous layer.
The introduction of contextual embeddings
Peters et al. (2018), Deep contextualized word representations (ELMo), were first to demonstrate this at scale, using representations from a bidirectional LSTM language model rather than a transformer, and showing consistent gains across six diverse NLP tasks over static embeddings of the time. Devlin et al. (2018), BERT, replaced the LSTM backbone with a bidirectional transformer trained via masked language modelling, and became the dominant architecture within roughly a year.
The anisotropy problem
Ethayarajh (2019), How Contextual are Contextualized Word Representations?, measured self-similarity: the average cosine similarity between a word's contextual vectors across different sentences it appears in. If contextualization worked perfectly and orthogonally to everything else, an unrelated word's self-similarity should be near the corpus's average random-pair similarity. Instead, all layers of BERT, ELMo and GPT-2 showed a strong anisotropy effect: the embedding space occupies a narrow cone, so any two random vectors in the space have surprisingly high cosine similarity, deflating the effective separation between distinct word senses. This directly explains why the developer block's measured gap (0.69–0.81) is real but numerically compressed rather than dramatic — the finding generalizes to essentially all transformer hidden states, not a quirk of the specific model used there.
Later work — Timkey & van Schijndel (2021), All Bark and No Bite: Rogue Dimensions in Transformer Language Models — traced much of this anisotropy to a small number of extreme-magnitude "rogue" dimensions that dominate the raw cosine similarity calculation, and showed that removing or standardizing them substantially improves the measured separation between genuinely different meanings, without retraining the model.
Quantifying context-sensitivity across layers
Studies probing per-layer representations (Tenney, Das & Pavlick, 2019, BERT Rediscovers the Classical NLP Pipeline) found early transformer layers behave closer to static, syntax-dominated representations, while later layers encode progressively more context-dependent, semantic information — consistent with the developer block's single-layer measurement being a lower bound on how much separation is achievable by pooling or selecting from other layers.
Complexity, precisely stated
Computing a static embedding lookup is O(1) per word, a table read. Computing a contextual embedding requires a full forward pass through the encoder, O(n^2 * d) per sentence for sequence length n and model width d, dominated by the attention score matrix. This cost difference — constant-time lookup versus quadratic-in-length neural inference — is the central practical trade-off between the two approaches, and is why static embeddings remain the right choice wherever throughput matters more than sense disambiguation.
Key references
- Peters, M. E. et al. (2018). Deep contextualized word representations. NAACL. arXiv:1802.05365
- Devlin, J. et al. (2018). BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. arXiv:1810.04805
- Ethayarajh, K. (2019). How Contextual are Contextualized Word Representations? EMNLP-IJCNLP. arXiv:1909.00512
- Timkey, W. & van Schijndel, M. (2021). All Bark and No Bite: Rogue Dimensions in Transformer Language Models Obscure Representational Quality. EMNLP. arXiv:2109.04404
- Tenney, I., Das, D. & Pavlick, E. (2019). BERT Rediscovers the Classical NLP Pipeline. ACL. arXiv:1905.05950
Current state and open problems
Contextual embeddings are the default representation for nearly every modern NLP system, but the anisotropy finding above means naive cosine similarity on raw hidden states is a measurably imperfect similarity metric — one reason sentence-transformers exists as a separate line of work, fine-tuning representations specifically to make cosine similarity behave well, rather than relying on it emerging for free from language-model pretraining. The open research question is less "static versus contextual" — that is settled — and more about which layer, which pooling strategy, and which fine-tuning objective produce representations whose geometry, not only whose downstream task accuracy, faithfully reflects meaning.
What to learn next
- Sentence-transformers — fine-tuning contextual representations specifically for similarity.
- Pooling strategies — combining per-token contextual vectors into one representation.
- Attention — the underlying mechanism that makes contextualization possible.