Embeddings
An embedding is a list of numbers that stands for a word or a sentence, arranged so that things with similar meaning end up close together.
- 16 min read
- 3 reading levels
- Published
Read these first
On this page 6
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
An embedding is a list of numbers that stands for a word. Words with similar meaning get similar lists.
Think about seating guests at a wedding. You do not scatter people randomly. Cousins go on one table. Office colleagues go on another. The uncles who talk only about property go in the corner.
Nobody wrote down a rule for that seating. You arranged the hall so that distance means something. Sit close, you have something in common. Sit far apart, you do not.
An embedding does exactly this to words. It gives every word a seat in a very large hall. Neighbours in that hall share meaning.
Why it exists
Go back to how a model stores words. After tokenization, every piece of text is a number from a list. "chai" might be piece number 4021. "coffee" might be 9887.
Those numbers are name tags, nothing more. Number 4021 is not smaller or warmer than 9887. The model has no way to know that chai and coffee belong together. Every word is a stranger to every other word.
That was the wall NLP hit for decades. You could teach a model everything about "chai" and it would learn nothing about "coffee". Every word needed its own training examples.
Embeddings broke the wall. Give words a seat in a hall instead of a name tag, and knowledge spreads. Teach the model about chai, and coffee — sitting nearby — gets some of it for free.
How it works
Every word gets a long list of numbers. Each position in the list is like a dial the model can turn.
Nobody decides what the dials mean. The model works them out by reading text and noticing which words show up in similar surroundings.
THE HALL (a made-up, tiny version)
hot-drinks corner vehicles corner
--------------------- -------------------
chai bike
coffee bus
tea train
Words that appear in similar sentences
end up seated near each other.The rule the model follows is old and surprisingly powerful. A word is known by the company it keeps.
Look at these three sentences.
I drink chai every morning
I drink coffee every morning
I drink tea every morningChai, coffee and tea are surrounded by the same words. The model has never tasted anything. It notices the surroundings match, and seats them together. That is the whole trick.
This bit trips up almost everyone
Embeddings do not know what things are. They know how words are used.
So "good" and "bad" often sit close together, which feels wrong. But look at how they are used. "The food was good." "The food was bad." Identical surroundings. Same table at the wedding.
If this feels confusing the first time, that is normal. Read the two paragraphs above again. Nearly everyone stumbles here, and it is the single most useful thing to understand about embeddings.
Where you have already seen it
- Search that understands you. Type "cheap phone under 15000" and get results titled "best budget smartphones". No word matched. The meaning did.
- Google Photos. Search "beach" and it finds beach photos you never labelled. Photos get embeddings too.
- YouTube and Spotify. Videos and songs get seats in a hall. Yours sits near what you already watched.
- Chatbots that read your PDF. Your question is given a seat, and the nearest paragraphs are pulled in as the answer material.
Remember this
- An embedding is a list of numbers standing in for a word, a sentence, an image or a song.
- Similar meaning gives a similar list. Distance in the hall is the measure of similarity.
- Embeddings capture how words are used, not what they are — which is why opposites can sit close.
What to learn next
- Word2Vec — the method that made embeddings practical.
- Vector databases — where embeddings get stored and searched.
- What is RAG? — the most common use of embeddings today.
Developer — Code and libraries.
Two short programs. The first shows exactly what an embedding is and how similarity is measured. The second learns similarity from a corpus, so you can see it is not hand-crafted.
Both run in under a second on any laptop.
Setup
pip install numpyThat is all. No model download, no network access, no GPU.
Part 1 — what an embedding literally is
An embedding table is a plain 2-D array. Row i is the vector for vocabulary item i. Looking up a word is an array index and nothing more.
Here the three dials are hand-written so you can check the arithmetic yourself. A real model learns them.
import numpy as np
vocab = ["chai", "coffee", "tea", "bike", "train"]
# Three made-up meaning dials: hot-drink-ness, vehicle-ness, morning-ness.
table = np.array([
[0.9, 0.0, 0.8], # chai
[0.9, 0.0, 0.7], # coffee
[0.8, 0.0, 0.4], # tea
[0.0, 0.9, 0.1], # bike
[0.0, 0.8, 0.2], # train
])
def similarity(a, b):
va, vb = table[vocab.index(a)], table[vocab.index(b)]
return float(va @ vb / (np.linalg.norm(va) * np.linalg.norm(vb)))
print("chai vs coffee:", round(similarity("chai", "coffee"), 3))
print("chai vs tea :", round(similarity("chai", "tea"), 3))
print("chai vs bike :", round(similarity("chai", "bike"), 3))
print("bike vs train :", round(similarity("bike", "train"), 3))chai vs coffee: 0.998 chai vs tea : 0.966 chai vs bike : 0.073 bike vs train : 0.991
The function is cosine similarity, which measures the angle between two vectors and ignores their length. It runs from 1.0 for identical direction, through 0.0 for perpendicular, to -1.0 for opposite.
Ignoring length is the point. A word that appears a million times gets a longer vector than a rare word, and you do not want frequency masquerading as meaning.
Part 2 — learning it from text instead
Nothing above was learned. This is. It reads five sentences and works out which words behave alike, using only counting.
import numpy as np
from collections import defaultdict
sentences = [
"i drink chai every morning",
"i drink coffee every morning",
"i drink tea every morning",
"she rides a bike to work",
"he rides a bike to college",
]
counts = defaultdict(lambda: defaultdict(int))
for s in sentences:
words = s.split()
for i, target in enumerate(words):
for j in range(max(0, i - 2), min(len(words), i + 3)):
if j != i:
counts[target][words[j]] += 1 # who sits near whom
vocab = sorted({w for s in sentences for w in s.split()})
M = np.array([[counts[a][b] for b in vocab] for a in vocab], dtype=float)
def cos(a, b):
va, vb = M[vocab.index(a)], M[vocab.index(b)]
return float(va @ vb / (np.linalg.norm(va) * np.linalg.norm(vb)))
print("vocabulary:", len(vocab), "words")
for a, b in [("chai", "coffee"), ("chai", "tea"), ("chai", "bike"), ("bike", "rides")]:
print(f"{a:6} vs {b:6} -> {cos(a, b):.2f}")vocabulary: 15 words chai vs coffee -> 1.00 chai vs tea -> 1.00 chai vs bike -> 0.00 bike vs rides -> 0.34
Nobody told the program that chai and coffee are both drinks. It counted neighbours.
Line by line
The window. range(max(0, i - 2), min(len(words), i + 3)) collects the two words on each side of the target. This window size is a real hyperparameter. Small windows (2 or so) capture grammatical role. Large windows (10 or so) capture topic. Neither is correct — they answer different questions.
Each row of M is an embedding. Row chai records which words appeared near chai across the whole corpus. It is a crude embedding, but it is a genuine one, and it is what every method in this area started from.
Why chai vs coffee is exactly 1.00. Those three sentences are identical apart from the drink. So the rows for chai, coffee and tea are byte-for-byte identical, and the angle between them is zero.
Treat that perfect score with suspicion rather than pride. It reflects a toy corpus, not a good model. Real text never gives identical contexts, and real similarities land somewhere between 0.6 and 0.9.
Why chai vs bike is exactly 0.00. The two rows share no non-zero position at all. Drinks and bikes never appear near the same words in this corpus. In real text this almost never happens, because common words like "the" appear near everything.
Why bike vs rides is only 0.34. This one is worth pausing on. bike and rides sit next to each other in every sentence they appear in — yet they score low.
That is correct behaviour. Distributional similarity asks whether two words appear in the same surroundings, not whether they appear together. Chai and coffee are similar because they are interchangeable. Bike and rides are related but not interchangeable, and this measure does not confuse the two.
From counting to real embeddings
The rows above have one number per vocabulary word. With a 50,000-word vocabulary, each row is 50,000 numbers and nearly all of them are zero.
Real embeddings compress that into a few hundred dense numbers, using either matrix factorisation or a small neural network. Compression is what forces generalisation. With only 300 numbers to spend, the model has to put chai and coffee near each other. It cannot afford to store them separately.
The counting stays the same. Only the storage changes. See Word2Vec for the version that made this practical at corpus scale.
Common mistakes
Comparing vectors from two different models. Every model builds its own hall with its own seating. A vector from model A and a vector from model B are two different coordinate systems, and the cosine between them is meaningless. If you change embedding models, you re-embed your entire corpus. Every time. No exceptions.
Using the wrong distance function. Many modern embedding models are trained with normalised vectors and expect cosine. Use raw dot product on unnormalised vectors and long vectors win regardless of meaning. Check the model card, then match it.
Averaging word vectors to represent a sentence. It is quick and it discards word order entirely. "Dog bites man" and "man bites dog" produce the identical average. Use a sentence-embedding model when the sentence matters.
Expecting facts. An embedding encodes usage patterns, not knowledge. It cannot tell you the capital of Karnataka. It can tell you that "Bengaluru" is used the way city names are used.
Ignoring what the model absorbed. Embeddings trained on web text learn the associations in that text, including the prejudiced ones. Bolukbasi et al. (2016) documented occupation and gender associations in published word vectors. If your embeddings feed a hiring or lending system, test for this before shipping, not after.
Try it yourself
Add "she rides a scooter to work" to sentences and compare bike with scooter.
Then predict the result before running it. bike appears in two sentences and scooter in one, so their raw counts differ in size. Cosine similarity ignores size and looks only at direction. Work out what you expect, then check.
After that, change the window from 2 to 1 and re-run everything. Watch which scores move. That single number changes what "similar" means, and it is a decision every embedding pipeline makes whether you notice it or not.
What to learn next
- Word2Vec — the shallow objective that made dense embeddings cheap.
- Vector databases — searching millions of these vectors quickly.
- Attention — how a word's vector changes to suit its sentence.
Researcher — Mathematics and papers.
The distributional hypothesis
The theoretical grounding predates the field. Harris (1954), Distributional Structure, argued that difference in meaning correlates with difference in distribution. Firth (1957) gave the slogan version: a word is characterised by the company it keeps.
Every method below is an operationalisation of that single claim. It is an empirical hypothesis, and it has a known boundary. Antonyms share distributions almost perfectly. That is why cosine similarity between "hot" and "cold" is high in essentially every unsupervised embedding space.
Count-based methods
Build a co-occurrence matrix M of shape |V| x |C|, where |V| is vocabulary size and |C| is the context vocabulary. Raw counts are dominated by frequency, so the standard correction is positive pointwise mutual information:
PMI(w, c) = log( P(w, c) / (P(w) * P(c)) )
PPMI(w, c) = max(PMI(w, c), 0)P(w, c)is the joint probability of wordwoccurring with contextc.P(w)andP(c)are the marginals.- The
maxwith 0 discards negative associations, which are poorly estimated from finite corpora.
Latent semantic analysis (Deerwester et al., 1990) then applies truncated SVD, M ~= U_k * S_k * V_k^T, and takes U_k * S_k as the embedding. k is the retained rank, typically 100 to 500.
Cost is O(|V| * |C| * k) for truncated SVD via randomised methods, and the full matrix must be materialised or streamed.
Predictive methods
Mikolov et al. (2013) introduced skip-gram with negative sampling. For a centre word w and context word c:
L = -log sigma(v_c . v_w) - sum over i = 1..k of E[ log sigma(-v_{n_i} . v_w) ]v_wis the centre-word vector,v_cthe context vector for the observed pair.sigma(x) = 1 / (1 + exp(-x))is the logistic function.n_iareknegative samples drawn from a noise distribution, in practice the unigram distribution raised to the power 0.75.kis typically 5 to 20 for small corpora, 2 to 5 for large ones.
The 0.75 exponent is an empirical choice with no derivation behind it. It downweights frequent words less aggressively than uniform sampling and more than the raw unigram distribution.
Levy & Goldberg (2014), Neural Word Embedding as Implicit Matrix Factorization (NeurIPS), settled what this objective is doing. Skip-gram with negative sampling implicitly factorises the shifted PPMI matrix, PMI(w, c) - log k. Count-based and predictive methods are therefore the same object, reached by different computational routes. That result closed a noisy debate in the literature.
GloVe (Pennington, Socher & Manning, 2014) makes the factorisation explicit, with a weighted least-squares objective over log co-occurrence counts.
fastText (Bojanowski et al., 2017) represents a word as the sum of its character n-gram vectors. This gives non-zero vectors for unseen words and helps substantially for morphologically rich languages — Tamil, Turkish, Finnish, and the Indic family generally.
Contextual embeddings
Static embeddings assign one vector per word type. That is a modelling error whenever a word has more than one sense.
ELMo (Peters et al., 2018) and BERT (Devlin et al., 2018) produce one vector per word token, computed from the whole sentence. "bank" in a river sentence and "bank" in a finance sentence receive different vectors.
Ethayarajh (2019), How Contextual are Contextualized Word Representations? (arXiv:1909.00512), measured the geometry and found something awkward. Contextual representations occupy a narrow cone rather than filling the space, so the expected cosine similarity between two randomly chosen words is far above zero. Raw cosine on these vectors is therefore not directly interpretable, and anisotropy correction — mean-centring, or whitening — measurably improves downstream similarity tasks.
Sentence embeddings
Averaging BERT token vectors gives poor sentence similarity, worse than averaging GloVe vectors. Reimers & Gurevych (2019), Sentence-BERT (arXiv:1908.10084), fixed this with a siamese architecture trained on natural language inference pairs. Cosine in the output space is then trained to mean semantic similarity.
Current models train with contrastive objectives on large mined pair corpora. The standard loss is InfoNCE:
L = -log( exp(sim(q, p) / t) / sum over j of exp(sim(q, d_j) / t) )qis the query embedding,pthe positive passage embedding.d_jranges over the positive and all in-batch negatives.simis cosine similarity.tis the temperature, typically 0.01 to 0.07. Lowtsharpens the distribution and increases the penalty on hard negatives.
Two practical facts dominate results here. Large batch sizes matter, because in-batch negatives scale with batch size. And hard negative mining — negatives that a first-pass retriever ranked highly — matters more than architecture.
Retrieval cost
Exact nearest-neighbour search over n vectors of dimension d costs O(n * d) per query. At n = 10^8 and d = 768 that is untenable.
HNSW (Malkov & Yashunin, 2018) builds a navigable small-world graph with O(log n) expected search hops, giving high recall at a fraction of exact cost. Memory is the binding constraint: the graph plus full-precision vectors runs roughly 4 * d bytes per vector plus edge storage.
Product quantisation and its variants trade recall for memory, typically 8 to 32 times compression. Matryoshka representation learning (Kusupati et al., 2022, arXiv:2205.13147) trains so that leading prefixes of the vector are independently usable. You can retrieve with the first 64 dimensions, then re-rank with all 768.
Key references
- Harris, Z. (1954). Distributional Structure. Word 10(2-3).
- Deerwester, S. et al. (1990). Indexing by Latent Semantic Analysis. JASIS 41(6).
- Mikolov, T. et al. (2013). Distributed Representations of Words and Phrases and their Compositionality. arXiv:1310.4546
- Pennington, J., Socher, R. & Manning, C. (2014). GloVe: Global Vectors for Word Representation. EMNLP.
- Levy, O. & Goldberg, Y. (2014). Neural Word Embedding as Implicit Matrix Factorization. NeurIPS 27.
- Bolukbasi, T. et al. (2016). Man is to Computer Programmer as Woman is to Homemaker? arXiv:1607.06520
- Bojanowski, P. et al. (2017). Enriching Word Vectors with Subword Information. arXiv:1607.04606
- Peters, M. et al. (2018). Deep contextualized word representations. arXiv:1802.05365
- Reimers, N. & Gurevych, I. (2019). Sentence-BERT. arXiv:1908.10084
- Malkov, Y. & Yashunin, D. (2018). Efficient and robust approximate nearest neighbor search using HNSW graphs. arXiv:1603.09320
Current state and honest caveats
The analogy result is weaker than folklore suggests. The king - man + woman ~= queen demonstration is real but fragile. Standard evaluation code excludes the three input words from the candidate set. Without that exclusion, the nearest vector to the computed point is frequently king itself. Linzen (2016) and Rogers et al. (2017) documented this. Treat vector arithmetic as a suggestive property, not as evidence of compositional structure.
Benchmark saturation. MTEB is the standard evaluation suite, and the top of the leaderboard is now tightly clustered and contaminated by training on evaluation-adjacent data. Differences of one point are not meaningful. Evaluate on your own retrieval set.
Dimension is not quality. A well-trained 384-dimension model routinely beats a poorly trained 1536-dimension one, at a quarter of the storage and latency. Measure recall on your data before paying for width.
Multilingual gaps persist. Cross-lingual alignment for Indic scripts remains materially behind English and Chinese, driven by corpus availability and by the tokenizer fertility problem described in tokenization.
What to learn next
- Word2Vec — the objective and its implicit factorisation in full.
- Attention — how static vectors became contextual ones.
- Vector databases — indexing, quantisation and recall trade-offs.