Word2Vec
Word2Vec learns a number list for every word by playing a guessing game about which words sit near which, then throws the game away and keeps the numbers.
- 18 min read
- 3 reading levels
- Updated
Read these first
On this page 10
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
Word2Vec is a way to learn number lists for words by guessing which words sit near each other.
The analogy you have already lived
A friend reads a sentence out loud, but mumbles one word on purpose. "Pass me the ____." You guess. Sometimes you are right, sometimes you are corrected.
Play that game a few thousand times and something happens without you deciding it. You start to feel which words belong in which slots. "Salt" and "sugar" both fit after "pass me the". "Bicycle" does not.
Word2Vec plays exactly this game, on billions of sentences, and keeps what it learned.
Why it exists
The embeddings lesson showed the counting approach. Build a big table. One row per word, one column per neighbour word. Count who appeared near whom.
That table works, and it is enormous. With a vocabulary of 50,000 words, each row holds 50,000 numbers. Almost all of them are zero. To shrink it you had to run heavy matrix maths over the whole table at once.
Two problems followed. It needed the entire table in memory before you could start. And it got slower in a hurry as the vocabulary grew.
In 2013 a team at Google published Word2Vec, which sidestepped both. It never builds the table. It walks through the text one window at a time, nudging numbers as it goes. It ran on a normal desktop machine, on a billion words, in under a day.
That is the whole reason this lesson exists. Word2Vec did not invent the idea that meaning lives in company. It made that idea affordable.
The trick that surprises people
The guessing game is not the point. The answers are thrown away.
What you keep is the number list the model built for each word while it was trying to guess well. The game was scaffolding. The scaffolding is demolished and the building stays.
This is called a pretext task — a fake job set for a model. Solving it forces the model to learn what you actually wanted. It shows up everywhere in modern AI, and Word2Vec is the cleanest example of it.
How it works
Slide a small window along the text. Inside each window, one word is the centre and the rest are its neighbours.
... i drink chai every morning ...
| | |
left CENTRE rightThere are two ways to play, and both were published together.
Skip-gram. Show the model the centre word. Ask it to guess the neighbours.
chai -> guess: drink, every, i, morningCBOW, short for continuous bag of words. Show the model the neighbours. Ask it to guess the missing centre.
drink _____ every -> guess: chaiSkip-gram is slower and better for rare words. CBOW is faster and smoother on common ones. Most people reach for skip-gram.
The second trick: do not check every word
Here is a problem the first version hit hard.
To score a guess properly, the model would have to compare its answer against every single word in the dictionary. With 50,000 words that is 50,000 comparisons, for every window, over billions of windows. Far too slow.
The fix is called negative sampling. Check the real neighbour, plus a small handful of random impostors. Skip the rest of the dictionary.
centre word: chai
real neighbour -> "yes, these belong together" (push closer)
5 random words -> "no, these do not" (push apart)Five fake comparisons instead of fifty thousand. Same lesson learned, ten thousand times cheaper. This single change is most of why Word2Vec was fast enough to matter.
Where you have already seen it
- Shopping sites. Treat a customer's browsing session as a sentence, and each product as a word. Word2Vec on that gives you "people who looked at this also looked at" without any ratings.
- Music and video queues. Playlists become sentences, songs become words. Nearby songs get similar number lists.
- Search suggestions. Typing "cheap flight" surfaces "budget airline" because those phrases live in similar company.
- Spelling and query correction. A misspelling sits near its correct form when both appear in similar surroundings.
What is honestly limited about it
Word2Vec gives each word one number list, forever. The word "bank" gets a single seat, blending the river bank and the money bank into one blurred point.
That is a real modelling error, not a small imperfection. It cannot be fixed by training longer. Fixing it needed a different idea, called attention. Those models give a word a fresh number list in every sentence it appears in.
Word2Vec also absorbs whatever is in the text it read. If the training text pairs certain jobs with certain genders, the vectors carry that pairing. Nobody put it there on purpose. It came in with the data.
Remember this
- Word2Vec learns word vectors by guessing neighbours, then discards the guessing machinery.
- Negative sampling is what made it fast: check a few random impostors, not the whole dictionary.
- Every word gets one vector, so words with several meanings end up blurred.
What to learn next
- Attention — how a word gets a different vector in every sentence.
- BERT — the model that made context-dependent vectors standard.
- Vector databases — storing and searching millions of vectors.
Developer — Code and libraries.
Word2Vec is short enough to write from scratch, and writing it once removes all the mystery. The version below is skip-gram with negative sampling, in plain NumPy.
It trains on ten sentences in about a second on any laptop. No GPU, no download.
Setup
pip install numpySkip-gram with negative sampling, complete
import numpy as np
sentences = [
"i drink chai every morning", "i drink coffee every morning",
"i drink tea every morning", "she drinks chai after lunch",
"he drinks coffee after lunch", "they drink tea after lunch",
"i ride a bike to work", "she rides a bike to college",
"he rides a scooter to work", "they ride a scooter to college",
]
tokens = [w for s in sentences for w in s.split()]
vocab = sorted(set(tokens))
idx = {w: i for i, w in enumerate(vocab)}
V, D, WINDOW, NEG = len(vocab), 8, 2, 5
# Every (centre word, nearby word) pair in the corpus. This is the whole training set.
pairs = []
for s in sentences:
ws = [idx[w] for w in s.split()]
for i, centre in enumerate(ws):
for j in range(max(0, i - WINDOW), min(len(ws), i + WINDOW + 1)):
if j != i:
pairs.append((centre, ws[j]))
rng = np.random.default_rng(0)
Wc = rng.normal(0, 0.1, (V, D)) # centre vectors — these become the embeddings
Wo = rng.normal(0, 0.1, (V, D)) # context vectors — thrown away when training ends
counts = np.bincount([idx[w] for w in tokens], minlength=V).astype(float)
noise = counts ** 0.75 # word2vec's own choice: rare words get sampled more often
noise /= noise.sum()
def sigmoid(x):
return 1.0 / (1.0 + np.exp(-np.clip(x, -30, 30))) # clip so exp cannot overflow
lr = 0.05
order = np.arange(len(pairs))
for epoch in range(300):
rng.shuffle(order)
for p in order:
centre, context = pairs[p]
targets = np.concatenate(([context], rng.choice(V, size=NEG, p=noise)))
labels = np.array([1.0] + [0.0] * NEG) # 1 = real neighbour, 0 = impostor
v = Wc[centre]
err = sigmoid(Wo[targets] @ v) - labels # gradient of the logistic loss
Wc[centre] -= lr * (err @ Wo[targets])
Wo[targets] -= lr * np.outer(err, v)
def cos(a, b):
x, y = Wc[idx[a]], Wc[idx[b]]
return float(x @ y / (np.linalg.norm(x) * np.linalg.norm(y)))
def nearest(word, n=3):
scores = [(cos(word, w), w) for w in vocab if w != word]
return [f"{w} {s:+.2f}" for s, w in sorted(scores, reverse=True)[:n]]
print(f"{V} words, {len(pairs)} training pairs, {D} numbers per word")
for a, b in [("chai", "coffee"), ("chai", "tea"), ("bike", "scooter"), ("chai", "bike")]:
print(f" {a:5} vs {b:8} {cos(a, b):+.2f}")
print()
for w in ("chai", "bike"):
print(f" nearest to {w}: {', '.join(nearest(w))}")21 words, 156 training pairs, 8 numbers per word chai vs coffee +0.94 chai vs tea +0.64 bike vs scooter +0.98 chai vs bike -0.31 nearest to chai: coffee +0.94, lunch +0.81, drinks +0.78 nearest to bike: scooter +0.98, she +0.76, to +0.70
Ten sentences, one second, and the drinks have found each other. Nobody told the program that chai and coffee are related.
Read the output honestly
The good result is real. bike and scooter land at +0.98. chai and bike land at -0.31, on opposite sides.
The messy result is more instructive. Look at the nearest neighbours of chai: coffee, then lunch, then drinks.
lunch is not a drink. It scored high because "after lunch" appears in the same three sentences as chai and coffee do. With ten sentences, that is enough to fuse them.
This is what a small corpus does. It cannot separate "words that mean similar things" from "words that happened to co-occur in my tiny sample". Word2Vec on ten sentences is a demonstration. Word2Vec on ten million sentences is a tool.
Line by line, for the parts that are not obvious
Two matrices, not one. Wc holds the vector a word has when it is the centre. Wo holds the vector it has when it is a neighbour. They are separate on purpose. If a word shared one vector for both roles, the model could score "chai near chai" highly and cheat. At the end, Wo is discarded.
counts ** 0.75 shapes the impostor lottery. Draw impostors by raw frequency and you get "the" almost every time, which teaches nothing. Draw them uniformly and rare junk words dominate. The 0.75 exponent sits between the two. It has no derivation behind it — the original authors tried values and this one worked best.
err = sigmoid(...) - labels is the entire gradient. The real neighbour has label 1, so its error is negative when the model already scores it highly, and the update is small. An impostor has label 0, so a high score gives a large positive error and a strong push apart.
np.clip(x, -30, 30) inside the sigmoid. Without it, a confident score produces np.exp of a large negative number and NumPy prints RuntimeWarning: overflow encountered in exp. The result is still usable, but a warning in your training loop is a habit worth not forming.
WINDOW = 2. This is the single most consequential number here. A small window learns grammatical role — words that could replace each other. A large window learns topic — words that appear in the same kind of document. Neither is right. They answer different questions.
In real work, use gensim
Writing it once is educational. Shipping it is not.
pip install gensimfrom gensim.models import Word2Vec
sentences = [s.split() for s in ["i drink chai every morning", "i drink coffee every morning"]]
model = Word2Vec(sentences, vector_size=100, window=5, min_count=1, sg=1, workers=4, seed=0)
print(model.wv.most_similar("chai"))No output block for this one. On two sentences the result is noise, and gensim's numbers depend on version, thread count and scheduling, so any figure printed here would be a lie. sg=1 selects skip-gram; sg=0 selects CBOW.
Common mistakes
Training on too little text. Word2Vec needs volume. Under roughly ten million tokens the vectors are unstable, and the similarity scores you read off them will not survive a re-run with a different seed. If you have a small corpus, use pretrained vectors or a sentence-embedding model instead.
Leaving min_count=1. Words seen once or twice get vectors built from almost no evidence, and they pollute your nearest-neighbour results. The gensim default of 5 exists for a reason. Raise it, do not lower it.
Comparing vectors across two training runs. Each run builds its own coordinate system, from its own random start. Vector 42 from run A and vector 42 from run B are unrelated. Re-embed everything whenever you retrain.
Expecting the king-minus-man-plus-woman trick to be robust. It works in demonstrations because the evaluation code excludes the three input words from the candidate list. Without that exclusion the nearest point is frequently one of the inputs. Treat the analogy result as a curiosity, not as proof of structure.
Reaching for Word2Vec in 2026 by default. For search and retrieval, a modern sentence-embedding model beats it comfortably. Word2Vec is still the right tool when your "words" are not words at all — product ids, page ids, song ids — where no pretrained model exists.
Try it yourself
Change WINDOW from 2 to 1 and re-run. Watch nearest to chai change. A one-word window mostly captures which words can stand in the same grammatical slot.
Then add "i drink chai every evening" to the sentences and check cos("morning", "evening") before and after. Predict the direction of the change first. Both words now appear in the same slot after "every", so they should move toward each other.
What to learn next
- Attention — one vector per word in a sentence, not one per word forever.
- BERT — pretraining with a harder guessing game.
- Text classification — putting word features to work on a real task.
Researcher — Mathematics and papers.
The two architectures
Mikolov et al. (2013a) proposed both in one paper. Given a corpus of tokens w_1 ... w_N and a window size c:
Skip-gram maximises the log probability of context words given the centre word:
(1/N) * sum over t = 1..N of sum over -c <= j <= c, j != 0 of log p(w_{t+j} | w_t)CBOW reverses the conditioning, predicting the centre from the summed context:
(1/N) * sum over t = 1..N of log p(w_t | w_{t-c} ... w_{t+c})Nis the number of tokens in the corpus.cis the maximum window radius.w_tis the token at positiont.
CBOW averages the context vectors before the prediction, so one update per window. Skip-gram makes 2c separate predictions per window, so it does more work per token and extracts more signal from rare words.
The softmax that had to go
The naive parameterisation is:
p(w_O | w_I) = exp(v'_{w_O} . v_{w_I}) / sum over w = 1..V of exp(v'_w . v_{w_I})v_wis the input (centre) vector for wordw;v'_wis the output (context) vector.Vis the vocabulary size..denotes the dot product.
The denominator costs O(V) per training example. At V = 10^6 and 10^10 training pairs this is not slow, it is impossible. Both published alternatives attack that denominator.
Hierarchical softmax
Arrange the vocabulary as the leaves of a binary Huffman tree. The probability of a word is the product of the binary decisions along the root-to-leaf path:
p(w | w_I) = product over j = 1..L(w)-1 of sigma( [[ n(w, j+1) = left-child(n(w, j)) ]] * v'_{n(w,j)} . v_{w_I} )n(w, j)is thej-th node on the path from root to the leaf forw.L(w)is the path length,O(log V)under a Huffman code.[[ x ]]is+1whenxis true and-1otherwise.sigmais the logistic function.
Cost falls from O(V) to O(log V). Huffman coding gives frequent words short paths, so the average is better than log V. Hierarchical softmax remains preferable for very large vocabularies and for rare-word quality.
Negative sampling
Mikolov et al. (2013b) replaced the multiclass problem with binary discrimination, a simplification of noise-contrastive estimation (Gutmann & Hyvärinen, 2012). The embeddings lesson states the objective and Levy & Goldberg's proof that it implicitly factorises shifted PPMI. Three implementation details are worth adding here, because all three affect results measurably and none appears in the paper's equations.
The noise distribution. Negatives are drawn from P_n(w) proportional to U(w)^0.75, where U(w) is the unigram frequency. The exponent is empirical. Levy, Goldberg & Dagan (2015) confirmed it transfers to other methods, which suggests it is correcting for frequency effects rather than for anything specific to skip-gram.
Frequent-word subsampling. Each token is discarded during corpus reading with probability:
P(discard w) = 1 - sqrt( t / f(w) )f(w)is the relative frequency ofwin the corpus.tis a threshold, typically10^-5.
This is not a speed optimisation with a side effect. Deleting frequent tokens effectively widens the window over the surviving tokens, because removed words no longer occupy window slots. Reported gains from subsampling are substantial, especially for rare words.
Dynamic window size. The released C code does not use a fixed window c. For each token it draws c' ~ Uniform{1, ..., c} and uses that. The effect is a distance-dependent weighting: a word at offset d is included with probability (c - d + 1) / c. Papers that report "window = 5" and implement a fixed window are not reproducing the original method.
Cost
Per training pair, with D the embedding dimension and k the number of negatives:
| Variant | Time per pair | Parameters |
|---|---|---|
| Full softmax | O(V D) | 2 V D |
| Hierarchical softmax | O(D log V) | 2 V D |
| Negative sampling | O(k D) | 2 V D |
With k = 5 and D = 300, negative sampling costs roughly 1800 multiply-accumulates per pair. Memory is dominated by the two V x D tables: at V = 10^6 and D = 300 in float32, that is 2.4 GB. Training is embarrassingly parallel under Hogwild-style lock-free asynchronous SGD (Recht et al., 2011), which the original C implementation uses.
Hyperparameters carry more weight than architecture
Levy, Goldberg & Dagan (2015), Improving Distributional Similarity with Lessons Learned from Word Embeddings (TACL), is the most useful paper in this area and the least cited relative to its value.
They transplanted word2vec's implementation choices — subsampling, dynamic windows, context distribution smoothing, adding context vectors to word vectors at the end — into PPMI and SVD baselines. Most of the reported advantage of word2vec over count-based methods disappeared.
The practical conclusion is uncomfortable and durable. Published comparisons between embedding methods frequently compare hyperparameter tuning budgets rather than methods.
Two specific settings from that work are worth applying:
- Context distribution smoothing with the same 0.75 exponent improves PPMI as much as it improves SGNS.
w + c, summing the input and output vectors instead of discarding the output vectors, helps on similarity tasks at no cost. Most implementations discardWoby default.
Where the line went after this
GloVe (Pennington et al., 2014) and fastText (Bojanowski et al., 2017) are covered in embeddings. fastText's subword decomposition matters most for the Indic languages, where a single lemma surfaces in many inflected forms and a type-level vocabulary fragments badly.
Beyond words. The skip-gram objective is indifferent to what a "token" is. Any sequence of discrete items works.
- item2vec (Barkan & Koenigstein, 2016) — purchase sessions as sentences.
- node2vec (Grover & Leskovec, 2016) — biased random walks on a graph as sentences, interpolating between breadth-first and depth-first exploration with two parameters
pandq. - DeepWalk (Perozzi et al., 2014) — uniform random walks, the precursor.
These remain live techniques. For a recommendation system over ids with no pretrained model available, an SGNS objective over interaction sequences is still a strong and very cheap baseline.
Key references
- Mikolov, T., Chen, K., Corrado, G. & Dean, J. (2013a). Efficient Estimation of Word Representations in Vector Space. arXiv:1301.3781
- Mikolov, T., Sutskever, I., Chen, K., Corrado, G. & Dean, J. (2013b). Distributed Representations of Words and Phrases and their Compositionality. arXiv:1310.4546
- Goldberg, Y. & Levy, O. (2014). word2vec Explained: Deriving Mikolov et al.'s Negative-Sampling Word-Embedding Method. arXiv:1402.3722
- Levy, O., Goldberg, Y. & Dagan, I. (2015). Improving Distributional Similarity with Lessons Learned from Word Embeddings. TACL 3.
- Gutmann, M. & Hyvärinen, A. (2012). Noise-Contrastive Estimation of Unnormalized Statistical Models. JMLR 13.
- Recht, B., Ré, C., Wright, S. & Niu, F. (2011). Hogwild!: A Lock-Free Approach to Parallelizing Stochastic Gradient Descent. NeurIPS 24.
- Grover, A. & Leskovec, J. (2016). node2vec: Scalable Feature Learning for Networks. arXiv:1607.00653
Honest assessment in 2026
Word2Vec is no longer a competitive text representation. Contextual encoders beat it on every language-understanding benchmark, and contrastively trained sentence encoders beat it on every retrieval benchmark, usually by a wide margin.
Its remaining value is threefold and genuine. It is the clearest existing worked example of a self-supervised pretext task. Its implicit-factorisation result connects the neural and count-based literatures in a way no later model does as cleanly. And the SGNS objective is still the default first attempt when the tokens are arbitrary ids rather than language.
The open question it left behind is unresolved: nobody has a principled account of why the 0.75 exponent, dynamic windows and subsampling help as much as they do. They are load-bearing empirical choices sitting underneath a decade of downstream work.
What to learn next
- Attention — the mechanism that replaced static vectors.
- BERT — masked pretraining and contextual representations.
- Embeddings — the representation view, count-based methods and modern encoders.