Word2Vec
In one sentence Word2Vec was the 2013 breakthrough that learned word embeddings from raw text, famously making king − man + woman land near queen.
Updated
Word2Vec is the classic technique that learns a vector for every word purely from which words appear near it in text.
Its premise is a proverb made algorithm: a person is known by the company they keep. "Chai" and "coffee" occur amid the same companions — cup, morning, hot, sugar — so any method scoring words by their neighbours will place them close together. Word2Vec (Mikolov et al., Google, 2013) trained a small network on billions of such neighbourhoods: given a word, predict its surrounding words (or the reverse). The prediction task is a pretext; the trained-up word vectors are the product — an early, pure case of self-supervised-learning.
What made it famous was the geometry that emerged unbidden. Directions in the learned latent-space turned out to carry meaning, supporting arithmetic on words:
king − man + woman ≈ queen
Delhi − India + Japan ≈ TokyoNobody programmed analogies; they precipitated out of co-occurrence statistics. This result convinced the field that meaning could genuinely live in vectors, opening the road to modern embeddings.
Its historical limitation defined the next chapter: one vector per word, fixed forever, so "bank" carries river and money in a single blend. Contextual models — BERT and its successors — compute a fresh vector for each occurrence based on the sentence around it, and superseded Word2Vec for nearly all production use. It remains worth knowing as the conceptual ancestor, a fast lightweight tool, and the cleanest possible introduction to what an embedding is.
Where to go next
- Full lesson: Embeddings
- Related terms: embedding, latent-space, bert, self-supervised-learning