The Embedding Family

GloVe

GloVe learns a number list for every word by counting, across an entire corpus, how often word pairs appear near each other, then compressing those counts into vectors.

Read these first

On this page 5
  1. Why it exists
  2. How it works
  3. Where you have already seen it
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

GloVe learns a number list for every word. It counts how often word pairs show up near each other, across a huge amount of text.

Think about a city's traffic patterns. Watch enough cars over enough months, and you can tell which neighbourhoods are close-knit. Lots of people drive between them. Others sit barely connected, even on the same map.

GloVe treats words the same way. It watches which words keep turning up near which other words, across millions of sentences. That traffic pattern becomes a map. Words that keep appearing together end up sitting close on it.

Why it exists

Word2Vec learns word vectors by playing a guessing game, one small window of text at a time. It never looks at the whole corpus at once. Only a few words to either side, over and over.

That works well, but it throws away something useful. Two words counted across the entire corpus tell you more than a small window ever sees. Two words might rarely sit next to each other, yet still be related across the whole collection of text.

GloVe — short for Global Vectors — uses that global count directly. It never has to peek through a small window alone.

How it works

First, count. For every pair of words, count how many times they appear near each other, anywhere in the whole corpus. This produces one giant table of counts.

   Counting across millions of sentences:

     "ice"  near "cold"    -> counted a huge number of times
     "ice"  near "steam"   -> counted rarely
     "steam" near "cold"   -> counted rarely
     "steam" near "hot"    -> counted a huge number of times

Then, compress. That giant table is too large and too sparse to use directly. So GloVe learns a short vector for each word instead. It builds those vectors so their relationship matches the ratio of the words' co-occurrence counts. Words whose counts behave alike end up with vectors that behave alike too.

The result looks similar to Word2Vec's output, one dense vector per word. But it was learned from global counts, not local windows alone.

Where you have already seen it

  • Early spell-check suggestions and autocomplete, before today's larger models, often used GloVe vectors under the hood.
  • Search relevance tools built before 2018, frequently comparing word meanings with GloVe.
  • University NLP courses, almost universally, teaching GloVe as the standard first pretrained embedding.

Remember this

  • GloVe counts how often word pairs appear together across an entire corpus, then compresses those counts into vectors.
  • It captures relationships Word2Vec's small local windows can miss, by using global statistics.
  • GloVe and Word2Vec are close cousins, both giving you one fixed vector per word. They get there by different routes.

What to learn next

Developer — Code and libraries.

Training GloVe from scratch needs a large corpus and real compute. Loading a small pretrained GloVe model, on the other hand, takes seconds and needs nothing but a CPU.

Setup

bash
pip install gensim

Loading real GloVe vectors

glove_load.py
import gensim.downloader as api

# downloads once (~66MB) and caches locally — this is the smallest
# standard pretrained GloVe model, 400,000 words, 50 numbers per word
model = api.load("glove-wiki-gigaword-50")

print("vocabulary size:", len(model))
print("vector shape:", model["king"].shape)
print()
print("closest words to 'king':")
for word, score in model.most_similar("king", topn=5):
    print(f"  {score:.3f}  {word}")
Output
vocabulary size: 400000
vector shape: (50,)

closest words to 'king':
  0.824  prince
  0.784  queen
  0.775  ii
  0.774  emperor
  0.767  son

The first run downloads about 66MB and caches it — later runs load from disk in a few seconds. This is a real, published model trained on Wikipedia and newswire text, not a toy.

Line by line

api.load("glove-wiki-gigaword-50") returns a KeyedVectors object — essentially a dictionary from word to 50-number vector, plus fast similarity search built in.

most_similar ranks every word in the vocabulary by cosine similarity to "king", and returns the top few. "prince", "queen" and "emperor" showing up is a direct result of these words appearing in similar surroundings across the training corpus — not because the model knows anything about royalty.

A limitation, shown honestly

glove_bank.py
import gensim.downloader as api

model = api.load("glove-wiki-gigaword-50")
print("closest words to 'bank':")
for word, score in model.most_similar("bank", topn=8):
    print(f"  {score:.3f}  {word}")
Output
closest words to 'bank':
  0.870  banks
  0.800  securities
  0.797  banking
  0.785  investment
  0.781  exchange
  0.767  financial
  0.765  credit
  0.752  lender

Every single neighbour is about money. Not one is about a river. "Bank" has exactly one vector in GloVe, and that vector was pulled almost entirely toward the financial sense, because that sense is far more common in the training text. The riverbank sense is not blended in — it is effectively lost. See static vs contextual embeddings for the fix.

Common mistakes

Assuming a bigger GloVe model is always better for your task. Larger GloVe models (up to 300 dimensions, trained on 840 billion tokens) capture finer distinctions, but cost more memory and load time. For prototyping and coursework, the 50-dimensional model shown here is usually enough.

Comparing GloVe similarity scores across different pretrained models. A cosine similarity of 0.8 in one GloVe model is not directly comparable to 0.8 in another — different training corpora and dimensions produce differently-scaled spaces.

Expecting GloVe to handle a word it never saw during training. Ask this exact model for a misspelling, a brand-new slang word, or most non-English words, and it raises a KeyError. This is GloVe's core limitation, and it is precisely what FastText was built to fix.

Try it yourself

Run model.most_similar("apple", topn=8) and look at the results. Because GloVe gives every word exactly one vector, "apple" the fruit and Apple the company are blended into a single meaning — check whether the neighbours lean toward fruit, technology, or an odd mix of both.

What to learn next

Researcher — Mathematics and papers.

The objective function

GloVe (Pennington, Socher & Manning, 2014) fits word vectors to directly model the log of global co-occurrence counts:

text
J = sum over i,j in V of  f(X_ij) * (w_i . w~_j + b_i + b~_j - log(X_ij))^2
  • X_ij is the number of times word j occurs in the context of word i, counted across the entire corpus.
  • w_i is the target-word vector for word i; w~_j is a separate context-word vector for word j. Two vector sets are learned per word; the final representation is usually w_i + w~_i.
  • b_i, b~_j are learned scalar bias terms.
  • f(X_ij) is a weighting function that down-weights very rare co-occurrences (noisy) and caps the weight of very frequent ones (so extremely common pairs like "the" and "of" do not dominate the loss).

The weighting function used in the paper:

text
f(x) = (x / x_max)^alpha   if x < x_max
f(x) = 1                   otherwise
  • x_max = 100 and alpha = 3/4 in the original paper, chosen empirically.

Why log co-occurrence, specifically

The central design argument in the paper is that ratios of co-occurrence probabilities encode meaning better than raw probabilities. For probe word k, the ratio P(k|ice) / P(k|steam) is large when k relates to ice specifically (e.g. "solid"), small when k relates to steam specifically (e.g. "gas"), and near 1 when k relates to both or neither (e.g. "water", "fashion"). Fitting w_i . w~_j to log(X_ij) makes vector dot-product differences correspond to log-ratios, which is exactly the quantity the paper argues carries the meaningful signal.

Relationship to matrix factorization

GloVe can be read as a weighted, log-transformed, biased factorization of the word-word co-occurrence matrix — closely related to older count-based methods like truncated SVD / LSA on a PPMI-weighted co-occurrence matrix (Levy & Goldberg, 2014, showed Word2Vec's skip-gram with negative sampling is implicitly factorizing a PMI-shifted matrix too). GloVe, Word2Vec and PPMI-SVD are, in this sense, three different optimization routes to a related underlying object: a low-rank representation of word co-occurrence statistics.

Complexity

Building the co-occurrence matrix is O(corpus size * window size), once. Storage of the matrix is O(min(V^2, number of non-zero pairs)) — non-zero pairs dominate for real corpora, since most word pairs never co-occur. Training the factorization is O(nnz * iterations), where nnz is the number of non-zero co-occurrence entries, typically far below V^2. This global-count-then-factorize structure is why GloVe training, unlike Word2Vec, requires building and holding a large sparse matrix before optimization begins.

Known limitations

GloVe assigns exactly one vector per word type, regardless of context — the same structural limitation as Word2Vec, demonstrated concretely in the developer block's "bank" example. It also has no mechanism for out-of-vocabulary words: any word absent from the training corpus's vocabulary has no vector at all, a gap FastText addresses through subword composition.

Key references

  • Pennington, J., Socher, R. & Manning, C. D. (2014). GloVe: Global Vectors for Word Representation. EMNLP.
  • Levy, O. & Goldberg, Y. (2014). Neural Word Embedding as Implicit Matrix Factorization. NeurIPS. Shows Word2Vec SGNS is implicitly factorizing a shifted PMI matrix, connecting it formally to GloVe's approach.
  • Levy, O., Goldberg, Y. & Dagan, I. (2015). Improving Distributional Similarity with Lessons Learned from Word Embeddings. TACL. Shows much of the reported quality gap between different embedding methods comes from hyperparameter choices, not the underlying objective.

Current state and open problems

Static word vectors like GloVe and Word2Vec have been superseded, on nearly every benchmark that measures contextual understanding, by embeddings derived from transformer encoders (see static vs contextual embeddings). GloVe remains in active use for exactly the settings where its limitations do not matter: extremely low-latency, CPU-only pipelines, interpretable feature engineering, and as a well-understood pedagogical baseline. Its co-occurrence-counting idea also persists inside modern systems in disguised form — several efficient retrieval and recommendation systems still build explicit co-occurrence statistics as a first-stage signal, precisely because they are cheap to compute and easy to audit, properties dense transformer embeddings do not share.

What to learn next