FastText and subword vectors
FastText builds a word's vector from the vectors of its smaller pieces, so it can still make a sensible guess about a word it has never seen before.
- 10 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
FastText builds a word's vector from the vectors of its smaller pieces. That way, it can still handle a word it has never seen.
Think about guessing what an unfamiliar English word means, from its pieces. You have never heard "unfriendliness" before. But you can break it down: "un-", "friend", "-li-", "-ness". Each piece is familiar, even though the whole word is new to you.
FastText gives a computer that same trick. GloVe and Word2Vec learn one vector per whole word. FastText instead learns vectors for small pieces of words, and builds a word's vector by combining its pieces.
Why it exists
GloVe and Word2Vec share one sharp limitation. If a word never appeared in training, they have no vector for it at all. A misspelling, a slang term, a person's name — all become a dead end. So does an unfamiliar word-ending from a language rich in them.
This hits some languages far harder than others. English mostly builds meaning by adding whole new words. Tamil, Turkish and Finnish often glue extra pieces onto a word to change its meaning. That produces far more distinct word-forms from the same root. A whole-word-only model has to see every one of those forms separately, to learn anything about them at all.
FastText — built at Facebook AI Research and published in 2017 — fixes this by working one level below the whole word.
How it works
Break every word into overlapping chunks of letters, called character n-grams, plus the whole word itself.
Word: "chai"
Character chunks (n-grams of length 3, with boundary markers):
<ch cha hai ai>
plus the whole word: <chai>Each chunk gets its own vector, learned the same way Word2Vec learns word vectors — by predicting nearby words during training. A word's final vector is the sum of the vectors of all its chunks.
Now, when FastText meets a brand-new word — a misspelling, a rare name — it breaks that word into chunks too. Most of those chunks it has seen before, inside other words. It builds a reasonable vector on the spot, from familiar pieces alone.
Where you have already seen it
- Spell-tolerant search. Search boxes that still find the right result even when you type "recieve" instead of "receive".
- Facebook's own products, where FastText was first deployed at large scale for language identification and text classification.
- NLP tools for morphologically rich languages — Turkish, Finnish, Tamil, Hindi — where FastText consistently outperforms whole-word-only methods.
Remember this
- FastText builds a word's vector from vectors of its smaller letter-chunks, not one fixed vector per whole word.
- This lets it produce a sensible vector for a word it has never seen, unlike GloVe or plain Word2Vec.
- It helps most for languages that build many word-forms from one root. It also helps with typos and rare words in any language.
What to learn next
- GloVe — the whole-word approach FastText extends with subword pieces.
- Word2Vec — the prediction game FastText's chunk-training is directly built on.
- Tokenization — a related but different way of breaking words into pieces, used for a different purpose.
Developer — Code and libraries.
Training FastText on a tiny made-up corpus, from scratch, shows the out-of-vocabulary trick directly — no pretrained download needed.
Setup
pip install gensimTraining FastText and testing an unseen word
from gensim.models import FastText
sentences = [
"the chai was hot and sweet".split(),
"i drink chai every morning".split(),
"the coffee was hot and bitter".split(),
"i drink coffee every morning".split(),
"the tea was warm and sweet".split(),
"i drink tea every evening".split(),
"chai and coffee are both drinks".split(),
"the milk was warm".split(),
] * 30 # repeated so the tiny vocabulary gets enough training signal
model = FastText(sentences, vector_size=20, window=3, min_count=1,
sg=1, epochs=20, seed=42, workers=1)
print("'chai' is in the trained vocabulary:", "chai" in model.wv.key_to_index)
print("'chaii' (a misspelling) is in the trained vocabulary:", "chaii" in model.wv.key_to_index)
print()
print("vector for the never-seen word 'chaii':", model.wv["chaii"].shape)
print()
print("words most similar to the misspelling 'chaii':")
for word, score in model.wv.most_similar("chaii", topn=3):
print(f" {score:.3f} {word}")'chai' is in the trained vocabulary: True 'chaii' (a misspelling) is in the trained vocabulary: False vector for the never-seen word 'chaii': (20,) words most similar to the misspelling 'chaii': 0.988 chai 0.986 was 0.984 hot
The model was never trained on the word "chaii" — it is not in the vocabulary at all. Yet asking for its vector works without error, and its closest neighbour is correctly "chai". This is the whole mechanism, working exactly as intended, on a tiny corpus with no download required.
workers=1 and seed=42 are set here so this exact run reproduces. On different hardware or a different gensim version, the precise similarity scores may shift slightly — the ranking, with "chai" first, is the robust part.
Line by line
sg=1 selects skip-gram mode, the same prediction objective Word2Vec uses — FastText is best understood as Word2Vec's skip-gram training applied to character n-grams instead of whole words.
"chaii" in model.wv.key_to_index is False, confirming this word never appeared during training. model.wv["chaii"] still returns a vector, because gensim automatically breaks it into character n-grams and sums their learned vectors — none of that logic is visible in this script, it happens inside the library.
sentences * 30 repeats the tiny corpus 30 times. Real FastText training uses corpora with billions of words; a handful of made-up sentences need repetition to give the model any real signal to learn from.
Common mistakes
Assuming FastText handles every unseen word equally well. A misspelling of a known word ("chaii" from "chai") shares almost all its character chunks with the real word, so it works well. A genuinely novel word built from unfamiliar chunks — a name in an unfamiliar script, for instance — gets a much weaker, less reliable vector.
Setting the n-gram range too narrow or too wide. gensim's min_n and max_n parameters (default 3 to 6) control chunk size. Too narrow, and chunks stop carrying much meaning on their own. Too wide, and rare words rarely share chunks with anything else, defeating the purpose.
Forgetting that subword vectors add real training and memory cost. Storing and training vectors for every character chunk, not only every word, uses meaningfully more memory than GloVe or plain Word2Vec on the same vocabulary.
Try it yourself
Change "chaii" to "xyzxyz" — a string sharing no real chunks with anything in the tiny training corpus — and check its nearest neighbours. Expect the results to look far less meaningful than the "chaii" example, since there is little genuine shared structure for the model to fall back on.
What to learn next
- Word2Vec — the prediction objective FastText's training is built directly on top of.
- Detecting the language of a text — a task FastText's own classifier variant was originally built for.
- King minus man plus woman — testing whether FastText vectors support the same analogy trick as GloVe.
Researcher — Mathematics and papers.
The model
FastText (Bojanowski, Grave, Joulin & Mikolov, 2017) extends the skip-gram with negative sampling (SGNS) objective from Word2Vec by representing a word as a sum over its character n-gram embeddings:
v_w = sum over g in G_w of z_gG_wis the set of character n-grams for wordw, typically length 3 to 6, plus a special whole-word n-gram (so a word retains a distinct identity beyond the sum of its parts).z_gis the learned embedding for n-gramg.- Word boundary markers (
<and>) are added before extraction, so that, for example, the trigram<in(word-initial "in") is distinguished fromin<or a mid-wordin.
The scoring function for a target-context word pair (w, c) becomes:
s(w, c) = sum over g in G_w of z_g . v_cv_cis the context word's vector, learned the standard SGNS way (see Word2Vec for the negative-sampling objective this plugs into unchanged).
Training otherwise follows SGNS exactly: negative sampling, subsampling of frequent words, and a sliding context window, all inherited from Mikolov et al. (2013).
Why this specifically fixes OOV handling
Any word, seen or unseen, can be decomposed into character n-grams using a fixed, deterministic rule — no lookup table is required for the decomposition step itself. An out-of-vocabulary word's vector is therefore always computable, as the sum of whichever of its n-gram embeddings were learned during training (unlearned n-grams, novel to this word too, are absent from the sum entirely). This degrades gracefully with how unfamiliar the word's substrings are, rather than failing outright.
Effect on morphologically rich languages
Bojanowski et al.'s original evaluation showed the largest gains over plain Word2Vec on languages with productive morphology — Czech, German, Russian, Finnish — measured via word similarity and analogy benchmarks in each language, precisely because these languages generate far more distinct word-forms per lemma than English, and whole-word models see each form as an entirely separate, disconnected token.
Complexity
Training cost per token is O(|G_w| * d) for embedding dimension d, versus O(d) for plain Word2Vec — a constant-factor slowdown proportional to average n-grams per word (in practice, a handful). Storage is O(|hash table size| * d), since gensim and the original implementation hash n-grams into a fixed-size bucket table (default 2 million buckets) rather than storing one embedding per unique n-gram, trading a small amount of hash-collision noise for bounded memory regardless of corpus size.
Relationship to later subword methods
FastText's n-gram decomposition is a fixed, character-level, non-learned segmentation, distinct from the learned subword vocabularies used by BPE and WordPiece tokenizers in transformer models. Where BPE learns a vocabulary of variable-length merged units used for the model's actual input tokens, FastText's n-grams are a fixed-scheme feature extractor sitting underneath a still-word-level, non-contextual embedding. The two ideas solve a related problem — handling the long tail of rare and unseen word-forms — through different mechanisms.
Key references
- Mikolov, T. et al. (2013). Efficient Estimation of Word Representations in Vector Space. arXiv:1301.3781. The SGNS objective FastText extends.
- Bojanowski, P., Grave, E., Joulin, A. & Mikolov, T. (2017). Enriching Word Vectors with Subword Information. TACL 5, 135–146. arXiv:1607.04606
- Joulin, A., Grave, E., Bojanowski, P. & Mikolov, T. (2016). Bag of Tricks for Efficient Text Classification. arXiv:1607.01759. A sibling paper applying the same team's fast training techniques to supervised text classification, distinct from the word-vector paper above but often confused with it.
Current state and open problems
FastText's static, non-contextual vectors carry the same core limitation as GloVe once a downstream task needs word-sense disambiguation (see static vs contextual embeddings) — subword composition fixes coverage, not polysemy. Its practical niche today is where transformer-based contextual embeddings are too expensive to run: extremely low-resource languages with little training text, edge and on-device deployments, and any pipeline needing microsecond-latency lookups over millions of words. Meta's continued release of pretrained FastText vectors for 157 languages remains one of the more complete openly available multilingual static-embedding resources, though quality varies significantly by how much training text existed for each language — a gap that mirrors the tokenizer fairness problem seen elsewhere in NLP.
What to learn next
- Word2Vec — the base training objective FastText builds subword composition on top of.
- Detecting the language of a text — a task these per-language vectors were built to support.
- BERT — the contextual, subword-tokenized approach that superseded static embeddings for most tasks.