Classical NLP That Still Works
TF-IDF
TF-IDF scores a word higher when it is frequent in one document but rare across all others, so common words stop drowning out the words that actually matter.
- 9 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
TF-IDF gives a word a high score when it shows up a lot in one document, but rarely anywhere else.
Think about a classroom where the teacher says "okay" every few minutes. It means almost nothing — everyone says it, in every class. But if one student keeps saying "photosynthesis", that word tells you something specific about what that student is talking about.
TF-IDF is a way of teaching a computer to notice "photosynthesis" and ignore "okay". Common-everywhere words get pushed down. Rare-but-repeated words get pushed up.
Why it exists
Bag of words counts every word equally. That is its biggest weakness. Words like "the", "was" and "is" show up constantly and dominate the count. Yet they carry almost no information about what a document is actually about.
Meanwhile, a word like "chai", "monsoon" or "GST" might appear only two or three times. Yet that word is what the document is actually about. It gets buried under dozens of "the"s.
TF-IDF was invented to fix exactly this. It multiplies two scores together, and each one does a different job.
How it works
TF (how often in THIS document)
combined with
IDF (how rare across ALL documents)
gives you
the TF-IDF scoreTF, term frequency: how many times the word appears in this one document. Higher means the document is talking about this word a lot.
IDF, inverse document frequency: how rare the word is across your whole collection of documents. A word that appears in every single document — like "the" — gets a very low IDF, close to zero. A word that appears in only a few documents gets a high IDF.
Word "the": shows up in almost every document -> IDF near zero
Word "chai": shows up in only a handful of them -> IDF is high
In one document that repeats "chai" several times, and "the" even more:
"chai" -> frequent here, AND rare everywhere else
-> both scores are high -> TF-IDF gives it a big score
"the" -> frequent here, but common everywhere else too
-> the second score cancels the first out -> TF-IDF gives it a tiny scoreA word only scores high when both numbers are high at once: it must be frequent here and rare elsewhere. Frequent everywhere gets cancelled out. Rare everywhere but also rare here gets no boost either.
Where you have already seen it
- Search engines, for decades. Matching your query to the documents where those exact words are distinctively frequent, before neural search existed.
- "Find similar articles." News sites comparing TF-IDF vectors to recommend related stories.
- Highlighting important words in a document. Tools that auto-highlight "key terms" are often doing nothing more than ranking words by TF-IDF.
Remember this
- TF-IDF rewards a word that is frequent in one document but rare across the whole collection.
- It fixes bag of words' biggest flaw: common filler words no longer dominate.
- It still throws away word order, exactly like bag of words does.
What to learn next
- BM25 — the search-ranking formula built directly on top of this idea.
- N-gram language models — another classical way of scoring text, built on counting instead.
- Keyword extraction — using this exact scoring to pull out a document's most important words.
Developer — Code and libraries.
One program is enough to see the whole idea. Compare it against plain counts to see exactly what changes.
Setup
pip install scikit-learnTF-IDF with scikit-learn
from sklearn.feature_extraction.text import TfidfVectorizer
import numpy as np
docs = [
"the chai was hot",
"the coffee was hot",
"the chai was sweet",
]
vec = TfidfVectorizer()
X = vec.fit_transform(docs)
terms = vec.get_feature_names_out()
np.set_printoptions(precision=3, suppress=True)
print("vocabulary:", terms)
print()
print("tf-idf matrix:")
print(X.toarray())
print()
row0 = dict(zip(terms, X.toarray()[0]))
ranked = sorted(row0.items(), key=lambda kv: -kv[1])
print("doc 0 terms ranked by weight:", ranked)vocabulary: ['chai' 'coffee' 'hot' 'sweet' 'the' 'was']
tf-idf matrix:
[[0.558 0. 0.558 0. 0.434 0.434]
[0. 0.663 0.504 0. 0.391 0.391]
[0.504 0. 0. 0.663 0.391 0.391]]
doc 0 terms ranked by weight: [('chai', 0.5584778353707552), ('hot', 0.5584778353707552), ('the', 0.4337078595086741), ('was', 0.4337078595086741), ('coffee', 0.0), ('sweet', 0.0)]Line by line
"the" and "was" score 0.434 in doc 0 — lower than "chai" and "hot" at 0.558, even though every word here appears exactly once. "the" and "was" appear in all three documents, so their IDF is lower. "chai" and "hot" appear in only two of three, so they score higher. This is the entire mechanism, visible in four numbers.
Row 0 sums to more than 1, and rows are not raw TF times raw IDF. scikit-learn L2-normalizes every row by default, so each document's vector has length 1. This makes documents of different lengths comparable by cosine similarity, without long documents automatically scoring higher purely from having more words.
scikit-learn's exact IDF formula is smoothed: idf(t) = ln((1 + n) / (1 + df(t))) + 1, where n is the number of documents and df(t) is how many documents contain term t. The +1s inside the logarithm stop any word from ever getting a zero or undefined score, even one appearing in every document.
Common mistakes
Computing IDF on too small or too skewed a document collection. IDF only means something relative to a collection. Score one document's TF-IDF using a 3-document vocabulary and a 3-million-document vocabulary, and you get very different numbers for the same word. Always fit the vectorizer on a corpus that represents what "common" and "rare" should mean for your task.
Refitting the vectorizer on new data at prediction time. As with bag of words, fit_transform learns document frequencies from whatever you pass it. A word can flip from "rare and important" to "common and unimportant" purely because your new batch of documents talks about it constantly. Fit once on training data, then always call .transform(), never .fit_transform() again, on new text.
Assuming TF-IDF understands meaning. "Bank" scores the same whether it means a riverbank or a savings account. TF-IDF is entirely about frequency statistics, not word meaning — see static vs contextual embeddings for representations that do capture this.
Try it yourself
Add a fourth document, "the movie was great", to the docs list and re-run. "the" and "was" now appear in all four documents — watch their weight drop further relative to "chai", "hot", "coffee" and "sweet", none of which are shared by this new unrelated sentence.
What to learn next
- BM25 — TF-IDF's term-frequency part, refined with saturation and length normalization.
- Truncated SVD and LSA — compressing a TF-IDF matrix down to catch word relationships it misses on its own.
- Text classification baselines — TF-IDF vectors feeding directly into a classifier.
Researcher — Mathematics and papers.
The formula
tf-idf(t, d) = tf(t, d) * idf(t)tis a term,da document.tf(t, d)is the frequency oftind. Raw count is the simplest choice; log-scaled1 + log(count)and other variants are common, since raw count grows the weight of a repeated word faster than its actual informativeness usually does.idf(t)measures corpus-wide rarity oft.
The original formulation (Spärck Jones, 1972):
idf(t) = log(N / df(t))Nis total document count.df(t)is document frequency: the number of documents containingtat least once.
scikit-learn's default, smoothed variant:
idf(t) = ln((1 + N) / (1 + df(t))) + 1The +1 terms prevent division by zero when df(t) = 0 (unseen term) and keep a term appearing in every document from receiving exactly zero weight. TfidfVectorizer then L2-normalizes each document row, so sum_t tfidf(t, d)^2 = 1.
Why L2 normalization, specifically
Cosine similarity between two L2-normalized vectors reduces to their dot product: cos(d_1, d_2) = d_1 . d_2 when ||d_1|| = ||d_2|| = 1. Pre-normalizing at index-build time turns every similarity query into a single dot product rather than a normalize-then-dot-product at query time, which matters at retrieval scale.
Relationship to information theory
log(N / df(t)) is, up to the base of the logarithm, the number of bits of information contained in learning that document d contains term t, under a model where a term's prior probability of occurring in a random document is df(t) / N. A term that occurs in half the corpus contributes exactly 1 bit; a term occurring in one document out of a million contributes roughly 20 bits. IDF is, in this sense, a self-information weighting, not an arbitrary heuristic.
Known weaknesses, precisely stated
TF-IDF has no notion of term saturation: two occurrences of a word are worth exactly twice one occurrence, indefinitely, with no diminishing returns. It has no document-length normalization for the TF component itself — the L2 step normalizes the whole vector, not tf in isolation, which is a coarser fix. BM25 directly repairs both of these, which is why it displaced raw TF-IDF as the standard IR scoring function.
TF-IDF also treats every term as independent (a bag-of-words assumption applied to weighting, not only counting), so it cannot represent that "New" and "York" co-occurring means something "New" and "York" separately do not. n-gram extensions and, later, learned dense representations were the two lines of work that addressed this from different directions.
Key references
- Luhn, H. P. (1957). A Statistical Approach to Mechanized Encoding and Searching of Literary Information. IBM Journal of Research and Development. The term-frequency half.
- Spärck Jones, K. (1972). A Statistical Interpretation of Term Specificity and Its Application in Retrieval. Journal of Documentation 28(1), 11–21. The inverse-document-frequency half.
- Salton, G. & Buckley, C. (1988). Term-weighting Approaches in Automatic Text Retrieval. Information Processing & Management 24(5), 513–523. Systematic comparison of TF-IDF weighting schemes.
- Robertson, S. (2004). Understanding Inverse Document Frequency: On Theoretical Arguments for IDF. Journal of Documentation 60(5), 503–520. The information-theoretic reading of IDF.
Current state and open problems
TF-IDF is superseded by BM25 for ranking and by dense embeddings for semantic matching, but it remains in active use as a cheap, interpretable feature for classification baselines, as the sparse half of hybrid sparse-plus-dense retrieval systems, and as the scoring function inside every default configuration of Lucene-family search engines (Elasticsearch, Solr, OpenSearch) prior to explicit BM25 configuration. Its main remaining research relevance is as the interpretable baseline that any new retrieval or classification method is expected to clear before being taken seriously.
What to learn next
- BM25 — the direct successor, fixing TF-IDF's saturation and length-normalization gaps.
- Sentence-transformers — dense retrieval that captures meaning TF-IDF cannot.
- Truncated SVD and LSA — factorizing the TF-IDF matrix to recover latent topic structure.