Classical NLP That Still Works

The baseline you must beat before reaching for BERT

A TF-IDF vector plus a linear classifier trains in seconds and often gets most of the way to what a large language model would score, so it earns the right to be tried first.

Read these first

On this page 5
  1. Why it exists
  2. How it works
  3. Where you have already seen it
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Before reaching for a big AI model, build the simplest classifier that could work, and see how far it gets.

Think about fixing a leaking tap. Most people try tightening the washer first — five minutes, one tool. Nobody starts by replacing the entire plumbing system, in case a small fix was enough.

Text classification has the same discipline. A simple model — counting words and drawing a line between categories — often comes surprisingly close. A huge, expensive model may only beat it by a little. Skipping straight to the huge model means you never find that out.

Why it exists

Deep learning models like BERT are genuinely more capable on hard problems. But they cost more to train, more to run, and take longer to build and deploy. Think of everyday tasks: is this email spam, is this review positive, is this ticket urgent. The gap turns out to be small on tasks like these. A cheap baseline and an expensive model often land close together.

Without a baseline, you cannot know that gap. A team might spend weeks fine-tuning a large model and ship it. They may never learn that a linear classifier, trained in ten seconds, would have scored almost the same. The baseline is not a lesser first draft. It is the measuring stick every fancier idea has to justify itself against.

How it works

The recipe is short. Turn each piece of text into numbers using TF-IDF. Then hand those numbers to a simple classifier that draws a boundary between categories.

   "brilliant film, great acting"  ->  TF-IDF vector  ->  [ Linear classifier ] -> positive
   "terrible film, awful acting"   ->  TF-IDF vector  ->  [ Linear classifier ] -> negative

The classifier is not learning to "read" in any deep sense. It is learning that certain words push a decision one way, and other words push it the other way. "Brilliant" and "awful" become strong signals almost immediately, because they show up disproportionately in one category or the other.

That is a shallow trick. It is also astonishingly effective for a huge range of everyday classification problems.

Where you have already seen it

  • Early spam filters. "Free", "winner", "click now" pushed strongly toward spam, learned from counts alone.
  • Star-rating prediction from review text, before recommendation systems got more sophisticated.
  • Support ticket routing. Words like "refund" or "broken" reliably point toward the right team, without needing deep language understanding.

Remember this

  • Always try a simple, fast baseline before reaching for a large model. It tells you exactly what the large model needs to beat.
  • TF-IDF plus a linear classifier is the standard baseline for text classification, and it trains in seconds.
  • If the baseline already scores 80%+, a bigger model may only add a few points. And it costs a lot more to run.

What to learn next

  • TF-IDF — the feature representation this baseline is built on.
  • Text classification — the fuller picture, including where BERT-style models genuinely pull ahead.
  • BERT — what you reach for once the baseline shows you need more.

Developer — Code and libraries.

One program: a full baseline pipeline, trained and evaluated on a small, honest, labelled dataset, with real mistakes shown so you can see exactly where a bag-of-words model runs out of understanding.

Setup

bash
pip install scikit-learn

A complete TF-IDF + logistic regression baseline

baseline.py
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import train_test_split
from sklearn.metrics import accuracy_score

positive = [
    "this movie was fantastic and moving",
    "i loved every minute of this film",
    "best purchase i have made all year",
    "the acting was superb throughout",
    "a delightful and touching story",
    "brilliant direction and great music",
    "a charming and funny little movie",
    "the cast did an amazing job",
    "wonderful cinematography and pacing",
    "a joy to watch from start to end",
    "genuinely great film, would watch again",
    "superb performance from the whole cast",
    "a beautiful and moving experience",
    "excellent movie, highly recommend it",
    "fantastic story with a great ending",
    "loved the music and the acting both",
    "brilliant film, one of the best this year",
    "a wonderful and delightful watch",
    "amazing visuals and a moving story",
    "great film, superb acting throughout",
]
negative = [
    "this was a complete waste of time",
    "i hated the plot and the acting",
    "worst film i have seen this year",
    "boring, slow and far too long",
    "terrible dialogue and flat characters",
    "i want my money back after this",
    "dull, predictable and forgettable",
    "i regret watching this film",
    "awful acting ruined the story",
    "i would not recommend this to anyone",
    "boring plot and terrible acting",
    "a complete waste of money",
    "dull and forgettable from start to end",
    "awful film, do not watch this",
    "terrible movie, i want my time back",
    "slow, boring and badly acted",
    "worst acting i have ever seen",
    "a forgettable and dull experience",
    "hated this film, complete waste of time",
    "predictable, boring and poorly acted",
]
texts = positive + negative
labels = [1] * len(positive) + [0] * len(negative)

X_train, X_test, y_train, y_test = train_test_split(
    texts, labels, test_size=10, random_state=42, stratify=labels
)

vec = TfidfVectorizer()
Xtr = vec.fit_transform(X_train)
Xte = vec.transform(X_test)

clf = LogisticRegression()
clf.fit(Xtr, y_train)
pred = clf.predict(Xte)

print("train size:", len(X_train), " test size:", len(X_test))
print("test accuracy:", accuracy_score(y_test, pred))
print()
for text, true, p in zip(X_test, y_test, pred):
    mark = "OK" if true == p else "WRONG"
    print(f"  {mark:5} true={true} pred={p}  {text!r}")
Output
train size: 30  test size: 10
test accuracy: 0.8

  OK    true=0 pred=0  'dull, predictable and forgettable'
  OK    true=0 pred=0  'boring plot and terrible acting'
  OK    true=1 pred=1  'superb performance from the whole cast'
  WRONG true=0 pred=1  'dull and forgettable from start to end'
  OK    true=1 pred=1  'a wonderful and delightful watch'
  OK    true=0 pred=0  'terrible movie, i want my time back'
  OK    true=1 pred=1  'wonderful cinematography and pacing'
  OK    true=1 pred=1  'this movie was fantastic and moving'
  WRONG true=1 pred=0  'i loved every minute of this film'
  OK    true=0 pred=0  'i regret watching this film'

random_state=42 fixes the train/test split and LogisticRegression's solver is deterministic on fixed input, so this exact split, this exact 80% and these exact two mistakes will reproduce on any machine running the code above unchanged. A different random_state, or a few more hand-written sentences added to the lists, will shift the split and can move the accuracy by a point or two — the pattern of "a strong baseline with a handful of explainable mistakes" is what to expect, not this precise number.

Line by line

80% accuracy, from a training set of 30 sentences, in well under a second. Compare that to the majority-class baseline — always guessing the more common label — which on a balanced test set of 10 would score 50%. This simple model beats a coin flip by 30 points with no architecture design and no GPU.

The two wrong predictions are informative, not embarrassing. "Dull and forgettable from start to end" was predicted positive. Every individual word here except "dull" and "forgettable" is neutral, and the sentence is short — the model had thin evidence to work with, illustrating why baselines need enough data per class to be reliable, not only enough data overall.

stratify=labels keeps the same 50/50 balance of positive and negative examples in both the train and test split. Without it, an unlucky split could put almost all of one class into the test set, making accuracy numbers meaningless.

Common mistakes

Reporting accuracy without a baseline to compare against. "92% accuracy" sounds impressive until you learn 90% of your data was one class, and always guessing that class alone would score 90%. Always report the majority-class baseline alongside your model's score.

Testing on the same data you trained on. A classifier can memorise training examples and still fail completely on new text. train_test_split exists specifically to catch this — evaluate only on data the model never saw during training.

Assuming this baseline handles negation. Bag-of-words-style models see "not", "boring" and "not boring" as three separate, disconnected signals, not one inverted meaning. "This is not a bad film" can confuse a linear text classifier in a way it would never confuse a human reader. This specific weakness is one of the more compelling arguments for reaching for a model that reads word order, once your baseline's errors start clustering around negation.

Try it yourself

Swap LogisticRegression() for from sklearn.naive_bayes import MultinomialNB and MultinomialNB(), keeping everything else the same. Naive Bayes is the other classic baseline choice — it is faster to train and often performs similarly on small text datasets like this one.

What to learn next

Researcher — Mathematics and papers.

What "baseline" means precisely

A baseline is a deliberately simple, fast, well-understood model used to establish a floor — the performance level any more complex method must clear to justify its added cost. Two baselines matter for classification specifically:

  • Majority-class baseline: predict the most frequent training label for every input. Accuracy equals the majority class's prevalence in the test set.
  • Linear-model-on-sparse-features baseline: TF-IDF (or bag of words) plus logistic regression, linear SVM, or Naive Bayes.

A result is only meaningful relative to both.

Why linear models on sparse text features are strong

Text classification features (TF-IDF weighted unigrams/bigrams) are extremely high-dimensional (10^4 to 10^6) and extremely sparse (a document touches a few dozen to a few hundred dimensions). In this regime, linear separability is common: many real-world label boundaries in text genuinely correspond, to good approximation, to a weighted sum of word-presence signals. This is precisely the setting where linear models are strong and where the added representational capacity of a deep network buys comparatively little, because the fundamental predictive signal — vocabulary co-occurring with a label — is itself close to linear.

Joachims (1998) formalized this observation for SVMs specifically, showing that text classification meets several theoretical conditions favouring linear SVMs: high dimensionality, few irrelevant features after TF-IDF weighting, and document vectors that are sparse.

Measured comparisons

Wang & Manning (2012), Baselines and Bigrams: Simple, Good Sentiment and Topic Classification, showed that Naive Bayes with bigram features, plus a simple SVM variant, was competitive with or better than considerably more complex methods of that era on several standard sentiment benchmarks — a result frequently cited as a caution against skipping the baseline step.

More recently, on modern benchmarks, fine-tuned transformer encoders (BERT-family) reliably outperform TF-IDF-plus-linear baselines, typically by 5 to 15 accuracy points on harder semantic tasks (natural language inference, nuanced sentiment, sarcasm), and by a much smaller margin — sometimes under 2 points — on coarse topic or spam-style classification where lexical signal alone is close to sufficient. The size of that gap is precisely the number a baseline exists to measure before committing to the more expensive model.

Regularization and the bias-variance trade-off here

LogisticRegression's default C=1.0 sets inverse L2 regularization strength. With a 10^4-dimensional sparse feature space and only tens or hundreds of training examples, the model is in a regime where regularization strength materially affects generalization — too little, and the model memorises rare words that happened to appear once in one class; too much, and genuinely useful rare words get suppressed toward zero weight. Cross-validating C is standard practice; see cross-validation strategies.

Key references

  • Joachims, T. (1998). Text Categorization with Support Vector Machines: Learning with Many Relevant Features. ECML.
  • Wang, S. & Manning, C. D. (2012). Baselines and Bigrams: Simple, Good Sentiment and Topic Classification. ACL.
  • Zhang, Y. & Wallace, B. (2015). A Sensitivity Analysis of (and Practitioners' Guide to) Convolutional Neural Networks for Sentence Classification. Shows how much classification results vary with hyperparameters even for models well past the linear baseline — a caution about over-trusting any single reported number, baseline or otherwise.

Current state and open problems

The open question in practice is not whether baselines are worth building — that is settled — but how to make the comparison honest under distribution shift. A TF-IDF baseline and a fine-tuned BERT model can score similarly on an in-distribution test set and diverge sharply on out-of-domain text, adversarial phrasing, or a shifted label definition, because the deep model may be capturing more transferable structure even where the aggregate accuracy numbers look close. A single held-out accuracy number, from either model, is not sufficient evidence for a production decision on its own.

What to learn next

  • Naive Bayes — the probabilistic sibling baseline to logistic regression here.
  • Support vector machines — the other classical linear classifier commonly paired with TF-IDF.
  • BERT — what you reach for once the baseline has told you it is worth the cost.