Natural Language Processing

Text classification

Text classification sorts a piece of writing into one of a few named boxes, and the hard parts are almost never the model — they are the labels, the duplicates and the number you report.

Read these first

On this page 10
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. How it works
  5. Two ways to be wrong
  6. The trap that catches everyone
  7. The part nobody warns you about
  8. Where you have already seen it
  9. Remember this
  10. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

Text classification is sorting a piece of writing into one of a few named boxes.

The analogy you have already lived

Think about the day's post arriving at your house. You stand at the door and deal the envelopes into three piles. Bills. Letters that matter. Junk.

You do not read a single envelope fully. A window with a printed address goes to bills. A handwritten name goes to the letters pile. A glossy offer goes to junk.

A few clues, a fast decision, three piles. That is text classification, and you have been doing it your whole life.

Why it exists

A shop receiving thirty complaints a day can have a person read them. A shop receiving thirty thousand cannot.

Every large service faces the same wall. Somebody has to look at each message and decide where it goes, and there are not enough somebodies. So the sorting gets handed to a model, and the humans handle only what the model is unsure about.

That last part matters. A good system does not replace the person. It shrinks their pile.

How it works

   "my refund has not reached my upi account"
                     |
                     v
        [ turn words into numbers ]
                     |
                     v
             [ the model ]
                     |
                     v
       billing   0.81
       delivery  0.14      <- a score for every box
       technical 0.05
                     |
                     v
              billing

Notice the middle. The model does not hand you a box. It hands you a score for every box, and something afterwards picks the winner.

That is not a detail. It is the most useful lever you have, and most beginners never touch it. Suppose wrongly routing a fraud complaint is expensive. You can demand a much higher score before acting. Everything below that goes to a human.

Two ways to be wrong

Every classifier makes two different kinds of mistake, and they hurt differently.

Think about a spam filter. It can let a spam email into your inbox. Annoying. Or it can throw your exam result letter into the spam folder. A disaster.

Same model, same accuracy number, wildly different pain. Which mistake you can afford is a business decision, not a technical one. The model evaluation lesson goes through how to measure each kind separately, and you should read it before reporting any number.

The trap that catches everyone

Imagine ninety-four out of every hundred emails are genuine, and six are spam.

Now build the laziest possible machine. It ignores the email completely and answers "not spam" every single time.

That machine is right ninety-four times out of a hundred. Ninety-four percent accuracy. It has learned nothing, reads nothing, and catches zero spam.

This is why a single accuracy number is close to worthless on its own. Whenever somebody quotes you accuracy without saying how common each box is, the number is not yet meaningful.

The part nobody warns you about

Here is what really limits text classification, and it is not the model.

Take a hundred customer complaints. Give the same hundred to two experienced support staff. Ask each to sort them into billing, delivery and technical.

They will disagree on a good number of them. Not because either is careless. Because "the payment went through but the order never shipped" belongs in two boxes at once. And "the app crashed during checkout so I got charged twice" belongs in three.

If two humans agree only eighty times out of a hundred, no model can be judged above roughly eighty. The rest is not a model failure. It is a disagreement about what the answer is.

This is confusing the first time and worth sitting with. Before you spend a month improving a model, spend a day checking whether your labels agree with each other. It is the cheapest large improvement available in this entire field.

Where you have already seen it

  • Your spam folder, the most widely deployed classifier ever built.
  • Support chat, routing your complaint to the right team before a human sees it.
  • App store and YouTube comments, filtered for abuse before publishing.
  • Language detection, deciding whether to offer you a translation.
  • Bank SMS parsing, sorting a message into "money in", "money out" or "promotion".

Remember this

  • A classifier gives you a score for every box, and you decide where to cut.
  • Accuracy alone lies, especially when one box is far more common than the others.
  • The quality of your labels sets the ceiling. No model can beat the disagreement in its own training data.

What to learn next

  • Model evaluation — measuring each kind of mistake separately.
  • Named entity recognition — labelling words inside a sentence rather than the whole sentence.
  • BERT — the encoder people fine-tune when word counting is not enough.

Developer — Code and libraries.

The what is NLP lesson built a working classifier in ten lines. This lesson is about everything after that: measuring it honestly, reading where it fails, and knowing when a bigger model is worth the money.

Setup

bash
pip install scikit-learn

That is the whole dependency list. Everything here runs on CPU in a couple of seconds.

A real evaluation, not a demo

Forty-two support tickets, three departments. Small on purpose so you can read every row.

tickets.py
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.linear_model import LogisticRegression
from sklearn.pipeline import make_pipeline
from sklearn.model_selection import StratifiedKFold, cross_val_predict
from sklearn.metrics import classification_report, confusion_matrix

BILLING = [
    "i was charged twice for the same order",
    "my refund has not reached my upi account",
    "the invoice shows the wrong gst amount",
    "you deducted money but the order failed",
    "please cancel the subscription and refund me",
    "i paid online but the payment is not showing",
    "card was debited twice yesterday",
    "the coupon discount was not applied to my bill",
    "money was deducted and no refund yet",
    "i want a refund for the cancelled order",
    "wrong amount charged on my credit card",
    "the bill has an extra delivery charge",
    "refund was promised but never credited",
    "i have been billed for a plan i never bought",
]
DELIVERY = [
    "the parcel has not arrived after ten days",
    "courier marked delivered but i received nothing",
    "my package is stuck at the hub since monday",
    "the delivery boy never called me",
    "shipment shows out for delivery for three days",
    "the box arrived crushed and wet",
    "wrong address on the shipping label",
    "the item was delivered to my neighbour",
    "parcel is delayed with no tracking update",
    "courier refused to deliver to my pin code",
    "the package was returned to the seller",
    "delivery was attempted when nobody was home",
    "my order shipped but tracking never moved",
    "the courier lost my shipment",
]
TECHNICAL = [
    "the app crashes when i open my orders",
    "login page keeps showing a server error",
    "the website will not load on chrome",
    "otp never arrives on my phone",
    "app freezes on the payment screen",
    "i cannot upload a photo in the review form",
    "the search bar returns a blank page",
    "password reset link gives an error",
    "the app logs me out every few minutes",
    "checkout button does nothing when clicked",
    "the site shows a 500 error on my profile",
    "images do not load in the product page",
    "the app keeps crashing after the update",
    "i get an error when i try to sign in",
]

texts  = BILLING + DELIVERY + TECHNICAL
labels = (["billing"] * len(BILLING) + ["delivery"] * len(DELIVERY)
          + ["technical"] * len(TECHNICAL))
names  = sorted(set(labels))

model = make_pipeline(TfidfVectorizer(), LogisticRegression(max_iter=1000))

# Every ticket gets predicted by a model that never saw it during fitting.
cv   = StratifiedKFold(n_splits=6, shuffle=True, random_state=0)
pred = cross_val_predict(model, texts, labels, cv=cv)

print(classification_report(labels, pred, zero_division=0))
print("rows = true label, columns = predicted label, order:", names)
print(confusion_matrix(labels, pred, labels=names))
Output
              precision    recall  f1-score   support

     billing       0.62      0.57      0.59        14
    delivery       0.53      0.64      0.58        14
   technical       0.92      0.79      0.85        14

    accuracy                           0.67        42
   macro avg       0.69      0.67      0.67        42
weighted avg       0.69      0.67      0.67        42

rows = true label, columns = predicted label, order: ['billing', 'delivery', 'technical']
[[ 8  5  1]
 [ 5  9  0]
 [ 0  3 11]]

Read that output before doing anything else

0.67 accuracy is a weak result, and printing it is the point. Forty-two examples is not enough data. The number tells you that, honestly, before you ship anything.

Now the confusion matrix, which is far more useful than the accuracy line.

                    predicted
                billing delivery technical
   billing          8       5        1
   delivery         5       9        0
   technical        0       3       11

The story is right there. technical is nearly solved — 11 of 14 correct, and zero technical tickets were ever called billing. Meanwhile billing and delivery bleed into each other, 5 in each direction.

That is not random noise. Read "the bill has an extra delivery charge" and "the package was returned to the seller". Those two categories genuinely overlap in the real world, and the model is telling you so.

The action that follows is concrete: do not tune the model. Either merge the two categories, or write a rule about which wins when both apply, and relabel accordingly.

Why cross-validation instead of a single split

With 42 rows, a single train/test split holds out about 10 tickets. Move random_state and your reported accuracy swings by fifteen points. That number is noise wearing a lab coat.

cross_val_predict with 6 folds fits six models. Each ticket is predicted exactly once, by a model that never saw it. You get a prediction for all 42 rows and a far steadier estimate.

Under a few thousand labelled examples, use cross-validation. Above that, a held-out test set is fine and much faster.

Look at what the model learned

Add these lines to the end of tickets.py and run it again. The report above prints first; this is what appears after it.

tickets.py (continued)
import numpy as np

model.fit(texts, labels)          # cross_val_predict never leaves a fitted model behind
vec = model.named_steps["tfidfvectorizer"]
clf = model.named_steps["logisticregression"]
words = vec.get_feature_names_out()

for i, cls in enumerate(clf.classes_):
    top = np.argsort(clf.coef_[i])[-5:][::-1]
    print(f"{cls:10} strongest words: {', '.join(words[j] for j in top)}")
Output
billing    strongest words: refund, amount, bill, card, twice
delivery   strongest words: courier, delivery, to, arrived, shipment
technical  strongest words: app, error, page, in, on

Most of those are sensible. refund, courier, app — a human would pick the same words.

Then look at to, in and on. Those are function words with no topical meaning at all. They earned high weights because, in 42 rows, they happened to land unevenly. That is memorisation of an accident.

Printing the strongest features takes three lines and catches this class of problem instantly. Do it on every text classifier you build, every time.

TF-IDF, in one paragraph

TfidfVectorizer counts words like CountVectorizer does, then divides each count by how common that word is across all documents.

The effect: a word appearing in every ticket ("the", "my") gets crushed toward zero. A word appearing in a handful ("gst", "otp") keeps its weight. It is a cheap, unsupervised way of saying "rare words carry more information", and it usually beats raw counts by a small but free margin.

The imbalance trap, in code

lazy.py
from sklearn.dummy import DummyClassifier
from sklearn.metrics import accuracy_score, precision_recall_fscore_support

# 200 support tickets. Only 12 are genuine fraud reports.
y_true = ["normal"] * 188 + ["fraud"] * 12
X = [[0]] * len(y_true)                      # no features at all, on purpose

lazy = DummyClassifier(strategy="most_frequent").fit(X, y_true)
pred = lazy.predict(X)

p, r, f, _ = precision_recall_fscore_support(
    y_true, pred, labels=["fraud"], zero_division=0)

print(f"accuracy of a model that always says 'normal': {accuracy_score(y_true, pred):.3f}")
print(f"fraud precision {p[0]:.3f}  fraud recall {r[0]:.3f}  fraud f1 {f[0]:.3f}")
Output
accuracy of a model that always says 'normal': 0.940
fraud precision 0.000  fraud recall 0.000  fraud f1 0.000

94% accurate. Catches nothing. Run DummyClassifier as your first model on every project and report its score alongside your real one. If your model cannot beat it by a wide margin, you do not have a model.

For a genuine fix, pass class_weight="balanced" to LogisticRegression. It makes errors on the rare class count proportionally more during fitting. It is one keyword and it is often the biggest single improvement available.

The bug that will actually bite you: near-duplicate leakage

This one is specific to text and it is everywhere.

Support tickets, product reviews and news articles contain near-duplicates. The same complaint filed twice. A forwarded email quoted inside a reply. A template message with one field changed.

A random train/test split puts one copy in train and the other in test. Your model memorises it and scores it correctly. Your reported accuracy is inflated, sometimes enormously, and production accuracy is not.

The fix is to deduplicate before splitting, not after. Append this to tickets.py as well; again, only the final line is shown.

tickets.py (continued)
import hashlib

# Two tickets arrive again, one of them with different capitalisation.
noisy_texts  = texts  + ["I was charged twice for the same order", "the courier lost my shipment"]
noisy_labels = labels + ["billing", "delivery"]

seen, unique_texts, unique_labels = set(), [], []
for t, y in zip(noisy_texts, noisy_labels):
    key = hashlib.md5(" ".join(sorted(t.lower().split())).encode()).hexdigest()
    if key not in seen:                       # same words in any order = the same ticket
        seen.add(key)
        unique_texts.append(t); unique_labels.append(y)

print(f"{len(noisy_texts)} tickets in, {len(unique_texts)} after removing duplicates")
Output
44 tickets in, 42 after removing duplicates

That catches exact and reordered duplicates, including ones that differ only by case. Near-duplicates need fuzzy matching — MinHash, or cosine similarity over TF-IDF vectors with a threshold. Either is a morning's work and routinely worth several points of honest accuracy.

Preprocessing choices that are not obvious

Lowercasing. On by default in scikit-learn. Usually right for topic classification. Wrong when case is signal — "HELP" versus "help" carries urgency, and it destroys the main clue for named entity recognition.

Stopword removal. Frequently harmful. "not" and "never" are on most stopword lists, and removing them turns "this is not working" into "working". For sentiment, leave stopwords in.

Stemming and lemmatisation. Cutting "running" to "run". Modest gains on small datasets, near-zero once you have enough data, and destructive for languages with rich morphology if the stemmer was built for English.

Character n-grams. TfidfVectorizer(analyzer="char_wb", ngram_range=(3, 5)) splits into overlapping character chunks instead of words. This is the single best trick for noisy, misspelt or code-mixed text. Hinglish tickets, where a sentence switches between English and Romanised Hindi mid-way, are exactly where word features fall apart and character features hold up.

Choosing your model, honestly

ApproachLabels neededCost per 1M docsLatencyWhen it is right
TF-IDF + logistic regression~500near zerounder 1 mskeyword-driven tasks, always run this first
Fine-tuned encoder (BERT)~2,000GPU hours once, then cheap5–20 msmeaning matters, volume is high, labels are stable
LLM with a prompt0high, per call300 ms+no labels yet, categories change often

A pattern that works well in practice: use an LLM to label a few thousand examples, check a sample by hand, then train an encoder on those labels for serving. You get the LLM's understanding at the encoder's cost.

Common mistakes

Calling fit on your test data. fit_transform on train, transform everywhere else. Using a Pipeline makes this structurally impossible, which is the main reason to use one.

Reporting accuracy without the class balance. Always print collections.Counter(labels) next to your score. Without it the score cannot be interpreted.

Choosing the threshold at 0.5 because it is the default. 0.5 is a convention, not a decision. Pick the cut based on what each mistake costs — model evaluation covers how.

Tuning the model when the labels are the problem. Before your second modelling idea, sample 50 rows the model got wrong and read them. A large share will turn out to be labelled wrongly, or genuinely ambiguous. That reading session is worth more than a week of hyperparameter search.

Assuming one label per document. A ticket can be billing and delivery. If that is your reality, this is multi-label classification, and single-label metrics will mislead you. Use MultiOutputClassifier and per-label scores.

Try it yourself

Add class_weight="balanced" to the LogisticRegression in tickets.py and re-run. The macro average will move. Note the direction and work out why.

Then swap TfidfVectorizer() for TfidfVectorizer(analyzer="char_wb", ngram_range=(3, 5)). Predict first: character features shine on noisy text but you have clean text here, so they may not help. Check whether you were right.

Finally, write five new tickets in Hinglish — "paise cut ho gaye but order fail", "delivery boy ne call hi nahi kiya" — and predict them with each version. That comparison tells you more about your real users than any benchmark.

What to learn next

Researcher — Mathematics and papers.

The feature map

Classical text classification factorises into a fixed feature map followed by a linear model. The map is the interesting part.

TF-IDF. For term t in document d over a corpus D:

text
tfidf(t, d) = tf(t, d) * idf(t)

idf(t) = log( (1 + N) / (1 + df(t)) ) + 1          (scikit-learn, smooth_idf=True)
  • tf(t, d) is the raw count of t in d.
  • N = |D| is the number of documents.
  • df(t) is the number of documents containing t.
  • The +1 terms prevent division by zero and stop idf reaching exactly zero.

Rows are then L2-normalised by default, so document length does not dominate the dot product. Note that scikit-learn's idf differs from the textbook log(N / df(t)); papers reporting "TF-IDF baseline" without naming the variant are not reproducible.

Character n-grams. char_wb extracts n-grams padded within word boundaries. Robust to misspelling, morphology and code-switching, at the cost of a much larger and denser feature space. For Romanised Indic text — where a single word has half a dozen accepted spellings — this frequently outperforms word features outright.

Why linear models remain the baseline

For high-dimensional sparse features, the data is close to linearly separable, and regularised linear models are near-optimal. Multinomial logistic regression:

text
P(y = k | x) = exp(w_k . x + b_k) / sum over j of exp(w_j . x + b_j)
  • x is the feature vector for one document.
  • w_k, b_k are the weight vector and bias for class k.
  • The denominator sums over all classes, giving a proper distribution.

Fitted by minimising regularised negative log likelihood. With L2 regularisation the objective is strictly convex, so there is a unique optimum and no seed sensitivity.

Two facts worth carrying:

Naive Bayes is not obsolete. Wang & Manning (2012), Baselines and Bigrams, showed that multinomial naive Bayes beats SVMs on short snippet tasks, and that NBSVM — an SVM over log-count ratio features — was competitive with far heavier methods. Their broader point holds: the bigram baseline is frequently underreported.

Linear SVMs and logistic regression differ mainly in calibration. Hinge loss gives no probability. If you need a score to threshold, use logistic regression, or wrap the SVM in CalibratedClassifierCV.

Averaging schemes, precisely

The choice between micro, macro and weighted averaging changes your headline number substantially, and it is routinely left unstated.

text
micro:     pool all TP, FP, FN across classes, then compute the metric once
macro:     compute the metric per class, then take the unweighted mean
weighted:  compute the metric per class, then average weighted by support
  • In single-label multi-class problems, micro-F1 equals accuracy exactly. Reporting both as separate achievements is a common error.
  • Macro-F1 gives a rare class the same weight as a common one. Use it when rare classes matter.
  • Weighted-F1 tracks the majority class and hides rare-class failure. It is the default in many reports and it is usually the wrong choice.

State which you used. A macro-F1 of 0.62 and a weighted-F1 of 0.88 can describe the identical model.

Imbalance, with the trade-off stated

Options, roughly in order of how often they are the right answer:

  1. Cost-sensitive reweighting. class_weight="balanced" sets w_k = n / (K * n_k), where n is total samples, K the number of classes and n_k the count of class k. Cheap, no data distortion, no resampling artefacts.
  2. Threshold adjustment. Train normally, move the operating point afterwards. This decouples the fitting problem from the decision problem, which is the theoretically cleaner separation.
  3. Resampling. SMOTE (Chawla et al., 2002) interpolates synthetic minority points. In high-dimensional sparse text space, interpolation between TF-IDF vectors produces documents that do not correspond to any text, and reported gains over simple reweighting are inconsistent.
  4. Focal loss (Lin et al., 2017). Down-weights easy examples by (1 - p_t)^gamma. Designed for extreme foreground/background imbalance in detection; gains on text classification are modest.

The important caveat: resampling and reweighting both destroy calibration. Your predicted probabilities no longer match observed frequencies. If a downstream system consumes those probabilities, correct them back — Saerens et al. (2002) give the prior-shift adjustment.

The label-quality ceiling

Model performance is bounded by annotation consistency, and this bound is measurable.

Report inter-annotator agreement before reporting model scores. Cohen's kappa for two annotators, Fleiss' kappa for more, Krippendorff's alpha when annotators are missing labels:

text
kappa = (p_o - p_e) / (1 - p_e)
  • p_o is observed agreement.
  • p_e is agreement expected by chance from the marginal distributions.

Values above 0.8 are strong; 0.6 to 0.8 is workable; below 0.6 means your task definition is underspecified and no modelling will repair it.

Bowman & Dahl (2021), What Will it Take to Fix Benchmarking in NLU?, is the sharpest statement of the surrounding problem: benchmark ceilings are often annotation ceilings, and models exceeding human agreement are usually exploiting annotation artefacts rather than understanding.

Northcutt et al. (2021) audited ten widely used benchmarks and found pervasive label errors — around 3.4% on average across test sets, and higher in some. Their confident-learning method estimates the joint distribution of noisy and true labels and is worth running on any dataset before trusting a leaderboard position on it.

Current landscape

Fine-tuned encoders — DeBERTa-v3, ModernBERT — remain the strongest option per unit of inference cost when a few thousand labels exist. See BERT.

Zero-shot and few-shot LLM classification is now a serious baseline, and its failure modes are specific. Label order in the prompt affects predictions (Lu et al., 2022, on prompt-order sensitivity). Verbaliser choice — the surface words used for each class — affects results more than most model choices. And the returned probabilities are poorly calibrated, so thresholding them is unsafe without a calibration set.

Distillation is the pattern that reconciles the two. Generate labels with an LLM, verify a sample, fine-tune an encoder. Reported results are close to the teacher at a fraction of serving cost, and the encoder gives you a calibrated score the LLM does not.

Extreme multi-label classification — hundreds of thousands of labels, as in product tagging — is its own field. Standard softmax is intractable; the practical approaches are label trees and approximate nearest-neighbour retrieval over label embeddings.

Key references

  • Wang, S. & Manning, C. (2012). Baselines and Bigrams: Simple, Good Sentiment and Topic Classification. ACL.
  • Joachims, T. (1998). Text Categorization with Support Vector Machines. ECML.
  • Chawla, N. et al. (2002). SMOTE: Synthetic Minority Over-sampling Technique. JAIR 16.
  • Saerens, M., Latinne, P. & Decaestecker, C. (2002). Adjusting the Outputs of a Classifier to New a Priori Probabilities. Neural Computation 14(1).
  • Lin, T.-Y. et al. (2017). Focal Loss for Dense Object Detection. arXiv:1708.02002
  • Northcutt, C., Athalye, A. & Mueller, J. (2021). Pervasive Label Errors in Test Sets. arXiv:2103.14749
  • Bowman, S. & Dahl, G. (2021). What Will it Take to Fix Benchmarking in Natural Language Understanding? arXiv:2104.02145
  • Lu, Y. et al. (2022). Fantastically Ordered Prompts and Where to Find Them. arXiv:2104.08786

Open problems

Distribution shift is unsolved in practice. Text drifts continuously — new product names, new slang, new complaint types. Detecting that a deployed text classifier has degraded, without fresh labels, remains genuinely hard. Confidence-based drift signals are unreliable precisely because miscalibration and drift arrive together.

Evaluation on code-mixed text. Benchmarks for Hinglish and other code-switched varieties are small, inconsistently annotated, and not representative of production traffic. A model reported at 0.9 on a public code-mixed set will not reproduce that on your own tickets, and the field has no good answer for why.

What to learn next