How it works
How spam filters work
A spam filter turns each email into measurable features, scores it with a classifier trained on millions of labelled messages, and keeps learning as spammers adapt.
- 3 min read
- Updated
On this page 6
Billions of spam emails are sent every day. You see almost none of them. The wall between you and that flood is a classifier — a model that reads every incoming message and makes one call: inbox or spam. Here is how that call gets made.
The pipeline at a glance
incoming email
|
v
[1. extract features] words, sender, links, headers
|
v
[2. score] trained classifier -> 0.97 "spammy"
|
v
[3. threshold] above the line? -> spam folder
|
v
[4. learn] your "report spam" clicks retrain itStage 1 — the email becomes features
A model cannot read an email the way you do. The filter first converts the message into features — measurable facts about it. The words it contains, chopped into tokens. Who sent it, and whether that server really speaks for that domain. How many links, and where they point. Whether the subject line is all capitals.
Think of a security guard at a gate. The guard does not know each visitor personally. They check observable things: the badge, the vehicle, the time of day. Features are those observable things, for email.
Stage 2 — a classifier turns features into a score
The features go into a classifier — a model trained to sort inputs into categories, here spam or not-spam. It was trained on millions of emails that people already labelled, largely through "report spam" clicks.
Training amounts to learning weights: how strongly each feature pushes toward spam or ham (the traditional word for legitimate mail). "Free" plus a shortened link plus a mismatched sender pushes hard one way. A sender you reply to weekly pushes hard the other. The output is a single probability, say 0.97, meaning: messages that looked like this were spam 97 times in 100.
The classic first-generation filter was Naive Bayes, which scores each word independently and multiplies the evidence together. Modern filters layer on stronger models, but the shape — features in, probability out — is unchanged.
Stage 3 — the threshold makes the decision
A probability is not a decision. The filter compares the score against a threshold: above it, spam folder; below it, inbox.
Where to set that line is a genuine trade-off, because the two mistakes are not equally bad. Spam in your inbox is annoying. A job offer in your spam folder is a small disaster. So providers set the threshold cautiously, accepting more leaked spam to avoid burying real mail. Some add a middle band — the "suspicious" warning banner — for scores near the line.
Stage 4 — the filter keeps learning
Spam filtering never finishes, because the opponent adapts. Spammers misspell trigger words, hide text in images, and hijack reputable servers. Each trick erodes yesterday's features.
So the loop closes continuously. Every "report spam" and every "not spam" rescue becomes a fresh labelled example, and the classifier retrains on the new data. This is why a trick that floods inboxes this month quietly stops working next month — millions of users voted, and the weights moved.
This adversarial back-and-forth makes spam filtering unusual among ML problems: the data distribution shifts because the model got better. A fraud model at a bank lives the same life.
Which lessons teach each stage
- Stage 1, chopping text into usable pieces: Tokenization
- Stage 2, models that sort into categories: Classification and Logistic regression
- Stage 3, weighing the two kinds of mistake: Model evaluation