AI Projects

Build a spam detector

Train a model that reads a message and decides whether it is spam, using twenty-eight hand-typed examples and about forty lines of Python.

Read these first

On this page 9
  1. The short answer
  2. The analogy you have already lived
  3. Why this project exists
  4. What you are actually going to build
  5. How it works
  6. Where you have already seen this
  7. The honest part
  8. Remember this
  9. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

A spam detector is a program that reads a message and sorts it into one of two trays: real, or junk.

The analogy you have already lived

Think about the pile of post that arrives at your house. You do not read every letter to sort it. You glance at the envelope. A handwritten address from a familiar name goes in one pile. A shiny "YOU HAVE WON" envelope goes in another.

You never studied envelope-sorting. You learned it by seeing hundreds of envelopes and finding out later which ones mattered.

A spam detector learns the same way, from messages somebody has already sorted for it.

Why this project exists

Roughly half of all email sent in the world is junk. Text message scams cost people real money every day in India — fake KYC alerts, fake lottery wins, fake loan approvals.

Somebody has to sort that. Doing it by hand does not scale, because a mail service handles billions of messages a day.

Writing rules by hand does not work either. Block the word "free" and you block your friend asking if you are free on Sunday. Block "prize" and you block your school newsletter. Every rule you write breaks something honest.

So we stop writing rules. We show the machine solved examples, and let it work out the pattern.

What you are actually going to build

You will give a program twenty-eight short messages. Fourteen are scams. Fourteen are ordinary messages from friends and family. Each one already carries a label saying which it is.

The program counts words. It notices that "claim" and "congratulations" turn up often in the junk pile and almost never in the real pile. It notices the reverse for "tomorrow" and "meeting".

Then you hand it a message it has never seen, and it makes a guess with a confidence attached.

How it works

   "WIN a free iPhone now! Click to claim"
                 |
                 v
       [ chop into separate words ]
                 |
                 v
     win   free   iphone   click   claim
                 |
                 v
   [ how often does each word show up in
     the junk pile, versus the real pile? ]
                 |
                 v
        "spam"   (99% sure)  ->  junk tray

There is no understanding of English anywhere in that picture. The program has no idea what an iPhone is. It counts words, and counting is enough to be useful.

Where you have already seen this

  • Gmail's spam folder. The most-used machine learning system on earth.
  • Your phone's "Spam" tab in the messages app, catching fake delivery alerts.
  • WhatsApp's "forwarded many times" label, which uses different signals for a related job.
  • Comment filters on YouTube that hide crypto-scam replies before you see them.

The honest part

This model can be fooled, and you will fool it yourself later in this lesson. A real message asking "Are you free tomorrow evening?" gets flagged as spam by the model you are about to build. The word "free" carries too much weight, because our training pile is tiny.

That is not a bug in the code. It is the model correctly learning from a small, unrepresentative sample. Real spam filters see millions of messages, which is the actual fix.

There is a second honest problem. Spammers read about spam filters too. They change their wording every week to slip past. A spam filter is never finished — it is a fight that continues for as long as the service exists.

Remember this

  • A spam detector learns from labelled examples, not from rules a person wrote.
  • It works by counting which words lean junk and which lean real.
  • A small training pile gives a confident model with bad judgement — more honest examples is the real fix.

What to learn next

Developer — Code and libraries.

The problem, stated precisely

Given a short text message, output a probability that it is spam. This is binary text classification: two possible labels, and text as the input.

Our plan:

  1. Turn each message into a count of its words — a bag of words.
  2. Fit multinomial naive Bayes, a model that learns how likely each word is under each label.
  3. Measure it honestly, then break it on purpose.

Naive Bayes is the right first choice here. It trains in milliseconds, needs very little data to be reasonable, and you can read its learned weights and understand them. It was the algorithm behind the first practical spam filters.

Setup

bash
pip install scikit-learn numpy

That is everything. No GPU, no downloads, no dataset file. The whole thing runs in under two seconds on any laptop.

The dataset

Our dataset is twenty-eight messages typed directly into the file. That is deliberately, almost absurdly small — small enough that you can read every training example and check the model's reasoning against your own.

When you want a real one, the SMS Spam Collection from the UCI repository has 5,574 labelled messages and downloads as a 200 KB zip. There is a link at the end of this block.

The full build

spam_detector.py
import numpy as np
from sklearn.feature_extraction.text import CountVectorizer
from sklearn.naive_bayes import MultinomialNB
from sklearn.model_selection import train_test_split, cross_val_score
from sklearn.pipeline import make_pipeline
from sklearn.metrics import classification_report, confusion_matrix

# 28 messages typed by hand. 1 = spam, 0 = a real message ("ham" is the usual word).
DATA = [
    ("WIN a free iPhone now! Click this link to claim your prize", 1),
    ("Congratulations you have won 50000 rupees claim now", 1),
    ("URGENT your account will be closed click here to verify", 1),
    ("Free recharge offer click the link and get 1GB data free", 1),
    ("You are selected for a lottery prize claim it before midnight", 1),
    ("Get a personal loan approved instantly no documents needed", 1),
    ("Claim your free cashback now limited offer click fast", 1),
    ("Your KYC is pending click this link to update immediately", 1),
    ("Win win win! Free tickets waiting for you claim now", 1),
    ("Cheap medicine online free home delivery order now", 1),
    ("Congratulations! Your number won a free holiday package", 1),
    ("Click here now to double your money in one week guaranteed", 1),
    ("Free free free unlimited offer claim your reward today", 1),
    ("Urgent! Verify your bank account or it will be blocked", 1),
    ("Hey are we still meeting at the cafe tomorrow", 0),
    ("Please send me the notes from yesterday class", 0),
    ("Mummy said dinner is ready come home early", 0),
    ("The train is delayed by an hour I will reach late", 0),
    ("Can you review my code before the standup tomorrow", 0),
    ("Happy birthday! Have a wonderful year ahead", 0),
    ("I have submitted the assignment on the portal", 0),
    ("Let me know when you reach the office", 0),
    ("The meeting is moved to three o clock today", 0),
    ("Thanks for helping me with the project yesterday", 0),
    ("Are you coming for the cricket match on Sunday", 0),
    ("I will call you after my class finishes", 0),
    ("Sending the photos from the trip in a while", 0),
    ("Your order has been delivered at your doorstep", 0),
]

texts = [t for t, _ in DATA]
labels = [y for _, y in DATA]

# stratify keeps the spam/ham ratio identical in both halves of the split
X_train, X_test, y_train, y_test = train_test_split(
    texts, labels, test_size=0.25, random_state=42, stratify=labels
)

vec = CountVectorizer()          # turns each message into word counts
X_train_counts = vec.fit_transform(X_train)

model = MultinomialNB()          # counts words, then applies Bayes' rule
model.fit(X_train_counts, y_train)

print("vocabulary size:", len(vec.vocabulary_))
print("training matrix:", X_train_counts.shape, "(messages, words)")
print()

pred = model.predict(vec.transform(X_test))
print(classification_report(y_test, pred, target_names=["ham", "spam"], zero_division=0))
print("confusion matrix [[TN FP] [FN TP]]:")
print(confusion_matrix(y_test, pred))
Output
vocabulary size: 112
training matrix: (21, 112) (messages, words)

              precision    recall  f1-score   support

         ham       1.00      1.00      1.00         4
        spam       1.00      1.00      1.00         3

    accuracy                           1.00         7
   macro avg       1.00      1.00      1.00         7
weighted avg       1.00      1.00      1.00         7

confusion matrix [[TN FP] [FN TP]]:
[[4 0]
 [0 3]]

Do not celebrate that 1.00

A perfect score on seven test messages tells you almost nothing. Seven items means every message is worth about fourteen percentage points. One different message and the number moves a long way.

This is the most common self-deception in applied machine learning, and it costs people real jobs. Read the support column before the accuracy column, every single time. It shows how many test examples each number was computed from.

Here is the same recipe measured properly.

spam_detector.py (add to the bottom)
# The same recipe re-fitted five times on five different splits.
pipe = make_pipeline(CountVectorizer(), MultinomialNB())
scores = cross_val_score(pipe, texts, labels, cv=5)
print("accuracy on each of 5 folds:", np.round(scores, 3))
print("mean accuracy:", round(float(scores.mean()), 3))
print()

# Which words is the model actually leaning on?
log_ratio = model.feature_log_prob_[1] - model.feature_log_prob_[0]
names = np.array(vec.get_feature_names_out())
order = np.argsort(log_ratio)
print("most spam-ish words:", list(names[order[-8:]][::-1]))
print("most ham-ish words: ", list(names[order[:8]]))
print()

# Messages the model has never seen.
new = [
    "Free prize waiting click now to claim",
    "Can you send me the meeting notes please",
    "Are you free tomorrow evening",
]
for msg, p in zip(new, model.predict_proba(vec.transform(new))[:, 1]):
    print(f"{'SPAM' if p > 0.5 else 'ok':4}  p(spam)={p:.3f}  {msg}")
Output
accuracy on each of 5 folds: [0.833 0.833 1.    1.    0.8  ]
mean accuracy: 0.893

most spam-ish words: ['free', 'now', 'claim', 'win', 'click', 'offer', 'congratulations', 'won']
most ham-ish words:  ['the', 'yesterday', 'is', 'my', 'me', 'at', 'class', 'meeting']

SPAM  p(spam)=1.000  Free prize waiting click now to claim
ok    p(spam)=0.001  Can you send me the meeting notes please
SPAM  p(spam)=0.770  Are you free tomorrow evening

Read that output — it contains the whole lesson

The five folds disagree: 0.833, 0.833, 1.0, 1.0, 0.8. The honest headline number is 0.893, not 1.00. Cross-validation refits the model on five different splits and averages. It is more work and it is the number you report.

The learned words are readable, and they make sense. "claim", "congratulations", "won" lean spam. "meeting", "class", "yesterday" lean real. You can audit this model by eye, which is rare and valuable. A deep network gives you no such window.

The third message got it wrong. "Are you free tomorrow evening" scored 0.770 and was flagged as spam. That is a false positive — a real message thrown into the junk tray.

Look at why. In our training pile "free" appears only in scams. The model has no counter-example of an innocent "free". It is reasoning correctly from bad evidence.

For a spam filter, a false positive is far worse than a false negative. A scam text you see and ignore costs you a second. A job offer silently binned costs you a job.

Line by line, the parts that trip people up

CountVectorizer() builds the vocabulary from the training text and converts each message into a row of counts. Position is discarded entirely — "dog bites man" and "man bites dog" become the same row. That is the "bag" in bag of words.

vec.fit_transform(X_train) then vec.transform(X_test) — note the different methods. fit_transform learns the vocabulary and applies it. transform only applies it. Calling fit_transform on the test set lets test words leak into the vocabulary, which is data leakage: the model is quietly told something about data it is meant to be graded on.

stratify=labels forces the same spam-to-ham ratio in the train and test halves. Leave it out with a dataset this small and you can land a test set with no spam in it at all.

model.feature_log_prob_ holds, for each label, the learned log-probability of every word. Subtracting one row from the other gives a per-word score for how much that word tilts the decision.

confusion_matrix returns [[TN, FP], [FN, TP]]. Top-right is the count of real messages wrongly called spam — for this application, that is the cell to watch.

Common mistakes

Reporting accuracy on an imbalanced dataset. Real inboxes might be 5% spam. A model that labels everything "ham" scores 95% accuracy and catches zero spam. Use precision, recall and the confusion matrix. Accuracy alone hides total failure.

Fitting the vectorizer on all the data before splitting. Your scores go up and your model gets worse in production. Fit on train only, or wrap everything in a Pipeline as the second code block does, so scikit-learn handles it for you.

Assuming a 0.5 threshold is correct. It is a default, not a decision. Since false positives hurt more here, raise the bar: model.predict_proba(X)[:, 1] > 0.9. Choosing this number is a product decision, not a maths one.

Testing on messages you wrote while looking at the training data. You will unconsciously write things the model handles. Ask a friend for ten real texts from their phone instead.

How to make this genuinely good

In rough order of payoff per hour of work:

  1. Get more data. Download the SMS Spam Collection — 5,574 real labelled messages, about 200 KB. Nothing else on this list matters as much.
  2. Switch to TfidfVectorizer. It down-weights words that appear in nearly every message, so "the" stops competing with "claim".
  3. Add character n-grams: TfidfVectorizer(analyzer="char_wb", ngram_range=(3, 5)). This catches deliberate misspellings like "fr33" and "cl1ck" that word-based features miss completely.
  4. Try LinearSVC or LogisticRegression. On a few thousand messages these usually beat naive Bayes, and logistic regression gives better-calibrated probabilities.
  5. Tune the threshold on a validation set against a stated goal, such as "keep false positives under one in a thousand".
  6. Save the model with joblib.dump(pipe, "spam.joblib") and load it in a small FastAPI endpoint.

Try it yourself

Add three real messages containing the word "free" to the ham side of DATA. Use "are you free later", "the event is free to attend" and "free seats at the back". Re-run.

Watch "free" fall out of the most-spam-ish list, and watch that false positive turn into a correct answer. You have now performed the single most effective fix in machine learning: giving the model the counter-example it was missing.

What to learn next

Researcher — Mathematics and papers.

Formulas here are written in plain text, since the site renders no maths typesetting library.

The model

Multinomial naive Bayes models a document as a sequence of draws from a per-class word distribution. For a document d with word counts x = (x_1, ..., x_V) over a vocabulary of size V, and class c:

                                    V
log p(c | d)  =  log p(c)  +   sum     x_t * log p(t | c)   +  const
                                  t=1
  • p(c) is the class prior, estimated as the fraction of training documents in class c.
  • p(t | c) is the probability of term t under class c.
  • x_t is the count of term t in the document.
  • const is log p(d), identical across classes, so it drops out of the argmax.

Parameters are estimated by smoothed maximum likelihood:

p(t | c)  =  ( N_tc + alpha ) / ( sum_s N_sc  +  alpha * V )

N_tc is the total count of term t across all training documents of class c. alpha is the additive smoothing constant, defaulting to 1.0 in scikit-learn (Laplace smoothing). Smoothing is not optional: an unsmoothed zero count sends the whole log-probability to negative infinity, so one unseen word would veto a class outright.

Why "naive"

The model assumes term occurrences are conditionally independent given the class. This is plainly false for natural language — "New" and "York" are strongly dependent — and it means the resulting probability estimates are badly calibrated, typically pushed toward 0 or 1.

The ranking, however, survives. Domingos and Pazzani (1997), On the Optimality of the Simple Bayesian Classifier under Zero-One Loss, show that the argmax can be correct over a much wider region than the region where the independence assumption holds. Naive Bayes is a good classifier and a poor probability estimator, and the distinction matters when you threshold on predict_proba.

Event models

McCallum and Nigam (1998), A Comparison of Event Models for Naive Bayes Text Classification, distinguish two formulations that are routinely conflated:

  • Multinomial: features are term counts. Repeated terms contribute repeatedly. Better on larger vocabularies; this is MultinomialNB.
  • Bernoulli: features are binary presence indicators, and crucially the model includes an explicit factor for terms that are absent. This is BernoulliNB, and it can win on very short documents where non-occurrence is informative.

ComplementNB (Rennie et al., 2003, Tackling the Poor Assumptions of Naive Bayes Text Classifiers) estimates parameters from the complement of each class and is measurably more robust under class imbalance. On skewed spam corpora it is often the better default.

Complexity

Let N be the number of documents, L the mean document length, V the vocabulary size and C the number of classes.

  • Training: O(N * L) to accumulate counts, plus O(C * V) to normalise. It is a single pass, it parallelises across documents, and it streams — no document need be held in memory twice.
  • Prediction: O(L_d * C) per document, touching only terms present in that document.
  • Memory: O(C * V) for the parameter table.

This is why naive Bayes survived as a production spam filter through the 2000s at volumes where anything iterative was infeasible.

Practical adjustments that matter

TF-IDF and length normalisation. Rennie et al. document that raw counts overweight repeated terms, and recommend a log transform log(1 + x_t), IDF weighting, and L2 normalisation per document. Applying these to naive Bayes is theoretically inconsistent with the generative story and empirically improves it substantially.

NBSVM. Wang and Manning (2012), Baselines and Bigrams: Simple, Good Sentiment and Topic Classification, use naive Bayes log-count ratios as feature values inside a linear SVM. It remains a stubbornly strong baseline on short-text classification, and it is worth beating before reaching for a transformer.

The hashing trick. HashingVectorizer maps terms into a fixed 2^b space with a signed hash, removing the vocabulary dictionary entirely. This gives constant memory and permits online updating, at the cost of hash collisions and no inverse mapping back to words.

What a production filter actually looks like

Content classification is a minority of the system. A modern mail provider layers:

  1. Connection-level reputation — IP and domain history, which rejects the bulk of volume before the body is ever parsed.
  2. Authentication — SPF, DKIM and DMARC, verifying the sender is who the envelope claims.
  3. Behavioural signals — how recipients across the fleet interact with similar messages, aggregated within minutes.
  4. Content models — gradient-boosted trees or transformer encoders over text, URLs, attachments and rendered HTML.
  5. Near-duplicate detection — locality-sensitive hashing to catch a campaign after its first few thousand deliveries.

Two properties dominate the engineering. First, the adversary adapts: this is adversarial classification, formalised by Dalvi et al. (2004), where the data-generating distribution responds to your classifier. Retraining cadence is a first-class design parameter, not an afterthought.

Second, the loss is asymmetric by orders of magnitude. Filters are tuned to false-positive rates around one in ten thousand or better, accepting substantially more false negatives in exchange. Optimising accuracy, or even F1, targets the wrong quantity entirely.

Attacks to be aware of. Good-word attacks append innocuous high-ham-weight tokens to shift the score (Lowd and Meek, 2005). Poisoning attacks inject crafted training examples through the "report spam" channel itself. Both are why no serious filter trains on unfiltered user feedback without robust aggregation.

Papers

  • Sahami, Dumais, Heckerman and Horvitz, A Bayesian Approach to Filtering Junk E-Mail, AAAI Workshop 1998 — the paper that started this application.
  • McCallum and Nigam, A Comparison of Event Models for Naive Bayes Text Classification, AAAI Workshop 1998.
  • Domingos and Pazzani, On the Optimality of the Simple Bayesian Classifier under Zero-One Loss, Machine Learning 29, 1997.
  • Rennie, Shih, Teevan and Karger, Tackling the Poor Assumptions of Naive Bayes Text Classifiers, ICML 2003.
  • Dalvi, Domingos, Mausam, Sanghai and Verma, Adversarial Classification, KDD 2004.
  • Lowd and Meek, Good Word Attacks on Statistical Spam Filters, CEAS 2005.
  • Metsis, Androutsopoulos and Paliouras, Spam Filtering with Naive Bayes — Which Naive Bayes?, CEAS 2006.
  • Wang and Manning, Baselines and Bigrams, ACL 2012.

What to learn next