Baselines and Choosing a Model

Reading your own errors

An hour spent reading and categorising the examples your model got wrong beats a week of blind tuning, because errors cluster into fixable families.

On this page 5
  1. Why it exists
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Error analysis means personally reading the examples your model got wrong, sorting them into named groups, and fixing the biggest group first.

A cricket coach does not tell a batsman "your average is 32, raise it". The coach rewatches every dismissal and finds the pattern: seven of the last ten were the short ball outside off. Now practice has a target. The average was a number; the dismissals were information.

Your model's wrong answers are its dismissals. The metric says how much it fails. Only reading the failures says how — and "how" is the part you can fix.

Why it exists

Facing a disappointing score, the reflex is to try a bigger model, new parameters, another architecture — blind tuning. It sometimes works, slowly, by accident.

Reading errors works on purpose. Model mistakes are rarely random; they cluster into families with shared causes. One family might be mislabelled data. Another, a missing feature. Another, a group of cases nobody realised existed. Each family has a different fix, and no global knob addresses any of them.

How it works

The procedure fits in an afternoon:

1. collect 50-100 wrong predictions from validation data
2. READ each one, like a human, one at a time
3. tag each with a short cause: "negation", "blurry photo", "label wrong"
4. count the tags

   negation      ████████████  24   ← fix this first
   label wrong   ██████  13
   sarcasm       ████  8
   rare words    ██  5

5. fix the biggest bar, retrain, repeat

The counting step converts complaints into a to-do list ordered by value. If "label wrong" is a big bar, your best move is cleaning data, and no model change would have found that.

A real example you have seen

Every serious exam student does this with mock tests: sort lost marks into "silly mistakes", "didn't know the concept", "ran out of time". Each bucket demands a different fix — care, study, or speed practice. Students who skip the sorting repeat the same losses. So do teams.

Remember this

  • Read the wrong answers yourself. The metric cannot tell you what it saw.
  • Tag errors with causes, then count — the biggest bar is your next task.
  • Different bars need different fixes: data cleaning, features, model. Tuning fixes almost none of them.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install scikit-learn

Outputs verified with scikit-learn 1.7 and numpy 1.26.

A sentiment model, and the reading of its sins

Restaurant reviews, a classic small classifier, and the part most people skip: looking.

read_the_errors.py
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.linear_model import LogisticRegression
from sklearn.pipeline import make_pipeline

train_texts = [
    "loved the food", "great service and taste", "delicious biryani",
    "amazing flavours", "fresh and tasty", "wonderful experience",
    "best dosa in town", "really enjoyed the meal", "superb paneer",
    "excellent thali", "crispy and hot", "tasty and cheap",
    "terrible food", "cold and stale", "worst service ever",
    "awful experience", "rude staff", "dirty tables",
    "bland and oily", "overpriced and slow", "burnt roti",
    "smelly fish", "waited an hour", "disgusting chutney",
]
train_labels = [1] * 12 + [0] * 12

test_texts = [
    "tasty food and quick service",      # plain positive
    "stale bread and cold curry",        # plain negative
    "not tasty at all",                  # negation
    "not bad, would come back",          # negation
    "the food was fine i guess",         # lukewarm
    "great, another hour in the queue",  # sarcasm
]
test_labels = [1, 0, 0, 1, 1, 0]

model = make_pipeline(TfidfVectorizer(), LogisticRegression())
model.fit(train_texts, train_labels)

probs = model.predict_proba(test_texts)[:, 1]
print("p(positive)  true  verdict  text")
for text, label, p in zip(test_texts, test_labels, probs):
    verdict = "ok" if (p >= 0.5) == label else "WRONG"
    print(f"{p:11.2f}  {label:4d}  {verdict:7s}  {text}")
Output
p(positive)  true  verdict  text
       0.56     1  ok       tasty food and quick service
       0.42     0  ok       stale bread and cold curry
       0.61     0  WRONG    not tasty at all
       0.49     1  WRONG    not bad, would come back
       0.55     1  ok       the food was fine i guess
       0.57     0  WRONG    great, another hour in the queue

The walkthrough

The metric says "50% wrong on hard cases". The reading says why. Both negation cases failed, in both directions: "not tasty" scored positive 0.61, "not bad" scored negative 0.49. One tag — "negation" — covers two of three errors. The sarcastic "great, another hour" failed for the related reason: "great" votes positive, and context cannot overrule it.

The cause is architectural, and the reading found it in one minute. TF-IDF treats each word as an independent vote. "not" is one vote among five, not a sign-flipper. No amount of hyperparameter tuning changes that — the fix is representation: word pairs (ngram_range=(1, 2)), or a model that reads order, like BERT.

The probabilities add a second dimension to the reading. All three errors sit near 0.5 — the model is hesitantly wrong, not confidently wrong. Hesitant errors suit a review band; confident errors (0.95, wrong) usually point at label problems or leakage. Sort your reading by confidence and the worst diseases float to the top.

Six test rows is a toy, and the method is not. At work the loop is identical with 100 rows: read, tag, count, fix the biggest tag.

Common mistakes

Sampling errors non-randomly. Reading only the first N errors, alphabetically or by date, biases the tag counts. Shuffle before sampling, or the biggest bar may be an artifact of ordering.

Tags that describe symptoms, not causes. "Model said positive, was negative" is the error itself, restated. "Contains negation" is a cause with a fix attached. If you cannot name the fix, keep staring — or hand the case to a domain expert.

Fixing a tiny bar because it is interesting. Sarcasm is fascinating and 8 cases; negation is boring and 24. The count exists to defend you from your own curiosity.

Never suspecting the label. In mature datasets, a fat error bar often turns out to be annotation mistakes — the model was right and the "truth" was wrong. Every reading session doubles as a label audit.

Try it yourself

Fix the biggest bar: change TfidfVectorizer() to TfidfVectorizer(ngram_range=(1, 2)), retrain on the same data, and re-print the table. "not tasty" and "not bad" become learnable pairs. Check which errors flip — and whether anything else broke, because fixes have side effects and the table shows those too.

What to learn next

Researcher — Mathematics and papers.

The methodology, formalised

Error analysis estimates the conditional error decomposition: partition the error set $E = {i : \hat{y}_i \neq y_i}$ by cause tags $c(i) \in {1..K}$ and prioritise by expected reduction:

$$ \Delta_k = \frac{|E_k|}{n} \cdot \rho_k $$

Where $|E_k|$ is the count of errors with tag $k$, $n$ the evaluation size, and $\rho_k \in [0,1]$ the plausible fix rate for that cause. Ng's Machine Learning Yearning (2018) prescribes exactly this ceiling analysis: the tag count bounds the metric gain of any fix, preventing engineering effort on categories that cannot repay it.

Tooling and systematisation

The manual practice has research descendants:

  • Errudite (Wu et al., 2019, ACL) — a domain-specific language for reproducible, dataset-scale error grouping in NLP, replacing ad-hoc spreadsheets.
  • Domino (Eyuboglu et al., 2022, ICLR) — discovers coherent error slices in embedding space and names them with natural-language descriptions.
  • CheckList (Ribeiro et al., 2020, ACL) — inverts the flow: pre-registered behavioural test categories (negation among them) probed with generated minimal pairs; the negation failure in the developer block is a canonical CheckList finding for bag-of-words models.
  • Label-error detection at scale: confident learning (Northcutt et al., 2021, JAIR) estimates the joint distribution of noisy and true labels from model confidences — automating the "never suspect the label" audit; their companion study found ~3.3% label errors across ten standard benchmark test sets.

Confidence-stratified reading

Partitioning errors by predicted confidence targets distinct pathologies: high-confidence errors correlate with label noise and spurious features (the model has strong, wrong evidence), near-threshold errors with genuine ambiguity or insufficient capacity. This is the practical use of calibration information before any recalibration: miscalibrated confidence still ranks (Guo et al., 2017, ICML, on modern networks' overconfidence). Influence-function methods (Koh & Liang, 2017, ICML) extend the reading upstream — for a given error, which training points most caused it — closing the loop from symptom to responsible data.

Limits of the afternoon method

Tag taxonomies are analyst-dependent (inter-annotator agreement applies to your own tags), counts on 100 sampled errors carry multinomial noise (a 24-vs-13 gap on 50 tags is ~2 standard errors), and slice discovery by manual reading misses non-salient groupings — the systematic complement is slice-based evaluation, which starts from metadata instead of eyeballs.

What to learn next