Imbalanced, Multi-class and Multi-label

Macro, micro and weighted averaging

The same predictions can score 0.80 or 0.54 depending on one argument — because micro averaging counts rows and macro averaging counts classes.

Read these first

On this page 5
  1. Why it exists
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Macro, micro and weighted are three ways to average per-class scores into one number — and on imbalanced data they disagree loudly.

Think of rating a thali restaurant. Micro says: taste every spoonful on the plate and average — the giant mound of rice dominates the verdict. Macro says: give the rice, the dal and the tiny pickle one vote each, no matter the portion size. Same thali, two honest, different verdicts.

Why it exists

With many classes, each class earns its own score. A support-ticket model might do well on "billing" tickets, decently on "bug" tickets and terribly on rare "refund" tickets. Three scores. Your dashboard wants one.

Averaging them requires a decision nobody warns you about:

  • Micro averaging pools every individual prediction first, then scores the pool. Big classes dominate, because they contribute more predictions.
  • Macro averaging scores each class separately, then averages those scores equally. The rare class counts as much as the common one.
  • Weighted averaging is macro, but each class's score is weighted by its size. A compromise — mostly leaning micro.

How it works

per-class scores:   billing 0.87   bug 0.75   refund 0.00

micro    → pool every row first        → 0.80  (billing's bulk hides the refund disaster)
macro    → average the three scores    → 0.54  (the refund failure screams)
weighted → average, sized by class     → 0.76  (mostly hides it again)

One model, one set of predictions, three "F1 scores". None is wrong. They answer different questions: "how are the rows doing?" versus "how are the classes doing?"

A real example you have seen

Voice assistants understand common commands ("set an alarm") almost perfectly and rare ones ("split this bill three ways") poorly. Judged per attempt — micro — they look excellent, because most attempts are common commands. Judged per skill — macro — the gaps show. Which number a team reports shapes what gets fixed.

Remember this

  • Micro: every row votes. Big classes dominate.
  • Macro: every class votes once. Rare-class failures become visible.
  • On imbalanced data, report macro (or the full per-class table), or rare failures stay hidden.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install scikit-learn

Outputs verified with scikit-learn 1.7.2, CPU only.

Three averages, one confusion

averaging.py
from sklearn.metrics import classification_report, f1_score

# 20 support tickets: 0 = billing (common), 1 = bug, 2 = refund (rare)
y_true = [0,0,0,0,0,0,0,0,0,0,0,0,0,0,1,1,1,1,2,2]
y_pred = [0,0,0,0,0,0,0,0,0,0,0,0,0,1,1,1,1,0,0,0]  # refunds all missed

for avg in ["micro", "macro", "weighted"]:
    print(f"{avg:8s} F1 = {f1_score(y_true, y_pred, average=avg):.2f}")
print()
print(classification_report(y_true, y_pred,
      target_names=["billing", "bug", "refund"], zero_division=0))
Output
micro    F1 = 0.80
macro    F1 = 0.54
weighted F1 = 0.76

              precision    recall  f1-score   support

     billing       0.81      0.93      0.87        14
         bug       0.75      0.75      0.75         4
      refund       0.00      0.00      0.00         2

    accuracy                           0.80        20
   macro avg       0.52      0.56      0.54        20
weighted avg       0.72      0.80      0.76        20

The walkthrough

0.80 versus 0.54 — from the same predictions. The model never once identified a refund ticket. Micro F1 barely notices (two rows out of twenty). Macro F1 charges a full third of its value. If refunds matter to your business, only one of these numbers is telling you the truth.

Micro F1 equals accuracy here. For single-label multiclass problems, every wrong prediction is one false positive and one false negative at the same time, so pooled precision, pooled recall, micro F1 and accuracy all collapse into the same number. Seeing micro F1 = accuracy is expected, not a bug.

support is the row count per class — 14, 4, 2. Weighted averaging uses exactly these as weights, which is why it lands near micro.

zero_division=0. Refund has no predicted samples, so its precision divides by zero. Without the argument, sklearn prints the score as 0 and warns. Set it explicitly to declare the behaviour you intend.

Common mistakes

Reporting one number without naming the average. "F1 = 0.80" is meaningless in a many-class paper or dashboard. Name the averaging every time; better, show the per-class table.

Optimising macro F1 with untouched thresholds. Macro F1 rewards fixing rare classes. Combine it with class weights or per-class thresholds; otherwise the model keeps grinding on the common classes where easy gains live.

Using weighted averaging to "handle imbalance". It does the opposite — it re-buries rare classes under their small weights. Weighted exists to summarise, not to protect.

Averaging averages across datasets. Macro F1 from two test sets cannot be averaged into a combined macro F1; class supports differ. Recompute from pooled predictions, or report per-dataset.

Try it yourself

Fix two predictions so one refund ticket is caught. Recompute all three averages and note which moved most. Then compute balanced_accuracy_score — macro-averaged recall — and match it against the report's macro recall of 0.56.

What to learn next

Researcher — Mathematics and papers.

Definitions

With per-class counts TP_k, FP_k, FN_k for classes k = 1…K, per-class precision P_k = TP_k/(TP_k+FP_k) and recall R_k = TP_k/(TP_k+FN_k):

  • Micro: P_micro = ΣTP_k / Σ(TP_k+FP_k), similarly R_micro; F1_micro is their harmonic mean. Pool counts, then compute.
  • Macro: F1_macro = (1/K) Σ_k F1_k. Compute per class, then average unweighted.
  • Weighted: Σ_k (n_k/n) F1_k with class supports n_k. Average weighted by prevalence.

For single-label multiclass, Σ FP_k = Σ FN_k (each error is misassignment), giving P_micro = R_micro = F1_micro = accuracy. The identity breaks for multi-label problems, where micro metrics regain independent meaning.

Estimator subtleties

Two macro-F1 conventions exist: averaging per-class F1 (sklearn's) versus the harmonic mean of macro-P and macro-R. Opitz and Burst (2019), Macro F1 and macro F1, show gaps up to 0.2 between them on skewed data — check which one a paper used before comparing numbers. Macro F1 with empty predicted classes inherits the 0/0 convention (zero_division), which alone can shift reported scores by several points on long-tailed benchmarks.

As a decision-theoretic target, micro-F1 optimisation reduces to accuracy's plug-in rule, while macro-F1 has no per-row decomposition: the optimal classifier depends on the joint distribution through per-class thresholds. The Optimal Thresholding characterisation (Ye et al., 2012, Optimizing F-measures: a tale of two approaches, ICML; Koyejo et al., 2014, NeurIPS) shows population-optimal macro-F1 decisions threshold each class's posterior at half its optimal F1 value — the theory behind per-class threshold tuning.

Choosing an average, defensibly

  • Rows are the unit of value (ad clicks, per-query cost): micro.
  • Classes are the unit of value (diseases, intents, long-tail products): macro.
  • Chance-corrected alternatives: Cohen's κ and Matthews correlation (multiclass MCC; Gorodkin, 2004) resist both imbalance and the free lunch of majority guessing, at the cost of interpretability.
  • Long-tailed benchmarks (iNaturalist, LVIS) report per-group macro metrics (head/mid/tail splits) precisely because single averages hide tail collapse.

Statistical comparison of two models should pair predictions per row (McNemar) or bootstrap the full metric; macro-F1 differences on small rare-class supports have wide intervals — the refund class above has n = 2, so its F1 estimate is essentially binary.

What to learn next