Scoping an ML Project

Choosing the label

The label is the definition of truth your model learns from, and two reasonable definitions of the same word can disagree on most of your data.

On this page 5
  1. Why it exists
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The label is your written definition of the truth, and the model will learn exactly that definition — including its flaws.

Put two workers at a mango crate and tell them "keep the ripe ones". One judges by colour, the other by softness. By evening you have two different piles from the same crate, and both workers followed your instruction.

"Ripe" was never defined. In machine learning, the label is the answer attached to each training example. It is that definition, written down so precisely that a machine can apply it millions of times.

Why it exists

Words like churn, fraud, spam, and quality feel obvious until you must compute them. Does a customer churn when they cancel? When they stop opening the app? When they stop paying? Each choice creates a different column of 0s and 1s, and therefore a different model.

The model never learns "churn". It learns your column. If the column is a bad translation of the business meaning, the model is fluent in the wrong language.

How it works

Take one word and watch it split:

"churn"
   ├─ definition A: pressed the cancel button
   │     → clean signal, but misses people who drift away silently
   └─ definition B: 30 days without activity
         → catches drifters, but needs a 30-day wait to compute

Neither is wrong. They disagree, they arrive at different times, and they cost different amounts to compute. Choosing between them is a real decision, not paperwork.

Good label checks before training anything:

  • Can two people apply the definition and get the same answer?
  • When does the label become known — instantly, or after a wait?
  • Who wrote the historical labels, and did they have a reason to bend them?

A real example you have seen

Email spam. One person's "spam" is another person's newsletter they forgot subscribing to. Gmail's label comes from millions of people pressing "report spam" — a definition with moods, mistakes, and revenge-reporting baked in. The model inherits all of it.

Remember this

  • The label is the definition of truth. The model learns the definition, not the intention.
  • Two reasonable definitions of the same word can disagree on much of your data.
  • Ask when the label is knowable, who produced it, and whether two people would agree.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install pandas

Outputs verified with pandas 2.2.

Two honest definitions, one dataset

Both definitions of churn below are defensible. Watch how far apart they land.

two_labels.py
import pandas as pd

today = pd.Timestamp("2025-09-01")
customers = pd.DataFrame({
    "customer": ["asha", "bilal", "chen", "deepa", "evan",
                 "fatima", "gopal", "hana", "ivan", "jaya"],
    "cancelled": [0, 1, 0, 0, 1, 0, 0, 0, 1, 0],
    "last_active": pd.to_datetime([
        "2025-08-30", "2025-08-29", "2025-06-02", "2025-07-10", "2025-05-15",
        "2025-08-25", "2025-06-20", "2025-08-31", "2025-08-28", "2025-07-01"]),
})

# Definition A: they pressed the cancel button.
customers["churn_a"] = customers["cancelled"]

# Definition B: no activity for 30 days or more.
gap = (today - customers["last_active"]).dt.days
customers["churn_b"] = (gap >= 30).astype(int)

print(pd.crosstab(customers["churn_a"], customers["churn_b"],
                  rownames=["pressed cancel"], colnames=["30 days inactive"]))
agree = (customers["churn_a"] == customers["churn_b"]).mean()
print(f"\nthe two labels agree on {agree:.0%} of customers")
print(customers[customers["churn_a"] != customers["churn_b"]][
    ["customer", "cancelled", "last_active"]].to_string(index=False))
Output
30 days inactive  0  1
pressed cancel        
0                 3  4
1                 2  1

the two labels agree on 40% of customers
customer  cancelled last_active
   bilal          1  2025-08-29
    chen          0  2025-06-02
   deepa          0  2025-07-10
   gopal          0  2025-06-20
    ivan          1  2025-08-28
    jaya          0  2025-07-01

The walkthrough

40% agreement. Same customers, same word, and the two truth columns disagree on six people out of ten. Any model comparison across these labels is meaningless — they are different tasks.

Bilal and Ivan cancelled while still active. Definition A calls them churned; definition B says they are fine. Perhaps they cancelled auto-renew but kept using the service until expiry. The business meaning of these two rows is a product question, not a data question.

Chen, Gopal, and Jaya drifted away without cancelling. Definition A misses them entirely. If the retention team only calls "churned" customers, definition A means the silent leavers never get a call.

The crosstab is the tool. Before committing to a label, compute it two or three defensible ways and cross-tabulate. Disagreement cells are not noise — each one is a policy question wearing a data costume.

Common mistakes

Using the label that is easiest to query. cancelled is one column; inactivity needs a window and a wait. Teams pick the easy one, then wonder why the model misses silent churn.

Labels made by the process you want to change. Training "who is a fraudster" on "who our old rules flagged" teaches the model the old rules, blind spots included.

Thresholds that manufacture borderline noise. Turning "days inactive" into a 0/1 at exactly 30 days makes day-29 and day-31 customers opposite classes. Where possible, predict the continuous quantity and let the decision layer apply the threshold.

Never versioning the label. Definitions change — someone widens the window from 30 to 45 days. If the label logic is not versioned alongside the data, old and new training sets silently mix two truths.

Try it yourself

Add definition C: cancelled OR 45 days inactive. Compute the three-way agreement. Then decide which single definition you would ship if the retention team can call 50 customers a week — and write one sentence defending it.

What to learn next

Researcher — Mathematics and papers.

Label choice as proxy selection

The business cares about an unobservable construct $Y^*$ (true customer abandonment, true fraud intent). Any computable label $Y$ is a proxy with its own error structure:

$$ P(Y \neq Y^* \mid X = x) = \alpha(x) \cdot \mathbb{1}[Y^=1] + \beta(x) \cdot \mathbb{1}[Y^=0] $$

Where:

  • $Y^*$ — the latent construct of interest.
  • $Y$ — the operationalised label actually trained on.
  • $\alpha(x)$ — the miss rate: true positives the definition fails to capture.
  • $\beta(x)$ — the false-alarm rate of the definition itself.
  • $\mathbb{1}[\cdot]$ — the indicator function, 1 when the condition holds, else 0.

The dangerous case is feature-dependent label error: $\alpha(x)$ varying with $x$. Class-conditional noise only shifts decision thresholds (Natarajan et al., 2013, Learning with Noisy Labels, NeurIPS), but instance-dependent noise biases the learned decision boundary itself, and no reweighting fixes it without a noise model.

Documented proxy failures

  • Obermeyer et al. (2019), Dissecting racial bias in an algorithm used to manage the health of populations (Science): healthcare cost used as a proxy for healthcare need. Cost is measurable; need is the construct. Because less money was historically spent on Black patients at equal need, the proxy encoded the disparity and the model amplified it.
  • Jacobs and Wallach (2021), Measurement and Fairness (FAccT), import measurement-theory vocabulary — construct validity, reliability — and argue most fairness disputes are disputes about $Y \neq Y^*$.

Inter-annotator agreement as a label ceiling

When labels come from human annotators, chance-corrected agreement bounds achievable performance. Cohen's kappa for two raters:

$$ \kappa = \frac{p_o - p_e}{1 - p_e} $$

Where $p_o$ is observed agreement and $p_e$ the agreement expected by chance from the raters' marginals. A model evaluated against a single rater cannot meaningfully exceed the raters' own agreement; apparent super-human scores usually mean the test labels share one rater's quirks. Aggregation schemes (majority vote, Dawid-Skene, 1979) trade annotator noise for systematic smoothing of genuine ambiguity — Aroyo and Welty (2015), Truth Is a Lie, argue disagreement is often signal about the construct, not error.

Practical mitigations

  • Predict continuous outcomes where possible; thresholds belong in the decision layer where costs live (see choosing the metric).
  • Hold out a small, expensively adjudicated "gold" set to estimate $\alpha, \beta$ for the cheap label.
  • Version label definitions as code; a label change is a dataset migration, not an edit.

What to learn next