Scoping an ML Project

Can this even be learned?

Some prediction tasks are impossible no matter the model, and a one-day feasibility probe saves months of tuning against a ceiling nobody measured.

On this page 5
  1. Why it exists
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Before committing to a project, spend a day checking whether the answer is predictable at all from the data you have.

You can guess what is cooking in the kitchen from the smell in the corridor. You cannot guess next week's lottery numbers from last week's, no matter how long you study them. The difference is not effort or talent — one situation carries a signal and the other does not.

Prediction tasks divide the same way. Some outcomes are written, faintly, in the data you hold. Others are decided by information you do not have, and no model reaches it.

Why it exists

Teams regularly spend six months tuning models on a task that was 70%-predictable at best. Every model lands near 70%, each fancier attempt adds nothing, and nobody can say why. The ceiling was there from day one — unmeasured.

There is a name for this ceiling: the part of the outcome that your features genuinely cannot see. Two customers with identical histories, one leaves, one stays. Whatever separated them lives outside your data.

How it works

A feasibility probe is a deliberately quick experiment:

1. take the data you already have  (no new pipelines)
2. train the most ordinary model you know
3. compare against the dumbest guess
4. ask a person to do the task on 50 examples

ordinary model ties the dumb guess    → signal is missing: stop, get better data
ordinary model beats the dumb guess   → signal exists: the project is real
a person scores like the model        → you may already be near the ceiling

The human check matters more than it looks. Suppose experienced staff cannot predict which customer leaves by reading their file. Then the file probably does not contain the answer, and the model reads the same file.

A real example you have seen

Weather apps predict rain hours ahead quite well, and two weeks ahead barely better than the season's average. Nothing is wrong with the models. The atmosphere itself limits how far ahead the signal reaches. Good forecasters know their ceiling; good ML teams measure theirs.

Remember this

  • Every task has a ceiling — the accuracy limit set by what the data cannot see.
  • A one-day probe with an ordinary model tells you whether a signal exists.
  • If a person with the same data cannot do the task, be suspicious the model can.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install scikit-learn

Outputs verified with scikit-learn 1.7 and numpy 1.26.

Watching a ceiling refuse to move

The experiment: same task, but we corrupt a known fraction of labels. Then we send a model, and a model ten times bigger, to fight the corruption.

ceiling_probe.py
import numpy as np
from sklearn.datasets import make_classification
from sklearn.ensemble import HistGradientBoostingClassifier
from sklearn.model_selection import train_test_split

for noise in [0.0, 0.2, 0.4]:
    # flip_y randomly re-assigns that fraction of labels
    X, y = make_classification(n_samples=6000, n_features=10,
                               n_informative=8, n_redundant=0,
                               flip_y=noise, random_state=0)
    X_tr, X_te, y_tr, y_te = train_test_split(X, y, random_state=0)
    small = HistGradientBoostingClassifier(max_iter=50, random_state=0)
    big = HistGradientBoostingClassifier(max_iter=500, random_state=0)
    small.fit(X_tr, y_tr)
    big.fit(X_tr, y_tr)
    print(f"label noise {noise:.0%}:  "
          f"small model {small.score(X_te, y_te):.3f}   "
          f"10x bigger model {big.score(X_te, y_te):.3f}")
Output
label noise 0%:  small model 0.963   10x bigger model 0.975
label noise 20%:  small model 0.850   10x bigger model 0.855
label noise 40%:  small model 0.745   10x bigger model 0.728

The walkthrough

At 0% noise, more model buys more accuracy — 0.963 to 0.975. There is unclaimed signal, and capacity claims it.

At 20% noise, ten times the model buys 0.005. The ceiling is doing the talking. The randomness in the labels is not learnable, by anything, ever.

At 40% noise, the bigger model is worse. It spends capacity memorising corrupted labels. Past the ceiling, extra effort turns actively harmful — this is overfitting to noise.

Your real data has an unknown noise setting. You cannot read it off a dial, but this dynamic — every model landing in the same band — is the visible symptom. When three different model families cluster within a point of each other, believe the cluster.

Common mistakes

Blaming the model when the ceiling is at fault. Six weeks of hyperparameter search cannot move 0.85 to 0.95 if the ceiling is 0.86. Cheap test: double your training data and retrain. If accuracy barely moves, more data of the same kind is not the answer — different features are.

Skipping the human benchmark. Give 50 examples to someone who knows the domain. Their score is a rough, cheap ceiling estimate. Ceilings and humans are both imperfect, but a huge model-human gap in either direction is information.

Probing with the fanciest model instead of the fastest. The probe's job is a yes/no answer today, not a leaderboard entry. Gradient boosting on whatever table exists answers "is there signal" in an afternoon.

Trusting one split. Feasibility probes run on small data, and small data makes noisy measurements. Repeat across three seeds before declaring signal or its absence — the lesson on noisy validation shows how big that noise really is.

Try it yourself

Add a DummyClassifier(strategy="most_frequent") to the loop as a floor. Then re-run with n_informative=2 instead of 8. Watch both models sink toward the dummy as the features stop carrying signal — the other way a project dies.

What to learn next

Researcher — Mathematics and papers.

The Bayes error rate

For classification, the irreducible floor is the Bayes error:

$$ \varepsilon^* = \mathbb{E}{x}\left[1 - \max{k} P(Y = k \mid X = x)\right] $$

Where:

  • $\varepsilon^*$ — the Bayes error rate, the minimum achievable expected error.
  • $P(Y = k \mid X = x)$ — the true conditional class distribution given features $x$.
  • The expectation runs over the feature distribution.

No classifier, of any capacity, achieves expected error below $\varepsilon^$ on the same feature set. Crucially, $\varepsilon^$ is a property of $(X, Y)$ jointly: adding features can lower it, more model cannot.

Estimating the unestimable

$\varepsilon^*$ is not directly observable, but it can be bracketed:

  • Cover and Hart (1967), Nearest Neighbor Pattern Classification: the asymptotic 1-NN error $\varepsilon_{NN}$ satisfies $\varepsilon^* \leq \varepsilon_{NN} \leq 2\varepsilon^(1 - \varepsilon^)$ for binary problems — a classical sandwich bound.
  • Human-level performance as a proxy: for perception tasks where humans are near-Bayes (vision, speech), the human-model gap estimates remaining avoidable bias. This drives the diagnostic in Ng's Machine Learning Yearning (2018): compare training error to human error (avoidable bias) and validation error to training error (variance), and direct effort at the larger gap.
  • Modern estimators go through a fitted generative model: Theisen et al. (2021), Evaluating State-of-the-Art Classification Models Against Bayes Optimality (NeurIPS), exploit the invariance of the Bayes error under invertible maps to compute it exactly for a learned normalizing flow, then compare real classifiers against it. The estimate is only as good as the fitted flow, so all such methods inherit their own model assumptions.

Label noise interacts with capacity

Under symmetric binary label noise with flip rate $\eta < 0.5$, the clean Bayes rule stays optimal, but it is now scored against corrupted labels. Its measured accuracy is $(1 - \eta)(1 - \varepsilon^{}) + \eta \varepsilon^{}$: it agrees with the clean label with probability $1 - \varepsilon^{}$ and that label survives the flip with probability $1 - \eta$, plus the symmetric case where both go wrong. At $\varepsilon^{} = 0$ this is exactly $1 - \eta$ — the practical reading: measured accuracy saturates near $1 - \eta$ regardless of model. Worse, deep networks fit noisy labels perfectly given capacity: Zhang et al. (2017), Understanding Deep Learning Requires Rethinking Generalization (ICLR), showed standard architectures reach zero training error on fully randomised labels. Memorisation tends to follow signal learning in time (Arpit et al., 2017, A Closer Look at Memorization in Deep Networks, ICML) — one theoretical motivation for early stopping under noise.

Learning-curve extrapolation

Feasibility also asks "would more data help?". Empirically, generalisation error follows a power law in dataset size $n$: $\varepsilon(n) \approx a n^{-b} + \varepsilon_\infty$ (Hestness et al., 2017, Deep Learning Scaling is Predictable, Empirically). Fitting this on subsets of current data extrapolates the value of collection — and estimates the asymptote $\varepsilon_\infty$, another view of the ceiling. The scaling-law programme (Kaplan et al., 2020) industrialised exactly this reasoning.

What to learn next