Can this even be learned?
Some prediction tasks are impossible no matter the model, and a one-day feasibility probe saves months of tuning against a ceiling nobody measured.
- 8 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Before committing to a project, spend a day checking whether the answer is predictable at all from the data you have.
You can guess what is cooking in the kitchen from the smell in the corridor. You cannot guess next week's lottery numbers from last week's, no matter how long you study them. The difference is not effort or talent — one situation carries a signal and the other does not.
Prediction tasks divide the same way. Some outcomes are written, faintly, in the data you hold. Others are decided by information you do not have, and no model reaches it.
Why it exists
Teams regularly spend six months tuning models on a task that was 70%-predictable at best. Every model lands near 70%, each fancier attempt adds nothing, and nobody can say why. The ceiling was there from day one — unmeasured.
There is a name for this ceiling: the part of the outcome that your features genuinely cannot see. Two customers with identical histories, one leaves, one stays. Whatever separated them lives outside your data.
How it works
A feasibility probe is a deliberately quick experiment:
1. take the data you already have (no new pipelines)
2. train the most ordinary model you know
3. compare against the dumbest guess
4. ask a person to do the task on 50 examples
ordinary model ties the dumb guess → signal is missing: stop, get better data
ordinary model beats the dumb guess → signal exists: the project is real
a person scores like the model → you may already be near the ceilingThe human check matters more than it looks. Suppose experienced staff cannot predict which customer leaves by reading their file. Then the file probably does not contain the answer, and the model reads the same file.
A real example you have seen
Weather apps predict rain hours ahead quite well, and two weeks ahead barely better than the season's average. Nothing is wrong with the models. The atmosphere itself limits how far ahead the signal reaches. Good forecasters know their ceiling; good ML teams measure theirs.
Remember this
- Every task has a ceiling — the accuracy limit set by what the data cannot see.
- A one-day probe with an ordinary model tells you whether a signal exists.
- If a person with the same data cannot do the task, be suspicious the model can.
What to learn next
- Choosing the metric that matches the decision — feasibility says possible; the metric says worthwhile.
- The baselines you must beat first — the probe's dumb guess, formalised.
- Overfitting and underfitting — what capacity does when it runs out of signal.
Developer — Code and libraries.
Setup
pip install scikit-learnOutputs verified with scikit-learn 1.7 and numpy 1.26.
Watching a ceiling refuse to move
The experiment: same task, but we corrupt a known fraction of labels. Then we send a model, and a model ten times bigger, to fight the corruption.
import numpy as np
from sklearn.datasets import make_classification
from sklearn.ensemble import HistGradientBoostingClassifier
from sklearn.model_selection import train_test_split
for noise in [0.0, 0.2, 0.4]:
# flip_y randomly re-assigns that fraction of labels
X, y = make_classification(n_samples=6000, n_features=10,
n_informative=8, n_redundant=0,
flip_y=noise, random_state=0)
X_tr, X_te, y_tr, y_te = train_test_split(X, y, random_state=0)
small = HistGradientBoostingClassifier(max_iter=50, random_state=0)
big = HistGradientBoostingClassifier(max_iter=500, random_state=0)
small.fit(X_tr, y_tr)
big.fit(X_tr, y_tr)
print(f"label noise {noise:.0%}: "
f"small model {small.score(X_te, y_te):.3f} "
f"10x bigger model {big.score(X_te, y_te):.3f}")label noise 0%: small model 0.963 10x bigger model 0.975 label noise 20%: small model 0.850 10x bigger model 0.855 label noise 40%: small model 0.745 10x bigger model 0.728
The walkthrough
At 0% noise, more model buys more accuracy — 0.963 to 0.975. There is unclaimed signal, and capacity claims it.
At 20% noise, ten times the model buys 0.005. The ceiling is doing the talking. The randomness in the labels is not learnable, by anything, ever.
At 40% noise, the bigger model is worse. It spends capacity memorising corrupted labels. Past the ceiling, extra effort turns actively harmful — this is overfitting to noise.
Your real data has an unknown noise setting. You cannot read it off a dial, but this dynamic — every model landing in the same band — is the visible symptom. When three different model families cluster within a point of each other, believe the cluster.
Common mistakes
Blaming the model when the ceiling is at fault. Six weeks of hyperparameter search cannot move 0.85 to 0.95 if the ceiling is 0.86. Cheap test: double your training data and retrain. If accuracy barely moves, more data of the same kind is not the answer — different features are.
Skipping the human benchmark. Give 50 examples to someone who knows the domain. Their score is a rough, cheap ceiling estimate. Ceilings and humans are both imperfect, but a huge model-human gap in either direction is information.
Probing with the fanciest model instead of the fastest. The probe's job is a yes/no answer today, not a leaderboard entry. Gradient boosting on whatever table exists answers "is there signal" in an afternoon.
Trusting one split. Feasibility probes run on small data, and small data makes noisy measurements. Repeat across three seeds before declaring signal or its absence — the lesson on noisy validation shows how big that noise really is.
Try it yourself
Add a DummyClassifier(strategy="most_frequent") to the loop as a floor. Then re-run with n_informative=2 instead of 8. Watch both models sink toward the dummy as the features stop carrying signal — the other way a project dies.
What to learn next
- Choosing the metric that matches the decision — feasibility says possible; the metric says worthwhile.
- The baselines you must beat first — the probe's dumb guess, formalised.
- Overfitting and underfitting — what capacity does when it runs out of signal.
Researcher — Mathematics and papers.
The Bayes error rate
For classification, the irreducible floor is the Bayes error:
$$ \varepsilon^* = \mathbb{E}{x}\left[1 - \max{k} P(Y = k \mid X = x)\right] $$
Where:
- $\varepsilon^*$ — the Bayes error rate, the minimum achievable expected error.
- $P(Y = k \mid X = x)$ — the true conditional class distribution given features $x$.
- The expectation runs over the feature distribution.
No classifier, of any capacity, achieves expected error below $\varepsilon^$ on the same feature set. Crucially, $\varepsilon^$ is a property of $(X, Y)$ jointly: adding features can lower it, more model cannot.
Estimating the unestimable
$\varepsilon^*$ is not directly observable, but it can be bracketed:
- Cover and Hart (1967), Nearest Neighbor Pattern Classification: the asymptotic 1-NN error $\varepsilon_{NN}$ satisfies $\varepsilon^* \leq \varepsilon_{NN} \leq 2\varepsilon^(1 - \varepsilon^)$ for binary problems — a classical sandwich bound.
- Human-level performance as a proxy: for perception tasks where humans are near-Bayes (vision, speech), the human-model gap estimates remaining avoidable bias. This drives the diagnostic in Ng's Machine Learning Yearning (2018): compare training error to human error (avoidable bias) and validation error to training error (variance), and direct effort at the larger gap.
- Modern estimators go through a fitted generative model: Theisen et al. (2021), Evaluating State-of-the-Art Classification Models Against Bayes Optimality (NeurIPS), exploit the invariance of the Bayes error under invertible maps to compute it exactly for a learned normalizing flow, then compare real classifiers against it. The estimate is only as good as the fitted flow, so all such methods inherit their own model assumptions.
Label noise interacts with capacity
Under symmetric binary label noise with flip rate $\eta < 0.5$, the clean Bayes rule stays optimal, but it is now scored against corrupted labels. Its measured accuracy is $(1 - \eta)(1 - \varepsilon^{}) + \eta \varepsilon^{}$: it agrees with the clean label with probability $1 - \varepsilon^{}$ and that label survives the flip with probability $1 - \eta$, plus the symmetric case where both go wrong. At $\varepsilon^{} = 0$ this is exactly $1 - \eta$ — the practical reading: measured accuracy saturates near $1 - \eta$ regardless of model. Worse, deep networks fit noisy labels perfectly given capacity: Zhang et al. (2017), Understanding Deep Learning Requires Rethinking Generalization (ICLR), showed standard architectures reach zero training error on fully randomised labels. Memorisation tends to follow signal learning in time (Arpit et al., 2017, A Closer Look at Memorization in Deep Networks, ICML) — one theoretical motivation for early stopping under noise.
Learning-curve extrapolation
Feasibility also asks "would more data help?". Empirically, generalisation error follows a power law in dataset size $n$: $\varepsilon(n) \approx a n^{-b} + \varepsilon_\infty$ (Hestness et al., 2017, Deep Learning Scaling is Predictable, Empirically). Fitting this on subsets of current data extrapolates the value of collection — and estimates the asymptote $\varepsilon_\infty$, another view of the ceiling. The scaling-law programme (Kaplan et al., 2020) industrialised exactly this reasoning.
What to learn next
- Choosing the metric that matches the decision — feasibility says possible; the metric says worthwhile.
- The baselines you must beat first — the probe's dumb guess, formalised.
- Overfitting and underfitting — what capacity does when it runs out of signal.