Why data quality decides everything
Your data sets a ceiling on how good a model can ever be, and the model choice only decides how close you get to that ceiling.
- 12 min read
- 3 reading levels
- Updated
Read these first
On this page 7
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The data you feed a model sets a ceiling on how good it can ever get. The model you pick only decides how close you come to that ceiling.
Think about cooking dinner for your family. You can hire a five-star chef and hand over your kitchen. If the dal is stale, the milk has turned and the salt jar is full of sugar, the meal will be bad.
Swapping the chef does not fix any of that. Fixing the ingredients does.
Models work the same way. A better algorithm cannot recover information that was never in the data.
Why anyone needs telling this
Model names are exciting. Data work is not.
Papers get written about new architectures. Nobody writes a paper about a mismatched column of incomes. So teams reach for the exciting lever first, and keep pulling it for months.
The uncomfortable finding, over and over, is that the boring lever moves further. Teams that stop tuning and start reading their own rows tend to gain more, faster, for less money.
What bad data actually looks like
It is rarely dramatic. It looks like this:
- Half your incomes were typed in rupees, half in thousands of rupees.
- The same customer appears three times because an upload was retried.
- A blank cell means "not asked" in one region and "answered zero" in another.
- Someone changed a form field last March and told nobody.
Not one of those is a hard problem. Every one of them quietly costs you more accuracy than switching to a fancier model would give you.
How the two levers compare
dirty data fixed data
| |
+------------+------------+ +-----------+-----------+
| | | |
simple model fancy model simple model fancy model
| | | |
poor slightly GOOD also fine
less poorMoving down that picture is a model change. Moving across is a data change. Across is the bigger step, and it is nearly always the cheaper one.
Somewhere you have seen this
A delivery app that keeps sending your parcel to the wrong locality is rarely running a weak model. Somebody stored two different pin codes for the same street, and both are still in the system.
A hospital dashboard reporting impossible blood pressure readings has the same shape of problem. A machine wrote a zero when a nurse skipped a measurement, and the average was taken over those zeros.
What is honestly hard here
Cleaning data is dull, and it never finishes. Nobody claps for it.
It is also hard to know when you are done, because there is no score for "the data is now good". You have to invent your own checks, and then keep running them forever.
If you want a role where you are praised for a clever idea every week, this is not it. If you want the work that decides whether the product survives contact with real users, this is exactly it.
Remember this
- Data sets the ceiling. The model only decides how close you get to it.
- Most real data problems are boring: wrong units, repeated rows, silent blanks.
- Fixing one of those usually beats every model you were considering.
What to learn next
- Collecting data — where the ceiling gets set in the first place.
- Cleaning data — the specific bugs, and how to catch them.
- Feature engineering — raising the ceiling by building better columns.
Developer — Code and libraries.
Assertions are cheap. Here is the claim measured, on a dataset small enough to fit on this page.
We build a loan-repayment table carrying one realistic bug: four rows in ten had income typed in rupees, and the rest in thousands of rupees. Nobody noticed, because both columns are numbers and both look plausible.
Then we compare the two levers. Upgrade the model, or fix the column.
Setup
pip install numpy pandas scikit-learnThis runs on a CPU in about two seconds and downloads nothing. The numbers below come from scikit-learn 1.7 and NumPy 1.26; a different version can move the last digit.
The experiment
import numpy as np
import pandas as pd
from sklearn.ensemble import RandomForestClassifier
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import roc_auc_score
from sklearn.model_selection import train_test_split
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler
rng = np.random.default_rng(0)
n = 600
income_k = rng.gamma(4.0, 15.0, n) + 10 # true income, in thousands of rupees
years_job = rng.integers(0, 20, n)
truth = 0.06 * income_k + 0.15 * years_job - 4.0 # the rule the world actually follows
repaid = rng.binomial(1, 1 / (1 + np.exp(-truth)))
# The bug: four rows in ten were typed in rupees, the rest in thousands of rupees.
mixed = np.where(rng.random(n) < 0.4, income_k * 1000, income_k)
dirty = pd.DataFrame({"income": mixed, "years_job": years_job, "repaid": repaid})
# The fix: one line. Anything above 1000 was entered in rupees, not thousands.
fixed = dirty.assign(income=np.where(dirty["income"] > 1000,
dirty["income"] / 1000,
dirty["income"]))
def auc(frame, model):
"""Train on 70% of the rows, score the held-out 30%."""
X_tr, X_te, y_tr, y_te = train_test_split(
frame[["income", "years_job"]], frame["repaid"],
test_size=0.3, random_state=42, stratify=frame["repaid"])
model.fit(X_tr, y_tr)
return round(roc_auc_score(y_te, model.predict_proba(X_te)[:, 1]), 3)
logreg = lambda: make_pipeline(StandardScaler(), LogisticRegression(max_iter=1000))
forest = lambda: RandomForestClassifier(n_estimators=300, random_state=0)
print("dirty data + logistic regression :", auc(dirty, logreg()))
print("dirty data + random forest :", auc(dirty, forest()))
print("fixed data + logistic regression :", auc(fixed, logreg()))
print("fixed data + random forest :", auc(fixed, forest()))dirty data + logistic regression : 0.672 dirty data + random forest : 0.694 fixed data + logistic regression : 0.851 fixed data + random forest : 0.752
Read the four numbers slowly
The model upgrade bought 0.022. Logistic regression to a 300-tree random forest, same dirty data: 0.672 to 0.694. That is the lever everybody pulls first.
The data fix bought 0.179. Same simple model, one corrected column: 0.672 to 0.851. Eight times the gain, from three lines of pandas.
The fancy model on clean data is worse than the simple one. 0.752 against 0.851. The real relationship here is close to a straight line, and the forest spends its capacity carving that line into steps. A more powerful model is not a safer default.
AUC is the score used throughout: the chance that a randomly picked repayer is ranked above a randomly picked defaulter. 0.5 is a coin toss and 1.0 is perfect. It is used here because it does not depend on where you set the approval cut-off.
Why mixed units hurt this much
The model is asked to learn one relationship between income and repayment. In the dirty column, two different relationships sit stacked on top of each other, a thousand times apart in scale.
No single straight line fits both. Fitting one throws away most of the signal in a genuinely useful column.
The random forest recovers a little, because a tree can split at income > 1000 and then treat each group separately. It has to spend depth doing that, and what it learns is your bug rather than the world.
Common mistakes
Fixing the column only in the training script. The same rupee rows will arrive tomorrow at prediction time. If the fix lives in a notebook rather than in the pipeline, production still sees the bug. That gap has a name and a lesson: feature stores.
Cleaning with a threshold you never checked. income > 1000 works here because the true range sits well below it. On another table that line would cut a real population in half. Print the histogram before you pick a cut-off.
Treating one AUC as proof. With 600 rows a single split moves by a few points on the random seed. Re-run with several values of random_state before believing a gap of 0.01. A gap of 0.18, as above, is safe.
Cleaning with the test set included. Compute a median, a threshold or a scaler over all rows and then split, and information has leaked backwards. Split first. See train, test and validation splits.
Try it yourself
Change 0.4 to 0.05, so one row in twenty carries the bug. Predict what happens before you run it.
Most people expect the damage to fall to about an eighth. It falls much further, because a linear model copes with a few extreme rows better than with a mixed population. Then set it to 0.5 and watch the dirty score sink towards a coin toss.
What to learn next
- Collecting data — where the ceiling gets set in the first place.
- Cleaning data — the specific bugs, and how to catch them.
- Feature engineering — raising the ceiling by building better columns.
Researcher — Mathematics and papers.
Stating the ceiling precisely
For a classification problem with input $X$ and label $Y$, the Bayes error rate is the lowest error achievable by any function of $X$:
$$ R^{*} = \mathbb{E}{X}!\left[ 1 - \max{y} \; p(Y = y \mid X) \right] $$
Where $p(Y=y \mid X)$ is the true conditional distribution of the label given the features, and the expectation runs over the marginal distribution of $X$.
$R^{}$ is a property of the joint distribution, not of any model. Model selection, capacity and optimisation all operate on the gap $R(\hat{f}) - R^{}$. Data work operates on $R^{*}$ itself, by changing which random variable $X$ is.
That is the formal content of "data sets the ceiling". Adding an informative column lowers $R^{}$. Removing measurement error lowers $R^{}$. No amount of architecture search does either.
Measurement error as attenuation
The units bug above is a case of measurement error, and the classical version has a closed form. Suppose the true model is $y = \beta x + \varepsilon$ but you observe $\tilde{x} = x + u$, with $u$ independent of $x$ and $\varepsilon$, $\mathbb{E}[u]=0$ and $\operatorname{Var}(u) = \sigma_u^2$. Then ordinary least squares converges to:
$$ \hat{\beta} \;\xrightarrow{\;p\;}\; \beta \cdot \frac{\sigma_x^2}{\sigma_x^2 + \sigma_u^2} \;=\; \beta \lambda $$
Where $\sigma_x^2$ is the variance of the true regressor and $\lambda \in (0,1]$ is the reliability ratio. The estimated effect is biased towards zero by exactly the fraction of variance that is noise.
The mixed-units case is worse than classical error, because $u$ is not independent of $x$: it equals $999x$ on a randomly chosen 40% of rows. The contaminated variable becomes a two-component mixture, so a single linear coefficient is misspecified rather than only attenuated.
Two consequences follow that practitioners routinely get wrong:
- More rows do not fix it. Attenuation is bias, not variance. A larger $n$ estimates the wrong quantity more precisely.
- Regularisation makes it look better and be worse. A shrunk coefficient on a noisy feature produces a smoother validation curve while discarding real signal.
Label noise, briefly
Under symmetric random label noise with flip rate $\rho < 0.5$, several surrogate losses remain classification-calibrated and the Bayes-optimal classifier is unchanged (Natarajan et al., 2013, Learning with Noisy Labels, NeurIPS). That result is why the experiments in labelling data show random noise doing surprisingly little damage.
Asymmetric and instance-dependent noise carries no such guarantee. It shifts the decision boundary, and the shift does not vanish as $n \to \infty$. Systematic annotation error is categorically more dangerous than careless annotation error at the same rate.
Evidence that data work dominates in practice
- Sambasivan et al. (2021), "Everyone wants to do the model work, not the data work": Data Cascades in High-Stakes AI, CHI. Interviews with 53 practitioners across India, East and West Africa and the US. Documents compounding downstream failures traced to upstream data decisions, and the systematic undervaluing of data work.
- Northcutt, Athalye and Mueller (2021), Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks, NeurIPS Datasets and Benchmarks. Estimated a 3.4% average label error rate across ten widely used test sets, and showed model rankings reversing once errors are corrected. Some benchmark leadership was an artefact of fitting label noise.
- Sculley et al. (2015), Hidden Technical Debt in Machine Learning Systems, NeurIPS. The modelling code is a small box inside a large diagram of data plumbing. Introduced CACE — changing anything changes everything — and undeclared consumers.
- Kapoor and Narayanan (2023), Leakage and the Reproducibility Crisis in ML-based Science, Patterns. Surveyed 17 scientific fields and found data leakage affecting hundreds of published papers, inflating reported performance in every case.
- Recht et al. (2019), Do ImageNet Classifiers Generalize to ImageNet?, ICML. Rebuilding the test set by the original protocol dropped accuracy 11 to 14 points, with rankings largely preserved. That gap is a property of the collection process, not of test-set overfitting.
The counter-argument, stated fairly
"Data-centric AI" gets pushed to a claim that is false: that architecture never matters. It plainly does. Transformers over recurrent networks was not a data fix, and neither were residual connections.
The defensible claim is narrower and is about marginal return. Given a mature architecture and a fixed budget, expected gain per engineer-hour is higher on data than on model search for most applied problems. Hooker (2021), Moving beyond "algorithmic bias is a data problem", Patterns, argues the converse in the fairness setting: model choices such as compression amplify harm on underrepresented groups independently of the data. Both are true. Which lever pays depends on where you currently stand.
Measuring the ceiling in your own problem
Practical estimators of how much room is left:
- Human agreement rate. Have two competent annotators label the same held-out sample. Their disagreement rate is a workable upper bound on achievable accuracy for a model given the same information.
- Feature ablation. Fit with and without a candidate column. The gap estimates that column's contribution to lowering $R^{*}$, conditional on the rest.
- Learning curves. Fit at 25%, 50% and 100% of rows. A curve that has flattened means more rows will not help, so the bottleneck is features or labels rather than volume.
The third is the one most often skipped, and it decides whether next quarter goes on collection or on cleaning.
Reading
- Sambasivan et al., Data Cascades in High-Stakes AI, CHI 2021.
- Northcutt et al., Pervasive Label Errors in Test Sets, NeurIPS D&B 2021 — arxiv.org/abs/2103.14749
- Sculley et al., Hidden Technical Debt in Machine Learning Systems, NeurIPS 2015.
- Natarajan et al., Learning with Noisy Labels, NeurIPS 2013.
- Kapoor and Narayanan, Leakage and the Reproducibility Crisis in ML-based Science, Patterns 2023 — arxiv.org/abs/2207.07048
What to learn next
- Collecting data — where the ceiling gets set in the first place.
- Cleaning data — the specific bugs, and how to catch them.
- Feature engineering — raising the ceiling by building better columns.