Asserting on your data before you train
Ten lines of checks — nulls, duplicates, ranges, label values — run before every training job catch the silent data problems that would otherwise become silent model problems.
- 8 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Check your ingredients before you cook — a model trained on bad data fails silently, so the checking must happen before training.
No cook pours milk into the chai without a sniff. The sniff takes one second, costs nothing, and prevents a ruined pot. Crucially, the cook sniffs every time — not only on days the milk seems suspicious, because spoiled milk does not announce itself.
Training data spoils the same way. A broken export, a column of nulls, ages of minus three, a label spelled "maybe" where only yes and no belong. And here is the frightening part: the training will not crash. The model will train on nonsense, produce a slightly worse score, and nobody will know why — or even that — something went wrong.
Why it exists
Code bugs crash loudly. Data bugs whisper. A model is a compression of its data — feed it corrupted data and you get a corrupted model with no error message, ever.
Damage usually shows up weeks later, as a drooping metric. By then the bad data has been overwritten by the next export, and the investigation has nothing to hold. The only reliable moment to catch a data problem is before training consumes it — with checks that run automatically, every single time, like the sniff.
How it works
data arrives → ✅ VALIDATION GATE → training
│
├─ nulls where none allowed?
├─ duplicate rows?
├─ values outside sane ranges?
├─ labels outside the allowed set?
└─ columns frozen at one value?
│
any failure → STOP, loudly, before trainingEach check is one line of "this must be true about the data" — an assertion. The gate turns silent data problems into loud, immediate, pre-training failures that name the exact problem.
A real example you have seen
Airport security screens every bag, on every flight. Not because most bags are dangerous, but because screening is cheap and the rare miss is enormously costly. Nobody suggests screening "only when suspicious". Your validation gate is that screening line for data, and it earns its keep on the one day in fifty it fires.
Remember this
- Data bugs do not crash — they quietly become model bugs.
- A validation gate runs every training, and fails loudly before training.
- Each check is one sentence about the data that must always be true.
What to learn next
- Duplicate rows across your splits — the leak your duplicate check now prevents.
- Why data quality matters — the wider case for treating data as the product.
- Monitoring and drift — the same checks, running in production forever.
Developer — Code and libraries.
Setup
pip install pandasVerified with pandas 2.2.3, numpy 1.26.4, Python 3.10, CPU. Deterministic — your output should match.
A validation gate in thirty lines
Five checks, run against a small table with five planted problems. Every problem is the kind that survives a df.head() glance.
import numpy as np
import pandas as pd
def validate(df, label_col, allowed_labels):
problems = []
if df.isna().any().any():
bad = df.columns[df.isna().any()].tolist()
problems.append(f"missing values in {bad}")
if df.duplicated().any():
problems.append(f"{df.duplicated().sum()} duplicate rows")
unexpected = set(df[label_col].dropna()) - set(allowed_labels)
if unexpected:
problems.append(f"unexpected labels: {sorted(unexpected)}")
numeric = df.select_dtypes("number")
flat = numeric.columns[numeric.nunique() <= 1].tolist()
if flat:
problems.append(f"constant columns: {flat}")
if "age" in df and ((df.age < 0) | (df.age > 120)).any():
problems.append("age outside the range 0-120")
return problems
df = pd.DataFrame({
"age": [34, 61, -3, 45, 45],
"plan": ["basic", "pro", "pro", "basic", "basic"],
"spend": [1.0, 1.0, 1.0, 1.0, 1.0],
"churned": ["yes", "no", "no", "maybe", np.nan],
})
df = pd.concat([df, df.iloc[[0]]], ignore_index=True)
issues = validate(df, "churned", allowed_labels=["yes", "no"])
print(f"{len(issues)} problems found:")
for issue in issues:
print(" -", issue)5 problems found: - missing values in ['churned'] - 1 duplicate rows - unexpected labels: ['maybe'] - constant columns: ['spend'] - age outside the range 0-120
In a real pipeline, the caller raises on any problem: assert not issues, issues — training never starts.
The walkthrough
Each check maps to a downstream disaster you have now prevented. Nulls in the label: rows silently dropped or, worse, imputed. The duplicate: split contamination. The label "maybe": a surprise third class, and every metric quietly redefined. The constant spend column: a dead feature — and, more interestingly, a symptom, because columns go constant when an upstream join or export breaks. The age of −3: a corrupted record, or a unit change upstream.
Collect all problems, then fail. The function gathers every issue instead of raising at the first, so one failed run gives the whole repair list. Debugging a data drop at one-problem-per-run is misery.
The checks encode domain knowledge, cheaply. "Age 0–120" and "labels ∈ {yes, no}" are facts you know and the machine does not. Ten such facts, written once, guard every future run. Start with the facts that have already burned you once.
Where does the gate live? In tests/test_data.py for CI, and as the first call in train.py — data can break between commits, so the training-time gate is the one that matters most. This is the file the repo layout reserved for exactly this.
Growing the gate
- Distribution checks come next: "positive share between 5% and 40%", "mean spend within 3× of last month's". These catch drift, not corruption — see monitoring and drift for the production version.
- Leakage checks: the single-feature AUC scan as an assertion — "no single feature predicts the label above 0.9 alone".
- Frameworks — Great Expectations, pandera, Deequ — give you declarative schemas, reports and integrations. Adopt one when your handwritten gate crosses ~50 checks; the thinking transfers unchanged.
Common mistakes
Validating once, during development. The gate's whole value is running on every future dataset, on schedule, unattended. Data spoils on Tuesdays too.
Checks that log instead of stop. A warning nobody reads is a check that does not exist. Fail the run; make a human decide to override.
Only checking structure, never values. Schema tools confirm age is an integer. They are happy with −3. Value-range and label-set checks are where the real catches live.
Cleaning instead of failing. Auto-dropping bad rows inside the gate hides the problem and changes your data silently — the very disease this lesson treats. The gate reports; cleaning is a separate, visible, versioned step in data cleaning.
Try it yourself
Add two checks: row count at least 100 (guards against a truncated export), and no column more than 20% null. Break the data to trigger each, and confirm one run reports both together.
What to learn next
- Duplicate rows across your splits — the leak your duplicate check now prevents.
- Why data quality matters — the wider case for treating data as the product.
- Monitoring and drift — the same checks, running in production forever.
Researcher — Mathematics and papers.
Validation as schema plus expectations
The industrial formulation is Breck et al. (2019), Data Validation for Machine Learning, MLSys — the system behind TFX Data Validation at Google. Key design points that generalise beyond their stack:
- Schema as a living artefact: types, domains, ranges and presence constraints, versioned alongside code, with the system proposing schema updates when legitimate drift occurs — validation as an evolving contract, not a static filter.
- Skew detection across environments: training/serving feature distributions compared continuously — the validation-time twin of training-serving skew.
- The anomaly taxonomy: missingness, domain violations, distribution shifts, and feature crosses whose joint behaviour breaks while marginals look fine — the last being invisible to per-column gates like this lesson's, and the argument for eventually testing joint statistics.
Schelter et al. (2018), Automating Large-Scale Data Quality Verification, VLDB (Deequ, on Spark): declarative constraints compiled to distributed metric computations, with incremental evaluation over growing data — the constraint-as-code idea at warehouse scale. Polyzotis et al. (2019), Data Validation for Machine Learning (SIGMOD tutorial lineage) survey the research space.
Where data errors actually rank
Breck et al. (2017), The ML Test Score, IEEE Big Data: of their 28-item production-readiness rubric, seven items are data tests, and their field observation is that data tests are the least-implemented category while data issues dominate incident lists. Sculley et al. (2015) (NeurIPS) explain the structural reason: ML systems have data dependencies with none of the tooling (compilers, type checkers) that code dependencies enjoy — unstable upstream signals, underutilised features, and silent consumers. A validation gate is a hand-built type checker for a data dependency.
The empirical alarm bell: "silent degradation" appears across incident postmortem literature (and in ML incident response folklore) precisely because the failure mode has no exception — the model's loss surface absorbs the corruption. Detection latency is bounded below by evaluation frequency unless a gate exists upstream of training.
Statistical checks and their thresholds
Distribution checks need a distance and a threshold: common choices are population stability index, KL divergence, Kolmogorov–Smirnov distance per feature. The uncomfortable truth from practice (echoed in Breck 2019): fixed statistical thresholds either alarm constantly or never — production systems converge on relative checks (against a trailing window), per-feature severity tiers, and human-in-the-loop triage for medium severities. Alert fatigue is the failure mode of over-eager gates; the design problem is precision-recall on anomalies, not maximal sensitivity.
A final framing: the gate is where domain constraints become executable. Every physical or business invariant (ages, prices, category sets, monotonic timestamps, conservation between columns) is a theorem about valid data; the gate is its proof obligation, discharged per batch. Datasets with documented invariants — datasheets (Gebru et al., 2021, CACM) — make gates writable by people other than the original collector, which is the collaboration property everything else in this section is built on.
What to learn next
- Duplicate rows across your splits — the leak your duplicate check now prevents.
- Why data quality matters — the wider case for treating data as the product.
- Monitoring and drift — the same checks, running in production forever.