Collecting data
Who you collect from decides who your model works for, and a sample drawn from the wrong crowd cannot be repaired later by any model.
- 12 min read
- 3 reading levels
- Updated
Read these first
On this page 7
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Collecting data means deciding who and what gets recorded — and that decision quietly decides who your model will work for.
Picture a shopkeeper who wants to know what the neighbourhood likes to eat. He asks the customers standing in his own shop, and only those. His notebook fills up with answers.
But the people who never walk into his shop never appear in it. He learns what his existing customers want, and calls it what the neighbourhood wants.
Every dataset ever collected has this shape. Somebody chose where to stand and hold the notebook.
Why this is the step that matters most
Cleaning fixes wrong values. Collection decides which values exist at all.
A missing column can sometimes be recovered. A missing group of people cannot. If nobody from a village is in your training data, no amount of cleaning invents them.
That is why collection sits so early in this section. Mistakes made here cannot be undone downstream, at any price.
The four questions to answer before you collect anything
Who is in the sample? Not who you meant to reach. Who actually replied, opened the app, or had a smartphone.
What is the label, exactly? "Fraud" could mean the bank reversed the payment, or the customer complained, or an analyst ticked a box. Those three are different datasets.
When was it recorded? Data from festival week does not describe an ordinary week.
What is missing on purpose? Every form has a field somebody chose not to add.
How a sample goes wrong
the population you serve
############################ (100 people)
|
| you advertise the survey inside your app
v
#### (the 15 who use the app daily)
|
v
[ train the model ]
|
v
works well for those 15
guesses for the other 85Nothing in that picture is a bug in the code. Every line ran correctly. The model is honest about what it was shown.
Somewhere you have seen this
Voice assistants that struggle with Indian English were not built by careless engineers. They were built with recordings from people who did not sound like you.
Face unlock that works less well on darker skin has the same history. Early face datasets were heavily lighter-skinned, and the models inherited that.
A credit model trained only on salaried customers will misread a shopkeeper with irregular income. It never met one.
The honest part
You can never collect a perfect sample. Budgets, consent, geography and time all bite.
The goal is not perfection. It is knowing precisely who is missing, writing it down, and refusing to make claims about them.
A dataset with a documented gap is professional. A dataset with an unknown gap is a liability.
Remember this
- Collection decides who the model works for. Cleaning cannot fix it later.
- Ask who replied, what the label really means, when it was taken, and what was never asked.
- Write down who is missing. An admitted gap is safe; a hidden one is not.
What to learn next
- Labelling data — turning collected rows into trustworthy answers.
- Bias in datasets — what a skewed sample does to real people.
- Train, test and validation splits — building a test set that resembles the world.
Developer — Code and libraries.
Here is the effect measured. Same model, same number of rows, same code. The only change is who the rows came from.
The setup is a lending population where two thirds of people live outside metros, and repayment follows a different rule in each group. That is realistic: in one group a steady income predicts repayment, in the other how long you have owned your phone predicts it better.
Setup
pip install numpy pandas scikit-learnTwo ways to spend the same survey budget
import numpy as np
import pandas as pd
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import roc_auc_score
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler
rng = np.random.default_rng(7)
N = 4000
# The population you actually lend to. Two thirds of it is not in a metro.
rural = rng.binomial(1, 0.65, N)
income = np.where(rural == 1, rng.gamma(3.0, 8.0, N) + 5, rng.gamma(4.0, 15.0, N) + 10)
phone_age = rng.integers(0, 6, N)
# Repayment follows a different rule in each group. Nobody told you that.
truth = np.where(rural == 1,
0.003 * income + 0.9 * phone_age - 2.2,
0.075 * income - 0.02 * phone_age - 4.2)
repaid = rng.binomial(1, 1 / (1 + np.exp(-truth)))
pop = pd.DataFrame({"income": income, "phone_age": phone_age, "rural": rural, "repaid": repaid})
test = pop.sample(1000, random_state=1) # the real world you will be judged on
pool = pop.drop(test.index)
# Collection A: you advertised the survey inside your app. App users are mostly urban.
app_users = np.where(pool["rural"] == 1, 0.03, 1.0)
biased = pool.sample(600, weights=app_users, random_state=2)
# Collection B: same 600 rows, same budget, drawn at random from the same population.
representative = pool.sample(600, random_state=3)
def evaluate(train):
model = make_pipeline(StandardScaler(), LogisticRegression(max_iter=1000))
model.fit(train[["income", "phone_age"]], train["repaid"])
p = model.predict_proba(test[["income", "phone_age"]])[:, 1]
is_rural = test["rural"] == 1
return (round(roc_auc_score(test["repaid"], p), 3),
round(roc_auc_score(test["repaid"][is_rural], p[is_rural]), 3),
round(roc_auc_score(test["repaid"][~is_rural], p[~is_rural]), 3))
print("rural share of the population :", round(pop["rural"].mean(), 3))
print("rural share of the app survey :", round(biased["rural"].mean(), 3))
print("rural share of the random sample:", round(representative["rural"].mean(), 3))
print()
print(f"{'training set':16s}{'rows':>6}{'AUC all':>10}{'AUC rural':>11}{'AUC urban':>11}")
for name, sample in [("app survey", biased), ("random sample", representative)]:
a, r, u = evaluate(sample)
print(f"{name:16s}{len(sample):>6}{a:>10}{r:>11}{u:>11}")rural share of the population : 0.646 rural share of the app survey : 0.082 rural share of the random sample: 0.645 training set rows AUC all AUC rural AUC urban app survey 600 0.63 0.508 0.887 random sample 600 0.768 0.813 0.701
What those two rows say
0.508 for rural customers. That is a coin toss, for two thirds of the population. The model is not slightly worse for them; it carries no usable information about them at all.
The biased model is genuinely better for urban customers. 0.887 against 0.701. This is the part people miss. Sampling bias does not produce a bad model. It produces a model that is excellent for the people you sampled and useless for everyone else.
Overall AUC hides it. 0.630 versus 0.768 looks like an ordinary gap you might shrug at. Splitting by group is what makes the failure visible.
Nothing was fixable downstream. Same algorithm, same 600 rows, same features, same code. The gap is entirely in who was sampled.
Line by line, the parts that matter
pool.sample(600, weights=app_users) gives rural rows a selection weight of 0.03 against 1.0 for urban. That is the survey being advertised where urban people are. The result: the sample is 8% rural in a 65% rural population.
test is drawn from the population before either sample. Your test set has to look like the world, not like your collection channel. A test set drawn from the same biased channel would have shown 0.887 and you would have shipped.
stratify is deliberately absent here, because the point is the grouping variable rural that the model never sees. It is in the frame for measurement only.
Common mistakes
Testing on data from the same channel that trained. The single most common way a biased dataset ships. Your offline score will be excellent and your live score will not be. Hold out a slice collected differently, even a small manual one.
Reporting one aggregate number. Always break the score down by the groups you can identify: region, device, language, age band, acquisition channel. It costs three lines and catches this class of failure.
Assuming more rows fix it. Collect 60,000 app-user rows instead of 600 and the rural AUC stays near 0.5. Bias is not noise. It does not average out.
Dropping the group column so you cannot be accused of bias. Removing rural from the features does not remove the bias, it removes your ability to measure it. Keep such columns for evaluation even when you exclude them from the model.
Try it yourself
Change 0.03 to 0.65 in app_users, so the survey reaches rural people a little less rather than almost never. Watch how far the rural AUC recovers.
Then set the weight back to 0.03 and instead raise the biased sample to n=3000. Confirm for yourself that ten times the data does not close the gap.
What to learn next
- Labelling data — turning collected rows into trustworthy answers.
- Bias in datasets — what a skewed sample does to real people.
- Train, test and validation splits — building a test set that resembles the world.
Researcher — Mathematics and papers.
Selection as conditioning on a collider
Let $S \in {0,1}$ indicate whether a unit enters the sample. Your training distribution is $p(x, y \mid S = 1)$; the deployment distribution is $p(x, y)$. The two coincide only when $S \perp (X, Y)$.
Three regimes, in increasing severity:
Covariate shift. $p_{\text{train}}(x) \neq p_{\text{test}}(x)$ but $p(y \mid x)$ is shared. Selection depends on $X$ alone: $S \perp Y \mid X$. This is correctable in principle by importance weighting with $w(x) = p_{\text{test}}(x) / p_{\text{train}}(x)$, provided the support condition $p_{\text{train}}(x) > 0$ wherever $p_{\text{test}}(x) > 0$ holds.
Label shift. $p(y)$ changes but $p(x \mid y)$ is shared. Correctable via the confusion-matrix estimator of Saerens et al. (2002) or BBSE (Lipton et al., 2018, Detecting and Correcting for Label Shift with Black Box Predictors).
Concept shift. $p(y \mid x)$ itself differs across strata. This is the developer example: the rural and urban conditionals are different functions. Reweighting cannot fix it, because the target function is not shared. The only remedies are collecting the missing stratum or fitting per-stratum models.
Distinguishing these three before choosing a correction is the whole game, and it is frequently skipped.
The support condition is the binding constraint
Importance weighting is often quoted as the general fix for biased sampling. Its variance is governed by:
$$ \operatorname{Var}!\left[\hat{R}w\right] \;\propto\; \mathbb{E}{p_{\text{train}}}!\left[ w(X)^2 \ell(X)^2 \right] $$
Where $w(X)$ is the density ratio and $\ell$ the loss. As $p_{\text{train}}(x) \to 0$ in a region where $p_{\text{test}}(x)$ is positive, $w$ diverges and the estimator's variance explodes. The effective sample size, $\left(\sum_i w_i\right)^2 / \sum_i w_i^2$, collapses.
In the developer example with an 8% rural sample, weights would exceed 20 on the sparse stratum. The reweighted estimate would be nominally unbiased and practically worthless. Reweighting cannot manufacture coverage.
Heckman correction, and why it is rarely used here
Heckman (1979), Sample Selection Bias as a Specification Error, Econometrica, models selection and outcome jointly, estimating a selection equation and adding the inverse Mills ratio as a regressor in the outcome equation. Consistency requires a valid exclusion restriction: a variable that drives selection but not the outcome.
That is a strong requirement. In ML settings such an instrument is seldom available and seldom defensible, which is why the technique appears far more often in econometrics than in applied machine learning.
Sampling designs worth knowing
- Stratified sampling. Partition by a known variable, sample within strata. Variance of a stratum-weighted mean estimator is $\sum_h W_h^2 \sigma_h^2 / n_h$, minimised by Neyman allocation $n_h \propto W_h \sigma_h$. When strata differ in conditional structure, oversampling small strata is the correct move even though it distorts marginals — reweight at evaluation time, not collection time.
- Cluster sampling. Cheaper in the field, but the design effect $\text{deff} = 1 + (m-1)\rho$ inflates variance, where $m$ is cluster size and $\rho$ the intra-cluster correlation. Ten villages of 100 people is nowhere near 1,000 independent observations.
- Active learning. Uncertainty sampling and query-by-committee (Settles, 2009, Active Learning Literature Survey) reduce labelling cost, but the resulting pool is deliberately non-representative. Keep a separately drawn random slice for unbiased evaluation.
Documentation as a first-class artefact
- Gebru et al. (2021), Datasheets for Datasets, CACM. A standard questionnaire covering motivation, composition, collection process, preprocessing, uses and distribution.
- Bender and Friedman (2018), Data Statements for NLP, TACL. Focused on speaker demographics, curation rationale and dialect, which are exactly the axes that go undocumented.
- Mitchell et al. (2019), Model Cards for Model Reporting, FAccT. Requires disaggregated evaluation across groups — the practice the developer section demonstrates.
Empirical work on dataset-level bias
Torralba and Efros (2011), Unbiased Look at Dataset Bias, CVPR, ran a "name that dataset" experiment: a classifier trained to identify which corpus an image came from reached far above chance. They also measured cross-dataset generalisation, finding consistent performance drops when training on one corpus and testing on another.
Buolamwini and Gebru (2018), Gender Shades, FAccT, evaluated commercial gender classifiers disaggregated by skin type and gender, reporting error rates up to 34.7% for darker-skinned women against 0.8% for lighter-skinned men. The aggregate accuracy of those systems looked strong.
Both results share a structure: aggregate metrics concealed the failure, and disaggregated evaluation exposed it.
Practical protocol
- Write the target population down before collecting anything.
- Record the selection mechanism, including the channel and any incentive offered.
- Draw a small probability sample separately, at whatever cost, and reserve it purely for evaluation.
- Report every metric disaggregated by the strata you can observe.
- Track the effective sample size per stratum, not only the total row count.
Reading
- Heckman, Sample Selection Bias as a Specification Error, Econometrica 1979.
- Torralba and Efros, Unbiased Look at Dataset Bias, CVPR 2011.
- Buolamwini and Gebru, Gender Shades, FAccT 2018.
- Gebru et al., Datasheets for Datasets, CACM 2021 — arxiv.org/abs/1803.09010
- Lipton, Wang and Smola, Detecting and Correcting for Label Shift with Black Box Predictors, ICML 2018 — arxiv.org/abs/1802.03916
What to learn next
- Labelling data — turning collected rows into trustworthy answers.
- Bias in datasets — what a skewed sample does to real people.
- Train, test and validation splits — building a test set that resembles the world.