Baselines and Choosing a Model
Evaluating by slice, not by average
One average hides every group the model fails on, so cut the test set into named slices and report the worst one — a 2% segment scoring 0.229 moves the overall number by almost nothing and can still be the reason the product fails.
- 13 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
One overall score is an average, and an average can look healthy while a whole group of your users is being failed completely.
A restaurant with a 4.4-star average sounds good. Read the reviews and it turns out every single vegetarian gave it one star.
The average is not wrong. It is answering a question nobody asked. Diners do not experience the average; each one experiences their own meal.
Your model is the same. It does not serve "the test set", the pile of saved examples you score it on. It serves a shopkeeper in Nagpur on an old phone, and either it works for her or it does not.
Why it exists
Model scores get reported as one number because one number is easy to put in a slide. Teams then spend months moving that number from 0.82 to 0.84.
Meanwhile, a group making up 2% of users can be getting answers that are worse than a coin toss. Because they are 2%, fixing them moves the headline number by almost nothing, and breaking them moved it by almost nothing too. Nobody sees it.
Those small groups are rarely unimportant. They are new markets, cheap devices, a second language, night-shift customers, the rare disease. The whole reason to build the model is often the group the average is hiding.
Slice-based evaluation means cutting the test set into named groups and reporting a score for each one.
How it works
ONE NUMBER SLICED
accuracy 0.826 android 1477 rows 0.835 fine
ios 575 rows 0.852 fine
(looks fine, ship it) kaios 48 rows 0.229 BROKEN
small amounts 0.891
medium 0.712 <- worse than it looks
large 0.976Same model, same test set, same predictions. One report says ship it. The other says a whole class of phone gets answers that are worse than guessing.
The extra work is small: pick the groups that matter, count each one, score each one, and look at the worst.
A real example you have seen
Speech recognition on your phone. National accuracy figures look excellent. Then someone speaks Marathi-accented English, or speaks while a fan is running, and it fails repeatedly.
Those users are a slice. Their experience is not "94% accurate". It is "this does not work for me", every time.
Remember this
- An average can be healthy while a whole group is completely failed.
- Slice by the things that matter: device, region, language, time, customer size, rare classes.
- Report the worst slice next to the average, always. A small slice can still be the whole product.
What to learn next
- Deciding a model is good enough — what to do once the worst slice is acceptable.
- Macro, micro and weighted averaging — the metric that treats every class as a slice.
- Simpson's paradox — the same arithmetic, reversing a conclusion.
Developer — Code and libraries.
Setup
pip install scikit-learn pandasVerified with scikit-learn 1.7.2, pandas 2.2.3, numpy 1.26.4, Python 3.10, CPU. Runs in about five seconds. Data is generated in memory.
A slice scan in one function
The data has a 2% segment whose behaviour is the reverse of everyone else's — a realistic pattern, since a cheap-device or new-market population often has genuinely different economics. The device column is deliberately not a feature, which is the usual situation: you slice by things the model never saw.
import numpy as np, pandas as pd
from sklearn.ensemble import HistGradientBoostingClassifier
from sklearn.model_selection import train_test_split
rng = np.random.default_rng(7)
n = 6000
device = rng.choice(["android", "ios", "kaios"], n, p=[0.70, 0.28, 0.02])
hour = rng.integers(0, 24, n)
amount = rng.gamma(2.0, 60.0, n).round(2)
base = 0.05*amount - 4.0 # the ordinary rule
flip = np.where(device == "kaios", -1.0, 1.0) # this segment behaves oppositely
y = (rng.random(n) < 1/(1+np.exp(-flip*base))).astype(int)
df = pd.DataFrame({"device": device, "hour": hour, "amount": amount, "y": y})
FEATURES = ["hour", "amount"] # device is NOT a feature, only a slice
Xtr, Xte, ytr, yte, dtr, dte = train_test_split(
df[FEATURES], df.y, df.device, test_size=0.35, random_state=0)
model = HistGradientBoostingClassifier(random_state=0).fit(Xtr, ytr)
pred = model.predict(Xte)
test = pd.DataFrame({"device": dte.to_numpy(), "hour": Xte.hour.to_numpy(),
"amount": Xte.amount.to_numpy(), "y": yte.to_numpy(), "pred": pred})
test["night"] = np.where(test.hour < 6, "00-06", "06-24")
test["size"] = pd.cut(test.amount, [0, 50, 150, 1e9], labels=["small", "medium", "large"])
overall = (test.pred == test.y).mean()
print(f"overall accuracy {overall:.3f} on {len(test)} rows\n")
rows = []
for col in ["device", "night", "size"]:
for value, g in test.groupby(col, observed=True):
acc = (g.pred == g.y).mean()
se = (acc*(1-acc)/len(g))**0.5 # how much of the gap could be luck
rows.append((col, str(value), len(g), acc, acc - overall, 1.96*se))
print("slice column value rows accuracy gap vs overall +/- 95%")
for col, value, k, acc, gap, band in sorted(rows, key=lambda r: r[4]):
print(f"{col:12s} {value:9s} {k:5d} {acc:8.3f} {gap:+14.3f} {band:7.3f}")overall accuracy 0.826 on 2100 rows slice column value rows accuracy gap vs overall +/- 95% device kaios 48 0.229 -0.597 0.119 size medium 1064 0.712 -0.113 0.027 night 00-06 524 0.824 -0.001 0.033 night 06-24 1576 0.826 +0.000 0.019 device android 1477 0.835 +0.009 0.019 device ios 575 0.852 +0.026 0.029 size small 414 0.891 +0.066 0.030 size large 622 0.976 +0.150 0.012
The walkthrough
0.229 accuracy on kaios is worse than always answering "no". On this slice the model is not weak, it is inverted. The segment's rule reverses the majority's, and the model never saw the column that would have told it. Predicting the opposite of the model would score 0.771 on those users.
Forty-eight rows, and the finding is still solid. The ±0.119 band puts the interval at roughly 0.11 to 0.35, nowhere near the 0.826 average or the 0.5 coin-toss line. Small slices need their error bars printed, and a small slice is not automatically unreliable. The noise band either overlaps the average or it does not. Here it does not.
Compare that with the night rows. Gaps of −0.001 and +0.000 sit inside bands of ±0.033 and ±0.019. That slicing dimension carries no signal at all, and reporting it as a finding would be noise dressed as insight.
size = medium is the finding people miss. It has 1,064 rows and is 0.113 below average — a large slice, a real gap, and a completely different cause from the kaios one. Medium amounts sit near the decision boundary where the task is genuinely hard, while large amounts are easy (0.976). That is not a bug; it is where the remaining errors live, and it tells you which examples to buy more of.
The overall number barely moves. Fix kaios perfectly and overall accuracy rises by about 0.018 — invisible in a report, decisive for those users. This is the structural reason slice evaluation has to be a separate ritual rather than something you hope the headline metric reveals.
Choosing slices
Write the list before you look at any results, so the slices are the ones that matter rather than the ones that flatter.
- Who the users are: device, operating system, region, language, new versus returning, account age, customer size.
- What the input looks like: input length, missing-field count, image resolution, background noise, time of day.
- What the label is: one slice per class, always. A rare class is a slice with a support of forty, and macro averaging is what happens when you weight those slices equally on purpose.
- When it happened: the most recent month as its own slice, which is the earliest visible sign of drift.
- Where the model is unsure: the band around the threshold, which is the slice a human review queue will actually receive.
Then add one rule: every slice needs a minimum support before it is reported as a finding. Fifty rows is a reasonable floor for accuracy; a rare positive class needs far more, because the metric moves in steps of one over the positive count.
Common mistakes
Slicing after seeing the results. Cutting the data forty ways and reporting the worst is the multiple-testing trap — with enough slices, something always looks broken. Fix the slice list first, and treat anything discovered afterwards as a hypothesis to confirm on fresh data.
Reporting a slice with nine rows. Its band spans most of the interval, and next month it will say something different. Report it as "not enough data", which is itself a finding worth acting on.
Assuming the slice column must be a feature. It usually should not be. Slicing by device tells you where the model fails; adding device as a feature is a decision with fairness and stability consequences that deserves its own discussion.
Only slicing on protected attributes. Fairness slices matter, and they are a subset. Input length, missing data and rare classes break models constantly and belong on the same report.
Letting slices exist only in a notebook. The value comes from running them on every candidate model, so a regression on kaios blocks a release instead of being discovered by a customer. Put the scan in the evaluation script with a floor for the worst slice.
Try it yourself
Give the model the column and rerun the scan. device holds strings, so add a numeric code first — insert df["device_code"] = pd.Categorical(df.device).codes and set FEATURES = ["device_code", "hour", "amount"]. The kaios slice repairs itself, from 0.229 to 0.750, and overall accuracy moves from 0.826 to 0.838. Then reduce that segment to 0.5% of the data (p=[0.71, 0.285, 0.005]) and try again — the model has the column and still fails (accuracy 0.167 on twelve test rows), because the sixteen matching training rows cannot teach a reversed rule. That second result is the argument for class weights and targeted data collection over yet another architecture.
What to learn next
- Deciding a model is good enough — what to do once the worst slice is acceptable.
- Macro, micro and weighted averaging — the metric that treats every class as a slice.
- Simpson's paradox — the same arithmetic, reversing a conclusion.
Researcher — Mathematics and papers.
Why aggregates conceal, formally
For a partition of the test distribution into slices $S_1,\dots,S_K$ with masses $\pi_k$ and slice-conditional risks $R_k$, the reported risk is $R = \sum_k \pi_k R_k$. The sensitivity of the headline number to a slice is $\partial R/\partial R_k = \pi_k$, so a slice's visibility in the aggregate is exactly its mass, entirely independent of its importance. A segment at $\pi = 0.02$ can move from perfect to inverted while $R$ changes by at most 0.02 — smaller than the seed-to-seed variation of most training pipelines, and therefore undetectable in a headline comparison even in principle.
This is the same arithmetic that produces Simpson's paradox (Simpson's paradox): an aggregate comparison can reverse the direction of every slice-level comparison when slice masses differ between the two conditions. In model comparison this appears as a candidate that wins overall while losing on every named slice, purely through a shift in which slices the test set happens to weight — which is why fair comparison requires matched slice composition, not only matched folds.
Automated slice discovery
Hand-written slice lists are limited by imagination, and a literature exists on finding underperforming subgroups automatically.
Chung et al. (2019), Slice Finder: Automated Data Slicing for Model Validation (ICDE), search the lattice of conjunctions over categorical and binned features for slices that are large, interpretable, and significantly worse than the complement, with explicit multiple-hypothesis control via a marginal false discovery rate — the necessary correction given that the lattice contains exponentially many candidates. Sagadeeva and Boehm (2021), SliceLine (SIGMOD), give a linear-algebra formulation with monotonic upper bounds enabling pruning, which makes exhaustive search tractable at realistic scale.
Eyuboglu et al. (2022), Domino: Discovering Systematic Errors with Cross-Modal Embeddings (ICLR), attack the harder case where no attribute columns exist: cluster errors in an embedding space and describe each cluster in natural language. This addresses the failure that hand-written lists structurally cannot — a coherent error mode nobody had a column for. Errudite (Wu et al., 2019, ACL) provides the interactive counterpart, insisting error groups be specified as reproducible, counterfactually testable queries rather than as impressions from reading examples, which is the disciplined version of error analysis.
The statistical hazard across all of these is the one named in the mistakes list. With $K$ slices tested, the expected number of spurious "significant" gaps at level $\alpha$ is $\alpha K$ under the null, and $K$ grows combinatorially with conjunction depth. Benjamini–Hochberg control over the discovered set, plus confirmation on a held-out slice sample, is the minimum defensible protocol.
Optimising the worst slice
Once slices are named, one may optimise them rather than only measure them. Distributionally robust optimisation minimises the worst-case risk over an uncertainty set of distributions; with known group labels, Group DRO (Sagawa et al., 2020, Distributionally Robust Neural Networks for Group Shifts, ICLR) minimises $\max_k R_k$ directly and shows that strong regularisation is essential — over-parameterised models otherwise achieve low worst-group training loss while generalising poorly on the small groups. Without group labels, Just Train Twice (Liu et al., 2021, ICML) upweights examples misclassified by an initial model, recovering much of the benefit.
The cost is explicit and should be stated when proposing it: minimising the worst slice generally lowers average performance, since the objective reallocates capacity toward the hardest population. Whether that trade is correct is a product decision, not a modelling one, and it belongs in the success metric agreed at scoping time rather than being resolved silently by a loss function.
A final measurement caution: slice metrics inherit every property of the metric being sliced. Slicing a calibration-sensitive metric across groups with different base rates conflates discrimination with calibration (reliability diagrams), and the impossibility results in fair classification (Kleinberg, Mullainathan and Raghavan, 2017; Chouldechova, 2017) establish that calibration within groups, equal false-positive rates and equal false-negative rates cannot generally hold simultaneously when base rates differ. Reporting several slice metrics side by side is therefore not redundancy — it is the only honest presentation available.
What to learn next
- Deciding a model is good enough — what to do once the worst slice is acceptable.
- Macro, micro and weighted averaging — the metric that treats every class as a slice.
- Simpson's paradox — the same arithmetic, reversing a conclusion.