Model cards and documentation
A model card is the label on the tin — what the model is for, who it was tested on, where it fails — generated from the evaluation run, not written from memory.
- 13 min read
- 3 reading levels
- Updated
Read these first
On this page 9
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
A model card is a short document that says what a model is for, who it was tested on, and where it fails.
The analogy you have already lived
Look at a packet of biscuits. There is a list of ingredients, a batch number, a manufacturing date, and a line saying "may contain traces of nuts".
Nobody reads it on a good day. The person with a nut allergy reads it every single time, and that line is why they can buy the packet at all.
A model card is that label. Most people skip it. The person about to use your model for something it was never built for needs it, and cannot get it anywhere else.
Why it exists
Somebody will use your model for a purpose you never imagined. That is not a hypothetical; it is the normal life of a useful model.
A model trained to sort blurry photos gets used for identity checks. A model trained on adult chest scans gets pointed at children. A model trained on data from one city gets deployed in another.
None of those people were careless. They had no way to know, because nobody wrote it down.
A model card is the difference between "we did not know" and "we told you".
What goes on the card
The original 2019 proposal by Mitchell and colleagues at Google lists nine sections. Five of them carry almost all the value.
What this is. The version, the date, the type of model, who owns it.
What it is for. The uses you support, in plain words.
What it is not for. The uses you refuse. This section prevents more harm than any other.
How it performs, broken down. Not one accuracy number. A table, one row per group of users, with the size of each group next to it.
What is wrong with it. Known failures, groups it serves badly, conditions that break it.
The picture
┌────────────────────────────────────────────┐
│ MODEL CARD: signup-risk v0.3 │
├────────────────────────────────────────────┤
│ For : ranking signups for human review │
│ NOT : automatic rejection, credit │
├────────────────────────────────────────────┤
│ slice n accuracy recall │
│ everyone 900 0.874 0.865 │
│ new device 594 0.867 0.887 │
│ old device 306 0.889 0.808 ← │
├────────────────────────────────────────────┤
│ Known limits: users on old devices are │
│ missed more often. Do not use offline. │
└────────────────────────────────────────────┘The arrow is the whole point of the document. One group is being served worse, it is written down, and now a decision can be made about it.
Where you have already seen one
Every model page on Hugging Face has a card. Google publishes cards for some of its APIs. Regulators in several places are moving toward requiring something like this for higher-risk systems.
What is honestly hard here
A card written by hand goes stale the moment the model is retrained.
The fix is to stop writing cards and start generating them, from the same script that computes your evaluation metrics. If the numbers are typed by a person, they will eventually be wrong.
And a card is a claim, not a proof. It is only as honest as the team writing it.
Remember this
- A model card says what the model is for, what it is not for, and where it fails.
- Metrics must be broken down by group, with the group size shown.
- Generate it from your evaluation code, so it cannot drift out of date.
What to learn next
- AI regulation — the rules that increasingly expect this document.
- Deploying responsibly — what happens after the card is written.
- CI/CD for machine learning — wiring card generation into the pipeline.
Developer — Code and libraries.
Generate the card, never type it
The rule that makes documentation survive contact with a real team: the card is an output of the evaluation script. It is written in the same run that computes the numbers, so it cannot disagree with them.
Setup
pip install numpy scikit-learnAn evaluation script that emits its own card
import hashlib, platform
import numpy as np
import sklearn
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import accuracy_score, recall_score, precision_score
from sklearn.model_selection import train_test_split
SEED = 0
rng = np.random.default_rng(SEED)
N = 3000
# two feature columns plus a group column kept OUT of training, used ONLY for reporting
region = np.where(rng.random(N) < 0.2, "rural", "urban")
device = np.where(rng.random(N) < 0.35, "old", "new")
signal = rng.normal(0, 1, N) + np.where(device == "old", -0.4, 0.0)
history = rng.normal(0, 1, N)
y = (1.3 * signal + 0.7 * history + rng.normal(0, 0.6, N) > 0).astype(int)
X = np.c_[signal, history]
idx = np.arange(N)
tr, te = train_test_split(idx, test_size=0.3, random_state=SEED, stratify=y)
model = LogisticRegression(max_iter=1000).fit(X[tr], y[tr])
pred = model.predict(X[te])
def slice_rows(mask, name):
n = int(mask.sum())
if n < 30:
return f"| {name} | {n} | too few rows to report | | |"
yt, yp = y[te][mask], pred[mask]
return (f"| {name} | {n} | {accuracy_score(yt, yp):.3f} "
f"| {recall_score(yt, yp, zero_division=0):.3f} "
f"| {precision_score(yt, yp, zero_division=0):.3f} |")
data_hash = hashlib.sha256(np.ascontiguousarray(X).tobytes()).hexdigest()[:12]
rows = [slice_rows(np.ones(len(te), bool), "ALL")]
for col, vals in (("region", ("urban", "rural")), ("device", ("new", "old"))):
arr = {"region": region, "device": device}[col][te]
for v in vals:
rows.append(slice_rows(arr == v, f"{col}={v}"))
for r in ("urban", "rural"): # the intersection, which marginals hide
for d in ("new", "old"):
rows.append(slice_rows((region[te] == r) & (device[te] == d), f"{r}+{d}"))
card = f"""# Model card: signup-risk-classifier v0.3
## Details
- Type: logistic regression (scikit-learn {sklearn.__version__}), 2 features
- Trained: {len(tr)} rows | Evaluated: {len(te)} held-out rows
- Random seed: {SEED} | Python {platform.python_version()}
- Feature-matrix fingerprint: sha256:{data_hash}
## Intended use
- IN SCOPE: ranking new signups for manual review.
- OUT OF SCOPE: automated rejection, credit decisions, anything irreversible.
- The model outputs a score. A human makes the decision.
## Metrics, disaggregated
| slice | n | accuracy | recall | precision |
|---|---|---|---|---|
{chr(10).join(rows)}
## Known limitations
- `region` and `device` are NOT model inputs. They are held for auditing only.
- Users on old devices are under-served: see the device=old row above.
- No data after the training cut-off. Re-evaluate before reusing.
## Ethical and safety notes
- Trained on synthetic data for teaching. Do not deploy.
- Contact: the owning team, listed in the service registry.
"""
print(card)# Model card: signup-risk-classifier v0.3 ## Details - Type: logistic regression (scikit-learn 1.7.2), 2 features - Trained: 2100 rows | Evaluated: 900 held-out rows - Random seed: 0 | Python 3.10.11 - Feature-matrix fingerprint: sha256:22862d73dfd7 ## Intended use - IN SCOPE: ranking new signups for manual review. - OUT OF SCOPE: automated rejection, credit decisions, anything irreversible. - The model outputs a score. A human makes the decision. ## Metrics, disaggregated | slice | n | accuracy | recall | precision | |---|---|---|---|---| | ALL | 900 | 0.874 | 0.865 | 0.867 | | region=urban | 725 | 0.873 | 0.865 | 0.865 | | region=rural | 175 | 0.880 | 0.866 | 0.877 | | device=new | 594 | 0.867 | 0.887 | 0.856 | | device=old | 306 | 0.889 | 0.808 | 0.898 | | urban+new | 491 | 0.874 | 0.895 | 0.860 | | urban+old | 234 | 0.872 | 0.783 | 0.878 | | rural+new | 103 | 0.835 | 0.852 | 0.836 | | rural+old | 72 | 0.944 | 0.893 | 0.962 | ## Known limitations - `region` and `device` are NOT model inputs. They are held for auditing only. - Users on old devices are under-served: see the device=old row above. - No data after the training cut-off. Re-evaluate before reusing. ## Ethical and safety notes - Trained on synthetic data for teaching. Do not deploy. - Contact: the owning team, listed in the service registry.
Your scikit-learn and Python lines will show your own versions. Every number in the metrics table is fixed by SEED and will match.
What the generated table is telling you
Accuracy is not the interesting column. device=old has higher accuracy than device=new — 0.889 against 0.867 — and lower recall, 0.808 against 0.887. It misses more of the positive cases while looking better overall. Report accuracy alone and this is invisible.
The intersection is worse than either marginal. urban+old recall is 0.783, lower than device=old at 0.808 and lower than region=urban at 0.865. Marginal slices average intersections away, which is why the card generates both.
rural+old looks excellent and should not be trusted. Recall 0.893 on n=72. A 95% interval on a proportion from 72 samples is roughly plus or minus 7 points. Printing n in every row is what lets a reader apply that scepticism themselves.
The n < 30 guard is doing real work. It prints the count and refuses the metric rather than emitting a number computed from a handful of rows. A refusal is information; a meaningless 1.000 is not.
Line by line
region and device are built but never enter X. That is the pattern from bias in datasets: keep the sensitive attribute for measurement, keep it out of the features.
hashlib.sha256(...) fingerprints the exact data the numbers came from. When someone asks in nine months whether the card matches the model in production, this line answers it. Hash your model file the same way.
platform.python_version() and sklearn.__version__ are recorded because metrics move between library versions. A card without an environment is not reproducible.
train_test_split(..., stratify=y) keeps class balance identical across the split, so slice metrics are comparable.
Common mistakes
Writing the card after launch. By then the numbers are from memory and the limitations section says "none known", which is never true.
Reporting only marginal groups. Compute the intersections. The worst cell is the one your incident report will be about.
Omitting n. A metric without a sample size cannot be judged, and small slices produce impressive numbers by accident.
Leaving "intended use" vague. "General purpose classification" gives a reader nothing to check against. Name the decision the model feeds and the decision it must not feed.
Never updating it. Wire card generation into your training pipeline and fail the build when the card is missing. See CI/CD for machine learning.
Documenting the model and skipping the data. A model card describes the model; a datasheet describes the dataset — how it was collected, who labelled it, who is missing. Both are needed, and the datasheet is usually harder to write.
Try it yourself
Add a threshold variable, set it to 0.35 instead of the implicit 0.5, and regenerate. Watch recall rise and precision fall in every row, by different amounts per slice. Then put the threshold on the card, because a card that omits it does not describe the deployed system.
What to learn next
- AI regulation — the rules that increasingly expect this document.
- Deploying responsibly — what happens after the card is written.
- CI/CD for machine learning — wiring card generation into the pipeline.
Researcher — Mathematics and papers.
The proposals and what each covers
Model Cards (Mitchell et al., FAT* 2019) specify nine sections: model details, intended use, factors, metrics, evaluation data, training data, quantitative analyses, ethical considerations, caveats and recommendations. The methodological contribution is the insistence on disaggregated evaluation across intersectional groups, motivated directly by Buolamwini and Gebru's Gender Shades finding that aggregate accuracy concealed a 30-point error gap.
Datasheets for Datasets (Gebru et al., 2018) covers the dataset: motivation, composition, collection process, preprocessing, uses, distribution, maintenance. Adapted from the electronics component datasheet, where the analogy is exact — the part works within stated operating conditions and not outside them.
Data Statements (Bender and Friedman, TACL 2018) target NLP specifically, with schedules for speaker and annotator demographics, curation rationale and language variety.
FactSheets (Arnold et al., IBM 2019) extend toward supplier's declarations of conformity, adding provenance, safety and lineage attributes for procurement contexts.
System Cards describe the deployed system rather than the model — the pre- and post-processing, the human review layer, the thresholds. The distinction matters because harm is a property of the system, and a model card that omits the threshold and the human loop describes something nobody deployed.
The empirical case for disaggregation
Aggregate metrics are averages over the deployment population, so they are dominated by the majority slice. The relevant results:
- Buolamwini and Gebru (2018): commercial gender classifiers at high aggregate accuracy with error rates of 0.8% on lighter-skinned men and up to 34.7% on darker-skinned women. The worst cell was intersectional, not visible on either marginal.
- Koenecke et al. (2020): commercial speech recognisers with roughly double the word error rate for Black speakers versus white speakers, aggregate performance reported as strong.
- Oakden-Rayner et al. (2020): "hidden stratification" in medical imaging — clinically critical subclasses with far worse performance than the aggregate, occurring at rates too low to move the headline number.
Oakden-Rayner et al. give the sharpest framing: the dangerous slices are precisely those rare enough not to affect aggregate metrics and severe enough to matter clinically. Slice discovery, not slice reporting, is the harder open problem, and automated approaches (Eyuboglu et al., 2022, Domino) remain research tools.
Reporting uncertainty
Slice metrics are estimates from small samples and should carry intervals. For a proportion $\hat{p}$ from $n$ samples, the Wilson score interval is preferred over the normal approximation, which is badly behaved near 0 and 1 and at small $n$:
$$ \frac{\hat{p} + \frac{z^2}{2n} \pm z\sqrt{\frac{\hat{p}(1-\hat{p})}{n} + \frac{z^2}{4n^2}}}{1 + \frac{z^2}{n}} $$
with $z$ the standard normal quantile (1.96 for 95%). For composite metrics such as F1 or AUC, bootstrap over the slice.
With many slices, multiple comparisons apply. Reporting the worst of twenty slices without correction will find an apparent disparity in a fair model at a predictable rate. Pre-register the slices, or report corrected intervals.
Documentation as a governance artefact
The regulatory direction is toward mandated technical documentation for higher-risk systems. The EU AI Act's Annex IV lists the technical documentation requirements for high-risk systems, and the NIST AI RMF's Measure and Manage functions assume the existence of documented evaluation. Model cards and datasheets are close in shape to what these frameworks call for, and adopting them early is cheaper than retrofitting.
Two honest limitations. Documentation is self-reported, so it constrains a careless team more than an adversarial one; the enforcement mechanism is external audit, which requires access these documents alone do not grant. And Bommasani et al. (2023), the Foundation Model Transparency Index, found published disclosure for large models to be sparse and inconsistent — the practice lags the proposal by a wide margin.
Papers
- Mitchell et al., Model Cards for Model Reporting, FAT* 2019 — arxiv.org/abs/1810.03993
- Gebru et al., Datasheets for Datasets, CACM 2021 — arxiv.org/abs/1803.09010
- Bender and Friedman, Data Statements for NLP, TACL 2018
- Oakden-Rayner et al., Hidden Stratification Causes Clinically Meaningful Failures, CHIL 2020 — arxiv.org/abs/1909.12475
- Koenecke et al., Racial disparities in automated speech recognition, PNAS 2020
- Raji et al., Closing the AI Accountability Gap, FAccT 2020 — arxiv.org/abs/2001.00973
- Eyuboglu et al., Domino: Discovering Systematic Errors with Cross-Modal Embeddings, ICLR 2022
What to learn next
- AI regulation — the rules that increasingly expect this document.
- Deploying responsibly — what happens after the card is written.
- CI/CD for machine learning — wiring card generation into the pipeline.