Disability and dataset bias
Disability-related dataset bias happens because disabled people appear rarely in most training data and vary enormously within any one category, so overall accuracy can hide poor performance for exactly the group most in need.
- 11 min read
- 3 reading levels
- Published
Read these first
On this page 6
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Disability-related dataset bias happens because disabled people appear rarely in the data most AI systems are trained on.
Think about a class photo taken on a day several students were sick and absent. The photo looks complete — it shows a full class of smiling faces. But it silently leaves out exactly the students who were not there that day. A dataset can do the same thing. A training set built mostly from typical, non-disabled examples looks complete, and quietly represents disabled people poorly or not at all.
Why it exists
Machine learning models learn from whatever data they are shown. Most large datasets — photos scraped from the internet, speech recordings, medical records — are dominated by non-disabled people. Disabled people are a smaller share of the population to begin with. They are further underrepresented in exactly the kinds of data that get collected easily and shared widely online.
This creates a specific, well-documented pattern, distinct from the general bias problem covered in bias in datasets. Disabled people are not one uniform group a model can learn to recognise once it sees enough examples. A wheelchair user, someone with a guide dog, someone with an atypical gait, and someone using a walking frame — these are all different. Each one looks, sounds, and moves in a distinct way.
A model needs real diversity within its disability-related examples, not one example type standing in for all of it. That kind of rich, representative data is exactly what most large datasets lack.
How it works
A dataset built mostly from typical, non-disabled examples
|
v train a model on it
The model learns patterns that fit the majority well
|
v test it on disabled users specifically
Performance is often noticeably worse for this group
|
v but the OVERALL accuracy number can still look fine
because disabled users were always a small share of the test data tooWhere you have already seen it
Pedestrian-detection systems are usually trained mostly on images of standing adults. Researchers have repeatedly found they perform measurably worse on disabled pedestrians — people using wheelchairs, canes, or walking with an atypical gait. That is a real safety concern for any system meant to detect people near a vehicle or a robot. Voice assistants trained mostly on typical speech, discussed in an earlier lesson on atypical speech recognition, show the same underlying pattern in a different modality.
An honest warning
An overall accuracy number can hide exactly this problem. A model that performs excellently for 95% of users and poorly for the remaining 5% can still report a very high overall accuracy score. That 5% barely moves the average.
Anyone building or evaluating a system that might be used by disabled people has to check performance broken down by group, not only in aggregate. Otherwise this exact failure mode goes completely unnoticed, until it causes real harm.
Remember this
- Disabled people are underrepresented in most large training datasets, both in raw numbers and in how much real variation is captured.
- Disability is not one uniform category — a model needs genuine diversity within disability-related examples, not one example standing in for all of it.
- A strong overall accuracy score can hide poor performance for a small group entirely — always check performance broken down by group, not only the average.
What to learn next
- Evaluating assistive AI with real users — the evaluation practices that catch exactly the gap this lesson describes.
- Bias in datasets — the general version of the problem, covering other underrepresented groups.
- Fairness metrics — ways to measure performance gaps between groups directly, instead of relying on one overall number.
Developer — Code and libraries.
This example makes the hidden-behind-the-average problem directly visible: one model, one dataset, two very different pictures depending on which number you look at.
Setup
pip install scikit-learn numpyMinimal runnable code
A toy pedestrian-detector training set with two overlapping but distinguishable feature clusters, standing in for real visual features — one for standing pedestrians, one for wheelchair users, deliberately skewed so wheelchair users are a small minority of the data.
import numpy as np
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import recall_score
from sklearn.model_selection import train_test_split
rng = np.random.default_rng(0)
# A toy pedestrian-detector training set: two overlapping but distinguishable
# feature clusters (a stand-in for real visual features), one for standing
# pedestrians and one for wheelchair users. The SHAPE of the problem here is
# real, even though these exact numbers are invented: public vision datasets
# are widely reported to under-represent mobility aids, canes and guide dogs
# relative to standing adults -- see dataset-bias-and-shortcut-learning.
def make_features(n, centre, rng):
return centre + rng.normal(0, 1.0, size=(n, 2))
n_standing, n_wheelchair = 950, 50 # an illustrative skew, chosen to be severe -- real datasets vary
X = np.vstack([
make_features(n_standing, [0, 0], rng),
make_features(n_wheelchair, [3, 1], rng),
])
y = np.array(["standing"] * n_standing + ["wheelchair"] * n_wheelchair)
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.3, random_state=0, stratify=y)
plain = LogisticRegression().fit(X_train, y_train)
plain_recall = recall_score(y_test, plain.predict(X_test), pos_label="wheelchair")
print(f"trained on the natural 95:5 skew -> wheelchair-user recall: {plain_recall:.0%}")
balanced = LogisticRegression(class_weight="balanced").fit(X_train, y_train)
balanced_recall = recall_score(y_test, balanced.predict(X_test), pos_label="wheelchair")
print(f"same data, class_weight='balanced' -> wheelchair-user recall: {balanced_recall:.0%}")
overall_accuracy = plain.score(X_test, y_test)
print(f"\noverall accuracy of the first (unbalanced) model: {overall_accuracy:.0%}")
print("that overall number looks fine -- it hides exactly the group the model is failing.")trained on the natural 95:5 skew -> wheelchair-user recall: 60% same data, class_weight='balanced' -> wheelchair-user recall: 93% overall accuracy of the first (unbalanced) model: 97% that overall number looks fine -- it hides exactly the group the model is failing.
Walkthrough
The unbalanced model reaches 97% overall accuracy — a number that would pass almost any casual review as "working well." Look one line up, though: its recall for wheelchair users specifically — how often it correctly identifies someone who is actually a wheelchair user — is only 60%. It misses roughly two out of every five wheelchair users in this test set, while the overall number stays high, because wheelchair users are only 5% of the data and a majority-class model can afford to get most of that small minority wrong without denting its aggregate score much at all.
class_weight="balanced" tells the model to weight errors on the minority class more heavily during training, correcting for the raw imbalance in the data. With that one change, recall for wheelchair users jumps to 93%, at essentially no engineering cost beyond the choice to check for and address the gap in the first place.
Common mistakes
Reporting only overall accuracy. This is the exact trap the developer example is built to demonstrate — a model can look excellent in aggregate while badly failing a specific, smaller group.
Assuming the skew in this toy example is exaggerated for effect. The severity of underrepresentation varies by dataset and category, and the direction of the problem — disability-related categories being a small, often overlooked minority in most general-purpose datasets — is a widely documented, real pattern, not a contrivance invented for this lesson.
Fixing the skew in the model without also improving the data. class_weight="balanced" helps a model trained on the data it already has. It does nothing to fix a dataset that lacks real diversity within the minority class itself — five copies of the same kind of wheelchair user teach a model far less than five genuinely different examples.
Checking group-level performance once and considering the problem solved. Real deployments benefit from ongoing monitoring broken down by group, since a system that performs well at launch can drift as real-world usage patterns shift.
Try it yourself
Change n_wheelchair from 50 to 500, keeping n_standing at 950, and rerun both models. Compare how much the unbalanced model's recall improves purely from having more minority-class examples available, even without class_weight="balanced" — a direct illustration of why collecting more representative data matters as much as any modelling technique.
What to learn next
- Fairness metrics — formal ways to measure and report performance gaps between groups.
- Class weights — the general imbalance-correction technique used in the code above.
- Evaluating assistive AI with real users — why testing with real, diverse users catches gaps a held-out test set alone can miss.
Researcher — Mathematics and papers.
The formal setting
Let a dataset be drawn from a mixture of subpopulations G_1, ..., G_k (e.g. standing pedestrians, wheelchair users, cane users), with true proportions π_1, ..., π_k in the real deployment population and observed training proportions p̂_1, ..., p̂_k, where p̂_i ≪ π_i for disability-related groups is a widely reported empirical pattern rather than a universal law. A model trained via standard empirical risk minimisation optimises:
θ_hat = argmin_θ (1/n) Σ_i L(f_θ(x_i), y_i)- This objective weights every training example equally, so groups with small
p̂_icontribute proportionally little to the overall loss, and the optimiser has correspondingly little pressure to fit them well - Overall accuracy on a test set drawn from the same skewed distribution inherits this bias, since a small group also contributes little to the test metric — exactly why the developer example's 97% overall accuracy coexists with 60% minority-group recall
Beyond underrepresentation: intra-category diversity
A subtler, disability-specific issue beyond raw underrepresentation is captured by Trewin (2018), who argues that standard group-fairness framings (parity of some metric across a small number of discrete groups) fit disability categories poorly, because disability is not one attribute with a small number of values — it spans an extremely wide range of conditions, severities, and individual adaptations, many of which co-occur, and disability status is often not disclosed due to legitimate privacy and discrimination concerns, making standard fairness-auditing techniques that assume known group membership difficult to apply directly.
Documented real-world instances
The AI Now Institute's "Disability, Bias, and AI" report (Whittaker et al., 2019) documents this pattern across multiple deployed system categories: hiring algorithms that penalise atypical speech or movement patterns picked up in video interviews, benefits-eligibility algorithms trained on historical data that embeds prior discriminatory decisions, and computer vision systems for autonomous vehicles and robotics whose pedestrian-detection training data underrepresents wheelchair users, cane users, and people with atypical gaits — the exact category the developer example simplifies into a toy classification problem.
Complexity and cost of remediation
For a dataset with minority group proportion p̂ in a total of n examples:
| Remediation approach | What it addresses | What it does not address |
|---|---|---|
| Class-weighted or cost-sensitive loss (developer example) | Underweighting of minority-class errors during training | Lack of diversity within the minority class itself |
| Oversampling minority examples | Same as above, via resampling instead of loss weighting | Same limitation — resampling existing examples adds no new information |
| Targeted data collection | Both underrepresentation and lack of diversity | Cost, time, and the ethical requirement of informed, well-compensated participation from disabled contributors |
| Post-hoc threshold adjustment per group | Calibration differences between groups at deployment | Nothing about the underlying training data or representation |
Only targeted, genuinely diverse data collection addresses both problems at once, which is why it remains the recommended remediation in the literature despite being the most costly option, and why algorithmic fixes like class weighting are described as mitigations rather than solutions.
Papers
- Buolamwini, J. and Gebru, T. (2018). Gender Shades: Intersectional Accuracy Disparities in Commercial Gender Classification. FAT* (now FAccT). The foundational demonstration of the "good overall accuracy, poor subgroup accuracy" pattern this lesson's developer example reproduces in miniature, in a different domain.
- Trewin, S. (2018). AI Fairness for People with Disabilities: Point of View. arXiv:1811.10670. Argues standard group-fairness frameworks fit disability poorly, for the reasons described above.
- Whittaker, M. et al. (2019). Disability, Bias, and AI. AI Now Institute report. Documents real, deployed instances of the underrepresentation pattern this lesson addresses.
Current state
Fairness evaluation broken down by disability status remains far less standardised in industry practice than evaluation broken down by race or gender, in part because of the genuine measurement difficulty Trewin (2018) raises: disability is highly heterogeneous, often undisclosed, and does not map cleanly onto the small-number-of-discrete-groups framing most fairness metrics assume. Participatory design and data collection practices that directly involve disabled contributors, rather than attempting to infer disability status from existing data after the fact, are increasingly recommended as the more tractable and more ethical path forward, and are the direct subject of the next lesson.
What to learn next
- Evaluating assistive AI with real users — the participatory evaluation practices that directly address the measurement gap this lesson raises.
- Fairness metrics — the formal metrics literature, and their known limitations for disability specifically.
- Bias in datasets — the general underrepresentation problem this lesson specialises.