AI for Accessibility

Disability and dataset bias

Disability-related dataset bias happens because disabled people appear rarely in most training data and vary enormously within any one category, so overall accuracy can hide poor performance for exactly the group most in need.

Read these first

On this page 6
  1. Why it exists
  2. How it works
  3. Where you have already seen it
  4. An honest warning
  5. Remember this
  6. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Disability-related dataset bias happens because disabled people appear rarely in the data most AI systems are trained on.

Think about a class photo taken on a day several students were sick and absent. The photo looks complete — it shows a full class of smiling faces. But it silently leaves out exactly the students who were not there that day. A dataset can do the same thing. A training set built mostly from typical, non-disabled examples looks complete, and quietly represents disabled people poorly or not at all.

Why it exists

Machine learning models learn from whatever data they are shown. Most large datasets — photos scraped from the internet, speech recordings, medical records — are dominated by non-disabled people. Disabled people are a smaller share of the population to begin with. They are further underrepresented in exactly the kinds of data that get collected easily and shared widely online.

This creates a specific, well-documented pattern, distinct from the general bias problem covered in bias in datasets. Disabled people are not one uniform group a model can learn to recognise once it sees enough examples. A wheelchair user, someone with a guide dog, someone with an atypical gait, and someone using a walking frame — these are all different. Each one looks, sounds, and moves in a distinct way.

A model needs real diversity within its disability-related examples, not one example type standing in for all of it. That kind of rich, representative data is exactly what most large datasets lack.

How it works

A dataset built mostly from typical, non-disabled examples
        |
        v  train a model on it
The model learns patterns that fit the majority well
        |
        v  test it on disabled users specifically
Performance is often noticeably worse for this group
        |
        v  but the OVERALL accuracy number can still look fine
because disabled users were always a small share of the test data too

Where you have already seen it

Pedestrian-detection systems are usually trained mostly on images of standing adults. Researchers have repeatedly found they perform measurably worse on disabled pedestrians — people using wheelchairs, canes, or walking with an atypical gait. That is a real safety concern for any system meant to detect people near a vehicle or a robot. Voice assistants trained mostly on typical speech, discussed in an earlier lesson on atypical speech recognition, show the same underlying pattern in a different modality.

An honest warning

An overall accuracy number can hide exactly this problem. A model that performs excellently for 95% of users and poorly for the remaining 5% can still report a very high overall accuracy score. That 5% barely moves the average.

Anyone building or evaluating a system that might be used by disabled people has to check performance broken down by group, not only in aggregate. Otherwise this exact failure mode goes completely unnoticed, until it causes real harm.

Remember this

  • Disabled people are underrepresented in most large training datasets, both in raw numbers and in how much real variation is captured.
  • Disability is not one uniform category — a model needs genuine diversity within disability-related examples, not one example standing in for all of it.
  • A strong overall accuracy score can hide poor performance for a small group entirely — always check performance broken down by group, not only the average.

What to learn next

Developer — Code and libraries.

This example makes the hidden-behind-the-average problem directly visible: one model, one dataset, two very different pictures depending on which number you look at.

Setup

bash
pip install scikit-learn numpy

Minimal runnable code

A toy pedestrian-detector training set with two overlapping but distinguishable feature clusters, standing in for real visual features — one for standing pedestrians, one for wheelchair users, deliberately skewed so wheelchair users are a small minority of the data.

databias.py
import numpy as np
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import recall_score
from sklearn.model_selection import train_test_split

rng = np.random.default_rng(0)

# A toy pedestrian-detector training set: two overlapping but distinguishable
# feature clusters (a stand-in for real visual features), one for standing
# pedestrians and one for wheelchair users. The SHAPE of the problem here is
# real, even though these exact numbers are invented: public vision datasets
# are widely reported to under-represent mobility aids, canes and guide dogs
# relative to standing adults -- see dataset-bias-and-shortcut-learning.
def make_features(n, centre, rng):
    return centre + rng.normal(0, 1.0, size=(n, 2))

n_standing, n_wheelchair = 950, 50   # an illustrative skew, chosen to be severe -- real datasets vary
X = np.vstack([
    make_features(n_standing, [0, 0], rng),
    make_features(n_wheelchair, [3, 1], rng),
])
y = np.array(["standing"] * n_standing + ["wheelchair"] * n_wheelchair)

X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.3, random_state=0, stratify=y)

plain = LogisticRegression().fit(X_train, y_train)
plain_recall = recall_score(y_test, plain.predict(X_test), pos_label="wheelchair")
print(f"trained on the natural 95:5 skew  -> wheelchair-user recall: {plain_recall:.0%}")

balanced = LogisticRegression(class_weight="balanced").fit(X_train, y_train)
balanced_recall = recall_score(y_test, balanced.predict(X_test), pos_label="wheelchair")
print(f"same data, class_weight='balanced' -> wheelchair-user recall: {balanced_recall:.0%}")

overall_accuracy = plain.score(X_test, y_test)
print(f"\noverall accuracy of the first (unbalanced) model: {overall_accuracy:.0%}")
print("that overall number looks fine -- it hides exactly the group the model is failing.")
Output
trained on the natural 95:5 skew  -> wheelchair-user recall: 60%
same data, class_weight='balanced' -> wheelchair-user recall: 93%

overall accuracy of the first (unbalanced) model: 97%
that overall number looks fine -- it hides exactly the group the model is failing.

Walkthrough

The unbalanced model reaches 97% overall accuracy — a number that would pass almost any casual review as "working well." Look one line up, though: its recall for wheelchair users specifically — how often it correctly identifies someone who is actually a wheelchair user — is only 60%. It misses roughly two out of every five wheelchair users in this test set, while the overall number stays high, because wheelchair users are only 5% of the data and a majority-class model can afford to get most of that small minority wrong without denting its aggregate score much at all.

class_weight="balanced" tells the model to weight errors on the minority class more heavily during training, correcting for the raw imbalance in the data. With that one change, recall for wheelchair users jumps to 93%, at essentially no engineering cost beyond the choice to check for and address the gap in the first place.

Common mistakes

Reporting only overall accuracy. This is the exact trap the developer example is built to demonstrate — a model can look excellent in aggregate while badly failing a specific, smaller group.

Assuming the skew in this toy example is exaggerated for effect. The severity of underrepresentation varies by dataset and category, and the direction of the problem — disability-related categories being a small, often overlooked minority in most general-purpose datasets — is a widely documented, real pattern, not a contrivance invented for this lesson.

Fixing the skew in the model without also improving the data. class_weight="balanced" helps a model trained on the data it already has. It does nothing to fix a dataset that lacks real diversity within the minority class itself — five copies of the same kind of wheelchair user teach a model far less than five genuinely different examples.

Checking group-level performance once and considering the problem solved. Real deployments benefit from ongoing monitoring broken down by group, since a system that performs well at launch can drift as real-world usage patterns shift.

Try it yourself

Change n_wheelchair from 50 to 500, keeping n_standing at 950, and rerun both models. Compare how much the unbalanced model's recall improves purely from having more minority-class examples available, even without class_weight="balanced" — a direct illustration of why collecting more representative data matters as much as any modelling technique.

What to learn next

Researcher — Mathematics and papers.

The formal setting

Let a dataset be drawn from a mixture of subpopulations G_1, ..., G_k (e.g. standing pedestrians, wheelchair users, cane users), with true proportions π_1, ..., π_k in the real deployment population and observed training proportions p̂_1, ..., p̂_k, where p̂_i ≪ π_i for disability-related groups is a widely reported empirical pattern rather than a universal law. A model trained via standard empirical risk minimisation optimises:

θ_hat = argmin_θ  (1/n) Σ_i  L(f_θ(x_i), y_i)
  • This objective weights every training example equally, so groups with small p̂_i contribute proportionally little to the overall loss, and the optimiser has correspondingly little pressure to fit them well
  • Overall accuracy on a test set drawn from the same skewed distribution inherits this bias, since a small group also contributes little to the test metric — exactly why the developer example's 97% overall accuracy coexists with 60% minority-group recall

Beyond underrepresentation: intra-category diversity

A subtler, disability-specific issue beyond raw underrepresentation is captured by Trewin (2018), who argues that standard group-fairness framings (parity of some metric across a small number of discrete groups) fit disability categories poorly, because disability is not one attribute with a small number of values — it spans an extremely wide range of conditions, severities, and individual adaptations, many of which co-occur, and disability status is often not disclosed due to legitimate privacy and discrimination concerns, making standard fairness-auditing techniques that assume known group membership difficult to apply directly.

Documented real-world instances

The AI Now Institute's "Disability, Bias, and AI" report (Whittaker et al., 2019) documents this pattern across multiple deployed system categories: hiring algorithms that penalise atypical speech or movement patterns picked up in video interviews, benefits-eligibility algorithms trained on historical data that embeds prior discriminatory decisions, and computer vision systems for autonomous vehicles and robotics whose pedestrian-detection training data underrepresents wheelchair users, cane users, and people with atypical gaits — the exact category the developer example simplifies into a toy classification problem.

Complexity and cost of remediation

For a dataset with minority group proportion p̂ in a total of n examples:

Remediation approachWhat it addressesWhat it does not address
Class-weighted or cost-sensitive loss (developer example)Underweighting of minority-class errors during trainingLack of diversity within the minority class itself
Oversampling minority examplesSame as above, via resampling instead of loss weightingSame limitation — resampling existing examples adds no new information
Targeted data collectionBoth underrepresentation and lack of diversityCost, time, and the ethical requirement of informed, well-compensated participation from disabled contributors
Post-hoc threshold adjustment per groupCalibration differences between groups at deploymentNothing about the underlying training data or representation

Only targeted, genuinely diverse data collection addresses both problems at once, which is why it remains the recommended remediation in the literature despite being the most costly option, and why algorithmic fixes like class weighting are described as mitigations rather than solutions.

Papers

  • Buolamwini, J. and Gebru, T. (2018). Gender Shades: Intersectional Accuracy Disparities in Commercial Gender Classification. FAT* (now FAccT). The foundational demonstration of the "good overall accuracy, poor subgroup accuracy" pattern this lesson's developer example reproduces in miniature, in a different domain.
  • Trewin, S. (2018). AI Fairness for People with Disabilities: Point of View. arXiv:1811.10670. Argues standard group-fairness frameworks fit disability poorly, for the reasons described above.
  • Whittaker, M. et al. (2019). Disability, Bias, and AI. AI Now Institute report. Documents real, deployed instances of the underrepresentation pattern this lesson addresses.

Current state

Fairness evaluation broken down by disability status remains far less standardised in industry practice than evaluation broken down by race or gender, in part because of the genuine measurement difficulty Trewin (2018) raises: disability is highly heterogeneous, often undisclosed, and does not map cleanly onto the small-number-of-discrete-groups framing most fairness metrics assume. Participatory design and data collection practices that directly involve disabled contributors, rather than attempting to infer disability status from existing data after the fact, are increasingly recommended as the more tractable and more ethical path forward, and are the direct subject of the next lesson.

What to learn next