Evaluating assistive AI with real users
Evaluating assistive AI means judging it by whether it actually helps the person using it, since standard accuracy metrics can rank a technically better model below one that genuinely works better for the task.
- 11 min read
- 3 reading levels
- Published
Read these first
On this page 6
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Evaluating assistive AI means judging it by whether it actually helps the person relying on it, not by an accuracy number alone.
Think about judging a walking stick. You would not judge it by how solidly it is built in a lab test. You would judge it by whether it actually helps someone avoid tripping on a real, uneven street. A number on a spec sheet is not the real test.
Assistive AI needs the same kind of judgment. The right question is not "how accurate is this model on a test set." It is "does this actually work for the person using it, doing the real task they need it for."
Why it exists
Ordinary machine learning evaluation, covered in model evaluation, asks how often a model gets the right answer on a held-out test set. That is a reasonable question for many tasks. It is often the wrong question for assistive technology, for a reason repeated throughout this whole section. The cost of different mistakes is rarely equal, and a single accuracy number blends every kind of mistake together, as if they cost the same.
Picture two navigation aids: one misses one obstacle out of a hundred, the other gives one false alarm out of a hundred. Both could report the identical 99% accuracy score. One of those failures is a minor annoyance. The other could be someone walking into something dangerous.
Evaluating assistive AI properly means measuring the things that actually matter for the specific task. It means testing with the actual people the technology is meant to serve.
How it works
Standard evaluation: "How accurate is this model on a test set?"
|
v ONE number, blending every kind of mistake together
Often misses which mistakes matter most, and to whom
Assistive AI evaluation: "Does this help someone complete the REAL task?"
|
v test with real, diverse users doing the real task
v weigh different mistakes by their real cost
v measure task success, not only prediction accuracy
A far more honest picture of whether the system actually worksWhere you have already seen it
Accessibility researchers regularly run studies where blind or low-vision participants use a navigation aid or reading tool in realistic conditions. Success is measured by whether the person completed the task safely and confidently — not by the underlying model's accuracy score in isolation.
The Deaf community's long-standing principle "nothing about us without us" reflects the same core idea in a broader form. The people a technology is built for need to be genuinely involved in judging whether it actually works.
An honest warning
It is tempting to ship an assistive AI feature based only on a strong benchmark accuracy score. Benchmark evaluation is fast, cheap, and easy to run repeatedly.
Real evaluation with real, diverse users is slower and more effortful. But it is also the only evaluation that reliably catches one kind of failure. That failure is a model that looks excellent on paper, and fails the specific people it was meant to help.
Remember this
- Standard accuracy metrics blend every kind of mistake together, even when different mistakes have very different real-world costs.
- Assistive AI should be evaluated by whether it helps someone complete the real task, tested with real, diverse users doing that real task.
- A strong benchmark score is not proof of real-world usefulness. It is a hypothesis that still needs testing with the people the technology is meant to serve.
What to learn next
- Choosing a threshold from costs — the general technique behind weighing different mistakes by their real cost, demonstrated directly in this lesson's code.
- Disability and dataset bias — why testing with real, diverse users catches problems that a held-out test set alone can miss.
- ROC vs precision-recall curves — a deeper look at the trade-offs standard accuracy hides.
Developer — Code and libraries.
This example compares a model chosen by raw accuracy against the same model tuned to minimise a realistic cost — the same underlying idea threaded through several earlier lessons in this section, made explicit and measurable here.
Setup
pip install scikit-learn numpyMinimal runnable code
A toy fall-detection wearable predicts "fall" versus "no fall" from a sensor reading. Falls are made rare in the synthetic data, which is realistic for this kind of alert system, even though the specific numbers here are invented for the demonstration.
import numpy as np
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import accuracy_score
rng = np.random.default_rng(0)
# A toy fall-detection wearable: predict "fall" vs "no fall" from a sensor
# reading. Falls are rare, which is realistic for this kind of alert system,
# even though these particular numbers are invented for the demonstration.
n_no_fall, n_fall = 400, 20
X_no_fall = rng.normal(0, 1, size=(n_no_fall, 1))
X_fall = rng.normal(2.2, 1, size=(n_fall, 1))
X = np.vstack([X_no_fall, X_fall])
y = np.array([0] * n_no_fall + [1] * n_fall)
model = LogisticRegression().fit(X, y)
probabilities = model.predict_proba(X)[:, 1]
# a false negative (a real fall, missed) is far worse than a false alarm --
# someone lying injured and unable to call for help, versus an unnecessary check-in
COST_FALSE_NEGATIVE = 50
COST_FALSE_POSITIVE = 1
def expected_cost(threshold):
predicted_fall = probabilities >= threshold
false_negatives = np.sum((predicted_fall == 0) & (y == 1))
false_positives = np.sum((predicted_fall == 1) & (y == 0))
return false_negatives * COST_FALSE_NEGATIVE + false_positives * COST_FALSE_POSITIVE
thresholds = np.arange(0.1, 0.95, 0.05)
best_accuracy_threshold = max(thresholds, key=lambda t: accuracy_score(y, probabilities >= t))
best_cost_threshold = min(thresholds, key=expected_cost)
print(f"threshold that maximises raw accuracy: {best_accuracy_threshold:.2f} "
f"(accuracy={accuracy_score(y, probabilities >= best_accuracy_threshold):.1%}, "
f"cost={expected_cost(best_accuracy_threshold)})")
print(f"threshold that minimises the real cost: {best_cost_threshold:.2f} "
f"(accuracy={accuracy_score(y, probabilities >= best_cost_threshold):.1%}, "
f"cost={expected_cost(best_cost_threshold)})")threshold that maximises raw accuracy: 0.40 (accuracy=96.7%, cost=553) threshold that minimises the real cost: 0.10 (accuracy=91.9%, cost=230)
Walkthrough
The threshold that maximises raw accuracy, 0.40, reaches an impressive-looking 96.7% — and, checked against the real cost of its mistakes, actually costs more than double the alternative. The threshold that minimises real cost, 0.10, reports a noticeably lower accuracy of 91.9%, and yet performs meaningfully better by the metric that actually reflects what a missed fall versus a false alarm costs in the real world.
COST_FALSE_NEGATIVE = 50 and COST_FALSE_POSITIVE = 1 encode the same asymmetry running through this entire section: missing something dangerous is treated as far worse than a harmless extra check-in. Lowering the decision threshold to 0.10 makes the model flag "fall" more readily, catching more true falls at the cost of more false alarms — exactly the direction this specific cost structure calls for, and exactly the opposite of what "maximise accuracy" alone would recommend.
Common mistakes
Choosing a decision threshold by accuracy alone. The entire lesson of this example: the accuracy-best threshold and the cost-best threshold are different values here, and accuracy alone would have picked the worse one for real use.
Assuming the cost numbers here are the correct ones for every situation. COST_FALSE_NEGATIVE = 50 and COST_FALSE_POSITIVE = 1 are illustrative, chosen to make the trade-off visible. Real cost values need to come from genuine discussion with the people affected, not an arbitrary round number picked by an engineer alone.
Treating this kind of cost-weighted analysis as a substitute for testing with real users. Cost-weighted threshold selection is a genuinely useful technique, and it still relies on a model of the world — the assumed costs, the assumed distribution of falls versus non-falls — that only real-world testing with real users can actually validate.
Optimising a single global threshold when the right cost trade-off might differ by person. A younger, more mobile user and an elderly user living alone might reasonably want different sensitivity settings for the same underlying model, given how differently a missed fall could affect each of them.
Try it yourself
Change COST_FALSE_NEGATIVE to 10 and rerun, simulating a situation where a missed fall is considered less catastrophic — perhaps because a caregiver checks in frequently regardless. Watch how the cost-optimal threshold shifts, and think about what that shift implies about who gets to decide these cost values in a real deployment.
What to learn next
- Choosing a threshold from costs — the general technique this lesson's code applies directly.
- ROC vs precision-recall curves — visualising the full trade-off between false positives and false negatives across every possible threshold.
- Model evaluation — the general evaluation practices this lesson specialises for assistive AI's specific cost structure.
Researcher — Mathematics and papers.
The formal setting
Given a binary classifier producing predicted probability p(x) for a positive outcome, a decision threshold τ converts this into a hard prediction. Standard accuracy treats every misclassification identically:
Accuracy(τ) = (TP(τ) + TN(τ)) / NA cost-sensitive framing instead defines an expected cost function using distinct costs for each error type:
Cost(τ) = FN(τ) · c_FN + FP(τ) · c_FPc_FN, c_FP— the real-world cost of a false negative and false positive respectively, which need not be equal and, in assistive AI, are very often wildly unequal- The threshold minimising
Accuracyand the threshold minimisingCostcoincide only in the special casec_FN = c_FP— otherwise, as the developer example shows directly, they diverge, and accuracy-optimal is not cost-optimal
This is a direct instance of the general cost-sensitive decision theory covered in choosing a threshold from costs, specialised here to the particular, often extreme cost asymmetries that recur throughout assistive AI: a missed obstacle, a missed fall, a wrongly withheld image description.
Beyond cost-weighted metrics: task-based evaluation
Cost-weighted accuracy is still a proxy metric computed on held-out data, not a direct measurement of real-world usefulness. Accessibility research has established a complementary, more direct standard: task-based evaluation with representative users, measuring whether a participant actually completes a realistic task successfully, at what time cost, and with what subjective confidence — rather than evaluating the underlying model in isolation from the person using it (Sears and Hanson, 2012, on the broader methodological argument for representative users in accessibility research specifically).
Participatory evaluation and its rationale
Bragg et al. (2019), discussed in the sign language recognition lesson, and Whittaker et al. (2019), discussed in the dataset bias lesson, both converge on the same methodological conclusion from different angles: technology built for a disabled population, evaluated only by researchers or engineers outside that population, systematically misses failure modes and priorities the target users would have identified immediately. This is the technical justification behind the disability-rights principle "nothing about us without us" — it names not only an ethical claim but a genuine measurement gap that non-participatory evaluation reliably produces.
Complexity and practical trade-offs
| Evaluation approach | What it measures | Cost | What it misses |
|---|---|---|---|
| Held-out accuracy | Prediction correctness on historical data | Low, fully automatable | Real-world cost asymmetry, usability, distribution shift to real conditions |
| Cost-weighted threshold selection (developer example) | Prediction correctness, weighted by assumed real costs | Low, automatable once costs are estimated | Whether the assumed costs are actually correct |
| Task-based user study | Real task success, with real users | High, requires recruiting representative participants and running real sessions | Scale — hard to run continuously or on every model update |
Mature deployments typically combine all three: automated, cost-weighted metrics for continuous monitoring, paired with periodic task-based studies with real, representative users to validate that the automated metrics are still tracking what actually matters.
Papers
- Sears, A. and Hanson, V. (2012). Representing Users in Accessibility Research. ACM Transactions on Accessible Computing 4(2). The methodological case for evaluating with genuinely representative users rather than convenience samples.
- Bragg, D. et al. (2019). Sign Language Recognition, Generation, and Translation: An Interdisciplinary Perspective. ASSETS.
- Whittaker, M. et al. (2019). Disability, Bias, and AI. AI Now Institute report.
Current state
Task-based, participatory evaluation with representative users is increasingly treated as the accepted standard in accessibility research specifically, even though it remains more expensive and slower than automated benchmark evaluation, and is consequently still inconsistently applied across the wider AI industry building assistive features. The persistent, honest tension this section has returned to repeatedly — a technically strong model is not the same thing as a technology that genuinely works for the person relying on it — is not a solved problem, and every lesson in this section has been, in one way or another, about taking that gap seriously rather than assuming a good benchmark score closes it.
What to learn next
- Choosing a threshold from costs — the full general treatment of the cost-sensitive decision theory used throughout this lesson.
- Disability and dataset bias — the representation problem that makes representative user testing especially necessary in this domain.
- Alt text generation — the first lesson in this section, worth revisiting with this lesson's evaluation lens in mind.