AI for Accessibility

Generating alt text that is actually useful

Alt text generation describes an image in words for a screen reader, and the honest version says less rather than guessing wrong with confidence.

On this page 6
  1. Why it exists
  2. How it works
  3. Where you have already seen it
  4. An honest warning
  5. Remember this
  6. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Alt text generation writes a short description of an image, for someone who cannot see the image.

Think about describing a photo over a phone call to a friend who cannot see your screen. You would not read out every pixel. You would say what matters — "it's a chart showing sales going up." And you would be honest when you were not sure what you were looking at.

A screen reader does the same job for a blind or low-vision person browsing the web. It reads aloud a short text description, called alt text, in place of every image.

Why it exists

Millions of people browse the internet using a screen reader — software that reads a page aloud instead of displaying it. When an image has no description attached, most screen readers say only "image" or read the filename, something like "IMG4821.jpg." That tells the listener nothing about what is actually there.

Website builders are supposed to write a short description for every image by hand. In practice, a huge number of images on the internet have no description at all. Writing one for every single image, on every page, takes real time nobody has budgeted for. Automatic alt text generation exists to fill that gap — not perfectly, but far better than silence or a meaningless filename.

How it works

An image  ->  A model looks at it  ->  "A chart showing data"  (with a confidence level)

A well-built system does one more thing that matters as much as the description itself. When it is not confident, it says so, or says nothing misleading at all — rather than inventing a detailed description that might be wrong.

Where you have already seen it

Facebook and Instagram generate automatic descriptions for photos that have no caption. Screen reader users hear these read aloud as "may show" statements — deliberately hedged language, not a confident claim. Google Photos and several browser extensions offer similar automatic descriptions for images that would otherwise be silent.

An honest warning

Alt text is read aloud with no image beside it to double-check against. If a system confidently describes a chart's numbers wrong, or misreads a warning sign, the listener has no way to catch the error themselves. They are trusting the description completely. A wrong, confident description can be worse than no description at all, because it is treated as fact.

Remember this

  • Alt text generation gives screen reader users a short spoken description of an image that would otherwise be silent or meaningless.
  • Millions of real images online have no description, which is the practical problem automatic alt text solves.
  • A confident, wrong description can mislead a listener who has no way to check it — honest uncertainty matters as much as accuracy.

What to learn next

  • Live captioning — the same honesty-under-uncertainty problem, for spoken audio instead of images.
  • Image captioning — the general computer vision technique alt text generation builds on.
  • Vision-language models — the modern model family behind most current alt text systems.

Developer — Code and libraries.

Real alt text systems use a trained vision-language model. This example builds a small, honest version of the decision every such system has to make: what to say, and when to say nothing instead.

Setup

bash
pip install scikit-learn

Minimal runnable code

Each "image" here is described by two simple features standing in for what a real vision model would extract: how much of the image is plain background, and how many straight edges it contains. The model classifies the image type, and the confidence score decides whether a description is safe to give.

alttext.py
from sklearn.neighbors import KNeighborsClassifier

# Each row is one image, described by two simple features:
# [fraction of pixels that are pure white background, count of straight edges]
# Real systems use a vision model for this; the features stand in for that.
images = [
    [0.85, 40], [0.80, 55], [0.90, 35], [0.83, 48], [0.78, 50],   # chart
    [0.10, 5],  [0.15, 8],  [0.05, 3],  [0.20, 10], [0.12, 6],    # photo
    [0.92, 60], [0.88, 70], [0.95, 65], [0.90, 58], [0.86, 62],   # receipt
]
labels = (["chart"] * 5) + (["photo"] * 5) + (["receipt"] * 5)

model = KNeighborsClassifier(n_neighbors=5)
model.fit(images, labels)

TEMPLATES = {
    "chart": "A chart showing data. Exact values are not available from the image alone.",
    "photo": "A photograph. Automatic description is not confident enough to describe its content.",
    "receipt": "A receipt or document with printed text. Use an OCR reader for the exact text.",
}

def describe(image_features, confidence_floor=0.7):
    probabilities = model.predict_proba([image_features])[0]
    best_index = probabilities.argmax()
    best_label = model.classes_[best_index]
    confidence = probabilities[best_index]
    if confidence < confidence_floor:
        return f"Image (type unclear, best guess {best_label} at {confidence:.0%}). Description withheld to avoid misleading you."
    return f"{TEMPLATES[best_label]} ({confidence:.0%} confident)"

test_images = {
    "unambiguous chart":   [0.87, 45],
    "unambiguous receipt": [0.90, 65],
    "ambiguous image":     [0.5, 25],
}
for name, features in test_images.items():
    print(f"{name:20s} -> {describe(features)}")
Output
unambiguous chart    -> A chart showing data. Exact values are not available from the image alone. (100% confident)
unambiguous receipt  -> A receipt or document with printed text. Use an OCR reader for the exact text. (100% confident)
ambiguous image      -> Image (type unclear, best guess photo at 60%). Description withheld to avoid misleading you.

Walkthrough

KNeighborsClassifier looks at the 5 most similar training examples to a new image and votes among their labels — this naturally produces graded confidence scores, unlike some other classifiers that tend to output near-100% confidence even when genuinely unsure.

describe is the important part. It does not always return a description. When confidence falls below confidence_floor, it explicitly says the type is unclear and withholds the templated description, rather than guessing and presenting the guess as fact. The third test case, deliberately built to sit between two categories, triggers exactly this path.

TEMPLATES are intentionally cautious rather than richly descriptive — "a chart showing data" instead of inventing specific numbers a simple classifier has no way of actually reading. A real system pairs this kind of type detection with actual text or number reading (see OCR) before making any claim about specific content.

Common mistakes

Always returning a description, regardless of confidence. This is the single most damaging habit in this domain — a listener cannot see the image to check the claim, so a wrong confident answer is trusted as fact.

Confusing "the model is confident" with "the model is correct." A model can be consistently, confidently wrong on a whole category it was never trained to recognise. Confidence measures agreement with training data, not truth.

Describing specific numbers or text a classifier never actually read. The chart template deliberately avoids inventing values. A real chart-reading system needs a separate step that actually extracts the data, and should say so honestly when it cannot.

Testing only on images similar to training data. KNeighborsClassifier looks confident on anything resembling its 15 training examples. A genuinely novel image type would need to be added to training data before the model could recognise it at all.

Try it yourself

Add a fourth category, "screenshot", with a handful of training examples of your own choosing for the two features. Add a matching entry to TEMPLATES, and test an image that sits ambiguously between screenshot and chart. Watch whether confidence_floor=0.7 correctly withholds a description for it.

What to learn next

Researcher — Mathematics and papers.

The formal setting

Alt text generation is a specific case of image captioning: learning p(caption | image), typically parameterised by an encoder that maps the image to a representation and a decoder that generates text conditioned on it (see image captioning for the general architecture). What distinguishes alt text from a general caption is the objective it is optimised against: usefulness to a listener with no other access to the image, not fluency or descriptiveness for its own sake.

The abstractness problem

W3C accessibility guidance (Web Content Accessibility Guidelines, WCAG 2.1, Success Criterion 1.1.1) specifies that alt text should serve the same purpose the image serves in context — a decorative image needs no description at all, a functional image (a button) needs to describe the action, and an informative image needs to convey the information, not an exhaustive visual description. This is a context-dependent objective that pure caption fluency metrics (BLEU, CIDEr) do not directly measure, which is why alt text quality evaluation increasingly relies on task-based human evaluation rather than caption-similarity scores alone (MacLeod et al., 2017; Morris et al., 2018 on how blind users actually rate automatically generated Facebook alt text).

Selective prediction and abstention

The developer example's confidence-floor mechanism is a simplified instance of selective prediction (El-Yaniv and Wiener, 2010): a classifier that may abstain from predicting when uncertain, trading coverage for reliability. Formally, given a confidence function κ(x) and threshold τ:

predict(x) =  ŷ(x)     if κ(x) ≥ τ
              abstain   if κ(x) < τ
  • ŷ(x) — the model's top prediction for input x
  • κ(x) — a scalar confidence score, ideally calibrated so κ(x) approximates the true probability of correctness
  • τ — the abstention threshold, chosen to trade off coverage (how often the model answers) against risk (how often it is wrong when it does)

The risk-coverage trade-off this induces is exactly the mechanism used in real captioning systems that report "low confidence" or refuse to caption an image outright — see also choosing a threshold from costs for the general treatment of this trade-off.

Complexity and cost

For an image with h × w pixels and a captioning model with vision encoder cost O(hw) and a decoder generating L tokens autoregressively:

ComponentTypical cost
Vision encoding (CNN or ViT)O(hw) per image, once
Autoregressive caption decodingO(L) sequential steps, each attending over the full encoded image
Confidence estimation (softmax over vocabulary, per token)O(V) per token, V = vocabulary size

Real-time or bulk alt text generation at web scale is dominated by the decoding step, which is why production systems favour smaller, distilled captioning models or non-autoregressive variants over the largest general-purpose vision-language models.

Papers

  • W3C (2018). Web Content Accessibility Guidelines (WCAG) 2.1, Success Criterion 1.1.1: Non-text Content. The governing accessibility standard for alt text.
  • MacLeod, H. et al. (2017). Understanding Blind People's Experiences with Computer-Generated Captions of Social Media Images. CHI. Direct evidence from blind users on what makes automatic alt text useful or harmful.
  • Morris, M. et al. (2018). "With Most of It Being Pictures Now, I Rarely Use It": Understanding Twitter's Evolving Accessibility to Blind Users. CHI.
  • El-Yaniv, R. and Wiener, Y. (2010). On the Foundations of Noise-Free Selective Classification. Journal of Machine Learning Research 11. The formal selective prediction framework.

Current state

Large vision-language models have substantially improved raw caption fluency and detail, but MacLeod et al.'s and Morris et al.'s core finding — that blind users specifically want honest hedging language and object-level detail relevant to their task, not maximal descriptiveness — remains only partially addressed by systems optimised for general caption quality metrics. Evaluating alt text usefulness directly with blind and low-vision users, rather than only with automated caption-similarity scores, is an active and necessary methodological direction, not a solved problem.

What to learn next

What to learn next

These follow on from what you just read.

  • AI for Accessibility

    Live captioning

    Live captioning turns speech into text as someone is still talking, and has to decide how to show a guess on screen that might still change.

  • AI for Accessibility

    Speech recognition for atypical speech

    Atypical speech recognition adapts a general speech model to one person's distinct speech pattern, since a model trained on typical speech alone often fails people with conditions like dysarthria or a stutter.

  • AI for Accessibility

    Sign language recognition

    Sign language recognition reads hand shape, movement, and facial expression together, and stays a research problem because sign languages are full languages with their own grammar, not spelled-out speech.