AI for Accessibility

Predictive text for AAC devices

Predictive text for AAC devices turns every avoided keystroke into real time saved for someone who cannot speak, which changes what counts as a good prediction.

Read these first

On this page 6
  1. Why it exists
  2. How it works
  3. Where you have already seen it
  4. An honest warning
  5. Remember this
  6. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Predictive text for AAC devices suggests the next word for someone building sentences one slow selection at a time.

Think about texting an entire message using only two buttons on an old television remote. One button moves a highlighted choice forward, the other selects it. Every single letter would take several presses.

That is close to daily reality for many people who use AAC — augmentative and alternative communication — devices. These devices are controlled by eye gaze, a single switch, or a few slow, deliberate movements, because speaking is difficult or not possible for them. Predictive text exists to cut down how many of those slow selections a full sentence actually costs.

Why it exists

Ordinary phone keyboards predict the next word to save a little time for someone who can already type quickly. AAC users are in a different situation. For someone selecting words by eye gaze or a single switch, every keystroke can cost real seconds, not fractions of one. A sentence that takes a typical typist five seconds might take an AAC user several minutes, letter by letter.

That difference changes what "good prediction" means. A phone keyboard's next-word suggestion is a small convenience. An AAC device's next-word suggestion is different. It can mean the difference between finishing a thought at a normal conversational pace, or falling behind it while everyone else waits.

How it works

Person selects: "I"
        |
        v
Predicted next words, ranked: "want", "need", "am"
        |
        v
Person selects "want" with ONE tap, instead of spelling w-a-n-t letter by letter
        |
        v
Predicted next words, ranked: "more", "to", "help"
        |
        v
... and so on, one word at a time, mostly by selecting rather than spelling

Where you have already seen it

AAC apps such as Proloquo2Go and TouchChat include word prediction built directly into their communication boards, alongside picture-based vocabulary. Many of these systems also learn a person's own frequently used words over time, adapting the suggestions to that individual's own way of talking.

An honest warning

A predicted word that is wrong costs an AAC user real time to notice, undo, and correct. That time is exactly the scarce resource this whole feature exists to save.

Predictions also need to be stable and placed consistently. Re-learning where a word tends to appear on the screen is itself a skill many long-time AAC users develop. A system that reorders suggestions unpredictably can slow someone down, even while trying to help them.

Remember this

  • AAC predictive text exists because every keystroke costs real seconds for someone selecting words by eye gaze or a switch. Ordinary typing only costs fractions of a second.
  • The value of a correct prediction is measured in real time saved during an actual conversation, not typing convenience.
  • A wrong prediction, or an unpredictable change to where suggestions appear, costs the user real time and effort to work around.

What to learn next

  • Plain-language rewriting — another accessibility task built on language modelling, aimed at readability instead of typing speed.
  • What is NLP? — the general field this lesson's word-prediction technique belongs to.
  • Tokenization — how text gets broken into the units a prediction model actually works with.

Developer — Code and libraries.

This example builds a small, real word predictor from scratch, and measures the exact thing that matters in this domain: how many keystrokes it actually saves.

Setup

bash
python3 --version

No installation is needed — this example uses the Python standard library only.

Minimal runnable code

A tiny corpus of short, everyday sentences trains a bigram model — predicting the next word from only the word immediately before it. A real system trains on far more text, often personalised to one person's own frequently used words.

aac.py
from collections import Counter, defaultdict

# A tiny corpus of short, everyday AAC-style sentences. A real system trains
# on far more text, often personalised to one person's own core vocabulary --
# the words used in almost every sentence they build (I, want, more, help, go).
CORPUS = [
    "i want more food",
    "i want more water",
    "i want to go home",
    "i want help please",
    "can you help me please",
    "i need help now",
    "i want more time",
    "can i go home now",
]

bigram_counts = defaultdict(Counter)
for sentence in CORPUS:
    words = sentence.split()
    for a, b in zip(words, words[1:]):
        bigram_counts[a][b] += 1

def suggest_next(previous_word, top_k=3):
    counts = bigram_counts.get(previous_word)
    if not counts:
        return []
    return [word for word, _ in counts.most_common(top_k)]

def build_sentence(words, top_k=3):
    # count characters saved: a suggestion in the top-k costs ONE selection,
    # not one keystroke per letter of the word
    keystrokes_letter_by_letter = sum(len(w) + 1 for w in words)  # +1 for the space
    keystrokes_with_prediction = 0
    for i, word in enumerate(words):
        if i == 0:
            keystrokes_with_prediction += len(word) + 1
            continue
        suggestions = suggest_next(words[i - 1], top_k)
        if word in suggestions:
            keystrokes_with_prediction += 1   # one tap selects the whole word
        else:
            keystrokes_with_prediction += len(word) + 1
    return keystrokes_letter_by_letter, keystrokes_with_prediction

sentence = "i want more food".split()
print("next-word suggestions after 'want':", suggest_next("want"))
print("next-word suggestions after 'more':", suggest_next("more"))

letters, predicted = build_sentence(sentence)
saved = (1 - predicted / letters) * 100
print(f"\ntyping '{' '.join(sentence)}':")
print(f"  letter by letter: {letters} keystrokes")
print(f"  with prediction:  {predicted} keystrokes")
print(f"  keystrokes saved: {saved:.0f}%")
Output
next-word suggestions after 'want': ['more', 'to', 'help']
next-word suggestions after 'more': ['food', 'water', 'time']

typing 'i want more food':
  letter by letter: 17 keystrokes
  with prediction:  5 keystrokes
  keystrokes saved: 71%

Walkthrough

bigram_counts records, for every word that appeared in the corpus, which word followed it and how often. suggest_next("want") looks up everything that followed "want" across the whole corpus and returns the three most common — "more," "to," "help" — in the exact order they should be offered to the user, most likely first.

build_sentence is where the domain-specific idea lives: it compares two ways of producing the same sentence. Typed letter by letter, "i want more food" costs 17 keystrokes, counting a space after each word. Using prediction, the first word must still be typed in full, but "want" and "more" and "food" all happen to appear as top-3 suggestions given the word before them, so each costs a single selection instead of several letters. The result: 5 keystrokes instead of 17, a 71% reduction — a concrete, measurable version of the time savings this whole feature exists to deliver.

Common mistakes

Measuring a predictor by next-word accuracy alone. A model can have excellent accuracy while still costing a user a large number of keystrokes if its correct predictions rarely land within a short, quickly selectable top-k list. Keystroke savings, as computed here, is the metric that reflects real-world benefit.

Training only on generic text, never on the individual's own words. "Help," "please," and specific names matter far more to one particular AAC user's real conversations than to a generic corpus. Real systems adapt to an individual's own vocabulary over time, similar in spirit to the personalisation covered in speech recognition for atypical speech.

Reordering the suggestion list unpredictably between sessions. A user who has learned that "more" tends to appear as the first or second suggestion after "want" loses that muscle memory if the ranking shifts around for no clear reason.

Assuming a bigger, more accurate language model always helps in this setting. A model that considers far more context might predict better in the abstract, while being too slow to run instantly on a low-power device, or offering suggestions the user finds harder to predict and rely on — accuracy is not the only cost that matters here.

Try it yourself

Add "i need help please" to CORPUS, then rerun and check the new suggestions after "need." Try building the sentence "i need help please" using build_sentence and see how the keystroke savings change now that the corpus has seen a closer match to it.

What to learn next

  • What is NLP? — the general field of language modelling this lesson's bigram predictor is a small instance of.
  • Tokenization — the general technique of breaking text into predictable units, which real AAC systems build on for word and phrase prediction.
  • Speech recognition for atypical speech — another accessibility technology built around personalising a general model to one individual.

Researcher — Mathematics and papers.

The formal setting

An AAC word predictor is a language model p(w_t | w_1, ..., w_{t-1}), most directly approximated by an n-gram model conditioning on only the previous n-1 words. The developer example uses n=2 (a bigram model):

p(w_t | w_{t-1})  ≈  count(w_{t-1}, w_t) / count(w_{t-1})
  • count(w_{t-1}, w_t) — how often word w_t immediately followed w_{t-1} in the training corpus
  • count(w_{t-1}) — how often w_{t-1} appeared at all

The right objective is not perplexity

Standard language modelling is evaluated by perplexity, a measure of how well the model predicts the true next word across a whole corpus. AAC prediction is more usefully evaluated by keystroke savings (KSR), the exact quantity the developer example computes directly:

KSR = 1 - (keystrokes with prediction) / (keystrokes without prediction)

This metric, standard in the AAC research literature (Trnka and McCoy, 2007; Wandmacher and Antoine, 2007, among others establishing it as the field's evaluation norm), directly captures what matters operationally: real selections saved, not next-word ranking quality in the abstract. A model can have lower perplexity than another while achieving worse keystroke savings, if its correct predictions do not reliably land within the small top_k list a user can realistically scan and select from in a low-bandwidth interface — the interface constraint changes which model is actually better for this task.

Motor planning stability

A documented, AAC-specific finding is that experienced users of icon- or grid-based AAC systems develop motor planning: reaching for a word's position becomes a learned physical habit, similar to touch-typing. Light and Drager (2007) and subsequent AAC usability research document that unpredictable rearrangement of prediction lists between uses can measurably slow down or frustrate experienced users, even when the underlying prediction accuracy improves — a genuine tension between model quality in the abstract and usability in the specific, individual, motor-learned context AAC devices are used in.

Complexity and cost

For a vocabulary of size V and an n-gram model:

ComponentCost
Training (counting)O(corpus length), a single pass
StorageO(V^n) in the worst case, mitigated heavily by the fact almost all n-grams never occur
Prediction at inferenceO(1) lookup, or O(log V) with a sorted structure — critical, since AAC devices are often low-power and predictions must feel instantaneous

n-gram models remain genuinely competitive in this specific domain precisely because of this last row: their inference cost is negligible on low-power hardware, which matters more here than the modest accuracy gains a much larger neural language model might offer at meaningfully higher latency and power cost.

Papers

  • Trnka, K. and McCoy, K. (2007). Corpus Studies in Word Prediction. ASSETS. Establishes keystroke-savings-oriented evaluation for AAC prediction.
  • Wandmacher, T. and Antoine, J. (2007). Methods to Integrate a Language Model with Semantic Information for a Word Prediction Component. EMNLP-CoNLL.
  • Light, J. and Drager, K. (2007). AAC Technologies for Young Children with Complex Communication Needs: State of the Science and Future Research Directions. Augmentative and Alternative Communication 23(3). Discusses motor planning and interface stability in AAC design.
  • Higginbotham, D. and Caves, K. (2002). AAC Performance and Usage: Ecological Momentary Assessment. Discusses real conversational-rate measurement for AAC users.

Current state

Modern AAC prediction systems increasingly incorporate neural language models and personalisation, while the field's usability research consistently reinforces that raw model accuracy is not the primary lever for real-world benefit — interface design, prediction stability, and adapting to an individual's specific, evolving vocabulary matter as much or more. Deploying a new prediction model for an existing AAC user is itself a design decision with real cost, since retraining someone's motor habits around a changed suggestion layout is not free, however much the underlying model has improved.

What to learn next

  • What is NLP? — the general field this lesson's n-gram model is a minimal instance of.
  • Word2Vec — a step beyond bag-of-counts n-grams toward representing word meaning, relevant to smarter AAC vocabulary suggestions.
  • Evaluating assistive AI with real users — why keystroke savings and usability, not perplexity, is the correct evaluation lens here.