AI for Accessibility

Live captioning

Live captioning turns speech into text as someone is still talking, and has to decide how to show a guess on screen that might still change.

Read these first

On this page 6
  1. Why it exists
  2. How it works
  3. Where you have already seen it
  4. An honest warning
  5. Remember this
  6. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Live captioning turns speech into text while the person is still talking, not after they finish.

Think about a court stenographer typing every word as a witness speaks, with the words appearing on a screen a heartbeat later. Nobody waits for the witness to finish their whole answer before the stenographer starts typing. Live captioning software does the same job automatically. It turns a lecture, a video call, or a news broadcast into text almost as fast as it is spoken. That helps Deaf and hard-of-hearing viewers, or anyone in a noisy or silent environment.

Why it exists

Ordinary speech recognition can wait for someone to finish a whole sentence, then take its time producing an accurate transcript. Live captioning cannot wait. A Deaf viewer in a video call needs the words on screen within roughly a second of them being spoken. Otherwise the conversation becomes impossible to follow in real time.

That time pressure creates a problem ordinary transcription does not have. A speech recognition system re-guesses its best transcript every time a little more audio arrives. Its early guesses about a word are sometimes wrong, corrected only once more audio provides context. A system that displayed every single guess immediately would show words appearing, then flickering, then changing, over and over, which is disorienting and hard to read.

How it works

Speaker talks  ->  a little audio arrives  ->  best guess so far:  "the museum"
Speaker keeps talking  ->  more audio arrives  ->  revised guess:  "the museum close"
Speaker keeps talking  ->  more audio arrives  ->  revised guess:  "the museum closes at"

A well-built captioning system does not flash every one of these guesses onto the screen as it changes. It waits for a word to be confirmed by more than one guess in a row, before locking it into place. What the viewer reads never gets silently rewritten behind their eyes.

Where you have already seen it

YouTube and video call platforms such as Google Meet and Zoom offer live captions during any video, generated as the speaker talks. Television news broadcasts have used live human-typed or automatic captioning for decades, which is where much of the audience for this technology first encountered it.

An honest warning

Live captioning trades some accuracy for speed, by necessity. A system given more time to think would often produce a more accurate transcript. Live captioning cannot afford that time, so a small amount of error is the accepted cost of keeping captions genuinely live. Any high-stakes use — a legal deposition, a medical consultation — still needs a qualified human captioner or interpreter, not an automatic system alone.

Remember this

  • Live captioning re-guesses the transcript continuously as someone speaks, rather than waiting until they finish.
  • Displaying every raw guess immediately causes distracting flicker, as early guesses get revised.
  • A careful system only locks a word onto the screen once it has been confirmed, trading a small delay for a caption that never rewrites itself.

What to learn next

Developer — Code and libraries.

The core engineering problem here is not speech recognition itself — it is deciding what to put on screen from a stream of changing guesses. This example makes that decision, and the flicker it prevents, directly visible.

Setup

bash
python3 --version

No installation is needed — this example is standard Python, focused on the display logic rather than the speech model itself.

Minimal runnable code

hypotheses simulates what a streaming speech recognition model might output as more audio arrives: a growing, occasionally self-correcting guess at the sentence so far. Word 3 flips from "close" to "closes" partway through, exactly the kind of mid-sentence correction that causes flicker.

livecaption.py
# A streaming ASR model does not wait for the speaker to finish. Every time a
# little more audio arrives, it re-guesses the WHOLE sentence so far, and
# earlier words can change. This list simulates eight such re-guesses.
hypotheses = [
    "the",
    "the museum",
    "the museum close",
    "the museum closes at",
    "the museum closes at five",
    "the museum closes at five thirty",
    "the museum closes at five thirty p m",
    "the museum closes at five thirty p m today",
]

def naive_display(hypotheses):
    # shows whatever the model currently believes -- simple, and it flickers
    return [h for h in hypotheses]

def longest_common_prefix(a, b):
    n = min(len(a), len(b))
    i = 0
    while i < n and a[i] == b[i]:
        i += 1
    return a[:i]

def stable_prefix_display(hypotheses):
    # LocalAgreement policy: only commit words that agree between THIS
    # hypothesis and the PREVIOUS one. Once committed, a word is never
    # taken back or changed on screen again.
    committed = []
    shown_history = []
    previous = []
    for current in hypotheses:
        current_words = current.split()
        if previous:
            agreed = longest_common_prefix(previous, current_words)
            if len(agreed) > len(committed):
                committed = agreed
        previous = current_words
        shown_history.append(list(committed))
    return shown_history

print("naive: latest guess shown every update (word 3 flips from 'close' to 'closes')")
for h in naive_display(hypotheses):
    print(" ", h)

print()
print("stable-prefix: only committed words are shown, never rewritten")
for words in stable_prefix_display(hypotheses):
    print(" ", " ".join(words) if words else "(waiting)")
Output
naive: latest guess shown every update (word 3 flips from 'close' to 'closes')
  the
  the museum
  the museum close
  the museum closes at
  the museum closes at five
  the museum closes at five thirty
  the museum closes at five thirty p m
  the museum closes at five thirty p m today

stable-prefix: only committed words are shown, never rewritten
  (waiting)
  the
  the museum
  the museum
  the museum closes at
  the museum closes at five
  the museum closes at five thirty
  the museum closes at five thirty p m

Walkthrough

naive_display prints exactly what a streaming model currently believes, at every update. Watch the third line: "close." The very next update silently rewrites it to "closes." A viewer reading along would have already read "close," and now has to notice and mentally correct it — this is the flicker problem in miniature.

stable_prefix_display implements a real, widely used technique called LocalAgreement: it only commits a word to the display once it agrees between two consecutive hypotheses in a row. longest_common_prefix finds how much of the start of two word lists matches exactly. Because "close" and "closes" never agree with each other at that position, that word is never displayed at all — the viewer waits one update longer, and sees "closes" directly once it stabilises. Compare the two outputs: the naive version briefly showed a wrong word; the stable version showed nothing at that position rather than something wrong.

Common mistakes

Displaying every raw hypothesis update immediately. This is the naive approach shown here, and it is the most common way a live captioning system feels distracting or untrustworthy to watch, even when its final transcript is accurate.

Committing words too eagerly. Committing after only one hypothesis, with no agreement check at all, brings back the exact flicker problem stable-prefix display exists to prevent.

Committing words too conservatively. Waiting for many consecutive agreements before committing increases accuracy but adds real delay, which defeats the purpose of captions advertised as "live." Real systems tune this trade-off carefully against a target latency, often under one second.

Testing only on hypotheses that never get revised. The interesting behaviour of a stable-prefix policy only shows up when an early guess turns out wrong. Always include a test case with a genuine correction, as this example does.

Try it yourself

Add a second correction to hypotheses — for example, revise "five thirty" to "five forty five" a few updates later, in the same way "close" became "closes." Rerun and check that stable_prefix_display never shows either wrong version on screen, at the cost of a short additional delay before showing the corrected words.

What to learn next

Researcher — Mathematics and papers.

The formal setting

Streaming automatic speech recognition produces, at each time step t, a hypothesis h_t over the transcript of all audio seen so far. Unlike offline ASR, h_t is not guaranteed to be a prefix-consistent extension of h_{t-1} — later context can revise earlier tokens, since the model's belief about early words can change once more acoustic and language context becomes available. The display policy problem is choosing a committed output sequence c_t at each step such that c_t is a prefix of c_{t+1} for all t — once shown, never retracted — while minimising the lag between when a word is spoken and when it is committed.

The developer example implements LocalAgreement-n in its simplest form (n=2): a token is committed once it appears identically at the same position across n consecutive hypotheses. This policy, and variants of it, underlie practical streaming wrappers around offline-trained models such as Whisper-Streaming (Macháček, Dabre and Bojar, 2023), which uses LocalAgreement to adapt an inherently non-streaming model (Whisper was trained on fixed 30-second windows) into a usable live captioning system without retraining.

Formally, with hypotheses h_1, ..., h_t as token sequences:

commit_t = longest prefix p such that p is a prefix of both h_{t-1} and h_t
c_t = commit_t   if |commit_t| > |c_{t-1}|,  otherwise c_t = c_{t-1}
  • |·| — sequence length in tokens
  • c_t — the committed, displayed sequence at time t, monotonically non-decreasing in length by construction

Natively streaming architectures — trained end-to-end to emit tokens incrementally, such as RNN-Transducer models (Graves, 2012) — avoid needing this post-hoc reconciliation step, since their hypothesis at each step is constructed incrementally by design rather than re-decoded from scratch. They trade this structural advantage against typically higher word error rates than the largest offline-trained models, at a given model size.

Cost and latency budget

For a captioning system targeting end-to-end latency L (time from speech to displayed text):

L  =  L_audio_chunk  +  L_inference  +  L_commit_delay
  • L_audio_chunk — how much audio must accumulate before the model processes it (a fixed engineering choice, often 200ms-1s)
  • L_inference — compute time for one forward pass of the model
  • L_commit_delay — additional lag introduced by the display policy, e.g. one extra hypothesis cycle under LocalAgreement-2

Real deployed systems, including Google's live captioning features, report target end-to-end latencies in the few-hundred-millisecond to low-single-second range, with word error rate and latency treated as a tuned trade-off rather than independent targets to separately maximise.

Papers

  • Graves, A. (2012). Sequence Transduction with Recurrent Neural Networks. ICML Workshop on Representation Learning. The RNN-Transducer architecture underlying many natively streaming ASR systems.
  • Macháček, D., Dabre, R. and Bojar, O. (2023). Turning Whisper into Real-Time Transcription System. IJCNLP-AACL. The LocalAgreement-based streaming wrapper referenced above.
  • Radford, A. et al. (2023). Robust Speech Recognition via Large-Scale Weak Supervision. ICML. The original Whisper paper, describing the offline, fixed-window architecture that streaming wrappers adapt.

Current state

Natively streaming architectures and streaming wrappers around large offline models both remain in active production use, chosen based on whether raw accuracy or engineering simplicity matters more for a given deployment. Captioning accuracy for accented, overlapping, or atypical speech remains measurably worse than for the speech types most training data represents — the subject of the next lesson — and is an acknowledged, active limitation rather than a solved problem across every major commercial system.

What to learn next