Speech and Audio AI

Speech recognition

Speech recognition turns spoken audio into written text, and its central difficulty is that nobody tells the model which sound belongs to which letter.

On this page 10
  1. The short answer
  2. The analogy you have already lived
  3. Why this is harder than it looks
  4. How it works
  5. The alignment problem, and the clever fix
  6. Where you have already used it
  7. How wrong is it allowed to be?
  8. What is honestly hard here
  9. Remember this
  10. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

Speech recognition is turning spoken sound into written words.

The technical name is ASR, short for automatic speech recognition. You may also see it called speech-to-text.

The analogy you have already lived

Think of the first time you heard a language you did not speak. It sounded like one long continuous noise. You could not tell where one word ended and the next began.

Now think of your own language. You hear neat, separate words with tidy gaps between them.

Here is the surprise: those gaps do not exist. Your ear is not finding them; your brain is inserting them, from years of knowing which sounds form words. Record yourself saying "how are you" and look at the waveform — there is no silence between "how" and "are".

Speech recognition has to do that inserting from scratch.

Why this is harder than it looks

Nothing marks a boundary. "Ice cream" and "I scream" produce nearly identical sound. Only meaning separates them.

Everyone sounds different. Age, accent, a cold, speed, how far you sit from the microphone. The word "eight" from a child in Chennai and a man in Glasgow share almost nothing physically.

Sounds bleed into each other. The way you say "t" changes depending on what follows it. Say "top" then "stop" and feel the difference in your mouth.

The world is noisy. Fans, traffic, a television, two people talking at once.

Words that sound the same are spelled differently. "There", "their" and "they're" are one sound and three spellings. No amount of listening resolves that; only context does.

How it works

   audio
     |
     v
  [ chop into 20 ms frames ]         about 50 slices per second
     |
     v
  [ features per frame ]             a fingerprint of each slice
     |
     v
  [ neural network ]                 what sound is in this slice?
     |
     v
  frame guesses:  _ _ h h _ e e l l l _ _ l l o o _ _
     |
     v
  [ collapse repeats, drop gaps ]
     |
     v
  "hello"
     |
     v
  [ language knowledge ]             fix spellings that sound alike
     |
     v
  "Hello"

The step people find surprising is the collapsing. The model does not output words. It outputs a guess for every twentieth of a second, and those guesses get squashed into text afterwards.

The alignment problem, and the clever fix

Here is the problem that held this field back for years.

To train a model you need to tell it the right answer. But if your training clip is one second of audio labelled "hello", which frame is the "h"? Which frames are the "e"? Nobody knows, and labelling that by hand for thousands of hours is impossible.

The fix was to stop trying. Let the model output a letter for every frame, plus one extra symbol meaning "nothing here". Then define a rule for turning any frame sequence into text: squash repeated letters together, then delete the "nothing" symbols.

Now many different frame sequences all produce "hello". Training rewards the model for making all of them likely at once, without ever saying which one is correct. The alignment stops being something you supply and becomes something the model settles on by itself.

That trick has a name, CTC, and you will see it working in the developer section.

Where you have already used it

  • YouTube auto-captions on a video nobody subtitled.
  • Voice typing on your phone's keyboard.
  • "Hey Google, set a timer" — the wake word wakes it, then ASR reads the rest.
  • Automatic subtitles in video calls.
  • Call-centre recordings turned into searchable text.

How wrong is it allowed to be?

Accuracy is measured by word error rate, or WER. Count the words it got wrong. Add the ones it invented and the ones it skipped. Divide by how many words were actually spoken.

Lower is better. Zero is perfect. A rough guide:

  • Under 5 percent — reads well, occasional oddity.
  • 5 to 15 percent — usable, and you notice the errors.
  • Over 25 percent — frustrating to read.

What is honestly hard here

Modern speech recognition is very good on clear, accented-the-way-the-training-data-was speech. It is measurably worse on accents that were rare in its training audio. The same holds for children, elderly speakers, and people with speech differences.

Published audits have repeatedly found large accuracy gaps between speaker groups for commercial systems. This is not a rumour and it is not settled. If you build on ASR, test it on the people who will actually use it, not on a benchmark.

Two other honest limits. Systems degrade sharply when two people speak at once, because most were trained on one voice at a time. And they punctuate by guessing, since speech contains no full stops.

Remember this

  • ASR guesses a sound for every short slice of audio, then squashes those guesses into words.
  • The hard part is alignment: nobody says which slice holds which letter, so the model works it out itself.
  • Accuracy is measured as word error rate, and it varies a lot between different speakers.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install numpy

NumPy only. This section builds the two pieces of ASR you can understand completely: the decoder that turns frame guesses into text, and the metric that says how good it was. Running a full model is the next lesson, Whisper, which needs a download.

Doing the decoder by hand first is worth it. Once you have written twelve lines of CTC collapse, the rest of ASR stops feeling like magic.

What the network actually emits

An acoustic model does not emit words. For every frame — usually 20 or 30 milliseconds — it emits a score for every character in its alphabet, plus one extra symbol.

That extra symbol is the blank. It does not mean silence. It means "no new character here", and it exists for one specific reason you will see in a moment.

ctc_decode.py
import numpy as np

VOCAB = "_abcdefghijklmnopqrstuvwxyz "   # index 0 is the CTC blank
BLANK = 0

# What an acoustic model emits: the character it favours in each 20 ms frame.
# Sounds last several frames, so characters repeat. Blanks mark the gaps.
frame_choices = "__hh_eelll__lloo__" + "__  __" + "ww__oorr__lll_dd__"
frames = np.array([VOCAB.index(c) for c in frame_choices])

def ctc_greedy_decode(idx, blank=BLANK):
    out, previous = [], -1
    for f in idx:
        if f != previous and f != blank:    # a repeat means the sound was held longer
            out.append(VOCAB[f])
        previous = f
    return "".join(out)

print("frames the model favoured :", frame_choices)
print("frame count               :", len(frames), f"({len(frames)*0.02:.2f} s at 20 ms a frame)")
print("naive: drop blanks only   :", repr("".join(VOCAB[f] for f in frames if f != BLANK)))
print("CTC collapse then drop    :", repr(ctc_greedy_decode(frames)))
Output
frames the model favoured : __hh_eelll__lloo____  __ww__oorr__lll_dd__
frame count               : 42 (0.84 s at 20 ms a frame)
naive: drop blanks only   : 'hheellllloo  wwoorrllldd'
CTC collapse then drop    : 'hello world'

Why the blank symbol has to exist

Compare the last two lines of that output. That comparison is the entire idea.

Dropping blanks alone gives hheellllloo, because every sound lasts several frames. So you must collapse repeats. But collapsing repeats naively destroys real double letters — hello would become helo.

Look at the frames in the middle: lll__ll. Three frames of "l", then two blanks, then two more frames of "l". The collapse rule works on runs, and the blank breaks the run into two. The output correctly gets two l's.

That is the whole job of the blank. It is a separator that lets the model say "this is a new letter, not a continuation of the last one". Without it, no CTC system could write "hello", "coffee" or "address".

Note also that a genuine space is an ordinary character in the alphabet, at index 27. It is not the blank. Mixing those two up is a classic bug, and the symptom is output with all the spaces missing.

Measuring how wrong it was

WER is edit distance at the word level. The implementation is dynamic programming, and it is short enough to write from memory.

wer.py
import numpy as np

def wer(reference, hypothesis):
    """Word error rate: edits needed to fix the guess, divided by reference length."""
    r, h = reference.split(), hypothesis.split()
    d = np.zeros((len(r) + 1, len(h) + 1), dtype=int)
    d[:, 0] = np.arange(len(r) + 1)           # cost of deleting every reference word
    d[0, :] = np.arange(len(h) + 1)           # cost of inserting every guessed word
    for i in range(1, len(r) + 1):
        for j in range(1, len(h) + 1):
            cost = 0 if r[i-1] == h[j-1] else 1
            d[i, j] = min(d[i-1, j] + 1,       # deletion
                          d[i, j-1] + 1,       # insertion
                          d[i-1, j-1] + cost)  # substitution, or a free match
    return d[len(r), len(h)] / len(r)

REF = "please book two tickets to mumbai for friday"
tests = [
    ("perfect",        "please book two tickets to mumbai for friday"),
    ("one wrong word", "please book two tickets to mumbai for sunday"),
    ("a dropped word", "please book two tickets mumbai for friday"),
    ("an extra word",  "please book two more tickets to mumbai for friday"),
    ("badly wrong",    "please cook two chickens to bombay for friday"),
]
print(f"reference: {REF}\n")
print(f"{'case':<18}{'WER':>8}   hypothesis")
for label, hyp in tests:
    print(f"{label:<18}{wer(REF, hyp):>8.3f}   {hyp}")
Output
reference: please book two tickets to mumbai for friday

case                   WER   hypothesis
perfect              0.000   please book two tickets to mumbai for friday
one wrong word       0.125   please book two tickets to mumbai for sunday
a dropped word       0.125   please book two tickets mumbai for friday
an extra word        0.125   please book two more tickets to mumbai for friday
badly wrong          0.375   please cook two chickens to bombay for friday

What that table does not tell you

All three single-error rows score 0.125, and they are not equally bad.

Changing "friday" to "sunday" books the wrong day. Adding "more" is harmless. WER weights every word the same, so a wrong date and a stray filler word count identically.

This matters in practice. A voice booking system with 8 percent WER concentrated on dates and numbers is unusable. The same 8 percent spread over "um" and "the" is fine. Always look at which words are wrong, not only how many.

WER can also exceed 1.0, because insertions are unbounded. A model that hallucinates a paragraph over a two-word clip can score 5.0. That is not a bug in the metric, and it does happen — see the note on Whisper and silence in the next lesson.

Line by line, the parts that are not obvious

d[:, 0] = np.arange(...) seeds the table with the cost of turning the reference into an empty string, one deletion per word. The first row is the mirror case. Forget these and every distance comes out too small.

min over three options is the standard edit-distance recurrence. Deletion and insertion cost one; a substitution costs one; matching costs nothing.

/ len(r) normalises by the reference, never by the hypothesis. Dividing by the hypothesis length would reward a system for producing more words.

Text normalisation happens before any of this. Real WER tools lowercase, strip punctuation and expand numbers, because "Rs. 500" against "rupees five hundred" is a formatting difference, not a hearing mistake. Comparing WER figures computed under different normalisation rules is meaningless, and it happens in published comparisons more often than it should.

Common mistakes

Treating the blank as silence. It is not. It appears in the middle of loud speech, whenever the model is between characters. Feeding blanks into a voice-activity detector produces nonsense.

Greedy decoding when the task is hard. Taking the top character per frame ignores that a slightly less likely path may spell a real word. Beam search decoding with a language model typically cuts WER by a fifth or more on noisy audio. Greedy is correct for learning, and rarely correct for shipping.

Comparing WER across different normalisers. Covered above, and worth repeating, because it invalidates most casual benchmark comparisons you will read.

Assuming a benchmark number transfers. A model reporting 3 percent WER on clean read audiobooks may sit at 25 percent on your call-centre recordings. Read speech and spontaneous speech are different problems.

Forgetting the 16 kHz rule. Nearly all ASR models expect mono 16 kHz float audio. See how computers hear sound for why resampling carelessly damages the signal.

Try it yourself

Change one blank in frame_choices — the pair between the two runs of "l" — into an "l" and rerun the decoder. Watch "hello" turn into "helo". Then add a matched counter to wer that also reports substitutions, insertions and deletions separately. That breakdown is what you actually need to debug a real system.

What to learn next

Researcher — Mathematics and papers.

Three eras, briefly

Generative HMM-GMM (1980s to about 2012). Model $P(X \mid W)$ with hidden Markov models over context-dependent phone states, emissions by Gaussian mixtures over MFCC-39, decoded through a weighted finite-state transducer composing the acoustic model, lexicon and n-gram language model. Recognition maximises $P(W \mid X) \propto P(X \mid W) P(W)$.

Hybrid DNN-HMM (2012 onward). Replace GMM emissions with a neural network estimating $P(s \mid x_t)$ and divide by the state prior to get a scaled likelihood. Hinton et al. (2012), Deep Neural Networks for Acoustic Modeling in Speech Recognition, reported 20 to 30 percent relative WER reductions and effectively ended the GMM era.

End-to-end (2014 onward). Map audio directly to characters or subwords, discarding the phone lexicon. Three families dominate, described below.

CTC

Graves et al. (2006) introduced connectionist temporal classification. Given input frames $x_{1:T}$ and target $y_{1:U}$ with $U \le T$, define an extended alphabet $\mathcal{A}' = \mathcal{A} \cup {\varnothing}$ and the collapse map $\mathcal{B}: \mathcal{A}'^{T} \to \mathcal{A}^{\le T}$ that merges repeats then deletes blanks. The likelihood marginalises over every alignment:

$$ P(y \mid x) = \sum_{\pi \in \mathcal{B}^{-1}(y)} \prod_{t=1}^{T} P(\pi_t \mid x) $$

Where $\pi$ is one frame-level alignment, $\pi_t$ its symbol at frame $t$, and $\mathcal{B}^{-1}(y)$ the set of alignments collapsing to $y$. That set is exponentially large, so the sum is computed by a forward-backward recursion over the extended target $y' = (\varnothing, y_1, \varnothing, y_2, \dots, y_U, \varnothing)$ of length $2U+1$:

$$ \alpha_t(s) = P(y'_s \mid x_t) \cdot \begin{cases} \alpha_{t-1}(s) + \alpha_{t-1}(s-1) & \text{if } y'_s = \varnothing \text{ or } y's = y'{s-2}\ \alpha_{t-1}(s) + \alpha_{t-1}(s-1) + \alpha_{t-1}(s-2) & \text{otherwise} \end{cases} $$

The three-term case is the transition that permits skipping a blank between distinct characters; the two-term case forbids skipping when it would merge a genuine double letter. Cost is $O(T \cdot U)$ per utterance. Loss is $-\log \alpha_T(2U+1) + \alpha_T(2U)$, summing the two valid end states.

The conditional independence assumption is CTC's defining weakness: outputs are independent given the encoder states, so the model cannot represent $P(y_u \mid y_{<u})$ internally. It has no implicit language model, which is why external LM fusion helps CTC systems far more than it helps attention systems.

RNN-Transducer

Graves (2012) added a prediction network over previous non-blank outputs and a joint network combining it with the encoder:

$$ P(k \mid t, u) = \operatorname{softmax}\big(W \tanh(W_e h_t^{\text{enc}} + W_p g_u^{\text{pred}} + b)\big) $$

Where $h_t^{\text{enc}}$ is the encoder state at frame $t$ and $g_u^{\text{pred}}$ the prediction state after $u$ emitted symbols. The lattice is two-dimensional, and the forward recursion costs $O(T \cdot U \cdot |\mathcal{A}|)$ in memory, which is the practical obstacle to training it.

RNN-T removes the conditional independence assumption while remaining strictly streaming — it emits without seeing future frames. That combination is why it powers most on-device dictation, including Google's fully on-device recogniser (He et al., 2019).

Attention encoder-decoder

Listen, Attend and Spell (Chan et al., 2015) applies sequence-to-sequence attention directly to speech, modelling $P(y_u \mid y_{<u}, x)$ with full access to the encoder. It has the strongest implicit language model of the three and is not naturally streaming, since attention is over the complete utterance. Whisper is in this family.

The Conformer encoder (Gulati et al., 2020) interleaves self-attention with convolution to capture global and local structure together, and remains the standard encoder across all three loss families.

Hybrid CTC-attention training (Watanabe et al., 2017) attaches a CTC loss to the encoder alongside the attention decoder, typically weighted around $0.3$. The CTC branch enforces monotonic alignment, which markedly stabilises early training and reduces the attention decoder's tendency to loop or truncate.

Benchmarks, with the caveats attached

LibriSpeech (Panayotov et al., 2015) is 960 hours of read audiobooks. Strong supervised systems reach roughly 2 percent WER on test-clean and 4 to 5 percent on test-other.

Treat those numbers with suspicion when planning a product. Read speech from volunteers recording deliberately is close to the easiest possible condition. On conversational telephone speech (Switchboard, Fisher), spontaneous meetings (AMI), or accented and code-switched audio, WER for the same architectures is several times higher. Whisper's contribution was largely robustness across conditions rather than a lower number on test-clean, where it is outperformed by LibriSpeech-specialised models.

Fairness, measured

Koenecke et al. (2020), Racial disparities in automated speech recognition, PNAS 117(14), measured five commercial ASR systems on matched interview audio and reported an average WER of 0.35 for Black speakers against 0.19 for white speakers, with the gap traced to acoustic modelling rather than language modelling. Subsequent work has found comparable disparities across accents, ages and speakers with dysarthria.

Report per-group WER with confidence intervals. Aggregate WER conceals exactly the failures that make a deployed system unusable for some of its users. See fairness metrics for the general framing.

Reading

  • Graves et al. (2006), Connectionist Temporal Classification, ICML — the CTC paper.
  • Graves (2012), Sequence Transduction with Recurrent Neural Networks — arxiv.org/abs/1211.3711
  • Chan et al. (2015), Listen, Attend and Spell — arxiv.org/abs/1508.01211
  • Gulati et al. (2020), Conformer — arxiv.org/abs/2005.08100
  • Hannun (2017), Sequence Modeling with CTC, Distill — distill.pub/2017/ctc — the clearest visual explanation available.
  • Koenecke et al. (2020), Racial disparities in automated speech recognition, PNAS 117(14).

What to learn next

What to learn next

These follow on from what you just read.

  • Speech and Audio AI

    Whisper

    Whisper is OpenAI's open-weight speech recognition model family, strong across languages and noisy audio, and it runs on a laptop CPU if you pick a small enough size.

  • Speech and Audio AI

    Text to speech

    Text to speech turns written words into audio, and the work splits into deciding how the words should sound and then generating the actual waveform.

  • Speech and Audio AI

    Speaker identification

    Speaker identification works out who is talking rather than what they said, by turning each voice into a fingerprint vector and comparing distances between fingerprints.