How it works

How speech recognition works

A microphone turns your voice into thousands of numbers per second, a spectrogram turns those into a picture of sound, and a model reads that picture into words.

On this page 7
  1. The pipeline at a glance
  2. Stage 1 — sound becomes numbers
  3. Stage 2 — numbers become a picture
  4. Stage 3 — the encoder extracts what matters
  5. Stage 4 — the decoder writes the words
  6. Stage 5 — cleanup
  7. Which lessons teach each stage

You say "set a timer for ten minutes" and words appear on the screen. Between your mouth and that text sits a pipeline that turns air pressure into numbers, numbers into a picture, and the picture into words. Here it is, end to end.

The pipeline at a glance

 your voice (vibrating air)
        |
        v
 [1. microphone + sampling]  ->  16,000 numbers per second
        |
        v
 [2. spectrogram]            ->  a picture: which pitches, when
        |
        v
 [3. encoder]                ->  compact features of the sound
        |
        v
 [4. decoder]                ->  "set a timer for ten minutes"
        |
        v
 [5. cleanup]                ->  punctuation, casing, formatting

Stage 1 — sound becomes numbers

Sound is vibrating air. A microphone turns the vibration into an electrical wiggle, and a converter measures that wiggle thousands of times per second. Each measurement is one number. This is sampling, and speech systems typically sample 16,000 times per second.

The result is a waveform: one very long list of numbers. A single second of your voice is already 16,000 numbers, and none of them individually means anything.

Stage 2 — numbers become a picture

Raw waveforms are hard for models to read, so the pipeline converts them into a spectrogram — a chart showing which frequencies are present at each moment. Time runs left to right, pitch runs bottom to top, brightness shows loudness.

Think of sheet music, but written automatically from the sound itself. Every vowel and consonant leaves a distinctive shape in it. An "s" is a bright smear up high. An "o" is stacked bands down low. From here on, recognising speech is close to reading an image.

Stage 3 — the encoder extracts what matters

The spectrogram passes through an encoder — a neural network, in modern systems a transformer — that compresses each slice of time into features that matter for speech. It learns to keep "which sound is being made" and discard background hiss, room echo and the particular pitch of your voice.

This is where accents, noise and speaker differences get absorbed. The encoder saw thousands of hours of varied speech in training, so a Delhi accent and a Dublin accent both map to the same features for the same word.

Stage 4 — the decoder writes the words

A decoder reads those features and writes text one token at a time — a token being a small chunk of text, usually part of a word. At each step it weighs two sources of evidence: what the audio says, and what would make sense next in the language.

The second source does real work. "Recognise speech" and "wreck a nice beach" sound nearly identical. The audio alone cannot settle it; knowing which phrase people actually write can. This built-in feel for likely word sequences is a language model, and it is why transcripts read like sentences rather than sound-alike soup.

Whisper, the best-known open model, is this exact encoder–decoder design trained on 680,000 hours of speech from the internet.

Stage 5 — cleanup

Raw decoder output is lowercase and unpunctuated. A final pass adds capitals, commas and question marks, expands "10" or spells it out per house style, and formats dates and numbers. Voice assistants tap the same stage to mark where a command ends.

The honest caveat: every stage passes its errors downstream. Heavy noise in stage 1 becomes a smeared spectrogram in stage 2, uncertain features in stage 3, and a guessed word in stage 4. The guess arrives with no hint that it was one.

Which lessons teach each stage