Speech and Audio AI

Whisper

Whisper is OpenAI's open-weight speech recognition model family, strong across languages and noisy audio, and it runs on a laptop CPU if you pick a small enough size.

Read these first

On this page 10
  1. The short answer
  2. The analogy you have already lived
  3. Before anything else: what it costs you
  4. Why it landed the way it did
  5. What it does beyond plain transcription
  6. How it works
  7. What is honestly hard here
  8. Where you have already seen it
  9. Remember this
  10. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

Whisper is a free, downloadable model that turns recorded speech into text, in about a hundred languages.

OpenAI released it in September 2022, with the weights published openly. You can run it on your own machine, offline, without sending audio to anyone.

The analogy you have already lived

Think of a friend who has spent twenty years travelling. They have heard your language spoken by farmers, professors, children and drunks. They have listened over bad phone lines and in noisy markets.

Ask them to write down what someone said, and they manage. Even through a thick accent. Even with a fan running. Not because they are cleverer than a language teacher, but because they have heard so much more variety.

Whisper is that friend. Its strength is not architecture. It is having listened to an enormous, messy range of audio.

Before anything else: what it costs you

This lesson asks you to download a file. Here is the honest size, before you start, so you can decide on your data plan.

SizeDownloadRuns comfortably on
tinyabout 75 MBany laptop, any phone
baseabout 140 MBany laptop
smallabout 470 MBa laptop, slowly
mediumabout 1.5 GBa good laptop or a GPU
largeabout 3 GBa GPU, realistically

Those figures are for the original released checkpoints. The same model repackaged in another format can be larger or smaller. A compressed version is smaller. A full-precision copy is roughly twice the size. Check before a big download.

You download it once, and it stays on your disk. Start with tiny or base. They are genuinely useful, and nothing in this lesson needs more.

If you are on a metered connection, the first developer example downloads nothing at all.

Why it landed the way it did

Speech systems before Whisper were usually trained on clean, carefully prepared recordings. They scored beautifully on tests made of similar clean recordings. Then they fell apart on a real voice note from a scooter.

The Whisper team took a different bet. Instead of a small, clean dataset, they collected around 680,000 hours of audio from the web. Whatever text already accompanied it became the label. Much of it was imperfect. They kept going anyway, and filtered out the worst.

That is roughly seventy-seven years of continuous audio. The result was a model noticeably harder to break with real-world messiness.

What it does beyond plain transcription

Whisper was trained to do several jobs with the same weights. A marker at the start of its output chooses which.

  • Transcribe — write down what was said, in the language it was said in.
  • Translate — write it down in English, whatever language went in.
  • Detect the language — work out which language it is hearing.
  • Timestamp — say roughly when each chunk was spoken, which is what makes subtitles possible.

How it works

   audio (any length)
        |
        v
  cut into 30-second chunks         always exactly 30 seconds
        |
        v
  turn each chunk into a spectrogram
        |
        v
  [ encoder ]  reads the whole chunk at once
        |
        v
  [ decoder ]  writes the text, one piece at a time,
               looking back at the audio as it goes
        |
        v
  "<|en|><|transcribe|> Hello, how are you?"

The last line is the trick that makes one model do four jobs. Those markers in angle brackets are part of what the model writes. Change which marker you ask for, and the same weights do a different task.

What is honestly hard here

Three real problems, none of them secret.

It invents text during silence. Give Whisper a recording with a long quiet gap and it will sometimes write a plausible sentence that nobody said. Often it produces something like a subtitle-credit line, because its training audio contained many of those. This failure is well documented and it has caused real trouble where transcripts were used without review.

Not all hundred languages work equally well. The training audio was dominated by English. Languages with less audio on the web get noticeably worse accuracy. Its own list of supported languages includes some it handles poorly.

Timestamps are approximate. They are good enough for subtitles and not good enough for precise editing.

Read that first point twice. One tool occasionally mishears a word. Another occasionally writes fluent sentences nobody spoke. Those are different kinds of tool.

Where you have already seen it

  • Subtitle tools in video editors that caption a clip in one click.
  • Meeting-notes apps that produce a transcript afterwards.
  • Podcast search that lets you find a phrase inside an audio episode.
  • Offline dictation apps on phones that do not send your voice anywhere.

Remember this

  • Whisper is open-weight: you download it and run it yourself, offline.
  • Its strength comes from the sheer variety of its training audio, not from a new architecture.
  • It can hallucinate text during silence, so never use a transcript unchecked where being wrong matters.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install numpy librosa faster-whisper

faster-whisper reimplements Whisper on CTranslate2 and is several times quicker than the reference implementation on CPU, with lower memory use. The reference openai-whisper package works identically for these examples if you prefer it.

Downloads, stated plainly. The libraries above are tens of megabytes. Model weights are separate and are fetched on first use: roughly 75 MB for tiny, 140 MB for base, 470 MB for small, 1.5 GB for medium and 3 GB for large-v3. faster-whisper fetches int8-quantised conversions, which are smaller again for the same name.

The first example below downloads nothing. Run it even if you are not ready to fetch weights.

What Whisper actually eats

Before touching the model, build its input by hand. Whisper's front-end is fixed and public, and reproducing it removes most of the mystery.

whisper_frontend.py
import numpy as np, librosa

SR, N_FFT, HOP, N_MELS, WINDOW_S = 16000, 400, 160, 80, 30

t = np.arange(int(SR * 1.4)) / SR
clip = 0.3 * np.sin(2*np.pi*220*t) * np.hanning(t.size)     # 1.4 s of stand-in "speech"

print("clip                 :", clip.size, "samples =", clip.size / SR, "s")

# Whisper always works on exactly 30 seconds. Shorter audio is padded with silence.
padded = np.pad(clip, (0, SR * WINDOW_S - clip.size))
print("padded to 30 s       :", padded.size, "samples")

mel = librosa.feature.melspectrogram(y=padded, sr=SR, n_fft=N_FFT,
                                     hop_length=HOP, n_mels=N_MELS)
log_mel = np.log10(np.maximum(mel, 1e-10))[:, :-1]   # drop the extra centring frame
print("log-mel input        :", log_mel.shape, " (80 mel bins, 3000 frames)")
print("frames per second    :", SR / HOP)
print("encoder positions    :", log_mel.shape[1] // 2, "after the stride-2 conv")
print("this clip is", round(100 * (1 - clip.size / padded.size), 1),
      "% padding -- and it costs the same as 30 s of real speech")
Output
clip                 : 22400 samples = 1.4 s
padded to 30 s       : 480000 samples
log-mel input        : (80, 3000)  (80 mel bins, 3000 frames)
frames per second    : 100.0
encoder positions    : 1500 after the stride-2 conv
this clip is 95.3 % padding -- and it costs the same as 30 s of real speech

Every number in that output is a design decision

16000 Hz, mono. Not negotiable. Feed 44.1 kHz and you get either an error or quiet nonsense. Resample properly — see the aliasing warning in how computers hear sound.

400-sample window, 160-sample hop. That is 25 ms windows every 10 ms, giving exactly 100 frames per second, the standard speech convention from audio features.

80 mel bins. Fixed for every model up to large-v2. Note that large-v3 changed to 128 bins, so front-end code written for earlier models does not transfer to it unchanged.

Always 30 seconds. This is the single most important line for your bill. A 1.4-second clip is padded to 30 seconds and the encoder processes all 3000 frames regardless. Transcribing one short clip costs the same as transcribing thirty seconds. If you are processing many short utterances, batch them into 30-second groups yourself, or you are paying twenty times over.

[:, :-1] trims one frame. Centred STFT produces 3001 frames from 480000 samples, and Whisper drops the last one to land on 3000. Small detail, and it will cause a shape mismatch if you feed a hand-built spectrogram to the model.

Running the real thing

This block downloads weights on first run. It is written for the smallest useful model.

transcribe.py
import numpy as np, soundfile as sf
from faster_whisper import WhisperModel

# int8 on CPU: the smallest, fastest configuration that still works well.
model = WhisperModel("tiny", device="cpu", compute_type="int8")

segments, info = model.transcribe(
    "your_audio.wav",
    beam_size=5,
    vad_filter=True,          # drop silence before it reaches the model
)

print(f"detected language: {info.language}  (confidence {info.language_probability:.2f})")
for s in segments:
    print(f"[{s.start:6.2f} -> {s.end:6.2f}]  {s.text.strip()}")

No output block here, deliberately. The text depends on your recording, the model size, the release version and the beam settings. Pasting a specific transcript would teach you to expect something that will not happen on your machine.

To try it without a recording of your own, record five seconds on your phone, send it to yourself, and convert it with ffmpeg -i note.opus -ar 16000 -ac 1 your_audio.wav.

The silence hallucination, and how to blunt it

Feed Whisper thirty seconds of near-silence or non-speech noise and it may emit a confident, fluent sentence that nobody said. Reported examples cluster around subtitle-credit phrases, because its training data was full of them.

Three defences, in order of how much they help:

Voice activity detection first. vad_filter=True in the call above runs a small detector and removes non-speech regions before Whisper sees them. This is the single most effective fix, because the model cannot hallucinate over audio it never receives.

Threshold tuning. no_speech_threshold and log_prob_threshold let the decoder discard low-confidence segments. hallucination_silence_threshold targets this failure directly.

Turn off cross-chunk conditioning. condition_on_previous_text=False stops a repetition loop in one 30-second chunk from seeding the next. It costs a little context-driven accuracy and it prevents the worst runaway failures.

None of these eliminate the problem. Where a wrong transcript causes real harm — medical notes, legal records, anything a decision rests on — a human reads the output. That is not caution for its own sake; it follows from what the model does.

Choosing a size honestly

Bigger is more accurate and much slower. On CPU, roughly:

  • tiny and base transcribe faster than real time on a modern laptop. Good for drafts, search indexes, and anything where you will read the result anyway.
  • small is around real time on CPU. Usually the best accuracy-per-rupee point for batch work.
  • medium and large are painful on CPU and want a GPU.

Those are ranges, not promises. Your thread count, whether you use int8, and clip length all move them substantially. Measure on your own machine with your own audio before planning around any of it.

Also look at distil-whisper, a distilled family that is substantially faster and smaller than large-v2 with a small accuracy cost, and whisper turbo, a pruned-decoder variant of large-v3 that is far quicker at transcription. Turbo was trained for transcription rather than translation, so do not reach for it when you need the translate task.

Common mistakes

Sending short clips one at a time. Demonstrated above: every call costs a full 30 seconds of encoder work. Batch short utterances.

Trusting the language detector on the first chunk. It reads only the opening 30 seconds. A video that starts with music gets misdetected. Pass language="hi" when you know the answer.

Using word timestamps for precise editing. word_timestamps=True estimates alignment from attention weights. It is good enough for karaoke-style captions and not for frame-accurate cuts.

Assuming the same code path for large-v3. It uses 128 mel bins rather than 80. Hand-built front-ends silently break.

Fine-tuning before trying prompts. initial_prompt biases spelling toward domain vocabulary — product names, drug names, place names — at zero training cost. Try that before reaching for fine-tuning.

Try it yourself

Run the front-end script with WINDOW_S = 30 on a 25-second clip and a 35-second clip, and work out how many chunks each becomes. Then transcribe one recording with vad_filter=True and again with False, and compare the outputs over any quiet stretch. That difference is the hallucination behaviour, visible on your own audio.

What to learn next

Researcher — Mathematics and papers.

The model

Radford et al. (2022), Robust Speech Recognition via Large-Scale Weak Supervision, arXiv:2212.04356.

A plain encoder-decoder transformer (Vaswani et al., 2017) over log-mel input. The audio encoder applies two 1-D convolutions — the second with stride 2 — to the $80 \times 3000$ log-mel, halving the time axis to 1500 positions, adds sinusoidal position embeddings, then runs standard pre-activation transformer blocks. The text decoder uses learned position embeddings and cross-attends to the encoder output. Byte-level BPE tokeniser, GPT-2's vocabulary for the English-only models and a multilingual extension otherwise.

ModelLayers (enc/dec)WidthHeadsParameters
tiny4 / 4384639 M
base6 / 6512874 M
small12 / 1276812244 M
medium24 / 24102416769 M
large32 / 321280201550 M

The architectural claim of the paper is deliberately modest. Nothing here is novel. The contribution is the data and the multitask output format.

The multitask format

Every task is expressed as a token sequence the decoder produces:

<|startoftranscript|> <|language|> <|task|> <|timestamp|> text... <|endoftext|>

Where <|task|> is <|transcribe|> or <|translate|>, and <|notimestamps|> replaces the timestamp token when they are not wanted. Timestamps are quantised to 20 ms and emitted as vocabulary tokens interleaved with text.

This is the paper's real design idea. Language identification, voice activity detection, transcription, translation and segmentation become one autoregressive decoding problem over a shared vocabulary, rather than a pipeline of separate models. It also explains the hallucination behaviour: the decoder is a language model, and a language model handed uninformative encoder states falls back on its prior.

The data argument

680,000 hours of paired audio and transcript scraped from the web, of which 117,000 hours are non-English and 125,000 hours are X-to-English translation. No human transcription was commissioned. Automatic filters removed machine-generated transcripts — detected via the absence of expected punctuation and casing patterns, since ASR output looks distinctive — and an audio language detector removed mismatched pairs.

The central empirical result is about robustness under distribution shift, not about any single benchmark. Whisper is outperformed on LibriSpeech test-clean by supervised models trained on LibriSpeech. Averaged across out-of-distribution sets, it makes roughly 55 percent fewer errors than a wav2vec 2.0 model matched on test-clean performance. The paper's argument is that in-distribution benchmark numbers had stopped predicting real-world behaviour, and that scaled weak supervision buys generalisation that clean supervised training does not.

Zero-shot evaluation is what makes this comparison meaningful: Whisper never saw the training split of any evaluation set.

Scaling and the multilingual picture

Performance improves roughly log-linearly with model size across languages, but the intercept varies enormously by language, tracking hours of training audio for that language. English holds around 438,000 hours; many supported languages have a few hundred. Reported WER on low-resource languages is correspondingly poor, and the published list of supported languages includes several the paper itself notes perform badly.

Radford et al. also report that on English, further scaling shows diminishing returns and the model approaches human-level agreement rates on some sets — while remaining far from human on others. Read the per-language tables before promising anything about a specific language.

Known failure modes, with mechanisms

Hallucination on non-speech. The decoder's language-model prior dominates when encoder states carry no linguistic content. Koenecke et al. (2024), Careless Whisper: Speech-to-Text Hallucination Harms, ACM FAccT, measured this systematically on aphasic speech recordings, finding hallucinations in roughly 1 percent of transcriptions, with a substantial fraction containing invented violent content, fabricated personal information or spurious authority claims. The harms are concentrated on speakers with longer speech disfluencies, which is precisely the population least able to absorb the error.

Repetition loops. Autoregressive decoding without a repetition penalty can enter a fixed point. Mitigated by temperature fallback on high compression ratio, which the reference implementation ships by default.

Cross-chunk error propagation. condition_on_previous_text feeds the previous chunk's output as decoder context. It improves coherence and propagates failures.

Fixed 30-second windows. Long-form audio is handled by sequential chunking with timestamp-based boundary selection, which is heuristic and a known source of drift and dropped segments.

Derived and competing systems

Distil-Whisper (Gandhi, von Platen and Rush, 2023) applies pseudo-labelling and layer-dropping to reduce the decoder to two layers, reporting roughly 6x speedup and 49 percent fewer parameters, within 1 percent WER on out-of-distribution audio.

Whisper large-v3-turbo prunes the decoder to 4 layers, trading translation quality for large transcription speedups.

WhisperX (Bain et al., 2023) adds forced alignment with a phoneme model and VAD-based chunking, giving genuinely word-level timestamps that Whisper's attention heuristic does not provide.

Competing open models worth benchmarking against: NVIDIA's Parakeet and Canary families, Meta's MMS and Seamless models for wide language coverage, and wav2vec 2.0 or HuBERT fine-tunes when you have in-domain labelled data. On a narrow domain with a few hundred labelled hours, a fine-tuned smaller model routinely beats zero-shot Whisper.

Reading

What to learn next

What to learn next

These follow on from what you just read.

  • Speech and Audio AI

    Text to speech

    Text to speech turns written words into audio, and the work splits into deciding how the words should sound and then generating the actual waveform.

  • Speech and Audio AI

    Speaker identification

    Speaker identification works out who is talking rather than what they said, by turning each voice into a fingerprint vector and comparing distances between fingerprints.

  • Speech and Audio AI

    Audio classification

    Audio classification puts a label on a sound, and the hard part is not building the model but making it survive noise it never met during training.