Text to speech
Text to speech turns written words into audio, and the work splits into deciding how the words should sound and then generating the actual waveform.
- 18 min read
- 3 reading levels
- Updated
Read these first
On this page 10
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
Text to speech takes written words and produces audio of someone saying them.
The short name is TTS. It is speech recognition run backwards, and it is a harder problem than it sounds.
The analogy you have already lived
Think of being asked to read a paragraph aloud in class. You did two separate jobs, and you probably never noticed.
First you worked out how it should sound. Where to pause. Which word to stress. Whether the last sentence was a question, so your voice should rise. You did this before opening your mouth.
Then you actually made the sound, using your lungs, vocal cords, tongue and lips.
Machines split the work the same way. One part decides how it should sound. Another part produces the audio. Almost every TTS system you meet has these two halves.
Why the first half is harder than the second
Making a voice sound realistic is largely solved. Deciding how to read something is not.
Take these:
- "I never said she stole my money." Seven words. Stress a different one each time and you get seven different meanings. Nothing in the writing says which.
- "I read the book." Is that "reed" or "red"? Same letters, two words, and only the tense decides.
- "Dr. Rao lives on Oxford St." The first is "Doctor". The second is "Street". Identical abbreviation.
- "Turn left in 100 m." Is that "one hundred metres", "one zero zero", or "a hundred"?
A human resolves these without effort by understanding the sentence. A machine has to be taught. Mistakes here make a synthetic voice sound wrong even when the audio is flawless.
Why it exists
Screen readers gave blind and low-vision people access to computers. That was the original reason, and it remains the most important one.
Then it spread. Announcements at railway stations. Navigation directions. Audiobooks for text nobody recorded. Reading a message aloud while you drive. Voice assistants answering out loud.
And a use that matters enormously to the people it serves: voice banking. Someone losing their speech to illness records their voice while they still can. A synthesiser can then speak in their own voice.
How it works
"The train is arriving."
|
v
[ normalise ] expand numbers, dates, abbreviations
|
v
[ pronounce ] letters -> sounds (phonemes)
|
v
[ plan the delivery ] how long each sound lasts, where the pitch rises
|
v
[ acoustic model ] produce a spectrogram of the intended speech
|
v
[ vocoder ] turn that picture back into a wave
|
v
audio you can playThe last box has a name worth remembering. A vocoder is the part that converts a picture of sound into actual sound. It exists because drawing a spectrogram is far easier for the acoustic model. Producing sixteen thousand samples a second directly is much harder.
How voices got good
Concatenative (1990s to 2010s). Record one person for many hours, chop it into tiny pieces, and glue pieces together for new sentences. It sounded like a real human, because it was, and it lurched at every join. Old railway announcements are this.
Parametric. Model the voice with a small set of controls. Smooth, flexible, and unmistakably robotic.
Neural (2016 onward). Learn the whole mapping from data. WaveNet in 2016 was the break: it modelled the waveform one sample at a time and closed most of the gap to human speech. It was also so slow that generating one second took minutes. Everything since has been about keeping that quality while running fast enough to be useful.
Where you have already heard it
- Google Maps reading out your turns.
- Screen readers on phones and computers.
- Railway and airport announcements, increasingly synthetic rather than recorded.
- Audiobooks of books nobody paid a narrator to read.
- Video narration on social media, in that recognisable app voice.
What is honestly hard here
Modern TTS sounds excellent for a plain sentence read neutrally. It is still weak at:
- Long-form emotion. A whole chapter read with genuine feeling, varying naturally, remains hard.
- Names and rare words. Your surname, your village, a medicine name. Systems guess from spelling and often guess wrong.
- Code-mixed text. An English sentence with Hindi words in it, which is how a great many people actually write.
- Knowing what it does not know. A TTS system mispronounces confidently. It never hesitates.
And one more thing, which gets its own lesson: a system that can copy a voice from a short sample can copy your voice. See voice cloning and its risks.
Remember this
- TTS splits into deciding how it should sound and producing the audio.
- The vocoder turns a spectrogram back into a playable wave.
- Deciding pronunciation and emphasis from plain text is the unsolved half, not the audio.
What to learn next
- Voice cloning and its risks — what happens when a system needs only seconds of your voice.
- Speaker identification — how a machine tells one voice from another.
- Real-time audio pipelines — fitting synthesis into a live conversation.
Developer — Code and libraries.
Setup
pip install numpy scipyEverything verified in this section runs on NumPy and SciPy, on a CPU, with no model download. You will build a working speech synthesiser in about fifteen lines and write a WAV file you can play.
Real neural TTS costs a download, and honest sizes are listed near the end before you are asked to fetch anything.
Stage one: normalising the text
Before any audio exists, the text has to be turned into what a person would actually say. This stage is unglamorous and it is where most audible errors are born.
import re
ONES = ["zero","one","two","three","four","five","six","seven","eight","nine","ten",
"eleven","twelve","thirteen","fourteen","fifteen","sixteen","seventeen",
"eighteen","nineteen"]
TENS = ["","","twenty","thirty","forty","fifty","sixty","seventy","eighty","ninety"]
def say_number(n):
if n < 20: return ONES[n]
if n < 100: return TENS[n//10] + ("" if n%10 == 0 else " " + ONES[n%10])
if n < 1000: return ONES[n//100] + " hundred" + ("" if n%100 == 0 else " and " + say_number(n%100))
return say_number(n//1000) + " thousand" + ("" if n%1000 == 0 else " " + say_number(n%1000))
ABBREV = {"dr": "doctor", "mr": "mister", "st": "saint", "km": "kilometres"}
def normalise(text):
text = re.sub(r"\bRs\.?\s*(\d+)", # money, before anything else
lambda m: say_number(int(m.group(1))) + " rupees", text)
text = re.sub(r"\b([A-Za-z]{2})\.", # then two-letter abbreviations
lambda m: ABBREV.get(m.group(1).lower(), m.group(1)), text)
text = re.sub(r"\b(\d+)\b", lambda m: say_number(int(m.group(1))), text)
return re.sub(r"\s+", " ", text).strip()
tests = [
"The fare is Rs. 250.",
"Dr. Rao lives 3 km away.",
"Meet me at St. Mary's church.",
"Turn left onto Oxford St. and stop.",
"I read the book last night.",
]
for line in tests:
print(f"in : {line}")
print(f"out: {normalise(line)}\n")in : The fare is Rs. 250. out: The fare is two hundred and fifty rupees. in : Dr. Rao lives 3 km away. out: doctor Rao lives three km away. in : Meet me at St. Mary's church. out: Meet me at saint Mary's church. in : Turn left onto Oxford St. and stop. out: Turn left onto Oxford saint and stop. in : I read the book last night. out: I read the book last night.
Three of those five are wrong, and that is the point of showing them.
Line 2: 3 km became "three km". The abbreviation rule requires a trailing full stop, and "km" has none. A rule gap, fixable with more rules.
Line 4: Oxford St. became "Oxford saint". It should be "Street". The identical token means opposite things depending on whether a place name precedes it. No regex resolves that; it needs to know that Oxford is a road here and a person elsewhere.
Line 5: read passed through untouched, which looks fine and is not. A synthesiser must choose "reed" or "red", and "last night" is the only clue. This is a homograph: same spelling, different pronunciation, disambiguated by grammar and meaning alone.
This is why modern systems use a learned front-end rather than a rule table. Rules handle the common 95 percent and the remaining 5 percent is where listeners notice.
Stage two: making sound from nothing
Now the audio. Speech is well modelled as a source and a filter: vocal cords buzzing, and a mouth shaping the buzz. Both are short.
import numpy as np, wave, os
from scipy.signal import lfilter
SR = 16000
VOWELS = { # the first three mouth resonances, in Hz
"a": (730, 1090, 2440), # as in "father"
"i": (270, 2290, 3010), # as in "see"
"u": (300, 870, 2240), # as in "boot"
}
def glottal_pulses(n, f0):
"""The source: vocal cords snapping shut f0 times a second."""
src = np.zeros(n)
src[::int(SR / f0)] = 1.0
return src
def resonator(x, fc, bw=90.0):
"""The filter: a two-pole resonance that rings at fc, one mouth cavity."""
r = np.exp(-np.pi * bw / SR)
th = 2 * np.pi * fc / SR
return lfilter([1 - 2*r*np.cos(th) + r*r], [1.0, -2*r*np.cos(th), r*r], x)
def say(vowel, seconds=0.35, f0=120):
n = int(SR * seconds)
x = glottal_pulses(n, f0)
for fc in VOWELS[vowel]: # cascade the resonances: source -> filter
x = resonator(x, fc)
fade = np.minimum(1.0, np.minimum(np.arange(n), n - np.arange(n)) / (0.03 * SR))
return x * fade / (np.abs(x).max() + 1e-9)
audio = 0.7 * np.concatenate([say("a"), say("i"), say("u")])
pcm = np.int16(np.clip(audio, -1, 1) * 32767)
with wave.open("vowels.wav", "wb") as f:
f.setnchannels(1); f.setsampwidth(2); f.setframerate(SR)
f.writeframes(pcm.tobytes())
print("synthesised", round(audio.size / SR, 2), "s of speech-like audio")
print("wrote vowels.wav:", os.path.getsize("vowels.wav"), "bytes")
print("pitch of the voice:", 120, "Hz")
print("lines of code doing the actual synthesis: about 15")synthesised 1.05 s of speech-like audio wrote vowels.wav: 33644 bytes pitch of the voice: 120 Hz lines of code doing the actual synthesis: about 15
Play vowels.wav. It is robotic and it is recognisably three vowels in a human-sounding voice. Fifteen lines, no training, no data.
This is a formant synthesiser, the technology behind Stephen Hawking's voice. It is worth building once, because it makes the source-filter idea physical rather than abstract.
Proving the filter does what it claims
The resonator chain is the "mouth". Here is its exact frequency response, so you can check it boosts the right pitches.
import numpy as np
from scipy.signal import freqz
SR = 16000
VOWELS = {"a": (730,1090,2440), "i": (270,2290,3010), "u": (300,870,2240)}
def reso_coeffs(fc, bw=90.0):
r = np.exp(-np.pi*bw/SR); th = 2*np.pi*fc/SR
return [1 - 2*r*np.cos(th) + r*r], [1.0, -2*r*np.cos(th), r*r]
def tract_response(formants, n=2048):
"""Combined gain of the whole mouth-resonance chain, in dB."""
total = np.ones(n)
for fc in formants:
b, a = reso_coeffs(fc)
w, h = freqz(b, a, worN=n, fs=SR)
total = total * np.abs(h)
return w, 20*np.log10(total + 1e-12)
probes = [270, 300, 730, 870, 1090, 2240, 2290, 2440, 3010]
print("gain of each vowel's mouth shape, in dB, at nine probe pitches")
print(" " + "".join(f"{p:>7}" for p in probes))
for v, formants in VOWELS.items():
w, db = tract_response(formants)
row = [db[np.abs(w - p).argmin()] for p in probes]
print(f" '{v}' " + "".join(f"{x:7.1f}" for x in row))gain of each vowel's mouth shape, in dB, at nine probe pitches
270 300 730 870 1090 2240 2290 2440 3010
'a' 1.9 2.4 24.3 17.2 22.2 -11.1 -9.7 -1.5 -31.8
'i' 10.0 7.9 -14.2 -16.9 -19.6 -3.1 0.7 -9.6 -5.2
'u' 10.5 11.9 -2.4 4.1 -14.1 -19.9 -24.2 -36.5 -55.6Read across the rows. The a shape boosts 730 Hz by 24 dB and cuts 3010 Hz by 32 dB. The i shape does the reverse: it cuts 730 Hz by 14 dB and passes the high region.
A difference of 24 dB against −14 dB is a factor of about 400 in amplitude at the same pitch, from nothing but the shape of your mouth. That ratio is what makes vowels distinguishable, and it is the same structure the MFCCs in audio features were designed to capture.
Real neural TTS, with the costs stated first
The synthesiser above is a teaching tool. For a natural voice you need a trained model, and that means a download. Approximate figures, before you commit:
| System | Download | Notes |
|---|---|---|
| Piper | roughly 20–110 MB per voice | ONNX, CPU real-time, many languages, designed for low-power devices |
| Kokoro-82M | roughly 310 MB | 82 M parameters, strong quality for the size, CPU-usable |
| SpeechT5 + vocoder | several hundred MB | Convenient through the Hugging Face pipeline |
| Coqui XTTS-v2 | well over 1 GB | Multilingual, voice cloning from a short sample |
Piper is the right first choice on a laptop or a Raspberry Pi. A single voice is smaller than a photo album, and it runs faster than real time on a CPU. See edge AI for why that size matters on a device.
A minimal Piper call, once you have downloaded a voice file:
import wave
from piper import PiperVoice
voice = PiperVoice.load("en_US-lessac-medium.onnx") # downloaded separately
with wave.open("hello.wav", "wb") as f:
voice.synthesize_wav("The train is arriving on platform three.", f)No output block, deliberately. What comes out is audio, not text, and its quality depends on the voice file, the version and your machine. An invented sample line would teach you nothing true.
Common mistakes
Skipping normalisation and blaming the model. The most common complaint about a TTS system — "it read the date wrong" — is a front-end bug, not an audio bug.
Feeding raw markup. HTML tags, markdown asterisks and emoji all get read aloud or mangled. Strip them before synthesis.
Ignoring the sample rate of the voice. Piper voices are typically 16 kHz or 22.05 kHz. Write the WAV at the wrong rate and the speech plays too fast or too slow, sounding like a different person.
Synthesising one long block. Split at sentence boundaries. It lowers latency for the first audible word, and it stops one bad sentence from spoiling a whole paragraph.
Assuming pronunciation control exists. If you need a specific pronunciation of a name, check whether your system supports a lexicon or phoneme input before you build on it. Many do not, and there is no workaround afterwards.
Try it yourself
Add a fourth vowel to VOWELS — "e" as in "bed" is roughly (530, 1840, 2480) — and listen to whether it is distinguishable. Then make f0 slide from 100 to 140 across the word by generating the pulse train with a changing period. That slide is intonation, and adding it is what makes the difference between a machine reading and a person speaking.
What to learn next
- Voice cloning and its risks — what happens when a system needs only seconds of your voice.
- Speaker identification — how a machine tells one voice from another.
- Real-time audio pipelines — fitting synthesis into a live conversation.
Researcher — Mathematics and papers.
The classical pipeline, formally
$$ \text{text} \xrightarrow{\text{TN}} \text{normalised} \xrightarrow{\text{G2P}} \text{phonemes} \xrightarrow{\text{duration}} \text{aligned} \xrightarrow{\text{acoustic}} \mathbf{M} \xrightarrow{\text{vocoder}} \mathbf{x} $$
Where $\mathbf{M} \in \mathbb{R}^{F \times T}$ is a mel spectrogram with $F$ bins over $T$ frames, and $\mathbf{x} \in \mathbb{R}^{N}$ the waveform with $N = T \cdot H$ for hop $H$. The vocoder's job is an upsampling of roughly $256\times$ in time while inventing phase.
Source-filter theory
Fant (1960) established the model the developer block implements. Voiced speech is an excitation $e(t)$ convolved with the vocal tract response $h(t)$ and lip radiation $r(t)$:
$$ s(t) = e(t) * h(t) * r(t) \quad \Longrightarrow \quad S(f) = E(f)\,H(f)\,R(f) $$
$E(f)$ carries $f_0$ and its harmonics, decaying at roughly $-12$ dB/octave. $H(f)$ has resonant peaks — formants — determined by tract geometry. $R(f)$ is approximately a $+6$ dB/octave differentiator, giving the net $-6$ dB/octave tilt that pre-emphasis in MFCC extraction removes.
Each two-pole resonator in the code realises one formant with poles at $r e^{\pm i\theta}$, where $\theta = 2\pi f_c / f_s$ sets centre frequency and $r = e^{-\pi B / f_s}$ sets bandwidth $B$. Klatt (1980) built a full synthesiser on cascaded and parallel resonators of exactly this form.
Neural acoustic models
Tacotron 2 (Shen et al., 2018) is an attention-based encoder-decoder emitting mel frames autoregressively, conditioned on character embeddings, paired with a WaveNet vocoder. It reached a mean opinion score of 4.53 against 4.58 for recorded human speech. Its weakness is the attention mechanism: with no monotonicity constraint it can skip or repeat words, and failure is catastrophic rather than graceful.
FastSpeech 2 (Ren et al., 2021) removes attention from the decoder entirely. A variance adaptor predicts per-phoneme duration, pitch and energy, and a length regulator expands the phoneme sequence to frame length. Generation is then non-autoregressive and fully parallel, giving large speedups and eliminating the skip-repeat failure mode. Duration targets come from forced alignment, originally from a teacher model and later from Montreal Forced Aligner or an internal aligner.
VITS (Kim, Kong and Son, 2021) collapses the pipeline into one end-to-end model: a conditional variational autoencoder with normalising flows, adversarial training, and a monotonic alignment search that learns alignment without external supervision. It outputs waveform directly, with no separate vocoder and no mel intermediate. It also introduces a stochastic duration predictor, so the same text produces different rhythms on different runs — a deliberate step toward one-to-many modelling.
The one-to-many problem is the field's structural difficulty: a single text has infinitely many valid renderings differing in prosody, and any model trained with a deterministic regression loss regresses to their mean, which is precisely the flat, lifeless delivery listeners complain about.
Vocoders
| Family | Example | Mechanism | Speed |
|---|---|---|---|
| Autoregressive | WaveNet (2016) | Dilated causal convolutions, sample-by-sample categorical | Far slower than real time |
| Flow | WaveGlow, Parallel WaveNet | Invertible transform of noise, parallel sampling | Real time on GPU |
| GAN | HiFi-GAN (2020), BigVGAN | Transposed convolutions, multi-period and multi-scale discriminators | Hundreds of times real time |
| Diffusion | DiffWave, WaveGrad | Iterative denoising | Quality-for-steps tradeoff |
HiFi-GAN (Kong, Kim and Bae, 2020) is the practical default. Its key idea is the multi-period discriminator: several discriminators each viewing the waveform reshaped by a different prime period (2, 3, 5, 7, 11), which captures the periodic structure of voiced speech that a plain 1-D discriminator misses. The generator loss combines adversarial, feature-matching and mel-spectrogram reconstruction terms.
Neural codec language models
The current frontier reframes TTS as next-token prediction over discrete audio tokens produced by a neural codec such as EnCodec or SoundStream.
VALL-E (Wang et al., 2023) treats TTS as conditional language modelling over EnCodec residual-vector-quantiser tokens, trained on 60,000 hours. It performs zero-shot voice cloning from a 3-second enrolment and preserves the speaker's acoustic environment and emotion — because it learned to continue the prompt, not to read text neutrally.
This capability is the direct subject of voice cloning and its risks, and the safety discussion is not separable from the modelling discussion.
Evaluation
MOS — listeners rate naturalness 1 to 5. Report confidence intervals and the number of listeners; MOS is not comparable across studies with different rater pools, instructions or reference conditions, and cross-paper MOS comparisons are routinely made and routinely invalid.
CMOS — comparative MOS on a $[-3, +3]$ scale against a reference. Far more reliable than absolute MOS for A/B decisions.
Intelligibility — run ASR over synthesised speech and measure WER. Cheap, automatic, and it catches skipped or repeated words that MOS raters forgive.
Speaker similarity — cosine similarity between speaker embeddings of reference and synthesis, using the models from speaker identification.
Real-time factor — synthesis time divided by audio duration. Below 1.0 is real time; below 0.1 is comfortable for interactive use.
No automatic metric correlates well with human judgement of prosody. This remains the honest bottleneck in TTS evaluation.
Reading
- Fant (1960), Acoustic Theory of Speech Production — the source-filter model.
- Klatt (1980), Software for a cascade/parallel formant synthesizer, JASA 67(3).
- van den Oord et al. (2016), WaveNet — arxiv.org/abs/1609.03499
- Shen et al. (2018), Natural TTS Synthesis by Conditioning WaveNet on Mel Spectrogram Predictions — arxiv.org/abs/1712.05884
- Ren et al. (2021), FastSpeech 2 — arxiv.org/abs/2006.04558
- Kim, Kong and Son (2021), VITS — arxiv.org/abs/2106.06103
- Kong, Kim and Bae (2020), HiFi-GAN — arxiv.org/abs/2010.05646
- Wang et al. (2023), Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers (VALL-E) — arxiv.org/abs/2301.02111
What to learn next
- Voice cloning and its risks — what happens when a system needs only seconds of your voice.
- Speaker identification — how a machine tells one voice from another.
- Real-time audio pipelines — fitting synthesis into a live conversation.