Speech and Audio AI

Text to speech

Text to speech turns written words into audio, and the work splits into deciding how the words should sound and then generating the actual waveform.

On this page 10
  1. The short answer
  2. The analogy you have already lived
  3. Why the first half is harder than the second
  4. Why it exists
  5. How it works
  6. How voices got good
  7. Where you have already heard it
  8. What is honestly hard here
  9. Remember this
  10. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

Text to speech takes written words and produces audio of someone saying them.

The short name is TTS. It is speech recognition run backwards, and it is a harder problem than it sounds.

The analogy you have already lived

Think of being asked to read a paragraph aloud in class. You did two separate jobs, and you probably never noticed.

First you worked out how it should sound. Where to pause. Which word to stress. Whether the last sentence was a question, so your voice should rise. You did this before opening your mouth.

Then you actually made the sound, using your lungs, vocal cords, tongue and lips.

Machines split the work the same way. One part decides how it should sound. Another part produces the audio. Almost every TTS system you meet has these two halves.

Why the first half is harder than the second

Making a voice sound realistic is largely solved. Deciding how to read something is not.

Take these:

  • "I never said she stole my money." Seven words. Stress a different one each time and you get seven different meanings. Nothing in the writing says which.
  • "I read the book." Is that "reed" or "red"? Same letters, two words, and only the tense decides.
  • "Dr. Rao lives on Oxford St." The first is "Doctor". The second is "Street". Identical abbreviation.
  • "Turn left in 100 m." Is that "one hundred metres", "one zero zero", or "a hundred"?

A human resolves these without effort by understanding the sentence. A machine has to be taught. Mistakes here make a synthetic voice sound wrong even when the audio is flawless.

Why it exists

Screen readers gave blind and low-vision people access to computers. That was the original reason, and it remains the most important one.

Then it spread. Announcements at railway stations. Navigation directions. Audiobooks for text nobody recorded. Reading a message aloud while you drive. Voice assistants answering out loud.

And a use that matters enormously to the people it serves: voice banking. Someone losing their speech to illness records their voice while they still can. A synthesiser can then speak in their own voice.

How it works

   "The train is arriving."
            |
            v
  [ normalise ]      expand numbers, dates, abbreviations
            |
            v
  [ pronounce ]      letters -> sounds (phonemes)
            |
            v
  [ plan the delivery ]   how long each sound lasts, where the pitch rises
            |
            v
  [ acoustic model ]  produce a spectrogram of the intended speech
            |
            v
  [ vocoder ]        turn that picture back into a wave
            |
            v
   audio you can play

The last box has a name worth remembering. A vocoder is the part that converts a picture of sound into actual sound. It exists because drawing a spectrogram is far easier for the acoustic model. Producing sixteen thousand samples a second directly is much harder.

How voices got good

Concatenative (1990s to 2010s). Record one person for many hours, chop it into tiny pieces, and glue pieces together for new sentences. It sounded like a real human, because it was, and it lurched at every join. Old railway announcements are this.

Parametric. Model the voice with a small set of controls. Smooth, flexible, and unmistakably robotic.

Neural (2016 onward). Learn the whole mapping from data. WaveNet in 2016 was the break: it modelled the waveform one sample at a time and closed most of the gap to human speech. It was also so slow that generating one second took minutes. Everything since has been about keeping that quality while running fast enough to be useful.

Where you have already heard it

  • Google Maps reading out your turns.
  • Screen readers on phones and computers.
  • Railway and airport announcements, increasingly synthetic rather than recorded.
  • Audiobooks of books nobody paid a narrator to read.
  • Video narration on social media, in that recognisable app voice.

What is honestly hard here

Modern TTS sounds excellent for a plain sentence read neutrally. It is still weak at:

  • Long-form emotion. A whole chapter read with genuine feeling, varying naturally, remains hard.
  • Names and rare words. Your surname, your village, a medicine name. Systems guess from spelling and often guess wrong.
  • Code-mixed text. An English sentence with Hindi words in it, which is how a great many people actually write.
  • Knowing what it does not know. A TTS system mispronounces confidently. It never hesitates.

And one more thing, which gets its own lesson: a system that can copy a voice from a short sample can copy your voice. See voice cloning and its risks.

Remember this

  • TTS splits into deciding how it should sound and producing the audio.
  • The vocoder turns a spectrogram back into a playable wave.
  • Deciding pronunciation and emphasis from plain text is the unsolved half, not the audio.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install numpy scipy

Everything verified in this section runs on NumPy and SciPy, on a CPU, with no model download. You will build a working speech synthesiser in about fifteen lines and write a WAV file you can play.

Real neural TTS costs a download, and honest sizes are listed near the end before you are asked to fetch anything.

Stage one: normalising the text

Before any audio exists, the text has to be turned into what a person would actually say. This stage is unglamorous and it is where most audible errors are born.

normalise.py
import re

ONES = ["zero","one","two","three","four","five","six","seven","eight","nine","ten",
        "eleven","twelve","thirteen","fourteen","fifteen","sixteen","seventeen",
        "eighteen","nineteen"]
TENS = ["","","twenty","thirty","forty","fifty","sixty","seventy","eighty","ninety"]

def say_number(n):
    if n < 20:   return ONES[n]
    if n < 100:  return TENS[n//10] + ("" if n%10 == 0 else " " + ONES[n%10])
    if n < 1000: return ONES[n//100] + " hundred" + ("" if n%100 == 0 else " and " + say_number(n%100))
    return say_number(n//1000) + " thousand" + ("" if n%1000 == 0 else " " + say_number(n%1000))

ABBREV = {"dr": "doctor", "mr": "mister", "st": "saint", "km": "kilometres"}

def normalise(text):
    text = re.sub(r"\bRs\.?\s*(\d+)",                       # money, before anything else
                  lambda m: say_number(int(m.group(1))) + " rupees", text)
    text = re.sub(r"\b([A-Za-z]{2})\.",                     # then two-letter abbreviations
                  lambda m: ABBREV.get(m.group(1).lower(), m.group(1)), text)
    text = re.sub(r"\b(\d+)\b", lambda m: say_number(int(m.group(1))), text)
    return re.sub(r"\s+", " ", text).strip()

tests = [
    "The fare is Rs. 250.",
    "Dr. Rao lives 3 km away.",
    "Meet me at St. Mary's church.",
    "Turn left onto Oxford St. and stop.",
    "I read the book last night.",
]
for line in tests:
    print(f"in : {line}")
    print(f"out: {normalise(line)}\n")
Output
in : The fare is Rs. 250.
out: The fare is two hundred and fifty rupees.

in : Dr. Rao lives 3 km away.
out: doctor Rao lives three km away.

in : Meet me at St. Mary's church.
out: Meet me at saint Mary's church.

in : Turn left onto Oxford St. and stop.
out: Turn left onto Oxford saint and stop.

in : I read the book last night.
out: I read the book last night.

Three of those five are wrong, and that is the point of showing them.

Line 2: 3 km became "three km". The abbreviation rule requires a trailing full stop, and "km" has none. A rule gap, fixable with more rules.

Line 4: Oxford St. became "Oxford saint". It should be "Street". The identical token means opposite things depending on whether a place name precedes it. No regex resolves that; it needs to know that Oxford is a road here and a person elsewhere.

Line 5: read passed through untouched, which looks fine and is not. A synthesiser must choose "reed" or "red", and "last night" is the only clue. This is a homograph: same spelling, different pronunciation, disambiguated by grammar and meaning alone.

This is why modern systems use a learned front-end rather than a rule table. Rules handle the common 95 percent and the remaining 5 percent is where listeners notice.

Stage two: making sound from nothing

Now the audio. Speech is well modelled as a source and a filter: vocal cords buzzing, and a mouth shaping the buzz. Both are short.

formant_synth.py
import numpy as np, wave, os
from scipy.signal import lfilter

SR = 16000
VOWELS = {                      # the first three mouth resonances, in Hz
    "a": (730, 1090, 2440),     # as in "father"
    "i": (270, 2290, 3010),     # as in "see"
    "u": (300,  870, 2240),     # as in "boot"
}

def glottal_pulses(n, f0):
    """The source: vocal cords snapping shut f0 times a second."""
    src = np.zeros(n)
    src[::int(SR / f0)] = 1.0
    return src

def resonator(x, fc, bw=90.0):
    """The filter: a two-pole resonance that rings at fc, one mouth cavity."""
    r  = np.exp(-np.pi * bw / SR)
    th = 2 * np.pi * fc / SR
    return lfilter([1 - 2*r*np.cos(th) + r*r], [1.0, -2*r*np.cos(th), r*r], x)

def say(vowel, seconds=0.35, f0=120):
    n = int(SR * seconds)
    x = glottal_pulses(n, f0)
    for fc in VOWELS[vowel]:            # cascade the resonances: source -> filter
        x = resonator(x, fc)
    fade = np.minimum(1.0, np.minimum(np.arange(n), n - np.arange(n)) / (0.03 * SR))
    return x * fade / (np.abs(x).max() + 1e-9)

audio = 0.7 * np.concatenate([say("a"), say("i"), say("u")])
pcm = np.int16(np.clip(audio, -1, 1) * 32767)
with wave.open("vowels.wav", "wb") as f:
    f.setnchannels(1); f.setsampwidth(2); f.setframerate(SR)
    f.writeframes(pcm.tobytes())

print("synthesised", round(audio.size / SR, 2), "s of speech-like audio")
print("wrote vowels.wav:", os.path.getsize("vowels.wav"), "bytes")
print("pitch of the voice:", 120, "Hz")
print("lines of code doing the actual synthesis: about 15")
Output
synthesised 1.05 s of speech-like audio
wrote vowels.wav: 33644 bytes
pitch of the voice: 120 Hz
lines of code doing the actual synthesis: about 15

Play vowels.wav. It is robotic and it is recognisably three vowels in a human-sounding voice. Fifteen lines, no training, no data.

This is a formant synthesiser, the technology behind Stephen Hawking's voice. It is worth building once, because it makes the source-filter idea physical rather than abstract.

Proving the filter does what it claims

The resonator chain is the "mouth". Here is its exact frequency response, so you can check it boosts the right pitches.

tract_response.py
import numpy as np
from scipy.signal import freqz

SR = 16000
VOWELS = {"a": (730,1090,2440), "i": (270,2290,3010), "u": (300,870,2240)}

def reso_coeffs(fc, bw=90.0):
    r = np.exp(-np.pi*bw/SR); th = 2*np.pi*fc/SR
    return [1 - 2*r*np.cos(th) + r*r], [1.0, -2*r*np.cos(th), r*r]

def tract_response(formants, n=2048):
    """Combined gain of the whole mouth-resonance chain, in dB."""
    total = np.ones(n)
    for fc in formants:
        b, a = reso_coeffs(fc)
        w, h = freqz(b, a, worN=n, fs=SR)
        total = total * np.abs(h)
    return w, 20*np.log10(total + 1e-12)

probes = [270, 300, 730, 870, 1090, 2240, 2290, 2440, 3010]
print("gain of each vowel's mouth shape, in dB, at nine probe pitches")
print("      " + "".join(f"{p:>7}" for p in probes))
for v, formants in VOWELS.items():
    w, db = tract_response(formants)
    row = [db[np.abs(w - p).argmin()] for p in probes]
    print(f"  '{v}' " + "".join(f"{x:7.1f}" for x in row))
Output
gain of each vowel's mouth shape, in dB, at nine probe pitches
          270    300    730    870   1090   2240   2290   2440   3010
  'a'     1.9    2.4   24.3   17.2   22.2  -11.1   -9.7   -1.5  -31.8
  'i'    10.0    7.9  -14.2  -16.9  -19.6   -3.1    0.7   -9.6   -5.2
  'u'    10.5   11.9   -2.4    4.1  -14.1  -19.9  -24.2  -36.5  -55.6

Read across the rows. The a shape boosts 730 Hz by 24 dB and cuts 3010 Hz by 32 dB. The i shape does the reverse: it cuts 730 Hz by 14 dB and passes the high region.

A difference of 24 dB against −14 dB is a factor of about 400 in amplitude at the same pitch, from nothing but the shape of your mouth. That ratio is what makes vowels distinguishable, and it is the same structure the MFCCs in audio features were designed to capture.

Real neural TTS, with the costs stated first

The synthesiser above is a teaching tool. For a natural voice you need a trained model, and that means a download. Approximate figures, before you commit:

SystemDownloadNotes
Piperroughly 20–110 MB per voiceONNX, CPU real-time, many languages, designed for low-power devices
Kokoro-82Mroughly 310 MB82 M parameters, strong quality for the size, CPU-usable
SpeechT5 + vocoderseveral hundred MBConvenient through the Hugging Face pipeline
Coqui XTTS-v2well over 1 GBMultilingual, voice cloning from a short sample

Piper is the right first choice on a laptop or a Raspberry Pi. A single voice is smaller than a photo album, and it runs faster than real time on a CPU. See edge AI for why that size matters on a device.

A minimal Piper call, once you have downloaded a voice file:

piper_speak.py
import wave
from piper import PiperVoice

voice = PiperVoice.load("en_US-lessac-medium.onnx")   # downloaded separately
with wave.open("hello.wav", "wb") as f:
    voice.synthesize_wav("The train is arriving on platform three.", f)

No output block, deliberately. What comes out is audio, not text, and its quality depends on the voice file, the version and your machine. An invented sample line would teach you nothing true.

Common mistakes

Skipping normalisation and blaming the model. The most common complaint about a TTS system — "it read the date wrong" — is a front-end bug, not an audio bug.

Feeding raw markup. HTML tags, markdown asterisks and emoji all get read aloud or mangled. Strip them before synthesis.

Ignoring the sample rate of the voice. Piper voices are typically 16 kHz or 22.05 kHz. Write the WAV at the wrong rate and the speech plays too fast or too slow, sounding like a different person.

Synthesising one long block. Split at sentence boundaries. It lowers latency for the first audible word, and it stops one bad sentence from spoiling a whole paragraph.

Assuming pronunciation control exists. If you need a specific pronunciation of a name, check whether your system supports a lexicon or phoneme input before you build on it. Many do not, and there is no workaround afterwards.

Try it yourself

Add a fourth vowel to VOWELS — "e" as in "bed" is roughly (530, 1840, 2480) — and listen to whether it is distinguishable. Then make f0 slide from 100 to 140 across the word by generating the pulse train with a changing period. That slide is intonation, and adding it is what makes the difference between a machine reading and a person speaking.

What to learn next

Researcher — Mathematics and papers.

The classical pipeline, formally

$$ \text{text} \xrightarrow{\text{TN}} \text{normalised} \xrightarrow{\text{G2P}} \text{phonemes} \xrightarrow{\text{duration}} \text{aligned} \xrightarrow{\text{acoustic}} \mathbf{M} \xrightarrow{\text{vocoder}} \mathbf{x} $$

Where $\mathbf{M} \in \mathbb{R}^{F \times T}$ is a mel spectrogram with $F$ bins over $T$ frames, and $\mathbf{x} \in \mathbb{R}^{N}$ the waveform with $N = T \cdot H$ for hop $H$. The vocoder's job is an upsampling of roughly $256\times$ in time while inventing phase.

Source-filter theory

Fant (1960) established the model the developer block implements. Voiced speech is an excitation $e(t)$ convolved with the vocal tract response $h(t)$ and lip radiation $r(t)$:

$$ s(t) = e(t) * h(t) * r(t) \quad \Longrightarrow \quad S(f) = E(f)\,H(f)\,R(f) $$

$E(f)$ carries $f_0$ and its harmonics, decaying at roughly $-12$ dB/octave. $H(f)$ has resonant peaks — formants — determined by tract geometry. $R(f)$ is approximately a $+6$ dB/octave differentiator, giving the net $-6$ dB/octave tilt that pre-emphasis in MFCC extraction removes.

Each two-pole resonator in the code realises one formant with poles at $r e^{\pm i\theta}$, where $\theta = 2\pi f_c / f_s$ sets centre frequency and $r = e^{-\pi B / f_s}$ sets bandwidth $B$. Klatt (1980) built a full synthesiser on cascaded and parallel resonators of exactly this form.

Neural acoustic models

Tacotron 2 (Shen et al., 2018) is an attention-based encoder-decoder emitting mel frames autoregressively, conditioned on character embeddings, paired with a WaveNet vocoder. It reached a mean opinion score of 4.53 against 4.58 for recorded human speech. Its weakness is the attention mechanism: with no monotonicity constraint it can skip or repeat words, and failure is catastrophic rather than graceful.

FastSpeech 2 (Ren et al., 2021) removes attention from the decoder entirely. A variance adaptor predicts per-phoneme duration, pitch and energy, and a length regulator expands the phoneme sequence to frame length. Generation is then non-autoregressive and fully parallel, giving large speedups and eliminating the skip-repeat failure mode. Duration targets come from forced alignment, originally from a teacher model and later from Montreal Forced Aligner or an internal aligner.

VITS (Kim, Kong and Son, 2021) collapses the pipeline into one end-to-end model: a conditional variational autoencoder with normalising flows, adversarial training, and a monotonic alignment search that learns alignment without external supervision. It outputs waveform directly, with no separate vocoder and no mel intermediate. It also introduces a stochastic duration predictor, so the same text produces different rhythms on different runs — a deliberate step toward one-to-many modelling.

The one-to-many problem is the field's structural difficulty: a single text has infinitely many valid renderings differing in prosody, and any model trained with a deterministic regression loss regresses to their mean, which is precisely the flat, lifeless delivery listeners complain about.

Vocoders

FamilyExampleMechanismSpeed
AutoregressiveWaveNet (2016)Dilated causal convolutions, sample-by-sample categoricalFar slower than real time
FlowWaveGlow, Parallel WaveNetInvertible transform of noise, parallel samplingReal time on GPU
GANHiFi-GAN (2020), BigVGANTransposed convolutions, multi-period and multi-scale discriminatorsHundreds of times real time
DiffusionDiffWave, WaveGradIterative denoisingQuality-for-steps tradeoff

HiFi-GAN (Kong, Kim and Bae, 2020) is the practical default. Its key idea is the multi-period discriminator: several discriminators each viewing the waveform reshaped by a different prime period (2, 3, 5, 7, 11), which captures the periodic structure of voiced speech that a plain 1-D discriminator misses. The generator loss combines adversarial, feature-matching and mel-spectrogram reconstruction terms.

Neural codec language models

The current frontier reframes TTS as next-token prediction over discrete audio tokens produced by a neural codec such as EnCodec or SoundStream.

VALL-E (Wang et al., 2023) treats TTS as conditional language modelling over EnCodec residual-vector-quantiser tokens, trained on 60,000 hours. It performs zero-shot voice cloning from a 3-second enrolment and preserves the speaker's acoustic environment and emotion — because it learned to continue the prompt, not to read text neutrally.

This capability is the direct subject of voice cloning and its risks, and the safety discussion is not separable from the modelling discussion.

Evaluation

MOS — listeners rate naturalness 1 to 5. Report confidence intervals and the number of listeners; MOS is not comparable across studies with different rater pools, instructions or reference conditions, and cross-paper MOS comparisons are routinely made and routinely invalid.

CMOS — comparative MOS on a $[-3, +3]$ scale against a reference. Far more reliable than absolute MOS for A/B decisions.

Intelligibility — run ASR over synthesised speech and measure WER. Cheap, automatic, and it catches skipped or repeated words that MOS raters forgive.

Speaker similarity — cosine similarity between speaker embeddings of reference and synthesis, using the models from speaker identification.

Real-time factor — synthesis time divided by audio duration. Below 1.0 is real time; below 0.1 is comfortable for interactive use.

No automatic metric correlates well with human judgement of prosody. This remains the honest bottleneck in TTS evaluation.

Reading

What to learn next

What to learn next

These follow on from what you just read.

  • Speech and Audio AI

    Speaker identification

    Speaker identification works out who is talking rather than what they said, by turning each voice into a fingerprint vector and comparing distances between fingerprints.

  • Speech and Audio AI

    Audio classification

    Audio classification puts a label on a sound, and the hard part is not building the model but making it survive noise it never met during training.

  • Speech and Audio AI

    Noise reduction

    Noise reduction removes background sound from a recording by editing its spectrogram, and pushing too hard trades hiss for a worse artefact called musical noise.