Speech and Audio AI

Waveforms and spectrograms

A waveform shows loudness over time and a spectrogram shows which pitches are present at each moment, which is the picture almost every audio model actually reads.

Read these first

On this page 9
  1. The short answer
  2. The analogy you have already lived
  3. Why the waveform is not enough
  4. The idea that fixes it
  5. From spectrum to spectrogram
  6. The trade nobody can escape
  7. Where you have already seen one
  8. Remember this
  9. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

A waveform shows how loud a sound is at each instant. A spectrogram shows which pitches are in it at each instant.

Most audio models never look at the waveform at all — they look at the spectrogram.

The analogy you have already lived

Think of a plate of dal. Looking at it from the side tells you how much there is. It tells you nothing about what went into it.

Now taste it. You pick out salt, then chilli, then a hint of garlic. Several separate flavours, all in the same spoonful at once.

A waveform is the side view: how much sound. A spectrogram is the tasting: which ingredients, in what proportion, moment by moment.

Why the waveform is not enough

In how computers hear sound you saw that audio is a list of numbers. Draw those numbers on a graph and you get a wiggly line. That line is the waveform.

The waveform is honest but nearly unreadable. Two people saying different words produce squiggles that look about the same to your eye and to a model.

Here is the deeper problem. When a flute and a drum play together, the microphone stores one number per instant, not two. Both sounds are mashed into a single wiggle. The waveform cannot tell you they were ever separate.

The idea that fixes it

There is a piece of mathematics that takes a chunk of wiggle and answers one question: which steady pitches, mixed together, would produce this exact wiggle?

Hand it a mixture and it hands back the recipe. Two hundred units of low hum, fifty units of a middle tone, none of anything higher.

That answer is called a spectrum — a list of how much of each pitch is present.

From spectrum to spectrogram

One spectrum describes a whole clip at once. That is not enough, because speech changes constantly. The word "cat" is not one sound; it is three, in order.

So you chop the recording into short overlapping slices, around twenty to thirty milliseconds each. You take a spectrum of each slice. Then you stand the spectra side by side, in time order.

That stack of spectra is a spectrogram.

                      time  ->
   high pitch  |                    ########
               |          ########
   low pitch   |########
               +---------------------------
                  "aa"      "ee"      "oo"

Read it like a chart. Left to right is time. Bottom to top is pitch, low to high. Dark or bright means loud.

Suddenly a sound is a picture. And we already have excellent tools for reading pictures. See convolutional neural networks, built for exactly this shape of data.

The trade nobody can escape

Here is the part that surprises people. You cannot have sharp timing and sharp pitch at the same time.

Use short slices and you know exactly when things happened. Each slice is then too brief to pin down pitch precisely. Use long slices and you get precise pitch, but you can no longer say exactly when the sound started.

This is not a limitation of our tools or our budget. It is a property of waves themselves. Physicists meet the same wall, and there is no clever trick that gets around it.

Read that paragraph twice. It confuses almost everyone the first time, and it explains a great many design choices later in this section.

Where you have already seen one

  • The bouncing bars on a music player are a live spectrum, redrawn many times a second.
  • Voice notes on WhatsApp show a simplified waveform as the little bumpy line.
  • Shazam recognises a song by matching bright spots in its spectrogram against a database.
  • Video call software finds background hum in the spectrogram and removes that band.

Remember this

  • A waveform is loudness over time. A spectrogram is pitch content over time.
  • Chopping audio into short overlapping slices, then taking a spectrum of each, produces the spectrogram.
  • Sharper timing costs you pitch detail, and the reverse. That trade is unavoidable.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install numpy

NumPy alone. Everything here runs on a CPU in about a second, with no audio files downloaded — the signals are generated in the script.

One spectrum: which pitches are in this clip?

np.fft.rfft takes a real signal and returns the strength of each pitch. The r means "real input", which halves the work and returns only the positive frequencies you care about.

one_spectrum.py
import numpy as np

SR = 8000
t = np.arange(SR) / SR                                  # one second of time stamps
x = np.sin(2*np.pi*440*t) + 0.5*np.sin(2*np.pi*1200*t)  # two tones, mixed into one signal

spec  = np.abs(np.fft.rfft(x))          # strength of every pitch
freqs = np.fft.rfftfreq(x.size, 1/SR)   # which pitch each number stands for

top = np.argsort(spec)[-2:][::-1]
print("waveform: ", np.round(x[:6], 3))
print("loudest pitches found:")
for i in sorted(top):
    print(f"  {freqs[i]:7.1f} Hz   strength {spec[i]:8.1f}")
Output
waveform:  [0.    0.743 1.113 1.015 0.688 0.488]
loudest pitches found:
    440.0 Hz   strength   4000.0
   1200.0 Hz   strength   2000.0

Look at what happened. The waveform row is six meaningless numbers. The spectrum recovered both original tones and their exact loudness ratio — 4000 against 2000, matching the amplitudes 1.0 and 0.5 that went in.

np.fft.rfftfreq(n, 1/SR) is the piece people forget. rfft returns strengths with no labels attached. This function builds the matching list of frequencies, so you know that index 440 means 440 Hz.

The problem with one spectrum

That clip held two tones for the whole second, so a single spectrum described it perfectly. Speech does not work that way.

If a voice says "aa" then "oo", one spectrum over the whole clip reports both, with no hint of the order. Reverse the recording and you get an identical spectrum. That is a serious loss of information.

Chopping into slices: the short-time Fourier transform

The fix is to take many spectra, each over a short slice. That is the STFT, short-time Fourier transform.

spectrogram.py
import numpy as np

SR, N_FFT, HOP = 8000, 256, 128

# A three-note signal: the pitch steps up twice. One spectrum cannot show this.
t = np.arange(SR) / SR
pitch = np.where(t < 0.33, 500.0, np.where(t < 0.66, 1000.0, 2000.0))
x = 0.6 * np.sin(2 * np.pi * np.cumsum(pitch) / SR)   # cumsum keeps the phase continuous

window = np.hanning(N_FFT)                            # fade each slice in and out at its edges
frames = np.array([x[i:i + N_FFT] * window
                   for i in range(0, len(x) - N_FFT, HOP)])
spec = np.abs(np.fft.rfft(frames, axis=1))            # one column of pitches per slice
db   = np.clip(20 * np.log10(spec + 1e-10) - 20 * np.log10(spec.max()), -60, 0)

freqs = np.fft.rfftfreq(N_FFT, 1 / SR)
chars = " .:-=+*#%@"
print("        time ->")
for r in range(72, -1, -8):                           # every 8th bin, high pitch on top
    line = "".join(chars[int((db[c, r] + 60) / 60 * 9)]
                   for c in range(0, spec.shape[0], 2))
    print(f"{freqs[r]:5.0f} Hz |{line}|")
print("         +" + "-" * len(line) + "+")
Output
        time ->
 2250 Hz |          .         =          |
 2000 Hz |          .         #@@@@@@@@@@|
 1750 Hz |          .         =          |
 1500 Hz |          :         =          |
 1250 Hz |          -         =          |
 1000 Hz |          %%%%%%%%%%%          |
  750 Hz |          +         =          |
  500 Hz |%%%%%%%%%%*         -          |
  250 Hz |          =         -          |
    0 Hz |          -         :          |
         +-------------------------------+

That is a spectrogram, printed with text. The three notes appear as three horizontal bars, each higher than the last. No plotting library was needed to see it.

Reading the details in that output

The vertical smears at the two transitions are real, not a bug. An abrupt pitch change is a sudden event, and sudden events contain energy at every pitch at once. A slice that straddles the boundary sees a bit of everything. Handclaps and consonants like "t" and "k" look the same way.

np.cumsum(pitch) / SR builds the phase by accumulating it. Writing np.sin(2*np.pi*pitch*t) instead would make the wave jump discontinuously at each note change, adding a loud click. Accumulating keeps the wave continuous through the change.

np.hanning(N_FFT) is a window function: a bell shape that fades each slice to zero at both ends. Chopping a wave with a hard edge creates a false click, and that click spreads fake energy across every pitch. This is called spectral leakage, and the window suppresses it. Skip the window and every horizontal bar in that picture gets a fuzzy halo.

20 * np.log10(...) converts to decibels. Hearing is roughly logarithmic: the gap between a whisper and speech feels like the gap between speech and a shout, though the raw numbers differ by thousands. Nearly every audio model is fed log-scaled magnitudes for this reason.

+ 1e-10 prevents log10(0) from returning negative infinity. Any silent bin becomes -inf, then NaN on the next arithmetic step, and the whole array is poisoned. This tiny constant is standard practice, not a hack.

The resolution trade, measured

Change only the window length and watch what it costs you.

window_tradeoff.py
import numpy as np

SR     = 8000
LADDER = [2500, 2000, 1500, 1000, 500, 0]     # the rows we print, in Hz
CHARS  = " .:-=+*#%@"

t = np.arange(SR) / SR
pitch = np.where(t < 0.33, 500.0, np.where(t < 0.66, 1000.0, 2000.0))
x = 0.6 * np.sin(2 * np.pi * np.cumsum(pitch) / SR)

def ascii_spec(x, n_fft, cols=32):
    hop = n_fft // 2
    w   = np.hanning(n_fft)
    fr  = np.array([x[i:i + n_fft] * w for i in range(0, len(x) - n_fft, hop)])
    sp  = np.abs(np.fft.rfft(fr, axis=1))
    db  = np.clip(20*np.log10(sp + 1e-10) - 20*np.log10(sp.max()), -60, 0)
    freqs = np.fft.rfftfreq(n_fft, 1 / SR)
    idx = np.linspace(0, len(db) - 1, cols).astype(int)     # sample columns evenly
    print(f"window {n_fft:>4} samples = {n_fft/SR*1000:5.1f} ms | "
          f"bin spacing {SR/n_fft:5.1f} Hz | {len(db):>3} frames")
    for f in LADDER:
        r = np.abs(freqs - f).argmin()
        line = "".join(CHARS[int((db[c, r] + 60) / 60 * 9)] for c in idx)
        print(f"  {f:5d} Hz |{line}|")
    print()

ascii_spec(x, 64)
ascii_spec(x, 256)
ascii_spec(x, 1024)
Output
window   64 samples =   8.0 ms | bin spacing 125.0 Hz | 248 frames
   2500 Hz |                                |
   2000 Hz |                     @@@@@@@@@@@|
   1500 Hz |                                |
   1000 Hz |           %%%%%%%%%%           |
    500 Hz |%%%%%%%%%%%                     |
      0 Hz |                                |

window  256 samples =  32.0 ms | bin spacing  31.2 Hz |  61 frames
   2500 Hz |                     -          |
   2000 Hz |                     #@@@@@@@@@@|
   1500 Hz |          .          =          |
   1000 Hz |          =%%%%%%%%%%%          |
    500 Hz |%%%%%%%%%%%          -          |
      0 Hz |          :          :          |

window 1024 samples = 128.0 ms | bin spacing   7.8 Hz |  14 frames
   2500 Hz |                      ..        |
   2000 Hz |                      **%%%@@@@@|
   1500 Hz |          ..          ::        |
   1000 Hz |          ##%%%%%%%%%%%%===     |
    500 Hz |%%%%%%%%%%%%...       ..        |
      0 Hz |          ::                    |

Compare the top and bottom pictures carefully.

At 64 samples the note changes are knife-sharp, but bin spacing is 125 Hz, so pitch is described coarsely. At 1024 samples bin spacing is 7.8 Hz, sixteen times finer — and the notes now bleed into each other in time. Look at the 500 Hz row in the last picture: it still shows energy after the note has already changed.

The 256-sample window in the middle is the usual compromise, and 20 to 30 ms is the standard choice for speech. That is not tradition. It is roughly the length over which a vocal tract holds still.

Common mistakes

Forgetting the window function. Slicing without np.hanning leaks energy across all frequencies and blurs every horizontal line. There is no case where a rectangular window is the right default.

Taking log of zero. Silence gives -inf, then NaN, then a model that trains to NaN loss twenty minutes later. Always add a small epsilon, or use np.log(x + 1e-10).

Feeding raw magnitudes to a network. Linear magnitudes span an enormous range, so the loudest frame dominates the gradient and quiet detail is ignored. Use decibels or a log.

Throwing away phase without realising it. np.abs() discards the timing offset of each pitch. Magnitude is enough to recognise audio, and not enough to rebuild it faithfully. Reconstruction is a separate, harder problem.

Mismatched hop between training and inference. A model trained on HOP=160 and served with HOP=128 sees a time axis stretched by 25 percent. It does not crash; accuracy drops and nobody can find why.

Try it yourself

Set HOP = 256 in the spectrogram script, making the slices touch but not overlap, and count the frames. Then add + 0.3 * np.random.default_rng(0).standard_normal(x.size) to the signal and see how much noise the bars survive. Noise robustness is the whole game in real audio work.

What to learn next

Researcher — Mathematics and papers.

The discrete Fourier transform

For a finite sequence $x[n]$ of length $N$:

$$ X[k] = \sum_{n=0}^{N-1} x[n]\, e^{-i 2\pi k n / N}, \qquad k = 0, \dots, N-1 $$

Where $x[n]$ is the sampled signal, $N$ is the transform length, $k$ indexes the frequency bins, and $X[k] \in \mathbb{C}$. Bin $k$ corresponds to the physical frequency $f_k = k f_s / N$, so bin spacing is $\Delta f = f_s / N$.

The FFT (Cooley and Tukey, 1965) computes this in $O(N \log N)$ rather than $O(N^2)$ by recursively factoring the transform. For real input, Hermitian symmetry $X[N-k] = \overline{X[k]}$ means only $\lfloor N/2 \rfloor + 1$ bins are independent, which is what rfft returns.

The short-time Fourier transform

$$ X[m, k] = \sum_{n=0}^{N-1} x[n + mH]\, w[n]\, e^{-i 2\pi k n / N} $$

Where $m$ is the frame index, $H$ is the hop size in samples, $w[n]$ is the analysis window of length $N$, and $N$ is the FFT length. The spectrogram is $|X[m,k]|^2$; the magnitude spectrogram $|X[m,k]|$ is what most pipelines use.

Frame rate is $f_s / H$. For $f_s = 16000$ and $H = 160$ you get 100 frames per second, the near-universal convention in speech systems, chosen so one frame equals 10 ms.

Window functions

The rectangular window has a narrow main lobe ($2 f_s / N$ between nulls) and dreadful sidelobes, the first at $-13$ dB, decaying at 6 dB per octave. Tapered windows trade main-lobe width for sidelobe suppression:

WindowMain lobe widthPeak sidelobeSidelobe rolloff
Rectangular$2 f_s/N$−13 dB−6 dB/octave
Hann$4 f_s/N$−31 dB−18 dB/octave
Hamming$4 f_s/N$−43 dB−6 dB/octave
Blackman$6 f_s/N$−58 dB−18 dB/octave

Hann is the default in speech work because its fast rolloff matters more than its first-sidelobe height when the signal is harmonic.

For invertibility, the window and hop must satisfy the constant overlap-add condition, $\sum_m w[n - mH] = \text{const}$ for all $n$. Hann with $H = N/2$ satisfies it exactly. The weaker NOLA condition (nonzero overlap-add) is what scipy.signal.istft actually requires.

The uncertainty limit

Define the time spread $\sigma_t$ and frequency spread $\sigma_f$ of a window as the standard deviations of $|w(t)|^2$ and $|W(f)|^2$. Then:

$$ \sigma_t \, \sigma_f \geq \frac{1}{4\pi} $$

This is the Gabor limit (Gabor, 1946), mathematically identical to the Heisenberg inequality — both are properties of Fourier pairs, not of physics or of measurement quality. Equality holds only for a Gaussian window.

The consequence is concrete. A 25 ms window cannot resolve frequency more finely than roughly 40 Hz, whatever zero-padding you apply. Zero-padding interpolates the spectrum onto a denser grid; it adds no information and does not improve resolution. This is the single most common misunderstanding in applied audio work.

Multi-resolution methods sidestep it only by changing the question: wavelets and constant-Q transforms (Brown, 1991) vary the window length with frequency, buying fine timing at high frequencies and fine pitch at low ones.

The mel scale

Linear frequency bins mismatch human hearing, which discriminates far better at low frequencies. The mel scale (Stevens, Volkmann and Newman, 1937) applies a roughly logarithmic warp. The common HTK form is:

$$ m = 2595 \log_{10}!\left(1 + \frac{f}{700}\right) $$

A mel spectrogram projects the linear magnitude spectrogram through a bank of overlapping triangular filters spaced evenly in mel. For 16 kHz speech, 80 mel bins from a 400-sample window is the modern default, used by Whisper, Tacotron 2 and most neural vocoders. It compresses 201 linear bins to 80 with negligible loss of what matters for speech. This is covered further in audio features and MFCCs.

Phase, and why discarding it hurts

$|X[m,k]|$ throws away $\arg X[m,k]$. Recognition tolerates this well; reconstruction does not.

Griffin and Lim (1984) recover phase by alternating projection between the set of spectrograms with the target magnitude and the set of consistent STFTs of real signals. It converges to a local optimum and produces a characteristic metallic artefact. Modern neural vocoders (HiFi-GAN, Kong et al., 2020) learn the magnitude-to-waveform map directly and comprehensively outperform it.

Consistency is the underlying issue: an arbitrary array of magnitudes is generally not the STFT of any real signal, because overlapping frames impose constraints that an unconstrained array violates.

Learned front-ends

Several lines of work replace the fixed STFT with learned filters. SincNet (Ravanelli and Bengio, 2018) parameterises the first convolutional layer as band-pass sinc filters with learnable cutoffs. LEAF (Zeghidour et al., 2021) learns filterbank, pooling and compression jointly. wav2vec 2.0 (Baevski et al., 2020) uses a plain strided convolutional encoder over raw waveform.

The honest summary: learned front-ends match but rarely decisively beat log-mel on large supervised datasets, while costing more compute and more tuning. Whisper, trained on 680k hours, still uses a fixed 80-bin log-mel front-end. Treat "learn the front-end" as a research direction, not a default.

Reading

  • Cooley and Tukey (1965), An Algorithm for the Machine Calculation of Complex Fourier Series, Math. Comp. 19(90).
  • Gabor (1946), Theory of Communication, J. IEE 93(26) — the time-frequency uncertainty relation.
  • Harris (1978), On the Use of Windows for Harmonic Analysis with the DFT, Proc. IEEE 66(1) — still the reference table for window properties.
  • Griffin and Lim (1984), Signal Estimation from Modified Short-Time Fourier Transform, IEEE TASSP 32(2).
  • Ravanelli and Bengio (2018), Speaker Recognition from Raw Waveform with SincNet — arxiv.org/abs/1808.00158

What to learn next