Speech and Audio AI

Audio features and MFCCs

Audio features squeeze a second of sound into a handful of numbers that describe its character, and MFCCs are the classic recipe for doing that to speech.

On this page 9
  1. The short answer
  2. The analogy you have already lived
  3. Why this had to be invented
  4. The two ideas behind MFCCs
  5. How it works
  6. Where you have already used this
  7. What is honestly hard here
  8. Remember this
  9. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

An audio feature is a small number describing some quality of a sound. It saves a model from reading every raw measurement.

MFCCs are the most famous set of them for speech.

The analogy you have already lived

Think about describing a friend to someone who has never met them. You do not read out every strand of hair. You say: tall, curly hair, loud laugh, always in a blue jacket.

Five short facts, and the other person can pick your friend out of a crowd. You threw away almost everything and kept the parts that distinguish.

Audio features do that to sound. One second of speech is sixteen thousand numbers. Features boil it down to a few dozen that carry the character.

Why this had to be invented

Feeding raw audio to a model is wasteful and fragile.

It is enormous. Ten seconds of speech at the usual rate is one hundred and sixty thousand numbers. A picture of the same size would be a large photograph.

Most of it does not matter. Say the same word twice and the raw numbers differ almost completely. Start a hundredth of a second later and every value shifts. Yet a human hears the same word.

It hides the useful part. What makes "sh" sound different from "ss" is the shape of energy across pitches. That shape is buried in the raw wiggle.

So for decades, engineers built a compression step. Keep what makes sounds distinguishable, discard the rest, and hand a model something small and stable.

The two ideas behind MFCCs

MFCC stands for mel-frequency cepstral coefficients. The name is unhelpful. The two ideas underneath it are not.

Idea one: your ears are not a ruler. You can easily hear the difference between a two hundred and a three hundred cycle tone. You cannot hear the difference between eight thousand and eight thousand one hundred, although the gap is the same size.

Hearing is stretched at the bottom and squashed at the top. So we stop measuring every pitch equally. We group pitches into bands: narrow down low, wide up high. That grouping is the mel scale, named after "melody".

Idea two: separate the voice from the mouth. Speech is made in two parts. Your vocal cords buzz, giving pitch. Your mouth, tongue and lips then shape that buzz into a vowel or consonant.

The buzz decides whether you sound high or low. The shaping decides which sound it is. For recognising words, the shaping is what matters, and the buzz is a distraction.

MFCCs pull those two apart and keep the shaping.

How it works

  raw audio
      |
      v
  spectrogram                which pitches, moment by moment
      |
      v
  group into mel bands       many pitches -> about 40 bands, ear-shaped
      |
      v
  take the logarithm         because loudness is felt logarithmically
      |
      v
  one more transform         separates mouth-shape from vocal-cord buzz
      |
      v
  keep the first 13 numbers  a fingerprint of this instant of sound

You end up with roughly thirteen numbers for every ten milliseconds of audio. That is a hundred small fingerprints per second, instead of sixteen thousand raw measurements.

Where you have already used this

  • Old voice dialling on phones, before neural networks, ran almost entirely on MFCCs.
  • Shazam-style song matching uses related fingerprints from the same family of ideas.
  • Bird call identification apps classify a chirp from features like these.
  • Call-centre systems that detect a speaker's emotion start from features of this kind.

What is honestly hard here

MFCCs were designed in the 1980s, when computers were slow and data was scarce. Hand-designed compression was the only option.

Today the picture has changed. Large speech models often skip MFCCs. They read the mel spectrogram or the raw waveform, and learn their own compression. When you have enough data, learning beats hand-designing.

So MFCCs are no longer the default for large systems. They remain excellent when data is small or compute is tight. They also help when you need something you can inspect. Do not let anyone tell you they are obsolete, and do not assume they are state of the art.

Remember this

  • A feature is a small number describing a quality of the sound, standing in for thousands of raw measurements.
  • The mel scale groups pitches the way ears do: fine detail low down, coarse detail high up.
  • MFCCs separate the shape of your mouth from the pitch of your voice, and keep the mouth shape.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install numpy librosa

librosa is the standard Python audio-analysis library. It pulls in scipy and soundfile, totalling a few tens of megabytes. Everything below is CPU-only and generates its own audio, so nothing is downloaded at run time.

Start with the crudest features that work

Before MFCCs, understand what a feature is at all. Two of the simplest are loudness and how often the wave crosses zero.

crude_features.py
import numpy as np

SR = 16000
rng = np.random.default_rng(0)
t = np.arange(SR) / SR

sounds = {
    "low tone (100 Hz)":  0.5 * np.sin(2*np.pi*100*t),
    "high tone (4000 Hz)":0.5 * np.sin(2*np.pi*4000*t),
    "white noise":        0.5 * rng.standard_normal(SR),
    "near silence":       0.001 * rng.standard_normal(SR),
}

print(f"{'sound':<20} {'loudness (RMS)':>15} {'zero crossings/s':>18}")
for name, x in sounds.items():
    rms = float(np.sqrt(np.mean(x**2)))                       # average size of the wiggle
    zcr = int(np.sum(np.abs(np.diff(np.sign(x))) > 0))        # how often it flips sign
    print(f"{name:<20} {rms:15.4f} {zcr:18d}")
Output
sound                 loudness (RMS)   zero crossings/s
low tone (100 Hz)             0.3536                200
high tone (4000 Hz)           0.3536               8000
white noise                   0.4987               7984
near silence                  0.0010               8072

Two numbers, and already something useful. The 100 Hz tone crosses zero 200 times per second, exactly twice per cycle. The 4000 Hz tone crosses 8000 times. Zero-crossing rate is a cheap pitch estimate, and it costs one subtraction per sample.

Now look at the failure. White noise and the 4000 Hz tone have nearly identical zero-crossing rates, 7984 against 8000. These two features cannot tell a whistle from a hiss. That gap is what MFCCs fill.

Also note 0.3536 for both tones: that is 0.5 / sqrt(2), the RMS of any sine of amplitude 0.5. Loudness ignores pitch entirely, which is the point of having more than one feature.

The mel scale, made concrete

Before computing MFCCs, look at what "ear-shaped bands" actually means.

mel_bands.py
import numpy as np, librosa

SR, N_MELS = 16000, 12
centres = librosa.mel_frequencies(n_mels=N_MELS, fmin=0, fmax=SR/2)

print("centre frequency of each mel band, in Hz:")
print(np.round(centres).astype(int))
print("gap between neighbouring bands:")
print(np.round(np.diff(centres)).astype(int))
Output
centre frequency of each mel band, in Hz:
[   0  274  548  823 1105 1466 1945 2581 3425 4544 6029 8000]
gap between neighbouring bands:
[ 274  274  274  282  361  479  636  844 1119 1485 1971]

Read the second row. Down at the bottom, bands are 274 Hz wide. Up at the top, they are 1971 Hz wide, seven times coarser.

That is the mel scale doing its job. Detail is spent where your ear can use it. Below 1000 Hz the spacing is nearly linear; above it, nearly logarithmic.

MFCCs on two synthetic vowels

Real speech needs no download either. A vowel is a buzzing set of vocal cords passed through the resonances of your mouth, and both parts are a few lines of NumPy.

mfcc_vowels.py
import numpy as np, librosa

SR = 16000
t = np.arange(SR) / SR

def vowel(f0, formants, bw=90.0):
    """Source-filter vowel: buzzing vocal cords shaped by mouth resonances."""
    x = np.zeros_like(t)
    for k in range(1, int(SR/2 // f0)):                 # every harmonic of the buzz
        f = k * f0
        gain = sum(1.0 / (1.0 + ((f - fc)/bw)**2) for fc in formants)   # resonance peaks
        x += gain * np.sin(2*np.pi*f*t) / k
    return 0.5 * x / np.abs(x).max()

AH   = vowel(120, [730, 1090, 2440])    # the vowel in "father"
EE   = vowel(120, [270, 2290, 3010])    # the vowel in "see"
AH_Q = 0.25 * AH                        # same vowel, quarter the volume

def fingerprint(x):
    return librosa.feature.mfcc(y=x, sr=SR, n_mfcc=13,
                                n_fft=400, hop_length=160).mean(axis=1)

def cos(a, b):
    return float(a @ b / (np.linalg.norm(a) * np.linalg.norm(b)))

fa, fe, fq = fingerprint(AH), fingerprint(EE), fingerprint(AH_Q)
print("with coefficient 0 (which is loudness):")
print(f"   ah vs ee            {cos(fa, fe): .3f}")
print(f"   ah vs quiet ah      {cos(fa, fq): .3f}")
print("dropping coefficient 0:")
print(f"   ah vs ee            {cos(fa[1:], fe[1:]): .3f}")
print(f"   ah vs quiet ah      {cos(fa[1:], fq[1:]): .3f}")
Output
with coefficient 0 (which is loudness):
   ah vs ee             0.908
   ah vs quiet ah       0.988
dropping coefficient 0:
   ah vs ee             0.742
   ah vs quiet ah       1.000

This output is the whole lesson. Read the bottom two lines carefully.

The same vowel at a quarter of the volume scores 1.000 — a perfect match. The volume changed completely and the fingerprint did not move at all. That is exactly what you want, because "ah" said quietly is still "ah".

Two different vowels at identical volume score 0.742, far apart. The fingerprint responds to what the mouth was doing and ignores how loud it was.

Now look at the top two lines, where coefficient 0 is kept. The quiet "ah" scores only 0.988, and the two different vowels score a misleadingly high 0.908. Coefficient 0 is essentially total energy, and it drowns out the coefficients that carry the actual identity of the sound.

This is why production pipelines routinely drop MFCC coefficient 0, or replace it with a separately normalised energy term. Keep it by accident and your model spends its capacity learning about microphone gain.

Line by line, the parts that are not obvious

formants are the resonant frequencies of the vocal tract. The first two do most of the work of identifying a vowel, and the values used here are near the textbook averages for an adult male speaker. Changing f0 from 120 to 210 makes the same vowels in a higher voice, and the MFCCs barely move — try it.

n_fft=400, hop_length=160 at 16 kHz means 25 ms windows advancing 10 ms at a time. This exact pair is the near-universal convention in speech processing, and matching it is what lets features from different tools line up.

.mean(axis=1) collapses 101 per-frame fingerprints into one clip-level fingerprint. This works here because the sound is steady throughout. For real speech, averaging over a whole sentence destroys the word order, and you must keep the frames.

n_mfcc=13 is the traditional count. The higher coefficients describe increasingly fine ripples in the spectrum, which are mostly noise and speaker quirks.

Common mistakes

Keeping coefficient 0 without thinking. Demonstrated above. It encodes loudness, which usually depends on microphone distance rather than content.

Averaging over too long a window. mfcc.mean(axis=1) across a ten-second clip gives one blurred fingerprint of everything said. Fine for "which language is this", useless for "what words were spoken".

Normalising test data with training statistics computed per file. Cepstral mean normalisation, subtracting the per-utterance mean, is standard and helps enormously with channel differences. Apply it consistently in training and serving, or the two will disagree.

Reaching for MFCCs when a neural model is on the table. If you are training a CNN or a transformer on a decent amount of audio, feed it log-mel spectrograms with 40 to 80 bins instead. The final decorrelating transform in MFCC exists to help models that assume uncorrelated inputs, and neural networks do not need that help.

Ignoring deltas. Classic systems append the frame-to-frame change and the change of the change, tripling the feature count to 39. librosa.feature.delta(mfcc) computes them. They add the movement information that a single frame cannot carry.

Try it yourself

Add a third vowel, "oo", with formants [300, 870, 2240], and print the full three-by-three similarity matrix using coefficients 1 upward. Then change f0 to 210 for one vowel only and check whether the same-vowel similarity survives a change of voice pitch. Whether it does is the honest measure of whether the feature is doing its job.

What to learn next

Researcher — Mathematics and papers.

The MFCC pipeline, stated precisely

Given a frame $x[n]$ of length $N$ with window $w[n]$:

1. Pre-emphasis (optional, common in classical systems):

$$ x'[n] = x[n] - \alpha\, x[n-1], \qquad \alpha \approx 0.97 $$

A first-order high-pass that flattens the roughly $-6$ dB/octave spectral tilt of voiced speech, boosting the high-frequency formants that carry consonant identity.

2. Power spectrum:

$$ P[k] = \frac{1}{N}\left| \sum_{n=0}^{N-1} x'[n]\,w[n]\, e^{-i2\pi kn/N} \right|^2 $$

3. Mel filterbank. Map hertz to mel with the HTK formula:

$$ m(f) = 2595 \log_{10}!\left(1 + \frac{f}{700}\right), \qquad f(m) = 700\left(10^{m/2595} - 1\right) $$

Place $M$ triangular filters with centres uniformly spaced in mel between $f_{\min}$ and $f_{\max}$, each spanning its two neighbouring centres. Filter $j$ has response $H_j[k]$, and the band energy is:

$$ E_j = \sum_{k} H_j[k]\, P[k], \qquad j = 1, \dots, M $$

Typical $M$ is 26 to 40 for MFCC, and 80 for neural log-mel front-ends.

4. Log compression: $L_j = \log(E_j + \epsilon)$. This approximates loudness perception and, importantly, converts the multiplicative source-filter relationship into an additive one.

5. Discrete cosine transform, type II:

$$ c_i = \sum_{j=1}^{M} L_j \cos!\left[\frac{\pi i}{M}\left(j - \tfrac{1}{2}\right)\right], \qquad i = 0, \dots, M-1 $$

Keep $i = 0, \dots, 12$. Here $c_i$ is the $i$-th cepstral coefficient, $L_j$ the log mel band energy, and $M$ the number of filters.

Why the DCT is there

Speech production is well modelled as a source convolved with a filter: $s(t) = e(t) * h(t)$, where $e$ is the glottal excitation and $h$ the vocal tract impulse response. In the frequency domain this is a product, $|S(f)| = |E(f)| \cdot |H(f)|$, and the logarithm turns it into a sum:

$$ \log|S(f)| = \log|E(f)| + \log|H(f)| $$

The excitation term is a fast-varying ripple, with period equal to $f_0$ across frequency. The vocal tract term is a slowly varying envelope. Taking a further transform along the frequency axis separates them by rate of variation: low quefrency coefficients carry the envelope, high ones carry the pitch harmonics. Truncating to 13 coefficients is a low-pass operation in quefrency, keeping the vocal tract and discarding the pitch.

This is the cepstrum, named by Bogert, Healy and Tukey (1963) by reversing the first syllable of "spectrum" — hence the companion coinages quefrency, liftering and alanysis. The DCT stands in for the inverse Fourier transform because the log power spectrum is real and even, which makes the two equivalent up to scaling.

The DCT has a second, more practical role: it approximately decorrelates the filterbank energies. Adjacent mel bands overlap and are strongly correlated, and diagonal-covariance Gaussian mixture models — the acoustic model of choice from the 1980s until roughly 2012 — depend on that decorrelation to be viable. When the acoustic model became a neural network, the constraint vanished, and with it the main reason for step 5.

Deltas and liftering

Static coefficients ignore dynamics. The standard remedy appends regression estimates of the first and second time derivatives:

$$ d_t = \frac{\sum_{\theta=1}^{\Theta} \theta\,(c_{t+\theta} - c_{t-\theta})}{2\sum_{\theta=1}^{\Theta}\theta^2}, \qquad \Theta = 2 $$

Giving 13 static, 13 delta and 13 delta-delta coefficients, the 39-dimensional vector that dominated ASR for two decades.

Cepstral liftering rescales $c_i \leftarrow (1 + \frac{L}{2}\sin\frac{\pi i}{L}) c_i$ with $L = 22$, equalising the variance across coefficient index. It matters for GMM systems and is irrelevant for neural ones.

Cepstral mean and variance normalisation subtracts the per-utterance mean and divides by the per-utterance standard deviation. Because a linear channel is additive in the log-spectral domain, subtracting the mean removes a fixed channel response exactly. This remains one of the highest-value preprocessing steps in existence for mismatched-channel conditions.

Where MFCCs stand today

The empirical picture, stated without hedging:

  • For GMM-HMM systems, MFCC-39 was the correct choice, and alternatives measurably lost.
  • For neural acoustic models with more than a few hundred hours of data, log-mel filterbank energies (typically 80 bins, no DCT) consistently match or beat MFCCs. The DCT discards information that a network can exploit. Whisper, Conformer and wav2vec 2.0 fine-tuning recipes all use log-mel or raw waveform.
  • For low-resource, low-compute, or interpretable settings — keyword spotting on a microcontroller, a classifier over a few hundred labelled clips, forensic analysis needing an auditable pipeline — MFCCs remain a strong and defensible baseline.

Reporting an MFCC baseline alongside a log-mel one costs almost nothing and is good practice. Claiming MFCCs are obsolete is as wrong as claiming they are current best practice.

Alternatives worth knowing

PLP (Hermansky, 1990) applies equal-loudness pre-emphasis and intensity-loudness cube-root compression before an all-pole model, and is more robust than MFCC under some noise conditions.

Gammatone filterbank features replace triangular filters with gammatone filters fitted to auditory nerve responses.

Learned front-ends — SincNet, LEAF, and the convolutional encoder in wav2vec 2.0 — are covered in waveforms and spectrograms.

Self-supervised representations are the genuine successor. wav2vec 2.0, HuBERT and WavLM produce contextual frame embeddings that outperform every hand-designed feature on nearly every downstream task, at the cost of a large pretrained encoder in your pipeline.

Reading

  • Davis and Mermelstein (1980), Comparison of Parametric Representations for Monosyllabic Word Recognition, IEEE TASSP 28(4) — the MFCC paper.
  • Bogert, Healy and Tukey (1963), The Quefrency Alanysis of Time Series for Echoes — where the cepstrum comes from.
  • Hermansky (1990), Perceptual Linear Predictive Analysis of Speech, JASA 87(4).
  • Stevens, Volkmann and Newman (1937), A Scale for the Measurement of the Psychological Magnitude Pitch, JASA 8(3) — the original mel experiments.
  • Baevski et al. (2020), wav2vec 2.0 — arxiv.org/abs/2006.11477

What to learn next

What to learn next

These follow on from what you just read.

  • Speech and Audio AI

    Speech recognition

    Speech recognition turns spoken audio into written text, and its central difficulty is that nobody tells the model which sound belongs to which letter.

  • Speech and Audio AI

    Whisper

    Whisper is OpenAI's open-weight speech recognition model family, strong across languages and noisy audio, and it runs on a laptop CPU if you pick a small enough size.

  • Speech and Audio AI

    Text to speech

    Text to speech turns written words into audio, and the work splits into deciding how the words should sound and then generating the actual waveform.