Audio classification
Audio classification puts a label on a sound, and the hard part is not building the model but making it survive noise it never met during training.
- 15 min read
- 3 reading levels
- Updated
Read these first
On this page 8
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
Audio classification is putting a label on a sound.
Dog barking, glass breaking, siren, baby crying, cough. The model listens and picks one label from a list it knows.
The analogy you have already lived
You are in another room and you hear a sound from the kitchen. Instantly you know: that was the pressure cooker whistle. Not the kettle, not the doorbell.
You did not see anything. You did not think it through. One second of sound was enough. You have heard that whistle hundreds of times, and nothing else in the house sounds like it.
That is the whole task. A machine hears a short clip and names it.
Why it is different from speech recognition
Speech recognition produces a sentence, and the order of things matters enormously. "Dog bites man" and "man bites dog" use identical words.
Classification produces one label. Order matters much less. A dog barking at the start of a clip and a dog barking at the end are both "dog".
That makes it an easier problem in shape. It is why audio classification works well on small models and cheap hardware.
How it works
1 second of audio
|
v
[ spectrogram ] a picture of pitch over time
|
v
[ model ] reads the picture, same as an image classifier
|
v
"glass breaking" (87% sure)That second box is worth pausing on. Once a sound is a spectrogram, it is a picture. The whole toolkit built for image classification applies directly. A convolutional neural network trained on sound pictures works the same way as one trained on photographs.
This is why audio classification improved so quickly. It inherited a decade of computer vision work almost unchanged.
Where you have already seen it
- Bird-song apps that name a bird from a recording of its call.
- Smart home alerts for a smoke alarm or glass breaking while you are out.
- Baby monitors that tell crying apart from other noise.
- Content moderation flagging gunshots or explosions in uploaded video.
- Machine monitoring in factories, hearing a bearing that has started to fail.
- Health screening research on cough sounds.
What is honestly hard here
Here is the thing that catches almost everyone, and it is worth saying plainly.
A model can score perfectly in your test and be useless in the real world.
The reason is that lab recordings are clean, and life is not. Train on clear recordings of dogs and sirens, test on clear recordings, and you get an excellent number. Now deploy it on a phone in a busy street. It can fall all the way to guessing at random.
Not a little worse. All the way to random. You will see that happen, with real numbers, in the developer section.
The fix is not a bigger model. The fix is to add noise, echo and distortion to your training audio on purpose. The model then meets bad conditions before your users do. That is called augmentation, and it usually matters more than the choice of model.
Two other honest limits. Sounds overlap constantly in real life, as when a dog barks while a car passes. A model forced to pick exactly one label handles that badly. And some sounds are genuinely ambiguous: a door slam, a book dropped and a distant firework can look nearly identical.
Remember this
- Audio classification turns a sound into a spectrogram, then treats it as a picture.
- A perfect score on clean test audio predicts almost nothing about performance in a noisy room.
- Augmentation — training on deliberately degraded audio — is usually the highest-value thing you can do.
What to learn next
- Noise reduction — cleaning the audio instead of training around the noise.
- Wake-word detection — classification with a brutal false-alarm budget.
- Data augmentation — the same idea across every kind of data.
Developer — Code and libraries.
Setup
pip install numpy librosa scikit-learnCPU only, no downloads, and the whole script finishes in well under a minute. We generate four kinds of sound, extract features, train a classifier, and then run the experiment that actually matters.
The dataset, generated in twelve lines
Four classes with genuinely different structure: a steady tone, a rising tone, broadband noise, and short impacts.
import numpy as np, librosa
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import accuracy_score
SR, DUR = 16000, 1.0
n = int(SR*DUR); t = np.arange(n)/SR
def make_clip(kind, rng):
if kind == "whistle": # a steady tone
return np.sin(2*np.pi*rng.uniform(700,1600)*t)
if kind == "siren": # a tone that slides upward
f0, f1 = rng.uniform(400,600), rng.uniform(1200,1800)
return np.sin(2*np.pi*np.cumsum(f0 + (f1-f0)*t/DUR)/SR)
if kind == "hiss": # energy at every pitch
return rng.standard_normal(n)
x = np.zeros(n) # knock: a few sharp taps
for _ in range(rng.integers(3,7)):
s = rng.integers(0, n-1600)
x[s:s+1600] += np.exp(-np.arange(1600)/120)*rng.standard_normal(1600)
return x
CLASSES = ["hiss","knock","siren","whistle"]
def add_noise(x, snr_db, rng):
"""Mix in background noise at a chosen signal-to-noise ratio."""
if snr_db is None: return x
sig = np.mean(x**2)
noise = rng.standard_normal(x.size)
noise *= np.sqrt(sig/(10**(snr_db/10)) / np.mean(noise**2))
return x + noise
def features(x):
x = x/(np.abs(x).max()+1e-9)
mel = librosa.feature.melspectrogram(y=x, sr=SR, n_fft=512, hop_length=256, n_mels=32)
log = librosa.power_to_db(mel)
return np.concatenate([log.mean(axis=1), log.std(axis=1)]) # 64 numbers per clip
def build(n_each, snr_choices, seed):
rng = np.random.default_rng(seed)
X, y = [], []
for kind in CLASSES:
for _ in range(n_each):
snr = snr_choices[rng.integers(len(snr_choices))]
X.append(features(add_noise(make_clip(kind, rng), snr, rng)))
y.append(kind)
return np.array(X), np.array(y)
# Model A: trained only on clean audio, the way most first projects are built.
Xtr_clean, ytr = build(60, [None], seed=0)
clf_clean = LogisticRegression(max_iter=3000, random_state=0).fit(Xtr_clean, ytr)
# Model B: identical model, trained on the same sounds at five noise levels.
Xtr_aug, ytr_aug = build(60, [None, 20, 10, 5, 0], seed=1)
clf_aug = LogisticRegression(max_iter=3000, random_state=0).fit(Xtr_aug, ytr_aug)
print("accuracy on 200 unseen clips at each noise level")
print(f"{'test SNR':>10}{'trained clean':>15}{'trained with noise':>21}")
for snr in [None, 20, 10, 5, 0, -5]:
Xte, yte = build(50, [snr], seed=100)
a = accuracy_score(yte, clf_clean.predict(Xte))
b = accuracy_score(yte, clf_aug.predict(Xte))
label = "clean" if snr is None else f"{snr} dB"
print(f"{label:>10}{a:15.3f}{b:21.3f}")accuracy on 200 unseen clips at each noise level
test SNR trained clean trained with noise
clean 1.000 1.000
20 dB 0.315 1.000
10 dB 0.250 1.000
5 dB 0.250 1.000
0 dB 0.250 1.000
-5 dB 0.250 0.830This table is the entire lesson
Model A scored a perfect 1.000 on clean audio. In a first project that number goes in the report, and the work looks finished.
Then look at the next row. At 20 dB signal-to-noise — noise a hundred times quieter than the signal, the kind of background hum you stop noticing in a room — accuracy fell to 0.315.
By 10 dB it is at 0.250. With four classes, 0.250 is exactly what you get by guessing. The model has not degraded. It has stopped working entirely, while remaining completely confident.
Model B is the same logistic regression, on the same sounds. The only difference is that its training clips were mixed with noise at five levels. It holds 1.000 all the way down to 0 dB, where the noise is as loud as the signal, and still reaches 0.830 at −5 dB — a level it never saw during training.
One change to the data. No change to the model. That is the ordering of priorities in applied audio work.
Why normalisation does not rescue Model A
The natural first guess is that noise only shifts the overall level, so subtracting a per-clip mean should undo it. It does not. Adding that step gives 0.290 at 20 dB instead of 0.315 — no better.
The reason is worth understanding. Noise does not offset the spectrum, it fills in the quiet parts of it. A whistle is a single loud band with near-silence everywhere else, and that silence is most of what identifies it. Add broadband noise and the silence disappears, so the pattern the classifier learned genuinely is not present any more.
No normalisation recovers information that was covered up. Only having seen covered-up examples during training helps.
Line by line, the parts that are not obvious
log.mean(axis=1) and log.std(axis=1) collapse time away, giving 32 average band energies plus 32 measures of how much each band fluctuated. The standard deviation is what separates knock from hiss: both have broad spectra, and only one is bursty. Dropping it costs a lot of accuracy — try it.
np.cumsum(...) for the siren accumulates phase so the frequency slide is continuous. Writing np.sin(2*np.pi*f*t) with a varying f produces a different, discontinuous signal, as explained in waveforms and spectrograms.
The SNR formula. Signal-to-noise ratio in decibels is 10 * log10(signal_power / noise_power). Scaling the noise by sqrt(sig / 10**(snr/10) / noise_power) sets that ratio exactly. Getting this right matters, because "we tested at 10 dB" is meaningless if your definition of 10 dB differs from everyone else's.
Separate seeds for train and test. seed=100 for evaluation guarantees the test clips were never in training. With generated data it is easy to leak by accident, and a leak looks exactly like success.
Common mistakes
Reporting clean-set accuracy as the headline. Demonstrated above. Always report accuracy across a noise sweep, and quote the worst realistic condition rather than the best.
Splitting randomly when clips share a source. If ten clips come from one recording session, a random split puts some in training and some in testing, and the model learns the room rather than the sound. Split by recording session, by device, or by speaker — never by clip.
Ignoring class imbalance. Real sound datasets are wildly imbalanced: thousands of hours of speech and eleven minutes of glass breaking. Accuracy is the wrong metric there. Use per-class recall and average precision. See imbalanced data.
Forcing one label when sounds overlap. A dog barking during a siren is both. That is multi-label classification: use a sigmoid per class with binary cross-entropy, not a softmax. Choosing softmax by habit makes overlapping sounds unlearnable.
Augmenting the test set with the same noise used in training. Then you have measured nothing. Hold out noise types too — train on one set of background recordings, test on another.
Try it yourself
Delete log.std(axis=1) from features, leaving only the means, and rerun. Watch which classes collapse into each other and work out why. Then add a fifth noise level of -10 to Model B's training mix and see whether the −5 dB result improves. It should — and how much it improves tells you whether you are short of data or short of model capacity.
Scaling up beyond this
When you outgrow logistic regression, the standard path is a small CNN over log-mel spectrograms, then a pretrained audio model. PANNs, AST (audio spectrogram transformer) and BEATs are all pretrained on AudioSet and fine-tune well on a few hundred labelled clips. CLAP gives audio and text a shared embedding space, so you can classify by describing a sound in words with no labelled examples at all.
All of them are downloads in the hundreds of megabytes, and all of them are unnecessary until the experiment above stops being your bottleneck.
What to learn next
- Noise reduction — cleaning the audio instead of training around the noise.
- Wake-word detection — classification with a brutal false-alarm budget.
- Data augmentation — the same idea across every kind of data.
Researcher — Mathematics and papers.
Task formulations
Clip-level tagging assigns labels to a whole recording: $f: \mathbb{R}^{N} \to {0,1}^{C}$. Sound event detection additionally localises onsets and offsets, producing $(c, t_{\text{on}}, t_{\text{off}})$ triples. Acoustic scene classification labels the environment rather than any event, and is single-label by construction.
Polyphony makes multi-label the default. With $C$ classes and independent presence, the target is a binary vector and the loss is binary cross-entropy over sigmoids:
$$ \mathcal{L} = -\frac{1}{C}\sum_{c=1}^{C}\left[y_c \log \hat{y}_c + (1-y_c)\log(1-\hat{y}_c)\right] $$
A softmax over classes imposes $\sum_c \hat{y}_c = 1$, which is a factually wrong constraint whenever two sounds co-occur. This is the most common architectural error in the field.
Datasets and what they encode
AudioSet (Gemmeke et al., 2017) is the field-defining resource: roughly 2 million 10-second YouTube clips, 527 classes in an ontology, weakly labelled at clip level with no temporal boundaries. Labels are noisy, the distribution is severely long-tailed, and clips are distributed as YouTube IDs, so the set decays as videos are removed — an underappreciated reproducibility problem.
ESC-50 (2000 clips, 50 classes) and UrbanSound8K are small benchmarks with predefined folds. Both ship fold assignments specifically because clips share source recordings; ignoring the folds and splitting randomly inflates accuracy substantially and is a recurring flaw in published results.
DCASE challenges provide the community's evaluation infrastructure, including cross-device and cross-city conditions that measure exactly the generalisation gap the developer block demonstrates.
Weak labels and multiple-instance learning
AudioSet-style supervision gives a clip label with no timing. The standard treatment is multiple-instance learning: the clip is a bag of frame instances, positive if any instance is positive. A pooling function maps frame-level scores $y_c(t)$ to a clip score:
$$ \hat{y}_c = \sum_t w_c(t)\, y_c(t), \qquad w_c(t) = \frac{\exp(\lambda\, y_c(t))}{\sum_{t'} \exp(\lambda\, y_c(t'))} $$
Max pooling ($\lambda \to \infty$) localises sharply and trains unstably, since gradient reaches one frame. Mean pooling ($\lambda = 0$) trains stably and smears localisation. Learned attention pooling interpolates and is the usual choice. McFee, Salamon and Bello (2018) analyse this trade in detail.
Architectures
PANNs (Kong et al., 2020) established the pretrain-on-AudioSet-then-transfer recipe with VGG-style CNNs plus a Wavegram-Logmel hybrid, reporting 0.439 mAP on AudioSet tagging.
AST (Gong, Chung and Glass, 2021) applies a ViT directly to spectrogram patches, initialised from ImageNet weights — cross-modal transfer from pictures to sound, which works better than it has any right to and confirms how much of the gain is spectrogram-as-image.
BEATs (Chen et al., 2022) adds self-supervised pretraining with an acoustic tokeniser, iteratively refined against the classifier.
CLAP (Elizalde et al., 2022; Wu et al., 2023) contrastively aligns audio and text encoders, enabling zero-shot classification by embedding class descriptions. It is the audio analogue of CLIP, and it is the practical answer when you have classes but no labelled examples.
Augmentation, which is where the gains are
The developer block shows the effect at toy scale. The production toolkit:
- Noise mixing at controlled SNR, using real background recordings (MUSAN, FSD50K, DEMAND) rather than Gaussian noise. Gaussian noise is spectrally flat and unrepresentative; models overfit to it.
- Room impulse response convolution for reverberation. Far-field performance without this is poor and no amount of clean data substitutes.
- SpecAugment (Park et al., 2019): mask contiguous bands of time and frequency in the spectrogram. Extremely cheap, applied on the fly, and among the highest value-per-line changes available.
- Mixup (Zhang et al., 2018): train on convex combinations of pairs, $\tilde{x} = \lambda x_i + (1-\lambda)x_j$ with matching label interpolation. It is unusually well-suited to audio, because mixing two waveforms is physically what happens when two sounds occur together.
- Time and pitch shifting, applied independently so the model does not couple them.
Evaluation
For multi-label tagging, report mAP (mean average precision) — the mean over classes of area under the precision-recall curve. It is threshold-free and handles imbalance sensibly, which accuracy and ROC-AUC do not. d-prime is common in the AudioSet literature and is a monotone transform of ROC-AUC.
For sound event detection, event-based F1 with a collar tolerance, and segment-based F1 at a fixed resolution, measure different things and both should be reported. The sed_eval toolkit (Mesaros et al., 2016) is the reference implementation, and hand-rolled event matching is a reliable source of incomparable numbers.
Report per-class metrics. A macro-averaged mAP of 0.45 on a long-tailed set typically hides near-zero performance on the rarest half of the classes, which are frequently the classes anyone actually cares about detecting.
Reading
- Gemmeke et al. (2017), AudioSet: An ontology and human-labeled dataset for audio events, ICASSP.
- Kong et al. (2020), PANNs — arxiv.org/abs/1912.10211
- Gong, Chung and Glass (2021), AST: Audio Spectrogram Transformer — arxiv.org/abs/2104.01778
- Park et al. (2019), SpecAugment — arxiv.org/abs/1904.08779
- Wu et al. (2023), Large-Scale Contrastive Language-Audio Pretraining — arxiv.org/abs/2211.06687
- Mesaros, Heittola and Virtanen (2016), Metrics for Polyphonic Sound Event Detection, Applied Sciences 6(6).
What to learn next
- Noise reduction — cleaning the audio instead of training around the noise.
- Wake-word detection — classification with a brutal false-alarm budget.
- Data augmentation — the same idea across every kind of data.