Speaker identification
Speaker identification works out who is talking rather than what they said, by turning each voice into a fingerprint vector and comparing distances between fingerprints.
- 16 min read
- 3 reading levels
- Updated
Read these first
On this page 10
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
Speaker identification works out who is speaking, not what they said.
It is a completely different question from speech recognition, and it uses different machinery.
The analogy you have already lived
Your phone rings from an unknown number. The person says one word — "hello" — and you already know it is your mother.
You did not analyse anything. One syllable was enough. It works when she has a cold. It works on a terrible line. It works on a word you never heard her say.
That is because a voice carries two separate things at once: the words, and the person. Your ear reads both. Speaker identification is a machine reading the second one and ignoring the first.
Why a voice is personal in the first place
Your voice depends on your body.
The pitch comes from how fast your vocal cords vibrate, which depends on their length and thickness. The character comes from the shape of your throat, mouth and nose. These act as resonating chambers, the way a large drum sounds deeper than a small one.
Nobody else has quite your dimensions. Add your habits — your speed, your rhythm, the way you end sentences — and a voice becomes recognisable.
Three different jobs, often confused
Verification — "is this the person they claim to be?" One voice against one stored voice. Answer: yes or no. This is a bank asking you to say a passphrase.
Identification — "which of my known people is this?" One voice against a list. Answer: a name, or nobody.
Diarisation — "who spoke when?" Take a recording of a meeting and label the segments by speaker. Answer: a timeline. This is what turns a meeting recording into a readable transcript.
How it works
audio of a voice
|
v
[ features per short slice ]
|
v
[ neural network ]
|
v
a fingerprint: a list of a few hundred numbers
|
v
compare with a stored fingerprint
|
v
similar enough? -> yes / noThe fingerprint is called an embedding — a list of numbers positioned so that similar things sit close together. It is the same idea as the word embeddings in embeddings, applied to voices instead of words.
The model is trained on one simple rule. Two clips of the same person should land close together. Two clips of different people should land far apart. Nothing else is specified. Where each voice ends up is learned.
The part that makes it useful
Once voices are points in space, everything becomes distance.
Verification is "are these two points close?" Identification is "which stored point is closest?" Diarisation is "group these points into clusters, one per speaker".
And critically, a new person needs no retraining. Record them once, store their point, done. The model never learned that specific person and still recognises them. It learned what makes voices different, not a fixed list of people.
Where you have already seen it
- Bank phone lines saying "my voice is my password".
- Smart speakers answering "who am I?" and giving different people different calendars.
- Meeting transcripts labelled Speaker 1 and Speaker 2.
- Podcast editors that separate the host from the guest automatically.
What is honestly hard here
Your voice is not a secret. You use it in public constantly. Anyone can record it. That makes it a poor password: unlike a password, you cannot change it after it leaks.
It can be copied. Systems now exist that imitate a voice from a few seconds of audio. They can score high against a naive verifier. Voice alone should never be the only thing standing between a stranger and your money. That whole topic gets its own lesson: voice cloning and its risks.
Accuracy is not equal for everyone. Performance depends on how many voices like yours were in the training data. Audits have repeatedly found differences across accent, age and gender.
Conditions matter enormously. Enrol on a laptop, verify on a phone in a bus, and scores drop hard. A cold changes your voice. So does tiredness.
Remember this
- Speaker identification asks who, and speech recognition asks what.
- A voice becomes an embedding — a point in space — and everything after that is distance.
- A voice is public and copyable, so treat it as a convenience, not as a password.
What to learn next
- Voice cloning and its risks — why a voice fingerprint is not a lock.
- Audio classification — the same embedding idea applied to sounds instead of people.
- Embeddings — the general idea of turning things into comparable points.
Developer — Code and libraries.
Setup
pip install numpy scipy librosaNo downloads, no GPU. Rather than fetch a trained speaker model, we build three synthetic speakers with different physical vocal tracts, then run the real verification pipeline over them. Every number below is reproducible on your machine.
The point is not that MFCC averages make a good speaker model — they do not. The point is that the pipeline is short, and one specific step in it does far more work than beginners expect.
Three voices and one passphrase
Each synthetic speaker gets a pitch, a throat-length scale, and a breathiness. All three say the same passphrase, which mirrors real text-dependent verification: "say your passphrase to log in".
import numpy as np, librosa
from scipy.signal import lfilter
SR = 16000
rng = np.random.default_rng(0)
def resonator(x, fc, bw):
r = np.exp(-np.pi*bw/SR); th = 2*np.pi*fc/SR
return lfilter([1-2*r*np.cos(th)+r*r], [1.0, -2*r*np.cos(th), r*r], x)
PASSPHRASE = [(730,1090,2440), (270,2290,3010), (300,870,2240), (530,1840,2480)]
def utterance(f0, scale, bw, seconds=1.6):
"""The same passphrase, in a voice with pitch f0, throat length scale, breathiness bw."""
n = int(SR*seconds/len(PASSPHRASE)); out = []
for formants in PASSPHRASE:
f0v = f0 * (1 + 0.04*rng.standard_normal()) # nobody repeats a pitch exactly
src = np.zeros(n); src[::max(1, int(SR/f0v))] = 1.0
x = src
for fc in formants:
x = resonator(x, fc*scale, bw)
out.append(x / (np.abs(x).max() + 1e-9))
return np.concatenate(out)
SPEAKERS = {"Asha": (215, 1.18, 60), "Bilal": (110, 0.88, 130), "Chen": (150, 1.00, 55)}
def raw_embed(x):
m = librosa.feature.mfcc(y=x, sr=SR, n_mfcc=20, n_fft=400, hop_length=160)
return m[1:].mean(axis=1) # drop coefficient 0: it is loudness
enrol = {n: np.mean([raw_embed(utterance(*p)) for _ in range(3)], axis=0)
for n, p in SPEAKERS.items()} # three enrolment clips each
test = {n: raw_embed(utterance(*p)) for n, p in SPEAKERS.items()} # one unseen clip each
mu = np.mean(list(enrol.values()), axis=0) # the "average voice" of the population
def fp(v, centred):
v = v - mu if centred else v
return v / np.linalg.norm(v)
for label, centred in (("raw cosine", False), ("after subtracting the average voice", True)):
print(label + " (rows = enrolled, columns = new recording)")
print(" " + "".join(f"{n:>8}" for n in SPEAKERS))
scores = {}
for a in SPEAKERS:
row = [float(fp(enrol[a], centred) @ fp(test[b], centred)) for b in SPEAKERS]
scores[a] = row
print(f"{a:>7} " + "".join(f"{x:8.3f}" for x in row))
same = min(scores[a][i] for i, a in enumerate(SPEAKERS))
diff = max(scores[a][j] for i, a in enumerate(SPEAKERS) for j in range(3) if j != i)
print(f" worst genuine match {same:.3f} | best impostor {diff:.3f} | margin {same-diff:.3f}\n")raw cosine (rows = enrolled, columns = new recording)
Asha Bilal Chen
Asha 0.999 0.932 0.964
Bilal 0.928 1.000 0.985
Chen 0.961 0.988 0.998
worst genuine match 0.998 | best impostor 0.988 | margin 0.010
after subtracting the average voice (rows = enrolled, columns = new recording)
Asha Bilal Chen
Asha 0.996 -0.965 0.052
Bilal -0.963 0.998 -0.235
Chen -0.421 0.157 0.676
worst genuine match 0.676 | best impostor 0.157 | margin 0.518Read the margins, not the diagonals
Both tables get every answer right — the largest score in each row is the correct speaker. A careless write-up would report "100 percent accuracy" for both and move on.
Look at the margins instead.
Raw cosine: margin 0.010. The genuine matches sit at 0.998 and the best impostor at 0.988. To accept genuine users you would set a threshold below 0.998; that same threshold accepts an impostor at 0.988. A threshold of 0.9 — which sounds strict — accepts every impostor in the table. The system is one bad recording away from letting the wrong person in.
After centring: margin 0.518. Genuine matches sit above 0.676 and impostors below 0.157. Any threshold between them works, and there is room for a noisy recording to move a score without changing the decision.
One line of code produced that difference. Subtracting the population mean removes everything the three voices have in common — the passphrase itself, the channel, the general shape of speech — leaving only what differs between speakers. Real systems do this and more: centring, then length normalisation, then a learned projection.
The general lesson transfers well beyond audio. A similarity score is meaningless without knowing the impostor distribution. Report the gap, never the raw number.
Line by line, the parts that are not obvious
m[1:] drops MFCC coefficient 0, which is loudness. Keep it and the model learns how far each speaker sat from the microphone, which does not generalise. This was demonstrated in audio features.
Three enrolment clips, averaged. One clip captures one moment. Averaging several suppresses the day-to-day variation and is standard practice — most commercial systems ask for three to five enrolment phrases for exactly this reason.
rng.standard_normal() on the pitch makes each utterance slightly different, so the test clip is never identical to an enrolment clip. Without that, the diagonal would be exactly 1.0 and the demonstration would prove nothing.
Cosine similarity, not Euclidean distance. After length normalisation the two are monotonically related, and cosine is the convention in every speaker-recognition toolkit.
Choosing a threshold, and what it costs
A verification system has one dial and two ways of being wrong:
- False accept — a stranger gets in. Lower the threshold and this rises.
- False reject — the real user is locked out. Raise the threshold and this rises.
Moving the dial trades one for the other; it never removes both. The point where the two rates are equal is the equal error rate, or EER, the standard headline metric.
EER is the wrong number to design with. The costs are not symmetric. Locking a customer out of their own bank account is annoying; letting a stranger in is fraud. Pick the threshold from the cost of each error in your setting, then report both rates at that threshold. Quoting EER alone hides the decision that actually matters.
For real work, use a trained model
The MFCC average above is a teaching device. Production systems use a network trained on thousands of speakers to produce embeddings directly. The usual open choices are ECAPA-TDNN and x-vector, both available through SpeechBrain or NeMo, at a few tens of megabytes for ECAPA.
The pipeline shape does not change. Embed, normalise, compare, threshold. You are swapping raw_embed for a stronger function and keeping everything else.
Common mistakes
Using voice as a sole authentication factor. It is public, recordable and now cloneable. Treat it as one signal among several, never as the lock itself.
Enrolling and testing on different equipment. A model that has not been trained across channels drops sharply when enrolment is on a laptop and verification is on a phone. Enrol on the device people will actually use.
Reporting one aggregate accuracy figure. Break results down by accent, age and gender. An overall EER of 2 percent can hide 8 percent for one group. See fairness metrics.
Ignoring how much speech you have. Below about three seconds, scores get noticeably unreliable. Measure accuracy against clip duration and set a minimum.
Confusing diarisation with identification. Diarisation clusters unknown voices and hands you "Speaker 1" and "Speaker 2". Attaching real names needs enrolled fingerprints, which is a separate step.
Try it yourself
Add a fourth speaker whose parameters sit between two existing ones, say (180, 1.09, 58), and watch the margin shrink. Then set bw identical for all speakers and rerun. How much of the separation came from breathiness alone? That is a small experiment in which physical cue a feature is actually using.
What to learn next
- Voice cloning and its risks — why a voice fingerprint is not a lock.
- Audio classification — the same embedding idea applied to sounds instead of people.
- Embeddings — the general idea of turning things into comparable points.
Researcher — Mathematics and papers.
Problem definitions
Verification is a hypothesis test. Given enrolment $\mathcal{E}$ and test $\mathcal{T}$, decide between $H_0$ (same speaker) and $H_1$ (different), by comparing a log-likelihood ratio against a threshold:
$$ s = \log \frac{p(\mathcal{E}, \mathcal{T} \mid H_0)}{p(\mathcal{E}, \mathcal{T} \mid H_1)} \gtrless \tau $$
Identification is closed-set or open-set classification over $N$ enrolled speakers. Diarisation is unsupervised segmentation plus clustering, evaluated by diarisation error rate — the sum of missed speech, false alarm and speaker confusion time, divided by total speech.
Four generations
GMM-UBM (Reynolds et al., 2000). Train a universal background model on pooled speech, then MAP-adapt the means per speaker. Scoring is a log-likelihood ratio against the UBM.
Supervector and i-vector (Dehak et al., 2011). Stack the adapted GMM means into a supervector $M$ and posit a low-rank total variability subspace:
$$ M = m + T w $$
Where $m$ is the UBM supervector, $T \in \mathbb{R}^{CF \times R}$ the total variability matrix, and $w \in \mathbb{R}^{R}$ (typically $R = 400$–600) the i-vector. Unlike earlier joint factor analysis, $T$ deliberately mixes speaker and channel variability, deferring their separation to the backend. This is the first true embedding in speaker recognition, and its architecture — a fixed-length vector plus a probabilistic backend — still describes the field.
x-vector (Snyder et al., 2018). A time-delay neural network over frames, a statistics pooling layer taking the mean and standard deviation across time, then affine layers whose penultimate activation is the embedding. Pooling is the crucial piece: it maps variable-length input to a fixed vector, and the standard deviation term carries genuinely useful information beyond the mean. Trained with speaker-classification cross-entropy and heavy augmentation (MUSAN noise, RIR reverberation), x-vectors decisively beat i-vectors given enough data.
ECAPA-TDNN (Desplanques et al., 2020). Adds squeeze-excitation to the Res2Net-style blocks, aggregates features across multiple layers, and replaces plain pooling with attentive statistics pooling, so frames contribute according to learned importance. It is the standard strong open baseline, reporting roughly 0.9 percent EER on VoxCeleb1 test in the original paper.
Margin losses
Plain softmax separates training classes without making embeddings metrically useful at test time on unseen speakers. Angular margin losses fix this directly. AAM-softmax (ArcFace, Deng et al., 2019) normalises weights and features and inserts an additive angular margin:
$$ \mathcal{L} = -\frac{1}{N}\sum_{i=1}^{N} \log \frac{e^{s\cos(\theta_{y_i} + m)}} {e^{s\cos(\theta_{y_i} + m)} + \sum_{j \neq y_i} e^{s\cos\theta_j}} $$
Where $\theta_j$ is the angle between embedding $i$ and class weight $j$, $m$ the angular margin (typically 0.2), $s$ the scale (typically 30), and $y_i$ the true speaker. The margin forces intra-class compactness and inter-class separation on the hypersphere, which is exactly the geometry cosine scoring assumes at test time. This is the single largest architecture-independent gain of the last several years.
Backends
Cosine scoring after centring and length normalisation is the standard for margin-trained embeddings — precisely the two operations demonstrated in the developer block.
PLDA (Prince and Elder, 2007; Ioffe, 2006) models an embedding as $\phi_{ij} = \mu + F h_i + \epsilon_{ij}$, where $h_i$ is a speaker factor and $\epsilon_{ij}$ within-speaker noise, and computes an exact likelihood ratio. PLDA dominated the i-vector era. With margin-trained embeddings its advantage largely disappears, and cosine scoring is both simpler and competitive.
Adaptive score normalisation (AS-norm) standardises a score against its distribution over an imposter cohort. It typically yields a 10 to 20 percent relative EER reduction for a few lines of code and is the highest-value cheap addition to a deployed system.
Evaluation
- EER — the operating point where false accept equals false reject. A summary statistic, not a design target.
- minDCF — minimum detection cost, $C_{\text{det}} = C_{\text{miss}} P_{\text{target}} P_{\text{miss}} + C_{\text{fa}} (1 - P_{\text{target}}) P_{\text{fa}}$. NIST SRE conventionally uses $P_{\text{target}} = 0.01$ or $0.05$. This is the metric that reflects asymmetric costs.
- DET curves — false reject against false accept on a normal deviate scale, which makes the low-error region readable.
VoxCeleb 1 and 2 (Nagrani et al., 2017, 2018) are the standard public benchmarks, built from YouTube interview audio. Note the domain: celebrity interviews, largely English, in reasonable acoustic conditions. Telephone, far-field and multilingual performance are substantially worse, and VoxCeleb numbers do not transfer.
Spoofing, and why this section is not optional
Speaker verification systems are attacked by replay, by voice conversion and by synthesis. The ASVspoof challenge series (2015 onward) benchmarks countermeasures, and the ASVspoof 2019 and 2021 logical-access conditions specifically target neural TTS and voice conversion.
Two findings matter for anyone deploying this. First, a verification system without a countermeasure accepts good synthetic speech at high rates — the embedding captures vocal-tract characteristics, which synthesis reproduces. Second, countermeasures generalise poorly to attack types absent from their training data, which is the usual state of affairs against an adversary who reads the literature.
The current framing is SASV — spoofing-aware speaker verification — treating verification and countermeasure as one jointly optimised system rather than two stacked filters. Treat published EERs on clean benchmarks as an upper bound on what you will see against an adversary. See adversarial attacks.
Reading
- Reynolds, Quatieri and Dunn (2000), Speaker Verification Using Adapted Gaussian Mixture Models, Digital Signal Processing 10.
- Dehak et al. (2011), Front-End Factor Analysis for Speaker Verification, IEEE TASLP 19(4).
- Snyder et al. (2018), X-vectors: Robust DNN Embeddings for Speaker Recognition, ICASSP.
- Desplanques et al. (2020), ECAPA-TDNN — arxiv.org/abs/2005.07143
- Deng et al. (2019), ArcFace — arxiv.org/abs/1801.07698
- Nagrani et al. (2017), VoxCeleb — arxiv.org/abs/1706.08612
- Wu et al. (2015 onward), ASVspoof challenge series — asvspoof.org
What to learn next
- Voice cloning and its risks — why a voice fingerprint is not a lock.
- Audio classification — the same embedding idea applied to sounds instead of people.
- Embeddings — the general idea of turning things into comparable points.