Speech and Audio AI

Voice cloning and its risks

Voice cloning copies a person's voice from seconds of audio, which helps people who are losing their speech and also breaks voice as a way of proving who you are.

On this page 9
  1. The short answer
  2. The analogy you have already lived
  3. Start with what it is genuinely good for
  4. Now the honest part
  5. Keeping perspective
  6. What actually protects you
  7. What consent means here
  8. Remember this
  9. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

Voice cloning makes a computer speak in a specific person's voice, learned from a short recording of them.

Early systems needed hours of studio audio. Current ones work from a few seconds.

The analogy you have already lived

You have signed your name thousands of times. Your signature is yours, and for a long time it was proof.

Now imagine a machine that watches you sign once and reproduces your signature perfectly, forever, for anyone who asks.

Your signature has not changed. What changed is what it proves. It still identifies you, and it no longer confirms that you were there.

That is what happened to the human voice.

Start with what it is genuinely good for

This technology exists because people wanted it, and some of those reasons are among the best things AI does.

Voice banking. Someone diagnosed with motor neurone disease records their voice while they still can. When speech goes, their communication device speaks in their voice — not a stranger's. People describe this as keeping part of themselves.

Voices for people who never had one. Children with conditions affecting speech can be given a synthetic voice of their own. It is built to match their age, region and family.

Translation that keeps the speaker. A lecture delivered in one language, heard in another, still sounding like the person who gave it.

Dubbing and audiobooks. An author narrating their own book without spending three weeks in a studio.

Recovering a recording. Repairing a few words in a documentary instead of recalling everyone.

These are real, and they are why a blanket "this technology is bad" is the wrong response.

Now the honest part

The same capability breaks something we relied on without noticing.

Recognising a voice used to be evidence. Your mother's voice on the phone meant your mother. That is no longer safe to assume, and most people have not updated.

Three specific harms, described plainly.

Family emergency scams. A caller sounds exactly like your son or grandson. They are in trouble, they need money now, please do not tell anyone. The voice is the whole attack — it removes your ability to doubt. Losses from this pattern are substantial and it targets older people hardest.

Business fraud. A finance employee gets a call from someone who sounds like their director, urgently authorising a transfer. Large sums have been lost this way.

Non-consensual use of someone's voice. Making a real person appear to say things they never said. Political statements, endorsements, sexual content, harassment. The harm lands on someone who was never asked.

None of this needs a criminal genius. That is the uncomfortable part: the tools are ordinary, cheap and widely available.

Keeping perspective

This is a real problem and it is not the end of trust.

Voice fraud has existed for as long as telephones. People impersonated relatives long before AI, using pressure and a bad line. Cloning makes the attack easier and more convincing. It does not create a new category of crime.

Most cloned audio is still detectable by a careful listener, especially over a longer conversation. Detection tools exist, though they are unreliable and should not be leaned on.

And the defences that work best are not technical.

What actually protects you

Agree a family code word. One word, agreed in person, never sent by message. Anyone claiming an emergency must say it. This costs nothing and defeats the entire attack.

Hang up and call back. On the number you already have, not one you were given. A cloned voice cannot answer their real phone.

Ask something only they know. Not a fact from social media. Something from a shared memory.

Be suspicious of urgency and secrecy. "Right now" and "don't tell anyone" are the two constants of this scam. Real emergencies survive a five-minute check.

Never treat a voice as a password. If a service offers voice as your only login, use something else as well.

If you build or use this, the rule is short: cloning someone's voice needs that person's clear, informed, specific permission.

Clear means they said yes. Informed means they know what it will say and where. Specific means permission for this use, not for everything forever. And they should be able to withdraw it.

"It was on the internet" is not consent. Neither is a checkbox nobody read.

Courts in several countries, including India, have started treating a person's voice as something they own and can protect. Rules differ by country and are changing quickly. Check where you are.

Remember this

  • Voice cloning is genuinely valuable for people losing their speech, and genuinely dangerous as a fraud tool.
  • A familiar voice is no longer proof of who is speaking. Agree a code word with your family.
  • Copying a real person's voice needs their clear, specific, withdrawable permission.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install numpy scipy librosa

There is no cloning code in this lesson. That is a deliberate choice, and not because the topic is untouchable.

The useful thing to understand as a developer is not how to copy a voice. It is what copying a voice breaks — and the demonstration below shows that with numbers, using nothing but the synthesiser from text to speech and the verifier from speaker identification.

Building a voice verifier that looks solid, then breaking it

The scenario: a bank enrols a customer's voice, then accepts or rejects future callers by similarity.

voice_is_not_a_password.py
import numpy as np, librosa
from scipy.signal import lfilter

SR = 16000

def resonator(x, fc, bw):
    r = np.exp(-np.pi*bw/SR); th = 2*np.pi*fc/SR
    return lfilter([1-2*r*np.cos(th)+r*r], [1.0, -2*r*np.cos(th), r*r], x)

PHRASE = [(730,1090,2440), (270,2290,3010), (300,870,2240), (530,1840,2480)]

def utterance(f0, scale, bw, rng, seconds=1.6):
    """A voice with pitch f0, throat length scale, breathiness bw."""
    n = int(SR*seconds/len(PHRASE)); out = []
    for formants in PHRASE:
        f0v = f0 * (1 + 0.04*rng.standard_normal())
        src = np.zeros(n); src[::max(1, int(SR/f0v))] = 1.0
        x = src
        for fc in formants:
            x = resonator(x, fc*scale, bw)
        out.append(x/(np.abs(x).max() + 1e-9))
    return np.concatenate(out)

def embed(x):
    m = librosa.feature.mfcc(y=x, sr=SR, n_mfcc=20, n_fft=400, hop_length=160)
    return m[1:].mean(axis=1)

TARGET = (150, 1.00, 55)                       # the customer the bank enrolled
OTHERS = [(215,1.18,60), (110,0.88,130), (180,1.05,70), (130,0.95,90)]
THRESHOLD = 0.60

R = lambda s: np.random.default_rng(s)
enrol = np.mean([embed(utterance(*TARGET, rng=R(i))) for i in range(3)], axis=0)
mu = np.mean([enrol] + [embed(utterance(*o, rng=R(50+i))) for i, o in enumerate(OTHERS)], axis=0)
def fp(v):
    v = v - mu
    return v/np.linalg.norm(v)

print("Step 1 -- does the verifier work at all?")
genuine  = [float(fp(enrol) @ fp(embed(utterance(*TARGET, rng=R(200+i))))) for i in range(5)]
impostor = [float(fp(enrol) @ fp(embed(utterance(*o, rng=R(300+i))))) for i, o in enumerate(OTHERS)]
print("  the real customer, 5 fresh recordings:", np.round(genuine, 3))
print("  four different real people           :", np.round(impostor, 3))
print(f"  threshold {THRESHOLD}: every genuine attempt passes, every real impostor fails.\n")

print("Step 2 -- an attacker has a few seconds of the customer's public audio.")
stolen = embed(utterance(*TARGET, rng=R(99)))
best, params = -1.0, None
for f0 in np.arange(100, 240, 10):
    for scale in np.arange(0.85, 1.25, 0.03):
        for bw in (50, 70, 90, 120):
            s = float(fp(stolen) @ fp(embed(utterance(f0, scale, bw, rng=R(7)))))
            if s > best: best, params = s, (int(f0), round(float(scale), 2), bw)
print("  tuned a plain synthesiser to match that clip:", params)

final = float(fp(enrol) @ fp(embed(utterance(*params, rng=R(11)))))
print(f"  its score against the enrolled voice: {final:.3f}")
print("  verdict:", "ACCEPTED" if final >= THRESHOLD else "rejected")
Output
Step 1 -- does the verifier work at all?
  the real customer, 5 fresh recordings: [0.976 0.88  0.94  0.794 0.836]
  four different real people           : [-0.179 -0.031  0.287  0.052]
  threshold 0.6: every genuine attempt passes, every real impostor fails.

Step 2 -- an attacker has a few seconds of the customer's public audio.
  tuned a plain synthesiser to match that clip: (150, 1.0, 50)
  its score against the enrolled voice: 0.876
  verdict: ACCEPTED

Read step 1 and step 2 together

Step 1 is a well-behaved biometric system. The real customer scores between 0.794 and 0.976 across five separate recordings. Four different real people score between −0.179 and 0.287. The gap is enormous, and any threshold in between separates them cleanly. On this evidence you would deploy it.

Step 2 defeats it without any machine learning. A grid search over three physical parameters — pitch, throat length, breathiness — found a setting scoring 0.876 against the enrolment. That sits comfortably inside the genuine range, above three of the five real attempts by the actual customer.

The attacker never had the customer's password, or their phone, or their cooperation. They had a few seconds of audio, which is what a voicemail greeting is.

The system did exactly what it was designed to do. It answered "does this sound like the enrolled person". It has no mechanism for answering "is this a human being at all". Those are different questions, and only the second one is authentication.

Note what this demonstration is not: the audio produced here is buzzy synthetic vowels, nowhere near a usable imitation of anyone. It does not need to be. The verifier is comparing statistics, and statistics were all that had to be matched.

What this means for anything you build

Voice is an identifier, not an authenticator. It is a username. It is publicly observable, permanently recorded in a hundred places, and impossible to change after a breach. Treat it as one factor among several, and never as the only one.

If you are asked to build voice login, the correct response is to say what it can and cannot do, in writing, before it ships.

Countermeasures, and their honest limits

Systems that need to resist this add a spoofing countermeasure — a separate model asking "is this audio synthetic?" rather than "whose voice is this?"

They work, with two caveats that matter more than the headline numbers.

They generalise poorly to new attacks. A countermeasure trained on the synthesis methods of last year performs substantially worse on this year's. The ASVspoof challenge series measures exactly this, and the drop against unseen attack types is consistently large.

Your adversary reads the same papers. Published accuracy is measured against fixed attack sets. A motivated attacker adapts, and published numbers become an upper bound rather than an expectation.

Add liveness checks that do not depend on the audio being real: a challenge phrase generated at call time, a callback to a registered number, a second channel entirely.

If you build a cloning product, build these too

Not aspirations. Things with code behind them.

Verifiable consent. The person whose voice it is records a consent statement, in their own voice, naming the use. Store it with the model. No consent record, no model.

Refusal for public figures. Maintain a blocklist and check enrolment audio against it. Expect to be tested by people trying to get around it.

Audible or embedded disclosure. Label synthetic output. Perceptual watermarking — AudioSeal and WavMark are the current open options — survives ordinary compression and re-encoding, though not a determined removal attempt. Also write plain metadata, since a large share of downstream handling is honest and only uninformed.

Rate limits and logging. Bulk generation against many identities is the signature of abuse. Log who cloned what, and keep it.

A revocation path. A person must be able to say stop, and something must actually happen. That means knowing which outputs came from which voice.

A model card. State what the system does, what it refuses, its known failure modes, and who to contact. See model cards.

Common mistakes

Assuming a detector solves it. Detection is unreliable, degrades with compression, and is adversarially fragile. Provenance at generation time is far more dependable than forensics afterwards.

Treating consent as a checkbox. Consent must be specific, informed and withdrawable. A terms-of-service clause covering "all content you upload" is not consent to clone a third party who never used your product.

Ignoring the voice of a person who is not your user. The most common misuse is cloning someone who never touched the product. Your consent flow must cover the subject, not only the uploader.

Publishing only clean-condition security numbers. Report performance against adaptive attacks, or state plainly that you did not test them.

Try it yourself

Add a fifth real person whose parameters are close to the target and check whether the threshold still separates real people. Then narrow the attacker's grid to a single guess at f0 and see how much audio-derived information the attack actually needed. The answer — very little — is the security-relevant result.

What to learn next

Researcher — Mathematics and papers.

Capability, stated precisely

Zero-shot and few-shot voice cloning is a solved capability at research scale, and the enrolment requirement has collapsed by roughly four orders of magnitude in a decade.

SV2TTS (Jia et al., 2018) conditioned a Tacotron 2 synthesiser on a speaker embedding from an independently trained verification encoder, achieving zero-shot cloning from about 5 seconds of enrolment. The transfer-learning framing — borrow the encoder from a verification system — is why cloning and verification are technically the same problem viewed from two sides.

YourTTS (Casanova et al., 2022) extended VITS with speaker conditioning for zero-shot multilingual cloning, including cross-lingual transfer.

VALL-E (Wang et al., 2023) reframed TTS as language modelling over EnCodec residual-vector-quantiser tokens, trained on 60,000 hours. From a 3-second prompt it preserves speaker identity, prosody, emotion and acoustic environment, because it learned continuation rather than neutral reading. That last property is what makes output convincing: the room matches.

The engineering consequence is that no meaningful barrier remains at the model level. Open weights, consumer hardware and seconds of audio suffice. Governance cannot depend on capability being scarce.

Attack surface against speaker verification

ASV systems face replay, voice conversion and synthesis. The developer block demonstrates the general principle at toy scale: verification scores a similarity, and any process that maximises that similarity succeeds, whether or not it produces perceptually convincing speech.

ASVspoof (2015, 2017, 2019, 2021, 2024) is the benchmark series. Two conditions matter: logical access (synthesis and conversion injected into the channel) and physical access (replay). Countermeasures are evaluated by EER and, since 2019, by the t-DCF — tandem detection cost function — which scores the countermeasure jointly with the ASV system it protects rather than in isolation. Evaluating a countermeasure alone optimises the wrong objective, since its errors interact with the verifier's.

The consistent finding across editions is that countermeasure performance degrades substantially on attack types absent from training, and that the degradation is larger than architecture differences among countermeasures. SASV — spoofing-aware speaker verification — is the current framing, jointly optimising verification and countermeasure.

Detection, and why it is the weaker lever

Synthetic-speech detection exploits artefacts: vocoder phase inconsistency, unnatural long-term prosody, missing breath and articulation noise, and spectral discontinuities at frame boundaries.

Three structural problems limit it.

Codec and channel destruction. Telephone bandwidth, Opus, re-recording over a speaker all remove the fine artefacts detectors rely on. Detection accuracy in the wild is far below laboratory figures, and the deployment conditions that matter most are the worst ones.

Adversarial adaptation. Detection is a classification problem with an adversary who can query it. Improvements are transient by construction.

Base rates. Even at 99 percent accuracy, screening a population where synthetic audio is rare produces overwhelmingly false positives. Accusing real recordings of being fake is its own serious harm, and it is the failure mode that scales.

Provenance is the stronger lever precisely because it does not require distinguishing real from fake after the fact. Audio watermarking — AudioSeal (San Roman et al., 2024) and WavMark (Chen et al., 2023) — embeds a detectable, localisable signal at generation, surviving common compression and editing while remaining imperceptible. C2PA content credentials attach signed provenance metadata.

Both have the same honest limitation: they are robust to incidental processing, not to a determined removal attempt, and they only cover generators that choose to participate. They raise the cost of casual misuse and do not stop a motivated adversary.

This area is moving quickly, and anything written here should be checked against the current position in your jurisdiction.

  • EU AI Act. Article 50 imposes transparency obligations on providers and deployers of systems generating synthetic audio, including machine-readable marking and disclosure of deepfake content. The Act entered into force in 2024, with obligations phased in over the following years.
  • Tennessee ELVIS Act (2024). Amended the state's right of publicity to cover voice explicitly, extending protection to a person's voice as a property interest and creating liability for tools whose primary purpose is unauthorised replication.
  • NO FAKES Act. A proposed US federal bill creating a digital replication right. Introduced and reintroduced; not enacted at the time of writing.
  • India. No dedicated statute specific to voice cloning. Courts have granted personality-rights injunctions covering voice and likeness for public figures, and existing law on impersonation, cheating and IT offences applies.
  • Denmark has proposed treating likeness and voice under a copyright-style framework.

The general direction is convergent: disclosure obligations on generators, and a property or personality interest in one's own voice. The specifics differ substantially.

Research directions worth following

Robust watermarking under adversarial removal — current schemes assume incidental processing, and the adversarial case is largely open.

Generalisable countermeasures — one-class and self-supervised formulations that model bona fide speech rather than known attacks, avoiding the retrain-per-attack treadmill.

Provenance across the pipeline — cryptographic signing at capture, so authenticity is established positively rather than inferred negatively.

Consent infrastructure — machine-readable, revocable, auditable consent for biometric identity. Largely an unsolved systems problem rather than a modelling one, and the bottleneck in practice.

Reading

What to learn next