Speech and Audio AI

Noise reduction

Noise reduction removes background sound from a recording by editing its spectrogram, and pushing too hard trades hiss for a worse artefact called musical noise.

On this page 9
  1. The short answer
  2. The analogy you have already lived
  3. Why this is possible at all
  4. How it works
  5. The artefact nobody warns you about
  6. Where you have already used this
  7. What is honestly hard here
  8. Remember this
  9. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

Noise reduction removes unwanted background sound while keeping the sound you want.

Almost every method works the same way: learn what the noise looks like, then subtract it.

The analogy you have already lived

Think of washing a stain out of a white shirt. You scrub gently and some of the stain goes. You scrub harder and more goes — but the fabric starts to thin, and eventually you have a hole.

There is a point where continuing to attack the stain damages the shirt more than the stain ever did.

Noise reduction works exactly like this. Gentle cleaning removes hiss and leaves the voice intact. Aggressive cleaning removes almost all the hiss, and leaves the voice sounding underwater and strange.

Most people's instinct is to turn the dial all the way up. That is the wrong instinct, and the developer section measures precisely why.

Why this is possible at all

A recording of a voice in a noisy room contains both, added together. There is no separate noise track to delete.

So how can it be separated? Because noise and speech look different in a spectrogram. That is the picture of pitch over time from waveforms and spectrograms.

Steady noise stays the same. A fan, an air conditioner, a hum from a bad cable. The same energy at the same pitches, second after second.

Speech moves constantly. It starts and stops. Its pitches leap around. It has gaps.

That difference is the whole opening. Listen to the parts where nobody is talking, and you learn what the noise is. Then remove that amount from everywhere.

How it works

   noisy recording
        |
        v
  [ make a spectrogram ]
        |
        v
  [ during the silence, measure the noise ]   ->  a noise profile
        |
        v
  [ subtract that profile from every moment ]
        |
        v
  [ turn the spectrogram back into sound ]
        |
        v
   cleaner recording

That third step is where all the difficulty lives. Subtract too little and hiss remains. Subtract too much and you damage the voice.

The artefact nobody warns you about

Push the subtraction hard and the hiss does go away. What replaces it is worse.

You get short, random, tinkling tones. They come and go, at unrelated pitches, like tiny bells or water dripping in a metal pipe. Engineers call this musical noise, and once you have heard it you cannot stop hearing it.

Here is why it appears. Noise is random, so at any instant a pitch band is a little above or below its average. Subtract the average and most bands go to zero. The ones that happened to be loud survive as isolated blips. Those surviving blips are scattered, brief and tonal — which is the recipe for something that sounds like music.

Steady hiss is easy for your ear to ignore. Random tinkling is not. A recording with more hiss often sounds better than one with musical noise.

Where you have already used this

  • Video calls with a "suppress background noise" setting.
  • Phone calls, where noise cancellation runs on every call whether you asked or not.
  • Noise-cancelling headphones, which do a related job with microphones and speakers instead of files.
  • Voice notes cleaned up automatically before they are sent.
  • Old recordings restored by removing tape hiss or record crackle.

What is honestly hard here

Steady noise is manageable. Sudden noise is not. A fan is predictable. A door slamming, a horn, a child shouting — those arrive without warning and look like speech to the algorithm.

Another voice is the hardest case of all. A second person talking has exactly the properties of speech, because it is speech. Separating overlapping voices is a distinct and much harder problem.

Cleaning can hurt what comes next. This one surprises people. Aggressively denoised audio often gives worse speech recognition accuracy than the noisy original. The cleaning removed real speech detail along with the noise. If a machine is the listener, test whether cleaning helps before assuming it does.

Remember this

  • Noise reduction learns the noise from quiet moments, then subtracts it from everywhere.
  • Subtracting harder gives a quieter background and a more damaged voice, always.
  • Musical noise — random tinkling — is what over-cleaning sounds like, and it is worse than hiss.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install numpy scipy

No downloads, no GPU, no audio files. We generate a voice-like signal, bury it in noise, and build spectral subtraction from scratch — then measure exactly what turning the dial up costs.

The classic algorithm, in full

Spectral subtraction is the oldest practical method and still the right one to understand first. Everything modern is a refinement of this shape.

denoise.py
import numpy as np
from scipy.signal import stft, istft, lfilter
from scipy.stats import kurtosis

SR = 16000
rng = np.random.default_rng(0)

def resonator(x, fc, bw=90.0):
    r = np.exp(-np.pi*bw/SR); th = 2*np.pi*fc/SR
    return lfilter([1-2*r*np.cos(th)+r*r], [1.0,-2*r*np.cos(th), r*r], x)

def speech(seconds):
    """A vowel-like signal: vocal-cord pulses through three mouth resonances."""
    n = int(SR*seconds)
    src = np.zeros(n); src[::int(SR/120)] = 1.0
    for fc in (730, 1090, 2440):
        src = resonator(src, fc)
    return src/np.abs(src).max()

LEAD  = 0.5                                  # half a second of noise before speech starts
clean = np.concatenate([np.zeros(int(SR*LEAD)), speech(1.5)])
noisy = clean + 0.15*rng.standard_normal(clean.size)
sp = slice(int(SR*LEAD), None)               # where the speech is
si = slice(0, int(SR*LEAD))                  # where it is supposed to be silent

def snr_db(est):
    return 10*np.log10(np.mean(clean[sp]**2) / np.mean((est[sp]-clean[sp])**2))

f, t, Z = stft(noisy, fs=SR, nperseg=512, noverlap=384)
mag, phase = np.abs(Z), np.angle(Z)
profile = mag[:, t < LEAD].mean(axis=1, keepdims=True)    # learn noise from the lead-in

print(f"{'setting':<26}{'SNR (dB)':>10}{'residual hiss':>15}{'burstiness':>12}")
print(f"{'noisy input':<26}{snr_db(noisy):10.2f}{np.sqrt(np.mean(noisy[si]**2)):15.4f}"
      f"{kurtosis(noisy[si]):12.2f}")
for alpha, floor in ((1.0, 0.02), (2.0, 0.02), (4.0, 0.02), (4.0, 0.15)):
    m = np.maximum(mag - alpha*profile, floor*mag)        # subtract, but never below a floor
    _, est = istft(m*np.exp(1j*phase), fs=SR, nperseg=512, noverlap=384)
    est = est[:clean.size]
    print(f"{'subtract '+str(alpha)+'x, floor '+str(floor):<26}{snr_db(est):10.2f}"
          f"{np.sqrt(np.mean(est[si]**2)):15.4f}{kurtosis(est[si]):12.2f}")
Output
setting                     SNR (dB)  residual hiss  burstiness
noisy input                     6.68         0.1504       -0.07
subtract 1.0x, floor 0.02      12.75         0.0477        0.03
subtract 2.0x, floor 0.02      11.55         0.0088       21.54
subtract 4.0x, floor 0.02       7.34         0.0048       87.53
subtract 4.0x, floor 0.15       7.55         0.0227        0.08

Read those three columns against each other

This table shows the trap that catches every newcomer to audio cleanup.

SNR is best at 1x subtraction: 12.75 dB, up from 6.68. That is the honest measure of "how close is this to the original voice", and gentle cleaning wins it.

Residual hiss keeps falling as you subtract harder. At 4x the background is thirty times quieter than the input. If you judged by "is the silence silent", 4x looks like the best setting by a wide margin.

Burstiness exposes what is actually happening. That column is the kurtosis of the supposedly silent region — a measure of how spiky a signal is. Gaussian hiss sits near zero, and the input measures −0.07. At 2x it jumps to 21.5, and at 4x to 87.5.

That climb is musical noise appearing. The background got quieter and got spikier: no longer a smooth hiss but isolated random blips. Your ear notices blips and ignores hiss, so the 4x version sounds considerably worse despite a quieter and numerically better-looking background.

Notice that SNR also drops at 4x, to 7.34 — barely above the untouched input. Aggressive subtraction damaged the speech about as much as the original noise did.

The floor is the fix

Compare the last two rows. Same 4x subtraction, floor raised from 0.02 to 0.15.

Burstiness collapses from 87.53 back to 0.08. The residual hiss rises from 0.0048 to 0.0227 — louder background, deliberately.

That is the trade, made explicit. The spectral floor forbids any band from being pushed below a fraction of its original value, so instead of isolated survivors on a silent background you keep a low, even hiss for them to sit in. Slightly more noise, and it stops sounding like wind chimes.

Every real denoiser has this control. It is called noise floor, residual level, or reduction depth. Do not set it to zero, however tempting the silence looks.

Line by line, the parts that are not obvious

mag[:, t < LEAD] learns the noise profile from a stretch known to contain no speech. In production nobody hands you that stretch, so you need a voice activity detector, or a minimum-statistics tracker that follows the lowest energy seen in each band over a sliding window. Estimating the noise is the hard part of real noise reduction, not subtracting it.

np.exp(1j*phase) reuses the original noisy phase. We modify magnitudes only, because phase is far harder to estimate and the ear is relatively insensitive to it. At low SNR this reuse becomes a real quality ceiling, and it is one reason neural denoisers that predict complex masks outperform this method.

np.maximum(..., floor*mag) prevents negative magnitudes. Subtracting an average from a fluctuating quantity produces negative values roughly half the time in quiet bands, and a magnitude cannot be negative. Clipping at zero is what creates musical noise; clipping at a floor is what tames it.

noverlap=384 with nperseg=512 is 75 percent overlap. Denoising modifies frames independently, and heavy overlap smooths the boundaries between them. At 50 percent overlap you can hear frame edges.

Common mistakes

Judging by the silence. Demonstrated above. Always listen to the speech, not the gaps. A denoiser evaluated only on quiet passages will be tuned into uselessness.

Denoising before speech recognition without measuring. This feels self-evidently right and frequently is not. Modern ASR models are trained on noisy audio and often handle the original better than a cleaned version with suppressed consonants. Measure word error rate both ways before shipping it.

Assuming the noise is stationary. Spectral subtraction assumes the noise profile holds. When someone turns on a fan mid-recording, everything after that point is under-cleaned. Update the estimate continuously.

Cleaning each channel of stereo separately. It destroys the tiny inter-channel differences your ear uses for direction, producing a flat, disembodied result. Use a multi-channel method, or process a mono downmix.

Reporting SNR alone. SNR does not track perceived quality well. Use PESQ or STOI as well, or run a listening test. The table above shows why: the 4x row wins on background quietness and loses on everything a listener cares about.

Try it yourself

Add alpha=0.5 to the loop and see whether gentler-than-1x is better still. Then replace the Gaussian noise with a hum, 0.15*np.sin(2*np.pi*50*np.arange(clean.size)/SR), and watch spectral subtraction handle it far better — a hum lives in one band and does not fluctuate, so there is nothing to over-subtract.

When to reach for a neural denoiser

Spectral subtraction is the right tool for steady hum and hiss when you can estimate the noise. For non-stationary noise, babble, or anything where quality matters, trained models are substantially better.

RNNoise is tiny, runs in real time on very modest hardware, and is what many voice-call systems use. DeepFilterNet and the models from the Microsoft DNS Challenge are stronger and heavier. The Python package noisereduce implements a refined stationary and non-stationary spectral gate and is a reasonable step up from the code above without any model download.

What to learn next

Researcher — Mathematics and papers.

Problem statement

Given $y[n] = x[n] + d[n]$ with speech $x$ and additive noise $d$ assumed uncorrelated, recover an estimate $\hat{x}$. In the STFT domain, writing $Y_{k,m}$, $X_{k,m}$, $D_{k,m}$ for bin $k$ and frame $m$, every method in this family computes a real or complex gain:

$$ \hat{X}{k,m} = G{k,m}\, Y_{k,m} $$

The methods differ only in how $G$ is derived.

Spectral subtraction

Boll (1979) proposed, in the generalised power form:

$$ |\hat{X}{k,m}|^{\gamma} = \max\left(|Y{k,m}|^{\gamma} - \alpha\, |\hat{D}{k}|^{\gamma},\ \beta\, |Y{k,m}|^{\gamma}\right) $$

Where $\gamma = 1$ gives magnitude subtraction and $\gamma = 2$ power subtraction, $\alpha \geq 1$ is the over-subtraction factor, $\beta$ the spectral floor, and $|\hat{D}_k|$ the noise estimate. The developer block implements $\gamma = 1$ and sweeps $\alpha$ and $\beta$.

Berouti, Schwartz and Makhoul (1979) made $\alpha$ a decreasing function of the frame SNR, subtracting more where the signal is weak, which is the standard practical form.

Musical noise, explained properly. $|Y_{k,m}|$ fluctuates around its expectation even where only noise is present. Subtracting the mean leaves a residual whose half-wave-rectified values are non-zero at randomly scattered $(k,m)$ points. Isolated, short-duration, narrowband survivors are perceptually tonal. The floor $\beta$ works because it converts isolated spikes on a zero background into spikes on a masking background, and auditory masking then hides them. The kurtosis rise in the developer output is a direct measurement of this residual becoming spiky, and Uemura et al. (2009) proposed kurtosis ratio as a musical-noise metric on exactly that reasoning.

Wiener filtering

The MMSE-optimal linear gain under Gaussian assumptions:

$$ G^{\text{W}}{k,m} = \frac{\xi{k,m}}{1 + \xi_{k,m}}, \qquad \xi_{k,m} = \frac{\mathbb{E}\left[|X_{k,m}|^2\right]}{\mathbb{E}\left[|D_{k,m}|^2\right]} $$

Where $\xi$ is the a priori SNR — the ratio of clean speech power to noise power, which is unknown because it depends on the quantity being estimated. The a posteriori SNR $\gamma_{k,m} = |Y_{k,m}|^2 / \mathbb{E}[|D_{k,m}|^2]$ is observable.

Ephraim and Malah (1984) resolved this with the decision-directed estimator:

$$ \hat{\xi}{k,m} = \mu \frac{|\hat{X}{k,m-1}|^2}{\mathbb{E}[|D_{k}|^2]} + (1-\mu)\max(\gamma_{k,m} - 1,\ 0) $$

with $\mu \approx 0.98$. This recursive smoothing is the single most important idea in classical speech enhancement, and it is why MMSE-STSA produces far less musical noise than spectral subtraction: heavy temporal smoothing of $\hat{\xi}$ suppresses precisely the frame-to-frame fluctuations that create isolated survivors.

The MMSE short-time spectral amplitude estimator itself minimises $\mathbb{E}[(|X| - |\hat{X}|)^2]$ under a Gaussian model, giving a gain in terms of confluent hypergeometric functions. Ephraim and Malah (1985) extended it to the log-spectral amplitude, which better matches perception and remains a competitive classical baseline forty years later.

Noise estimation

The genuinely hard subproblem. Options:

  • VAD-gated averaging. Update the estimate only in non-speech frames. Fails when the VAD fails, and VADs fail at low SNR.
  • Minimum statistics (Martin, 2001). Track the minimum of smoothed noisy power over a sliding window and compensate for the bias of the minimum operator. Requires no VAD and handles non-stationary noise, at the cost of a lag on the order of the window.
  • MMSE-based noise power estimation (Gerkmann and Hendriks, 2012). Estimate noise power with a speech-presence-probability weighting. Lower tracking delay than minimum statistics and the usual modern choice.

Neural approaches

Masking in the T-F domain. Predict a mask applied to the noisy STFT. The ideal ratio mask is:

$$ \text{IRM}{k,m} = \left(\frac{|X{k,m}|^2}{|X_{k,m}|^2 + |D_{k,m}|^2}\right)^{1/2} $$

Real-valued masks leave phase untouched, capping quality. The complex ideal ratio mask (Williamson, Wang and Wang, 2016) predicts real and imaginary parts, enabling phase correction and measurably outperforming magnitude masks.

Time-domain separation. Conv-TasNet (Luo and Mesgarani, 2019) replaces the STFT with a learned encoder-decoder and estimates masks in that latent space, sidestepping the phase problem entirely. It substantially outperformed STFT masking on separation and reframed the field.

Complex-domain recurrent networks. DCCRN (Hu et al., 2020) uses complex convolutions and a complex LSTM, and won the DNS Challenge 2020 real-time track.

Efficiency-first. RNNoise (Valin, 2018) combines a small GRU predicting band gains with classical pitch filtering, running in well under 1 percent of a mobile core. DeepFilterNet (Schröter et al., 2022) achieves full-band enhancement in real time on a CPU with a per-frame deep filter over a small complex neighbourhood.

The DNS Challenge series (Reddy et al., 2020 onward) is the field's benchmark and provides the standard training data and real-recording test sets.

Evaluation

  • SI-SDR (Le Roux et al., 2019). Scale-invariant signal-to-distortion ratio, which fixes SDR's exploitable sensitivity to rescaling. The default objective metric for separation and enhancement.
  • PESQ (ITU-T P.862) and POLQA (P.863). Intrusive perceptual quality, mapping to a MOS-like scale. PESQ was designed for codec and network degradation, and it correlates imperfectly with neural-enhancement artefacts — it can reward outputs listeners dislike.
  • STOI (Taal et al., 2011). Predicts intelligibility rather than quality, and the two dissociate: enhancement often raises quality while lowering or not changing intelligibility.
  • DNSMOS (Reddy et al., 2021). A non-intrusive learned MOS estimator, usable on real recordings with no clean reference.

The enhancement-for-ASR caveat is real and well documented. Front-end enhancement frequently degrades downstream recognition, because the enhancement objective targets perceptual quality while the recogniser was trained on noisy speech and relies on cues the enhancer removes. Joint training, or feeding both streams, is the standard remedy. Never assume a denoiser helps a machine listener — measure WER end to end.

Reading

  • Boll (1979), Suppression of Acoustic Noise in Speech Using Spectral Subtraction, IEEE TASSP 27(2).
  • Ephraim and Malah (1984), Speech Enhancement Using a Minimum Mean-Square Error Short-Time Spectral Amplitude Estimator, IEEE TASSP 32(6).
  • Martin (2001), Noise Power Spectral Density Estimation Based on Optimal Smoothing and Minimum Statistics, IEEE TSAP 9(5).
  • Luo and Mesgarani (2019), Conv-TasNet — arxiv.org/abs/1809.07454
  • Valin (2018), A Hybrid DSP/Deep Learning Approach to Real-Time Full-Band Speech Enhancement — arxiv.org/abs/1709.08243
  • Le Roux et al. (2019), SDR — Half-baked or Well Done? — arxiv.org/abs/1811.02508

What to learn next