How computers hear sound
A computer hears by measuring air pressure thousands of times a second and storing each measurement as a number, so all audio is really one long list of numbers.
- 14 min read
- 3 reading levels
- Updated
Read these first
On this page 10
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
A computer hears by measuring the air thousands of times a second and writing down each measurement as a number.
That is the whole trick. Sound goes in, a very long list of numbers comes out.
The analogy you have already lived
Stand at the edge of a lake and watch the water against a jetty post. The water rises and falls. Now imagine a ruler taped to that post, and a friend shouting "now!" every second.
Each time your friend shouts, you read the ruler and write down the height. At the end you have no water. You have a column of numbers in a notebook.
If your readings are close enough together, you can draw the wave back out of them. A microphone does exactly this, about forty-four thousand times every second.
What sound actually is
Sound is not a thing that travels. It is a squeeze that travels.
When you clap, your hands shove the air next to them. That squeezed air shoves the air beyond it, and so on, until a patch of squeezed air reaches your ear. Your eardrum gets pushed, then pulled, then pushed again.
A microphone has a thin sheet inside it that gets pushed and pulled the same way. That movement becomes a small electrical signal, and the signal gets measured.
The three words worth knowing
Sample — one measurement of air pressure at one instant. One reading of the ruler.
Sample rate — how many measurements are taken each second. Phone calls use eight thousand. Music CDs use forty-four thousand one hundred. Films usually use forty-eight thousand.
Bit depth — how finely each measurement is written down. Sixteen bits means each reading is stored as a whole number between about minus thirty-three thousand and plus thirty-three thousand. More bits means a more careful reading and a bigger file.
How it works
your voice
|
v
air gets squeezed and released
|
v
[ microphone ] thin sheet pushed and pulled
|
v
measure the push, thousands of times a second
|
v
[ 0.00, 0.17, 0.32, 0.43, 0.49, 0.49, 0.44, ... ]
|
v
a file on diskEverything else in this section works on that list of numbers. Speech recognition, noise removal, wake words, voice assistants — every one of them starts here.
Why the sample rate is not a free choice
Take too few measurements and the wave you draw back is the wrong wave.
Picture a ceiling fan filmed on a phone. Sometimes the blades look frozen. Sometimes they look like they are turning slowly backwards. The fan is fine. The camera is taking pictures too slowly to keep up with it.
The same thing happens to sound. A high whistle recorded at too low a rate comes back as a low hum that was never there. That mistake has a name: aliasing, meaning a fast wave pretending to be a slow one.
The rule that follows is short. To capture a sound, measure more than twice as often as its highest pitch. Human hearing stops around twenty thousand cycles a second, which is why music is recorded at forty-four thousand.
Phone calls at eight thousand samples a second can only carry pitches up to four thousand. That is why "s" and "f" sound alike on a bad call. The detail that separates them was never recorded in the first place.
Where you have already seen this
- WhatsApp voice notes sound thinner than a music file, because fewer measurements are kept per second.
- A ceiling fan on video appearing to spin backwards is aliasing, in pictures instead of sound.
- Spotify's quality setting is a choice about how much of that number list to send you.
- "Recording in progress" on a call means a copy of those numbers is being kept.
What is honestly hard here
Nothing here is hard to follow, but one part is hard to believe: there is no sound inside the computer. None. There is a list of numbers, and a speaker that can push air around in the pattern those numbers describe.
Read that twice if it feels strange. Almost everyone has to.
Remember this
- A microphone measures air pressure over and over, and each measurement is one sample.
- Sample rate is how many measurements per second, and a higher rate keeps more of the high pitches.
- Every audio model in this section reads that list of numbers and nothing else.
What to learn next
- Waveforms and spectrograms — turning that list of numbers into a picture of pitch over time.
- Audio features and MFCCs — squeezing a second of sound into a handful of useful numbers.
- NumPy — the array library every example in this section leans on.
Developer — Code and libraries.
Setup
pip install numpyNumPy is the only requirement. The wave module used below ships with Python. This runs on any laptop in well under a second, with no GPU and no download.
Making sound from nothing
You do not need a microphone or a dataset to work with audio. A tone is a sine wave, and NumPy makes sine waves. Most examples in this section are built this way, so nobody has to spend mobile data to follow along.
import numpy as np
SR = 8000 # samples per second: how often we measure the air
DUR = 0.01 # ten milliseconds, long enough to see the shape
FREQ = 440 # hertz: the pitch of the A above middle C
t = np.arange(int(SR * DUR)) / SR # the clock time of each measurement
wave = 0.5 * np.sin(2 * np.pi * FREQ * t) # air pressure at each of those times
print("numbers captured:", wave.size)
print("first eight:", np.round(wave[:8], 3))
print("quietest:", round(float(wave.min()), 3), " loudest:", round(float(wave.max()), 3))
print("seconds between two numbers:", 1 / SR)numbers captured: 80 first eight: [0. 0.169 0.319 0.43 0.491 0.494 0.438 0.331] quietest: -0.5 loudest: 0.5 seconds between two numbers: 0.000125
There it is. That array is the sound. Not a picture of it, not a summary of it — the thing itself, exactly as your machine stores it.
Look at the shape of the numbers. They climb from zero, peak near 0.494, and turn back down. Across the full ten milliseconds this pattern repeats about four and a half times, because 440 cycles per second means 4.4 cycles in ten milliseconds.
Line by line, the parts that are not obvious
np.arange(int(SR * DUR)) / SR — this builds the timestamp of every sample. Sample 0 happens at 0 seconds, sample 1 at 0.000125 seconds, and so on. Dividing the sample index by the sample rate turns "which measurement" into "at what time".
0.5 * — amplitude, which means loudness. Audio is conventionally kept inside the range minus one to plus one. Going outside that range flattens the peaks and sounds harsh, so 0.5 leaves comfortable headroom.
2 * np.pi * FREQ * t — one full cycle of a sine is 2 * np.pi radians. Multiplying by frequency and time answers "how far around the circle are we at this instant".
Writing it to a real file
The array holds floating-point values. A WAV file holds whole numbers, so there is one conversion step.
import wave, os
import numpy as np
SR = 8000
t = np.arange(int(SR * 1.0)) / SR
tone = 0.5 * np.sin(2 * np.pi * 440 * t)
# 16-bit audio stores whole numbers from -32768 to 32767, so stretch onto that range.
pcm16 = np.int16(np.clip(tone, -1.0, 1.0) * 32767)
with wave.open("tone.wav", "wb") as f:
f.setnchannels(1) # mono: one microphone, one column of numbers
f.setsampwidth(2) # 2 bytes = 16 bits per measurement
f.setframerate(SR)
f.writeframes(pcm16.tobytes())
print("first eight whole numbers:", pcm16[:8])
print("file size on disk:", os.path.getsize("tone.wav"), "bytes")
print("numbers x bytes each:", pcm16.size, "x 2 =", pcm16.size * 2, "bytes of audio")
with wave.open("tone.wav", "rb") as f:
print("read back ->", f.getnchannels(), "channel,",
f.getframerate(), "Hz,", f.getsampwidth() * 8, "bit,",
f.getnframes(), "frames")first eight whole numbers: [ 0 5549 10443 14101 16093 16181 14356 10834] file size on disk: 16044 bytes numbers x bytes each: 8000 x 2 = 16000 bytes of audio read back -> 1 channel, 8000 Hz, 16 bit, 8000 frames
Open tone.wav in any media player and you will hear a one-second beep. The extra 44 bytes on disk are the WAV header, which records sample rate, channel count and bit depth so a player knows how to read the numbers.
That header is why raw audio arrays are useless on their own. The values [0, 5549, 10443, ...] mean nothing until something states that they arrive 8000 times a second.
Watching aliasing happen
These are the most useful twelve lines in the lesson. Play four tones through an 8000 Hz recorder, then ask the recording what it heard.
import numpy as np
SR = 8000
t = np.arange(SR) / SR
for freq in (1000, 3000, 6000, 7000):
x = np.sin(2 * np.pi * freq * t)
spec = np.abs(np.fft.rfft(x)) # strength of each pitch
peak = np.fft.rfftfreq(x.size, 1 / SR)[spec.argmax()] # the loudest pitch present
print(f"played {freq:5d} Hz -> recording contains {peak:6.0f} Hz")played 1000 Hz -> recording contains 1000 Hz played 3000 Hz -> recording contains 3000 Hz played 6000 Hz -> recording contains 2000 Hz played 7000 Hz -> recording contains 1000 Hz
The first two are honest. The last two are lies, and the file has no way of knowing.
At 8000 samples per second the highest recordable pitch is 4000 Hz, half the sample rate. That halfway point is the Nyquist frequency. Anything above it folds back down: 6000 arrives as 2000, and 7000 arrives as 1000.
Real recording hardware puts an analogue low-pass filter in front of the converter to remove those pitches before they can fold. Downsampling in software without filtering first reproduces this bug exactly, and it happens often.
Common mistakes
Resampling by throwing samples away. Writing audio[::3] to go from 48 kHz to 16 kHz folds every pitch above 8 kHz down into your speech band, permanently. Use scipy.signal.resample_poly or librosa.resample, which filter before they decimate.
Forgetting that most models want 16 kHz mono. Whisper, wav2vec 2.0 and nearly every speech model expect a single channel at 16000 Hz. Feeding 44.1 kHz stereo gives you either a crash or, worse, quiet nonsense.
Mixing up the float and integer conventions. Some libraries hand you floats between minus one and plus one; others hand you int16 between minus 32768 and 32767. Multiply an int16 array by 32767 by mistake and it overflows into loud noise. Print arr.dtype and arr.max() before trusting any audio array.
Assuming stereo means two useful signals. Plenty of recordings duplicate one channel twice. Average them with audio.mean(axis=1) rather than assuming the second channel adds information.
Try it yourself
Change SR to 44100 in the first script, keep DUR at 0.01, and predict the new wave.size before running it. Then add a second sine at 660 Hz to the first and write the sum to a WAV. Two tones playing at once are stored as one list of numbers, not two. Hearing that is the point.
What to learn next
- Waveforms and spectrograms — turning that list of numbers into a picture of pitch over time.
- Audio features and MFCCs — squeezing a second of sound into a handful of useful numbers.
- NumPy — the array library every example in this section leans on.
Researcher — Mathematics and papers.
The sampling theorem
Continuous-time pressure $x(t)$ becomes a discrete sequence under ideal sampling at interval $T_s$:
$$ x[n] = x(nT_s), \qquad f_s = \frac{1}{T_s} $$
Where $n \in \mathbb{Z}$ is the sample index, $T_s$ is the sampling period in seconds, and $f_s$ is the sampling rate in hertz.
Sampling multiplies the signal by a Dirac comb, and multiplication in time is convolution in frequency. The spectrum of the sampled signal is therefore the original spectrum repeated at every multiple of $f_s$:
$$ X_s(f) = f_s \sum_{k=-\infty}^{\infty} X(f - k f_s) $$
Nyquist–Shannon (Shannon, 1949, building on Nyquist, 1928): if $X(f) = 0$ for all $|f| \geq f_s/2$, the copies do not overlap and $x(t)$ is exactly recoverable by sinc interpolation:
$$ x(t) = \sum_{n=-\infty}^{\infty} x[n] \,\operatorname{sinc}!\left(\frac{t - nT_s}{T_s}\right) $$
If the bandlimit condition fails, adjacent copies overlap, and the overlapping energy is indistinguishable from genuine in-band content. A component at $f_0 > f_s/2$ appears at $|f_0 - k f_s|$ for whichever integer $k$ brings it into $[0, f_s/2]$ — exactly the folding printed in the developer block.
The theorem is an idealisation twice over. It assumes strictly bandlimited signals, which necessarily have infinite time support, and it assumes infinitely long sinc reconstruction. Real converters approximate both.
Quantisation
Rounding each sample to one of $2^b$ levels introduces an error $e[n]$. Under the standard high-resolution assumptions — $e[n]$ uniform on $[-\Delta/2, \Delta/2]$, white, and independent of the signal — the step size and error power are:
$$ \Delta = \frac{2A}{2^b}, \qquad \sigma_e^2 = \frac{\Delta^2}{12} $$
Where $A$ is the peak amplitude of the full-scale range and $b$ is the bit depth. For a full-scale sinusoid this yields the familiar bound:
$$ \mathrm{SNR}_{\mathrm{dB}} = 6.02\,b + 1.76 $$
So 16-bit audio gives about 98 dB of signal-to-noise ratio and 24-bit about 146 dB.
The independence assumption fails at low amplitudes, where the error correlates with the signal and becomes audible as harmonic distortion rather than hiss. Dither — adding roughly one least-significant-bit of noise before quantising — decorrelates the error at a cost of about 3 dB in noise floor. Lipshitz, Wannamaker and Vanderkooy (1992), Quantization and Dither: A Theoretical Survey, is the definitive treatment.
Storage cost and companding
Linear PCM at rate $f_s$, depth $b$ bits and $C$ channels costs $f_s \cdot b \cdot C$ bits per second: 1.41 Mbit/s for CD stereo, 256 kbit/s for 16 kHz 16-bit mono.
Telephony historically used companding rather than more bits. μ-law (ITU-T G.711, North America and Japan) and A-law (Europe) apply a logarithmic amplitude map before 8-bit quantisation:
$$ F(x) = \operatorname{sgn}(x)\,\frac{\ln(1 + \mu |x|)}{\ln(1 + \mu)}, \qquad \mu = 255 $$
This delivers roughly the perceptual quality of 12-bit linear PCM in 8 bits, because the quantisation steps grow with amplitude and perceived loudness is approximately logarithmic. The same transform reappears in modern generative audio: WaveNet (van den Oord et al., 2016) modelled 8-bit μ-law samples as a 256-way categorical distribution specifically to keep the output softmax tractable.
Rate choices and why they are what they are
| Rate | Nyquist | Where it is used | Rationale |
|---|---|---|---|
| 8 kHz | 4 kHz | G.711 telephony | Intelligibility survives; fricative confusion does not |
| 16 kHz | 8 kHz | Whisper, wav2vec 2.0, most ASR | Captures fricative energy, halves compute versus 32 kHz |
| 22.05 kHz | 11 kHz | Many TTS vocoders | Exactly half of 44.1 kHz, cheap to convert |
| 44.1 kHz | 22.05 kHz | CD audio | Above the roughly 20 kHz hearing limit with filter margin |
| 48 kHz | 24 kHz | Film, broadcast, most DAWs | Integer-friendly with video frame rates |
The odd 44.1 kHz figure is an artefact of storing digital audio on video tape: 3 samples per line, 245 usable lines, 60 fields per second gives 44100.
Resampling
Rate conversion by a rational factor $L/M$ is upsample by $L$ with zero stuffing, low-pass filter at $\min(\pi/L, \pi/M)$, then decimate by $M$. scipy.signal.resample_poly implements this as a polyphase filter bank, which skips computing the samples that decimation would discard and cuts cost by a factor of about $M$.
Naive slicing (x[::M]) omits the filter, and it is among the most common preprocessing bugs in audio ML pipelines. It fails silently: no exception, no warning, only permanently corrupted high-frequency content that surfaces later as an unexplained accuracy gap between training data and production audio.
Reading
- Shannon (1949), Communication in the Presence of Noise, Proc. IRE 37(1) — the sampling theorem as usually cited.
- Nyquist (1928), Certain Topics in Telegraph Transmission Theory — the earlier bandwidth result it builds on.
- Oppenheim and Schafer, Discrete-Time Signal Processing, 3rd ed., chapter 4 — sampling and reconstruction done properly.
- Smith, The Scientist and Engineer's Guide to Digital Signal Processing — free online, unusually readable.
- van den Oord et al. (2016), WaveNet — arxiv.org/abs/1609.03499
What to learn next
- Waveforms and spectrograms — turning that list of numbers into a picture of pitch over time.
- Audio features and MFCCs — squeezing a second of sound into a handful of useful numbers.
- NumPy — the array library every example in this section leans on.