Speech and Audio AI

Real-time audio pipelines

A real-time audio pipeline processes sound in small blocks as it arrives, and the thing that makes a voice assistant feel slow is usually waiting to decide you finished talking.

On this page 10
  1. The short answer
  2. The analogy you have already lived
  3. Why this changes how everything is built
  4. How it works
  5. The thing that actually makes it feel slow
  6. Where you have already used one
  7. What feels fast to a person
  8. What is honestly hard here
  9. Remember this
  10. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

A real-time audio pipeline handles sound in tiny pieces as it arrives, instead of waiting for a whole recording.

It is what separates a voice assistant that feels alive from one that feels broken.

The analogy you have already lived

Think of a good interpreter at a meeting. They do not wait for the speaker to finish a five-minute answer and then translate it. They start translating a few words behind, continuously, and finish moments after the speaker stops.

Now think of a bad one, who listens to the whole thing, thinks, then begins. Everyone stares at the table. The information is the same. The experience is completely different.

Every part of a voice assistant has to work like the good interpreter. Listen in small pieces, start working immediately, stay a little behind and never fall further behind.

Why this changes how everything is built

In every previous lesson, we had a complete audio file. We could look at the whole thing, measure the noise from a quiet section, take the average, look ahead.

In real time you have none of that. You get twenty milliseconds of sound and must do something with it. Then the next twenty arrive, ready or not.

That single constraint rules out many otherwise good methods. An algorithm needing the whole recording cannot be used live, no matter how well it works offline.

How it works

   microphone
       |
       v
  20 ms blocks, arriving forever
       |
       v
  [ ring buffer ]     keeps the last second, throws away older audio
       |
       v
  [ is anyone talking? ]     a cheap check on every block
       |
       v
  [ recognition ]     runs on speech, skipped during silence
       |
       v
  [ has the person stopped? ]   <-- this decision is the slow part
       |
       v
  [ answer, then speak it ]

Two ideas do most of the work here.

The ring buffer is a fixed-size store that overwrites its oldest content. Memory usage stays flat no matter how long the conversation runs. A pipeline that keeps everything eventually falls over, and it always happens to a real user rather than in testing.

Voice activity detection, or VAD, is a cheap check asking "is this block speech or silence?" Skipping the expensive work during silence is what keeps the whole thing affordable.

The thing that actually makes it feel slow

Here is what surprises people who build their first voice assistant.

When you stop speaking, the machine does not know you stopped. A pause could mean "I am finished" or "I am thinking of the word". The only way to tell is to wait and see whether more speech arrives.

That wait is called endpointing. It is usually the single largest part of the delay before you get an answer. It costs no computing power at all. It is pure waiting.

And it cannot be set to zero. Wait too little and the assistant interrupts you mid-sentence, which feels much worse than a short pause. Wait too long and every reply feels sluggish.

People optimise the wrong thing here constantly. They buy a faster model to save a hundred milliseconds. Meanwhile half a second goes on waiting to see if the user finished. The developer section measures exactly this.

Where you have already used one

  • Voice assistants on speakers and phones.
  • Live captions appearing while someone is still speaking.
  • Video call noise removal, running on every block as it arrives.
  • Voice typing, showing words as you say them.
  • Live translation on a phone call.

What feels fast to a person

Rough guidance for a spoken reply:

  • Under half a second feels immediate, like a person.
  • Half a second to one second feels normal.
  • One to two seconds feels slow but usable.
  • Beyond two seconds feels broken, and people start talking over it.

Human conversation has gaps of around two tenths of a second between turns. That is the standard you are unconsciously being measured against.

What is honestly hard here

Interrupting is expected. People talk over an assistant that is speaking. Handling that means listening while your own speaker is playing, and separating your own voice from theirs. That is called echo cancellation and it is genuinely difficult.

Consistency matters more than speed. A system that answers in 0.8 seconds every time feels better than one averaging 0.5 with occasional three-second stalls. People adapt to a rhythm and notice when it breaks.

Everything gets harder at once. Noise removal, recognition and endpointing all become harder without the ability to look ahead.

Remember this

  • Real-time means processing small blocks as they arrive, with no ability to see the future.
  • A ring buffer keeps memory flat, and VAD keeps cost low by skipping silence.
  • Endpointing — deciding the person stopped — is usually the biggest part of the delay.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install numpy

NumPy only. Nothing here needs a microphone, so it runs identically everywhere and in continuous integration. Real capture uses sounddevice or pyaudio, which need hardware and cannot be tested this way — one reason to keep pipeline logic separate from capture.

A streaming VAD with endpointing

This is the core loop of every voice assistant, minus the models. It processes 20 ms blocks in order and never looks ahead.

streaming_vad.py
import numpy as np
from collections import deque

SR, BLOCK = 16000, 320                      # 20 ms blocks
rng = np.random.default_rng(0)

def speechy(n, f0=140):
    t = np.arange(n)/SR
    return 0.5*np.sin(2*np.pi*f0*t)*(1 + 0.5*np.sin(2*np.pi*4*t))

# "book two tickets ... to mumbai" -- one sentence with a thinking pause in it
PARTS = [("sil",1.0), ("talk",1.2), ("sil",0.35), ("talk",0.8), ("sil",1.2)]
stream = np.concatenate([speechy(int(SR*d)) if k=="talk" else np.zeros(int(SR*d))
                         for k, d in PARTS])
stream = stream + 0.02*rng.standard_normal(stream.size)

def run(end_silence_blocks, speech_th=0.05, preroll=10):
    ring = deque(maxlen=preroll)            # keeps 200 ms of lookback, nothing older
    state, silent_run, segments, start = "idle", 0, [], None
    n_blocks = len(stream)//BLOCK
    for i in range(n_blocks):
        blk = stream[i*BLOCK:(i+1)*BLOCK]
        ring.append(blk)
        voiced = float(np.sqrt(np.mean(blk**2))) > speech_th
        if state == "idle" and voiced:
            state, start, silent_run = "speaking", max(0, i - len(ring) + 1), 0
        elif state == "speaking":
            silent_run = 0 if voiced else silent_run + 1
            if silent_run >= end_silence_blocks:
                segments.append((start, i - end_silence_blocks + 1))
                state, start = "idle", None
    if state == "speaking":
        segments.append((start, n_blocks))
    return segments

print(f"stream: {len(stream)/SR:.2f} s, blocks of {BLOCK/SR*1000:.0f} ms")
print("the speaker pauses for 350 ms in the middle of one sentence\n")
print(f"{'endpoint wait':>14}{'turns found':>13}   segments (s)")
for blocks in (5, 10, 20, 25, 40):
    segs = run(blocks)
    txt = "  ".join(f"{a*BLOCK/SR:.2f}-{b*BLOCK/SR:.2f}" for a, b in segs)
    print(f"{blocks*BLOCK/SR*1000:>11.0f} ms{len(segs):>13}   {txt}")
Output
stream: 4.55 s, blocks of 20 ms
the speaker pauses for 350 ms in the middle of one sentence

 endpoint wait  turns found   segments (s)
        100 ms            2   0.82-2.20  2.36-3.36
        200 ms            2   0.82-2.20  2.36-3.36
        400 ms            1   0.82-3.36
        500 ms            1   0.82-3.36
        800 ms            1   0.82-3.36

That table is the endpointing dilemma

The user spoke one sentence with a 350 ms thinking pause in the middle.

At 100 and 200 ms the pipeline finds two turns. It decided the user finished, cut them off, and would have answered half a question. Anyone who has been interrupted by a smart speaker has met this setting.

At 400 ms and above it correctly finds one turn. The cost is that every reply now waits at least 400 ms after the user genuinely stops.

There is no setting that is right for everyone. Fast speakers want 300 ms. Someone composing a difficult sentence needs 800. Production systems make this adaptive — longer after a hesitation sound, shorter after a complete-sounding phrase, and shorter still when the recogniser's partial text already parses as a finished request.

Line by line, the parts that are not obvious

The pre-roll ring buffer. Segments begin at 0.82 s, though speech starts at 1.0 s. That is deliberate. By the time a VAD is confident, the first consonant is already past, so the buffer hands back the preceding 200 ms. Without pre-roll your recogniser receives "ook two tickets" and no amount of model quality repairs it. deque(maxlen=...) gives this for free and keeps memory constant however long the session runs.

silent_run reset on every voiced block. The counter tracks consecutive silence, not total silence. Getting this wrong makes long utterances endpoint at random.

Energy-based VAD is the weakest usable choice. It fails in a noisy room, where the noise floor exceeds speech_th. Production systems use a trained VAD such as Silero or WebRTC's. The state machine around it stays exactly the same, which is why it is worth building once by hand.

No lookahead anywhere. The loop only ever reads block i. That constraint is what makes this code deployable, and it is what rules out the noise estimation used in noise reduction, which learned its profile from a lead-in the live pipeline does not have.

Where the time actually goes

Now account for the whole response. This is arithmetic, and it is the most useful thing in this lesson.

latency_budget.py
BLOCK_MS = 20

STAGES = [                       # (name, milliseconds, does it run while the user talks?)
    ("microphone block",            BLOCK_MS, True),
    ("network to server",                 40, True),
    ("streaming ASR, per block",          15, True),
    ("endpoint wait (silence)",          400, False),
    ("ASR finalisation",                  60, False),
    ("language model, first token",      280, False),
    ("TTS, first audio chunk",           120, False),
    ("network back to device",            40, False),
    ("playback buffer",                   60, False),
]

overlapped = sum(ms for _, ms, live in STAGES if live)
blocking   = sum(ms for _, ms, live in STAGES if not live)

print(f"{'stage':<30}{'ms':>7}   counts toward the wait?")
for name, ms, live in STAGES:
    print(f"{name:<30}{ms:>7}   {'no (runs while talking)' if live else 'YES'}")
print(f"\n{'work hidden under speech':<30}{overlapped:>7} ms")
print(f"{'time the user actually waits':<30}{blocking:>7} ms")

print("\nwhere the wait goes:")
for name, ms, live in STAGES:
    if not live:
        bar = "#" * round(ms/blocking*40)
        print(f"  {name:<28}{ms:>5} ms {100*ms/blocking:5.1f}%  {bar}")

print(f"\nhalving the language model to 140 ms  -> {blocking-140} ms wait")
print(f"halving the endpoint wait to 200 ms   -> {blocking-200} ms wait")
Output
stage                              ms   counts toward the wait?
microphone block                   20   no (runs while talking)
network to server                  40   no (runs while talking)
streaming ASR, per block           15   no (runs while talking)
endpoint wait (silence)           400   YES
ASR finalisation                   60   YES
language model, first token       280   YES
TTS, first audio chunk            120   YES
network back to device             40   YES
playback buffer                    60   YES

work hidden under speech           75 ms
time the user actually waits      960 ms

where the wait goes:
  endpoint wait (silence)       400 ms  41.7%  #################
  ASR finalisation               60 ms   6.2%  ##
  language model, first token   280 ms  29.2%  ############
  TTS, first audio chunk        120 ms  12.5%  #####
  network back to device         40 ms   4.2%  ##
  playback buffer                60 ms   6.2%  ##

halving the language model to 140 ms  -> 820 ms wait
halving the endpoint wait to 200 ms   -> 760 ms wait

Those numbers are plausible placeholders for a cloud assistant, not measurements. Replace every one with a figure from your own system. The structure is the transferable part, not the values.

Two conclusions worth taking seriously

Streaming is what makes work free. The first three stages total 75 ms and cost the user nothing, because they finish while the person is still talking. A batch pipeline that waits for the recording adds all of it to the wait. Streaming does not make anything faster; it moves work to where nobody is waiting.

The biggest single term is a timer. Endpointing is 41.7 percent of the wait and uses no computation. Halving the language model — the expensive, difficult optimisation everyone reaches for — saves 140 ms. Halving the endpoint wait saves 200 ms and needs no new hardware.

Measure before optimising, or you will spend a month on the second-biggest term.

The failure that ends pipelines

If processing one 20 ms block takes longer than 20 ms, you are producing backlog faster than you clear it. The queue grows without bound and latency climbs steadily until something breaks.

This does not show up in a quick test, because a short run hides slow growth. It appears after twenty minutes, in production. Two defences:

Measure the real-time factor per block and log it. Anything above about 0.5 has no safety margin for a garbage collection pause or a busy machine.

Drop, never queue. When you fall behind, discard blocks and carry on with current audio. Stale audio is worthless — nobody wants a correct answer to what they said forty seconds ago. An unbounded queue turns a small slowdown into total failure.

Common mistakes

Choosing the block size carelessly. Small blocks mean low latency and high per-block overhead. Large blocks are efficient and add delay. 10 to 30 ms is the usual range; 20 ms is a sound default.

Doing work in the audio callback. The capture callback must return before the next block arrives. Copy the data into a queue and process on another thread. Model inference inside a callback causes dropouts.

Forgetting pre-roll. Demonstrated above. Words get their beginnings clipped and accuracy falls for reasons nobody can locate.

Reporting mean latency. Users experience the bad cases. Report p95 and p99. A pipeline averaging 500 ms with a p99 of four seconds feels unreliable, correctly.

Ignoring barge-in. People interrupt. If you do not keep listening while speaking, your assistant talks over its user. That needs acoustic echo cancellation so your own output is removed from the microphone signal.

Assuming the offline model works live. Whisper processes 30-second windows, so it is not a streaming recogniser without a chunking layer around it. See Whisper. Streaming systems use RNN-T or chunked-attention models built for it.

Try it yourself

Change PARTS to give the speaker a 700 ms pause and find the endpoint setting that keeps one turn. Then add a min_speech_blocks rule requiring three consecutive voiced blocks before entering the speaking state, and check whether it stops a single noise burst from triggering the pipeline. That one rule removes a large share of spurious activations.

What to learn next

  • Model serving — running these models behind an API that holds its latency.
  • Latency and throughput — the general trade this whole lesson is one case of.
  • Whisper — why an offline recogniser needs work before it can stream.

Researcher — Mathematics and papers.

Latency decomposition

Total response latency for a spoken-dialogue turn:

$$ L = \underbrace{L_{\text{buf}} + L_{\text{net}}^{\uparrow} + L_{\text{ASR}}^{\text{stream}}}_{\text{overlapped with speech}}

  • \underbrace{L_{\text{end}} + L_{\text{ASR}}^{\text{final}} + L_{\text{LLM}}^{\text{TTFT}} + L_{\text{TTS}}^{\text{first}} + L_{\text{net}}^{\downarrow} + L_{\text{play}}}_{\text{perceived}} $$

Only the second group is perceived, because the first completes before the user stops speaking. This decomposition is the reason streaming architectures win: they do not reduce total work, they relocate it under the user's own speech.

The binding constraint for sustainability is per-block real-time factor:

$$ \text{RTF} = \frac{t_{\text{process}}}{t_{\text{audio}}} < 1 $$

RTF $\geq 1$ makes the queue a non-stationary process with unbounded expected length. Under variable service times, queueing theory gives the sharper practical bound: expected delay grows as $\rho/(1-\rho)$ with utilisation $\rho$, so latency degrades severely well before $\rho = 1$. Target $\rho \lesssim 0.5$, and use a bounded queue with a drop policy so overload degrades gracefully rather than catastrophically.

Streaming recognition

Offline attention encoder-decoders attend over the full utterance and cannot stream. The streaming-capable families:

RNN-T emits without future context by construction, and is the dominant on-device choice (He et al., 2019).

Monotonic chunkwise attention (MoChA, Chiu and Raffel, 2018) restricts attention to a bounded chunk, giving a tunable latency-accuracy trade.

Chunked and cached Conformer processes fixed chunks with cached left context and limited right context. Right-context lookahead $R$ buys accuracy at exactly $R$ milliseconds of latency, and the exchange rate is the key design curve. Emformer (Shi et al., 2021) makes this efficient by caching memory across chunks.

Whisper is not streaming. Its 30-second fixed window makes chunking an external heuristic, with boundary artefacts and re-decoding cost. Treat streaming Whisper wrappers as a compromise, not a streaming model.

Latency-quality can also be shaped by the loss. FastEmit (Yu et al., 2021) adds a regulariser penalising blank emission in RNN-T, encouraging earlier token emission and cutting partial-result latency substantially at a small WER cost.

Endpointing

The developer block shows fixed-threshold endpointing and its failure. Better formulations:

VAD-based with adaptive thresholds conditioned on observed speech rate and hesitation.

End-of-query classifiers (Maas et al., 2018; Chang et al., 2019) predict turn completion from acoustic and decoder state, using the partial hypothesis. A partial that parses as a complete request licenses a shorter wait than one ending in a determiner.

End-of-utterance as a token. Train the recogniser to emit an explicit </s>, folding endpointing into decoding and removing the separate timer.

The metric is a two-sided cost: early cutoffs truncate user speech, and late endpoints add latency. Report both — median endpoint latency and cutoff rate — because either one alone is optimised to nothing by moving the threshold.

Full-duplex and barge-in

Half-duplex — listen, then speak, then listen — is a simplification human conversation does not obey. Turn-taking gaps in natural conversation centre near 200 ms (Stivers et al., 2009), which is shorter than most pipelines' endpoint wait alone. Overlap is normal, not exceptional.

Barge-in requires acoustic echo cancellation: an adaptive filter estimating the loudspeaker-to-microphone path and subtracting the known reference. Difficulty comes from nonlinear loudspeaker distortion, clock drift between playback and capture, and echo path changes when anything in the room moves. Residual echo suppression is a separate post-filter stage, and it is where most of the practical engineering sits.

Recent full-duplex dialogue models fold listening and speaking into one continuously running model, learning when to yield rather than being told by a timer. This is an active area rather than settled practice.

Transport

Opus (RFC 6716) is the codec for interactive voice: 2.5 to 60 ms frames, 6 to 510 kbit/s, seamless speech-music switching, and built-in forward error correction and packet-loss concealment. Below roughly 30 ms frames, per-packet header overhead grows significant.

Jitter buffers trade latency for loss concealment. Adaptive buffers track observed jitter and are the standard, but note the interaction: an adaptive buffer that grows during congestion silently adds to $L_{\text{net}}$ and to perceived latency, so it must be monitored as part of the budget rather than treated as transport-layer detail.

WebRTC bundles echo cancellation, VAD, jitter buffering and Opus, and reimplementing any of it is rarely justified.

Measurement

Instrument per-stage timestamps on a single request identifier and export histograms, not averages. Report p50, p95 and p99 for each stage plus end to end. Tail latency is what users perceive, and a stage with a well-behaved mean and a heavy tail — garbage collection, cold model load, a cache miss — is invisible in averages and painfully visible to a user.

Track RTF distribution per block, queue depth over time, and drop counts. Rising queue depth is the leading indicator of the failure described in the developer block, and it is observable minutes before latency becomes visible.

Reading

What to learn next

  • Model serving — running these models behind an API that holds its latency.
  • Latency and throughput — the general trade this whole lesson is one case of.
  • Whisper — why an offline recogniser needs work before it can stream.

What to learn next

These follow on from what you just read.

  • Time Series and Forecasting

    What is time series data?

    Time series data is data where the order of the rows is part of the information, so you can never shuffle it into a random train and test split.

  • Time Series and Forecasting

    Trend, seasonality and noise

    Almost every time series is three things added together — a slow direction, a repeating pattern and leftover randomness — and separating them is the first thing you do.

  • Time Series and Forecasting

    Stationarity

    A stationary series keeps the same average and the same wobble forever, and most forecasting models quietly assume yours does too.