Multimodal AI

Audio-visual learning

Audio-visual learning uses the fact that sound and picture happen at the same moment as a free training signal, with no human labels needed.

On this page 9
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. How it works
  5. What it is used for
  6. Where you have already seen it
  7. What is honestly hard here
  8. Remember this
  9. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

Audio-visual learning uses sound and picture as answers for each other.

Anything you can see happening, you can usually also hear. That pairing is the training signal.

The analogy you have already lived

Think about watching a badly dubbed film. The lips finish moving and the words arrive a moment later. You notice instantly, and it is unpleasant.

Nobody taught you the rule. You learned it from a lifetime of seeing a door slam and hearing it at the same instant.

That pairing is everywhere, and it is free. Every video ever made contains a picture and a sound that agree with each other about what happened.

A model can learn from that agreement with no human labelling at all.

Why it exists

Labelling data is the expensive part of AI. Somebody has to watch a clip and type what is in it.

But videos come with their own answer key attached. The sound tells you something about the picture, and the picture tells you something about the sound.

So you can set the model a question that grades itself. Do this sound and this picture come from the same moment? The answer is already known, for every video in the world, at no cost.

This is called self-supervised learning — training on a question that grades itself. The answer comes from the data's own shape, not from a person.

How it works

The classic setup is a matching game across time.

   video frames  ->  [ vision encoder ]  ->  vector
                                                \
                                                 compare
                                                /
   sound clip    ->  [ audio encoder  ]  ->  vector

   same moment  -> should score HIGH
   different moment or different video -> should score LOW

Train that, and something surprising happens. Nobody said the word "guitar". Yet the vision encoder learns to notice guitars, because guitar sounds keep arriving with guitar pictures.

The model learns object categories from correlation alone. That result is what made this area interesting.

What it is used for

Lip sync checking. Measure whether the sound lines up with the mouth, and by how much.

Separating voices. When two people speak at once, watching their mouths helps decide which sound belongs to whom. This works better than audio alone, and it is not a small improvement.

Finding what made a sound. Highlight the region of the picture that a sound came from, with no boxes ever drawn by a human.

Better recognition in noise. In a loud room, adding lip movement to speech recognition recovers accuracy that audio alone loses.

Where you have already seen it

  • Video calls that reduce background noise while keeping the speaker's voice.
  • Automatic captions that stay aligned with the speaker.
  • Editing tools that flag when audio has drifted out of sync.
  • Phone cameras that focus audio on whoever is on screen.

What is honestly hard here

Silence and off-screen sound. A great deal of what you hear was made by something outside the frame. Music added later has no visual source at all. The assumption that sound and picture agree is often false.

Small delays are real. Sound and light do not arrive together in the recording. A microphone twenty metres from the stage records sound noticeably late. A model must tolerate this without treating it as a mismatch.

Edited video breaks it. Cuts, dubbing and background music all sever the natural link. Datasets need heavy filtering before the free supervision is actually free.

Remember this

  • Sound and picture from the same moment are a free training signal.
  • Models learn objects from correlation, with no labels written by anyone.
  • The link breaks on edited video, added music and off-screen sounds.

What to learn next

Developer — Code and libraries.

The core of audio-visual learning is measuring whether two streams agree over time. This example builds a synthetic clip with a known offset and recovers it. The same measurement then picks the right soundtrack out of three.

Setup

bash
pip install numpy

No audio files, no downloads. Runs instantly.

Finding the sync offset

av_sync.py
import numpy as np

rng = np.random.default_rng(2)
FPS, SECONDS = 25, 8
N = FPS * SECONDS

def event_track(times, length=N, decay=0.5):
    """A spike at each event time that fades away, like a clap or a bounce."""
    track = np.zeros(length)
    for t in times:
        idx = int(t * FPS)
        for k in range(12):
            if idx + k < length:
                track[idx + k] += decay ** k
    return track

CLAPS = [0.6, 2.1, 3.0, 5.4, 6.8]
audio_energy = event_track(CLAPS) + 0.15 * rng.normal(size=N)     # loudness per frame
TRUE_OFFSET = 6                                                    # video runs 6 frames late
mouth_motion = np.roll(event_track(CLAPS), TRUE_OFFSET) + 0.15 * rng.normal(size=N)

def standardise(x):
    return (x - x.mean()) / x.std()

def best_offset(a, v, search=25):
    a, v = standardise(a), standardise(v)
    scores = {d: float(np.corrcoef(a[: N - abs(d)], np.roll(v, -d)[: N - abs(d)])[0, 1])
              for d in range(-search, search + 1)}
    return max(scores, key=scores.get), scores

found, scores = best_offset(audio_energy, mouth_motion)
print(f"true offset: {TRUE_OFFSET} frames   recovered: {found} frames "
      f"({found / FPS * 1000:.0f} ms)")
print("correlation near the peak:",
      {d: round(scores[d], 2) for d in range(found - 2, found + 3)})

# The real training task: given one video, which of three audio tracks belongs to it?
tracks = {
    "matching clip":  audio_energy,
    "different claps": event_track([0.2, 1.4, 4.9, 6.1]) + 0.15 * rng.normal(size=N),
    "room noise":      0.15 * rng.normal(size=N),
}
print("\nwhich audio goes with this video?")
for name, track in tracks.items():
    d, sc = best_offset(track, mouth_motion)
    print(f"  {name:>16}: best correlation {sc[d]:+.2f} at offset {d:+d}")
Output
true offset: 6 frames   recovered: 6 frames (240 ms)
correlation near the peak: {4: 0.2, 5: 0.39, 6: 0.54, 7: 0.2, 8: 0.03}

which audio goes with this video?
     matching clip: best correlation +0.54 at offset +6
   different claps: best correlation +0.29 at offset +23
        room noise: best correlation +0.17 at offset -11

What the two blocks demonstrate

The offset was recovered exactly, and the peak is sharp. It reads 0.54 at the right offset, and 0.20 one frame either side. That sharpness is what makes sync detection possible at all. A broad, flat peak would mean the offset is unidentifiable.

The matching track wins, and the margin is what matters. The correct audio scores 0.54; wrong claps score 0.29; noise scores 0.17. Notice that the wrong tracks do not score zero. Any two spiky signals correlate somewhat at their best alignment. Searching over 51 offsets gives 51 chances to find an accidental match.

That is the whole problem of contrastive audio-visual training in one number. The model has to learn a margin, not a threshold, and the negatives are never at zero.

Look also at where the wrong tracks peaked: offsets of +23 and -11, at the far edges of the search window. A correct match peaks near zero drift; a spurious one peaks wherever it happens to. Checking where the peak sits is a cheap extra filter.

Line by line, the parts that are not obvious

event_track builds an exponentially decaying spike. That is what an energy envelope looks like after a clap, a footstep or a bounce. Real pipelines compute this from a spectrogram; see waveforms and spectrograms.

standardise removes mean and scale from both signals. Without it, a loud recording and a dim video would correlate badly for reasons that have nothing to do with sync.

np.roll(v, -d)[: N - abs(d)] shifts the video track by d frames and trims the wrapped section. Forgetting the trim is a real bug: the wrapped tail creates fake correlations at large offsets.

np.corrcoef(...)[0, 1] pulls the off-diagonal element of the 2-by-2 correlation matrix, which is the Pearson correlation between the two series.

The search range of +/- 25 frames is one second either way at 25 fps. Broadcast standards keep audio within roughly 40 ms ahead to 60 ms behind the picture, so a one-second window is generous. Widening it always finds a higher correlation, and always by accident.

How real systems scale this up

The correlation above is hand-written. A trained model replaces it with two encoders and a learned similarity. That is CLIP again, with time as the axis negatives are drawn along.

  • SyncNet trains on short windows of mouth crops against audio, with mismatched offsets as negatives. It is still the basis of most lip-sync tooling.
  • L3 / Look, Listen and Learn trains on whole frames against one-second audio clips, with negatives from other videos. Its vision features transfer to object recognition, which was the striking result.
  • CLAP aligns audio with text rather than with images, giving zero-shot sound classification the way CLIP gives zero-shot image classification.
  • ImageBind binds six modalities to the image space, using image-paired data only. Audio and text end up aligned without ever being paired directly.

Those checkpoints are large. CLAP-style models run to a gigabyte or more, and ImageBind is several gigabytes. Check the model card before pulling one on a limited connection, and start with the correlation approach above if all you need is sync.

Common mistakes

Choosing negatives from the same video. Two moments from one clip share background, lighting and room acoustics. The model learns to recognise the recording, not the event. Draw most negatives from other videos.

Ignoring the silent majority. A large share of video frames have no informative sound. Training on them teaches nothing and dilutes the gradient. Filter by audio energy variance before training.

Assuming the sound came from the frame. Voiceover, background music and off-screen noise all violate the core assumption. Datasets like AudioSet and VGGSound apply heavy filtering for exactly this reason.

Sub-frame precision expectations. At 25 fps, one frame is 40 ms. You cannot recover an offset finer than that from frame-level features. Use a finer audio hop and interpolate the correlation peak if you need better.

Try it yourself

Add + 0.6 * event_track([1.0, 4.0]) to mouth_motion only, simulating two visual events with no sound — someone entering the frame. Rerun and watch the peak correlation drop while the recovered offset stays correct. That robustness is exactly what makes cross-correlation the right first tool here.

What to learn next

Researcher — Mathematics and papers.

The correspondence objective

Audio-visual self-supervision comes in two closely related forms.

Audio-Visual Correspondence (AVC). Binary classification: does this audio clip come from this video? Positives are aligned pairs, negatives are drawn from different videos. Arandjelovic and Zisserman (2017), Look, Listen and Learn, showed that a network trained only on this task learns visual features competitive with supervised pretraining on object recognition, and audio features that set the state of the art on ESC-50 at the time.

Audio-Visual Temporal Synchronisation (AVTS). Harder: positives are aligned, negatives are the same video shifted in time. This forces temporal precision rather than semantic co-occurrence, and it is what lip-sync models optimise. Korbar, Tran and Torresani (2018) showed curriculum learning matters here — starting with easy negatives from other videos and moving to hard shifted negatives.

The distinction is worth keeping straight. AVC learns "what", AVTS learns "when", and a model trained for one is not automatically good at the other.

Sound source localisation

Because the correspondence network scores whole clips, its spatial attention can be read out as a localisation map with no bounding boxes ever supplied. Objects that Sound (Arandjelovic and Zisserman, 2017) made this explicit with an architecture whose similarity is computed per spatial location, so the argmax over locations is a free detector.

The known failure is that the map often highlights context rather than the source — the stage rather than the guitar. Later work adds hard negative mining within the image and explicit background modelling to sharpen it.

Audio-visual speech

This is the area with the largest measured gains from adding vision.

Speech separation. Looking to Listen (Ephrat et al., 2018) conditions a separation network on face crops, giving large SDR improvements over audio-only separation on overlapping speech, and solving speaker assignment for free — the output is tied to a face, so there is no permutation ambiguity.

Speech recognition in noise. AV-HuBERT (Shi et al., 2022) masks and predicts cluster assignments across both streams. Under babble noise, audio-visual word error rate is a fraction of audio-only, and the gap widens as signal-to-noise ratio falls. See speech recognition for the audio-only baseline this improves on.

Lip reading. Video-only recognition is possible and poor, because visemes — visually distinct mouth shapes — are far fewer than phonemes. Several phonemes map to one viseme, so the mapping is not invertible from vision alone.

Binding many modalities

ImageBind (Girdhar et al., 2023) trains audio, depth, thermal and IMU encoders to align with a frozen image encoder, using only image-paired data for each. The result is emergent alignment between modality pairs that were never trained together — audio to text works despite no audio-text pairs being used.

The image modality acts as a hub. This is efficient and it inherits every limitation of the image space, including the modality gap discussed in what is multimodal AI.

Datasets

  • AudioSet (Gemmeke et al., 2017): about 2 million 10-second YouTube clips with 527 sound classes, weakly labelled. The standard pretraining corpus, and the label noise is substantial.
  • VGGSound (Chen et al., 2020): roughly 200,000 clips, curated so the sound source is visible in the frame. Smaller and cleaner, and better suited to correspondence training for exactly that reason.
  • LRS2 / LRS3: BBC and TED talks with aligned transcripts, the standard audio-visual speech benchmarks.
  • AVE (Tian et al., 2018): audio-visual event localisation with temporal annotations.

What breaks the free-supervision assumption

Measured failure sources, in rough order of frequency in web video: post-production music, voiceover narration, off-screen sound sources, shot cuts within a clip, and constant systematic audio delay from encoding pipelines. Any of these makes a positive pair effectively negative.

This is why curation dominates. VGGSound's contribution over AudioSet is filtering, not scale, and models trained on it transfer better per hour of video despite being ten times smaller.

Papers

What to learn next