AI glossary

Multimodal

In one sentence A multimodal model handles more than one kind of input or output — text, images, audio, video — inside a single model.

By Updated

Multimodal means a model that works across multiple kinds of data — reading images, hearing audio, or producing them — rather than text alone.

A "modality" is one channel of information: text, images, audio, video. Humans are natively multimodal — you watch a cooking video and effortlessly fuse the visuals, the sizzle, the speech and the on-screen captions into one understanding. Classic ML built a separate specialist per channel: one model reads, another sees, a third listens, and none can discuss what another perceived. A multimodal model merges the channels the way you do: show it a photo of your electricity bill and ask about it in words.

The engineering insight is unification into one sequence. Everything becomes tokens or token-like embeddings in a shared space: an image is cut into patches and each patch embedded, audio becomes frame embeddings, and these slot into the same transformer sequence as words. The model attends across modalities exactly as it attends across words — which is how it connects "the rash in the photo" to your written question about it.

Today this is the default for frontier models — GPT-4o, Gemini and Claude accept images natively; several handle audio and video. Everyday uses already in your pocket: Google Lens, photographing a maths problem for help, voice assistants that see the camera feed, and screen-reading accessibility tools. Caveats carry over from text: models can hallucinate visual details with full confidence — reading a blurry meter reading wrongly, but fluently.

Where to go next