Generative AI and LLMs Section 074

Multimodal AI

Models that handle text, images, audio and video together instead of one at a time.

12 of 12 lessons published Three reading levels on every lesson

Start with “What is multimodal AI?”

Lessons in order

Work top to bottom. Each lesson assumes the one above it.

  1. What is multimodal AI?
  2. Multimodal embeddings
  3. CLIP
  4. Vision-language models
  5. Image captioning
  6. Visual question answering
  7. Text to image
  8. Text to video
  9. Document AI
  10. Cross-modal retrieval
  11. Audio-visual learning
  12. Multimodal RAG