Glossary · Term

multimodal

← all terms

Definition

Plain language

AI systems that handle more than one kind of input — like text and images together — instead of just one.

As stated in the literature

Models trained on or operating over multiple data modalities (text, image, audio, video, action), often using a unified representation space across modalities.

Why it matters: Many real tasks span text plus images, audio, or video, and a single multimodal model avoids the seams of stitching specialists together.

For example, a multimodal model can look at a photo of a fridge and tell you what meals you could make with what's inside.

Heard on the show

“The Planner is a multimodal model.”
Episode 247 — One Edited Photo, an Honest Caption, and a RAG System That Believes It

Mentioned in 9 episodes

  1. 247
    One Edited Photo, an Honest Caption, and a RAG System That Believes It
  2. 209
    How 2.6 Billion Doodles Exposed the Culture Words Quietly Delete
  3. 167
    How Teaching an AI to Predict, Not Act, Made It a Better Actor
  4. 157
    When an AI Coding Agent Drives a Phone Through the Terminal, No Screen Needed
  5. 132
    The Agent Failed — But Did the Instructions Deserve to Be Followed?
  6. 127
    What Diffusion Language Models Were Missing: A Map, Not an Algorithm
  7. 080
    How a Two-Agent Trick Unlocked Large-Scale Training for Computer-Use Agents
  8. 066
    Why Giving an AI Agent More Tools Can Make It Worse at Using a Computer
  9. 027
    When AI Agents Build the Serving Stack: A Bet on Bespoke Infrastructure

Related terms