Definition
Plain language
AI systems that handle more than one kind of input — like text and images together — instead of just one.
As stated in the literature
Models trained on or operating over multiple data modalities (text, image, audio, video, action), often using a unified representation space across modalities.
Why it matters: Many real tasks span text plus images, audio, or video, and a single multimodal model avoids the seams of stitching specialists together.
For example, a multimodal model can look at a photo of a fridge and tell you what meals you could make with what's inside.
Heard on the show
“The Planner is a multimodal model.”Episode 247 — One Edited Photo, an Honest Caption, and a RAG System That Believes It