Definition
Plain language
AI systems that handle more than one kind of input — like text and images together — instead of just one.
As stated in the literature
Models trained on or operating over multiple data modalities (text, image, audio, video, action), often using a unified representation space across modalities.
Why it matters: Many real tasks span text plus images, audio, or video, and a single multimodal model avoids the seams of stitching specialists together.
For example, a multimodal model can look at a photo of a fridge and tell you what meals you could make with what's inside.
Heard on the show
“… Imagine you want to deploy a model — let's say one of those new multimodal ones that interleaves autoregressive text and diffusion image steps in a single forward pass — …”Episode 027 — When AI Agents Build the Serving Stack: A Bet on Bespoke Infrastructure