Glossary · Term

multimodal

← all terms

Definition

Plain language

AI systems that handle more than one kind of input — like text and images together — instead of just one.

As stated in the literature

Models trained on or operating over multiple data modalities (text, image, audio, video, action), often using a unified representation space across modalities.

Why it matters: Many real tasks span text plus images, audio, or video, and a single multimodal model avoids the seams of stitching specialists together.

For example, a multimodal model can look at a photo of a fridge and tell you what meals you could make with what's inside.

Heard on the show

“… Imagine you want to deploy a model — let's say one of those new multimodal ones that interleaves autoregressive text and diffusion image steps in a single forward pass — …”
Episode 027 — When AI Agents Build the Serving Stack: A Bet on Bespoke Infrastructure

Related terms