Definition
Multimodal models handle more than one modality — text and images, audio, video, action streams — usually by projecting them into a shared representation space. The frontier question is how cleanly capabilities transfer from one modality to another.
Episodes covering this
Worth reading next
Papers we haven't done a deep dive on yet, but would recommend on this topic.
- VideoAgent: Long-form Video Understanding with Large Language Model as Agent
- OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
- A Neural Representation of Sketch Drawings (Sketch-RNN)
- Emu3: Next-Token Prediction is All You Need
- Multimodal Neurons in Artificial Neural Networks