Definition
Plain language
AI search that pulls up pictures as well as text and answers from them.
As stated in the literature
Retrieval-augmented generation over a corpus containing images or visual documents, where retrieved images are placed directly into a vision-language model's context; the retrieval step may use a shared image-text embedding space or generated captions.
Also called: multimodal retrieval, multi-modal RAG
Why it matters: It extends AI search to everything that isn't plain text — charts, scans, product photos — while making the answer depend on whether the retrieved picture is honest.
For example, asking an assistant which plant is in your garden makes it search a photo library, pull up a matching leaf image, and answer from what it sees.
Heard on the show
“So the reframe is: this isn't proof that multimodal RAG is broken.”Episode 247 — One Edited Photo, an Honest Caption, and a RAG System That Believes It