Glossary · Term

multimodal RAG

← all terms

Definition

Plain language

AI search that pulls up pictures as well as text and answers from them.

As stated in the literature

Retrieval-augmented generation over a corpus containing images or visual documents, where retrieved images are placed directly into a vision-language model's context; the retrieval step may use a shared image-text embedding space or generated captions.

Also called: multimodal retrieval, multi-modal RAG

Why it matters: It extends AI search to everything that isn't plain text — charts, scans, product photos — while making the answer depend on whether the retrieved picture is honest.

For example, asking an assistant which plant is in your garden makes it search a photo library, pull up a matching leaf image, and answer from what it sees.

Heard on the show

“So the reframe is: this isn't proof that multimodal RAG is broken.”
Episode 247 — One Edited Photo, an Honest Caption, and a RAG System That Believes It

Related terms