Definition
Plain language
Automatically writing a sentence or two describing what's in a picture.
As stated in the literature
Generating a natural-language description of an image, typically with a vision-language model; in multimodal retrieval it produces the text that gets indexed, and because it runs before any user question it is query-agnostic and misses locally salient edits.
Also called: captioner, captioners, captioning, caption pipeline
Why it matters: When captions are what gets indexed, anything the caption leaves out is effectively invisible to search — including a small but decisive alteration in the picture.
For example, a system looks at a photo and writes "a red bicycle leaning against a brick wall" so that sentence can be stored and searched later.
Heard on the show
“Not the retriever, not the captioner, not the model answering.”Episode 247 — One Edited Photo, an Honest Caption, and a RAG System That Believes It