Glossary · Term

LLaVA

← all terms

Definition

Plain language

An openly available AI system that can look at a picture and talk about it.

As stated in the literature

Open-source vision-language model pairing a CLIP-style vision encoder with an instruction-tuned LLM via a learned projection layer; widely used as a small multimodal research baseline.

Why it matters: Because its weights are openly available, it lets researchers inspect and modify a working image-and-text model rather than guessing at a closed service's behavior.

For example, a researcher can run LLaVA on a laptop-class setup, show it a photo of a fridge, and ask what meals the ingredients could make.

Heard on the show

“On a smaller model, LLaVA, thirty-nine prompts change their answer once the canvas is attached, and every single one changes toward harm.”
Episode 277 — The Blank White Square That Swings AI Refusal Rates Fifty Points

Related terms