Glossary · Term

vision-language model

← all terms

Definition

Plain language

An AI that can take in both pictures and text and reason about them together.

As stated in the literature

A model trained over paired image and text inputs that produces text outputs; the backbone for GUI and computer-use agents that read screenshots, and for unified multimodal generation systems.

Also called: vision-language models, VLM, VLMs

Why it matters: It lets AI handle tasks that mix pictures and words, like reading a screenshot or describing an image, which text-only systems simply cannot do.

For example, you can show such a model a photo of your open refrigerator and ask what meals you could make from what's inside.

Heard on the show

“You can make a vision-language model better at spotting tiny details in big images by never letting it see the details.”
Episode 242 — Making a Vision Model Better by Showing It Blurry Images

Mentioned in 6 episodes

  1. 242
    Making a Vision Model Better by Showing It Blurry Images
  2. 156
    Why More Human Demonstrations Made a Computer-Use Agent Worse
  3. 115
    Teaching a Phone Agent to Reason Silently, And Keeping It Honest
  4. 111
    How a 4B Web Agent Beat Models 60x Its Size on 500 Demonstrations
  5. 027
    When AI Agents Build the Serving Stack: A Bet on Bespoke Infrastructure
  6. 017
    When the Agent Grades Its Own Homework: A Brutal New Benchmark for AI Workers

Related terms