Definition
Plain language
One small tile of a picture, after the picture has been chopped into a grid for the model to read.
As stated in the literature
Fixed-size region of an input image mapped to one or more visual tokens by the vision encoder; patch count scales with resolution, so larger images consume proportionally more of the context window.
Also called: image patches, patches, visual patch, visual patches
Why it matters: Patch counts determine how much of a model's limited attention budget an image consumes, so a high-resolution picture can crowd out the text you actually wanted answered.
For example, a photo handed to a model might be sliced into a grid of small squares, each one becoming a piece the model reads much like a word.
Heard on the show
“The model patches its own weaknesses across phases.”Episode 008 — Why Long-Horizon AI Agents Get Stuck, and a Milestone-Based Fix That Helps