Glossary · Term

vision encoder

← all terms

Definition

Plain language

The part of an AI that turns an image into a compact set of numbers capturing what's in it.

As stated in the literature

A pretrained network mapping images to embeddings; in Orca a frozen vision encoder supplies the target latents, tethering the learned world-state to an existing semantic space.

Also called: vision encoders

Why it matters: It gives the rest of a system a compact, meaningful handle on images, so downstream parts can reason about what's in a picture without processing raw pixels.

For example, a vision encoder turns a photo of a cat on a couch into a list of numbers that capture 'cat,' 'couch,' and their arrangement.

Heard on the show

“Against the real next screenshot, run through the model's own frozen vision encoder, used as a fixed target.”
Episode 115 — Teaching a Phone Agent to Reason Silently, And Keeping It Honest

Related terms