Glossary · Term

CLIP

← all terms

Definition

Plain language

A system that turns pictures and sentences into numbers in the same space, so a photo and its description land near each other.

As stated in the literature

Contrastive Language-Image Pretraining: dual image and text encoders trained to align matching pairs in a shared embedding space, widely used as the retrieval backbone for image search and multimodal RAG.

Also called: CLIP-style

Why it matters: It lets a search system match words to pictures without anyone hand-labeling every image, which is what makes image search and picture-fetching AI assistants possible.

For example, typing "a golden retriever on a beach" can pull up matching photos because the sentence and the images end up as nearby points in the same numeric space.

Heard on the show

“… Under the shared-embedding pipeline — where pictures and text queries live in one numeric space, CLIP-style — ninety-nine percent of poisoned images sit within a tiny distance of their clean originals. …”
Episode 247 — One Edited Photo, an Honest Caption, and a RAG System That Believes It

Related concepts

Related terms