Glossary · Term

V-JEPA

← all terms

Definition

Plain language

An approach to learning from video by predicting the meaning of what comes next instead of redrawing the exact pixels.

As stated in the literature

A joint-embedding predictive architecture that trains on video by predicting future frames in latent space rather than reconstructing pixels; the lineage Orca extends with language steering and backward prediction.

Also called: JEPA

Why it matters: Predicting meaning rather than exact pixels lets a model learn what actually matters in a scene without wasting effort reproducing every visual detail.

For example, instead of redrawing the exact next video frame, it predicts that 'the ball will be mid-air on the right' at the meaning level.

Heard on the show

“And that's the V-JEPA lineage, right?”
Episode 213 — A Model Learned to Control a Robot by Watching Video It Never Acted On

Mentioned in 1 episode

  1. 213
    A Model Learned to Control a Robot by Watching Video It Never Acted On

Related terms