Definition
Plain language
An approach to learning from video by predicting the meaning of what comes next instead of redrawing the exact pixels.
As stated in the literature
A joint-embedding predictive architecture that trains on video by predicting future frames in latent space rather than reconstructing pixels; the lineage Orca extends with language steering and backward prediction.
Also called: JEPA
Why it matters: Predicting meaning rather than exact pixels lets a model learn what actually matters in a scene without wasting effort reproducing every visual detail.
For example, instead of redrawing the exact next video frame, it predicts that 'the ball will be mid-air on the right' at the meaning level.
Heard on the show
“And that's the V-JEPA lineage, right?”Episode 213 — A Model Learned to Control a Robot by Watching Video It Never Acted On