Course progress Course outline 34 of 34 lessons available
Part 0 — Get the map
Part 1 — Why a small objective can learn to see
- 05 Chapter 4 — Video can set its own homework available now
- 06 Chapter 5 — Keep the meaning; do not repaint every pixel available now
- 07 Chapter 6 — Match the cards, but do not leave every card blank available now
- 08 Chapter 7 — SIGReg checks a cloud by looking at its shadows available now
- 09 Chapter 8 — The whole LeVJEPA objective on one line available now
Part 2 — Send a video through one encoder
- 10 Chapter 9 — How global and local views are paired available now
- 11 Chapter 10 — Cut a video into space-time tiles available now
- 12 Chapter 11 — One encoder, one projector, one summary card available now
- 13 Chapter 12 — One complete trip through the model available now
- 14 Chapter 13 — Why throwing away 95% can help available now
- 15 Chapter 14 — Same-frame teamwork, no peeking into tomorrow available now
- 16 Chapter 15 — RoPE, single-frame tubelets, and unexpectedly useful patch features available now
Part 3 — Read the experiments, not just the headline
- 17 Chapter 16 — What the four ablation ladders actually test available now
- 18 Chapter 17 — Equal epochs are not equal bills available now
- 19 Chapter 18 — What ImageNet, K400, and SSv2 are really asking available now
- 20 Chapter 19 — Keep the paper’s results in a ledger available now
- 21 Chapter 20 — Claims the evidence does not yet earn available now
Part 4 — From the official repository to your own experiment
- 22 Chapter 21 — A map of the official repository available now
- 23 Chapter 22 — Ten long walks become a training set available now
- 24 Chapter 23 — Read the defaults, then start training available now
- 25 Chapter 24 — Run a smoke test that cannot flatter you available now
- 26 Chapter 25 — Skip training: extract features from the public checkpoint available now
- 27 Chapter 26 — Freeze the encoder and test your own videos available now
Part 5 — Put the representation back on the world-model road
Appendices — A backpack for the trail
Names are labels, not bags of spare parts
- FamilyJEPA is a blueprint
- MethodI, V, Le, and LeV differ
- Versionnewer is not stronger evidence
- Evidencepaper, code, weights, reproduction
A hammer, screwdriver, and drill may all live in a carpenter’s toolbox. That does not put a battery inside the hammer. JEPA is similarly a research blueprint and family of methods. Members choose different encoders, target branches, predictors, regularizers, and data.
The most common reading error is not a bad calculation. It is quietly installing I-JEPA’s EMA target encoder, V-JEPA’s predictor, and V-JEPA 2-AC’s action model inside LeVJEPA. This glossary gives every term a small “do not smuggle this in” label.
Start with the names themselves
- The official spelling is LeVJEPA, not
LEVJEPA,Le-VJEPA, orV-JEPA 3. - The LeJEPA paper expands its name as Latent-Euclidean JEPA. It contributes the general objective and SIGReg.
- The LeVJEPA paper does not supply a formal long name to expand. The safe description is “a video encoder trained with the LeJEPA objective.”
- All these papers are collaborations. Saying LeVJEPA follows Yann LeCun’s research road describes an intellectual lineage; it does not erase the work of the other authors.
The working glossary
| Term | One-line meaning | Small picture | Common mix-up |
|---|---|---|---|
| JEPA | An architectural idea that models compatibility or prediction in representation space | guess meaning on a map, not every brick | It does not prescribe one loss or collapse remedy. |
| I-JEPA | Predicts representations of masked target regions from image context | infer the middle of a puzzle | It is a still-image method, not a planner. |
| V-JEPA | Extends latent feature prediction to video space-time blocks | infer a hidden stretch of a film | It uses a predictor and EMA-target route; it is not LeVJEPA. |
| V-JEPA 2 | A large-scale video JEPA representation model | an observer trained on much video | Action-free pretraining does not itself control a robot. |
| V-JEPA 2-AC | Trains an action-conditioned predictor above a frozen encoder | add an “if I move” sandbox | It is a V-JEPA 2 post-training stage, not a LeVJEPA component. |
| LeJEPA | Joins a JEPA pairing objective with SIGReg in a lean general recipe | match same-answer cards without allowing blanks | Its theory has assumptions and is not video-specific. |
| LeVJEPA | Uses the LeJEPA collapse-control objective to train a video encoder | describe one film through large and small windows | It is encoder pretraining, not an action world model. |
| LeWorldModel | Trains next-embedding prediction, SIGReg, and action-conditioned dynamics for control | a sandbox that can try actions | It is a neighbouring branch, not another name for LeVJEPA. |
| self-supervised | Builds a training signal from the data itself | the video writes its own matching question | Humans still choose data, transformations, and objective. |
| representation / embedding | A numeric vector produced for an input | a numeric address card | Individual coordinates usually have no hand-written meaning. |
| encoder | Turns pixels into tokens and representations | a reader writing notes | LeVJEPA’s encoder has no built-in class answer. |
| global view | A broad spatial crop of one clip | a wide window | It is still an augmented crop, not the entire original recording. |
| local view | A smaller spatial crop from the same temporal window | a cardboard viewing tube | Moving it to another time changes the pairing contract. |
| augmentation | Changes a view while intending to preserve its relevant identity | crop a corner or alter brightness | An aggressive transform can erase information the task needs. |
| invariance | Views of the same clip should receive nearby representations | recognize the same act from two windows | Alone, it permits one constant answer. |
| complete collapse | Nearly every input maps to the same vector | the whole class hands in one blank card | A low invariance loss may hide it. |
| dimensional collapse | Embeddings occupy only a low-dimensional subspace | 256 streets squeezed into one alley | Samples can differ and still collapse this way. |
| SIGReg | Constrains embedding distributions through Gaussianity of random projections | inspect a city from many shadows | It does not test physical meaning or controllability. |
| isotropic Gaussian | A zero-centered Gaussian reference with equal scale in every direction | a cloud with no favored compass point | It is neither a uniform sphere nor the topology of the real world. |
| random projection | Maps a high-dimensional point to one scalar with a random unit vector | shine a flashlight from one direction | Finitely many directions are an approximation. |
| characteristic function | The complex-exponential expectation that uniquely describes a distribution | a frequency fingerprint | The code compares only finitely many frequency nodes. |
| Epps–Pulley statistic | A smooth discrepancy between empirical and Gaussian characteristic functions | compare two fingerprints | A training loss is not a formal pass/fail normality certificate. |
| patch | A 16×16 pixel region in one frame | a mosaic tile | It becomes a transformer token only after embedding. |
| token | The vector made from an embedded patch | a block stamped with coordinates | Training randomly drops 95% of patch tokens. |
| tubelet | A spatial patch extended across several frames | stack matching tiles into a tube | LeVJEPA defaults to tau=1, so it does no temporal stacking here. |
[CLS] | A learned token that summarizes the clip | a summary card | Patch tokens cannot read it, preventing future information from flowing back. |
| projector | Maps [CLS] into the 256-D space constrained by SIGReg | a temporary training translation desk | It is discarded after pretraining and is not a classifier. |
| LayerNorm | Normalizes features within each token | press each card onto a constrained surface | Its output geometry is one reason SIGReg follows a projector. |
| RoPE | Encodes relative time and space positions through rotations in attention | stamp time, row, and column on a block | It does not teach causal intervention by itself. |
| block-causal attention | Bidirectional within a frame, current-and-past only across frames | classmates share a page but cannot open tomorrow’s page | [CLS] may see the clip; patches cannot read [CLS]. |
| target encoder | A separate branch that supplies target features during training | a second teacher writing answers | LeVJEPA’s objective has none. |
| stop-gradient | Blocks backpropagation through a branch | a one-way valve | Both LeVJEPA view branches receive gradients. |
| predictor | Estimates a target representation from context | the part that guesses | LeVJEPA aligns a shared encoder’s view summaries directly and has no predictor. |
| Polyak/EMA weights | A moving average copy of training weights used for evaluation | a smoothed copy of the scorebook | The official copy does not make targets or enter the loss. |
| frozen probe | Trains a small readout while leaving the encoder fixed | test the notes without rewriting them | Decodability does not prove the model actively uses the information. |
| attentive probe | Pools tokens with a learned query before classifying | a marker who learns where to look | It is stronger than plain linear mean pooling; scores are not interchangeable. |
| epoch | One pass over a defined dataset | finish one lap of a workbook | The work per sample differs between methods. |
| FLOPs | An estimate of floating-point operations | accounting units for computation | It is not wall time or energy consumption. |
| world model | Predicts how a world state may evolve | an internal sandbox | A temporally causal encoder is only a possible foundation. |
| MPC | Plans a horizon, acts briefly, observes, and replans | take one step and reopen the map | LeVJEPA itself contains no MPC. |
One road with two important forks
| Public date | Primary work | What it advances | What it does not establish |
|---|---|---|---|
| 2006 | Energy-based learning tutorial | Low energy for compatible configurations, high energy for incompatible ones | A JEPA or video recipe |
| 2022-06-27 | A Path Towards Autonomous Machine Intelligence v0.9.2 | A blueprint joining perception, configurable world model, cost, actor, memory, and hierarchical JEPA | A completed autonomous system |
| 2023 | I-JEPA | Representation prediction for large target blocks in still images | Actions or planning |
| 2024 | V-JEPA | Latent prediction for video space-time blocks | A control model |
| 2025-06 | V-JEPA 2 / 2-AC | Large-scale observation pretraining, then a separate action-conditioned stage with MPC | General long-horizon autonomy |
| 2025-11 | LeJEPA v3 | An explicit, scalable SIGReg route around collapse | An assumption-free global optimum for every task |
| 2026-03 | V-JEPA 2.1 | Stronger dense token representations on the original V-JEPA branch | A move to the LeJEPA/SIGReg branch |
| 2026-03 | LeWorldModel v3 | Action-conditioned latent dynamics and planning on the LeJEPA branch | Broad real-world validation from small control studies |
| 2026-08-27 | LeVJEPA v1 | Lean video-encoder pretraining with sparse observation and block-causal attention | An action predictor or planner |
This is a tree, not a sequence of automatic upgrades. V-JEPA 2.1 continues the asymmetric feature-prediction branch. LeWorldModel and LeVJEPA both draw on LeJEPA/SIGReg but pursue action dynamics and general video representation, respectively.
Put an evidence label on every claim
| Label | What it can support | What it cannot support automatically |
|---|---|---|
| method definition | A paper explicitly specifies a graph or objective | The implementation is bug-free or reproducible. |
| author-reported | Authors report a value under a named protocol | An independent team will obtain it. |
| author code | A concrete interface or default can be inspected | Every setting matches every paper experiment. |
| public weights | Feature extraction and downstream evaluation can be run | The training process has been reproduced independently. |
| independent reproduction | A third party repeats a stated scope | The claim holds for all models, data, and tasks. |
| assumption-bound theory | A conclusion holds under listed assumptions | An unconditional guarantee in the physical world. |
At this course’s 2026-09-05 (Asia/Tokyo) evidence cutoff, LeVJEPA remains a newly released v1 preprint. Its results are author-reported. This review found no independent reproduction that would justify a stronger label.
Sources for the family tree
The lineage above was cross-checked against LeCun’s 2022 AMI paper, the I-JEPA CVPR record, V-JEPA at TMLR/OpenReview, the V-JEPA 2 publication page, LeJEPA v3, and LeVJEPA v1. Public dates establish order, not strength of evidence.
Keep the labels on the drawers
- JEPA is a family blueprint; each member’s computation graph must be checked separately.
- LeVJEPA has no official long name to invent. It is a video encoder trained with the LeJEPA objective.
- A later date, public code, and independent reproduction are three different facts.
Name the right machine
- Why can V-JEPA’s predictor not be drawn inside LeVJEPA?
- Why may we say “no EMA target encoder” even though the official implementation keeps EMA weights?
- What is the central task boundary between LeVJEPA and LeWorldModel?
Check the answers
- They use different training graphs; LeVJEPA directly aligns global and local
[CLS]outputs from a shared encoder and uses SIGReg against collapse. - The EMA copy is a smoothed evaluation checkpoint. It neither produces targets nor participates in the loss.
- LeVJEPA learns general video representations; LeWorldModel explicitly learns action-conditioned next-latent states and uses them for planning.