Course progress Course outline 34 of 34 lessons available
Part 0 — Get the map
Part 1 — Why a small objective can learn to see
- 05 Chapter 4 — Video can set its own homework available now
- 06 Chapter 5 — Keep the meaning; do not repaint every pixel available now
- 07 Chapter 6 — Match the cards, but do not leave every card blank available now
- 08 Chapter 7 — SIGReg checks a cloud by looking at its shadows available now
- 09 Chapter 8 — The whole LeVJEPA objective on one line available now
Part 2 — Send a video through one encoder
- 10 Chapter 9 — How global and local views are paired available now
- 11 Chapter 10 — Cut a video into space-time tiles available now
- 12 Chapter 11 — One encoder, one projector, one summary card available now
- 13 Chapter 12 — One complete trip through the model available now
- 14 Chapter 13 — Why throwing away 95% can help available now
- 15 Chapter 14 — Same-frame teamwork, no peeking into tomorrow available now
- 16 Chapter 15 — RoPE, single-frame tubelets, and unexpectedly useful patch features available now
Part 3 — Read the experiments, not just the headline
- 17 Chapter 16 — What the four ablation ladders actually test available now
- 18 Chapter 17 — Equal epochs are not equal bills available now
- 19 Chapter 18 — What ImageNet, K400, and SSv2 are really asking available now
- 20 Chapter 19 — Keep the paper’s results in a ledger available now
- 21 Chapter 20 — Claims the evidence does not yet earn available now
Part 4 — From the official repository to your own experiment
- 22 Chapter 21 — A map of the official repository available now
- 23 Chapter 22 — Ten long walks become a training set available now
- 24 Chapter 23 — Read the defaults, then start training available now
- 25 Chapter 24 — Run a smoke test that cannot flatter you available now
- 26 Chapter 25 — Skip training: extract features from the public checkpoint available now
- 27 Chapter 26 — Freeze the encoder and test your own videos available now
Part 5 — Put the representation back on the world-model road
Appendices — A backpack for the trail
The road, seen from above
- 2006: energydo these belong together?
- 2022: blueprintperceive, predict, act
- 2023: imagesI-JEPA
- 2024–25: videothe V-JEPA line
- 2025–26: leanerLeJEPA → LeVJEPA
Imagine LeCun building a town out of blocks. The first useful block is a compatibility meter: good pairings sit low on an energy landscape; bad pairings sit high. Years later comes a plan for the whole town—eyes, memory, a world model, goals, costs, and an actor. I-JEPA, V-JEPA, and LeVJEPA test pieces of the foundation. None is the whole town.
Five stops, five different questions
2006: learning with energy. An energy function scores whether an input and a proposed answer fit. Low means compatible; high means incompatible. This lets a system compare abstract states without committing to a probability model that generates every pixel.
2022: the AMI blueprint. LeCun’s position paper arranges perception, a world model, short-term memory, an actor, intrinsic cost, and a configurator in a hierarchical architecture. JEPA appears as a way to predict in representation space. The paper is a research program, not one already-complete model.
2023–2025: pictures, then moving pictures. I-JEPA predicts representations of masked image regions from visible context. V-JEPA carries feature prediction into video. V-JEPA 2 scales action-free video pretraining, then separately trains an action-conditioned model, V-JEPA 2-AC, on robot trajectories. The planning evidence belongs to that later action-conditioned and MPC pipeline, not to a passive video encoder alone.
2025–2026: put collapse prevention in the loss. LeJEPA argues for an isotropic Gaussian target and introduces SIGReg, replacing teacher–student anti-collapse machinery with an explicit distribution constraint. LeVJEPA brings that objective to video: one shared encoder, a small projector, invariance, and SIGReg. Its first arXiv version appeared on 2026-08-27.
The recurring direction is now easier to see: score compatibility; predict in a high-level space; find stable ways to learn from unlabeled images and video; only then connect perception to action-conditioned prediction and planning. LeVJEPA stands at the “efficient video perception” stop. Block-causal attention does not smuggle in actions, costs, or search.
Put the cards in order
Arrange these labels by public date: AMI blueprint, I-JEPA image experiment, LeJEPA/SIGReg, LeVJEPA video objective. The order is 2022 → 2023 → 2025 → 2026. Under the final card, add: encoder, not complete agent.
Primary trail markers
- A Tutorial on Energy-Based Learning (2006, author-hosted PDF)
- A Path Towards Autonomous Machine Intelligence (2022)
- I-JEPA paper and Meta’s introduction
- V-JEPA paper and Meta’s introduction
- V-JEPA 2 paper and Meta’s project article
- LeJEPA v3 and LeVJEPA v1
Three things worth remembering
- LeCun’s long road runs from representation and prediction toward agents that can reason and plan.
- JEPA names a family of ideas; I-JEPA, V-JEPA, LeJEPA, and LeVJEPA are not interchangeable implementations.
- LeVJEPA supplies a video-encoding foundation, not the completed destination.
Can you tell the stops apart?
- What question does an energy function ask?
- Is the 2022 AMI paper a finished product specification or a research blueprint?
- Where does V-JEPA 2’s robot-planning evidence come from?
Answers
- Whether an input and proposed output are compatible.
- A research blueprint.
- The separately action-conditioned V-JEPA 2-AC stage and its planning loop, not the action-free video encoder by itself.