JEPA4Japan · tutorials

Chapter 2 — A walk along Yann LeCun’s research road

684 words 4 min read #LeVJEPA#JEPA#self-supervised video#SIGReg

Travel from energy-based learning and the 2022 AMI blueprint through JEPA, I-JEPA, V-JEPA, LeJEPA, and LeVJEPA.

Course progress Course outline 34 of 34 lessons available

Part 0 — Get the map

  1. 01 Chapter 0 — Before you begin: what this course promises available now
  2. 02 Chapter 1 — One video, two windows available now
  3. 03 Chapter 2 — A walk along Yann LeCun’s research road Current lesson
  4. 04 Chapter 3 — The JEPA family, without the name soup available now

Part 1 — Why a small objective can learn to see

  1. 05 Chapter 4 — Video can set its own homework available now
  2. 06 Chapter 5 — Keep the meaning; do not repaint every pixel available now
  3. 07 Chapter 6 — Match the cards, but do not leave every card blank available now
  4. 08 Chapter 7 — SIGReg checks a cloud by looking at its shadows available now
  5. 09 Chapter 8 — The whole LeVJEPA objective on one line available now

Part 2 — Send a video through one encoder

  1. 10 Chapter 9 — How global and local views are paired available now
  2. 11 Chapter 10 — Cut a video into space-time tiles available now
  3. 12 Chapter 11 — One encoder, one projector, one summary card available now
  4. 13 Chapter 12 — One complete trip through the model available now
  5. 14 Chapter 13 — Why throwing away 95% can help available now
  6. 15 Chapter 14 — Same-frame teamwork, no peeking into tomorrow available now
  7. 16 Chapter 15 — RoPE, single-frame tubelets, and unexpectedly useful patch features available now

Part 3 — Read the experiments, not just the headline

  1. 17 Chapter 16 — What the four ablation ladders actually test available now
  2. 18 Chapter 17 — Equal epochs are not equal bills available now
  3. 19 Chapter 18 — What ImageNet, K400, and SSv2 are really asking available now
  4. 20 Chapter 19 — Keep the paper’s results in a ledger available now
  5. 21 Chapter 20 — Claims the evidence does not yet earn available now

Part 4 — From the official repository to your own experiment

  1. 22 Chapter 21 — A map of the official repository available now
  2. 23 Chapter 22 — Ten long walks become a training set available now
  3. 24 Chapter 23 — Read the defaults, then start training available now
  4. 25 Chapter 24 — Run a smoke test that cannot flatter you available now
  5. 26 Chapter 25 — Skip training: extract features from the public checkpoint available now
  6. 27 Chapter 26 — Freeze the encoder and test your own videos available now

Part 5 — Put the representation back on the world-model road

  1. 28 Chapter 27 — The important boundary: an encoder is not a planner available now
  2. 29 Chapter 28 — How LeVJEPA might feed a future world model available now
  3. 30 Chapter 29 — Ten projects, from first experiment to paper-sized question available now

Appendices — A backpack for the trail

  1. 31 Appendix A — The smallest useful math kit available now
  2. 32 Appendix B — The complete tensor-shape table available now
  3. 33 Appendix C — Glossary and paper timeline available now
  4. 34 Appendix D — Reproduction and review checklist available now

The road, seen from above

  1. 2006: energydo these belong together?
  2. 2022: blueprintperceive, predict, act
  3. 2023: imagesI-JEPA
  4. 2024–25: videothe V-JEPA line
  5. 2025–26: leanerLeJEPA → LeVJEPA
This is not a race toward bigger models. It is a chain of questions about learning, predicting, and eventually acting in an abstract space.

Imagine LeCun building a town out of blocks. The first useful block is a compatibility meter: good pairings sit low on an energy landscape; bad pairings sit high. Years later comes a plan for the whole town—eyes, memory, a world model, goals, costs, and an actor. I-JEPA, V-JEPA, and LeVJEPA test pieces of the foundation. None is the whole town.

Five stops, five different questions

2006: learning with energy. An energy function scores whether an input and a proposed answer fit. Low means compatible; high means incompatible. This lets a system compare abstract states without committing to a probability model that generates every pixel.

2022: the AMI blueprint. LeCun’s position paper arranges perception, a world model, short-term memory, an actor, intrinsic cost, and a configurator in a hierarchical architecture. JEPA appears as a way to predict in representation space. The paper is a research program, not one already-complete model.

2023–2025: pictures, then moving pictures. I-JEPA predicts representations of masked image regions from visible context. V-JEPA carries feature prediction into video. V-JEPA 2 scales action-free video pretraining, then separately trains an action-conditioned model, V-JEPA 2-AC, on robot trajectories. The planning evidence belongs to that later action-conditioned and MPC pipeline, not to a passive video encoder alone.

2025–2026: put collapse prevention in the loss. LeJEPA argues for an isotropic Gaussian target and introduces SIGReg, replacing teacher–student anti-collapse machinery with an explicit distribution constraint. LeVJEPA brings that objective to video: one shared encoder, a small projector, invariance, and SIGReg. Its first arXiv version appeared on 2026-08-27.

The recurring direction is now easier to see: score compatibility; predict in a high-level space; find stable ways to learn from unlabeled images and video; only then connect perception to action-conditioned prediction and planning. LeVJEPA stands at the “efficient video perception” stop. Block-causal attention does not smuggle in actions, costs, or search.

Put the cards in order

Arrange these labels by public date: AMI blueprint, I-JEPA image experiment, LeJEPA/SIGReg, LeVJEPA video objective. The order is 2022 → 2023 → 2025 → 2026. Under the final card, add: encoder, not complete agent.

Primary trail markers

Three things worth remembering

  1. LeCun’s long road runs from representation and prediction toward agents that can reason and plan.
  2. JEPA names a family of ideas; I-JEPA, V-JEPA, LeJEPA, and LeVJEPA are not interchangeable implementations.
  3. LeVJEPA supplies a video-encoding foundation, not the completed destination.

Can you tell the stops apart?

  1. What question does an energy function ask?
  2. Is the 2022 AMI paper a finished product specification or a research blueprint?
  3. Where does V-JEPA 2’s robot-planning evidence come from?
Answers
  1. Whether an input and proposed output are compatible.
  2. A research blueprint.
  3. The separately action-conditioned V-JEPA 2-AC stage and its planning loop, not the action-free video encoder by itself.