JEPA4Japan · tutorials

Chapter 3 — The JEPA family, without the name soup

719 words 4 min read #LeVJEPA#JEPA#self-supervised video#SIGReg

Tell JEPA, I-JEPA, V-JEPA, V-JEPA 2, LeJEPA, LeVJEPA, and LeWorldModel apart by their actual computation graphs.

Course progress Course outline 34 of 34 lessons available

Part 0 — Get the map

  1. 01 Chapter 0 — Before you begin: what this course promises available now
  2. 02 Chapter 1 — One video, two windows available now
  3. 03 Chapter 2 — A walk along Yann LeCun’s research road available now
  4. 04 Chapter 3 — The JEPA family, without the name soup Current lesson

Part 1 — Why a small objective can learn to see

  1. 05 Chapter 4 — Video can set its own homework available now
  2. 06 Chapter 5 — Keep the meaning; do not repaint every pixel available now
  3. 07 Chapter 6 — Match the cards, but do not leave every card blank available now
  4. 08 Chapter 7 — SIGReg checks a cloud by looking at its shadows available now
  5. 09 Chapter 8 — The whole LeVJEPA objective on one line available now

Part 2 — Send a video through one encoder

  1. 10 Chapter 9 — How global and local views are paired available now
  2. 11 Chapter 10 — Cut a video into space-time tiles available now
  3. 12 Chapter 11 — One encoder, one projector, one summary card available now
  4. 13 Chapter 12 — One complete trip through the model available now
  5. 14 Chapter 13 — Why throwing away 95% can help available now
  6. 15 Chapter 14 — Same-frame teamwork, no peeking into tomorrow available now
  7. 16 Chapter 15 — RoPE, single-frame tubelets, and unexpectedly useful patch features available now

Part 3 — Read the experiments, not just the headline

  1. 17 Chapter 16 — What the four ablation ladders actually test available now
  2. 18 Chapter 17 — Equal epochs are not equal bills available now
  3. 19 Chapter 18 — What ImageNet, K400, and SSv2 are really asking available now
  4. 20 Chapter 19 — Keep the paper’s results in a ledger available now
  5. 21 Chapter 20 — Claims the evidence does not yet earn available now

Part 4 — From the official repository to your own experiment

  1. 22 Chapter 21 — A map of the official repository available now
  2. 23 Chapter 22 — Ten long walks become a training set available now
  3. 24 Chapter 23 — Read the defaults, then start training available now
  4. 25 Chapter 24 — Run a smoke test that cannot flatter you available now
  5. 26 Chapter 25 — Skip training: extract features from the public checkpoint available now
  6. 27 Chapter 26 — Freeze the encoder and test your own videos available now

Part 5 — Put the representation back on the world-model road

  1. 28 Chapter 27 — The important boundary: an encoder is not a planner available now
  2. 29 Chapter 28 — How LeVJEPA might feed a future world model available now
  3. 30 Chapter 29 — Ten projects, from first experiment to paper-sized question available now

Appendices — A backpack for the trail

  1. 31 Appendix A — The smallest useful math kit available now
  2. 32 Appendix B — The complete tensor-shape table available now
  3. 33 Appendix C — Glossary and paper timeline available now
  4. 34 Appendix D — Reproduction and review checklist available now

One surname, different tools

  1. JEPAthe family idea
  2. I-JEPApredict image-region features
  3. V-JEPApredict video features
  4. LeJEPASIGReg blocks collapse
  5. LeVJEPAa lean video encoder
Shared names do not imply shared parts. Ask what enters, what is compared, where gradients go, and what the trained system can do.

A workshop has a jigsaw board, a video reader, a card sorter, and a route planner. Their labels all include “JEPA.” The label alone is no help. Four questions are: What does it eat? What does it compare? Which pieces learn? What can the finished tool do?

Those four questions clear up most of the family confusion.

Meet the relatives by their computation graphs

JEPA is the broad idea: use one part of the available information to constrain or predict another part in a learned representation space. The idea does not prescribe one encoder, mask, or anti-collapse mechanism.

I-JEPA works with still images. A context encoder plus predictor estimates target-encoder representations of selected image regions. V-JEPA moves this pattern to video: a context encoder and predictor estimate masked space-time target features, while an EMA target encoder and stop-gradient make the training graph asymmetric.

V-JEPA 2 scales video pretraining. Keep its action-free encoder/predictor separate from V-JEPA 2-AC, the later model trained with robot trajectories. Evidence about planning belongs to the action-conditioned predictor and MPC loop.

LeJEPA does not mean “large JEPA.” The paper expands the name as Latent-Euclidean JEPA and describes its recipe as lean. It combines same-source view matching with SIGReg, using an explicit distribution constraint to rule out collapse instead of teacher–student asymmetry.

LeVJEPA applies that lean objective to video. Global and local views all travel through one encoder and one projector; their [CLS] embeddings enter the loss. There is no target encoder, predictor, stopped gradient, or masked-token reconstruction in that objective. The EMA encoder maintained by the training code is an evaluation copy outside the loss graph.

LeWorldModel (LeWM) adds action-conditioned latent dynamics and a planning loop. It asks, “What might follow if I take this action?” LeVJEPA asks, “How can I encode the video I am observing?” The two ideas can meet, but one is not simply the next version number of the other. A causal encoder only prevents a frame from reading future frames; it does not invent actions.

A practical name test follows. Masked target tokens + predictor + EMA teacher looks like V-JEPA. Shared encoder/projector + same-clip [CLS] matching + SIGReg is the LeVJEPA studied here.

Repair a mislabeled sketch

Suppose a diagram titled “LeVJEPA” shows online encoder → predictor → stopped-gradient EMA teacher. Cross out the predictor and target branch. Feed every view through one encoder and projector; compute invariance and SIGReg from [CLS]. Draw EMA outside the objective and label it evaluation checkpoint only.

Family records

Keep these labels straight

  1. JEPA is a representation-space prediction blueprint, not one fixed graph.
  2. LeVJEPA’s objective contains shared encoder/projector paths, [CLS] matching, and SIGReg.
  3. Planning results from V-JEPA 2-AC or LeWM cannot be credited to LeVJEPA.

Name test

  1. Does “Le” in LeJEPA mean “large”?
  2. Is there a predictor in the LeVJEPA training objective?
  3. Is block-causal attention enough to turn the encoder into a planner?
Answers
  1. No. The paper expands LeJEPA as Latent-Euclidean JEPA; “lean” describes the compact recipe.
  2. No.
  3. No. Planning still needs action-conditioned dynamics, an objective or cost, and search and control.