JEPA4Japan · tutorials

Chapter 27 — The important boundary: an encoder is not a planner

874 words 4 min read #LeVJEPA#JEPA#self-supervised video#SIGReg

Distinguish video encoding, action-conditioned dynamics, costs, search, and closed-loop control.

Course progress Course outline 34 of 34 lessons available

Part 0 — Get the map

  1. 01 Chapter 0 — Before you begin: what this course promises available now
  2. 02 Chapter 1 — One video, two windows available now
  3. 03 Chapter 2 — A walk along Yann LeCun’s research road available now
  4. 04 Chapter 3 — The JEPA family, without the name soup available now

Part 1 — Why a small objective can learn to see

  1. 05 Chapter 4 — Video can set its own homework available now
  2. 06 Chapter 5 — Keep the meaning; do not repaint every pixel available now
  3. 07 Chapter 6 — Match the cards, but do not leave every card blank available now
  4. 08 Chapter 7 — SIGReg checks a cloud by looking at its shadows available now
  5. 09 Chapter 8 — The whole LeVJEPA objective on one line available now

Part 2 — Send a video through one encoder

  1. 10 Chapter 9 — How global and local views are paired available now
  2. 11 Chapter 10 — Cut a video into space-time tiles available now
  3. 12 Chapter 11 — One encoder, one projector, one summary card available now
  4. 13 Chapter 12 — One complete trip through the model available now
  5. 14 Chapter 13 — Why throwing away 95% can help available now
  6. 15 Chapter 14 — Same-frame teamwork, no peeking into tomorrow available now
  7. 16 Chapter 15 — RoPE, single-frame tubelets, and unexpectedly useful patch features available now

Part 3 — Read the experiments, not just the headline

  1. 17 Chapter 16 — What the four ablation ladders actually test available now
  2. 18 Chapter 17 — Equal epochs are not equal bills available now
  3. 19 Chapter 18 — What ImageNet, K400, and SSv2 are really asking available now
  4. 20 Chapter 19 — Keep the paper’s results in a ledger available now
  5. 21 Chapter 20 — Claims the evidence does not yet earn available now

Part 4 — From the official repository to your own experiment

  1. 22 Chapter 21 — A map of the official repository available now
  2. 23 Chapter 22 — Ten long walks become a training set available now
  3. 24 Chapter 23 — Read the defaults, then start training available now
  4. 25 Chapter 24 — Run a smoke test that cannot flatter you available now
  5. 26 Chapter 25 — Skip training: extract features from the public checkpoint available now
  6. 27 Chapter 26 — Freeze the encoder and test your own videos available now

Part 5 — Put the representation back on the world-model road

  1. 28 Chapter 27 — The important boundary: an encoder is not a planner Current lesson
  2. 29 Chapter 28 — How LeVJEPA might feed a future world model available now
  3. 30 Chapter 29 — Ten projects, from first experiment to paper-sized question available now

Appendices — A backpack for the trail

  1. 31 Appendix A — The smallest useful math kit available now
  2. 32 Appendix B — The complete tensor-shape table available now
  3. 33 Appendix C — Glossary and paper timeline available now
  4. 34 Appendix D — Reproduction and review checklist available now

Five boxes, only one supplied

  1. Encodewhat is here?
  2. Actwhat might I do?
  3. Predictwhat happens next?
  4. Scorewhich future is better?
  5. Replanmove, look, repeat
Published LeVJEPA supplies the first box. Block-causal attention does not conjure the other four.

A guide can describe a street perfectly and still be unable to drive through it. Ask, “What is beside the bus?” and the guide answers. Ask, “If I turn now, will I hit the curb?” and four new problems appear: the proposed turn, its effect, what counts as bad, and how to choose among alternatives.

That is the line between a video encoder and a planner.

Three meanings of causal that should not share a label

LeVJEPA accepts a clip and returns a [CLS] summary plus patch representations. Its block_causal mask controls information flow. A patch token may read its own frame and earlier frames, but not later frames. [CLS] is a read-only summary slot: it may read the whole clip, while patch tokens cannot read information back from it.

This makes the patch stream temporally usable without future leakage. It does not show intervention-level causality. In passive video, a light switching on before a door opens may reflect a hidden common cause. A past-only attention mask cannot answer, “Would pressing this button open the door?”

Keep these claims separate:

  1. Causal mask: later frames cannot leak into earlier patch states.
  2. Action-conditioned prediction: a model estimates a future under a proposed action.
  3. Causal discovery: a system identifies intervention relationships rather than correlations.

Action-labeled trajectories can support the second claim, but even then the predictor might ignore the action and copy visual momentum. Shuffled-action, no-action, and counterfactual tests are needed to find that shortcut.

What planning would have to add

A minimal planner needs four pieces that LeVJEPA does not provide:

  • an action a_t;
  • dynamics such as F(z_t, a_t) -> z_hat_t+1;
  • a goal or cost C(z_hat, goal);
  • a search procedure and a closed loop that observes again after acting.

If several futures are plausible, it also needs uncertainty or a stochastic latent variable. This sketch is an interface checklist, not an official LeVJEPA API:

z_now = encoder(history)             # supplied by LeVJEPA
z_next = dynamics(z_now, action)     # proposed addition; train on actions
score = cost(z_next, goal)           # proposed addition; define and validate
action = search(dynamics, cost)[0]   # act once, observe, then plan again

Two nearby projects make the missing pieces concrete. V-JEPA 2-AC freezes a V-JEPA 2 encoder, trains a separate action-conditioned predictor on fewer than 62 hours of robot trajectories, and then searches toward an image goal with model-predictive control. LeWorldModel instead trains an encoder and action-conditioned predictor jointly from pixels. These are useful comparisons. Their components and results do not travel backward into LeVJEPA.

A classification or retrieval probe is not a quiet substitute for planning, either. It shows that a reader can extract some information from a representation. It does not show that the information survives long rollouts, that costs are calibrated, or that a search procedure finds safe actions. A defensible planning claim reports representation quality, one-step dynamics, long-horizon error, search budget, and closed-loop success separately.

The pinned repository’s paper.md contains a draft Push-T world-model section, but its result table still says TODO and omits the episode, seed, and planner protocol needed to interpret a score. There is no LeVJEPA Push-T number to quote. LeVJEPA v1 reports representation, efficiency, and probe results; using it as a world-model foundation remains a research proposal.

Audit a planning diagram

Take any diagram captioned “LeVJEPA plans.” Point to the action tensor, the next-state predictor, the cost, the action search, and the moment when a fresh observation triggers replanning. If four of those are absent, the diagram contains an encoder, not a planner.

The relevant records

Keep the boundary visible

  1. Block-causal attention limits where information comes from; it does not learn the causal effect of an action.
  2. Planning still needs action-conditioned dynamics, a cost, search, and closed-loop replanning.
  3. The upstream Push-T table is an unfinished draft, not a reported LeVJEPA result.

Point to the missing box

  1. Does native LeVJEPA receive robot actions?
  2. Why is [CLS] not a causal state for every instant in the clip?
  3. May a V-JEPA 2-AC planning result be reported as a LeVJEPA result?
Check the answers
  1. No.
  2. It may read the complete input clip; the per-frame patch tokens are the states restricted to temporal prefixes.
  3. No. The model, training stages, and added action predictor are different.