Course progress Course outline 34 of 34 lessons available
Part 0 — Get the map
Part 1 — Why a small objective can learn to see
- 05 Chapter 4 — Video can set its own homework available now
- 06 Chapter 5 — Keep the meaning; do not repaint every pixel available now
- 07 Chapter 6 — Match the cards, but do not leave every card blank available now
- 08 Chapter 7 — SIGReg checks a cloud by looking at its shadows available now
- 09 Chapter 8 — The whole LeVJEPA objective on one line available now
Part 2 — Send a video through one encoder
- 10 Chapter 9 — How global and local views are paired available now
- 11 Chapter 10 — Cut a video into space-time tiles available now
- 12 Chapter 11 — One encoder, one projector, one summary card available now
- 13 Chapter 12 — One complete trip through the model available now
- 14 Chapter 13 — Why throwing away 95% can help available now
- 15 Chapter 14 — Same-frame teamwork, no peeking into tomorrow available now
- 16 Chapter 15 — RoPE, single-frame tubelets, and unexpectedly useful patch features available now
Part 3 — Read the experiments, not just the headline
- 17 Chapter 16 — What the four ablation ladders actually test available now
- 18 Chapter 17 — Equal epochs are not equal bills available now
- 19 Chapter 18 — What ImageNet, K400, and SSv2 are really asking available now
- 20 Chapter 19 — Keep the paper’s results in a ledger available now
- 21 Chapter 20 — Claims the evidence does not yet earn available now
Part 4 — From the official repository to your own experiment
- 22 Chapter 21 — A map of the official repository available now
- 23 Chapter 22 — Ten long walks become a training set available now
- 24 Chapter 23 — Read the defaults, then start training available now
- 25 Chapter 24 — Run a smoke test that cannot flatter you available now
- 26 Chapter 25 — Skip training: extract features from the public checkpoint available now
- 27 Chapter 26 — Freeze the encoder and test your own videos available now
Part 5 — Put the representation back on the world-model road
Appendices — A backpack for the trail
Five boxes, only one supplied
- Encodewhat is here?
- Actwhat might I do?
- Predictwhat happens next?
- Scorewhich future is better?
- Replanmove, look, repeat
A guide can describe a street perfectly and still be unable to drive through it. Ask, “What is beside the bus?” and the guide answers. Ask, “If I turn now, will I hit the curb?” and four new problems appear: the proposed turn, its effect, what counts as bad, and how to choose among alternatives.
That is the line between a video encoder and a planner.
Three meanings of causal that should not share a label
LeVJEPA accepts a clip and returns a [CLS] summary plus patch representations. Its block_causal mask controls information flow. A patch token may read its own frame and earlier frames, but not later frames. [CLS] is a read-only summary slot: it may read the whole clip, while patch tokens cannot read information back from it.
This makes the patch stream temporally usable without future leakage. It does not show intervention-level causality. In passive video, a light switching on before a door opens may reflect a hidden common cause. A past-only attention mask cannot answer, “Would pressing this button open the door?”
Keep these claims separate:
- Causal mask: later frames cannot leak into earlier patch states.
- Action-conditioned prediction: a model estimates a future under a proposed action.
- Causal discovery: a system identifies intervention relationships rather than correlations.
Action-labeled trajectories can support the second claim, but even then the predictor might ignore the action and copy visual momentum. Shuffled-action, no-action, and counterfactual tests are needed to find that shortcut.
What planning would have to add
A minimal planner needs four pieces that LeVJEPA does not provide:
- an action
a_t; - dynamics such as
F(z_t, a_t) -> z_hat_t+1; - a goal or cost
C(z_hat, goal); - a search procedure and a closed loop that observes again after acting.
If several futures are plausible, it also needs uncertainty or a stochastic latent variable. This sketch is an interface checklist, not an official LeVJEPA API:
z_now = encoder(history) # supplied by LeVJEPA
z_next = dynamics(z_now, action) # proposed addition; train on actions
score = cost(z_next, goal) # proposed addition; define and validate
action = search(dynamics, cost)[0] # act once, observe, then plan again
Two nearby projects make the missing pieces concrete. V-JEPA 2-AC freezes a V-JEPA 2 encoder, trains a separate action-conditioned predictor on fewer than 62 hours of robot trajectories, and then searches toward an image goal with model-predictive control. LeWorldModel instead trains an encoder and action-conditioned predictor jointly from pixels. These are useful comparisons. Their components and results do not travel backward into LeVJEPA.
A classification or retrieval probe is not a quiet substitute for planning, either. It shows that a reader can extract some information from a representation. It does not show that the information survives long rollouts, that costs are calibrated, or that a search procedure finds safe actions. A defensible planning claim reports representation quality, one-step dynamics, long-horizon error, search budget, and closed-loop success separately.
The pinned repository’s paper.md contains a draft Push-T world-model section, but its result table still says TODO and omits the episode, seed, and planner protocol needed to interpret a score. There is no LeVJEPA Push-T number to quote. LeVJEPA v1 reports representation, efficiency, and probe results; using it as a world-model foundation remains a research proposal.
Audit a planning diagram
Take any diagram captioned “LeVJEPA plans.” Point to the action tensor, the next-state predictor, the cost, the action search, and the moment when a fresh observation triggers replanning. If four of those are absent, the diagram contains an encoder, not a planner.
The relevant records
- LeVJEPA v1: method and limitations
- The unfinished Push-T draft in pinned
paper.md - V-JEPA 2 and V-JEPA 2-AC
- LeWorldModel v3
Keep the boundary visible
- Block-causal attention limits where information comes from; it does not learn the causal effect of an action.
- Planning still needs action-conditioned dynamics, a cost, search, and closed-loop replanning.
- The upstream Push-T table is an unfinished draft, not a reported LeVJEPA result.
Point to the missing box
- Does native LeVJEPA receive robot actions?
- Why is
[CLS]not a causal state for every instant in the clip? - May a V-JEPA 2-AC planning result be reported as a LeVJEPA result?
Check the answers
- No.
- It may read the complete input clip; the per-frame patch tokens are the states restricted to temporal prefixes.
- No. The model, training stages, and added action predictor are different.