Course progress Course outline 34 of 34 lessons available
Part 0 — Get the map
Part 1 — Why a small objective can learn to see
- 05 Chapter 4 — Video can set its own homework available now
- 06 Chapter 5 — Keep the meaning; do not repaint every pixel available now
- 07 Chapter 6 — Match the cards, but do not leave every card blank available now
- 08 Chapter 7 — SIGReg checks a cloud by looking at its shadows available now
- 09 Chapter 8 — The whole LeVJEPA objective on one line available now
Part 2 — Send a video through one encoder
- 10 Chapter 9 — How global and local views are paired available now
- 11 Chapter 10 — Cut a video into space-time tiles available now
- 12 Chapter 11 — One encoder, one projector, one summary card available now
- 13 Chapter 12 — One complete trip through the model available now
- 14 Chapter 13 — Why throwing away 95% can help available now
- 15 Chapter 14 — Same-frame teamwork, no peeking into tomorrow available now
- 16 Chapter 15 — RoPE, single-frame tubelets, and unexpectedly useful patch features available now
Part 3 — Read the experiments, not just the headline
- 17 Chapter 16 — What the four ablation ladders actually test available now
- 18 Chapter 17 — Equal epochs are not equal bills available now
- 19 Chapter 18 — What ImageNet, K400, and SSv2 are really asking available now
- 20 Chapter 19 — Keep the paper’s results in a ledger available now
- 21 Chapter 20 — Claims the evidence does not yet earn available now
Part 4 — From the official repository to your own experiment
- 22 Chapter 21 — A map of the official repository available now
- 23 Chapter 22 — Ten long walks become a training set available now
- 24 Chapter 23 — Read the defaults, then start training available now
- 25 Chapter 24 — Run a smoke test that cannot flatter you available now
- 26 Chapter 25 — Skip training: extract features from the public checkpoint available now
- 27 Chapter 26 — Freeze the encoder and test your own videos available now
Part 5 — Put the representation back on the world-model road
Appendices — A backpack for the trail
Add one rung, test it, then climb
- Freeze sightkeep a known baseline
- Add actionswhat did the agent do?
- Roll forwardpredict latent futures
- Price themgoal, risk, constraints
- Use MPCone move, then look again
Suppose you have built a good pair of eyes and want to add a tabletop imagination. Lock the eyes first. Teach the new box only this: “Given what I see now and this push, what should I see next?” If the experiment fails, you know which box to open.
Training everything end to end on day one makes a failure much harder to locate. The encoder may drift, the dynamics may ignore actions, or the planner may simply search badly—and one success rate cannot tell those stories apart.
A conservative first build
Start with the released encoder frozen. Feed it an observation history shaped [B,3,T,224,224]. Use patch states associated with the current temporal slot, z_t:[B,N,D], rather than the [CLS] token that can see the entire supplied clip. Add recorded actions a:[B,H,A] and a causal, action-conditioned dynamics model:
(z_t, a_t) -> z_hat_t+1 -> z_hat_t+2 -> ... -> z_hat_t+H
Encode the real future frames with the same frozen encoder to make targets. Train both one-step and multi-step latent losses. Everything after the frozen encoder—the action input, predictor, rollout loss, and control dataset—is a proposed addition, not part of LeVJEPA v1.
Token alignment needs its own contract. Token i can be compared across time only when camera, crop, frame rate, and patch grid give it a stable meaning. With camera motion and occlusion, state whether the model predicts every patch, only visible patches, aligned object features, or a pooled state. A mean error over a mostly static background can look excellent while every moving object is wrong, so report those errors separately.
The first baseline is almost embarrassingly simple: copy the current state forward. Adjacent video frames are similar. A dynamics model that cannot beat copy-last has not earned its extra machinery. Add no-action and shuffled-action models next; they reveal whether the predictor actually listens to a_t.
Close the loop only after prediction works
For image-goal control, encode the goal as z_goal. A search method such as the cross-entropy method (CEM) samples K action sequences, the learned dynamics rolls each sequence through latent space, and a cost ranks the futures. Model-predictive control executes only the first action, observes the world again, and repeats:
# Proposed pseudocode, not an upstream repository function.
while not done:
z0 = encode_causal(observation_history)
candidates = cem.sample(K, horizon=H)
futures = dynamics.rollout(z0, candidates) # [K, H, N, D]
scores = cost(futures, z_goal, constraints)
env.step(candidates[scores.argmin(), 0])
That final observation matters. It lets the controller correct model error and unexpected changes instead of trusting a long imagined path.
Six gates before a robot demo
- Interface: shapes are asserted and future-leakage tests pass.
- One step: the predictor beats copy-last, no-action, and shuffled-action baselines.
- Rollout: error versus horizon and uncertainty calibration are reported.
- Counterfactual: different plausible actions produce distinguishable futures.
- Simulation: data, compute, CEM samples, horizon, and baselines are fixed before comparing success.
- Closed loop: only then test distribution shifts, stopping rules, human takeover, and constraint violations.
Stop at the first failed gate and diagnose it. A larger robot run will not repair an ignored action input.
A deterministic MSE predictor can blur several possible futures into one. Compare an ensemble, stochastic latent, or energy-ranked candidates, then test coverage and calibration rather than admiring a sample. Longer time scales may warrant a slow latent layer and subgoals, but each addition must beat the single-scale, short-horizon baseline. If the encoder is later unfrozen, track its original frozen probes and embedding spectrum so a narrow control dataset does not erase general visual features or collapse the representation.
The paper notes that block-causal structure permits reuse of past computation. The pinned repository and released Transformers interface do not ship a complete incremental KV-cache API. A streaming implementation is therefore new engineering and should be checked token by token against recomputing the full prefix. “Cacheable in principle” and “a verified cache is available” are different statements.
Sources for the build, not claims of completion
- LeCun’s 2022 autonomous-machine blueprint
- LeVJEPA v1: block-causal design and future directions
- V-JEPA 2-AC and LeWorldModel v3 as primary design comparisons
- Pinned LeVJEPA implementation
Three build rules
- A world-model extension must add action-conditioned prediction, a cost, search, and MPC explicitly.
- Freeze the encoder and try to falsify each layer before end-to-end tuning or long-horizon control.
- A cache that the architecture permits is not a streaming interface the public code already supplies.
Before adding the next rung
- What are the two indispensable inputs to the smallest dynamics predictor?
- Why compare it with copy-last?
- Why does MPC execute only the first action in its chosen sequence?
Check the answers
- A current or historical latent state and a candidate action.
- Neighboring frames are naturally similar; failing to beat copying means the model has not learned useful dynamics.
- It can observe again after acting, correct prediction errors or disturbances, and replan.