Course progress Course outline 48 of 48 lessons available
Part 0 — Reading Guide: What Exactly Are We Going to Learn?
Part 1 — World Models: An Internal Sandbox for the Agent
Part 2 — Turning Images into State: The LeWM Architecture
- 06 Chapter 5 — Trajectory Data: To the Model, the World Is Not an Image Collection available now
- 07 Chapter 6 — The Visual Encoder: Issuing a “State Passport” for Every Frame available now
- 08 Chapter 7 — The Dynamics Predictor: Moving Time Forward in the Mind available now
- 09 Chapter 8 — A Complete Forward Pass: Follow One Batch from Start to Finish available now
Part 3 — Preventing the Model from Cheating: Prediction Loss and SIGReg
- 10 Chapter 9 — The Most Dangerous Shortcut: Representation Collapse available now
- 11 Chapter 10 — Prediction Loss: How the Model Learns the Next Step available now
- 12 Chapter 11 — The Intuition Behind SIGReg: Letting Representation Space “Breathe” available now
- 13 Chapter 12 — Keep the Mathematics Minimal but Sufficient available now
- 14 Chapter 13 — The Original LeWM’s End-to-End Training Mechanism available now
- 15 Chapter 14 — Train a Model That Does Not Collapse Immediately available now
Part 4 — Putting the Model into Action: Planning in Latent Space
- 16 Chapter 15 — Goal-Conditioned Planning: From “Where Am I?” to “Where Do I Want to Go?” available now
- 17 Chapter 16 — Latent Euclidean Distance: Convenient, but Not Necessarily Reliable available now
- 18 Chapter 17 — CEM: Searching for Actions Through an Elimination Tournament available now
- 19 Chapter 18 — MPC: Do Not Trust the Model for Too Long at Once available now
- 20 Chapter 19 — Long-Horizon Rollouts: How Small Errors Snowball into Major Failures available now
- 21 Chapter 20 — Implement a Minimal LeWM Planner from Scratch available now
Part 5 — Engineering Reproduction: From Paper to Running System
- 22 Chapter 21 — The Official Repository and Experimental Environment available now
- 23 Chapter 22 — First Experiment: A TwoRoom Smoke Test available now
- 24 Chapter 23 — Second Experiment: Reproducing PushT available now
- 25 Chapter 24 — How to Evaluate a World Model Fairly available now
- 26 Chapter 25 — Failure-Diagnosis Manual available now
Part 6 — What Has LeWM Actually Learned?
- 27 Chapter 26 — Linear Probes: Which Physical Variables Are Encoded in the Latent State? available now
- 28 Chapter 27 — Give Latent Space a “Health Check” available now
- 29 Chapter 28 — Violation of Expectation: Is the Model Surprised by “Impossible Events”? available now
- 30 Chapter 29 — How to Discuss “Understanding the World” Rigorously available now
Part 7 — Why “Accurate Prediction” Can Still Produce “Poor Planning”
- 31 Chapter 30 — The Gap Between the Training Objective and the Planning Objective available now
- 32 Chapter 31 — Global Non-Collapse Does Not Guarantee Preservation of Task-Relevant Dynamics available now
- 33 Chapter 32 — When Is an Isotropic Gaussian Prior Too Strong? available now
- 34 Chapter 33 — Long-Horizon Planning: Predict Farther or Plan More Intelligently? available now
- 35 Chapter 34 — From Positional Distance to Task Progress available now
- 36 Chapter 35 — Multi-Task Learning, Real Robots, and Visual Distractions available now
- 37 Chapter 36 — Theoretical Boundaries: When Can the True State Be Identified? available now
Part 8 — From Reproducer to Researcher
Appendices
- 41 Appendix A — The Minimum Necessary Mathematical Toolkit available now
- 42 Appendix B — PyTorch Implementation Quick Reference available now
- 43 Appendix C — Complete Tensor-Shape Table available now
- 44 Appendix D — Experiment Configuration Cards available now
- 45 Appendix E — Paper Timeline and Evidence Levels available now
- 46 Appendix F — Glossary available now
- 47 Appendix G — Reproduction Checklist available now
- 48 Appendix H — Expert-Review Checklist available now
The big picture
- Trainreal past + real next
- Learnencoder + dynamics
- Freezemodel stops changing
- Planimagined futures + feedback

A tiny story
In the daytime control room, LeWM studies recorded journeys. It may look at the real next frame, compare its guess with that answer, and update its encoder and predictor. At night, the trained model helps choose an action. The unexecuted future is blank, so every candidate future must come from the model itself.
Mixing the rooms creates plausible nonsense: a training graph with a goal and CEM, a planner that sees the true future, or a planner that bends model weights until an impossible route looks cheap.
The real rule
The two lanes are short enough to memorize:
TRAIN
recorded frames + actions -> encode -> predict
encoded real next frames -> connected targets
prediction loss + weighted SIGReg -> update the learned components
PLAN
current + goal -> frozen encoder
candidate actions -> autoregressive latent rollouts -> terminal cost
CEM changes the action proposal -> execute K -> observe -> repeat
The first lane has an answer key but no task goal. The second has a task goal but no answer key for an action that has not happened.
Symbolically, images move from [B,T,C,H,W] to latent states [B,T,D]. In the cited experiments images are 224×224, frame skip is five, and sampled subtrajectories contain four observation positions in the frozen default. The v3 appendix reports history one for TwoRoom and three for PushT and OGBench-Cube; the frozen global config uses history_size: 3 and TwoRoom does not override it. Preserve both labels.
During training, real encoded states provide causal context—“teacher forcing” here does not mean a teacher network. The shifted target stays connected to the shared encoder. During planning, predictions are appended and reused. Logging or inference .detach() calls do not turn v3 training into stop-gradient.
SIGReg acts on training embeddings. It makes the exact constant-code shortcut costly, but it is not a planning cost and does not certify physics. The paper method default for its weight is 0.1; the frozen YAML uses 0.09. Neither is a universal optimum.
The trick that fools us
Suppose planning updates both candidate actions and world-model weights. The terminal cost falls quickly. Did the action improve? Maybe not—the optimizer may simply bend the predictor until its endpoint resembles the goal.
The baseline invariant is:
model_parameter_change = 0
candidate_action_change = allowed
CEM refines only the action distribution. The goal latent bypasses dynamics and meets imagined endpoints at terminal squared Euclidean cost. The paper’s Appendix B prose returns the final proposal mean, while Algorithm 2 also allows the best sampled sequence or the first action of that mean. The repository delegates CEM to an unpinned stable-worldmodel dependency, so record the resolved return convention.
MPC plans H model steps, executes K, then observes. The reported setup uses H=5, K=5, and five environment steps per model step: 25 environment steps before feedback. Replanning can limit stale imagination; it cannot repair a missing wall or bad goal geometry.
Evidence receipt
LeWorldModel v3, frozen train.py, and frozen jepa.py support the two-lane method, connected training graph, vectorized latent rollout, and terminal criterion.
Allowed claim: LeWM jointly trains an encoder and action-conditioned predictor with prediction plus SIGReg, then freezes the model while CEM/MPC search and execute actions. Do not claim that CEM finds a global optimum, latent distance equals reachability, or MPC always executes one action.
Quick check
- Which information exists in training but not in an unexecuted planning branch?
- What changes during CEM, and what must stay frozen?
- Why does a goal detach during frozen planning say nothing about target detach during training?