Course progress Course outline 48 of 48 lessons available
Part 0 — Reading Guide: What Exactly Are We Going to Learn?
Part 1 — World Models: An Internal Sandbox for the Agent
Part 2 — Turning Images into State: The LeWM Architecture
- 06 Chapter 5 — Trajectory Data: To the Model, the World Is Not an Image Collection available now
- 07 Chapter 6 — The Visual Encoder: Issuing a “State Passport” for Every Frame available now
- 08 Chapter 7 — The Dynamics Predictor: Moving Time Forward in the Mind available now
- 09 Chapter 8 — A Complete Forward Pass: Follow One Batch from Start to Finish available now
Part 3 — Preventing the Model from Cheating: Prediction Loss and SIGReg
- 10 Chapter 9 — The Most Dangerous Shortcut: Representation Collapse available now
- 11 Chapter 10 — Prediction Loss: How the Model Learns the Next Step Current lesson
- 12 Chapter 11 — The Intuition Behind SIGReg: Letting Representation Space “Breathe” available now
- 13 Chapter 12 — Keep the Mathematics Minimal but Sufficient available now
- 14 Chapter 13 — The Original LeWM’s End-to-End Training Mechanism available now
- 15 Chapter 14 — Train a Model That Does Not Collapse Immediately available now
Part 4 — Putting the Model into Action: Planning in Latent Space
- 16 Chapter 15 — Goal-Conditioned Planning: From “Where Am I?” to “Where Do I Want to Go?” available now
- 17 Chapter 16 — Latent Euclidean Distance: Convenient, but Not Necessarily Reliable available now
- 18 Chapter 17 — CEM: Searching for Actions Through an Elimination Tournament available now
- 19 Chapter 18 — MPC: Do Not Trust the Model for Too Long at Once available now
- 20 Chapter 19 — Long-Horizon Rollouts: How Small Errors Snowball into Major Failures available now
- 21 Chapter 20 — Implement a Minimal LeWM Planner from Scratch available now
Part 5 — Engineering Reproduction: From Paper to Running System
- 22 Chapter 21 — The Official Repository and Experimental Environment available now
- 23 Chapter 22 — First Experiment: A TwoRoom Smoke Test available now
- 24 Chapter 23 — Second Experiment: Reproducing PushT available now
- 25 Chapter 24 — How to Evaluate a World Model Fairly available now
- 26 Chapter 25 — Failure-Diagnosis Manual available now
Part 6 — What Has LeWM Actually Learned?
- 27 Chapter 26 — Linear Probes: Which Physical Variables Are Encoded in the Latent State? available now
- 28 Chapter 27 — Give Latent Space a “Health Check” available now
- 29 Chapter 28 — Violation of Expectation: Is the Model Surprised by “Impossible Events”? available now
- 30 Chapter 29 — How to Discuss “Understanding the World” Rigorously available now
Part 7 — Why “Accurate Prediction” Can Still Produce “Poor Planning”
- 31 Chapter 30 — The Gap Between the Training Objective and the Planning Objective available now
- 32 Chapter 31 — Global Non-Collapse Does Not Guarantee Preservation of Task-Relevant Dynamics available now
- 33 Chapter 32 — When Is an Isotropic Gaussian Prior Too Strong? available now
- 34 Chapter 33 — Long-Horizon Planning: Predict Farther or Plan More Intelligently? available now
- 35 Chapter 34 — From Positional Distance to Task Progress available now
- 36 Chapter 35 — Multi-Task Learning, Real Robots, and Visual Distractions available now
- 37 Chapter 36 — Theoretical Boundaries: When Can the True State Be Identified? available now
Part 8 — From Reproducer to Researcher
Appendices
- 41 Appendix A — The Minimum Necessary Mathematical Toolkit available now
- 42 Appendix B — PyTorch Implementation Quick Reference available now
- 43 Appendix C — Complete Tensor-Shape Table available now
- 44 Appendix D — Experiment Configuration Cards available now
- 45 Appendix E — Paper Timeline and Evidence Levels available now
- 46 Appendix F — Glossary available now
- 47 Appendix G — Reproduction Checklist available now
- 48 Appendix H — Expert-Review Checklist available now
The big picture
- Guessnext latent
- Encodewhat really came next
- Measuresquare coordinate gaps
- Rememberlocal ruler only

A tiny story
A driving student can make many individual turns accurately and still choose the wrong route across town. LeWM’s prediction loss is like a ruler for one local turn. It asks whether the next latent guess is near the shared encoder’s representation of the next recorded observation.
That ruler is essential. Without it, the predictor has no direct reason to follow recorded transitions. But its coordinates are learned—not meters, degrees, or velocities.
The technical backpack
The frozen implementation uses mean squared error:
observed_states = encode(recorded_frames)
guesses = predict(observed_states_without_last, recorded_actions_without_last)
targets = observed_states_without_first # still connected
L_pred = mean((guesses - targets)^2)
Squaring prevents positive and negative gaps from cancelling; averaging combines batch, time positions, and latent coordinates. If an important physical variable is absent from the representation, MSE has no coordinate with which to complain.
The target is learned. It is not an oracle state, an EMA teacher, or a stopped answer. Prediction gradients reach the shared encoder through context and target uses. This flexibility also admits the constant-code collapse shortcut, which is why SIGReg has a different job.
With the frozen four-observation default, all shifted pairs count:
context/actions: z0+a0 z1+a1 z2+a2
predictions: p1 p2 p3
targets: z1 z2 z3
They are three one-step targets at different context positions—not a direct one-, two-, and three-step rollout from one start. num_preds=1 sets the offset. With frameskip=5, “one model step” spans five environment actions. The paper reports TwoRoom history one, while the frozen global default is history three.
During teacher forcing, every context state comes from a recorded observation. During planning, predicted states are appended and become inputs. The first small displacement can change the next prediction. Error may grow, shrink, rotate, cancel, or saturate; it need not rise monotonically.
The trick that fools us
Build a dataset with 95 straight transitions and five rare doorway turns. A predictor can handle the majority, ignore the turn action, and still report a pleasing average.
Do not throw away MSE. Unpack it:
- report held-out teacher-forced one-step error by region and prediction slot;
- replay recorded actions autoregressively and plot error by horizon;
- compare teacher-forced and autoregressive contexts at the same targets;
- hold history fixed and swap supported actions;
- use privileged-state probes only with clear limits;
- report cost ranking and closed-loop success separately.
Training MSE and terminal planning cost can both contain squared latent gaps yet use different tensors and reductions. One is a mean over shifted transitions; the other is a summed terminal score per candidate. Their raw values are not interchangeable.
Evidence receipt
LeWorldModel v3, frozen train.py, and frozen JEPA.rollout support teacher-forced one-step training and autoregressive planning.
In one single-seed TwoRoom reproduction, one-step accuracy ranked short-horizon behavior across three checkpoints but did not rank long-horizon planning success. That supports measuring the later links separately under that protocol; it does not establish a universal law or identify the cause.
Allowed claim: low connected one-step MSE shows local latent agreement on evaluated transitions. Do not promote it to meaningful state, reachability, stable rollout, effective CEM, or successful control.
Quick check
- Why is the encoded next observation useful without being fixed ground truth?
- What changes between teacher-forced and autoregressive context?
- Which metric would reveal a rare action-dependent failure hidden by the average?