Course progress Course outline 48 of 48 lessons available
Part 0 — Reading Guide: What Exactly Are We Going to Learn?
Part 1 — World Models: An Internal Sandbox for the Agent
Part 2 — Turning Images into State: The LeWM Architecture
- 06 Chapter 5 — Trajectory Data: To the Model, the World Is Not an Image Collection available now
- 07 Chapter 6 — The Visual Encoder: Issuing a “State Passport” for Every Frame available now
- 08 Chapter 7 — The Dynamics Predictor: Moving Time Forward in the Mind available now
- 09 Chapter 8 — A Complete Forward Pass: Follow One Batch from Start to Finish available now
Part 3 — Preventing the Model from Cheating: Prediction Loss and SIGReg
- 10 Chapter 9 — The Most Dangerous Shortcut: Representation Collapse available now
- 11 Chapter 10 — Prediction Loss: How the Model Learns the Next Step available now
- 12 Chapter 11 — The Intuition Behind SIGReg: Letting Representation Space “Breathe” available now
- 13 Chapter 12 — Keep the Mathematics Minimal but Sufficient available now
- 14 Chapter 13 — The Original LeWM’s End-to-End Training Mechanism available now
- 15 Chapter 14 — Train a Model That Does Not Collapse Immediately available now
Part 4 — Putting the Model into Action: Planning in Latent Space
- 16 Chapter 15 — Goal-Conditioned Planning: From “Where Am I?” to “Where Do I Want to Go?” available now
- 17 Chapter 16 — Latent Euclidean Distance: Convenient, but Not Necessarily Reliable Current lesson
- 18 Chapter 17 — CEM: Searching for Actions Through an Elimination Tournament available now
- 19 Chapter 18 — MPC: Do Not Trust the Model for Too Long at Once available now
- 20 Chapter 19 — Long-Horizon Rollouts: How Small Errors Snowball into Major Failures available now
- 21 Chapter 20 — Implement a Minimal LeWM Planner from Scratch available now
Part 5 — Engineering Reproduction: From Paper to Running System
- 22 Chapter 21 — The Official Repository and Experimental Environment available now
- 23 Chapter 22 — First Experiment: A TwoRoom Smoke Test available now
- 24 Chapter 23 — Second Experiment: Reproducing PushT available now
- 25 Chapter 24 — How to Evaluate a World Model Fairly available now
- 26 Chapter 25 — Failure-Diagnosis Manual available now
Part 6 — What Has LeWM Actually Learned?
- 27 Chapter 26 — Linear Probes: Which Physical Variables Are Encoded in the Latent State? available now
- 28 Chapter 27 — Give Latent Space a “Health Check” available now
- 29 Chapter 28 — Violation of Expectation: Is the Model Surprised by “Impossible Events”? available now
- 30 Chapter 29 — How to Discuss “Understanding the World” Rigorously available now
Part 7 — Why “Accurate Prediction” Can Still Produce “Poor Planning”
- 31 Chapter 30 — The Gap Between the Training Objective and the Planning Objective available now
- 32 Chapter 31 — Global Non-Collapse Does Not Guarantee Preservation of Task-Relevant Dynamics available now
- 33 Chapter 32 — When Is an Isotropic Gaussian Prior Too Strong? available now
- 34 Chapter 33 — Long-Horizon Planning: Predict Farther or Plan More Intelligently? available now
- 35 Chapter 34 — From Positional Distance to Task Progress available now
- 36 Chapter 35 — Multi-Task Learning, Real Robots, and Visual Distractions available now
- 37 Chapter 36 — Theoretical Boundaries: When Can the True State Be Identified? available now
Part 8 — From Reproducer to Researcher
Appendices
- 41 Appendix A — The Minimum Necessary Mathematical Toolkit available now
- 42 Appendix B — PyTorch Implementation Quick Reference available now
- 43 Appendix C — Complete Tensor-Shape Table available now
- 44 Appendix D — Experiment Configuration Cards available now
- 45 Appendix E — Paper Timeline and Evidence Levels available now
- 46 Appendix F — Glossary available now
- 47 Appendix G — Reproduction Checklist available now
- 48 Appendix H — Expert-Review Checklist available now
The big picture
- EncodeMake two latent dots.
- MeasureSmaller gap looks better.
- CheckA wall may block the short line.
- TestDoes the ranking prefer the real route?

A tiny story
Two children stand very close on opposite sides of a wall. A ruler says, “Tiny gap!” Their feet say, “Find the door.”
Original LeWorldModel v3 uses a similarly simple planning ruler. It compares the final imagined latent (z_H) with the separately encoded goal latent (z_g):
[ J = \lVert zH-z_g\rVert_2^2 = \sum_i (z{H,i}-z_{g,i})^2 . ]
The ruler is fast, needs no reward labels, and gives one score per candidate plan. Those are excellent baseline properties. They do not make the score shortest-path length, travel time, success probability, safety, or task progress.
The real rule
- The goal is used by the terminal cost; it is not a secret future observation fed into the dynamics predictor.
- Squaring a nonnegative Euclidean gap preserves its ordering within one fixed representation, but the source-faithful LeWM criterion is the summed squared gap.
- Euclidean distance is symmetric: (d(A,B)=d(B,A)). Reachability may be directed or depend on an action budget.
- Raw cost scale depends on latent dimension and coordinate scale. Compare rankings within a named checkpoint before comparing naked magnitudes across models.
- SIGReg can make a population broadly Gaussian without making nearby latent points connected by valid actions.
The planner needs an ordering: candidate A should cost less than candidate B when A is the better decision. A high average correlation on easy pairs can hide a few decisive inversions near a wall or contact boundary.
The trick that can fool us
Make four start–goal pairs with the same straight-line gap:
- both points in one open room;
- points close across a wall;
- points aligned with the doorway;
- points in opposite rooms, far from the doorway.
Use simulator geometry only as an evaluation oracle. Compare latent ranking, shortest feasible route, budgeted reachability, the first action selected under a fixed CEM budget, and the executed result.
First score real encoded endpoints. Then score model-generated endpoints. If the ranking is already wrong with oracle dynamics, inspect representation or cost geometry. If it becomes wrong only after rollout, inspect dynamics drift. If ranking is good but the plan fails, search, support, action conversion, or execution may be responsible.
Experiment receipt
The baseline claim comes from LeWorldModel v3 and frozen author commit 8edfeb3. It supports this narrow wording: LeWM ranks terminal latents with a simple symmetric squared distance whose agreement with route length or reachability must be tested.
The later Objective Bottleneck study reports weak latent-L2/spatial-distance alignment and behavior changes after a cost substitution in one independently implemented TwoRoom system. It is a same-system follow-up, not a second independent replication and not a universal reachability metric.
Record checkpoint, preprocessing, latent width, pair generator, oracle definition, action budget, CEM budget, and seeds. The wall picture is a tutorial counterexample; it is not evidence that every LeWM checkpoint fails it.
Quick check
- Why is squared latent distance a useful baseline?
- What can a symmetric endpoint ruler never express by itself?
- Which oracle swap separates a bad cost ranking from a bad rollout?