Course progress Course outline 48 of 48 lessons available
Part 0 — Reading Guide: What Exactly Are We Going to Learn?
Part 1 — World Models: An Internal Sandbox for the Agent
Part 2 — Turning Images into State: The LeWM Architecture
- 06 Chapter 5 — Trajectory Data: To the Model, the World Is Not an Image Collection available now
- 07 Chapter 6 — The Visual Encoder: Issuing a “State Passport” for Every Frame available now
- 08 Chapter 7 — The Dynamics Predictor: Moving Time Forward in the Mind available now
- 09 Chapter 8 — A Complete Forward Pass: Follow One Batch from Start to Finish available now
Part 3 — Preventing the Model from Cheating: Prediction Loss and SIGReg
- 10 Chapter 9 — The Most Dangerous Shortcut: Representation Collapse available now
- 11 Chapter 10 — Prediction Loss: How the Model Learns the Next Step available now
- 12 Chapter 11 — The Intuition Behind SIGReg: Letting Representation Space “Breathe” available now
- 13 Chapter 12 — Keep the Mathematics Minimal but Sufficient available now
- 14 Chapter 13 — The Original LeWM’s End-to-End Training Mechanism available now
- 15 Chapter 14 — Train a Model That Does Not Collapse Immediately available now
Part 4 — Putting the Model into Action: Planning in Latent Space
- 16 Chapter 15 — Goal-Conditioned Planning: From “Where Am I?” to “Where Do I Want to Go?” available now
- 17 Chapter 16 — Latent Euclidean Distance: Convenient, but Not Necessarily Reliable available now
- 18 Chapter 17 — CEM: Searching for Actions Through an Elimination Tournament available now
- 19 Chapter 18 — MPC: Do Not Trust the Model for Too Long at Once available now
- 20 Chapter 19 — Long-Horizon Rollouts: How Small Errors Snowball into Major Failures available now
- 21 Chapter 20 — Implement a Minimal LeWM Planner from Scratch available now
Part 5 — Engineering Reproduction: From Paper to Running System
- 22 Chapter 21 — The Official Repository and Experimental Environment available now
- 23 Chapter 22 — First Experiment: A TwoRoom Smoke Test Current lesson
- 24 Chapter 23 — Second Experiment: Reproducing PushT available now
- 25 Chapter 24 — How to Evaluate a World Model Fairly available now
- 26 Chapter 25 — Failure-Diagnosis Manual available now
Part 6 — What Has LeWM Actually Learned?
- 27 Chapter 26 — Linear Probes: Which Physical Variables Are Encoded in the Latent State? available now
- 28 Chapter 27 — Give Latent Space a “Health Check” available now
- 29 Chapter 28 — Violation of Expectation: Is the Model Surprised by “Impossible Events”? available now
- 30 Chapter 29 — How to Discuss “Understanding the World” Rigorously available now
Part 7 — Why “Accurate Prediction” Can Still Produce “Poor Planning”
- 31 Chapter 30 — The Gap Between the Training Objective and the Planning Objective available now
- 32 Chapter 31 — Global Non-Collapse Does Not Guarantee Preservation of Task-Relevant Dynamics available now
- 33 Chapter 32 — When Is an Isotropic Gaussian Prior Too Strong? available now
- 34 Chapter 33 — Long-Horizon Planning: Predict Farther or Plan More Intelligently? available now
- 35 Chapter 34 — From Positional Distance to Task Progress available now
- 36 Chapter 35 — Multi-Task Learning, Real Robots, and Visual Distractions available now
- 37 Chapter 36 — Theoretical Boundaries: When Can the True State Be Identified? available now
Part 8 — From Reproducer to Researcher
Appendices
- 41 Appendix A — The Minimum Necessary Mathematical Toolkit available now
- 42 Appendix B — PyTorch Implementation Quick Reference available now
- 43 Appendix C — Complete Tensor-Shape Table available now
- 44 Appendix D — Experiment Configuration Cards available now
- 45 Appendix E — Paper Timeline and Evidence Levels available now
- 46 Appendix F — Glossary available now
- 47 Appendix G — Reproduction Checklist available now
- 48 Appendix H — Expert-Review Checklist available now
The big picture
- DataAre episodes and actions real?
- StateDo useful differences survive?
- FutureDoes fixed-action rollout stay on track?
- PlanCan cost, search, and execution use it?

A tiny story
To test a smoke alarm, use a small puff of smoke—not a burning house. TwoRoom is that small puff.
It has two continuous rooms, one dividing wall, one doorway, a moving agent, and a goal across the wall. The paper reports 10,000 episodes averaging 92 environment steps. A noisy heuristic steers toward the doorway, then toward the target. So the data form a corridor of support; they are not a uniform survey of every position and action.
Passing this test means the assembled data-to-planning path behaves under one transparent environment. It does not prove contact physics, real-robot safety, broad generalization, or exact paper reproduction.
The real rule
Freeze the artifact and episode split before tuning. Verify time order, action/image alignment, doorway crossings, support, frame skip, and that no window crosses an episode boundary.
Then walk through seven gates in order:
- Data: valid episodes, actions, ranges, and support.
- Representation: spread, covariance spectrum, duplicates, neighbors, and action sensitivity.
- One step: held-out recorded transitions plus action swaps.
- Self-fed: fixed recorded actions, with rollout error by horizon.
- Cost: candidate endpoints separate and wall topology is not misranked.
- Search: CEM proposals and elites actually change.
- Closed loop: denormalized actions reach the environment; new observations trigger replanning.
Stop at the first failed gate and preserve the checkpoint and episode. “First observable break” is safer than “one true cause,” because deeper faults can coexist.
The paper card says 10 epochs, TwoRoom history 1, and 10 non-PushT CEM refinements. The frozen-code card says maximum 100 epochs, global history 3 with no TwoRoom override, and shared CEM refinements 30. The inspected official checkpoint also records history 3. A small run is allowed, but label it TEACHING SMOKE TEST and disclose parameters, updates or epochs, batch/history, precision, device, time, and data exposure.
The trick that can fool us
Suppose average one-step error is low because most transitions happen in open floor. Put start and goal close on opposite sides of the wall. Terminal latent distance may prefer a straight path that the environment blocks.
Run controlled replacements:
- oracle shortest-path cost with learned dynamics;
- oracle dynamics with learned cost;
- a forced doorway waypoint;
- a replay of the selected actions without new search randomness.
A wall collision alone does not identify dynamics. Oracle-cost repair points toward geometry or ranking. Oracle-dynamics repair points toward rollout. A waypoint only shows that supplied topology can help. Action rescaling, sparse doorway data, and finite search remain possible.
The same caution applies to SIGReg. LeWM v3 reports weaker TwoRoom planning than the named baselines and suggests a possible tension between low-diversity, low-intrinsic-dimensional data and a high-dimensional isotropic Gaussian target. That is a hypothesis consistent with the result, not an isolated causal finding.
Experiment receipt
The inspected official data revision is 6903a2d. The frozen evaluation card uses 50 episodes, a same-trajectory goal 25 environment steps ahead, a 50-step action budget, five model blocks per horizon and execution interval, and five raw actions per block. The shared solver card uses 300 candidates, 30 elites, initial variance 1, and 30 refinements. With (H=K=5), feedback comes after the full 25-step planned segment.
The independent TwoRoom reproduction reports 94% for its checkpoint and 84% for the released checkpoint on identical episodes, with one training seed per configuration. It also reports that one-step error does not order long-horizon success across three checkpoints. Its discussion of 100/150 refers to older LeWM v1/v2; v3 now uses the repository’s 25/50 goal/budget card.
Record data revision, split, checkpoint, paper/code/run cards, seven-gate traces, all seeds, and exact action units. A healthy global latent cloud or one successful route still does not prove useful topology everywhere.
Quick check
- Why does the doorway heuristic limit what the data support?
- Why must fixed-action rollout come before planner-selected rollout?
- What can—and cannot—a wall collision diagnose?