Course progress Course outline 48 of 48 lessons available
Part 0 — Reading Guide: What Exactly Are We Going to Learn?
Part 1 — World Models: An Internal Sandbox for the Agent
Part 2 — Turning Images into State: The LeWM Architecture
- 06 Chapter 5 — Trajectory Data: To the Model, the World Is Not an Image Collection available now
- 07 Chapter 6 — The Visual Encoder: Issuing a “State Passport” for Every Frame available now
- 08 Chapter 7 — The Dynamics Predictor: Moving Time Forward in the Mind available now
- 09 Chapter 8 — A Complete Forward Pass: Follow One Batch from Start to Finish available now
Part 3 — Preventing the Model from Cheating: Prediction Loss and SIGReg
- 10 Chapter 9 — The Most Dangerous Shortcut: Representation Collapse available now
- 11 Chapter 10 — Prediction Loss: How the Model Learns the Next Step available now
- 12 Chapter 11 — The Intuition Behind SIGReg: Letting Representation Space “Breathe” available now
- 13 Chapter 12 — Keep the Mathematics Minimal but Sufficient available now
- 14 Chapter 13 — The Original LeWM’s End-to-End Training Mechanism available now
- 15 Chapter 14 — Train a Model That Does Not Collapse Immediately available now
Part 4 — Putting the Model into Action: Planning in Latent Space
- 16 Chapter 15 — Goal-Conditioned Planning: From “Where Am I?” to “Where Do I Want to Go?” available now
- 17 Chapter 16 — Latent Euclidean Distance: Convenient, but Not Necessarily Reliable available now
- 18 Chapter 17 — CEM: Searching for Actions Through an Elimination Tournament available now
- 19 Chapter 18 — MPC: Do Not Trust the Model for Too Long at Once available now
- 20 Chapter 19 — Long-Horizon Rollouts: How Small Errors Snowball into Major Failures available now
- 21 Chapter 20 — Implement a Minimal LeWM Planner from Scratch available now
Part 5 — Engineering Reproduction: From Paper to Running System
- 22 Chapter 21 — The Official Repository and Experimental Environment available now
- 23 Chapter 22 — First Experiment: A TwoRoom Smoke Test available now
- 24 Chapter 23 — Second Experiment: Reproducing PushT available now
- 25 Chapter 24 — How to Evaluate a World Model Fairly available now
- 26 Chapter 25 — Failure-Diagnosis Manual available now
Part 6 — What Has LeWM Actually Learned?
- 27 Chapter 26 — Linear Probes: Which Physical Variables Are Encoded in the Latent State? available now
- 28 Chapter 27 — Give Latent Space a “Health Check” Current lesson
- 29 Chapter 28 — Violation of Expectation: Is the Model Surprised by “Impossible Events”? available now
- 30 Chapter 29 — How to Discuss “Understanding the World” Rigorously available now
Part 7 — Why “Accurate Prediction” Can Still Produce “Poor Planning”
- 31 Chapter 30 — The Gap Between the Training Objective and the Planning Objective available now
- 32 Chapter 31 — Global Non-Collapse Does Not Guarantee Preservation of Task-Relevant Dynamics available now
- 33 Chapter 32 — When Is an Isotropic Gaussian Prior Too Strong? available now
- 34 Chapter 33 — Long-Horizon Planning: Predict Farther or Plan More Intelligently? available now
- 35 Chapter 34 — From Positional Distance to Task Progress available now
- 36 Chapter 35 — Multi-Task Learning, Real Robots, and Visual Distractions available now
- 37 Chapter 36 — Theoretical Boundaries: When Can the True State Be Identified? available now
Part 8 — From Reproducer to Researcher
Appendices
- 41 Appendix A — The Minimum Necessary Mathematical Toolkit available now
- 42 Appendix B — PyTorch Implementation Quick Reference available now
- 43 Appendix C — Complete Tensor-Shape Table available now
- 44 Appendix D — Experiment Configuration Cards available now
- 45 Appendix E — Paper Timeline and Evidence Levels available now
- 46 Appendix F — Glossary available now
- 47 Appendix G — Reproduction Checklist available now
- 48 Appendix H — Expert-Review Checklist available now
The big picture
- SpreadIs the cloud alive?
- NeighborsAre close states useful neighbors?
- PathsDo actions and events bend the path?
- PicturesTurn pretty views into tests.

A tiny story
A patient can have perfect blood pressure and still be unable to climb stairs. The measurement is real; it is simply not the whole body.
A latent cloud can have broad variance while losing contact, velocity, or wall side. It can draw a beautiful smooth path by remembering only a slow background. Freeze the specimen first: checkpoint, preprocessing, history, attachment point, episodes, and sampling rule.
The real rule
Use several instruments and keep their questions separate:
- Mean and coordinate variance: catches constants, dead axes, scale drift, and reload bugs.
- Covariance spectrum and effective dimension: shows where population spread lives. A low value can be collapse or a legitimately simple task.
- Pairwise distance and duplicates: catches pile-ups, but not reachability.
- Nearest neighbors: show original observations, temporal distance, actions, state differences, and next outcomes; exclude self and adjacent duplicates unless duplication is the test.
- Trajectory continuity: place visual change, latent change, and task-variable change on the same time axis. Separate real encoded states from self-fed predictions.
Slice every reading by doorway, wall, free motion, approach, first contact, and maintained contact. Broad nuisance variation can hide a missing task distinction.
LeWM v3 reports that PushT latent trajectories become straighter during training under its chosen post-hoc measure and are straighter than PLDM in that analysis. Baseline LeWM has no explicit straightness loss. Pair straightness with representation spread, action sensitivity, and matched trajectory regions: a collapsed line is perfectly straight.
The trick that can fool us
t-SNE and UMAP are mapmakers that flatten a learned high-dimensional world. Repeat seeds, neighborhood settings, and reasonable metrics; keep every map, audit original-space neighbors, and do not name axes “position” or “time.” Islands, bridges, empty gaps, areas, and far-cluster distances can be projection artifacts.
LeWM v3 uses t-SNE on a structured PushT state grid and reports a qualitative sampled neighborhood pattern. UMAP is a tutorial extension. Neither reveals true topology, reachability, causality, or understanding.
An auxiliary decoder is another viewer. The paper describes a lightweight transformer decoder from the 192-dimensional final [CLS] representation; a 224-pixel, patch-16 example uses 196 learned patch queries. Its chronology wording is inconsistent across sections. Use one explicit tutorial protocol: freeze a named checkpoint, train the decoder afterward on an episode split, and keep reconstruction outside baseline LeWM training and planning.
A sharp image shows decoder-dependent reconstructability, not complete state or model use. The paper’s PushT reconstruction-loss ablation has lower planning success than baseline under its tested setup; it does not prove reconstruction is always harmful.
Experiment receipt
For every neighbor gallery or projection, save sample identities, original latent distances, labels added after fitting, algorithm settings, seed, and the original observations. For every trajectory, mark door crossing, contact, camera change, and action; use the same reference projection across episodes.
Do not confuse baseline’s emergent post-hoc straightening with the separate Temporal Straightening for Latent Planning, which uses an explicit curvature regularizer, a stop-gradient target, and gradient-based planning. Those ingredients do not belong to original LeWM v3.
The source boundary is LeWorldModel v3 and frozen core commit 8edfeb3. The frozen repository does not include the complete t-SNE, straightening, decoder, or VoE scripts and settings, so exact reproduction of those analyses is not established here.
Quick check
- Why can low effective dimension be healthy in one task and harmful in another?
- What turns a nearest-neighbor gallery into a dynamics audit?
- What does a sharp post-hoc decoder establish—and what does it not?