Course progress Course outline 48 of 48 lessons available
Part 0 — Reading Guide: What Exactly Are We Going to Learn?
Part 1 — World Models: An Internal Sandbox for the Agent
Part 2 — Turning Images into State: The LeWM Architecture
- 06 Chapter 5 — Trajectory Data: To the Model, the World Is Not an Image Collection available now
- 07 Chapter 6 — The Visual Encoder: Issuing a “State Passport” for Every Frame available now
- 08 Chapter 7 — The Dynamics Predictor: Moving Time Forward in the Mind available now
- 09 Chapter 8 — A Complete Forward Pass: Follow One Batch from Start to Finish available now
Part 3 — Preventing the Model from Cheating: Prediction Loss and SIGReg
- 10 Chapter 9 — The Most Dangerous Shortcut: Representation Collapse available now
- 11 Chapter 10 — Prediction Loss: How the Model Learns the Next Step available now
- 12 Chapter 11 — The Intuition Behind SIGReg: Letting Representation Space “Breathe” available now
- 13 Chapter 12 — Keep the Mathematics Minimal but Sufficient available now
- 14 Chapter 13 — The Original LeWM’s End-to-End Training Mechanism available now
- 15 Chapter 14 — Train a Model That Does Not Collapse Immediately available now
Part 4 — Putting the Model into Action: Planning in Latent Space
- 16 Chapter 15 — Goal-Conditioned Planning: From “Where Am I?” to “Where Do I Want to Go?” available now
- 17 Chapter 16 — Latent Euclidean Distance: Convenient, but Not Necessarily Reliable available now
- 18 Chapter 17 — CEM: Searching for Actions Through an Elimination Tournament available now
- 19 Chapter 18 — MPC: Do Not Trust the Model for Too Long at Once available now
- 20 Chapter 19 — Long-Horizon Rollouts: How Small Errors Snowball into Major Failures available now
- 21 Chapter 20 — Implement a Minimal LeWM Planner from Scratch available now
Part 5 — Engineering Reproduction: From Paper to Running System
- 22 Chapter 21 — The Official Repository and Experimental Environment available now
- 23 Chapter 22 — First Experiment: A TwoRoom Smoke Test available now
- 24 Chapter 23 — Second Experiment: Reproducing PushT available now
- 25 Chapter 24 — How to Evaluate a World Model Fairly available now
- 26 Chapter 25 — Failure-Diagnosis Manual available now
Part 6 — What Has LeWM Actually Learned?
- 27 Chapter 26 — Linear Probes: Which Physical Variables Are Encoded in the Latent State? available now
- 28 Chapter 27 — Give Latent Space a “Health Check” available now
- 29 Chapter 28 — Violation of Expectation: Is the Model Surprised by “Impossible Events”? available now
- 30 Chapter 29 — How to Discuss “Understanding the World” Rigorously available now
Part 7 — Why “Accurate Prediction” Can Still Produce “Poor Planning”
- 31 Chapter 30 — The Gap Between the Training Objective and the Planning Objective available now
- 32 Chapter 31 — Global Non-Collapse Does Not Guarantee Preservation of Task-Relevant Dynamics Current lesson
- 33 Chapter 32 — When Is an Isotropic Gaussian Prior Too Strong? available now
- 34 Chapter 33 — Long-Horizon Planning: Predict Farther or Plan More Intelligently? available now
- 35 Chapter 34 — From Positional Distance to Task Progress available now
- 36 Chapter 35 — Multi-Task Learning, Real Robots, and Visual Distractions available now
- 37 Chapter 36 — Theoretical Boundaries: When Can the True State Be Identified? available now
Part 8 — From Reproducer to Researcher
Appendices
- 41 Appendix A — The Minimum Necessary Mathematical Toolkit available now
- 42 Appendix B — PyTorch Implementation Quick Reference available now
- 43 Appendix C — Complete Tensor-Shape Table available now
- 44 Appendix D — Experiment Configuration Cards available now
- 45 Appendix E — Paper Timeline and Evidence Levels available now
- 46 Appendix F — Glossary available now
- 47 Appendix G — Reproduction Checklist available now
- 48 Appendix H — Expert-Review Checklist available now
The big picture
- Change appearanceKeep physical transition fixed.
- Change physicsKeep appearance nearly fixed.
- Change actionKeep the full history fixed.

A tiny story
A passport office gives everyone a different card. The archive never collapses. But each card records shirt color and wallpaper while omitting direction of motion and which gate can be crossed.
Latents can do that too. Background texture can fill many dimensions while a small velocity or contact variable disappears. Global variance, rank, and a pretty projection rule out some failures; they do not prove task-relevant dynamics survived.
The real rule
Use matched interventions.
- Hold simulator state and transition fixed; change only an appearance factor the environment guarantees is dynamically irrelevant.
- Hold appearance nearly fixed; change one physical variable known to alter the next outcome.
- Hold the complete observable history fixed; change only the proposed action, reset the simulator, and compare every predicted branch with its matching real transition.
“Appearance” is not automatically nuisance. Traffic-light color, shadow, or texture can signal real control state. The experiment—not intuition—must guarantee invariance.
Observation sufficiency comes first. One frame cannot reveal two opposite velocities that render identically. Give enough history to make motion observable before diagnosing representation failure. If history contains the clue but latents and predicted futures still collide, the test is stronger.
Action sensitivity has three levels: does output change, does it change in the right direction, and does the response depend on state–action interaction? Shuffled, zero, opposing, and withheld actions help, but can create out-of-support pairs. Nonzero response alone is not correct dynamics.
The trick that can fool us
Make the same PushT pose appear under three harmless floor textures. Then make a nearly identical view with opposite motion or different contact. If the harmless texture changes prediction more than the physical intervention, the model prioritizes the wrong distinction in this controlled environment.
PhyLatent names three operational failures:
- physical invariance collapse: appearance-only movement dominates a selected physical transition;
- physical identifiability collapse: states far apart under selected physical variables map close;
- counterfactual dynamics collapse: materially different action consequences compress together.
These are definitions inside one recent paper, not a universal taxonomy or theorem-level identification. Thresholds, simulator variables, scale, observability, and pair selection matter. PhyLatent’s described baseline uses a stopped future target, unlike original connected-target LeWM v3, and its proposed training uses privileged physical targets.
Experiment receipt
Later LeWM-family proposals make different supervision purchases:
- PSG-JEPA keeps LeWM forward loss and SIGReg, then adds training-only heads for proprioceptive state and multi-horizon joint-angle change (q_{t+k}-q_t). Privileged labels are absent at test time.
- SCALE adds (1-\mathrm{corr}(d_z,d_q)), aligning squared latent pair distances with squared distances in a standardized selected simulator state. Its regression control uses the same labels without shaping pairwise geometry. Authors report improvement in all 15 task–solver averages; each training configuration has one seed, while evaluation sets measure episode sampling.
- AC-MTM removes SIGReg and keeps connected forward MSE plus a training-only Action-NCE inverse head that identifies the recorded continuous-action block among in-batch actions. It has no target network, stop-gradient, pretrained encoder, or reconstruction loss; the inverse head is discarded at test time. In a matched three-training-seed study it matches SIGReg on average over four standard tasks but trails on PushT, and reports 80% versus 58% on its OGBench Visual Scene trajectory-goal stress test. The 52% random baseline is one 50-episode run.
PSG-JEPA and SCALE use privileged simulator labels. AC-MTM instead requires informative action variation and observable consequences. None is an ingredient of original v3, and none is independently reproduced here.
For every new diagnostic, save simulator reset state, history, intervention, real outcomes, appearance control, support, checkpoint, thresholds, and seeds. Cross any representation improvement with fixed-planner ranking and success; a better probe can remain unread by the cost.
Quick check
- How can a globally non-collapsed representation lose a control variable?
- What must be true before an appearance change is called irrelevant?
- Why must action branches be compared with real matching outcomes?