Course progress Course outline 48 of 48 lessons available
Part 0 — Reading Guide: What Exactly Are We Going to Learn?
Part 1 — World Models: An Internal Sandbox for the Agent
Part 2 — Turning Images into State: The LeWM Architecture
- 06 Chapter 5 — Trajectory Data: To the Model, the World Is Not an Image Collection available now
- 07 Chapter 6 — The Visual Encoder: Issuing a “State Passport” for Every Frame available now
- 08 Chapter 7 — The Dynamics Predictor: Moving Time Forward in the Mind available now
- 09 Chapter 8 — A Complete Forward Pass: Follow One Batch from Start to Finish available now
Part 3 — Preventing the Model from Cheating: Prediction Loss and SIGReg
- 10 Chapter 9 — The Most Dangerous Shortcut: Representation Collapse available now
- 11 Chapter 10 — Prediction Loss: How the Model Learns the Next Step available now
- 12 Chapter 11 — The Intuition Behind SIGReg: Letting Representation Space “Breathe” available now
- 13 Chapter 12 — Keep the Mathematics Minimal but Sufficient available now
- 14 Chapter 13 — The Original LeWM’s End-to-End Training Mechanism available now
- 15 Chapter 14 — Train a Model That Does Not Collapse Immediately available now
Part 4 — Putting the Model into Action: Planning in Latent Space
- 16 Chapter 15 — Goal-Conditioned Planning: From “Where Am I?” to “Where Do I Want to Go?” available now
- 17 Chapter 16 — Latent Euclidean Distance: Convenient, but Not Necessarily Reliable available now
- 18 Chapter 17 — CEM: Searching for Actions Through an Elimination Tournament available now
- 19 Chapter 18 — MPC: Do Not Trust the Model for Too Long at Once available now
- 20 Chapter 19 — Long-Horizon Rollouts: How Small Errors Snowball into Major Failures available now
- 21 Chapter 20 — Implement a Minimal LeWM Planner from Scratch available now
Part 5 — Engineering Reproduction: From Paper to Running System
- 22 Chapter 21 — The Official Repository and Experimental Environment available now
- 23 Chapter 22 — First Experiment: A TwoRoom Smoke Test available now
- 24 Chapter 23 — Second Experiment: Reproducing PushT available now
- 25 Chapter 24 — How to Evaluate a World Model Fairly available now
- 26 Chapter 25 — Failure-Diagnosis Manual available now
Part 6 — What Has LeWM Actually Learned?
- 27 Chapter 26 — Linear Probes: Which Physical Variables Are Encoded in the Latent State? Current lesson
- 28 Chapter 27 — Give Latent Space a “Health Check” available now
- 29 Chapter 28 — Violation of Expectation: Is the Model Surprised by “Impossible Events”? available now
- 30 Chapter 29 — How to Discuss “Understanding the World” Rigorously available now
Part 7 — Why “Accurate Prediction” Can Still Produce “Poor Planning”
- 31 Chapter 30 — The Gap Between the Training Objective and the Planning Objective available now
- 32 Chapter 31 — Global Non-Collapse Does Not Guarantee Preservation of Task-Relevant Dynamics available now
- 33 Chapter 32 — When Is an Isotropic Gaussian Prior Too Strong? available now
- 34 Chapter 33 — Long-Horizon Planning: Predict Farther or Plan More Intelligently? available now
- 35 Chapter 34 — From Positional Distance to Task Progress available now
- 36 Chapter 35 — Multi-Task Learning, Real Robots, and Visual Distractions available now
- 37 Chapter 36 — Theoretical Boundaries: When Can the True State Be Identified? available now
Part 8 — From Reproducer to Researcher
Appendices
- 41 Appendix A — The Minimum Necessary Mathematical Toolkit available now
- 42 Appendix B — PyTorch Implementation Quick Reference available now
- 43 Appendix C — Complete Tensor-Shape Table available now
- 44 Appendix D — Experiment Configuration Cards available now
- 45 Appendix E — Paper Timeline and Evidence Levels available now
- 46 Appendix F — Glossary available now
- 47 Appendix G — Reproduction Checklist available now
- 48 Appendix H — Expert-Review Checklist available now
The big picture
- FreezeKeep one named representation unchanged.
- Split episodesNo neighboring-frame leakage.
- Fit a readoutLinear ruler, then controlled MLP.
- Say only “recoverable”Use needs a different test.

A tiny story
An airport scanner sees a key inside a suitcase. That proves the scanner can recover key information from this view. It does not prove the traveler knows the key is there or will use it.
A probe works the same way. Freeze a trained representation, collect its latent vectors, and train a small readout from those vectors to simulator labels. A successful readout shows statistical accessibility for this representation, readout, split, and distribution. It does not prove the predictor or planner uses the information.
The real rule
Name the attachment point: encoder output, projected state, predictor input, or predictor output are different objects. Freeze checkpoint, evaluation mode, crop, resize, normalization, pooling, projector, and history construction. If probe gradients update the encoder, the experiment has become supervised fine-tuning.
Split complete episodes before extracting embeddings or fitting target normalization. Train on training episodes, choose hyperparameters on validation episodes, and evaluate once on untouched test episodes. Adjacent frames on opposite sides of a random frame split are near duplicates, not independent generalization.
LeWM v3 reports:
- TwoRoom: two-dimensional agent position;
- PushT: agent location, block location, block angle;
- OGBench-Cube: joint position/velocity, end-effector position/yaw, gripper state, block position/quaternion/yaw.
Target distance is a tutorial extension, not a v3 probe target. The paper reports MSE and Pearson correlation for linear and nonlinear probes. Correlation can stay high with scale bias; average MSE can hide a bad dimension or rare contact slice.
A linear probe tests simple affine accessibility. An MLP tests accessibility to a richer hypothesis class and may exploit shortcuts. Report capacity, regularization, optimization and tuning budget for both.
The trick that can fool us
Imagine every TwoRoom episode crosses the door at roughly the same time, and a faint corner mark brightens with frame index. A random-frame split lets an MLP—and perhaps a linear probe—predict position from time.
Now split by episode, randomize policy speed, and remove the mark. If performance falls, the original result still meant exactly this: position was recoverable where time and position were coupled.
Use controls: constant-mean predictor, simple pixels or privileged observable baseline, untrained encoder, shuffled labels, and training-versus-test curves. Slice TwoRoom by open floor, wall, and doorway; slice PushT by free motion, approach, first contact, sliding, and rotation where labels are reliable.
To test use, intervene after probing. Perturb a variable-associated subspace and compare with matched-size random directions, norm-preserving noise, and no intervention on the same episodes and candidate actions. A behavior change can still come from entanglement or an out-of-distribution latent, so even this supports “behaviorally implicated under this intervention,” not automatic causality.
Experiment receipt
Later results make the boundary concrete:
- Objective Bottleneck reports terminal latent L2 versus true endpoint distance at Pearson (r=0.426), while a ridge position probe reaches held-out (R^2=0.9922), in one TwoRoom follow-up system.
- SCALE reports a state-regression control that can match or exceed full-embedding decodability on some tasks without matching distance alignment; it improves 13 of 15 task–solver averages versus 15 of 15 for SCALE.
- Decision-Metric Alignment reports probe (R^2) differences below 0.03 across four non-collapsed PushT variants, while Plan–Real Spearman moves from +0.280 to about +0.410–+0.420 and online success spans 49.3%–92.7%. Each configuration has one training run; three seeds measure evaluation variation.
These are separate, author-reported protocols. They show that decodability can separate from planner-facing geometry; they do not identify one universal mechanism. The original v3 probe implementation details are insufficient for an exact reconstruction of every reported probe.
Quick check
- What must remain frozen for a probe to describe the original representation?
- What does MLP success establish beyond linear success—and what does it not?
- Which intervention moves a claim from recoverability toward behavioral relevance?