Course progress Course outline 48 of 48 lessons available
Part 0 — Reading Guide: What Exactly Are We Going to Learn?
Part 1 — World Models: An Internal Sandbox for the Agent
Part 2 — Turning Images into State: The LeWM Architecture
- 06 Chapter 5 — Trajectory Data: To the Model, the World Is Not an Image Collection available now
- 07 Chapter 6 — The Visual Encoder: Issuing a “State Passport” for Every Frame available now
- 08 Chapter 7 — The Dynamics Predictor: Moving Time Forward in the Mind available now
- 09 Chapter 8 — A Complete Forward Pass: Follow One Batch from Start to Finish available now
Part 3 — Preventing the Model from Cheating: Prediction Loss and SIGReg
- 10 Chapter 9 — The Most Dangerous Shortcut: Representation Collapse available now
- 11 Chapter 10 — Prediction Loss: How the Model Learns the Next Step available now
- 12 Chapter 11 — The Intuition Behind SIGReg: Letting Representation Space “Breathe” available now
- 13 Chapter 12 — Keep the Mathematics Minimal but Sufficient available now
- 14 Chapter 13 — The Original LeWM’s End-to-End Training Mechanism available now
- 15 Chapter 14 — Train a Model That Does Not Collapse Immediately Current lesson
Part 4 — Putting the Model into Action: Planning in Latent Space
- 16 Chapter 15 — Goal-Conditioned Planning: From “Where Am I?” to “Where Do I Want to Go?” available now
- 17 Chapter 16 — Latent Euclidean Distance: Convenient, but Not Necessarily Reliable available now
- 18 Chapter 17 — CEM: Searching for Actions Through an Elimination Tournament available now
- 19 Chapter 18 — MPC: Do Not Trust the Model for Too Long at Once available now
- 20 Chapter 19 — Long-Horizon Rollouts: How Small Errors Snowball into Major Failures available now
- 21 Chapter 20 — Implement a Minimal LeWM Planner from Scratch available now
Part 5 — Engineering Reproduction: From Paper to Running System
- 22 Chapter 21 — The Official Repository and Experimental Environment available now
- 23 Chapter 22 — First Experiment: A TwoRoom Smoke Test available now
- 24 Chapter 23 — Second Experiment: Reproducing PushT available now
- 25 Chapter 24 — How to Evaluate a World Model Fairly available now
- 26 Chapter 25 — Failure-Diagnosis Manual available now
Part 6 — What Has LeWM Actually Learned?
- 27 Chapter 26 — Linear Probes: Which Physical Variables Are Encoded in the Latent State? available now
- 28 Chapter 27 — Give Latent Space a “Health Check” available now
- 29 Chapter 28 — Violation of Expectation: Is the Model Surprised by “Impossible Events”? available now
- 30 Chapter 29 — How to Discuss “Understanding the World” Rigorously available now
Part 7 — Why “Accurate Prediction” Can Still Produce “Poor Planning”
- 31 Chapter 30 — The Gap Between the Training Objective and the Planning Objective available now
- 32 Chapter 31 — Global Non-Collapse Does Not Guarantee Preservation of Task-Relevant Dynamics available now
- 33 Chapter 32 — When Is an Isotropic Gaussian Prior Too Strong? available now
- 34 Chapter 33 — Long-Horizon Planning: Predict Farther or Plan More Intelligently? available now
- 35 Chapter 34 — From Positional Distance to Task Progress available now
- 36 Chapter 35 — Multi-Task Learning, Real Robots, and Visual Distractions available now
- 37 Chapter 36 — Theoretical Boundaries: When Can the True State Be Identified? available now
Part 8 — From Reproducer to Researcher
Appendices
- 41 Appendix A — The Minimum Necessary Mathematical Toolkit available now
- 42 Appendix B — PyTorch Implementation Quick Reference available now
- 43 Appendix C — Complete Tensor-Shape Table available now
- 44 Appendix D — Experiment Configuration Cards available now
- 45 Appendix E — Paper Timeline and Evidence Levels available now
- 46 Appendix F — Glossary available now
- 47 Appendix G — Reproduction Checklist available now
- 48 Appendix H — Expert-Review Checklist available now
The big picture
- Inspect dataorder, actions, transforms
- Inspect graphshapes + gradients
- Inspect healthfinite, spread, action use
- Then scalerollout before long run

A tiny story
Run one lowers total loss while every image becomes nearly the same latent. Run two produces non-finite values in reduced precision. Run three loads model weights but silently restarts its scheduler and data order. All three dashboards looked reassuring.
A smoke test is not a success certificate. It removes specific wiring failures cheaply, before they consume a long training budget.
The technical backpack
Start by recording provenance: code commit, resolved configuration, dependency versions, data hash, hardware, precision, initialization, loader/split state, worker policy, SIGReg randomness, evaluation randomness, and deterministic-kernel exceptions. The visible seed constructs split/loader generators; it does not establish global determinism. Important stable-pretraining and stable-worldmodel revisions are not pinned by the inspected snapshot.
Inspect one trajectory before and after transforms. Frozen defaults use history three, offset one, frame skip five, and four observation positions. Three shifted pairs contribute. Non-pixel statistics are computed from the loaded dataset before splitting, not train-only. Boundary action NaN values are replaced with zero; that zero is a data convention, not necessarily a physical no-op.
Batch size and projection count solve different problems. Frozen batch size is 128; each SIGReg call sees that forward-pass population per time slice. More projections inspect more directions through the same crowd. Gradient accumulation may enlarge the optimizer batch without enlarging the population used in one SIGReg call.
The frozen training start point is AdamW with learning rate 5e-5, weight decay 1e-3, warmup-cosine scheduling, gradient clipping at 1.0, and bfloat16 mixed precision. Weight decay is not latent Gaussianity; clipping cannot reconnect a missing branch. Log actual learning rate, module-level pre-clip gradients, clipping frequency, and finite intermediates.
The paper describes ten training epochs; frozen YAML says max_epochs: 100. A runtime override can reconcile a run, but the two source facts must remain separate. Report update count and processed windows, not only “epochs.”
Use this order:
provenance
-> raw and transformed sequence
-> forward shapes + finite raw losses
-> backward signals in encoder/action/predictor paths
-> tiny-subset fit + latent spread + action swaps
-> held-out one-step + autoregressive error by horizon
-> then scale workers, epochs, precision, projections, or devices
A model-weight file can support evaluation without supporting exact continuation. Resuming may also require optimizer, scheduler, step, precision, random, and data-order states.
The trick that fools us
Run four cheap interventions on fixed inputs: shuffle actions, repeat one action, shrink the batch until SIGReg becomes noisy, and compare one bfloat16 update with one full-precision diagnostic update. Each changes one suspected mechanism and must remain labelled as a probe, not a new baseline.
Never compress health into one score. Log raw prediction, raw SIGReg, weighted total, learning rate, gradient norms, non-finite counts, latent moments and spectrum, action swaps, held-out one-step error, and recorded-action rollout by horizon. A broad cloud can coexist with action-deaf dynamics; a good one-step curve can coexist with drift.
Evidence receipt
LeWorldModel v3, frozen train.py, utils.py, and module.py support the source settings.
The single-seed TwoRoom reproduction found that dense action gathering, runtime action width, ImageNet pixel normalization, and action z-scoring were needed for its predictor to descend within the ten-epoch budget. It also found a validation artifact up to 300× when projector BatchNorm running variance was far below moving activation scale; the released author checkpoint did not meet that first condition and was unaffected. These are narrow diagnostic findings, not universal factors or evidence that author training fails.
Quick check
- Why can more projections not replace more batch examples?
- What does a tiny-subset fit establish—and what does it not?
- Which states distinguish a model export from a resumable experiment?