Course progress Course outline 48 of 48 lessons available
Part 0 — Reading Guide: What Exactly Are We Going to Learn?
Part 1 — World Models: An Internal Sandbox for the Agent
Part 2 — Turning Images into State: The LeWM Architecture
- 06 Chapter 5 — Trajectory Data: To the Model, the World Is Not an Image Collection Current lesson
- 07 Chapter 6 — The Visual Encoder: Issuing a “State Passport” for Every Frame available now
- 08 Chapter 7 — The Dynamics Predictor: Moving Time Forward in the Mind available now
- 09 Chapter 8 — A Complete Forward Pass: Follow One Batch from Start to Finish available now
Part 3 — Preventing the Model from Cheating: Prediction Loss and SIGReg
- 10 Chapter 9 — The Most Dangerous Shortcut: Representation Collapse available now
- 11 Chapter 10 — Prediction Loss: How the Model Learns the Next Step available now
- 12 Chapter 11 — The Intuition Behind SIGReg: Letting Representation Space “Breathe” available now
- 13 Chapter 12 — Keep the Mathematics Minimal but Sufficient available now
- 14 Chapter 13 — The Original LeWM’s End-to-End Training Mechanism available now
- 15 Chapter 14 — Train a Model That Does Not Collapse Immediately available now
Part 4 — Putting the Model into Action: Planning in Latent Space
- 16 Chapter 15 — Goal-Conditioned Planning: From “Where Am I?” to “Where Do I Want to Go?” available now
- 17 Chapter 16 — Latent Euclidean Distance: Convenient, but Not Necessarily Reliable available now
- 18 Chapter 17 — CEM: Searching for Actions Through an Elimination Tournament available now
- 19 Chapter 18 — MPC: Do Not Trust the Model for Too Long at Once available now
- 20 Chapter 19 — Long-Horizon Rollouts: How Small Errors Snowball into Major Failures available now
- 21 Chapter 20 — Implement a Minimal LeWM Planner from Scratch available now
Part 5 — Engineering Reproduction: From Paper to Running System
- 22 Chapter 21 — The Official Repository and Experimental Environment available now
- 23 Chapter 22 — First Experiment: A TwoRoom Smoke Test available now
- 24 Chapter 23 — Second Experiment: Reproducing PushT available now
- 25 Chapter 24 — How to Evaluate a World Model Fairly available now
- 26 Chapter 25 — Failure-Diagnosis Manual available now
Part 6 — What Has LeWM Actually Learned?
- 27 Chapter 26 — Linear Probes: Which Physical Variables Are Encoded in the Latent State? available now
- 28 Chapter 27 — Give Latent Space a “Health Check” available now
- 29 Chapter 28 — Violation of Expectation: Is the Model Surprised by “Impossible Events”? available now
- 30 Chapter 29 — How to Discuss “Understanding the World” Rigorously available now
Part 7 — Why “Accurate Prediction” Can Still Produce “Poor Planning”
- 31 Chapter 30 — The Gap Between the Training Objective and the Planning Objective available now
- 32 Chapter 31 — Global Non-Collapse Does Not Guarantee Preservation of Task-Relevant Dynamics available now
- 33 Chapter 32 — When Is an Isotropic Gaussian Prior Too Strong? available now
- 34 Chapter 33 — Long-Horizon Planning: Predict Farther or Plan More Intelligently? available now
- 35 Chapter 34 — From Positional Distance to Task Progress available now
- 36 Chapter 35 — Multi-Task Learning, Real Robots, and Visual Distractions available now
- 37 Chapter 36 — Theoretical Boundaries: When Can the True State Be Identified? available now
Part 8 — From Reproducer to Researcher
Appendices
- 41 Appendix A — The Minimum Necessary Mathematical Toolkit available now
- 42 Appendix B — PyTorch Implementation Quick Reference available now
- 43 Appendix C — Complete Tensor-Shape Table available now
- 44 Appendix D — Experiment Configuration Cards available now
- 45 Appendix E — Paper Timeline and Evidence Levels available now
- 46 Appendix F — Glossary available now
- 47 Appendix G — Reproduction Checklist available now
- 48 Appendix H — Expert-Review Checklist available now
The big picture
- Seeobservation now
- Doaligned action block
- See againobserved consequence

A tiny story
Make two photo albums from one TwoRoom journey. The first keeps every frame in order and writes the controls between them. The second shuffles the same pictures and throws the controls away. Both show what the rooms look like. Only the first can teach what followed after a particular action.
An episode boundary matters too. Joining the final frame of one run to the first frame of the next teaches a fake teleport. A world model needs transitions, not merely a visual inventory.
The real rule
LeWM’s world-model training is offline and reward-free. It replays recorded episodes and uses no task reward or goal in the base objective. That does not mean action-free, goal-free deployment, or unlimited extrapolation.
In the frozen data configuration, frameskip=5. One model transition groups five consecutive lower-level environment actions into a dense block; it does not keep only the last action or repeat one action five times. train.py z-scores non-pixel columns including action, and the action encoder width is set programmatically to:
frameskip × environment action dimension
The frozen defaults are history_size=3 and num_preds=1. A sampled window therefore contains four observation positions. The first three latents form causal context; the sequence shifted by one supplies three one-step targets:
z0 -> z1
z1 -> z2
z2 -> z3
num_preds=1 is the offset, not “use only the last pair.” The v3 paper reports history one for TwoRoom and three for PushT/OGBench-Cube, whereas the frozen global configuration uses three. Record which source the run follows.
Coverage is more than an x-y heatmap. A transition depends on what is seen, recent motion or contact, and the proposed action. Visiting every position slowly does not cover fast doorway motion. Seeing a PushT object at one angle while still does not cover the same image during a collision. A fluent prediction outside this combined support is still unsupported.
The trick that fools us
Collector A records millions of beautiful frames while circling one room. Collector B records fewer frames but approaches both doors, collides, pauses, crosses in both directions, and varies speed. A has more images; B may contain much better cross-room dynamics evidence.
Now shift B’s action log by one time index. Its visitation map still looks excellent, but every consequence is paired with the wrong control. Or let a training window cross a reset: the model learns teleportation. Shape checks alone may miss both errors.
Before training, replay complete episodes, inspect timestamps and action blocks, count full doorway crossings, condition action histograms on region and motion, and list missing combinations.
Evidence receipt
LeWorldModel v3, frozen train.py, and the frozen TwoRoom data configuration support the trajectory interface and configuration facts. The paper reports 10,000 TwoRoom episodes averaging 92 environment steps, collected by a noisy doorway-travel heuristic. That is a dataset description, not a universal minimum or proof of complete coverage.
The independent TwoRoom reproduction identified dense action gathering, runtime action width, ImageNet pixel normalization, and action z-scoring as reproduction-critical conventions not all visible in YAML alone. It remains one single-seed environment.
Quick check
- Why can shuffled copies of the same frames not teach action-conditioned dynamics?
- What false event appears when a window crosses an episode reset?
- Why does full x-y visitation still fail to prove transition coverage?