Course progress Course outline 48 of 48 lessons available
Part 0 — Reading Guide: What Exactly Are We Going to Learn?
Part 1 — World Models: An Internal Sandbox for the Agent
Part 2 — Turning Images into State: The LeWM Architecture
- 06 Chapter 5 — Trajectory Data: To the Model, the World Is Not an Image Collection available now
- 07 Chapter 6 — The Visual Encoder: Issuing a “State Passport” for Every Frame available now
- 08 Chapter 7 — The Dynamics Predictor: Moving Time Forward in the Mind available now
- 09 Chapter 8 — A Complete Forward Pass: Follow One Batch from Start to Finish available now
Part 3 — Preventing the Model from Cheating: Prediction Loss and SIGReg
- 10 Chapter 9 — The Most Dangerous Shortcut: Representation Collapse available now
- 11 Chapter 10 — Prediction Loss: How the Model Learns the Next Step available now
- 12 Chapter 11 — The Intuition Behind SIGReg: Letting Representation Space “Breathe” available now
- 13 Chapter 12 — Keep the Mathematics Minimal but Sufficient available now
- 14 Chapter 13 — The Original LeWM’s End-to-End Training Mechanism available now
- 15 Chapter 14 — Train a Model That Does Not Collapse Immediately available now
Part 4 — Putting the Model into Action: Planning in Latent Space
- 16 Chapter 15 — Goal-Conditioned Planning: From “Where Am I?” to “Where Do I Want to Go?” available now
- 17 Chapter 16 — Latent Euclidean Distance: Convenient, but Not Necessarily Reliable available now
- 18 Chapter 17 — CEM: Searching for Actions Through an Elimination Tournament available now
- 19 Chapter 18 — MPC: Do Not Trust the Model for Too Long at Once available now
- 20 Chapter 19 — Long-Horizon Rollouts: How Small Errors Snowball into Major Failures available now
- 21 Chapter 20 — Implement a Minimal LeWM Planner from Scratch available now
Part 5 — Engineering Reproduction: From Paper to Running System
- 22 Chapter 21 — The Official Repository and Experimental Environment Current lesson
- 23 Chapter 22 — First Experiment: A TwoRoom Smoke Test available now
- 24 Chapter 23 — Second Experiment: Reproducing PushT available now
- 25 Chapter 24 — How to Evaluate a World Model Fairly available now
- 26 Chapter 25 — Failure-Diagnosis Manual available now
Part 6 — What Has LeWM Actually Learned?
- 27 Chapter 26 — Linear Probes: Which Physical Variables Are Encoded in the Latent State? available now
- 28 Chapter 27 — Give Latent Space a “Health Check” available now
- 29 Chapter 28 — Violation of Expectation: Is the Model Surprised by “Impossible Events”? available now
- 30 Chapter 29 — How to Discuss “Understanding the World” Rigorously available now
Part 7 — Why “Accurate Prediction” Can Still Produce “Poor Planning”
- 31 Chapter 30 — The Gap Between the Training Objective and the Planning Objective available now
- 32 Chapter 31 — Global Non-Collapse Does Not Guarantee Preservation of Task-Relevant Dynamics available now
- 33 Chapter 32 — When Is an Isotropic Gaussian Prior Too Strong? available now
- 34 Chapter 33 — Long-Horizon Planning: Predict Farther or Plan More Intelligently? available now
- 35 Chapter 34 — From Positional Distance to Task Progress available now
- 36 Chapter 35 — Multi-Task Learning, Real Robots, and Visual Distractions available now
- 37 Chapter 36 — Theoretical Boundaries: When Can the True State Be Identified? available now
Part 8 — From Reproducer to Researcher
Appendices
- 41 Appendix A — The Minimum Necessary Mathematical Toolkit available now
- 42 Appendix B — PyTorch Implementation Quick Reference available now
- 43 Appendix C — Complete Tensor-Shape Table available now
- 44 Appendix D — Experiment Configuration Cards available now
- 45 Appendix E — Paper Timeline and Evidence Levels available now
- 46 Appendix F — Glossary available now
- 47 Appendix G — Reproduction Checklist available now
- 48 Appendix H — Expert-Review Checklist available now
The big picture
- PaperWhat experiment was described?
- CoreWhich frozen LeWM code?
- WorkshopWhich dependency revisions?
- ReceiptWhat actually ran?

A tiny story
You receive an official engine, an official toolbox, and an official key. The key still may not fit the engine because the toolbox changed after the key was made.
LeWorldModel is similarly layered. The paper is the design description. The small frozen repository is the core engine. stable-pretraining and stable-worldmodel provide training, environments, planning, and evaluation. Data and checkpoints are separate artifacts. “Official + official” does not automatically recreate the authors’ historical workshop.
The real rule
The frozen LeWM core is commit 8edfeb3. jepa.py/module.py own the model and SIGReg, train.py joins data and losses, and eval.py joins a frozen model, CEM or random actions, environments, and metrics.
The core has no dependency lockfile. This course inspected stable-worldmodel@addbab4 and stable-pretraining@9aa93f8. Those revisions make current behavior inspectable; they are not LeWM-supplied historical pins.
A five-part health check is enough to call a run mechanically ready:
- record Python, PyTorch, accelerator stack, and operating system;
- record all source revisions;
- hash data and checkpoint files and record decompressed names;
- verify an episode-safe batch, finite forward pass, and short environment interaction;
- archive the resolved configuration, hardware, logs, and outputs.
The README names Python 3.10. Hydra must be fully composed before diagnosing missing fields: launcher fragments supply wandb and cache_dir. The resolved run sheet—not scattered YAML—is the executable experiment record.
Original v3 remains one shared, fully trainable visual encoder plus an action-conditioned predictor and SIGReg. Targets stay connected. There is no stop-gradient target, EMA teacher, pretrained visual backbone, reward, goal, or CEM supervision in training.
The trick that can fool us
Paper, code, and run values are three different cards:
| Setting | Paper card | Frozen-code card |
|---|---|---|
| Named-task training | 10 epochs | maximum 100 |
| TwoRoom history | 1 | global default 3 |
| SIGReg weight | 0.1 unless stated otherwise | 0.09 |
| Non-PushT CEM refinements | 10 | shared config 30 |
| PushT training data | trajectory description | YAML names Lance; public artifact is compressed HDF5 |
Do not average these cards or choose the convenient value silently.
Another trap is checkpoint identity. The inspected official repositories provide weights.pt and config.json targeting classes under stable_worldmodel.wm.lewm; the frozen core constructs jepa.JEPA and documents conversion. Record class path, architecture, preprocessing, source and dependency revisions, hash, and conversion.
The training script fits non-image normalizers on the full loaded dataset before its delegated random split, so this checked-in path is not train-only normalization. Episode grouping also depends on the resolved stable-pretraining revision. A successful load cannot erase those scientific differences.
Experiment receipt
Official artifact revisions inspected here are TwoRoom data 6903a2d, PushT data 655cd44, TwoRoom model 77adaae, and PushT model 22b330c. Dataset receipts must include episode/time fields, preprocessing, normalization split, temporal window, frame skip, loader, format conversion, and a test that no window crosses an episode boundary.
The paper describes roughly 15 million parameters trained on one GPU in a few hours and reports planning speedups up to 48× under its setup. The frozen trainer targets GPU, automatic device choice, and bfloat16, but the exact historical GPU and timing stack are missing. Preserve those statements as author-reported and configuration-dependent.
An independent TwoRoom study with public commit efa9e5d reports four consequential conventions outside released YAML: five raw actions per frame-skip block, action-encoder width 10, ImageNet normalization, and per-coordinate action z-scoring. Under its matched 50-episode protocol it reports 94% for its checkpoint, 84% for the released author checkpoint, and position-probe Pearson 0.9988. It is one environment and one training seed per configuration—not the four-task table or exact historical environment.
Quick check
- What does “mechanically ready” prove, and what does it not prove?
- Why must a resolved Hydra configuration be saved?
- Why can two official artifacts still fail to recreate the paper environment?