Course progress Course outline 48 of 48 lessons available
Part 0 — Reading Guide: What Exactly Are We Going to Learn?
Part 1 — World Models: An Internal Sandbox for the Agent
Part 2 — Turning Images into State: The LeWM Architecture
- 06 Chapter 5 — Trajectory Data: To the Model, the World Is Not an Image Collection available now
- 07 Chapter 6 — The Visual Encoder: Issuing a “State Passport” for Every Frame available now
- 08 Chapter 7 — The Dynamics Predictor: Moving Time Forward in the Mind available now
- 09 Chapter 8 — A Complete Forward Pass: Follow One Batch from Start to Finish available now
Part 3 — Preventing the Model from Cheating: Prediction Loss and SIGReg
- 10 Chapter 9 — The Most Dangerous Shortcut: Representation Collapse available now
- 11 Chapter 10 — Prediction Loss: How the Model Learns the Next Step available now
- 12 Chapter 11 — The Intuition Behind SIGReg: Letting Representation Space “Breathe” available now
- 13 Chapter 12 — Keep the Mathematics Minimal but Sufficient available now
- 14 Chapter 13 — The Original LeWM’s End-to-End Training Mechanism available now
- 15 Chapter 14 — Train a Model That Does Not Collapse Immediately available now
Part 4 — Putting the Model into Action: Planning in Latent Space
- 16 Chapter 15 — Goal-Conditioned Planning: From “Where Am I?” to “Where Do I Want to Go?” available now
- 17 Chapter 16 — Latent Euclidean Distance: Convenient, but Not Necessarily Reliable available now
- 18 Chapter 17 — CEM: Searching for Actions Through an Elimination Tournament available now
- 19 Chapter 18 — MPC: Do Not Trust the Model for Too Long at Once available now
- 20 Chapter 19 — Long-Horizon Rollouts: How Small Errors Snowball into Major Failures available now
- 21 Chapter 20 — Implement a Minimal LeWM Planner from Scratch available now
Part 5 — Engineering Reproduction: From Paper to Running System
- 22 Chapter 21 — The Official Repository and Experimental Environment available now
- 23 Chapter 22 — First Experiment: A TwoRoom Smoke Test available now
- 24 Chapter 23 — Second Experiment: Reproducing PushT available now
- 25 Chapter 24 — How to Evaluate a World Model Fairly available now
- 26 Chapter 25 — Failure-Diagnosis Manual available now
Part 6 — What Has LeWM Actually Learned?
- 27 Chapter 26 — Linear Probes: Which Physical Variables Are Encoded in the Latent State? available now
- 28 Chapter 27 — Give Latent Space a “Health Check” available now
- 29 Chapter 28 — Violation of Expectation: Is the Model Surprised by “Impossible Events”? available now
- 30 Chapter 29 — How to Discuss “Understanding the World” Rigorously available now
Part 7 — Why “Accurate Prediction” Can Still Produce “Poor Planning”
- 31 Chapter 30 — The Gap Between the Training Objective and the Planning Objective available now
- 32 Chapter 31 — Global Non-Collapse Does Not Guarantee Preservation of Task-Relevant Dynamics available now
- 33 Chapter 32 — When Is an Isotropic Gaussian Prior Too Strong? available now
- 34 Chapter 33 — Long-Horizon Planning: Predict Farther or Plan More Intelligently? available now
- 35 Chapter 34 — From Positional Distance to Task Progress available now
- 36 Chapter 35 — Multi-Task Learning, Real Robots, and Visual Distractions Current lesson
- 37 Chapter 36 — Theoretical Boundaries: When Can the True State Be Identified? available now
Part 8 — From Reproducer to Researcher
Appendices
- 41 Appendix A — The Minimum Necessary Mathematical Toolkit available now
- 42 Appendix B — PyTorch Implementation Quick Reference available now
- 43 Appendix C — Complete Tensor-Shape Table available now
- 44 Appendix D — Experiment Configuration Cards available now
- 45 Appendix E — Paper Timeline and Evidence Levels available now
- 46 Appendix F — Glossary available now
- 47 Appendix G — Reproduction Checklist available now
- 48 Appendix H — Expert-Review Checklist available now
The big picture
- New sceneryPixels changed; physics may not.
- New playThe same picture may need another goal.
- Real stageNoise, delay, contact, and safety arrive.
A tiny story: the robot goes on tour
A robot rehearses one play: under a warm lamp, the blue cup always goes left. On tour the curtain turns red, the camera moves, and another play asks the identical cup to go right. Then a hand hides the moment of contact.
Was the change harmless scenery, a different task, a different physical state, or missing evidence? Calling all four “domain shift” erases the answer the model needs. Multi-task training can expose shortcuts, but it can also create new ones: each task may have its own tablecloth, camera, or episode length.

The matrix changes one known factor at a time. It does not declare every background, light, or viewpoint irrelevant.
The technical backpack
Use paired controls. Replay the same simulator state and action while changing one texture, light, or camera variable known not to alter dynamics; then include a positive control that really changes future consequences. Invariance to everything is collapse. Sensitivity to everything is an appearance shortcut. The desired pattern is selective and must also survive closed-loop planning.
If one image belongs to two tasks with different correct actions, no current-frame encoder can guess an omitted task. Supply a goal image, token, or language condition; or use history when the missing variable leaves a trace. A token can become a dataset-ID shortcut, and history cannot recover information never observed.
Later evidence must keep its identity:
| Later route | Extra signal or setting | Narrow boundary |
|---|---|---|
| TC-LeWM v2 | temporally centered SIGReg; multi-task LIBERO with a from-scratch multi-view ViT-S/16 and task-conditioned flow-matching behavior cloning | not baseline CEM/MPC, TwoRoom, or robot deployment |
| Depth-regularized JEPA v1 | depth supervision and training-only overparameterization; reported 18M model on real agricultural-robot video | offline visual-odometry probes, surprise, and latent rollouts—not online planning or safety |
| PhyLatent v1 | simulator physical-state supervision | privileged training, not pixels-only recovery |
| PSG-JEPA v1 | training-only proprioceptive-state and multi-horizon joint-change grounding | author-reported Mobile ALOHA downstream policy evidence, not original-LeWM MPC or a safety certificate |
| stable-worldmodel v1 | shared data, training, solver, MPC, and evaluation plumbing | infrastructure can expose choices; it cannot certify neutral or independently reproduced science |
Original LeWM v3 remains fully end-to-end: no stop-gradient target, EMA teacher, or pretrained visual encoder. Later auxiliary heads must not redraw its graph. A larger predictor is also only a hypothesis: compare parameter- and compute-matched widths, several seeds, one-step prediction, self-fed rollout, action sensitivity, nuisance invariance, and planning.
Try to break the idea
Make two episodes whose current image and action history are identical. Task A says “cup left”; Task B says “cup right.” Hide the instruction. Consistent choice of one side reveals a dataset prior, not recovered context. Then add, one at a time, a goal image, task token, earlier instruction frame, and a background correlated with task during training but swapped at test.
For deployment, climb an evidence ladder: original simulator → controlled visual changes → sensor noise/missing frames → delayed and imperfect actuation → hardware/object/calibration variation → safety-constrained runs with abort and recovery rules. Passing one rung does not grant the next. MPC feedback limits open-loop trust but does not make the executed step safe.
Experiment receipt and evidence boundary
- Sources: LeWorldModel v3; TC-LeWM v2; depth-regularized JEPA v1; PhyLatent v1; PSG-JEPA v1 and its project page; stable-worldmodel v1.
- Frozen code: LeWM
8edfeb3; PSG-JEPA8de96b8, dated 2026-08-14; stable-worldmodeladdbab4, dated 2026-08-18, while formal release remained0.1.1from 2026-06-06. No author-linked code was identified for TC-LeWM, the depth paper, or PhyLatent. - Dates: depth 2026-07-15; TC-LeWM v2 2026-07-31; PhyLatent 2026-08-06; PSG-JEPA 2026-08-07; evidence cutoff 2026-08-20, Asia/Tokyo.
- These later outcomes are author-reported; independent reproduction is not established for them or their robot results. Report the exact method, controller, data, intervention, and safety protocol. “Real video” is not “online robot planning,” and “robot policy experiment” is not “safe deployment.”
Three quick questions
- Why is a camera change not automatically a nuisance?
- What does the identical-frame, opposite-task example prove about missing context?
- Which new evidence is required before a simulator robustness claim becomes a robot safety claim?