Course progress Course outline 48 of 48 lessons available
Part 0 — Reading Guide: What Exactly Are We Going to Learn?
Part 1 — World Models: An Internal Sandbox for the Agent
Part 2 — Turning Images into State: The LeWM Architecture
- 06 Chapter 5 — Trajectory Data: To the Model, the World Is Not an Image Collection available now
- 07 Chapter 6 — The Visual Encoder: Issuing a “State Passport” for Every Frame available now
- 08 Chapter 7 — The Dynamics Predictor: Moving Time Forward in the Mind available now
- 09 Chapter 8 — A Complete Forward Pass: Follow One Batch from Start to Finish available now
Part 3 — Preventing the Model from Cheating: Prediction Loss and SIGReg
- 10 Chapter 9 — The Most Dangerous Shortcut: Representation Collapse available now
- 11 Chapter 10 — Prediction Loss: How the Model Learns the Next Step available now
- 12 Chapter 11 — The Intuition Behind SIGReg: Letting Representation Space “Breathe” available now
- 13 Chapter 12 — Keep the Mathematics Minimal but Sufficient available now
- 14 Chapter 13 — The Original LeWM’s End-to-End Training Mechanism available now
- 15 Chapter 14 — Train a Model That Does Not Collapse Immediately available now
Part 4 — Putting the Model into Action: Planning in Latent Space
- 16 Chapter 15 — Goal-Conditioned Planning: From “Where Am I?” to “Where Do I Want to Go?” available now
- 17 Chapter 16 — Latent Euclidean Distance: Convenient, but Not Necessarily Reliable available now
- 18 Chapter 17 — CEM: Searching for Actions Through an Elimination Tournament available now
- 19 Chapter 18 — MPC: Do Not Trust the Model for Too Long at Once available now
- 20 Chapter 19 — Long-Horizon Rollouts: How Small Errors Snowball into Major Failures available now
- 21 Chapter 20 — Implement a Minimal LeWM Planner from Scratch Current lesson
Part 5 — Engineering Reproduction: From Paper to Running System
- 22 Chapter 21 — The Official Repository and Experimental Environment available now
- 23 Chapter 22 — First Experiment: A TwoRoom Smoke Test available now
- 24 Chapter 23 — Second Experiment: Reproducing PushT available now
- 25 Chapter 24 — How to Evaluate a World Model Fairly available now
- 26 Chapter 25 — Failure-Diagnosis Manual available now
Part 6 — What Has LeWM Actually Learned?
- 27 Chapter 26 — Linear Probes: Which Physical Variables Are Encoded in the Latent State? available now
- 28 Chapter 27 — Give Latent Space a “Health Check” available now
- 29 Chapter 28 — Violation of Expectation: Is the Model Surprised by “Impossible Events”? available now
- 30 Chapter 29 — How to Discuss “Understanding the World” Rigorously available now
Part 7 — Why “Accurate Prediction” Can Still Produce “Poor Planning”
- 31 Chapter 30 — The Gap Between the Training Objective and the Planning Objective available now
- 32 Chapter 31 — Global Non-Collapse Does Not Guarantee Preservation of Task-Relevant Dynamics available now
- 33 Chapter 32 — When Is an Isotropic Gaussian Prior Too Strong? available now
- 34 Chapter 33 — Long-Horizon Planning: Predict Farther or Plan More Intelligently? available now
- 35 Chapter 34 — From Positional Distance to Task Progress available now
- 36 Chapter 35 — Multi-Task Learning, Real Robots, and Visual Distractions available now
- 37 Chapter 36 — Theoretical Boundaries: When Can the True State Be Identified? available now
Part 8 — From Reproducer to Researcher
Appendices
- 41 Appendix A — The Minimum Necessary Mathematical Toolkit available now
- 42 Appendix B — PyTorch Implementation Quick Reference available now
- 43 Appendix C — Complete Tensor-Shape Table available now
- 44 Appendix D — Experiment Configuration Cards available now
- 45 Appendix E — Paper Timeline and Evidence Levels available now
- 46 Appendix F — Glossary available now
- 47 Appendix G — Reproduction Checklist available now
- 48 Appendix H — Expert-Review Checklist available now
The big picture
- FreezeEncoder and dynamics do not learn now.
- SearchCEM proposes complete plans.
- RankCompare imagined endpoints with the goal.
- ExecuteRun K blocks, observe, repeat.

A tiny story
Before building a toy railway, label every connector. Which wire carries a real observation? Which carries an imagined latent? Which action reaches the motor?
A minimal LeWM planner should do the same. “From scratch” here means building clear interfaces around a trained, frozen checkpoint. It does not mean inventing a new LeWM implementation or pretending tutorial pseudocode is the authors’ exact runtime.
The real rule
Keep these objects separate:
current images [B, T_context, C, H_image, W_image]
current latents [B, T_context, D]
goal latent [B, D]
candidate actions[B, N, H_plan, A_block]
imagined latents [B, N, T_rollout, D]
candidate costs [B, N]
execution prefix [B, K, A_block]
The separately encoded goal enters the terminal cost, not the dynamics predictor. Correct vectorization may flatten (B\times N) for throughput, but it must preserve each environment, candidate, and goal identity. Test a tiny loop against the vectorized path; permuting candidates should only permute the matching outputs.
Baseline LeWM scores the declared final predicted latent with squared Euclidean distance to the goal. CEM refits an action proposal while the model stays frozen. Record both the best sampled sequence and the final proposal mean. Appendix B and the inspected stable-worldmodel snapshots return the final mean; changing execution to best_seen is a labelled variant.
For the reported setup, (H=5) model blocks, (K=5), and each block contains five environment actions. Only the selected prefix crosses into reality. Predicted latents never become fake measurements in the real observation buffer.
The trick that can fool us
Choose a TwoRoom goal close across a wall. Run the same frozen task with a fixed candidate and real-action budget under four named cells:
- learned dynamics + learned latent cost;
- oracle dynamics + learned cost;
- learned dynamics + oracle cost, only if imagined states have a valid scoring bridge;
- both oracles, with the same finite search.
Add an equal-budget random policy and, separately, a support-constrained proposal. These are different questions. Oracle dynamics tests rollout replacement. Oracle cost tests ranking replacement. Both oracles still do not make CEM optimal. A support constraint changes where search looks.
If oracle dynamics repairs the episode, rollout is implicated. If oracle cost repairs it, ranking is implicated. If both fail, inspect candidate coverage, action representation, reachability, and execution. One final frame cannot identify a cause.
Experiment receipt
At every MPC cycle, save current and goal identities, preprocessing, checkpoint, (H), (K), action-block semantics, all seeds, sampled and elite costs, proposal spread, support warnings, selected normalized actions, raw environment actions, dashed imagined latents, solid returned observations, termination, and timing. A fixed projection may visualize the two paths, but native latent distances remain the measurement; simulator coordinates belong in a separate privileged panel.
The source boundary is LeWorldModel v3, frozen jepa.py and eval.py, commit 8edfeb3. The core did not pin its historical solver dependency, although both inspected solver revisions return the final mean. Oracles and the failure dossier are tutorial diagnostics, not baseline features or deployable pixels-only results.
Quick check
- Why must the goal latent stay outside baseline dynamics rollout?
- Which assertion can reveal a vectorization identity mix-up?
- What different bottlenecks do oracle dynamics and oracle cost test?