Course progress Course outline 48 of 48 lessons available
Part 0 — Reading Guide: What Exactly Are We Going to Learn?
Part 1 — World Models: An Internal Sandbox for the Agent
Part 2 — Turning Images into State: The LeWM Architecture
- 06 Chapter 5 — Trajectory Data: To the Model, the World Is Not an Image Collection available now
- 07 Chapter 6 — The Visual Encoder: Issuing a “State Passport” for Every Frame available now
- 08 Chapter 7 — The Dynamics Predictor: Moving Time Forward in the Mind available now
- 09 Chapter 8 — A Complete Forward Pass: Follow One Batch from Start to Finish available now
Part 3 — Preventing the Model from Cheating: Prediction Loss and SIGReg
- 10 Chapter 9 — The Most Dangerous Shortcut: Representation Collapse available now
- 11 Chapter 10 — Prediction Loss: How the Model Learns the Next Step available now
- 12 Chapter 11 — The Intuition Behind SIGReg: Letting Representation Space “Breathe” available now
- 13 Chapter 12 — Keep the Mathematics Minimal but Sufficient available now
- 14 Chapter 13 — The Original LeWM’s End-to-End Training Mechanism available now
- 15 Chapter 14 — Train a Model That Does Not Collapse Immediately available now
Part 4 — Putting the Model into Action: Planning in Latent Space
- 16 Chapter 15 — Goal-Conditioned Planning: From “Where Am I?” to “Where Do I Want to Go?” available now
- 17 Chapter 16 — Latent Euclidean Distance: Convenient, but Not Necessarily Reliable available now
- 18 Chapter 17 — CEM: Searching for Actions Through an Elimination Tournament available now
- 19 Chapter 18 — MPC: Do Not Trust the Model for Too Long at Once available now
- 20 Chapter 19 — Long-Horizon Rollouts: How Small Errors Snowball into Major Failures available now
- 21 Chapter 20 — Implement a Minimal LeWM Planner from Scratch available now
Part 5 — Engineering Reproduction: From Paper to Running System
- 22 Chapter 21 — The Official Repository and Experimental Environment available now
- 23 Chapter 22 — First Experiment: A TwoRoom Smoke Test available now
- 24 Chapter 23 — Second Experiment: Reproducing PushT available now
- 25 Chapter 24 — How to Evaluate a World Model Fairly available now
- 26 Chapter 25 — Failure-Diagnosis Manual available now
Part 6 — What Has LeWM Actually Learned?
- 27 Chapter 26 — Linear Probes: Which Physical Variables Are Encoded in the Latent State? available now
- 28 Chapter 27 — Give Latent Space a “Health Check” available now
- 29 Chapter 28 — Violation of Expectation: Is the Model Surprised by “Impossible Events”? available now
- 30 Chapter 29 — How to Discuss “Understanding the World” Rigorously available now
Part 7 — Why “Accurate Prediction” Can Still Produce “Poor Planning”
- 31 Chapter 30 — The Gap Between the Training Objective and the Planning Objective available now
- 32 Chapter 31 — Global Non-Collapse Does Not Guarantee Preservation of Task-Relevant Dynamics available now
- 33 Chapter 32 — When Is an Isotropic Gaussian Prior Too Strong? available now
- 34 Chapter 33 — Long-Horizon Planning: Predict Farther or Plan More Intelligently? available now
- 35 Chapter 34 — From Positional Distance to Task Progress Current lesson
- 36 Chapter 35 — Multi-Task Learning, Real Robots, and Visual Distractions available now
- 37 Chapter 36 — Theoretical Boundaries: When Can the True State Be Identified? available now
Part 8 — From Reproducer to Researcher
Appendices
- 41 Appendix A — The Minimum Necessary Mathematical Toolkit available now
- 42 Appendix B — PyTorch Implementation Quick Reference available now
- 43 Appendix C — Complete Tensor-Shape Table available now
- 44 Appendix D — Experiment Configuration Cards available now
- 45 Appendix E — Paper Timeline and Evidence Levels available now
- 46 Appendix F — Glossary available now
- 47 Appendix G — Reproduction Checklist available now
- 48 Appendix H — Expert-Review Checklist available now
The big picture
- Looks closeA ruler likes the wall.
- Make progressA good route may first move away.
- Can reachRules and budget still matter.
A tiny story: the wrong airport terminal
A traveler walks straight toward a gate painted on the other side of glass. The ruler says the traveler is closer. The trip says the traveler is stuck. The useful move is to walk away, turn around the partition, pass security, and approach from the correct side.
TwoRoom has the same lesson: crossing a doorway can require increasing straight-line distance. PushT adds stages—make contact, rotate, reposition, then push—so two states with equal goal distance may have opposite futures. “Closer in latent space” is not automatically “more complete.”

The detour separates straight-line proximity from route completion. It is a teaching case, not proof that every task has one monotone progress scale.
The technical backpack
Keep these labels apart:
- Distance compares two points with a chosen ruler. Baseline LeWM uses terminal squared Euclidean distance between predicted endpoint and encoded goal.
- Progress orders states toward task completion. It may be non-monotone in raw distance and may depend on task stage.
- Reachability asks whether allowed actions and dynamics can reach a state, often within a directed time or action budget.
ProWorld v1 is a later proposal, not part of original LeWM v3. It builds weak goal-conditioned order from hindsight temporal pairs, uses a Lorentz-model hyperbolic latent space, and scores intermediate as well as terminal progress. Hyperbolic space offers growing room for a branching hierarchy, but geometry is an inductive bias—not proof that the world truly has that hierarchy. Temporal order can also lie about progress when trajectories backtrack, detour, switch subgoals, or contain idle segments.
Other later methods ask neighboring but different questions: RC-aux learns budget-conditioned directed reachability; Temporal-Distance JEPA learns a directed temporal metric with rollout consistency and heuristic cross-trajectory negatives; TRM is a post-hoc horizon-matched reachability metric; Traj-LeWM learns a goal-conditioned trajectory cost used beside endpoint scoring; SCALE aligns latent pair distances with privileged state distance; DA-LeWM studies agreement between predicted and realized decision rankings. Do not merge their supervision, planner role, or evidence.
Try to break the idea
Create two PushT states at equal terminal distance. In one, the pusher has stable contact at a useful angle; in the other, the block is wedged. Then create a TwoRoom pair where the state that is farther by the ruler is already aligned with the doorway.
Ask every candidate score three questions: Which state ranks nearer? Which is reachable within the same action budget? Which is farther along a declared task phase? A method fails the lesson if it reports one scalar and silently answers all three.
For ProWorld-like supervision, add trajectories with deliberate detours and backtracking. If hindsight time labels the wrong state as “more progressed,” the coarse temporal proxy has met its falsifier. A task-specific phase label may repair this benchmark, but then report the extra prior; do not rename it general world understanding.
Experiment receipt and evidence boundary
- Baseline and proposal: LeWorldModel v3, frozen code
8edfeb3; ProWorld v1, released 2026-08-03. No ProWorld implementation was identified by the cutoff. - Related routes: RC-aux v1, code
ecb4496; Temporal-Distance JEPA v2, codeb4c17ca; TRM v1; Traj-LeWM v1, code67577fa; SCALE v1; DA-LeWM v1. - Dates: RC-aux 2026-05-08; TRM 2026-05-21; Temporal-Distance v2 2026-07-29; ProWorld 2026-08-03; Traj-LeWM 2026-08-14; SCALE 2026-08-17; DA-LeWM 2026-08-19. Evidence cutoff: 2026-08-20, Asia/Tokyo.
- All these later results are author-reported; independent reproduction is not established here. Permitted wording is “proposes,” “weak temporal order,” and “under its tested setup”—never “true progress,” “guaranteed reachability,” or “hyperbolic proof of hierarchy.”
Three quick questions
- How can moving farther from a goal be real progress?
- What information does a reachability claim need that a distance does not?
- Which detour experiment would expose a misleading temporal progress label?