Course progress Course outline 48 of 48 lessons available
Part 0 — Reading Guide: What Exactly Are We Going to Learn?
Part 1 — World Models: An Internal Sandbox for the Agent
Part 2 — Turning Images into State: The LeWM Architecture
- 06 Chapter 5 — Trajectory Data: To the Model, the World Is Not an Image Collection available now
- 07 Chapter 6 — The Visual Encoder: Issuing a “State Passport” for Every Frame available now
- 08 Chapter 7 — The Dynamics Predictor: Moving Time Forward in the Mind available now
- 09 Chapter 8 — A Complete Forward Pass: Follow One Batch from Start to Finish available now
Part 3 — Preventing the Model from Cheating: Prediction Loss and SIGReg
- 10 Chapter 9 — The Most Dangerous Shortcut: Representation Collapse available now
- 11 Chapter 10 — Prediction Loss: How the Model Learns the Next Step available now
- 12 Chapter 11 — The Intuition Behind SIGReg: Letting Representation Space “Breathe” available now
- 13 Chapter 12 — Keep the Mathematics Minimal but Sufficient available now
- 14 Chapter 13 — The Original LeWM’s End-to-End Training Mechanism available now
- 15 Chapter 14 — Train a Model That Does Not Collapse Immediately available now
Part 4 — Putting the Model into Action: Planning in Latent Space
- 16 Chapter 15 — Goal-Conditioned Planning: From “Where Am I?” to “Where Do I Want to Go?” available now
- 17 Chapter 16 — Latent Euclidean Distance: Convenient, but Not Necessarily Reliable available now
- 18 Chapter 17 — CEM: Searching for Actions Through an Elimination Tournament available now
- 19 Chapter 18 — MPC: Do Not Trust the Model for Too Long at Once available now
- 20 Chapter 19 — Long-Horizon Rollouts: How Small Errors Snowball into Major Failures available now
- 21 Chapter 20 — Implement a Minimal LeWM Planner from Scratch available now
Part 5 — Engineering Reproduction: From Paper to Running System
- 22 Chapter 21 — The Official Repository and Experimental Environment available now
- 23 Chapter 22 — First Experiment: A TwoRoom Smoke Test available now
- 24 Chapter 23 — Second Experiment: Reproducing PushT available now
- 25 Chapter 24 — How to Evaluate a World Model Fairly available now
- 26 Chapter 25 — Failure-Diagnosis Manual available now
Part 6 — What Has LeWM Actually Learned?
- 27 Chapter 26 — Linear Probes: Which Physical Variables Are Encoded in the Latent State? available now
- 28 Chapter 27 — Give Latent Space a “Health Check” available now
- 29 Chapter 28 — Violation of Expectation: Is the Model Surprised by “Impossible Events”? available now
- 30 Chapter 29 — How to Discuss “Understanding the World” Rigorously available now
Part 7 — Why “Accurate Prediction” Can Still Produce “Poor Planning”
- 31 Chapter 30 — The Gap Between the Training Objective and the Planning Objective available now
- 32 Chapter 31 — Global Non-Collapse Does Not Guarantee Preservation of Task-Relevant Dynamics available now
- 33 Chapter 32 — When Is an Isotropic Gaussian Prior Too Strong? available now
- 34 Chapter 33 — Long-Horizon Planning: Predict Farther or Plan More Intelligently? available now
- 35 Chapter 34 — From Positional Distance to Task Progress available now
- 36 Chapter 35 — Multi-Task Learning, Real Robots, and Visual Distractions available now
- 37 Chapter 36 — Theoretical Boundaries: When Can the True State Be Identified? available now
Part 8 — From Reproducer to Researcher
Appendices
- 41 Appendix A — The Minimum Necessary Mathematical Toolkit available now
- 42 Appendix B — PyTorch Implementation Quick Reference available now
- 43 Appendix C — Complete Tensor-Shape Table available now
- 44 Appendix D — Experiment Configuration Cards available now
- 45 Appendix E — Paper Timeline and Evidence Levels available now
- 46 Appendix F — Glossary available now
- 47 Appendix G — Reproduction Checklist available now
- 48 Appendix H — Expert-Review Checklist available now
The big picture
- One questionName a failure and a prediction.
- One small testChange one mechanism in a sandbox.
- One way to loseWrite the falsifier before the score.
A tiny story: twelve boats
Twelve boats wait in a harbor. Some test planning geometry, some uncertainty, some representation structure, and some search or transfer. None is “the winner.” Each carries a small launch card: question, one changed part, measurements, and a condition that sends it back to port.

These are experiment plans, not new results. Equal berths do not claim equal difficulty, promise, or evidence.
The technical backpack: twelve launch cards
All cards begin with a frozen LeWM baseline, episode-disjoint data, declared seeds, matched training and planning budgets, raw logs, and positive/null/negative reporting. Simulator state, shortest path, rewards, depth, or language annotations are labeled privileged whenever they train or evaluate a component.
| # | Project | Change and measurement | A result that weakens the idea |
|---|---|---|---|
| 1 | Latent distance vs true shortest path | In TwoRoom, correlate latent pair distance with oracle geodesic distance; stratify same-room/across-wall and near/far pairs | correlation vanishes on held-out layouts, or better correlation does not improve candidate ranking |
| 2 | Directed reachability | Learn a budget-conditioned, ordered score from trajectory offsets and constructed negatives; test withheld directions and horizons | the score is symmetric, follows dataset frequency, or fails matched oracle-budget ranking |
| 3 | Predictive uncertainty calibration | Predict spread or error bins and compare confidence with later self-fed rollout error | confidence is sharp but uncalibrated, or only tracks horizon length |
| 4 | Ensemble epistemic uncertainty | Train matched-seed or bootstrap models; compare disagreement inside and outside support | members agree on wrong unsupported shortcuts, or disagreement merely reflects optimization noise |
| 5 | Multi-step training | Add one declared rollout-consistency or multi-horizon loss while keeping inference and budgets fixed | one-step loss improves but long rollout/planning worsens, or gain comes from extra updates |
| 6 | Temporally hierarchical latents | Learn coarse states or action chunks and audit reconstruction, reuse, support, and two-level cost | high-level error, unsupported subgoals, or overhead exceeds saved low-level computation |
| 7 | Local Gaussian structure, global topology | Regularize local neighborhoods or subspaces while measuring global spread and doorway continuity separately | anti-collapse weakens, or topology gains disappear without extra capacity |
| 8 | Action-controllable subspace | Split action-sensitive change from content; test interventions, nuisance changes, and weakly controlled variables | the “control” space encodes background/task ID or drops slow task-critical state |
| 9 | Counterfactual-dynamics probes | At matched states, compare predictions for recorded and withheld actions against simulator transitions | branch separation exists without correct consequences, or only recorded actions are accurate |
| 10 | Support-constrained CEM | Penalize support distance or search near empirical action/macro-action anchors; report constraint strength | a strict gate blocks valid novel detours, while a loose gate still admits model exploits |
| 11 | Goal images to language goals | Add a declared language-to-goal interface; hold out object–relation combinations and compare oracle goal encodings | success depends on memorized phrases or paired-goal leakage rather than composition |
| 12 | Cross-environment latent alignment | Align environments with paired anchors, then test held-out dynamics, views, and tasks | alignment improves a probe but harms transition prediction or planning in either world |
Projects 1–2 distinguish distance from directed, budgeted feasibility. Projects 3–4 distinguish predictive spread, ensemble disagreement, support distance, and actual error. Projects 5–6 distinguish longer training targets from hierarchy. Projects 7–9 ask what representation structure actions really need. Projects 10–12 move the pressure to search, goals, and transfer. Combining cards before each wins its own test destroys the diagnosis.
Try to break the idea
Support-constrained CEM is a useful self-attack. In one TwoRoom layout, an unconstrained optimizer takes an impossible shortcut. A moderate support rule guides it through the doorway. Then introduce a valid novel detour absent from the training trajectories. An overly strict rule blocks the only successful path.
Sweep the support penalty and show both failure families. Use an oracle feasible-route label only for evaluation. If the best threshold changes with goal type, report the interaction; do not hide it in one average. The project succeeds scientifically even if no universal threshold exists, because it has mapped the trade-off between model exploitation and over-conservative search.
Every card needs a removal test. Remove the wall for geometry, equalize action coverage for controllability, inject controlled out-of-support samples for uncertainty, shuffle privileged labels for grounding, remove paired anchors for alignment, and hold out language combinations for composition. A result that survives only the easy condition has not earned the broad claim.
Experiment receipt and evidence boundary
- These cards are tutorial-generated hypotheses, not reported improvements. Direct inspirations include LeWorldModel v3, RC-aux, TRM, Sub-JEPA, SMWM, VLWM, Temporal-Distance JEPA, Fast-LeWM, Hi-LeWM, and PSG-JEPA.
- Later diagnostics and decision routes include TwoRoom reproduction, VIScore, ACPC, Objective Bottleneck, Traj-LeWM, SCALE, AC-MTM, and DA-LeWM. Their evaluations remain author-reported except the explicitly bounded TwoRoom reimplementation.
- Public frozen implementations were identified for LeWM, RC-aux, Temporal-Distance JEPA, Sub-JEPA, Fast-LeWM, Hi-LeWM, stable-worldmodel, passive-identifiability, SMWM, PSG-JEPA, tinylab, VIScore, ACPC, Traj-LeWM, and AC-MTM. None was identified at cutoff for TRM, VLWM, SCALE, DA-LeWM, PhyLatent, TC-LeWM, QQWorld, ProWorld, controlled-identifiability, or the 2024 “HWM” course-shorthand route. Public code is not independent reproduction.
- Source versions span 2024-06-01 through DA-LeWM v1 on 2026-08-19; evidence cutoff is 2026-08-20, Asia/Tokyo. Permitted verbs are “test,” “hypothesis,” “falsifier,” and “supports under this locked setting”—not “will improve,” “solves uncertainty,” “pixels-only” when privileged labels train the model, or “proves compositional understanding.”
Three quick questions
- What five fields turn a research idea into an executable launch card?
- Why are ensemble disagreement and actual rollout error different instruments?
- How can a support constraint both help and hurt the same planner?