Course progress Course outline 48 of 48 lessons available
Part 0 — Reading Guide: What Exactly Are We Going to Learn?
Part 1 — World Models: An Internal Sandbox for the Agent
Part 2 — Turning Images into State: The LeWM Architecture
- 06 Chapter 5 — Trajectory Data: To the Model, the World Is Not an Image Collection available now
- 07 Chapter 6 — The Visual Encoder: Issuing a “State Passport” for Every Frame available now
- 08 Chapter 7 — The Dynamics Predictor: Moving Time Forward in the Mind available now
- 09 Chapter 8 — A Complete Forward Pass: Follow One Batch from Start to Finish available now
Part 3 — Preventing the Model from Cheating: Prediction Loss and SIGReg
- 10 Chapter 9 — The Most Dangerous Shortcut: Representation Collapse available now
- 11 Chapter 10 — Prediction Loss: How the Model Learns the Next Step available now
- 12 Chapter 11 — The Intuition Behind SIGReg: Letting Representation Space “Breathe” available now
- 13 Chapter 12 — Keep the Mathematics Minimal but Sufficient available now
- 14 Chapter 13 — The Original LeWM’s End-to-End Training Mechanism available now
- 15 Chapter 14 — Train a Model That Does Not Collapse Immediately available now
Part 4 — Putting the Model into Action: Planning in Latent Space
- 16 Chapter 15 — Goal-Conditioned Planning: From “Where Am I?” to “Where Do I Want to Go?” available now
- 17 Chapter 16 — Latent Euclidean Distance: Convenient, but Not Necessarily Reliable available now
- 18 Chapter 17 — CEM: Searching for Actions Through an Elimination Tournament available now
- 19 Chapter 18 — MPC: Do Not Trust the Model for Too Long at Once available now
- 20 Chapter 19 — Long-Horizon Rollouts: How Small Errors Snowball into Major Failures available now
- 21 Chapter 20 — Implement a Minimal LeWM Planner from Scratch available now
Part 5 — Engineering Reproduction: From Paper to Running System
- 22 Chapter 21 — The Official Repository and Experimental Environment available now
- 23 Chapter 22 — First Experiment: A TwoRoom Smoke Test available now
- 24 Chapter 23 — Second Experiment: Reproducing PushT available now
- 25 Chapter 24 — How to Evaluate a World Model Fairly available now
- 26 Chapter 25 — Failure-Diagnosis Manual available now
Part 6 — What Has LeWM Actually Learned?
- 27 Chapter 26 — Linear Probes: Which Physical Variables Are Encoded in the Latent State? available now
- 28 Chapter 27 — Give Latent Space a “Health Check” available now
- 29 Chapter 28 — Violation of Expectation: Is the Model Surprised by “Impossible Events”? available now
- 30 Chapter 29 — How to Discuss “Understanding the World” Rigorously available now
Part 7 — Why “Accurate Prediction” Can Still Produce “Poor Planning”
- 31 Chapter 30 — The Gap Between the Training Objective and the Planning Objective available now
- 32 Chapter 31 — Global Non-Collapse Does Not Guarantee Preservation of Task-Relevant Dynamics available now
- 33 Chapter 32 — When Is an Isotropic Gaussian Prior Too Strong? available now
- 34 Chapter 33 — Long-Horizon Planning: Predict Farther or Plan More Intelligently? available now
- 35 Chapter 34 — From Positional Distance to Task Progress available now
- 36 Chapter 35 — Multi-Task Learning, Real Robots, and Visual Distractions available now
- 37 Chapter 36 — Theoretical Boundaries: When Can the True State Be Identified? available now
Part 8 — From Reproducer to Researcher
Appendices
- 41 Appendix A — The Minimum Necessary Mathematical Toolkit available now
- 42 Appendix B — PyTorch Implementation Quick Reference available now
- 43 Appendix C — Complete Tensor-Shape Table available now
- 44 Appendix D — Experiment Configuration Cards available now
- 45 Appendix E — Paper Timeline and Evidence Levels available now
- 46 Appendix F — Glossary available now
- 47 Appendix G — Reproduction Checklist available now
- 48 Appendix H — Expert-Review Checklist available now
The big picture
- Known coastWe can name the working pipeline.
- Dotted coastSeveral explanations still fit.
- Next voyageChoose a test that separates them.
A tiny story: the unfinished atlas
An atlas has solid coastlines where ships have measured the shore and dotted lines where several shapes remain possible. A careful cartographer does not fill dotted regions with the newest traveler’s favorite drawing.
The course’s solid spine is modest: observe pixels, compress them into a useful latent state, imagine action-conditioned futures, search and act with CEM/MPC, then diagnose where imagination is unreliable. The frontier asks which geometry, target, uncertainty, hierarchy, support rule, or successor paradigm best serves that loop. Newer papers supply candidates and author-reported evidence; they do not turn the dotted coast into consensus.

The map organizes questions. It does not rank papers or predict which route will win.
The technical backpack: ten open questions
| # | Open question | Competing hypotheses | A useful discriminator |
|---|---|---|---|
| 1 | Which geometry is suitable for planning? | Euclidean proximity, directed reachability, task progress, trajectory cost, or uncertainty-aware decision ranking may dominate in different tasks | matched open-room, wall, contact, and budget sweeps with oracle rankings |
| 2 | Does non-collapse require global distribution matching? | full Gaussian pressure, subspace/local matching, temporal centering, or transition-derived action objectives may suffice | equal-data anti-collapse tests plus topology, slow-variable, tail, and planning checks |
| 3 | How should several possible futures be represented? | a conditional mean may be enough for control, or explicit branches/distributions may be necessary | calibrated multimodal environments where the mean is physically unreal |
| 4 | How do we know imagination is out of support? | ensemble disagreement, predictive spread, nearest-data distance, action-conditional consistency, or realized error may be best | controlled support removal and later ground-truth rollout errors |
| 5 | Can controllability be learned from offline behavior? | inverse actions and action-sensitive branches may identify useful control, or missing state-conditioned interventions may leave ambiguity | hold state fixed and reveal/withhold action directions |
| 6 | Do long goals require explicit hierarchy? | prefix prediction, flat MPC with a better cost, learned chunks, or latent subgoals may each be enough | matched model-call, action, latency, and support budgets across horizon |
| 7 | How large should latent state be? | larger width preserves rare distinctions, or wastes capacity and fights low intrinsic dimension | width/rank sweeps with fixed compute, collapse, rollout, and planning diagnostics |
| 8 | Do task-agnostic and task-effective representations conflict? | invariance may aid transfer, while task context or privileged grounding may be essential | factor-isolated nuisance/task interventions and held-out task compositions |
| 9 | How can a model move from predicting to explaining? | probes and counterfactual predictions may reveal mechanisms, or only interventions with causal assumptions can support explanation | preregistered interventions that separate correlation, use, and causal effect |
| 10 | What might replace or absorb the LeWM line? | branching prediction, unified JEPA objectives, amortized/search-free control, value-shaped geometry, or other world-model paradigms | one shared load test: prediction, counterfactuals, support, planning, transfer, compute, and safety |
The questions interact but must not collapse into one scoreboard. A model can predict well and rank plans badly. It can be useful under one planner without being identifiable. It can have public code without independent reproduction. It can be faster on one GPU budget while weaker on support-shifted goals.
Try to break the idea
Create two systems, Cedar and Comet. They have the same average one-step loss and the same average success. Cedar succeeds on nearby goals but fails every across-wall goal. Comet succeeds on across-wall goals but fails after visual changes. One average makes them look identical; their worlds are not.
Now disclose goal type, horizon, support, visual intervention, planner budget, and failure trajectory. The next experiment differs: Cedar needs a geometry/cost test; Comet needs a nuisance/representation test. This is why an open question needs a discriminator, not another aggregate leaderboard.
For any proposed successor, freeze a load-test ledger before admiring its name: training information, target graph, future representation, action interface, model calls, search budget, hardware, data support, transfer split, and safety mechanism. A new paradigm can win one column and lose another. “Will replace LeWM” remains speculation until the declared evidence exists.
Experiment receipt and evidence boundary
- Baseline coordinates: LeWorldModel v3 and LeJEPA v3. Direct updates include TRM, SMWM, VLWM, PSG-JEPA, TwoRoom reproduction, VIScore, ACPC, Objective Bottleneck, Traj-LeWM, SCALE, AC-MTM, and DA-LeWM.
- Neighboring routes include Branch-JEPA v3, UniJEPA, Qantara, and INTACT. They are not LeWM versions, and adjacency does not establish replacement.
- LeJEPA v3 was revised 2025-11-14; LeWM v3 on 2026-06-03; the latest named direct update is DA-LeWM v1 on 2026-08-19. Evidence cutoff: 2026-08-20, Asia/Tokyo. Baseline code is frozen at
8edfeb3; other inspected revisions stay frozen in Appendix E rather than moving withmain. - Author-linked code was verified for the baseline and, among the latest items, SMWM, PSG-JEPA, tinylab, VIScore, ACPC, Traj-LeWM, AC-MTM, and INTACT; none was identified for TRM, VLWM, SCALE, or DA-LeWM. Independent reproduction is established only for the stated TwoRoom environment and protocols. Every other result keeps its author’s task, supervision, planner, compute, and assumptions.
Three quick questions
- What turns a broad unknown into a scientific open question?
- Why can equal average loss and success hide different world-model failures?
- Which evidence fields must a proposed successor share before “better” is meaningful?