Course progress Course outline 48 of 48 lessons available
Part 0 — Reading Guide: What Exactly Are We Going to Learn?
Part 1 — World Models: An Internal Sandbox for the Agent
Part 2 — Turning Images into State: The LeWM Architecture
- 06 Chapter 5 — Trajectory Data: To the Model, the World Is Not an Image Collection available now
- 07 Chapter 6 — The Visual Encoder: Issuing a “State Passport” for Every Frame available now
- 08 Chapter 7 — The Dynamics Predictor: Moving Time Forward in the Mind available now
- 09 Chapter 8 — A Complete Forward Pass: Follow One Batch from Start to Finish available now
Part 3 — Preventing the Model from Cheating: Prediction Loss and SIGReg
- 10 Chapter 9 — The Most Dangerous Shortcut: Representation Collapse available now
- 11 Chapter 10 — Prediction Loss: How the Model Learns the Next Step available now
- 12 Chapter 11 — The Intuition Behind SIGReg: Letting Representation Space “Breathe” available now
- 13 Chapter 12 — Keep the Mathematics Minimal but Sufficient Current lesson
- 14 Chapter 13 — The Original LeWM’s End-to-End Training Mechanism available now
- 15 Chapter 14 — Train a Model That Does Not Collapse Immediately available now
Part 4 — Putting the Model into Action: Planning in Latent Space
- 16 Chapter 15 — Goal-Conditioned Planning: From “Where Am I?” to “Where Do I Want to Go?” available now
- 17 Chapter 16 — Latent Euclidean Distance: Convenient, but Not Necessarily Reliable available now
- 18 Chapter 17 — CEM: Searching for Actions Through an Elimination Tournament available now
- 19 Chapter 18 — MPC: Do Not Trust the Model for Too Long at Once available now
- 20 Chapter 19 — Long-Horizon Rollouts: How Small Errors Snowball into Major Failures available now
- 21 Chapter 20 — Implement a Minimal LeWM Planner from Scratch available now
Part 5 — Engineering Reproduction: From Paper to Running System
- 22 Chapter 21 — The Official Repository and Experimental Environment available now
- 23 Chapter 22 — First Experiment: A TwoRoom Smoke Test available now
- 24 Chapter 23 — Second Experiment: Reproducing PushT available now
- 25 Chapter 24 — How to Evaluate a World Model Fairly available now
- 26 Chapter 25 — Failure-Diagnosis Manual available now
Part 6 — What Has LeWM Actually Learned?
- 27 Chapter 26 — Linear Probes: Which Physical Variables Are Encoded in the Latent State? available now
- 28 Chapter 27 — Give Latent Space a “Health Check” available now
- 29 Chapter 28 — Violation of Expectation: Is the Model Surprised by “Impossible Events”? available now
- 30 Chapter 29 — How to Discuss “Understanding the World” Rigorously available now
Part 7 — Why “Accurate Prediction” Can Still Produce “Poor Planning”
- 31 Chapter 30 — The Gap Between the Training Objective and the Planning Objective available now
- 32 Chapter 31 — Global Non-Collapse Does Not Guarantee Preservation of Task-Relevant Dynamics available now
- 33 Chapter 32 — When Is an Isotropic Gaussian Prior Too Strong? available now
- 34 Chapter 33 — Long-Horizon Planning: Predict Farther or Plan More Intelligently? available now
- 35 Chapter 34 — From Positional Distance to Task Progress available now
- 36 Chapter 35 — Multi-Task Learning, Real Robots, and Visual Distractions available now
- 37 Chapter 36 — Theoretical Boundaries: When Can the True State Be Identified? available now
Part 8 — From Reproducer to Researcher
Appendices
- 41 Appendix A — The Minimum Necessary Mathematical Toolkit available now
- 42 Appendix B — PyTorch Implementation Quick Reference available now
- 43 Appendix C — Complete Tensor-Shape Table available now
- 44 Appendix D — Experiment Configuration Cards available now
- 45 Appendix E — Paper Timeline and Evidence Levels available now
- 46 Appendix F — Glossary available now
- 47 Appendix G — Reproduction Checklist available now
- 48 Appendix H — Expert-Review Checklist available now
The big picture
- Dynamics gaugedid next-latent agree?
- Population gaugedid the cloud collapse?
- lambdabalance the two pressures

A tiny story
One workshop inspector checks whether the action-conditioned guess matches the latent of what really happened next. A second inspector checks whether the crowd of observation embeddings has shrunk into an easy, degenerate shape.
Listen only to the first and every observation can become one constant answer. Listen only to the second and the cloud can look beautifully spread while forgetting which action reaches the doorway. The two inspectors do not grade the same thing.
The technical backpack
The prediction term is coordinate-wise mean squared error against a connected next-observation embedding:
z = shared_encoder(observation_window)
z_hat = action_conditioned_predictor(z_context, recorded_actions)
L_pred = mean((z_hat - z_next_connected)^2)
Because the target is produced by the shared trainable encoder and is not detached, L_pred can update the predictor, action path, context representation, and target representation.
SIGReg acts on observation embeddings at each relative time position. Random unit directions create one-dimensional shadows; finite cosine/sine fingerprints are compared with the analytic standard-Gaussian fingerprint. It directly updates the visual encoder and encoder projector, not the action encoder or dynamics predictor.
The complete objective is:
L_total = L_pred + lambda * L_SIGReg
lambda is a volume knob for the population pressure. It does not alter what SIGReg measures, choose latent width, set projection count, tune the optimizer, or control planning horizon. Too quiet leaves the trivial shortcut available; too loud can favor global Gaussian shape over transition relationships.
Version labels are part of the mathematics:
| Source | Value | Safe reading |
|---|---|---|
| LeWM v3 method | 0.1 | paper’s stated default |
| frozen YAML | 0.09 | checked-in executable setting |
| PushT Appendix G sweep | peak near 0.09; >80% over tested 0.01–0.2; sharp drop at 0.5 | author-reported under that setup |
The paper describes projection count and the regularization weight as two introduced settings and calls lambda the only effective one to tune after its reported projection ablation. That does not mean the model, optimizer, data, and planner have only one hyperparameter.
Frozen SIGReg uses 1,024 projections and 17 knots from 0 through 3. Batch size controls samples in each empirical shadow; projection count controls how many directions inspect those same samples. One cannot replace the other.
The trick that fools us
Train three matched conditions: no SIGReg, a clearly labelled paper or code weight, and a deliberately overwhelming weight. Record raw L_pred and L_SIGReg separately—not only their weighted sum—plus variance, effective rank, neighbors, action swaps, one-step error, and autoregressive error by horizon.
Expected mechanisms are conditional. Removing SIGReg makes constant agreement compatible with the objective, but one run need not collapse. Overwhelming SIGReg can hurt dynamics, but no universal threshold follows. A spread cloud with action-insensitive predictions is not globally collapsed; it is dynamically weak. Good training gauges with bad planning point to a later link such as rollout, cost, or search.
Evidence receipt
LeWorldModel v3, frozen train.py, frozen SIGReg, and frozen lewm.yaml support the two-term objective and source-specific values. LeJEPA v3 supplies assumption-bounded Gaussian motivation, not a universal physical-state guarantee.
Allowed claim: ideal Gaussian matching excludes the exact constant representation, while prediction supplies temporal/action relationships. Low total loss proves neither physics nor planning.
Quick check
- Which shortcut does each gauge oppose or constrain?
- Why can increasing
lambdaimprove population shape while harming dynamics? - Why are
0.1and0.09both correct but not interchangeable facts?