Course progress Course outline 48 of 48 lessons available
Part 0 — Reading Guide: What Exactly Are We Going to Learn?
Part 1 — World Models: An Internal Sandbox for the Agent
Part 2 — Turning Images into State: The LeWM Architecture
- 06 Chapter 5 — Trajectory Data: To the Model, the World Is Not an Image Collection available now
- 07 Chapter 6 — The Visual Encoder: Issuing a “State Passport” for Every Frame available now
- 08 Chapter 7 — The Dynamics Predictor: Moving Time Forward in the Mind available now
- 09 Chapter 8 — A Complete Forward Pass: Follow One Batch from Start to Finish available now
Part 3 — Preventing the Model from Cheating: Prediction Loss and SIGReg
- 10 Chapter 9 — The Most Dangerous Shortcut: Representation Collapse available now
- 11 Chapter 10 — Prediction Loss: How the Model Learns the Next Step available now
- 12 Chapter 11 — The Intuition Behind SIGReg: Letting Representation Space “Breathe” Current lesson
- 13 Chapter 12 — Keep the Mathematics Minimal but Sufficient available now
- 14 Chapter 13 — The Original LeWM’s End-to-End Training Mechanism available now
- 15 Chapter 14 — Train a Model That Does Not Collapse Immediately available now
Part 4 — Putting the Model into Action: Planning in Latent Space
- 16 Chapter 15 — Goal-Conditioned Planning: From “Where Am I?” to “Where Do I Want to Go?” available now
- 17 Chapter 16 — Latent Euclidean Distance: Convenient, but Not Necessarily Reliable available now
- 18 Chapter 17 — CEM: Searching for Actions Through an Elimination Tournament available now
- 19 Chapter 18 — MPC: Do Not Trust the Model for Too Long at Once available now
- 20 Chapter 19 — Long-Horizon Rollouts: How Small Errors Snowball into Major Failures available now
- 21 Chapter 20 — Implement a Minimal LeWM Planner from Scratch available now
Part 5 — Engineering Reproduction: From Paper to Running System
- 22 Chapter 21 — The Official Repository and Experimental Environment available now
- 23 Chapter 22 — First Experiment: A TwoRoom Smoke Test available now
- 24 Chapter 23 — Second Experiment: Reproducing PushT available now
- 25 Chapter 24 — How to Evaluate a World Model Fairly available now
- 26 Chapter 25 — Failure-Diagnosis Manual available now
Part 6 — What Has LeWM Actually Learned?
- 27 Chapter 26 — Linear Probes: Which Physical Variables Are Encoded in the Latent State? available now
- 28 Chapter 27 — Give Latent Space a “Health Check” available now
- 29 Chapter 28 — Violation of Expectation: Is the Model Surprised by “Impossible Events”? available now
- 30 Chapter 29 — How to Discuss “Understanding the World” Rigorously available now
Part 7 — Why “Accurate Prediction” Can Still Produce “Poor Planning”
- 31 Chapter 30 — The Gap Between the Training Objective and the Planning Objective available now
- 32 Chapter 31 — Global Non-Collapse Does Not Guarantee Preservation of Task-Relevant Dynamics available now
- 33 Chapter 32 — When Is an Isotropic Gaussian Prior Too Strong? available now
- 34 Chapter 33 — Long-Horizon Planning: Predict Farther or Plan More Intelligently? available now
- 35 Chapter 34 — From Positional Distance to Task Progress available now
- 36 Chapter 35 — Multi-Task Learning, Real Robots, and Visual Distractions available now
- 37 Chapter 36 — Theoretical Boundaries: When Can the True State Be Identified? available now
Part 8 — From Reproducer to Researcher
Appendices
- 41 Appendix A — The Minimum Necessary Mathematical Toolkit available now
- 42 Appendix B — PyTorch Implementation Quick Reference available now
- 43 Appendix C — Complete Tensor-Shape Table available now
- 44 Appendix D — Experiment Configuration Cards available now
- 45 Appendix E — Paper Timeline and Evidence Levels available now
- 46 Appendix F — Glossary available now
- 47 Appendix G — Reproduction Checklist available now
- 48 Appendix H — Expert-Review Checklist available now
The big picture
- Latent crowdone time, many examples
- Cast shadowsrandom 1D directions
- Fingerprintcosine + sine probes
- Comparestandard Gaussian

A tiny story
“Do not put everyone at one address” is too vague. A city can avoid one point yet squeeze every resident onto one road, stretch into one dominant direction, or make an X-shaped pattern whose simple statistics look healthy.
SIGReg—Sketched Isotropic Gaussian Regularization—gives the crowd a reference shape. It asks many one-dimensional shadows of the embedding population to resemble shadows from an isotropic standard Gaussian.
The technical backpack
At each relative time position, frozen LeWM treats the B embeddings as one population. Time is not mixed into the crowd. A random unit vector turns every D=192 embedding into one signed scalar shadow. Repeat across directions and compare each one-dimensional sample with the analytic Gaussian fingerprint.
The Gaussian words matter:
- standard: each unit-direction shadow is centered at zero with unit scale;
- isotropic: the reference has no preferred direction;
- Gaussian: it is not a uniform ball. Density peaks at the origin, while high-dimensional radial probability mass concentrates roughly near radius
sqrt(D).
The Epps–Pulley construction uses an empirical characteristic function. At finite frequency probes, it averages cosine and sine responses and compares them with the known standard-Gaussian responses. LeWM minimizes a smooth discrepancy; it does not compute a p-value or declare that a batch “passed normality.”
[T,B,D]
-> for each time: B-point cloud
-> random 1D projections
-> cosine/sine responses at finite knots
-> gap to analytic N(0,1) fingerprint
-> average across directions and time
Frozen code samples 1,024 directions per call and uses 17 knots from 0 through 3 with a Gaussian weighting window. Directions are resampled, so the training value is stochastic. The class is marked single-GPU and does not gather a global population across devices. The paper appendix presents an illustrative integration differently; keep paper and executable facts separate.
The Cramér–Wold principle identifies multivariate distributions if all one-dimensional projections match. Frozen training sees finite batches, finite directions, and finite knots. It is an approximation, not a Gaussianity certificate.
The trick that fools us
Inspect only horizontal and vertical axes of an X-shaped cloud. Both shadows can look harmless while a diagonal shadow reveals dependence. Random views reduce permanent blind spots, but 1,024 finite views can still miss a narrow defect.
Now inspect a real corridor trajectory. A thin cloud may reflect a genuinely low-dimensional world rather than a broken encoder. Forcing it into a 192-dimensional isotropic target can create tension. The LeWM paper suggests this mismatch as a possible contributor in TwoRoom; it is an author hypothesis, not an established cause.
More projections do not create more observations. Looking at a small crowd from a thousand angles does not make it a large crowd. Batch size and projection count address different uncertainty.
Evidence receipt
The mechanism comes from LeWorldModel v3, frozen SIGReg, LeJEPA v3, the Epps–Pulley paper, and Cramér–Wold.
SIGReg needs batch examples but no labelled negatives and no learned teacher because its reference is analytic. Matching that reference excludes the exact constant representation at the ideal target. It does not prove physics, action sensitivity, reachability, or finite-run success.
Quick check
- Why are nonzero variance and identity covariance still incomplete checks?
- Why do finite random projections not turn Cramér–Wold into a proof?
- What does SIGReg use instead of negative pairs or an EMA teacher?