Course
LeWorldModel: From Pixels to Imagination
A picture-first course about giving a robot a small rehearsal room, using it to choose actions, and catching it when its imagined world is wrong.
- modules
- 10 modules
- lessons
- 48 lessons
- estimated
- 24h estimated
- For
- For: Curious readers who want the simple story first, with Python, mathematics, and evidence receipts kept close by.
The whole course in one picture
- See camera + action
- Pack make a small state
- Imagine try actions inside
- Act pick, move, look again
- Check find the first lie
Imagine a tiny robot facing a wall. It sees a goal on the other side. Running straight into the wall is easy. Choosing the doorway requires a little make-believe world where the robot can try actions before using them for real.
That is the course spine:
observe → compress → imagine with actions → search → move a little → observe again → diagnose
TwoRoom makes mistakes easy to see. PushT adds contact and rotation, so the simple maze story cannot pretend to explain every control problem.
Three truth cards
- The original model: LeWM v3 trains one shared encoder and predictor end to end with prediction loss and SIGReg. It has no stop-gradient target, EMA teacher, or pretrained visual encoder.
- The planner: after training, the model is frozen. CEM tries action sequences; MPC executes a declared prefix and then looks again.
- The guardrail: non-collapse is not physics, latent closeness is not reachability, and one successful episode is not proof of understanding.
How to use this map
Read the big pictures first. Open the short technical backpack when you need an equation, tensor shape, configuration, or source. Every later method stays in the research frontier; it never travels backward and silently changes original LeWM v3.
Course progress
Course outline
Part 0 — Reading Guide: What Exactly Are We Going to Learn?
Establish the model family, evidence boundaries, reading paths, teaching environments, and complete system readers will build.
Part 1 — World Models: An Internal Sandbox for the Agent
Build the core intuition for predicting decision-relevant futures in latent space instead of reproducing every pixel.
- 02 Chapter 1 — Why an Agent Needs to “Imagine the Future” Connect future prediction to action, contrast reactive policies with world models, and locate failures in perception, imagination, and action selection. available now
- 03 Chapter 2 — Why Not Predict the Next Image Directly? See why pixel error can miss task-relevant structure and why useful latent states must balance compression with predictability. available now
- 04 Chapter 3 — The JEPA Idea: Predict Meaning, Not a Replica of the Image Understand encoders, target representations, and predictors, then compare JEPA with autoencoders, contrastive learning, and generative video models. available now
- 05 Chapter 4 — Understand LeWM in One Diagram Trace image histories and actions through the encoder, dynamics predictor, SIGReg, CEM planner, and MPC loop before introducing formulas. available now
Part 2 — Turning Images into State: The LeWM Architecture
Follow ordered visual trajectories through the encoder and action-conditioned dynamics predictor, from raw tensors to predicted latent states.
- 06 Chapter 5 — Trajectory Data: To the Model, the World Is Not an Image Collection Represent observations and actions as offline, reward-free episodes while treating data coverage as a boundary on reliable imagination. available now
- 07 Chapter 6 — The Visual Encoder: Issuing a “State Passport” for Every Frame Track images through patches, a Vision Transformer, the [CLS] summary, projection, normalization, and the resulting latent-vector shapes. available now
- 08 Chapter 7 — The Dynamics Predictor: Moving Time Forward in the Mind Condition latent dynamics on actions with AdaLN and causal context, then contrast teacher-forced training with autoregressive planning rollouts. available now
- 09 Chapter 8 — A Complete Forward Pass: Follow One Batch from Start to Finish Walk through observation encoding, action conditioning, target construction, both loss terms, gradient flow, hooks, and shape assertions in a minimal example. available now
Part 3 — Preventing the Model from Cheating: Prediction Loss and SIGReg
Explain representation collapse, the two-part learning objective, and the fully end-to-end training mechanism of the original LeWM v3.
- 10 Chapter 9 — The Most Dangerous Shortcut: Representation Collapse Learn how constant embeddings can satisfy prediction loss, why collapse differs from overfitting, and how to diagnose it with latent statistics and neighbors. available now
- 11 Chapter 10 — Prediction Loss: How the Model Learns the Next Step Interpret next-latent-state MSE, its limits for planning, and the gap between accurate teacher-forced steps and unstable long rollouts. available now
- 12 Chapter 11 — The Intuition Behind SIGReg: Letting Representation Space “Breathe” Use isotropic Gaussian targets, random projections, and the Epps–Pulley statistic to understand non-collapse without negatives or a teacher network. available now
- 13 Chapter 12 — Keep the Mathematics Minimal but Sufficient Derive only the prediction and SIGReg machinery needed to read the full objective, including the precise role of the loss trade-off coefficient. available now
- 14 Chapter 13 — The Original LeWM’s End-to-End Training Mechanism Show why LeWM v3 uses neither stop-gradient, an EMA teacher, nor a pretrained visual encoder, and distinguish its training graph from related JEPA methods. available now
- 15 Chapter 14 — Train a Model That Does Not Collapse Immediately Turn the objective into a stable training workflow covering initialization, data slicing, memory, optimization, projections, checkpoints, smoke tests, and required metrics. available now
Part 4 — Putting the Model into Action: Planning in Latent Space
Use the learned dynamics for goal-conditioned action search while exposing the assumptions and failure modes hidden inside the planning cost.
- 16 Chapter 15 — Goal-Conditioned Planning: From “Where Am I?” to “Where Do I Want to Go?” Encode current and goal images together, roll out candidate action sequences, and select plans by terminal latent cost without requiring rewards during model training. available now
- 17 Chapter 16 — Latent Euclidean Distance: Convenient, but Not Necessarily Reliable Test when visual similarity and Euclidean distance fail to represent directed reachability, obstacles, route length, or a trustworthy planning cost. available now
- 18 Chapter 17 — CEM: Searching for Actions Through an Elimination Tournament Implement iterative candidate sampling, latent simulation, elite selection, and distribution updates while balancing horizon, population, and iterations. available now
- 19 Chapter 18 — MPC: Do Not Trust the Model for Too Long at Once Close the control loop by executing only a short action segment, observing again, and replanning to limit model error and environmental disturbance. available now
- 20 Chapter 19 — Long-Horizon Rollouts: How Small Errors Snowball into Major Failures Diagnose accumulated error, model-generated distribution shift, latent drift, hidden uncertainty, and planners that exploit model flaws. available now
- 21 Chapter 20 — Implement a Minimal LeWM Planner from Scratch Build vectorized CEM rollouts and terminal costs, execute the first action segment, compare random and oracle baselines, and visualize one failure. available now
Part 5 — Engineering Reproduction: From Paper to Running System
Reproduce the official implementation and evaluate it with transparent data, compute, planning, and failure-analysis protocols.
- 22 Chapter 21 — The Official Repository and Experimental Environment Navigate the official code, Hydra configuration, HDF5 trajectories, single-GPU workflow, checkpoints, and training-to-evaluation call graph. available now
- 23 Chapter 22 — First Experiment: A TwoRoom Smoke Test Train a small model, test collapse and one-step versus multi-step prediction, run CEM, and visualize cross-room planning failures. available now
- 24 Chapter 23 — Second Experiment: Reproducing PushT Move from geometric navigation to contact dynamics through data preparation, model training, goal-image planning, metrics, and failure classification. available now
- 25 Chapter 24 — How to Evaluate a World Model Fairly Report one-step, multi-step, closed-loop, and planning metrics under matched data and compute budgets, strong baselines, multiple seeds, and explicit conditions. available now
- 26 Chapter 25 — Failure-Diagnosis Manual Trace failures layer by layer from data coverage and latent collapse through ignored actions, unstable statistics, misleading costs, and a stalled planner. available now
Part 6 — What Has LeWM Actually Learned?
Examine what latent representations encode while keeping probes, visualizations, and surprise experiments within their legitimate evidentiary limits.
- 27 Chapter 26 — Linear Probes: Which Physical Variables Are Encoded in the Latent State? Probe position, velocity, angle, and goal distance with proper baselines while separating decodability, model use, correlation, and causation. available now
- 28 Chapter 27 — Give Latent Space a “Health Check” Inspect moments, covariance spectra, effective dimension, neighbors, trajectory continuity, and path geometry without overreading t-SNE, UMAP, or decoders. available now
- 29 Chapter 28 — Violation of Expectation: Is the Model Surprised by “Impossible Events”? Compare normal dynamics with controlled violations, define surprise through latent prediction error, and state what the resulting evidence can and cannot prove. available now
- 30 Chapter 29 — How to Discuss “Understanding the World” Rigorously Replace broad claims with tests of predictability, decodability, controllability, generalization, counterfactual consistency, causality, and planning utility. available now
Part 7 — Why “Accurate Prediction” Can Still Produce “Poor Planning”
Organize later work by the failure mechanisms separating predictive accuracy from useful, supported, long-horizon control.
- 31 Chapter 30 — The Gap Between the Training Objective and the Planning Objective Compare local MSE with global reachability, directed temporal distance, multi-timescale prediction, and rollout consistency without rewriting original LeWM. available now
- 32 Chapter 31 — Global Non-Collapse Does Not Guarantee Preservation of Task-Relevant Dynamics Test state identifiability, physical invariants, and counterfactual action sensitivity while treating PhyLatent as a recent diagnostic proposal, not the LeWM baseline. available now
- 33 Chapter 32 — When Is an Isotropic Gaussian Prior Too Strong? Study low-dimensional geometry, neighborhood distortion, temporal drift, and proposed subspace, temporal-centering, and tail-matching alternatives to SIGReg. available now
- 34 Chapter 33 — Long-Horizon Planning: Predict Farther or Plan More Intelligently? Compare action-prefix prediction, parallel futures, macro-actions, and hierarchical planning while confronting unsupported latent subgoals and out-of-distribution search. available now
- 35 Chapter 34 — From Positional Distance to Task Progress Separate distance, reachability, and staged task progress, including progress-aware and hyperbolic approaches without disguising task priors as general understanding. available now
- 36 Chapter 35 — Multi-Task Learning, Real Robots, and Visual Distractions Examine task identity, visual variation, latent drift, auxiliary depth, predictor capacity, partial observability, robot noise, and the simulation-to-reality evidence gap. available now
- 37 Chapter 36 — Theoretical Boundaries: When Can the True State Be Identified? Use assumption-bounded identifiability results to guide experiments while keeping stationarity, additive noise, linearity, observability, and multimodality explicit. available now
Part 8 — From Reproducer to Researcher
Turn diagnosed failure mechanisms into controlled experiments, executable projects, and precisely framed open questions.
- 38 Chapter 37 — Design a Credible LeWM Improvement Experiment Isolate one failure mechanism with counterexamples, matched capacity and compute, ablations, oracle conditions, and honest reporting of negative results. available now
- 39 Chapter 38 — Twelve Executable Research Projects Develop concrete studies of reachability, uncertainty, multi-step objectives, temporal hierarchy, topology, controllability, support constraints, language goals, and transfer. available now
- 40 Chapter 39 — Open Questions in LeWM Research Frame unresolved questions about planning geometry, non-collapse, multimodal futures, out-of-distribution imagination, hierarchy, latent capacity, and future paradigms. available now
Appendices
Keep the mathematical, implementation, configuration, evidence, terminology, reproduction, and review references close to the main learning path.
- 41 Appendix A — The Minimum Necessary Mathematical Toolkit Review vectors, batches, covariance, Gaussian geometry, random projections, gradients, MSE, autoregressive error, Monte Carlo estimation, and CEM probability. available now
- 42 Appendix B — PyTorch Implementation Quick Reference Find concise patterns for data loading, Transformer inputs, causal masks, AdaLN, SIGReg, vectorized rollout, CEM, mixed precision, and checkpoints. available now
- 43 Appendix C — Complete Tensor-Shape Table Trace tensors from [B,T,C,H,W] through [B,T,D] to [B,N,H,D], including dimension meanings, broadcasting rules, and common errors. available now
- 44 Appendix D — Experiment Configuration Cards Compare TwoRoom, Reacher, PushT, OGBench-Cube, teaching, paper-reproduction, and resource-constrained single-GPU configurations. available now
- 45 Appendix E — Paper Timeline and Evidence Levels Track which system layer each paper changes, its target failure, version and release date, public code, evidence strength, reproduction status, and limits on generalization. available now
- 46 Appendix F — Glossary Define core terms with a formal definition, one-sentence explanation, everyday analogy, and common misconception. available now
- 47 Appendix G — Reproduction Checklist Record software, hardware, data hashes, seeds, parameter counts, budgets, baselines, episodes, failure cases, raw logs, and checkpoints. available now
- 48 Appendix H — Expert-Review Checklist Audit the training graph, variant boundaries, probe interpretation, benchmark generalization, causal language, theoretical assumptions, negative results, and code paths. available now