JEPA4Japan · tutorials

Chapter 0 — Before You Begin

686 words 4 min read #LeWorldModel#World Models#JEPA

Distinguish LeWM from LeJEPA, fix the paper and code baselines, separate confirmed facts from research frontiers, and preview the TwoRoom-to-PushT learning path.

Course progress Course outline 48 of 48 lessons available

Part 0 — Reading Guide: What Exactly Are We Going to Learn?

  1. 01 Chapter 0 — Before You Begin Current lesson

Part 1 — World Models: An Internal Sandbox for the Agent

  1. 02 Chapter 1 — Why an Agent Needs to “Imagine the Future” available now
  2. 03 Chapter 2 — Why Not Predict the Next Image Directly? available now
  3. 04 Chapter 3 — The JEPA Idea: Predict Meaning, Not a Replica of the Image available now
  4. 05 Chapter 4 — Understand LeWM in One Diagram available now

Part 2 — Turning Images into State: The LeWM Architecture

  1. 06 Chapter 5 — Trajectory Data: To the Model, the World Is Not an Image Collection available now
  2. 07 Chapter 6 — The Visual Encoder: Issuing a “State Passport” for Every Frame available now
  3. 08 Chapter 7 — The Dynamics Predictor: Moving Time Forward in the Mind available now
  4. 09 Chapter 8 — A Complete Forward Pass: Follow One Batch from Start to Finish available now

Part 3 — Preventing the Model from Cheating: Prediction Loss and SIGReg

  1. 10 Chapter 9 — The Most Dangerous Shortcut: Representation Collapse available now
  2. 11 Chapter 10 — Prediction Loss: How the Model Learns the Next Step available now
  3. 12 Chapter 11 — The Intuition Behind SIGReg: Letting Representation Space “Breathe” available now
  4. 13 Chapter 12 — Keep the Mathematics Minimal but Sufficient available now
  5. 14 Chapter 13 — The Original LeWM’s End-to-End Training Mechanism available now
  6. 15 Chapter 14 — Train a Model That Does Not Collapse Immediately available now

Part 4 — Putting the Model into Action: Planning in Latent Space

  1. 16 Chapter 15 — Goal-Conditioned Planning: From “Where Am I?” to “Where Do I Want to Go?” available now
  2. 17 Chapter 16 — Latent Euclidean Distance: Convenient, but Not Necessarily Reliable available now
  3. 18 Chapter 17 — CEM: Searching for Actions Through an Elimination Tournament available now
  4. 19 Chapter 18 — MPC: Do Not Trust the Model for Too Long at Once available now
  5. 20 Chapter 19 — Long-Horizon Rollouts: How Small Errors Snowball into Major Failures available now
  6. 21 Chapter 20 — Implement a Minimal LeWM Planner from Scratch available now

Part 5 — Engineering Reproduction: From Paper to Running System

  1. 22 Chapter 21 — The Official Repository and Experimental Environment available now
  2. 23 Chapter 22 — First Experiment: A TwoRoom Smoke Test available now
  3. 24 Chapter 23 — Second Experiment: Reproducing PushT available now
  4. 25 Chapter 24 — How to Evaluate a World Model Fairly available now
  5. 26 Chapter 25 — Failure-Diagnosis Manual available now

Part 6 — What Has LeWM Actually Learned?

  1. 27 Chapter 26 — Linear Probes: Which Physical Variables Are Encoded in the Latent State? available now
  2. 28 Chapter 27 — Give Latent Space a “Health Check” available now
  3. 29 Chapter 28 — Violation of Expectation: Is the Model Surprised by “Impossible Events”? available now
  4. 30 Chapter 29 — How to Discuss “Understanding the World” Rigorously available now

Part 7 — Why “Accurate Prediction” Can Still Produce “Poor Planning”

  1. 31 Chapter 30 — The Gap Between the Training Objective and the Planning Objective available now
  2. 32 Chapter 31 — Global Non-Collapse Does Not Guarantee Preservation of Task-Relevant Dynamics available now
  3. 33 Chapter 32 — When Is an Isotropic Gaussian Prior Too Strong? available now
  4. 34 Chapter 33 — Long-Horizon Planning: Predict Farther or Plan More Intelligently? available now
  5. 35 Chapter 34 — From Positional Distance to Task Progress available now
  6. 36 Chapter 35 — Multi-Task Learning, Real Robots, and Visual Distractions available now
  7. 37 Chapter 36 — Theoretical Boundaries: When Can the True State Be Identified? available now

Part 8 — From Reproducer to Researcher

  1. 38 Chapter 37 — Design a Credible LeWM Improvement Experiment available now
  2. 39 Chapter 38 — Twelve Executable Research Projects available now
  3. 40 Chapter 39 — Open Questions in LeWM Research available now

Appendices

  1. 41 Appendix A — The Minimum Necessary Mathematical Toolkit available now
  2. 42 Appendix B — PyTorch Implementation Quick Reference available now
  3. 43 Appendix C — Complete Tensor-Shape Table available now
  4. 44 Appendix D — Experiment Configuration Cards available now
  5. 45 Appendix E — Paper Timeline and Evidence Levels available now
  6. 46 Appendix F — Glossary available now
  7. 47 Appendix G — Reproduction Checklist available now
  8. 48 Appendix H — Expert-Review Checklist available now

The big picture

  1. Observepixels + actions
  2. Compressmake latent states
  3. Imaginetry action futures
  4. Actsearch, move, check
LeWM is a learned rehearsal room. It is not the camera, the goal, the driver, or a perfect physics oracle.

A camera, a direct policy, and a LeWM rehearsal sandbox are different tools.

This course has one spine: observe → compress → imagine → search → act → diagnose. Every later box belongs somewhere on that line.

A tiny story

Put three tools on a robot’s workbench. A camera records one journey. A direct policy chooses an action, possibly with rich memory, without explicitly searching many futures at decision time. A world-model planner rehearses proposed actions inside a learned model before choosing.

LeWorldModel, or LeWM, is the third tool. It learns from ordered visual–action trajectories. An encoder turns an observation into a compact latent “passport.” An action-conditioned predictor advances those passports. After training, CEM searches action sequences through the frozen model, and MPC executes a configured prefix before looking again.

The passport is not a literal map. Two nearby latent points may sit on opposite sides of a real wall. A compact state can preserve wallpaper and forget the doorway.

The real rule

Keep four names separate:

  • A world model is a job: represent enough change to reason about consequences.
  • JEPA is a broad idea: predict a learned target representation from a context representation.
  • LeJEPA studies predictive representation learning with SIGReg and non-collapse.
  • LeWM adds time, actions, latent dynamics, and a planning loop.

Our baseline is LeWorldModel v3, revised 2026-06-03, plus the official code frozen at 8edfeb336732b5f3ce7b8b210d0ba370a09e2cac, dated 2026-05-22. LeJEPA v3 and the JEPA position paper v0.9.2 explain related ideas, not interchangeable implementations.

Original LeWM v3 uses one shared trainable visual encoder. Its next-observation targets stay connected to that encoder. It uses no training-time stop-gradient, no EMA teacher, and no pretrained visual encoder. Baseline world-model training needs no task reward or goal, but it absolutely needs ordered actions. The goal arrives later, during planning.

LEARN:     observations + actions -> latent dynamics + non-collapse pressure
FREEZE:    encoder and predictor
ACT:       current + goal -> CEM plans -> execute a prefix -> observe again
DIAGNOSE:  data -> representation -> dynamics -> cost -> search -> execution

The trick that fools us

Imagine a predictor that cannot hear actions. Most recorded TwoRoom trips follow the same doorway route, so history alone may predict the common continuation and produce a respectable loss. During planning, however, “up,” “right,” and “stay” all create nearly the same imagined future. CEM is holding an audition for a judge with earplugs.

That counterexample proves only this: predicting common chronology is not enough for an action-conditioned planning model. A different prediction after an action swap is necessary evidence, but it still does not prove the direction, size, or long-term effect is correct.

Evidence receipt

The course’s evidence cutoff is 2026-08-20. arXiv:2608.10145v1 independently reimplemented one single-seed TwoRoom system and reproduced a position probe and repository-protocol planning; it did not reproduce PushT, OGBench-Cube, or multi-seed robustness. arXiv:2608.12959v1 reused that system, so it is a follow-up in the same evidence chain, not a second independent replication.

TwoRoom makes walls, detours, and bad costs easy to see. PushT adds contact, rotation, and sliding. Neither benchmark proves general physics or safe real-robot control. Tutorial stories and diagrams explain mechanisms; they are not paper results.

Quick check

  1. Why is LeWM a rehearsal room rather than a camera or policy?
  2. Which original-v3 facts rule out an EMA-teacher diagram?
  3. Why can a reward-free world model still need a goal when it acts?