JEPA4Japan · tutorials

Chapter 13 — The Original LeWM’s End-to-End Training Mechanism

710 words 4 min read #LeWorldModel#World Models#JEPA

Show why LeWM v3 uses neither stop-gradient, an EMA teacher, nor a pretrained visual encoder, and distinguish its training graph from related JEPA methods.

Course progress Course outline 48 of 48 lessons available

Part 0 — Reading Guide: What Exactly Are We Going to Learn?

  1. 01 Chapter 0 — Before You Begin available now

Part 1 — World Models: An Internal Sandbox for the Agent

  1. 02 Chapter 1 — Why an Agent Needs to “Imagine the Future” available now
  2. 03 Chapter 2 — Why Not Predict the Next Image Directly? available now
  3. 04 Chapter 3 — The JEPA Idea: Predict Meaning, Not a Replica of the Image available now
  4. 05 Chapter 4 — Understand LeWM in One Diagram available now

Part 2 — Turning Images into State: The LeWM Architecture

  1. 06 Chapter 5 — Trajectory Data: To the Model, the World Is Not an Image Collection available now
  2. 07 Chapter 6 — The Visual Encoder: Issuing a “State Passport” for Every Frame available now
  3. 08 Chapter 7 — The Dynamics Predictor: Moving Time Forward in the Mind available now
  4. 09 Chapter 8 — A Complete Forward Pass: Follow One Batch from Start to Finish available now

Part 3 — Preventing the Model from Cheating: Prediction Loss and SIGReg

  1. 10 Chapter 9 — The Most Dangerous Shortcut: Representation Collapse available now
  2. 11 Chapter 10 — Prediction Loss: How the Model Learns the Next Step available now
  3. 12 Chapter 11 — The Intuition Behind SIGReg: Letting Representation Space “Breathe” available now
  4. 13 Chapter 12 — Keep the Mathematics Minimal but Sufficient available now
  5. 14 Chapter 13 — The Original LeWM’s End-to-End Training Mechanism Current lesson
  6. 15 Chapter 14 — Train a Model That Does Not Collapse Immediately available now

Part 4 — Putting the Model into Action: Planning in Latent Space

  1. 16 Chapter 15 — Goal-Conditioned Planning: From “Where Am I?” to “Where Do I Want to Go?” available now
  2. 17 Chapter 16 — Latent Euclidean Distance: Convenient, but Not Necessarily Reliable available now
  3. 18 Chapter 17 — CEM: Searching for Actions Through an Elimination Tournament available now
  4. 19 Chapter 18 — MPC: Do Not Trust the Model for Too Long at Once available now
  5. 20 Chapter 19 — Long-Horizon Rollouts: How Small Errors Snowball into Major Failures available now
  6. 21 Chapter 20 — Implement a Minimal LeWM Planner from Scratch available now

Part 5 — Engineering Reproduction: From Paper to Running System

  1. 22 Chapter 21 — The Official Repository and Experimental Environment available now
  2. 23 Chapter 22 — First Experiment: A TwoRoom Smoke Test available now
  3. 24 Chapter 23 — Second Experiment: Reproducing PushT available now
  4. 25 Chapter 24 — How to Evaluate a World Model Fairly available now
  5. 26 Chapter 25 — Failure-Diagnosis Manual available now

Part 6 — What Has LeWM Actually Learned?

  1. 27 Chapter 26 — Linear Probes: Which Physical Variables Are Encoded in the Latent State? available now
  2. 28 Chapter 27 — Give Latent Space a “Health Check” available now
  3. 29 Chapter 28 — Violation of Expectation: Is the Model Surprised by “Impossible Events”? available now
  4. 30 Chapter 29 — How to Discuss “Understanding the World” Rigorously available now

Part 7 — Why “Accurate Prediction” Can Still Produce “Poor Planning”

  1. 31 Chapter 30 — The Gap Between the Training Objective and the Planning Objective available now
  2. 32 Chapter 31 — Global Non-Collapse Does Not Guarantee Preservation of Task-Relevant Dynamics available now
  3. 33 Chapter 32 — When Is an Isotropic Gaussian Prior Too Strong? available now
  4. 34 Chapter 33 — Long-Horizon Planning: Predict Farther or Plan More Intelligently? available now
  5. 35 Chapter 34 — From Positional Distance to Task Progress available now
  6. 36 Chapter 35 — Multi-Task Learning, Real Robots, and Visual Distractions available now
  7. 37 Chapter 36 — Theoretical Boundaries: When Can the True State Be Identified? available now

Part 8 — From Reproducer to Researcher

  1. 38 Chapter 37 — Design a Credible LeWM Improvement Experiment available now
  2. 39 Chapter 38 — Twelve Executable Research Projects available now
  3. 40 Chapter 39 — Open Questions in LeWM Research available now

Appendices

  1. 41 Appendix A — The Minimum Necessary Mathematical Toolkit available now
  2. 42 Appendix B — PyTorch Implementation Quick Reference available now
  3. 43 Appendix C — Complete Tensor-Shape Table available now
  4. 44 Appendix D — Experiment Configuration Cards available now
  5. 45 Appendix E — Paper Timeline and Evidence Levels available now
  6. 46 Appendix F — Glossary available now
  7. 47 Appendix G — Reproduction Checklist available now
  8. 48 Appendix H — Expert-Review Checklist available now

The big picture

  1. One encodercontext and future
  2. Actionscondition the predictor
  3. Connected targetno stop sign
  4. SIGRegoppose the constant cheat
Original LeWM v3 trains one live coordinate system end to end—without an EMA teacher.

Context and future observations use one shared connected encoder in the v3 training graph.

A tiny story

An easy but wrong sketch shows a context encoder, a pale target tower, a stop sign, and an EMA arrow. It looks familiar because BYOL, I-JEPA, and V-JEPA use related target mechanisms in named versions. It is not original LeWM v3.

LeWM is more like a mapmaker and route-rule learner editing one live map together. SIGReg is a zoning rule that prevents every address from becoming one lot. Zoning can spread a city without making its roads correct.

The technical backpack

The v3 training topology is:

observations -> one shared trainable encoder -> all latents
context latents + recorded actions -> predicted future latent
predicted future <-> connected future latent
observation latents -> SIGReg
one total objective -> joint update

The word target means “what the prediction is compared with.” It does not specify who owns the weights or how they update. Frozen train.py slices future latents from the same encoded sequence and does not detach them before the prediction loss. Thus prediction pressure reaches the shared encoder through both context and future uses, as well as the action encoder and predictor.

The Gaussian reference is analytic, not a learned teacher. The recorded next observation is data, not a teacher network. Original v3 also sets visual pretraining off; “from scratch” means the visual encoder exists but does not begin from an external pretrained checkpoint.

Prediction alone permits every place to receive the same address. At the ideal SIGReg target, a repeated point cannot match a non-degenerate isotropic Gaussian. The joint objective therefore opposes the easiest exact constant solution. It does not make collapse impossible in finite optimization or ensure that a spread representation preserves rare task variables.

“End to end” stops at the training boundary. Dataset construction is not learned. During later CEM planning, encoder and predictor weights are frozen and no planning gradient returns to retrain them.

Named graphs remain different: BYOL v3, I-JEPA v3, and V-JEPA v1 use stop-gradient/EMA target mechanisms in their cited recipes; LeWorldModel v3 uses a shared connected encoder plus SIGReg. This is a mechanism comparison, not a performance ranking across unequal data and tasks.

The trick that fools us

Three .detach() calls can look identical in a keyword search while meaning different things:

  • future target detached before training loss: changes the method;
  • metric copy detached after loss: logging bookkeeping;
  • goal detached during frozen planning: no model update was possible anyway.

Audit value, timing, and phase. Then run a tiny topology probe:

A: future target connected
B: future target detached before comparison
C: target connected + SIGReg

Hooks should show B losing only the target-side prediction route, while C adds an encoder-side population route. One batch reveals connectivity, not which variant will train or plan better.

Evidence receipt

The graph comes from LeWorldModel v3, frozen train.py, and frozen model/lewm.yaml. Comparisons are tied to BYOL v3, I-JEPA v3, and V-JEPA v1.

Allowed claim: LeWM v3 is fully end to end and structurally differs from those named stopped-target/EMA recipes. Do not claim EMA is obsolete, SIGReg universally replaces teachers, from-scratch training is inherently better, or connectivity proves learned physics.

Quick check

  1. Why does “target” not imply “teacher”?
  2. Which learned paths receive prediction pressure, and which receive SIGReg directly?
  3. Which detach changes v3: logging, frozen goal, or future target before loss?