JEPA4Japan · tutorials

Chapter 19 — Long-Horizon Rollouts: How Small Errors Snowball into Major Failures

683 words 4 min read #LeWorldModel#World Models#JEPA

Diagnose accumulated error, model-generated distribution shift, latent drift, hidden uncertainty, and planners that exploit model flaws.

Course progress Course outline 48 of 48 lessons available

Part 0 — Reading Guide: What Exactly Are We Going to Learn?

  1. 01 Chapter 0 — Before You Begin available now

Part 1 — World Models: An Internal Sandbox for the Agent

  1. 02 Chapter 1 — Why an Agent Needs to “Imagine the Future” available now
  2. 03 Chapter 2 — Why Not Predict the Next Image Directly? available now
  3. 04 Chapter 3 — The JEPA Idea: Predict Meaning, Not a Replica of the Image available now
  4. 05 Chapter 4 — Understand LeWM in One Diagram available now

Part 2 — Turning Images into State: The LeWM Architecture

  1. 06 Chapter 5 — Trajectory Data: To the Model, the World Is Not an Image Collection available now
  2. 07 Chapter 6 — The Visual Encoder: Issuing a “State Passport” for Every Frame available now
  3. 08 Chapter 7 — The Dynamics Predictor: Moving Time Forward in the Mind available now
  4. 09 Chapter 8 — A Complete Forward Pass: Follow One Batch from Start to Finish available now

Part 3 — Preventing the Model from Cheating: Prediction Loss and SIGReg

  1. 10 Chapter 9 — The Most Dangerous Shortcut: Representation Collapse available now
  2. 11 Chapter 10 — Prediction Loss: How the Model Learns the Next Step available now
  3. 12 Chapter 11 — The Intuition Behind SIGReg: Letting Representation Space “Breathe” available now
  4. 13 Chapter 12 — Keep the Mathematics Minimal but Sufficient available now
  5. 14 Chapter 13 — The Original LeWM’s End-to-End Training Mechanism available now
  6. 15 Chapter 14 — Train a Model That Does Not Collapse Immediately available now

Part 4 — Putting the Model into Action: Planning in Latent Space

  1. 16 Chapter 15 — Goal-Conditioned Planning: From “Where Am I?” to “Where Do I Want to Go?” available now
  2. 17 Chapter 16 — Latent Euclidean Distance: Convenient, but Not Necessarily Reliable available now
  3. 18 Chapter 17 — CEM: Searching for Actions Through an Elimination Tournament available now
  4. 19 Chapter 18 — MPC: Do Not Trust the Model for Too Long at Once available now
  5. 20 Chapter 19 — Long-Horizon Rollouts: How Small Errors Snowball into Major Failures Current lesson
  6. 21 Chapter 20 — Implement a Minimal LeWM Planner from Scratch available now

Part 5 — Engineering Reproduction: From Paper to Running System

  1. 22 Chapter 21 — The Official Repository and Experimental Environment available now
  2. 23 Chapter 22 — First Experiment: A TwoRoom Smoke Test available now
  3. 24 Chapter 23 — Second Experiment: Reproducing PushT available now
  4. 25 Chapter 24 — How to Evaluate a World Model Fairly available now
  5. 26 Chapter 25 — Failure-Diagnosis Manual available now

Part 6 — What Has LeWM Actually Learned?

  1. 27 Chapter 26 — Linear Probes: Which Physical Variables Are Encoded in the Latent State? available now
  2. 28 Chapter 27 — Give Latent Space a “Health Check” available now
  3. 29 Chapter 28 — Violation of Expectation: Is the Model Surprised by “Impossible Events”? available now
  4. 30 Chapter 29 — How to Discuss “Understanding the World” Rigorously available now

Part 7 — Why “Accurate Prediction” Can Still Produce “Poor Planning”

  1. 31 Chapter 30 — The Gap Between the Training Objective and the Planning Objective available now
  2. 32 Chapter 31 — Global Non-Collapse Does Not Guarantee Preservation of Task-Relevant Dynamics available now
  3. 33 Chapter 32 — When Is an Isotropic Gaussian Prior Too Strong? available now
  4. 34 Chapter 33 — Long-Horizon Planning: Predict Farther or Plan More Intelligently? available now
  5. 35 Chapter 34 — From Positional Distance to Task Progress available now
  6. 36 Chapter 35 — Multi-Task Learning, Real Robots, and Visual Distractions available now
  7. 37 Chapter 36 — Theoretical Boundaries: When Can the True State Be Identified? available now

Part 8 — From Reproducer to Researcher

  1. 38 Chapter 37 — Design a Credible LeWM Improvement Experiment available now
  2. 39 Chapter 38 — Twelve Executable Research Projects available now
  3. 40 Chapter 39 — Open Questions in LeWM Research available now

Appendices

  1. 41 Appendix A — The Minimum Necessary Mathematical Toolkit available now
  2. 42 Appendix B — PyTorch Implementation Quick Reference available now
  3. 43 Appendix C — Complete Tensor-Shape Table available now
  4. 44 Appendix D — Experiment Configuration Cards available now
  5. 45 Appendix E — Paper Timeline and Evidence Levels available now
  6. 46 Appendix F — Glossary available now
  7. 47 Appendix G — Reproduction Checklist available now
  8. 48 Appendix H — Expert-Review Checklist available now

The big picture

  1. Start realEncode a true history.
  2. PredictMake the next latent.
  3. Feed it backThe estimate becomes context.
  4. MeasureTrack drift at every horizon.
A copy of a copy can leave the road, even when the first copy looks good.

A chain of repeated copies gradually drifts from the original.

A tiny story

Make one photocopy of a picture and it may look fine. Copy the copy many times and small distortions can matter.

LeWM rollout is autoregressive. The first prediction starts from encoded real observations. Later predictions can consume earlier predicted latents. Training and one-step tests usually see more real contexts than a long planning rollout does. This is the train–rollout context gap.

The photocopy story has a limit: error need not grow smoothly or forever. Dynamics can amplify, rotate, cancel, or damp a discrepancy. Measure the curve; do not assume exponential or monotonic drift.

The real rule

Keep two lanes for the same held-out episode and the same recorded action sequence:

  • Recorded-context lane: encode the real observation again at every step.
  • Self-fed lane: after the initial history, feed predictions back into the next prediction.

Compare each predicted latent with the frozen encoder’s latent for the observation that actually followed. Plot error by model horizon, not only at the endpoint. Slice by door crossings, contact onset, contact loss, action magnitude, and empirical support.

Original LeWM v3 emits one latent estimate. It does not output a calibrated confidence or a distribution of possible futures. Deterministic output does not mean the environment is certain.

One-step accuracy is still useful. It simply answers a smaller question. Good recorded-context prediction can coexist with poor self-fed rollout, a bad terminal cost, or a broken execution adapter.

The trick that can fool us

CEM is paid to find low predicted cost. If one unsupported action sequence makes the learned rollout unrealistically optimistic, more candidates or refinement rounds can discover that loophole more reliably.

Test this claim instead of naming every failure “model exploitation”:

  1. Freeze a bank of action sequences.
  2. Compare imagined cost with real executed outcome.
  3. Label action-support and horizon.
  4. Increase search budget while keeping the model, task, and real-action budget fixed.

Exploitation becomes plausible if predicted best cost improves, selected sequences move farther from support, and real outcomes worsen in a paired way. A final crash alone cannot separate dynamics, cost, search, support, and execution.

If oracle cost fixes candidate ranking but rollout endpoints remain wrong, repair dynamics. If oracle dynamics plus learned cost still prefers the wrong route, repair cost or representation geometry. Shortening the horizon reduces rollout depth; it does not fix a ruler that ranks across-wall states incorrectly.

Experiment receipt

Store the checkpoint, preprocessing, initial real history, aligned recorded actions, action-block semantics, teacher-forced and self-fed latents, per-horizon errors, support labels, and all seeds. Mark the first qualitative divergence, not only the last ugly frame.

The mechanism is grounded in frozen LeWM rollout code at author commit 8edfeb3 and the limitation discussion in LeWorldModel v3. The model-exploitation idea also has neighboring precedent in Model-Based Planning with Energy-Based Models.

The independent TwoRoom reimplementation reports that one-step error orders short-horizon but not long-horizon success across three checkpoints in one environment, with one training seed per configuration. It leaves the cause open; it does not prove universal drift or exploitation.

Quick check

  1. What changes between recorded-context and self-fed prediction?
  2. Why is one deterministic latent not calibrated certainty?
  3. Which paired trend would make model exploitation a credible hypothesis?