JEPA4Japan · tutorials

Chapter 18 — MPC: Do Not Trust the Model for Too Long at Once

676 words 4 min read #LeWorldModel#World Models#JEPA

Close the control loop by executing only a short action segment, observing again, and replanning to limit model error and environmental disturbance.

Course progress Course outline 48 of 48 lessons available

Part 0 — Reading Guide: What Exactly Are We Going to Learn?

  1. 01 Chapter 0 — Before You Begin available now

Part 1 — World Models: An Internal Sandbox for the Agent

  1. 02 Chapter 1 — Why an Agent Needs to “Imagine the Future” available now
  2. 03 Chapter 2 — Why Not Predict the Next Image Directly? available now
  3. 04 Chapter 3 — The JEPA Idea: Predict Meaning, Not a Replica of the Image available now
  4. 05 Chapter 4 — Understand LeWM in One Diagram available now

Part 2 — Turning Images into State: The LeWM Architecture

  1. 06 Chapter 5 — Trajectory Data: To the Model, the World Is Not an Image Collection available now
  2. 07 Chapter 6 — The Visual Encoder: Issuing a “State Passport” for Every Frame available now
  3. 08 Chapter 7 — The Dynamics Predictor: Moving Time Forward in the Mind available now
  4. 09 Chapter 8 — A Complete Forward Pass: Follow One Batch from Start to Finish available now

Part 3 — Preventing the Model from Cheating: Prediction Loss and SIGReg

  1. 10 Chapter 9 — The Most Dangerous Shortcut: Representation Collapse available now
  2. 11 Chapter 10 — Prediction Loss: How the Model Learns the Next Step available now
  3. 12 Chapter 11 — The Intuition Behind SIGReg: Letting Representation Space “Breathe” available now
  4. 13 Chapter 12 — Keep the Mathematics Minimal but Sufficient available now
  5. 14 Chapter 13 — The Original LeWM’s End-to-End Training Mechanism available now
  6. 15 Chapter 14 — Train a Model That Does Not Collapse Immediately available now

Part 4 — Putting the Model into Action: Planning in Latent Space

  1. 16 Chapter 15 — Goal-Conditioned Planning: From “Where Am I?” to “Where Do I Want to Go?” available now
  2. 17 Chapter 16 — Latent Euclidean Distance: Convenient, but Not Necessarily Reliable available now
  3. 18 Chapter 17 — CEM: Searching for Actions Through an Elimination Tournament available now
  4. 19 Chapter 18 — MPC: Do Not Trust the Model for Too Long at Once Current lesson
  5. 20 Chapter 19 — Long-Horizon Rollouts: How Small Errors Snowball into Major Failures available now
  6. 21 Chapter 20 — Implement a Minimal LeWM Planner from Scratch available now

Part 5 — Engineering Reproduction: From Paper to Running System

  1. 22 Chapter 21 — The Official Repository and Experimental Environment available now
  2. 23 Chapter 22 — First Experiment: A TwoRoom Smoke Test available now
  3. 24 Chapter 23 — Second Experiment: Reproducing PushT available now
  4. 25 Chapter 24 — How to Evaluate a World Model Fairly available now
  5. 26 Chapter 25 — Failure-Diagnosis Manual available now

Part 6 — What Has LeWM Actually Learned?

  1. 27 Chapter 26 — Linear Probes: Which Physical Variables Are Encoded in the Latent State? available now
  2. 28 Chapter 27 — Give Latent Space a “Health Check” available now
  3. 29 Chapter 28 — Violation of Expectation: Is the Model Surprised by “Impossible Events”? available now
  4. 30 Chapter 29 — How to Discuss “Understanding the World” Rigorously available now

Part 7 — Why “Accurate Prediction” Can Still Produce “Poor Planning”

  1. 31 Chapter 30 — The Gap Between the Training Objective and the Planning Objective available now
  2. 32 Chapter 31 — Global Non-Collapse Does Not Guarantee Preservation of Task-Relevant Dynamics available now
  3. 33 Chapter 32 — When Is an Isotropic Gaussian Prior Too Strong? available now
  4. 34 Chapter 33 — Long-Horizon Planning: Predict Farther or Plan More Intelligently? available now
  5. 35 Chapter 34 — From Positional Distance to Task Progress available now
  6. 36 Chapter 35 — Multi-Task Learning, Real Robots, and Visual Distractions available now
  7. 37 Chapter 36 — Theoretical Boundaries: When Can the True State Be Identified? available now

Part 8 — From Reproducer to Researcher

  1. 38 Chapter 37 — Design a Credible LeWM Improvement Experiment available now
  2. 39 Chapter 38 — Twelve Executable Research Projects available now
  3. 40 Chapter 39 — Open Questions in LeWM Research available now

Appendices

  1. 41 Appendix A — The Minimum Necessary Mathematical Toolkit available now
  2. 42 Appendix B — PyTorch Implementation Quick Reference available now
  3. 43 Appendix C — Complete Tensor-Shape Table available now
  4. 44 Appendix D — Experiment Configuration Cards available now
  5. 45 Appendix E — Paper Timeline and Evidence Levels available now
  6. 46 Appendix F — Glossary available now
  7. 47 Appendix G — Reproduction Checklist available now
  8. 48 Appendix H — Expert-Review Checklist available now

The big picture

  1. ObserveEncode real evidence.
  2. PlanCEM imagines H blocks.
  3. ExecuteCommit to K blocks.
  4. Look againDiscard the stale suffix and replan.
CEM searches inside imagination; MPC brings back real observations.

An inner CEM search sits inside an outer observe, execute, and replan loop.

A tiny story

A map says “turn right,” but roadworks push you left. A sensible navigator looks again and draws a new route.

Model Predictive Control (MPC) does the same kind of thing. It imagines a finite plan, executes a declared prefix, receives a new real observation, and searches again. Feedback limits how long a prediction remains the only story. It does not guarantee that the encoder, predictor, cost, or search will interpret the new evidence correctly.

The real rule

Use three different clocks:

  • (H): model-level action blocks imagined in each candidate.
  • (K): model-level blocks actually executed before replanning, with (K\le H).
  • raw environment timesteps inside one action block.

In the reported and frozen LeWM configuration, (H=5), (K=5), and each block contains five consecutive environment actions. So the whole five-block plan spans 25 environment timesteps and is executed before the next planning call. This is still MPC across the episode, but it is not the often-drawn (K=1) feedback loop.

freeze the trained encoder and predictor
while the task is active:
  encode the real observation history and separate goal
  run inner CEM over H-block candidates
  select a plan using the recorded return convention
  execute its first K blocks in the environment
  receive a new real observation and plan again

CEM and MPC are not synonyms. CEM changes an action proposal within one planning call. MPC is the outer observe–plan–execute–observe policy. Neither updates LeWM weights during evaluation, and unrealized candidates never receive real feedback.

The trick that can fool us

Call the new observation an epistemic reset, but keep the limit visible. It resets dependence on the old imagined endpoint. It does not reveal hidden velocity, contact force, or a privileged simulator state. A systematic across-wall cost error can simply be chosen again.

Use one controlled shove after the first action block. Compare:

  1. an episode-open-loop controller;
  2. a tutorial (K=1) controller;
  3. the reported-style (K=5) controller.

Hold the start, goal, checkpoint, action bounds, total real-action budget, and seed schedule fixed. Record when each controller first sees the shove, how its next plan changes, planning calls, latency, and executed path. A shorter (K) receives fresher feedback but also more CEM calls. That is a feedback–compute trade-off, not a free win.

Experiment receipt

For each cycle, save the selected imagined plan, solid executed prefix, dashed unused suffix, returned observations, prefix forecast error, and replacement plan. Log both normalized model actions and raw environment actions; an adapter bug can imitate model failure.

The source boundary is LeWorldModel v3 and frozen TwoRoom evaluation config, commit 8edfeb3. The independent TwoRoom reimplementation also matches the five-block horizon and five-block execution interval in its own system.

The careful claim is: MPC can reduce reliance on long open-loop prediction when the deviation is observable and the remaining model, cost, search, and action path are useful. It is not a safety proof, a perfect-state reset, or evidence that shorter (K) is universally better.

Quick check

  1. Why must a report separate model blocks from environment timesteps?
  2. What does a new observation reset, and what can remain wrong?
  3. How do the inner CEM clock and outer MPC clock differ?