JEPA4Japan · tutorials

Chapter 33 — Long-Horizon Planning: Predict Farther or Plan More Intelligently?

818 words 4 min read #LeWorldModel#World Models#JEPA

Compare action-prefix prediction, parallel futures, macro-actions, and hierarchical planning while confronting unsupported latent subgoals and out-of-distribution search.

Course progress Course outline 48 of 48 lessons available

Part 0 — Reading Guide: What Exactly Are We Going to Learn?

  1. 01 Chapter 0 — Before You Begin available now

Part 1 — World Models: An Internal Sandbox for the Agent

  1. 02 Chapter 1 — Why an Agent Needs to “Imagine the Future” available now
  2. 03 Chapter 2 — Why Not Predict the Next Image Directly? available now
  3. 04 Chapter 3 — The JEPA Idea: Predict Meaning, Not a Replica of the Image available now
  4. 05 Chapter 4 — Understand LeWM in One Diagram available now

Part 2 — Turning Images into State: The LeWM Architecture

  1. 06 Chapter 5 — Trajectory Data: To the Model, the World Is Not an Image Collection available now
  2. 07 Chapter 6 — The Visual Encoder: Issuing a “State Passport” for Every Frame available now
  3. 08 Chapter 7 — The Dynamics Predictor: Moving Time Forward in the Mind available now
  4. 09 Chapter 8 — A Complete Forward Pass: Follow One Batch from Start to Finish available now

Part 3 — Preventing the Model from Cheating: Prediction Loss and SIGReg

  1. 10 Chapter 9 — The Most Dangerous Shortcut: Representation Collapse available now
  2. 11 Chapter 10 — Prediction Loss: How the Model Learns the Next Step available now
  3. 12 Chapter 11 — The Intuition Behind SIGReg: Letting Representation Space “Breathe” available now
  4. 13 Chapter 12 — Keep the Mathematics Minimal but Sufficient available now
  5. 14 Chapter 13 — The Original LeWM’s End-to-End Training Mechanism available now
  6. 15 Chapter 14 — Train a Model That Does Not Collapse Immediately available now

Part 4 — Putting the Model into Action: Planning in Latent Space

  1. 16 Chapter 15 — Goal-Conditioned Planning: From “Where Am I?” to “Where Do I Want to Go?” available now
  2. 17 Chapter 16 — Latent Euclidean Distance: Convenient, but Not Necessarily Reliable available now
  3. 18 Chapter 17 — CEM: Searching for Actions Through an Elimination Tournament available now
  4. 19 Chapter 18 — MPC: Do Not Trust the Model for Too Long at Once available now
  5. 20 Chapter 19 — Long-Horizon Rollouts: How Small Errors Snowball into Major Failures available now
  6. 21 Chapter 20 — Implement a Minimal LeWM Planner from Scratch available now

Part 5 — Engineering Reproduction: From Paper to Running System

  1. 22 Chapter 21 — The Official Repository and Experimental Environment available now
  2. 23 Chapter 22 — First Experiment: A TwoRoom Smoke Test available now
  3. 24 Chapter 23 — Second Experiment: Reproducing PushT available now
  4. 25 Chapter 24 — How to Evaluate a World Model Fairly available now
  5. 26 Chapter 25 — Failure-Diagnosis Manual available now

Part 6 — What Has LeWM Actually Learned?

  1. 27 Chapter 26 — Linear Probes: Which Physical Variables Are Encoded in the Latent State? available now
  2. 28 Chapter 27 — Give Latent Space a “Health Check” available now
  3. 29 Chapter 28 — Violation of Expectation: Is the Model Surprised by “Impossible Events”? available now
  4. 30 Chapter 29 — How to Discuss “Understanding the World” Rigorously available now

Part 7 — Why “Accurate Prediction” Can Still Produce “Poor Planning”

  1. 31 Chapter 30 — The Gap Between the Training Objective and the Planning Objective available now
  2. 32 Chapter 31 — Global Non-Collapse Does Not Guarantee Preservation of Task-Relevant Dynamics available now
  3. 33 Chapter 32 — When Is an Isotropic Gaussian Prior Too Strong? available now
  4. 34 Chapter 33 — Long-Horizon Planning: Predict Farther or Plan More Intelligently? Current lesson
  5. 35 Chapter 34 — From Positional Distance to Task Progress available now
  6. 36 Chapter 35 — Multi-Task Learning, Real Robots, and Visual Distractions available now
  7. 37 Chapter 36 — Theoretical Boundaries: When Can the True State Be Identified? available now

Part 8 — From Reproducer to Researcher

  1. 38 Chapter 37 — Design a Credible LeWM Improvement Experiment available now
  2. 39 Chapter 38 — Twelve Executable Research Projects available now
  3. 40 Chapter 39 — Open Questions in LeWM Research available now

Appendices

  1. 41 Appendix A — The Minimum Necessary Mathematical Toolkit available now
  2. 42 Appendix B — PyTorch Implementation Quick Reference available now
  3. 43 Appendix C — Complete Tensor-Shape Table available now
  4. 44 Appendix D — Experiment Configuration Cards available now
  5. 45 Appendix E — Paper Timeline and Evidence Levels available now
  6. 46 Appendix F — Glossary available now
  7. 47 Appendix G — Reproduction Checklist available now
  8. 48 Appendix H — Expert-Review Checklist available now

The big picture

  1. Slow chainEach imagined step waits for the last.
  2. Bigger stepsPrefixes or hierarchy may help.
  3. Wrong destinationMore horizon cannot fix a bad cost.
“Long horizon” hides separate problems: sequential compute, accumulated error, reusable temporal structure, unsupported subgoals, and a misleading terminal metric.

A tiny story: four travel offices

A courier must cross a country. Office A draws every road one after another. Office B asks, in parallel, “Where would the first 1, 2, 3, or 4 instructions end?” Office C plans with reusable train segments. Office D invents a magical station that is close to the goal but connected to no track.

All four advertise “long-range planning,” but they repair different things. Baseline LeWorldModel v3 rolls latent states forward autoregressively: candidate plans can run side by side, yet time inside each plan remains sequential. Later predicted states consume earlier predicted states, so compute and model error can accumulate.

One action sequence is evaluated as a self-fed chain and as several action-prefix predictions anchored at the observed state.

The fan changes the prediction interface; it does not make several stochastic futures, guarantee calibration, or remove action search.

The technical backpack

RouteExact interventionBoundary that must stay attached
Baseline v3one-step, self-fed latent rollout inside CEM/MPCcandidate parallelism does not remove time dependence
Fast-LeWMpredict several prefix horizons in parallel from one observed anchor“multiple futures” means several horizons, not multimodal outcomes; encoding, sampling, scoring, and model error remain
VLWM, Beyond the Next Stepvariable-length direct prediction with chunk schedulesnot Fast-LeWM; the reported 13% average uses training seed 3072 and post-hoc best P1/P2/P3 selection for each (dataset, Δ)
2024 hierarchical RSSM studylearned coarser temporal levels and abstract actions“HWM” is course shorthand, not the paper’s method name and not a LeWM variant
Hi-LeWMhigh-level CEM proposes latent subgoals above a frozen low-level LeWMunconstrained macro-actions and subgoals can leave training support; hierarchy is not automatically better

A useful hierarchy needs reusable chunks, a high-level model whose errors grow more slowly than the saved low-level work, subgoals the lower controller can actually reach, and enough search budget at both levels. Empirical-macro search can keep proposals nearer observed anchors, but being near data is not a certificate of reachability or safety.

Before adding any horizon, test the terminal metric. The independent TwoRoom follow-up reports that, for its audited checkpoints and protocol, changing the cost alone greatly improved a distant-goal result. That identifies a planner-interface bottleneck there; it does not prove every long-horizon failure is a cost failure.

Try to break the idea

Give a perfect low-level driver an impossible high-level station: the station looks close to the goal in latent Euclidean distance but lies outside the set of encoded macro-actions and across a wall. Increase high-level horizon and CEM population. If the planner selects the impossible station more confidently, more search has amplified support mismatch.

Now compare four locked conditions: baseline rollout, prefix prediction, hierarchy with free latent subgoals, and hierarchy constrained to empirical macro-actions. Use identical data, encoder, low-level model, action budget, candidate accounting, goals, and seeds. Log wall-clock latency, model calls, rollout error, support distance, subgoal feasibility, and closed-loop success. A speedup without matched hardware and budgets is not a portable factor; a success gain without feasible subgoals is not evidence for hierarchy.

Experiment receipt and evidence boundary

Three quick questions

  1. Which part of baseline rollout remains sequential even when candidates are vectorized?
  2. Why can a reachable-looking latent subgoal be unusable by the low-level controller?
  3. What should you diagnose before paying for a longer horizon?