JEPA4Japan · tutorials

Chapter 15 — Goal-Conditioned Planning: From “Where Am I?” to “Where Do I Want to Go?”

761 words 4 min read #LeWorldModel#World Models#JEPA

Encode current and goal images together, roll out candidate action sequences, and select plans by terminal latent cost without requiring rewards during model training.

Course progress Course outline 48 of 48 lessons available

Part 0 — Reading Guide: What Exactly Are We Going to Learn?

  1. 01 Chapter 0 — Before You Begin available now

Part 1 — World Models: An Internal Sandbox for the Agent

  1. 02 Chapter 1 — Why an Agent Needs to “Imagine the Future” available now
  2. 03 Chapter 2 — Why Not Predict the Next Image Directly? available now
  3. 04 Chapter 3 — The JEPA Idea: Predict Meaning, Not a Replica of the Image available now
  4. 05 Chapter 4 — Understand LeWM in One Diagram available now

Part 2 — Turning Images into State: The LeWM Architecture

  1. 06 Chapter 5 — Trajectory Data: To the Model, the World Is Not an Image Collection available now
  2. 07 Chapter 6 — The Visual Encoder: Issuing a “State Passport” for Every Frame available now
  3. 08 Chapter 7 — The Dynamics Predictor: Moving Time Forward in the Mind available now
  4. 09 Chapter 8 — A Complete Forward Pass: Follow One Batch from Start to Finish available now

Part 3 — Preventing the Model from Cheating: Prediction Loss and SIGReg

  1. 10 Chapter 9 — The Most Dangerous Shortcut: Representation Collapse available now
  2. 11 Chapter 10 — Prediction Loss: How the Model Learns the Next Step available now
  3. 12 Chapter 11 — The Intuition Behind SIGReg: Letting Representation Space “Breathe” available now
  4. 13 Chapter 12 — Keep the Mathematics Minimal but Sufficient available now
  5. 14 Chapter 13 — The Original LeWM’s End-to-End Training Mechanism available now
  6. 15 Chapter 14 — Train a Model That Does Not Collapse Immediately available now

Part 4 — Putting the Model into Action: Planning in Latent Space

  1. 16 Chapter 15 — Goal-Conditioned Planning: From “Where Am I?” to “Where Do I Want to Go?” Current lesson
  2. 17 Chapter 16 — Latent Euclidean Distance: Convenient, but Not Necessarily Reliable available now
  3. 18 Chapter 17 — CEM: Searching for Actions Through an Elimination Tournament available now
  4. 19 Chapter 18 — MPC: Do Not Trust the Model for Too Long at Once available now
  5. 20 Chapter 19 — Long-Horizon Rollouts: How Small Errors Snowball into Major Failures available now
  6. 21 Chapter 20 — Implement a Minimal LeWM Planner from Scratch available now

Part 5 — Engineering Reproduction: From Paper to Running System

  1. 22 Chapter 21 — The Official Repository and Experimental Environment available now
  2. 23 Chapter 22 — First Experiment: A TwoRoom Smoke Test available now
  3. 24 Chapter 23 — Second Experiment: Reproducing PushT available now
  4. 25 Chapter 24 — How to Evaluate a World Model Fairly available now
  5. 26 Chapter 25 — Failure-Diagnosis Manual available now

Part 6 — What Has LeWM Actually Learned?

  1. 27 Chapter 26 — Linear Probes: Which Physical Variables Are Encoded in the Latent State? available now
  2. 28 Chapter 27 — Give Latent Space a “Health Check” available now
  3. 29 Chapter 28 — Violation of Expectation: Is the Model Surprised by “Impossible Events”? available now
  4. 30 Chapter 29 — How to Discuss “Understanding the World” Rigorously available now

Part 7 — Why “Accurate Prediction” Can Still Produce “Poor Planning”

  1. 31 Chapter 30 — The Gap Between the Training Objective and the Planning Objective available now
  2. 32 Chapter 31 — Global Non-Collapse Does Not Guarantee Preservation of Task-Relevant Dynamics available now
  3. 33 Chapter 32 — When Is an Isotropic Gaussian Prior Too Strong? available now
  4. 34 Chapter 33 — Long-Horizon Planning: Predict Farther or Plan More Intelligently? available now
  5. 35 Chapter 34 — From Positional Distance to Task Progress available now
  6. 36 Chapter 35 — Multi-Task Learning, Real Robots, and Visual Distractions available now
  7. 37 Chapter 36 — Theoretical Boundaries: When Can the True State Be Identified? available now

Part 8 — From Reproducer to Researcher

  1. 38 Chapter 37 — Design a Credible LeWM Improvement Experiment available now
  2. 39 Chapter 38 — Twelve Executable Research Projects available now
  3. 40 Chapter 39 — Open Questions in LeWM Research available now

Appendices

  1. 41 Appendix A — The Minimum Necessary Mathematical Toolkit available now
  2. 42 Appendix B — PyTorch Implementation Quick Reference available now
  3. 43 Appendix C — Complete Tensor-Shape Table available now
  4. 44 Appendix D — Experiment Configuration Cards available now
  5. 45 Appendix E — Paper Timeline and Evidence Levels available now
  6. 46 Appendix F — Glossary available now
  7. 47 Appendix G — Reproduction Checklist available now
  8. 48 Appendix H — Expert-Review Checklist available now

The big picture

  1. Currentstart the rollout
  2. Actionsmake candidate futures
  3. Goalwait at the finish line
  4. RankCEM refines plans
The goal is a planning-time postcard. It ranks imagined endpoints; it does not train or steer the baseline dynamics.

A sealed world model receives a current observation and a separate destination postcard at planning time.

A tiny story

The world model has finished learning an atlas from reward-free observation–action journeys. A traveler now arrives with a current photograph and a destination postcard. The atlas can predict consequences, but prediction alone has no preference. The postcard supplies one.

The latent atlas may still distort walls, direction, time, and data support. A goal image is not a success detector, route, reward used in training, or proof that the destination is reachable.

The technical backpack

Current and goal observations use the same frozen visual encoder, so their coordinates can be compared. Their roles remain different:

current history -> initial condition for dynamics
candidate actions -> controls for dynamics
goal image -> terminal comparison only

Swapping the goal while keeping current history and candidate actions fixed must not change baseline rollouts. It should change only their ranking. Goal preprocessing, camera, environment, and support still need to match the checkpoint’s input contract. A shared encoder cannot repair mismatched crops or normalization.

CEM samples complete action sequences over horizon H. Every candidate begins from the same current latent history and receives its own actions. The frozen predictor rolls each sequence autoregressively. The baseline emits latent tokens, not videos, and supplies one deterministic path without calibrated confidence.

For every candidate, the frozen criterion takes the final predicted latent and computes:

J(a0:H-1) = sum_d (z_hat_H[d] - z_goal[d])^2

Only the terminal token is scored. There is no explicit wall, collision, intermediate safety, control effort, route length, or support penalty. Squared Euclidean distance is cheap and parameter-free after representation learning; it is not reachability by definition.

CEM uses low-cost samples as elites and refits its action proposal. “Best” always means under this frozen model, terminal cost, finite samples, and recorded solver convention. Appendix B prose describes returning the final proposal mean; Algorithm 2 also permits the best sample or the first action of that mean. The LeWM repository’s external solver revision is unpinned, so exact reproduction must record it.

Reward-free training and goal-directed action are therefore compatible: dynamics are learned without a task reward, then an inference-time objective tells search which imagined future to prefer.

The trick that fools us

Roll one fixed candidate population once. Then swap among three supported goal images. The terminal latents must stay identical while costs reorder. If rollouts change, the goal leaked into dynamics or randomness was not held fixed.

Now use an impossible or support-poor goal. Every candidate still receives a number, and CEM still returns a plan unless an explicit abstention rule is added. Numerical ranking is not feasibility.

Protocol details can dominate results. LeWM v1/v2 used TwoRoom goal offset/budget 100/150; current v3 revised these to 25/50, matching the repository on those fields. One independent audit reported 14% versus 84% on the same released weights under the two protocol readings. In a paired comparison, changing to each episode’s recorded target with a 150-step budget moved 84% to 8%; because both goal construction and budget changed, that is not a one-variable causal result.

Evidence receipt

LeWorldModel v3, frozen jepa.py, and eval.py support frozen goal-image planning and terminal latent cost.

The same-system Research Frontier follow-up arXiv:2608.12959v1 reported that terminal latent L2 saturated and then decreased with true separation at long range in one TwoRoom evidence chain, and that alternative objectives changed behavior. It does not add a reachability head to original LeWM or establish a universal repair.

Quick check

  1. Which two inputs enter dynamics, and which input enters only the cost?
  2. Why can an unreachable goal still produce a CEM plan?
  3. What should remain unchanged when only the goal postcard is swapped?