JEPA4Japan · tutorials

Chapter 1 — Why an Agent Needs to “Imagine the Future”

655 words 3 min read #LeWorldModel#World Models#JEPA

Connect future prediction to action, contrast reactive policies with world models, and locate failures in perception, imagination, and action selection.

Course progress Course outline 48 of 48 lessons available

Part 0 — Reading Guide: What Exactly Are We Going to Learn?

  1. 01 Chapter 0 — Before You Begin available now

Part 1 — World Models: An Internal Sandbox for the Agent

  1. 02 Chapter 1 — Why an Agent Needs to “Imagine the Future” Current lesson
  2. 03 Chapter 2 — Why Not Predict the Next Image Directly? available now
  3. 04 Chapter 3 — The JEPA Idea: Predict Meaning, Not a Replica of the Image available now
  4. 05 Chapter 4 — Understand LeWM in One Diagram available now

Part 2 — Turning Images into State: The LeWM Architecture

  1. 06 Chapter 5 — Trajectory Data: To the Model, the World Is Not an Image Collection available now
  2. 07 Chapter 6 — The Visual Encoder: Issuing a “State Passport” for Every Frame available now
  3. 08 Chapter 7 — The Dynamics Predictor: Moving Time Forward in the Mind available now
  4. 09 Chapter 8 — A Complete Forward Pass: Follow One Batch from Start to Finish available now

Part 3 — Preventing the Model from Cheating: Prediction Loss and SIGReg

  1. 10 Chapter 9 — The Most Dangerous Shortcut: Representation Collapse available now
  2. 11 Chapter 10 — Prediction Loss: How the Model Learns the Next Step available now
  3. 12 Chapter 11 — The Intuition Behind SIGReg: Letting Representation Space “Breathe” available now
  4. 13 Chapter 12 — Keep the Mathematics Minimal but Sufficient available now
  5. 14 Chapter 13 — The Original LeWM’s End-to-End Training Mechanism available now
  6. 15 Chapter 14 — Train a Model That Does Not Collapse Immediately available now

Part 4 — Putting the Model into Action: Planning in Latent Space

  1. 16 Chapter 15 — Goal-Conditioned Planning: From “Where Am I?” to “Where Do I Want to Go?” available now
  2. 17 Chapter 16 — Latent Euclidean Distance: Convenient, but Not Necessarily Reliable available now
  3. 18 Chapter 17 — CEM: Searching for Actions Through an Elimination Tournament available now
  4. 19 Chapter 18 — MPC: Do Not Trust the Model for Too Long at Once available now
  5. 20 Chapter 19 — Long-Horizon Rollouts: How Small Errors Snowball into Major Failures available now
  6. 21 Chapter 20 — Implement a Minimal LeWM Planner from Scratch available now

Part 5 — Engineering Reproduction: From Paper to Running System

  1. 22 Chapter 21 — The Official Repository and Experimental Environment available now
  2. 23 Chapter 22 — First Experiment: A TwoRoom Smoke Test available now
  3. 24 Chapter 23 — Second Experiment: Reproducing PushT available now
  4. 25 Chapter 24 — How to Evaluate a World Model Fairly available now
  5. 26 Chapter 25 — Failure-Diagnosis Manual available now

Part 6 — What Has LeWM Actually Learned?

  1. 27 Chapter 26 — Linear Probes: Which Physical Variables Are Encoded in the Latent State? available now
  2. 28 Chapter 27 — Give Latent Space a “Health Check” available now
  3. 29 Chapter 28 — Violation of Expectation: Is the Model Surprised by “Impossible Events”? available now
  4. 30 Chapter 29 — How to Discuss “Understanding the World” Rigorously available now

Part 7 — Why “Accurate Prediction” Can Still Produce “Poor Planning”

  1. 31 Chapter 30 — The Gap Between the Training Objective and the Planning Objective available now
  2. 32 Chapter 31 — Global Non-Collapse Does Not Guarantee Preservation of Task-Relevant Dynamics available now
  3. 33 Chapter 32 — When Is an Isotropic Gaussian Prior Too Strong? available now
  4. 34 Chapter 33 — Long-Horizon Planning: Predict Farther or Plan More Intelligently? available now
  5. 35 Chapter 34 — From Positional Distance to Task Progress available now
  6. 36 Chapter 35 — Multi-Task Learning, Real Robots, and Visual Distractions available now
  7. 37 Chapter 36 — Theoretical Boundaries: When Can the True State Be Identified? available now

Part 8 — From Reproducer to Researcher

  1. 38 Chapter 37 — Design a Credible LeWM Improvement Experiment available now
  2. 39 Chapter 38 — Twelve Executable Research Projects available now
  3. 40 Chapter 39 — Open Questions in LeWM Research available now

Appendices

  1. 41 Appendix A — The Minimum Necessary Mathematical Toolkit available now
  2. 42 Appendix B — PyTorch Implementation Quick Reference available now
  3. 43 Appendix C — Complete Tensor-Shape Table available now
  4. 44 Appendix D — Experiment Configuration Cards available now
  5. 45 Appendix E — Paper Timeline and Evidence Levels available now
  6. 46 Appendix F — Glossary available now
  7. 47 Appendix G — Reproduction Checklist available now
  8. 48 Appendix H — Expert-Review Checklist available now

The big picture

  1. Seewall + doorway
  2. Branchtry several actions
  3. Comparewhich future helps?
  4. Replanlook again
Recognition describes now. Action requires comparing futures—including a detour that first looks worse.

A short line hits a wall while a longer route reaches the goal through a doorway.

A tiny story

A robot sees the wall, doorway, and goal perfectly. It still drives into the wall because it always chooses the move that looks closest to the goal right now. The useful route first moves away, reaches the doorway, crosses, and turns back.

Seeing nouns is not enough. Acting needs a purpose, a model of consequences, and time. A direct policy can still be powerful and use history; the distinction is that a world-model planner explicitly rolls out candidate futures and searches them at decision time.

The real rule

In this course, “imagination” means a conditional computation, not a human mental movie:

current latent history + proposed action -> predicted next latent

LeWM’s encoder describes the current situation in latent space. Its predictor asks, “What follows if I apply this action?” A useful decision simulator needs three things:

  • Action sensitivity: consequentially different actions can change the prediction.
  • Rollout: a predicted state can be advanced again.
  • A decision interface: a cost can compare imagined outcomes with a goal.

The core model predicts latent states, not future videos. The paper’s auxiliary decoder is for inspection, not the training target or planner. One deterministic prediction can also hide several plausible futures; the baseline interface supplies no calibrated uncertainty distribution.

MPC makes a plan provisional: propose sequences, imagine endpoints, score them, execute a configured prefix, then observe again. In the reported LeWM v3 setup, the horizon is five model steps and all five are executed before replanning. With five environment actions per model step, feedback arrives after 25 environment steps. A one-action loop is valid MPC, but it is not that reported timing.

The baseline endpoint score is squared Euclidean distance to the encoded goal. That is a convenient latent ruler, not reachability, path length, or travel time.

The trick that fools us

Remove look-ahead and a greedy agent hugs the wall because every immediate move toward the goal seems good. Remove feedback and an initially sensible route can become stale after an observable block appears. Add MPC and the agent may reroute—but only if the observation, representation, model, cost, and search are useful.

The same collision can begin in different places:

  1. Perception: the latent omits the wall or motion.
  2. Imagination: the predictor passes through the wall or drifts.
  3. Selection: the cost prefers the wrong endpoint or CEM misses the door.

Replanning cannot repair a systematic model that redraws the same fake shortcut every cycle. A final failure does not identify its cause.

Evidence receipt

LeWorldModel v3 and the frozen official repository support the pixel encoder, action-conditioned latent predictor, autoregressive rollout, terminal latent cost, CEM search, and configurable MPC prefix. The repository delegates parts of planning and evaluation to stable-worldmodel; the exact solver convention must be recorded rather than guessed.

The TwoRoom detour, blocked-door story, and three-part diagnosis are teaching constructions. They show why look-ahead and feedback can matter; they do not prove that every direct policy fails or every model-based controller succeeds.

Quick check

  1. Why can moving away from the goal be the correct first action?
  2. What action-swap result would warn that the predictor ignores control?
  3. Which error can MPC keep repeating even after every new observation?