JEPA4Japan · tutorials

Chapter 7 — The Dynamics Predictor: Moving Time Forward in the Mind

697 words 4 min read #LeWorldModel#World Models#JEPA

Condition latent dynamics on actions with AdaLN and causal context, then contrast teacher-forced training with autoregressive planning rollouts.

Course progress Course outline 48 of 48 lessons available

Part 0 — Reading Guide: What Exactly Are We Going to Learn?

  1. 01 Chapter 0 — Before You Begin available now

Part 1 — World Models: An Internal Sandbox for the Agent

  1. 02 Chapter 1 — Why an Agent Needs to “Imagine the Future” available now
  2. 03 Chapter 2 — Why Not Predict the Next Image Directly? available now
  3. 04 Chapter 3 — The JEPA Idea: Predict Meaning, Not a Replica of the Image available now
  4. 05 Chapter 4 — Understand LeWM in One Diagram available now

Part 2 — Turning Images into State: The LeWM Architecture

  1. 06 Chapter 5 — Trajectory Data: To the Model, the World Is Not an Image Collection available now
  2. 07 Chapter 6 — The Visual Encoder: Issuing a “State Passport” for Every Frame available now
  3. 08 Chapter 7 — The Dynamics Predictor: Moving Time Forward in the Mind Current lesson
  4. 09 Chapter 8 — A Complete Forward Pass: Follow One Batch from Start to Finish available now

Part 3 — Preventing the Model from Cheating: Prediction Loss and SIGReg

  1. 10 Chapter 9 — The Most Dangerous Shortcut: Representation Collapse available now
  2. 11 Chapter 10 — Prediction Loss: How the Model Learns the Next Step available now
  3. 12 Chapter 11 — The Intuition Behind SIGReg: Letting Representation Space “Breathe” available now
  4. 13 Chapter 12 — Keep the Mathematics Minimal but Sufficient available now
  5. 14 Chapter 13 — The Original LeWM’s End-to-End Training Mechanism available now
  6. 15 Chapter 14 — Train a Model That Does Not Collapse Immediately available now

Part 4 — Putting the Model into Action: Planning in Latent Space

  1. 16 Chapter 15 — Goal-Conditioned Planning: From “Where Am I?” to “Where Do I Want to Go?” available now
  2. 17 Chapter 16 — Latent Euclidean Distance: Convenient, but Not Necessarily Reliable available now
  3. 18 Chapter 17 — CEM: Searching for Actions Through an Elimination Tournament available now
  4. 19 Chapter 18 — MPC: Do Not Trust the Model for Too Long at Once available now
  5. 20 Chapter 19 — Long-Horizon Rollouts: How Small Errors Snowball into Major Failures available now
  6. 21 Chapter 20 — Implement a Minimal LeWM Planner from Scratch available now

Part 5 — Engineering Reproduction: From Paper to Running System

  1. 22 Chapter 21 — The Official Repository and Experimental Environment available now
  2. 23 Chapter 22 — First Experiment: A TwoRoom Smoke Test available now
  3. 24 Chapter 23 — Second Experiment: Reproducing PushT available now
  4. 25 Chapter 24 — How to Evaluate a World Model Fairly available now
  5. 26 Chapter 25 — Failure-Diagnosis Manual available now

Part 6 — What Has LeWM Actually Learned?

  1. 27 Chapter 26 — Linear Probes: Which Physical Variables Are Encoded in the Latent State? available now
  2. 28 Chapter 27 — Give Latent Space a “Health Check” available now
  3. 29 Chapter 28 — Violation of Expectation: Is the Model Surprised by “Impossible Events”? available now
  4. 30 Chapter 29 — How to Discuss “Understanding the World” Rigorously available now

Part 7 — Why “Accurate Prediction” Can Still Produce “Poor Planning”

  1. 31 Chapter 30 — The Gap Between the Training Objective and the Planning Objective available now
  2. 32 Chapter 31 — Global Non-Collapse Does Not Guarantee Preservation of Task-Relevant Dynamics available now
  3. 33 Chapter 32 — When Is an Isotropic Gaussian Prior Too Strong? available now
  4. 34 Chapter 33 — Long-Horizon Planning: Predict Farther or Plan More Intelligently? available now
  5. 35 Chapter 34 — From Positional Distance to Task Progress available now
  6. 36 Chapter 35 — Multi-Task Learning, Real Robots, and Visual Distractions available now
  7. 37 Chapter 36 — Theoretical Boundaries: When Can the True State Be Identified? available now

Part 8 — From Reproducer to Researcher

  1. 38 Chapter 37 — Design a Credible LeWM Improvement Experiment available now
  2. 39 Chapter 38 — Twelve Executable Research Projects available now
  3. 40 Chapter 39 — Open Questions in LeWM Research available now

Appendices

  1. 41 Appendix A — The Minimum Necessary Mathematical Toolkit available now
  2. 42 Appendix B — PyTorch Implementation Quick Reference available now
  3. 43 Appendix C — Complete Tensor-Shape Table available now
  4. 44 Appendix D — Experiment Configuration Cards available now
  5. 45 Appendix E — Paper Timeline and Evidence Levels available now
  6. 46 Appendix F — Glossary available now
  7. 47 Appendix G — Reproduction Checklist available now
  8. 48 Appendix H — Expert-Review Checklist available now

The big picture

  1. Same presentone latent history
  2. Action Aturn the control knob
  3. Action Bturn it differently
  4. New branchescompare next latents
The predictor adds a verb to the latent state: what follows under this proposed action?

The same TwoRoom state branches into different predicted successors under different actions.

A tiny story

Place the agent below a doorway. The current image alone does not choose tomorrow. “Left,” “right,” and “stay” can lead to different successors. The encoder supplies a compact noun—the situation. The dynamics predictor supplies a verb—the action-conditioned change.

History helps when one frame hides motion, but it cannot reveal every hidden force, contact mode, or occluded object. The output is a deterministic latent estimate, not a promise that the real future is single-valued.

The technical backpack

The conceptual interface is:

state_history  = encode(observed_frames)     # [B,T,D]
action_history = encode(action_blocks)       # [B,T,D]
next_guesses   = causal_predict(state_history, action_history)

In frozen data configs, frameskip=5: one model step contains five consecutive environment actions. The dense block is flattened and mapped to width D. The action embedder uses a width-one temporal convolution plus a small multilayer map, so neighboring time positions are not mixed there; temporal mixing happens in the causal Transformer.

The paper describes a six-layer predictor with 10% dropout. Actions enter each conditional Transformer block through AdaLN. The action pathway produces shift, scale, and residual-gate controls for both attention and feed-forward branches—six controls in total. Its final modulation layer starts at zero, so action modulation is neutral at initialization and can grow during learning.

AdaLN installs a control cable; it does not prove the trained model listens. A causal mask similarly blocks peeking at later tokens but does not prove causal physics. A future-mutation check should show that changing a later input leaves earlier outputs unchanged.

With the frozen four-observation default, training aligns all three pairs:

context:      z0  z1  z2
actions:      a0  a1  a2
predictions:  p1  p2  p3
targets:      z1  z2  z3

This is teacher forcing: context comes from recorded encoded observations. It is not a teacher encoder. During planning, the rollout appends p1, then predicts from a history that increasingly contains its own outputs. The frozen code retains only the most recent configured history window.

The trick that fools us

Mute the action channel by replacing every action block with zero while leaving data, model size, and budget matched. A behavior policy that usually takes the same route may still let the muted model predict an average continuation well.

Now choose a supported doorway context where left and right have different recorded successors. Hold the visual history fixed and swap actions. Compare predicted separation with the separation between encoded real successors. Add a free-space control where two tiny actions genuinely behave similarly.

Three metrics must stay separate: teacher-forced one-step error, action-swap sensitivity, and open-loop error by horizon. An action can affect step one yet vanish later; low one-step error can coexist with drift; a good rollout can still be ranked by a bad planner cost.

Evidence receipt

LeWorldModel v3 and frozen module.py support the causal attention and AdaLN path. Frozen jepa.py supports autoregressive append-and-truncate rollout.

Allowed claim: v3 provides an action-conditioning route, causal training context, and autoregressive planning. Do not say AdaLN proves controllability, teacher forcing implies an EMA teacher, every LeWM step means five raw actions, or low one-step error guarantees planning.

Quick check

  1. What does AdaLN make possible, and what does it not prove?
  2. Why is teacher forcing not a teacher–student architecture here?
  3. Which three diagnostics separate local prediction, action use, and rollout stability?