JEPA4Japan · tutorials

Chapter 2 — Why Not Predict the Next Image Directly?

694 words 4 min read #LeWorldModel#World Models#JEPA

See why pixel error can miss task-relevant structure and why useful latent states must balance compression with predictability.

Course progress Course outline 48 of 48 lessons available

Part 0 — Reading Guide: What Exactly Are We Going to Learn?

  1. 01 Chapter 0 — Before You Begin available now

Part 1 — World Models: An Internal Sandbox for the Agent

  1. 02 Chapter 1 — Why an Agent Needs to “Imagine the Future” available now
  2. 03 Chapter 2 — Why Not Predict the Next Image Directly? Current lesson
  3. 04 Chapter 3 — The JEPA Idea: Predict Meaning, Not a Replica of the Image available now
  4. 05 Chapter 4 — Understand LeWM in One Diagram available now

Part 2 — Turning Images into State: The LeWM Architecture

  1. 06 Chapter 5 — Trajectory Data: To the Model, the World Is Not an Image Collection available now
  2. 07 Chapter 6 — The Visual Encoder: Issuing a “State Passport” for Every Frame available now
  3. 08 Chapter 7 — The Dynamics Predictor: Moving Time Forward in the Mind available now
  4. 09 Chapter 8 — A Complete Forward Pass: Follow One Batch from Start to Finish available now

Part 3 — Preventing the Model from Cheating: Prediction Loss and SIGReg

  1. 10 Chapter 9 — The Most Dangerous Shortcut: Representation Collapse available now
  2. 11 Chapter 10 — Prediction Loss: How the Model Learns the Next Step available now
  3. 12 Chapter 11 — The Intuition Behind SIGReg: Letting Representation Space “Breathe” available now
  4. 13 Chapter 12 — Keep the Mathematics Minimal but Sufficient available now
  5. 14 Chapter 13 — The Original LeWM’s End-to-End Training Mechanism available now
  6. 15 Chapter 14 — Train a Model That Does Not Collapse Immediately available now

Part 4 — Putting the Model into Action: Planning in Latent Space

  1. 16 Chapter 15 — Goal-Conditioned Planning: From “Where Am I?” to “Where Do I Want to Go?” available now
  2. 17 Chapter 16 — Latent Euclidean Distance: Convenient, but Not Necessarily Reliable available now
  3. 18 Chapter 17 — CEM: Searching for Actions Through an Elimination Tournament available now
  4. 19 Chapter 18 — MPC: Do Not Trust the Model for Too Long at Once available now
  5. 20 Chapter 19 — Long-Horizon Rollouts: How Small Errors Snowball into Major Failures available now
  6. 21 Chapter 20 — Implement a Minimal LeWM Planner from Scratch available now

Part 5 — Engineering Reproduction: From Paper to Running System

  1. 22 Chapter 21 — The Official Repository and Experimental Environment available now
  2. 23 Chapter 22 — First Experiment: A TwoRoom Smoke Test available now
  3. 24 Chapter 23 — Second Experiment: Reproducing PushT available now
  4. 25 Chapter 24 — How to Evaluate a World Model Fairly available now
  5. 26 Chapter 25 — Failure-Diagnosis Manual available now

Part 6 — What Has LeWM Actually Learned?

  1. 27 Chapter 26 — Linear Probes: Which Physical Variables Are Encoded in the Latent State? available now
  2. 28 Chapter 27 — Give Latent Space a “Health Check” available now
  3. 29 Chapter 28 — Violation of Expectation: Is the Model Surprised by “Impossible Events”? available now
  4. 30 Chapter 29 — How to Discuss “Understanding the World” Rigorously available now

Part 7 — Why “Accurate Prediction” Can Still Produce “Poor Planning”

  1. 31 Chapter 30 — The Gap Between the Training Objective and the Planning Objective available now
  2. 32 Chapter 31 — Global Non-Collapse Does Not Guarantee Preservation of Task-Relevant Dynamics available now
  3. 33 Chapter 32 — When Is an Isotropic Gaussian Prior Too Strong? available now
  4. 34 Chapter 33 — Long-Horizon Planning: Predict Farther or Plan More Intelligently? available now
  5. 35 Chapter 34 — From Positional Distance to Task Progress available now
  6. 36 Chapter 35 — Multi-Task Learning, Real Robots, and Visual Distractions available now
  7. 37 Chapter 36 — Theoretical Boundaries: When Can the True State Be Identified? available now

Part 8 — From Reproducer to Researcher

  1. 38 Chapter 37 — Design a Credible LeWM Improvement Experiment available now
  2. 39 Chapter 38 — Twelve Executable Research Projects available now
  3. 40 Chapter 39 — Open Questions in LeWM Research available now

Appendices

  1. 41 Appendix A — The Minimum Necessary Mathematical Toolkit available now
  2. 42 Appendix B — PyTorch Implementation Quick Reference available now
  3. 43 Appendix C — Complete Tensor-Shape Table available now
  4. 44 Appendix D — Experiment Configuration Cards available now
  5. 45 Appendix E — Paper Timeline and Evidence Levels available now
  6. 46 Appendix F — Glossary available now
  7. 47 Appendix G — Reproduction Checklist available now
  8. 48 Appendix H — Expert-Review Checklist available now

The big picture

  1. Many pixelstexture, light, noise
  2. One tiny cluedoor open or shut?
  3. Make a codekeep useful differences
  4. Test itcompact is not correct
A big pixel error can be harmless; a tiny doorway error can reverse the right action.

One forecast gets harmless textures wrong while another makes a tiny but crucial doorway error.

A tiny story

Forecast A gets the floor texture and shadows wrong, but keeps the agent, wall, and doorway in the right places. Forecast B looks almost perfect, except for a thin strip that closes the doorway. A pixel score may prefer B. A controller should not.

Pixels are not useless. A small red light, cable, wet patch, or contact edge can matter enormously. The lesson is narrower: the place where we measure error decides which differences training pressures the model to preserve.

The real rule

LeWM does not generate the next image for its core objective. It encodes the recorded next observation, predicts that observation’s code from earlier codes and actions, and compares the two codes.

observations: [B,T,C,H,W] -> codes: [B,T,D]
used actions: [B,T-1,A]
prediction and target: [B,T-1,D] <-> [B,T-1,D]

This creates room for selective abstraction. Different lighting can share a useful code if it does not change consequences. But the word “latent” guarantees nothing. A learned code may keep wallpaper and discard the doorway. The core rollout is a chain of vectors, not a hidden video; rendering requires a separate diagnostic decoder.

A control-useful representation should preserve situation differences, action effects, temporal clues, and a comparison useful to the planner. It should also be treated cautiously outside the state–action support of the offline data. Those are desiderata to test, not properties promised by compression.

The frozen path uses one shared trainable encoder. The encoded next observation is connected to it—no EMA teacher and no stop-gradient. SIGReg later opposes the easiest constant-code shortcut, but SIGReg does not label physics or reachability.

The trick that fools us

Turn a “compression dial.” At one end is the full image with every nuisance. At the other is one constant vector for every image. The constant is tiny and perfectly predictable: the predictor can always return it. It is also useless.

A good diagnostic changes one factor at a time. Keep geometry fixed while changing texture and lighting; then keep appearance similar while opening or closing the doorway. The desired pattern is invariance to the declared nuisance and sensitivity to the route-changing cue. Reverse the task—for example, make floor texture indicate friction—to prove that “nuisance” was conditional, not an intrinsic property of pixels.

A pretty cluster plot is not enough. Pair it with held-out prediction, action swaps, and planning outcomes.

Evidence receipt

LeWorldModel v3 and frozen train.py support compact next-embedding prediction, connected targets, SIGReg, and an auxiliary decoder outside the core objective.

A narrow comparison in I-JEPA v3 reported 66.9 top-1 with representation targets versus 40.7 with pixel targets for ViT-L/16 on ImageNet-1K with a 1%-label linear evaluation. The pretraining schedules differed—500 versus 800 epochs. That is a named static-image result, not proof that latent prediction always beats pixel prediction in control.

Permitted claim: latent prediction allows, but does not guarantee, selective abstraction. Do not claim that pixel prediction is inherently inferior, the smallest code is best, or latent error equals task error.

Quick check

  1. How can many wrong pixels matter less than one wrong door edge?
  2. Why is a constant code easy to predict but useless to control?
  3. What controlled pair would test whether the code kept doorway geometry?