JEPA4Japan · tutorials

Chapter 3 — The JEPA Idea: Predict Meaning, Not a Replica of the Image

632 words 3 min read #LeWorldModel#World Models#JEPA

Understand encoders, target representations, and predictors, then compare JEPA with autoencoders, contrastive learning, and generative video models.

Course progress Course outline 48 of 48 lessons available

Part 0 — Reading Guide: What Exactly Are We Going to Learn?

  1. 01 Chapter 0 — Before You Begin available now

Part 1 — World Models: An Internal Sandbox for the Agent

  1. 02 Chapter 1 — Why an Agent Needs to “Imagine the Future” available now
  2. 03 Chapter 2 — Why Not Predict the Next Image Directly? available now
  3. 04 Chapter 3 — The JEPA Idea: Predict Meaning, Not a Replica of the Image Current lesson
  4. 05 Chapter 4 — Understand LeWM in One Diagram available now

Part 2 — Turning Images into State: The LeWM Architecture

  1. 06 Chapter 5 — Trajectory Data: To the Model, the World Is Not an Image Collection available now
  2. 07 Chapter 6 — The Visual Encoder: Issuing a “State Passport” for Every Frame available now
  3. 08 Chapter 7 — The Dynamics Predictor: Moving Time Forward in the Mind available now
  4. 09 Chapter 8 — A Complete Forward Pass: Follow One Batch from Start to Finish available now

Part 3 — Preventing the Model from Cheating: Prediction Loss and SIGReg

  1. 10 Chapter 9 — The Most Dangerous Shortcut: Representation Collapse available now
  2. 11 Chapter 10 — Prediction Loss: How the Model Learns the Next Step available now
  3. 12 Chapter 11 — The Intuition Behind SIGReg: Letting Representation Space “Breathe” available now
  4. 13 Chapter 12 — Keep the Mathematics Minimal but Sufficient available now
  5. 14 Chapter 13 — The Original LeWM’s End-to-End Training Mechanism available now
  6. 15 Chapter 14 — Train a Model That Does Not Collapse Immediately available now

Part 4 — Putting the Model into Action: Planning in Latent Space

  1. 16 Chapter 15 — Goal-Conditioned Planning: From “Where Am I?” to “Where Do I Want to Go?” available now
  2. 17 Chapter 16 — Latent Euclidean Distance: Convenient, but Not Necessarily Reliable available now
  3. 18 Chapter 17 — CEM: Searching for Actions Through an Elimination Tournament available now
  4. 19 Chapter 18 — MPC: Do Not Trust the Model for Too Long at Once available now
  5. 20 Chapter 19 — Long-Horizon Rollouts: How Small Errors Snowball into Major Failures available now
  6. 21 Chapter 20 — Implement a Minimal LeWM Planner from Scratch available now

Part 5 — Engineering Reproduction: From Paper to Running System

  1. 22 Chapter 21 — The Official Repository and Experimental Environment available now
  2. 23 Chapter 22 — First Experiment: A TwoRoom Smoke Test available now
  3. 24 Chapter 23 — Second Experiment: Reproducing PushT available now
  4. 25 Chapter 24 — How to Evaluate a World Model Fairly available now
  5. 26 Chapter 25 — Failure-Diagnosis Manual available now

Part 6 — What Has LeWM Actually Learned?

  1. 27 Chapter 26 — Linear Probes: Which Physical Variables Are Encoded in the Latent State? available now
  2. 28 Chapter 27 — Give Latent Space a “Health Check” available now
  3. 29 Chapter 28 — Violation of Expectation: Is the Model Surprised by “Impossible Events”? available now
  4. 30 Chapter 29 — How to Discuss “Understanding the World” Rigorously available now

Part 7 — Why “Accurate Prediction” Can Still Produce “Poor Planning”

  1. 31 Chapter 30 — The Gap Between the Training Objective and the Planning Objective available now
  2. 32 Chapter 31 — Global Non-Collapse Does Not Guarantee Preservation of Task-Relevant Dynamics available now
  3. 33 Chapter 32 — When Is an Isotropic Gaussian Prior Too Strong? available now
  4. 34 Chapter 33 — Long-Horizon Planning: Predict Farther or Plan More Intelligently? available now
  5. 35 Chapter 34 — From Positional Distance to Task Progress available now
  6. 36 Chapter 35 — Multi-Task Learning, Real Robots, and Visual Distractions available now
  7. 37 Chapter 36 — Theoretical Boundaries: When Can the True State Be Identified? available now

Part 8 — From Reproducer to Researcher

  1. 38 Chapter 37 — Design a Credible LeWM Improvement Experiment available now
  2. 39 Chapter 38 — Twelve Executable Research Projects available now
  3. 40 Chapter 39 — Open Questions in LeWM Research available now

Appendices

  1. 41 Appendix A — The Minimum Necessary Mathematical Toolkit available now
  2. 42 Appendix B — PyTorch Implementation Quick Reference available now
  3. 43 Appendix C — Complete Tensor-Shape Table available now
  4. 44 Appendix D — Experiment Configuration Cards available now
  5. 45 Appendix E — Paper Timeline and Evidence Levels available now
  6. 46 Appendix F — Glossary available now
  7. 47 Appendix G — Reproduction Checklist available now
  8. 48 Appendix H — Expert-Review Checklist available now

The big picture

  1. Encodemake a context code
  2. Conditionadd mask, time, or action
  3. Predictaim for a target code
  4. Name the graphfamily is not recipe
JEPA predicts in representation space. The word “target” does not automatically mean “teacher.”

Four workshops reconstruct pixels, match pairs, generate frames, or predict target representations.

A tiny story

A painter copies every brick and reflection. A mapmaker keeps roads and junctions. To predict where a traveler goes next, the mapmaker advances a point on the map instead of repainting the street.

That is the useful meaning of “predict meaning, not a replica.” But a learned map has no human cartographer assigning symbols. Coordinate 17 is not labelled “doorway.” The model may keep paint color and forget a one-way rule. “Meaning” here means a learned representation whose usefulness still needs evidence.

The real rule

The family-level JEPA skeleton is deliberately neutral:

context_code = encode_context(context)
target_code = encode_target(target)
predicted_code = predict(context_code, optional_condition)
loss = compare(predicted_code, target_code)

That sketch does not specify one or two encoders, shared or separate weights, stop-gradient, EMA, negatives, or a stochastic latent. Those choices belong to named methods.

LeWM v3 specializes the idea like this:

image-history codes + recorded actions -> next-image code prediction

Current and next frames use one shared trainable visual encoder. The action-conditioned predictor estimates the next representation. Prediction loss remains connected through both context and target uses of that encoder, and SIGReg supplies a separate population-shape pressure. Original v3 has no EMA target teacher, no target stop-gradient, and no pretrained visual encoder.

An autoencoder reconstructs an input-space target. A generative video model outputs frames or visual tokens. Contrastive objectives compare selected matches and mismatches. LeJEPA and LeWM use prediction plus SIGReg without negative pairs. These are different contracts, not a universal performance ranking; a JEPA could, in principle, be trained with other losses.

The trick that fools us

Hold one observation fixed and ask for two actions:

up = predict(current, action="up")
right = predict(current, action="right")
check distance(up, right)

If the two outputs stay almost identical where the real outcomes differ, the model may predict the usual continuation while ignoring control. CEM then evaluates many action sequences through almost the same branch. More search cannot create a choice that the predictor erased.

Nonzero separation is not yet correct dynamics. Compare direction and size with supported recorded successors, and include a control state where two small actions truly have similar effects.

Evidence receipt

The broad architecture comes from A Path Towards Autonomous Machine Intelligence v0.9.2. I-JEPA v3 is a source-specific image recipe with an EMA target encoder. LeJEPA v3 motivates prediction plus SIGReg. LeWorldModel v3 and frozen train.py are the authority for LeWM’s connected action-conditioned graph.

Do not import I-JEPA’s teacher into LeWM, call every JEPA non-contrastive by definition, or infer semantics from architecture alone. A controlled probe can support a narrow claim that one distinction is decodable, predictive, or useful for planning.

Quick check

  1. Why does “target representation” not prove a teacher network exists?
  2. Which graph features make LeWM v3 different from I-JEPA’s recipe?
  3. Why can good average prediction coexist with useless CEM search?