JEPA4Japan · tutorials

Chapter 29 — How to Discuss “Understanding the World” Rigorously

705 words 4 min read #LeWorldModel#World Models#JEPA

Replace broad claims with tests of predictability, decodability, controllability, generalization, counterfactual consistency, causality, and planning utility.

Course progress Course outline 48 of 48 lessons available

Part 0 — Reading Guide: What Exactly Are We Going to Learn?

  1. 01 Chapter 0 — Before You Begin available now

Part 1 — World Models: An Internal Sandbox for the Agent

  1. 02 Chapter 1 — Why an Agent Needs to “Imagine the Future” available now
  2. 03 Chapter 2 — Why Not Predict the Next Image Directly? available now
  3. 04 Chapter 3 — The JEPA Idea: Predict Meaning, Not a Replica of the Image available now
  4. 05 Chapter 4 — Understand LeWM in One Diagram available now

Part 2 — Turning Images into State: The LeWM Architecture

  1. 06 Chapter 5 — Trajectory Data: To the Model, the World Is Not an Image Collection available now
  2. 07 Chapter 6 — The Visual Encoder: Issuing a “State Passport” for Every Frame available now
  3. 08 Chapter 7 — The Dynamics Predictor: Moving Time Forward in the Mind available now
  4. 09 Chapter 8 — A Complete Forward Pass: Follow One Batch from Start to Finish available now

Part 3 — Preventing the Model from Cheating: Prediction Loss and SIGReg

  1. 10 Chapter 9 — The Most Dangerous Shortcut: Representation Collapse available now
  2. 11 Chapter 10 — Prediction Loss: How the Model Learns the Next Step available now
  3. 12 Chapter 11 — The Intuition Behind SIGReg: Letting Representation Space “Breathe” available now
  4. 13 Chapter 12 — Keep the Mathematics Minimal but Sufficient available now
  5. 14 Chapter 13 — The Original LeWM’s End-to-End Training Mechanism available now
  6. 15 Chapter 14 — Train a Model That Does Not Collapse Immediately available now

Part 4 — Putting the Model into Action: Planning in Latent Space

  1. 16 Chapter 15 — Goal-Conditioned Planning: From “Where Am I?” to “Where Do I Want to Go?” available now
  2. 17 Chapter 16 — Latent Euclidean Distance: Convenient, but Not Necessarily Reliable available now
  3. 18 Chapter 17 — CEM: Searching for Actions Through an Elimination Tournament available now
  4. 19 Chapter 18 — MPC: Do Not Trust the Model for Too Long at Once available now
  5. 20 Chapter 19 — Long-Horizon Rollouts: How Small Errors Snowball into Major Failures available now
  6. 21 Chapter 20 — Implement a Minimal LeWM Planner from Scratch available now

Part 5 — Engineering Reproduction: From Paper to Running System

  1. 22 Chapter 21 — The Official Repository and Experimental Environment available now
  2. 23 Chapter 22 — First Experiment: A TwoRoom Smoke Test available now
  3. 24 Chapter 23 — Second Experiment: Reproducing PushT available now
  4. 25 Chapter 24 — How to Evaluate a World Model Fairly available now
  5. 26 Chapter 25 — Failure-Diagnosis Manual available now

Part 6 — What Has LeWM Actually Learned?

  1. 27 Chapter 26 — Linear Probes: Which Physical Variables Are Encoded in the Latent State? available now
  2. 28 Chapter 27 — Give Latent Space a “Health Check” available now
  3. 29 Chapter 28 — Violation of Expectation: Is the Model Surprised by “Impossible Events”? available now
  4. 30 Chapter 29 — How to Discuss “Understanding the World” Rigorously Current lesson

Part 7 — Why “Accurate Prediction” Can Still Produce “Poor Planning”

  1. 31 Chapter 30 — The Gap Between the Training Objective and the Planning Objective available now
  2. 32 Chapter 31 — Global Non-Collapse Does Not Guarantee Preservation of Task-Relevant Dynamics available now
  3. 33 Chapter 32 — When Is an Isotropic Gaussian Prior Too Strong? available now
  4. 34 Chapter 33 — Long-Horizon Planning: Predict Farther or Plan More Intelligently? available now
  5. 35 Chapter 34 — From Positional Distance to Task Progress available now
  6. 36 Chapter 35 — Multi-Task Learning, Real Robots, and Visual Distractions available now
  7. 37 Chapter 36 — Theoretical Boundaries: When Can the True State Be Identified? available now

Part 8 — From Reproducer to Researcher

  1. 38 Chapter 37 — Design a Credible LeWM Improvement Experiment available now
  2. 39 Chapter 38 — Twelve Executable Research Projects available now
  3. 40 Chapter 39 — Open Questions in LeWM Research available now

Appendices

  1. 41 Appendix A — The Minimum Necessary Mathematical Toolkit available now
  2. 42 Appendix B — PyTorch Implementation Quick Reference available now
  3. 43 Appendix C — Complete Tensor-Shape Table available now
  4. 44 Appendix D — Experiment Configuration Cards available now
  5. 45 Appendix E — Paper Timeline and Evidence Levels available now
  6. 46 Appendix F — Glossary available now
  7. 47 Appendix G — Reproduction Checklist available now
  8. 48 Appendix H — Expert-Review Checklist available now

The big picture

  1. Name the capabilityPredict? Decode? Control?
  2. Name the testMetric, baseline, intervention.
  3. Name the conditionsData, support, budget, seeds.
  4. Name the boundaryWhat still does not follow?
Replace one overloaded noun with a falsifiable verb.

Seven witnesses bring different evidence and cannot speak for one universal “understanding” claim.

A tiny story

Seven witnesses say, “The model understands the world.” One saw low prediction error. One decoded position. One changed an action. One tested a new combination. One saw a teleport spike. One ran a causal intervention. One watched the planner succeed.

They may all report real observations. They are not reporting the same capability.

The real rule

Use seven separate coordinates:

  1. Predictability: forecasts held-out latent consequences under named actions, horizons, and support.
  2. Decodability: a specified readout recovers labels from a frozen representation.
  3. Controllability: actions steer predicted and executed outcomes under a budget.
  4. Compositional generalization: familiar parts work in an unseen combination.
  5. Counterfactual consistency: paired changes cause the right sensitivities and invariances.
  6. Causal identification: variables and relations are recovered under explicit intervention assumptions.
  7. Planning utility: the learned model and cost improve action choice under fixed task and search contracts.

These are coordinates, not a ladder. One does not automatically grant another, and several positive tests do not become proof by vote.

Random held-out episodes are not necessarily compositional. A genuine composition split holds out start–goal, object–action, layout, or phase combinations while preserving the components elsewhere. Same-trajectory goals 25 steps ahead are useful control tasks, not arbitrary-goal composition.

Causality needs named variables, valid interventions, confounder assumptions, identifying invariances, support controls, and alternative structures. A decodable direction plus a behavior-changing ablation can still be entangled or out of distribution.

The trick that can fool us

Imagine a lookup navigator storing many familiar trajectory fragments. It often succeeds on same-trajectory future goals. A probe decodes position because background and time correlate with place. A last-frame predictor spikes on teleportation.

Planning utility, decodability, and VoE sensitivity are all positive. Now randomize policy speed, change the background, and require familiar actions in a new layout. The system fails.

This does not describe LeWM. It proves a logical point: several diagnostics can share shortcuts. Triangulation is strongest when tests have different failure modes and each narrows a particular alternative.

Audit the familiar phrases:

  • “position is inside” → “position is recoverable by this readout”;
  • “the decoder shows the imagined future” → “an auxiliary decoder reconstructs a view from a predicted latent”;
  • “t-SNE shows the learned map” → “this sampled projection suggests a neighborhood pattern”;
  • “the planner understands control” → “the planner achieved a budgeted result under this protocol”;
  • “the model knows teleporting is impossible” → “mismatch increased under this perturbation.”

Experiment receipt

Rebuild every claim with five fields:

  1. capability;
  2. test, metric, baseline, and intervention;
  3. model/paper version, data, environment, support, horizon, budgets, and seeds;
  4. result and evidence level—primary method fact, author-reported result, independent reproduction, or tutorial observation;
  5. alternatives and explicit non-implications.

Use a wording ladder. Say recoverable after a probe; generalizes across the tested split after a declared shift; behaviorally implicated under this intervention after matched controls. Reserve necessary, sufficient, causal, and understands for designs and assumptions that support them.

LeWorldModel v3 supplies author-reported evidence for action-conditioned latent prediction, budgeted planning in TwoRoom, Reacher, PushT, and OGBench-Cube, selected-variable probes, qualitative latent views, PushT straightening, and sensitivity to tested teleport interventions. Frozen commit 8edfeb3 does not include complete scripts for every analysis. Together these results do not establish complete state, broad composition, calibrated futures, causal physics, real-world reliability, or human-like understanding.

Quick check

  1. Why are the seven capabilities coordinates rather than stages?
  2. What makes a split genuinely compositional?
  3. How does a five-field scientific claim improve on replacing “proves” with “suggests”?