JEPA4Japan · tutorials

Chapter 36 — Theoretical Boundaries: When Can the True State Be Identified?

828 words 4 min read #LeWorldModel#World Models#JEPA

Use assumption-bounded identifiability results to guide experiments while keeping stationarity, additive noise, linearity, observability, and multimodality explicit.

Course progress Course outline 48 of 48 lessons available

Part 0 — Reading Guide: What Exactly Are We Going to Learn?

  1. 01 Chapter 0 — Before You Begin available now

Part 1 — World Models: An Internal Sandbox for the Agent

  1. 02 Chapter 1 — Why an Agent Needs to “Imagine the Future” available now
  2. 03 Chapter 2 — Why Not Predict the Next Image Directly? available now
  3. 04 Chapter 3 — The JEPA Idea: Predict Meaning, Not a Replica of the Image available now
  4. 05 Chapter 4 — Understand LeWM in One Diagram available now

Part 2 — Turning Images into State: The LeWM Architecture

  1. 06 Chapter 5 — Trajectory Data: To the Model, the World Is Not an Image Collection available now
  2. 07 Chapter 6 — The Visual Encoder: Issuing a “State Passport” for Every Frame available now
  3. 08 Chapter 7 — The Dynamics Predictor: Moving Time Forward in the Mind available now
  4. 09 Chapter 8 — A Complete Forward Pass: Follow One Batch from Start to Finish available now

Part 3 — Preventing the Model from Cheating: Prediction Loss and SIGReg

  1. 10 Chapter 9 — The Most Dangerous Shortcut: Representation Collapse available now
  2. 11 Chapter 10 — Prediction Loss: How the Model Learns the Next Step available now
  3. 12 Chapter 11 — The Intuition Behind SIGReg: Letting Representation Space “Breathe” available now
  4. 13 Chapter 12 — Keep the Mathematics Minimal but Sufficient available now
  5. 14 Chapter 13 — The Original LeWM’s End-to-End Training Mechanism available now
  6. 15 Chapter 14 — Train a Model That Does Not Collapse Immediately available now

Part 4 — Putting the Model into Action: Planning in Latent Space

  1. 16 Chapter 15 — Goal-Conditioned Planning: From “Where Am I?” to “Where Do I Want to Go?” available now
  2. 17 Chapter 16 — Latent Euclidean Distance: Convenient, but Not Necessarily Reliable available now
  3. 18 Chapter 17 — CEM: Searching for Actions Through an Elimination Tournament available now
  4. 19 Chapter 18 — MPC: Do Not Trust the Model for Too Long at Once available now
  5. 20 Chapter 19 — Long-Horizon Rollouts: How Small Errors Snowball into Major Failures available now
  6. 21 Chapter 20 — Implement a Minimal LeWM Planner from Scratch available now

Part 5 — Engineering Reproduction: From Paper to Running System

  1. 22 Chapter 21 — The Official Repository and Experimental Environment available now
  2. 23 Chapter 22 — First Experiment: A TwoRoom Smoke Test available now
  3. 24 Chapter 23 — Second Experiment: Reproducing PushT available now
  4. 25 Chapter 24 — How to Evaluate a World Model Fairly available now
  5. 26 Chapter 25 — Failure-Diagnosis Manual available now

Part 6 — What Has LeWM Actually Learned?

  1. 27 Chapter 26 — Linear Probes: Which Physical Variables Are Encoded in the Latent State? available now
  2. 28 Chapter 27 — Give Latent Space a “Health Check” available now
  3. 29 Chapter 28 — Violation of Expectation: Is the Model Surprised by “Impossible Events”? available now
  4. 30 Chapter 29 — How to Discuss “Understanding the World” Rigorously available now

Part 7 — Why “Accurate Prediction” Can Still Produce “Poor Planning”

  1. 31 Chapter 30 — The Gap Between the Training Objective and the Planning Objective available now
  2. 32 Chapter 31 — Global Non-Collapse Does Not Guarantee Preservation of Task-Relevant Dynamics available now
  3. 33 Chapter 32 — When Is an Isotropic Gaussian Prior Too Strong? available now
  4. 34 Chapter 33 — Long-Horizon Planning: Predict Farther or Plan More Intelligently? available now
  5. 35 Chapter 34 — From Positional Distance to Task Progress available now
  6. 36 Chapter 35 — Multi-Task Learning, Real Robots, and Visual Distractions available now
  7. 37 Chapter 36 — Theoretical Boundaries: When Can the True State Be Identified? Current lesson

Part 8 — From Reproducer to Researcher

  1. 38 Chapter 37 — Design a Credible LeWM Improvement Experiment available now
  2. 39 Chapter 38 — Twelve Executable Research Projects available now
  3. 40 Chapter 39 — Open Questions in LeWM Research available now

Appendices

  1. 41 Appendix A — The Minimum Necessary Mathematical Toolkit available now
  2. 42 Appendix B — PyTorch Implementation Quick Reference available now
  3. 43 Appendix C — Complete Tensor-Shape Table available now
  4. 44 Appendix D — Experiment Configuration Cards available now
  5. 45 Appendix E — Paper Timeline and Evidence Levels available now
  6. 46 Appendix F — Glossary available now
  7. 47 Appendix G — Reproduction Checklist available now
  8. 48 Appendix H — Expert-Review Checklist available now

The big picture

  1. Inside assumptionsA theorem gives a precise promise.
  2. Rotated answerThe world may be recovered up to orthogonal coordinates.
  3. Outside the shoreThe theorem is silent, not magically false or true.
Identifiability is an if–then contract. A probe, a good loss, or a benchmark score cannot erase the “if.”

A tiny story: two honest maps

Two cartographers draw the same island. One turns the paper 90 degrees. Every distance and transition still agrees, but “up” no longer means north. Both maps can be correct.

That is why “identified up to an orthogonal transformation” is meaningful but limited. The learned coordinates are linearly related to the hidden coordinates through a rotation or reflection; individual axes need not receive human names. This is stronger than arbitrary nonlinear warping and weaker than recovering “x position” in coordinate 1.

An island marks the linear, Gaussian, stationary, fully observed, exact-constraint, matched-dimension population setting; practical cases sit outside its shoreline.

Outside means “not directly covered by these theorems,” not “impossible to learn.”

The technical backpack

The passive result in When Does LeJEPA Learn a World Model? is a population/global-optimum statement for a specified class: observable state through the stipulated observation map, stationary additive-noise dynamics, Gaussian latent structure, exact Gaussian representation constraint, and matched latent/representation dimension. Under its separation conditions it gives linear recovery up to orthogonal coordinates. Its planning statement additionally needs transformed dynamics to agree and a cost that respects the orthogonal ambiguity. The encoder theorem does not by itself show that an action-conditioned transition was learned.

The controlled v2 paper has a different, stronger contract: invertible observation, stationary linear controlled latent dynamics with independent Gaussian noise, jointly Gaussian state/action behavior, exact Gaussian representation, matched dimension, a sufficiently expressive continuous predictor, population optimization, and positive state-conditioned action excitation. Under its conditions, it identifies the state and controlled conditional-mean transition up to a common orthogonal change. It does not recover every stochastic future.

Three common assumption gaps are different:

  • Partial observability: two hidden states can make the same current image. History helps only when the hidden variable leaves a trace.
  • Multimodal futures: squared-error point prediction may land between real modes; a conditional mean is not a full future distribution.
  • Missing interventions: seeing Left and Right overall is insufficient if each action occurs in only one state. Counterfactual action branches remain unconstrained.

SIGReg in finite networks pressures finite minibatches, sampled projections, and numerical knots toward a Gaussian target. That is not exact population enforcement. The separate free-energy paper assumes constant encoder noise and successful exact isotropic-Gaussian enforcement; it is not an identifiability theorem and leaves empirical validation for later work.

Try to break the idea

Build two states. In State A, the behavior policy always chooses Right; in State B, it always chooses Left. The whole dataset contains both actions equally often, yet neither state contains both branches. A predictor can fit every recorded transition and invent anything for A+Left or B+Right.

Add a second dataset with both actions at both states while keeping observation, state distribution, model, and updates fixed. Include an oracle-state condition. If withheld-action error remains with oracle observations, the issue is transition coverage in this construction, not representation ambiguity. This test supports a conditional-excitation mechanism; it does not diagnose every LeWM failure.

Before citing a theorem, ask: Is observation invertible enough? Does the transition/noise family match? Are stationarity, dimension, Gaussian constraint, and global optimum actually available? After conditioning on state, does every needed action direction vary? Unknown answers mean “use as guidance,” not “claim coverage.”

Experiment receipt and evidence boundary

Three quick questions

  1. What does orthogonal ambiguity preserve, and what semantic names can it change?
  2. Why is marginal action diversity weaker than state-conditioned action excitation?
  3. What is the correct conclusion when a benchmark lies outside a theorem’s assumptions?