JEPA4Japan · tutorials

Chapter 25 — Failure-Diagnosis Manual

706 words 4 min read #LeWorldModel#World Models#JEPA

Trace failures layer by layer from data coverage and latent collapse through ignored actions, unstable statistics, misleading costs, and a stalled planner.

Course progress Course outline 48 of 48 lessons available

Part 0 — Reading Guide: What Exactly Are We Going to Learn?

  1. 01 Chapter 0 — Before You Begin available now

Part 1 — World Models: An Internal Sandbox for the Agent

  1. 02 Chapter 1 — Why an Agent Needs to “Imagine the Future” available now
  2. 03 Chapter 2 — Why Not Predict the Next Image Directly? available now
  3. 04 Chapter 3 — The JEPA Idea: Predict Meaning, Not a Replica of the Image available now
  4. 05 Chapter 4 — Understand LeWM in One Diagram available now

Part 2 — Turning Images into State: The LeWM Architecture

  1. 06 Chapter 5 — Trajectory Data: To the Model, the World Is Not an Image Collection available now
  2. 07 Chapter 6 — The Visual Encoder: Issuing a “State Passport” for Every Frame available now
  3. 08 Chapter 7 — The Dynamics Predictor: Moving Time Forward in the Mind available now
  4. 09 Chapter 8 — A Complete Forward Pass: Follow One Batch from Start to Finish available now

Part 3 — Preventing the Model from Cheating: Prediction Loss and SIGReg

  1. 10 Chapter 9 — The Most Dangerous Shortcut: Representation Collapse available now
  2. 11 Chapter 10 — Prediction Loss: How the Model Learns the Next Step available now
  3. 12 Chapter 11 — The Intuition Behind SIGReg: Letting Representation Space “Breathe” available now
  4. 13 Chapter 12 — Keep the Mathematics Minimal but Sufficient available now
  5. 14 Chapter 13 — The Original LeWM’s End-to-End Training Mechanism available now
  6. 15 Chapter 14 — Train a Model That Does Not Collapse Immediately available now

Part 4 — Putting the Model into Action: Planning in Latent Space

  1. 16 Chapter 15 — Goal-Conditioned Planning: From “Where Am I?” to “Where Do I Want to Go?” available now
  2. 17 Chapter 16 — Latent Euclidean Distance: Convenient, but Not Necessarily Reliable available now
  3. 18 Chapter 17 — CEM: Searching for Actions Through an Elimination Tournament available now
  4. 19 Chapter 18 — MPC: Do Not Trust the Model for Too Long at Once available now
  5. 20 Chapter 19 — Long-Horizon Rollouts: How Small Errors Snowball into Major Failures available now
  6. 21 Chapter 20 — Implement a Minimal LeWM Planner from Scratch available now

Part 5 — Engineering Reproduction: From Paper to Running System

  1. 22 Chapter 21 — The Official Repository and Experimental Environment available now
  2. 23 Chapter 22 — First Experiment: A TwoRoom Smoke Test available now
  3. 24 Chapter 23 — Second Experiment: Reproducing PushT available now
  4. 25 Chapter 24 — How to Evaluate a World Model Fairly available now
  5. 26 Chapter 25 — Failure-Diagnosis Manual Current lesson

Part 6 — What Has LeWM Actually Learned?

  1. 27 Chapter 26 — Linear Probes: Which Physical Variables Are Encoded in the Latent State? available now
  2. 28 Chapter 27 — Give Latent Space a “Health Check” available now
  3. 29 Chapter 28 — Violation of Expectation: Is the Model Surprised by “Impossible Events”? available now
  4. 30 Chapter 29 — How to Discuss “Understanding the World” Rigorously available now

Part 7 — Why “Accurate Prediction” Can Still Produce “Poor Planning”

  1. 31 Chapter 30 — The Gap Between the Training Objective and the Planning Objective available now
  2. 32 Chapter 31 — Global Non-Collapse Does Not Guarantee Preservation of Task-Relevant Dynamics available now
  3. 33 Chapter 32 — When Is an Isotropic Gaussian Prior Too Strong? available now
  4. 34 Chapter 33 — Long-Horizon Planning: Predict Farther or Plan More Intelligently? available now
  5. 35 Chapter 34 — From Positional Distance to Task Progress available now
  6. 36 Chapter 35 — Multi-Task Learning, Real Robots, and Visual Distractions available now
  7. 37 Chapter 36 — Theoretical Boundaries: When Can the True State Be Identified? available now

Part 8 — From Reproducer to Researcher

  1. 38 Chapter 37 — Design a Credible LeWM Improvement Experiment available now
  2. 39 Chapter 38 — Twelve Executable Research Projects available now
  3. 40 Chapter 39 — Open Questions in LeWM Research available now

Appendices

  1. 41 Appendix A — The Minimum Necessary Mathematical Toolkit available now
  2. 42 Appendix B — PyTorch Implementation Quick Reference available now
  3. 43 Appendix C — Complete Tensor-Shape Table available now
  4. 44 Appendix D — Experiment Configuration Cards available now
  5. 45 Appendix E — Paper Timeline and Evidence Levels available now
  6. 46 Appendix F — Glossary available now
  7. 47 Appendix G — Reproduction Checklist available now
  8. 48 Appendix H — Expert-Review Checklist available now

The big picture

  1. Freeze one failureSame episode, seeds, candidates.
  2. Walk backwardExecution to data.
  3. Replace one layerMeasured or oracle counterpart.
  4. Name the smallest claimKeep alternatives open.
The same motionless agent can have many different causes.

A troubleshooting tree walks backward from environment action to valid data.

A tiny story

A doctor hears, “The patient cannot move.” That could mean muscle, nerve, medicine, pain, or a disconnected monitor. Replacing everything may remove the symptom while destroying the diagnosis.

A still LeWM agent can mean collapsed representation, action-deaf dynamics, flat cost, poor CEM samples, zeroed denormalization, or a wall blocking a valid command. Start with one frozen failing specimen, not a favorite theory.

The real rule

Walk backward from the last physical fact:

  1. Did the environment receive the intended raw action?
  2. Did denormalized actions vary and survive clipping?
  3. Did CEM’s proposal and elites change?
  4. Did candidate costs vary?
  5. Did rollouts respond correctly to action?
  6. Did representations keep the needed distinction?
  7. Were data windows, timing, and actions valid?

Log both sides of every interface. “Selected action” is incomplete unless it says normalized, clipped, repeated, block-level, or environment units. Freeze planner randomness and the candidate batch.

For action use, hold one history fixed and compare recorded, swapped, zeroed, and permuted actions against matching environment transitions. Output change is stronger evidence than a nonzero gradient, but change in the wrong direction is still wrong.

For self-fed failure, replay the same recorded actions in recorded-context and self-fed lanes. Mark the first divergence and slice by doorway, contact, and support. Do not involve CEM until this simpler comparison is interpretable.

The trick that can fool us

Two final videos show an unmoving dot.

Agent One: candidates and rollouts vary; CEM updates correctly; costs are flat because the current embedding was duplicated into the goal slot. Oracle cost separates the frozen candidates.

Agent Two: costs separate and CEM chooses a strong action; the wrong normalization column converts it to zero before environment execution.

The symptom matches; the earliest observable breaks do not.

Other discriminating tests:

  • flat cost: verify goal identity, broadcasting, terminal spread, shuffled goals, and oracle task progress;
  • coverage vs bug: cross a support atlas with known-transition replay and a tiny repeated-transition fit;
  • tiny overfit: add complexity in order—one transition, two actions from one context, a short sequence, then doorway/contact;
  • SIGReg dominance: log each loss and module gradient, spread, prediction, and a controlled weight sweep;
  • small batch: repeat statistics across batches and random projections, count distinct examples per relative time position, and vary composition.

More projection directions cannot create missing observations. Gradient accumulation does not enlarge the population that a batch-wise SIGReg test actually saw.

Experiment receipt

Archive source revisions, checkpoint hash, composed config, dataset indices, raw observations, normalized and raw actions, candidate batch, costs, and random states before repair.

The frozen SIGReg weight is 0.09; the method text says 0.1 unless stated otherwise. The PushT high-weight ablation is task-specific, not a universal safe interval. A large printed loss or gradient is not “dominance” until held-out dynamics and behavior change under controlled weights.

The diagnostic tree is a tutorial construction around the original paths in LeWorldModel v3 and frozen jepa.py, train.py, and eval.py, commit 8edfeb3. Oracle repair implicates the replaced layer under that test; it never proves every unchanged layer correct.

Quick check

  1. Why must fixed-action replay come before planner tuning?
  2. What can disprove “CEM is broken” for a motionless agent?
  3. How does an earliest observable break differ from a unique root cause?