JEPA4Japan · tutorials

Chapter 37 — Design a Credible LeWM Improvement Experiment

788 words 4 min read #LeWorldModel#World Models#JEPA

Isolate one failure mechanism with counterexamples, matched capacity and compute, ablations, oracle conditions, and honest reporting of negative results.

Course progress Course outline 48 of 48 lessons available

Part 0 — Reading Guide: What Exactly Are We Going to Learn?

  1. 01 Chapter 0 — Before You Begin available now

Part 1 — World Models: An Internal Sandbox for the Agent

  1. 02 Chapter 1 — Why an Agent Needs to “Imagine the Future” available now
  2. 03 Chapter 2 — Why Not Predict the Next Image Directly? available now
  3. 04 Chapter 3 — The JEPA Idea: Predict Meaning, Not a Replica of the Image available now
  4. 05 Chapter 4 — Understand LeWM in One Diagram available now

Part 2 — Turning Images into State: The LeWM Architecture

  1. 06 Chapter 5 — Trajectory Data: To the Model, the World Is Not an Image Collection available now
  2. 07 Chapter 6 — The Visual Encoder: Issuing a “State Passport” for Every Frame available now
  3. 08 Chapter 7 — The Dynamics Predictor: Moving Time Forward in the Mind available now
  4. 09 Chapter 8 — A Complete Forward Pass: Follow One Batch from Start to Finish available now

Part 3 — Preventing the Model from Cheating: Prediction Loss and SIGReg

  1. 10 Chapter 9 — The Most Dangerous Shortcut: Representation Collapse available now
  2. 11 Chapter 10 — Prediction Loss: How the Model Learns the Next Step available now
  3. 12 Chapter 11 — The Intuition Behind SIGReg: Letting Representation Space “Breathe” available now
  4. 13 Chapter 12 — Keep the Mathematics Minimal but Sufficient available now
  5. 14 Chapter 13 — The Original LeWM’s End-to-End Training Mechanism available now
  6. 15 Chapter 14 — Train a Model That Does Not Collapse Immediately available now

Part 4 — Putting the Model into Action: Planning in Latent Space

  1. 16 Chapter 15 — Goal-Conditioned Planning: From “Where Am I?” to “Where Do I Want to Go?” available now
  2. 17 Chapter 16 — Latent Euclidean Distance: Convenient, but Not Necessarily Reliable available now
  3. 18 Chapter 17 — CEM: Searching for Actions Through an Elimination Tournament available now
  4. 19 Chapter 18 — MPC: Do Not Trust the Model for Too Long at Once available now
  5. 20 Chapter 19 — Long-Horizon Rollouts: How Small Errors Snowball into Major Failures available now
  6. 21 Chapter 20 — Implement a Minimal LeWM Planner from Scratch available now

Part 5 — Engineering Reproduction: From Paper to Running System

  1. 22 Chapter 21 — The Official Repository and Experimental Environment available now
  2. 23 Chapter 22 — First Experiment: A TwoRoom Smoke Test available now
  3. 24 Chapter 23 — Second Experiment: Reproducing PushT available now
  4. 25 Chapter 24 — How to Evaluate a World Model Fairly available now
  5. 26 Chapter 25 — Failure-Diagnosis Manual available now

Part 6 — What Has LeWM Actually Learned?

  1. 27 Chapter 26 — Linear Probes: Which Physical Variables Are Encoded in the Latent State? available now
  2. 28 Chapter 27 — Give Latent Space a “Health Check” available now
  3. 29 Chapter 28 — Violation of Expectation: Is the Model Surprised by “Impossible Events”? available now
  4. 30 Chapter 29 — How to Discuss “Understanding the World” Rigorously available now

Part 7 — Why “Accurate Prediction” Can Still Produce “Poor Planning”

  1. 31 Chapter 30 — The Gap Between the Training Objective and the Planning Objective available now
  2. 32 Chapter 31 — Global Non-Collapse Does Not Guarantee Preservation of Task-Relevant Dynamics available now
  3. 33 Chapter 32 — When Is an Isotropic Gaussian Prior Too Strong? available now
  4. 34 Chapter 33 — Long-Horizon Planning: Predict Farther or Plan More Intelligently? available now
  5. 35 Chapter 34 — From Positional Distance to Task Progress available now
  6. 36 Chapter 35 — Multi-Task Learning, Real Robots, and Visual Distractions available now
  7. 37 Chapter 36 — Theoretical Boundaries: When Can the True State Be Identified? available now

Part 8 — From Reproducer to Researcher

  1. 38 Chapter 37 — Design a Credible LeWM Improvement Experiment Current lesson
  2. 39 Chapter 38 — Twelve Executable Research Projects available now
  3. 40 Chapter 39 — Open Questions in LeWM Research available now

Appendices

  1. 41 Appendix A — The Minimum Necessary Mathematical Toolkit available now
  2. 42 Appendix B — PyTorch Implementation Quick Reference available now
  3. 43 Appendix C — Complete Tensor-Shape Table available now
  4. 44 Appendix D — Experiment Configuration Cards available now
  5. 45 Appendix E — Paper Timeline and Evidence Levels available now
  6. 46 Appendix F — Glossary available now
  7. 47 Appendix G — Reproduction Checklist available now
  8. 48 Appendix H — Expert-Review Checklist available now

The big picture

  1. Name one faultWrite the predicted failure pattern.
  2. Change one partLock data, budgets, and evaluation.
  3. Keep every resultPositive, null, and negative all count.
An improvement is credible when the intervention tests a mechanism and a falsifier could make the explanation lose.

A tiny story: the mechanic who changed everything

A car pulls left. The mechanic replaces the tires, steering, engine, road map, and driver, then the car goes straight. Something helped—but the repair taught us almost nothing.

A world-model system has the same separable parts: data and preprocessing, representation, action-conditioned dynamics, planning cost, CEM search, and execution/evaluation. If all move together, a higher success number cannot tell us which idea worked.

A LeWM control system keeps representation, dynamics, search, and execution locked while only the cost block changes.

One locked comparison localizes a claim. It does not prove the same repair works in every environment.

The technical backpack

Write the claim before running:

  1. Failure mechanism: for example, terminal Euclidean cost ranks across-wall endpoints too favorably.
  2. Predicted interaction: a reachability-aware cost should help wall-separated goals more than matched open-room goals.
  3. Falsifier: the gain stays equally large after the wall is removed, disappears with oracle dynamics, or comes only from extra CEM budget.
  4. Locked contract: dataset/hash, preprocessing, encoder, predictor, checkpoint rule, seeds, starts/goals, action budget, CEM population/elites/refinements, executed prefix, episode count, and metric.

Use an ablation ladder: baseline → add the proposed component → remove it → shuffle or randomize its signal → sweep its strength. Match both parameter count and compute where possible. Training budget and planning budget are separate currencies; extra epochs cannot silently pay for fewer candidate plans, and extra CEM candidates cannot be called a representation gain.

Oracle substitutions localize the bottleneck:

DynamicsCostWhat the cell asks
learnedlearneddoes the full system work?
oraclelearneddoes failure remain when rollout error is removed?
learnedoracledoes a better ranking rescue the same model?
oracleoracleis search/execution itself limiting under simulator knowledge?

Oracle state, shortest path, or dynamics are simulator-only diagnostic tools. An oracle gain shows headroom, not that the learned replacement is attainable.

Try to break the idea

Use matched TwoRoom layouts: one has the separating wall and doorway; one removes the wall. Reuse the same start, goal, candidate bank, model checkpoint, and budgets. A topological-cost hypothesis predicts a larger improvement in the wall case. If the “repair” gains equally in open space, it may simply rescale scores or alter optimization.

Report distributions across seeds and goal types, not one best run. Save representative trajectories and classify failures: encoder aliasing, action-insensitive prediction, self-fed rollout drift, wrong cost ranking, CEM miss, or execution mismatch. A null result means “no supported gain under these conditions,” not “the idea can never work.” A negative result can be the strongest clue when it matches the predicted failure boundary.

Experiment receipt and evidence boundary

  • Sources that motivate diagnostics: LeWorldModel v3, RC-aux v1, TwoRoom reproduction v1, VIScore v2, ACPC v1, Objective Bottleneck v1, DA-LeWM v1, and stable-worldmodel v1.
  • Frozen public snapshots: LeWM 8edfeb3, RC-aux ecb4496, tinylab efa9e5d, VIScore bbb60fc, ACPC 90d4276, stable-worldmodel addbab4, release 0.1.1. DA-LeWM code was not identified.
  • Cited diagnostic versions span 2026-08-10 through 2026-08-19; cutoff is 2026-08-20, Asia/Tokyo. Apart from the stated TwoRoom reimplementation, source evaluations are author-reported. The experiment proposed here remains a tutorial construction until someone runs it.
  • Permitted conclusions are “isolates within this design,” “supports the mechanism under locked conditions,” or “failed to improve here.” Avoid “proves the cause,” “learns universal reachability,” “the platform guarantees fairness,” or “TwoRoom establishes robotics generalization.”

Three quick questions

  1. Why is changing representation, predictor, and planner together scientifically weak?
  2. What different bottlenecks do the four oracle cells isolate?
  3. What result in the wall-removal test would weaken the topology explanation?