JEPA4Japan · tutorials

Chapter 38 — Twelve Executable Research Projects

1,129 words 6 min read #LeWorldModel#World Models#JEPA

Develop concrete studies of reachability, uncertainty, multi-step objectives, temporal hierarchy, topology, controllability, support constraints, language goals, and transfer.

Course progress Course outline 48 of 48 lessons available

Part 0 — Reading Guide: What Exactly Are We Going to Learn?

  1. 01 Chapter 0 — Before You Begin available now

Part 1 — World Models: An Internal Sandbox for the Agent

  1. 02 Chapter 1 — Why an Agent Needs to “Imagine the Future” available now
  2. 03 Chapter 2 — Why Not Predict the Next Image Directly? available now
  3. 04 Chapter 3 — The JEPA Idea: Predict Meaning, Not a Replica of the Image available now
  4. 05 Chapter 4 — Understand LeWM in One Diagram available now

Part 2 — Turning Images into State: The LeWM Architecture

  1. 06 Chapter 5 — Trajectory Data: To the Model, the World Is Not an Image Collection available now
  2. 07 Chapter 6 — The Visual Encoder: Issuing a “State Passport” for Every Frame available now
  3. 08 Chapter 7 — The Dynamics Predictor: Moving Time Forward in the Mind available now
  4. 09 Chapter 8 — A Complete Forward Pass: Follow One Batch from Start to Finish available now

Part 3 — Preventing the Model from Cheating: Prediction Loss and SIGReg

  1. 10 Chapter 9 — The Most Dangerous Shortcut: Representation Collapse available now
  2. 11 Chapter 10 — Prediction Loss: How the Model Learns the Next Step available now
  3. 12 Chapter 11 — The Intuition Behind SIGReg: Letting Representation Space “Breathe” available now
  4. 13 Chapter 12 — Keep the Mathematics Minimal but Sufficient available now
  5. 14 Chapter 13 — The Original LeWM’s End-to-End Training Mechanism available now
  6. 15 Chapter 14 — Train a Model That Does Not Collapse Immediately available now

Part 4 — Putting the Model into Action: Planning in Latent Space

  1. 16 Chapter 15 — Goal-Conditioned Planning: From “Where Am I?” to “Where Do I Want to Go?” available now
  2. 17 Chapter 16 — Latent Euclidean Distance: Convenient, but Not Necessarily Reliable available now
  3. 18 Chapter 17 — CEM: Searching for Actions Through an Elimination Tournament available now
  4. 19 Chapter 18 — MPC: Do Not Trust the Model for Too Long at Once available now
  5. 20 Chapter 19 — Long-Horizon Rollouts: How Small Errors Snowball into Major Failures available now
  6. 21 Chapter 20 — Implement a Minimal LeWM Planner from Scratch available now

Part 5 — Engineering Reproduction: From Paper to Running System

  1. 22 Chapter 21 — The Official Repository and Experimental Environment available now
  2. 23 Chapter 22 — First Experiment: A TwoRoom Smoke Test available now
  3. 24 Chapter 23 — Second Experiment: Reproducing PushT available now
  4. 25 Chapter 24 — How to Evaluate a World Model Fairly available now
  5. 26 Chapter 25 — Failure-Diagnosis Manual available now

Part 6 — What Has LeWM Actually Learned?

  1. 27 Chapter 26 — Linear Probes: Which Physical Variables Are Encoded in the Latent State? available now
  2. 28 Chapter 27 — Give Latent Space a “Health Check” available now
  3. 29 Chapter 28 — Violation of Expectation: Is the Model Surprised by “Impossible Events”? available now
  4. 30 Chapter 29 — How to Discuss “Understanding the World” Rigorously available now

Part 7 — Why “Accurate Prediction” Can Still Produce “Poor Planning”

  1. 31 Chapter 30 — The Gap Between the Training Objective and the Planning Objective available now
  2. 32 Chapter 31 — Global Non-Collapse Does Not Guarantee Preservation of Task-Relevant Dynamics available now
  3. 33 Chapter 32 — When Is an Isotropic Gaussian Prior Too Strong? available now
  4. 34 Chapter 33 — Long-Horizon Planning: Predict Farther or Plan More Intelligently? available now
  5. 35 Chapter 34 — From Positional Distance to Task Progress available now
  6. 36 Chapter 35 — Multi-Task Learning, Real Robots, and Visual Distractions available now
  7. 37 Chapter 36 — Theoretical Boundaries: When Can the True State Be Identified? available now

Part 8 — From Reproducer to Researcher

  1. 38 Chapter 37 — Design a Credible LeWM Improvement Experiment available now
  2. 39 Chapter 38 — Twelve Executable Research Projects Current lesson
  3. 40 Chapter 39 — Open Questions in LeWM Research available now

Appendices

  1. 41 Appendix A — The Minimum Necessary Mathematical Toolkit available now
  2. 42 Appendix B — PyTorch Implementation Quick Reference available now
  3. 43 Appendix C — Complete Tensor-Shape Table available now
  4. 44 Appendix D — Experiment Configuration Cards available now
  5. 45 Appendix E — Paper Timeline and Evidence Levels available now
  6. 46 Appendix F — Glossary available now
  7. 47 Appendix G — Reproduction Checklist available now
  8. 48 Appendix H — Expert-Review Checklist available now

The big picture

  1. One questionName a failure and a prediction.
  2. One small testChange one mechanism in a sandbox.
  3. One way to loseWrite the falsifier before the score.
A research idea becomes executable when it has a locked baseline, measurable intervention, privileged-information label, budget, and result that could weaken it.

A tiny story: twelve boats

Twelve boats wait in a harbor. Some test planning geometry, some uncertainty, some representation structure, and some search or transfer. None is “the winner.” Each carries a small launch card: question, one changed part, measurements, and a condition that sends it back to port.

Twelve numbered project boats occupy four equal fleets around shared diagnostic buoys, with no maturity ranking.

These are experiment plans, not new results. Equal berths do not claim equal difficulty, promise, or evidence.

The technical backpack: twelve launch cards

All cards begin with a frozen LeWM baseline, episode-disjoint data, declared seeds, matched training and planning budgets, raw logs, and positive/null/negative reporting. Simulator state, shortest path, rewards, depth, or language annotations are labeled privileged whenever they train or evaluate a component.

#ProjectChange and measurementA result that weakens the idea
1Latent distance vs true shortest pathIn TwoRoom, correlate latent pair distance with oracle geodesic distance; stratify same-room/across-wall and near/far pairscorrelation vanishes on held-out layouts, or better correlation does not improve candidate ranking
2Directed reachabilityLearn a budget-conditioned, ordered score from trajectory offsets and constructed negatives; test withheld directions and horizonsthe score is symmetric, follows dataset frequency, or fails matched oracle-budget ranking
3Predictive uncertainty calibrationPredict spread or error bins and compare confidence with later self-fed rollout errorconfidence is sharp but uncalibrated, or only tracks horizon length
4Ensemble epistemic uncertaintyTrain matched-seed or bootstrap models; compare disagreement inside and outside supportmembers agree on wrong unsupported shortcuts, or disagreement merely reflects optimization noise
5Multi-step trainingAdd one declared rollout-consistency or multi-horizon loss while keeping inference and budgets fixedone-step loss improves but long rollout/planning worsens, or gain comes from extra updates
6Temporally hierarchical latentsLearn coarse states or action chunks and audit reconstruction, reuse, support, and two-level costhigh-level error, unsupported subgoals, or overhead exceeds saved low-level computation
7Local Gaussian structure, global topologyRegularize local neighborhoods or subspaces while measuring global spread and doorway continuity separatelyanti-collapse weakens, or topology gains disappear without extra capacity
8Action-controllable subspaceSplit action-sensitive change from content; test interventions, nuisance changes, and weakly controlled variablesthe “control” space encodes background/task ID or drops slow task-critical state
9Counterfactual-dynamics probesAt matched states, compare predictions for recorded and withheld actions against simulator transitionsbranch separation exists without correct consequences, or only recorded actions are accurate
10Support-constrained CEMPenalize support distance or search near empirical action/macro-action anchors; report constraint strengtha strict gate blocks valid novel detours, while a loose gate still admits model exploits
11Goal images to language goalsAdd a declared language-to-goal interface; hold out object–relation combinations and compare oracle goal encodingssuccess depends on memorized phrases or paired-goal leakage rather than composition
12Cross-environment latent alignmentAlign environments with paired anchors, then test held-out dynamics, views, and tasksalignment improves a probe but harms transition prediction or planning in either world

Projects 1–2 distinguish distance from directed, budgeted feasibility. Projects 3–4 distinguish predictive spread, ensemble disagreement, support distance, and actual error. Projects 5–6 distinguish longer training targets from hierarchy. Projects 7–9 ask what representation structure actions really need. Projects 10–12 move the pressure to search, goals, and transfer. Combining cards before each wins its own test destroys the diagnosis.

Try to break the idea

Support-constrained CEM is a useful self-attack. In one TwoRoom layout, an unconstrained optimizer takes an impossible shortcut. A moderate support rule guides it through the doorway. Then introduce a valid novel detour absent from the training trajectories. An overly strict rule blocks the only successful path.

Sweep the support penalty and show both failure families. Use an oracle feasible-route label only for evaluation. If the best threshold changes with goal type, report the interaction; do not hide it in one average. The project succeeds scientifically even if no universal threshold exists, because it has mapped the trade-off between model exploitation and over-conservative search.

Every card needs a removal test. Remove the wall for geometry, equalize action coverage for controllability, inject controlled out-of-support samples for uncertainty, shuffle privileged labels for grounding, remove paired anchors for alignment, and hold out language combinations for composition. A result that survives only the easy condition has not earned the broad claim.

Experiment receipt and evidence boundary

  • These cards are tutorial-generated hypotheses, not reported improvements. Direct inspirations include LeWorldModel v3, RC-aux, TRM, Sub-JEPA, SMWM, VLWM, Temporal-Distance JEPA, Fast-LeWM, Hi-LeWM, and PSG-JEPA.
  • Later diagnostics and decision routes include TwoRoom reproduction, VIScore, ACPC, Objective Bottleneck, Traj-LeWM, SCALE, AC-MTM, and DA-LeWM. Their evaluations remain author-reported except the explicitly bounded TwoRoom reimplementation.
  • Public frozen implementations were identified for LeWM, RC-aux, Temporal-Distance JEPA, Sub-JEPA, Fast-LeWM, Hi-LeWM, stable-worldmodel, passive-identifiability, SMWM, PSG-JEPA, tinylab, VIScore, ACPC, Traj-LeWM, and AC-MTM. None was identified at cutoff for TRM, VLWM, SCALE, DA-LeWM, PhyLatent, TC-LeWM, QQWorld, ProWorld, controlled-identifiability, or the 2024 “HWM” course-shorthand route. Public code is not independent reproduction.
  • Source versions span 2024-06-01 through DA-LeWM v1 on 2026-08-19; evidence cutoff is 2026-08-20, Asia/Tokyo. Permitted verbs are “test,” “hypothesis,” “falsifier,” and “supports under this locked setting”—not “will improve,” “solves uncertainty,” “pixels-only” when privileged labels train the model, or “proves compositional understanding.”

Three quick questions

  1. What five fields turn a research idea into an executable launch card?
  2. Why are ensemble disagreement and actual rollout error different instruments?
  3. How can a support constraint both help and hurt the same planner?