JEPA4Japan · tutorials

Chapter 23 — Second Experiment: Reproducing PushT

803 words 4 min read #LeWorldModel#World Models#JEPA

Move from geometric navigation to contact dynamics through data preparation, model training, goal-image planning, metrics, and failure classification.

Course progress Course outline 48 of 48 lessons available

Part 0 — Reading Guide: What Exactly Are We Going to Learn?

  1. 01 Chapter 0 — Before You Begin available now

Part 1 — World Models: An Internal Sandbox for the Agent

  1. 02 Chapter 1 — Why an Agent Needs to “Imagine the Future” available now
  2. 03 Chapter 2 — Why Not Predict the Next Image Directly? available now
  3. 04 Chapter 3 — The JEPA Idea: Predict Meaning, Not a Replica of the Image available now
  4. 05 Chapter 4 — Understand LeWM in One Diagram available now

Part 2 — Turning Images into State: The LeWM Architecture

  1. 06 Chapter 5 — Trajectory Data: To the Model, the World Is Not an Image Collection available now
  2. 07 Chapter 6 — The Visual Encoder: Issuing a “State Passport” for Every Frame available now
  3. 08 Chapter 7 — The Dynamics Predictor: Moving Time Forward in the Mind available now
  4. 09 Chapter 8 — A Complete Forward Pass: Follow One Batch from Start to Finish available now

Part 3 — Preventing the Model from Cheating: Prediction Loss and SIGReg

  1. 10 Chapter 9 — The Most Dangerous Shortcut: Representation Collapse available now
  2. 11 Chapter 10 — Prediction Loss: How the Model Learns the Next Step available now
  3. 12 Chapter 11 — The Intuition Behind SIGReg: Letting Representation Space “Breathe” available now
  4. 13 Chapter 12 — Keep the Mathematics Minimal but Sufficient available now
  5. 14 Chapter 13 — The Original LeWM’s End-to-End Training Mechanism available now
  6. 15 Chapter 14 — Train a Model That Does Not Collapse Immediately available now

Part 4 — Putting the Model into Action: Planning in Latent Space

  1. 16 Chapter 15 — Goal-Conditioned Planning: From “Where Am I?” to “Where Do I Want to Go?” available now
  2. 17 Chapter 16 — Latent Euclidean Distance: Convenient, but Not Necessarily Reliable available now
  3. 18 Chapter 17 — CEM: Searching for Actions Through an Elimination Tournament available now
  4. 19 Chapter 18 — MPC: Do Not Trust the Model for Too Long at Once available now
  5. 20 Chapter 19 — Long-Horizon Rollouts: How Small Errors Snowball into Major Failures available now
  6. 21 Chapter 20 — Implement a Minimal LeWM Planner from Scratch available now

Part 5 — Engineering Reproduction: From Paper to Running System

  1. 22 Chapter 21 — The Official Repository and Experimental Environment available now
  2. 23 Chapter 22 — First Experiment: A TwoRoom Smoke Test available now
  3. 24 Chapter 23 — Second Experiment: Reproducing PushT Current lesson
  4. 25 Chapter 24 — How to Evaluate a World Model Fairly available now
  5. 26 Chapter 25 — Failure-Diagnosis Manual available now

Part 6 — What Has LeWM Actually Learned?

  1. 27 Chapter 26 — Linear Probes: Which Physical Variables Are Encoded in the Latent State? available now
  2. 28 Chapter 27 — Give Latent Space a “Health Check” available now
  3. 29 Chapter 28 — Violation of Expectation: Is the Model Surprised by “Impossible Events”? available now
  4. 30 Chapter 29 — How to Discuss “Understanding the World” Rigorously available now

Part 7 — Why “Accurate Prediction” Can Still Produce “Poor Planning”

  1. 31 Chapter 30 — The Gap Between the Training Objective and the Planning Objective available now
  2. 32 Chapter 31 — Global Non-Collapse Does Not Guarantee Preservation of Task-Relevant Dynamics available now
  3. 33 Chapter 32 — When Is an Isotropic Gaussian Prior Too Strong? available now
  4. 34 Chapter 33 — Long-Horizon Planning: Predict Farther or Plan More Intelligently? available now
  5. 35 Chapter 34 — From Positional Distance to Task Progress available now
  6. 36 Chapter 35 — Multi-Task Learning, Real Robots, and Visual Distractions available now
  7. 37 Chapter 36 — Theoretical Boundaries: When Can the True State Be Identified? available now

Part 8 — From Reproducer to Researcher

  1. 38 Chapter 37 — Design a Credible LeWM Improvement Experiment available now
  2. 39 Chapter 38 — Twelve Executable Research Projects available now
  3. 40 Chapter 39 — Open Questions in LeWM Research available now

Appendices

  1. 41 Appendix A — The Minimum Necessary Mathematical Toolkit available now
  2. 42 Appendix B — PyTorch Implementation Quick Reference available now
  3. 43 Appendix C — Complete Tensor-Shape Table available now
  4. 44 Appendix D — Experiment Configuration Cards available now
  5. 45 Appendix E — Paper Timeline and Evidence Levels available now
  6. 46 Appendix F — Glossary available now
  7. 47 Appendix G — Reproduction Checklist available now
  8. 48 Appendix H — Expert-Review Checklist available now

The big picture

  1. See historyApproaching or separating?
  2. Keep contactWhere and how did we touch?
  3. ImagineSlide, rotate, or lose contact?
  4. PlanChoose a supported push.
PushT needs contact-relevant state, not just recognition of a T.

Three pushes at different contact points can slide or rotate the same T-shaped block.

A tiny story

Push a piano near its center and it may slide. Push near a corner and it may rotate. One photograph can miss whether the mover is approaching, maintaining contact, or separating.

PushT turns that story into a simulated task. A blue agent can push, but not pull, a T-shaped block toward a target pose. Position alone is not enough: orientation, contact side, approach direction, and access to the next useful push matter. A successful simulated episode is evidence of budgeted control in this setup—not generic rigid-body physics or safe robot skill.

The real rule

The paper reports 20,000 expert episodes averaging 196 environment steps and 10 training epochs. “Expert” does not mean every useful contact mode is covered. Build a support atlas from the frozen artifact: contact location, approach direction, block position and angle, action magnitude and direction, contact duration and loss, and episode boundaries. Sparse means “little recorded evidence,” not “physically impossible.”

Use episode-disjoint splits. Record artifact revision and hash, decompression or conversion, loader, image preprocessing, action/state normalization, history, frame skip, and action width. The frozen training YAML names pusht_expert_train.lance; official public revision 655cd44 contains compressed HDF5, while the README and evaluation path also describe HDF5.

Keep four reproduction targets separate:

  1. checkpoint loads and produces finite output;
  2. checkpoint runs a locked evaluation;
  3. a fresh model trains and enters that evaluation;
  4. several training seeds produce a defined statistical result.

The paper card uses history 3, 10 epochs, and SIGReg weight 0.1 unless stated otherwise. Frozen code also uses history 3, but maximum 100 epochs and weight 0.09. The official checkpoint revision inspected here is 22b330c.

The trick that can fool us

Make two short histories ending in nearly identical frames. In one, the pusher has just made inward contact and starts rotation. In the other, it is sliding past and separating. Apply the same next action.

If full history produces appropriate different predictions, recent visual context helps. If predictions differ but disagree with real outcomes, dynamics is wrong. If rollout is useful but CEM chooses a bad contact, inspect cost or search. Replacing both histories with their last frame deliberately tests missing observability.

Classify the first divergence, not only the final frame: no contact, wrong-side contact, translation/rotation confusion, contact loss, false progress, support exploitation, or cost failure. A decoder can look plausible while hiding angle error; a probe can decode angle without proving the predictor or cost uses it.

Oracle cost with learned dynamics tests ranking. Oracle dynamics with learned cost tests rollout. Both oracles with the same finite CEM budget test search and action parameterization. These are privileged tutorial diagnostics, not deployable LeWM baselines.

Experiment receipt

The goal is not arbitrary. The protocol chooses a start from a recorded trajectory and its state 25 environment steps later as the goal, with a 50-step action budget. Planning uses five model blocks, five environment actions per block, and executes five blocks before replanning. CEM uses 300 candidates, 30 elites, 30 refinements, and initial variance 1. One plan therefore spans—and the frozen interval executes—25 environment steps.

LeWM v3 reports three training seeds on the same 50 trajectories:

MethodAuthor-reported PushT success
LeWM96.0 ± 2.83
DINO-WM92.0 ± 1.63
PLDM78.0 ± 5.0

Those are differences of 4 and 18 percentage points under this protocol, not universal percent improvements. The caption calls the accompanying quantity “variance”; the ± notation alone does not justify renaming it SD, SE, or CI. Pair success with position and orientation error, overlap or task measure, closest approach and reversal, actions/replans, timed hardware receipt, and first failure.

DINO-WM uses a pretrained DINOv2 encoder; original LeWM trains its visual encoder end to end on task trajectories. Keep pretraining and modality provenance visible.

Quick check

  1. Why can two similar PushT frames need different next predictions?
  2. What does a sparse support-atlas cell mean?
  3. Why must 96.0 ± 2.83 stay attached to the exact protocol?