JEPA4Japan · tutorials

Chapter 24 — How to Evaluate a World Model Fairly

728 words 4 min read #LeWorldModel#World Models#JEPA

Report one-step, multi-step, closed-loop, and planning metrics under matched data and compute budgets, strong baselines, multiple seeds, and explicit conditions.

Course progress Course outline 48 of 48 lessons available

Part 0 — Reading Guide: What Exactly Are We Going to Learn?

  1. 01 Chapter 0 — Before You Begin available now

Part 1 — World Models: An Internal Sandbox for the Agent

  1. 02 Chapter 1 — Why an Agent Needs to “Imagine the Future” available now
  2. 03 Chapter 2 — Why Not Predict the Next Image Directly? available now
  3. 04 Chapter 3 — The JEPA Idea: Predict Meaning, Not a Replica of the Image available now
  4. 05 Chapter 4 — Understand LeWM in One Diagram available now

Part 2 — Turning Images into State: The LeWM Architecture

  1. 06 Chapter 5 — Trajectory Data: To the Model, the World Is Not an Image Collection available now
  2. 07 Chapter 6 — The Visual Encoder: Issuing a “State Passport” for Every Frame available now
  3. 08 Chapter 7 — The Dynamics Predictor: Moving Time Forward in the Mind available now
  4. 09 Chapter 8 — A Complete Forward Pass: Follow One Batch from Start to Finish available now

Part 3 — Preventing the Model from Cheating: Prediction Loss and SIGReg

  1. 10 Chapter 9 — The Most Dangerous Shortcut: Representation Collapse available now
  2. 11 Chapter 10 — Prediction Loss: How the Model Learns the Next Step available now
  3. 12 Chapter 11 — The Intuition Behind SIGReg: Letting Representation Space “Breathe” available now
  4. 13 Chapter 12 — Keep the Mathematics Minimal but Sufficient available now
  5. 14 Chapter 13 — The Original LeWM’s End-to-End Training Mechanism available now
  6. 15 Chapter 14 — Train a Model That Does Not Collapse Immediately available now

Part 4 — Putting the Model into Action: Planning in Latent Space

  1. 16 Chapter 15 — Goal-Conditioned Planning: From “Where Am I?” to “Where Do I Want to Go?” available now
  2. 17 Chapter 16 — Latent Euclidean Distance: Convenient, but Not Necessarily Reliable available now
  3. 18 Chapter 17 — CEM: Searching for Actions Through an Elimination Tournament available now
  4. 19 Chapter 18 — MPC: Do Not Trust the Model for Too Long at Once available now
  5. 20 Chapter 19 — Long-Horizon Rollouts: How Small Errors Snowball into Major Failures available now
  6. 21 Chapter 20 — Implement a Minimal LeWM Planner from Scratch available now

Part 5 — Engineering Reproduction: From Paper to Running System

  1. 22 Chapter 21 — The Official Repository and Experimental Environment available now
  2. 23 Chapter 22 — First Experiment: A TwoRoom Smoke Test available now
  3. 24 Chapter 23 — Second Experiment: Reproducing PushT available now
  4. 25 Chapter 24 — How to Evaluate a World Model Fairly Current lesson
  5. 26 Chapter 25 — Failure-Diagnosis Manual available now

Part 6 — What Has LeWM Actually Learned?

  1. 27 Chapter 26 — Linear Probes: Which Physical Variables Are Encoded in the Latent State? available now
  2. 28 Chapter 27 — Give Latent Space a “Health Check” available now
  3. 29 Chapter 28 — Violation of Expectation: Is the Model Surprised by “Impossible Events”? available now
  4. 30 Chapter 29 — How to Discuss “Understanding the World” Rigorously available now

Part 7 — Why “Accurate Prediction” Can Still Produce “Poor Planning”

  1. 31 Chapter 30 — The Gap Between the Training Objective and the Planning Objective available now
  2. 32 Chapter 31 — Global Non-Collapse Does Not Guarantee Preservation of Task-Relevant Dynamics available now
  3. 33 Chapter 32 — When Is an Isotropic Gaussian Prior Too Strong? available now
  4. 34 Chapter 33 — Long-Horizon Planning: Predict Farther or Plan More Intelligently? available now
  5. 35 Chapter 34 — From Positional Distance to Task Progress available now
  6. 36 Chapter 35 — Multi-Task Learning, Real Robots, and Visual Distractions available now
  7. 37 Chapter 36 — Theoretical Boundaries: When Can the True State Be Identified? available now

Part 8 — From Reproducer to Researcher

  1. 38 Chapter 37 — Design a Credible LeWM Improvement Experiment available now
  2. 39 Chapter 38 — Twelve Executable Research Projects available now
  3. 40 Chapter 39 — Open Questions in LeWM Research available now

Appendices

  1. 41 Appendix A — The Minimum Necessary Mathematical Toolkit available now
  2. 42 Appendix B — PyTorch Implementation Quick Reference available now
  3. 43 Appendix C — Complete Tensor-Shape Table available now
  4. 44 Appendix D — Experiment Configuration Cards available now
  5. 45 Appendix E — Paper Timeline and Evidence Levels available now
  6. 46 Appendix F — Glossary available now
  7. 47 Appendix G — Reproduction Checklist available now
  8. 48 Appendix H — Expert-Review Checklist available now

The big picture

  1. Same taskLock starts, goals, and success.
  2. Show resourcesData, pretraining, parameters.
  3. Match budgetsActions, candidates, horizon, refinements.
  4. Show the spreadSeeds, episodes, and failures.
A scoreboard is fair only when every runner’s course and backpack are visible.

Evaluation climbs from mechanics and prediction to cost, search, control, and robustness.

A tiny story

Three runners show times of 42, 48, and 51 seconds. One used a shorter track, one rode a bicycle for half the race, and one started the clock after warming up. Every number can be correct while the ranking is meaningless.

World-model scores need their receipts too. A raw latent MSE depends on dimension, scale, normalization, and a learned target. A success rate depends on starts, goals, tolerance, termination, action budget, search budget, and environment version. Wall time depends on hardware and batching.

The real rule

Keep evaluation layers separate:

  1. mechanical validity and episode-safe data;
  2. one-step prediction on recorded contexts;
  3. fixed-action self-fed rollout by horizon;
  4. candidate cost ranking;
  5. search progress;
  6. closed-loop control;
  7. robustness and support shift.

LeWM v3 evaluates TwoRoom, Reacher, PushT, and OGBench-Cube. In its reported protocol, starts come from offline trajectories, goals are 25 environment steps later in the same trajectory, and the real-action budget is 50. That makes the goal recorded as reachable under one behavior; it does not make arbitrary visual goals solved.

Lock or disclose two budgets. The environment budget includes real actions, termination, and replanning frequency. The planning-compute budget includes candidates, horizon, actions per model block, elites, refinements, model calls, batch shape, and parallelism. Hardware, precision, warm-up, synchronization, compilation, transfer, and the timed region belong beside wall time.

Useful comparison views ask different questions: matched task data, matched trainable parameters, matched task-training compute, native method settings, or a compute frontier. Name the view instead of calling one of them “fair” in the abstract.

The trick that can fool us

DINO-WM arrives with a frozen pretrained DINOv2 representation; original LeWM trains its visual encoder end to end on task trajectories. Report task-training cost and upstream provenance separately. Also name every input modality, including conditions with proprioception.

Use baselines as diagnostic tools:

  • random policy tests task easiness by chance;
  • goal-conditioned behavior cloning tests direct imitation;
  • oracle cost + learned dynamics tests ranking;
  • oracle dynamics + learned cost tests rollout;
  • both oracles + the same finite search test candidate coverage and action parameterization.

An oracle receives privileged simulator information. It is not a deployable competitor.

Separate training seed, evaluation episode, and planner seed. Three trained checkpoints on the same 50 trajectories are not 50 independent training runs. Define every ± mark: SD, SE, bootstrap interval, confidence interval, or another quantity, plus its resampling unit. LeWM’s PushT caption calls its quantity “variance”; do not silently reinterpret it.

Experiment receipt

Publish per-episode identity, seeds, success/timeout, final and best task distance, actions, replans, time, first failure, support slice, crashes, and exclusions. Means can hide rare catastrophic exploitation or steady near-success.

LeWM v3 reports planning speedups up to 48× over foundation-model-based world models under its selected setup, averaged in the figure over 50 runs. The exact historical hardware receipt is incomplete, so the multiplier is not an intrinsic property.

The later VIScore v2 multiplies Veracity, Influence, and Sobriety. Its authors report cross-task Spearman correlation at least 0.75 on their tested pools. This is association, not deterministic success prediction: constants and calibration were selected on development runs, Cube is excluded from pooled calibration claims, and the estimator assumes rollout-error direction is approximately stable across horizons.

Source claims remain attached to LeWorldModel v3, frozen commit 8edfeb3, and each named baseline version. A complete report asks, “Under which resources and protocol?” before asking, “Who won?”

Quick check

  1. Why can raw latent errors from two encoders be incomparable?
  2. Which two budgets must every planner comparison expose?
  3. What different questions do oracle cost and oracle dynamics answer?