JEPA4Japan · tutorials

Appendix H — Expert-Review Checklist

1,264 words 6 min read #LeWorldModel#World Models#JEPA

Audit the training graph, variant boundaries, probe interpretation, benchmark generalization, causal language, theoretical assumptions, negative results, and code paths.

Course progress Course outline 48 of 48 lessons available

Part 0 — Reading Guide: What Exactly Are We Going to Learn?

  1. 01 Chapter 0 — Before You Begin available now

Part 1 — World Models: An Internal Sandbox for the Agent

  1. 02 Chapter 1 — Why an Agent Needs to “Imagine the Future” available now
  2. 03 Chapter 2 — Why Not Predict the Next Image Directly? available now
  3. 04 Chapter 3 — The JEPA Idea: Predict Meaning, Not a Replica of the Image available now
  4. 05 Chapter 4 — Understand LeWM in One Diagram available now

Part 2 — Turning Images into State: The LeWM Architecture

  1. 06 Chapter 5 — Trajectory Data: To the Model, the World Is Not an Image Collection available now
  2. 07 Chapter 6 — The Visual Encoder: Issuing a “State Passport” for Every Frame available now
  3. 08 Chapter 7 — The Dynamics Predictor: Moving Time Forward in the Mind available now
  4. 09 Chapter 8 — A Complete Forward Pass: Follow One Batch from Start to Finish available now

Part 3 — Preventing the Model from Cheating: Prediction Loss and SIGReg

  1. 10 Chapter 9 — The Most Dangerous Shortcut: Representation Collapse available now
  2. 11 Chapter 10 — Prediction Loss: How the Model Learns the Next Step available now
  3. 12 Chapter 11 — The Intuition Behind SIGReg: Letting Representation Space “Breathe” available now
  4. 13 Chapter 12 — Keep the Mathematics Minimal but Sufficient available now
  5. 14 Chapter 13 — The Original LeWM’s End-to-End Training Mechanism available now
  6. 15 Chapter 14 — Train a Model That Does Not Collapse Immediately available now

Part 4 — Putting the Model into Action: Planning in Latent Space

  1. 16 Chapter 15 — Goal-Conditioned Planning: From “Where Am I?” to “Where Do I Want to Go?” available now
  2. 17 Chapter 16 — Latent Euclidean Distance: Convenient, but Not Necessarily Reliable available now
  3. 18 Chapter 17 — CEM: Searching for Actions Through an Elimination Tournament available now
  4. 19 Chapter 18 — MPC: Do Not Trust the Model for Too Long at Once available now
  5. 20 Chapter 19 — Long-Horizon Rollouts: How Small Errors Snowball into Major Failures available now
  6. 21 Chapter 20 — Implement a Minimal LeWM Planner from Scratch available now

Part 5 — Engineering Reproduction: From Paper to Running System

  1. 22 Chapter 21 — The Official Repository and Experimental Environment available now
  2. 23 Chapter 22 — First Experiment: A TwoRoom Smoke Test available now
  3. 24 Chapter 23 — Second Experiment: Reproducing PushT available now
  4. 25 Chapter 24 — How to Evaluate a World Model Fairly available now
  5. 26 Chapter 25 — Failure-Diagnosis Manual available now

Part 6 — What Has LeWM Actually Learned?

  1. 27 Chapter 26 — Linear Probes: Which Physical Variables Are Encoded in the Latent State? available now
  2. 28 Chapter 27 — Give Latent Space a “Health Check” available now
  3. 29 Chapter 28 — Violation of Expectation: Is the Model Surprised by “Impossible Events”? available now
  4. 30 Chapter 29 — How to Discuss “Understanding the World” Rigorously available now

Part 7 — Why “Accurate Prediction” Can Still Produce “Poor Planning”

  1. 31 Chapter 30 — The Gap Between the Training Objective and the Planning Objective available now
  2. 32 Chapter 31 — Global Non-Collapse Does Not Guarantee Preservation of Task-Relevant Dynamics available now
  3. 33 Chapter 32 — When Is an Isotropic Gaussian Prior Too Strong? available now
  4. 34 Chapter 33 — Long-Horizon Planning: Predict Farther or Plan More Intelligently? available now
  5. 35 Chapter 34 — From Positional Distance to Task Progress available now
  6. 36 Chapter 35 — Multi-Task Learning, Real Robots, and Visual Distractions available now
  7. 37 Chapter 36 — Theoretical Boundaries: When Can the True State Be Identified? available now

Part 8 — From Reproducer to Researcher

  1. 38 Chapter 37 — Design a Credible LeWM Improvement Experiment available now
  2. 39 Chapter 38 — Twelve Executable Research Projects available now
  3. 40 Chapter 39 — Open Questions in LeWM Research available now

Appendices

  1. 41 Appendix A — The Minimum Necessary Mathematical Toolkit available now
  2. 42 Appendix B — PyTorch Implementation Quick Reference available now
  3. 43 Appendix C — Complete Tensor-Shape Table available now
  4. 44 Appendix D — Experiment Configuration Cards available now
  5. 45 Appendix E — Paper Timeline and Evidence Levels available now
  6. 46 Appendix F — Glossary available now
  7. 47 Appendix G — Reproduction Checklist available now
  8. 48 Appendix H — Expert-Review Checklist Current lesson

The big picture

  1. Right machine?Trace the graph and version.
  2. Right evidence?Probe, cause, theorem, and benchmark differ.
  3. Right scope?Pin conditions, failures, and reproducible paths.
Expert review asks whether every word in a claim is attached to the correct model, evidence type, condition, and inspectable path.

A tiny story: eight inspectors

A manuscript is fluent, its figures are polished, and its table has neat decimals. Eight inspectors still stop it: one finds an invented teacher network; another finds a later preprint smuggled into v3; others ask what a probe proves, where benchmark conditions went, whether “caused” is justified, which theorem assumptions apply, where failed seeds are, and why the code link points to moving main.

Eight inspectors surround one manuscript, each owning a separate gate for graph, versions, probes, scope, causality, theory, negative evidence, and reproducible paths.

A check mark is not evidence. Every pass points to a primary source, frozen implementation, controlled experiment, explicit assumption, raw result, or declared tutorial construction.

Before review, rewrite “LeWM works” as a checkable sentence with subject, mechanism, model version, environment, data, planner/budget, evidence type, and conclusion strength.

The eight-gate checklist

H.1 Is the original training graph accurate?

  • One shared trainable visual encoder processes context and shifted next observations.
  • The shifted target is encoded by that same model and is not detached; prediction gradients can reach both connected encoder routes.
  • Actions align with transitions and condition the predictor; SIGReg regularizes observation-embedding populations.
  • Original v3 has no stop-gradient target, EMA teacher, or pretrained visual encoder.
  • Goal images, CEM, rewards, probes, and optional decoders stay outside the baseline training lane.

Passing this gate proves graph identity and possible gradient routes—not non-collapse, physics, rollout accuracy, or planning success.

H.2 Are baseline and later routes separated?

  • Give every later method a passport: title, exact version/date, changed layer, targeted failure, evidence level, code status, and independent-reproduction status.
  • Use LeJEPA only as motivation/theory where cited; never import a later detach, auxiliary head, cost, hierarchy, or teacher into v3.
  • Distinguish “defines/proposes,” “authors report,” “theorem under assumptions,” “independently reproduced,” and “tutorial construction.”
  • Treat recency, acceptance, and public author code as separate facts—not an evidence ladder.

H.3 Are probes and pictures interpreted with restraint?

  • Record probe family/capacity, train/test split, preprocessing fit, episode leakage controls, shuffled labels, and simple/raw-feature baselines.
  • Say “linearly decodable under this protocol,” not “the model learned/uses/understands the variable.”
  • Treat t-SNE/UMAP, reconstructed images, latent trajectories, and VoE surprise as conditional views, not proof of semantics or intuitive physics.
  • If claiming use or cause, add an intervention and disclose its side effects.

H.4 Do benchmark claims keep their conditions?

  • Attach environment/task, data, model, planner, action/search/compute budgets, pretraining, seed policy, episode ledger, metric, and hardware timing.
  • Inspect every speed multiplier’s denominator: model calls, batching, rendering, candidates, horizon, GPU, and whether pretraining is counted.
  • Keep scope distinct: TwoRoom topology, PushT 2-D contact, simulated Reacher, and simulated OGBench-Cube do not by themselves establish general robotics, physical understanding, or safety.
  • Preserve author-reported weaknesses and possible explanations as “possible,” not diagnosed causes; inspect goal/failure distributions, not only means.

H.5 Does language distinguish correlation from causation?

  • Circle jumps such as “correlates with velocity, therefore uses velocity” or “SIGReg changed success, therefore Gaussianity caused physics.”
  • Challenge the nearest alternative: hold physical state fixed while changing appearance; swap cost with dynamics fixed; swap dynamics with cost fixed; use oracle cells to localize bottlenecks.
  • Remember that an ablation can change capacity, optimization, support, and search queries together. Match causal wording to the intervention actually isolated.

H.6 Are theorem assumptions attached?

  • Name stationarity, transition/noise family, observation adequacy/invertibility, Gaussian/exact constraints, matched dimension, predictor expressivity, population/global optimum, signal separation, and state-conditioned action excitation where required.
  • Keep the conclusion narrow: orthogonal recovery does not produce human-named axes; controlled theory identifies a conditional mean, not every multimodal future.
  • List benchmark gaps separately: nonlinearity, partial observability, missing action support, drift, multimodality, finite data, approximate optimization, and misspecification.
  • Say “diagnostic lens, not direct coverage” when assumptions are unknown or false.

H.7 Are positive, negative, null, and missing results visible?

  • Show all declared seeds or the preregistered exclusion rule, uncertainty, per-episode outcomes, divergence/collapse/timeouts, unsuccessful ablations, and representative failures chosen by a stated policy.
  • Use open for untested, null under this setup for measured non-separation, and failed under this setup for a missed criterion.
  • Do not generalize a negative result beyond its task, method, support, budget, and protocol—or hide it because it weakens the story.

H.8 Does every implementation claim have a reproducible path?

  • End at a frozen first-party file/SHA, pinned artifact plus hash/format, or clearly labeled tutorial construction and divergence.
  • Pin dependencies separately, save composed configuration, data/checkpoint conversion, episode ledger, and analysis revision; avoid moving main when a SHA exists.
  • Run the archive smoke test: load one checkpoint, trace one window to its episode, regenerate one summary. This tests handoff, not independent reproduction.
  • Preserve history: current v3 TwoRoom is 25/50; the reproduction’s audited paper-side 100/150 belongs to v1/v2.

Evidence is coordinates, not a staircase

Decodability, predictability, controllability, generalization, causal identification, and planning utility ask different questions. A decoded position can be ignored by the predictor. Recorded-action accuracy can coexist with counterfactual failure. A planner can exploit a task shortcut without an identifiable state.

Six evidence tests occupy separate coordinates with bounded overlaps and no automatic staircase to understanding.

Ask which coordinates were tested and what remains open. Several indirect signals do not merge into the word “understanding.”

Try to break the claim

“LeWM stabilizes its pretrained visual teacher with EMA and stop-gradient, then proves through probes that its Gaussian state understands physics and makes CEM universally faster.”

This fluent sentence fails every gate. The graph is wrong; later mechanisms are imported; probes are promoted to understanding; speed lost hardware/budget conditions; correlation became causality; no theorem assumptions appear; failures vanished; and no frozen path is given.

Red-flag phrases trigger questions, not automatic rejection: proves understanding, guarantees physics, universally faster, distance is reachability, and the ablation shows the cause. Strong wording survives only when the operational definition, alternatives, budgets, assumptions, and direct evidence survive with it.

Evidence receipt

Baseline identity comes from LeWorldModel v3, frozen official code, and LeJEPA v3 for cited motivation. Theory comes from passive identifiability v1 and controlled identifiability v2. Independent evidence is limited to the TwoRoom reproduction v1 and its documented protocols. LeWM v3 was revised 2026-06-03, frozen code dated 2026-05-22, and this checklist’s cutoff is 2026-08-20. Passing all gates means the dated claim is accurately attributed, scoped, calibrated, balanced, and inspectable—not that the model is correct forever or that specialist peer review is unnecessary.

Three quick questions

  1. Which direct facts expose an invented EMA teacher in a v3 diagram?
  2. Why does a strong probe not establish that the planner uses the variable?
  3. What is the strongest conclusion after all eight gates pass?