JEPA4Japan · tutorials

Appendix G — Reproduction Checklist

1,279 words 6 min read #LeWorldModel#World Models#JEPA

Record software, hardware, data hashes, seeds, parameter counts, budgets, baselines, episodes, failure cases, raw logs, and checkpoints.

Course progress Course outline 48 of 48 lessons available

Part 0 — Reading Guide: What Exactly Are We Going to Learn?

  1. 01 Chapter 0 — Before You Begin available now

Part 1 — World Models: An Internal Sandbox for the Agent

  1. 02 Chapter 1 — Why an Agent Needs to “Imagine the Future” available now
  2. 03 Chapter 2 — Why Not Predict the Next Image Directly? available now
  3. 04 Chapter 3 — The JEPA Idea: Predict Meaning, Not a Replica of the Image available now
  4. 05 Chapter 4 — Understand LeWM in One Diagram available now

Part 2 — Turning Images into State: The LeWM Architecture

  1. 06 Chapter 5 — Trajectory Data: To the Model, the World Is Not an Image Collection available now
  2. 07 Chapter 6 — The Visual Encoder: Issuing a “State Passport” for Every Frame available now
  3. 08 Chapter 7 — The Dynamics Predictor: Moving Time Forward in the Mind available now
  4. 09 Chapter 8 — A Complete Forward Pass: Follow One Batch from Start to Finish available now

Part 3 — Preventing the Model from Cheating: Prediction Loss and SIGReg

  1. 10 Chapter 9 — The Most Dangerous Shortcut: Representation Collapse available now
  2. 11 Chapter 10 — Prediction Loss: How the Model Learns the Next Step available now
  3. 12 Chapter 11 — The Intuition Behind SIGReg: Letting Representation Space “Breathe” available now
  4. 13 Chapter 12 — Keep the Mathematics Minimal but Sufficient available now
  5. 14 Chapter 13 — The Original LeWM’s End-to-End Training Mechanism available now
  6. 15 Chapter 14 — Train a Model That Does Not Collapse Immediately available now

Part 4 — Putting the Model into Action: Planning in Latent Space

  1. 16 Chapter 15 — Goal-Conditioned Planning: From “Where Am I?” to “Where Do I Want to Go?” available now
  2. 17 Chapter 16 — Latent Euclidean Distance: Convenient, but Not Necessarily Reliable available now
  3. 18 Chapter 17 — CEM: Searching for Actions Through an Elimination Tournament available now
  4. 19 Chapter 18 — MPC: Do Not Trust the Model for Too Long at Once available now
  5. 20 Chapter 19 — Long-Horizon Rollouts: How Small Errors Snowball into Major Failures available now
  6. 21 Chapter 20 — Implement a Minimal LeWM Planner from Scratch available now

Part 5 — Engineering Reproduction: From Paper to Running System

  1. 22 Chapter 21 — The Official Repository and Experimental Environment available now
  2. 23 Chapter 22 — First Experiment: A TwoRoom Smoke Test available now
  3. 24 Chapter 23 — Second Experiment: Reproducing PushT available now
  4. 25 Chapter 24 — How to Evaluate a World Model Fairly available now
  5. 26 Chapter 25 — Failure-Diagnosis Manual available now

Part 6 — What Has LeWM Actually Learned?

  1. 27 Chapter 26 — Linear Probes: Which Physical Variables Are Encoded in the Latent State? available now
  2. 28 Chapter 27 — Give Latent Space a “Health Check” available now
  3. 29 Chapter 28 — Violation of Expectation: Is the Model Surprised by “Impossible Events”? available now
  4. 30 Chapter 29 — How to Discuss “Understanding the World” Rigorously available now

Part 7 — Why “Accurate Prediction” Can Still Produce “Poor Planning”

  1. 31 Chapter 30 — The Gap Between the Training Objective and the Planning Objective available now
  2. 32 Chapter 31 — Global Non-Collapse Does Not Guarantee Preservation of Task-Relevant Dynamics available now
  3. 33 Chapter 32 — When Is an Isotropic Gaussian Prior Too Strong? available now
  4. 34 Chapter 33 — Long-Horizon Planning: Predict Farther or Plan More Intelligently? available now
  5. 35 Chapter 34 — From Positional Distance to Task Progress available now
  6. 36 Chapter 35 — Multi-Task Learning, Real Robots, and Visual Distractions available now
  7. 37 Chapter 36 — Theoretical Boundaries: When Can the True State Be Identified? available now

Part 8 — From Reproducer to Researcher

  1. 38 Chapter 37 — Design a Credible LeWM Improvement Experiment available now
  2. 39 Chapter 38 — Twelve Executable Research Projects available now
  3. 40 Chapter 39 — Open Questions in LeWM Research available now

Appendices

  1. 41 Appendix A — The Minimum Necessary Mathematical Toolkit available now
  2. 42 Appendix B — PyTorch Implementation Quick Reference available now
  3. 43 Appendix C — Complete Tensor-Shape Table available now
  4. 44 Appendix D — Experiment Configuration Cards available now
  5. 45 Appendix E — Paper Timeline and Evidence Levels available now
  6. 46 Appendix F — Glossary available now
  7. 47 Appendix G — Reproduction Checklist Current lesson
  8. 48 Appendix H — Expert-Review Checklist available now

The big picture

  1. Name itVersion, data, seeds, and episodes.
  2. Meter itTraining and planning get separate budgets.
  3. Hand it overLogs, failures, commands, and checkpoints travel together.
A success score is one observation. A reproduction package is the chain of identity, procedure, raw outcomes, and known gaps that lets someone inspect it.

A tiny story: the suitcase with one number

A researcher reaches the border carrying a card that says 82% success. The officer asks: Which code? Which bytes? Which seeds? Which planner budget? Which episodes? Where are the failed runs? The card cannot answer.

The number may be honest. It is simply not a complete experiment receipt. Use the ten compartments below before a run, during it, and before publication. A blank marked “not verified” is better than a remembered guess.

An evidence suitcase has ten compartments for software, data, seeds, parameters, training, planning, baselines, episodes, failures, and raw artifacts.

Completeness supports inspection; it does not promise bitwise identity across hardware and nondeterministic kernels.

Copyable ten-compartment run receipt

G.1 Software and hardware

  • Paper/version, repository URL, exact SHA, local diff, launch command, and behavior-changing environment variables.
  • OS, Python, framework, CUDA/cuDNN, GPU model/count/memory, precision, compilation/fused kernels, and deterministic settings.
  • Lockfile or exported resolved environment, including first-party dependencies. LeWM core 8edfeb3 does not freeze moving stable-pretraining or stable-worldmodel by itself.
  • Protocol history. Current v3 TwoRoom says goal offset/budget 25/50; the independent audit’s paper-side 100/150 belongs to older v1/v2 history.

G.2 Data identity and hash

  • Artifact URI/revision, downloaded filename, cryptographic hash, license/access rule, decompression/conversion, and exact bytes consumed by the loader.
  • Format, fields, crop/resize, normalization, fitted preprocessing split, frame skip/action grouping, history, window length/stride, and invalid-window rule.
  • Episode IDs and episode-disjoint train/validation/test split before windowing; trace every sampled window back to one episode.
  • Resolve format discrepancies explicitly: the frozen PushT YAML names Lance, while the inspected public artifact/README path is compressed HDF5/HDF5-oriented.

For the audited TwoRoom reimplementation, record dense actions across each frame-skip block, programmatic action-encoder width, ImageNet pixel normalization, and action z-scoring. Those are TwoRoom audit facts, not universal LeWM requirements.

G.3 Seeds and stochastic policy

  • List seeds for split/window sampling, initialization, augmentation, dropout, loader workers, environment resets, collection policy, CEM sampling, and goal selection—or the exact derivation from a master seed.
  • State remaining nondeterminism and the rule for failed/excluded runs. Keep seed identities, not only “five seeds.”
  • Pair evaluation and planner seeds across variants where the comparison calls for paired noise.

G.4 Parameters and training graph

  • Total/trainable parameter counts by visual backbone, projectors, action encoder, predictor, and optional heads; mark frozen and pretrained parts.
  • Confirm baseline v3 identity: trainable visual encoder, connected shifted target, no EMA teacher, no pretrained visual encoder.
  • Record which loss reaches which parameters. Similar totals do not mean equal architecture or prior compute.

G.5 Training budget

  • Optimizer updates, batch/context sizes, data windows/exposures, epochs, optimizer/schedule, validation cadence, early stopping, precision, wall time/device-hours, and checkpoint-selection rule.
  • Pretraining data/compute, or unreported—never silently zero.
  • Resolve paper/code differences: named LeWM experiments report 10 epochs; the frozen global maximum is 100. Archive the composed value actually run.

G.6 Planning budget

  • Candidate population, elites, refinement rounds, proposal initialization, action bounds, model horizon, environment actions per model step, executed prefix, replanning interval, total action allowance, vectorization, and per-decision latency.
  • Keep this meter separate from training. Equal parameters do not equal search; equal wall time does not equal sampled actions.
  • Translate clocks. In the cited frozen cards, 5 environment actions form one model step; horizon and executed prefix are both 5 model steps, so the full five-block sequence is executed before replanning.

G.7 Baselines and oracle conditions

  • For every baseline: implementation/SHA, checkpoint, data/preprocessing, pretraining, capacity, training/planning budget, tuning space, and the same episode ledger.
  • Label random policy, behavior cloning, oracle dynamics, and oracle cost by role. Oracles localize headroom; they are not deployable methods or proof a learned replacement is attainable.
  • State whether a schedule/hyperparameter was selected before evaluation or post hoc. Retuning only the favored method is not a matched comparison.

G.8 Evaluation episodes

  • Environment/task revision, episode/start/goal IDs, goal rule, count, action budget, success definition, termination/timeout/reset rules, planner seed, and failure handling.
  • Reuse paired starts/goals/randomness where appropriate. Report per-seed/per-episode outcomes, uncertainty, timeouts, and goal-type strata—not only the mean.
  • Freeze the episode ledger before reading the headline result.

G.9 Failures, nulls, and negatives

  • Save raw successes and failures using a declared selection rule: all failures in a fixed suite, first per seed, uniform sample, or preregistered strata.
  • Classify the earliest observable break: data support, collapse/missing variable, action insensitivity, one-step mismatch, self-fed drift, cost misranking, CEM miss, execution/interface error, or timeout.
  • Preserve aligned frames, observations, actions, predicted diagnostics, candidate costs, selected sequence, executed prefix, and termination reason.
  • Archive null and negative ablations. “No gain here” is evidence under this setup, not impossibility.

G.10 Raw archive and command record

  • Source/diff, dependency lock, command, environment variables, composed configuration, data manifest/hashes, stdout/stderr, structured metrics/events, checkpoints, optimizer/scheduler/scaler/RNG state when resumption matters, episode ledger, raw predictions/actions, and analysis/figure script revision.
  • README with directory map, missing-item inventory, access restrictions, and meanings of not collected, lost, private, too large, and available on request.
  • Separate raw evidence (frames/actions/IDs/costs), derived evidence (tables/statistics), and presentation (cropped figures/videos), with a traceable path between them.
  • Smoke-test a clean handoff: load one checkpoint, trace one window to its episode, regenerate one summary, and test training resumption separately from inference.

Try to break the receipt

Two laboratories use the same commit, seed, GPU, and checkpoint. Their losses match exactly. Then both discover the sampler deterministically paired the last frame of one episode with the first action of the next.

Reproducibility repeated the procedure; it did not make the procedure correct. Add mechanical tests for episode boundaries, target shifts, causal visibility, action sensitivity, gradient recipients, and the separation between training target and planning goal. Conversely, two correct stochastic runs may differ numerically while supporting the same distributional conclusion.

Now imagine two 80% scores. One used pretrained vision, 300 candidates, 30 refinements, five seeds, and same-trajectory goals. The other trained end-to-end, used 64 candidates, five refinements, one seed, and a broader goal split. Equal numbers are not equal experiments. Lock the receipt before comparing them.

Evidence receipt

Method identity comes from LeWorldModel v3, frozen LeWM code, first-party dependencies/artifacts, and the TwoRoom reproduction v1 with tinylab efa9e5d. LeWM v3 was revised 2026-06-03; the core commit is dated 2026-05-22; reproduction v1 appeared 2026-08-10; checklist cutoff is 2026-08-20. Independent reproduction remains TwoRoom-only. A complete checklist preserves identity and gaps; it cannot guarantee bitwise equality, repair unsupported data, or make incomparable runs fair after the fact.

Three quick questions

  1. Why must training and planning budgets use separate meters?
  2. What does a data hash prove, and what does it leave open?
  3. Which artifacts let a stranger regenerate one result without guessing?