JEPA4Japan · tutorials

Appendix B — PyTorch Implementation Quick Reference

1,107 words 6 min read #LeWorldModel#World Models#JEPA

Find concise patterns for data loading, Transformer inputs, causal masks, AdaLN, SIGReg, vectorized rollout, CEM, mixed precision, and checkpoints.

Course progress Course outline 48 of 48 lessons available

Part 0 — Reading Guide: What Exactly Are We Going to Learn?

  1. 01 Chapter 0 — Before You Begin available now

Part 1 — World Models: An Internal Sandbox for the Agent

  1. 02 Chapter 1 — Why an Agent Needs to “Imagine the Future” available now
  2. 03 Chapter 2 — Why Not Predict the Next Image Directly? available now
  3. 04 Chapter 3 — The JEPA Idea: Predict Meaning, Not a Replica of the Image available now
  4. 05 Chapter 4 — Understand LeWM in One Diagram available now

Part 2 — Turning Images into State: The LeWM Architecture

  1. 06 Chapter 5 — Trajectory Data: To the Model, the World Is Not an Image Collection available now
  2. 07 Chapter 6 — The Visual Encoder: Issuing a “State Passport” for Every Frame available now
  3. 08 Chapter 7 — The Dynamics Predictor: Moving Time Forward in the Mind available now
  4. 09 Chapter 8 — A Complete Forward Pass: Follow One Batch from Start to Finish available now

Part 3 — Preventing the Model from Cheating: Prediction Loss and SIGReg

  1. 10 Chapter 9 — The Most Dangerous Shortcut: Representation Collapse available now
  2. 11 Chapter 10 — Prediction Loss: How the Model Learns the Next Step available now
  3. 12 Chapter 11 — The Intuition Behind SIGReg: Letting Representation Space “Breathe” available now
  4. 13 Chapter 12 — Keep the Mathematics Minimal but Sufficient available now
  5. 14 Chapter 13 — The Original LeWM’s End-to-End Training Mechanism available now
  6. 15 Chapter 14 — Train a Model That Does Not Collapse Immediately available now

Part 4 — Putting the Model into Action: Planning in Latent Space

  1. 16 Chapter 15 — Goal-Conditioned Planning: From “Where Am I?” to “Where Do I Want to Go?” available now
  2. 17 Chapter 16 — Latent Euclidean Distance: Convenient, but Not Necessarily Reliable available now
  3. 18 Chapter 17 — CEM: Searching for Actions Through an Elimination Tournament available now
  4. 19 Chapter 18 — MPC: Do Not Trust the Model for Too Long at Once available now
  5. 20 Chapter 19 — Long-Horizon Rollouts: How Small Errors Snowball into Major Failures available now
  6. 21 Chapter 20 — Implement a Minimal LeWM Planner from Scratch available now

Part 5 — Engineering Reproduction: From Paper to Running System

  1. 22 Chapter 21 — The Official Repository and Experimental Environment available now
  2. 23 Chapter 22 — First Experiment: A TwoRoom Smoke Test available now
  3. 24 Chapter 23 — Second Experiment: Reproducing PushT available now
  4. 25 Chapter 24 — How to Evaluate a World Model Fairly available now
  5. 26 Chapter 25 — Failure-Diagnosis Manual available now

Part 6 — What Has LeWM Actually Learned?

  1. 27 Chapter 26 — Linear Probes: Which Physical Variables Are Encoded in the Latent State? available now
  2. 28 Chapter 27 — Give Latent Space a “Health Check” available now
  3. 29 Chapter 28 — Violation of Expectation: Is the Model Surprised by “Impossible Events”? available now
  4. 30 Chapter 29 — How to Discuss “Understanding the World” Rigorously available now

Part 7 — Why “Accurate Prediction” Can Still Produce “Poor Planning”

  1. 31 Chapter 30 — The Gap Between the Training Objective and the Planning Objective available now
  2. 32 Chapter 31 — Global Non-Collapse Does Not Guarantee Preservation of Task-Relevant Dynamics available now
  3. 33 Chapter 32 — When Is an Isotropic Gaussian Prior Too Strong? available now
  4. 34 Chapter 33 — Long-Horizon Planning: Predict Farther or Plan More Intelligently? available now
  5. 35 Chapter 34 — From Positional Distance to Task Progress available now
  6. 36 Chapter 35 — Multi-Task Learning, Real Robots, and Visual Distractions available now
  7. 37 Chapter 36 — Theoretical Boundaries: When Can the True State Be Identified? available now

Part 8 — From Reproducer to Researcher

  1. 38 Chapter 37 — Design a Credible LeWM Improvement Experiment available now
  2. 39 Chapter 38 — Twelve Executable Research Projects available now
  3. 40 Chapter 39 — Open Questions in LeWM Research available now

Appendices

  1. 41 Appendix A — The Minimum Necessary Mathematical Toolkit available now
  2. 42 Appendix B — PyTorch Implementation Quick Reference Current lesson
  3. 43 Appendix C — Complete Tensor-Shape Table available now
  4. 44 Appendix D — Experiment Configuration Cards available now
  5. 45 Appendix E — Paper Timeline and Evidence Levels available now
  6. 46 Appendix F — Glossary available now
  7. 47 Appendix G — Reproduction Checklist available now
  8. 48 Appendix H — Expert-Review Checklist available now

The big picture

  1. Keep episodes wholevalid windows first
  2. Name every axisreshape with meaning
  3. Test each routemask, action, gradient
  4. Save the voyagestate plus receipt
PyTorch moves tensors. Our job is to keep their story intact.

Episode-aware windows are formed before the loader batches them.

A tiny story

Think of PyTorch as a very fast luggage system. It will batch, reshape, broadcast, and multiply whatever we hand it. It cannot know that two neighboring frames belong to different episodes, that one H means image height and another means planning horizon, or that a target was supposed to stay connected to a shared encoder.

So this appendix is a cockpit card, not a line-by-line replacement for the official implementation. Its question is simple: which invariant should still be true after every fast tensor move?

The technical backpack

Data and shape contracts

Dataset answers “what is sample 417?”; DataLoader schedules and batches those samples. The dataset or sampler must preserve episode membership, time order, action alignment, and window length. A loader may shuffle whole windows. It must not manufacture a transition by mixing frames inside one.

Storage is not scientific identity. HDF5, Lance, folders, or video can carry the same trajectory contract, while two HDF5 files can differ in preprocessing, splits, frame skip, or action alignment. Record both container and contract.

Before a Transformer call, use a verbal boarding pass:

batch = independent examples
time = ordered latent positions
feature = learned coordinates
visibility = this position may read only itself and earlier positions

Equal sizes do not prove equal roles. If candidate count and time length are both four, swapping them can leave the printed shape unchanged. Carry recognizable episode, time, and candidate markers through a disposable reshape test.

Masking and action conditioning

A causal mask blocks a direct look at later tokens. Change a later context token and verify that earlier outputs stay unchanged. Passing this proves a computational visibility rule, not causal physics.

AdaLN gives actions a route into a Transformer block through shift, scale, and gate modulation. In the frozen implementation, its final modulation generator starts neutral through zero initialization. That does not guarantee the trained model uses actions. Hold latent history fixed, swap consequential actions, and inspect the predictions.

The causal curtain and AdaLN action panel have different jobs.

Original LeWM v3 uses one shared trainable visual encoder. The shifted next-observation target is connected: there is no stop-gradient, EMA teacher, or pretrained frozen encoder. A reimplementation that detaches the target has the same forward shape and a different learning system.

SIGReg reference route

Read the frozen SIGReg implementation as five rooms:

collect projected observation embeddings
keep time positions separate; inspect each one across the batch
sample random one-dimensional directions
compare empirical fingerprints with the analytic standard-Gaussian fingerprint
average discrepancies into a differentiable training penalty

Flattening batch and time changes the population question. Gathering across devices changes the empirical population; the frozen LeWM class is explicitly single-GPU. Detaching embeddings changes gradients. Calling the result a p-value changes its role. Finite directions, minibatches, and integration points make this an approximation, not a Gaussianity certificate. Log the penalty beside coordinate spread, covariance spectrum, effective rank, and held-out neighbors.

Rollout and CEM route

Candidate plans can be stacked and evaluated side by side on a GPU. Time within each candidate remains autoregressive: step two consumes the state predicted after step one. Keep environment batch, candidate, model horizon, and latent width distinct.

sample sequences from the current action proposal
roll out candidates with the frozen world model
score terminal latents against the encoded goal
select the configured elite action sequences
refit only the action proposal, then repeat

Elite indices must still point to the action sequences that earned their scores. CEM changes the proposal, not encoder or predictor weights. Continuous bounded actions need declared clipping/update semantics; discrete actions need a categorical proposal. Vectorization saves execution time; it does not remove temporal dependence or model drift.

Precision, checkpoints, and reproducibility

Mixed precision changes how arithmetic executes, not the scientific definition of prediction loss or SIGReg. Watch for non-finite values and compare a short full-precision smoke run. This control need not match bit for bit; it should expose implausible scale or a fragile accumulation path.

A weight file supports inference. A training-resumption checkpoint should also identify optimizer, training position, composed configuration, and any evolving scheduler or gradient-scaler state. Preserve RNG state when continuity requires it, or say it is absent. Also record data identity, code revision, hardware/software, precision, windows, budgets, checkpoint rule, planner settings, goals, seeds, and evaluation episodes.

same data and preprocessing
same model, optimizer, and training budget
same planner and evaluation budget
same episodes and goal rule
declare every intentional difference before reading the score

The trick that fools us

Two labs use the same commit, configuration, seed, checkpoint, and GPU. Their losses match exactly. Then they discover that both samplers pair the last observation of one episode with the first action of the next. The bug is perfectly reproducible.

Repetition confirms a procedure; it does not certify meaning. Challenge episode boundaries, one-step shifts, causal visibility, action sensitivity, gradient recipients, candidate identity, and the separation between world-model training and planning. Conversely, two correct stochastic runs may differ numerically while supporting the same bounded conclusion.

Evidence receipt

PyTorch documentation and Hydra documentation define framework abstractions. LeWorldModel v3 and frozen official code 8edfeb336732b5f3ce7b8b210d0ba370a09e2cac define the inspected graph and interfaces. stable-worldmodel is evidence only for its later active platform, not retroactive evidence about original v3.

Allowed claim: these patterns preserve named data, shape, conditioning, rollout, optimization, precision, and recovery invariants under a recorded configuration. They do not show that a loader understands episodes, a causal mask discovers physics, AdaLN guarantees action use, mixed precision is harmless, or one seed proves reproducibility. Independent reproduction is not established here.

Quick check

  1. Which layer must reject a window that crosses an episode boundary?
  2. What does the later-token mutation test prove—and what does it leave open?
  3. During CEM, which object changes and which model parts stay frozen?