JEPA4Japan · tutorials

Appendix D — Reproduction and review checklist

1,843 words 9 min read #LeVJEPA#JEPA#self-supervised video#SIGReg

Record code hashes, environment, data, budget, seeds, metrics, failure cases, licenses, and claim levels.

Course progress Course outline 34 of 34 lessons available

Part 0 — Get the map

  1. 01 Chapter 0 — Before you begin: what this course promises available now
  2. 02 Chapter 1 — One video, two windows available now
  3. 03 Chapter 2 — A walk along Yann LeCun’s research road available now
  4. 04 Chapter 3 — The JEPA family, without the name soup available now

Part 1 — Why a small objective can learn to see

  1. 05 Chapter 4 — Video can set its own homework available now
  2. 06 Chapter 5 — Keep the meaning; do not repaint every pixel available now
  3. 07 Chapter 6 — Match the cards, but do not leave every card blank available now
  4. 08 Chapter 7 — SIGReg checks a cloud by looking at its shadows available now
  5. 09 Chapter 8 — The whole LeVJEPA objective on one line available now

Part 2 — Send a video through one encoder

  1. 10 Chapter 9 — How global and local views are paired available now
  2. 11 Chapter 10 — Cut a video into space-time tiles available now
  3. 12 Chapter 11 — One encoder, one projector, one summary card available now
  4. 13 Chapter 12 — One complete trip through the model available now
  5. 14 Chapter 13 — Why throwing away 95% can help available now
  6. 15 Chapter 14 — Same-frame teamwork, no peeking into tomorrow available now
  7. 16 Chapter 15 — RoPE, single-frame tubelets, and unexpectedly useful patch features available now

Part 3 — Read the experiments, not just the headline

  1. 17 Chapter 16 — What the four ablation ladders actually test available now
  2. 18 Chapter 17 — Equal epochs are not equal bills available now
  3. 19 Chapter 18 — What ImageNet, K400, and SSv2 are really asking available now
  4. 20 Chapter 19 — Keep the paper’s results in a ledger available now
  5. 21 Chapter 20 — Claims the evidence does not yet earn available now

Part 4 — From the official repository to your own experiment

  1. 22 Chapter 21 — A map of the official repository available now
  2. 23 Chapter 22 — Ten long walks become a training set available now
  3. 24 Chapter 23 — Read the defaults, then start training available now
  4. 25 Chapter 24 — Run a smoke test that cannot flatter you available now
  5. 26 Chapter 25 — Skip training: extract features from the public checkpoint available now
  6. 27 Chapter 26 — Freeze the encoder and test your own videos available now

Part 5 — Put the representation back on the world-model road

  1. 28 Chapter 27 — The important boundary: an encoder is not a planner available now
  2. 29 Chapter 28 — How LeVJEPA might feed a future world model available now
  3. 30 Chapter 29 — Ten projects, from first experiment to paper-sized question available now

Appendices — A backpack for the trail

  1. 31 Appendix A — The smallest useful math kit available now
  2. 32 Appendix B — The complete tensor-shape table available now
  3. 33 Appendix C — Glossary and paper timeline available now
  4. 34 Appendix D — Reproduction and review checklist Current lesson

Leave enough footprints to walk the route again

  1. Identitypaper version + code hash
  2. Datasource + split + fingerprint
  3. Trainingconfig + budget + seeds
  4. Evaluationprotocol + baselines + failures
  5. Receiptlogs + weights + narrow claims
“It ran” is a memory. A reproduction record lets someone else identify exactly what ran.

Two cooks both write “tomato soup.” One used canned tomatoes and simmered for two hours; the other used fresh tomatoes and a microwave for five minutes. The name alone cannot reproduce either meal.

“I trained LeVJEPA” is just as underspecified. It could mean the public Walking Tours default, the paper’s controlled 20% K710 comparison, the thousand-step teaching notebook, or loading released weights without pretraining. Each can be useful. Each supports a different claim.

Keep blank boxes blank until you measure them. A visible TODO is a better scientific record than a confident guess.

A. Identity and version

  • Record title, arXiv ID, version, and download date: 2608.27395v1, not “the latest paper.”
  • Record the project-page and code-repository URLs.
  • Pin a full 40-character commit. This course audits 3ea0dda16030bc0fd6472bf809fc0a0ad836a812 from 2026-09-02.
  • Save git status --short; archive a diff if local changes exist, and give that run a distinct name.
  • Record the checkpoint repository, revision, and file hashes.
  • Check the repository’s MIT license separately from the CC BY-NC 4.0 notices covering V-JEPA-derived module.py and the weights. Do not decide commercial eligibility from a tutorial.

The smallest version receipt looks like this:

git rev-parse HEAD
git status --short
uv --version
uv run python - <<'PY'
import platform, torch
print(platform.platform())
print(torch.__version__, torch.version.cuda)
print(torch.cuda.get_device_name(0) if torch.cuda.is_available() else "no CUDA")
PY

B. Environment and hardware

  • Save operating system, Python, PyTorch, CUDA, driver, and uv.lock hash.
  • Record GPU model and count, memory, CPU cores, RAM, and storage medium.
  • Record precision mode, such as bf16-mixed, and distributed strategy.
  • Measure peak memory, average step time, data-wait time, and total wall time.
  • If comparing FLOPs, name the counting tool, input shape, and whether targets, predictors, or probes are included.

An author-reported 12-hour RTX 5080 run is not a timer for a different 16 GB GPU. Chip, software stack, and data I/O all alter wall time.

C. Data receipt

  • Record dataset name, version, license, original URL list, and download date.
  • Hash raw videos or their manifest so later replacement or deletion is detectable.
  • Record decode size, stored frame rate, JPEG quality, episode length, and Lance build command.
  • Record valid episode count, frames per episode, filtered short segments, and total bytes.
  • State T=16, storage at 15 fps, and frame_stride=2: sampling is about 7.5 fps and covers roughly 2.1 seconds.
  • Split by independent source group—original video, person, or scene—not by random clips. Treat an episode as a split unit only when it is genuinely independent and has no near-neighbor source leakage.
  • Save a contact sheet from the first batch showing that global and local views share one temporal window.
  • Record color space, normalization mean and standard deviation, crop ranges, and photometric transforms.

A processed-data note might read:

source videos: 10
stored fps: 15
episodes: ...
frames: ...
train/val/test split unit: source-video
Lance bytes: ...
manifest sha256: ...

The ellipses mean “not measured yet,” not “use a default.”

D. Model and objective

  • State ViT size, patch_size=16, tubelet_size=1, and num_frames=16.
  • State global resolution 224, local resolution 96, local-view count V, and crop scales.
  • State token_drop_rate=0.95 and attn_mode=block_causal.
  • State projector D -> 2048 -> 256, with BatchNorm and GELU.
  • State L = L_inv + 0.02 * L_SIGReg, 17 nodes, 1,024 random directions, and the pinned default normalize_by_n=false.
  • Confirm gradients enter both global and local branches; do not add an undeclared detach.
  • Confirm there is no training-time target encoder or predictor.
  • If saving Polyak/EMA weights, record decay, update schedule, and whether raw or EMA weights are evaluated; state that the copy does not enter the loss.

E. Optimizer and budget

  • Record per-device batch, device and node counts, accumulation, and the actual effective batch.
  • Record optimizer, learning rate, betas, weight decay, and exclusions for bias or normalization parameters.
  • Record optimizer-step warmup and whether the later schedule is flat or decayed.
  • Report epochs, optimizer steps, clips seen, tokens per step, estimated total FLOPs, and wall time together.
  • Fix and report seeds for data sampling, view transforms, token dropping, initialization, and probes.
  • Run at least three training seeds or label the study honestly as one run; never present the best seed as stability evidence.
  • Save the fully resolved Hydra configuration, not only command-line overrides.

The repository comment’s effective batch of 3072 comes from:

2 nodes * 8 GPUs * 96 clips/device * 2 gradient accumulation = 3072

One GPU, batch 8, no accumulation gives an effective batch of 8. It is a smoke test, not a reproduction of the paper recipe.

F. Minimum training dashboard

Archive at least:

  • pred_loss, sigreg_loss, and weighted and unweighted total loss;
  • per-dimension CLS mean and standard deviation, covariance spectrum, minimum and maximum variance, and effective rank;
  • positive global–local distance and a shuffled-batch negative-control distance;
  • the same validation probe on raw-encoder and EMA checkpoints;
  • gradient norms, learning rate, weight decay, throughput, and peak memory;
  • a space-time plot of one batch’s randomly retained tokens;
  • a block-causal unit test in which changing future frames cannot change earlier patch tokens;
  • alerts for NaN/Inf, zero gradients, duplicate samples, and data stalls.

Total loss alone can hide pred_loss falling while SIGReg explodes, or the reverse. Keep the terms separate.

G. Climb the validation ladder in order

  1. Interface: one synthetic clip completes a forward pass; every shape and dtype is asserted.
  2. Objective: constant embeddings give low invariance but high SIGReg; shuffled pairs raise invariance.
  3. Causality: modifying only future frames leaves earlier patch outputs unchanged, while [CLS] may change.
  4. Tiny data: a very small set overfits in an interpretable way while collapse statistics remain visible.
  5. Teaching run: the official workshop runs and its images, curves, and config are saved.
  6. Single GPU: model or batch is scaled to available resources and labeled as such.
  7. Recipe: only after data, budget, effective batch, model, and evaluation protocol align should paper numbers be compared.

Stop and repair the first failed level. A large run does not cure swapped axes or future leakage.

H. Evaluation and fair comparison

  • Freeze the encoder and train only the declared linear or attentive probe.
  • Use the same split, input frames and temporal positions, probe capacity, optimization budget, and view sampling for every method.
  • Report ImageNet-1K, K400, and SSv2 separately because they favour different information.
  • Report mean, standard deviation, seed count, confidence interval, or the raw per-seed results.
  • For epoch-matched and FLOP-matched comparisons, state exactly what was held fixed; never blend them into one “faster” claim.
  • On new data, include at least one of random initialization, supervised pretraining, or a budget-matched alternative as a baseline.
  • Save both successes and failures from prediction, retrieval, or feature maps; do not publish only the prettiest PCA.
  • Treat probe decodability as “this probe can extract the information,” not “a world model will use it.”

I. Give each sentence an evidence grade

Before publishing, prefix working claims with one of these labels:

[method fact]       directly defined by a paper or pinned code
[author-reported]   measured by the authors under a named protocol
[this run]          directly present in this run's logs or artifacts
[assumption-bound]  valid only under the listed theoretical assumptions
[inference]         supported indirectly by several observations
[hypothesis]        awaiting an experiment that can refute or support it

For example:

  • [method fact] The public implementation uniformly drops 95% of patch tokens during training.
  • [author-reported] In its ImageNet ablation, accuracy rises from 33.9 to 47.6 as the drop rate moves from 0 to 0.95.
  • [hypothesis] A sampler preserving cross-frame correspondences will outperform independent random dropping on fine, fast actions.

These sentences have three different evidential strengths.

A receipt you can save with the run

experiment: levjepa-walkingtours-smoke-001
paper: arxiv:2608.27395v1
code_commit: 3ea0dda16030bc0fd6472bf809fc0a0ad836a812
local_diff_sha256: null
data:
  name: walking_tours
  manifest_sha256: TODO
  split_unit: source_video
model:
  name: vit_tiny
  frames: 16
  patch: 16
  tubelet: 1
  token_drop: 0.95
  attention: block_causal
objective:
  sigreg_weight: 0.02
  projections: 1024
  knots: 17
budget:
  devices: 1
  batch_per_device: 8
  accumulation: 1
  optimizer_steps: 100
seeds: [0]
status: smoke_test_not_paper_reproduction
artifacts:
  resolved_config: TODO
  metrics_log: TODO
  checkpoint_sha256: TODO

Keeping TODO is more honest than filling an unmeasured field.

The publication gate

Call a result a reproduction only when every answer is yes:

  1. Are paper version and code hash exact?
  2. Are data and splits traceable and free of known leakage?
  3. Do model, objective, effective batch, steps, and budget align?
  4. Are probe and baseline protocols fair?
  5. Can every seed and raw artifact be retrieved?
  6. Are results inside a declared tolerance?
  7. Does the conclusion cover only what was actually repeated?

Otherwise use the narrower name that fits: installation check, forward-pass smoke test, teaching run, partial reproduction, or downstream evaluation of an author checkpoint.

Records for the checklist

The pinned official repository defines installation, data, training, and license boundaries. Pinned conf/config.yaml defines the public default recipe. LeVJEPA v1 defines the paper experiments, and the model card defines the public checkpoint interface. The audit ladder here is this course’s recommendation, not a certification standard claimed by the paper.

Sign only what the receipt supports

  1. “It runs,” “the author checkpoint runs,” and “the paper result was independently reproduced” are different statements.
  2. A reproducible record needs data fingerprints, effective batch, budget, logs, and failures—not only a config file.
  3. Reproduction ends by shrinking the conclusion to the evidence, not by finding an impressive number.

Final inspection

  1. Is a one-GPU, batch-8, 100-step run a reproduction of the paper recipe?
  2. Why split by independent source group rather than random clips or adjacent Walking Tours episodes?
  3. Why do public code and weights not establish independent reproduction?
Check the answers
  1. No. It is a valuable smoke test, but its data, effective batch, step count, and budget do not match the paper protocol.
  2. Neighbouring clips can share almost identical frames; adjacent episodes from one long walk still share source and scene. Either split can leak near-duplicates into the test set.
  3. Both artifacts come from the authors. They improve inspectability but do not show that an independent team repeated the experiment.