JEPA4Japan · tutorials

Chapter 24 — Run a smoke test that cannot flatter you

779 words 4 min read #LeVJEPA#JEPA#self-supervised video#SIGReg

Test tiny-batch overfitting, view agreement, SIGReg statistics, future leakage, and memory before spending a full run.

Course progress Course outline 34 of 34 lessons available

Part 0 — Get the map

  1. 01 Chapter 0 — Before you begin: what this course promises available now
  2. 02 Chapter 1 — One video, two windows available now
  3. 03 Chapter 2 — A walk along Yann LeCun’s research road available now
  4. 04 Chapter 3 — The JEPA family, without the name soup available now

Part 1 — Why a small objective can learn to see

  1. 05 Chapter 4 — Video can set its own homework available now
  2. 06 Chapter 5 — Keep the meaning; do not repaint every pixel available now
  3. 07 Chapter 6 — Match the cards, but do not leave every card blank available now
  4. 08 Chapter 7 — SIGReg checks a cloud by looking at its shadows available now
  5. 09 Chapter 8 — The whole LeVJEPA objective on one line available now

Part 2 — Send a video through one encoder

  1. 10 Chapter 9 — How global and local views are paired available now
  2. 11 Chapter 10 — Cut a video into space-time tiles available now
  3. 12 Chapter 11 — One encoder, one projector, one summary card available now
  4. 13 Chapter 12 — One complete trip through the model available now
  5. 14 Chapter 13 — Why throwing away 95% can help available now
  6. 15 Chapter 14 — Same-frame teamwork, no peeking into tomorrow available now
  7. 16 Chapter 15 — RoPE, single-frame tubelets, and unexpectedly useful patch features available now

Part 3 — Read the experiments, not just the headline

  1. 17 Chapter 16 — What the four ablation ladders actually test available now
  2. 18 Chapter 17 — Equal epochs are not equal bills available now
  3. 19 Chapter 18 — What ImageNet, K400, and SSv2 are really asking available now
  4. 20 Chapter 19 — Keep the paper’s results in a ledger available now
  5. 21 Chapter 20 — Claims the evidence does not yet earn available now

Part 4 — From the official repository to your own experiment

  1. 22 Chapter 21 — A map of the official repository available now
  2. 23 Chapter 22 — Ten long walks become a training set available now
  3. 24 Chapter 23 — Read the defaults, then start training available now
  4. 25 Chapter 24 — Run a smoke test that cannot flatter you Current lesson
  5. 26 Chapter 25 — Skip training: extract features from the public checkpoint available now
  6. 27 Chapter 26 — Freeze the encoder and test your own videos available now

Part 5 — Put the representation back on the world-model road

  1. 28 Chapter 27 — The important boundary: an encoder is not a planner available now
  2. 29 Chapter 28 — How LeVJEPA might feed a future world model available now
  3. 30 Chapter 29 — Ten projects, from first experiment to paper-sized question available now

Appendices — A backpack for the trail

  1. 31 Appendix A — The smallest useful math kit available now
  2. 32 Appendix B — The complete tensor-shape table available now
  3. 33 Appendix C — Glossary and paper timeline available now
  4. 34 Appendix D — Reproduction and review checklist available now

Five gates before the expensive run

  1. Shapeslabel every box
  2. Finite valuesreject NaN and Inf
  3. Gradientsparameters really move
  4. No peekingfuture cannot alter past patches
  5. Resource logmemory and throughput
A smoke test finds obvious leaks in the pipeline. It does not validate the paper’s conclusions.

A new toy train takes two laps before a cross-country journey. Count the carriages. See whether the wheels turn. Verify the red signal stops it. Two laps do not prove national reliability; they catch crossed wires before the costly trip.

Gate 1: shapes

The default loader emits global [B,16,3,224,224] and local [B,10,16,3,96,96]. main.py converts them to encoder order [B,C,T,H,W]. With a 5% training keep-set, 3136 global patches become about 157 plus [CLS] = 158 tokens. Each local’s 576 becomes about 29 plus [CLS] = 30.

Gates 2 and 3: numbers and gradients

pred_loss, sigreg_loss, and total loss must be finite. After backward(), both encoder and projector should have at least one finite, nonzero gradient. Do not demand monotonically falling loss over two batches; random crops, projections, and small-batch noise will wobble.

For a tiny-batch overfit test, freeze the batch and seed, and record feature variance too. Falling MSE alone can mean collapse.

Check dtype boundaries. With normalize_on_gpu=true, workers send uint8 to save roughly fourfold host/IPC memory; main.to_float_normalized() divides by 255 and applies ImageNet normalization. Normalizing a float batch twice quietly ruins it. Log dtype, range, shape, and finiteness before the encoder: uint8 should lie in [0,255]; the model input should be finite float. Global receives spatial crop and normalization; locals add color jitter, grayscale, and flip. Their time indices must match exactly.

Gate 4: future leakage

This test changes only the final two frames. Patch tokens from the first two should remain equal. [CLS] is deliberately excluded because it reads the full clip.

uv run python - <<'PY'
import torch
from module import vit_tiny
torch.manual_seed(0)
m = vit_tiny(img_size=32, patch_size=16, num_frames=4, tubelet_size=1,
             use_rope=True, token_drop_rate=0, attn_mode="block_causal").eval()
x = torch.randn(1, 3, 4, 32, 32)
y = x.clone(); y[:, :, 2:] = torch.randn_like(y[:, :, 2:])
with torch.inference_mode():
    a, b = m(x), m(y)
d = (a[:, 1:9] - b[:, 1:9]).abs().max().item()  # 2 frames × 4 patches
print(f"causal prefix max diff: {d:.3g}")
assert d < 1e-5
PY

Gate 5: resources—and useful failure modes

Record peak device memory, step time, and actual token counts. For OOM diagnosis, reduce batch size, local crops, or switch to vit_tiny, while labeling the changed experiment. For stuck workers, set loader.num_workers=0. For NaNs, check dtype/normalization first, then learning rate, mixed precision, SIGReg batch size, and distributed synchronization. The official workshop notebook runs views, model, both losses, and distribution plots on CUDA, MPS, or CPU.

Turn success criteria into assertions: exact shapes; finite losses and gradients; over 50–100 fixed-batch steps, a lower moving-average loss without all feature standard deviations tending to zero; causal-prefix difference below tolerance; reproducible first-step loss under the same seed. SIGReg resamples directions, so a single value need not match a reference digit for digit. Persistent explosion, NaN, or a near-zero embedding spectrum is the real alarm.

Debugging sources

The smoke-test contract

  1. The loader uses [B,T,C,H,W]; the encoder uses [B,C,T,H,W].
  2. A causality test compares prefix patches; whole-clip [CLS] may change.
  3. Two successful batches prove plumbing, not accuracy or independent reproduction.

Diagnose it

  1. Roughly how many tokens does the global training view output by default?
  2. Why watch representation variance when MSE falls?
  3. Must [CLS] stay fixed after future frames change?
Answers
  1. About 158: 157 retained patches plus [CLS].
  2. A constant representation also lowers invariance MSE.
  3. No. [CLS] summarizes the full clip; frame patch prefixes carry the causal guarantee.