JEPA4Japan · tutorials

Chapter 23 — Read the defaults, then start training

862 words 4 min read #LeVJEPA#JEPA#self-supervised video#SIGReg

Separate the paper’s comparison recipe from repository defaults, including batch size, learning rate, warmup, EMA checkpoints, and Hydra overrides.

Course progress Course outline 34 of 34 lessons available

Part 0 — Get the map

  1. 01 Chapter 0 — Before you begin: what this course promises available now
  2. 02 Chapter 1 — One video, two windows available now
  3. 03 Chapter 2 — A walk along Yann LeCun’s research road available now
  4. 04 Chapter 3 — The JEPA family, without the name soup available now

Part 1 — Why a small objective can learn to see

  1. 05 Chapter 4 — Video can set its own homework available now
  2. 06 Chapter 5 — Keep the meaning; do not repaint every pixel available now
  3. 07 Chapter 6 — Match the cards, but do not leave every card blank available now
  4. 08 Chapter 7 — SIGReg checks a cloud by looking at its shadows available now
  5. 09 Chapter 8 — The whole LeVJEPA objective on one line available now

Part 2 — Send a video through one encoder

  1. 10 Chapter 9 — How global and local views are paired available now
  2. 11 Chapter 10 — Cut a video into space-time tiles available now
  3. 12 Chapter 11 — One encoder, one projector, one summary card available now
  4. 13 Chapter 12 — One complete trip through the model available now
  5. 14 Chapter 13 — Why throwing away 95% can help available now
  6. 15 Chapter 14 — Same-frame teamwork, no peeking into tomorrow available now
  7. 16 Chapter 15 — RoPE, single-frame tubelets, and unexpectedly useful patch features available now

Part 3 — Read the experiments, not just the headline

  1. 17 Chapter 16 — What the four ablation ladders actually test available now
  2. 18 Chapter 17 — Equal epochs are not equal bills available now
  3. 19 Chapter 18 — What ImageNet, K400, and SSv2 are really asking available now
  4. 20 Chapter 19 — Keep the paper’s results in a ledger available now
  5. 21 Chapter 20 — Claims the evidence does not yet earn available now

Part 4 — From the official repository to your own experiment

  1. 22 Chapter 21 — A map of the official repository available now
  2. 23 Chapter 22 — Ten long walks become a training set available now
  3. 24 Chapter 23 — Read the defaults, then start training Current lesson
  4. 25 Chapter 24 — Run a smoke test that cannot flatter you available now
  5. 26 Chapter 25 — Skip training: extract features from the public checkpoint available now
  6. 27 Chapter 26 — Freeze the encoder and test your own videos available now

Part 5 — Put the representation back on the world-model road

  1. 28 Chapter 27 — The important boundary: an encoder is not a planner available now
  2. 29 Chapter 28 — How LeVJEPA might feed a future world model available now
  3. 30 Chapter 29 — Ten projects, from first experiment to paper-sized question available now

Appendices — A backpack for the trail

  1. 31 Appendix A — The smallest useful math kit available now
  2. 32 Appendix B — The complete tensor-shape table available now
  3. 33 Appendix C — Glossary and paper timeline available now
  4. 34 Appendix D — Reproduction and review checklist available now

A config is a receipt, not decoration

  1. Hydra recipeevery knob has a name
  2. Many views1 global + 10 local
  3. ViT-Bsparse and block causal
  4. Two lossesλ stays 0.02
  5. Two weight setsonline + evaluation EMA
Change one training knob and you have made a new recipe. Save it as such.

A banquet recipe says “make 3,072 servings.” That number assumes two kitchens, eight ovens each, 96 servings per oven, and two rounds of accumulation. One home oven can test whether the batter bakes. It cannot honestly claim to have reproduced the banquet.

Read the public default end to end

The pinned conf/config.yaml trains ViT-B/16 on ten Walking Tours videos: 16 frames, stride 2, tubelet 1, 95% random dropping, and block_causal attention. Each clip produces one 224 global crop and ten 96 local crops. The projector is 768 → 2048 → 256. SIGReg uses 17 knots and 1,024 random projections; its weight remains 0.02.

AdamW uses learning rate 4e-4, weight decay 0.04, and betas (0.9, 0.95). The first 1,200 optimizer steps warm from 1e-4 to 4e-4; because end_lr == lr, the schedule stays flat afterward. The default runs 26 epochs, noted as roughly ten thousand optimizer steps.

Check the graph beside the recipe. Global and ten locals use one encoder. Locals reshape from [B,V,T,C,H,W] to [B×V,C,T,H,W]; eleven [CLS] summaries pass through one projector to [B,11,256]. The code evaluates (global_emb - embeddings).pow(2).mean()—including one harmless global-self zero—and sends [11,B,256] to SIGReg. Both losses update the shared encoder. There is no teacher forward.

Batch size is the easiest number to misquote. The comment assumes 2 nodes × 8 GPUs × 96 clips × accumulation 2 = 3,072. Move to different hardware without changing accumulation and the experiment has changed. EMA tracks only the encoder, updates every 32 optimizer steps at decay 0.9999, and saves as state_dict_ema. Online weights live in state_dict; the average neither enters the loss nor acts as a target.

Once data exists, a single-device, two-batch smoke run can look like this:

uv run python main.py \
  trainer.devices=1 trainer.num_nodes=1 trainer.strategy=auto \
  trainer.max_epochs=1 trainer.limit_train_batches=2 \
  loader.batch_size=2 loader.num_workers=0 accumulate_grad_batches=1 \
  model.name=vit_tiny augmentation.local_crops_number=2

This is a tiny effective-batch-2 pipeline check, not the default or a paper reproduction. Hydra writes run metadata under outputs/<date>/<time>/. But the pinned checkpoint.dirpath=checkpoints resolves from the launch directory because hydra.job.chdir=false, so checkpoints normally land in repository-root checkpoints/. A config comment suggesting the Hydra run directory conflicts with that behavior. Only explicit hydra.job.chdir=true makes a relative path follow the run directory.

Before a real launch, run uv run python main.py --cfg job and save the resolved YAML. Confirm data.train. Walking Tours has no held-out validation set: val points to train while trainer.limit_val_batches=0 disables validation. Training loss is therefore no generalization metric. Checkpoints save every 2,000 train steps and at the end; inspect both state_dict and state_dict_ema, and record actual world size.

Keep four recipes separate: the paper’s 20% K710, 240-epoch control; public ViT-B Walking Tours, 26-epoch default; 12-hour ViT-Tiny consumer demo on eight videos; released ViT-L trained on 1.8M VideoMix clips. Method identity does not make training conditions identical.

Finally, resume.weights_only=true loads weights but restarts optimizer and schedule. Set it to false to resume optimizer state and step count, and retain the original resolved config.

Configuration sources

Keep the receipt

  1. Effective batch 3,072 assumes 2×8 GPUs, 96 clips per GPU, and accumulation 2.
  2. A two-batch local command checks plumbing; it inherits no paper result.
  3. Shared method does not mean shared data, model, schedule, or budget.

Recipe questions

  1. With one eight-GPU node and 96 clips per GPU, what accumulation keeps batch 3,072?
  2. Does state_dict_ema provide an online encoder target?
  3. Can the Walking Tours default reproduce the 20% K710 table?
Answers
  1. Accumulation 4.
  2. No. It is an evaluation weight average.
  3. No. The data and training protocol differ.