JEPA4Japan · tutorials

Chapter 16 — What the four ablation ladders actually test

709 words 4 min read #LeVJEPA#JEPA#self-supervised video#SIGReg

Read controlled comparisons of drop rate, random versus tube patterns, local-view count, tubelet size, and attention topology.

Course progress Course outline 34 of 34 lessons available

Part 0 — Get the map

  1. 01 Chapter 0 — Before you begin: what this course promises available now
  2. 02 Chapter 1 — One video, two windows available now
  3. 03 Chapter 2 — A walk along Yann LeCun’s research road available now
  4. 04 Chapter 3 — The JEPA family, without the name soup available now

Part 1 — Why a small objective can learn to see

  1. 05 Chapter 4 — Video can set its own homework available now
  2. 06 Chapter 5 — Keep the meaning; do not repaint every pixel available now
  3. 07 Chapter 6 — Match the cards, but do not leave every card blank available now
  4. 08 Chapter 7 — SIGReg checks a cloud by looking at its shadows available now
  5. 09 Chapter 8 — The whole LeVJEPA objective on one line available now

Part 2 — Send a video through one encoder

  1. 10 Chapter 9 — How global and local views are paired available now
  2. 11 Chapter 10 — Cut a video into space-time tiles available now
  3. 12 Chapter 11 — One encoder, one projector, one summary card available now
  4. 13 Chapter 12 — One complete trip through the model available now
  5. 14 Chapter 13 — Why throwing away 95% can help available now
  6. 15 Chapter 14 — Same-frame teamwork, no peeking into tomorrow available now
  7. 16 Chapter 15 — RoPE, single-frame tubelets, and unexpectedly useful patch features available now

Part 3 — Read the experiments, not just the headline

  1. 17 Chapter 16 — What the four ablation ladders actually test Current lesson
  2. 18 Chapter 17 — Equal epochs are not equal bills available now
  3. 19 Chapter 18 — What ImageNet, K400, and SSv2 are really asking available now
  4. 20 Chapter 19 — Keep the paper’s results in a ledger available now
  5. 21 Chapter 20 — Claims the evidence does not yet earn available now

Part 4 — From the official repository to your own experiment

  1. 22 Chapter 21 — A map of the official repository available now
  2. 23 Chapter 22 — Ten long walks become a training set available now
  3. 24 Chapter 23 — Read the defaults, then start training available now
  4. 25 Chapter 24 — Run a smoke test that cannot flatter you available now
  5. 26 Chapter 25 — Skip training: extract features from the public checkpoint available now
  6. 27 Chapter 26 — Freeze the encoder and test your own videos available now

Part 5 — Put the representation back on the world-model road

  1. 28 Chapter 27 — The important boundary: an encoder is not a planner available now
  2. 29 Chapter 28 — How LeVJEPA might feed a future world model available now
  3. 30 Chapter 29 — Ten projects, from first experiment to paper-sized question available now

Appendices — A backpack for the trail

  1. 31 Appendix A — The smallest useful math kit available now
  2. 32 Appendix B — The complete tensor-shape table available now
  3. 33 Appendix C — Glossary and paper timeline available now
  4. 34 Appendix D — Reproduction and review checklist available now

Change one rung at a time

  1. Fix the floordata and model
  2. Turn one knobask one question
  3. Reuse the examfrozen probe
  4. Write the differencekeep tables separate
  5. Stop at the edgelocal result, not universal law
An ablation is a ladder: hold the frame still, replace one rung, and only then attribute the change.

To improve a paper airplane, you could swap paper, clip the wings, add a weight, and throw harder. If it flies farther, you learn almost nothing. A useful test fixes the paper and throw, then changes one piece. LeVJEPA Section 4 asks questions in that spirit: how much to see, how to sample it, how many small windows to open, how many frames to fuse, and whether patches may read the future.

First, separate two meanings of “default”

IdentityData and durationMain settingsWhat it can answer
Paper ablation defaultViT-B/16, 20% K710; Section 4 unless statedV=4, random dropping; frozen attentive-probe IN1K top-1 (v1 Figure 4 once says linear)controlled local design comparisons inside the paper
Runnable repository defaultWalking Tours, 26 epochs, roughly 10k optimizer stepsViT-B/16, V=10, 95%, tubelet=1, block causal, batch 3072a public runnable recipe, not the 240/1085-epoch tables

They share many components. Running the second does not reproduce the first table.

Climb the experimental ladder

Every value below is author-reported in arXiv v1 and belongs to its own protocol.

QuestionKnob changedReported valuesNarrow conclusion earned
Is sparse viewing only an approximation?drop rate 0 → 95%IN1K 33.9 → 47.6; 90% gives 47.4IN1K improves rather than falls across this sweep
Random or tube pattern?structure of retained positionsrandom/tube: IN1K 50.7/39.6; SSv2 28.8/26.4 in the stated tubelet=2 settingwhen there is no fill-in task, hiding one location forever performs worse here
How many local views?V=4 → 10 → 1247.6 → 50.2 → 49.8improvement through ten, then saturation in this range; each local adds about 29 retained patches
Fuse two frames first?τ=2,ρ=.90 vs τ=1,ρ=.95, matched token/evaluation slotsIN1K 47.4/50.7; SSv2 28.8/30.4the objective does not require input-level two-frame aggregation
Must attention see the future?full vs block causalIN1K 50.7/51.2no measurable causal-mask penalty in this frozen-probe setting

This is not a menu whose best values can be added. Rungs may come from different sub-settings. Motion is more sensitive to extreme sparsity: under a short schedule, SSv2 falls beyond ρ>0.3. An ablation describes the neighborhood actually tested, not every future dataset, scale, and task.

Edit an overclaim

“95% dropping is best for every task because 47.6 exceeds 33.9” fails twice. Those are ImageNet values from one sweep, and the same paper reports short-schedule SSv2 damage at high dropping. A defensible sentence is: “The authors report 95% as best in the stated IN1K ablation; motion tasks show a conditional cost.”

Records

A good ablation habit

  1. Paper-ablation defaults and runnable-repository defaults are different experiment identities.
  2. Read every value with its changed knob, fixed conditions, data, and probe.
  3. An ablation supports a local empirical conclusion, not a universal theorem.

Ladder check

  1. How many local views does the controlled paper default use?
  2. Does the repository’s 26-epoch default reproduce a 240-epoch paper table?
  3. Does 51.2 versus 50.7 prove block-causal attention is better everywhere?
Answers
  1. V=4, unless otherwise stated.
  2. No. Data, schedule, and experimental purpose differ.
  3. No. It is an author-reported result in one IN1K frozen-probe protocol.