Course progress Course outline 34 of 34 lessons available
Part 0 — Get the map
Part 1 — Why a small objective can learn to see
- 05 Chapter 4 — Video can set its own homework available now
- 06 Chapter 5 — Keep the meaning; do not repaint every pixel available now
- 07 Chapter 6 — Match the cards, but do not leave every card blank available now
- 08 Chapter 7 — SIGReg checks a cloud by looking at its shadows available now
- 09 Chapter 8 — The whole LeVJEPA objective on one line available now
Part 2 — Send a video through one encoder
- 10 Chapter 9 — How global and local views are paired available now
- 11 Chapter 10 — Cut a video into space-time tiles available now
- 12 Chapter 11 — One encoder, one projector, one summary card available now
- 13 Chapter 12 — One complete trip through the model available now
- 14 Chapter 13 — Why throwing away 95% can help available now
- 15 Chapter 14 — Same-frame teamwork, no peeking into tomorrow available now
- 16 Chapter 15 — RoPE, single-frame tubelets, and unexpectedly useful patch features available now
Part 3 — Read the experiments, not just the headline
- 17 Chapter 16 — What the four ablation ladders actually test Current lesson
- 18 Chapter 17 — Equal epochs are not equal bills available now
- 19 Chapter 18 — What ImageNet, K400, and SSv2 are really asking available now
- 20 Chapter 19 — Keep the paper’s results in a ledger available now
- 21 Chapter 20 — Claims the evidence does not yet earn available now
Part 4 — From the official repository to your own experiment
- 22 Chapter 21 — A map of the official repository available now
- 23 Chapter 22 — Ten long walks become a training set available now
- 24 Chapter 23 — Read the defaults, then start training available now
- 25 Chapter 24 — Run a smoke test that cannot flatter you available now
- 26 Chapter 25 — Skip training: extract features from the public checkpoint available now
- 27 Chapter 26 — Freeze the encoder and test your own videos available now
Part 5 — Put the representation back on the world-model road
Appendices — A backpack for the trail
Change one rung at a time
- Fix the floordata and model
- Turn one knobask one question
- Reuse the examfrozen probe
- Write the differencekeep tables separate
- Stop at the edgelocal result, not universal law
To improve a paper airplane, you could swap paper, clip the wings, add a weight, and throw harder. If it flies farther, you learn almost nothing. A useful test fixes the paper and throw, then changes one piece. LeVJEPA Section 4 asks questions in that spirit: how much to see, how to sample it, how many small windows to open, how many frames to fuse, and whether patches may read the future.
First, separate two meanings of “default”
| Identity | Data and duration | Main settings | What it can answer |
|---|---|---|---|
| Paper ablation default | ViT-B/16, 20% K710; Section 4 unless stated | V=4, random dropping; frozen attentive-probe IN1K top-1 (v1 Figure 4 once says linear) | controlled local design comparisons inside the paper |
| Runnable repository default | Walking Tours, 26 epochs, roughly 10k optimizer steps | ViT-B/16, V=10, 95%, tubelet=1, block causal, batch 3072 | a public runnable recipe, not the 240/1085-epoch tables |
They share many components. Running the second does not reproduce the first table.
Climb the experimental ladder
Every value below is author-reported in arXiv v1 and belongs to its own protocol.
| Question | Knob changed | Reported values | Narrow conclusion earned |
|---|---|---|---|
| Is sparse viewing only an approximation? | drop rate 0 → 95% | IN1K 33.9 → 47.6; 90% gives 47.4 | IN1K improves rather than falls across this sweep |
| Random or tube pattern? | structure of retained positions | random/tube: IN1K 50.7/39.6; SSv2 28.8/26.4 in the stated tubelet=2 setting | when there is no fill-in task, hiding one location forever performs worse here |
| How many local views? | V=4 → 10 → 12 | 47.6 → 50.2 → 49.8 | improvement through ten, then saturation in this range; each local adds about 29 retained patches |
| Fuse two frames first? | τ=2,ρ=.90 vs τ=1,ρ=.95, matched token/evaluation slots | IN1K 47.4/50.7; SSv2 28.8/30.4 | the objective does not require input-level two-frame aggregation |
| Must attention see the future? | full vs block causal | IN1K 50.7/51.2 | no measurable causal-mask penalty in this frozen-probe setting |
This is not a menu whose best values can be added. Rungs may come from different sub-settings. Motion is more sensitive to extreme sparsity: under a short schedule, SSv2 falls beyond ρ>0.3. An ablation describes the neighborhood actually tested, not every future dataset, scale, and task.
Edit an overclaim
“95% dropping is best for every task because 47.6 exceeds 33.9” fails twice. Those are ImageNet values from one sweep, and the same paper reports short-schedule SSv2 damage at high dropping. A defensible sentence is: “The authors report 95% as best in the stated IN1K ablation; motion tasks show a conditional cost.”
Records
- LeVJEPA v1: all Section 4 ablations
- Official project page: token, view, tubelet, and attention plots
- Pinned Walking Tours default configuration
A good ablation habit
- Paper-ablation defaults and runnable-repository defaults are different experiment identities.
- Read every value with its changed knob, fixed conditions, data, and probe.
- An ablation supports a local empirical conclusion, not a universal theorem.
Ladder check
- How many local views does the controlled paper default use?
- Does the repository’s 26-epoch default reproduce a 240-epoch paper table?
- Does 51.2 versus 50.7 prove block-causal attention is better everywhere?
Answers
V=4, unless otherwise stated.- No. Data, schedule, and experimental purpose differ.
- No. It is an author-reported result in one IN1K frozen-probe protocol.