Course progress Course outline 34 of 34 lessons available
Part 0 — Get the map
Part 1 — Why a small objective can learn to see
- 05 Chapter 4 — Video can set its own homework available now
- 06 Chapter 5 — Keep the meaning; do not repaint every pixel available now
- 07 Chapter 6 — Match the cards, but do not leave every card blank available now
- 08 Chapter 7 — SIGReg checks a cloud by looking at its shadows available now
- 09 Chapter 8 — The whole LeVJEPA objective on one line available now
Part 2 — Send a video through one encoder
- 10 Chapter 9 — How global and local views are paired available now
- 11 Chapter 10 — Cut a video into space-time tiles available now
- 12 Chapter 11 — One encoder, one projector, one summary card available now
- 13 Chapter 12 — One complete trip through the model available now
- 14 Chapter 13 — Why throwing away 95% can help available now
- 15 Chapter 14 — Same-frame teamwork, no peeking into tomorrow available now
- 16 Chapter 15 — RoPE, single-frame tubelets, and unexpectedly useful patch features available now
Part 3 — Read the experiments, not just the headline
- 17 Chapter 16 — What the four ablation ladders actually test available now
- 18 Chapter 17 — Equal epochs are not equal bills Current lesson
- 19 Chapter 18 — What ImageNet, K400, and SSv2 are really asking available now
- 20 Chapter 19 — Keep the paper’s results in a ledger available now
- 21 Chapter 20 — Claims the evidence does not yet earn available now
Part 4 — From the official repository to your own experiment
- 22 Chapter 21 — A map of the official repository available now
- 23 Chapter 22 — Ten long walks become a training set available now
- 24 Chapter 23 — Read the defaults, then start training available now
- 25 Chapter 24 — Run a smoke test that cannot flatter you available now
- 26 Chapter 25 — Skip training: extract features from the public checkpoint available now
- 27 Chapter 26 — Freeze the encoder and test your own videos available now
Part 5 — Put the representation back on the world-model road
Appendices — A backpack for the trail
Laps and fuel are different kinds of fairness
- Same lapsepoch matched
- Different fuelFLOPs per sample
- Same allowanceFLOP matched
- Efficient bus loops more1085 epochs
Two school buses circle a city 240 times. One burns a tank per lap; the other burns eight. Their distance matches, their bill does not. Give them equal fuel instead and the efficient bus can make many more laps.
Video pretraining comparisons have the same trap. An epoch counts passes through data. FLOPs estimate computation. Never let one impersonate the other.
The two experimental contracts
As a rough accounting identity, total training computation is sample presentations multiplied by work per presentation. LeVJEPA encodes sparse tokens during training and does not run a target encoder or predictor, so each presentation is relatively cheap.
| Contract | Held fixed | Allowed to vary | Exact v1 setting |
|---|---|---|---|
| Epoch-matched | data, epochs, effective batch, evaluation | total FLOPs | 20% K710, 240 epochs, batch 3072; baselines retrained from official implementations and recommended settings |
| FLOP-matched | data, model size, total pretraining FLOPs, frozen evaluation | epochs and view budget | ViT-B, 20% K710; LeVJEPA gets 1085 epochs and V=10 |
In the epoch-matched figure, the authors report LeVJEPA using 5.6–20.8× less compute than V-JEPA 2 across model sizes. For ViT-B, accuracy differs by under one point while total work is 4.8 versus 36.4 ExaFLOPs. At ViT-L, LeVJEPA uses 5.6× less compute and scores 1.9 points higher. This answers: “After the same 240 data laps, who paid less?”
For the FLOP-matched table, lower per-sample cost buys longer training: 1085 epochs and ten local views. The authors report IN1K 61.0, SSv2 40.4, and K400 44.6. This answers: “Given the same compute bill, who learned more under these probes?” It does not make 1085 epochs free, and the greater number of data presentations remains part of the contract.
Keep the public default on a third card: Walking Tours, 26 epochs, V=10, about 10k optimizer steps, effective batch 3072. It shares the architecture but is neither the 20% K710 240-epoch comparison nor the 1085-epoch FLOP match.
Training sparsity also does not promise 20× cheaper inference. Evaluation restores all tokens. Key/value caching under causal attention is a different potential streaming benefit.
Two 7.6s that mean different things
36.4 ÷ 4.8 ≈ 7.6 is the ViT-B epoch-matched compute ratio. In the FLOP-matched table, LeVJEPA’s IN1K score exceeds the next-highest result by 7.6 percentage points. Same digits, different units and experiments. A fair results card needs data, epochs, total FLOPs, and probe.
Budget records
- LeVJEPA v1: epoch- and FLOP-matched protocols
- Official project page: both budgets and 1085 epochs
- Pinned config: the separate 26-epoch Walking Tours default
Balance the books
- Equal epochs do not imply equal compute.
- Under equal FLOPs, LeVJEPA trades low per-step cost for 1085 epochs and
V=10. - The repository’s 26-epoch Walking Tours run is neither controlled paper protocol.
Budget quiz
- Which comparison uses 240 epochs and batch 3072?
- How long does FLOP-matched LeVJEPA train?
- Does 95% training dropping mean evaluation reads only 5%?
Answers
- Epoch-matched.
- 1085 epochs, with
V=10. - No. Evaluation and inference use the full token set.