Course progress Course outline 34 of 34 lessons available
Part 0 — Get the map
Part 1 — Why a small objective can learn to see
- 05 Chapter 4 — Video can set its own homework available now
- 06 Chapter 5 — Keep the meaning; do not repaint every pixel available now
- 07 Chapter 6 — Match the cards, but do not leave every card blank available now
- 08 Chapter 7 — SIGReg checks a cloud by looking at its shadows available now
- 09 Chapter 8 — The whole LeVJEPA objective on one line available now
Part 2 — Send a video through one encoder
- 10 Chapter 9 — How global and local views are paired available now
- 11 Chapter 10 — Cut a video into space-time tiles available now
- 12 Chapter 11 — One encoder, one projector, one summary card available now
- 13 Chapter 12 — One complete trip through the model available now
- 14 Chapter 13 — Why throwing away 95% can help available now
- 15 Chapter 14 — Same-frame teamwork, no peeking into tomorrow available now
- 16 Chapter 15 — RoPE, single-frame tubelets, and unexpectedly useful patch features available now
Part 3 — Read the experiments, not just the headline
- 17 Chapter 16 — What the four ablation ladders actually test available now
- 18 Chapter 17 — Equal epochs are not equal bills available now
- 19 Chapter 18 — What ImageNet, K400, and SSv2 are really asking available now
- 20 Chapter 19 — Keep the paper’s results in a ledger Current lesson
- 21 Chapter 20 — Claims the evidence does not yet earn available now
Part 4 — From the official repository to your own experiment
- 22 Chapter 21 — A map of the official repository available now
- 23 Chapter 22 — Ten long walks become a training set available now
- 24 Chapter 23 — Read the defaults, then start training available now
- 25 Chapter 24 — Run a smoke test that cannot flatter you available now
- 26 Chapter 25 — Skip training: extract features from the public checkpoint available now
- 27 Chapter 26 — Freeze the encoder and test your own videos available now
Part 5 — Put the representation back on the world-model road
Appendices — A backpack for the trail
A score needs four tickets
- Copy the valueno conclusion yet
- Attach protocoldata and budget
- Attach probehow features were read
- Name the reporterauthor-reported
- Then quote itlimits travel with the number
Three voices shout “61.0,” “20.8 times,” and “69.5.” Nobody says which exam was taken. Open a ledger and give every number four ticket slots: training data, compute budget, evaluation probe, and reporter. Incomplete entries stay in pencil.
First stamp every page: the following are author-reported results from the 2026-08-27 LeVJEPA arXiv v1 preprint. This course has not repeated the large training runs. A downloadable checkpoint is not an independent reproduction.
Ledger A: equal data passes, unequal compute
The authors retrain methods from official implementations and recommended settings on the same 20% K710, for 240 epochs, effective batch 3072, matching evaluation token counts.
| Scale | Reported LeVJEPA comparison with V-JEPA 2 | Condition that must travel with it |
|---|---|---|
| ViT-S | 20.8× less total pretraining compute, comparable or better accuracy | epoch-matched, not FLOP-matched |
| ViT-B | 4.8 vs 36.4 ExaFLOPs, accuracy within one point | same data and 240 epochs |
| ViT-L | 5.6× less compute and 1.9 points higher IN1K | same controlled setting; LeVJEPA-L still uses under half V-JEPA 2-S compute |
The 5.6–20.8× range refers to total pretraining FLOPs across S/B/L. It is not a multiplier for wall-clock speed, memory, inference, or accuracy.
Ledger B: equal total FLOPs
This table fixes ViT-B, 20% K710, and total pretraining computation. Cheap examples let LeVJEPA run 1085 epochs with V=10. IN1K and SSv2 use frozen attentive-probe top-1; K400 uses frozen mean pooling plus linear-probe top-1.
| Method | IN1K | SSv2 | K400 |
|---|---|---|---|
| VideoMAEv2 | 53.4 | 43.6 | 37.4 |
| V-JEPA 2 | 51.6 | 42.5 | 40.7 |
| LeVJEPA | 61.0 | 40.4 | 44.6 |
LeVJEPA’s 61.0 is 7.6 percentage points above the next IN1K result; it also leads K400. On SSv2 it is 3.2 points below the best 43.6. “Competitive motion information” is more faithful than “best at motion.”
Ledger C: video pretraining versus image pretraining
For ViT-B, the comparison uses frames from the same source videos and equal total FLOPs. DINOv2 trains on single frames with its official implementation—11.7M frame samples and 11,400 optimizer steps—while LeVJEPA uses its 240-epoch setting. Both receive frozen attentive probes.
| Method | IN1K | SSv2 |
|---|---|---|
| DINOv2 | 53.8 | 16.9 |
| LeVJEPA | 50.7 | 30.4 |
DINOv2 leads IN1K by 3.1 points; LeVJEPA’s SSv2 is about 1.8× the DINOv2 value. Neither method “wins everything,” and two benchmarks do not establish general world understanding.
Ledger D: a consumer run and VideoMix scale-up
| Experiment | Recipe | Author-reported result | Boundary |
|---|---|---|---|
| Consumer | ViT-Tiny; one RTX 5080 16GB; 12 hours; eight Walking Tours, about 620k frames and roughly 5M processed clips | frozen IN1K 8.9 → 25.2; batch 128 below 8GB, same-size V-JEPA fills memory at batch 28 | feasibility example, not a cluster-scale SOTA comparison |
| VideoMix | ViT-L/16; 100 epochs; K710 + SSv2 + Walking Tours + PE-Video; model card lists 1,806,869 clips | arXiv v1, repository README, and model card say IN1K 69.5, SSv2 55.0, frozen attentive probe | larger mixed data; not a single-variable comparison with 20% K710 |
One correction card must remain visible. At the 2026-09-05 cutoff, the official project page still displays 67.5/55.0, while arXiv v1, the official README, and model card report 69.5/55.0. Because this course is anchored to v1, it uses 69.5 and records the disagreement rather than silently choosing a convenient value.
Patch PCA and query-similarity pictures are qualitative entries, not segmentation or tracking scores. None of these ledgers contains action-conditioned prediction or planning experiments.
Complete the 61.0 receipt
Write the whole record: author-reported; arXiv v1; ViT-B; 20% K710; equal total FLOPs; 1085 epochs; V=10; IN1K frozen attentive-probe top-1. If those labels do not fit in your sentence, the bare number should not leave the ledger.
Original records
- LeVJEPA v1: comparison tables and scale-up experiments
- Official project page: tables and current 67.5 text
- Pinned README: 69.5/55.0, consumer run, and default recipe
- Official VideoMix model card: 1,806,869 clips and weights
Close the ledger
- Every result here is author-reported preprint evidence, not this course’s reproduction.
- 61.0/40.4/44.6 belongs to the FLOP-matched table; 69.5/55.0 belongs to the larger VideoMix ViT-L.
- Strong appearance or K400 results do not imply the strongest SSv2 result, much less planning.
Receipt check
- Do 61.0, 40.4, and 44.6 all use the same probe?
- Does 69.5/55.0 come from 20% K710 alone?
- May this course call the tables an independent reproduction?
Answers
- No. The first two use attentive probes; K400 uses mean pooling and a linear probe.
- No. VideoMix includes K710, SSv2, Walking Tours, and PE-Video.
- No. They are author-reported in arXiv v1.