JEPA4Japan · tutorials

Chapter 19 — Keep the paper’s results in a ledger

928 words 5 min read #LeVJEPA#JEPA#self-supervised video#SIGReg

Attach 5.6–20.8×, 61.0, 40.4, 44.6, 69.5, and 55.0 to the right model, data, budget, and evaluation protocol.

Course progress Course outline 34 of 34 lessons available

Part 0 — Get the map

  1. 01 Chapter 0 — Before you begin: what this course promises available now
  2. 02 Chapter 1 — One video, two windows available now
  3. 03 Chapter 2 — A walk along Yann LeCun’s research road available now
  4. 04 Chapter 3 — The JEPA family, without the name soup available now

Part 1 — Why a small objective can learn to see

  1. 05 Chapter 4 — Video can set its own homework available now
  2. 06 Chapter 5 — Keep the meaning; do not repaint every pixel available now
  3. 07 Chapter 6 — Match the cards, but do not leave every card blank available now
  4. 08 Chapter 7 — SIGReg checks a cloud by looking at its shadows available now
  5. 09 Chapter 8 — The whole LeVJEPA objective on one line available now

Part 2 — Send a video through one encoder

  1. 10 Chapter 9 — How global and local views are paired available now
  2. 11 Chapter 10 — Cut a video into space-time tiles available now
  3. 12 Chapter 11 — One encoder, one projector, one summary card available now
  4. 13 Chapter 12 — One complete trip through the model available now
  5. 14 Chapter 13 — Why throwing away 95% can help available now
  6. 15 Chapter 14 — Same-frame teamwork, no peeking into tomorrow available now
  7. 16 Chapter 15 — RoPE, single-frame tubelets, and unexpectedly useful patch features available now

Part 3 — Read the experiments, not just the headline

  1. 17 Chapter 16 — What the four ablation ladders actually test available now
  2. 18 Chapter 17 — Equal epochs are not equal bills available now
  3. 19 Chapter 18 — What ImageNet, K400, and SSv2 are really asking available now
  4. 20 Chapter 19 — Keep the paper’s results in a ledger Current lesson
  5. 21 Chapter 20 — Claims the evidence does not yet earn available now

Part 4 — From the official repository to your own experiment

  1. 22 Chapter 21 — A map of the official repository available now
  2. 23 Chapter 22 — Ten long walks become a training set available now
  3. 24 Chapter 23 — Read the defaults, then start training available now
  4. 25 Chapter 24 — Run a smoke test that cannot flatter you available now
  5. 26 Chapter 25 — Skip training: extract features from the public checkpoint available now
  6. 27 Chapter 26 — Freeze the encoder and test your own videos available now

Part 5 — Put the representation back on the world-model road

  1. 28 Chapter 27 — The important boundary: an encoder is not a planner available now
  2. 29 Chapter 28 — How LeVJEPA might feed a future world model available now
  3. 30 Chapter 29 — Ten projects, from first experiment to paper-sized question available now

Appendices — A backpack for the trail

  1. 31 Appendix A — The smallest useful math kit available now
  2. 32 Appendix B — The complete tensor-shape table available now
  3. 33 Appendix C — Glossary and paper timeline available now
  4. 34 Appendix D — Reproduction and review checklist available now

A score needs four tickets

  1. Copy the valueno conclusion yet
  2. Attach protocoldata and budget
  3. Attach probehow features were read
  4. Name the reporterauthor-reported
  5. Then quote itlimits travel with the number
“61.0” alone says nothing. Who obtained it, on which data and budget, with which probe—that is the result.

Three voices shout “61.0,” “20.8 times,” and “69.5.” Nobody says which exam was taken. Open a ledger and give every number four ticket slots: training data, compute budget, evaluation probe, and reporter. Incomplete entries stay in pencil.

First stamp every page: the following are author-reported results from the 2026-08-27 LeVJEPA arXiv v1 preprint. This course has not repeated the large training runs. A downloadable checkpoint is not an independent reproduction.

Ledger A: equal data passes, unequal compute

The authors retrain methods from official implementations and recommended settings on the same 20% K710, for 240 epochs, effective batch 3072, matching evaluation token counts.

ScaleReported LeVJEPA comparison with V-JEPA 2Condition that must travel with it
ViT-S20.8× less total pretraining compute, comparable or better accuracyepoch-matched, not FLOP-matched
ViT-B4.8 vs 36.4 ExaFLOPs, accuracy within one pointsame data and 240 epochs
ViT-L5.6× less compute and 1.9 points higher IN1Ksame controlled setting; LeVJEPA-L still uses under half V-JEPA 2-S compute

The 5.6–20.8× range refers to total pretraining FLOPs across S/B/L. It is not a multiplier for wall-clock speed, memory, inference, or accuracy.

Ledger B: equal total FLOPs

This table fixes ViT-B, 20% K710, and total pretraining computation. Cheap examples let LeVJEPA run 1085 epochs with V=10. IN1K and SSv2 use frozen attentive-probe top-1; K400 uses frozen mean pooling plus linear-probe top-1.

MethodIN1KSSv2K400
VideoMAEv253.443.637.4
V-JEPA 251.642.540.7
LeVJEPA61.040.444.6

LeVJEPA’s 61.0 is 7.6 percentage points above the next IN1K result; it also leads K400. On SSv2 it is 3.2 points below the best 43.6. “Competitive motion information” is more faithful than “best at motion.”

Ledger C: video pretraining versus image pretraining

For ViT-B, the comparison uses frames from the same source videos and equal total FLOPs. DINOv2 trains on single frames with its official implementation—11.7M frame samples and 11,400 optimizer steps—while LeVJEPA uses its 240-epoch setting. Both receive frozen attentive probes.

MethodIN1KSSv2
DINOv253.816.9
LeVJEPA50.730.4

DINOv2 leads IN1K by 3.1 points; LeVJEPA’s SSv2 is about 1.8× the DINOv2 value. Neither method “wins everything,” and two benchmarks do not establish general world understanding.

Ledger D: a consumer run and VideoMix scale-up

ExperimentRecipeAuthor-reported resultBoundary
ConsumerViT-Tiny; one RTX 5080 16GB; 12 hours; eight Walking Tours, about 620k frames and roughly 5M processed clipsfrozen IN1K 8.9 → 25.2; batch 128 below 8GB, same-size V-JEPA fills memory at batch 28feasibility example, not a cluster-scale SOTA comparison
VideoMixViT-L/16; 100 epochs; K710 + SSv2 + Walking Tours + PE-Video; model card lists 1,806,869 clipsarXiv v1, repository README, and model card say IN1K 69.5, SSv2 55.0, frozen attentive probelarger mixed data; not a single-variable comparison with 20% K710

One correction card must remain visible. At the 2026-09-05 cutoff, the official project page still displays 67.5/55.0, while arXiv v1, the official README, and model card report 69.5/55.0. Because this course is anchored to v1, it uses 69.5 and records the disagreement rather than silently choosing a convenient value.

Patch PCA and query-similarity pictures are qualitative entries, not segmentation or tracking scores. None of these ledgers contains action-conditioned prediction or planning experiments.

Complete the 61.0 receipt

Write the whole record: author-reported; arXiv v1; ViT-B; 20% K710; equal total FLOPs; 1085 epochs; V=10; IN1K frozen attentive-probe top-1. If those labels do not fit in your sentence, the bare number should not leave the ledger.

Original records

Close the ledger

  1. Every result here is author-reported preprint evidence, not this course’s reproduction.
  2. 61.0/40.4/44.6 belongs to the FLOP-matched table; 69.5/55.0 belongs to the larger VideoMix ViT-L.
  3. Strong appearance or K400 results do not imply the strongest SSv2 result, much less planning.

Receipt check

  1. Do 61.0, 40.4, and 44.6 all use the same probe?
  2. Does 69.5/55.0 come from 20% K710 alone?
  3. May this course call the tables an independent reproduction?
Answers
  1. No. The first two use attentive probes; K400 uses mean pooling and a linear probe.
  2. No. VideoMix includes K710, SSv2, Walking Tours, and PE-Video.
  3. No. They are author-reported in arXiv v1.