JEPA4Japan · tutorials

Chapter 20 — Claims the evidence does not yet earn

778 words 4 min read #LeVJEPA#JEPA#self-supervised video#SIGReg

Mark the preprint status, missing independent reproduction, limited scale, motion weakness, dense-task gap, and theoretical approximation.

Course progress Course outline 34 of 34 lessons available

Part 0 — Get the map

  1. 01 Chapter 0 — Before you begin: what this course promises available now
  2. 02 Chapter 1 — One video, two windows available now
  3. 03 Chapter 2 — A walk along Yann LeCun’s research road available now
  4. 04 Chapter 3 — The JEPA family, without the name soup available now

Part 1 — Why a small objective can learn to see

  1. 05 Chapter 4 — Video can set its own homework available now
  2. 06 Chapter 5 — Keep the meaning; do not repaint every pixel available now
  3. 07 Chapter 6 — Match the cards, but do not leave every card blank available now
  4. 08 Chapter 7 — SIGReg checks a cloud by looking at its shadows available now
  5. 09 Chapter 8 — The whole LeVJEPA objective on one line available now

Part 2 — Send a video through one encoder

  1. 10 Chapter 9 — How global and local views are paired available now
  2. 11 Chapter 10 — Cut a video into space-time tiles available now
  3. 12 Chapter 11 — One encoder, one projector, one summary card available now
  4. 13 Chapter 12 — One complete trip through the model available now
  5. 14 Chapter 13 — Why throwing away 95% can help available now
  6. 15 Chapter 14 — Same-frame teamwork, no peeking into tomorrow available now
  7. 16 Chapter 15 — RoPE, single-frame tubelets, and unexpectedly useful patch features available now

Part 3 — Read the experiments, not just the headline

  1. 17 Chapter 16 — What the four ablation ladders actually test available now
  2. 18 Chapter 17 — Equal epochs are not equal bills available now
  3. 19 Chapter 18 — What ImageNet, K400, and SSv2 are really asking available now
  4. 20 Chapter 19 — Keep the paper’s results in a ledger available now
  5. 21 Chapter 20 — Claims the evidence does not yet earn Current lesson

Part 4 — From the official repository to your own experiment

  1. 22 Chapter 21 — A map of the official repository available now
  2. 23 Chapter 22 — Ten long walks become a training set available now
  3. 24 Chapter 23 — Read the defaults, then start training available now
  4. 25 Chapter 24 — Run a smoke test that cannot flatter you available now
  5. 26 Chapter 25 — Skip training: extract features from the public checkpoint available now
  6. 27 Chapter 26 — Freeze the encoder and test your own videos available now

Part 5 — Put the representation back on the world-model road

  1. 28 Chapter 27 — The important boundary: an encoder is not a planner available now
  2. 29 Chapter 28 — How LeVJEPA might feed a future world model available now
  3. 30 Chapter 29 — Ten projects, from first experiment to paper-sized question available now

Appendices — A backpack for the trail

  1. 31 Appendix A — The smallest useful math kit available now
  2. 32 Appendix B — The complete tensor-shape table available now
  3. 33 Appendix C — Glossary and paper timeline available now
  4. 34 Appendix D — Reproduction and review checklist available now

Give every sentence the right-sized boat

  1. Safe to sayobserved in protocol
  2. Add conditionsauthor-reported
  3. Evidence stopsdo not upgrade
  4. Test nextreproduction and new tasks
Good research writing does not inflate a claim. It builds a sentence exactly large enough for its evidence.

A paper boat floats beautifully in a bathtub. We may say, “It floated in this bath under this load.” We may not book it for an ocean crossing. Water, weight, and experimenter tell the reader where trust ends and the next test begins.

The claim traffic light

At the 2026-09-05 cutoff, LeVJEPA remains arXiv v1 with author-published code and weights.

TopicEvidence permitsEvidence does not permit
Publication“author-reported arXiv v1 result”“peer-reviewed confirmation” or “independently reproduced”
Simpler graphno target encoder, predictor, or stop-gradient in the objective; EMA only for evaluation weights“no EMA exists at all,” or drawing evaluation EMA as a teacher
Efficiencyon 20% K710 for 240 epochs, authors report 5.6–20.8× fewer total pretraining FLOPs than V-JEPA 2“20.8× faster on every machine” or equal inference savings
Causal attentionpatch tokens cannot read future frames; stated IN1K probe gives block causal 51.2 vs full 50.7“lossless for every task,” “understands causality,” or “plans”
MotionFLOP-match SSv2 is 40.4 versus best 43.6; longer training relieves some high-drop damage“best on motion across the board”
Dense featuresselected PCA/cosine maps show semantic and spatial organization“segmentation, detection, or tracking validated”
Data and scalecontrolled tests through ViT-L on limited K710; separate 1.806M-clip VideoMix scale-up“proven at internet scale or arbitrarily larger model/batch”
World modelscausal encoder may support streaming perception or a future world model“LeVJEPA is an action-conditioned world model, controller, or planner”

Three fine-print boundaries people miss

Theory. LeVJEPA invokes LeJEPA’s result about isotropic Gaussians and collapse under stated assumptions. Practical SIGReg uses finite random sketches and numerical integration—officially 1024 directions and 17 nodes on [0,3]. We may call it a theoretically grounded anti-collapse objective. We may not claim every finite batch is proved exactly Gaussian.

Recipe. The ablation default is 20% K710 with V=4. Epoch matching uses 240 epochs and batch 3072. FLOP-matched LeVJEPA alone uses 1085 epochs and V=10. The repository default is Walking Tours for 26 epochs with V=10. These identities do not commute.

Version. ArXiv v1, the README, and the model card report VideoMix IN1K at 69.5; the project page still says 67.5. This edition follows v1’s 69.5, records the conflict, and pins the repository commit. A future update should save version, commit, and access date again.

The paper’s own open gaps deserve equally plain language: short schedules plus aggressive dropping lose motion clues; highly sparse sampling that preserves temporal correspondence remains open; behavior beyond limited corpora and ViT-L, and the interaction of SIGReg with very large models/batches, remain unknown; dense downstream tasks lack quantitative evaluation. Public weights prove that files are available, not that a third party reproduced the training curves.

Rewrite the advertisement

“LeVJEPA proves video models understand causality without loss and can plan” outruns the evidence. Try: “In one arXiv v1 IN1K frozen-probe setting, the authors report block-causal encoding matching or slightly exceeding full attention. The model has no action-conditioned dynamics or planner; planning remains future work.”

Boundary documents

The honest boundary

  1. These are preprint author reports—not this course’s reproduction or a peer-reviewed conclusion.
  2. Causal patch features are not causal understanding, action prediction, control, or planning.
  3. Motion, internet scale, very large models/batches, and dense tasks retain explicit evidence gaps.

Red-light check

  1. Does downloading public weights reproduce pretraining?
  2. Does the block-causal encoder contain an action-conditioned predictor?
  3. Does a crisp patch PCA image prove tracking performance?
Answers
  1. No. Retrieving author weights and independently retraining and verifying curves are different acts.
  2. No. LeVJEPA has no action input, dynamics predictor, or planner.
  3. No. The paper supplies qualitative structure and does not evaluate segmentation or tracking.