JEPA4Japan · tutorials
All tutorials

Course

LeVJEPA: Let Video Be Its Own Teacher

A picture-first, evidence-conscious guide to learning video representations with one encoder, SIGReg, random token dropping, and block-causal attention.

modules
7 modules
lessons
34 lessons
estimated
18h estimated
For
For: Readers new to self-supervised video learning, plus engineers and researchers who want to inspect the paper, code, training recipe, and limits of the evidence.
Start with lesson 1

The course in one glance

  1. One clipone wide view, a few close-ups
  2. One encoderwrite a card for each view
  3. Pull togethersame clip, same idea
  4. Spread outSIGReg prevents one-answer collapse
  5. Look lessdrop 95% of video tokens
  6. Respect timethe present cannot read the future
LeVJEPA gives one encoder two clear jobs: agree about the same clip, yet keep the whole class of clips richly different.

Picture yourself watching the same street performer through two windows. The wide window shows the drummer, the crowd, and the road. A cardboard tube shows only a moving hand. A useful mind should recognize one event in both views. But if every event is summarized as “something moved,” the agreement is worthless.

That small tension is the center of LeVJEPA. An invariance loss pulls summaries of the same clip together. SIGReg stops every summary from piling onto one point. Training keeps only a small random share of the video tokens, and attention lets each frame use the present and past, never the future. The graph is short. Understanding why it works—and exactly what the evidence supports—takes more care.

Pick a route

Four promises this course keeps

  1. Names stay attached to the right method. LeVJEPA is a 2026 video-representation method. It is not LeJEPA, V-JEPA 2, or LeWorldModel.
  2. The computation graph stays honest. There is no target encoder, predictor, or stop-gradient in the training objective. The official code keeps a Polyak/EMA copy of the encoder for evaluation checkpoints; that copy is not a teacher branch.
  3. Every number keeps its receipt. “Cheaper” and “better” mean little without the dataset, model size, epoch or FLOP budget, and probe protocol.
  4. A research destination is not a reported result. A causal video encoder may be useful beneath a streaming agent or world model. LeVJEPA itself has no action-conditioned predictor, cost function, or planner.

Evidence line

The evidence cutoff for this edition is 2026-09-05 (Asia/Tokyo). The factual backbone is LeVJEPA arXiv v1, the official project page, the official repository, and the public model card. The paper appeared on 2026-08-27 and remains a preprint at this cutoff. Experimental values in this course are author-reported. Public code and weights make scrutiny possible; they do not, by themselves, constitute an independent reproduction.

Course progress

Course outline

34 of 34 lessons available

Part 0 — Get the map

Name the method, mark its boundaries, and place it on Yann LeCun’s research road before opening the engine.

  1. 01 Chapter 0 — Before you begin: what this course promises Fix the paper version and evidence date, choose a reading route, and see what you will be able to explain, run, and question. available now
  2. 02 Chapter 1 — One video, two windows Meet global and local views through a street scene, and learn why both windows must cover the same 16 frames. available now
  3. 03 Chapter 2 — A walk along Yann LeCun’s research road Travel from energy-based learning and the 2022 AMI blueprint through JEPA, I-JEPA, V-JEPA, LeJEPA, and LeVJEPA. available now
  4. 04 Chapter 3 — The JEPA family, without the name soup Tell JEPA, I-JEPA, V-JEPA, V-JEPA 2, LeJEPA, LeVJEPA, and LeWorldModel apart by their actual computation graphs. available now

Part 1 — Why a small objective can learn to see

Start with label-free learning and representation matching, then meet the collapse problem and the second force that prevents it.

  1. 05 Chapter 4 — Video can set its own homework See where supervision comes from when nobody labels the clips—and what that free signal cannot promise. available now
  2. 06 Chapter 5 — Keep the meaning; do not repaint every pixel Compare pixel reconstruction, feature prediction, and view invariance, then pin down what LeVJEPA does and does not predict. available now
  3. 07 Chapter 6 — Match the cards, but do not leave every card blank Understand why invariance is useful, why a constant answer satisfies it, and how dimensions can quietly collapse. available now
  4. 08 Chapter 7 — SIGReg checks a cloud by looking at its shadows Build an intuition for random one-dimensional projections, characteristic functions, and a Gaussian reference distribution. available now
  5. 09 Chapter 8 — The whole LeVJEPA objective on one line Unpack invariance plus SIGReg, the fixed λ=0.02 trade-off, two-way gradients, and the boundary of the theory. available now

Part 2 — Send a video through one encoder

Follow real tensors through crops, video tokens, a ViT, a projector, sparse observation, and causal attention.

  1. 10 Chapter 9 — How global and local views are paired Build one 224-pixel global view and several 96-pixel local views while preserving the same 16-frame time window. available now
  2. 11 Chapter 10 — Cut a video into space-time tiles Count the tokens made from 16 frames, 16×16 patches, per-frame tubelets, and positional coordinates. available now
  3. 12 Chapter 11 — One encoder, one projector, one summary card Separate the jobs of the shared ViT, [CLS], and two-layer projector—and see why LayerNorm makes the projection space matter. available now
  4. 13 Chapter 12 — One complete trip through the model Track global and local views, embeddings, both losses, gradient flow, and the pieces discarded after training. available now
  5. 14 Chapter 13 — Why throwing away 95% can help Separate random token dropping from tube masking, then weigh compute savings, augmentation, and lost motion clues. available now
  6. 15 Chapter 14 — Same-frame teamwork, no peeking into tomorrow Read the block-causal attention pattern, including within-frame attention, past-only time flow, and the [CLS] exception. available now
  7. 16 Chapter 15 — RoPE, single-frame tubelets, and unexpectedly useful patch features Connect 3D rotary position, tubelet size 1, and the dense features that emerge even though only [CLS] is supervised. available now

Part 3 — Read the experiments, not just the headline

Audit ablations, compute budgets, probes, and author-reported numbers before accepting an impressive score.

  1. 17 Chapter 16 — What the four ablation ladders actually test Read controlled comparisons of drop rate, random versus tube patterns, local-view count, tubelet size, and attention topology. available now
  2. 18 Chapter 17 — Equal epochs are not equal bills Distinguish epoch-matched from FLOP-matched tests, and see how an efficient method can process more batches on one budget. available now
  3. 19 Chapter 18 — What ImageNet, K400, and SSv2 are really asking Separate appearance from motion evidence, and understand frozen encoders, linear probes, and attentive probes. available now
  4. 20 Chapter 19 — Keep the paper’s results in a ledger Attach 5.6–20.8×, 61.0, 40.4, 44.6, 69.5, and 55.0 to the right model, data, budget, and evaluation protocol. available now
  5. 21 Chapter 20 — Claims the evidence does not yet earn Mark the preprint status, missing independent reproduction, limited scale, motion weakness, dense-task gap, and theoretical approximation. available now

Part 4 — From the official repository to your own experiment

Confirm the interface with public weights, then move through data, training, diagnosis, and evaluation in affordable steps.

  1. 22 Chapter 21 — A map of the official repository Locate main.py, module.py, data/loader.py, Hydra configs, two notebooks, the license, and the pinned code snapshot. available now
  2. 23 Chapter 22 — Ten long walks become a training set Follow Walking Tours through download, 15 fps extraction, Lance storage, episode boundaries, random clips, and data costs. available now
  3. 24 Chapter 23 — Read the defaults, then start training Separate the paper’s comparison recipe from repository defaults, including batch size, learning rate, warmup, EMA checkpoints, and Hydra overrides. available now
  4. 25 Chapter 24 — Run a smoke test that cannot flatter you Test tiny-batch overfitting, view agreement, SIGReg statistics, future leakage, and memory before spending a full run. available now
  5. 26 Chapter 25 — Skip training: extract features from the public checkpoint Load Hugging Face weights safely, use the right normalization and tensor layout, and avoid mistaking an encoder for a classifier. available now
  6. 27 Chapter 26 — Freeze the encoder and test your own videos Build linear and attentive probes, retrieval checks, and dense-feature diagnostics with leakage-resistant splits. available now

Part 5 — Put the representation back on the world-model road

Separate the encoder LeVJEPA has built from the prediction, action, and planning machinery that still has to be added.

  1. 28 Chapter 27 — The important boundary: an encoder is not a planner Distinguish video encoding, action-conditioned dynamics, costs, search, and closed-loop control. available now
  2. 29 Chapter 28 — How LeVJEPA might feed a future world model Add action prediction, latent rollouts, uncertainty, multiple time scales, and MPC one testable layer at a time. available now
  3. 30 Chapter 29 — Ten projects, from first experiment to paper-sized question Turn motion-aware sampling, streaming caches, SIGReg, dense tasks, scale, world-model interfaces, and reproducibility into executable studies. available now

Appendices — A backpack for the trail

Keep the mathematics, tensor shapes, vocabulary, history, and reproduction paperwork close without blocking the main walk.

  1. 31 Appendix A — The smallest useful math kit Review vectors, MSE, Gaussians, random projections, characteristic functions, Cramér–Wold, FLOPs, and gradients. available now
  2. 32 Appendix B — The complete tensor-shape table Check every step from [B,T,C,H,W] to views, sparse tokens, [CLS], projector output, and SIGReg input. available now
  3. 33 Appendix C — Glossary and paper timeline Give each term a one-line meaning, a concrete analogy, and a common mistake, then anchor the JEPA lineage in primary sources. available now
  4. 34 Appendix D — Reproduction and review checklist Record code hashes, environment, data, budget, seeds, metrics, failure cases, licenses, and claim levels. available now