Course progress Course outline 34 of 34 lessons available
Part 0 — Get the map
Part 1 — Why a small objective can learn to see
- 05 Chapter 4 — Video can set its own homework available now
- 06 Chapter 5 — Keep the meaning; do not repaint every pixel available now
- 07 Chapter 6 — Match the cards, but do not leave every card blank Current lesson
- 08 Chapter 7 — SIGReg checks a cloud by looking at its shadows available now
- 09 Chapter 8 — The whole LeVJEPA objective on one line available now
Part 2 — Send a video through one encoder
- 10 Chapter 9 — How global and local views are paired available now
- 11 Chapter 10 — Cut a video into space-time tiles available now
- 12 Chapter 11 — One encoder, one projector, one summary card available now
- 13 Chapter 12 — One complete trip through the model available now
- 14 Chapter 13 — Why throwing away 95% can help available now
- 15 Chapter 14 — Same-frame teamwork, no peeking into tomorrow available now
- 16 Chapter 15 — RoPE, single-frame tubelets, and unexpectedly useful patch features available now
Part 3 — Read the experiments, not just the headline
- 17 Chapter 16 — What the four ablation ladders actually test available now
- 18 Chapter 17 — Equal epochs are not equal bills available now
- 19 Chapter 18 — What ImageNet, K400, and SSv2 are really asking available now
- 20 Chapter 19 — Keep the paper’s results in a ledger available now
- 21 Chapter 20 — Claims the evidence does not yet earn available now
Part 4 — From the official repository to your own experiment
- 22 Chapter 21 — A map of the official repository available now
- 23 Chapter 22 — Ten long walks become a training set available now
- 24 Chapter 23 — Read the defaults, then start training available now
- 25 Chapter 24 — Run a smoke test that cannot flatter you available now
- 26 Chapter 25 — Skip training: extract features from the public checkpoint available now
- 27 Chapter 26 — Freeze the encoder and test your own videos available now
Part 5 — Put the representation back on the world-model road
Appendices — A backpack for the trail
The perfect score that teaches nothing
- Wide dog viewwrite one card
- Close-up earmake the cards agree
- Cheap trickwrite “thing” every time
- Collapseagreement without information
Suppose the teacher checks only whether the wide and narrow cards agree for each video. A flawless cheat is to write the number 0 for the cat, the train, and the pouring cup. Every pair matches. The loss is zero. The cards distinguish nothing.
The optimizer has not malfunctioned. The assignment allowed this answer. Once we say “pull paired views together,” we must also say how the entire class avoids crowding onto one point.
Two kinds of collapse
Invariance means a representation should stay stable under changes we consider irrelevant. LeVJEPA measures squared distance between a global embedding z_0 and each local z_v. Crops and color can differ; if the views still cover the same 16 frames, their summaries should be near. Gradients flow through both ends, so the global and local paths move together rather than acting as teacher and student.
With that term alone, the constant function z(x) = c is an exact solution: every difference is zero. This is complete collapse.
There is a quieter failure too. A vector may have hundreds of coordinates while only two independent directions vary. Other coordinates stay fixed or copy one another. This dimensional collapse leaves a representation that looks large on paper but carries little usable variation.
Other self-supervised systems fight these shortcuts with negatives, variance/covariance penalties, or asymmetric EMA-teacher and predictor designs. LeVJEPA chooses SIGReg. It looks at the cloud formed by a batch of embeddings and pushes that cloud toward a zero-mean, equal-scale, uncorrelated isotropic Gaussian. A dot or a flattened line cannot satisfy that shape.
The division of labor is clean. Invariance says how different views of one sample should relate. SIGReg asks what many samples together look like. Invariance alone permits one answer. SIGReg alone could make a beautiful cloud without matching the views of any video. The fixed weighted sum supplies both pressures.
Non-collapse is not understanding. A healthy cloud may encode wallpaper, camera model, or another shortcut. Dataset audits and downstream evaluations are still required.
Four points on paper
Write (0,0), (0,0), (0,0), (0,0). Any paired-view distance can be zero, yet the points vary in neither direction. Now write (−1,0), (1,0), (0,−1), (0,1). The cloud at least spreads out. Its shape alone still cannot tell you whether the points mean “running,” “pouring,” or “which camera filmed this.”
Evidence
- LeJEPA v3: isotropic Gaussian target and anti-collapse theory
- LeVJEPA v1: bidirectional invariance and SIGReg
- Pinned LeVJEPA implementation
The compact version
- Invariance alone gives a constant encoder a zero-loss escape route.
- SIGReg constrains the batch distribution to resist complete and dimensional collapse.
- “Did not collapse” means “did not degenerate,” not “understood the world.”
Check the shortcut
- Why can a constant encoder fool the invariance loss?
- Is gradient stopped on the global path?
- Does a well-spread cloud alone prove the model learned actions?
Answers
- Both sides of every pair produce exactly the same constant.
- No. Both sides receive gradients.
- No. Motion-specific evaluation is still necessary.