Course progress Course outline 34 of 34 lessons available
Part 0 — Get the map
Part 1 — Why a small objective can learn to see
- 05 Chapter 4 — Video can set its own homework available now
- 06 Chapter 5 — Keep the meaning; do not repaint every pixel available now
- 07 Chapter 6 — Match the cards, but do not leave every card blank available now
- 08 Chapter 7 — SIGReg checks a cloud by looking at its shadows Current lesson
- 09 Chapter 8 — The whole LeVJEPA objective on one line available now
Part 2 — Send a video through one encoder
- 10 Chapter 9 — How global and local views are paired available now
- 11 Chapter 10 — Cut a video into space-time tiles available now
- 12 Chapter 11 — One encoder, one projector, one summary card available now
- 13 Chapter 12 — One complete trip through the model available now
- 14 Chapter 13 — Why throwing away 95% can help available now
- 15 Chapter 14 — Same-frame teamwork, no peeking into tomorrow available now
- 16 Chapter 15 — RoPE, single-frame tubelets, and unexpectedly useful patch features available now
Part 3 — Read the experiments, not just the headline
- 17 Chapter 16 — What the four ablation ladders actually test available now
- 18 Chapter 17 — Equal epochs are not equal bills available now
- 19 Chapter 18 — What ImageNet, K400, and SSv2 are really asking available now
- 20 Chapter 19 — Keep the paper’s results in a ledger available now
- 21 Chapter 20 — Claims the evidence does not yet earn available now
Part 4 — From the official repository to your own experiment
- 22 Chapter 21 — A map of the official repository available now
- 23 Chapter 22 — Ten long walks become a training set available now
- 24 Chapter 23 — Read the defaults, then start training available now
- 25 Chapter 24 — Run a smoke test that cannot flatter you available now
- 26 Chapter 25 — Skip training: extract features from the public checkpoint available now
- 27 Chapter 26 — Freeze the encoder and test your own videos available now
Part 5 — Put the representation back on the world-model road
Appendices — A backpack for the trail
Inspect the cloud by its shadows
- Feature cloudmany dimensions
- Turn the lightrandom unit direction
- Read the shadowdoes it look bell-shaped?
- Try againmany directions, rounder cloud
There is a tangled pile of blocks on a table, but its full shape is hard to see. Shine a lamp from different sides. A shadow that shrinks to a dot reveals a flattened direction. Many smooth, centered, bell-shaped shadows give you more confidence that the pile has not been squeezed into a crack.
An afternoon permits only finitely many lamp positions. One small pile cannot check every direction in the universe.
From the lamp to SIGReg
Each view produces a K-dimensional embedding. A batch of them forms the cloud. SIGReg stands for Sketched Isotropic Gaussian Regularization. “Sketched” means checking lower-dimensional summaries. The intended shape is the standard isotropic Gaussian N(0, I): centered at zero, unit scale in every unit direction, with no direction crushed flat.
The Cramér–Wold idea supplies the bridge. If a high-dimensional random variable has the right one-dimensional distribution along every direction, then its joint distribution is determined; in this case it is the standard isotropic Gaussian.
An implementation cannot test every direction. It samples unit vectors a_m and compresses embedding z_i to the scalar inner product <z_i, a_m>. It then compares an empirical characteristic function—a collection of sine-and-cosine fingerprints—with the analytic fingerprint of a standard normal. The paper uses an Epps–Pulley-style statistic, approximates an integral numerically, and aggregates across directions and views.
Why does this notice collapse? If every embedding is equal, its shadows have almost no width, unlike a unit-variance Gaussian. If the cloud lives in a low-dimensional sheet, directions perpendicular to that sheet expose near-zero variance. The penalty sends gradients through the embeddings into both projector and encoder. It needs neither negative examples nor a teacher branch.
Now the boundary. The theorem concerns an ideal distribution under its assumptions. A training step sees a finite batch, finitely many random directions, and finitely many quadrature points; optimization may not reach its target either. A low observed SIGReg loss is not proof of an exactly Gaussian population in every direction. It certainly is not proof that coordinates correspond to true physical state. It is a theoretically motivated pressure against collapse.
Try a deceptive cloud
Draw ten dots on one horizontal line. A horizontal lamp sees a broad shadow; a vertical lamp sees one point. One direction can miss dimensional collapse. More random directions improve the chance of finding it, but finite inspection remains an approximation.
Paper trail
- LeJEPA v3: SIGReg, Cramér–Wold, and the risk argument
- Pinned LeJEPA repository: finite-projection implementation
- LeVJEPA v1: view-wise SIGReg with quadrature
- Pinned LeVJEPA
module.py
Three shadows to remember
- SIGReg uses random one-dimensional projections to test whether high-dimensional embeddings approach an isotropic Gaussian.
- Constant points and flattened clouds leave abnormal shadows in some directions.
- Finite batches, directions, and quadrature approximate the theory; a low loss is not a full distributional proof.
Shine the light
- What does “sketched” mean here?
- How could a horizontal-only check miss collapse?
- Does low SIGReg establish that the encoder learned physics?
Answers
- Using random low-dimensional projections as compact views of a high-dimensional distribution.
- The cloud may be broad horizontally but crushed vertically.
- No. It only supports non-degeneration under that particular check.