Course progress Course outline 34 of 34 lessons available
Part 0 — Get the map
Part 1 — Why a small objective can learn to see
- 05 Chapter 4 — Video can set its own homework available now
- 06 Chapter 5 — Keep the meaning; do not repaint every pixel available now
- 07 Chapter 6 — Match the cards, but do not leave every card blank available now
- 08 Chapter 7 — SIGReg checks a cloud by looking at its shadows available now
- 09 Chapter 8 — The whole LeVJEPA objective on one line available now
Part 2 — Send a video through one encoder
- 10 Chapter 9 — How global and local views are paired available now
- 11 Chapter 10 — Cut a video into space-time tiles available now
- 12 Chapter 11 — One encoder, one projector, one summary card available now
- 13 Chapter 12 — One complete trip through the model available now
- 14 Chapter 13 — Why throwing away 95% can help available now
- 15 Chapter 14 — Same-frame teamwork, no peeking into tomorrow available now
- 16 Chapter 15 — RoPE, single-frame tubelets, and unexpectedly useful patch features available now
Part 3 — Read the experiments, not just the headline
- 17 Chapter 16 — What the four ablation ladders actually test available now
- 18 Chapter 17 — Equal epochs are not equal bills available now
- 19 Chapter 18 — What ImageNet, K400, and SSv2 are really asking available now
- 20 Chapter 19 — Keep the paper’s results in a ledger available now
- 21 Chapter 20 — Claims the evidence does not yet earn available now
Part 4 — From the official repository to your own experiment
- 22 Chapter 21 — A map of the official repository available now
- 23 Chapter 22 — Ten long walks become a training set available now
- 24 Chapter 23 — Read the defaults, then start training available now
- 25 Chapter 24 — Run a smoke test that cannot flatter you available now
- 26 Chapter 25 — Skip training: extract features from the public checkpoint available now
- 27 Chapter 26 — Freeze the encoder and test your own videos available now
Part 5 — Put the representation back on the world-model road
Appendices — A backpack for the trail
Unpack the equations into a small toolkit
- Vectora numeric address card
- Distancehow far are two cards?
- Projectionlook at one shadow
- Gaussiana cloud with no favorite direction
- Gradientwhich way should weights move?
Suppose every clip receives a 256-slot address card. Two windows onto the same clip should get nearby addresses. Across many clips, however, the addresses must not all point to the same mailbox. The first rule is written as a distance; the second looks at the shape of the crowd from many directions.
When an equation appears, ask three questions before doing algebra: What goes in? Which axes are averaged? Does the result constrain one pair or a whole batch?
1. Vectors and mean squared distance
An embedding z = [z1, z2, ..., zK] is a K-dimensional vector. The public LeVJEPA recipe uses K=256 in its projection space. Squared Euclidean distance adds the squared coordinate differences:
||a - b||² = sum_k (a_k - b_k)²
Code such as .pow(2).mean() also averages over coordinates, samples, and views, so its numerical scale differs from a single-vector sum. The invariance term can be read as:
L_inv = mean over samples, views, and coordinates of (global_z - view_z)²
It pulls local summaries toward the global summary from the same temporal window. On its own, it has a perfect but useless answer: emit one constant vector for every input. That makes every distance zero. SIGReg exists because agreement alone cannot reject that shortcut.
2. A batch is a crowd, not a bag of unrelated pairs
Invariance compares paired samples. SIGReg inspects a population distribution. With B clips, V+1 views per clip, and projection width K, SIGReg logically receives:
[V+1, B, K]
For each view position, it studies the B vectors across different clips. In distributed training, the official implementation all-reduces the empirical characteristic-function statistics, so they describe the global batch. A tiny batch gives a noisy picture of the population. Adding more projection directions does not create more people for that picture.
3. Standard, isotropic, and Gaussian
The one-dimensional standard Gaussian is N(0,1): mean zero, variance one. Its K-dimensional isotropic counterpart is N(0,I_K): the same scale in every direction, with no privileged axis.
It is not a uniform coating on a sphere. A Gaussian has higher density near its center, while in high dimensions much of its probability mass lies at a typical radius around sqrt(K). Nor is matching the mean and variance of every coordinate sufficient. Two joint distributions may share those moments while hiding bends, branches, or X-shaped dependencies.
LeJEPA’s optimality argument for Gaussian embeddings belongs to the downstream risk and assumptions analysed in that paper. It is not a law that every useful representation of the physical world must be Gaussian.
4. Random projections and Cramér–Wold
A unit vector a is a flashlight direction. The one-dimensional shadow of high-dimensional point z is their dot product:
s = <z, a> = sum_k z_k a_k
The Cramér–Wold principle says that two high-dimensional probability distributions are the same if their one-dimensional projections match in every direction. SIGReg turns that idea into a trainable approximation. It samples M=1024 directions per step; it does not inspect infinitely many.
The careful claim is therefore: random projections provide a scalable distribution constraint. The careless claim is: 1,024 projections prove this batch is exactly Gaussian. Finite batches, directions, and integration nodes leave approximation error.
5. Characteristic functions are distribution fingerprints
A distribution’s characteristic function can be written as:
phi(t) = E[exp(i*t*s)]
= E[cos(t*s)] + i*E[sin(t*s)]
For a standard Gaussian, that fingerprint has the analytic form exp(-t²/2). The Epps–Pulley statistic compares a projected sample’s empirical fingerprint with this curve. The pinned implementation evaluates t from 0 to 3 at 17 nodes, uses trapezoidal integration and a Gaussian weight, and samples 1,024 directions. Sine and cosine are bounded, so a single enormous outlier cannot make the objective or gradient grow without limit.
6. The objective and where its gradients travel
L = L_inv + lambda * L_SIGReg, lambda = 0.02
lambda is the balance between the two losses. It is what the paper calls the objective’s sole hyperparameter. Training still has learning rates, batch sizes, view counts, drop rates, and many other settings; “one objective hyperparameter” does not mean “one knob in the entire experiment.”
Neither the global nor local embedding is detached. The invariance term backpropagates through both view branches, and SIGReg updates the shared encoder and projector too. After pretraining, the projector is discarded; downstream work uses encoder features.
7. FLOPs measure work, not seconds
FLOPs estimate floating-point computation and help compare algorithmic budgets. Wall-clock time also depends on hardware, kernels, memory traffic, I/O, communication, and implementation quality. Transformer attention has a roughly quadratic sequence-length component, while its MLP work is roughly linear in length. Reducing 3,136 patch tokens to about 157 can cut a great deal of work and memory, but it does not make every operation exactly 20 times faster.
Work one tiny example
Take global embedding g=[1,0] and local embeddings l1=[0,0], l2=[1,1].
- Each squared distance is 1, so their average paired error is 1.
- Change all three vectors to
[0,0]; invariance falls to 0. - If every clip in the dataset also maps to
[0,0], every projected shadow has zero variance and looks nothing likeN(0,1). SIGReg objects.
Three samples cannot estimate a Gaussian reliably. The example shows only how the two losses divide their jobs.
Where the tools come from
LeVJEPA v1 defines the two losses, lambda=0.02, and the implementation approximation. LeJEPA v3 develops SIGReg and the assumption-bound downstream-risk argument. The pinned SIGReg implementation fixes the 1,024 directions, 17 nodes, and distributed aggregation. The statistical roots are the Epps–Pulley paper and the Cramér–Wold principle.
Pack the toolkit
- Invariance constrains matched samples; SIGReg constrains a population of embeddings.
- Random projections approximate “all directions.” They do not issue a Gaussian certificate.
lambda=0.02is a loss weight, not the only setting in training.
Check the toolkit
- Why are per-coordinate means and variances not enough?
- Do a larger batch and more projection directions solve the same problem?
- Why can a low SIGReg loss not prove that the model understands motion?
Check the answers
- Joint distributions can have the same first two coordinate-wise moments but very different dependencies and shapes.
- No. More directions inspect more shadows; a larger batch supplies more samples in each empirical distribution.
- SIGReg checks the population shape, not temporal correspondence, actions, object identity, or downstream usefulness.