Course progress Course outline 34 of 34 lessons available
Part 0 — Get the map
Part 1 — Why a small objective can learn to see
- 05 Chapter 4 — Video can set its own homework available now
- 06 Chapter 5 — Keep the meaning; do not repaint every pixel available now
- 07 Chapter 6 — Match the cards, but do not leave every card blank available now
- 08 Chapter 7 — SIGReg checks a cloud by looking at its shadows available now
- 09 Chapter 8 — The whole LeVJEPA objective on one line available now
Part 2 — Send a video through one encoder
- 10 Chapter 9 — How global and local views are paired available now
- 11 Chapter 10 — Cut a video into space-time tiles available now
- 12 Chapter 11 — One encoder, one projector, one summary card Current lesson
- 13 Chapter 12 — One complete trip through the model available now
- 14 Chapter 13 — Why throwing away 95% can help available now
- 15 Chapter 14 — Same-frame teamwork, no peeking into tomorrow available now
- 16 Chapter 15 — RoPE, single-frame tubelets, and unexpectedly useful patch features available now
Part 3 — Read the experiments, not just the headline
- 17 Chapter 16 — What the four ablation ladders actually test available now
- 18 Chapter 17 — Equal epochs are not equal bills available now
- 19 Chapter 18 — What ImageNet, K400, and SSv2 are really asking available now
- 20 Chapter 19 — Keep the paper’s results in a ledger available now
- 21 Chapter 20 — Claims the evidence does not yet earn available now
Part 4 — From the official repository to your own experiment
- 22 Chapter 21 — A map of the official repository available now
- 23 Chapter 22 — Ten long walks become a training set available now
- 24 Chapter 23 — Read the defaults, then start training available now
- 25 Chapter 24 — Run a smoke test that cannot flatter you available now
- 26 Chapter 25 — Skip training: extract features from the public checkpoint available now
- 27 Chapter 26 — Freeze the encoder and test your own videos available now
Part 5 — Put the representation back on the world-model road
Appendices — A backpack for the trail
A reader, a summary card, and a temporary sorting room
- Many tilesvideo tokens
- One ViTshared by all views
- Summary cardread `[CLS]`
- Temporary roomproject to 256 dimensions
A librarian reads a full page and several magnified cutouts with the same brain. She writes a “what this page is about” card for each. Her final writing rules place those cards on a constrained surface, awkward for the distribution inspector. So the library adds a temporary sorting room that converts cards into a more useful format. Graduation day keeps the librarian and removes the room.
Those roles belong to the ViT, [CLS], and projector.
The parts list
| Part | Input → output | Training job | After pretraining |
|---|---|---|---|
Shared video ViT Eθ | view tokens → tokens of width d | every global/local path reuses one parameter set | kept as the downstream encoder |
[CLS] | one learned token at the sequence front | gathers the clip; the loss reads it directly and dropping never removes it | kept as a clip summary |
Two-layer projector hφ | d → 2048 → K, with K=256 | shared Linear → BatchNorm → GELU → Linear | discarded |
For one view:
z_v = hφ(Eθ(x_v)[CLS]) ∈ R^256
Why not apply SIGReg directly to the encoder output? The ViT ends in LayerNorm, which constrains the scale and geometry of [CLS]. The paper argues that this makes an isotropic-Gaussian objective awkward to optimize there. The projector creates a training interface not subject to exactly that final output geometry.
It is not a V-JEPA predictor. It does not use one path to guess another, and it has no target encoder or stopped-gradient partner.
The pinned implementation matches the appendix. A ViT-B has d=768; summaries may have shape [B, V+1, 768] before projection and [B, V+1, 256] after it. Downstream work returns to encoder representations instead of carrying the 256-dimensional projector along.
Sketch the sharing
Draw a global arrow and a local arrow into the same Eθ box. From each [CLS], point into the same hφ. Leave “second EMA target encoder,” “predictor,” and “stop-gradient” outside. For B=8, V=4, d=768, check [8,5,768] → [8,5,256].
Component receipts
- LeVJEPA v1: method and architecture
- Appendix B: two-layer projector and 256-dimensional output
- Pinned projector implementation
- Pinned shared forward paths
The division of labor
- Every view passes through one ViT and reads its own
[CLS]clip summary. - The
d → 2048 → 256projector provides a workable SIGReg space beyond final LayerNorm geometry. - The projector is discarded and is not a predictor; the objective has no target encoder or stop-gradient.
Parts quiz
- Do global and local views own separate ViT weights?
- Why introduce the projector?
- Does downstream evaluation keep the 256-dimensional head?
Answers
- No. They share one encoder parameter set.
- It gives SIGReg a space not constrained in the same way by the encoder’s final LayerNorm.
- No. Downstream evaluation reads encoder features.