JEPA4Japan · tutorials

Chapter 11 — One encoder, one projector, one summary card

592 words 3 min read #LeVJEPA#JEPA#self-supervised video#SIGReg

Separate the jobs of the shared ViT, [CLS], and two-layer projector—and see why LayerNorm makes the projection space matter.

Course progress Course outline 34 of 34 lessons available

Part 0 — Get the map

  1. 01 Chapter 0 — Before you begin: what this course promises available now
  2. 02 Chapter 1 — One video, two windows available now
  3. 03 Chapter 2 — A walk along Yann LeCun’s research road available now
  4. 04 Chapter 3 — The JEPA family, without the name soup available now

Part 1 — Why a small objective can learn to see

  1. 05 Chapter 4 — Video can set its own homework available now
  2. 06 Chapter 5 — Keep the meaning; do not repaint every pixel available now
  3. 07 Chapter 6 — Match the cards, but do not leave every card blank available now
  4. 08 Chapter 7 — SIGReg checks a cloud by looking at its shadows available now
  5. 09 Chapter 8 — The whole LeVJEPA objective on one line available now

Part 2 — Send a video through one encoder

  1. 10 Chapter 9 — How global and local views are paired available now
  2. 11 Chapter 10 — Cut a video into space-time tiles available now
  3. 12 Chapter 11 — One encoder, one projector, one summary card Current lesson
  4. 13 Chapter 12 — One complete trip through the model available now
  5. 14 Chapter 13 — Why throwing away 95% can help available now
  6. 15 Chapter 14 — Same-frame teamwork, no peeking into tomorrow available now
  7. 16 Chapter 15 — RoPE, single-frame tubelets, and unexpectedly useful patch features available now

Part 3 — Read the experiments, not just the headline

  1. 17 Chapter 16 — What the four ablation ladders actually test available now
  2. 18 Chapter 17 — Equal epochs are not equal bills available now
  3. 19 Chapter 18 — What ImageNet, K400, and SSv2 are really asking available now
  4. 20 Chapter 19 — Keep the paper’s results in a ledger available now
  5. 21 Chapter 20 — Claims the evidence does not yet earn available now

Part 4 — From the official repository to your own experiment

  1. 22 Chapter 21 — A map of the official repository available now
  2. 23 Chapter 22 — Ten long walks become a training set available now
  3. 24 Chapter 23 — Read the defaults, then start training available now
  4. 25 Chapter 24 — Run a smoke test that cannot flatter you available now
  5. 26 Chapter 25 — Skip training: extract features from the public checkpoint available now
  6. 27 Chapter 26 — Freeze the encoder and test your own videos available now

Part 5 — Put the representation back on the world-model road

  1. 28 Chapter 27 — The important boundary: an encoder is not a planner available now
  2. 29 Chapter 28 — How LeVJEPA might feed a future world model available now
  3. 30 Chapter 29 — Ten projects, from first experiment to paper-sized question available now

Appendices — A backpack for the trail

  1. 31 Appendix A — The smallest useful math kit available now
  2. 32 Appendix B — The complete tensor-shape table available now
  3. 33 Appendix C — Glossary and paper timeline available now
  4. 34 Appendix D — Reproduction and review checklist available now

A reader, a summary card, and a temporary sorting room

  1. Many tilesvideo tokens
  2. One ViTshared by all views
  3. Summary cardread `[CLS]`
  4. Temporary roomproject to 256 dimensions
The encoder learns to look, `[CLS]` gathers the clip, and the projector gives the training loss a suitable workspace.

A librarian reads a full page and several magnified cutouts with the same brain. She writes a “what this page is about” card for each. Her final writing rules place those cards on a constrained surface, awkward for the distribution inspector. So the library adds a temporary sorting room that converts cards into a more useful format. Graduation day keeps the librarian and removes the room.

Those roles belong to the ViT, [CLS], and projector.

The parts list

PartInput → outputTraining jobAfter pretraining
Shared video ViT Eθview tokens → tokens of width devery global/local path reuses one parameter setkept as the downstream encoder
[CLS]one learned token at the sequence frontgathers the clip; the loss reads it directly and dropping never removes itkept as a clip summary
Two-layer projector hφd → 2048 → K, with K=256shared Linear → BatchNorm → GELU → Lineardiscarded

For one view:

z_v = hφ(Eθ(x_v)[CLS]) ∈ R^256

Why not apply SIGReg directly to the encoder output? The ViT ends in LayerNorm, which constrains the scale and geometry of [CLS]. The paper argues that this makes an isotropic-Gaussian objective awkward to optimize there. The projector creates a training interface not subject to exactly that final output geometry.

It is not a V-JEPA predictor. It does not use one path to guess another, and it has no target encoder or stopped-gradient partner.

The pinned implementation matches the appendix. A ViT-B has d=768; summaries may have shape [B, V+1, 768] before projection and [B, V+1, 256] after it. Downstream work returns to encoder representations instead of carrying the 256-dimensional projector along.

Sketch the sharing

Draw a global arrow and a local arrow into the same Eθ box. From each [CLS], point into the same hφ. Leave “second EMA target encoder,” “predictor,” and “stop-gradient” outside. For B=8, V=4, d=768, check [8,5,768] → [8,5,256].

Component receipts

The division of labor

  1. Every view passes through one ViT and reads its own [CLS] clip summary.
  2. The d → 2048 → 256 projector provides a workable SIGReg space beyond final LayerNorm geometry.
  3. The projector is discarded and is not a predictor; the objective has no target encoder or stop-gradient.

Parts quiz

  1. Do global and local views own separate ViT weights?
  2. Why introduce the projector?
  3. Does downstream evaluation keep the 256-dimensional head?
Answers
  1. No. They share one encoder parameter set.
  2. It gives SIGReg a space not constrained in the same way by the encoder’s final LayerNorm.
  3. No. Downstream evaluation reads encoder features.