JEPA4Japan · tutorials

Chapter 8 — The whole LeVJEPA objective on one line

647 words 3 min read #LeVJEPA#JEPA#self-supervised video#SIGReg

Unpack invariance plus SIGReg, the fixed λ=0.02 trade-off, two-way gradients, and the boundary of the theory.

Course progress Course outline 34 of 34 lessons available

Part 0 — Get the map

  1. 01 Chapter 0 — Before you begin: what this course promises available now
  2. 02 Chapter 1 — One video, two windows available now
  3. 03 Chapter 2 — A walk along Yann LeCun’s research road available now
  4. 04 Chapter 3 — The JEPA family, without the name soup available now

Part 1 — Why a small objective can learn to see

  1. 05 Chapter 4 — Video can set its own homework available now
  2. 06 Chapter 5 — Keep the meaning; do not repaint every pixel available now
  3. 07 Chapter 6 — Match the cards, but do not leave every card blank available now
  4. 08 Chapter 7 — SIGReg checks a cloud by looking at its shadows available now
  5. 09 Chapter 8 — The whole LeVJEPA objective on one line Current lesson

Part 2 — Send a video through one encoder

  1. 10 Chapter 9 — How global and local views are paired available now
  2. 11 Chapter 10 — Cut a video into space-time tiles available now
  3. 12 Chapter 11 — One encoder, one projector, one summary card available now
  4. 13 Chapter 12 — One complete trip through the model available now
  5. 14 Chapter 13 — Why throwing away 95% can help available now
  6. 15 Chapter 14 — Same-frame teamwork, no peeking into tomorrow available now
  7. 16 Chapter 15 — RoPE, single-frame tubelets, and unexpectedly useful patch features available now

Part 3 — Read the experiments, not just the headline

  1. 17 Chapter 16 — What the four ablation ladders actually test available now
  2. 18 Chapter 17 — Equal epochs are not equal bills available now
  3. 19 Chapter 18 — What ImageNet, K400, and SSv2 are really asking available now
  4. 20 Chapter 19 — Keep the paper’s results in a ledger available now
  5. 21 Chapter 20 — Claims the evidence does not yet earn available now

Part 4 — From the official repository to your own experiment

  1. 22 Chapter 21 — A map of the official repository available now
  2. 23 Chapter 22 — Ten long walks become a training set available now
  3. 24 Chapter 23 — Read the defaults, then start training available now
  4. 25 Chapter 24 — Run a smoke test that cannot flatter you available now
  5. 26 Chapter 25 — Skip training: extract features from the public checkpoint available now
  6. 27 Chapter 26 — Freeze the encoder and test your own videos available now

Part 5 — Put the representation back on the world-model road

  1. 28 Chapter 27 — The important boundary: an encoder is not a planner available now
  2. 29 Chapter 28 — How LeVJEPA might feed a future world model available now
  3. 30 Chapter 29 — Ten projects, from first experiment to paper-sized question available now

Appendices — A backpack for the trail

  1. 31 Appendix A — The smallest useful math kit available now
  2. 32 Appendix B — The complete tensor-shape table available now
  3. 33 Appendix C — Glossary and paper timeline available now
  4. 34 Appendix D — Reproduction and review checklist available now

Two rules and one volume knob

  1. Shared machineencoder + projector
  2. Pull pairs closeinvariance
  3. Shape the batchSIGReg
  4. Set 0.02one trade-off weight
The objective says two things: paired views should agree, and the population must not collapse.

Imagine a classroom with two marks for every set of cards. The first is awarded clip by clip: do the wide and close-up cards agree? The second looks across the room: do all cards form a healthy spread, or did everyone copy one answer? The second score is turned down to 0.02 and added to the first. No hidden teacher supplies perfect cards. Both sides revise together.

The shortest honest accounting

z_v = projector(shared_encoder(view_v)[CLS])
L_total = L_inv + 0.02 * L_SIGReg

view_0 is global; view_1 … view_V are local views of the same 16 frames. The paper writes the invariance term as:

L_inv = (1 / (V + 1)) * sum_{v=0..V} ||z_0 - z_v||²

The v = 0 term is zero, so the working constraint is each local embedding’s distance from the global one. Both arguments of that squared distance remain connected to the graph. Gradients from global and local paths update the same encoder/projector parameters.

L_SIGReg measures how the projected view embeddings differ from a standard normal along sampled directions. λ = 0.02 is the objective’s only trade-off hyperparameter, and the paper reports using that public default throughout rather than tuning it per experiment. “Only” does not erase the learning rate, batch size, crop count, token drop ratio, or model size.

The training box contains one shared Vision Transformer, one small shared projector, and a [CLS] readout per view. The projector gives SIGReg a space not pinned to the encoder’s final LayerNorm geometry. It is thrown away after pretraining; downstream tasks consume encoder features.

Outside that box: no target encoder, predictor, stop-gradient, or reconstruction of missing tokens. The official code does update a Polyak/EMA encoder after parameter steps and saves state_dict_ema for evaluation. Because that copy never supplies any z_v or loss target, it belongs outside the objective diagram.

One more absence matters. The equation contains no action a_t, future-state prediction, reward, cost, or search. Elegance does not turn the encoder into a planner.

Check the arithmetic

Suppose a log line shows L_inv = 0.30 and L_SIGReg = 5.0. Then L_total = 0.30 + 0.02 × 5.0 = 0.40. If someone computes 0.02 × L_inv + L_SIGReg, the knob is attached to the wrong term.

Objective records

Fold the equation into three lines

  1. The objective is L_inv + 0.02 × L_SIGReg.
  2. Global and local paths share the encoder/projector, and gradients flow through both sides.
  3. EMA serves evaluation; there is no objective teacher, predictor, stop-gradient, or planner.

Final check for this part

  1. Which loss does λ = 0.02 multiply?
  2. Does the projector stay for downstream use?
  3. How can the repository contain EMA weights while the objective has no EMA teacher?
Answers
  1. SIGReg.
  2. No. It is discarded after pretraining.
  3. The EMA copy is saved for evaluation; it neither makes targets nor participates in either loss.