JEPA4Japan · tutorials

Chapter 12 — One complete trip through the model

624 words 3 min read #LeVJEPA#JEPA#self-supervised video#SIGReg

Track global and local views, embeddings, both losses, gradient flow, and the pieces discarded after training.

Course progress Course outline 34 of 34 lessons available

Part 0 — Get the map

  1. 01 Chapter 0 — Before you begin: what this course promises available now
  2. 02 Chapter 1 — One video, two windows available now
  3. 03 Chapter 2 — A walk along Yann LeCun’s research road available now
  4. 04 Chapter 3 — The JEPA family, without the name soup available now

Part 1 — Why a small objective can learn to see

  1. 05 Chapter 4 — Video can set its own homework available now
  2. 06 Chapter 5 — Keep the meaning; do not repaint every pixel available now
  3. 07 Chapter 6 — Match the cards, but do not leave every card blank available now
  4. 08 Chapter 7 — SIGReg checks a cloud by looking at its shadows available now
  5. 09 Chapter 8 — The whole LeVJEPA objective on one line available now

Part 2 — Send a video through one encoder

  1. 10 Chapter 9 — How global and local views are paired available now
  2. 11 Chapter 10 — Cut a video into space-time tiles available now
  3. 12 Chapter 11 — One encoder, one projector, one summary card available now
  4. 13 Chapter 12 — One complete trip through the model Current lesson
  5. 14 Chapter 13 — Why throwing away 95% can help available now
  6. 15 Chapter 14 — Same-frame teamwork, no peeking into tomorrow available now
  7. 16 Chapter 15 — RoPE, single-frame tubelets, and unexpectedly useful patch features available now

Part 3 — Read the experiments, not just the headline

  1. 17 Chapter 16 — What the four ablation ladders actually test available now
  2. 18 Chapter 17 — Equal epochs are not equal bills available now
  3. 19 Chapter 18 — What ImageNet, K400, and SSv2 are really asking available now
  4. 20 Chapter 19 — Keep the paper’s results in a ledger available now
  5. 21 Chapter 20 — Claims the evidence does not yet earn available now

Part 4 — From the official repository to your own experiment

  1. 22 Chapter 21 — A map of the official repository available now
  2. 23 Chapter 22 — Ten long walks become a training set available now
  3. 24 Chapter 23 — Read the defaults, then start training available now
  4. 25 Chapter 24 — Run a smoke test that cannot flatter you available now
  5. 26 Chapter 25 — Skip training: extract features from the public checkpoint available now
  6. 27 Chapter 26 — Freeze the encoder and test your own videos available now

Part 5 — Put the representation back on the world-model road

  1. 28 Chapter 27 — The important boundary: an encoder is not a planner available now
  2. 29 Chapter 28 — How LeVJEPA might feed a future world model available now
  3. 30 Chapter 29 — Ten projects, from first experiment to paper-sized question available now

Appendices — A backpack for the trail

  1. 31 Appendix A — The smallest useful math kit available now
  2. 32 Appendix B — The complete tensor-shape table available now
  3. 33 Appendix C — Glossary and paper timeline available now
  4. 34 Appendix D — Reproduction and review checklist available now

Follow one batch all the way around

  1. Many windows1 global + V local
  2. Keep 5%sample sparse clues
  3. Encode togetherread `[CLS]`
  4. Score twicematch + spread
  5. Update both sidesgradients everywhere
There is no concealed teacher path: all views enter one encoder and projector, then meet invariance and SIGReg.

A camera crew films one match as a wide shot and several close-ups. The machine samples a handful of visual scraps from each, then writes a summary. One marker checks that all summaries describe the same match. Another checks that summaries across many matches do not repeat one sentence. Both sets of corrections travel backward—even the wide card can change. Nothing carries a “do not edit” stamp.

Station-by-station ledger

Here is the paper architecture in the order the official code runs it. Always state V: controlled paper settings commonly use V=4; the public Walking Tours default uses V=10.

StationGlobal viewEach local viewNote
Input16×224×22416×96×96same time interval
Patch tokens before dropping313657616×14×14; 16×6×6
After 95% dropping15729code keeps round(N×0.05)
After [CLS]15830[CLS] is never dropped
Encoder summary[B,1,d]combined as [B,V,d]local views are folded into the batch
Projector output[B,V+1,256]same tensorsummaries concatenate before projection

The code broadcasts the global embedding against all V+1 embeddings and averages squared error; the global-against-itself term is zero. SIGReg rearranges the tensor to [view, batch, 256] and examines the batch distribution one view at a time. The objective is MSE + 0.02 × SIGReg. Both global and local gradients update the shared encoder and projector.

That ends the training graph. There is no training target encoder, masked-query predictor, or stop-gradient.

Separately, the implementation maintains Polyak weights with decay 0.9999, updated every 32 optimizer steps, and saves them for evaluation. This copy makes no forward pass and no target. After pretraining, the projector is removed; reported evaluation and released weights use the encoder’s EMA copy.

Sanity-check a tiny batch

Let B=2 and V=4. There are 2×5=10 view summaries; projection yields [2,5,256]. At optimizer step 31, the every-32-step EMA has not reached its next update point. Draw no arrow from EMA into either loss.

Forward-pass records

The trip in three lines

  1. After 95% dropping, global/local views retain 157/29 patch tokens, then each receives [CLS].
  2. Invariance and SIGReg share [B,V+1,256]; gradients reach global and local paths.
  3. The 0.9999, every-32-step EMA is only for evaluation, never a target encoder.

Trace test

  1. How many patch tokens does the 224 global view have before dropping?
  2. Is the global summary stopped?
  3. Does the EMA copy participate in a training forward pass?
Answers
  1. 3136, or 16 × 14 × 14.
  2. No. Gradients flow through both sides.
  3. No. Its weights are periodically averaged and saved for evaluation.