JEPA4Japan · tutorials

Chapter 18 — What ImageNet, K400, and SSv2 are really asking

657 words 3 min read #LeVJEPA#JEPA#self-supervised video#SIGReg

Separate appearance from motion evidence, and understand frozen encoders, linear probes, and attentive probes.

Course progress Course outline 34 of 34 lessons available

Part 0 — Get the map

  1. 01 Chapter 0 — Before you begin: what this course promises available now
  2. 02 Chapter 1 — One video, two windows available now
  3. 03 Chapter 2 — A walk along Yann LeCun’s research road available now
  4. 04 Chapter 3 — The JEPA family, without the name soup available now

Part 1 — Why a small objective can learn to see

  1. 05 Chapter 4 — Video can set its own homework available now
  2. 06 Chapter 5 — Keep the meaning; do not repaint every pixel available now
  3. 07 Chapter 6 — Match the cards, but do not leave every card blank available now
  4. 08 Chapter 7 — SIGReg checks a cloud by looking at its shadows available now
  5. 09 Chapter 8 — The whole LeVJEPA objective on one line available now

Part 2 — Send a video through one encoder

  1. 10 Chapter 9 — How global and local views are paired available now
  2. 11 Chapter 10 — Cut a video into space-time tiles available now
  3. 12 Chapter 11 — One encoder, one projector, one summary card available now
  4. 13 Chapter 12 — One complete trip through the model available now
  5. 14 Chapter 13 — Why throwing away 95% can help available now
  6. 15 Chapter 14 — Same-frame teamwork, no peeking into tomorrow available now
  7. 16 Chapter 15 — RoPE, single-frame tubelets, and unexpectedly useful patch features available now

Part 3 — Read the experiments, not just the headline

  1. 17 Chapter 16 — What the four ablation ladders actually test available now
  2. 18 Chapter 17 — Equal epochs are not equal bills available now
  3. 19 Chapter 18 — What ImageNet, K400, and SSv2 are really asking Current lesson
  4. 20 Chapter 19 — Keep the paper’s results in a ledger available now
  5. 21 Chapter 20 — Claims the evidence does not yet earn available now

Part 4 — From the official repository to your own experiment

  1. 22 Chapter 21 — A map of the official repository available now
  2. 23 Chapter 22 — Ten long walks become a training set available now
  3. 24 Chapter 23 — Read the defaults, then start training available now
  4. 25 Chapter 24 — Run a smoke test that cannot flatter you available now
  5. 26 Chapter 25 — Skip training: extract features from the public checkpoint available now
  6. 27 Chapter 26 — Freeze the encoder and test your own videos available now

Part 5 — Put the representation back on the world-model road

  1. 28 Chapter 27 — The important boundary: an encoder is not a planner available now
  2. 29 Chapter 28 — How LeVJEPA might feed a future world model available now
  3. 30 Chapter 29 — Ten projects, from first experiment to paper-sized question available now

Appendices — A backpack for the trail

  1. 31 Appendix A — The smallest useful math kit available now
  2. 32 Appendix B — The complete tensor-shape table available now
  3. 33 Appendix C — Glossary and paper timeline available now
  4. 34 Appendix D — Reproduction and review checklist available now

Three exam rooms and a frozen student

  1. Name the objectImageNet-1K
  2. Name the actionKinetics-400
  3. Read the motionSomething-v2
  4. Freeze the braintrain only the examiner
The exams ask different questions. A frozen encoder does not learn, but a supervised probe still does.

One child enters three rooms. A photo exam asks, “What is this?” An action exam asks, “What is happening?” A motion-sensitive exam separates “putting the cup on the table” from “taking it off.” The child may not study between rooms. Each room may train a small interpreter to translate what the child already knows into class names.

That is a frozen-probe evaluation—not zero-shot recognition.

Match each score to its reader

BenchmarkRole in the paperv1 readoutEvidence it can support
ImageNet-1K (IN1K)static objects, appearance-heavyfrozen attentive probeobject information accessible to a small trained reader
Something-Something-v2 (SSv2)temporal order and motion relationfrozen attentive probemotion/order information, not complete physical understanding
Kinetics-400 (K400)action recognitionfrozen mean-pooled tokens + linear probeaction-class information under a weaker adaptation protocol

Frozen means the encoder parameters do not change. The probe and classifier still learn from labeled downstream data. So this is neither zero-shot prediction nor end-to-end fine-tuning. Unfreezing the encoder would answer a separate question about adaptability.

The attentive probe is more than one linear layer. Appendix C specifies a learned query reading every frozen output token through one cross-attention layer; a residual adds the query, followed by a two-layer MLP, GELU, LayerNorm, and a linear classifier. It can learn where to look. The result measures information present in the representation, not whether that information was already linearly arranged.

To control probe cost on larger K400, the paper instead mean-pools all output tokens and trains a linear classifier, describing it as strictly weaker adaptation. Thus K400 and IN1K/SSv2 numbers in the FLOP-matched table do not use identical heads.

The authors also adjust temporal input length so video methods with different tubelets expose the same number of tokens to the probe. For static ImageNet, one image is repeated along time before entering the video encoder.

High IN1K does not imply strong motion. High SSv2 does not imply planning. Read all three exams to see the reported shape: under equal FLOPs, LeVJEPA is strong on appearance and K400 but remains below the best SSv2 value in that table.

Sort four claims

Label each as frozen probe, fine-tuning, zero-shot, or invalid: (1) freeze encoder, train cross-attention and classifier; (2) update encoder too; (3) answer without training any head; (4) infer robot planning from a high SSv2 score. The order is frozen probe, fine-tuning, zero-shot, invalid.

Exam protocols

Read the score correctly

  1. IN1K emphasizes appearance, SSv2 motion/order, and K400 action class; none substitutes for the others.
  2. A frozen probe trains a labeled readout but not the encoder; it is not zero-shot or fine-tuning.
  3. IN1K/SSv2 use an attentive probe; K400 uses mean pooling and a linear probe.

Hand in the exam

  1. Does frozen evaluation update the encoder?
  2. Is an attentive probe just one linear classifier?
  3. Why not claim K400 and IN1K use the same readout?
Answers
  1. No. Only the probe and classifier update.
  2. No. It includes cross-attention, a residual, a two-layer MLP, GELU, LayerNorm, and classification.
  3. K400 uses mean pooling plus a linear probe; IN1K uses an attentive probe.