JEPA4Japan · tutorials

Chapter 26 — Freeze the encoder and test your own videos

778 words 4 min read #LeVJEPA#JEPA#self-supervised video#SIGReg

Build linear and attentive probes, retrieval checks, and dense-feature diagnostics with leakage-resistant splits.

Course progress Course outline 34 of 34 lessons available

Part 0 — Get the map

  1. 01 Chapter 0 — Before you begin: what this course promises available now
  2. 02 Chapter 1 — One video, two windows available now
  3. 03 Chapter 2 — A walk along Yann LeCun’s research road available now
  4. 04 Chapter 3 — The JEPA family, without the name soup available now

Part 1 — Why a small objective can learn to see

  1. 05 Chapter 4 — Video can set its own homework available now
  2. 06 Chapter 5 — Keep the meaning; do not repaint every pixel available now
  3. 07 Chapter 6 — Match the cards, but do not leave every card blank available now
  4. 08 Chapter 7 — SIGReg checks a cloud by looking at its shadows available now
  5. 09 Chapter 8 — The whole LeVJEPA objective on one line available now

Part 2 — Send a video through one encoder

  1. 10 Chapter 9 — How global and local views are paired available now
  2. 11 Chapter 10 — Cut a video into space-time tiles available now
  3. 12 Chapter 11 — One encoder, one projector, one summary card available now
  4. 13 Chapter 12 — One complete trip through the model available now
  5. 14 Chapter 13 — Why throwing away 95% can help available now
  6. 15 Chapter 14 — Same-frame teamwork, no peeking into tomorrow available now
  7. 16 Chapter 15 — RoPE, single-frame tubelets, and unexpectedly useful patch features available now

Part 3 — Read the experiments, not just the headline

  1. 17 Chapter 16 — What the four ablation ladders actually test available now
  2. 18 Chapter 17 — Equal epochs are not equal bills available now
  3. 19 Chapter 18 — What ImageNet, K400, and SSv2 are really asking available now
  4. 20 Chapter 19 — Keep the paper’s results in a ledger available now
  5. 21 Chapter 20 — Claims the evidence does not yet earn available now

Part 4 — From the official repository to your own experiment

  1. 22 Chapter 21 — A map of the official repository available now
  2. 23 Chapter 22 — Ten long walks become a training set available now
  3. 24 Chapter 23 — Read the defaults, then start training available now
  4. 25 Chapter 24 — Run a smoke test that cannot flatter you available now
  5. 26 Chapter 25 — Skip training: extract features from the public checkpoint available now
  6. 27 Chapter 26 — Freeze the encoder and test your own videos Current lesson

Part 5 — Put the representation back on the world-model road

  1. 28 Chapter 27 — The important boundary: an encoder is not a planner available now
  2. 29 Chapter 28 — How LeVJEPA might feed a future world model available now
  3. 30 Chapter 29 — Ten projects, from first experiment to paper-sized question available now

Appendices — A backpack for the trail

  1. 31 Appendix A — The smallest useful math kit available now
  2. 32 Appendix B — The complete tensor-shape table available now
  3. 33 Appendix C — Glossary and paper timeline available now
  4. 34 Appendix D — Reproduction and review checklist available now

Split families before cutting clips

  1. Group firstpeople and sources do not cross
  2. Freeze encoderread it, do not teach it
  3. Train a small probemeasure readable information
  4. Add controlscatch shortcuts
  5. Report the protocolmean, spread, failures
Frozen evaluation changes the exam reader, not the student. A leaky split lets the reader memorize the answer sheet.

Randomly cut one long recording into train and test windows and the score may look wonderful. The test question could simply be the second after a training question, with the same person, wall, and watermark. Group by source video, scene, subject, or date before extracting windows.

Four questions need four readouts

Fix preprocessing, model revision, and frame rate first. For [B,3,16,224,224], freeze every encoder parameter and assert requires_grad=False.

A minimal linear probe maps [CLS] features [B,1024] through one linear layer. Alongside accuracy, report balanced accuracy or macro-F1 when classes are uneven.

The paper’s attentive probe gives a learned query cross-attention access to all 3,137 frozen tokens, adds a query residual, then applies a two-layer MLP, GELU, LayerNorm, and classifier. This is higher capacity and must not be called a linear probe.

For retrieval, L2-normalize [CLS], rank cosine similarity, and report Recall@1/5 and mAP. For dense study, remove [CLS] and reshape out.last_hidden_state[:,1:] to [B,16,14,14,1024]. Query-patch maps, small segmentation or tracking heads, and metrics such as mIoU, J&F, or point error answer different questions. Remember: patch tokens have the per-frame causal guarantee; whole-clip [CLS] does not.

@torch.inference_mode()
def features(model, raw_bcthw):
    mean = raw_bcthw.new_tensor([.485,.456,.406])[None,:,None,None,None]
    std = raw_bcthw.new_tensor([.229,.224,.225])[None,:,None,None,None]
    out = model(pixel_values=(raw_bcthw - mean) / std)
    cls = torch.nn.functional.normalize(out.pooler_output, dim=-1)
    patches = out.last_hidden_state[:, 1:].reshape(-1, 16, 14, 14, 1024)
    return cls.cpu(), patches.cpu()

This assumes [0,1] input already cropped to 224. Make conversion explicit otherwise.

Split source groups into train/validation/test, then sample windows inside each. Fit normalization and class weights on train only; choose hyperparameters with train/validation; keep test sealed. Include at least three negative controls: shuffled training labels should fall to chance; repeated single frames reveal reliance on appearance; reversed or shuffled frames test sensitivity to time. Compare with a randomly initialized same-shape encoder or simple color/motion baseline.

For a linear probe, caching train features and optimizing only torch.nn.Linear(1024, n_classes) is enough. Verify an encoder parameter digest remains unchanged each epoch. Select with validation, touch test once, run at least three seeds, and report mean, standard deviation, counts per class, and a confusion matrix. Small tests benefit from bootstrap confidence intervals. A stratified split cannot fix identity leakage; use group splitting when subjects recur.

Time tasks deserve an appearance-matched control: use exactly the same frames in forward and reverse order. Failure to beat chance suggests the probe is reading objects rather than sequence. For retrieval, remove same-source neighbors from the gallery. Score dense tasks on a fixed annotated set, not hand-picked pretty PCA clips. Record fps, seconds per window, crop, probe parameters, and revision.

The paper’s 69.5 and 55.0 belong to its own frozen attentive-probe protocol. A new dataset, split, or head is a new experiment—not an automatic reproduction or victory over those values.

Evaluation references

A clean evaluation

  1. Split source videos, people, or scenes before making clips to block near-neighbor leakage.
  2. Linear, attentive, retrieval, and dense probes answer different questions.
  3. [CLS] summarizes a full clip; use time-indexed patch tokens for causal per-frame state.

Leakage test

  1. Why not randomize adjacent two-second windows across train and test?
  2. What does far-above-chance accuracy with shuffled labels suggest?
  3. How do you reshape [B,3137,1024] after removing [CLS]?
Answers
  1. They nearly duplicate people, background, and motion.
  2. Leakage in splits, caches, or evaluation.
  3. [B,16,14,14,1024].