JEPA4Japan · tutorials

Chapter 25 — Skip training: extract features from the public checkpoint

734 words 4 min read #LeVJEPA#JEPA#self-supervised video#SIGReg

Load Hugging Face weights safely, use the right normalization and tensor layout, and avoid mistaking an encoder for a classifier.

Course progress Course outline 34 of 34 lessons available

Part 0 — Get the map

  1. 01 Chapter 0 — Before you begin: what this course promises available now
  2. 02 Chapter 1 — One video, two windows available now
  3. 03 Chapter 2 — A walk along Yann LeCun’s research road available now
  4. 04 Chapter 3 — The JEPA family, without the name soup available now

Part 1 — Why a small objective can learn to see

  1. 05 Chapter 4 — Video can set its own homework available now
  2. 06 Chapter 5 — Keep the meaning; do not repaint every pixel available now
  3. 07 Chapter 6 — Match the cards, but do not leave every card blank available now
  4. 08 Chapter 7 — SIGReg checks a cloud by looking at its shadows available now
  5. 09 Chapter 8 — The whole LeVJEPA objective on one line available now

Part 2 — Send a video through one encoder

  1. 10 Chapter 9 — How global and local views are paired available now
  2. 11 Chapter 10 — Cut a video into space-time tiles available now
  3. 12 Chapter 11 — One encoder, one projector, one summary card available now
  4. 13 Chapter 12 — One complete trip through the model available now
  5. 14 Chapter 13 — Why throwing away 95% can help available now
  6. 15 Chapter 14 — Same-frame teamwork, no peeking into tomorrow available now
  7. 16 Chapter 15 — RoPE, single-frame tubelets, and unexpectedly useful patch features available now

Part 3 — Read the experiments, not just the headline

  1. 17 Chapter 16 — What the four ablation ladders actually test available now
  2. 18 Chapter 17 — Equal epochs are not equal bills available now
  3. 19 Chapter 18 — What ImageNet, K400, and SSv2 are really asking available now
  4. 20 Chapter 19 — Keep the paper’s results in a ledger available now
  5. 21 Chapter 20 — Claims the evidence does not yet earn available now

Part 4 — From the official repository to your own experiment

  1. 22 Chapter 21 — A map of the official repository available now
  2. 23 Chapter 22 — Ten long walks become a training set available now
  3. 24 Chapter 23 — Read the defaults, then start training available now
  4. 25 Chapter 24 — Run a smoke test that cannot flatter you available now
  5. 26 Chapter 25 — Skip training: extract features from the public checkpoint Current lesson
  6. 27 Chapter 26 — Freeze the encoder and test your own videos available now

Part 5 — Put the representation back on the world-model road

  1. 28 Chapter 27 — The important boundary: an encoder is not a planner available now
  2. 29 Chapter 28 — How LeVJEPA might feed a future world model available now
  3. 30 Chapter 29 — Ten projects, from first experiment to paper-sized question available now

Appendices — A backpack for the trail

  1. 31 Appendix A — The smallest useful math kit available now
  2. 32 Appendix B — The complete tensor-shape table available now
  3. 33 Appendix C — Glossary and paper timeline available now
  4. 34 Appendix D — Reproduction and review checklist available now

Inspect a borrowed machine before plugging it in

  1. Read the coderemote Python will execute
  2. Pin the revisionweights and code stay paired
  3. Order the axes[B,C,T,H,W]
  4. Normalize onceImageNet mean/std
  5. Read featurestokens or `[CLS]`
A public checkpoint is borrowed machinery: read its manual and license before supplying power.

The download includes more than numbers. It includes custom assembly code. If time and color axes are swapped, the machine may still run and return plausible-shaped nonsense. Safe use begins with code review, a pinned version, and explicit tensor checks.

Pin and audit the release

LeVJEPA-VideoMix-Large is a 303.1M-parameter ViT-L/16 trained on 1,806,869 VideoMix clips. It uses 16 frames, tubelet 1, RoPE, and block_causal attention.

The repository requires trust_remote_code=True, meaning downloaded Python executes locally. Inspect config.json, configuration_levjepa.py, and modeling_levjepa.py; use an isolated environment without sensitive credentials; pin the audited model commit e831a03 instead of floating HEAD.

At minimum ask: Which files does auto_map import? Does code initiate network calls, read environment variables, or dynamically execute strings? Are weights in safetensors? Does the license permit your use? After a reviewed online download, test the same cache offline with local_files_only=True. revision pins both config and custom code.

import torch
from transformers import AutoModel

repo = "galilai-group/LeVJEPA-VideoMix-Large"
revision = "e831a0347737fcaa660b39c57d41c109de399845"
model = AutoModel.from_pretrained(
    repo, revision=revision, trust_remote_code=True
).eval()

# raw is already in [0,1]. Real video also needs matching resize/center crop to 224.
raw = torch.rand(1, 3, 16, 224, 224)  # [B,C,T,H,W]
mean = torch.tensor([0.485, 0.456, 0.406])[None, :, None, None, None]
std = torch.tensor([0.229, 0.224, 0.225])[None, :, None, None, None]
video = (raw - mean) / std
with torch.inference_mode():
    out = model(pixel_values=video)
print(out.last_hidden_state.shape, out.pooler_output.shape)

Expected shapes are torch.Size([1, 3137, 1024]) and torch.Size([1, 1024]): one [CLS] plus 16×14×14 patch tokens. Eval mode sets token dropping to zero, so all patches return. There is no classification head; pooler_output is the [CLS] representation, not class probabilities.

Full-token block-causal ViT-L attention is memory-hungry. Begin at batch 1. Do not quietly switch to full attention to make it fit. Add assertions for (1,3137,1024) and torch.isfinite. On CUDA or MPS, model, video, mean, and standard deviation need one device. Log peak memory before increasing batch. eval() disables training-only dropping and dropout; encoder LayerNorm does not keep BatchNorm-style running statistics. Repeated inference should agree within the hardware’s numerical tolerance.

Training samples roughly 7.5 fps, so 16 frames span about two seconds. Repeating an image with image.unsqueeze(2).repeat(1,1,16,1,1) can produce image features, but contains no true motion. Preserve model.config.attn_mode == "block_causal". Released tensors are the encoder’s EMA evaluation weights; they are not an EMA teacher and do not include the training projector.

The model card uses CC BY-NC 4.0. Review licensing before commercial use.

Checkpoint records

Before inference

  1. trust_remote_code=True executes downloaded code; audit it and pin a revision.
  2. Input is ImageNet-normalized [B,C,T,H,W], normally 16×224×224.
  3. Output is a frozen representation, not a classification answer; the release uses evaluation EMA weights.

Plug-in test

  1. Why are there 3137 output tokens?
  2. Can you switch attn_mode to full and still call it protocol-matched inference?
  3. Is pooler_output an ImageNet probability vector?
Answers
  1. 1 + 16×(224/16)×(224/16).
  2. No. Running successfully does not preserve the trained attention topology.
  3. No. It is a 1024-dimensional [CLS] feature; the model has no classifier.