Course progress Course outline 34 of 34 lessons available
Part 0 — Get the map
Part 1 — Why a small objective can learn to see
- 05 Chapter 4 — Video can set its own homework available now
- 06 Chapter 5 — Keep the meaning; do not repaint every pixel available now
- 07 Chapter 6 — Match the cards, but do not leave every card blank available now
- 08 Chapter 7 — SIGReg checks a cloud by looking at its shadows available now
- 09 Chapter 8 — The whole LeVJEPA objective on one line available now
Part 2 — Send a video through one encoder
- 10 Chapter 9 — How global and local views are paired available now
- 11 Chapter 10 — Cut a video into space-time tiles available now
- 12 Chapter 11 — One encoder, one projector, one summary card available now
- 13 Chapter 12 — One complete trip through the model available now
- 14 Chapter 13 — Why throwing away 95% can help available now
- 15 Chapter 14 — Same-frame teamwork, no peeking into tomorrow available now
- 16 Chapter 15 — RoPE, single-frame tubelets, and unexpectedly useful patch features available now
Part 3 — Read the experiments, not just the headline
- 17 Chapter 16 — What the four ablation ladders actually test available now
- 18 Chapter 17 — Equal epochs are not equal bills available now
- 19 Chapter 18 — What ImageNet, K400, and SSv2 are really asking available now
- 20 Chapter 19 — Keep the paper’s results in a ledger available now
- 21 Chapter 20 — Claims the evidence does not yet earn available now
Part 4 — From the official repository to your own experiment
- 22 Chapter 21 — A map of the official repository available now
- 23 Chapter 22 — Ten long walks become a training set available now
- 24 Chapter 23 — Read the defaults, then start training available now
- 25 Chapter 24 — Run a smoke test that cannot flatter you available now
- 26 Chapter 25 — Skip training: extract features from the public checkpoint Current lesson
- 27 Chapter 26 — Freeze the encoder and test your own videos available now
Part 5 — Put the representation back on the world-model road
Appendices — A backpack for the trail
Inspect a borrowed machine before plugging it in
- Read the coderemote Python will execute
- Pin the revisionweights and code stay paired
- Order the axes[B,C,T,H,W]
- Normalize onceImageNet mean/std
- Read featurestokens or `[CLS]`
The download includes more than numbers. It includes custom assembly code. If time and color axes are swapped, the machine may still run and return plausible-shaped nonsense. Safe use begins with code review, a pinned version, and explicit tensor checks.
Pin and audit the release
LeVJEPA-VideoMix-Large is a 303.1M-parameter ViT-L/16 trained on 1,806,869 VideoMix clips. It uses 16 frames, tubelet 1, RoPE, and block_causal attention.
The repository requires trust_remote_code=True, meaning downloaded Python executes locally. Inspect config.json, configuration_levjepa.py, and modeling_levjepa.py; use an isolated environment without sensitive credentials; pin the audited model commit e831a03 instead of floating HEAD.
At minimum ask: Which files does auto_map import? Does code initiate network calls, read environment variables, or dynamically execute strings? Are weights in safetensors? Does the license permit your use? After a reviewed online download, test the same cache offline with local_files_only=True. revision pins both config and custom code.
import torch
from transformers import AutoModel
repo = "galilai-group/LeVJEPA-VideoMix-Large"
revision = "e831a0347737fcaa660b39c57d41c109de399845"
model = AutoModel.from_pretrained(
repo, revision=revision, trust_remote_code=True
).eval()
# raw is already in [0,1]. Real video also needs matching resize/center crop to 224.
raw = torch.rand(1, 3, 16, 224, 224) # [B,C,T,H,W]
mean = torch.tensor([0.485, 0.456, 0.406])[None, :, None, None, None]
std = torch.tensor([0.229, 0.224, 0.225])[None, :, None, None, None]
video = (raw - mean) / std
with torch.inference_mode():
out = model(pixel_values=video)
print(out.last_hidden_state.shape, out.pooler_output.shape)
Expected shapes are torch.Size([1, 3137, 1024]) and torch.Size([1, 1024]): one [CLS] plus 16×14×14 patch tokens. Eval mode sets token dropping to zero, so all patches return. There is no classification head; pooler_output is the [CLS] representation, not class probabilities.
Full-token block-causal ViT-L attention is memory-hungry. Begin at batch 1. Do not quietly switch to full attention to make it fit. Add assertions for (1,3137,1024) and torch.isfinite. On CUDA or MPS, model, video, mean, and standard deviation need one device. Log peak memory before increasing batch. eval() disables training-only dropping and dropout; encoder LayerNorm does not keep BatchNorm-style running statistics. Repeated inference should agree within the hardware’s numerical tolerance.
Training samples roughly 7.5 fps, so 16 frames span about two seconds. Repeating an image with image.unsqueeze(2).repeat(1,1,16,1,1) can produce image features, but contains no true motion. Preserve model.config.attn_mode == "block_causal". Released tensors are the encoder’s EMA evaluation weights; they are not an EMA teacher and do not include the training projector.
The model card uses CC BY-NC 4.0. Review licensing before commercial use.
Checkpoint records
Before inference
trust_remote_code=Trueexecutes downloaded code; audit it and pin a revision.- Input is ImageNet-normalized
[B,C,T,H,W], normally 16×224×224. - Output is a frozen representation, not a classification answer; the release uses evaluation EMA weights.
Plug-in test
- Why are there 3137 output tokens?
- Can you switch
attn_modeto full and still call it protocol-matched inference? - Is
pooler_outputan ImageNet probability vector?
Answers
1 + 16×(224/16)×(224/16).- No. Running successfully does not preserve the trained attention topology.
- No. It is a 1024-dimensional
[CLS]feature; the model has no classifier.