Course progress Course outline 34 of 34 lessons available
Part 0 — Get the map
Part 1 — Why a small objective can learn to see
- 05 Chapter 4 — Video can set its own homework available now
- 06 Chapter 5 — Keep the meaning; do not repaint every pixel available now
- 07 Chapter 6 — Match the cards, but do not leave every card blank available now
- 08 Chapter 7 — SIGReg checks a cloud by looking at its shadows available now
- 09 Chapter 8 — The whole LeVJEPA objective on one line available now
Part 2 — Send a video through one encoder
- 10 Chapter 9 — How global and local views are paired available now
- 11 Chapter 10 — Cut a video into space-time tiles available now
- 12 Chapter 11 — One encoder, one projector, one summary card available now
- 13 Chapter 12 — One complete trip through the model available now
- 14 Chapter 13 — Why throwing away 95% can help available now
- 15 Chapter 14 — Same-frame teamwork, no peeking into tomorrow available now
- 16 Chapter 15 — RoPE, single-frame tubelets, and unexpectedly useful patch features available now
Part 3 — Read the experiments, not just the headline
- 17 Chapter 16 — What the four ablation ladders actually test available now
- 18 Chapter 17 — Equal epochs are not equal bills available now
- 19 Chapter 18 — What ImageNet, K400, and SSv2 are really asking available now
- 20 Chapter 19 — Keep the paper’s results in a ledger available now
- 21 Chapter 20 — Claims the evidence does not yet earn available now
Part 4 — From the official repository to your own experiment
- 22 Chapter 21 — A map of the official repository available now
- 23 Chapter 22 — Ten long walks become a training set available now
- 24 Chapter 23 — Read the defaults, then start training available now
- 25 Chapter 24 — Run a smoke test that cannot flatter you available now
- 26 Chapter 25 — Skip training: extract features from the public checkpoint available now
- 27 Chapter 26 — Freeze the encoder and test your own videos available now
Part 5 — Put the representation back on the world-model road
Appendices — A backpack for the trail
A shipping label for every tensor
- Video box
B,T,C,H,W - Viewsone global, V local
- Patches16×16 in each frame
- Sparse tokenstraining keeps 5%
- Summary cards
V+1,B,K
An airport receives a case arranged as passenger–time–color–height–width. The next machine expects passenger–color–time–height–width. If someone changes the label but not the order, the program may still run; it will simply treat frames as color channels.
That is why “a 5-D tensor” is not a useful specification. Name every axis, then assert shape, type, range, and normalization on the first batch.
Symbol passport
| Symbol | Meaning | Common value in the paper or public recipe |
|---|---|---|
B | batch per device | repository default 96; global effective batch is separate |
V | local views per clip | commonly 4 in controlled paper studies; repository default 10 |
T | input frames | 16 |
C | color channels | 3 |
Hg,Wg | global-view height and width | 224,224 |
Hl,Wl | local-view height and width | 96,96 |
P | spatial patch side | 16 |
tau | temporal tubelet length | default 1, so tokenization is frame by frame |
D | ViT hidden width | ViT-B 768; released ViT-L checkpoint 1024 |
K | projector output width | 256 |
rho | patch-token drop rate | 0.95 |
M | SIGReg projection directions | 1024 |
Q | SIGReg integration nodes | 17 |
“Common in paper studies” and “repository default” are different columns in an experiment ledger. In particular, V=4 and V=10 can both be correct when their protocols are named.
From a stored episode to several views
A raw clip sampled from one episode is usually:
[T,C,H_raw,W_raw]
The multi-view transform spatially crops the same T frames. After batching, the public loader produces:
| Name | Shape | Meaning |
|---|---|---|
global_frame | [B,T,C,224,224] | one near-global spatial crop without local photometric augmentation |
local_frames | [B,V,T,C,96,96] | V independent spatial crops with photometric augmentation |
Before the encoder, main.py changes the axis order:
global_video = rearrange(global_frame, "b t c h w -> b c t h w")
local_video = rearrange(local_frames, "b v t c h w -> (b v) c t h w")
The global input is now [B,C,T,224,224]; local input is [B*V,C,T,96,96]. Folding views into the batch lets one shared encoder process every local crop. The code unfolds that axis afterward.
Count the patches before dropping any
With tau=1, each token spans one frame and one 16×16 spatial patch:
global patches = 16 * (224/16) * (224/16) = 16 * 14 * 14 = 3136
local patches = 16 * ( 96/16) * ( 96/16) = 16 * 6 * 6 = 576
After patch embedding and before adding [CLS]:
| Branch | Patch tokens | Shape |
|---|---|---|
| global | 3136 | [B,3136,D] |
| local | 576 | [B*V,576,D] |
Changing tau to 2 would halve the temporal positions to 8 and make each token span two frames. LeVJEPA defaults to tau=1; quietly importing a tubelet-2 recipe from V-JEPA or VideoMAE changes the model.
What 95% dropping does to the sequence
The public code computes:
keep_len = round(N_patches * (1 - rho))
It samples an independent random keep-set for each item:
| Branch | Original patches | Kept patches | Length after [CLS] |
|---|---|---|---|
| global | 3136 | round(156.8)=157 | 158 |
| local | 576 | round(28.8)=29 | 30 |
Those are training shapes. Evaluation and inference do not drop patches. On a 16×224×224 input, the released ViT-L therefore returns all 3136+1=3137 tokens. Using the sparse training length as the public checkpoint’s output length will misalign every dense-feature operation.
Encoder, summary, and projection shapes
At the default drop rate during training:
global_tokens : [B, 158,D]
local_tokens : [B*V, 30,D]
global_cls : [B, 1,D]
local_cls : [B, V,D]
all cls : [B, V+1,D]
projected z : [B, V+1,K]
The code keeps the global projection as embeddings[:, :1] with shape [B,1,K], broadcasts it, and compares it with [B,V+1,K]. The global-versus-itself term is always zero but remains inside the mean, matching the paper formula and public code.
Before SIGReg, axes change once more:
[B,V+1,K] -> [V+1,B,K]
Each view position is then checked across B different clips. The views of one clip are not treated as additional independent population samples.
Inside SIGReg
Let its input be [S,B,K], where S=V+1:
| Operation | Shape | Meaning |
|---|---|---|
unit directions A | [K,M] | M=1024 flashlight directions |
projected z @ A | [S,B,M] | one shadow per sample and direction |
| multiply by nodes | [S,B,M,Q] | evaluate Q=17 frequencies |
| batch-mean cos and sin | [2,S,M,Q] | real and imaginary empirical fingerprint |
| error from Gaussian fingerprint | [S,M,Q] | each view, direction, and node |
| final loss | [] | weighted mean over those axes |
Distributed all-reduce aggregates [2,S,M,Q] characteristic-function statistics rather than transmitting every [B,K] embedding. That is why the communication size is largely independent of batch size and embedding width.
The released inference interface
The Hugging Face model expects:
pixel_values: [B,C,T,H,W]
For [1,3,16,224,224], its ViT-L/16 interface reports:
last_hidden_state: [1,3137,1024] # CLS + 16*14*14 patches
pooler_output: [1,1024] # CLS
Repeating one image across 16 frames makes a valid input tensor, but it does not create motion. It is only an adapter from a still image to a video encoder.
Five assertions worth keeping in the test suite
assert global_frame.ndim == 5 and global_frame.shape[1:3] == (16, 3)
assert local_frames.ndim == 6 and local_frames.shape[2:4] == (16, 3)
assert global_frame.shape[-2:] == (224, 224)
assert local_frames.shape[-2:] == (96, 96)
assert torch.isfinite(loss).all()
Also log dtype/min/max/mean/std before and after normalization. A perfectly shaped uint8 tensor in the range 0–255 is still wrong when the model expects normalized floats.
Shape records
LeVJEPA v1 defines the architecture and notation. Pinned main.py, module.py, and data/loader.py define the executable shapes. The released model card defines the checkpoint interface.
Three labels not to swap
- The loader uses
[B,T,C,H,W]; the encoder uses[B,C,T,H,W]. - The 95% drop happens during training; full inference returns 3,137 tokens.
- SIGReg’s batch axis contains different clips, not the views of one clip.
Inspect the manifest
- Why does the local branch enter the encoder as
[B*V,C,T,96,96]? - Why is the global training sequence 158 tokens while the released checkpoint returns 3,137?
- What goes wrong conceptually if
[B,V+1,K]is flattened into[B*(V+1),K]for SIGReg?
Check the answers
- The view axis is temporarily folded into the batch so every crop runs through the same encoder in parallel.
- Training keeps 157 patches and adds
[CLS]; evaluation keeps all 3,136 patches and adds[CLS]. - Strongly related views from one clip are counted as independent population members, changing the estimated distribution and effective sample size.