JEPA4Japan · tutorials

Appendix B — The complete tensor-shape table

1,299 words 6 min read #LeVJEPA#JEPA#self-supervised video#SIGReg

Check every step from [B,T,C,H,W] to views, sparse tokens, [CLS], projector output, and SIGReg input.

Course progress Course outline 34 of 34 lessons available

Part 0 — Get the map

  1. 01 Chapter 0 — Before you begin: what this course promises available now
  2. 02 Chapter 1 — One video, two windows available now
  3. 03 Chapter 2 — A walk along Yann LeCun’s research road available now
  4. 04 Chapter 3 — The JEPA family, without the name soup available now

Part 1 — Why a small objective can learn to see

  1. 05 Chapter 4 — Video can set its own homework available now
  2. 06 Chapter 5 — Keep the meaning; do not repaint every pixel available now
  3. 07 Chapter 6 — Match the cards, but do not leave every card blank available now
  4. 08 Chapter 7 — SIGReg checks a cloud by looking at its shadows available now
  5. 09 Chapter 8 — The whole LeVJEPA objective on one line available now

Part 2 — Send a video through one encoder

  1. 10 Chapter 9 — How global and local views are paired available now
  2. 11 Chapter 10 — Cut a video into space-time tiles available now
  3. 12 Chapter 11 — One encoder, one projector, one summary card available now
  4. 13 Chapter 12 — One complete trip through the model available now
  5. 14 Chapter 13 — Why throwing away 95% can help available now
  6. 15 Chapter 14 — Same-frame teamwork, no peeking into tomorrow available now
  7. 16 Chapter 15 — RoPE, single-frame tubelets, and unexpectedly useful patch features available now

Part 3 — Read the experiments, not just the headline

  1. 17 Chapter 16 — What the four ablation ladders actually test available now
  2. 18 Chapter 17 — Equal epochs are not equal bills available now
  3. 19 Chapter 18 — What ImageNet, K400, and SSv2 are really asking available now
  4. 20 Chapter 19 — Keep the paper’s results in a ledger available now
  5. 21 Chapter 20 — Claims the evidence does not yet earn available now

Part 4 — From the official repository to your own experiment

  1. 22 Chapter 21 — A map of the official repository available now
  2. 23 Chapter 22 — Ten long walks become a training set available now
  3. 24 Chapter 23 — Read the defaults, then start training available now
  4. 25 Chapter 24 — Run a smoke test that cannot flatter you available now
  5. 26 Chapter 25 — Skip training: extract features from the public checkpoint available now
  6. 27 Chapter 26 — Freeze the encoder and test your own videos available now

Part 5 — Put the representation back on the world-model road

  1. 28 Chapter 27 — The important boundary: an encoder is not a planner available now
  2. 29 Chapter 28 — How LeVJEPA might feed a future world model available now
  3. 30 Chapter 29 — Ten projects, from first experiment to paper-sized question available now

Appendices — A backpack for the trail

  1. 31 Appendix A — The smallest useful math kit available now
  2. 32 Appendix B — The complete tensor-shape table Current lesson
  3. 33 Appendix C — Glossary and paper timeline available now
  4. 34 Appendix D — Reproduction and review checklist available now

A shipping label for every tensor

  1. Video boxB,T,C,H,W
  2. Viewsone global, V local
  3. Patches16×16 in each frame
  4. Sparse tokenstraining keeps 5%
  5. Summary cardsV+1,B,K
A shape table is the conveyor belt’s manifest. Check the labels at every door.

An airport receives a case arranged as passenger–time–color–height–width. The next machine expects passenger–color–time–height–width. If someone changes the label but not the order, the program may still run; it will simply treat frames as color channels.

That is why “a 5-D tensor” is not a useful specification. Name every axis, then assert shape, type, range, and normalization on the first batch.

Symbol passport

SymbolMeaningCommon value in the paper or public recipe
Bbatch per devicerepository default 96; global effective batch is separate
Vlocal views per clipcommonly 4 in controlled paper studies; repository default 10
Tinput frames16
Ccolor channels3
Hg,Wgglobal-view height and width224,224
Hl,Wllocal-view height and width96,96
Pspatial patch side16
tautemporal tubelet lengthdefault 1, so tokenization is frame by frame
DViT hidden widthViT-B 768; released ViT-L checkpoint 1024
Kprojector output width256
rhopatch-token drop rate0.95
MSIGReg projection directions1024
QSIGReg integration nodes17

“Common in paper studies” and “repository default” are different columns in an experiment ledger. In particular, V=4 and V=10 can both be correct when their protocols are named.

From a stored episode to several views

A raw clip sampled from one episode is usually:

[T,C,H_raw,W_raw]

The multi-view transform spatially crops the same T frames. After batching, the public loader produces:

NameShapeMeaning
global_frame[B,T,C,224,224]one near-global spatial crop without local photometric augmentation
local_frames[B,V,T,C,96,96]V independent spatial crops with photometric augmentation

Before the encoder, main.py changes the axis order:

global_video = rearrange(global_frame, "b t c h w -> b c t h w")
local_video = rearrange(local_frames, "b v t c h w -> (b v) c t h w")

The global input is now [B,C,T,224,224]; local input is [B*V,C,T,96,96]. Folding views into the batch lets one shared encoder process every local crop. The code unfolds that axis afterward.

Count the patches before dropping any

With tau=1, each token spans one frame and one 16×16 spatial patch:

global patches = 16 * (224/16) * (224/16) = 16 * 14 * 14 = 3136
local patches  = 16 * ( 96/16) * ( 96/16) = 16 *  6 *  6 =  576

After patch embedding and before adding [CLS]:

BranchPatch tokensShape
global3136[B,3136,D]
local576[B*V,576,D]

Changing tau to 2 would halve the temporal positions to 8 and make each token span two frames. LeVJEPA defaults to tau=1; quietly importing a tubelet-2 recipe from V-JEPA or VideoMAE changes the model.

What 95% dropping does to the sequence

The public code computes:

keep_len = round(N_patches * (1 - rho))

It samples an independent random keep-set for each item:

BranchOriginal patchesKept patchesLength after [CLS]
global3136round(156.8)=157158
local576round(28.8)=2930

Those are training shapes. Evaluation and inference do not drop patches. On a 16×224×224 input, the released ViT-L therefore returns all 3136+1=3137 tokens. Using the sparse training length as the public checkpoint’s output length will misalign every dense-feature operation.

Encoder, summary, and projection shapes

At the default drop rate during training:

global_tokens : [B,       158,D]
local_tokens  : [B*V,      30,D]
global_cls    : [B,         1,D]
local_cls     : [B,         V,D]
all cls       : [B,       V+1,D]
projected z   : [B,       V+1,K]

The code keeps the global projection as embeddings[:, :1] with shape [B,1,K], broadcasts it, and compares it with [B,V+1,K]. The global-versus-itself term is always zero but remains inside the mean, matching the paper formula and public code.

Before SIGReg, axes change once more:

[B,V+1,K] -> [V+1,B,K]

Each view position is then checked across B different clips. The views of one clip are not treated as additional independent population samples.

Inside SIGReg

Let its input be [S,B,K], where S=V+1:

OperationShapeMeaning
unit directions A[K,M]M=1024 flashlight directions
projected z @ A[S,B,M]one shadow per sample and direction
multiply by nodes[S,B,M,Q]evaluate Q=17 frequencies
batch-mean cos and sin[2,S,M,Q]real and imaginary empirical fingerprint
error from Gaussian fingerprint[S,M,Q]each view, direction, and node
final loss[]weighted mean over those axes

Distributed all-reduce aggregates [2,S,M,Q] characteristic-function statistics rather than transmitting every [B,K] embedding. That is why the communication size is largely independent of batch size and embedding width.

The released inference interface

The Hugging Face model expects:

pixel_values: [B,C,T,H,W]

For [1,3,16,224,224], its ViT-L/16 interface reports:

last_hidden_state: [1,3137,1024]  # CLS + 16*14*14 patches
pooler_output:     [1,1024]       # CLS

Repeating one image across 16 frames makes a valid input tensor, but it does not create motion. It is only an adapter from a still image to a video encoder.

Five assertions worth keeping in the test suite

assert global_frame.ndim == 5 and global_frame.shape[1:3] == (16, 3)
assert local_frames.ndim == 6 and local_frames.shape[2:4] == (16, 3)
assert global_frame.shape[-2:] == (224, 224)
assert local_frames.shape[-2:] == (96, 96)
assert torch.isfinite(loss).all()

Also log dtype/min/max/mean/std before and after normalization. A perfectly shaped uint8 tensor in the range 0–255 is still wrong when the model expects normalized floats.

Shape records

LeVJEPA v1 defines the architecture and notation. Pinned main.py, module.py, and data/loader.py define the executable shapes. The released model card defines the checkpoint interface.

Three labels not to swap

  1. The loader uses [B,T,C,H,W]; the encoder uses [B,C,T,H,W].
  2. The 95% drop happens during training; full inference returns 3,137 tokens.
  3. SIGReg’s batch axis contains different clips, not the views of one clip.

Inspect the manifest

  1. Why does the local branch enter the encoder as [B*V,C,T,96,96]?
  2. Why is the global training sequence 158 tokens while the released checkpoint returns 3,137?
  3. What goes wrong conceptually if [B,V+1,K] is flattened into [B*(V+1),K] for SIGReg?
Check the answers
  1. The view axis is temporarily folded into the batch so every crop runs through the same encoder in parallel.
  2. Training keeps 157 patches and adds [CLS]; evaluation keeps all 3,136 patches and adds [CLS].
  3. Strongly related views from one clip are counted as independent population members, changing the estimated distribution and effective sample size.