JEPA4Japan · tutorials

Chapter 10 — Cut a video into space-time tiles

625 words 3 min read #LeVJEPA#JEPA#self-supervised video#SIGReg

Count the tokens made from 16 frames, 16×16 patches, per-frame tubelets, and positional coordinates.

Course progress Course outline 34 of 34 lessons available

Part 0 — Get the map

  1. 01 Chapter 0 — Before you begin: what this course promises available now
  2. 02 Chapter 1 — One video, two windows available now
  3. 03 Chapter 2 — A walk along Yann LeCun’s research road available now
  4. 04 Chapter 3 — The JEPA family, without the name soup available now

Part 1 — Why a small objective can learn to see

  1. 05 Chapter 4 — Video can set its own homework available now
  2. 06 Chapter 5 — Keep the meaning; do not repaint every pixel available now
  3. 07 Chapter 6 — Match the cards, but do not leave every card blank available now
  4. 08 Chapter 7 — SIGReg checks a cloud by looking at its shadows available now
  5. 09 Chapter 8 — The whole LeVJEPA objective on one line available now

Part 2 — Send a video through one encoder

  1. 10 Chapter 9 — How global and local views are paired available now
  2. 11 Chapter 10 — Cut a video into space-time tiles Current lesson
  3. 12 Chapter 11 — One encoder, one projector, one summary card available now
  4. 13 Chapter 12 — One complete trip through the model available now
  5. 14 Chapter 13 — Why throwing away 95% can help available now
  6. 15 Chapter 14 — Same-frame teamwork, no peeking into tomorrow available now
  7. 16 Chapter 15 — RoPE, single-frame tubelets, and unexpectedly useful patch features available now

Part 3 — Read the experiments, not just the headline

  1. 17 Chapter 16 — What the four ablation ladders actually test available now
  2. 18 Chapter 17 — Equal epochs are not equal bills available now
  3. 19 Chapter 18 — What ImageNet, K400, and SSv2 are really asking available now
  4. 20 Chapter 19 — Keep the paper’s results in a ledger available now
  5. 21 Chapter 20 — Claims the evidence does not yet earn available now

Part 4 — From the official repository to your own experiment

  1. 22 Chapter 21 — A map of the official repository available now
  2. 23 Chapter 22 — Ten long walks become a training set available now
  3. 24 Chapter 23 — Read the defaults, then start training available now
  4. 25 Chapter 24 — Run a smoke test that cannot flatter you available now
  5. 26 Chapter 25 — Skip training: extract features from the public checkpoint available now
  6. 27 Chapter 26 — Freeze the encoder and test your own videos available now

Part 5 — Put the representation back on the world-model road

  1. 28 Chapter 27 — The important boundary: an encoder is not a planner available now
  2. 29 Chapter 28 — How LeVJEPA might feed a future world model available now
  3. 30 Chapter 29 — Ten projects, from first experiment to paper-sized question available now

Appendices — A backpack for the trail

  1. 31 Appendix A — The smallest useful math kit available now
  2. 32 Appendix B — The complete tensor-shape table available now
  3. 33 Appendix C — Glossary and paper timeline available now
  4. 34 Appendix D — Reproduction and review checklist available now

Turn a flipbook into numbered coins

  1. 16 framesone short clip
  2. 16×16 tilescut every frame
  3. Make tokensone coin per tile
  4. Stamp addressestime, height, width
  5. Add `[CLS]`the clip’s summary card
With the default tubelet size, a patch token begins in one frame; neighboring frames are not fused at the input.

Lay out all 16 pages of a flipbook and draw the same grid on each. Every square becomes a numerical coin stamped with its page, row, and column. A special [CLS] coin sits at the mouth of the bag to gather a summary. The Transformer can now work with addressed tokens rather than raw pixels.

Count before thinking

Ignore batch and RGB channels for a moment. LeVJEPA’s default spatial patch is 16 × 16; temporal tubelet size is τ = 1. The 3D patch embedding spans only one frame at a time. It does not blend two neighboring frames into one initial token.

For a 16-frame 224 × 224 global view:

grid per frame = (224 / 16) × (224 / 16) = 14 × 14 = 196
patch tokens    = 16 × 196 = 3136

For a 16-frame 96 × 96 local view:

grid per frame = (96 / 16) × (96 / 16) = 6 × 6 = 36
patch tokens    = 16 × 36 = 576

Patch embedding maps each tile to width d, producing a sequence shaped like [T × Hpatch × Wpatch, d]. The published large-model recipe uses factorized 3D RoPE to encode relative time, vertical position, and horizontal position. This lets one encoder accept both grid sizes without interpolating a learned absolute position table. RoPE describes positions; the attention mask, not RoPE, controls future visibility.

During training, 95% of patch tokens are sampled away after embedding. With the pinned code’s rounding, the model keeps about 157 of 3136 global patches and 29 of 576 local patches. Only then is the always-kept learned [CLS] token prepended. The loss directly reads projected [CLS]; no decoder restores the missing patches. Evaluation and inference turn dropping off and use the full token set.

τ = 1 is the paper and released large model’s choice, not a law for every video Transformer. With τ = 2, 16 frames make eight temporal slots and half as many initial patch tokens, but adjacent frames have already been fused. A comparison must account for keep rate and evaluation length too.

Do one count yourself

Take 8 frames at 128 × 128, patch size 16, τ = 1. Each frame has 8 × 8 = 64 patches; the clip has 8 × 64 = 512. Add [CLS] to reach 513. Never include [CLS] in the 95% patch-dropping pool.

Tokenization records

Three numbers to retain

  1. τ = 1: one initial patch token spans one 16 × 16 area in one frame.
  2. A global clip has 3136 patch tokens and a local clip 576, before [CLS].
  3. Training discards 95% without reconstruction; evaluation uses all patches.

Token count test

  1. What does τ = 1 mean?
  2. Why are there 3136 global patch tokens?
  3. Can random dropping remove [CLS]?
Answers
  1. Each token’s initial temporal span is one frame.
  2. 16 × 14 × 14 = 3136.
  3. No. It is appended after patch sampling and retained as the clip summary.