JEPA4Japan · tutorials

Chapter 13 — Why throwing away 95% can help

649 words 3 min read #LeVJEPA#JEPA#self-supervised video#SIGReg

Separate random token dropping from tube masking, then weigh compute savings, augmentation, and lost motion clues.

Course progress Course outline 34 of 34 lessons available

Part 0 — Get the map

  1. 01 Chapter 0 — Before you begin: what this course promises available now
  2. 02 Chapter 1 — One video, two windows available now
  3. 03 Chapter 2 — A walk along Yann LeCun’s research road available now
  4. 04 Chapter 3 — The JEPA family, without the name soup available now

Part 1 — Why a small objective can learn to see

  1. 05 Chapter 4 — Video can set its own homework available now
  2. 06 Chapter 5 — Keep the meaning; do not repaint every pixel available now
  3. 07 Chapter 6 — Match the cards, but do not leave every card blank available now
  4. 08 Chapter 7 — SIGReg checks a cloud by looking at its shadows available now
  5. 09 Chapter 8 — The whole LeVJEPA objective on one line available now

Part 2 — Send a video through one encoder

  1. 10 Chapter 9 — How global and local views are paired available now
  2. 11 Chapter 10 — Cut a video into space-time tiles available now
  3. 12 Chapter 11 — One encoder, one projector, one summary card available now
  4. 13 Chapter 12 — One complete trip through the model available now
  5. 14 Chapter 13 — Why throwing away 95% can help Current lesson
  6. 15 Chapter 14 — Same-frame teamwork, no peeking into tomorrow available now
  7. 16 Chapter 15 — RoPE, single-frame tubelets, and unexpectedly useful patch features available now

Part 3 — Read the experiments, not just the headline

  1. 17 Chapter 16 — What the four ablation ladders actually test available now
  2. 18 Chapter 17 — Equal epochs are not equal bills available now
  3. 19 Chapter 18 — What ImageNet, K400, and SSv2 are really asking available now
  4. 20 Chapter 19 — Keep the paper’s results in a ledger available now
  5. 21 Chapter 20 — Claims the evidence does not yet earn available now

Part 4 — From the official repository to your own experiment

  1. 22 Chapter 21 — A map of the official repository available now
  2. 23 Chapter 22 — Ten long walks become a training set available now
  3. 24 Chapter 23 — Read the defaults, then start training available now
  4. 25 Chapter 24 — Run a smoke test that cannot flatter you available now
  5. 26 Chapter 25 — Skip training: extract features from the public checkpoint available now
  6. 27 Chapter 26 — Freeze the encoder and test your own videos available now

Part 5 — Put the representation back on the world-model road

  1. 28 Chapter 27 — The important boundary: an encoder is not a planner available now
  2. 29 Chapter 28 — How LeVJEPA might feed a future world model available now
  3. 30 Chapter 29 — Ten projects, from first experiment to paper-sized question available now

Appendices — A backpack for the trail

  1. 31 Appendix A — The smallest useful math kit available now
  2. 32 Appendix B — The complete tensor-shape table available now
  3. 33 Appendix C — Glossary and paper timeline available now
  4. 34 Appendix D — Reproduction and review checklist available now

Read a moving puzzle through a handful of pieces

  1. Full puzzle3136 patches
  2. Draw at randomnew places each pass
  3. Send only 5%157 patches
  4. Summarize the clipdo not fill the holes
Dropped tokens never enter the Transformer. The task is robust summarization from sparse clues, not completion.

Close your eyes and grab a few tiles from an animated picture book. This time you find a wheel and lamp; next time, a driver and patch of sky. Because the visible places keep changing, you have to piece together the event from shifting clues. A tube-shaped mask, by contrast, might hide the hero’s position in every frame.

Do not stretch the story into reconstruction. LeVJEPA has no mask token, decoder, or fill-in-the-pixels objective.

What the code removes

Sampling happens after patch embedding. For each sample, the code assigns random values across the full space-time grid and keeps the required subset. [CLS] is prepended afterward, so it always survives. Dropping runs only in training mode; evaluation and inference use the full sequence.

ViewPatches beforePatches after 95% dropWith [CLS]
16 frames, 224, patch 163136157158
16 frames, 96, patch 165762930

In a controlled sweep, the authors report ImageNet-1K rising from 33.9 with no dropping to 47.6 at 95%; 90% and 95% give 47.4 and 47.6. Their explanation has two parts: fewer tokens reduce work per step, while changing sparse observations act as augmentation. The paper’s 1 / (1 − ρ) = 20 is a theoretical reduction factor for feed-forward work at ρ=0.95, not a guarantee that every operator—or wall-clock training—becomes 20× faster.

The pattern matters too. Random keeping versus a same-location-through-time tube gives author-reported ImageNet 50.7 versus 39.6. In the stated tubelet=2 comparison, SSv2 is 28.8 versus 26.4. Yet extreme sparsity is not a free victory: in short training, SSv2 declines once dropping exceeds 0.3. Longer training helps in the paper, and motion-preserving sparse schemes remain an open question.

There is also a caption wrinkle. Section 4 broadly says “frozen attentive probe,” while the arXiv v1 Figure 4 caption says “linear probing”; the project page says attentive probing. Treat sweep numbers as reliable within-figure trends, not as free-standing values to compare across tables with different protocols.

Draw random versus tube

Sketch four frames with four cells each. Circle different cells in every frame for random keeping. Circle the same spatial cell in all four for a tube. Which scheme can hide an object sitting at one location for the entire clip? The tube can.

Evidence trail

What survives the drop

  1. Dropping is real removal during training; discarded tokens are neither encoded nor reconstructed.
  2. At 95%, global/local views keep 157/29 patches; [CLS] is separate and always present.
  3. The authors report better ImageNet behavior and efficiency, but short-run motion performance can fall.

Sparse-reading test

  1. Are removed positions replaced by mask tokens?
  2. How long is the global sequence including [CLS] after 95% dropping?
  3. Does the ImageNet trend prove 95% is best for every motion task?
Answers
  1. No. They do not enter the Transformer.
  2. 158 tokens.
  3. No. The paper reports a short-training SSv2 decline at high drop rates.