JEPA4Japan · tutorials

Chapter 6 — Match the cards, but do not leave every card blank

630 words 3 min read #LeVJEPA#JEPA#self-supervised video#SIGReg

Understand why invariance is useful, why a constant answer satisfies it, and how dimensions can quietly collapse.

Course progress Course outline 34 of 34 lessons available

Part 0 — Get the map

  1. 01 Chapter 0 — Before you begin: what this course promises available now
  2. 02 Chapter 1 — One video, two windows available now
  3. 03 Chapter 2 — A walk along Yann LeCun’s research road available now
  4. 04 Chapter 3 — The JEPA family, without the name soup available now

Part 1 — Why a small objective can learn to see

  1. 05 Chapter 4 — Video can set its own homework available now
  2. 06 Chapter 5 — Keep the meaning; do not repaint every pixel available now
  3. 07 Chapter 6 — Match the cards, but do not leave every card blank Current lesson
  4. 08 Chapter 7 — SIGReg checks a cloud by looking at its shadows available now
  5. 09 Chapter 8 — The whole LeVJEPA objective on one line available now

Part 2 — Send a video through one encoder

  1. 10 Chapter 9 — How global and local views are paired available now
  2. 11 Chapter 10 — Cut a video into space-time tiles available now
  3. 12 Chapter 11 — One encoder, one projector, one summary card available now
  4. 13 Chapter 12 — One complete trip through the model available now
  5. 14 Chapter 13 — Why throwing away 95% can help available now
  6. 15 Chapter 14 — Same-frame teamwork, no peeking into tomorrow available now
  7. 16 Chapter 15 — RoPE, single-frame tubelets, and unexpectedly useful patch features available now

Part 3 — Read the experiments, not just the headline

  1. 17 Chapter 16 — What the four ablation ladders actually test available now
  2. 18 Chapter 17 — Equal epochs are not equal bills available now
  3. 19 Chapter 18 — What ImageNet, K400, and SSv2 are really asking available now
  4. 20 Chapter 19 — Keep the paper’s results in a ledger available now
  5. 21 Chapter 20 — Claims the evidence does not yet earn available now

Part 4 — From the official repository to your own experiment

  1. 22 Chapter 21 — A map of the official repository available now
  2. 23 Chapter 22 — Ten long walks become a training set available now
  3. 24 Chapter 23 — Read the defaults, then start training available now
  4. 25 Chapter 24 — Run a smoke test that cannot flatter you available now
  5. 26 Chapter 25 — Skip training: extract features from the public checkpoint available now
  6. 27 Chapter 26 — Freeze the encoder and test your own videos available now

Part 5 — Put the representation back on the world-model road

  1. 28 Chapter 27 — The important boundary: an encoder is not a planner available now
  2. 29 Chapter 28 — How LeVJEPA might feed a future world model available now
  3. 30 Chapter 29 — Ten projects, from first experiment to paper-sized question available now

Appendices — A backpack for the trail

  1. 31 Appendix A — The smallest useful math kit available now
  2. 32 Appendix B — The complete tensor-shape table available now
  3. 33 Appendix C — Glossary and paper timeline available now
  4. 34 Appendix D — Reproduction and review checklist available now

The perfect score that teaches nothing

  1. Wide dog viewwrite one card
  2. Close-up earmake the cards agree
  3. Cheap trickwrite “thing” every time
  4. Collapseagreement without information
“Two views of one video should match” is a good rule. “Every video gets one card” is its bad shortcut.

Suppose the teacher checks only whether the wide and narrow cards agree for each video. A flawless cheat is to write the number 0 for the cat, the train, and the pouring cup. Every pair matches. The loss is zero. The cards distinguish nothing.

The optimizer has not malfunctioned. The assignment allowed this answer. Once we say “pull paired views together,” we must also say how the entire class avoids crowding onto one point.

Two kinds of collapse

Invariance means a representation should stay stable under changes we consider irrelevant. LeVJEPA measures squared distance between a global embedding z_0 and each local z_v. Crops and color can differ; if the views still cover the same 16 frames, their summaries should be near. Gradients flow through both ends, so the global and local paths move together rather than acting as teacher and student.

With that term alone, the constant function z(x) = c is an exact solution: every difference is zero. This is complete collapse.

There is a quieter failure too. A vector may have hundreds of coordinates while only two independent directions vary. Other coordinates stay fixed or copy one another. This dimensional collapse leaves a representation that looks large on paper but carries little usable variation.

Other self-supervised systems fight these shortcuts with negatives, variance/covariance penalties, or asymmetric EMA-teacher and predictor designs. LeVJEPA chooses SIGReg. It looks at the cloud formed by a batch of embeddings and pushes that cloud toward a zero-mean, equal-scale, uncorrelated isotropic Gaussian. A dot or a flattened line cannot satisfy that shape.

The division of labor is clean. Invariance says how different views of one sample should relate. SIGReg asks what many samples together look like. Invariance alone permits one answer. SIGReg alone could make a beautiful cloud without matching the views of any video. The fixed weighted sum supplies both pressures.

Non-collapse is not understanding. A healthy cloud may encode wallpaper, camera model, or another shortcut. Dataset audits and downstream evaluations are still required.

Four points on paper

Write (0,0), (0,0), (0,0), (0,0). Any paired-view distance can be zero, yet the points vary in neither direction. Now write (−1,0), (1,0), (0,−1), (0,1). The cloud at least spreads out. Its shape alone still cannot tell you whether the points mean “running,” “pouring,” or “which camera filmed this.”

Evidence

The compact version

  1. Invariance alone gives a constant encoder a zero-loss escape route.
  2. SIGReg constrains the batch distribution to resist complete and dimensional collapse.
  3. “Did not collapse” means “did not degenerate,” not “understood the world.”

Check the shortcut

  1. Why can a constant encoder fool the invariance loss?
  2. Is gradient stopped on the global path?
  3. Does a well-spread cloud alone prove the model learned actions?
Answers
  1. Both sides of every pair produce exactly the same constant.
  2. No. Both sides receive gradients.
  3. No. Motion-specific evaluation is still necessary.