JEPA4Japan · tutorials

Chapter 4 — Video can set its own homework

660 words 3 min read #LeVJEPA#JEPA#self-supervised video#SIGReg

See where supervision comes from when nobody labels the clips—and what that free signal cannot promise.

Course progress Course outline 34 of 34 lessons available

Part 0 — Get the map

  1. 01 Chapter 0 — Before you begin: what this course promises available now
  2. 02 Chapter 1 — One video, two windows available now
  3. 03 Chapter 2 — A walk along Yann LeCun’s research road available now
  4. 04 Chapter 3 — The JEPA family, without the name soup available now

Part 1 — Why a small objective can learn to see

  1. 05 Chapter 4 — Video can set its own homework Current lesson
  2. 06 Chapter 5 — Keep the meaning; do not repaint every pixel available now
  3. 07 Chapter 6 — Match the cards, but do not leave every card blank available now
  4. 08 Chapter 7 — SIGReg checks a cloud by looking at its shadows available now
  5. 09 Chapter 8 — The whole LeVJEPA objective on one line available now

Part 2 — Send a video through one encoder

  1. 10 Chapter 9 — How global and local views are paired available now
  2. 11 Chapter 10 — Cut a video into space-time tiles available now
  3. 12 Chapter 11 — One encoder, one projector, one summary card available now
  4. 13 Chapter 12 — One complete trip through the model available now
  5. 14 Chapter 13 — Why throwing away 95% can help available now
  6. 15 Chapter 14 — Same-frame teamwork, no peeking into tomorrow available now
  7. 16 Chapter 15 — RoPE, single-frame tubelets, and unexpectedly useful patch features available now

Part 3 — Read the experiments, not just the headline

  1. 17 Chapter 16 — What the four ablation ladders actually test available now
  2. 18 Chapter 17 — Equal epochs are not equal bills available now
  3. 19 Chapter 18 — What ImageNet, K400, and SSv2 are really asking available now
  4. 20 Chapter 19 — Keep the paper’s results in a ledger available now
  5. 21 Chapter 20 — Claims the evidence does not yet earn available now

Part 4 — From the official repository to your own experiment

  1. 22 Chapter 21 — A map of the official repository available now
  2. 23 Chapter 22 — Ten long walks become a training set available now
  3. 24 Chapter 23 — Read the defaults, then start training available now
  4. 25 Chapter 24 — Run a smoke test that cannot flatter you available now
  5. 26 Chapter 25 — Skip training: extract features from the public checkpoint available now
  6. 27 Chapter 26 — Freeze the encoder and test your own videos available now

Part 5 — Put the representation back on the world-model road

  1. 28 Chapter 27 — The important boundary: an encoder is not a planner available now
  2. 29 Chapter 28 — How LeVJEPA might feed a future world model available now
  3. 30 Chapter 29 — Ten projects, from first experiment to paper-sized question available now

Appendices — A backpack for the trail

  1. 31 Appendix A — The smallest useful math kit available now
  2. 32 Appendix B — The complete tensor-shape table available now
  3. 33 Appendix C — Glossary and paper timeline available now
  4. 34 Appendix D — Reproduction and review checklist available now

Homework made from the video itself

  1. Receive videono human labels
  2. Make a questionseveral views of one clip
  3. Write answersencode each view
  4. Take an examfreeze, then evaluate
Self-supervised means the data supplies the exercise. It does not mean the model automatically learns everything worth knowing.

Nobody at Mina’s school has time to label every frame in ten thousand videos. Her teacher tries something cheaper. He cuts one recording into a wide view and a close-up, then asks, “Do these tell the same story?” The video itself reveals the pairing. No one has to type pouring water or opening a door.

Mina could still cheat by memorizing the wallpaper. Free homework does not guarantee causal understanding.

Where the learning signal comes from

A supervision signal is simply a computable rule telling the parameters how to change. In conventional supervised learning, a person supplies a category. In self-supervised learning, the structure of the data supplies the rule.

LeVJEPA draws global and local views from one 16-frame interval and treats their common clip identity as a positive pair. One shared encoder makes [CLS] embeddings for them. No class words, hand-drawn boxes, or action labels enter pretraining. The objective does not require designated negative videos either.

Video offers a rich source of regularity. Objects move while remaining the same objects. Hands meet tools. Events have an order. If both views preserve the whole time interval, some of those clues can survive their different crops. Still, the loss does not explicitly calculate optical flow, track instances, or judge causation. It asks same-clip summaries to agree and SIGReg to preserve a rich batch distribution. What actually emerges depends on the footage, crop policy, sparse tokens, capacity, and optimization.

Keep training homework separate from the final exam. Pretraining never sees the category answers from ImageNet, Kinetics-400, or Something-Something-v2. Afterward, researchers freeze the encoder and train a light probe on its features. The resulting score shows that a particular protocol could read some appearance or motion information. It does not demonstrate human common sense or planning.

Block-causal attention gives another limited promise: a frame’s tokens use only their own and earlier frames. That supports streaming. Passive viewing still never says, “If I execute action A, this will happen.” Such action-conditioned data and predictors belong to later world-model work.

Make one pair by hand

Take a two-second clip. Do not name its action. Duplicate it. Keep a broad view in one copy and crop the hand in the other, while retaining exactly the same 16 instants. You have made a self-supervised positive pair. Replace the second copy with another clip—even another clip containing a hand—and it is no longer the pair LeVJEPA defines.

Where the pairing comes from

Three takeaways

  1. In self-supervision, answers come from data structure rather than human class labels.
  2. LeVJEPA pairs multiple views of the same 16 frames; it does not directly supervise optical flow or causality.
  3. A frozen-probe score is evidence for one task and protocol—not proof of general understanding or planning.

Quick check

  1. Does LeVJEPA pretraining require action-class labels?
  2. Does same-clip agreement guarantee causal knowledge?
  3. Why must we name pretraining and frozen evaluation separately?
Answers
  1. No.
  2. No. It only provides spatiotemporal structure that the encoder may exploit.
  3. Pretraining learns the representation; the later protocol measures which information can be read from it.