JEPA4Japan · tutorials

Chapter 5 — Keep the meaning; do not repaint every pixel

616 words 3 min read #LeVJEPA#JEPA#self-supervised video#SIGReg

Compare pixel reconstruction, feature prediction, and view invariance, then pin down what LeVJEPA does and does not predict.

Course progress Course outline 34 of 34 lessons available

Part 0 — Get the map

  1. 01 Chapter 0 — Before you begin: what this course promises available now
  2. 02 Chapter 1 — One video, two windows available now
  3. 03 Chapter 2 — A walk along Yann LeCun’s research road available now
  4. 04 Chapter 3 — The JEPA family, without the name soup available now

Part 1 — Why a small objective can learn to see

  1. 05 Chapter 4 — Video can set its own homework available now
  2. 06 Chapter 5 — Keep the meaning; do not repaint every pixel Current lesson
  3. 07 Chapter 6 — Match the cards, but do not leave every card blank available now
  4. 08 Chapter 7 — SIGReg checks a cloud by looking at its shadows available now
  5. 09 Chapter 8 — The whole LeVJEPA objective on one line available now

Part 2 — Send a video through one encoder

  1. 10 Chapter 9 — How global and local views are paired available now
  2. 11 Chapter 10 — Cut a video into space-time tiles available now
  3. 12 Chapter 11 — One encoder, one projector, one summary card available now
  4. 13 Chapter 12 — One complete trip through the model available now
  5. 14 Chapter 13 — Why throwing away 95% can help available now
  6. 15 Chapter 14 — Same-frame teamwork, no peeking into tomorrow available now
  7. 16 Chapter 15 — RoPE, single-frame tubelets, and unexpectedly useful patch features available now

Part 3 — Read the experiments, not just the headline

  1. 17 Chapter 16 — What the four ablation ladders actually test available now
  2. 18 Chapter 17 — Equal epochs are not equal bills available now
  3. 19 Chapter 18 — What ImageNet, K400, and SSv2 are really asking available now
  4. 20 Chapter 19 — Keep the paper’s results in a ledger available now
  5. 21 Chapter 20 — Claims the evidence does not yet earn available now

Part 4 — From the official repository to your own experiment

  1. 22 Chapter 21 — A map of the official repository available now
  2. 23 Chapter 22 — Ten long walks become a training set available now
  3. 24 Chapter 23 — Read the defaults, then start training available now
  4. 25 Chapter 24 — Run a smoke test that cannot flatter you available now
  5. 26 Chapter 25 — Skip training: extract features from the public checkpoint available now
  6. 27 Chapter 26 — Freeze the encoder and test your own videos available now

Part 5 — Put the representation back on the world-model road

  1. 28 Chapter 27 — The important boundary: an encoder is not a planner available now
  2. 29 Chapter 28 — How LeVJEPA might feed a future world model available now
  3. 30 Chapter 29 — Ten projects, from first experiment to paper-sized question available now

Appendices — A backpack for the trail

  1. 31 Appendix A — The smallest useful math kit available now
  2. 32 Appendix B — The complete tensor-shape table available now
  3. 33 Appendix C — Glossary and paper timeline available now
  4. 34 Appendix D — Reproduction and review checklist available now

Three kinds of homework

  1. Replace pixelspaint the missing squares
  2. Predict featuresguess the hidden region’s idea
  3. Match summariessame clip, nearby cards
LeVJEPA takes the third assignment. It neither restores discarded tokens nor repaints their pixels.

A teacher covers half a picture of a dog chasing a ball. One pupil must redraw every blade of grass and shadow. Another guesses a high-level feature for the hidden region. Mina gets a different task: compare a wide shot with a crop showing only the dog’s head, then write two nearby story cards.

All three exercises can be made without human labels. They are still different objectives.

Follow the answer tensor

Pixel reconstruction makes color values the answer. Methods such as VideoMAE hide most space-time patches and use a decoder to rebuild them. Neighboring video frames make low-level copying unusually easy, which is one reason structured tube masks make sense for that assignment.

Feature prediction changes the answer from pixels to a learned representation. In I-JEPA and V-JEPA, a context encoder and predictor estimate the encoded target regions. V-JEPA gets those targets from a stopped-gradient EMA target encoder. It avoids repainting every texture, but it most certainly has a “predict the hidden target” graph.

LeVJEPA view invariance asks something else. One global and several local views go through the same encoder and projector. The loss reads only their [CLS] embeddings and brings the local summaries toward the global summary. No predictor. No target encoder. No stop-gradient. No decoder. The global summary changes under gradient descent; it is not a fixed label.

This is why “token dropping” should not be casually renamed “masked prediction.” Randomly discarded LeVJEPA patch tokens never enter the encoder, and no term asks the model to recover them. The surviving sparse tokens are all the model sees on that pass. Dropping changes the compute bill and the difficulty of observation.

“Keep the meaning” also needs a footnote. The objective rewards information shared across crops and appearance perturbations; it does not announce which factors are semantic or which must be ignored. If a downstream task depends on a tiny motion and sparse observation regularly deletes its clue, the representation can suffer. Frozen probes, retrieval, and dense tasks must settle that question empirically.

A reliable diagnostic

Whenever you meet a self-supervised loss, ask: Where does the answer tensor come from? Original color values imply reconstruction. Hidden-region output from another encoder branch implies feature prediction. A distance between global and local [CLS] cards is the LeVJEPA-style view-invariance case.

Source shelf

Take these with you

  1. Reconstructing pixels, predicting hidden features, and matching view summaries are three separate training questions.
  2. LeVJEPA neither predicts discarded tokens nor rebuilds pixels.
  3. It brings same-clip [CLS] embeddings together; downstream evidence must tell us whether their meaning is useful.

Spot the graph

  1. Why does LeVJEPA need no decoder?
  2. Is there a reconstruction loss on randomly dropped tokens?
  3. What is the clearest graph-level difference between V-JEPA and LeVJEPA?
Answers
  1. It does not reconstruct pixels.
  2. No.
  3. V-JEPA uses a predictor, an EMA target, and stop-gradient. LeVJEPA’s objective uses a shared encoder/projector and SIGReg.