JEPA4Japan · tutorials

Chapter 1 — One video, two windows

597 words 3 min read #LeVJEPA#JEPA#self-supervised video#SIGReg

Meet global and local views through a street scene, and learn why both windows must cover the same 16 frames.

Course progress Course outline 34 of 34 lessons available

Part 0 — Get the map

  1. 01 Chapter 0 — Before you begin: what this course promises available now
  2. 02 Chapter 1 — One video, two windows Current lesson
  3. 03 Chapter 2 — A walk along Yann LeCun’s research road available now
  4. 04 Chapter 3 — The JEPA family, without the name soup available now

Part 1 — Why a small objective can learn to see

  1. 05 Chapter 4 — Video can set its own homework available now
  2. 06 Chapter 5 — Keep the meaning; do not repaint every pixel available now
  3. 07 Chapter 6 — Match the cards, but do not leave every card blank available now
  4. 08 Chapter 7 — SIGReg checks a cloud by looking at its shadows available now
  5. 09 Chapter 8 — The whole LeVJEPA objective on one line available now

Part 2 — Send a video through one encoder

  1. 10 Chapter 9 — How global and local views are paired available now
  2. 11 Chapter 10 — Cut a video into space-time tiles available now
  3. 12 Chapter 11 — One encoder, one projector, one summary card available now
  4. 13 Chapter 12 — One complete trip through the model available now
  5. 14 Chapter 13 — Why throwing away 95% can help available now
  6. 15 Chapter 14 — Same-frame teamwork, no peeking into tomorrow available now
  7. 16 Chapter 15 — RoPE, single-frame tubelets, and unexpectedly useful patch features available now

Part 3 — Read the experiments, not just the headline

  1. 17 Chapter 16 — What the four ablation ladders actually test available now
  2. 18 Chapter 17 — Equal epochs are not equal bills available now
  3. 19 Chapter 18 — What ImageNet, K400, and SSv2 are really asking available now
  4. 20 Chapter 19 — Keep the paper’s results in a ledger available now
  5. 21 Chapter 20 — Claims the evidence does not yet earn available now

Part 4 — From the official repository to your own experiment

  1. 22 Chapter 21 — A map of the official repository available now
  2. 23 Chapter 22 — Ten long walks become a training set available now
  3. 24 Chapter 23 — Read the defaults, then start training available now
  4. 25 Chapter 24 — Run a smoke test that cannot flatter you available now
  5. 26 Chapter 25 — Skip training: extract features from the public checkpoint available now
  6. 27 Chapter 26 — Freeze the encoder and test your own videos available now

Part 5 — Put the representation back on the world-model road

  1. 28 Chapter 27 — The important boundary: an encoder is not a planner available now
  2. 29 Chapter 28 — How LeVJEPA might feed a future world model available now
  3. 30 Chapter 29 — Ten projects, from first experiment to paper-sized question available now

Appendices — A backpack for the trail

  1. 31 Appendix A — The smallest useful math kit available now
  2. 32 Appendix B — The complete tensor-shape table available now
  3. 33 Appendix C — Glossary and paper timeline available now
  4. 34 Appendix D — Reproduction and review checklist available now

Two windows, one moment

  1. Choose 16 framesone short stretch of time
  2. Open widesee the scene
  3. Look closesee one detail
  4. Compare cardsone event should feel alike
The windows may change where you look. They must not change when you look.

A scooter rolls past a corner. Through a large window you see the rider, road, and storefront. Through a paper tube you see only the turning wheel. The pixels barely match, but both views show the same pass.

Now delay the paper tube by three seconds. The scooter is gone; a dog has entered the frame. Forcing the two summaries to agree would teach the wrong lesson.

That is the pairing rule in plain language: space may change; temporal identity may not. “The same source video” is too loose. Both views must come from the same 16-frame interval.

Under the picture

A view is a transformed version of a clip. An encoder turns pixels into features. An embedding is the compact numerical card we compare.

From one 16-frame clip, LeVJEPA makes V + 1 views: a larger global view x_0, and V more tightly cropped, more strongly augmented local views x_1 … x_V. All keep the same 16 time indices. Each passes through the very same encoder E and projector h; the [CLS] output becomes z_v.

“Shared” is literal parameter sharing. There are not several look-alike networks. Every path refers to one set of weights, so one update changes all paths.

Training pulls each local card toward the global card. It does not repaint the missing scene, predict a later clip, or send a separate predictor toward a frozen target. The global crop is a useful reference because it covers more space and receives milder photometric treatment. It is still a gradient-carrying branch, not a teacher. SIGReg acts on the summaries as well, so the whole batch cannot settle on one answer.

The pairing encourages the encoder to ignore crop and color accidents while retaining content shared by the views. It is not magic: a local crop that catches only blank wall offers weak evidence. Matching time is necessary, not sufficient, for every crop to be informative.

Draw it once

Draw a 16-box timeline and place two windows above it. Give the windows different heights and widths, but let both span boxes 1–16. Slide the small one to boxes 9–24. The moment you do, circle it: that is no longer the LeVJEPA positive pair described here.

Window records

Pocket summary

  1. Global and local views cover the same 16-frame time window.
  2. Every view uses the same encoder, projector, and [CLS] readout path.
  3. The global embedding is a gradient-carrying reference, not a frozen teacher answer.

Three quick questions

  1. May the local view come from another minute of the same video?
  2. What exactly does “shared encoder” mean?
  3. Must the model reconstruct pixels cropped out of the local view?
Check your answers
  1. No. It must share the global view’s 16 frames.
  2. All views use one trainable parameter set.
  3. No. The loss compares projected [CLS] summaries.