JEPA4Japan · tutorials

Chapter 14 — Same-frame teamwork, no peeking into tomorrow

631 words 3 min read #LeVJEPA#JEPA#self-supervised video#SIGReg

Read the block-causal attention pattern, including within-frame attention, past-only time flow, and the [CLS] exception.

Course progress Course outline 34 of 34 lessons available

Part 0 — Get the map

  1. 01 Chapter 0 — Before you begin: what this course promises available now
  2. 02 Chapter 1 — One video, two windows available now
  3. 03 Chapter 2 — A walk along Yann LeCun’s research road available now
  4. 04 Chapter 3 — The JEPA family, without the name soup available now

Part 1 — Why a small objective can learn to see

  1. 05 Chapter 4 — Video can set its own homework available now
  2. 06 Chapter 5 — Keep the meaning; do not repaint every pixel available now
  3. 07 Chapter 6 — Match the cards, but do not leave every card blank available now
  4. 08 Chapter 7 — SIGReg checks a cloud by looking at its shadows available now
  5. 09 Chapter 8 — The whole LeVJEPA objective on one line available now

Part 2 — Send a video through one encoder

  1. 10 Chapter 9 — How global and local views are paired available now
  2. 11 Chapter 10 — Cut a video into space-time tiles available now
  3. 12 Chapter 11 — One encoder, one projector, one summary card available now
  4. 13 Chapter 12 — One complete trip through the model available now
  5. 14 Chapter 13 — Why throwing away 95% can help available now
  6. 15 Chapter 14 — Same-frame teamwork, no peeking into tomorrow Current lesson
  7. 16 Chapter 15 — RoPE, single-frame tubelets, and unexpectedly useful patch features available now

Part 3 — Read the experiments, not just the headline

  1. 17 Chapter 16 — What the four ablation ladders actually test available now
  2. 18 Chapter 17 — Equal epochs are not equal bills available now
  3. 19 Chapter 18 — What ImageNet, K400, and SSv2 are really asking available now
  4. 20 Chapter 19 — Keep the paper’s results in a ledger available now
  5. 21 Chapter 20 — Claims the evidence does not yet earn available now

Part 4 — From the official repository to your own experiment

  1. 22 Chapter 21 — A map of the official repository available now
  2. 23 Chapter 22 — Ten long walks become a training set available now
  3. 24 Chapter 23 — Read the defaults, then start training available now
  4. 25 Chapter 24 — Run a smoke test that cannot flatter you available now
  5. 26 Chapter 25 — Skip training: extract features from the public checkpoint available now
  6. 27 Chapter 26 — Freeze the encoder and test your own videos available now

Part 5 — Put the representation back on the world-model road

  1. 28 Chapter 27 — The important boundary: an encoder is not a planner available now
  2. 29 Chapter 28 — How LeVJEPA might feed a future world model available now
  3. 30 Chapter 29 — Ten projects, from first experiment to paper-sized question available now

Appendices — A backpack for the trail

  1. 31 Appendix A — The smallest useful math kit available now
  2. 32 Appendix B — The complete tensor-shape table available now
  3. 33 Appendix C — Glossary and paper timeline available now
  4. 34 Appendix D — Reproduction and review checklist available now

Classrooms that cannot phone tomorrow

  1. Talk within a frametwo-way spatial attention
  2. Read the pastcausal across time
  3. Block the futurepatches keep the rule
  4. `[CLS]` receivesmail goes in, not back out
Patch tokens discuss their frame and its history. `[CLS]` reads the whole clip but cannot relay tomorrow back to yesterday.

Each video frame is a classroom. Children in one room may talk freely and read notes from earlier rooms, but nobody may call a future classroom. At the end of the hall, [CLS] is a master mailbox that receives notes from the whole building. Classrooms cannot read from it. If they could, the last room could mail tomorrow’s answer in, and an early room could retrieve it on the next layer.

The exact visibility rule

For a patch query in frame t and patch key in frame s, attention is allowed when s ≤ t. Thus all spatial patches at s=t see one another bidirectionally. Earlier frames are visible; later frames are not.

QueryMay readMay not readPurpose
frame t patchretained patches in frames 1…tframes t+1…endframe features contain no future
[CLS]all patches and itself—whole-clip summary
any patchother allowed patches[CLS]prevent a cross-layer future relay

“Causal” therefore describes the frame-level patch representations. Training [CLS] reads the entire clip and is not an online state at time t. After random dropping, the implementation recovers each retained patch’s original frame ID from token_ids; it does not infer time from the shortened sequence order.

In one controlled setting—ViT-B/16, 20% K710, tubelet=1, ρ=0.95, V=4, random dropping—the authors report frozen attentive-probe ImageNet top-1 of 50.7 with fully bidirectional attention and 51.2 with block-causal attention. That means no measurable penalty appeared under this protocol. It is not a theorem about all datasets and tasks.

The topology also permits a streaming implementation to cache past keys and values instead of re-encoding history. That is an architectural opportunity. The paper does not train action-conditioned dynamics, a cost, or a planner, so “streaming-friendly foundation” must not become “already plans.”

Color the mask

Draw a 3×3 frame-level attention matrix, query rows and key columns. Color the diagonal and lower triangle green, the upper triangle red. Add [CLS]: its query row is green; the column that would let patches read it is red. If you color that column green, trace the two-layer leak future → CLS → past.

Mask records

Three rules of the hallway

  1. Patches attend both ways within a frame and only backward or level across time.
  2. [CLS] sees the clip, but patches cannot read it; otherwise future information leaks across layers.
  3. 51.2 versus 50.7 is an author-reported probe comparison, not a universal no-cost theorem or planning evidence.

Leakage check

  1. Can a frame-3 patch read other patches in frame 3?
  2. Why may patches not read [CLS]?
  3. Is training [CLS] an online state containing only the first three frames?
Answers
  1. Yes. Attention is bidirectional within a frame.
  2. [CLS] has seen the full clip and could relay future information backward on later layers.
  3. No. This [CLS] reads the whole clip.