JEPA4Japan · tutorials

Chapter 15 — RoPE, single-frame tubelets, and unexpectedly useful patch features

615 words 3 min read #LeVJEPA#JEPA#self-supervised video#SIGReg

Connect 3D rotary position, tubelet size 1, and the dense features that emerge even though only [CLS] is supervised.

Course progress Course outline 34 of 34 lessons available

Part 0 — Get the map

  1. 01 Chapter 0 — Before you begin: what this course promises available now
  2. 02 Chapter 1 — One video, two windows available now
  3. 03 Chapter 2 — A walk along Yann LeCun’s research road available now
  4. 04 Chapter 3 — The JEPA family, without the name soup available now

Part 1 — Why a small objective can learn to see

  1. 05 Chapter 4 — Video can set its own homework available now
  2. 06 Chapter 5 — Keep the meaning; do not repaint every pixel available now
  3. 07 Chapter 6 — Match the cards, but do not leave every card blank available now
  4. 08 Chapter 7 — SIGReg checks a cloud by looking at its shadows available now
  5. 09 Chapter 8 — The whole LeVJEPA objective on one line available now

Part 2 — Send a video through one encoder

  1. 10 Chapter 9 — How global and local views are paired available now
  2. 11 Chapter 10 — Cut a video into space-time tiles available now
  3. 12 Chapter 11 — One encoder, one projector, one summary card available now
  4. 13 Chapter 12 — One complete trip through the model available now
  5. 14 Chapter 13 — Why throwing away 95% can help available now
  6. 15 Chapter 14 — Same-frame teamwork, no peeking into tomorrow available now
  7. 16 Chapter 15 — RoPE, single-frame tubelets, and unexpectedly useful patch features Current lesson

Part 3 — Read the experiments, not just the headline

  1. 17 Chapter 16 — What the four ablation ladders actually test available now
  2. 18 Chapter 17 — Equal epochs are not equal bills available now
  3. 19 Chapter 18 — What ImageNet, K400, and SSv2 are really asking available now
  4. 20 Chapter 19 — Keep the paper’s results in a ledger available now
  5. 21 Chapter 20 — Claims the evidence does not yet earn available now

Part 4 — From the official repository to your own experiment

  1. 22 Chapter 21 — A map of the official repository available now
  2. 23 Chapter 22 — Ten long walks become a training set available now
  3. 24 Chapter 23 — Read the defaults, then start training available now
  4. 25 Chapter 24 — Run a smoke test that cannot flatter you available now
  5. 26 Chapter 25 — Skip training: extract features from the public checkpoint available now
  6. 27 Chapter 26 — Freeze the encoder and test your own videos available now

Part 5 — Put the representation back on the world-model road

  1. 28 Chapter 27 — The important boundary: an encoder is not a planner available now
  2. 29 Chapter 28 — How LeVJEPA might feed a future world model available now
  3. 30 Chapter 29 — Ten projects, from first experiment to paper-sized question available now

Appendices — A backpack for the trail

  1. 31 Appendix A — The smallest useful math kit available now
  2. 32 Appendix B — The complete tensor-shape table available now
  3. 33 Appendix C — Glossary and paper timeline available now
  4. 34 Appendix D — Reproduction and review checklist available now

Three address stamps and an unexpected map

  1. Three compassestime, height, width
  2. One frame eachtubelet = 1
  3. Grade the summarysupervise `[CLS]` only
  4. Patches organizea qualitative finding
3D RoPE marks position, the input keeps neighboring frames separate, and unsupervised patch features show visible structure.

At an animation desk, every puzzle tile receives three stamps: frame number, row, column. The machine can now compare relative places. Neighboring pages are not glued together; each frame makes its own patches. Oddly, although the teacher grades only the whole-book summary, little tiles for fox, sofa, and background begin to gather into visible groups.

That last observation comes from paper visualizations. It is not a passed segmentation exam.

Three mechanisms, three separate claims

Factorized 3D RoPE. Each attention head’s channels are split into groups rotated by time, vertical, and horizontal coordinates. The appendix says no absolute position embedding is used. Relative position lets the encoder accept both 224 and 96 grids without interpolating one fixed learned table.

tubelet=1. A spatial patch is 16×16 and spans one frame in time. A 16-frame global view therefore has 16 × 14 × 14 = 3136 patches; a local view has 16 × 6 × 6 = 576. With tubelet=2, pairs of frames fuse into eight temporal slots. Under a matched training-token budget and eight evaluation slots, the authors report:

Patch embeddingDroppingIN1KSSv2
tubelet=290%47.428.8
tubelet=195%50.730.4

These are author-reported frozen attentive-probe top-1 values. They say tubelet=1 works better in that protocol, not that it wins everywhere.

Dense features. Only [CLS] receives a direct training loss; patch tokens have no auxiliary patch objective. Even so, the paper and project page show patch PCA maps and cosine similarity from a query patch. Regions on the same object group together and sometimes stay corresponding over time. This is useful qualitative evidence. The authors explicitly leave segmentation, tracking, and other true dense prediction tasks unevaluated.

Label one token

For 16 frames, count 16 time slots at tubelet=1 and eight at tubelet=2. Give a token the coordinate (t=5, y=3, x=9); the three RoPE channel groups use 5, 3, and 9 respectively. Then edit “PCA proves segmentation” into the defensible version: “The released examples show semantic organization; segmentation and tracking sufficiency remain unknown.”

Position-and-feature records

The careful reading

  1. 3D RoPE encodes relative time, vertical, and horizontal position without an absolute table.
  2. Default tubelet=1 means one token spans one frame; a global clip has 3136 patches.
  3. Patch organization is a qualitative emergence. No segmentation or tracking result upgrades it into a dense-capability claim.

Coordinate check

  1. Does tubelet=1 merge adjacent frames at the input?
  2. How does 3D RoPE distinguish “when” from “where”?
  3. Does a crisp patch PCA map establish segmentation accuracy?
Answers
  1. No. Its temporal span is one frame.
  2. Separate channel groups rotate using temporal, vertical, and horizontal coordinates.
  3. No. It is a qualitative visualization; segmentation and tracking were not evaluated.