Course progress Course outline 34 of 34 lessons available
Part 0 — Get the map
Part 1 — Why a small objective can learn to see
- 05 Chapter 4 — Video can set its own homework available now
- 06 Chapter 5 — Keep the meaning; do not repaint every pixel available now
- 07 Chapter 6 — Match the cards, but do not leave every card blank available now
- 08 Chapter 7 — SIGReg checks a cloud by looking at its shadows available now
- 09 Chapter 8 — The whole LeVJEPA objective on one line available now
Part 2 — Send a video through one encoder
- 10 Chapter 9 — How global and local views are paired Current lesson
- 11 Chapter 10 — Cut a video into space-time tiles available now
- 12 Chapter 11 — One encoder, one projector, one summary card available now
- 13 Chapter 12 — One complete trip through the model available now
- 14 Chapter 13 — Why throwing away 95% can help available now
- 15 Chapter 14 — Same-frame teamwork, no peeking into tomorrow available now
- 16 Chapter 15 — RoPE, single-frame tubelets, and unexpectedly useful patch features available now
Part 3 — Read the experiments, not just the headline
- 17 Chapter 16 — What the four ablation ladders actually test available now
- 18 Chapter 17 — Equal epochs are not equal bills available now
- 19 Chapter 18 — What ImageNet, K400, and SSv2 are really asking available now
- 20 Chapter 19 — Keep the paper’s results in a ledger available now
- 21 Chapter 20 — Claims the evidence does not yet earn available now
Part 4 — From the official repository to your own experiment
- 22 Chapter 21 — A map of the official repository available now
- 23 Chapter 22 — Ten long walks become a training set available now
- 24 Chapter 23 — Read the defaults, then start training available now
- 25 Chapter 24 — Run a smoke test that cannot flatter you available now
- 26 Chapter 25 — Skip training: extract features from the public checkpoint available now
- 27 Chapter 26 — Freeze the encoder and test your own videos available now
Part 5 — Put the representation back on the world-model road
Appendices — A backpack for the trail
One flipbook, several cutouts
- Fix 16 framesdo not move in time
- Global 224a wide scene window
- Local 96small detail windows
- Keep one groupsend all through shared weights
A teacher makes a 16-page flipbook. She photographs the whole desk for one edition, then prints close-up editions containing only the cup, hand, or table edge. Some close-ups are brighter or grayer. Every edition must still turn through the same 16 pages. One beginning on page 9 answers a different question.
Build the views carefully
For each training clip, LeVJEPA first fixes one 16-frame interval. It then makes 1 + V views. The global x_0 is 224 × 224 and covers more of the scene. Each local x_1 … x_V is 96 × 96, uses a tighter spatial crop, and receives photometric augmentation. The global view is the broad, mildly treated reference—not a label and not a stopped-gradient teacher.
V is an experimental setting, not a constant hidden in the method’s name. Unless stated otherwise, the paper’s controlled comparisons use four local views. In that setting, raising the count to ten improves the reported ImageNet probe; the curve approaches saturation around twelve. The public Walking Tours configuration defaults to ten. Always record the actual V; do not turn a repository default into a universal paper setting.
Why bother with a close-up? To align its [CLS] summary with the wide view, the encoder cannot rely entirely on one fixed coordinate or a complete background. It must find useful evidence in fragments. The wide view supplies a more stable, information-rich anchor than close-ups alone.
“Same time” is stronger than “same length.” Two clips can each contain 16 frames while starting at different stages of an action. LeVJEPA does not ask the first 16 frames to predict the next 16. Global and local views are simultaneous samples of one interval. Block-causal attention operates later, inside the encoder, and changes none of this pairing rule.
After augmentation, each view receives its own patch embedding and random token sample. Their keep-sets need not match. They share source times and network parameters, not necessarily visible patch coordinates.
Make a legal group
Let the source indices be 101–116. Use a 224 × 224 crop of those frames globally and three different 96 × 96 crops of those same frames locally. Change one local interval to 117–132. It remains 16 frames long, but the pairing is now invalid.
View-making records
- LeVJEPA v1: global/local views and local-count ablation
- Pinned default config: 224, 96, 16 frames, and ten local views
- Pinned data loader and multi-crop transform
Keep the window frame
- Global
224 × 224and local96 × 96views share exactly the same 16 frames. - Local-view count is part of the experiment; paper comparisons and repository defaults can differ.
- Views may crop, augment, and drop tokens independently while sharing encoder/projector weights.
Window check
- Can two 16-frame intervals with different start times form the positive pair?
- Is the repository’s local-view count a fixed constant for every paper experiment?
- Is the global view a stopped-gradient teacher?
Answers
- No.
- No; record the
Vused in each experiment. - No. It backpropagates along with the local paths.