JEPA4Japan · tutorials

Chapter 29 — Ten projects, from first experiment to paper-sized question

1,318 words 6 min read #LeVJEPA#JEPA#self-supervised video#SIGReg

Turn motion-aware sampling, streaming caches, SIGReg, dense tasks, scale, world-model interfaces, and reproducibility into executable studies.

Course progress Course outline 34 of 34 lessons available

Part 0 — Get the map

  1. 01 Chapter 0 — Before you begin: what this course promises available now
  2. 02 Chapter 1 — One video, two windows available now
  3. 03 Chapter 2 — A walk along Yann LeCun’s research road available now
  4. 04 Chapter 3 — The JEPA family, without the name soup available now

Part 1 — Why a small objective can learn to see

  1. 05 Chapter 4 — Video can set its own homework available now
  2. 06 Chapter 5 — Keep the meaning; do not repaint every pixel available now
  3. 07 Chapter 6 — Match the cards, but do not leave every card blank available now
  4. 08 Chapter 7 — SIGReg checks a cloud by looking at its shadows available now
  5. 09 Chapter 8 — The whole LeVJEPA objective on one line available now

Part 2 — Send a video through one encoder

  1. 10 Chapter 9 — How global and local views are paired available now
  2. 11 Chapter 10 — Cut a video into space-time tiles available now
  3. 12 Chapter 11 — One encoder, one projector, one summary card available now
  4. 13 Chapter 12 — One complete trip through the model available now
  5. 14 Chapter 13 — Why throwing away 95% can help available now
  6. 15 Chapter 14 — Same-frame teamwork, no peeking into tomorrow available now
  7. 16 Chapter 15 — RoPE, single-frame tubelets, and unexpectedly useful patch features available now

Part 3 — Read the experiments, not just the headline

  1. 17 Chapter 16 — What the four ablation ladders actually test available now
  2. 18 Chapter 17 — Equal epochs are not equal bills available now
  3. 19 Chapter 18 — What ImageNet, K400, and SSv2 are really asking available now
  4. 20 Chapter 19 — Keep the paper’s results in a ledger available now
  5. 21 Chapter 20 — Claims the evidence does not yet earn available now

Part 4 — From the official repository to your own experiment

  1. 22 Chapter 21 — A map of the official repository available now
  2. 23 Chapter 22 — Ten long walks become a training set available now
  3. 24 Chapter 23 — Read the defaults, then start training available now
  4. 25 Chapter 24 — Run a smoke test that cannot flatter you available now
  5. 26 Chapter 25 — Skip training: extract features from the public checkpoint available now
  6. 27 Chapter 26 — Freeze the encoder and test your own videos available now

Part 5 — Put the representation back on the world-model road

  1. 28 Chapter 27 — The important boundary: an encoder is not a planner available now
  2. 29 Chapter 28 — How LeVJEPA might feed a future world model available now
  3. 30 Chapter 29 — Ten projects, from first experiment to paper-sized question Current lesson

Appendices — A backpack for the trail

  1. 31 Appendix A — The smallest useful math kit available now
  2. 32 Appendix B — The complete tensor-shape table available now
  3. 33 Appendix C — Glossary and paper timeline available now
  4. 34 Appendix D — Reproduction and review checklist available now

Give an idea a finish line

  1. One questionwhat could change?
  2. Baselinedraw the start line
  3. Measuredecide before running
  4. Lose cleanlywhat would refute it?
  5. Budgettime, data, hardware
A useful project lets the idea lose—and tells you what the loss means.

“Try ten clever changes and keep the fastest one” is not yet a research plan. Before touching the code, write which old method must be beaten, on which data, by which measure, within which bill. Also write the result that would make you abandon the idea. That sentence turns wishful tinkering into an experiment.

The ten cards below grow from a small software test to a controlled planning study. Every card is a proposed project, not a result reported in LeVJEPA v1. Resource labels are rough: S = CPU or a few hours on one consumer GPU; M = roughly one GPU for 1–3 days; L = several days on 4–8 GPUs; XL = paper-scale controlled compute or robotics facilities.

Ten project cards

1. A causal-mask test suite (S)

  • Question: Does every tested frame count, resolution, and random keep-set prevent future leakage?
  • Smallest baseline: Pin build_block_causal_mask; treat full-prefix recomputation as the oracle.
  • Measure: Maximum absolute change in earlier patch tokens and pass rate over 100 randomized cases.
  • Refute it when: Changing only the future moves any earlier patch by more than a registered 1e-5, or a patch can read [CLS].

2. A real streaming KV cache (M)

  • Question: Can cached past states remain numerically equivalent while cutting latency?
  • Smallest baseline: Re-encode the complete 16-frame prefix whenever a frame arrives.
  • Measure: Token cosine similarity, maximum error, first-frame and incremental latency, throughput, and peak memory.
  • Refute it when: Outputs exceed the tolerance, or long-sequence throughput and memory both fail to improve on the same hardware. The public snapshot has no finished cache, so the implementation itself is part of the contribution.

3. Spend less on the SIGReg approximation (M)

  • Question: Can the default 1,024 projections and 17 integration nodes be reduced safely?
  • Smallest baseline: Freeze num_proj=1024, knots=17, and lambda=0.02; alter one item at a time.
  • Measure: Step time, embedding mean and covariance spectrum, held-out random-projection checks, and frozen probes.
  • Refute it when: Compute falls by less than 10%, or probe and spectrum changes exceed the preregistered tolerance. A small finite-projection loss is not proof that the full distribution is Gaussian.

4. Token dropping that notices motion (L)

  • Question: Do paired cross-frame locations or flow-guided samples preserve motion clues better than uniform dropping?
  • Smallest baseline: The 95% uniform random drop, plus no dropping as an upper-compute reference; match total FLOPs.
  • Measure: SSv2 and K400 frozen probes, ImageNet, total FLOPs, and clips per second.
  • Refute it when: SSv2 fails to improve across seeds, or the gain comes from seeing more tokens or compute while ImageNet drops sharply.

5. Move the local window in time (M)

  • Question: Does shifting local views by 1, 2, or 4 frames strengthen motion features, or merely break their shared meaning?
  • Smallest baseline: The official pairing in which global and local views cover exactly the same 16-frame window.
  • Measure: Invariance and SIGReg terms, temporal-order recognition, SSv2 probe, and a static ImageNet probe.
  • Refute it when: Temporal tasks do not improve, collapse appears, or appearance performance crosses the preregistered loss limit. A shifted-window variant is no longer the official LeVJEPA objective.

6. Test the unsupervised patch tokens (M)

  • Question: Are patch features useful for segmentation or tracking even though the loss directly supervises only [CLS]?
  • Smallest baseline: Freeze the released encoder; compare it with the same architecture at random initialization and simple color or optical-flow features.
  • Measure: Linear-segmentation mIoU, DAVIS J&F or point-tracking error, plus query-to-patch visualizations.
  • Refute it when: Quantitative results do not beat simple baselines, or the effect appears only in hand-picked pretty examples. Dense evaluation is explicitly unfinished work in the paper.

7. Diversity versus duration (L)

  • Question: At fixed clip count and FLOPs, does a varied mixture of shorter sources beat a few hours-long walks?
  • Smallest baseline: Walking Tours only; construct an equally sized 1.8M-clip mixture—or an affordable smaller matched pair.
  • Measure: IN1K, SSv2, and K400 frozen probes; cross-domain retrieval; duplicate-frame rate; and throughput.
  • Refute it when: The mixture advantage vanishes after strict deduplication or matched-FLOP accounting. Record licenses and sampling weights; do not call the public Walking Tours default a VideoMix recipe.

8. Continue pretraining in a small domain (M)

  • Question: On medical, industrial, or sports-camera video, is continued LeVJEPA training better than frozen transfer?
  • Smallest baseline: Frozen probes on the released ViT-L, plus random initialization as a cost reference.
  • Measure: Domain probes, retention of the original general probes, embedding spectrum, and GPU-hours.
  • Refute it when: The domain gain stays inside seed variation or general capability is catastrophically forgotten. Sensitive datasets need governance approval before the experiment begins.

9. Does action enter the latent dynamics? (L)

  • Question: Will a small predictor on frozen LeVJEPA features use actions rather than copy video momentum?
  • Smallest baseline: Copy-last, no-action, and shuffled-action models trained on the same offline trajectories.
  • Measure: One- and multi-step latent error, counterfactual action separability, and calibration error.
  • Refute it when: The true-action model cannot beat the no-action or shuffled-action controls, or rolls immediately diverge with horizon. This project adds a predictor; it is not a native LeVJEPA result.

10. A preregistered latent-MPC study (XL)

  • Question: Under a fixed simulated task, do LeVJEPA features improve data efficiency or planning success?
  • Smallest baseline: Hold predictor, data, CEM sample count, and horizon fixed; compare a random encoder, copy dynamics, and a public visual encoder. Begin in simulation, not on a physical robot.
  • Measure: Success with confidence intervals over at least 50 episodes, collisions or constraint violations, seconds per action, GPU-hours, and five seeds.
  • Refute it when: The success interval does not beat the strongest baseline, search is insensitive to replacing the model, or gains require substantially more planning compute. Move to hardware only after a safety gate.

The upstream paper.md Push-T table still says TODO. Projects 9 and 10 must create their own protocol and measurements; an empty upstream cell cannot become an imagined baseline. Every project should archive the pinned commit, resolved Hydra config, data manifest, seeds, hardware, raw logs, and failed runs. Reporting only the best seed removes the very evidence that makes a claim falsifiable.

Records behind the cards

Keep beside the lab notebook

  1. Every project needs a smallest baseline, measure, refutation condition, and resource bill.
  2. Matched epochs, samples, and FLOPs answer different questions; choose the ledger before running.
  3. The world-model projects are proposals. LeVJEPA v1 reports no planning success rate.

Approve the experiment

  1. Why write the refutation condition before seeing results?
  2. May a motion-aware sampler that keeps more tokens be compared directly with 95% random dropping on accuracy alone?
  3. Can the upstream Push-T TODO cell supply a baseline number?
Check the answers
  1. It prevents the finish line moving after the result and leaves negative findings informative.
  2. No. Match total FLOPs or show an explicit accuracy–compute curve.
  3. No. It contains neither a measured result nor the protocol needed to interpret one.