JEPA4Japan · tutorials

Chapter 21 — A map of the official repository

759 words 4 min read #LeVJEPA#JEPA#self-supervised video#SIGReg

Locate main.py, module.py, data/loader.py, Hydra configs, two notebooks, the license, and the pinned code snapshot.

Course progress Course outline 34 of 34 lessons available

Part 0 — Get the map

  1. 01 Chapter 0 — Before you begin: what this course promises available now
  2. 02 Chapter 1 — One video, two windows available now
  3. 03 Chapter 2 — A walk along Yann LeCun’s research road available now
  4. 04 Chapter 3 — The JEPA family, without the name soup available now

Part 1 — Why a small objective can learn to see

  1. 05 Chapter 4 — Video can set its own homework available now
  2. 06 Chapter 5 — Keep the meaning; do not repaint every pixel available now
  3. 07 Chapter 6 — Match the cards, but do not leave every card blank available now
  4. 08 Chapter 7 — SIGReg checks a cloud by looking at its shadows available now
  5. 09 Chapter 8 — The whole LeVJEPA objective on one line available now

Part 2 — Send a video through one encoder

  1. 10 Chapter 9 — How global and local views are paired available now
  2. 11 Chapter 10 — Cut a video into space-time tiles available now
  3. 12 Chapter 11 — One encoder, one projector, one summary card available now
  4. 13 Chapter 12 — One complete trip through the model available now
  5. 14 Chapter 13 — Why throwing away 95% can help available now
  6. 15 Chapter 14 — Same-frame teamwork, no peeking into tomorrow available now
  7. 16 Chapter 15 — RoPE, single-frame tubelets, and unexpectedly useful patch features available now

Part 3 — Read the experiments, not just the headline

  1. 17 Chapter 16 — What the four ablation ladders actually test available now
  2. 18 Chapter 17 — Equal epochs are not equal bills available now
  3. 19 Chapter 18 — What ImageNet, K400, and SSv2 are really asking available now
  4. 20 Chapter 19 — Keep the paper’s results in a ledger available now
  5. 21 Chapter 20 — Claims the evidence does not yet earn available now

Part 4 — From the official repository to your own experiment

  1. 22 Chapter 21 — A map of the official repository Current lesson
  2. 23 Chapter 22 — Ten long walks become a training set available now
  3. 24 Chapter 23 — Read the defaults, then start training available now
  4. 25 Chapter 24 — Run a smoke test that cannot flatter you available now
  5. 26 Chapter 25 — Skip training: extract features from the public checkpoint available now
  6. 27 Chapter 26 — Freeze the encoder and test your own videos available now

Part 5 — Put the representation back on the world-model road

  1. 28 Chapter 27 — The important boundary: an encoder is not a planner available now
  2. 29 Chapter 28 — How LeVJEPA might feed a future world model available now
  3. 30 Chapter 29 — Ten projects, from first experiment to paper-sized question available now

Appendices — A backpack for the trail

  1. 31 Appendix A — The smallest useful math kit available now
  2. 32 Appendix B — The complete tensor-shape table available now
  3. 33 Appendix C — Glossary and paper timeline available now
  4. 34 Appendix D — Reproduction and review checklist available now

Label the kitchen before turning on the stove

  1. READMEentrance and boundaries
  2. main.pyconnect the parts
  3. module.pyencoder and SIGReg
  4. data/long video into clips
  5. conf/the experiment recipe
Reading a repository is a kitchen tour: pin the address, then separate recipe, equipment, ingredients, and bill.

A toolbox is easier to trust once every drawer has a name. main.py is the switchboard. module.py contains the learning machinery. data/loader.py prepares the ingredients. conf/ stores the recipe. When something fails, this map tells you where to look before you start turning random screws.

Pin the evidence before installing it

This edition audits repository commit 3ea0dda dated 2026-09-02. A pinned checkout can be revisited; a floating main cannot.

git clone https://github.com/MLO-lab/LeVJEPA.git
cd LeVJEPA
git checkout 3ea0dda16030bc0fd6472bf809fc0a0ad836a812
git rev-parse HEAD
uv sync

The printed hash should match exactly. The project requires Python 3.12+ and locks dependencies in uv.lock. Use uv sync --extra data for dataset construction or uv sync --extra notebook for the notebooks. This is not an installed Python package; run commands from its repository root.

The directory map:

  • main.py joins loader, shared encoder, projector, SIGReg, optimizer, and Lightning trainer.
  • module.py implements the ViT, random token dropping, RoPE, block-causal mask, projector, and SIGReg.
  • data/loader.py samples clips from Lance and makes one global plus several local views.
  • callbacks.py schedules weight decay and saves evaluation-only EMA weights.
  • conf/config.yaml is the Hydra recipe; conf/data/ points to datasets.
  • scripts/ downloads and transcodes; slurm/ launches cluster jobs.
  • training_workshop.ipynb dismantles training with a small model; feature_visualization.ipynb inspects patch features from released weights.

For a first trace, stay on one road. main() calls build_video_loader(). Transforms emit a global [B,T,C,H,W] tensor and local [B,V,T,C,H,W]. multiview_forward() rearranges them to [B,C,T,H,W], folds locals into the batch, runs one shared encoder, and takes token 0 ([CLS]). One projector maps all summaries before view MSE and 0.02 × SIGReg. Breakpoints along this road distinguish data-layout, encoder, and loss failures—and show that the global output is never detached.

The notebooks are lessons, not paper evaluation scripts. The workshop shrinks the model, frame count, and sparsity; the feature notebook only visualizes frozen output. Record every teaching simplification if you begin there. paper.md is a manuscript copy and may retain draft TODOs, so use arXiv v1 and release pages for formal claims.

Licensing is file-specific. The repository root uses MIT, while module.py, adapted from Meta V-JEPA, explicitly carries CC BY-NC 4.0; released weights use CC BY-NC 4.0 too. Dependencies and linked Walking Tours videos keep their own terms. A root MIT file does not grant blanket commercial permission.

The smallest environment check

After checkout, git status --short should print nothing. Then run:

uv run python -c "from module import SIGReg, vit_tiny; print('imports ok')"

imports ok proves only that this environment imports. It says nothing about training accuracy.

Repository records

Keep the map folded

  1. A reproducible report records a commit, not merely “latest main.”
  2. main.py assembles, module.py computes, and conf/ varies the recipe.
  3. MIT does not cover every artifact; core adapted code and weights are noncommercial CC BY-NC 4.0.

Find the drawer

  1. Why record the full commit hash?
  2. Where is EMA bookkeeping, and is it a target encoder?
  3. Does a root MIT license make the released weights commercial-use software?
Answers
  1. It makes code, configuration, and conclusions exactly revisit-able.
  2. In callbacks.py; it averages evaluation weights and makes no training target.
  3. No. module.py and weights carry CC BY-NC 4.0, and data/dependency terms also apply.