Course progress Course outline 34 of 34 lessons available
Part 0 — Get the map
Part 1 — Why a small objective can learn to see
- 05 Chapter 4 — Video can set its own homework available now
- 06 Chapter 5 — Keep the meaning; do not repaint every pixel available now
- 07 Chapter 6 — Match the cards, but do not leave every card blank available now
- 08 Chapter 7 — SIGReg checks a cloud by looking at its shadows available now
- 09 Chapter 8 — The whole LeVJEPA objective on one line available now
Part 2 — Send a video through one encoder
- 10 Chapter 9 — How global and local views are paired available now
- 11 Chapter 10 — Cut a video into space-time tiles available now
- 12 Chapter 11 — One encoder, one projector, one summary card available now
- 13 Chapter 12 — One complete trip through the model available now
- 14 Chapter 13 — Why throwing away 95% can help available now
- 15 Chapter 14 — Same-frame teamwork, no peeking into tomorrow available now
- 16 Chapter 15 — RoPE, single-frame tubelets, and unexpectedly useful patch features available now
Part 3 — Read the experiments, not just the headline
- 17 Chapter 16 — What the four ablation ladders actually test available now
- 18 Chapter 17 — Equal epochs are not equal bills available now
- 19 Chapter 18 — What ImageNet, K400, and SSv2 are really asking available now
- 20 Chapter 19 — Keep the paper’s results in a ledger available now
- 21 Chapter 20 — Claims the evidence does not yet earn available now
Part 4 — From the official repository to your own experiment
- 22 Chapter 21 — A map of the official repository available now
- 23 Chapter 22 — Ten long walks become a training set available now
- 24 Chapter 23 — Read the defaults, then start training available now
- 25 Chapter 24 — Run a smoke test that cannot flatter you available now
- 26 Chapter 25 — Skip training: extract features from the public checkpoint available now
- 27 Chapter 26 — Freeze the encoder and test your own videos available now
Part 5 — Put the representation back on the world-model road
Appendices — A backpack for the trail
First, the whole picture
- Watch one clipwide view + close-ups
- Write summarieswith one shared encoder
- Match the clipkeep its meaning
- Keep varietydo not answer everything alike
Mina owns a folder of short videos with no titles. Her teacher opens each video twice: once through a wide window, once through a narrow tube. Mina writes a summary card for each view. Cards from the same clip should agree. Yet if she writes “there is stuff” on every card, she has technically agreed and learned nothing.
LeVJEPA is built around those two demands: recognize the same clip across views, without turning every clip into the same answer. The first job belongs to an invariance loss. The second belongs to SIGReg. We will unpack both names slowly.
What is actually in the machine
This course follows LeVJEPA arXiv v1, posted on 2026-08-27 and still a preprint at our evidence cutoff. During training, one 16-frame clip yields one global view and several local views. They cover the same moments in time, though their spatial crops and appearance changes differ. Every view passes through the same trainable encoder and the same small projector. The loss reads each view’s [CLS] summary token.
Three guardrails matter from the first page:
- The objective has no target encoder, predictor, or stop-gradient. Both the global and local paths receive gradients.
- The training code maintains a Polyak/EMA copy of the encoder and saves it for evaluation. That copy does not make targets and is not a teacher network.
- LeVJEPA produces a video encoder. It does not contain action-conditioned dynamics, a cost, or a search procedure. Calling it a possible foundation for a world model is a design direction, not a finished planning result.
You may hear that the method has “only one objective weight.” This means the total loss has one fixed trade-off, λ = 0.02. It does not mean that the data pipeline, optimizer, and architecture have no settings.
We will work in layers: picture first, then equation, tensors, code, and experiment. Reported benchmark values stay labeled as author-reported. A public checkpoint is useful evidence about availability, not proof that we independently reproduced training.
A two-minute check
Sort four sticky notes: shared encoder, SIGReg, EMA evaluation copy, action planner. The first two belong inside the objective. EMA sits beside training as an evaluation copy. The planner does not belong to LeVJEPA at all.
The evidence shelf
- LeVJEPA arXiv v1: version, date, and abstract
- Pinned repository: recipe, EMA note, and public weights
- Official LeVJEPA project page
Three lines to keep
- LeVJEPA v1 is a video-representation pretraining method posted on 2026-08-27.
- One encoder, one projector, and
[CLS]take part in the loss; no teacher, predictor, or stopped-gradient branch does. - An EMA evaluation copy is not an EMA teacher, and an encoder is not a planner.
Check yourself
- What are the two opposing forces in the course’s opening story?
- Does the EMA copy manufacture a training target?
- Can a bare LeVJEPA encoder plan robot actions?
Answers
- Invariance pulls summaries of the same clip together; SIGReg prevents representational collapse.
- No. It is an evaluation checkpoint only.
- No. Action-conditioned prediction, an objective or cost, search, and closed-loop control are still missing.