JEPA4Japan · tutorials

Appendix C — Glossary and paper timeline

1,894 words 9 min read #LeVJEPA#JEPA#self-supervised video#SIGReg

Give each term a one-line meaning, a concrete analogy, and a common mistake, then anchor the JEPA lineage in primary sources.

Course progress Course outline 34 of 34 lessons available

Part 0 — Get the map

  1. 01 Chapter 0 — Before you begin: what this course promises available now
  2. 02 Chapter 1 — One video, two windows available now
  3. 03 Chapter 2 — A walk along Yann LeCun’s research road available now
  4. 04 Chapter 3 — The JEPA family, without the name soup available now

Part 1 — Why a small objective can learn to see

  1. 05 Chapter 4 — Video can set its own homework available now
  2. 06 Chapter 5 — Keep the meaning; do not repaint every pixel available now
  3. 07 Chapter 6 — Match the cards, but do not leave every card blank available now
  4. 08 Chapter 7 — SIGReg checks a cloud by looking at its shadows available now
  5. 09 Chapter 8 — The whole LeVJEPA objective on one line available now

Part 2 — Send a video through one encoder

  1. 10 Chapter 9 — How global and local views are paired available now
  2. 11 Chapter 10 — Cut a video into space-time tiles available now
  3. 12 Chapter 11 — One encoder, one projector, one summary card available now
  4. 13 Chapter 12 — One complete trip through the model available now
  5. 14 Chapter 13 — Why throwing away 95% can help available now
  6. 15 Chapter 14 — Same-frame teamwork, no peeking into tomorrow available now
  7. 16 Chapter 15 — RoPE, single-frame tubelets, and unexpectedly useful patch features available now

Part 3 — Read the experiments, not just the headline

  1. 17 Chapter 16 — What the four ablation ladders actually test available now
  2. 18 Chapter 17 — Equal epochs are not equal bills available now
  3. 19 Chapter 18 — What ImageNet, K400, and SSv2 are really asking available now
  4. 20 Chapter 19 — Keep the paper’s results in a ledger available now
  5. 21 Chapter 20 — Claims the evidence does not yet earn available now

Part 4 — From the official repository to your own experiment

  1. 22 Chapter 21 — A map of the official repository available now
  2. 23 Chapter 22 — Ten long walks become a training set available now
  3. 24 Chapter 23 — Read the defaults, then start training available now
  4. 25 Chapter 24 — Run a smoke test that cannot flatter you available now
  5. 26 Chapter 25 — Skip training: extract features from the public checkpoint available now
  6. 27 Chapter 26 — Freeze the encoder and test your own videos available now

Part 5 — Put the representation back on the world-model road

  1. 28 Chapter 27 — The important boundary: an encoder is not a planner available now
  2. 29 Chapter 28 — How LeVJEPA might feed a future world model available now
  3. 30 Chapter 29 — Ten projects, from first experiment to paper-sized question available now

Appendices — A backpack for the trail

  1. 31 Appendix A — The smallest useful math kit available now
  2. 32 Appendix B — The complete tensor-shape table available now
  3. 33 Appendix C — Glossary and paper timeline Current lesson
  4. 34 Appendix D — Reproduction and review checklist available now

Names are labels, not bags of spare parts

  1. FamilyJEPA is a blueprint
  2. MethodI, V, Le, and LeV differ
  3. Versionnewer is not stronger evidence
  4. Evidencepaper, code, weights, reproduction
A glossary stops one paper’s machinery from sneaking into another paper under a shared surname.

A hammer, screwdriver, and drill may all live in a carpenter’s toolbox. That does not put a battery inside the hammer. JEPA is similarly a research blueprint and family of methods. Members choose different encoders, target branches, predictors, regularizers, and data.

The most common reading error is not a bad calculation. It is quietly installing I-JEPA’s EMA target encoder, V-JEPA’s predictor, and V-JEPA 2-AC’s action model inside LeVJEPA. This glossary gives every term a small “do not smuggle this in” label.

Start with the names themselves

  • The official spelling is LeVJEPA, not LEVJEPA, Le-VJEPA, or V-JEPA 3.
  • The LeJEPA paper expands its name as Latent-Euclidean JEPA. It contributes the general objective and SIGReg.
  • The LeVJEPA paper does not supply a formal long name to expand. The safe description is “a video encoder trained with the LeJEPA objective.”
  • All these papers are collaborations. Saying LeVJEPA follows Yann LeCun’s research road describes an intellectual lineage; it does not erase the work of the other authors.

The working glossary

TermOne-line meaningSmall pictureCommon mix-up
JEPAAn architectural idea that models compatibility or prediction in representation spaceguess meaning on a map, not every brickIt does not prescribe one loss or collapse remedy.
I-JEPAPredicts representations of masked target regions from image contextinfer the middle of a puzzleIt is a still-image method, not a planner.
V-JEPAExtends latent feature prediction to video space-time blocksinfer a hidden stretch of a filmIt uses a predictor and EMA-target route; it is not LeVJEPA.
V-JEPA 2A large-scale video JEPA representation modelan observer trained on much videoAction-free pretraining does not itself control a robot.
V-JEPA 2-ACTrains an action-conditioned predictor above a frozen encoderadd an “if I move” sandboxIt is a V-JEPA 2 post-training stage, not a LeVJEPA component.
LeJEPAJoins a JEPA pairing objective with SIGReg in a lean general recipematch same-answer cards without allowing blanksIts theory has assumptions and is not video-specific.
LeVJEPAUses the LeJEPA collapse-control objective to train a video encoderdescribe one film through large and small windowsIt is encoder pretraining, not an action world model.
LeWorldModelTrains next-embedding prediction, SIGReg, and action-conditioned dynamics for controla sandbox that can try actionsIt is a neighbouring branch, not another name for LeVJEPA.
self-supervisedBuilds a training signal from the data itselfthe video writes its own matching questionHumans still choose data, transformations, and objective.
representation / embeddingA numeric vector produced for an inputa numeric address cardIndividual coordinates usually have no hand-written meaning.
encoderTurns pixels into tokens and representationsa reader writing notesLeVJEPA’s encoder has no built-in class answer.
global viewA broad spatial crop of one clipa wide windowIt is still an augmented crop, not the entire original recording.
local viewA smaller spatial crop from the same temporal windowa cardboard viewing tubeMoving it to another time changes the pairing contract.
augmentationChanges a view while intending to preserve its relevant identitycrop a corner or alter brightnessAn aggressive transform can erase information the task needs.
invarianceViews of the same clip should receive nearby representationsrecognize the same act from two windowsAlone, it permits one constant answer.
complete collapseNearly every input maps to the same vectorthe whole class hands in one blank cardA low invariance loss may hide it.
dimensional collapseEmbeddings occupy only a low-dimensional subspace256 streets squeezed into one alleySamples can differ and still collapse this way.
SIGRegConstrains embedding distributions through Gaussianity of random projectionsinspect a city from many shadowsIt does not test physical meaning or controllability.
isotropic GaussianA zero-centered Gaussian reference with equal scale in every directiona cloud with no favored compass pointIt is neither a uniform sphere nor the topology of the real world.
random projectionMaps a high-dimensional point to one scalar with a random unit vectorshine a flashlight from one directionFinitely many directions are an approximation.
characteristic functionThe complex-exponential expectation that uniquely describes a distributiona frequency fingerprintThe code compares only finitely many frequency nodes.
Epps–Pulley statisticA smooth discrepancy between empirical and Gaussian characteristic functionscompare two fingerprintsA training loss is not a formal pass/fail normality certificate.
patchA 16×16 pixel region in one framea mosaic tileIt becomes a transformer token only after embedding.
tokenThe vector made from an embedded patcha block stamped with coordinatesTraining randomly drops 95% of patch tokens.
tubeletA spatial patch extended across several framesstack matching tiles into a tubeLeVJEPA defaults to tau=1, so it does no temporal stacking here.
[CLS]A learned token that summarizes the clipa summary cardPatch tokens cannot read it, preventing future information from flowing back.
projectorMaps [CLS] into the 256-D space constrained by SIGRega temporary training translation deskIt is discarded after pretraining and is not a classifier.
LayerNormNormalizes features within each tokenpress each card onto a constrained surfaceIts output geometry is one reason SIGReg follows a projector.
RoPEEncodes relative time and space positions through rotations in attentionstamp time, row, and column on a blockIt does not teach causal intervention by itself.
block-causal attentionBidirectional within a frame, current-and-past only across framesclassmates share a page but cannot open tomorrow’s page[CLS] may see the clip; patches cannot read [CLS].
target encoderA separate branch that supplies target features during traininga second teacher writing answersLeVJEPA’s objective has none.
stop-gradientBlocks backpropagation through a brancha one-way valveBoth LeVJEPA view branches receive gradients.
predictorEstimates a target representation from contextthe part that guessesLeVJEPA aligns a shared encoder’s view summaries directly and has no predictor.
Polyak/EMA weightsA moving average copy of training weights used for evaluationa smoothed copy of the scorebookThe official copy does not make targets or enter the loss.
frozen probeTrains a small readout while leaving the encoder fixedtest the notes without rewriting themDecodability does not prove the model actively uses the information.
attentive probePools tokens with a learned query before classifyinga marker who learns where to lookIt is stronger than plain linear mean pooling; scores are not interchangeable.
epochOne pass over a defined datasetfinish one lap of a workbookThe work per sample differs between methods.
FLOPsAn estimate of floating-point operationsaccounting units for computationIt is not wall time or energy consumption.
world modelPredicts how a world state may evolvean internal sandboxA temporally causal encoder is only a possible foundation.
MPCPlans a horizon, acts briefly, observes, and replanstake one step and reopen the mapLeVJEPA itself contains no MPC.

One road with two important forks

Public datePrimary workWhat it advancesWhat it does not establish
2006Energy-based learning tutorialLow energy for compatible configurations, high energy for incompatible onesA JEPA or video recipe
2022-06-27A Path Towards Autonomous Machine Intelligence v0.9.2A blueprint joining perception, configurable world model, cost, actor, memory, and hierarchical JEPAA completed autonomous system
2023I-JEPARepresentation prediction for large target blocks in still imagesActions or planning
2024V-JEPALatent prediction for video space-time blocksA control model
2025-06V-JEPA 2 / 2-ACLarge-scale observation pretraining, then a separate action-conditioned stage with MPCGeneral long-horizon autonomy
2025-11LeJEPA v3An explicit, scalable SIGReg route around collapseAn assumption-free global optimum for every task
2026-03V-JEPA 2.1Stronger dense token representations on the original V-JEPA branchA move to the LeJEPA/SIGReg branch
2026-03LeWorldModel v3Action-conditioned latent dynamics and planning on the LeJEPA branchBroad real-world validation from small control studies
2026-08-27LeVJEPA v1Lean video-encoder pretraining with sparse observation and block-causal attentionAn action predictor or planner

This is a tree, not a sequence of automatic upgrades. V-JEPA 2.1 continues the asymmetric feature-prediction branch. LeWorldModel and LeVJEPA both draw on LeJEPA/SIGReg but pursue action dynamics and general video representation, respectively.

Put an evidence label on every claim

LabelWhat it can supportWhat it cannot support automatically
method definitionA paper explicitly specifies a graph or objectiveThe implementation is bug-free or reproducible.
author-reportedAuthors report a value under a named protocolAn independent team will obtain it.
author codeA concrete interface or default can be inspectedEvery setting matches every paper experiment.
public weightsFeature extraction and downstream evaluation can be runThe training process has been reproduced independently.
independent reproductionA third party repeats a stated scopeThe claim holds for all models, data, and tasks.
assumption-bound theoryA conclusion holds under listed assumptionsAn unconditional guarantee in the physical world.

At this course’s 2026-09-05 (Asia/Tokyo) evidence cutoff, LeVJEPA remains a newly released v1 preprint. Its results are author-reported. This review found no independent reproduction that would justify a stronger label.

Sources for the family tree

The lineage above was cross-checked against LeCun’s 2022 AMI paper, the I-JEPA CVPR record, V-JEPA at TMLR/OpenReview, the V-JEPA 2 publication page, LeJEPA v3, and LeVJEPA v1. Public dates establish order, not strength of evidence.

Keep the labels on the drawers

  1. JEPA is a family blueprint; each member’s computation graph must be checked separately.
  2. LeVJEPA has no official long name to invent. It is a video encoder trained with the LeJEPA objective.
  3. A later date, public code, and independent reproduction are three different facts.

Name the right machine

  1. Why can V-JEPA’s predictor not be drawn inside LeVJEPA?
  2. Why may we say “no EMA target encoder” even though the official implementation keeps EMA weights?
  3. What is the central task boundary between LeVJEPA and LeWorldModel?
Check the answers
  1. They use different training graphs; LeVJEPA directly aligns global and local [CLS] outputs from a shared encoder and uses SIGReg against collapse.
  2. The EMA copy is a smoothed evaluation checkpoint. It neither produces targets nor participates in the loss.
  3. LeVJEPA learns general video representations; LeWorldModel explicitly learns action-conditioned next-latent states and uses them for planning.