JEPA4Japan · tutorials

Appendix F — Glossary

1,233 words 6 min read #LeWorldModel#World Models#JEPA

Define core terms with a formal definition, one-sentence explanation, everyday analogy, and common misconception.

Course progress Course outline 48 of 48 lessons available

Part 0 — Reading Guide: What Exactly Are We Going to Learn?

  1. 01 Chapter 0 — Before You Begin available now

Part 1 — World Models: An Internal Sandbox for the Agent

  1. 02 Chapter 1 — Why an Agent Needs to “Imagine the Future” available now
  2. 03 Chapter 2 — Why Not Predict the Next Image Directly? available now
  3. 04 Chapter 3 — The JEPA Idea: Predict Meaning, Not a Replica of the Image available now
  4. 05 Chapter 4 — Understand LeWM in One Diagram available now

Part 2 — Turning Images into State: The LeWM Architecture

  1. 06 Chapter 5 — Trajectory Data: To the Model, the World Is Not an Image Collection available now
  2. 07 Chapter 6 — The Visual Encoder: Issuing a “State Passport” for Every Frame available now
  3. 08 Chapter 7 — The Dynamics Predictor: Moving Time Forward in the Mind available now
  4. 09 Chapter 8 — A Complete Forward Pass: Follow One Batch from Start to Finish available now

Part 3 — Preventing the Model from Cheating: Prediction Loss and SIGReg

  1. 10 Chapter 9 — The Most Dangerous Shortcut: Representation Collapse available now
  2. 11 Chapter 10 — Prediction Loss: How the Model Learns the Next Step available now
  3. 12 Chapter 11 — The Intuition Behind SIGReg: Letting Representation Space “Breathe” available now
  4. 13 Chapter 12 — Keep the Mathematics Minimal but Sufficient available now
  5. 14 Chapter 13 — The Original LeWM’s End-to-End Training Mechanism available now
  6. 15 Chapter 14 — Train a Model That Does Not Collapse Immediately available now

Part 4 — Putting the Model into Action: Planning in Latent Space

  1. 16 Chapter 15 — Goal-Conditioned Planning: From “Where Am I?” to “Where Do I Want to Go?” available now
  2. 17 Chapter 16 — Latent Euclidean Distance: Convenient, but Not Necessarily Reliable available now
  3. 18 Chapter 17 — CEM: Searching for Actions Through an Elimination Tournament available now
  4. 19 Chapter 18 — MPC: Do Not Trust the Model for Too Long at Once available now
  5. 20 Chapter 19 — Long-Horizon Rollouts: How Small Errors Snowball into Major Failures available now
  6. 21 Chapter 20 — Implement a Minimal LeWM Planner from Scratch available now

Part 5 — Engineering Reproduction: From Paper to Running System

  1. 22 Chapter 21 — The Official Repository and Experimental Environment available now
  2. 23 Chapter 22 — First Experiment: A TwoRoom Smoke Test available now
  3. 24 Chapter 23 — Second Experiment: Reproducing PushT available now
  4. 25 Chapter 24 — How to Evaluate a World Model Fairly available now
  5. 26 Chapter 25 — Failure-Diagnosis Manual available now

Part 6 — What Has LeWM Actually Learned?

  1. 27 Chapter 26 — Linear Probes: Which Physical Variables Are Encoded in the Latent State? available now
  2. 28 Chapter 27 — Give Latent Space a “Health Check” available now
  3. 29 Chapter 28 — Violation of Expectation: Is the Model Surprised by “Impossible Events”? available now
  4. 30 Chapter 29 — How to Discuss “Understanding the World” Rigorously available now

Part 7 — Why “Accurate Prediction” Can Still Produce “Poor Planning”

  1. 31 Chapter 30 — The Gap Between the Training Objective and the Planning Objective available now
  2. 32 Chapter 31 — Global Non-Collapse Does Not Guarantee Preservation of Task-Relevant Dynamics available now
  3. 33 Chapter 32 — When Is an Isotropic Gaussian Prior Too Strong? available now
  4. 34 Chapter 33 — Long-Horizon Planning: Predict Farther or Plan More Intelligently? available now
  5. 35 Chapter 34 — From Positional Distance to Task Progress available now
  6. 36 Chapter 35 — Multi-Task Learning, Real Robots, and Visual Distractions available now
  7. 37 Chapter 36 — Theoretical Boundaries: When Can the True State Be Identified? available now

Part 8 — From Reproducer to Researcher

  1. 38 Chapter 37 — Design a Credible LeWM Improvement Experiment available now
  2. 39 Chapter 38 — Twelve Executable Research Projects available now
  3. 40 Chapter 39 — Open Questions in LeWM Research available now

Appendices

  1. 41 Appendix A — The Minimum Necessary Mathematical Toolkit available now
  2. 42 Appendix B — PyTorch Implementation Quick Reference available now
  3. 43 Appendix C — Complete Tensor-Shape Table available now
  4. 44 Appendix D — Experiment Configuration Cards available now
  5. 45 Appendix E — Paper Timeline and Evidence Levels available now
  6. 46 Appendix F — Glossary Current lesson
  7. 47 Appendix G — Reproduction Checklist available now
  8. 48 Appendix H — Expert-Review Checklist available now

The big picture

  1. FormalWhat the term names.
  2. PlainA short story to carry.
  3. BoundaryWhat it does not prove.
An analogy opens the door; the definition and misconception guardrail stop it from changing the machine.

A tiny story: the translator’s booth

One researcher says target and means the encoded next observation. Another hears goal and points to the planner’s image. A third hears teacher forcing and invents an EMA teacher that original LeWM does not have.

These words live in different lanes. Every row therefore keeps four windows: formal definition, plain explanation, analogy, and misconception.

A translator booth keeps formal definition, plain explanation, analogy, and misconception in four separate windows.

An analogy helps memory; it does not supply assumptions or evidence.

Model and representation terms

TermFormal definitionPlain explanationAnalogyMisconception
World modelLearned state/dynamics model predicting temporal or action consequences for reasoning/control.It rehearses “what happens next?”Rehearsal room.It must render video or contain complete readable physics.
Latent stateLearned vector compressing an observation/history for prediction, analysis, or planning.A compact internal state card.Passport, not photograph.Every coordinate is physical; one frame is automatically Markov-complete.
Target representationPredictive-loss reference; in v3, the same trainable encoder encodes the shifted real successor without detach.Encoded next observation to match.Training destination ticket.Goal image, EMA teacher, target network, or stop-gradient.
Action conditioningAction/action block modulates a transition predicted from latent context.Ask “what if I do this?”Left/right simulator buttons.Different symbols always produce different, correct causal branches.
Representation collapseDegeneracy that removes task-useful variation or effective structure.Different situations look identical inside.One passport for everyone.Any low dimension is collapse; collapse equals overfitting.
SIGRegFinite random 1-D projections compared with a standard Gaussian as anti-collapse pressure.Check many shadows of the latent crowd.Inspect a city from many angles.Finite shadows certify Gaussianity, topology, physics, or utility.

Prediction and control-loop terms

TermFormal definitionPlain explanationAnalogyMisconception
Prediction lossDiscrepancy between predicted future representation and encoded observed successor.Did predicted next meet real next?Forecast versus arrival.Low one-step loss guarantees rollout or planning.
Teacher forcingSequence training conditioned on recorded context, not only earlier predictions.Training walks from the real trail.Supplied preceding piano bar.Requires teacher network, EMA, or stop-gradient.
RolloutRepeated dynamics under actions, often feeding predictions into later steps.Build a future path step by step.Copies from earlier copies.Necessarily video, real data, or a feasible path.
Open-loop controlExecute a committed sequence without observation-based revision during that interval.Follow the plan before looking again.Printed driving route.It has no model or never works.
Closed-loop controlLater actions use observations after earlier actions affect the world.Look again and update.Navigation reroute.Feedback guarantees success or repairs wrong models/costs.

Planning terms

TermFormal definitionPlain explanationAnalogyMisconception
PlannerFinite-budget action-sequence selection using model, cost, and constraints.Chooses trials; model estimates outcomes.Travel agent.Planner equals world model; search restores discarded information.
CostScalar ranking rule for predicted outcomes/actions, usually lower-is-better.Says which candidate looks preferable.Judge’s scorecard.Necessarily training reward or feasibility.
Terminal latent distancev3 squared Euclidean cost between final predicted and encoded goal latents.Ruler from endpoint to goal.Straight map line.Equals shortest path, progress, temporal distance, or reachability.
Decision-metric alignmentAgreement between predicted and realized candidate rankings under a protocol.Does rehearsal choose reality’s winner?Two judges rank contestants.One correlation, probe, or one-step score guarantees it.
ReachabilityWhether permitted dynamics/actions reach a state, often directionally within budget.Can I get there under these rules?Adjacent rooms with locks.Symmetric, visual similarity, or certified by a learned score.
CEMSample action sequences, keep cost elites, refit proposal for finite rounds.Winners shape the next audition.Repeated tournament.Exhaustive search, global guarantee, or immune to model error.
MPCPlan, execute a configured prefix, observe, and plan again.Trust imagination briefly, then check.Sailor redraws each leg.Always one action; removes bias; is a learned controller.

Evidence and diagnosis terms

TermFormal definitionPlain explanationAnalogyMisconception
Support mismatchQueries lack adequate evidence in collected/training distributions.Planner asks about a barely learned neighborhood.Unsurveyed map blank.Outside means impossible; global action frequency proves local support.
Model exploitationSearch prefers actions because systematic model/cost error looks attractive.Optimizer finds a simulator loophole.Optimize a rules typo.Normal exploration, deception, or proof CEM is broken.
DecodabilityA specified probe/split can recover a variable from fixed representations.Can this reader extract it?Clerk reads coded passport.The model uses it; it is causal; it proves understanding.
ControllabilityState distinctions influenceable/reachable by admissible actions under named dynamics/horizon.What can controls change?Elevator floor buttons.Action sensitivity establishes controllability or observability.
IdentifiabilityUnder assumptions/equivalence, latent model recovery is unique up to allowed transforms.Is the hidden explanation uniquely recoverable?ID valid in one jurisdiction.Interpretability or true physics in arbitrary worlds.
Causal identificationIntervention effects/structure established by design or identification assumptions.What changed because one factor changed?Test one circuit switch.Correlation, probe, separated branches, or any ablation proves cause.
Planning utilityEmpirical decision usefulness under a declared task, planner, and budget.Does it help choose actions here?Map judged by trips.Intrinsic, task-free scalar.
Action-conditioned predictive consistencyMatched action-rollout diagnostics; ACPC uses divergence, Invariance Radius, Separation Rate.Stable for nuisance; different for consequence.Repaint sign versus move road.Consistency proves truth, calibrated uncertainty, or safety.
Independent reproductionEvidence outside authors reruns/reimplements a claim under documented version/protocol.Another investigator checks what they ran.Second lab repeats one assay.Author code or TwoRoom verifies every task/method.
Violation of ExpectationMatched surprise comparison for ordinary and deliberately perturbed event sequences.Does error react to the chosen violation?Costume versus motion-rule change.Surprise proves intuitive physics, causality, or broad anomaly understanding.

Try to break the vocabulary

“The teacher forces the target goal through a rollout, so low latent distance proves the state is reachable.”

Every noun is real; every connection is wrong. Teacher forcing supplies recorded context. A target representation is an encoded successor; a goal belongs to planning. A rollout imagines latents under actions. Terminal distance ranks endpoints but does not prove budgeted directed reachability.

Repair the sentence by attaching every word to one object, phase, assumption set, and evidence type. Two mechanisms hiding under one word signal a category error.

Evidence receipt

Definitions follow LeWorldModel v3, frozen official code, LeJEPA v3, passive identifiability v1, controlled identifiability v2, and TwoRoom reproduction v1. Appendix E pins frontier versions. Cutoff: 2026-08-20. Analogies are tutorial constructions; independent evidence remains TwoRoom-only under its protocol.

Three quick questions

  1. Where do target representation and goal image enter?
  2. Why can information be decodable but unused and non-causal?
  3. How do distance, reachability, and planning utility differ?