JEPA4Japan · tutorials

Appendix E — Paper Timeline and Evidence Levels

1,609 words 8 min read #LeWorldModel#World Models#JEPA

Track which system layer each paper changes, its target failure, version and release date, public code, evidence strength, reproduction status, and limits on generalization.

Course progress Course outline 48 of 48 lessons available

Part 0 — Reading Guide: What Exactly Are We Going to Learn?

  1. 01 Chapter 0 — Before You Begin available now

Part 1 — World Models: An Internal Sandbox for the Agent

  1. 02 Chapter 1 — Why an Agent Needs to “Imagine the Future” available now
  2. 03 Chapter 2 — Why Not Predict the Next Image Directly? available now
  3. 04 Chapter 3 — The JEPA Idea: Predict Meaning, Not a Replica of the Image available now
  4. 05 Chapter 4 — Understand LeWM in One Diagram available now

Part 2 — Turning Images into State: The LeWM Architecture

  1. 06 Chapter 5 — Trajectory Data: To the Model, the World Is Not an Image Collection available now
  2. 07 Chapter 6 — The Visual Encoder: Issuing a “State Passport” for Every Frame available now
  3. 08 Chapter 7 — The Dynamics Predictor: Moving Time Forward in the Mind available now
  4. 09 Chapter 8 — A Complete Forward Pass: Follow One Batch from Start to Finish available now

Part 3 — Preventing the Model from Cheating: Prediction Loss and SIGReg

  1. 10 Chapter 9 — The Most Dangerous Shortcut: Representation Collapse available now
  2. 11 Chapter 10 — Prediction Loss: How the Model Learns the Next Step available now
  3. 12 Chapter 11 — The Intuition Behind SIGReg: Letting Representation Space “Breathe” available now
  4. 13 Chapter 12 — Keep the Mathematics Minimal but Sufficient available now
  5. 14 Chapter 13 — The Original LeWM’s End-to-End Training Mechanism available now
  6. 15 Chapter 14 — Train a Model That Does Not Collapse Immediately available now

Part 4 — Putting the Model into Action: Planning in Latent Space

  1. 16 Chapter 15 — Goal-Conditioned Planning: From “Where Am I?” to “Where Do I Want to Go?” available now
  2. 17 Chapter 16 — Latent Euclidean Distance: Convenient, but Not Necessarily Reliable available now
  3. 18 Chapter 17 — CEM: Searching for Actions Through an Elimination Tournament available now
  4. 19 Chapter 18 — MPC: Do Not Trust the Model for Too Long at Once available now
  5. 20 Chapter 19 — Long-Horizon Rollouts: How Small Errors Snowball into Major Failures available now
  6. 21 Chapter 20 — Implement a Minimal LeWM Planner from Scratch available now

Part 5 — Engineering Reproduction: From Paper to Running System

  1. 22 Chapter 21 — The Official Repository and Experimental Environment available now
  2. 23 Chapter 22 — First Experiment: A TwoRoom Smoke Test available now
  3. 24 Chapter 23 — Second Experiment: Reproducing PushT available now
  4. 25 Chapter 24 — How to Evaluate a World Model Fairly available now
  5. 26 Chapter 25 — Failure-Diagnosis Manual available now

Part 6 — What Has LeWM Actually Learned?

  1. 27 Chapter 26 — Linear Probes: Which Physical Variables Are Encoded in the Latent State? available now
  2. 28 Chapter 27 — Give Latent Space a “Health Check” available now
  3. 29 Chapter 28 — Violation of Expectation: Is the Model Surprised by “Impossible Events”? available now
  4. 30 Chapter 29 — How to Discuss “Understanding the World” Rigorously available now

Part 7 — Why “Accurate Prediction” Can Still Produce “Poor Planning”

  1. 31 Chapter 30 — The Gap Between the Training Objective and the Planning Objective available now
  2. 32 Chapter 31 — Global Non-Collapse Does Not Guarantee Preservation of Task-Relevant Dynamics available now
  3. 33 Chapter 32 — When Is an Isotropic Gaussian Prior Too Strong? available now
  4. 34 Chapter 33 — Long-Horizon Planning: Predict Farther or Plan More Intelligently? available now
  5. 35 Chapter 34 — From Positional Distance to Task Progress available now
  6. 36 Chapter 35 — Multi-Task Learning, Real Robots, and Visual Distractions available now
  7. 37 Chapter 36 — Theoretical Boundaries: When Can the True State Be Identified? available now

Part 8 — From Reproducer to Researcher

  1. 38 Chapter 37 — Design a Credible LeWM Improvement Experiment available now
  2. 39 Chapter 38 — Twelve Executable Research Projects available now
  3. 40 Chapter 39 — Open Questions in LeWM Research available now

Appendices

  1. 41 Appendix A — The Minimum Necessary Mathematical Toolkit available now
  2. 42 Appendix B — PyTorch Implementation Quick Reference available now
  3. 43 Appendix C — Complete Tensor-Shape Table available now
  4. 44 Appendix D — Experiment Configuration Cards available now
  5. 45 Appendix E — Paper Timeline and Evidence Levels Current lesson
  6. 46 Appendix F — Glossary available now
  7. 47 Appendix G — Reproduction Checklist available now
  8. 48 Appendix H — Expert-Review Checklist available now

The big picture

  1. When?A version has a public date.
  2. Where?It changes one system layer.
  3. What evidence?Paper, code, theorem, and replication differ.
A newer paper is newer—not automatically stronger. Every claim keeps its exact version, changed layer, evidence type, code status, and reproduction scope.

A tiny story: the library with three shelves

A librarian puts every book on one vertical shelf. Yesterday’s preprint sits above an old theorem, and a public-code sticker looks like a replication medal. Readers naturally mistake height for strength.

This ledger uses three independent shelves: time, system layer, and evidence. A method definition can be verified from a primary paper while its experiments remain author-reported. A theorem can be rigorous inside assumptions without broad empirical evidence. Public author code improves inspection but is not independent reproduction.

The evidence cutoff is 2026-08-20, Asia/Tokyo. Dates are arXiv submission/revision dates unless labeled as code commits or releases. The only independent reproduction established here is the stated TwoRoom reimplementation and protocols in arXiv:2608.10145v1.

A system map keeps data, regularization, dynamics, cost, search, hierarchy, and evaluation in separate layers.

A later card may change one layer or add a head. It does not travel backward and redraw original LeWorldModel v3.

Baseline and theory anchors

Public version/dateIdentity and layerEvidence boundary
LeJEPA v3: v1 2025-11-11, v2 11-12, v3 11-14representation objective and SIGReg motivationprimary method and assumption-bounded theory; public repository, no snapshot pinned here; not a control system or independent reproduction
LeWorldModel v3: v1 2026-03-13, v2 03-24, v3 06-03baseline action-conditioned end-to-end model plus latent CEM/MPCfrozen code 8edfeb3, 2026-05-22; one shared trainable encoder, connected next target, no EMA teacher or pretrained visual encoder; benchmark outcomes author-reported
Passive identifiability v1, 2026-05-25population identifiability for a specified passive linear/Gaussian/stationary classassumption-bounded orthogonal recovery; code de7503f, 2026-05-27; not full benchmark coverage
Controlled identifiability v2: v1 2026-07-24, v2 07-27controlled conditional-mean identifiability with state-conditioned action excitationassumption-bounded theory; no author-linked code identified; not a theorem for arbitrary nonlinear or multimodal futures
SIGReg as Variational Free Energy v1, 2026-07-15theoretical interpretation under constant encoder noise and exact Gaussian enforcementnot identifiability; empirical validation left open; finite minibatch SIGReg is not exact enforcement

Direct frontier ledger

Every row below is a later route, never an ingredient silently added to original v3. “AR” means experimental outcomes are author-reported; “code” means inspectable author-linked implementation, not independent replication.

Version/dateChanged layer and status
RC-aux v1, 2026-05-08budget-conditioned directed reachability plus multi-horizon prediction; AR; code ecb4496, 2026-07-02; score is not a logical oracle
Latent Geometry Beyond Search v2: v1 05-09, v2 06-05amortized goal-conditioned inverse-dynamics planner on pretrained LeWM; AR 100–130× and 7/8 comparisons are protocol-bound; code 48c45b1
Sub-JEPA v1, 2026-05-10frozen random orthogonal subspace regularization; AR; code ef945ed, 2026-07-27; finite views do not certify topology
stable-worldmodel v1, 2026-05-20infrastructure/evaluation layer; code addbab4, 2026-08-18; formal release remained 0.1.1 from 2026-06-06; infrastructure does not certify fairness
TRM v1, 2026-05-21post-hoc horizon-matched trajectory reachability metric; AR; no author-linked code identified
Subspace-Decomposed JEPA v1, 2026-05-29progression/content subspaces; AR; code 1cc1210
SMWM v1, 2026-06-18inverse-dynamics regression replaces Gaussian anti-collapse; AR; code 9d22bcc, 2026-06-30
VLWM v1, 2026-06-19variable-length direct latent prediction; AR 13% average uses seed 3072 and post-hoc best P1/P2/P3 per (dataset, Δ); no code identified
Fast-LeWM v1, 2026-06-24parallel action-prefix horizon prediction; AR speed depends on hardware/budgets; code 492752d, 2026-08-06; not multimodal prediction
Hi-LeWM v2: v1 07-14, v2 07-15high-level latent subgoals above frozen low-level LeWM; AR support-mismatch failures and conditional gains; code 4bb21a2 plus Zenodo
Depth-Regularized JEPA v1, 2026-07-15depth supervision and training-only capacity on real agricultural video; AR offline probes/rollouts, not online planning; no code identified
Temporal-Distance JEPA v2: v1 07-28, v2 07-29directed temporal cost and rollout consistency; AR; code b4c17ca; later detach choices do not redefine v3
TC-LeWM v2: v1 07-29, v2 07-31temporally centered SIGReg; AR multi-task LIBERO behavior cloning, not baseline MPC; no code identified
INTACT v1, 2026-07-28search-free action interface on LeWM tasks; AR; code a3a7bf9
QQWorld v1, 2026-07-30quantile matching on projections, optional detached past-sample queue; AR; no code identified; its queue is not a v3 target detach
ProWorld v1, 2026-08-03hindsight progress order, hyperbolic geometry, intermediate cost; AR with detour/backtracking limits; no code identified
PhyLatent v1, 2026-08-06simulator-state grounding and diagnostics; AR; no code identified; its stopped-gradient baseline drawing cannot define original v3
PSG-JEPA v1, 2026-08-07training-only proprioception and joint-change grounding; AR Mobile ALOHA downstream-policy evidence, not v3 MPC or safety; code 8de96b8, 2026-08-14
TwoRoom reproduction v1, 2026-08-10independent only for its TwoRoom reimplementation/protocols; tinylab efa9e5d; historical 100/150 belongs to v1/v2, while current v3 says 25/50
VIScore v2: v1 08-11, v2 08-12Veracity, Influence, Sobriety diagnostics; AR; code bbb60fc
Objective Is the Bottleneck v1, 2026-08-13follow-up on audited checkpoints: terminal-cost change moves one offset-100 result from 26% to 98%; author analysis, TwoRoom/protocol-bound
ACPC v1, 2026-08-13pairwise divergence, Invariance Radius, Separation Rate, and planner-cost bounds; AR; code 90d4276
Traj-LeWM v1, 2026-08-14learned goal-conditioned trajectory cost plus endpoint score; AR; code 67577fa
SCALE v1, 2026-08-17aligns latent pair distance with privileged state distance; AR; no code identified
AC-MTM v1, 2026-08-18contrastive inverse-action anti-collapse; AR under informative observable continuous actions; code 43c96f6
DA-LeWM v1, 2026-08-19Plan–Real/CEM-stage rank agreement and decision-alignment objectives; AR; no code identified

Neighboring routes and the border

The 2024 hierarchical RSSM study appeared 2024-06-01; “HWM” is course shorthand, not its official name or a LeWM variant. Value-guided JEPA planning appeared 2025-12-28. Temporal Straightening v3 was revised 2026-08-11; its public code 9efb6fd predates v3. Qantara, Branch-JEPA v3, UniJEPA, monotone planning costs, and reinforced planning are neighboring routes, not LeWM releases.

ContactGuard, Calibrated Predictive Safety, UWM-JEPA, and Policy-Guided World Model Planning remain outside this direct ledger because they do not directly define or diagnose the baseline claims audited here. Exclusion is a scope boundary, not a quality judgment.

Try to break the catalog

Paper A has public author code and yesterday’s result. Paper B has an older theorem with explicit assumptions. Paper C has a documented result repeated by another group. Asking which is “most mature” has no single answer: A has inspectability, B has assumption-bounded proof, and C has independent empirical evidence under its protocol.

The safe procedure is mechanical: identify the exact claim; attach version and changed layer; separate method definition from outcome; record author code and independent reproduction separately; then narrow the sentence. Blank code or reproduction fields mean “not identified by the cutoff,” not “false.”

Evidence receipt

Paper identity and mechanisms above are primary-source facts. Experimental outcomes are author-reported unless the row explicitly says independent. Theory stays inside its assumptions. Code availability never upgrades an author result to replication. Workshop acceptance, recency, one benchmark, or one theorem cannot rewrite original v3 or become field consensus.

Three quick questions

  1. Why are date, public code, and independent reproduction three different fields?
  2. Which single row is independently reproduced here, and how narrow is that scope?
  3. What must a future update record before filling a blank code field?