JEPA4Japan · チュートリアル

付録H — 専門家査読チェックリスト

1,682文字 6分で読めます #LeWorldModel#World Models#JEPA

オリジナルの学習時の計算グラフ、後続の派生モデルとの区別、プローブの解釈、ベンチマークの一般化、因果、理論仮定、負の結果、再現に使うコードの場所を監査します。

コース進捗 コース目次 48レッスン中 48件を公開中

第0部 読み方ガイド:私たちは何を学ぶのか

  1. 01 第0章 — はじめる前に 公開中

第1部 世界モデル:エージェントの頭の中にある実験場

  1. 02 第1章 — なぜエージェントには「未来を想像する」力が必要なのか 公開中
  2. 03 第2章 — なぜ次の画像をそのまま予測しないのか 公開中
  3. 04 第3章 — JEPAの発想:画面の複製ではなく意味を予測する 公開中
  4. 05 第4章 — 1枚の図でLeWMを理解する 公開中

第2部 画面を状態に変える:LeWMのモデル構造

  1. 06 第5章 — 軌跡データ:モデルにとって世界は画像集ではない 公開中
  2. 07 第6章 — 視覚エンコーダー:各フレームに「状態パスポート」を発行する 公開中
  3. 08 第7章 — 動力学予測器:頭の中で時間を前へ進める 公開中
  4. 09 第8章 — 完全な順伝播:1バッチを最初から最後まで追う 公開中

第3部 モデルの抜け道を防ぐ:予測損失とSIGReg

  1. 10 第9章 — 最も危険な近道:表現崩壊 公開中
  2. 11 第10章 — 予測損失:モデルはどのように次の一歩を学ぶのか 公開中
  3. 12 第11章 — SIGRegの直感:表現空間に「呼吸」をさせる 公開中
  4. 13 第12章 — 必要最小限の数学 公開中
  5. 14 第13章 — オリジナルLeWMのエンドツーエンド学習の仕組み 公開中
  6. 15 第14章 — すぐに表現崩壊しないモデルを訓練する 公開中

第4部 モデルを行動に使う:潜在空間での計画

  1. 16 第15章 — 目標条件付き計画:「今いる場所」から「行きたい場所」へ 公開中
  2. 17 第16章 — 潜在ユークリッド距離:便利だが、常に信頼できるとは限らない 公開中
  3. 18 第17章 — CEM:勝ち抜き方式で行動を探索する 公開中
  4. 19 第18章 — MPC:モデルを一度に長く信じすぎない 公開中
  5. 20 第19章 — 長期ロールアウト:小さな誤差が大事故へ育つまで 公開中
  6. 21 第20章 — 最小のLeWMプランナーをゼロから実装する 公開中

第5部 エンジニアリング再現:論文から動くシステムへ

  1. 22 第21章 — 公式リポジトリと実験環境 公開中
  2. 23 第22章 — 最初の実験:TwoRoomのスモークテスト 公開中
  3. 24 第23章 — 2つ目の実験:PushTを再現する 公開中
  4. 25 第24章 — 世界モデルを公平に評価する方法 公開中
  5. 26 第25章 — 失敗診断マニュアル 公開中

第6部 LeWMは何を学んだのか

  1. 27 第26章 — 線形プローブ:潜在状態にはどの物理量が含まれるのか 公開中
  2. 28 第27章 — 潜在空間を「健康診断」する 公開中
  3. 29 第28章 — 期待違反:モデルは「あり得ない出来事」に驚くのか 公開中
  4. 30 第29章 — 「世界を理解する」を厳密に語るには 公開中

第7部 「予測が正確」でも「計画がうまくいかない」のはなぜか

  1. 31 第30章 — 訓練目的と計画目的のあいだにある亀裂 公開中
  2. 32 第31章 — 大域的には表現崩壊していなくても、タスクに必要な動力学が保たれるとは限らない 公開中
  3. 33 第32章 — 等方ガウス事前分布はいつ強すぎるのか 公開中
  4. 34 第33章 — 長期計画:より遠くを予測するか、より賢く計画するか 公開中
  5. 35 第34章 — 位置の距離からタスクの進捗へ 公開中
  6. 36 第35章 — マルチタスク、実ロボット、視覚的外乱 公開中
  7. 37 第36章 — 理論的な境界:真の状態はいつ同定できるのか 公開中

第8部 再現者から研究者へ

  1. 38 第37章 — 信頼できるLeWM改良実験を設計する 公開中
  2. 39 第38章 — 実行可能な12の研究課題 公開中
  3. 40 第39章 — LeWM研究の未解決問題 公開中

付録 数学・実装・再現・査読のための参照資料

  1. 41 付録A — 最低限必要な数学ツールキット 公開中
  2. 42 付録B — PyTorch実装クイックリファレンス 公開中
  3. 43 付録C — テンソル形状の完全一覧 公開中
  4. 44 付録D — 実験設定カード 公開中
  5. 45 付録E — 論文タイムラインとエビデンスレベル 公開中
  6. 46 付録F — 用語集 公開中
  7. 47 付録G — 再現チェックリスト 公開中
  8. 48 付録H — 専門家査読チェックリスト 現在のレッスン

まず大きな絵

  1. 正しい machine?graph と version を trace。
  2. 正しい evidence?probe、cause、theorem、benchmark は別。
  3. 正しい scope?condition、failure、reproducible path を pin。
expert review は、claim の全 word が正しい model、evidence type、condition、inspectable path に結ばれているかを尋ねます。

小さなお話:八人の inspector

manuscript は流暢で、figure は美しく、table の decimal も整っています。それでも八人の inspector が止めます。一人は発明された teacher network、次は v3 に密輸された後続 preprint、残りは probe が何を証明するか、benchmark condition はどこか、“caused” は妥当か、theorem assumption は何か、failed seed はどこか、code link がなぜ moving main なのかを尋ねます。

graph、version、probe、scope、causality、theory、negative evidence、reproducible path の別 gate を担当する八人の inspector。

check mark は evidence ではありません。各 pass は primary source、frozen implementation、controlled experiment、explicit assumption、raw result、declared tutorial construction のどれかを指します。

review 前に “LeWM works” を、subject、mechanism、model version、environment、data、planner/budget、evidence type、conclusion strength を持つ checkable sentence に書き換えます。

八つの gate checklist

H.1 Original training graph は正確か?

  • 一つの shared trainable visual encoder が context と shifted next observation を処理。
  • shifted target は同じ model が encode し detach しない。prediction gradient は connected encoder route の両方へ届く。
  • action は transition と align して predictor を condition。SIGReg は observation-embedding population を regularize。
  • Original v3 に stop-gradient target、EMA teacher、pretrained visual encoder はない。
  • goal image、CEM、reward、probe、optional decoder は baseline training lane の外。

この gate の pass が証明するのは graph identity と可能な gradient route であり、non-collapse、physics、rollout accuracy、planning success ではありません。

H.2 Baseline と後続 route は分離されているか?

  • 全 later method に passport を付ける。title、exact version/date、changed layer、targeted failure、evidence level、code status、independent-reproduction status。
  • LeJEPA は引用した motivation/theory にだけ使い、後続 detach、auxiliary head、cost、hierarchy、teacher を v3 に import しない。
  • “defines/proposes”“authors report”“theorem under assumptions”“independently reproduced”“tutorial construction” を区別。
  • recency、acceptance、public author code を別 fact とし、evidence ladder にしない。

H.3 Probe と picture を控えめに解釈しているか?

  • probe family/capacity、train/test split、preprocessing fit、episode leakage control、shuffled label、simple/raw-feature baseline を記録。
  • “この protocol で linearly decodable” と言い、“model が variable を learned/uses/understands” と言い換えない。
  • t-SNE/UMAP、reconstructed image、latent trajectory、VoE surprise を conditional view とし、semantics/intuitive physics の証明にしない。
  • use/cause を主張するなら intervention を足し、その side effect を開示。

H.4 Benchmark claim は condition 内にあるか?

  • environment/task、data、model、planner、action/search/compute budget、pretraining、seed policy、episode ledger、metric、hardware timing を付ける。
  • speed multiplier の denominator を検査。model call、batching、rendering、candidate、horizon、GPU、pretraining accounting。
  • scope を分離。TwoRoom topology、PushT 2-D contact、simulated Reacher、simulated OGBench-Cube だけでは general robotics、physical understanding、safety を確立しない。
  • author-reported weakness と possible explanation は “possible” のまま。mean だけでなく goal/failure distribution を調べる。

H.5 Language は correlation と causation を分けるか?

  • “velocity と correlate、したがって uses velocity” や “SIGReg が success を変えた、したがって Gaussianity が physics を caused” の jump を丸で囲む。
  • 最も近い alternative を challenge。physical state を固定して appearance を変える。dynamics 固定で cost を swap。cost 固定で dynamics を swap。oracle cell で bottleneck を局所化。
  • ablation は capacity、optimization、support、search query を一緒に変えうる。causal wording を実際に isolate した intervention に合わせる。

H.6 Theorem assumption は conclusion に付いているか?

  • 必要に応じ stationarity、transition/noise family、observation adequacy/invertibility、Gaussian/exact constraint、matched dimension、predictor expressivity、population/global optimum、signal separation、state-conditioned action excitation を明記。
  • conclusion を狭く保つ。orthogonal recovery は human-named axis を作らず、controlled theory は conditional mean を identify し全 multimodal future ではない。
  • benchmark gap を別々に列挙。nonlinearity、partial observability、missing action support、drift、multimodality、finite data、approximate optimization、misspecification。
  • assumption が unknown/false なら “diagnostic lens, not direct coverage” と言う。

H.7 Positive、negative、null、missing result は見えるか?

  • 全 declared seed または preregistered exclusion rule、uncertainty、per-episode outcome、divergence/collapse/timeout、unsuccessful ablation、stated policy で選んだ representative failure を示す。
  • 未 test は open、測定して非分離は null under this setup、criterion 未達は failed under this setup。
  • negative result を task、method、support、budget、protocol より広く generalize せず、物語を弱めるからと隠さない。

H.8 全 implementation claim に reproducible path があるか?

  • frozen first-party file/SHA、pinned artifact + hash/format、または clearly labeled tutorial construction + divergence のどれかで終わる。
  • dependency を別 pin、composed configuration、data/checkpoint conversion、episode ledger、analysis revision を保存。SHA があるのに moving main を使わない。
  • archive smoke test。checkpoint 一つを load、window 一つを episode へ trace、summary 一つを regenerate。これは handoff test で independent reproduction ではない。
  • history を保存。current v3 TwoRoom は 25/50、reproduction が audit した paper-side 100/150 は v1/v2。

Evidence は staircase ではなく coordinate

Decodability、predictability、controllability、generalization、causal identification、planning utility は別の質問です。decoded position を predictor が無視でき、recorded-action accuracy と counterfactual failure は共存でき、planner は identifiable state なしに task shortcut を exploit できます。

六つの evidence test を、understanding へ自動上昇しない separate coordinate として置いた図。

どの coordinate を test し、何が open かを尋ねます。複数 indirect signal を “understanding” 一語に混ぜません。

claim を壊してみる

“LeWM stabilizes its pretrained visual teacher with EMA and stop-gradient, then proves through probes that its Gaussian state understands physics and makes CEM universally faster.”

この流暢な sentence は全 gate に失敗します。graph は誤り、later mechanism を import、probe を understanding に昇格、speed から hardware/budget condition を除去、correlation を causality に変更、theorem assumption なし、failure 消失、frozen path なしです。

red-flag phrase は自動 rejection でなく question を起動します。proves understanding、guarantees physics、universally faster、distance is reachability、the ablation shows the cause。strong wording が残れるのは operational definition、alternative、budget、assumption、direct evidence も一緒に残るときだけです。

証拠レシート

baseline identity は LeWorldModel v3、frozen official code、cited motivation の LeJEPA v3 に基づきます。theory は passive identifiability v1 と controlled identifiability v2。independent evidence は TwoRoom reproduction v1 と documented protocol に限定します。LeWM v3 は 2026-06-03 改訂、frozen code は 2026-05-22、この checklist cutoff は 2026-08-20。全 gate pass が意味するのは dated claim が正しく attributed/scoped/calibrated/balanced/inspectable ということです。model が永遠に正しい、specialist peer review が不要、ではありません。

3つのクイック質問

  1. v3 図に発明された EMA teacher があると暴く direct fact は何ですか?
  2. strong probe が planner の variable use を確立しないのはなぜですか?
  3. 全八 gate を通過した後の最も強い conclusion は何ですか?