コース進捗 コース目次 48レッスン中 48件を公開中
第0部 読み方ガイド:私たちは何を学ぶのか
第1部 世界モデル:エージェントの頭の中にある実験場
第2部 画面を状態に変える:LeWMのモデル構造
第3部 モデルの抜け道を防ぐ:予測損失とSIGReg
第4部 モデルを行動に使う:潜在空間での計画
第5部 エンジニアリング再現:論文から動くシステムへ
第6部 LeWMは何を学んだのか
第7部 「予測が正確」でも「計画がうまくいかない」のはなぜか
第8部 再現者から研究者へ
付録 数学・実装・再現・査読のための参照資料
まず大きな絵
- predictこの action の次は何か。
- rankどの ending を好むべきか。
- searchCEM はどの plan を見たか。
- executeworld で何が起きたか。

小さなお話
courier は小さな一歩を全部正しく予測します。goal は壁越しに近いので、planner はずっと「goal へ歩く」を選びます。doorway へ一度離れず、壁を押し続けます。
この話に bad next-step prediction は不要です。training は「この action の次 latent は正しいか」、planning は「どの行動列を好むか」を聞きます。元の LeWM v3 は2つ目を terminal squared latent Euclidean distance で答えます。forecast と route judgment は別物です。
本当のルール
low mean squared prediction error がよい ordering を保証しない理由:
- common local motion が rare doorway/contact decision を隠す。
- nearby transition を予測できても wall-separated endpoints の配置は悪い。
- planning は recorded context だけでなく self-fed model-generated state を採点。
- Euclidean cost は symmetric で、direction、budget、support、uncertainty 引数がない。
全部を変える前に oracle substitution を行います。
oracle rollout + learned cost でも ranking が悪い
→ cost または representation geometry
learned rollout + oracle cost が失敗
→ multi-step dynamics と support
両 oracle は成功、deployed planner は失敗
→ search、observation、action conversion、execution
oracle は privileged simulator access を持つ laboratory instrument で、deployable pixels-only component ではありません。
後続2案は別の仕事を変えます。RC-aux は original one-step term を、step 1 も含む weighted multi-horizon open-loop prediction に置換し、budget-conditioned reachability head を加えます。Temporal-Distance JEPA は directed temporal cost と、(H=5) で各 predicted latent を recorded future の stop-gradient encoding に合わせる rollout-consistency loss を別々に学びます。この detach は later method のもので、original v3 target は connected です。
trajectory offset は behavior が実際にかけた時間で、oracle shortest time ではありません。cross-trajectory negative も false になれます。
だまされる反例
最初は定規だけ交換します。fixed TwoRoom candidate bank の exact simulator endpoint を使い、straight-line distance と wall-respecting graph distance で順位を付けます。wall を消して繰り返します。
wall が detour を作るときだけ graph distance が ranking を変えるなら、topological-cost explanation はこの toy removal test を生き残ります。exact endpoint でも両 cost が失敗すれば candidate coverage/task definition を調べます。その後 learned encoding、reverse direction、budget variation、最後に learned rollout を加えます。
later methods を全部同じ “planning-aware objective” と呼びません。
- Objective Bottleneck:1つの TwoRoom representation/predictor を freeze し planner cost を交換。
- Traj-LeWM:goal-conditioned latent trajectory cost、synthetic ranking negatives、各 epoch 後の failed endpoint-only execution mining。
- Decision-Metric Alignment:Plan–Real/CEM-stage Spearman を測り、training-only inverse-dynamics/goal-action heads を追加しつつ Euclidean MPC cost を維持。
それぞれ ruler、path、representation geometry を変えます。
実験のレシート
Objective Bottleneck follow-up は latent L2 vs true distance (r=0.426)、position ridge probe (R^2=0.9922) を報告します。自分の checkpoint の offset-100 success は temporal cost で26%→98%、released checkpoint は14%→34%、privileged decoded-position cost は70%。1環境、1 seed/checkpoint、非 v3 default long-goal protocol です。
Traj-LeWM は自分の LeWM comparison に対し PushT/Cube/Reacher/TwoRoom で3、14、7、7 percentage points gain、3 evaluation seeds の平均を報告します。failure mining があるため pure offline-data training ではありません。
Decision-Metric Alignment は各 config 1 training run、3 seeds は evaluation variation。rank diagnostic は simulator rollout が必要です。author-reported で independent reproduction ではありません。
baseline source は LeWorldModel v3、commit 8edfeb3。later stop-gradient、reachability、temporal、action-head objective を v3 へさかのぼらせません。
3つのクイック質問
- oracle dynamics でも TwoRoom planning が悪くなるのはなぜですか。
- multi-horizon prediction と reachability score は何を別々にしますか。
- predictor より ruler を先に変えるべきだと示す oracle cell はどれですか。