コース進捗 コース目次 48レッスン中 48件を公開中
第0部 読み方ガイド:私たちは何を学ぶのか
第1部 世界モデル:エージェントの頭の中にある実験場
第2部 画面を状態に変える:LeWMのモデル構造
第3部 モデルの抜け道を防ぐ:予測損失とSIGReg
第4部 モデルを行動に使う:潜在空間での計画
第5部 エンジニアリング再現:論文から動くシステムへ
第6部 LeWMは何を学んだのか
第7部 「予測が正確」でも「計画がうまくいかない」のはなぜか
第8部 再現者から研究者へ
付録 数学・実装・再現・査読のための参照資料
まず大きな絵
- 予測次に何が来るはずか。
- 1つ変えるcolor、state、support。
- mismatch を測るspike は sensitivity で、belief ではない。

小さなお話
ball が screen の後ろへ消え、stage の反対側へ現れます。子どもも驚き、motion detector も反応します。両方とも event に反応しましたが、physical expectation を持つとは限りません。
Violation-of-Expectation(VoE)test は、expected continuation を作り、intervention を入れ、prediction–observation mismatch を測ります。大きな teleport spike が最初に示すのは、その event が model prediction とより大きく違ったことです。model が teleportation を不可能と知る証拠ではありません。
本当のルール
同じ prehistory、actions、intervention time、evaluation window の paired triplets を使います。
- unperturbed continuation。
- appearance intervention。
- teleportation のような state discontinuity。
LeWM v3 の3環境:
- TwoRoom:agent color change vs agent teleport。
- PushT:block color change vs agent と block の両方を teleport。
- OGBench-Cube:cube color change vs cube teleport。
PushT は1-factor の “color vs physics” ではありません。teleport branch は2 objects と関係を変え、color branch は1 object の appearance を変えます。
論文は、3環境で teleportation surprise が paired test、threshold 0.01 未満で有意に高いと報告します。main figure の color effects は弱く non-significant です。慎重な文は「authors’ protocol の tested color intervention より tested teleport intervention に mismatch signal が強く反応した」です。
frozen public repository には authors’ exact reduction を復元する VoE code が足りません。course implementation は predicted next latent と actual next observation の frozen encoder latent の mean squared coordinate difference を tutorial scalar として定義できますが、tutorial diagnostic と明記します。
だまされる反例
last-frame predictor は current latent をコピーし、action を無視します。ordinary slow motion は modest error、teleportation は large spike、tiny color change は smaller spike。contact dynamics を学ばずに weakest ordering を通ります。
この説明を攻撃する controls:
- 全 branch で same actions。
- image-change magnitude を overlap。
- in-support fast motion と support-matched reset。
- matched displacement の one-object/two-object interventions。
- last-frame、stagnant-latent、action-shuffled、appearance-only、persistence baselines。
time trace 全体を見ます。intervention 前 response は leakage/alignment error を示します。one-step spike と recovery、lasting plateau は違いますが、自分で原因を診断しません。新観測で re-anchor することも、unsupported reset/self-fed drift が続くこともあります。
intervention strength、support bin、mismatch reduction、fixed-time/peak statistic、exclusion rule、episode-level unit を pre-register します。teleport test spike を見てから normalization を fit しません。
実験のレシート
last clean history、action block、intervention frame、before/after observations、affected objects、simulator reset states、policy、seed、checkpoint、attachment point、aggregation rule を保存します。dashed predicted latent と solid encoded observation を分けます。decoder image は protocol が明示しない限り別 readout です。
後続 ACPC v1 は clean/visually perturbed history に same action sequence を与え、rollout divergence を測ります。Invariance Radius(IR)と Separation Rate(SR)を導入し、multi-step prediction error と、shared candidate pool 条件で planner-cost change への samplewise bounds を示します。著者は4 tasks、3 seeds/condition、public code 90d4276 で trend を報告します。screen は observed source-task success labels を使い、blur/resize severity は1つ。ACPC で plan 選択すると control が改善するとは示していません。
VoE が示すのは named protocol 下の perturbation sensitivity です。object permanence、causal physics、human-like cognition、planner use、unseen violation への generalization を単独では証明しません。
3つのクイック質問
- VoE triplet で同じに保つものは何ですか。
- PushT の color vs teleport が1-factor physics test でないのはなぜですか。
- teleportation に “surprised” になれる単純 baseline は何ですか。