コース進捗 コース目次 48レッスン中 48件を公開中
第0部 読み方ガイド:私たちは何を学ぶのか
第1部 世界モデル:エージェントの頭の中にある実験場
第2部 画面を状態に変える:LeWMのモデル構造
第3部 モデルの抜け道を防ぐ:予測損失とSIGReg
第4部 モデルを行動に使う:潜在空間での計画
第5部 エンジニアリング再現:論文から動くシステムへ
第6部 LeWMは何を学んだのか
第7部 「予測が正確」でも「計画がうまくいかない」のはなぜか
第8部 再現者から研究者へ
付録 数学・実装・再現・査読のための参照資料
まず大きな絵
- spreadcloud は生きているか。
- neighbors近い state は役立つ隣人か。
- pathsaction と event が path に出るか。
- picturesきれいな図を test に変える。

小さなお話
血圧が完璧でも、階段を上れない患者はいます。測定は本物ですが、身体全部ではありません。
latent cloud も broad variance を持ちながら contact、velocity、wall side を失えます。slow background だけ覚えれば、美しい smooth path も描けます。まず specimen を凍結します。checkpoint、preprocessing、history、attachment point、episodes、sampling rule です。
本当のルール
複数 instrument を使い、問いを混ぜません。
- mean/coordinate variance: constant、dead axis、scale drift、reload bug。
- covariance spectrum/effective dimension: population spread がどこにあるか。低値は collapse または本当に単純な task。
- pairwise distance/duplicates: pile-up は見つけるが reachability は分からない。
- nearest neighbors: original observation、temporal distance、action、state difference、next outcome を並べる。self/adjacent duplicate は duplication test 以外で除く。
- trajectory continuity: visual change、latent change、task-variable change を同じ時間軸へ。real encoded と self-fed prediction を分ける。
doorway、wall、free motion、approach、first contact、maintained contact ごとに slice します。broad nuisance variation は missing task distinction を隠せます。
LeWM v3 は、選んだ post-hoc measure で PushT latent trajectory が training 中に straighter になり、PLDM より straight と報告します。baseline LeWM に explicit straightness loss はありません。spread、action sensitivity、matched trajectory regions と一緒に見ます。collapsed line も完全に straight です。
だまされる反例
t-SNE と UMAP は high-dimensional world を平らにする mapmaker です。seed、neighborhood setting、合理的 metric を変えてすべての map を残し、original-space neighbors を audit します。axis を “position” や “time” と名付けません。island、bridge、empty gap、area、far-cluster distance は projection artifact かもしれません。
LeWM v3 は structured PushT state grid に t-SNE を使い、qualitative sampled neighborhood pattern を報告します。UMAP は tutorial extension です。どちらも true topology、reachability、causality、understanding を示しません。
auxiliary decoder も viewer です。論文は192-dimensional final [CLS] からの lightweight transformer decoder を説明し、224-pixel/patch-16 の例では196 learned patch queries を使います。training chronology の記述は section 間で一致しません。tutorial では named checkpoint を freeze、episode split 上で後から decoder training、reconstruction を baseline training/planning 外に置きます。
sharp image が示すのは decoder-dependent reconstructability です。complete state や model use ではありません。paper の PushT reconstruction-loss ablation は tested setup で baseline より planning success が低いですが、reconstruction が常に有害という証拠ではありません。
実験のレシート
neighbor gallery/projection には sample identity、original latent distance、fit 後に加えた label、algorithm setting、seed、original observation を保存します。trajectory には door crossing、contact、camera change、action を印し、同じ reference projection を使います。
baseline の emergent post-hoc straightening と、別 method Temporal Straightening for Latent Planning を混ぜません。後者は explicit curvature regularizer、stop-gradient target、gradient-based planning を使います。元の LeWM v3 の ingredient ではありません。
source boundary は LeWorldModel v3 と frozen core commit 8edfeb3 です。frozen repository には完全な t-SNE、straightening、decoder、VoE scripts/settings がないため、exact reproduction は確立していません。
3つのクイック質問
- low effective dimension が task により healthy または harmful になるのはなぜですか。
- nearest-neighbor gallery を dynamics audit にする列は何ですか。
- sharp post-hoc decoder が示すこと、示さないことは何ですか。