コース進捗 コース目次 48レッスン中 48件を公開中
第0部 読み方ガイド:私たちは何を学ぶのか
第1部 世界モデル:エージェントの頭の中にある実験場
第2部 画面を状態に変える:LeWMのモデル構造
第3部 モデルの抜け道を防ぐ:予測損失とSIGReg
第4部 モデルを行動に使う:潜在空間での計画
第5部 エンジニアリング再現:論文から動くシステムへ
第6部 LeWMは何を学んだのか
第7部 「予測が正確」でも「計画がうまくいかない」のはなぜか
第8部 再現者から研究者へ
付録 数学・実装・再現・査読のための参照資料
まず大きな絵
- Formalterm が正確に指すもの。
- Plain持ち帰れる短いお話。
- Boundaryその語が証明しないもの。
小さなお話:翻訳ブース
ある researcher は target と言って training の encoded next observation を指します。別の人は goal と聞いて planner の image を指します。三人目は teacher forcing と聞いて、original LeWM にはない EMA teacher を発明します。
語は似ていますが別 lane に住みます。この glossary は各 term に四窓、formal definition、plain sentence、analogy、common misconception を与えます。

どの窓も他を置き換えません。analogy は記憶を助けますが、assumption や evidence を供給しません。
Model と representation の用語
| Term | Formal definition | 一言説明 | たとえ | よくある誤解 |
|---|---|---|---|---|
| World model | environment state/observation dynamics を表し、reasoning/control のため temporal/action consequence を予測する learned model。 | 「次に何が起こる?」を rehearsal します。 | rehearsal room。 | video を描くか、完全で人間可読な physics を持たねばならない。 |
| Latent state | prediction、analysis、planning 用に observation/history を圧縮する learned vector。 | model の小さな internal state card。 | 写真ではなく passport。 | 各 coordinate が物理量、または一 frame が自動的に完全 Markov state。 |
| Target representation | predictive representation loss の reference。original v3 では同じ trainable encoder が shifted real next observation を encode し、detach しない。 | prediction が合わせる encoded next observation。 | training の destination ticket。 | goal image、EMA teacher、target network、stop-gradient branch。 |
| Action conditioning | latent context からの predicted transition を action/action block で modulate すること。 | 「これをしたら?」を質問に入れます。 | route simulator の left/right button。 | 違う action symbol は常に違う state や正しい causal branch を作る。 |
| Representation collapse | constant code から task に不十分な effective structure まで、useful variation を失う degeneracy。 | 違う場面が内部で同じになります。 | 全 traveler に同じ passport。 | low-dimensional なら必ず collapse、または collapse と overfitting は同じ。 |
| SIGReg | Sketched Isotropic Gaussian Regularization。有限 random 1D projection を standard Gaussian reference と比べる anti-collapse pressure。 | latent crowd の多くの影を調べます。 | 町を多方向から検査。 | 有限 shadow が joint Gaussian、topology、physics、planning utility を保証。 |
Prediction と control-loop の用語
| Term | Formal definition | 一言説明 | たとえ | よくある誤解 |
|---|---|---|---|---|
| Prediction loss | predicted future representation と encoded observed successor の discrepancy。 | estimated next latent は real next latent に会えたか。 | forecast checkpoint と到着を比較。 | low one-step loss が long rollout/planning success を保証。 |
| Teacher forcing | model 自身の過去 prediction だけでなく recorded context を条件にする sequence training。 | training は real trail、rollout は自分の足跡から歩きます。 | 前の小節を渡す piano practice。 | teacher network、EMA、target encoder、stop-gradient が必要。 |
| Rollout | action sequence 下で dynamics を繰り返し、しばしば predicted state を後段へ渡すこと。 | future path を一歩ずつ作ります。 | 前の copy から次の copy。 | 必ず generated video、real trajectory、feasible path。 |
| Open-loop control | commitment interval 中に observation で修正せず action sequence を実行。 | もう一度見るまで plan に従います。 | 印刷 route で運転。 | model がない、または絶対に働かない。 |
| Closed-loop control | earlier action 後の observation を使って later action を選ぶこと。 | 見直して更新します。 | navigation app の reroute。 | feedback が成功を保証し wrong model/cost を直す。 |
Planning の用語
| Term | Formal definition | 一言説明 | たとえ | よくある誤解 |
|---|---|---|---|---|
| Planner | model、cost、constraint を使い、有限 budget で action sequence を提案・評価・選択する procedure。 | planner が試すものを選び、model が結果を見積もります。 | itinerary を比べる travel agent。 | planner と world model は一 component、または search が欠落情報を修復。 |
| Cost | predicted outcome/action sequence の scalar ranking rule。通常 lower-is-better。 | search にどの candidate が良く見えるか伝えます。 | judge の scorecard。 | 必ず training reward や physical feasibility。 |
| Terminal latent distance | baseline LeWM の final predicted latent と encoded goal latent 間 squared Euclidean cost。 | imagined endpoint と goal の定規。 | map 上の straight line。 | shortest path、progress、directed temporal distance、reachability と同じ。 |
| Decision-metric alignment | declared protocol 下で predicted candidate ranking と realized-outcome ranking が一致すること。 | rehearsal は reality と同じ winner を選ぶか。 | 二人の judge が同じ contestant を順位付け。 | 一 correlation、probe、one-step score が alignment を保証。 |
| Reachability | 許された dynamics/action で state に届くか。多くは方向付き time/action budget 内。 | 「この rules で着ける?」 | corridor と lock で隔てた隣室。 | symmetric、visual similarity、または learned score が証明。 |
| CEM | Cross-Entropy Method。action sequence を sample、cost で elite を残し、有限 round で proposal を refit。 | winner が次 audition を形作ります。 | 繰り返す tournament。 | 全 action を exhaust、global optimum を保証、model error を exploit できない。 |
| MPC | receding-horizon control。plan、configured prefix を execute、observe、replan。 | imagination を短く信じて reality を確認。 | sailor が各 leg を描き直す。 | 常に一 action を実行、rollout bias を除去、または learned controller そのもの。 |
Evidence と diagnosis の用語
| Term | Formal definition | 一言説明 | たとえ | よくある誤解 |
|---|---|---|---|---|
| Support mismatch | query した state/action/transition に、collected data/training distribution の十分な evidence がないこと。 | planner がほぼ未学習の neighborhood を尋ねます。 | route が未測量の空白を通る。 | outside support は physically impossible、または global action frequency が local support を証明。 |
| Model exploitation | search が systematic model/cost error、しばしば weak support のため魅力的な action を選ぶこと。 | optimizer が simulator の loophole を見つけます。 | rules の typo を最適化。 | 普通の exploration、deception、または CEM が壊れている証明。 |
| Decodability | 指定 probe/split で fixed representation から variable を recover できる度合い。 | この reader は情報を取り出せるか。 | clerk が coded passport field を読む。 | model が使う、causal、または understanding の証明。 |
| Controllability | named dynamics/horizon 下で admissible action が influence/reach できる state distinction。 | available control は何を変えられるか。 | reachable floor の elevator button。 | action-sensitive prediction が controllability/observability を確立。 |
| Identifiability | assumption と equivalence class の下で、observed data/control から latent variable/model を一意回復できること。 | allowed transform まで hidden explanation は一意か。 | 一 jurisdiction の rule で有効な ID。 | interpretability、または arbitrary world の true physics。 |
| Causal identification | experimental design / identification assumption 下で intervention effect/structure を確立。 | 一 factor を変えたから何が変わったか。 | circuit を固定して一 switch を試す。 | correlation、probe、separated branch、任意の一 ablation が cause を証明。 |
| Planning utility | declared task/planner/budget で action selection に役立つ empirical usefulness。 | この machinery はここで良い decision を助けるか。 | completed trip で map を評価。 | task-free intrinsic scalar。 |
| Action-conditioned predictive consistency | matched action-conditioned rollout の diagnostic。ACPC v1 は pairwise divergence、Invariance Radius、Separation Rate を含む。 | irrelevant change では安定し、consequential change では違うか。 | sign を塗り替える/road を動かす。 | internal consistency が truth、calibrated uncertainty、safety を証明。 |
| Independent reproduction | original author report 外の evidence が documented version/protocol で claim を rerun/reimplement。 | 別 investigator が実際に走らせた範囲を確認。 | 第二 lab が一つの named assay を反復。 | author code や一 TwoRoom reproduction が全 task/method を検証。 |
| Violation of Expectation | ordinary sequence と意図的 perturbed event sequence の model-derived surprise を matched comparison。 | prediction error は選んだ violation に反応するか。 | costume change と motion rule change。 | surprise peak が intuitive physics、causality、広い anomaly understanding を証明。 |
vocabulary を壊してみる
“The teacher forces the target goal through a rollout, so low latent distance proves the state is reachable.”
名詞は全部本物ですが、接続は全部誤りです。teacher forcing は recorded context を渡します。target representation は encoded successor で、goal は planning に属します。rollout は action 下の latent を imagine します。terminal distance は endpoint を順位付けしますが、budgeted directed reachability を証明しません。
training lane と planning lane を描き、各 word を一 object、一 phase、一 assumption set、一 evidence type に結びます。一 term が二 mechanism を同時に指すなら、流暢な prose が category error を隠しています。
証拠レシート
definition は LeWorldModel v3、frozen official code、LeJEPA v3、passive identifiability v1、controlled identifiability v2、TwoRoom reproduction v1 に従います。frontier definition の exact version は Appendix E にあります。evidence cutoff は 2026-08-20。analogy と壊れた sentence は tutorial construction、independent evidence は documented protocol の TwoRoom-only です。
3つのクイック質問
- target representation と goal image は system のどこに入りますか?
- information が decodable でも unused/non-causal でありうるのはなぜですか?
- distance、reachability、planning utility はどう違う質問をしますか?