コース進捗 コース目次 48レッスン中 48件を公開中
第0部 読み方ガイド:私たちは何を学ぶのか
第1部 世界モデル:エージェントの頭の中にある実験場
第2部 画面を状態に変える:LeWMのモデル構造
第3部 モデルの抜け道を防ぐ:予測損失とSIGReg
第4部 モデルを行動に使う:潜在空間での計画
第5部 エンジニアリング再現:論文から動くシステムへ
第6部 LeWMは何を学んだのか
第7部 「予測が正確」でも「計画がうまくいかない」のはなぜか
第8部 再現者から研究者へ
付録 数学・実装・再現・査読のための参照資料
まず大きな絵
- dataを見る順序、action、transform
- graphを見るshape + gradient
- 健康を見るfinite、spread、action use
- それから拡大long run前にrollout

小さなお話
run 1はtotal lossを下げましたが、全画像がほぼ同じlatentになりました。run 2はreduced precisionでnon-finite valueを出しました。run 3はmodel weightをloadしましたが、schedulerとdata orderを最初から再開しました。どのdashboardにも安心できる表示がありました。
smoke testは成功証明書ではありません。特定のwiring failureを安く取り除き、long training budgetを使う前に見つけます。
技術バックパック
まずprovenanceを記録します。code commit、resolved config、dependency version、data hash、hardware、precision、initialization、loader/split state、worker policy、SIGReg randomness、evaluation randomness、deterministic-kernel exceptionです。visible seedはsplit/loader generatorを作りますが、global determinismは示しません。重要なstable-pretrainingとstable-worldmodel revisionはinspected snapshotでpinされていません。
1 trajectoryをtransform前後で見ます。frozen defaultはhistory 3、offset 1、frame skip 5、4 observation positionsで、3 shifted pairsが寄与します。non-pixel statisticsはsplit前にloaded dataset全体から計算され、train-onlyではありません。boundary action NaNはzeroへ置換されますが、そのzeroはdata conventionで、必ずしもphysical no-opではありません。
batch sizeとprojection countは別問題を解きます。frozen batch sizeは128で、各SIGReg callはtime sliceごとにそのforward-pass populationを見ます。projectionを増やすと、同じcrowdを多directionから見るだけです。gradient accumulationでoptimizer batchを大きくしても、1 SIGReg callのpopulationが大きくなるとは限りません。
frozen training start pointはAdamW、learning rate 5e-5、weight decay 1e-3、warmup-cosine schedule、gradient clip 1.0、bfloat16 mixed precisionです。weight decayはlatent Gaussianityではなく、clippingはmissing branchを再接続しません。actual learning rate、module-level pre-clip gradient、clipping frequency、finite intermediateをlogします。
paperは10 training epochs、frozen YAMLはmax_epochs: 100を記述します。runtime overrideでrunが一致する場合はありますが、2 source factsを分けます。「epochs」だけでなくupdate countとprocessed windowsも報告します。
順序は次です。
provenance
-> raw/transformed sequence
-> forward shape + finite raw loss
-> encoder/action/predictor pathのbackward signal
-> tiny-subset fit + latent spread + action swap
-> held-out one-step + horizon別autoregressive error
-> その後にworker、epoch、precision、projection、deviceを拡大
model-weight fileでevaluationはできてもexact continuationはできない場合があります。resumeにはoptimizer、scheduler、step、precision、random、data-order stateも必要です。
だまされる仕掛け
fixed inputで4 cheap interventionsをします。action shuffle、1 action repeat、SIGRegがnoisyになるまでbatch縮小、1 bfloat16 updateと1 full-precision diagnostic updateの比較です。各probeは1 suspected mechanismだけを変え、新baselineではなくlabelled probeのままにします。
healthを1 scoreへ圧縮しません。raw prediction、raw SIGReg、weighted total、learning rate、gradient norm、non-finite count、latent moment/spectrum、action swap、held-out one-step error、recorded-action horizon rolloutをlogします。broad cloudとaction-deaf dynamics、good one-step curveとdriftは共存できます。
実験レシート
LeWorldModel v3、固定train.py、utils.py、module.pyがsource settingを裏づけます。
single-seed TwoRoom reproductionでは、dense action gathering、runtime action width、ImageNet pixel normalization、action z-scoringを加えると、同じ10-epoch budget内でpredictor lossが下降しました。またprojector BatchNorm running varianceがmoving activation scaleよりはるかに小さいとき、validation artifactが最大300×になることを報告しました。released author checkpointは最初の条件を満たさず影響を受けませんでした。これは狭いdiagnostic findingで、universal factorでもauthor training failureの証拠でもありません。
3問クイックチェック
- projectionを増やしてもbatch exampleを増やした代わりにならないのはなぜですか。
- tiny-subset fitは何を示し、何を示しませんか。
- model exportとresumable experimentを分けるstateは何ですか。