JEPA4Japan · tutorials

Chapter 17 — CEM: Searching for Actions Through an Elimination Tournament

722 words 4 min read #LeWorldModel#World Models#JEPA

Implement iterative candidate sampling, latent simulation, elite selection, and distribution updates while balancing horizon, population, and iterations.

Course progress Course outline 48 of 48 lessons available

Part 0 — Reading Guide: What Exactly Are We Going to Learn?

  1. 01 Chapter 0 — Before You Begin available now

Part 1 — World Models: An Internal Sandbox for the Agent

  1. 02 Chapter 1 — Why an Agent Needs to “Imagine the Future” available now
  2. 03 Chapter 2 — Why Not Predict the Next Image Directly? available now
  3. 04 Chapter 3 — The JEPA Idea: Predict Meaning, Not a Replica of the Image available now
  4. 05 Chapter 4 — Understand LeWM in One Diagram available now

Part 2 — Turning Images into State: The LeWM Architecture

  1. 06 Chapter 5 — Trajectory Data: To the Model, the World Is Not an Image Collection available now
  2. 07 Chapter 6 — The Visual Encoder: Issuing a “State Passport” for Every Frame available now
  3. 08 Chapter 7 — The Dynamics Predictor: Moving Time Forward in the Mind available now
  4. 09 Chapter 8 — A Complete Forward Pass: Follow One Batch from Start to Finish available now

Part 3 — Preventing the Model from Cheating: Prediction Loss and SIGReg

  1. 10 Chapter 9 — The Most Dangerous Shortcut: Representation Collapse available now
  2. 11 Chapter 10 — Prediction Loss: How the Model Learns the Next Step available now
  3. 12 Chapter 11 — The Intuition Behind SIGReg: Letting Representation Space “Breathe” available now
  4. 13 Chapter 12 — Keep the Mathematics Minimal but Sufficient available now
  5. 14 Chapter 13 — The Original LeWM’s End-to-End Training Mechanism available now
  6. 15 Chapter 14 — Train a Model That Does Not Collapse Immediately available now

Part 4 — Putting the Model into Action: Planning in Latent Space

  1. 16 Chapter 15 — Goal-Conditioned Planning: From “Where Am I?” to “Where Do I Want to Go?” available now
  2. 17 Chapter 16 — Latent Euclidean Distance: Convenient, but Not Necessarily Reliable available now
  3. 18 Chapter 17 — CEM: Searching for Actions Through an Elimination Tournament Current lesson
  4. 19 Chapter 18 — MPC: Do Not Trust the Model for Too Long at Once available now
  5. 20 Chapter 19 — Long-Horizon Rollouts: How Small Errors Snowball into Major Failures available now
  6. 21 Chapter 20 — Implement a Minimal LeWM Planner from Scratch available now

Part 5 — Engineering Reproduction: From Paper to Running System

  1. 22 Chapter 21 — The Official Repository and Experimental Environment available now
  2. 23 Chapter 22 — First Experiment: A TwoRoom Smoke Test available now
  3. 24 Chapter 23 — Second Experiment: Reproducing PushT available now
  4. 25 Chapter 24 — How to Evaluate a World Model Fairly available now
  5. 26 Chapter 25 — Failure-Diagnosis Manual available now

Part 6 — What Has LeWM Actually Learned?

  1. 27 Chapter 26 — Linear Probes: Which Physical Variables Are Encoded in the Latent State? available now
  2. 28 Chapter 27 — Give Latent Space a “Health Check” available now
  3. 29 Chapter 28 — Violation of Expectation: Is the Model Surprised by “Impossible Events”? available now
  4. 30 Chapter 29 — How to Discuss “Understanding the World” Rigorously available now

Part 7 — Why “Accurate Prediction” Can Still Produce “Poor Planning”

  1. 31 Chapter 30 — The Gap Between the Training Objective and the Planning Objective available now
  2. 32 Chapter 31 — Global Non-Collapse Does Not Guarantee Preservation of Task-Relevant Dynamics available now
  3. 33 Chapter 32 — When Is an Isotropic Gaussian Prior Too Strong? available now
  4. 34 Chapter 33 — Long-Horizon Planning: Predict Farther or Plan More Intelligently? available now
  5. 35 Chapter 34 — From Positional Distance to Task Progress available now
  6. 36 Chapter 35 — Multi-Task Learning, Real Robots, and Visual Distractions available now
  7. 37 Chapter 36 — Theoretical Boundaries: When Can the True State Be Identified? available now

Part 8 — From Reproducer to Researcher

  1. 38 Chapter 37 — Design a Credible LeWM Improvement Experiment available now
  2. 39 Chapter 38 — Twelve Executable Research Projects available now
  3. 40 Chapter 39 — Open Questions in LeWM Research available now

Appendices

  1. 41 Appendix A — The Minimum Necessary Mathematical Toolkit available now
  2. 42 Appendix B — PyTorch Implementation Quick Reference available now
  3. 43 Appendix C — Complete Tensor-Shape Table available now
  4. 44 Appendix D — Experiment Configuration Cards available now
  5. 45 Appendix E — Paper Timeline and Evidence Levels available now
  6. 46 Appendix F — Glossary available now
  7. 47 Appendix G — Reproduction Checklist available now
  8. 48 Appendix H — Expert-Review Checklist available now

The big picture

  1. SampleDraw complete action plans.
  2. ImagineRoll each plan through the frozen model.
  3. SelectKeep the lowest-cost elites.
  4. RefitMove the proposal and repeat.
CEM spends a finite audition budget; it does not try every plan.

A CEM round samples plans, imagines them, scores them, keeps elites, and refits the proposal.

A tiny story

A director cannot watch every possible dance. Instead, 300 dancers perform complete routines. The 30 routines with the best model-predicted score influence the next audition. After several rounds, the audition concentrates near promising routines.

That is the Cross-Entropy Method (CEM) used for LeWM planning. Here “cross-entropy” names a sampling optimizer, not the classification loss. The trained encoder and predictor stay frozen. Only the probability distribution used to sample action sequences changes.

The real rule

A candidate is a whole sequence, not one greedy action. A useful conceptual shape is [B, N, H, A]: environment batch, candidate, planning horizon, and action-block width.

initialize a proposal over complete action sequences
repeat for the declared number of rounds:
  sample candidates
  roll them through the frozen world model
  compute one terminal goal cost per candidate
  keep low-cost elites
  refit the proposal from those elites
return the plan using the declared solver convention

The named LeWM setup uses 300 candidates and 30 elites. The generic paper description and frozen shared YAML contain 30 refinement rounds, while Appendix D says PushT uses at most 30 and the other environments use 10. Keep those source cards separate.

The paper describes a multivariate Gaussian proposal for its continuous-control tasks. Exact clipping, covariance, smoothing, and bound handling belong to the resolved external solver revision; the LeWM core did not pin that dependency. A categorical discrete-action solver is possible in general, but is not established for this frozen baseline.

The trick that can fool us

Imagine two equally good corridors around an obstacle: one goes left, one goes right. The elites split into two groups. Their Gaussian average points straight into the obstacle.

So record the best sampled plan and the final proposal mean separately. LeWorldModel Appendix B uses the final mean; Algorithm 2 allows the best found sequence or the first action of the final mean. Both inspected stable-worldmodel revisions—tag 0.0.5 and snapshot addbab4—return the final mean. Because LeWM did not pin its historical dependency, this does not prove the authors ran either exact revision. Executing best_seen is a useful labelled variant, not a silent replacement.

More search is not always safer. CEM can miss a narrow good corridor, collapse too early, see flat costs, average modes, saturate action bounds, or discover a loophole in the learned model more efficiently. Elite means “best in this sampled round under this model and cost,” not reachable or successful in the real environment.

Experiment receipt

Log horizon, action width and transformation, proposal seed and scale, population, elites, rounds, candidate and elite cost distributions, proposal spread by horizon position, boundary fraction, support warnings, model calls, hardware, and time. Count real action budget separately from imagined predictor work.

The source boundary is LeWorldModel v3, frozen config 8edfeb3, and the separately inspected current solver snapshot. The independent TwoRoom reimplementation compares 10 and 30 rounds in one environment; it does not recover the historical dependency or validate every benchmark.

The defensible claim is small: CEM concentrates finite samples around low-model-cost action sequences. It has no global-optimum or real-success guarantee.

Quick check

  1. What changes during CEM, and what remains frozen?
  2. Why can two good elite families produce a bad proposal mean?
  3. Which paper and code cards disagree about non-PushT refinement rounds?