JEPA4Japan · tutorials

Chapter 35 — Multi-Task Learning, Real Robots, and Visual Distractions

855 words 4 min read #LeWorldModel#World Models#JEPA

Examine task identity, visual variation, latent drift, auxiliary depth, predictor capacity, partial observability, robot noise, and the simulation-to-reality evidence gap.

Course progress Course outline 48 of 48 lessons available

Part 0 — Reading Guide: What Exactly Are We Going to Learn?

  1. 01 Chapter 0 — Before You Begin available now

Part 1 — World Models: An Internal Sandbox for the Agent

  1. 02 Chapter 1 — Why an Agent Needs to “Imagine the Future” available now
  2. 03 Chapter 2 — Why Not Predict the Next Image Directly? available now
  3. 04 Chapter 3 — The JEPA Idea: Predict Meaning, Not a Replica of the Image available now
  4. 05 Chapter 4 — Understand LeWM in One Diagram available now

Part 2 — Turning Images into State: The LeWM Architecture

  1. 06 Chapter 5 — Trajectory Data: To the Model, the World Is Not an Image Collection available now
  2. 07 Chapter 6 — The Visual Encoder: Issuing a “State Passport” for Every Frame available now
  3. 08 Chapter 7 — The Dynamics Predictor: Moving Time Forward in the Mind available now
  4. 09 Chapter 8 — A Complete Forward Pass: Follow One Batch from Start to Finish available now

Part 3 — Preventing the Model from Cheating: Prediction Loss and SIGReg

  1. 10 Chapter 9 — The Most Dangerous Shortcut: Representation Collapse available now
  2. 11 Chapter 10 — Prediction Loss: How the Model Learns the Next Step available now
  3. 12 Chapter 11 — The Intuition Behind SIGReg: Letting Representation Space “Breathe” available now
  4. 13 Chapter 12 — Keep the Mathematics Minimal but Sufficient available now
  5. 14 Chapter 13 — The Original LeWM’s End-to-End Training Mechanism available now
  6. 15 Chapter 14 — Train a Model That Does Not Collapse Immediately available now

Part 4 — Putting the Model into Action: Planning in Latent Space

  1. 16 Chapter 15 — Goal-Conditioned Planning: From “Where Am I?” to “Where Do I Want to Go?” available now
  2. 17 Chapter 16 — Latent Euclidean Distance: Convenient, but Not Necessarily Reliable available now
  3. 18 Chapter 17 — CEM: Searching for Actions Through an Elimination Tournament available now
  4. 19 Chapter 18 — MPC: Do Not Trust the Model for Too Long at Once available now
  5. 20 Chapter 19 — Long-Horizon Rollouts: How Small Errors Snowball into Major Failures available now
  6. 21 Chapter 20 — Implement a Minimal LeWM Planner from Scratch available now

Part 5 — Engineering Reproduction: From Paper to Running System

  1. 22 Chapter 21 — The Official Repository and Experimental Environment available now
  2. 23 Chapter 22 — First Experiment: A TwoRoom Smoke Test available now
  3. 24 Chapter 23 — Second Experiment: Reproducing PushT available now
  4. 25 Chapter 24 — How to Evaluate a World Model Fairly available now
  5. 26 Chapter 25 — Failure-Diagnosis Manual available now

Part 6 — What Has LeWM Actually Learned?

  1. 27 Chapter 26 — Linear Probes: Which Physical Variables Are Encoded in the Latent State? available now
  2. 28 Chapter 27 — Give Latent Space a “Health Check” available now
  3. 29 Chapter 28 — Violation of Expectation: Is the Model Surprised by “Impossible Events”? available now
  4. 30 Chapter 29 — How to Discuss “Understanding the World” Rigorously available now

Part 7 — Why “Accurate Prediction” Can Still Produce “Poor Planning”

  1. 31 Chapter 30 — The Gap Between the Training Objective and the Planning Objective available now
  2. 32 Chapter 31 — Global Non-Collapse Does Not Guarantee Preservation of Task-Relevant Dynamics available now
  3. 33 Chapter 32 — When Is an Isotropic Gaussian Prior Too Strong? available now
  4. 34 Chapter 33 — Long-Horizon Planning: Predict Farther or Plan More Intelligently? available now
  5. 35 Chapter 34 — From Positional Distance to Task Progress available now
  6. 36 Chapter 35 — Multi-Task Learning, Real Robots, and Visual Distractions Current lesson
  7. 37 Chapter 36 — Theoretical Boundaries: When Can the True State Be Identified? available now

Part 8 — From Reproducer to Researcher

  1. 38 Chapter 37 — Design a Credible LeWM Improvement Experiment available now
  2. 39 Chapter 38 — Twelve Executable Research Projects available now
  3. 40 Chapter 39 — Open Questions in LeWM Research available now

Appendices

  1. 41 Appendix A — The Minimum Necessary Mathematical Toolkit available now
  2. 42 Appendix B — PyTorch Implementation Quick Reference available now
  3. 43 Appendix C — Complete Tensor-Shape Table available now
  4. 44 Appendix D — Experiment Configuration Cards available now
  5. 45 Appendix E — Paper Timeline and Evidence Levels available now
  6. 46 Appendix F — Glossary available now
  7. 47 Appendix G — Reproduction Checklist available now
  8. 48 Appendix H — Expert-Review Checklist available now

The big picture

  1. New sceneryPixels changed; physics may not.
  2. New playThe same picture may need another goal.
  3. Real stageNoise, delay, contact, and safety arrive.
Appearance, task identity, hidden state, and hardware shift are different problems. A colorful dataset does not separate them for us.

A tiny story: the robot goes on tour

A robot rehearses one play: under a warm lamp, the blue cup always goes left. On tour the curtain turns red, the camera moves, and another play asks the identical cup to go right. Then a hand hides the moment of contact.

Was the change harmless scenery, a different task, a different physical state, or missing evidence? Calling all four “domain shift” erases the answer the model needs. Multi-task training can expose shortcuts, but it can also create new ones: each task may have its own tablecloth, camera, or episode length.

A controlled theater keeps robot and object state fixed while changing one visual factor, then changes physics in a separate row.

The matrix changes one known factor at a time. It does not declare every background, light, or viewpoint irrelevant.

The technical backpack

Use paired controls. Replay the same simulator state and action while changing one texture, light, or camera variable known not to alter dynamics; then include a positive control that really changes future consequences. Invariance to everything is collapse. Sensitivity to everything is an appearance shortcut. The desired pattern is selective and must also survive closed-loop planning.

If one image belongs to two tasks with different correct actions, no current-frame encoder can guess an omitted task. Supply a goal image, token, or language condition; or use history when the missing variable leaves a trace. A token can become a dataset-ID shortcut, and history cannot recover information never observed.

Later evidence must keep its identity:

Later routeExtra signal or settingNarrow boundary
TC-LeWM v2temporally centered SIGReg; multi-task LIBERO with a from-scratch multi-view ViT-S/16 and task-conditioned flow-matching behavior cloningnot baseline CEM/MPC, TwoRoom, or robot deployment
Depth-regularized JEPA v1depth supervision and training-only overparameterization; reported 18M model on real agricultural-robot videooffline visual-odometry probes, surprise, and latent rollouts—not online planning or safety
PhyLatent v1simulator physical-state supervisionprivileged training, not pixels-only recovery
PSG-JEPA v1training-only proprioceptive-state and multi-horizon joint-change groundingauthor-reported Mobile ALOHA downstream policy evidence, not original-LeWM MPC or a safety certificate
stable-worldmodel v1shared data, training, solver, MPC, and evaluation plumbinginfrastructure can expose choices; it cannot certify neutral or independently reproduced science

Original LeWM v3 remains fully end-to-end: no stop-gradient target, EMA teacher, or pretrained visual encoder. Later auxiliary heads must not redraw its graph. A larger predictor is also only a hypothesis: compare parameter- and compute-matched widths, several seeds, one-step prediction, self-fed rollout, action sensitivity, nuisance invariance, and planning.

Try to break the idea

Make two episodes whose current image and action history are identical. Task A says “cup left”; Task B says “cup right.” Hide the instruction. Consistent choice of one side reveals a dataset prior, not recovered context. Then add, one at a time, a goal image, task token, earlier instruction frame, and a background correlated with task during training but swapped at test.

For deployment, climb an evidence ladder: original simulator → controlled visual changes → sensor noise/missing frames → delayed and imperfect actuation → hardware/object/calibration variation → safety-constrained runs with abort and recovery rules. Passing one rung does not grant the next. MPC feedback limits open-loop trust but does not make the executed step safe.

Experiment receipt and evidence boundary

  • Sources: LeWorldModel v3; TC-LeWM v2; depth-regularized JEPA v1; PhyLatent v1; PSG-JEPA v1 and its project page; stable-worldmodel v1.
  • Frozen code: LeWM 8edfeb3; PSG-JEPA 8de96b8, dated 2026-08-14; stable-worldmodel addbab4, dated 2026-08-18, while formal release remained 0.1.1 from 2026-06-06. No author-linked code was identified for TC-LeWM, the depth paper, or PhyLatent.
  • Dates: depth 2026-07-15; TC-LeWM v2 2026-07-31; PhyLatent 2026-08-06; PSG-JEPA 2026-08-07; evidence cutoff 2026-08-20, Asia/Tokyo.
  • These later outcomes are author-reported; independent reproduction is not established for them or their robot results. Report the exact method, controller, data, intervention, and safety protocol. “Real video” is not “online robot planning,” and “robot policy experiment” is not “safe deployment.”

Three quick questions

  1. Why is a camera change not automatically a nuisance?
  2. What does the identical-frame, opposite-task example prove about missing context?
  3. Which new evidence is required before a simulator robustness claim becomes a robot safety claim?