JEPA4Japan · tutorials

Chapter 6 — The Visual Encoder: Issuing a “State Passport” for Every Frame

710 words 4 min read #LeWorldModel#World Models#JEPA

Track images through patches, a Vision Transformer, the [CLS] summary, projection, normalization, and the resulting latent-vector shapes.

Course progress Course outline 48 of 48 lessons available

Part 0 — Reading Guide: What Exactly Are We Going to Learn?

  1. 01 Chapter 0 — Before You Begin available now

Part 1 — World Models: An Internal Sandbox for the Agent

  1. 02 Chapter 1 — Why an Agent Needs to “Imagine the Future” available now
  2. 03 Chapter 2 — Why Not Predict the Next Image Directly? available now
  3. 04 Chapter 3 — The JEPA Idea: Predict Meaning, Not a Replica of the Image available now
  4. 05 Chapter 4 — Understand LeWM in One Diagram available now

Part 2 — Turning Images into State: The LeWM Architecture

  1. 06 Chapter 5 — Trajectory Data: To the Model, the World Is Not an Image Collection available now
  2. 07 Chapter 6 — The Visual Encoder: Issuing a “State Passport” for Every Frame Current lesson
  3. 08 Chapter 7 — The Dynamics Predictor: Moving Time Forward in the Mind available now
  4. 09 Chapter 8 — A Complete Forward Pass: Follow One Batch from Start to Finish available now

Part 3 — Preventing the Model from Cheating: Prediction Loss and SIGReg

  1. 10 Chapter 9 — The Most Dangerous Shortcut: Representation Collapse available now
  2. 11 Chapter 10 — Prediction Loss: How the Model Learns the Next Step available now
  3. 12 Chapter 11 — The Intuition Behind SIGReg: Letting Representation Space “Breathe” available now
  4. 13 Chapter 12 — Keep the Mathematics Minimal but Sufficient available now
  5. 14 Chapter 13 — The Original LeWM’s End-to-End Training Mechanism available now
  6. 15 Chapter 14 — Train a Model That Does Not Collapse Immediately available now

Part 4 — Putting the Model into Action: Planning in Latent Space

  1. 16 Chapter 15 — Goal-Conditioned Planning: From “Where Am I?” to “Where Do I Want to Go?” available now
  2. 17 Chapter 16 — Latent Euclidean Distance: Convenient, but Not Necessarily Reliable available now
  3. 18 Chapter 17 — CEM: Searching for Actions Through an Elimination Tournament available now
  4. 19 Chapter 18 — MPC: Do Not Trust the Model for Too Long at Once available now
  5. 20 Chapter 19 — Long-Horizon Rollouts: How Small Errors Snowball into Major Failures available now
  6. 21 Chapter 20 — Implement a Minimal LeWM Planner from Scratch available now

Part 5 — Engineering Reproduction: From Paper to Running System

  1. 22 Chapter 21 — The Official Repository and Experimental Environment available now
  2. 23 Chapter 22 — First Experiment: A TwoRoom Smoke Test available now
  3. 24 Chapter 23 — Second Experiment: Reproducing PushT available now
  4. 25 Chapter 24 — How to Evaluate a World Model Fairly available now
  5. 26 Chapter 25 — Failure-Diagnosis Manual available now

Part 6 — What Has LeWM Actually Learned?

  1. 27 Chapter 26 — Linear Probes: Which Physical Variables Are Encoded in the Latent State? available now
  2. 28 Chapter 27 — Give Latent Space a “Health Check” available now
  3. 29 Chapter 28 — Violation of Expectation: Is the Model Surprised by “Impossible Events”? available now
  4. 30 Chapter 29 — How to Discuss “Understanding the World” Rigorously available now

Part 7 — Why “Accurate Prediction” Can Still Produce “Poor Planning”

  1. 31 Chapter 30 — The Gap Between the Training Objective and the Planning Objective available now
  2. 32 Chapter 31 — Global Non-Collapse Does Not Guarantee Preservation of Task-Relevant Dynamics available now
  3. 33 Chapter 32 — When Is an Isotropic Gaussian Prior Too Strong? available now
  4. 34 Chapter 33 — Long-Horizon Planning: Predict Farther or Plan More Intelligently? available now
  5. 35 Chapter 34 — From Positional Distance to Task Progress available now
  6. 36 Chapter 35 — Multi-Task Learning, Real Robots, and Visual Distractions available now
  7. 37 Chapter 36 — Theoretical Boundaries: When Can the True State Be Identified? available now

Part 8 — From Reproducer to Researcher

  1. 38 Chapter 37 — Design a Credible LeWM Improvement Experiment available now
  2. 39 Chapter 38 — Twelve Executable Research Projects available now
  3. 40 Chapter 39 — Open Questions in LeWM Research available now

Appendices

  1. 41 Appendix A — The Minimum Necessary Mathematical Toolkit available now
  2. 42 Appendix B — PyTorch Implementation Quick Reference available now
  3. 43 Appendix C — Complete Tensor-Shape Table available now
  4. 44 Appendix D — Experiment Configuration Cards available now
  5. 45 Appendix E — Paper Timeline and Evidence Levels available now
  6. 46 Appendix F — Glossary available now
  7. 47 Appendix G — Reproduction Checklist available now
  8. 48 Appendix H — Expert-Review Checklist available now

The big picture

  1. Frame224 × 224 pixels
  2. Patches256 visual pieces
  3. [CLS]one summary carrier
  4. Passport192-number latent
The encoder gives each frame a compact passport. It does not certify what the passport remembers.

A frame becomes patches, a CLS summary, and one projected latent passport.

A tiny story

A camera frame is too large to carry through hundreds of imagined action plans. LeWM therefore gives each frame a small “state passport.” The predictor can move this passport through time much more cheaply than repainting the whole scene.

A real passport works because people agreed which fields matter. LeWM receives no such list. Its training pressures decide what survives. It may keep the doorway, or keep an easy wall color that merely correlates with the doorway.

The technical backpack

The frozen entry point first applies ImageNet-statistics preprocessing and resizes to 224×224. The default ViT-Tiny divides the image into 14×14 patches. Because 224 / 14 = 16, there are 16×16 = 256 patch tokens. Add one learned [CLS] token and the sequence length becomes 257.

The default backbone has width 192, 12 Transformer layers, and 3 attention heads. Attention lets spatial pieces exchange information. A patch is not an independently recognized object, and [CLS] is not magically semantic; it becomes useful only through the objective.

The visual encoder handles frames independently. It flattens batch and time, processes B×T images with shared weights, and then restores time. Temporal reasoning belongs to the later dynamics predictor.

[B,T,3,224,224]
  -> [B×T,3,224,224]
  -> [B×T,257,192]
  -> take [CLS]: [B×T,192]
  -> projector: [B×T,192]
  -> restore time: [B,T,192]

The projector is not a simple dimension reducer. Frozen code maps 192 → 2048 → 192 with BatchNorm and GELU between linear layers. The paper’s “1-layer MLP” and the code’s two linear maps are safely described as a one-hidden-layer projection MLP. Its output is the latent used by prediction and SIGReg.

LayerNorm normalizes features within one token. Projector BatchNorm uses statistics across flattened batch-time examples for each feature. SIGReg is a separate population loss. “Normalization makes it Gaussian” wrongly merges three mechanisms.

Original v3 sets pretrained: false: ViT, projector, action path, and dynamics learn together from offline trajectories.

The trick that fools us

Suppose wall color perfectly predicts which door is open in training. The encoder can store the color and ignore the tiny door gap. Repaint the wall without moving the door, and the passport changes too much. Move the door while keeping the color, and it changes too little.

Test three controlled pairs: vary texture with fixed geometry; vary geometry with matched texture; and swap only the nuisance while the frozen planner runs. A compact code, nonzero variance, or a good reconstruction does not prove the correct fact was retained.

Mechanical assertions still matter. Do not turn time into channels, forget to restore the time axis, or pass patch-level output where one state per frame is expected. Shapes can reveal a wrong implementation, though never a good representation by themselves.

Evidence receipt

LeWorldModel v3, frozen jepa.py, model configuration, and the frozen projection module support the exact default path. The paper also reports a ResNet-18 ablation, so ViT-Tiny is a baseline choice, not LeWM’s definition.

Allowed claim: [CLS] plus the projector is the latent carrier optimized by LeWM. Do not call it a complete world state, say BatchNorm imposes SIGReg’s Gaussian target, or infer physics from compactness.

Quick check

  1. Why do 224-pixel sides and patch size 14 produce 256 patches?
  2. Which component handles time after frames are encoded independently?
  3. How would you expose a passport that follows wall color instead of doorway geometry?