JEPA4Japan · tutorials

Chapter 22 — Ten long walks become a training set

844 words 4 min read #LeVJEPA#JEPA#self-supervised video#SIGReg

Follow Walking Tours through download, 15 fps extraction, Lance storage, episode boundaries, random clips, and data costs.

Course progress Course outline 34 of 34 lessons available

Part 0 — Get the map

  1. 01 Chapter 0 — Before you begin: what this course promises available now
  2. 02 Chapter 1 — One video, two windows available now
  3. 03 Chapter 2 — A walk along Yann LeCun’s research road available now
  4. 04 Chapter 3 — The JEPA family, without the name soup available now

Part 1 — Why a small objective can learn to see

  1. 05 Chapter 4 — Video can set its own homework available now
  2. 06 Chapter 5 — Keep the meaning; do not repaint every pixel available now
  3. 07 Chapter 6 — Match the cards, but do not leave every card blank available now
  4. 08 Chapter 7 — SIGReg checks a cloud by looking at its shadows available now
  5. 09 Chapter 8 — The whole LeVJEPA objective on one line available now

Part 2 — Send a video through one encoder

  1. 10 Chapter 9 — How global and local views are paired available now
  2. 11 Chapter 10 — Cut a video into space-time tiles available now
  3. 12 Chapter 11 — One encoder, one projector, one summary card available now
  4. 13 Chapter 12 — One complete trip through the model available now
  5. 14 Chapter 13 — Why throwing away 95% can help available now
  6. 15 Chapter 14 — Same-frame teamwork, no peeking into tomorrow available now
  7. 16 Chapter 15 — RoPE, single-frame tubelets, and unexpectedly useful patch features available now

Part 3 — Read the experiments, not just the headline

  1. 17 Chapter 16 — What the four ablation ladders actually test available now
  2. 18 Chapter 17 — Equal epochs are not equal bills available now
  3. 19 Chapter 18 — What ImageNet, K400, and SSv2 are really asking available now
  4. 20 Chapter 19 — Keep the paper’s results in a ledger available now
  5. 21 Chapter 20 — Claims the evidence does not yet earn available now

Part 4 — From the official repository to your own experiment

  1. 22 Chapter 21 — A map of the official repository available now
  2. 23 Chapter 22 — Ten long walks become a training set Current lesson
  3. 24 Chapter 23 — Read the defaults, then start training available now
  4. 25 Chapter 24 — Run a smoke test that cannot flatter you available now
  5. 26 Chapter 25 — Skip training: extract features from the public checkpoint available now
  6. 27 Chapter 26 — Freeze the encoder and test your own videos available now

Part 5 — Put the representation back on the world-model road

  1. 28 Chapter 27 — The important boundary: an encoder is not a planner available now
  2. 29 Chapter 28 — How LeVJEPA might feed a future world model available now
  3. 30 Chapter 29 — Ten projects, from first experiment to paper-sized question available now

Appendices — A backpack for the trail

  1. 31 Appendix A — The smallest useful math kit available now
  2. 32 Appendix B — The complete tensor-shape table available now
  3. 33 Appendix C — Glossary and paper timeline available now
  4. 34 Appendix D — Reproduction and review checklist available now

Ten long walks become short training clips

  1. 10 linkslong city walks
  2. Downloadabout 25 GB at 720p
  3. Sample frames15 fps, short side 384
  4. Store in Lanceabout 33 GB
  5. Draw a clip16 frames, about 2.1 s
A long web video does not go straight into the model. It becomes a bounded frame store, then training samples one short interval.

Ten thick travel albums are awkward to lift onto a desk for every question. A better archive keeps photographs in order, stamps every 128 with one journey ID, then fetches 16 at a time without crossing into the next journey. That is roughly the job of the Walking Tours pipeline.

From URLs to Lance rows

The Walking Tours dataset publishes YouTube URLs, not video binaries. The pinned script downloads ten first-person videos, each roughly one to seven hours, preferring 720p60 without audio. The stated download is about 25 GB. Web sources may disappear or be region-blocked, so check terms and research use before downloading. Leave at least 70 GB for downloads plus the roughly 33 GB Lance store.

uv sync --extra data
bash scripts/download_walking_tours.sh
uv run python scripts/build_lance_walking_tours.py --workers 16

The builder derives a sampling step from each source’s true frame rate: nine are about 60 fps, while Wildlife is about 30 fps. All become roughly 15 fps. Frames are resized to short side 384 and saved at JPEG quality 90. Consecutive groups of 128 frames form about 8.5-second episodes.

Each Lance row stores episode_idx:int32, step_idx:int32, frame:binary, h/w:int16, and label:int16; because the set is unlabeled, label=-1. Rows must remain ordered by episode or the loader will infer bad boundaries.

Only complete 128-frame episodes are written. A short tail at the end of a source is omitted. A decoding failure increments failed and skips that episode; successful episodes are then numbered continuously. A zero process exit is therefore not enough. Save source URLs, downloaded sizes, observed fps, per-video episode counts, and failed. If a web video later changes, this log becomes the data fingerprint.

The training default draws num_frames=16 at frame_stride=2 from the 15 fps store: effective 7.5 fps over about 2.1 seconds. A start must remain inside one episode and only selected rows are read. The config comment records 6,067 episodes in the authors’ build; web sources drift, so your builder output—not that number—is authoritative.

Indices are start + [0,2,4,…,30]. The implementation uses span=16×2=32 to test length. With pad_short=false, short episodes are excluded; if padding is enabled, the final frame repeats. clips_per_video=200 samples each valid episode 200 times per epoch. These are repeated draws, not 200 new independent videos. Neighboring clips remain highly correlated, so downstream supervised splits must group by source.

Point to external data without editing code:

LEVJEPA_DATA_ROOT=/mnt/levjepa-data uv run python main.py --cfg job

The loader looks for $LEVJEPA_DATA_ROOT/walking_tours/train.lance. The builder refuses to overwrite an existing output. Keeping the old archive and selecting a new --out path is safer than deleting blindly.

Inspect the store before training

uv run python - <<'PY'
import lance
ds = lance.dataset("data/walking_tours/train.lance")
print(ds.count_rows())
print(ds.schema)
PY

Expect a positive row count and all six fields. Also require failed == 0 in the build log and multiple episode_idx values. One enormous episode signals broken indexing or ordering.

Data provenance

Three data facts

  1. Walking Tours supplies URLs; the script retrieves about 25 GB of video.
  2. The builder stores 15 fps, 128-frame Lance episodes; training draws 16 frames at stride 2.
  3. Roughly 33 GB and 6,067 episodes describe the authors’ snapshot, not an eternal guarantee.

Archive check

  1. Why not use one fixed frame step for Wildlife and the 60 fps videos?
  2. Roughly how much time does the default 16-frame clip cover?
  3. Why must Lance rows not be shuffled before writing?
Answers
  1. Wildlife is about 30 fps, so the same step would halve its stored sampling rate.
  2. About 2.1 seconds.
  3. The loader detects boundaries from runs of episode_idx; disorder would merge or split episodes incorrectly.