JEPA4Japan · チュートリアル

第4章:4枚の地図と1台のカメラ

4,560文字 15分で読めます #Canvas#Frontend Engineering#Infinite Canvas#ELI5

Local、Parent、World、Viewport、Screen座標を統一し、可逆なCamera、Zoom、Rotation、Minimapを完成させます。

コース進捗 コース目次 18レッスン中 18件を公開中

第I部:描く前に描画面を選ぶ——プロダクト、ピクセル、座標

  1. 01 第1章:まだ描かない——Canvasはプロダクト設計ではない 公開中
  2. 02 第2章:すぐに記憶を失うピクセルの紙 公開中
  3. 03 第3章:お絵描きを再現可能なレシピにする 公開中
  4. 04 第4章:4枚の地図と1台のカメラ 現在のレッスン

第II部:ピクセル世界に頭脳を与える——モデル、スケジューリング、入力、ツール

  1. 05 第5章:ピクセル世界に台帳を作る 公開中
  2. 06 第6章:ランプが点いたときだけ描き直す——Render SchedulerとReactの境界 公開中
  3. 07 第7章:マウス、指、ペンに同じ言葉を話してもらう 公開中
  4. 08 第8章:細い縁を調べる前に大きな箱を探す 公開中
  5. 09 第9章:ツールは信号機であり、Booleanの袋ではない 公開中

第III部:「ドラッグできる」から「信頼できる」へ——操作、文字、Asset、復旧

  1. 10 第10章:触って気持ちよいエディターにする 公開中
  2. 11 第11章:描かれた文字は編集できる文字ではない 公開中
  3. 12 第12章:借りた画像を勝手に箱へ詰めてはいけない 公開中
  4. 13 第13章:タイムマシンと古い箱 公開中
  5. 14 第14章:正しく見えることと、本当に正しいことは違う 公開中

第IV部:マスターの判断——Performance、Worker、GPU、SDK、共同編集、AI

  1. 15 第15章:1万人を一人ずつ探さない 公開中
  2. 16 第16章:受付を厨房へ入れない——WorkerとGPUへの更新 公開中
  3. 17 第17章:車を自作するか、実績あるシャーシを買うか 公開中
  4. 18 第18章:人とAIが同じ台帳を編集する 公開中

5歳児にもわかるゲームから始めよう

熊のぬいぐるみを靴箱へ入れ、箱の底の左上から(10, 5)cmの位置へ置きます。次に靴箱を部屋の床の(100, 40)へ置き、4分の1回転させます。Cameraを部屋へ向け、最後にPhoneの写真上で熊の鼻をTapします。

まず予想してください。Phone上の(220, 90)を、そのまま靴箱内の(220, 90)として使えるでしょうか。靴箱を回した後、「箱の右へ10cm」は、まだ「部屋の右へ10cm」でしょうか。写真を2倍に拡大すると、熊は部屋の中で本当に動いたのでしょうか。

  1. 箱の中の位置Local
  2. 部屋の中の箱Parent / World
  3. Cameraの映像Viewport
  4. Screen上のTapScreen
同じ熊に複数組の数値があります。座標を言う前に、どの地図かを示します。

本章で伝える唯一の真実は、Canvas操作のBugの大半は、本質的にはCoordinate Spaceの混乱であるということです。すべてのPointに「どの地図にあるか」を付け、すべての変換に戻り道を用意します。

おもちゃをCanvasに翻訳する

おもちゃの世界の地図Engineering座標例
熊自身の輪郭Local CoordinatesShape原点に対するResize handle
靴箱の中Parent CoordinatesGroupに対する子Shape
部屋全体World CoordinatesDocument内のStableなScene位置
CameraのFrameViewport CoordinatesCanvas Content領域の左上を原点にする
PhoneのGlass/PageScreen CoordinatesPointerのclientX/clientY
移動、箱の回転、拡大Affine TransformTranslation、Rotation、Scale
逆にたどるInverse MatrixScreen→WorldのHit Testing

比喩には限界があります。現実のCameraには通常Perspectiveがありますが、本章のCanvas Labは2次元Affine Cameraを使い、遠近による大きさの変化やDepthはありません。ViewportとScreenの間には、Page Scroll、CanvasのgetBoundingClientRect() Offset、CSS Transformが入る場合もあります。DOMMatrixの乗算方向はTestで証明し、言葉を読む順序から推測してはいけません。

まず誤った直感を捨てる

  • 「Pointer座標がそのままWorld座標だ」。clientX/YはViewport/Screenに属します。CameraをPan、Zoom、Rotationした瞬間に誤ります。
  • 「Zoomはscaleだけ変えればよい」。Zoomだけを変えるとCamera Centerを中心に拡大します。Pointerを中心にZoomするには、その下のWorld Pointが前後で同じでなければなりません。
  • 「Matrix乗算は交換できる」。Translation後のRotationと、Rotation後のTranslationは通常異なる結果になります。
  • 「Boundsは幅と高さにScaleを掛けるだけだ」。回転したRectangleのWorld AABBは4つのCornerを変換して極値を取ります。負のScaleでは向きも反転します。
  • 「Floating-pointの等価比較は===でよい」。Transformを繰り返すと微小な誤差が蓄積するため、Scaleに適したepsilonを使います。
  • 「Zoomは0まで下げられる」。Scaleが0になるとMatrixはInverseを持ちません。極端なScaleもPrecisionとUsabilityを損ないます。

本番用バックパック

前提となる契約

前提は第3章の再生可能なRendererです。次の規則を定めます。ShapeのLocal TransformはParentへ合成し、Scene Graphを祖先方向へたどってWorldになります。Camera State {x,y,zoom,rotation}のx,yはViewport Centerに対応するWorld Pointです。Viewport座標はCanvasのCSS PixelでDPRを含まず、DPRは第2章のHostが扱います。ScreenにはPointerのclientX/Yを使います。

正式な知識

2次元Affine TransformはTranslation、Scale、Rotation、Shearを表せます。Homogeneous Matrixによってそれらを合成できます。Canvas/DOMMatrixの2D成分は一般に(a,b,c,d,e,f)と表し、Pointをx'=ax+cy+e, y'=bx+dy+fで変換します。合成順序は契約です。本章ではtranslate(viewCenter) → rotate(camera.rotation) → scale(camera.zoom) → translate(-cameraCenter)でWorld→Viewportを構築し、実際の挙動をTestで固定します。ここでcamera.rotationは「World ContentがViewport内でどれだけ回転するか」と明示的に定義します。Productが現実のCameraの向きを保存するなら、ここへ代入する前に符号を反転します。2つの符号規約を混ぜてはいけません。

Local→ParentはShape自身のTransformです。Parent→WorldはAncestor Chainを乗算します。Group座標をあらかじめ子へBakeしながらParent Transformも残すと、2回適用されるため禁止します。World→ViewportはCamera Matrixが行い、Viewport→ScreenはCanvas DOMRect左上のOffsetを加えます。逆経路はMatrix inverseを使います。Determinantが0に近い、または結果が非有限なら明示的に失敗させ、正常に見える(0,0)を返してはいけません。

Zoom Around PointerのInvariantは次のとおりです。Zoom変更前にPointerのViewport PointをInverse TransformしてworldBeforeを得ます。新しいScaleをLimit内へClampした後、もう一度Inverse TransformしてworldAfterを得ます。Camera CenterへworldBefore-worldAfterを加えます。Scale Limitsは、たとえば0.05≤zoom≤32というProduct Constantです。Camera BoundsはCamera CenterまたはViewport World Boundsを制約します。Cameraが回転している場合、Viewportの4 CornerをInverse TransformしてBounding Polygonを作り、そのAABBを取ります。Zoomで割るだけではいけません。

Viewport World BoundsとTransformed Boundsはどちらも、「必要なCornerをすべて変換してからmin/maxを求める」という規則に従います。Minimapは別のCameraです。Document World Boundsを小さなRectangleへ比例Scaleし、主CameraのWorld Cornerを描きます。Infinite Gridは無限本のLineを保存しません。現在のViewport World Boundsから、最初から最後までの可視Grid Lineだけを計算します。Zoomが小さいときはLODを使えますが、本章ではまず正しさを保証します。

Precision戦略には、Zoomと座標の絶対値を制限すること、変換済みのInteraction結果を次のSourceとして繰り返し使わないこと、DocumentがAuthorityを持つLocal/World値を保存すること、almostEqual(a,b,epsilon)で比較することが含まれます。一般的なepsilonには1e-9 × max(1, |a|, |b|)を使えます。Visual Hit Toleranceは別の問題であり、後章ではScreen Pixelで定義します。

根拠と互換性

上記Sourceの確認日は2026-08-29です。DOMMatrixは現代的なBrowserで利用できますが、Node Unit Test環境には完全な実装がない場合があります。Geometry Coreには同じ式のPure TypeScript Matrix2Dを使い、Browser AdapterでDOMMatrixとのCross Testを行えます。Canvas ElementにCSS transformを適用した場合、DOMRect左上を引くだけでは足りず、CSS TransformもInverseにする必要があります。本TutorialのBaselineでは、Host ElementへのCSS rotation/scaleを禁止します。

この章のエンジニアリング増分

開始点: 第3章のCameraにはTranslationとZoomだけがあり、変換がRendererへ散らばっています。到達点: 単一Camera APIがLocal/Parent/World/Viewport/Screenの経路、Pointer中心Zoom、Rotation、Bounds、Minimap、Infinite Gridを提供します。

canvas-lab/src/lab/ch04/
  matrix2d.ts
  camera.ts
  camera.test.ts
  camera-demo.ts

完全なDOM非依存Math CoreとCameraです。

export type Point = Readonly<{ x: number; y: number }>;
export type Rect = Readonly<{ x: number; y: number; width: number; height: number }>;

export class Matrix2D {
  constructor(
    readonly a = 1,
    readonly b = 0,
    readonly c = 0,
    readonly d = 1,
    readonly e = 0,
    readonly f = 0,
  ) {}
  multiply(n: Matrix2D): Matrix2D {
    return new Matrix2D(
      this.a * n.a + this.c * n.b,
      this.b * n.a + this.d * n.b,
      this.a * n.c + this.c * n.d,
      this.b * n.c + this.d * n.d,
      this.a * n.e + this.c * n.f + this.e,
      this.b * n.e + this.d * n.f + this.f,
    );
  }
  transformPoint(p: Point): Point {
    return { x: this.a * p.x + this.c * p.y + this.e, y: this.b * p.x + this.d * p.y + this.f };
  }
  inverse(epsilon = 1e-12): Matrix2D {
    const det = this.a * this.d - this.b * this.c;
    if (!Number.isFinite(det) || Math.abs(det) <= epsilon)
      throw new Error('Matrix is not invertible');
    return new Matrix2D(
      this.d / det,
      -this.b / det,
      -this.c / det,
      this.a / det,
      (this.c * this.f - this.d * this.e) / det,
      (this.b * this.e - this.a * this.f) / det,
    );
  }
  static translation(x: number, y: number): Matrix2D {
    return new Matrix2D(1, 0, 0, 1, x, y);
  }
  static scale(value: number): Matrix2D {
    return new Matrix2D(value, 0, 0, value, 0, 0);
  }
  static rotation(radians: number): Matrix2D {
    const c = Math.cos(radians),
      s = Math.sin(radians);
    return new Matrix2D(c, s, -s, c, 0, 0);
  }
}

export type CameraState = { x: number; y: number; zoom: number; rotation: number };
export class Camera {
  private _state: CameraState;
  constructor(
    state: CameraState,
    public viewport: { width: number; height: number },
  ) {
    this.assertState(state);
    this._state = { ...state };
  }
  get state(): Readonly<CameraState> {
    return this._state;
  }
  trySetState(next: CameraState): boolean {
    try {
      this.assertState(next);
      this._state = { ...next };
      return true;
    } catch {
      return false;
    }
  }
  worldToViewportMatrix(): Matrix2D {
    const s = this._state;
    return Matrix2D.translation(this.viewport.width / 2, this.viewport.height / 2)
      .multiply(Matrix2D.rotation(s.rotation))
      .multiply(Matrix2D.scale(s.zoom))
      .multiply(Matrix2D.translation(-s.x, -s.y));
  }
  worldToViewport(point: Point): Point {
    return this.worldToViewportMatrix().transformPoint(point);
  }
  viewportToWorld(point: Point): Point {
    return this.worldToViewportMatrix().inverse().transformPoint(point);
  }
  worldToScreen(point: Point, canvasRect: Pick<DOMRect, 'left' | 'top'>): Point {
    const p = this.worldToViewport(point);
    return { x: p.x + canvasRect.left, y: p.y + canvasRect.top };
  }
  screenToWorld(point: Point, canvasRect: Pick<DOMRect, 'left' | 'top'>): Point {
    return this.viewportToWorld({ x: point.x - canvasRect.left, y: point.y - canvasRect.top });
  }
  panByViewportDelta(dx: number, dy: number): void {
    const invRotation = Matrix2D.rotation(-this._state.rotation);
    const worldDelta = invRotation.transformPoint({
      x: dx / this._state.zoom,
      y: dy / this._state.zoom,
    });
    this._state = {
      ...this._state,
      x: this._state.x - worldDelta.x,
      y: this._state.y - worldDelta.y,
    };
  }
  zoomAround(pointer: Point, factor: number): void {
    const before = this.viewportToWorld(pointer);
    const zoom = Math.min(32, Math.max(0.05, this._state.zoom * factor));
    this._state = { ...this._state, zoom };
    const after = this.viewportToWorld(pointer);
    this._state = {
      ...this._state,
      x: this._state.x + before.x - after.x,
      y: this._state.y + before.y - after.y,
    };
  }
  viewportWorldPolygon(): readonly Point[] {
    return [
      this.viewportToWorld({ x: 0, y: 0 }),
      this.viewportToWorld({ x: this.viewport.width, y: 0 }),
      this.viewportToWorld({ x: this.viewport.width, y: this.viewport.height }),
      this.viewportToWorld({ x: 0, y: this.viewport.height }),
    ];
  }
  viewportWorldBounds(): Rect {
    return boundsOf(this.viewportWorldPolygon());
  }
  clampCenter(bounds: Rect): void {
    this._state = {
      ...this._state,
      x: Math.min(bounds.x + bounds.width, Math.max(bounds.x, this._state.x)),
      y: Math.min(bounds.y + bounds.height, Math.max(bounds.y, this._state.y)),
    };
  }
  private assertState(s: CameraState): void {
    if (![s.x, s.y, s.zoom, s.rotation].every(Number.isFinite) || s.zoom < 0.05 || s.zoom > 32)
      throw new Error('Invalid camera state');
  }
}

export function boundsOf(points: readonly Point[]): Rect {
  if (points.length === 0) throw new Error('Cannot bound an empty point list');
  const xs = points.map((p) => p.x),
    ys = points.map((p) => p.y);
  const minX = Math.min(...xs),
    minY = Math.min(...ys);
  return { x: minX, y: minY, width: Math.max(...xs) - minX, height: Math.max(...ys) - minY };
}
export function transformedBounds(local: Rect, localToWorld: Matrix2D): Rect {
  return boundsOf(
    [
      { x: local.x, y: local.y },
      { x: local.x + local.width, y: local.y },
      { x: local.x + local.width, y: local.y + local.height },
      { x: local.x, y: local.y + local.height },
    ].map((p) => localToWorld.transformPoint(p)),
  );
}
export function composeLocalToWorld(
  local: Matrix2D,
  ancestorsRootToParent: readonly Matrix2D[],
): Matrix2D {
  return ancestorsRootToParent
    .reduce((world, ancestor) => world.multiply(ancestor), new Matrix2D())
    .multiply(local);
}
export function visibleGridLines(camera: Camera, spacing: number): { xs: number[]; ys: number[] } {
  if (!(spacing > 0)) throw new Error('Grid spacing must be positive');
  const b = camera.viewportWorldBounds(),
    xs: number[] = [],
    ys: number[] = [];
  for (let x = Math.floor(b.x / spacing) * spacing; x <= b.x + b.width; x += spacing) xs.push(x);
  for (let y = Math.floor(b.y / spacing) * spacing; y <= b.y + b.height; y += spacing) ys.push(y);
  return { xs, ys };
}
export function worldToMinimap(point: Point, world: Rect, map: Rect): Point {
  if (!(world.width > 0 && world.height > 0 && map.width > 0 && map.height > 0))
    throw new Error('World and minimap bounds must have positive area');
  const scale = Math.min(map.width / world.width, map.height / world.height);
  const ox = map.x + (map.width - world.width * scale) / 2;
  const oy = map.y + (map.height - world.height * scale) / 2;
  return { x: ox + (point.x - world.x) * scale, y: oy + (point.y - world.y) * scale };
}

Rendererはcamera.worldToViewportMatrix()だけを呼び、その6成分を第2章のDPRとともにctx.setTransformへ渡します。Pointer処理はscreenToWorldだけを呼びます。GroupのlocalToWorldはcomposeLocalToWorld(local, [root,...,parent])を使い、厳密にroot×…×parent×localの順序で生成します。Camera式を複製してはいけません。trySetStateは一時値を検証してからCommitするため、非可逆または非有限なInputが最後の有効なCameraを汚染しません。

Property TestでRotation、負座標、Zoomの反復、Pointer AnchorをCoverします。

import { describe, expect, it } from 'vitest';
import { Camera, Matrix2D, composeLocalToWorld, transformedBounds } from './camera';
const close = (a: number, b: number) =>
  expect(Math.abs(a - b)).toBeLessThan(1e-8 * Math.max(1, Math.abs(a), Math.abs(b)));

describe('coordinate contracts', () => {
  it('round-trips world ↔ screen under pan, zoom and rotation', () => {
    const camera = new Camera(
      { x: -240, y: 90, zoom: 3.25, rotation: Math.PI / 3 },
      { width: 900, height: 600 },
    );
    const rect = { left: 31, top: 47 } as DOMRect;
    for (const point of [
      { x: -999, y: 0 },
      { x: 4.5, y: -81.2 },
      { x: 10000, y: 8000 },
    ]) {
      const back = camera.screenToWorld(camera.worldToScreen(point, rect), rect);
      close(back.x, point.x);
      close(back.y, point.y);
    }
    for (const screen of [
      { x: 31, y: 47 },
      { x: 480, y: 310 },
      { x: 931, y: 647 },
    ]) {
      const back = camera.worldToScreen(camera.screenToWorld(screen, rect), rect);
      close(back.x, screen.x);
      close(back.y, screen.y);
    }
  });
  it('keeps the world point under the cursor during zoom', () => {
    const camera = new Camera(
      { x: 40, y: -20, zoom: 1, rotation: 0.4 },
      { width: 800, height: 500 },
    );
    const cursor = { x: 137, y: 392 },
      before = camera.viewportToWorld(cursor);
    for (let i = 0; i < 20; i += 1) camera.zoomAround(cursor, 1.08);
    const after = camera.viewportToWorld(cursor);
    close(after.x, before.x);
    close(after.y, before.y);
  });
  it('bounds all four corners after a flipped rotation', () => {
    const matrix = Matrix2D.rotation(Math.PI / 4).multiply(new Matrix2D(-2, 0, 0, 1, 10, 20));
    const b = transformedBounds({ x: 0, y: 0, width: 20, height: 10 }, matrix);
    expect(b.width).toBeGreaterThan(35);
    expect(b.height).toBeGreaterThan(35);
  });
  it('rejects a singular matrix', () =>
    expect(() => new Matrix2D(0, 0, 0, 0, 0, 0).inverse()).toThrow('not invertible'));
  it('composes nested groups from root to local without double-applying a transform', () => {
    const root = Matrix2D.translation(100, 40);
    const group = Matrix2D.rotation(Math.PI / 2);
    const shape = Matrix2D.translation(10, 0);
    const worldPoint = composeLocalToWorld(shape, [root, group]).transformPoint({ x: 2, y: 0 });
    close(worldPoint.x, 100);
    close(worldPoint.y, 52);
  });
  it('keeps the last valid camera when a replacement would be non-invertible', () => {
    const camera = new Camera({ x: 4, y: 8, zoom: 1, rotation: 0 }, { width: 800, height: 600 });
    expect(camera.trySetState({ x: 4, y: 8, zoom: 0, rotation: 0 })).toBe(false);
    expect(camera.state).toEqual({ x: 4, y: 8, zoom: 1, rotation: 0 });
  });
});

npx vitest run src/lab/ch04/camera.test.tsを実行し、6 passedを期待します。Browser DemoではInfinite Grid、主Viewの4 Cornerを表すMinimap内Polygonを描き、DragによるPanとCursor中心Zoomを可能にします。熊に戻ると、Tapするたびに「Phone→Camera→部屋→靴箱」を逆向きにたどります。地図を飛び越えて座標を推測してはいけません。

わざと壊す

注入する故障症状根拠修正Regression Test復旧
Rotation後にScreen Pointを直接Hitへ渡すTapがShapeからずれる同じPointの2変換経路で値が異なるscreenToWorldへ統一するrotation=π/3 round trip座標の直接転送を削除する
Zoom=1e-12または1e12Jitter、Overflow、非可逆determinant/非有限Log0.05–32に制限するBoundary Clamp Test有効なZoomへ戻す
Worldに負座標があるのにunsigned/誤ったfloorを使うGrid Lineが飛ぶ-1付近の最初のLineが誤るBoundsでMath.floorを使う0をまたぐPan SnapshotGrid関数を戻す
Nested Groupの順序を逆にする子Shapeが誤った原点を中心に回る手計算PointとMatrix Pointが不一致parent×local契約2層Groupの固定例乗算順序を戻す
連続Zoomで変換結果をModelへ上書きするRound TripがDriftする1,000回後にErrorが増える常にAuthority座標から計算する反復Property Test元Documentを再Loadする
Scale=0inverseがNaNを返すdeterminantが0先に拒否して旧Cameraを保持するsingular test最後の有効Stateへ戻す
Boundsで2 Cornerだけを変換する回転ShapeがClipされる残り2 CornerがAABB外4 Cornerすべてを変換するtransformed bounds invariant4 Corner Algorithmを戻す

注入時にはCamera、Matrixの6成分、Input/Output Space名、epsilonを記録します。修正後はNode Math Testを実行し、Browserで同じ視覚位置をTapします。復旧とは、悪いInputが旧Cameraを上書きしないことであり、NaNを0へ置換することではありません。

根拠を示して合格する

性質自動/手動根拠
worldToScreen(screenToWorld(p))≈p自動複数Pan/Zoom/Rotation/負座標のProperty Test
Zoom Around Cursor自動Zoom前後でPointer下のWorld Pointが等しい
Group順序が正しい自動手計算した2層Local→Parent→Worldの例
Cameraを復旧できる自動非可逆Inputを拒否し、前のStateを保持する
Grid/Minimap/Boundsが正しい手動Corner投影と可視BoundsのBrowser Screenshot
  • すべてのPointへLocal、Parent、World、Viewport、ScreenのいずれかをLabelできる。
  • Translation、Scale、Rotationの乗算順序を口頭で決めず、Testで固定した。
  • DOMMatrix/inverseとPure TypeScript CoreをCross-validationした。
  • Pan、Cursor中心Zoom、Camera Rotation、Camera Boundsを実装した。
  • Infinite GridとMinimapを可視World Boundsだけから生成する。
  • 極端なScale、負座標、Nested Group、非可逆Matrix、Floating-point epsilonのすべてにTestがある。

後続章で座標変換が必要になったら、本章のCamera/Matrix契約へ戻ります。別の式を再発明してはいけません。

5歳児に説明する

「Matrix」「Coordinate Space」「Inverse Transform」「Affine」という言葉を使わずに答えてください。Phoneの写真で熊をTapした数値を、そのまま靴箱内の位置として使えないのはなぜでしょう。箱の中の位置はどうやって取り戻せるでしょう。

専門用語を使わない合格回答を開く

指の数値はPhoneのGlass左上から測り、熊の数値は靴箱のCornerから測ります。その間に写真の拡大、Cameraの回転、靴箱の移動と回転があります。数値は同じ熊について話していても、別の地図を使っています。元へ戻るには、来た順序と反対にたどります。Phoneの縁を取り除き、写真を縮めて回転を戻し、部屋での位置と靴箱の回転を元へ戻します。その後でもう一度前向きに進み、元の指の位置へ戻れたら、経路は正しいとわかります。