JEPA4Japan · tutorials

Chapter 4: Four Maps and a Camera

3,162 words 15 min read #Canvas#Frontend Engineering#Infinite Canvas#ELI5

Unify local, parent, world, viewport, and screen coordinates while building an invertible camera, zoom, rotation, and minimap.

Course progress Course outline 18 of 18 lessons available

Part I: Choose the Surface Before You Draw—Product, Pixels, and Coordinates

  1. 01 Chapter 1: Do Not Draw Yet—Canvas Is Not a Product Architecture available now
  2. 02 Chapter 2: A Sheet of Pixels That Forgets available now
  3. 03 Chapter 3: Turn Drawing into a Replayable Recipe available now
  4. 04 Chapter 4: Four Maps and a Camera Current lesson

Part II: Give the Pixel World a Brain—Model, Scheduling, Input, and Tools

  1. 05 Chapter 5: Give the Pixel World a Registry available now
  2. 06 Chapter 6: Redraw Only When the Light Turns On—Render Scheduling and the React Boundary available now
  3. 07 Chapter 7: Mouse, Touch, and Pen Speak One Language available now
  4. 08 Chapter 8: Find the Big Box Before Inspecting the Edge available now
  5. 09 Chapter 9: Tools Are Traffic Lights, Not a Bag of Booleans available now

Part III: From “It Drags” to “It Is Trustworthy”—Interaction, Text, Assets, and Recovery

  1. 10 Chapter 10: Make the Editor Feel Right available now
  2. 11 Chapter 11: Drawn Text Is Not Editable Text available now
  3. 12 Chapter 12: Borrowed Images Cannot Be Packed Without Rules available now
  4. 13 Chapter 13: Time Machines and Old Boxes available now
  5. 14 Chapter 14: Looking Correct Is Not Being Correct available now

Part IV: Master-Level Decisions—Performance, Workers, GPU, SDKs, Collaboration, and AI

  1. 15 Chapter 15: Do Not Search Ten Thousand Children One by One available now
  2. 16 Chapter 16: Keep the Front Desk Out of the Kitchen—Worker and GPU Upgrades available now
  3. 17 Chapter 17: Build the Car or Buy a Proven Chassis? available now
  4. 18 Chapter 18: People and AI Edit the Same Ledger available now

Start with a game a five-year-old can understand

Put a teddy bear in a shoebox, measuring (10, 5) centimeters from the top-left corner of the box floor. Then place the shoebox on the room floor at (100, 40) and rotate it a quarter turn. Point a camera at the room, and finally tap the bear’s nose in the photo on your phone.

First predict: can (220, 90) on the phone be used directly as (220, 90) inside the shoebox? After the shoebox turns, is “10 centimeters to the right inside the box” still “10 centimeters to the right in the room”? If you enlarge the photo twofold, did the bear actually move in the room?

  1. Position inside the boxLocal
  2. Box in the roomParent / World
  3. Camera viewViewport
  4. Screen tapScreen
The same bear has several sets of numbers. Name the map before giving the coordinates.

The single truth of this chapter is: most Canvas interaction bugs are coordinate-space confusion in disguise. Every point needs to carry “which map is this on?”, and every conversion needs a route back.

Translate the toys into Canvas engineering

Map in the toy worldEngineering coordinatesExample
The bear’s own outlineLocal CoordinatesResize handle relative to a Shape origin
Inside the shoeboxParent CoordinatesChild Shape relative to its Group
The entire roomWorld CoordinatesStable scene position in the Document
Camera frameViewport CoordinatesTop-left of Canvas content as origin
Phone glass/pageScreen CoordinatesPointer clientX/clientY
Moving, turning the box, enlargingAffine TransformTranslation, Rotation, Scale
Tracing the route backwardInverse MatrixScreen→World hit testing

The analogy has limits. A real camera usually has perspective; this chapter’s Canvas Lab uses a two-dimensional affine camera with no depth or apparent size change by distance. Viewport and Screen may also differ because of page scrolling, the Canvas offset from getBoundingClientRect(), and CSS transforms. DOMMatrix multiplication order must be proven by tests, not guessed from the natural-language order in which words are read.

Kill the misleading intuitions first

  • “Pointer coordinates are World coordinates.” clientX/Y belong to Viewport/Screen. They become wrong as soon as the camera pans, zooms, or rotates.
  • “Zoom only changes scale.” Changing Zoom alone enlarges around the camera center. To zoom around the pointer, the World Point beneath it must remain the same before and after zooming.
  • “Matrix multiplication is commutative.” Translating then rotating usually differs from rotating then translating.
  • “Bounds are just width and height multiplied by Scale.” The World AABB of a rotated rectangle requires transforming all four corners and taking extrema. A negative Scale also flips direction.
  • “Use === for floating-point equality.” Repeated transforms accumulate tiny errors, so comparisons need an epsilon appropriate to their scale.
  • “Zoom can reach 0.” A zero Scale makes a matrix non-invertible; extreme Scales also damage precision and usability.

Production backpack

Prerequisite contract

The prerequisite is the replayable Renderer from Chapter 3. We define these rules: a Shape’s Local Transform composes into Parent and then follows the Scene Graph to World; in camera state {x,y,zoom,rotation}, x,y are the World Point corresponding to the center of the Viewport; Viewport coordinates are Canvas CSS Pixels and exclude DPR, which Chapter 2’s Host handles; Screen uses a Pointer’s clientX/Y.

Formal knowledge

A two-dimensional Affine Transform can represent Translation, Scale, Rotation, and Shear. Homogeneous matrices let us compose them. The 2D components of Canvas/DOMMatrix are commonly written (a,b,c,d,e,f), transforming a point as x'=ax+cy+e, y'=bx+dy+f. Composition order is a contract. This chapter builds World→Viewport with translate(viewCenter) → rotate(camera.rotation) → scale(camera.zoom) → translate(-cameraCenter) and locks its real behavior with tests. Here, camera.rotation explicitly means “how far world content rotates within the Viewport.” If the product stores the orientation of a physical camera instead, negate it before using it here; do not mix the two sign conventions.

Local→Parent is a Shape’s own Transform. Parent→World multiplies the ancestor chain. Do not pre-bake Group coordinates into a child and still retain the Parent Transform, or it will be applied twice. Camera Matrix performs World→Viewport. Viewport→Screen adds the top-left offset of the Canvas DOMRect. The reverse route uses the matrix inverse. If the determinant is near 0 or the result is non-finite, fail explicitly; do not return an apparently normal (0,0).

The invariant of Zoom Around Pointer is this: before changing Zoom, inverse-transform the pointer’s Viewport Point into worldBefore; clamp the new Scale, then inverse-transform it again into worldAfter; increase the Camera center by worldBefore-worldAfter. Scale Limits are product constants, for example 0.05≤zoom≤32. Camera Bounds constrain either the camera center or Viewport World Bounds. For a rotated camera, inverse-transform the four Viewport corners to get a bounding polygon and then take its AABB; merely dividing by Zoom is not enough.

Viewport World Bounds and Transformed Bounds both follow the rule “transform all required corners, then calculate min/max.” A Minimap is another Camera: scale Document World Bounds proportionally into a small rectangle, then draw the main camera’s World corners. An Infinite Grid does not store infinitely many lines. It calculates the first through last visible grid lines from the current Viewport World Bounds. At a small Zoom, LOD can help, but this chapter establishes correctness first.

Precision strategy includes limiting Zoom and absolute coordinate values, never feeding an already transformed interaction result back as the next source, storing authoritative Local/World values in the Document, and comparing with almostEqual(a,b,epsilon). A useful epsilon is 1e-9 × max(1, |a|, |b|). Visual hit tolerance is a different concern; a later chapter defines it in screen pixels.

Evidence and compatibility

Sources above were checked on 2026-08-29. DOMMatrix is available in modern browsers, but a Node unit-test environment may not provide a complete implementation. Geometry Core can use a pure TypeScript Matrix2D with the same formulas, then the browser adapter can cross-test it against DOMMatrix. When a CSS transform applies to the Canvas element, subtracting only the top-left DOMRect is insufficient; the CSS Transform must also be inverted. This tutorial’s baseline prohibits CSS rotation/scale on the Host element.

This chapter’s engineering increment

Starting point: Chapter 3’s Camera contains only pan and Zoom, with conversions scattered through the Renderer. Finish line: one Camera API supplies Local/Parent/World/Viewport/Screen routes, pointer-centered zoom, rotation, bounds, Minimap, and Infinite Grid.

canvas-lab/src/lab/ch04/
  matrix2d.ts
  camera.ts
  camera.test.ts
  camera-demo.ts

The complete DOM-free math core and Camera:

export type Point = Readonly<{ x: number; y: number }>;
export type Rect = Readonly<{ x: number; y: number; width: number; height: number }>;

export class Matrix2D {
  constructor(
    readonly a = 1,
    readonly b = 0,
    readonly c = 0,
    readonly d = 1,
    readonly e = 0,
    readonly f = 0,
  ) {}
  multiply(n: Matrix2D): Matrix2D {
    return new Matrix2D(
      this.a * n.a + this.c * n.b,
      this.b * n.a + this.d * n.b,
      this.a * n.c + this.c * n.d,
      this.b * n.c + this.d * n.d,
      this.a * n.e + this.c * n.f + this.e,
      this.b * n.e + this.d * n.f + this.f,
    );
  }
  transformPoint(p: Point): Point {
    return { x: this.a * p.x + this.c * p.y + this.e, y: this.b * p.x + this.d * p.y + this.f };
  }
  inverse(epsilon = 1e-12): Matrix2D {
    const det = this.a * this.d - this.b * this.c;
    if (!Number.isFinite(det) || Math.abs(det) <= epsilon)
      throw new Error('Matrix is not invertible');
    return new Matrix2D(
      this.d / det,
      -this.b / det,
      -this.c / det,
      this.a / det,
      (this.c * this.f - this.d * this.e) / det,
      (this.b * this.e - this.a * this.f) / det,
    );
  }
  static translation(x: number, y: number): Matrix2D {
    return new Matrix2D(1, 0, 0, 1, x, y);
  }
  static scale(value: number): Matrix2D {
    return new Matrix2D(value, 0, 0, value, 0, 0);
  }
  static rotation(radians: number): Matrix2D {
    const c = Math.cos(radians),
      s = Math.sin(radians);
    return new Matrix2D(c, s, -s, c, 0, 0);
  }
}

export type CameraState = { x: number; y: number; zoom: number; rotation: number };
export class Camera {
  private _state: CameraState;
  constructor(
    state: CameraState,
    public viewport: { width: number; height: number },
  ) {
    this.assertState(state);
    this._state = { ...state };
  }
  get state(): Readonly<CameraState> {
    return this._state;
  }
  trySetState(next: CameraState): boolean {
    try {
      this.assertState(next);
      this._state = { ...next };
      return true;
    } catch {
      return false;
    }
  }
  worldToViewportMatrix(): Matrix2D {
    const s = this._state;
    return Matrix2D.translation(this.viewport.width / 2, this.viewport.height / 2)
      .multiply(Matrix2D.rotation(s.rotation))
      .multiply(Matrix2D.scale(s.zoom))
      .multiply(Matrix2D.translation(-s.x, -s.y));
  }
  worldToViewport(point: Point): Point {
    return this.worldToViewportMatrix().transformPoint(point);
  }
  viewportToWorld(point: Point): Point {
    return this.worldToViewportMatrix().inverse().transformPoint(point);
  }
  worldToScreen(point: Point, canvasRect: Pick<DOMRect, 'left' | 'top'>): Point {
    const p = this.worldToViewport(point);
    return { x: p.x + canvasRect.left, y: p.y + canvasRect.top };
  }
  screenToWorld(point: Point, canvasRect: Pick<DOMRect, 'left' | 'top'>): Point {
    return this.viewportToWorld({ x: point.x - canvasRect.left, y: point.y - canvasRect.top });
  }
  panByViewportDelta(dx: number, dy: number): void {
    const invRotation = Matrix2D.rotation(-this._state.rotation);
    const worldDelta = invRotation.transformPoint({
      x: dx / this._state.zoom,
      y: dy / this._state.zoom,
    });
    this._state = {
      ...this._state,
      x: this._state.x - worldDelta.x,
      y: this._state.y - worldDelta.y,
    };
  }
  zoomAround(pointer: Point, factor: number): void {
    const before = this.viewportToWorld(pointer);
    const zoom = Math.min(32, Math.max(0.05, this._state.zoom * factor));
    this._state = { ...this._state, zoom };
    const after = this.viewportToWorld(pointer);
    this._state = {
      ...this._state,
      x: this._state.x + before.x - after.x,
      y: this._state.y + before.y - after.y,
    };
  }
  viewportWorldPolygon(): readonly Point[] {
    return [
      this.viewportToWorld({ x: 0, y: 0 }),
      this.viewportToWorld({ x: this.viewport.width, y: 0 }),
      this.viewportToWorld({ x: this.viewport.width, y: this.viewport.height }),
      this.viewportToWorld({ x: 0, y: this.viewport.height }),
    ];
  }
  viewportWorldBounds(): Rect {
    return boundsOf(this.viewportWorldPolygon());
  }
  clampCenter(bounds: Rect): void {
    this._state = {
      ...this._state,
      x: Math.min(bounds.x + bounds.width, Math.max(bounds.x, this._state.x)),
      y: Math.min(bounds.y + bounds.height, Math.max(bounds.y, this._state.y)),
    };
  }
  private assertState(s: CameraState): void {
    if (![s.x, s.y, s.zoom, s.rotation].every(Number.isFinite) || s.zoom < 0.05 || s.zoom > 32)
      throw new Error('Invalid camera state');
  }
}

export function boundsOf(points: readonly Point[]): Rect {
  if (points.length === 0) throw new Error('Cannot bound an empty point list');
  const xs = points.map((p) => p.x),
    ys = points.map((p) => p.y);
  const minX = Math.min(...xs),
    minY = Math.min(...ys);
  return { x: minX, y: minY, width: Math.max(...xs) - minX, height: Math.max(...ys) - minY };
}
export function transformedBounds(local: Rect, localToWorld: Matrix2D): Rect {
  return boundsOf(
    [
      { x: local.x, y: local.y },
      { x: local.x + local.width, y: local.y },
      { x: local.x + local.width, y: local.y + local.height },
      { x: local.x, y: local.y + local.height },
    ].map((p) => localToWorld.transformPoint(p)),
  );
}
export function composeLocalToWorld(
  local: Matrix2D,
  ancestorsRootToParent: readonly Matrix2D[],
): Matrix2D {
  return ancestorsRootToParent
    .reduce((world, ancestor) => world.multiply(ancestor), new Matrix2D())
    .multiply(local);
}
export function visibleGridLines(camera: Camera, spacing: number): { xs: number[]; ys: number[] } {
  if (!(spacing > 0)) throw new Error('Grid spacing must be positive');
  const b = camera.viewportWorldBounds(),
    xs: number[] = [],
    ys: number[] = [];
  for (let x = Math.floor(b.x / spacing) * spacing; x <= b.x + b.width; x += spacing) xs.push(x);
  for (let y = Math.floor(b.y / spacing) * spacing; y <= b.y + b.height; y += spacing) ys.push(y);
  return { xs, ys };
}
export function worldToMinimap(point: Point, world: Rect, map: Rect): Point {
  if (!(world.width > 0 && world.height > 0 && map.width > 0 && map.height > 0))
    throw new Error('World and minimap bounds must have positive area');
  const scale = Math.min(map.width / world.width, map.height / world.height);
  const ox = map.x + (map.width - world.width * scale) / 2;
  const oy = map.y + (map.height - world.height * scale) / 2;
  return { x: ox + (point.x - world.x) * scale, y: oy + (point.y - world.y) * scale };
}

The Renderer calls only camera.worldToViewportMatrix() and supplies its six components, together with the DPR from Chapter 2, to ctx.setTransform. Pointer handling calls only screenToWorld. A Group’s localToWorld uses composeLocalToWorld(local, [root,...,parent]), producing root×…×parent×local in that exact order; it must not copy the camera formula. trySetState validates a temporary value before committing it, so non-invertible or non-finite input cannot contaminate the last valid Camera.

Property tests cover rotation, negative coordinates, repeated Zoom, and pointer anchoring:

import { describe, expect, it } from 'vitest';
import { Camera, Matrix2D, composeLocalToWorld, transformedBounds } from './camera';
const close = (a: number, b: number) =>
  expect(Math.abs(a - b)).toBeLessThan(1e-8 * Math.max(1, Math.abs(a), Math.abs(b)));

describe('coordinate contracts', () => {
  it('round-trips world ↔ screen under pan, zoom and rotation', () => {
    const camera = new Camera(
      { x: -240, y: 90, zoom: 3.25, rotation: Math.PI / 3 },
      { width: 900, height: 600 },
    );
    const rect = { left: 31, top: 47 } as DOMRect;
    for (const point of [
      { x: -999, y: 0 },
      { x: 4.5, y: -81.2 },
      { x: 10000, y: 8000 },
    ]) {
      const back = camera.screenToWorld(camera.worldToScreen(point, rect), rect);
      close(back.x, point.x);
      close(back.y, point.y);
    }
    for (const screen of [
      { x: 31, y: 47 },
      { x: 480, y: 310 },
      { x: 931, y: 647 },
    ]) {
      const back = camera.worldToScreen(camera.screenToWorld(screen, rect), rect);
      close(back.x, screen.x);
      close(back.y, screen.y);
    }
  });
  it('keeps the world point under the cursor during zoom', () => {
    const camera = new Camera(
      { x: 40, y: -20, zoom: 1, rotation: 0.4 },
      { width: 800, height: 500 },
    );
    const cursor = { x: 137, y: 392 },
      before = camera.viewportToWorld(cursor);
    for (let i = 0; i < 20; i += 1) camera.zoomAround(cursor, 1.08);
    const after = camera.viewportToWorld(cursor);
    close(after.x, before.x);
    close(after.y, before.y);
  });
  it('bounds all four corners after a flipped rotation', () => {
    const matrix = Matrix2D.rotation(Math.PI / 4).multiply(new Matrix2D(-2, 0, 0, 1, 10, 20));
    const b = transformedBounds({ x: 0, y: 0, width: 20, height: 10 }, matrix);
    expect(b.width).toBeGreaterThan(35);
    expect(b.height).toBeGreaterThan(35);
  });
  it('rejects a singular matrix', () =>
    expect(() => new Matrix2D(0, 0, 0, 0, 0, 0).inverse()).toThrow('not invertible'));
  it('composes nested groups from root to local without double-applying a transform', () => {
    const root = Matrix2D.translation(100, 40);
    const group = Matrix2D.rotation(Math.PI / 2);
    const shape = Matrix2D.translation(10, 0);
    const worldPoint = composeLocalToWorld(shape, [root, group]).transformPoint({ x: 2, y: 0 });
    close(worldPoint.x, 100);
    close(worldPoint.y, 52);
  });
  it('keeps the last valid camera when a replacement would be non-invertible', () => {
    const camera = new Camera({ x: 4, y: 8, zoom: 1, rotation: 0 }, { width: 800, height: 600 });
    expect(camera.trySetState({ x: 4, y: 8, zoom: 0, rotation: 0 })).toBe(false);
    expect(camera.state).toEqual({ x: 4, y: 8, zoom: 1, rotation: 0 });
  });
});

Run npx vitest run src/lab/ch04/camera.test.ts; expect 6 passed. The browser Demo must draw an Infinite Grid and the main view’s four-corner polygon in the Minimap, support drag-to-Pan, and zoom around the cursor. Return to the teddy bear: every tap travels backward along “phone → camera → room → shoebox.” Never guess coordinates by jumping across maps.

Break it on purpose

InjectionSymptomEvidenceFixRegression testRecovery
Use Screen point directly for a hit after rotationTap misses the ShapeTwo conversion routes for the same point produce different valuesStandardize on screenToWorldrotation=π/3 round tripRemove direct coordinate forwarding
Zoom=1e-12 or 1e12Jitter, overflow, or non-invertibilityDeterminant/non-finite logsLimit to 0.05–32Boundary-clamping testRestore legal Zoom
World uses negative coordinates but grid uses unsigned/wrong floorGrid line jumpsFirst line near -1 is wrongUse Math.floor at boundsSnapshot panning across 0Restore Grid function
Nested Group order reversedChild Shape rotates around wrong originHand-calculated point differs from matrix pointparent×local contractFixed two-level Group exampleRestore multiplication order
Repeated Zoom overwrites model with transformed resultsRound Trip driftsError grows after 1,000 iterationsAlways calculate from authoritative coordinatesRepeated property testReload original Document
Scale=0inverse returns NaNDeterminant is 0Reject early and retain old Camerasingular testRestore last valid state
Bounds transforms only two cornersRotated Shape is clippedOther two corners fall outside AABBTransform all four cornerstransformed bounds invariantRestore four-corner algorithm

During injection, record the Camera, all six matrix components, input/output space names, and epsilon. After the fix, run the Node math tests and tap the same visual location in the browser. Recovery means bad input did not overwrite the old Camera; it does not mean replacing NaN with zero.

Pass with evidence

PropertyAutomated/manualEvidence
worldToScreen(screenToWorld(p))≈pAutomatedProperty tests across pan/zoom/rotation/negative coordinates
Zoom Around CursorAutomatedWorld Point under pointer is equal before and after Zoom
Group order is correctAutomatedHand-calculated two-level Local→Parent→World example
Camera can recoverAutomatedNon-invertible input rejected and previous state retained
Grid/Minimap/Bounds are correctManualBrowser screenshot of corner projections and visible bounds
  • I can label every point as Local, Parent, World, Viewport, or Screen.
  • Translation, Scale, and Rotation multiplication order is tested rather than verbally assumed.
  • DOMMatrix/inverse and the pure TypeScript core have been cross-validated.
  • Pan, cursor-centered Zoom, camera rotation, and Camera Bounds are implemented.
  • Infinite Grid and Minimap generate only from visible World Bounds.
  • Extreme Scale, negative coordinates, nested Groups, non-invertible matrices, and floating-point epsilon all have tests.

Whenever later chapters encounter coordinate conversion, return to this chapter’s Camera/Matrix contract instead of inventing another formula.

Explain it to a five-year-old

Without using the words “matrix,” “coordinate space,” “inverse transform,” or “affine,” answer: Why can you not use the number from tapping the bear in a phone photo directly as its position inside the shoebox? How do you recover the position inside the box?

Expand a good jargon-free answer

The finger's numbers are measured from the top-left of the phone glass, while the bear's numbers are measured from a corner of the shoebox. Between them, the photo was enlarged, the camera turned, and the shoebox moved and turned. The numbers describe the same bear but use different maps. To find the original place, retrace the trip in reverse: remove the phone's edge, shrink and turn back the photo, then undo the room position and the shoebox turn. Walk the route forward once more afterward. If it returns to the original finger position, the route is correct.