JEPA4Japan · 教程

第 4 章:四张地图与一台相机

4,354字 13分钟阅读 #Canvas#前端工程#无限画布#通俗讲解

统一局部、父级、世界、视口和屏幕坐标,同时构建可逆相机、缩放、旋转与小地图。

课程进度 课程大纲 已发布 18/18 课

从一个五岁孩子也能理解的游戏开始

把一只泰迪熊放进鞋盒,测得它距离盒底左上角 (10, 5) 厘米。然后把鞋盒放在房间地板上的 (100, 40) 处,并将它旋转四分之一圈。把相机对准房间,最后在手机照片中轻触小熊的鼻子。

先预测一下:手机上的 (220, 90) 能否直接用作鞋盒内部的 (220, 90)?鞋盒转动后,“在盒内向右 10 厘米”是否仍然是“在房间内向右 10 厘米”?如果把照片放大两倍,小熊在房间里真的移动了吗?

  1. 盒内位置局部
  2. 房间中的盒子父级 / 世界
  3. 相机视图视口
  4. 屏幕轻触屏幕
同一只小熊拥有多组数值。给出坐标之前,先说清楚使用的是哪张地图。

本章唯一的真理是:**大多数 Canvas 交互错误,本质上都是伪装起来的坐标空间混淆。**每个点都需要携带“它位于哪张地图上?”这一信息,而每次转换都需要一条返回路径。

把玩具转换成 Canvas 工程概念

玩具世界中的地图工程坐标示例
小熊自身的轮廓局部坐标相对于 Shape 原点的缩放手柄
鞋盒内部父级坐标Child Shape 相对于其 Group 的位置
整个房间世界坐标Document 中稳定的场景位置
相机画面视口坐标以 Canvas 内容的左上角为原点
手机玻璃屏幕/页面屏幕坐标Pointer 的 clientX/clientY
移动、转动盒子、放大仿射变换平移、旋转、缩放
反向追踪路径逆矩阵Screen→World 命中测试

这个类比有其局限。真实相机通常具有透视效果;本章的 Canvas Lab 使用没有深度、物体表观大小不会随距离变化的二维仿射相机。Viewport 和 Screen 也可能因页面滚动、Canvas 相对于 getBoundingClientRect() 的偏移以及 CSS 变换而有所不同。DOMMatrix 的乘法顺序必须通过测试证明,不能根据阅读词语时的自然语言顺序来猜测。

先消除误导性的直觉

  • “Pointer 坐标就是 World 坐标。” clientX/Y 属于 Viewport/Screen。相机一旦平移、缩放或旋转,它们就会出错。
  • “Zoom 只会改变缩放比例。” 单独改变 Zoom 会以相机中心为基准放大。若要围绕指针缩放,缩放前后指针下方的 World Point 必须保持不变。
  • “矩阵乘法满足交换律。” 先平移再旋转,通常不同于先旋转再平移。
  • “边界就是宽度和高度乘以 Scale。” 计算旋转矩形的 World AABB,需要变换全部四个角点并取极值。负 Scale 还会翻转方向。
  • “使用 === 判断浮点数是否相等。” 重复变换会累积微小误差,因此比较时需要使用与数值尺度相适应的 epsilon。
  • “Zoom 可以达到 0。” Scale 为零会使矩阵不可逆;极端的 Scale 也会损害精度和可用性。

生产环境工具包

前置条件约定

前置条件是第 3 章中可重放的 Renderer。我们定义以下规则:Shape 的 Local Transform 先合成到 Parent,然后沿 Scene Graph 进入 World;在相机状态 {x,y,zoom,rotation} 中,x,y 是与 Viewport 中心对应的 World Point;Viewport 坐标采用 Canvas CSS Pixels,并且不包含 DPR,DPR 由第 2 章的 Host 处理;Screen 使用 Pointer 的 clientX/Y。

形式化知识

二维 Affine Transform 可以表示 Translation、Scale、Rotation 和 Shear。齐次矩阵使我们能够组合这些变换。Canvas/DOMMatrix 的二维分量通常写作 (a,b,c,d,e,f),并按 x'=ax+cy+e, y'=bx+dy+f 变换点。组合顺序是一项约定。本章使用 translate(viewCenter) → rotate(camera.rotation) → scale(camera.zoom) → translate(-cameraCenter) 构建 World→Viewport,并通过测试锁定其实际行为。这里,camera.rotation 明确表示“世界内容在 Viewport 内旋转了多少”。如果乘积存储的是物理相机的朝向,那么在这里使用之前要将其取负;不要混用这两种符号约定。

Local→Parent 是 Shape 自身的 Transform。Parent→World 会将祖先链相乘。不要把 Group 坐标预先烘焙进 child,同时又保留 Parent Transform,否则变换会被应用两次。Camera Matrix 执行 World→Viewport。Viewport→Screen 会加上 Canvas DOMRect 的左上角偏移。反向路径使用矩阵的逆。如果行列式接近 0 或结果不是有限值,应明确失败;不要返回一个看似正常的 (0,0)。

Zoom Around Pointer 的不变量如下:改变 Zoom 之前,将指针的 Viewport Point 逆变换为 worldBefore;对新的 Scale 进行钳制,然后再次将其逆变换为 worldAfter;将 Camera 中心增加 worldBefore-worldAfter。Scale Limits 是产品常量,例如 0.05≤zoom≤32。Camera Bounds 约束相机中心或 Viewport World Bounds。对于旋转后的相机,应逆变换 Viewport 的四个角点以获得边界多边形,再取其 AABB;仅仅除以 Zoom 并不足够。

Viewport World Bounds 和 Transformed Bounds 都遵循“变换所有必需的角点,然后计算最小值/最大值”这一规则。Minimap 是另一台 Camera:将 Document World Bounds 按比例缩放到一个小矩形中,然后绘制主相机的 World 角点。Infinite Grid 不会存储无限多条线。它根据当前 Viewport World Bounds 计算第一条到最后一条可见网格线。Zoom 较小时,LOD 会有所帮助,但本章首先确保正确性。

精度策略包括限制缩放和绝对坐标值,绝不将已经变换过的交互结果作为下一次计算的源数据,在文档中存储权威的局部/世界值,并使用 almostEqual(a,b,epsilon) 进行比较。一个实用的 epsilon 是 1e-9 × max(1, |a|, |b|)。视觉命中容差是另一个问题;后续章节会以屏幕像素为单位定义它。

证据与兼容性

以上资料已于 2026-08-29 核查。现代浏览器支持 DOMMatrix,但 Node 单元测试环境可能没有提供完整实现。几何核心可以使用采用相同公式的纯 TypeScript Matrix2D,然后浏览器适配器可以将其与 DOMMatrix 进行交叉测试。当 CSS 变换应用于 Canvas 元素时,仅减去 DOMRect 的左上角并不够;还必须对 CSS 变换求逆。本教程的基线禁止在宿主元素上应用 CSS 旋转/缩放。

本章的工程增量

**起点:**第 3 章的相机只有平移和缩放,转换逻辑散落在渲染器各处。**终点:**由一个相机 API 提供局部/父级/世界/视口/屏幕之间的转换路径,以及以指针为中心的缩放、旋转、边界、小地图和无限网格。

canvas-lab/src/lab/ch04/
  matrix2d.ts
  camera.ts
  camera.test.ts
  camera-demo.ts

完整的无 DOM 数学核心和相机:

export type Point = Readonly<{ x: number; y: number }>;
export type Rect = Readonly<{ x: number; y: number; width: number; height: number }>;

export class Matrix2D {
  constructor(
    readonly a = 1,
    readonly b = 0,
    readonly c = 0,
    readonly d = 1,
    readonly e = 0,
    readonly f = 0,
  ) {}
  multiply(n: Matrix2D): Matrix2D {
    return new Matrix2D(
      this.a * n.a + this.c * n.b,
      this.b * n.a + this.d * n.b,
      this.a * n.c + this.c * n.d,
      this.b * n.c + this.d * n.d,
      this.a * n.e + this.c * n.f + this.e,
      this.b * n.e + this.d * n.f + this.f,
    );
  }
  transformPoint(p: Point): Point {
    return { x: this.a * p.x + this.c * p.y + this.e, y: this.b * p.x + this.d * p.y + this.f };
  }
  inverse(epsilon = 1e-12): Matrix2D {
    const det = this.a * this.d - this.b * this.c;
    if (!Number.isFinite(det) || Math.abs(det) <= epsilon)
      throw new Error('Matrix is not invertible');
    return new Matrix2D(
      this.d / det,
      -this.b / det,
      -this.c / det,
      this.a / det,
      (this.c * this.f - this.d * this.e) / det,
      (this.b * this.e - this.a * this.f) / det,
    );
  }
  static translation(x: number, y: number): Matrix2D {
    return new Matrix2D(1, 0, 0, 1, x, y);
  }
  static scale(value: number): Matrix2D {
    return new Matrix2D(value, 0, 0, value, 0, 0);
  }
  static rotation(radians: number): Matrix2D {
    const c = Math.cos(radians),
      s = Math.sin(radians);
    return new Matrix2D(c, s, -s, c, 0, 0);
  }
}

export type CameraState = { x: number; y: number; zoom: number; rotation: number };
export class Camera {
  private _state: CameraState;
  constructor(
    state: CameraState,
    public viewport: { width: number; height: number },
  ) {
    this.assertState(state);
    this._state = { ...state };
  }
  get state(): Readonly<CameraState> {
    return this._state;
  }
  trySetState(next: CameraState): boolean {
    try {
      this.assertState(next);
      this._state = { ...next };
      return true;
    } catch {
      return false;
    }
  }
  worldToViewportMatrix(): Matrix2D {
    const s = this._state;
    return Matrix2D.translation(this.viewport.width / 2, this.viewport.height / 2)
      .multiply(Matrix2D.rotation(s.rotation))
      .multiply(Matrix2D.scale(s.zoom))
      .multiply(Matrix2D.translation(-s.x, -s.y));
  }
  worldToViewport(point: Point): Point {
    return this.worldToViewportMatrix().transformPoint(point);
  }
  viewportToWorld(point: Point): Point {
    return this.worldToViewportMatrix().inverse().transformPoint(point);
  }
  worldToScreen(point: Point, canvasRect: Pick<DOMRect, 'left' | 'top'>): Point {
    const p = this.worldToViewport(point);
    return { x: p.x + canvasRect.left, y: p.y + canvasRect.top };
  }
  screenToWorld(point: Point, canvasRect: Pick<DOMRect, 'left' | 'top'>): Point {
    return this.viewportToWorld({ x: point.x - canvasRect.left, y: point.y - canvasRect.top });
  }
  panByViewportDelta(dx: number, dy: number): void {
    const invRotation = Matrix2D.rotation(-this._state.rotation);
    const worldDelta = invRotation.transformPoint({
      x: dx / this._state.zoom,
      y: dy / this._state.zoom,
    });
    this._state = {
      ...this._state,
      x: this._state.x - worldDelta.x,
      y: this._state.y - worldDelta.y,
    };
  }
  zoomAround(pointer: Point, factor: number): void {
    const before = this.viewportToWorld(pointer);
    const zoom = Math.min(32, Math.max(0.05, this._state.zoom * factor));
    this._state = { ...this._state, zoom };
    const after = this.viewportToWorld(pointer);
    this._state = {
      ...this._state,
      x: this._state.x + before.x - after.x,
      y: this._state.y + before.y - after.y,
    };
  }
  viewportWorldPolygon(): readonly Point[] {
    return [
      this.viewportToWorld({ x: 0, y: 0 }),
      this.viewportToWorld({ x: this.viewport.width, y: 0 }),
      this.viewportToWorld({ x: this.viewport.width, y: this.viewport.height }),
      this.viewportToWorld({ x: 0, y: this.viewport.height }),
    ];
  }
  viewportWorldBounds(): Rect {
    return boundsOf(this.viewportWorldPolygon());
  }
  clampCenter(bounds: Rect): void {
    this._state = {
      ...this._state,
      x: Math.min(bounds.x + bounds.width, Math.max(bounds.x, this._state.x)),
      y: Math.min(bounds.y + bounds.height, Math.max(bounds.y, this._state.y)),
    };
  }
  private assertState(s: CameraState): void {
    if (![s.x, s.y, s.zoom, s.rotation].every(Number.isFinite) || s.zoom < 0.05 || s.zoom > 32)
      throw new Error('Invalid camera state');
  }
}

export function boundsOf(points: readonly Point[]): Rect {
  if (points.length === 0) throw new Error('Cannot bound an empty point list');
  const xs = points.map((p) => p.x),
    ys = points.map((p) => p.y);
  const minX = Math.min(...xs),
    minY = Math.min(...ys);
  return { x: minX, y: minY, width: Math.max(...xs) - minX, height: Math.max(...ys) - minY };
}
export function transformedBounds(local: Rect, localToWorld: Matrix2D): Rect {
  return boundsOf(
    [
      { x: local.x, y: local.y },
      { x: local.x + local.width, y: local.y },
      { x: local.x + local.width, y: local.y + local.height },
      { x: local.x, y: local.y + local.height },
    ].map((p) => localToWorld.transformPoint(p)),
  );
}
export function composeLocalToWorld(
  local: Matrix2D,
  ancestorsRootToParent: readonly Matrix2D[],
): Matrix2D {
  return ancestorsRootToParent
    .reduce((world, ancestor) => world.multiply(ancestor), new Matrix2D())
    .multiply(local);
}
export function visibleGridLines(camera: Camera, spacing: number): { xs: number[]; ys: number[] } {
  if (!(spacing > 0)) throw new Error('Grid spacing must be positive');
  const b = camera.viewportWorldBounds(),
    xs: number[] = [],
    ys: number[] = [];
  for (let x = Math.floor(b.x / spacing) * spacing; x <= b.x + b.width; x += spacing) xs.push(x);
  for (let y = Math.floor(b.y / spacing) * spacing; y <= b.y + b.height; y += spacing) ys.push(y);
  return { xs, ys };
}
export function worldToMinimap(point: Point, world: Rect, map: Rect): Point {
  if (!(world.width > 0 && world.height > 0 && map.width > 0 && map.height > 0))
    throw new Error('World and minimap bounds must have positive area');
  const scale = Math.min(map.width / world.width, map.height / world.height);
  const ox = map.x + (map.width - world.width * scale) / 2;
  const oy = map.y + (map.height - world.height * scale) / 2;
  return { x: ox + (point.x - world.x) * scale, y: oy + (point.y - world.y) * scale };
}

渲染器只调用 camera.worldToViewportMatrix(),并将其六个分量以及第 2 章中的 DPR 一并提供给 ctx.setTransform。指针处理只调用 screenToWorld。Group 的 localToWorld 使用 composeLocalToWorld(local, [root,...,parent]),严格按照该顺序生成 root×…×parent×local;它不得照搬相机公式。trySetState 在提交临时值之前对其进行验证,因此不可逆或非有限的输入不会污染最后一个有效相机。

属性测试涵盖旋转、负坐标、重复缩放和指针锚定:

import { describe, expect, it } from 'vitest';
import { Camera, Matrix2D, composeLocalToWorld, transformedBounds } from './camera';
const close = (a: number, b: number) =>
  expect(Math.abs(a - b)).toBeLessThan(1e-8 * Math.max(1, Math.abs(a), Math.abs(b)));

describe('coordinate contracts', () => {
  it('round-trips world ↔ screen under pan, zoom and rotation', () => {
    const camera = new Camera(
      { x: -240, y: 90, zoom: 3.25, rotation: Math.PI / 3 },
      { width: 900, height: 600 },
    );
    const rect = { left: 31, top: 47 } as DOMRect;
    for (const point of [
      { x: -999, y: 0 },
      { x: 4.5, y: -81.2 },
      { x: 10000, y: 8000 },
    ]) {
      const back = camera.screenToWorld(camera.worldToScreen(point, rect), rect);
      close(back.x, point.x);
      close(back.y, point.y);
    }
    for (const screen of [
      { x: 31, y: 47 },
      { x: 480, y: 310 },
      { x: 931, y: 647 },
    ]) {
      const back = camera.worldToScreen(camera.screenToWorld(screen, rect), rect);
      close(back.x, screen.x);
      close(back.y, screen.y);
    }
  });
  it('keeps the world point under the cursor during zoom', () => {
    const camera = new Camera(
      { x: 40, y: -20, zoom: 1, rotation: 0.4 },
      { width: 800, height: 500 },
    );
    const cursor = { x: 137, y: 392 },
      before = camera.viewportToWorld(cursor);
    for (let i = 0; i < 20; i += 1) camera.zoomAround(cursor, 1.08);
    const after = camera.viewportToWorld(cursor);
    close(after.x, before.x);
    close(after.y, before.y);
  });
  it('bounds all four corners after a flipped rotation', () => {
    const matrix = Matrix2D.rotation(Math.PI / 4).multiply(new Matrix2D(-2, 0, 0, 1, 10, 20));
    const b = transformedBounds({ x: 0, y: 0, width: 20, height: 10 }, matrix);
    expect(b.width).toBeGreaterThan(35);
    expect(b.height).toBeGreaterThan(35);
  });
  it('rejects a singular matrix', () =>
    expect(() => new Matrix2D(0, 0, 0, 0, 0, 0).inverse()).toThrow('not invertible'));
  it('composes nested groups from root to local without double-applying a transform', () => {
    const root = Matrix2D.translation(100, 40);
    const group = Matrix2D.rotation(Math.PI / 2);
    const shape = Matrix2D.translation(10, 0);
    const worldPoint = composeLocalToWorld(shape, [root, group]).transformPoint({ x: 2, y: 0 });
    close(worldPoint.x, 100);
    close(worldPoint.y, 52);
  });
  it('keeps the last valid camera when a replacement would be non-invertible', () => {
    const camera = new Camera({ x: 4, y: 8, zoom: 1, rotation: 0 }, { width: 800, height: 600 });
    expect(camera.trySetState({ x: 4, y: 8, zoom: 0, rotation: 0 })).toBe(false);
    expect(camera.state).toEqual({ x: 4, y: 8, zoom: 1, rotation: 0 });
  });
});

运行 npx vitest run src/lab/ch04/camera.test.ts;预期得到 6 passed。浏览器演示必须绘制无限网格,并在小地图中绘制主视图的四角多边形,同时支持拖动平移和围绕光标缩放。回到泰迪熊的例子:每次点击都沿着“手机 → 相机 → 房间 → 鞋盒”反向传递。绝不要通过跨地图跳跃来猜测坐标。

故意破坏它

注入项症状证据修复回归测试恢复
旋转后直接使用屏幕点进行命中测试点击未命中形状同一点的两条转换路径产生不同的值统一使用 screenToWorldrotation=π/3 往返测试移除直接坐标转发
Zoom=1e-12 或 1e12抖动、溢出或不可逆行列式/非有限值日志限制在 0.05–32边界钳制测试恢复合法缩放值
世界使用负坐标,但网格使用无符号值/错误的向下取整网格线跳动-1 附近的第一条线有误在边界处使用 Math.floor跨越 0 平移的快照测试恢复网格函数
嵌套 Group 的顺序颠倒子形状围绕错误的原点旋转手工计算的点与矩阵计算的点不同parent×local 约定固定的两级 Group 示例恢复乘法顺序
重复缩放使用变换后的结果覆盖模型往返转换发生漂移1,000 次迭代后误差增大始终从权威坐标计算重复属性测试重新加载原始文档
Scale=0求逆返回 NaN行列式为 0提前拒绝并保留旧相机奇异性测试恢复最后一个有效状态
边界只变换两个角旋转后的形状被裁剪另外两个角落在 AABB 之外变换全部四个角变换后边界不变量恢复四角算法

注入期间,记录相机、矩阵的全部六个分量、输入/输出空间名称和 epsilon。修复后,运行 Node 数学测试,并在浏览器中点击同一视觉位置。恢复意味着错误输入没有覆盖旧相机;并不意味着用零替换 NaN。

用证据验收

属性自动/手动证据
worldToScreen(screenToWorld(p))≈p自动覆盖平移/缩放/旋转/负坐标的属性测试
围绕光标缩放自动缩放前后指针下方的世界点相同
Group 顺序正确自动手工计算的两级局部→父级→世界示例
相机可以恢复自动拒绝不可逆输入并保留先前状态
网格/小地图/边界正确手动展示角点投影和可见边界的浏览器截图
  • 我可以将每个点标记为局部、父级、世界、视口或屏幕。
  • 平移、缩放和旋转的乘法顺序经过测试,而不是靠口头假设。
  • DOMMatrix/inverse 与纯 TypeScript 核心已经过交叉验证。
  • 已实现平移、以光标为中心的缩放、相机旋转和相机边界。
  • 无限网格和小地图仅根据可见世界边界生成。
  • 极端缩放、负坐标、嵌套 Group、不可逆矩阵和浮点 epsilon 都有相应测试。

后续章节每当遇到坐标转换时,都应回到本章的相机/矩阵约定,而不是另造一个公式。

给五岁小孩讲明白

回答时不要使用“矩阵”“坐标空间”“逆变换”或“仿射”这些词:为什么不能直接把在手机照片上点击小熊得到的数字,当作它在鞋盒里的位置?怎样才能找回它在盒子里的位置?

展开一个不含术语的优质答案

手指的数字是从手机玻璃的左上角量起的,而小熊的数字是从鞋盒的一个角量起的。在这两者之间,照片被放大了,相机转动了,鞋盒也移动并转动了。这些数字描述的是同一只小熊,但使用的是不同的地图。要找到原来的位置,就反向沿原路返回:先去掉手机的边缘偏移,再把照片缩小并转回去,然后撤销房间中的位置变化和鞋盒的转动。之后再沿这条路线正向走一遍。如果最终回到手指原来的位置,就说明路线是正确的。