JEPA4Japan · tutorials

Chapter 11: Drawn Text Is Not Editable Text

3,444 words 16 min read #Canvas#Frontend Engineering#Infinite Canvas#ELI5

Let Canvas render previews while the DOM edits text, with correct handling for fonts, wrapping, graphemes, bidi, IME, caret, and selection.

Course progress Course outline 18 of 18 lessons available

Part I: Choose the Surface Before You Draw—Product, Pixels, and Coordinates

  1. 01 Chapter 1: Do Not Draw Yet—Canvas Is Not a Product Architecture available now
  2. 02 Chapter 2: A Sheet of Pixels That Forgets available now
  3. 03 Chapter 3: Turn Drawing into a Replayable Recipe available now
  4. 04 Chapter 4: Four Maps and a Camera available now

Part II: Give the Pixel World a Brain—Model, Scheduling, Input, and Tools

  1. 05 Chapter 5: Give the Pixel World a Registry available now
  2. 06 Chapter 6: Redraw Only When the Light Turns On—Render Scheduling and the React Boundary available now
  3. 07 Chapter 7: Mouse, Touch, and Pen Speak One Language available now
  4. 08 Chapter 8: Find the Big Box Before Inspecting the Edge available now
  5. 09 Chapter 9: Tools Are Traffic Lights, Not a Bag of Booleans available now

Part III: From “It Drags” to “It Is Trustworthy”—Interaction, Text, Assets, and Recovery

  1. 10 Chapter 10: Make the Editor Feel Right available now
  2. 11 Chapter 11: Drawn Text Is Not Editable Text Current lesson
  3. 12 Chapter 12: Borrowed Images Cannot Be Packed Without Rules available now
  4. 13 Chapter 13: Time Machines and Old Boxes available now
  5. 14 Chapter 14: Looking Correct Is Not Being Correct available now

Part IV: Master-Level Decisions—Performance, Workers, GPU, SDKs, Collaboration, and AI

  1. 15 Chapter 15: Do Not Search Ten Thousand Children One by One available now
  2. 16 Chapter 16: Keep the Front Desk Out of the Kitchen—Worker and GPU Upgrades available now
  3. 17 Chapter 17: Build the Car or Buy a Proven Chassis? available now
  4. 18 Chapter 18: People and AI Edit the Same Ledger available now

Start with a game a five-year-old can understand

The one truth in this chapter: Text Rendering and Text Editing are two completely different systems.

Write “Hello 👨‍👩‍👧‍👦” on a sheet of paper, take a photograph, and give both the original paper and the photo to a child. Both show the words. Ask the child to put a caret between “e” and “l,” select “Hello,” convert a reading into kanji with a Japanese input method, and have a reader speak the text aloud.

Predict what happens first. Which of those tasks can the photograph perform? If a program appends one character to the end of a string for every keydown, what happens during Chinese Pinyin composition? When Backspace deletes the family Emoji, should it remove one person or the whole family symbol?

  1. View the photo normallyfast and moves with the picture
  2. Switch to the original paper to editcaret, selection, language input
  3. Align the original paper exactlyposition, scale, rotation
  4. Take a new photo when finishedcommit or cancel
Answer first: if text is visible, does the browser know every editable position? No. Pixels have no Caret, Selection, or language-input protocol.

Canvas fillText() is like photographing letters into an image. A textarea or contenteditable element is like the original paper the browser understands. The Renderer draws a Text Shape when it is not being edited. A DOM Overlay takes over after entering Editing State. On commit, the Document updates and the Renderer redraws. On cancel, the Overlay is discarded and the original value stays unchanged.

Translate the toys into Canvas

Toy worldText systemOwner
Photograph of textCanvas PreviewFast display rendered with the Scene
Writable original papertextarea / contenteditableCaret, Selection, Clipboard, IME, Spellcheck
Box of printing typeFontFace / document.fontsFont loading and readiness notifications
Ruler for lettersTextMetrics / DOM MeasurementWidth, baseline, actual bounds, and wrapping
One “visible character”Grapheme ClusterThe user-operation unit for Emoji and combining characters
Right-to-left bookBidi / RTLText algorithms handle visual and logical order
Input-method candidate cardComposition EventsThe value during composition is not the final Document value
Clear tape over the paperDOM Overlay TransformAlign with Shape/Camera and compensate for Zoom
Finish writing or tear it upCommit / CancelOne History Transaction or zero writes

The analogy has limits. The DOM Overlay and Canvas are not the same layout engine. Even with an identical font string, wrapping, fallback, hinting, and baseline can differ by pixels. A Grapheme is a unit perceived by users, not necessarily a linguistic character or word. contenteditable supplies rich browser behavior, but it is not a safe Rich Text data model; pasted HTML still requires sanitization and normalization.

Kill the wrong instincts first

  • “Canvas draws text, so we can draw a caret too.” You would still have to rebuild Selection, IME, Bidi, Clipboard, virtual keyboards, assistive technology, and platform conventions. The cost is far beyond a caret.
  • “Appending event.key on keydown gives maximum control.” Composition emits intermediate keys, an Emoji can contain multiple code points, and Selection is not necessarily at the end.
  • “text.length is the character count.” UTF-16 code units, Unicode code points, and Grapheme Clusters differ. A family Emoji has a length far greater than 1.
  • “Once the font URL loads, text is ready to measure.” The font may not be decoded, and fallback may already have participated in layout. Wait for FontFaceSet and invalidate after font changes.
  • “Canvas measureText wraps automatically.” It measures only the supplied inline text. Paragraph layout or DOM measurement must perform wrapping.
  • “Placing a textarea at the Shape’s x/y aligns it.” Parent, World, and Camera matrices, rotation, transform origin, padding, line height, and Zoom compensation are still missing.
  • “Store Rich Text as innerHTML.” That mixes untrusted HTML, browser-private markup, and the business model.

Production backpack

Prerequisite contract

Chapter 4 supplies the Shape Local→World→Screen matrix. Chapter 6 provides a DOMOverlayLayer Portal with symmetric cleanup. Chapter 9 has an editing State and Commit/Cancel. Chapter 10 prevents Selection and Resize from stealing input during Text Editing. A Document Text Shape stores text, width, style tokens, and direction—not a DOM node.

Formal knowledge

Font Loading is more than downloading a file. FontFace represents a loadable font, and document.fonts is a FontFaceSet. Before layout or measurement, you can call await document.fonts.load('16px "Canvas Sans"', sample) and invalidate after loadingdone. On failure, use an explicit Font Fallback chain and record the failure without blocking editing. When fallback changes to the target font, Text Metrics, wrapping, Shape height, and Connector anchors may all change. Recompute derived layout instead of silently changing Document content.

Canvas measureText() returns TextMetrics. In addition to width, some browsers expose actual bounding boxes, font bounding boxes, and baseline-related values. textBaseline selects a drawing anchor line; it is not a CSS line box. Primitive Text Rendering explicitly sets font, direction, textAlign, and baseline, then calls fillText() for each laid-out line. Canvas has no paragraph-wrapping API. A simple Text Shape can break lines at a deterministic width. Production Rich Text is better served by authoritative measurement in a hidden DOM or a mature layout engine, with the measurement version included in the cache key.

Line Breaking is not split on spaces. CJK scripts have no required spaces, and punctuation restrictions, soft hyphens, and long words need rules. Wrapping distinguishes hard breaks \n from soft wraps. Intl.Segmenter(...,{granularity:'grapheme'}) avoids splitting most Grapheme Clusters. If the target Browser Matrix lacks it, bundle a segmentation library verified against Unicode fixtures; Array.from(string) separates code points and is not a Grapheme fallback. Complete Unicode line breaking also requires a line-breaking library for the product’s supported languages or DOM layout. Test Emoji, combining marks, skin-tone modifiers, and ZWJ sequences as user-perceived units.

The Bidi algorithm determines the visual arrangement of mixed LTR and RTL text. The Document stores the logical string and does not reorder characters into screen order. A Shape has direction: auto|ltr|rtl; the Overlay uses dir and CSS direction, while Canvas sets context direction. Never reverse Arabic manually. Font Fallback can occur per Grapheme, so measurement must use the same CSS font shorthand as rendering.

IME coordinates compositionstart/update/end with beforeinput/input. The DOM value can change during composition, but intermediate steps must not each create History. compositionend is not the only possible commit moment either; when the editing transaction ends, Commit uses the Overlay’s current value. keydown handles only control intentions such as Escape and explicit shortcuts, checking event.isComposing during composition. Leave Caret, Selection, the local undo stack, and Clipboard to the native control. Define the boundary between Text Shape Undo and the control’s internal Undo: for example, Ctrl/Cmd+Z goes to the control while Editing and to Document History only after exit.

textarea is suitable for plain text: it is stable and natively supports Selection, IME, and mobile keyboards. contenteditable suits Rich Text, but requires a normalized model, Paste Sanitization, Selection mapping, and browser-matrix tests. Choose spellcheck, autocorrect, and autocomplete based on product and privacy requirements; never send user strings to your own telemetry. A Screen Reader needs an accessible name, editing state, and focus. Outside editing, provide a DOM Inspector or Fallback Semantics rather than exposing only “Canvas.”

When Textarea Overlay opens, freeze the original value, create the DOM, set value/dir/lang/spellcheck, apply the complete screen matrix, focus, and restore Selection. Update the matrix whenever the Camera or a Parent Transform changes; aligning only at construction is insufficient. A Zoom-Compensated Editor can increase the DOM font size by screen scale or transform its outer element with a matrix. In either case, DOM wrap width must correspond to the Canvas layout’s world width. Apply Rotation with a CSS matrix, not by rotating only the position. Commit reads the control value, validates maximum length and the Rich Text schema, then emits one UpdateText Command. Do not secretly apply NFC/NFKC normalization here unless the Document Schema specifies it and the product explains it to users, because normalization changes the actual code-point sequence. Cancel emits no Command. Exit removes listeners and the node, then maps focus back to Canvas or the Inspector.

Evidence and compatibility (verified 2026-08-29)

The engineering increment for this chapter

Starting point: A Text Label uses only Canvas fillText(), with no Caret, IME, or Clipboard. Finish line: Canvas Preview when not editing; a DOM Overlay exactly covers the Text Shape during editing. Commit updates the Document once, while Cancel performs zero writes.

Add these files:

  • src/engine/text/TextLayout.ts: font cache, Grapheme-safe wrapping, and baseline;
  • src/engine/text/FontManager.ts: FontFaceSet readiness and invalidation;
  • src/ui/overlays/TextEditorOverlay.ts: textarea lifecycle and matrix;
  • src/engine/shapes/TextShapeRenderer.ts: draw Preview from layout lines;
  • src/engine/text/__tests__/layout.test.ts: CJK, Emoji, RTL, and fallback;
  • tests/browser/text-ime.spec.ts: composition, selection, clipboard, rotation, and Zoom.

Start with a complete pure function for testable primitive layout. It respects hard breaks, Graphemes, and maximum width without using split('') to tear Emoji apart:

export type Measure = (text: string) => number;
export type TextLine = Readonly<{ text: string; width: number }>;

export function wrapGraphemes(
  value: string,
  maxWidth: number,
  measure: Measure,
  locale = 'und',
): readonly TextLine[] {
  if (!(maxWidth > 0)) throw new Error('TEXT_WIDTH_MUST_BE_POSITIVE');
  const segmenter = new Intl.Segmenter(locale, { granularity: 'grapheme' });
  const output: TextLine[] = [];
  for (const paragraph of value.split('\n')) {
    const graphemes = Array.from(segmenter.segment(paragraph), (part) => part.segment);
    if (graphemes.length === 0) {
      output.push({ text: '', width: 0 });
      continue;
    }
    let line = '',
      width = 0;
    for (const grapheme of graphemes) {
      const candidate = line + grapheme;
      const candidateWidth = measure(candidate);
      if (line && candidateWidth > maxWidth) {
        output.push({ text: line, width });
        line = grapheme;
        width = measure(grapheme);
      } else {
        line = candidate;
        width = candidateWidth;
      }
    }
    output.push({ text: line, width });
  }
  return output;
}

export async function layoutCanvasText(
  context: CanvasRenderingContext2D,
  value: string,
  cssFont: string,
  width: number,
  locale: string,
  onFontFailure: (reason: unknown) => void = () => undefined,
): Promise<readonly TextLine[]> {
  try {
    const faces = await document.fonts.load(cssFont, value || 'M');
    if (faces.length === 0) onFontFailure(new Error(`no face matched ${cssFont}`));
  } catch (cause) {
    onFontFailure(cause);
  }
  context.save();
  try {
    context.font = cssFont;
    return wrapGraphemes(value, width, (text) => context.measureText(text).width, locale);
  } finally {
    context.restore();
  }
}

The complete plain-text Overlay below does not listen for ordinary characters on keydown, does not write the Document during composition, and passes the Shape→Screen matrix directly to CSS:

export type TextShape = Readonly<{
  id: string;
  text: string;
  width: number;
  minHeight: number;
  font: string;
  fontSize: number;
  lineHeight: number;
  color: string;
  direction: 'auto' | 'ltr' | 'rtl';
  language: string;
}>;
export type Matrix2D = Readonly<{
  a: number;
  b: number;
  c: number;
  d: number;
  e: number;
  f: number;
}>;
export interface TextEditPort {
  commit(id: string, before: string, after: string): void;
  focusCanvas(): void;
}

export class TextEditorOverlay {
  private readonly textarea: HTMLTextAreaElement;
  private composing = false;
  private blurPending = false;
  private closed = false;

  constructor(
    host: HTMLElement,
    private readonly shape: TextShape,
    screenMatrix: Matrix2D,
    private readonly port: TextEditPort,
  ) {
    const area = document.createElement('textarea');
    this.textarea = area;
    area.value = shape.text;
    area.dir = shape.direction;
    area.lang = shape.language;
    area.spellcheck = true;
    area.setAttribute('aria-label', 'Edit canvas text');
    area.style.position = 'absolute';
    area.style.left = '0';
    area.style.top = '0';
    area.style.width = `${shape.width}px`;
    area.style.minHeight = `${shape.minHeight}px`;
    area.style.margin = '0';
    area.style.padding = '0';
    area.style.border = '1px solid currentColor';
    area.style.background = 'Canvas';
    area.style.color = shape.color;
    area.style.font = shape.font;
    area.style.lineHeight = String(shape.lineHeight);
    area.style.resize = 'none';
    area.style.transformOrigin = '0 0';
    this.updateScreenMatrix(screenMatrix);
    area.addEventListener('compositionstart', this.onCompositionStart);
    area.addEventListener('compositionend', this.onCompositionEnd);
    area.addEventListener('keydown', this.onKeyDown);
    area.addEventListener('blur', this.onBlur);
    host.append(area);
    area.focus({ preventScroll: true });
    area.setSelectionRange(area.value.length, area.value.length);
  }

  private onCompositionStart = () => {
    this.composing = true;
  };
  private onCompositionEnd = () => {
    this.composing = false;
    if (this.blurPending) this.commit();
  };
  private onKeyDown = (event: KeyboardEvent) => {
    if (event.isComposing || this.composing) return;
    if (event.key === 'Escape') {
      event.preventDefault();
      this.cancel();
    }
    if (event.key === 'Enter' && (event.metaKey || event.ctrlKey)) {
      event.preventDefault();
      this.commit();
    }
  };
  private onBlur = () => {
    if (this.closed) return;
    if (this.composing) {
      this.blurPending = true;
      return;
    }
    this.commit();
  };

  updateScreenMatrix(matrix: Matrix2D) {
    if (this.closed) return;
    this.textarea.style.transform = `matrix(${matrix.a},${matrix.b},${matrix.c},${matrix.d},${matrix.e},${matrix.f})`;
  }

  commit() {
    if (this.closed) return;
    if (this.composing) {
      this.blurPending = true;
      return;
    }
    this.blurPending = false;
    const after = this.textarea.value;
    if (new TextEncoder().encode(after).byteLength > 100_000) {
      this.textarea.setCustomValidity('Text is too long');
      this.textarea.reportValidity();
      this.textarea.focus({ preventScroll: true });
      return;
    }
    this.textarea.setCustomValidity('');
    this.closed = true;
    this.cleanup();
    if (after !== this.shape.text) this.port.commit(this.shape.id, this.shape.text, after);
    this.port.focusCanvas();
  }
  cancel() {
    if (this.closed) return;
    this.closed = true;
    this.cleanup();
    this.port.focusCanvas();
  }
  private cleanup() {
    this.textarea.removeEventListener('compositionstart', this.onCompositionStart);
    this.textarea.removeEventListener('compositionend', this.onCompositionEnd);
    this.textarea.removeEventListener('keydown', this.onKeyDown);
    this.textarea.removeEventListener('blur', this.onBlur);
    this.textarea.remove();
  }
}

Canvas Preview uses the same font/lineHeight and cached lines, with an explicit baseline. direction:auto cannot be assigned directly to Canvas direction, which accepts only ltr/rtl/inherit. Authoritative DOM measurement and layout first resolves a resolvedDirection for the paragraph, then Preview reuses that value. fontSize must also be an independently validated number. Calling parseFloat() on "600 16px Canvas Sans" would mistake font weight 600 for the font size:

export function renderTextPreview(
  context: CanvasRenderingContext2D,
  shape: TextShape,
  lines: readonly TextLine[],
  resolvedDirection: 'ltr' | 'rtl',
) {
  context.save();
  try {
    context.font = shape.font;
    context.fillStyle = shape.color;
    context.textBaseline = 'alphabetic';
    context.direction = resolvedDirection;
    const fontSize = shape.fontSize;
    if (!(fontSize > 0) || !Number.isFinite(fontSize)) throw new Error('INVALID_FONT_SIZE');
    lines.forEach((line, index) =>
      context.fillText(line.text, 0, fontSize + index * fontSize * shape.lineHeight),
    );
  } finally {
    context.restore();
  }
}

The unit tests prove that a Grapheme never breaks inside a ZWJ sequence. Real-browser tests provide the final evidence for IME:

import { describe, expect, it } from 'vitest';
import { wrapGraphemes } from '../TextLayout';

describe('text layout', () => {
  it('treats a family emoji as one wrapping unit', () => {
    const family = '👨‍👩‍👧‍👦';
    const lines = wrapGraphemes(
      `A${family}B`,
      2,
      (text) =>
        Array.from(new Intl.Segmenter('und', { granularity: 'grapheme' }).segment(text)).length,
    );
    expect(lines.map((line) => line.text)).toEqual([`A${family}`, 'B']);
    expect(lines.flatMap((line) => [...line.text]).join('')).toBe(`A${family}B`);
  });

  it('preserves hard line breaks, including empty lines', () => {
    expect(wrapGraphemes('あ\n\nい', 10, (text) => text.length).map((line) => line.text)).toEqual([
      'あ',
      '',
      'い',
    ]);
  });
});

Run pnpm vitest run src/engine/text. The CJK, Emoji, empty-line, width-boundary, and font-cache tests should pass. Run pnpm playwright test tests/browser/text-ime.spec.ts --project=chromium. During composition, the History revision should remain unchanged; ending composition and committing should increase it by exactly 1. Then run Clipboard, Selection, and RTL tests in Safari/WebKit and Firefox. On mobile, manually verify that the virtual keyboard does not push the Overlay permanently outside the viewport.

Return to the photograph game. Renderer displays the new photograph, the DOM control is the original paper temporarily laid over it, FontManager is the box of printing type, and Composition is a candidate card not yet chosen. Only when the child says “finished” do we take a new photo and update the ledger. Cancel simply removes the original paper.

Break it on purpose

InjectionSymptomEvidenceFixRegression testRecovery
Chinese Pinyin compositionCandidate repeats; one Undo per keyComposition/input/History traceNative control owns intermediate values; Commit once on exitChromium IME event sequenceCancel Overlay; Document unchanged
Japanese conversionEnter exits editing too earlyisComposing=true keydownIgnore commit shortcuts during compositionWebKit/Chromium fixtureRestore Selection
Family EmojiBackspace leaves broken symbol or wrap splits itGrapheme segmentsIntl.Segmenter or DOM layoutZWJ/skin-tone/flag fixturesRebuild layout cache
Arabic RTLCharacters reverse and Caret is misplacedLogical value and visual screenshotDo not reverse string; set dir/directionRTL mixed-number testRestore logical text
Font loads lateText jumps and Connector anchor is wrongfonts.status and before/after metricsInvalidate on loadingdoneDelayed-font routeClear layout cache
Target font replaces fallbackLine count changes but Bounds do notLine-count and metrics revisionRecalculate derived height and anchorsFont-swap testOne deterministic redraw
400% Zoom + RotationOverlay separates from textCompare CSS and Shape screen matricesFull matrix and transform-originMulti-Zoom screenshots at 0/37/90°Rebuild Overlay
Paste Rich HTMLScript or private styles enter modelClipboard MIME/schema logPlain-text policy or sanitizer plus typed rich modelMalicious-HTML fixtureReject Paste atomically
Mobile Virtual KeyboardInput is covered or focus is lostVisualViewport/focus traceVisible-region scrolling and focus policyReal iOS/Android devicesPreserve editing session and reposition

Pass with evidence

Claim to proveAutomated evidenceRequired manual/device evidence
Keydown construction does not break IMEEvent replay and one CommandReal Chinese/Japanese input methods
Graphemes do not splitSegmenter fixturesDelete with the system Emoji keyboard
RTL/Bidi preserves logical valueValue snapshotScreen Reader and Caret direction
Font changes recoverDelayed load and cache invalidationFallback on different operating systems
Overlay alignsScreenshot diff across Zoom and rotationSelection at 400% Zoom
Commit/Cancel boundaryRevision assertions of 1/0Blur, Escape, and virtual keyboard
  • FontFace, document.fonts, fallback, loading failure, and remeasurement have a protocol.
  • Text Metrics, Line Breaking/Wrapping, Baseline, and Canvas Preview consume explicit layout data.
  • Grapheme, Emoji, Bidi, RTL, IME, Composition, Caret, and Selection are owned by the correct layers.
  • The textarea/contenteditable choice has an ADR; Rich Text, Spellcheck, and Paste have security policies.
  • DOM Measurement and Overlay remain aligned after Zoom, Rotation, and Parent Transform changes.
  • Commit is one Command; Cancel is zero Document writes.
  • Ordinary text is not assembled manually from keydown, so IME, Selection, Clipboard, Composition, and Screen Reader behavior remains intact.

Explain it to a five-year-old

Answer without saying “DOM,” “Canvas,” “input method,” “grapheme,” or “bidirectional text”:

  1. Why can you see words in a photograph but not place a little vertical line between two letters?
  2. While the child is still choosing a candidate character, why can every key press not update the shared ledger?
  3. Why does the family picture contain many tiny parts but usually need to act as one whole?
  4. New question: the box of printing type arrives one minute late and a line of text becomes two lines. What must be measured again?
Show the reference answer A photograph remembers only colors, not insertion points or selected regions. Editing needs writable paper the browser understands. A candidate that has not been chosen is only a draft, so record it once after the child confirms or finishes editing. A family picture is made from many symbols glued into one image the eye sees, and the glue must not be cut through. New printing type has different widths, so recalculate which letters fit on each line, the cutout's height, the cord positions, and the displayed photo—but do not change the text itself.