Fonts & measurement
Layout never asks the browser how wide text is. It goes through an injected TextMeasurer, and production uses HarfBuzz (WASM) over the actual font files. This is what makes Folio's line breaks deterministic and lets the same layout run in Node for PDF export or server-side page counts.
The TextMeasurer contract
interface TextMeasurer {
width(text: string, font: FontSpec): number;
/** Ascent/descent as fractions of font size. Defaults to 0.8 / 0.2. */
metrics?(font: FontSpec): { ascent: number; descent: number };
/** px offset at every UTF-16 boundary of `text`, from one shaping pass. */
positions?(text: string, font: FontSpec): number[];
/** Whether the font's primary face has an OpenType feature ('smcp', 'c2sc'). */
hasFeature?(font: FontSpec, tag: string): boolean;
}
interface FontSpec {
family: string; // a CSS family list: 'Geist, Noto Sans Arabic'
size: number; // px
weight: number;
style: 'normal' | 'italic';
features?: string; // OpenType features to shape with: 'smcp', 'smcp,c2sc'
letterSpacing?: number; // px after every grapheme cluster (may be negative)
}A measurer must be pure: same input, same output. Layout caches by it. positions is optional but recommended: without it the editor measures text prefixes to place the caret, which is quadratic per line.
FontEngine
import { FontEngine } from '@nextgensoftwares/folio-fonts';
const engine = await FontEngine.create({ fallback: ['Noto Sans Arabic'] });
engine.addFont({ family: 'Geist', data: geistBytes }); // variable: wght axis read from the font
engine.addFont({ family: 'Noto Sans Arabic', data: arabicRegular });
engine.addFont({ family: 'Noto Sans Arabic', data: arabicBold }); // weight from OS/2 usWeightClass
engine.addFont({ family: 'My Serif', data: bytes, style: 'italic', weight: [300, 700] });
engine.width('Hello', { family: 'Geist', size: 16, weight: 400, style: 'normal' });FontEngine implements TextMeasurer (width, metrics, positions), so you pass it straight to layoutDocument and FolioEditor. HarfBuzz is loaded lazily and once (loadHarfBuzz()), keeping every other entry point synchronous.
How a face is chosen
The engine mirrors what a browser does:
- Family stack.
font.familyis parsed as a CSS list (parseFamilyList), then thefallbackfamilies are appended. - Style, then weight. Within a family,
matchFaceprefers the requested style (italic falls back to normal and vice versa), then ranks weights by CSS Fonts §5.2: inside a face's range is exact; for 400–500 try heavier up to 500, then lighter, then heavier; below 400 lighter first; above 500 heavier first (rankByWeight). - Per-grapheme fallback. If the primary face covers every code point, the text is one run (fast path, no segmentation). Otherwise each grapheme cluster goes to the first face in the stack that covers all its code points. Combining marks and joiners stick to the previous run, and spaces stay in the primary font, as Blink does. So
abc مرحبا xbecomesGeist "abc ",Noto Sans Arabic "مرحبا",Geist " x". - Variable fonts. Variable faces get an instance per requested weight (clamped to the axis range).
Line boxes
metrics(font) returns the primary face's ascent, descent and line gap (the first loaded family in the stack, or its alias), as CSS does for line boxes with a non-normal line-height. With theme.leading: 'below' (Word documents) the line box follows Word instead: the line gap sits above the text, a "multiple" line's extra height below it, an "at least" line puts the text at the bottom and an "exactly" line puts the baseline at 80% of the box; a run in a taller font grows the line to its own single height. Pass the same function to the view (<FolioView metrics>) so painted baselines land where layout put them.
Speed
- Widths are cached per (family, weight, style, text), up to
cacheSizeentries (default 100,000), then reset. - Layout measures segments between break opportunities and sums them, instead of reshaping whole lines.
- HarfBuzz buffers are reused across calls.
Missing fonts: stand-ins and lazy loading
Documents name fonts the host rarely has (a Word file in Verdana, Calibri, Times New Roman). Swapping them silently for the app's own font changes every line break, so @nextgensoftwares/folio-fonts resolves them in three steps:
- The family itself, from the host's
loadFamily(family)hook (a bundle, a font CDN, the document's embedded faces registered withaddFaces). - A stand-in from
fontFallbacks(family)(OFFICE_FALLBACKS): metric-compatible open fonts where they exist (Liberation Serif/Sans/Mono for Times New Roman/Arial/Courier New, Carlito for Calibri, Caladea for Cambria; same advance widths, so the same line breaks as Word), else the closest look-alike (DejaVu Sans for Verdana). The stand-in is registered as an alias (engine.addAlias('Verdana', 'DejaVu Sans')): the document keeps its family names, the engine measures with the stand-in, andonFaceregisters the same bytes under the original name for drawing. - Missing: reported (
FontReport.missing), drawn with the engine'sfallback. Importers turn the report into one warning each for missing families and for stand-ins.
import { createFontLoader } from '@nextgensoftwares/folio-fonts';
const loader = createFontLoader({
engine,
loadFamily: async (family) => myCdn.faces(family), // [{ family, data, weight, style }] or null
onFace: async (face, name) => { // draw what layout measured
const f = new FontFace(name, face.data, { weight: String(face.weight ?? 400), style: face.style ?? 'normal' });
document.fonts.add(f);
await f.load();
},
});
const result = await importDocx(bytes, {
// Embedded faces first, then the styles' families (before conversion), then what runs used.
resolveFonts: async ({ used, embedded }) => {
await loader.addFaces(embedded);
return loader.ensure(used);
},
lineMetrics: (family) => engine.singleLine(family), // line boxes from the faces that draw
});Each family is fetched once, all four faces (regular, bold, italic, bold italic) when the source has them, so bold and italic text is never synthesized by the browser while HarfBuzz measured the regular face. Licences: Liberation, Carlito, Caladea (SIL OFL) and DejaVu (Bitstream Vera licence) allow bundling unmodified files with their licence text; the playground bundles them under public/fonts/office/.
Small caps, all caps and letter spacing
Three textStyle attrs (Word's Small caps, All caps and character spacing) are laid out by the engine, so screen, PDF and DOCX agree:
textTransform: 'uppercase': the run is drawn in capitals. The document keeps the text as typed; text items carry the capitals (only case mappings that keep the UTF-16 length, so offsets stay ProseMirror offsets:ßstays).fontVariant: 'small-caps' | 'all-small-caps': when the font's primary face has the OpenTypesmcp(andc2scfor all small caps) feature (TextMeasurer.hasFeature), the run is shaped with it (FontSpec.features) and painters ask for the same glyphs (CSSfont-feature-settings, canvasfontVariantCaps, the PDF writer's HarfBuzz features). Otherwise small caps are synthesized: lowercase letters (every cased letter for all small caps) become capitals atSMALL_CAPS_SCALE(80%, as LibreOffice and Word draw them), split into their own text items; the kerning pair across each split (Ha→H+ smallA) is kept (kernon the left piece: the full-size pair's kerning scaled to the pair's mean size).letterSpacing: a length ('1.5pt', px number), may be negative. It lands onFontSpec.letterSpacing:width()adds it once per grapheme cluster andpositions()includes it (caret and hit-testing stay exact); the DOM painter sets CSSletter-spacing, canvasletterSpacing(per grapheme where the context lacks it), the PDF adds it after each cluster's last glyph.
Text without these attrs keeps the exact same code path and cache keys.
Drawing with the same fonts
In the browser, the painter draws text with ordinary DOM text at the positions layout computed. Register the same font bytes with the browser so glyphs match the measurements:
const face = new FontFace('Geist', bytes, { weight: '100 900' });
document.fonts.add(face);
await face.load();If the browser drew with a different font, text would overflow or leave gaps, but line breaks and pagination would not change: they come from HarfBuzz.
Fidelity against Chrome
tools/chrome-compare.mjs lays out every textblock of a corpus with Folio and with headless Chrome, using the same font files, and compares line starts and heights:
pnpm build
PUPPETEER_FROM=/path/with/node_modules/puppeteer node tools/chrome-compare.mjs chapters.ndjsonThe input is NDJSON, one {"c": <ProseMirror doc>} per line; the tool adds a synthetic Arabic / mixed-direction / justified stress set. Current result: 205/206 blocks identical, heights within 0.01 px. The single miss is a line ending within half a pixel of the margin, where Chrome floors advances to 1/64 px.
Math and media
Two more measurers are injected the same way:
MathMeasurer.measure(latex, { display, font })returns{ width, height, depth }. In the browser, a host can render KaTeX off-screen and read the box (the playground'screateKatexMeasurer); load KaTeX's fonts first or boxes are measured in a fallback font. Without one, layout uses a rough estimate.MediaSizer(attrs)returns the intrinsic{ width, height }of a media node (e.g. from upload metadata or a media manifest).@nextgensoftwares/folio-plugin-mediashipsmediaSizer, which readsnaturalWidth/naturalHeight.