Document model & schema
ProseMirror-shaped JSON
A Folio document is plain JSON in the ProseMirror shape. There are no classes and no hidden state, so a document serializes without loss, survives a postMessage round-trip, and existing ProseMirror/Tiptap content loads without migration.
interface FolioNode {
type: string;
attrs?: Record<string, unknown>;
content?: FolioNode[];
text?: string; // only on `text` nodes
marks?: Mark[]; // only on inline nodes
}
interface FolioDocument extends FolioNode {
type: 'doc';
content: FolioNode[];
}{
"type": "doc",
"content": [
{ "type": "heading", "attrs": { "level": 1 }, "content": [{ "type": "text", "text": "Title" }] },
{ "type": "paragraph", "attrs": { "dir": "rtl", "textAlign": "justify" }, "content": [
{ "type": "text", "text": "نص ", "marks": [{ "type": "bold" }] },
{ "type": "mathInline", "attrs": { "latex": "x^2" } }
] }
]
}A node is addressed by its path of child indices from the root: [3, 0, 1] means doc.content[3].content[0].content[1]. Layout fragments carry the path of the node they came from, so anything on a page maps back to the document.
Nodes are immutable
This is the one rule every host and plugin must follow. An edit replaces the changed node and its ancestors with new objects; untouched subtrees stay the very same objects. Layout caches block measurements by object identity, so:
- unchanged blocks are never re-measured, and
- mutating a node in place serves stale layout.
The editor guarantees this for you (DocBridge converts ProseMirror nodes to Folio JSON memoized per node). If you build documents yourself, copy on write.
The standard schema
standardSchema is a superset of what a typical book editor needs: the StarterKit nodes and marks, tables, code, math, media and text styles, plus Word-native pagination attributes.
| Kind | Types |
|---|---|
| Blocks | paragraph (visualHeading), heading (level 1–6), blockquote, bulletList, orderedList (start, type), listItem, codeBlock (language), mathBlock (latex), resizableMedia, horizontalRule, table, tableRow, tableCell, tableHeader |
| Inline | text, hardBreak, mathInline (latex), image (src, alt, title, width/height px: a picture in a line of text, bottom on the baseline) |
| Marks | bold, italic, underline, strike, code, subscript, superscript (mutually exclusive), link (href, target, rel, class, title), textStyle (color, fontFamily, fontSize, fontVariant (small-caps / all-small-caps), textTransform (uppercase), letterSpacing (a length, may be negative)), highlight (color) |
Shared ("global") attributes:
| Attribute | On | Values |
|---|---|---|
textAlign | paragraph, heading | left, center, right, justify, start, end |
dir | paragraph, heading, lists, list items, quote, code, cells | ltr, rtl |
lineHeight, paddingInlineStart, paddingInlineEnd, textIndent | paragraph, heading | CSS-ish lengths (1.5, 150%, 24px, 12pt, 2em); indents may be negative (into the page margin, clamped at the paper edge) |
tabStops | paragraph, heading | [{ pos, align?, leader? }]: pos px from the paragraph box's start edge, align left (default) / center / right / decimal (logical: start/end in RTL), leader none / dot / hyphen / underscore / middleDot. Tabs are \t characters in text; without stops they advance to multiples of theme.tabInterval |
listStyleType | bulletList, orderedList | disc, circle, square, decimal, decimal-leading-zero, lower-/upper-alpha, lower-/upper-latin, lower-/upper-roman, lower-greek, none |
listLevel | bulletList, orderedList | the list's own marker (ListLevel, Word's numbering level): format (a listStyleType, lower-letter/upper-letter, ordinal, arabic-indic, bullet, none), text (bullet character or template: %1.%2., Chapter %1:, (%1); %k = the k-th enclosing list's counter), font, color, size (px), bold, image + imageWidth/imageHeight (picture bullet), indent, hanging (px), align (start/center/end). Null = listStyleType as before. It sits on the list node, not in a document numbering table, so a list lays out from its own node and ancestors (exact identity cache, incremental layout) |
pageBreakBefore, keepWithNext, keepLinesTogether | paragraph, heading | Word's w:pageBreakBefore, w:keepNext, w:keepLines |
columnBreakBefore | paragraph, heading | Word's column break: starts the next column (a new page in one-column text) |
spaceBefore, spaceAfter | paragraph, heading | Word's space before/after (lengths; null = the theme's). Override the theme margins, also on empty and list paragraphs; an explicit space before survives the document start and a forced page break. Written by the DOCX importer |
Table cells take colspan, rowspan, colwidth (array of px), backgroundColor, borders ({top, bottom, start, end} of {style: 'single'|'double'|'dashed'|'dotted'|'none', width (px), color}), padding ({top, bottom, start, end} px) and verticalAlign (top/middle/bottom). Tables take width (px, "NN%" or "auto"; null = full width), align (left/center/right/start/end), indent (px from the start side), borders (top, bottom, start, end, insideH, insideV), cellPadding, tableStyle (a built-in style: grid, plain, lines, banded, accent, boxed, or a theme-accent light/medium/dark with -2…-6), look (headerRow, firstColumn, lastRow, lastColumn, bandedRows, bandedColumns) and repeatHeader (see Table styling). All are optional; without them a table is Folio's full-width grid. resizableMedia takes src, mediaId, alt, title, mediaType (image / video / audio / iframe), width, height, alignment, float, borderRadius and objectFit.
Defining a schema
A schema is a SchemaSpec of nodes and marks. Content expressions use a subset of ProseMirror's syntax: names or (a|b) groups with *, + or ?.
import { Schema } from '@nextgensoftwares/folio-model';
const notes = new Schema({
nodes: {
doc: { content: 'block+' },
paragraph: { group: 'block', content: 'inline*' },
note: { group: 'block', content: 'paragraph+', attrs: { tone: { default: 'info' } } },
text: { group: 'inline', inline: true },
},
marks: { bold: {}, italic: {} },
});NodeSpec fields: group, content, inline, atom (a leaf treated as one unit, like math or media), marks ('_' = all, the default for textblocks; '' = none; or a space-separated list), code (preformatted, no marks) and attrs. An AttrSpec without default is required; validate rejects bad values.
Extending the standard schema
Hosts and plugins add types without forking:
const schema = standardSchema.extend({
nodes: { callout: { group: 'block', content: 'block+', attrs: { tone: { default: 'info' } } } },
globalAttrs: [{ types: ['resizableMedia'], attrs: { naturalWidth: { default: null }, naturalHeight: { default: null } } }],
});extend() returns a new Schema; the original is untouched.
Declare every attribute you store
The editor builds a ProseMirror schema from your Folio schema, and ProseMirror drops undeclared attributes. If a plugin stores naturalWidth on media, it must declare it (as above, or as @nextgensoftwares/folio-plugin-media does).
Validation and normalization
import { normalize, standardSchema, validate } from '@nextgensoftwares/folio-model';
const issues = validate(doc, standardSchema);
// [{ path: [1], message: 'heading: invalid level=9' }, …] never throws
const clean = normalize(doc, standardSchema);validatechecks node types, required and invalid attributes, content expressions, unknown and disallowed marks, and empty text nodes. It returns every issue with its path.normalizereturns a cleaned copy: default attributes filled in, empty text nodes dropped, adjacent text nodes with identical marks merged, disallowed marks removed. It doesn't repair structure; usevalidatefor that.
For big documents, validate per top-level block and cache by node identity, as the playground does: an edit then re-validates only the blocks it touched.
Helpers
attr(node, name, fallback), nodeAt(root, path), textContent(node), marksEqual(a, b), isTextNode(node), parseContentExpr(source) and matchContent(expr, childTypes, fits). See the API reference.