@nextgensoftwares/folio-import
The import contract every format implements, the dispatcher, schema fitting, and the shared DOM-free XML reader. Guide: Import.
Contract
| Export | Description |
|---|---|
type Importer | { format, extensions, mimeTypes, sniff(src), import(src, opts?) } |
type ImportSource | { bytes: Uint8Array, name?, type? } |
type ImportOptions | upload? (store assets as found), signal?, onProgress?(0..1), limits? (maxBytes, maxEntries, maxBlocks), importer-specific keys |
type ImportResult | doc, format, page? (px), theme? (partial), meta? (title, author, language, created, modified), comments? (plugin-comments thread shape), watermark? (ImportWatermark), assets, warnings |
type ImportWatermark | a document's watermark (Word's WordArt/picture watermark), shaped like @nextgensoftwares/folio-plugin-watermark options: { enabled: true, text?: { content, fontFamily?, fontSize? }, image?: { src, width?, height? }, color?, opacity?, angle? }; apply with watermarkPlugin(...).replaceOptions |
type ImportAsset | { ref, name, type, bytes }; the document points at ref until stored |
type AssetUploader | (asset) => Promise<src> |
type ImportWarning | { code, message, count? } |
type ImportFormat | 'docx' | 'odt' | 'rtf' | 'html' | 'markdown' | 'text' | 'epub' | 'pdf' or a host id |
Dispatch and helpers
| Export | Description |
|---|---|
importFile(src, importers, opts?) | imports with the first importer whose sniff claims the file |
UnsupportedFormatError | no importer claimed it |
acceptOf(importers) | the file input accept string |
WarningLog | warnings de-duplicated by code (add(code, message), list()) |
fitToSchema(doc, schema, log?), type FitResult | degrades what a schema can't hold (substitutes, unwraps, drops unknown marks/attrs, resets invalid values, wraps stray inline content) and reports each loss |
ProseMirror / Tiptap JSON (lossless migration)
For moving stored ProseMirror or Tiptap documents to Folio (see the migration guide). Unlike fitToSchema, nothing is ever dropped silently.
| Export | Signature | Description |
|---|---|---|
fromProseMirrorJson | (json: unknown, o?: FromPmOptions) => FromPmResult<FolioDocument> | { doc, issues, ok }. Unknown nodes/marks and undeclared attrs are kept verbatim and reported (errors under onUnknown: 'reject', the default; warnings under 'keep'); invalid values, missing required attrs (a link without href), content mismatches, transient nodes (default ['windowSpacer']) and unresolved media are errors. Only noise is normalised: null attrs, empty text, adjacent equal-mark text, mark order, key order. Deterministic, byte for byte |
FromPmOptions | schema? (default standardSchema), onUnknown?, rename? (e.g. TIPTAP_RENAMES), mediaPlaceholder? (default 'mediaPlaceholder'), resolveMedia?: MediaResolver ((mediaId, node) => attrs), transient?, ids?: { seed } (runs ensureBlockIds) | |
PmIssue | { kind: PmIssueKind, severity: 'error' | 'warning', path, type, attr?, message } | path is a JSON pointer into the source; no document text |
PmIssueKind | unknown-node, unknown-mark, undeclared-attr, invalid-attr, missing-attr, invalid-content, mark-not-allowed, transient-node, missing-media, invalid-json | |
toProseMirrorJson | (doc, o?: ToPmOptions) => PmJson | the inverse: renames inverted, resizableMedia with mediaId and no src collapses to { type: mediaPlaceholder, attrs: { mediaId } }, strip attrs removed (e.g. ['id']). cssStrings: true writes numeric lineHeight as '1.15' and numeric lengths (CSS_LENGTH_ATTRS: text indent, paddings, paragraph spacing, fontSize, letterSpacing) as '12px', which Tiptap/CSS hosts expect; DOCX imports write line-height multiples as numbers (CSS_UNITLESS_ATTRS) |
canonicalize / canonicalJson | (json, Schema | CanonicalOptions) => unknown / string | the comparison form: key order, null/default attrs, empty text, merged text, mark order and strip attrs normalised; defaults adds source-editor defaults |
diffJson | (a, b, { canonical?, limit? }?) => JsonDiff[] | { path, kind: 'added' | 'removed' | 'changed', nodeType }, values never included |
mediaAttrsFromItem / manifestResolver | (item: MediaManifestItem) => Attrs / (items) => resolver | manifest item → resizableMedia attrs (type → mediaType, embed → iframe, provider → data-type, posterPath → poster, presentation attrs, naturalWidth/Height); src stays null (resolved by mediaId at render time) |
PmJson, FromPmResult, ToPmOptions, MediaResolver, CanonicalOptions, DiffOptions, JsonDiff, MediaManifestItem, TIPTAP_RENAMES | types and the inlineMath/blockMath → mathInline/mathBlock renames |
const r = fromProseMirrorJson(stored, { schema, resolveMedia: manifestResolver(manifest.items), ids: { seed: chapterId } });
if (!r.ok) return review(r.issues); // never partial
const back = toProseMirrorJson(r.doc, { strip: ['id'] });
const lossless = diffJson(stored, back, { canonical: { schema } }).length === 0;Header/footer zone configs
mapZoneConfig(config: ZoneConfig | null, o: ZoneMappingOptions): ZoneMapping maps a three-slot header/footer (ZoneSide: enabled, left/center/rightContent, showPageNumber, pageNumberPosition, differentFirstPage) to { config, variables, settings }: config for LayoutOptions.headerFooter, variables for LayoutOptions.variables, settings for doc.attrs.headerFooter (fields). A missing side shows the default chrome (header: title, "Page N"; footer: "N of M") unless defaultEnabled: false; the page number fills its slot only when the slot's own text is empty; is o.title (the chapter), (default: title), , are variables; is LayoutOptions.totalPages (pass the book total for global numbering); "different first page" means page number 1, so it applies only when o.firstPageNumber (default 1) is 1. Pure; merging a book config with a chapter override stays with the host.
@nextgensoftwares/folio-import/xml
A namespace-aware XML parser for untrusted parts: DOCTYPEs are skipped and never expanded, only predefined entities and character references are decoded, depth and element counts are capped.
| Export | Description |
|---|---|
parseXml(src, { maxDepth?, maxNodes? }) | root XmlElement (name, local, ns, attrs, attrNs, children) |
XmlError, decodeEntities | malformed input / entity decoding |
is, child, elements, descendant, descendants, attr, textOf, isElement | queries by namespace URI + local name |