Import
Folio opens Word, OpenDocument, RTF, HTML, Markdown, plain text, EPUB and PDF files. Importing has three layers, so a host only pays for what it uses:
| Layer | Package | What it does |
|---|---|---|
| Contract | @nextgensoftwares/folio-import | Importer, ImportResult, importFile (pick the importer that claims a file), acceptOf, WarningLog, fitToSchema, and the shared XML reader @nextgensoftwares/folio-import/xml |
| Formats | @nextgensoftwares/folio-docx (Word) and @nextgensoftwares/folio-import-formats (everything else) | one importer per format, one entry per format (@nextgensoftwares/folio-import-formats/markdown, /html, /text, /odt, /rtf, /epub, /pdf) |
| Applying + UI | @nextgensoftwares/folio-plugin-import | applyImport (replace or insert, one undo step, progressive), toolbar items, drop handling, progress and warnings dialog |
Importers are DOM-free: they run in the browser, in a worker and in Node (a server can convert uploads with the same code). They never produce markup the editor would execute: scripts, styles, event handlers, frames and unsafe URLs are dropped, and only known node types with validated attributes come out.
Quick start
import { importFile, acceptOf } from '@nextgensoftwares/folio-import';
import { allImporters } from '@nextgensoftwares/folio-import-formats';
import { docxImporter } from '@nextgensoftwares/folio-docx';
import { applyImport, assetUploaderFrom } from '@nextgensoftwares/folio-plugin-import';
const importers = [docxImporter, ...allImporters()]; // stricter (magic-byte) importers first
const result = await importFile({ bytes, name: file.name, type: file.type }, importers, {
signal: controller.signal,
onProgress: (f) => bar.value = f,
});
// result.doc, result.page, result.theme, result.meta, result.comments, result.assets, result.warnings
await applyImport(editor, result, {
mode: 'replace', // or 'insert' at the caret
theme, setTheme, // page size/margins and text size go through the host's theme
upload: assetUploaderFrom(mediaUploader), // where images go (else data: URLs)
comments: { store, actor }, // plugin-comments' store
});allImporters() returns lazy importers: sniffing a file and building the file picker's accept cost nothing, and a format's parser is loaded (import()) the first time a file of that format comes in. Hosts that offer only some formats import those entries directly:
import { markdownImporter } from '@nextgensoftwares/folio-import-formats/markdown';
import { htmlImporter } from '@nextgensoftwares/folio-import-formats/html';
const importers = [htmlImporter, markdownImporter];The import UI
import { importPlugin, assetUploaderFrom } from '@nextgensoftwares/folio-plugin-import';
const importer = importPlugin({
importers: () => [docxImporter, ...allImporters()],
// "Open file…" / "Open as a new document": the host makes a new document.
openDocument: (prepared) => openAsNewDocument(prepared.doc, prepared.theme, prepared.comments),
// Settings for applying into the current editor ("Import…", "Insert here").
applyOptions: () => ({ theme: currentTheme, setTheme, upload: assetUploaderFrom(uploader), comments: { store, actor } }),
// Offer "Attach as file" for dropped documents (plugin-media keeps doing this for everything else).
attach: (editor, files, at) => media.input.insertFiles(editor, files, at),
// Schema additions importers rely on: DOCX bookmark ids, so internal links survive.
schema: docxImportSchema,
});
const plugins = [importer, media /* after import: it sees the drops import declines */];- Toolbar: "Open file…" (
import.open) and "Import…" (import.insert), groupfile. Route them anywhere like any item (a File menu, the mobile bar):route: { 'import.*': ['fileMenu'] }. - Drop: dropping an importable file (by extension or MIME type) on a page asks Open as a new document / Insert here / Attach as file. Images, video, audio and unknown files are not touched: plugin-media still turns them into media or file cards.
drop: falseturns the question off. - Progress: a dialog in
<FolioView>'soverlayslot shows reading, converting and applying with a cancel button. Cancelling while applying reverts what was inserted. - Warnings: "Imported with 3 notes", expandable to the list of what couldn't be carried over (de-duplicated, with counts).
Without openDocument, "Open file…" replaces the current document in place, as one undo step.
Applying a result
applyImport(editor, result, options):
- stores assets the importer didn't upload (
upload, a few at a time; without one they becomedata:URLs; failures drop the image with a warning) and rewrites the document'sasset:refs; - fits the document to the editor's schema (
fitToSchema): node types the host lacks are degraded (callouts → quotes, unknown containers unwrapped, fields → text), unknown marks and attributes removed, invalid values reset; every loss is a warning; - replaces the document or inserts at the selection (merging like a paste) with
applyProgressively: batches the page can paint between, one undo step, cancel reverts. In replace mode the document attributes travel too (header/footer settings, chapters); - calls
setThemewith the import's page size and margins and its theme hint: body size, paragraph spacing, line height, heading, list, table and quote styles (replace mode by default;applyPage,adoptTheme,adoptTextSize; fonts only withadoptFonts, since the host may not have them); - creates comment threads in the store (as the importing user, the original author's name kept in the text).
prepareImport does steps 1, 2 and 4 without an editor, for hosts that open imports as new documents (create the editor from prepared.doc).
Limits and safety
Every importer takes opts.limits (maxBytes, maxEntries for archives, maxBlocks), signal and onProgress, and yields to the event loop every ~30 ms on big files. Archives are read with caps per entry and in total (zip-bomb guard); XML is parsed without DTDs or entity expansion (no XXE, no "billion laughs"); images must be real PNG/JPEG/GIF/WebP/BMP (SVG is never imported as an image); links keep only http(s), mailto and tel.
| Format | Default maxBytes | Other caps |
|---|---|---|
| HTML | 100 MB | 2,000 zip entries (Google Docs zip) |
| Markdown | 50 MB | |
| Plain text | 100 MB | |
| ODT | 200 MB | 5,000 zip entries |
| RTF | 100 MB | 1,000 nested groups |
| EPUB | 200 MB | 10,000 zip entries |
| 200 MB | maxPages 2,000, 500 images |
All formats: maxBlocks 500,000 (the rest is cut with a warning).
Word (.docx)
@nextgensoftwares/folio-docx (docxImporter) reads .docx, .dotx, .docm and .dotm (macros are never read). Its goal is a book written in Word laying out in Folio as close to Word as possible. Word isn't available on Linux, so LibreOffice is the reference renderer for the fidelity numbers below.
- Styles: the full cascade (document defaults → table style → paragraph style chain → numbering level → direct formatting; character styles flip toggle properties). Normal and Heading 1–6 become
result.theme(fonts with metric-compatible fallbacks, sizes, colours, line height, paragraph spacing, heading spacing and keep-with-next); paragraphs and runs carry only what differs: alignment, direction, line height, indents (negative ones reach into the margin), tab stops (w:tabsthrough the style cascade, clears and leaders;w:tabruns stay\t;defaultTabStop→theme.tabInterval; bar tabs are dropped with a warning),spaceBefore/spaceAfter, keep/page-break controls, and bold/italic/underline/strike/ sub/sup/code/colour/size/family/highlight marks, plus all caps, small caps and character spacing (w:caps,w:smallCaps,w:spacing→ textStyletextTransform/fontVariant/letterSpacing; the text stays as typed). Outline levels → headings (set on a paragraph only →visualHeading), Title/Subtitle stay styled paragraphs, quote styles → blockquotes, code styles and monospace paragraphs → code blocks. - Lists:
numbering.xmllevels, overrides, starts, formats, restart and continue (instances of one definition share counters); flat numbered paragraphs become nested lists (level, then indent), interrupted lists continue throughstart. Numbered headings keep their number as text. Each level's marker becomes the list'slistLevel: bullet characters in their font, colour and size (Symbol/Wingdings mapped to Unicode), picture bullets (numPicBullet, as image assets),lvlTexttemplates ("1.1.", "Chapter 1:", "(a)", "1-") with each level's format,lvlJc, and the level's indent/hanging when they differ from the theme's. Levels outside the list's own nesting (e.g. a numbered heading's "2" in "2.1") are baked into the template. Still approximated (warning): number formats with no Folio equivalent (CJK counting, Hebrew...) as decimal, symbol-font characters with no known Unicode equivalent (kept in their font). - Tables: grid widths (
colwidth),gridSpan,vMerge, cell shading, repeated header rows, nested tables; table width (tblW: twips, percent, auto → the grid's width), alignment (jc) and indent (tblInd); table and cell borders (tblBorders/tcBorders, the table style chain's too: single, double, dashed, dotted, none, widths and colours; borderless tables stay borderless); table styles' conditional formatting (tblStylePrheader row, last row, first/last column, banded rows/columns pertblLook) resolved into cell fills, borders and run formatting; cell margins (tblCellMar/tcMar) and vertical alignment (vAlign). Fancy border styles (wave, 3-D, thin-thick) become the closest line (warning). - Pictures: inline and floating DrawingML/VML pictures with Word's wrap (square, tight → square, top-and-bottom, behind, in front), side, distances and offsets → plugin-media placement. EMF/WMF/TIFF go through
convertImageor stay assets with a placeholder. Embedded files (an .xlsx) become file attachments; OLE containers show their preview picture. Text boxes → outline callouts (or paragraphs); groups, charts and SmartArt keep their pictures or fallbacks, with a warning. - Sections: the first section is
result.page(size, orientation, margins, gutter); headers/footers (default, first, even; link to previous; PAGE/NUMPAGES/SECTIONPAGES/DATE/TITLE/STYLEREF fields as live field atoms) and page numbering become plugin-header-footer settings; later sections are section breaks (continuous ones stay on the page) with their own page size, orientation, margins (gutter on the left) and columns (w:cols: count, spacing, separator line, unequal widths, each column laid out at its own width). Column breaks becomecolumnBreakBefore; odd/even page starts becomepageBreakBefore: 'odd' | 'even'(layout adds the blank page Word would). - Fields: TOC → plugin-toc's table of contents (rebuilt by Folio, titled by its "TOC Heading"), HYPERLINK and REF/PAGEREF
\h→ links, the rest keep their cached result. Bookmarks become blockids and internal links#id(declaredocxImportSchemain the editor so they survive). - Footnotes and endnotes → numbered superscript links and a "Notes" section with back-links (a pluggable
NoteMappingfor a future footnote plugin). Comments (with replies and done state) → plugin-comments threads andcommentmarks. Tracked changes are accepted (or rejected withtrackedChanges: 'reject') and reported. Equations (OMML) → LaTeX for plugin-math. Embedded fonts are deobfuscated intofont/ttfassets (result.fonts.embedded);availableFontsreports missing ones.
Safety: archive caps (maxBytes 400 MB uncompressed, maxEntries 20,000, maxBlocks 500,000, XML depth 256), declared sizes checked before inflating; external relationships are never fetched (only hyperlinks are kept, and only http(s)/mailto/tel/ftp); no DTD or entity expansion.
Performance: document.xml is decoded and parsed in 1 MB chunks and each body block is converted and dropped, so memory never holds the XML tree. A ~500-page book imports in ~0.1 s in Chrome; a 22 MB real document (66 MB of XML, 920 pictures) in ~1.8 s. Big files: importDocxInWorker.
- Fonts: with
resolveFonts(e.g. @nextgensoftwares/folio-fonts'createFontLoader), the document's embedded faces and its styles' families load before conversion (solineMetricsreads the faces that draw), the rest after; families the host has no face for get a metric-compatible stand-in (OFFICE_FALLBACKS) and are reported, one warning for missing families, one for stand-ins.w:strikeandw:dstrikeare separate toggles. - Line positions: imports set
theme.leading: 'below'(wordLeading): the font's line gap above the text, extra "multiple" spacing below it, "at least" lines (lineHeightpx +lineRule: 'atLeast', a minimum a bigger run can't scale) with the text at the bottom, "exactly" lines with the baseline at 80%; a run in a taller font grows its line. An empty paragraph's line is sized by its paragraph mark (w:pPr/w:rPr), as Word and LibreOffice do. Heading styles map by the level paragraphs become (outline level first), so converters' "heading 3" at outline level 4 keeps its look; list paragraphs keep their right indent; list markers sit left-aligned at the hanging indent (listLevel.hanging). - Pictures in a line of text become inline
imageatoms (on the baseline, in the line, exported back as Word inline pictures); a picture alone in its paragraph stays a block whose distances are the paragraph's spacing plus the text's descent. Square-wrapped pictures keep their vertical offset (offsetY). Table rows keepw:trHeight(height,heightRule: 'exact'). - Headers and footers: each section's
w:pgMar/@w:headerand@w:footer(headerDistance/footerDistanceper section), parts inherited from the previous section when a section names none (never backwards: a cover section without parts shows none),titlePgper section; text boxes are imported once (DrawingML, not the VML fallback) and fields inside them (a page-anchored "Page N" box) get each page's value. A part with no content reserves no band.
Fidelity against LibreOffice (node tools/fidelity/import/run.mjs [file.docx] [--chunks N] [--fonts DIR]; --chunks compares each slice of N sections on its own, --fonts makes LibreOffice and Folio substitute Office fonts with a bundle; big files LibreOffice can't export directly go through ODT): prose, paragraph spacing, lists and tables match page counts with line positions within ~1 px. A 155-page converted e-book (149 sections): 155 pages in both, median word offset 0.9 px (was 158 pages, 95 px); with the playground's bundled stand-ins 151 vs LibreOffice's 152 (was 158). Known differences: a table break inside a row-spanning cell is avoided (Word and LibreOffice split there); fields in text boxes of the body keep their cached text; LibreOffice draws Arabic in Times New Roman with FreeSerif while hosts draw their own fallback.
HTML
.html, .htm, .xhtml, and Google Docs' "Download → Web page" zip (HTML plus images/). Parsed with htmlparser2 (MIT) into a light tree, no DOM, so it also runs in workers and Node.
- Headings, paragraphs, line breaks, block quotes, rules,
<pre>/<code>(language fromlanguage-xclasses → plugin-code), lists (nested, start, type,list-style-type; task checkboxes become ☐/☑ with a warning), definition lists, figures and captions. - Tables with
colspan/rowspan, header cells, background colours and pixel widths; captions become a centred paragraph. - Bold/italic/underline/strike/sub/sup/code/mark, links, and safe inline styles: colour, highlight, font family and size, alignment, direction, indents, line height, page breaks. Black text, white backgrounds and the font used by most of the text are dropped (they would pin the document to the source's look); the body font and size are reported in
result.theme. - Images:
data:URLs and zip entries become assets;http(s)images keep their URL (warning); local paths are dropped. Inline images move out of their paragraph (Folio images are blocks). - Math: KaTeX/MathJax output and MathML with a TeX annotation, and
data-latexelements, become equations (plugin-math). - Folio's own copies (
data-folioattributes) round-trip their layout attrs.
Google Docs
Three ways in, from best to quickest:
- File ▸ Download ▸ Microsoft Word (.docx) →
@nextgensoftwares/folio-docx(styles, comments, headers/footers, footnotes). - File ▸ Download ▸ OpenDocument (.odt) → the ODT importer, or Web page (.html, zipped) → the HTML importer (class styles, re-nested lists, unwrapped
google.com/url?q=links, images from the zip, Title/Subtitle as headings). - Copy and paste: the editor's HTML paste handles the clipboard flavour (the
<b id="docs-internal-guid-…">wrapper'sfont-weight:normalcancels its bold).
Markdown
CommonMark + GFM through markdown-it (MIT; linear-time, ~10 MB in under 2 s) with markdown-it-footnote and small built-in rules for math: headings, emphasis, strikethrough, inline code, links, images, <https://…> autolinks (bare URLs aren't linked), lists (task items as ☐/☑ with a warning), tables with column alignment, block quotes, rules, fenced code (language → plugin-code), math $…$ / $$…$$ (plugin-math), footnotes and inline ^[…] notes (as endnotes), YAML front matter (title, author, lang, date → meta). Raw HTML is never rendered: formatting tags are applied, everything else keeps only its text.
Plain text
Encoding detection: BOM (UTF-8, UTF-16 LE/BE), BOM-less UTF-16, strict UTF-8, then Windows-1252 (reported). Blank lines separate paragraphs; single line breaks stay as line breaks.
OpenDocument (.odt)
.odt, .ott and flat .fodt (LibreOffice, Google Docs, Collabora, OnlyOffice exports). The style cascade (defaults → parent chain → common → automatic styles) drives headings (outline level, Title/Subtitle and "Heading N" styles), paragraph alignment/indents/line height/keeps/page breaks, and character formatting. Also: lists (bullet/number per level, start, continue), tables (spans, covered cells, repeated rows/columns, column widths, backgrounds), images and text boxes (wrap → Folio placement: inline, square, top-and-bottom, behind, in front), the first page style (size, margins → result.page), headers/footers (default, first, even; page number/count, date, title fields) → plugin-header-footer, footnotes and endnotes → endnotes, comments with replies and ranges → plugin-comments, the table of contents → plugin-toc, meta.xml → meta. Charts, OLE objects and shapes are reported (shapes keep their text); tracked changes import as displayed.
RTF
A built-in tokenizer (groups, control words, \'hh with the document and per-font code pages, \u with \uc skipping, \bin). Paragraph and character formatting, font and colour tables, heading styles and outline levels, lists (Word list tables, WordPad \pn, nesting), tables (cell widths, horizontal and vertical merges, backgrounds), PNG/JPEG pictures, hyperlinks, footnotes (endnotes), headers/footers with PAGE/NUMPAGES fields, page size and margins, document info. WMF/EMF pictures, comments, text boxes and equations are reported.
EPUB
EPUB 2 and 3: container → OPF → spine; each chapter goes through the HTML importer and becomes an H1 section in reading order (titles from the nav or NCX when a chapter has no heading of its own; inner headings shift below it). The cover becomes the first image, a table of contents (plugin-toc) follows it, and every chapter starts on a new page. Metadata → meta. DRM-encrypted books are refused; font obfuscation is fine. Book CSS and fonts aren't imported.
PDF (reconstruction)
Best effort
A PDF has no paragraphs, headings or lists, only positioned text. The PDF importer rebuilds structure from the layout and always says so in a warning. It is never pixel-faithful: line breaks, spacing, fonts and page breaks follow your theme, not the PDF. There is no OCR.
pdf.js (Apache-2.0, an optional peer dependency, loaded on first use with scripting, XFA and network access off) extracts text with positions and fonts. Then: lines by baseline, reading order across columns, paragraphs (gaps, indents, font changes; joined across page and column breaks), de-hyphenation, headings by size relative to the body text (and short bold lines), bold/italic from font names, superscript footnote markers, bullet and numbered lists, links from annotations, images (decoded pixels re-encoded as PNG, placed in reading order), repeated headers/footers and page numbers removed, page size and estimated margins. Scanned PDFs (no text layer) get a clear warning and only their images. Tables come out as paragraphs. Pass pdfWorker: () => new Worker(...) to parse off the main thread.
Writing an importer
An importer is a plain object; defineImporter from @nextgensoftwares/folio-import-formats adds the context (limits, warnings, assets, abort, progress):
import { defineImporter } from '@nextgensoftwares/folio-import-formats';
export const csvImporter = defineImporter(
{ format: 'csv', extensions: ['csv'], mimeTypes: ['text/csv'], sniff: (s) => /\.csv$/i.test(s.name ?? '') },
async (src, ctx) => {
const rows = parse(new TextDecoder().decode(src.bytes));
await ctx.tick(0.5); // yields, honours abort, reports progress
if (rows.some((r) => r.length > 50)) ctx.warn('wide', 'Columns past 50 were dropped.');
return ctx.result('csv', [tableOf(rows)]);
},
{ maxBytes: 20 * 1024 * 1024 },
);