Loading & saving
@nextgensoftwares/folio-sync loads big documents progressively and saves edits, not documents. A save costs O(edit) no matter how big the book is, and the server never trusts a client to be implemented correctly.
The protocol
Reads are chunked and content-addressed. A small manifest describes the document at a version as an ordered list of chunks: runs of top-level blocks, cut at chapter boundaries. A chunk's id is the hash of its content, so its URL never changes meaning: HTTP caches, CDNs and the client's IndexedDB can keep it forever, and after an edit only the chunks whose hash changed are fetched again.
Writes are ProseMirror steps tagged with the version they apply to (the central-authority collaboration protocol). The server applies them and every other client pulls and rebases.
interface DocumentApi {
manifest(docId: string): Promise<Manifest>;
chunk(docId: string, ref: ChunkRef): Promise<Chunk>;
push(docId: string, version: number, steps: unknown[], clientID: string, opts?: PushOptions): Promise<PushResult>;
pull(docId: string, since: number): Promise<PullResult>;
reportLayout?(docId: string, report: LayoutReport): Promise<void>;
}The REST mapping used by createHttpApi({ baseUrl, headers }):
| Route | Returns | Caching |
|---|---|---|
GET /docs/:id/manifest | Manifest | no-cache, ETag |
GET /docs/:id/chunks/:hash | Chunk | immutable (client uses force-cache) |
POST /docs/:id/steps { version, steps, clientID, checksum, confirm? } | PushResult | 409 = behind, 422 = refused |
GET /docs/:id/steps?since=N | PullResult | no-store |
PUT /docs/:id/layout LayoutReport | 204 | optional |
The wire format is compact (attributes equal to their schema default are omitted; the client's schema fills them back), about 2× smaller.
Pages are an addressing layer
Page boundaries move with every edit, theme change and font change, so storing a document by page would re-split and invalidate on every keystroke. Folio stores by chunk and treats pages as addresses: after a client lays out the document it sends a layoutReport(manifest, layout) (which page each chunk starts on). The server attaches those page ranges to the next manifest (pageStart, pages, estimatedPages), so the next client can size its scrollbar before loading anything and open at any page: chunkForPage finds the chunk, and since chunks start at chapters, that chunk paginates on its own.
Progressive open
import { assembleDocument, IndexedDbStore, openDocument, SyncClient, createHttpApi } from '@nextgensoftwares/folio-sync';
const api = createHttpApi({ baseUrl: 'https://api.example.com', headers: () => ({ Authorization: `Bearer ${token}` }) });
const cache = new IndexedDbStore();
const outbox = await cache.getOutbox(docId);
const session = await openDocument(api, docId, { cache, ...(outbox ? { manifest: outbox.manifest } : {}) });
const { ready, complete } = assembleDocument(session, {
schema, measurer, layout, // the usual FolioEditor options
clientID,
...(outbox ? { outbox } : {}),
onProgress: (loaded, total) => setProgress(loaded / total),
});
const editor = await ready; // usable after the first chunk
showEditor(editor);
await complete; // everything (and any offline outbox) is in
const sync = new SyncClient(editor, api, docId, clientID, {
manifest: session.manifest,
outbox: cache,
...(outbox ? { confirmed: { steps: outbox.confirmed, clientIDs: outbox.confirmedClientIDs } } : {}),
}).start();
sync.subscribe((status) => render(status)); // synced | saving | offline | error | needs-confirmation | diverged- Manifest first (tiny). If the network is down, the cached manifest is used and the document opens offline (
session.offline). - Chunks: cache first (by hash: no request at all), then the network, up to 6 in parallel, delivered in document order.
- The editor is usable after the first chunk (
ready). The rest is appended in batches, at most every ~100 ms, with a yield between appends so the host can paint. - Loading is not an edit: appended content stays out of undo history and out of the collab plugin's unconfirmed steps (
appendLoaded), so it is never sent back to the server.
openDocument(api, docId, { firstChunk }) delivers a chosen chunk first (random access); signal aborts (in-flight requests are cancelled). A cold open asks for the manifest with the first chunk inlined (one round trip instead of two), and staleWhileRevalidate: true opens a warm document from the cached manifest with no request at all, revalidating in the background (the SyncClient's first pull catches up). A chunk that can't be loaded after retries stops the load with a ChunkLoadError; ready rejects if that happens before the first chunk. Measured behaviour over real links: Real networks.
Saving with SyncClient
- Debounced, batched pushes: 400 ms after the last change (
debounceMs), all pending steps go in one push. - Pull + rebase: each sync pulls remote steps first; if a push is behind, the next pull rebases local steps and it retries (up to 5 rounds).
- Polling every 2 s (
pollMs) for remote changes and retries. - Offline detection: transport
OfflineErrors (or fetchTypeErrors) set the state tooffline; edits keep accumulating and sync when back.
Offline outbox
With outbox (an IndexedDbStore), the client persists an OutboxRecord 150 ms after every change: the base manifest, the steps the server confirmed since that manifest, and the local steps it hasn't seen. Reopening replays it (restoreOutbox): cached chunks + confirmed steps + unconfirmed steps, so offline edits survive reloads and sync once online. When fully saved, the outbox is cleared.
For small documents, LocalStorageStore keeps whole-document snapshots (and the previous one, previous(docId), so one bad save can never destroy the only copy). Prefer IndexedDB for books.
Data safety
The server never trusts a client to be implemented correctly.
Result checksums
Every push carries checksum: the docChecksum of the client's document after its steps. The server applies the steps to its own copy and compares. Any mismatch means the client's base was wrong (a partial load, chunks appended as ordinary edits, a wrong version, tampering, any host bug), so nothing is applied (rejected: 'diverged'). The SyncClient then stops syncing for good and keeps the edits in the outbox for recovery; the host should reopen from the server.
docChecksum is an order-sensitive hash of the top-level blocks, with each block's hash memoized by ProseMirror node identity: ~60 ms cold for 35k blocks (warmed at start, off the save path), ~4 ms per save.
Chunk verification
Chunks are content-addressed, so the hash is also an integrity check: a corrupt cache entry is refetched; a corrupt response is refetched once past the HTTP cache (cache: 'reload'), then fails with IntegrityError instead of being loaded.
Mass-delete guard
The server refuses a push that would delete more than half of a large document (rejected: 'mass-delete'). The SyncClient holds it in needs-confirmation until the user confirms (sync.confirmPending(), which resends with confirm: true) or reopens to discard.
Other guarantees
- Loading never produces steps, so a failed, empty or partial load can't overwrite anything.
- Steps are positional, and a client's loaded prefix is identical to the server's, so edits made while the rest streams in apply to exactly the same content (tested).
- Out-of-band server changes must also be steps (e.g. replacing a chapter is one
ReplaceStep) so every client rebases; a raw database overwrite would bypass clients.
These behaviours are pinned by packages/sync/src/integrity.test.ts: a host that appends chunks as edits, a host that syncs before loading finished, a tampered cache entry and a tampered response are all caught.
Writing a backend
A complete Node server (persistence, WebSocket collaboration, compression, auth hook) is @nextgensoftwares/folio-server-node; see Real networks. The core it builds on is MemoryDocumentServer: it holds each document as a ProseMirror doc, cuts chunks at chapters (minChunkBlocks 100, maxChunkBlocks 1000), applies pushed steps, verifies checksums (requireChecksum, default true; never disable it in production) and enforces the mass-delete guard (massDeleteGuardSize). memoryApi(server, network) wraps it in a simulated network (NETWORKS.lan / wifi / 4g / 3g, offline switch) for tests and demos.
Backends reuse the same code. MemoryDocumentServer, docChecksum, hashString and the step logic are DOM-free; a Node backend (e.g. NestJS + Postgres) can import them so client and server compute identical checksums and chunk hashes. Persist the document and the step log; serve chunk bodies with Cache-Control: immutable.