Real networks
The playground's "Mock API" runs a MemoryDocumentServer inside the tab behind a simulated network. That proves the protocol, but it hides what real networks do: requests fail, time out and arrive corrupted, other people save while you load, first bytes are slow, fonts come over the same link, and many readers hit one server. This page is the answer to "does loading still work, and stay fast, against a real API?", measured against a real server over shaped links.
Short answer: yes, after the fixes listed below. Nothing in the loader depends on knowing the chapters ahead of time or on local data: the client learns the structure from the manifest, fetches content-addressed chunks over HTTP, verifies every one, and saves steps that the server re-validates. The real network surfaced two bugs (one that broke every save in the playground, one in the test proxy) and several latency costs, all fixed.
The reference server: @nextgensoftwares/folio-server-node
A Node 22 server on node:http with one dependency (ws, MIT). It is a reference, not a framework: hosts with NestJS/Express reuse its pieces.
┌──────────────── createFolioServer ─────────────────┐
HTTP/1.1 (+h2) │ routes ── PersistentDocumentServer ── DocStorage │
GET manifest │ │ (MemoryDocumentServer: FileStorage: │
GET chunks/:h │ │ validatePush, snapshot.json│
POST steps ────┼───┘ ChunkIndex, pulls) steps.log │
WS /collab ────┼── CollabGateway ── CollabRoom (same documents) │
└─────────────────────────────────────────────────────┘- One validation path.
validatePush(in@nextgensoftwares/folio-sync) is the storage-agnostic core shared by the in-memory server, the Node server and any custom backend: steps are parsed and applied to the server's own copy, the result must equal the client's checksum, and a push that deletes most of a large document needsconfirm. Collab pushes over WebSocket and REST pushes go through the samePersistentDocumentServer, so they are validated and persisted identically, and REST saves are broadcast to collaborators. - Untrusted input. Bodies are size-limited (413), JSON and shapes are checked (400), step counts are capped, document ids are restricted to
[A-Za-z0-9._-], chunk hashes to[0-9a-z], layout reports are bounded, errors never leak internals, and every route goes through anauthorizehook (read,write,create,collab). - Incremental chunking on write.
ChunkIndexkeeps each chunk's block objects. ProseMirror reuses unchanged blocks across versions, so after a push the index walks the 35,000 blocks once (pointer comparisons) and re-serializes and re-hashes only the chunks whose blocks changed. Rebuilding the stress book's manifest after a keystroke takes 1.5 ms (p50), versus 510 ms for a cold build (which is now done in the background at boot). - Superseded chunks stay fetchable (512 by default). A reader that got the manifest just before someone saved still downloads every chunk it lists; content addressing makes serving an old hash always correct.
- Caching and compression. Chunks:
ETag= hash,Cache-Control: private, max-age=31536000, immutable(publiconly for public documents, so CDNs can't bypass auth), 304 onIf-None-Match; brotli (quality 5) or gzip compressed once on zlib's threadpool and kept in a 64 MB LRU, so a hit is a buffer write. The manifest:no-cache+ ETag.?first=1inlines chunk 0. Keep-alive (65 s, above common load-balancer idle timeouts); HTTP/2 over TLS withtls: { key, cert }(HTTP/1.1 still accepted). - Persistence (
FileStorage): per document asnapshot.json(replaced by write + fsync + rename), an append-onlysteps.log(one JSON line per accepted step, fsynced before the push is acknowledged) andlayout.json. Boot = snapshot + replay of the log tail; a torn last line (crash mid-write, never acknowledged) is cut off, a corrupt line anywhere else refuses to load. Snapshots run in the background (every 2,000 steps or 10 s idle) and compact the log and the in-memory step tail, keeping the last 20,000 steps / 8 MB for clients pulling from behind (older:410, reopen).
Files + log, or SQLite?
DocStorage.append is synchronous on purpose: the push is acknowledged only after it returns. Files + log need no dependency and no native build, the log is trivially inspectable, and an append is one write + fsync (≈ 5 ms here, the largest part of a push). SQLite (better-sqlite3 or node:sqlite) gives transactions, WAL group commit and one file to back up, and is the better choice once a server holds thousands of documents or needs queries over them; it fits the same interface (append = one INSERT in a transaction). An asynchronous database (Postgres) needs a per-document async queue around validate → insert → commit; the validation itself (validatePush) is reusable as is.
Measurements
Stress book: 7,314 pages, 35,019 blocks, 127 chunks; 26.0 MB of compact JSON, 3.9 MB with brotli (4.2 MB gzip). Chrome (Playwright), playground in Vite dev mode on a shared desktop, server on the same machine. Network shaping: a TCP proxy in front of the API only (server/throttle-proxy.mjs; bandwidth shared fairly between connections, RTT/2 each way):
| profile | RTT | down | up |
|---|---|---|---|
| Wi-Fi | 30 ms | 3.75 MB/s | 1.9 MB/s |
| Slow 4G | 150 ms | 500 KB/s | 188 KB/s |
| Fast 3G (Chrome DevTools) | 562.5 ms | 189 KB/s | 86 KB/s |
Times are from the start of the open (folio:open) to: editable (first chunk in the editor), painted (the next frame), full (all 35,019 blocks appended and laid out). Cold = empty IndexedDB and HTTP cache.
Opening the stress book
| editable | painted | full | requests | transferred | |
|---|---|---|---|---|---|
| Mock, Wi-Fi sim (25 ms, 20 MB/s), cold | 1,101 ms | 1,322 ms | 7.5 s | (in tab) | (in tab) |
| Mock, Wi-Fi sim, warm | 1,080 ms | 1,247 ms | 5.6 s | ||
| Mock, 4G sim (70 ms, 4 MB/s), cold | 1,258 ms | 1,455 ms | 9.2 s | ||
| Mock, 3G sim (300 ms, 400 KB/s), cold | 1,994 ms | 2,173 ms | 24.0 s | ||
| Real, loopback, cold | 140 ms | 262 ms | 5.7 s | 129 | 3,881 KB |
| Real, loopback, warm | 152 ms | 345 ms | 3.6 s | 3 | 4 KB |
| Real, Wi-Fi, cold | 169 ms | 279 ms | 6.6 s | 129 | 3,881 KB |
| Real, Wi-Fi, warm | 140 ms | 365 ms | 3.8 s | 3 | 4 KB |
| Real, Slow 4G, cold | 357 ms | 490 ms | 11.8 s | 129 | 3,881 KB |
| Real, Slow 4G, warm | 149 ms | 457 ms | 3.6 s | 3 | 4 KB |
| Real, Fast 3G, cold | 860 ms | 967 ms | 26–33 s | 129 | 3,884 KB |
| Real, Fast 3G, warm | 159 ms | 450 ms | 3.7 s | 2 | 1 KB |
The mock is slower to first page than the real server: it generates the book and builds the first manifest (≈ 0.5 s of hashing) inside the tab. Its full-load numbers are of the same order because full load is dominated by laying out 35,000 blocks (≈ 3.5 s, CPU) on fast links and by bandwidth on slow ones (3.9 MB at 189 KB/s ≈ 21 s on Fast 3G). Warm opens transfer nothing but the manifest revalidation, on any network.
Before the latency fixes below, the same real runs measured: editable 254–288 ms (Wi-Fi), 559–591 ms (Slow 4G), 1,455–1,499 ms (Fast 3G) cold, and 720–777 ms warm on Fast 3G.
Saving, collaboration, offline
| loopback | Wi-Fi | Slow 4G | Fast 3G | |
|---|---|---|---|---|
| keystroke push round trip (p50 / p95) | 10 / 20 ms | 41 / 44 ms | 172 / 180 ms | 589 / 590 ms |
| collab: A types → B sees it (p50 / p95) | 25 ms | 56 / 68 ms | 177 / 190 ms | 592 / 607 ms |
| offline → reconnect, 100 queued keystrokes (collab) | 82 ms | 205 ms | 585 ms | 1,882 ms |
| offline flush, 200 queued steps in one REST push (18.5 KB) | 38 ms | 79 ms | 295 ms | 828 ms |
Push round trips are the link RTT plus ≈ 10 ms of server time. In the editor, saves are debounced (400 ms) and batched, so typing never waits on them.
Server costs (stress book)
| ingest through the write path (127 pushes, one per chapter, 55.8 MB of step JSON) | 3.6–3.9 s; per push p50 14 ms, p95 21–25 ms, max 64 ms |
| keystroke push on the full book (validate + checksum + append + fsync) | 9 ms p50 (fsync 5–13 ms of it) |
| manifest: cold build / after an edit / cached | 510 ms (at boot, background) / 1.5 ms / 0.01 ms |
| boot from snapshot (26 MB) + log replay | 340 ms; heap 200 MB, RSS 315 MB |
| heap after ingest (before the first snapshot compacts the seed steps) | 293 MB |
| 20 concurrent cold readers, each the whole book with 6 parallel requests | first chunk p50 114–146 ms; whole book 1.2–1.5 s each (78 MB served); heap ≈ 275 MB, RSS ≈ 400 MB |
Failures the mock can't show
| injected (Playwright route on one chunk, or the proxy) | result |
|---|---|
| chunk returns 500 every time | 4 attempts (backoff 0.3 s → 1.2 s), then a visible "Could not load chunk 41 (blocks 11028–11281)" error after 3 s; the partial editor is removed, nothing syncs |
| 503 twice, then OK | retried, document complete |
| connection reset once | retried, document complete |
| corrupted body every time | refetched once with cache: 'reload', then IntegrityError shown ("failed verification") |
| corrupted once (e.g. a bad HTTP-cache/CDN entry) | the reload refetch gets the good body; complete |
| 3 s to first byte on every chunk | slow (82 s for 127 chunks), never fails (timeout is 30 s per attempt) |
| another writer saves mid-load (Fast 3G, edit lands in chunk 0) | the old manifest's chunks are still served; the client loads v1498, pulls v1499, and its checksum equals the server's |
.../manifest 404 | the load stops with an error; nothing is shown as the document |
Unit tests pin all of these (packages/sync/src/network.test.ts, packages/server-node/src/*.test.ts), including hanging requests (timeout, then retry), abort (in-flight requests cancelled), 4xx not retried, torn and corrupt logs, restart recovery and collab over a real socket.
Fonts over the network
Fonts gate the first layout: Folio measures with the real font files, never with fallback metrics (a page laid out with the wrong metrics would break at different places, and cached measurements would be wrong afterwards). The playground's four fonts are 837 KB of TTF:
| Wi-Fi | Slow 4G | Fast 3G | |
|---|---|---|---|
| 4 fonts in parallel, uncompressed (Vite dev) | 267 ms | 1,798 ms | 4,987 ms |
| same, gzip (395 KB): estimate | ~150 ms | ~0.9 s | ~2.5 s |
On Fast 3G the fonts cost more than the document's first chunk. What helps: <link rel="preload" as="fetch" crossorigin> for the font files (they start with the HTML instead of after the JS), serving TTFs compressed (static hosts usually don't by default), starting the manifest request before awaiting the fonts (the playground does: the manifest + inlined first chunk download while the fonts do), and splitting rarely used faces (e.g. a bold Arabic face) into a lazy load. Instead of a fallback-metrics preview, show a skeleton sized from manifest.estimatedPages.
Realtime over WebSocket
The reference server hosts live collaboration at GET /docs/:id/collab (CollabGateway → CollabRoom, the same PersistentDocumentServer as the REST routes). In the playground, pick Data source "HTTP server" and turn on "Two writers, live". Open the same document in a second tab or window, on this machine or another one pointed at the same server: everyone on the same document joins one room automatically. ?bots=N adds scripted collaborators, each on its own socket.
Which channel saves. In live mode the CollabClient carries every save: steps go over the socket, are validated by validatePush, fsynced to steps.log and only then acknowledged and broadcast. No REST SyncClient runs on a live editor, so nothing polls while the socket is up. Mixing both on one editor would apply the same steps twice; the checksum would catch it, but it would stop syncing. REST saves from other (non-live) clients reach the room as broadcasts.
| over the real socket | how |
|---|---|
| presence, typing, idle/away, follow | presence messages, throttled (≤ 1 per 80 ms) and mapped through steps on arrival |
| version history, diffs, restore, blame | rpc messages (history.*); the server keeps one HistoryStore per document across room lifetimes, and replays REST saves made while no room was open |
| roles and room policy | role: 'reader' clients never push; the server can cap the room (maxMembers) and show readers as viewers or hide them (playground env: FOLIO_MAX_PEERS, FOLIO_READERS=hidden, FOLIO_READERS_SEE=0) |
| connection state | people panel: socket state, "reconnecting in Ns (attempt k)", server URL, pending edits, "Reconnect now", "Drop my connection" |
Server down, measured (two Playwright contexts on the kitchen-sink document; the server stopped for 75 s while both kept typing):
| both clients show offline, edits pending | within one heartbeat of the socket closing |
| reconnect attempts in the first 60 s (per client) | 6–7 (gaps 0.7, 2.1, 4.6, 9, 16 s: backoff 0.5 s → 30 s with jitter) |
| after the server came back | both reconnected on their next attempt (≈ 30 s later at the cap), rebased, pushed, and ended at the server's version with identical documents |
fatal refusals (diverged, ahead, gone, full) | the client stops and closes the socket, keeping its edits in the outbox; no reconnect loop |
On the Stress book (loopback, Vite dev, "Show authors" on in the receiving tab), a keystroke in A is in B's document in 74 ms p50 / 123 ms p95, with no long tasks in B while it streams in.
Offline edits also survive a reload: the live client writes the same IndexedDB outbox format as the REST SyncClient, in the same database, so either mode restores the other's unsaved edits.
What the real network changed (fixes)
- Every save in the playground was refused as
diverged(mock and real). The editor's schema has plugin attrs with non-null defaults that the server's base schema doesn't declare, anddocChecksumhashed them. The checksum now skips attrs at their schema default, like the compact wire format already did: a plugin-extended client and a base-schema server agree, while a non-default value the server can't hold still diverges. - A chunk failing before the first one loaded left
readypending forever; in-flight requests of an abandoned load were never cancelled and rejected unhandled. Nowreadyrejects, failures areChunkLoadError(index, block range, cause), leaving iteration aborts in-flight fetches, and the playground drops the partial editor on error. - The HTTP transport had no timeouts, retries or abort, and sent
Content-Type: application/jsonon GETs, which makes every cross-origin chunk request a non-simple request with its own CORS preflight (Chrome's preflight cache is per URL: one extra round trip per chunk). Now: per-attempt timeouts, retries with exponential backoff + full jitter on network errors, timeouts, 408/429/5xx (honouringRetry-After),signaleverywhere, and noContent-Typewithout a body. - A corrupt response cached as
immutablecould never be repaired: the loader retried through the same HTTP cache. It now refetches once withcache: 'reload'. - A save mid-load made the old manifest's chunks 404 (the server only knew the current version's chunks). Superseded chunks are now retained.
- The server re-serialized and re-hashed the whole book for every manifest after an edit (≈ 0.5 s for the stress book).
ChunkIndexmakes it incremental (1.5 ms), the per-document checksum is memoized (a push and the next manifest share one pass), manifests are warmed at boot, and compressed chunk hits skip serialization. - Latency: a cold open now asks for the manifest with chunk 0 inlined (
inlineFirst, one round trip instead of two: Fast 3G editable 1.46 s → 0.86 s), andstaleWhileRevalidateopens warm documents from the cached manifest with no request on the critical path (Fast 3G warm 0.75 s → 0.16 s); the SyncClient's first pull brings the editor up to date. - Tried and rejected: fetching chunk 0 alone before the others (
firstChunkAlone). Over HTTP/1.1 chunk 0 rides the manifest's warm connection and finishes first anyway; measured no first-page gain and one extra RTT on the full load. It stays as an option for HTTP/2, where all streams share one connection.
What a host must implement
- Endpoints (see Loading & saving): manifest (
no-cache, ETag, optional?first=1), chunk by hash (immutable, ETag, compressed; keep superseded chunks for a while), stepsPOST(409 behind, 422 refused, 400 malformed) andGET ?since=(cap the page size; 410 when history is gone), layoutPUT, optionally the page lookupGET /pages/:nand the collab WebSocket. - Validation: call
validatePush(or reuseMemoryDocumentServer) with the same ProseMirror schema family as the clients; persist the steps before acknowledging; never write documents except through steps. - Auth hook points: every route and the WebSocket upgrade (
authorize(req, docId, action)). Chunks are content-addressed, not secret: authorize the document, and useCache-Control: privateunless the document is public. Cross-origin APIs: anAuthorizationheader turns every chunk GET into a preflighted request (one extra RTT each, cached per URL by the browser); prefer a same-origin API path (reverse proxy) or cookie auth for reads. - CORS: allow the app's origin,
Access-Control-Max-Age, exposeETag;Timing-Allow-Originif you want transfer sizes in Resource Timing. - Client:
createHttpApi({ baseUrl }),openDocument(api, id, { cache, signal, staleWhileRevalidate: true }),assembleDocument, thenSyncClient; showChunkLoadError/IntegrityErrorand never keep a partially loaded editor as the document.
Honest limits
- Measured on one machine: server, proxy and browser share CPU; the proxy models bandwidth, RTT and fair sharing, not packet loss, TCP slow start or TLS handshakes. Absolute numbers will differ; the ratios and the round-trip counts are the point.
- Full load of the stress book is bandwidth-bound on slow links (3.9 MB); the first page is not. The layout of 35,000 blocks (~3.5 s CPU) dominates on fast links.
- The server keeps every open document in memory (≈ 200 MB heap for the stress book) and pushes are serialized per process: scale out by document (sticky routing), not by request.
- Version history on the reference server is in memory: after a restart it starts at the version the document is first opened live at (blame too).
- Collaboration runs on an HTTP/1.1 listener; WebSocket over HTTP/2 (RFC 8441) isn't implemented, put the h2 terminator in front (nginx) if you need both.
- Pull history is bounded (20,000 steps / 8 MB in memory); a client further behind gets
410and must reopen. After a restart, history older than the last snapshot's keep window is gone. - The Node server process has no clustering, rate limiting or quotas; put them in front (or in
authorize).