diff --git a/.changeset/apply-pnpm-patch-in-vite.md b/.changeset/apply-pnpm-patch-in-vite.md new file mode 100644 index 000000000..ad2e83a10 --- /dev/null +++ b/.changeset/apply-pnpm-patch-in-vite.md @@ -0,0 +1,9 @@ +--- +"@inkandswitch/patchwork": patch +--- + +The vite plugin applies this repo's pnpm patch of `@automerge/automerge-repo` (`patches/@automerge__automerge-repo@.patch`, shipped in the package as `dist/patches/`) to every copy of automerge-repo vite bundles for a site, in place of the two hand-written edits it carried before. A site installing from npm gets the same automerge-repo as this workspace: the real `detach` behind eviction, the `mesh` adapter role, and whatever else the patch holds. + +Covered: the page build's chunks, including the `/packages/@automerge/automerge-repo*.js` import-map chunks that patchwork's own protocol-handler worker imports from; the dev server's pre-bundled deps (an esbuild `onLoad` plugin in `optimizeDeps`); and module workers a site builds through vite (`new Worker(new URL("./x.ts", import.meta.url), { type: "module" })`), which vite bundles in a separate rollup pass with only `worker.plugins` — the `config()` plugin's worker config now sets `worker.plugins` to `[wasm(), patches({ complete: false })]`, so those bundles are patched too. A site that passes `worker: false` and writes its own `worker.plugins` has to add `patches({ complete: false })` to them itself. + +`patches()` fails the build if any file the patch edits never reached the bundler; `patches({ complete: false })` skips that check, for a worker bundle that may import none or only some of them. The patch is pinned to one automerge-repo version; a version bump fails the build until the patch is re-made against it. diff --git a/.changeset/evict-after-handoff.md b/.changeset/evict-after-handoff.md new file mode 100644 index 000000000..c6f75e2e6 --- /dev/null +++ b/.changeset/evict-after-handoff.md @@ -0,0 +1,7 @@ +--- +"@inkandswitch/patchwork-bootloader": patch +--- + +The automerge protocol handler worker evicts the documents its Repo loaded once no `automerge:` handoff has been in flight for five seconds. A page load is a burst of handoffs through the same folder documents; they now stay hot across the load and are released after it, instead of living in the SharedWorker for as long as any tab is open. A headless `automerge:/path` redirect waits up to three seconds for the folder to hold every head a connected Subduction peer has advertised, since a re-found folder comes back from IndexedDB before its sync round lands. + +Eviction goes through `Repo.removeFromCache`, which this workspace's pnpm patch of `@automerge/automerge-repo@2.6.0-subduction.48` makes real: `removeFromCache` awaits each source's `detach`, and the Subduction source's `detach` persists unsaved commits, runs one sync round if no peer has them, drops its entry and `heads-changed` listener, and unsubscribes the ephemeral topic, so the document can be collected and a later `find` attaches afresh. The Vite plugin ships and applies this patch for consumers installing `@inkandswitch/patchwork` from npm. diff --git a/.changeset/heads-announcements.md b/.changeset/heads-announcements.md new file mode 100644 index 000000000..1b7959b7d --- /dev/null +++ b/.changeset/heads-announcements.md @@ -0,0 +1,14 @@ +--- +"@inkandswitch/patchwork": patch +"@inkandswitch/patchwork-bootloader": minor +--- + +Every Repo on the origin — each tab's `createRepo()` and the automerge protocol handler worker's, in both the plain and the keyhive branch — passes `headsChannel`, named `-heads`. When a tab persists commits it authored, its Repo posts the document's new heads on that BroadcastChannel; a sibling with the document open reloads it from the shared IndexedDB into its handle, and ignores announcements for documents it does not have open or heads it already knows. The channel also carries what the sync server holds: every tab talks to that server as the same identity, so one tab's `onRemoteHeads` observation is a fact for all of them, and each keeps the newest one per peer rather than the last one to arrive. Only a real observation is announced, never a relay of a relay. Every node keeps the same signer and peer id, its own socket to the sync server and the shared database; what changes is how a write in one tab reaches the others. + +The siblings mesh is gone with it: `siblingAdapters()` and the `@inkandswitch/patchwork-bootloader/siblings` export are removed, and no Repo passes `subductionAdapters`. Under one peer id the Subduction core never pushes a commit back to the sending peer's other connections, so the mesh links only added sync rounds. + +The option lives in this workspace's pnpm patch of `@automerge/automerge-repo@2.6.0-subduction.48`, which also carries three fixes: a `find` resolves from local storage as soon as shared storage holds the document, before the first server round; a `heads-changed` that saved nothing new (a reload, for instance) opens no server round; a document still initializing hydrates from local storage when a sync round fails while disconnected. A reload whose storage read fails is retried a few times, then logged once and left until the next announcement. + +A reload updates the handle, not the Subduction node's resident tree, so a commit that reached a tab only over the channel is one that tab cannot push: its hash is in the set every push is filtered against, and it is not in the tree a round reads. A containment backstop watches for that. When an entry's handle holds heads no peer has been seen holding, and a settling delay passes without a sibling reporting that the server has them, the tab writes those commits to storage a second time — which is what puts them in its tree — and then opens a round that can carry them. Each commit costs at most one such duplicate write, and a tab that is offline still does the write, so the reconnect round finds the tree already correct. The write is owed to the commit, not to the round: the cap on heal rounds gates the rounds alone, and a commit stops counting as stranded only once it has been stored or some peer has been seen holding it. Where every tab is online and pushing its own edits, the sibling's report arrives inside the settling delay and none of this runs. + +Consumers installing from npm get the unpatched fork, where `headsChannel` is ignored and tabs meet only through the server, until the fork is republished with these changes and the catalog pin is bumped. diff --git a/.changeset/in-thread-indexeddb.md b/.changeset/in-thread-indexeddb.md new file mode 100644 index 000000000..749975f8d --- /dev/null +++ b/.changeset/in-thread-indexeddb.md @@ -0,0 +1,6 @@ +--- +"@inkandswitch/patchwork": patch +"@inkandswitch/patchwork-bootloader": patch +--- + +IndexedDB is opened on the node's own thread: `createRepo` and the automerge protocol handler worker use `IndexedDBStorageAdapter` in place of `IndexedDBWorkerStorageAdapter`. Measured in `sites/bench`, the worker adapter bought no main-thread time (same boot, same cold load of 40 documents, 1ms flush latency either way) and cost a dedicated worker per tab, about 30 MB across three. The origin-wide signer is unchanged. diff --git a/.changeset/lazy-keyhive.md b/.changeset/lazy-keyhive.md new file mode 100644 index 000000000..d60379f7c --- /dev/null +++ b/.changeset/lazy-keyhive.md @@ -0,0 +1,14 @@ +--- +"@inkandswitch/patchwork": patch +"@inkandswitch/patchwork-bootloader": patch +"@inkandswitch/patchwork-elements": patch +"@inkandswitch/patchwork-plugins": patch +--- + +`@automerge/automerge-repo-keyhive` is loaded only where keyhive is in use: `createRepo` and the automerge protocol handler worker `import()` it inside their keyhive branch, and `patchwork-elements` and `patchwork-plugins` no longer import it at runtime. Its entry module carries the keyhive wasm as a 3 MB base64 string, so the static imports put a 3.1 MB chunk in every tab's modulepreload list and in the worker whether or not the site enabled keyhive, at 7 to 10 MB of memory per tab, and 8 MB in the protocol-handler worker. The chunk is still emitted under `/packages/` and listed in the import map for tool code. Type imports are unchanged. + +`isKeyhiveDoc` in `patchwork-plugins`, and the keyhive access gates in `patchwork-elements`, decide from the document id's bytes: an id shorter than 32 bytes, or one whose bytes 16 through 31 are all zero, is a legacy document. They used to construct a keyhive `DocumentId` and take a throw as legacy, but that constructor is an ed25519 point decode and accepts about half of legacy padded ids, so about half of legacy documents went through `bestAccessForDoc`. This is the check behind ARK's `isUnprotectedDoc`, which it recommends over the deprecated `docIdFromAutomergeUrl`. + +When keyhive access to a document changes, `patchwork-elements` looks up the document's handle by its automerge document id before retrying. It used the keyhive `DocumentId` string, which is hex and never matched a handle, so an unavailable handle was never dropped before the retry. + +The vite plugin gives the worker chunks an empty module-preload dependency list. Vite wraps a dynamic import in a preload helper that touches `document` when it has dependencies to preload, and a worker has no `document`. diff --git a/.changeset/one-automerge-wasm.md b/.changeset/one-automerge-wasm.md new file mode 100644 index 000000000..7973ab515 --- /dev/null +++ b/.changeset/one-automerge-wasm.md @@ -0,0 +1,19 @@ +--- +"@inkandswitch/patchwork-bootloader": patch +"@inkandswitch/patchwork": patch +"@inkandswitch/patchwork-filesystem": patch +"@inkandswitch/patchwork-elements": patch +"@inkandswitch/patchwork-plugins": patch +"@inkandswitch/patchwork-providers": patch +"@inkandswitch/edge-handles": patch +--- + +Every tab and worker now runs one automerge wasm instance, streamed from `/automerge.wasm`. + +The bare `@automerge/automerge`, `@automerge/automerge-repo`, `@automerge/automerge-subduction` and `@keyhive/keyhive` specifiers resolve to their `/slim` builds everywhere: in the vite plugin's bundle, in the importmap a tool sees at runtime, and in the dev server's worker bundles. The fullfat entries embed and instantiate their own copy of the wasm on import, so a single value import of the bare name (there were four in our own packages) used to cost each tab a second automerge instance and a second, byte-identical `automerge.wasm` download. The `/packages/@automerge/automerge.js` chunk is no longer emitted; the bare name points at `/packages/@automerge/automerge/slim.js`. + +`initWasm` in the host and the protocol-handler worker hand the wasm-bindgen init a `Request` instead of buffering the bytes first, so both automerge and subduction go through `WebAssembly.instantiateStreaming`: no 5 MB transient copy, and the compiled module is eligible for Chrome's code cache. + +`@inkandswitch/patchwork-bootloader/externals` and `/externals-list` export the alias table as `slim`. + +`pnpm lint` (scripts/lint-slim-imports.mts, run in CI) fails on any import of a bare name in the table, type-only ones included, so the fullfat entries stay out of every bundle. diff --git a/.changeset/skip-recaching-unchanged.md b/.changeset/skip-recaching-unchanged.md new file mode 100644 index 000000000..2890294b8 --- /dev/null +++ b/.changeset/skip-recaching-unchanged.md @@ -0,0 +1,5 @@ +--- +"@inkandswitch/patchwork-bootloader": patch +--- + +The service worker no longer re-caches a passthrough response whose etag matches the copy it already holds. Every tab boot used to clone the whole bundle's responses and write them back to Cache Storage, holding a second copy of each body in the service worker's process until the write landed; with several tabs opening at once that peaked at a few hundred MB. diff --git a/.github/workflows/ci.yml b/.github/workflows/ci.yml index e1adfe6b7..812241f05 100644 --- a/.github/workflows/ci.yml +++ b/.github/workflows/ci.yml @@ -32,6 +32,9 @@ jobs: - name: install dependencies run: pnpm install --frozen-lockfile + - name: lint + run: pnpm lint + - name: build everything run: pnpm build diff --git a/core/bootloader/package.json b/core/bootloader/package.json index f1f34d6c7..47727c132 100644 --- a/core/bootloader/package.json +++ b/core/bootloader/package.json @@ -26,10 +26,6 @@ "import": "./dist/externals-list.js", "types": "./dist/externals-list.d.ts" }, - "./siblings": { - "import": "./dist/siblings.js", - "types": "./dist/siblings.d.ts" - }, "./signer": { "import": "./dist/signer.js", "types": "./dist/signer.d.ts" diff --git a/core/bootloader/src/automerge-protocol-handler-worker.ts b/core/bootloader/src/automerge-protocol-handler-worker.ts index 23f7b32fd..c2b38aea4 100644 --- a/core/bootloader/src/automerge-protocol-handler-worker.ts +++ b/core/bootloader/src/automerge-protocol-handler-worker.ts @@ -3,7 +3,7 @@ // does. // // It is a node like any tab's: the same IndexedDB, its own sync-server socket, -// and the siblings channel to the tabs. Resolving requests is its whole job. +// and the storage channel to the tabs. Resolving requests is its whole job. // When the service worker misses the cache for a request that looks like a URL // encoded URL, it broadcasts a HandoffRequestMessage on HANDOFF_CHANNEL; we // resolve the automerge URL, write the response into the service worker's @@ -11,8 +11,8 @@ // reply on the same channel. import { initializeWasm, hasHeads } from "@automerge/automerge/slim"; // eslint-disable-next-line -// @ts-ignore — initSync is a wasm-bindgen runtime helper not in the .d.ts -import { initSync as initSubductionSync } from "@automerge/automerge-subduction/slim"; +// @ts-ignore — the default init is a wasm-bindgen runtime helper not in the .d.ts +import initSubduction from "@automerge/automerge-subduction/slim"; import { Repo, @@ -21,21 +21,20 @@ import { stringifyAutomergeUrl, type AutomergeUrl, type DocHandle, + type DocumentId, type PeerId, + type StorageId, } from "@automerge/automerge-repo/slim"; import { resolvePath } from "@inkandswitch/patchwork-filesystem"; -import { IndexedDBWorkerStorageAdapter } from "@automerge/automerge-repo-storage-indexeddb/IndexedDBWorkerStorageAdapter"; +import { IndexedDBStorageAdapter } from "@automerge/automerge-repo-storage-indexeddb"; import { WebSocketWorkerClientAdapter } from "@automerge/automerge-repo-network-websocket"; -import { - initializeAutomergeRepoKeyhive, - initKeyhiveWasm, - type AutomergeRepoKeyhive, - type SyncServerSelection, +import type { + AutomergeRepoKeyhive, + SyncServerSelection, } from "@automerge/automerge-repo-keyhive"; import { DEFAULT_CLASSIC_SYNC_SERVER } from "./sync-config.js"; -import { siblingAdapters } from "./siblings.js"; import { loadOrCreateSigner } from "./signer.js"; import { keyhiveStorageName, storagePrefix } from "./storage.js"; import { startWorkerControl } from "./worker-control.js"; @@ -86,12 +85,10 @@ function getRepo(): Promise { async function buildRepo(): Promise { log("fetching wasm"); - const [automergeWasm, subductionWasm] = await Promise.all([ - fetch("/automerge.wasm").then((r) => r.arrayBuffer()), - fetch("/subduction.wasm").then((r) => r.arrayBuffer()), + await Promise.all([ + initializeWasm(new Request("/automerge.wasm")), + initSubduction(new Request("/subduction.wasm")), ]); - initSubductionSync(new Uint8Array(subductionWasm)); - await initializeWasm(new Uint8Array(automergeWasm)); log("wasm initialized"); const { repo, hive } = syncServer.keyhive @@ -104,14 +101,14 @@ async function buildRepo(): Promise { } async function buildPlainRepo(): Promise { - const storage = new IndexedDBWorkerStorageAdapter(); + const storage = new IndexedDBStorageAdapter(); return new Repo({ signer: await loadOrCreateSigner(storage), storage, peerId: `${storagePrefix}-resolver-${Math.random().toString(36).slice(2)}` as PeerId, subductionWebsocketEndpoints: [syncServer.url], - subductionAdapters: siblingAdapters(), + headsChannel: `${storagePrefix}-heads`, enableRemoteHeadsGossiping: true, }); } @@ -119,6 +116,8 @@ async function buildPlainRepo(): Promise { async function buildKeyhiveRepo( keyhiveSyncServer: SyncServerSelection ): Promise<{ repo: Repo; hive: AutomergeRepoKeyhive }> { + const { initKeyhiveWasm, initializeAutomergeRepoKeyhive } = + await import("@automerge/automerge-repo-keyhive"); initKeyhiveWasm(); const { hive, repo } = await initializeAutomergeRepoKeyhive({ // ARK injects an `idFactory` deriving document ids from keyhive. A site @@ -127,7 +126,7 @@ async function buildKeyhiveRepo( new Repo( syncServer.useIdFactory === false ? config : { ...config, idFactory } ), - storage: new IndexedDBWorkerStorageAdapter(keyhiveStorageName), + storage: new IndexedDBStorageAdapter(keyhiveStorageName), peerIdSuffix: `${storagePrefix}-resolver` + Math.random().toString(36).slice(2), automaticArchiveIngestion: true, @@ -136,9 +135,9 @@ async function buildKeyhiveRepo( // the matching peer id. Omitting it defaults to "subduction". syncServer: keyhiveSyncServer, repo: { - storage: new IndexedDBWorkerStorageAdapter(), + storage: new IndexedDBStorageAdapter(), subductionWebsocketEndpoints: [syncServer.url], - subductionAdapters: siblingAdapters(), + headsChannel: `${storagePrefix}-heads`, enableRemoteHeadsGossiping: true, }, }); @@ -244,6 +243,42 @@ function waitForHeads( }); } +const CATCH_UP_MS = 3_000; + +async function caughtUpWithPeers( + repo: Repo, + handle: DocHandle, + signal: AbortSignal +): Promise { + if (!repo.isSubductionConnected()) return; + const peers = (await repo.connectedSubductionPeerIds()) as StorageId[]; + const caughtUp = () => { + const states = peers.map((peer) => handle.isCaughtUpWith(peer)); + return ( + states.some((state) => state !== undefined) && !states.includes(false) + ); + }; + if (caughtUp() || signal.aborted) return; + await new Promise((resolve) => { + const done = () => { + clearTimeout(timer); + handle.off("remote-heads", check); + handle.off("heads-changed", check); + signal.removeEventListener("abort", done); + resolve(); + }; + const check = () => { + if (caughtUp()) done(); + }; + const timer = setTimeout(done, CATCH_UP_MS); + handle.on("remote-heads", check); + handle.on("heads-changed", check); + signal.addEventListener("abort", done); + if (signal.aborted) done(); + else check(); + }); +} + /** * Thrown instead of returning a Response when the request should fail as a * network error rather than resolve to something the caller can memoize. @@ -270,6 +305,7 @@ async function resolveAutomergeUrl( // the headless req if (!heads) { const folder = await repo.find(maybeAutomergeUrl, { signal }); + await caughtUpWithPeers(repo, folder, signal); const url = stringifyAutomergeUrl({ documentId, heads: folder.heads() }); const location = `/${encodeURIComponent(url)}${path.length ? `/${path.join("/")}` : ""}`; return Response.redirect(location, 307); @@ -325,8 +361,38 @@ function impatience(limit: number) { ); } +const EVICT_IDLE_MS = 5_000; +let handoffsInFlight = 0; +let evictTimer: ReturnType | undefined; + +function handoffStarted() { + handoffsInFlight++; + clearTimeout(evictTimer); +} + +function handoffFinished() { + if (--handoffsInFlight > 0) return; + clearTimeout(evictTimer); + evictTimer = setTimeout(() => { + evictLoaded().catch((error) => console.error("eviction failed", error)); + }, EVICT_IDLE_MS); +} + +async function evictLoaded() { + const repo = await repoPromise?.catch(() => null); + if (!repo) return; + const ids = Object.keys(repo.handles) as DocumentId[]; + log( + `evicting ${ids.length} document(s) after ${EVICT_IDLE_MS}ms without a handoff` + ); + for (const id of ids) { + if (handoffsInFlight > 0) return; + await repo.removeFromCache(id); + } +} + async function handleHandoffRequest(message: HandoffRequestMessage) { - const { id, cachename, request } = message; + const { id, request } = message; let handoff: URL; try { @@ -350,6 +416,17 @@ async function handleHandoffRequest(message: HandoffRequestMessage) { return; } + handoffStarted(); + try { + await respondToHandoff(message, handoff); + } finally { + handoffFinished(); + } +} + +async function respondToHandoff(message: HandoffRequestMessage, handoff: URL) { + const { id, cachename, request } = message; + let response: Response; try { log(`resolving handoff ${id} for ${handoff}`); diff --git a/core/bootloader/src/externals-list.ts b/core/bootloader/src/externals-list.ts index 7597f6750..a9a6419fe 100644 --- a/core/bootloader/src/externals-list.ts +++ b/core/bootloader/src/externals-list.ts @@ -37,3 +37,10 @@ const externals = [ "solid-js/jsx-runtime", ]; export default externals; + +export const slim: Record = { + "@automerge/automerge": "@automerge/automerge/slim", + "@automerge/automerge-repo": "@automerge/automerge-repo/slim", + "@automerge/automerge-subduction": "@automerge/automerge-subduction/slim", + "@keyhive/keyhive": "@keyhive/keyhive/slim", +}; diff --git a/core/bootloader/src/externals.ts b/core/bootloader/src/externals.ts index 3aee7ab1c..f5df191c2 100644 --- a/core/bootloader/src/externals.ts +++ b/core/bootloader/src/externals.ts @@ -4,7 +4,7 @@ import { fileURLToPath } from "node:url"; const require = createRequire(import.meta.url); -export { default } from "./externals-list.js"; +export { default, slim } from "./externals-list.js"; /** * pretend the import came from inside this package, so node_modules resolution diff --git a/core/bootloader/src/service-worker.ts b/core/bootloader/src/service-worker.ts index 4250cffcb..394d2cb1c 100644 --- a/core/bootloader/src/service-worker.ts +++ b/core/bootloader/src/service-worker.ts @@ -344,7 +344,10 @@ async function servePassthrough( CACHEABLE_STATUSES.includes(result.status) && /^https?:/.test(request.url) ) { - cacheInBackground(fetchEvent, cache, request, result.clone()); + const etag = result.headers.get("etag"); + if (!etag || etag !== cached?.headers.get("etag")) { + cacheInBackground(fetchEvent, cache, request, result.clone()); + } } else { log(`not caching status ${result.status} for ${request.url}`); } diff --git a/core/bootloader/src/siblings.ts b/core/bootloader/src/siblings.ts deleted file mode 100644 index 9843aafef..000000000 --- a/core/bootloader/src/siblings.ts +++ /dev/null @@ -1,32 +0,0 @@ -import type { RepoConfig } from "@automerge/automerge-repo/slim"; -import { BroadcastChannelNetworkAdapter } from "@automerge/automerge-repo-network-broadcastchannel"; -import { storagePrefix } from "./storage.js"; - -type SubductionAdapters = NonNullable; - -/** - * Every Repo on this origin — each tab's, and the automerge protocol handler - * worker's — is a full Subduction node with its own storage and its own - * sync-server socket. Siblings would still meet through the server, - * eventually; this meets them over a BroadcastChannel so an edit in one tab - * lands in the others in the time it takes to post a message, online or not. - * - * What crosses the channel is Subduction: transport frames between two nodes, - * authenticated by each one's signer. No classic automerge sync runs here, so - * the adapters go to `new Repo({ subductionAdapters })` rather than to the - * network subsystem. - * - * A BroadcastChannel is a mesh in which every node sees every other one, and - * Subduction's handshake has an initiator and a responder, so the role is - * "mesh": for each pair, the node whose peer id sorts lower speaks first. - */ -export function siblingAdapters(): SubductionAdapters { - const serviceName = `${storagePrefix}-siblings`; - return [ - { - adapter: new BroadcastChannelNetworkAdapter({ channelName: serviceName }), - serviceName, - role: "mesh", - }, - ]; -} diff --git a/core/elements/src/events.ts b/core/elements/src/events.ts index 3069d7e8e..7a29f6fda 100644 --- a/core/elements/src/events.ts +++ b/core/elements/src/events.ts @@ -1,4 +1,4 @@ -import type { AutomergeUrl } from "@automerge/automerge-repo"; +import type { AutomergeUrl } from "@automerge/automerge-repo/slim"; export interface OpenDocumentEventDetail { url: AutomergeUrl; diff --git a/core/elements/src/legacy-impl.ts b/core/elements/src/legacy-impl.ts index 87b7ab6c6..a8f7030cd 100644 --- a/core/elements/src/legacy-impl.ts +++ b/core/elements/src/legacy-impl.ts @@ -1,9 +1,10 @@ import { + parseAutomergeUrl, type AutomergeUrl, type DocHandle, type DocHandleChangePayload, type Repo, -} from "@automerge/automerge-repo"; +} from "@automerge/automerge-repo/slim"; import { getSuggestedImportUrl, getType, @@ -13,6 +14,7 @@ import { import { getFallbackTool, getRegistry, + isKeyhiveDoc, isLoadablePlugin, registerPlugins, type LoadedTool, @@ -20,10 +22,7 @@ import { type ToolElement, } from "@inkandswitch/patchwork-plugins"; import debug from "debug"; -import { - docIdFromAutomergeUrl, - type AutomergeRepoKeyhiveBase as AutomergeRepoKeyhive, -} from "@automerge/automerge-repo-keyhive"; +import type { AutomergeRepoKeyhiveBase as AutomergeRepoKeyhive } from "@automerge/automerge-repo-keyhive"; import { MountedEvent, UnmountedEvent } from "./events.js"; const log = debug("patchwork:elements:legacy"); @@ -322,15 +321,8 @@ export class LegacyImpl { this.#state = State.initializing; if (this.#element.hive && this.#docUrl) { - let isKeyhiveDoc = false; - try { - docIdFromAutomergeUrl(this.#docUrl); - isKeyhiveDoc = true; - } catch { - // Legacy (padded-zero) doc: skip keyhive gate - } - - if (isKeyhiveDoc) { + // Legacy (padded-zero) doc: skip keyhive gate + if (isKeyhiveDoc(this.#docUrl)) { const bestAccess = await this.#element.hive.bestAccessForDoc( this.#element.hive.active.individual.id, this.#docUrl @@ -541,7 +533,7 @@ export class LegacyImpl { retryingDocs.add(this.#docUrl); try { - const documentId = String(docIdFromAutomergeUrl(this.#docUrl)); + const { documentId } = parseAutomergeUrl(this.#docUrl); const handle = (this.#element.repo.handles as any)[documentId]; if (handle && handle.state === "unavailable") { this.#element.repo.delete(this.#docUrl); @@ -566,7 +558,7 @@ export class LegacyImpl { let hasAccess = false; let accessCheckSucceeded = false; try { - docIdFromAutomergeUrl(this.#docUrl); + if (!isKeyhiveDoc(this.#docUrl)) return; const bestAccess = await this.#element.hive.bestAccessForDoc( this.#element.hive.active.individual.id, this.#docUrl diff --git a/core/elements/src/patchwork-view.ts b/core/elements/src/patchwork-view.ts index ec18d0f64..51a66b307 100644 --- a/core/elements/src/patchwork-view.ts +++ b/core/elements/src/patchwork-view.ts @@ -1,7 +1,8 @@ -import type { AutomergeUrl, Repo } from "@automerge/automerge-repo"; +import type { AutomergeUrl, Repo } from "@automerge/automerge-repo/slim"; import type { AutomergeRepoKeyhiveBase as AutomergeRepoKeyhive } from "@automerge/automerge-repo-keyhive"; import { getRegistry, + isKeyhiveDoc, isLoadablePlugin, type LoadablePlugin, type LoadedPlugin, @@ -13,7 +14,6 @@ import { } from "@inkandswitch/patchwork-providers"; import { MountedEvent, UnmountedEvent } from "./events.js"; import { LegacyImpl } from "./legacy-impl.js"; -import { docIdFromAutomergeUrl } from "@automerge/automerge-repo-keyhive"; import debug from "debug"; const log = debug("patchwork:elements:patchwork-view"); @@ -314,15 +314,10 @@ export function registerPatchworkViewElement( this.#state = State.initializing; if (params.hive && this.url) { - let isKeyhiveDoc = false; - try { - docIdFromAutomergeUrl(this.url); - isKeyhiveDoc = true; - } catch { - // Legacy (padded-zero) doc: skip keyhive gate - } + // Legacy (padded-zero) doc: skip keyhive gate + const keyhiveDoc = isKeyhiveDoc(this.url); - if (isKeyhiveDoc) { + if (keyhiveDoc) { const bestAccess = await params.hive.bestAccessForDoc( params.hive.active.individual.id, this.url @@ -367,7 +362,7 @@ export function registerPatchworkViewElement( }); } - if (isKeyhiveDoc) { + if (keyhiveDoc) { // Access is confirmed, but the doc's content may not have synced // yet. const progress = params.repo.findWithProgress( @@ -512,7 +507,7 @@ export function registerPatchworkViewElement( let hasAccess = false; let accessCheckSucceeded = false; try { - docIdFromAutomergeUrl(this.url); + if (!isKeyhiveDoc(this.url)) return; const bestAccess = await params.hive.bestAccessForDoc( params.hive.active.individual.id, this.url diff --git a/core/filesystem/src/find-handle.ts b/core/filesystem/src/find-handle.ts index e3a1e089a..3c10bf757 100644 --- a/core/filesystem/src/find-handle.ts +++ b/core/filesystem/src/find-handle.ts @@ -1,4 +1,4 @@ -import { type DocHandle, type Repo } from "@automerge/automerge-repo"; +import { type DocHandle, type Repo } from "@automerge/automerge-repo/slim"; import type { FolderDoc } from "./types.js"; /** diff --git a/core/filesystem/src/types.ts b/core/filesystem/src/types.ts index b742de076..45898012b 100644 --- a/core/filesystem/src/types.ts +++ b/core/filesystem/src/types.ts @@ -1,4 +1,7 @@ -import type { AutomergeUrl, ImmutableString } from "@automerge/automerge-repo"; +import type { + AutomergeUrl, + ImmutableString, +} from "@automerge/automerge-repo/slim"; import type { HasPatchworkMetadata } from "./metadata.js"; // needed in serviceworker only right now? diff --git a/core/filesystem/src/urls.ts b/core/filesystem/src/urls.ts index 5613788bc..751c37fef 100644 --- a/core/filesystem/src/urls.ts +++ b/core/filesystem/src/urls.ts @@ -1,4 +1,7 @@ -import { type AutomergeUrl, type DocHandle } from "@automerge/automerge-repo"; +import { + type AutomergeUrl, + type DocHandle, +} from "@automerge/automerge-repo/slim"; // The origin to resolve service-worker module URLs against. `location.origin` // is the string "null" inside a srcdoc/sandboxed frame — an invalid URL base — diff --git a/core/filesystem/test/heads-channel.test.ts b/core/filesystem/test/heads-channel.test.ts new file mode 100644 index 000000000..7c0800b9c --- /dev/null +++ b/core/filesystem/test/heads-channel.test.ts @@ -0,0 +1,797 @@ +import { describe, it, expect, afterEach, vi } from "vitest"; +import * as Automerge from "@automerge/automerge"; +import { + documentIdToBinary, + encodeHeads, + Repo, + type Chunk, + type DocumentId, + type PeerId, + type RepoConfig, + type StorageAdapterInterface, + type StorageId, + type StorageKey, +} from "@automerge/automerge-repo"; +import { MemorySigner, SedimentreeId } from "@automerge/automerge-subduction"; + +interface Doc { + text: string; +} + +class SharedMemoryAdapter implements StorageAdapterInterface { + writes: string[] = []; + reads: string[] = []; + failRanges = 0; + /** Deliver a range read this long after snapshotting it. */ + rangeDelayMs = 0; + /** Set while writes are held open, so a save can be caught in flight. */ + saveGate: Promise | null = null; + releaseSaves = () => {}; + + blockSaves() { + this.saveGate = new Promise((resolve) => { + this.releaseSaves = () => { + this.saveGate = null; + resolve(); + }; + }); + } + + constructor(public data = new Map()) {} + + async load(key: StorageKey) { + this.reads.push(key.join("/")); + return this.data.get(key.join("/")); + } + async save(key: StorageKey, data: Uint8Array) { + if (this.saveGate) await this.saveGate; + this.data.set(key.join("/"), data); + this.writes.push(key.join("/")); + } + async remove(key: StorageKey) { + this.data.delete(key.join("/")); + } + async loadRange(prefix: StorageKey): Promise { + const p = prefix.join("/"); + this.reads.push(p); + if (this.failRanges > 0) { + this.failRanges--; + throw new Error("storage unavailable"); + } + const out: Chunk[] = []; + for (const [k, data] of this.data) { + if (k === p || k.startsWith(p + "/")) out.push({ key: k.split("/"), data }); + } + if (this.rangeDelayMs > 0) await pause(this.rangeDelayMs); + return out; + } + async removeRange(prefix: StorageKey) { + for (const c of await this.loadRange(prefix)) this.data.delete(c.key.join("/")); + } + async saveBatch(entries: Array<[StorageKey, Uint8Array]>) { + if (this.saveGate) await this.saveGate; + for (const [k, d] of entries) { + this.data.set(k.join("/"), d); + this.writes.push(k.join("/")); + } + } +} + +function pause(ms: number) { + return new Promise((r) => setTimeout(r, ms)); +} + +async function until( + cond: () => boolean | Promise, + ms: number, + label: string +) { + const deadline = Date.now() + ms; + while (!(await cond())) { + if (Date.now() > deadline) throw new Error(`timed out after ${ms}ms: ${label}`); + await pause(20); + } +} + +function withTimeout(p: Promise, ms: number, label: string): Promise { + return Promise.race([ + p, + new Promise((_, reject) => + setTimeout(() => reject(new Error(`timed out after ${ms}ms: ${label}`)), ms) + ), + ]); +} + +const succeededRound = async () => ({ + entries: () => [ + { success: true, stats: { commitsReceived: 0, fragmentsReceived: 0 } }, + ], +}); + +const channelName = () => `test-heads-${Math.random().toString(36).slice(2)}`; + +const sidOf = (storage: SharedMemoryAdapter) => + storage.writes.find((k) => k.startsWith("subduction/commits/"))!.split("/")[2]; + +const SERVER = "server" as StorageId; +const MAX_BACKSTOP_SYNCS = 3; + +const bogusHead = () => + [...crypto.getRandomValues(new Uint8Array(32))] + .map((b) => b.toString(16).padStart(2, "0")) + .join(""); + +/** + * A repo that reports itself connected without any real transport, so the + * containment backstop is allowed to run. `healInitialDelayMs` is also how + * long the backstop waits before believing a divergence, so it sets both + * how fast the test runs and how much room there is to interrupt it. + */ +const connectedWithin = (healInitialDelayMs: number): Partial => ({ + subductionAdapters: [ + { + adapter: { + on() {}, + connect() {}, + state: () => ({ value: "ready", watch: async function* () {} }), + } as never, + serviceName: "test", + }, + ], + subductionTimeouts: { healInitialDelayMs, healMaxDelayMs: 40, healMaxAttempts: 2 }, +}); + +const connected = connectedWithin(40); + +/** + * No adapter at all, so `isConnected()` is false. The containment + * backstop opens no round here; the only thing it can still do is put a + * sibling's commits back in this node's tree, ready for the reconnect. + */ +const offlineWithin = (healInitialDelayMs: number): Partial => ({ + subductionTimeouts: { healInitialDelayMs }, +}); + +/** The commit heads this repo's subduction tree actually holds. */ +async function treeCommits(repo: Repo, documentId: DocumentId) { + const bytes = new Uint8Array(32); + bytes.set(documentIdToBinary(documentId)!.subarray(0, 32)); + const commits = await ( + await repo.subduction + ).getCommits(SedimentreeId.fromBytes(bytes)); + return (commits ?? []).map((c) => c.commitId.toHexString()); +} + +const commitWrites = (storage: SharedMemoryAdapter, from: number) => + storage.writes.slice(from).filter((k) => k.startsWith("subduction/commits/")); + +const repos: Repo[] = []; +afterEach(async () => { + await Promise.all(repos.map((r) => r.shutdown().catch(() => {}))); + repos.length = 0; +}); + +async function siblings(channel: string | undefined, extra: Partial = {}) { + const data = new Map(); + const storageA = new SharedMemoryAdapter(data); + const storageB = new SharedMemoryAdapter(data); + const secret = crypto.getRandomValues(new Uint8Array(32)); + const mk = (storage: SharedMemoryAdapter) => { + const signer = MemorySigner.fromBytes(secret); + const repo = new Repo({ + storage, + signer, + peerId: signer.peerId().toString() as PeerId, + network: [], + headsChannel: channel, + ...extra, + }); + repos.push(repo); + return repo; + }; + const repoA = mk(storageA); + const repoB = mk(storageB); + const mkSibling = () => mk(new SharedMemoryAdapter(data)); + const syncA = vi + .spyOn(await repoA.subduction, "syncWithAllPeers") + .mockImplementation(succeededRound as any); + const syncB = vi + .spyOn(await repoB.subduction, "syncWithAllPeers") + .mockImplementation(succeededRound as any); + return { storageA, storageB, repoA, repoB, syncA, syncB, mkSibling }; +} + +async function shared(channel: string | undefined, extra: Partial = {}) { + const s = await siblings(channel, extra); + const a = s.repoA.create(); + a.change((d) => { + d.text = "one"; + }); + await s.repoA.flush(); + const b = await withTimeout(s.repoB.find(a.url), 5000, "B finds A's doc"); + expect(b.doc().text).toBe("one"); + await pause(500); + return { ...s, a, b }; +} + +describe("subduction heads channel", () => { + it("a sibling's change arrives through the channel: no server round, nothing written back", async () => { + const { storageB, repoA, repoB, syncA, syncB, a, b } = await shared(channelName()); + + const roundsA = syncA.mock.calls.length; + const roundsB = syncB.mock.calls.length; + const writeMark = storageB.writes.length; + + a.change((d) => { + d.text = "two"; + }); + await until(() => b.doc().text === "two", 5000, "B sees A's change"); + await pause(500); + + expect(syncB.mock.calls.length).toBe(roundsB); + expect(syncA.mock.calls.length).toBeGreaterThan(roundsA); + expect( + storageB.writes.slice(writeMark).filter((k) => k.startsWith("subduction/")) + ).toEqual([]); + + const roundsA2 = syncA.mock.calls.length; + const roundsB2 = syncB.mock.calls.length; + + b.change((d) => { + d.text = "three"; + }); + await until(() => a.doc().text === "three", 5000, "A sees B's change"); + await pause(500); + + expect(syncA.mock.calls.length).toBe(roundsA2); + expect(syncB.mock.calls.length).toBeGreaterThan(roundsB2); + void repoA; + void repoB; + }); + + it("without a channel the sibling keeps its stale view", async () => { + const { repoA, a, b } = await shared(undefined); + + a.change((d) => { + d.text = "two"; + }); + await repoA.flush(); + await pause(1000); + expect(b.doc().text).toBe("one"); + }); + + it("an announcement for a doc this node never attached reads nothing", async () => { + const { storageA, storageB, repoA } = await siblings(channelName()); + + const a = repoA.create(); + a.change((d) => { + d.text = "one"; + }); + await repoA.flush(); + const sid = sidOf(storageA); + const readMark = storageB.reads.length; + + a.change((d) => { + d.text = "two"; + }); + await repoA.flush(); + await pause(500); + + expect(storageB.reads.slice(readMark).filter((k) => k.includes(sid))).toEqual([]); + }); + + it("an announcement of heads this node already knows reads nothing", async () => { + const name = channelName(); + const { storageA, storageB, b } = await shared(name); + const sid = sidOf(storageA); + const readMark = storageB.reads.length; + + const announcer = new BroadcastChannel(name); + announcer.postMessage({ + kind: "saved", + sid, + heads: Automerge.getHeads(b.doc()), + }); + await pause(500); + announcer.close(); + + expect(storageB.reads.slice(readMark).filter((k) => k.includes(sid))).toEqual([]); + }); + + it("a detached doc shows a sibling's later change when found again", async () => { + const { repoA, repoB, a } = await shared(channelName()); + await repoB.removeFromCache(a.documentId); + + a.change((d) => { + d.text = "two"; + }); + await repoA.flush(); + await pause(500); + + const again = await withTimeout( + repoB.find(a.url), + 5000, + "B finds A's doc again" + ); + await until( + () => again.doc().text === "two", + 5000, + "B's re-found handle shows A's change" + ); + }); + + it("a reload whose storage read fails once still converges without a server", async () => { + const { storageB, syncB, a, b } = await shared(channelName()); + const roundsB = syncB.mock.calls.length; + const readMark = storageB.reads.length; + storageB.failRanges = 6; + + a.change((d) => { + d.text = "two"; + }); + await until(() => b.doc().text === "two", 5000, "B converges after a failed read"); + await pause(300); + + expect(storageB.failRanges).toBe(0); + expect(storageB.reads.length - readMark).toBeGreaterThan(6); + expect(syncB.mock.calls.length).toBe(roundsB); + }); + + it("a reload whose storage keeps failing stops after a few tries and says so once", async () => { + const { storageB, repoA, syncB, a, b } = await shared(channelName()); + const warn = vi.spyOn(console, "warn").mockImplementation(() => {}); + const roundsB = syncB.mock.calls.length; + const readMark = storageB.reads.length; + storageB.failRanges = Infinity; + + a.change((d) => { + d.text = "two"; + }); + await repoA.flush(); + await pause(1500); + const readsAfter = storageB.reads.length; + await pause(500); + + expect(storageB.reads.length).toBe(readsAfter); + const attempts = storageB.reads + .slice(readMark) + .filter((k) => k.startsWith("subduction/fragment-blobs/")).length; + expect(attempts).toBe(4); + expect(syncB.mock.calls.length).toBe(roundsB); + expect(b.doc().text).toBe("one"); + const said = warn.mock.calls.filter((call) => + call.some( + (arg) => typeof arg === "string" && arg.includes("reload from storage failed") + ) + ); + expect(said).toHaveLength(1); + warn.mockRestore(); + storageB.failRanges = 0; + }); + + it("a sibling's word about the server lands without a round of our own", async () => { + const name = channelName(); + const { storageA, syncA, syncB, a, b } = await shared(name); + const sid = sidOf(storageA); + const heads = encodeHeads(Automerge.getHeads(a.doc())); + const roundsA = syncA.mock.calls.length; + const roundsB = syncB.mock.calls.length; + + const announcer = new BroadcastChannel(name); + announcer.postMessage({ + kind: "remote", + sid, + storageId: SERVER, + heads, + timestamp: Date.now(), + }); + await pause(400); + announcer.close(); + + expect(a.getSyncInfo(SERVER)?.lastHeads).toEqual(heads); + expect(b.getSyncInfo(SERVER)?.lastHeads).toEqual(heads); + expect(syncA.mock.calls.length).toBe(roundsA); + expect(syncB.mock.calls.length).toBe(roundsB); + }); + + it("an older observation does not walk the view of the server backwards", async () => { + const name = channelName(); + const { storageA, repoA, a } = await shared(name); + const sid = sidOf(storageA); + const older = encodeHeads(Automerge.getHeads(a.doc())); + a.change((d) => { + d.text = "two"; + }); + await repoA.flush(); + const newer = encodeHeads(Automerge.getHeads(a.doc())); + expect(newer).not.toEqual(older); + + const surfaced: string[][] = []; + a.on("remote-heads", ({ heads }) => surfaced.push([...heads])); + + const now = Date.now(); + const announcer = new BroadcastChannel(name); + const say = (heads: string[], timestamp: number) => + announcer.postMessage({ + kind: "remote", + sid, + storageId: SERVER, + heads, + timestamp, + }); + + say(newer, now); + await pause(250); + say(older, now - 5000); + await pause(250); + say(newer, now + 5000); + await pause(250); + announcer.close(); + + expect(a.getSyncInfo(SERVER)?.lastHeads).toEqual(newer); + expect(surfaced).toEqual([newer]); + }); + + it("a relayed observation is not relayed again", async () => { + const name = channelName(); + const { storageA, a } = await shared(name); + const sid = sidOf(storageA); + const heads = encodeHeads(Automerge.getHeads(a.doc())); + + const heard: Array<{ kind?: string }> = []; + const listener = new BroadcastChannel(name); + listener.onmessage = ({ data }) => heard.push(data); + + const announcer = new BroadcastChannel(name); + const message = { + kind: "remote", + sid, + storageId: SERVER, + heads, + timestamp: Date.now(), + }; + announcer.postMessage(message); + await pause(400); + announcer.close(); + listener.close(); + + expect(heard.filter((m) => m.kind === "remote")).toHaveLength(1); + }); + + it("the backstop gives up once no round could catch the server up", async () => { + const name = channelName(); + const { storageA, syncA, a, b } = await shared(name, connected); + const sid = sidOf(storageA); + const announcer = new BroadcastChannel(name); + const nudge = () => + announcer.postMessage({ kind: "saved", sid, heads: [bogusHead()] }); + + // Tell both tabs the server holds everything they hold, then walk a + // recompute past it so the backstop budget is known to be untouched. + announcer.postMessage({ + kind: "remote", + sid, + storageId: SERVER, + heads: encodeHeads(Automerge.getHeads(a.doc())), + timestamp: Date.now(), + }); + nudge(); + await pause(400); + const mark = syncA.mock.calls.length; + + // B's commit reaches A's handle by way of the shared database, so A + // holds a hash it will never push: nothing A can do makes the server's + // known heads contain A's. + b.change((d) => { + d.text = "from b"; + }); + await until(() => a.doc().text === "from b", 5000, "A sees B's change"); + + const opened = () => syncA.mock.calls.length - mark; + const deadline = Date.now() + 15_000; + while (opened() < MAX_BACKSTOP_SYNCS && Date.now() < deadline) { + nudge(); + await pause(150); + } + expect(opened()).toBe(MAX_BACKSTOP_SYNCS); + + for (let i = 0; i < 6; i++) { + nudge(); + await pause(150); + } + await pause(600); + announcer.close(); + + expect(opened()).toBe(MAX_BACKSTOP_SYNCS); + }); + + it("a sibling's word about the server keeps the backstop quiet", async () => { + const name = channelName(); + const { storageA, syncA, a } = await shared(name, connected); + const sid = sidOf(storageA); + const announcer = new BroadcastChannel(name); + + announcer.postMessage({ + kind: "remote", + sid, + storageId: SERVER, + heads: encodeHeads(Automerge.getHeads(a.doc())), + timestamp: Date.now(), + }); + await pause(300); + const mark = syncA.mock.calls.length; + + for (let i = 0; i < 6; i++) { + announcer.postMessage({ kind: "saved", sid, heads: [bogusHead()] }); + await pause(150); + } + await pause(400); + announcer.close(); + + expect(syncA.mock.calls.length).toBe(mark); + }); + + it("a sibling's word arriving during the wait costs no round and no write", async () => { + const name = channelName(); + const { storageA, syncA, a, b } = await shared(name, connectedWithin(600)); + const sid = sidOf(storageA); + const announcer = new BroadcastChannel(name); + const serverHolds = (handle: { doc(): Doc }) => + announcer.postMessage({ + kind: "remote", + sid, + storageId: SERVER, + heads: encodeHeads(Automerge.getHeads(handle.doc())), + timestamp: Date.now(), + }); + + serverHolds(a); + announcer.postMessage({ kind: "saved", sid, heads: [bogusHead()] }); + await pause(900); + const mark = syncA.mock.calls.length; + const writeMark = storageA.writes.length; + + b.change((d) => { + d.text = "from b"; + }); + await until(() => a.doc().text === "from b", 5000, "A sees B's change"); + // Within the backstop's window: the server has it after all. + serverHolds(b); + await pause(1500); + announcer.close(); + + expect(syncA.mock.calls.length).toBe(mark); + expect(commitWrites(storageA, writeMark)).toEqual([]); + }); + + it("offline, a sibling's commit is written again so this tab can push it", async () => { + const { storageA, storageB, repoA, syncA, a, b } = await shared( + channelName(), + offlineWithin(300) + ); + const sid = sidOf(storageA); + const rounds = syncA.mock.calls.length; + const writeMark = storageA.writes.length; + + b.change((d) => { + d.text = "from b"; + }); + await until(() => a.doc().text === "from b", 5000, "A sees B's change"); + const hash = Automerge.getHeads(b.doc())[0]; + const bCommit = `subduction/commits/${sid}/${hash}`; + expect(storageB.writes).toContain(bCommit); + + // The reload put B's commit in A's handle and in A's knownHashes, but + // not in A's tree: as it stands A can never push it. + expect(await treeCommits(repoA, a.documentId)).not.toContain(hash); + expect(commitWrites(storageA, writeMark)).toEqual([]); + + // Settled, still diverged, and no round to fix it: re-ingest. + await until( + async () => (await treeCommits(repoA, a.documentId)).includes(hash), + 5000, + "A's tree gains B's commit" + ); + expect(commitWrites(storageA, writeMark)).toEqual([bCommit]); + + // Once. The set that named it is empty now and nothing refills it. + await pause(1200); + expect(commitWrites(storageA, writeMark)).toEqual([bCommit]); + // The backstop opened no round: the one round here is the one the + // save itself arms, the same one a local edit would have armed. + expect(syncA.mock.calls.length).toBe(rounds + 1); + }); + + it("the re-ingest happens once even when no round ever confirms it", async () => { + const name = channelName(); + const { storageA, syncA, a, b } = await shared(name, connected); + const sid = sidOf(storageA); + const announcer = new BroadcastChannel(name); + const nudge = () => + announcer.postMessage({ kind: "saved", sid, heads: [bogusHead()] }); + + announcer.postMessage({ + kind: "remote", + sid, + storageId: SERVER, + heads: encodeHeads(Automerge.getHeads(a.doc())), + timestamp: Date.now(), + }); + nudge(); + await pause(400); + const rounds = syncA.mock.calls.length; + const writeMark = storageA.writes.length; + + // `syncA` is a mock: it reports success and never tells A the server + // took anything, so containment is never observed however many rounds + // run. The write must not follow the rounds. + b.change((d) => { + d.text = "from b"; + }); + await until(() => a.doc().text === "from b", 5000, "A sees B's change"); + for (let i = 0; i < 10; i++) { + nudge(); + await pause(150); + } + await pause(600); + announcer.close(); + + expect(commitWrites(storageA, writeMark)).toHaveLength(1); + expect(syncA.mock.calls.length).toBeGreaterThan(rounds); + }); + + it("a save in flight when the backstop fires does not swallow the re-ingest", async () => { + const { storageA, repoA, a, b } = await shared( + channelName(), + offlineWithin(2000) + ); + + b.change((d) => { + d.text = "from b"; + }); + await until(() => a.doc().text === "from b", 5000, "A sees B's change"); + const hash = Automerge.getHeads(b.doc())[0]; + expect(await treeCommits(repoA, a.documentId)).not.toContain(hash); + + // A's own write is held open across the backstop's expiry. That save + // restores the save baseline as it finishes, so a re-ingest that + // cleared the baseline before waiting for it would find its own save + // returning on `#save`'s heads fast path, having stored nothing. + storageA.blockSaves(); + a.change((d) => { + d.text = "from b, and a"; + }); + await pause(2500); + storageA.releaseSaves(); + + await until( + async () => (await treeCommits(repoA, a.documentId)).includes(hash), + 8000, + "A's tree gains B's commit with a save already in flight" + ); + }); + + it("an announcement arriving during a reload is not lost with it", async () => { + const { storageA, repoA, a, b } = await shared( + channelName(), + offlineWithin(300) + ); + // The read snapshots and then takes its time, so B's second + // announcement lands while the first reload is still in flight. + storageA.rangeDelayMs = 500; + + b.change((d) => { + d.text = "one from b"; + }); + const first = Automerge.getHeads(b.doc())[0]; + await pause(250); + b.change((d) => { + d.text = "two from b"; + }); + const second = Automerge.getHeads(b.doc())[0]; + expect(second).not.toBe(first); + + await until( + () => a.doc().text === "two from b", + 8000, + "A sees both of B's changes" + ); + // The second announcement was set while the first reload was running. + // Cleared by that reload, it would come back as an ordinary load and + // file B's second commit as known and not stranded — unpushable, and + // invisible to the re-ingest. + await until( + async () => { + const tree = await treeCommits(repoA, a.documentId); + return tree.includes(first) && tree.includes(second); + }, + 8000, + "A's tree gains both of B's commits" + ); + storageA.rangeDelayMs = 0; + }); + + it("the re-ingest's announcement dies with three tabs listening", async () => { + const name = channelName(); + const { repoA, repoB, mkSibling, a, b } = await shared( + name, + offlineWithin(300) + ); + const repoC = mkSibling(); + vi.spyOn(await repoC.subduction, "syncWithAllPeers").mockImplementation( + succeededRound as never + ); + const c = await withTimeout(repoC.find(a.url), 5000, "C finds the doc"); + await pause(500); + + const heard: Array<{ kind?: string }> = []; + const listener = new BroadcastChannel(name); + listener.onmessage = ({ data }) => heard.push(data); + + b.change((d) => { + d.text = "from b"; + }); + await until( + () => a.doc().text === "from b" && c.doc().text === "from b", + 5000, + "A and C see B's change" + ); + await pause(2000); + listener.close(); + + // B's own save, then one re-ingest each from A and C. Each of those + // announces heads the other two already hold, so neither reloads and + // nothing announces again. + expect(heard.filter((m) => m.kind === "saved")).toHaveLength(3); + const hash = Automerge.getHeads(b.doc())[0]; + for (const [repo, label] of [ + [repoA, "A"], + [repoB, "B"], + [repoC, "C"], + ] as const) { + expect(await treeCommits(repo, a.documentId), label).toContain(hash); + } + }); + + it("nothing is posted on the channel after shutdown closes it", async () => { + const closed = new WeakSet(); + const afterClose: unknown[] = []; + const realClose = BroadcastChannel.prototype.close; + const realPost = BroadcastChannel.prototype.postMessage; + vi.spyOn(BroadcastChannel.prototype, "close").mockImplementation( + function (this: BroadcastChannel) { + closed.add(this); + return realClose.call(this); + } + ); + vi.spyOn(BroadcastChannel.prototype, "postMessage").mockImplementation( + function (this: BroadcastChannel, message: unknown) { + if (closed.has(this)) afterClose.push(message); + return realPost.call(this, message); + } + ); + + try { + const { repoA, a } = await shared(channelName()); + const done = repoA.shutdown(); + a.change((d) => { + d.text = "after the flush"; + }); + await done; + await pause(400); + expect(afterClose).toEqual([]); + } finally { + vi.restoreAllMocks(); + } + }); + + it("each source unrefs its channel so it does not keep a Node process alive", async () => { + const unref = vi.spyOn( + BroadcastChannel.prototype as unknown as { unref(): void }, + "unref" + ); + await siblings(channelName()); + expect(unref).toHaveBeenCalledTimes(2); + unref.mockRestore(); + }); +}); diff --git a/core/patchwork/package.json b/core/patchwork/package.json index 625a73b43..d9ff15827 100644 --- a/core/patchwork/package.json +++ b/core/patchwork/package.json @@ -94,7 +94,7 @@ "vite": "^7.3.5" }, "scripts": { - "build": "tsc && node build-head.js && cp src/global.css dist/global.css && cp src/client.d.ts dist/client.d.ts", + "build": "tsc && node build-head.js && cp src/global.css dist/global.css && cp src/client.d.ts dist/client.d.ts && rm -rf dist/patches && mkdir -p dist/patches && cp ../../patches/@automerge__automerge-repo@*.patch dist/patches/", "dev": "tsc -w --preserveWatchOutput" } } diff --git a/core/patchwork/src/head.ts b/core/patchwork/src/head.ts index 8bbe33f48..c74e180e0 100644 --- a/core/patchwork/src/head.ts +++ b/core/patchwork/src/head.ts @@ -1,9 +1,11 @@ -import externals from "@inkandswitch/patchwork-bootloader/externals-list"; +import externals, { + slim, +} from "@inkandswitch/patchwork-bootloader/externals-list"; const importmap: { imports: Record } = { imports: {} }; for (const name of externals) { - importmap.imports[name] = `/packages/${name}.js`; + importmap.imports[name] = `/packages/${slim[name] ?? name}.js`; } const script = document.createElement("script"); diff --git a/core/patchwork/src/index.ts b/core/patchwork/src/index.ts index 6a25d0b13..0ca61ec0f 100644 --- a/core/patchwork/src/index.ts +++ b/core/patchwork/src/index.ts @@ -66,8 +66,8 @@ declare global { interface Window { patchwork: Patchwork; repo: Repo; - Automerge: typeof import("@automerge/automerge"); - AutomergeRepo: typeof import("@automerge/automerge-repo"); + Automerge: typeof import("@automerge/automerge/slim"); + AutomergeRepo: typeof import("@automerge/automerge-repo/slim"); hive?: AutomergeRepoKeyhive; } } @@ -138,9 +138,9 @@ async function doSetup(options: PatchworkOptions): Promise { // `window.patchwork` handle is deliberately not set here — the caller does // `window.patchwork = await setup(...)`. window.repo = repo; - window.Automerge = Automerge as typeof import("@automerge/automerge"); + window.Automerge = Automerge as typeof import("@automerge/automerge/slim"); window.AutomergeRepo = - AutomergeRepo as typeof import("@automerge/automerge-repo"); + AutomergeRepo as typeof import("@automerge/automerge-repo/slim"); if (hive) window.hive = hive; (hive?.networkAdapter as any)?.syncKeyhive?.(); diff --git a/core/patchwork/src/repo.ts b/core/patchwork/src/repo.ts index aabd83806..46861fdd9 100644 --- a/core/patchwork/src/repo.ts +++ b/core/patchwork/src/repo.ts @@ -1,17 +1,14 @@ import { initializeWasm, Repo } from "@automerge/vanillajs/slim"; -import { IndexedDBWorkerStorageAdapter } from "@automerge/automerge-repo-storage-indexeddb/IndexedDBWorkerStorageAdapter"; -import { siblingAdapters } from "@inkandswitch/patchwork-bootloader/siblings"; +import { IndexedDBStorageAdapter } from "@automerge/automerge-repo-storage-indexeddb"; import { loadOrCreateSigner } from "@inkandswitch/patchwork-bootloader/signer"; import * as AutomergeRepo from "@automerge/automerge-repo/slim"; -import { - initKeyhiveWasm, - initializeAutomergeRepoKeyhive, - type AutomergeRepoKeyhive, - type SyncServerSelection, +import type { + AutomergeRepoKeyhive, + SyncServerSelection, } from "@automerge/automerge-repo-keyhive"; // eslint-disable-next-line -// @ts-ignore — initSync is a wasm-bindgen runtime helper not in the .d.ts -import { initSync as initSubductionSync } from "@automerge/automerge-subduction/slim"; +// @ts-ignore — the default init is a wasm-bindgen runtime helper not in the .d.ts +import initSubduction from "@automerge/automerge-subduction/slim"; import { keyhiveStorageName, storagePrefix, @@ -37,14 +34,10 @@ const syncServer = let wasmReady: Promise | undefined; export function initWasm(): Promise { if (!wasmReady) { - wasmReady = (async () => { - const [automergeWasm, subductionWasm] = await Promise.all([ - fetch("/automerge.wasm").then((r) => r.bytes()), - fetch("/subduction.wasm").then((r) => r.bytes()), - ]); - await initializeWasm(automergeWasm); - initSubductionSync(subductionWasm); - })(); + wasmReady = Promise.all([ + initializeWasm(new Request("/automerge.wasm")), + initSubduction(new Request("/subduction.wasm")), + ]).then(() => undefined); } return wasmReady; } @@ -56,13 +49,16 @@ export type TabRepo = { }; /** - * The tab's own node: this origin's IndexedDB, a socket to the sync server, - * and the siblings channel to every other Repo on the origin. Nothing is - * shared with other tabs except the database underneath. + * The tab's own node: this origin's IndexedDB, opened on this thread, a socket + * to the sync server, and the storage channel to every other Repo on the + * origin. Nothing is shared with other tabs except the database underneath + * and the identity every node on the origin signs with. */ export async function createRepo(): Promise { if (syncServer.keyhive) { log("setting up keyhive"); + const { initKeyhiveWasm, initializeAutomergeRepoKeyhive } = + await import("@automerge/automerge-repo-keyhive"); initKeyhiveWasm(); const { hive, repo } = await initializeAutomergeRepoKeyhive({ // ARK injects an `idFactory` deriving document ids from keyhive. A site @@ -73,16 +69,16 @@ export async function createRepo(): Promise { ? repoConfig : { ...repoConfig, idFactory } ), - storage: new IndexedDBWorkerStorageAdapter(keyhiveStorageName), + storage: new IndexedDBStorageAdapter(keyhiveStorageName), peerIdSuffix: storagePrefix + Math.random().toString(36).slice(2), automaticArchiveIngestion: true, cachingMode: "periodic", // ARK selects the relay via `syncServer`, defaulting to "subduction". syncServer: syncServer.keyhive, repo: { - storage: new IndexedDBWorkerStorageAdapter(), + storage: new IndexedDBStorageAdapter(), subductionWebsocketEndpoints: [syncServer.url], - subductionAdapters: siblingAdapters(), + headsChannel: `${storagePrefix}-heads`, enableRemoteHeadsGossiping: true, }, }); @@ -93,7 +89,7 @@ export async function createRepo(): Promise { // The signer is explicit rather than the Repo's internal default so that // every tab on this origin signs as the same peer, and so the identity the // tab presents to the server can be shown on window.patchwork. - const storage = new IndexedDBWorkerStorageAdapter(); + const storage = new IndexedDBStorageAdapter(); const signer = await loadOrCreateSigner(storage); const repo = new Repo({ signer, @@ -101,7 +97,7 @@ export async function createRepo(): Promise { peerId: `${storagePrefix}-tab-${crypto.randomUUID()}` as AutomergeRepo.PeerId, subductionWebsocketEndpoints: [syncServer.url], - subductionAdapters: siblingAdapters(), + headsChannel: `${storagePrefix}-heads`, enableRemoteHeadsGossiping: true, }); const signerIdentity = { diff --git a/core/patchwork/src/vite/config-plugin.ts b/core/patchwork/src/vite/config-plugin.ts index 5ee5c4dd1..c54ae50a2 100644 --- a/core/patchwork/src/vite/config-plugin.ts +++ b/core/patchwork/src/vite/config-plugin.ts @@ -1,5 +1,6 @@ import type { Plugin } from "vite"; import wasm from "vite-plugin-wasm"; +import { patches } from "./patches-plugin.js"; import { DEFAULT_STORAGE_PREFIX } from "@inkandswitch/patchwork-bootloader/storage"; import { DEFAULT_TITLE } from "../site-kit/options.js"; import type { PatchworkVitePluginOptions } from "./patchwork-plugin.js"; @@ -75,7 +76,7 @@ export function config(options: PatchworkVitePluginOptions = {}): Plugin { ? undefined : { format: options.worker?.format ?? "es", - plugins: () => [wasm()], + plugins: () => [wasm(), patches({ complete: false })], }, build: { target: "firefox150", diff --git a/core/patchwork/src/vite/dev-plugin.ts b/core/patchwork/src/vite/dev-plugin.ts index bb20adc24..57c7cc165 100644 --- a/core/patchwork/src/vite/dev-plugin.ts +++ b/core/patchwork/src/vite/dev-plugin.ts @@ -2,10 +2,10 @@ import type { Plugin } from "vite"; import { readFile } from "node:fs/promises"; import { fileURLToPath } from "node:url"; import * as esbuild from "esbuild"; -import { wasmAssets } from "@inkandswitch/patchwork-bootloader/externals"; +import { slim, wasmAssets } from "@inkandswitch/patchwork-bootloader/externals"; import type { PatchworkVitePluginOptions } from "./patchwork-plugin.js"; import { buildDefines } from "./config-plugin.js"; -import { builtins, devDependencyId } from "./importmap-plugin.js"; +import { builtins, chunks, devDependencyId } from "./importmap-plugin.js"; import { workers } from "./service-worker-plugin.js"; const PATCHWORK_CSS = "/@inkandswitch/patchwork/global.css"; @@ -23,9 +23,7 @@ const stylesheets: Record = { ), }; -const builtinPaths = new Map( - Object.entries(builtins).map(([id, fileName]) => [fileName, id]) -); +const builtinPaths = new Map(chunks.map((id) => [builtins[id]!, id])); /** * Workers are `type: "module"` scripts the browser fetches directly, so import @@ -40,7 +38,7 @@ function externalBuiltins(): esbuild.Plugin { build.onResolve({ filter: /.*/ }, (args) => { if (!(args.path in builtins)) return null; return { - path: `/@id/${devDependencyId(args.path)}`, + path: `/@id/${devDependencyId(slim[args.path] ?? args.path)}`, external: true, }; }); diff --git a/core/patchwork/src/vite/importmap-plugin.ts b/core/patchwork/src/vite/importmap-plugin.ts index 406fdc1b6..05c57b233 100644 --- a/core/patchwork/src/vite/importmap-plugin.ts +++ b/core/patchwork/src/vite/importmap-plugin.ts @@ -12,12 +12,16 @@ import { relative } from "node:path"; * packages as real dependencies) — reused here rather than duplicated. */ import externals, { + slim, resolveExternal, emitWasmAssets, } from "@inkandswitch/patchwork-bootloader/externals"; export const builtins = externals.reduce( - (builtins, name) => ((builtins[name] = `/packages/${name}.js`), builtins), + (builtins, name) => ( + (builtins[name] = `/packages/${slim[name] ?? name}.js`), + builtins + ), {} as Record ); @@ -25,6 +29,8 @@ export const builtins = externals.reduce( // bare-import it just like the other @inkandswitch/patchwork-* packages. builtins["@inkandswitch/patchwork"] = "/packages/@inkandswitch/patchwork.js"; +export const chunks = Object.keys(builtins).filter((id) => !(id in slim)); + /** * Node resolves a package's own name from within its own source when its * package.json has "exports" (self-reference) — the same mechanism @@ -82,7 +88,7 @@ export function importmap(options?: PatchworkVitePluginOptions): Plugin { config() { return { optimizeDeps: { - include: Object.keys(builtins).map(devDependencyId), + include: chunks.map(devDependencyId), }, }; }, @@ -91,10 +97,10 @@ export function importmap(options?: PatchworkVitePluginOptions): Plugin { }, async buildStart() { if (serve) return; - for (const [id, fileName] of Object.entries(builtins)) { + for (const id of chunks) { this.emitFile({ type: "chunk", - fileName: fileName.slice(1), + fileName: builtins[id]!.slice(1), id: await resolveBuiltin(this, id), preserveSignature: "strict", }); @@ -108,7 +114,7 @@ export function importmap(options?: PatchworkVitePluginOptions): Plugin { // point the site's own imports at the same copy we emit as a chunk, // otherwise rollup bundles a second one out of the site's node_modules // and you end up with two automerges racing to init the same wasm - return resolveBuiltin(this, id); + return resolveBuiltin(this, slim[id] ?? id); } if (id in importmap.imports) { return { id: importmap.imports[id], external: true }; @@ -127,7 +133,7 @@ export function importmap(options?: PatchworkVitePluginOptions): Plugin { await optimizer?.init(); const root = context.server.config.root; for (const id of Object.keys(builtins)) { - const dependency = devDependencyId(id); + const dependency = devDependencyId(slim[id] ?? id); const optimized = optimizer?.metadata.optimized[dependency]; activeImportmap.imports[id] = optimized ? `/${relative(root, optimized.file)}?v=${optimized.browserHash}` diff --git a/core/patchwork/src/vite/patches-plugin.ts b/core/patchwork/src/vite/patches-plugin.ts index 0d1943cae..0f3ffd719 100644 --- a/core/patchwork/src/vite/patches-plugin.ts +++ b/core/patchwork/src/vite/patches-plugin.ts @@ -1,73 +1,96 @@ -import { readFile } from "node:fs/promises"; +import { readdir, readFile } from "node:fs/promises"; import { join } from "node:path"; +import { fileURLToPath } from "node:url"; import type { Plugin } from "vite"; /** * Source patches to @automerge/automerge-repo, applied as it passes through * the bundler. * - * The same edits live in this repo's `patches/` as a pnpm patch, which is what + * The edits are this repo's `patches/` pnpm patch, which is what * makes our own typecheck see the widened `role` type. A pnpm patch only * exists in this repo's node_modules, though: a site that installs * @inkandswitch/patchwork resolves automerge-repo out of its own tree and - * would bundle it unpatched. These run there too. + * would bundle it unpatched. The plugin applies the patch there too. * - * Both edits are upstream-shaped and meant to be deleted once subduction takes + * The edits are upstream-shaped and meant to be deleted once subduction takes * them. Until then they are pinned to one automerge-repo version and every - * anchor has to match, so a dependency bump fails the build instead of quietly + * hunk has to match, so a dependency bump fails the build instead of quietly * un-patching it. */ const AUTOMERGE_REPO = "@automerge/automerge-repo"; -const VERSION = "2.6.0-subduction.48"; - -type Edit = { find: string; replace: string }; - -const PATCHES: Record = { - "dist/subduction/AdapterConnections.js": [ - { - find: ` if (role === "accept") { - await subduction.acceptTransport(transport, serviceName); - } - else { - await subduction.connectTransport(transport, serviceName); - }`, - replace: ` const initiate = role === "mesh" ? this.#localPeerId < peerId : role !== "accept"; - if (initiate) { - await subduction.connectTransport(transport, serviceName); - } - else { - await subduction.acceptTransport(transport, serviceName); - }`, - }, - ], - "dist/subduction/SubductionConnections.js": [ - { - find: ` if (state === "connecting") - return true;`, - replace: ` // "awaiting-reconnect" counts: the loop is between attempts, not - // given up, so a query should wait rather than report unavailable. - if (state === "connecting" || state === "awaiting-reconnect") - return true;`, - }, - ], +const PATCHES_DIR = fileURLToPath(new URL("../patches/", import.meta.url)); +const PATCH_PREFIX = `${AUTOMERGE_REPO.replace("/", "__")}@`; + +export type Hunk = { + header: string; + oldStart: number; + newStart: number; + lines: string[]; }; -function match(id: string): { file: string; root: string } | undefined { - const path = id.replace(/\\/g, "/").split("?")[0]; - for (const file of Object.keys(PATCHES)) { - const suffix = `/${AUTOMERGE_REPO}/${file}`; - if (path.endsWith(suffix)) { - return { - file, - root: path.slice(0, -suffix.length) + `/${AUTOMERGE_REPO}`, +type Patch = { version: string; files: Record }; + +export function parsePatch(text: string): Record { + const files: Record = {}; + let hunks: Hunk[] | undefined; + let hunk: Hunk | undefined; + for (const line of text.split("\n")) { + const file = /^diff --git a\/(\S+) b\//.exec(line); + if (file) { + hunks = /^dist\/.+\.js$/.test(file[1]) + ? (files[file[1]] = []) + : undefined; + hunk = undefined; + continue; + } + const header = /^@@ -(\d+)(?:,\d+)? \+(\d+)(?:,\d+)? @@/.exec(line); + if (header) { + hunk = { + header: line, + oldStart: +header[1], + newStart: +header[2], + lines: [], }; + hunks?.push(hunk); + continue; } + if (hunk && /^[ +-]/.test(line)) hunk.lines.push(line); } + return files; +} + +async function loadPatch(): Promise { + const names = await readdir(PATCHES_DIR).catch(() => [] as string[]); + const found = names.filter( + (name) => name.startsWith(PATCH_PREFIX) && name.endsWith(".patch") + ); + if (found.length !== 1) { + throw new Error( + `@inkandswitch/patchwork: expected exactly one ${PATCH_PREFIX}*.patch ` + + `in ${PATCHES_DIR}, found ${found.length}.` + ); + } + const version = found[0].slice(PATCH_PREFIX.length, -".patch".length); + const text = await readFile(join(PATCHES_DIR, found[0]), "utf8"); + return { version, files: parsePatch(text) }; +} + +function match(id: string): { file: string; root: string } | undefined { + const path = id.replace(/\\/g, "/").split("?")[0]; + const at = path.lastIndexOf(`/${AUTOMERGE_REPO}/dist/`); + if (at === -1 || !path.endsWith(".js")) return; + const root = path.slice(0, at) + `/${AUTOMERGE_REPO}`; + return { file: path.slice(root.length + 1), root }; } const versions = new Map>(); -async function assertVersion(root: string, file: string): Promise { +async function assertVersion( + root: string, + file: string, + expected: string +): Promise { let version = versions.get(root); if (!version) { version = readFile(join(root, "package.json"), "utf8").then( @@ -82,41 +105,64 @@ async function assertVersion(root: string, file: string): Promise { ); versions.set(root, version); } - if ((await version) !== VERSION) { + if ((await version) !== expected) { throw new Error( `@inkandswitch/patchwork: ${AUTOMERGE_REPO} is ${await version}, and the ` + - `source patches in patches-plugin.ts are written against ${VERSION}. ` + - `Re-check them against the new version (${file} is one of the files ` + - `they edit), then bump VERSION.` + `pnpm patch in patches/ is for ${expected}. Re-make it against the ` + + `new version (${file} is one of the files it edits).` ); } } -function apply(code: string, file: string): string { - return PATCHES[file].reduce((code, { find, replace }) => { - if (code.includes(replace)) return code; - const matches = code.split(find).length - 1; - if (matches !== 1) { +function without(hunk: Hunk, prefix: string): string[] { + return hunk.lines + .filter((line) => line[0] !== prefix) + .map((line) => line.slice(1)); +} + +export function apply(code: string, file: string, hunks: Hunk[]): string { + const lines = code.split("\n"); + const sits = (start: number, expected: string[]) => + expected.every((line, i) => lines[start + i] === line); + if (hunks.every((hunk) => sits(hunk.newStart - 1, without(hunk, "-")))) { + return code; + } + let delta = 0; + for (const hunk of hunks) { + const start = hunk.oldStart - 1 + delta; + const before = without(hunk, "+"); + const after = without(hunk, "-"); + if (!sits(start, before)) { throw new Error( - `@inkandswitch/patchwork: the source patch for ${AUTOMERGE_REPO}'s ` + - `${file} matched ${matches} times, expected 1. The file is the ` + - `version it says it is, so the patch needs rewriting against it.` + `@inkandswitch/patchwork: ${AUTOMERGE_REPO}'s ${file} doesn't match ` + + `the pnpm patch at "${hunk.header}". The file is the version it ` + + `says it is, so the patch needs re-making against it.` ); } - return code.replace(find, replace); - }, code); + lines.splice(start, before.length, ...after); + delta += after.length - before.length; + } + return lines.join("\n"); } -async function patchFile(id: string): Promise { +async function patchFile( + id: string, + patch: Patch +): Promise { const found = match(id); - if (!found) return; - await assertVersion(found.root, found.file); - return apply(await readFile(id, "utf8"), found.file); + const hunks = found && patch.files[found.file]; + if (!found || !hunks) return; + await assertVersion(found.root, found.file, patch.version); + return apply(await readFile(id, "utf8"), found.file, hunks); } -export function patches(): Plugin { +export function patches({ + complete = true, +}: { complete?: boolean } = {}): Plugin { const seen = new Set(); let serve = false; + let loading: Promise | undefined; + const patch = () => (loading ??= loadPatch()); return { name: "@patchwork/patches", @@ -143,9 +189,9 @@ export function patches(): Plugin { ): void; }) { build.onLoad( - { filter: /automerge-repo[\\/]dist[\\/]subduction[\\/]/ }, + { filter: /automerge-repo[\\/]dist[\\/]/ }, async ({ path }) => { - const contents = await patchFile(path); + const contents = await patchFile(path, await patch()); return contents ? { contents, loader: "js" } : undefined; } ); @@ -164,14 +210,18 @@ export function patches(): Plugin { async transform(code, id) { const found = match(id); if (!found) return; - await assertVersion(found.root, found.file); + const { version, files } = await patch(); + const hunks = files[found.file]; + if (!hunks) return; + await assertVersion(found.root, found.file, version); seen.add(found.file); - return apply(code, found.file); + return apply(code, found.file, hunks); }, - buildEnd() { - if (serve) return; - const missing = Object.keys(PATCHES).filter((file) => !seen.has(file)); + async buildEnd(error) { + if (serve || error || !complete) return; + const { files } = await patch(); + const missing = Object.keys(files).filter((file) => !seen.has(file)); if (missing.length) { throw new Error( `@inkandswitch/patchwork: ${AUTOMERGE_REPO}'s ${missing.join(", ")} ` + diff --git a/core/patchwork/src/vite/service-worker-plugin.ts b/core/patchwork/src/vite/service-worker-plugin.ts index 90750744f..3b1bb80b2 100644 --- a/core/patchwork/src/vite/service-worker-plugin.ts +++ b/core/patchwork/src/vite/service-worker-plugin.ts @@ -36,6 +36,16 @@ export function serviceworker(): Plugin { return { name: "@patchwork/service-worker", enforce: "pre", + config() { + return { + build: { + modulePreload: { + resolveDependencies: (_url, deps, { hostId }) => + workers.some(({ fileName }) => fileName === hostId) ? [] : deps, + }, + }, + }; + }, configResolved(config) { serve = config.command === "serve"; }, diff --git a/core/plugins/src/datatypes.ts b/core/plugins/src/datatypes.ts index 6497bf5dc..7f906c77b 100644 --- a/core/plugins/src/datatypes.ts +++ b/core/plugins/src/datatypes.ts @@ -4,7 +4,7 @@ import { stringifyAutomergeUrl, type DocHandle, type Repo, -} from "@automerge/automerge-repo"; +} from "@automerge/automerge-repo/slim"; import type { LoadablePlugin, LoadedPlugin, diff --git a/core/plugins/src/keyhive.ts b/core/plugins/src/keyhive.ts index bef4401ad..c95700af3 100644 --- a/core/plugins/src/keyhive.ts +++ b/core/plugins/src/keyhive.ts @@ -1,10 +1,15 @@ -import { docIdFromAutomergeUrl } from "@automerge/automerge-repo-keyhive"; -import type { AutomergeUrl } from "@automerge/automerge-repo/slim"; +import { + parseAutomergeUrl, + type AutomergeUrl, +} from "@automerge/automerge-repo/slim"; export function isKeyhiveDoc(url: AutomergeUrl): boolean { try { - docIdFromAutomergeUrl(url); - return true; + const { binaryDocumentId } = parseAutomergeUrl(url); + return ( + binaryDocumentId.length >= 32 && + binaryDocumentId.subarray(16, 32).some((byte) => byte !== 0) + ); } catch { return false; } diff --git a/core/plugins/src/tools.ts b/core/plugins/src/tools.ts index 0537822ad..12e4dd672 100644 --- a/core/plugins/src/tools.ts +++ b/core/plugins/src/tools.ts @@ -1,4 +1,8 @@ -import type { AutomergeUrl, DocHandle, Repo } from "@automerge/automerge-repo"; +import type { + AutomergeUrl, + DocHandle, + Repo, +} from "@automerge/automerge-repo/slim"; import type { LoadablePlugin, LoadedPlugin, diff --git a/package.json b/package.json index 5a076f7f6..fafbfc1ef 100644 --- a/package.json +++ b/package.json @@ -14,6 +14,7 @@ "dev": "pnpm -r --filter ...{sites/${SITE:-patchwork.inkandswitch.com}}... --parallel dev", "format": "prettier --write \"**/*.{js,jsx,ts,tsx,json,css,md,html,yaml,yml}\"", "format:check": "prettier --check \"**/*.{js,jsx,ts,tsx,json,css,md,html,yaml,yml}\"", + "lint": "node scripts/lint-slim-imports.mts", "link-cli": "pnpm --ignore-workspace --dir scripts/module-settings install && pnpm link --global", "preinstall": "npx only-allow pnpm", "preview": "pnpm --filter patchwork.inkandswitch.com preview", diff --git a/packages/edge-handles/src/edge-handle.ts b/packages/edge-handles/src/edge-handle.ts index ce477135e..e08df3223 100644 --- a/packages/edge-handles/src/edge-handle.ts +++ b/packages/edge-handles/src/edge-handle.ts @@ -40,7 +40,7 @@ import { type DocHandle, type Repo, type SubChangeFn, -} from "@automerge/automerge-repo"; +} from "@automerge/automerge-repo/slim"; // ─── subscriber set ──────────────────────────────────────────────────────── diff --git a/packages/providers/core/src/index.ts b/packages/providers/core/src/index.ts index 820e820a4..31baceff2 100644 --- a/packages/providers/core/src/index.ts +++ b/packages/providers/core/src/index.ts @@ -1,4 +1,4 @@ -import type { Repo } from "@automerge/automerge-repo"; +import type { Repo } from "@automerge/automerge-repo/slim"; export type JSONValue = | string diff --git a/packages/providers/core/src/overlay-handle.ts b/packages/providers/core/src/overlay-handle.ts index f523c075d..199e94dcc 100644 --- a/packages/providers/core/src/overlay-handle.ts +++ b/packages/providers/core/src/overlay-handle.ts @@ -4,7 +4,7 @@ import { type AutomergeUrl, type DocHandle, type DocumentId, -} from "@automerge/automerge-repo"; +} from "@automerge/automerge-repo/slim"; import { forwardingProxy } from "./forwarding-proxy.js"; diff --git a/packages/providers/core/src/overlay-repo.ts b/packages/providers/core/src/overlay-repo.ts index e543648cc..c5195cb6f 100644 --- a/packages/providers/core/src/overlay-repo.ts +++ b/packages/providers/core/src/overlay-repo.ts @@ -10,7 +10,7 @@ import { type DocumentProgress, type QueryState, type Repo, -} from "@automerge/automerge-repo"; +} from "@automerge/automerge-repo/slim"; import { subscribe } from "./index.js"; import { forwardingProxy } from "./forwarding-proxy.js"; diff --git a/packages/providers/core/src/repo-provider.ts b/packages/providers/core/src/repo-provider.ts index 294fb90cd..b42563cc6 100644 --- a/packages/providers/core/src/repo-provider.ts +++ b/packages/providers/core/src/repo-provider.ts @@ -1,4 +1,4 @@ -import type { AutomergeUrl, Repo } from "@automerge/automerge-repo"; +import type { AutomergeUrl, Repo } from "@automerge/automerge-repo/slim"; import { accept, type SubscribeEvent } from "./index.js"; import type { DocHandleDescriptor } from "./overlay-repo.js"; diff --git a/packages/providers/core/src/types.ts b/packages/providers/core/src/types.ts index 0198bd51e..3c129f7d9 100644 --- a/packages/providers/core/src/types.ts +++ b/packages/providers/core/src/types.ts @@ -3,7 +3,7 @@ import type { DocHandle, DocumentId, DocumentProgress, -} from "@automerge/automerge-repo"; +} from "@automerge/automerge-repo/slim"; /** * Minimal repo surface that both the real `Repo` and overlay repos (e.g. diff --git a/packages/providers/frameworks/react/src/index.ts b/packages/providers/frameworks/react/src/index.ts index f124433d8..73e9b7c26 100644 --- a/packages/providers/frameworks/react/src/index.ts +++ b/packages/providers/frameworks/react/src/index.ts @@ -1,6 +1,10 @@ import { useEffect, useState } from "react"; import { useDocument, useDocHandle } from "@automerge/automerge-repo-react-hooks"; -import type { AutomergeUrl, Doc, DocHandle } from "@automerge/automerge-repo"; +import type { + AutomergeUrl, + Doc, + DocHandle, +} from "@automerge/automerge-repo/slim"; import * as Providers from "@inkandswitch/patchwork-providers"; import type { JSONValue, Selector } from "@inkandswitch/patchwork-providers"; diff --git a/packages/providers/frameworks/solid/src/index.ts b/packages/providers/frameworks/solid/src/index.ts index af06e4c49..a4600e395 100644 --- a/packages/providers/frameworks/solid/src/index.ts +++ b/packages/providers/frameworks/solid/src/index.ts @@ -6,7 +6,11 @@ import { type Accessor, } from "solid-js"; import { createStore, reconcile, type Store } from "solid-js/store"; -import type { AutomergeUrl, Doc, DocHandle } from "@automerge/automerge-repo"; +import type { + AutomergeUrl, + Doc, + DocHandle, +} from "@automerge/automerge-repo/slim"; import * as Providers from "@inkandswitch/patchwork-providers"; import type { JSONArray, diff --git a/patches/@automerge__automerge-repo@2.6.0-subduction.48.patch b/patches/@automerge__automerge-repo@2.6.0-subduction.48.patch index 8779f3ee0..5744da8b0 100644 --- a/patches/@automerge__automerge-repo@2.6.0-subduction.48.patch +++ b/patches/@automerge__automerge-repo@2.6.0-subduction.48.patch @@ -1,7 +1,29 @@ +diff --git a/dist/DocumentSource.d.ts b/dist/DocumentSource.d.ts +index 08d36a6efcf9b93ba170dfe8a6a54e51a1b8f92d..fafdbda628f968bb266588d5a45192f5fb965cde 100644 +--- a/dist/DocumentSource.d.ts ++++ b/dist/DocumentSource.d.ts +@@ -34,7 +34,7 @@ export interface DocumentSource { + * appropriate to participate in the query's availability tracking. */ + attach(query: DocumentQuery): void; + /** Called when a document is removed from the repo. */ +- detach(documentId: DocumentId): void; ++ detach(documentId: DocumentId): void | Promise; + shareConfigChanged(): void; + /** + * Optional: drain any pending writes for the given documents (or all diff --git a/dist/Repo.d.ts b/dist/Repo.d.ts -index 15d4ca65aa613fb567335f09c8fcb8cb9c4a954a..0f89be0e8b7e19ebb706d88685593259ddf67c89 100644 +index 15d4ca65aa613fb567335f09c8fcb8cb9c4a954a..656c8061a9e657ea2677a7308b2bce87d96e8d32 100644 --- a/dist/Repo.d.ts +++ b/dist/Repo.d.ts +@@ -36,7 +36,7 @@ export declare class Repo extends EventEmitter { + /** maps peer id to to persistence information (storageId, isEphemeral), access by collection synchronizer */ + /** @hidden */ + peerMetadataByPeerId: Record; +- constructor({ storage, network, peerId, sharePolicy, shareConfig, isEphemeral, enableRemoteHeadsGossiping, denylist, saveDebounceRate, idFactory, signer, subductionPolicy, subductionWebsocketEndpoints, subductionAdapters, subductionTimeouts, subductionCompactionDeleteConcurrency, subductionBlobInterceptor, }?: RepoConfig); ++ constructor({ storage, network, peerId, sharePolicy, shareConfig, isEphemeral, enableRemoteHeadsGossiping, denylist, saveDebounceRate, idFactory, signer, subductionPolicy, subductionWebsocketEndpoints, subductionAdapters, subductionTimeouts, subductionCompactionDeleteConcurrency, subductionBlobInterceptor, headsChannel, }?: RepoConfig); + /** Returns all the handles we have cached. */ + get handles(): Record>; + /** Returns a list of all connected peer ids */ @@ -259,8 +259,11 @@ export interface RepoConfig { adapter: NetworkAdapterInterface; serviceName: string; @@ -16,6 +38,49 @@ index 15d4ca65aa613fb567335f09c8fcb8cb9c4a954a..0f89be0e8b7e19ebb706d88685593259 }[]; /** * Tunable timeouts for the Subduction sync engine and its +@@ -284,6 +287,7 @@ export interface RepoConfig { + * `storage` when only one of the two Repos runs the interceptor. + */ + subductionBlobInterceptor?: BlobInterceptor; ++ headsChannel?: string; + } + /** A function that determines whether we should share a document with a peer + * +diff --git a/dist/Repo.js b/dist/Repo.js +index adfddd9b3926510604274217c42f97f104e119a3..151de25da17d3579a2befa796893898c61ed5f06 100644 +--- a/dist/Repo.js ++++ b/dist/Repo.js +@@ -57,7 +57,7 @@ export class Repo extends EventEmitter { + #subductionEphemeralCount = 0; + #idFactory; + #peerId; +- constructor({ storage, network = [], peerId = randomPeerId(), sharePolicy, shareConfig, isEphemeral = storage === undefined, enableRemoteHeadsGossiping = false, denylist = [], saveDebounceRate = 100, idFactory, signer, subductionPolicy, subductionWebsocketEndpoints, subductionAdapters, subductionTimeouts, subductionCompactionDeleteConcurrency, subductionBlobInterceptor, } = {}) { ++ constructor({ storage, network = [], peerId = randomPeerId(), sharePolicy, shareConfig, isEphemeral = storage === undefined, enableRemoteHeadsGossiping = false, denylist = [], saveDebounceRate = 100, idFactory, signer, subductionPolicy, subductionWebsocketEndpoints, subductionAdapters, subductionTimeouts, subductionCompactionDeleteConcurrency, subductionBlobInterceptor, headsChannel, } = {}) { + super(); + this.#peerId = peerId; + this.#remoteHeadsGossipingEnabled = enableRemoteHeadsGossiping; +@@ -119,6 +119,7 @@ export class Repo extends EventEmitter { + timeouts: subductionTimeouts, + compactionDeleteConcurrency: subductionCompactionDeleteConcurrency, + blobInterceptor: subductionBlobInterceptor, ++ headsChannel, + onRemoteHeadsChanged: (documentId, storageId, heads, timestamp) => { + const handle = this.#queries[documentId]?.handle; + if (handle) { +@@ -702,11 +703,10 @@ export class Repo extends EventEmitter { + * @param documentId - documentId of the DocHandle to remove from handleCache, if present in cache. + */ + async removeFromCache(documentId) { +- for (const source of this.#sources.values()) { +- source.detach(documentId); +- } ++ const detached = [...this.#sources.values()].map(source => source.detach(documentId)); + delete this.#queries[documentId]; + this.#syncStateTracker.delete(documentId); ++ await Promise.all(detached); + } + async shutdown() { + // Quiesce Subduction first — stops reconnect loops, flushes pending diff --git a/dist/subduction/AdapterConnections.d.ts b/dist/subduction/AdapterConnections.d.ts index 55231beca6494ed4ab0eaeb45383b822bcea5fe3..419c4ec3d17dc9f521292d780d293f5b1aafa8ce 100644 --- a/dist/subduction/AdapterConnections.d.ts @@ -64,11 +129,897 @@ index 4760a1899a8c0127db4840d5de090f6825211d97..5b9d19d4dd1cfb0cb242f140d5401af5 return true; } return false; +diff --git a/dist/subduction/source.d.ts b/dist/subduction/source.d.ts +index c1f722a51f7e273ee8161e4569b1602c2e3c3d11..799f9ebffae9d4bac393f62de760affcff062339 100644 +--- a/dist/subduction/source.d.ts ++++ b/dist/subduction/source.d.ts +@@ -119,16 +119,17 @@ export interface SubductionSourceOptions { + */ + compactionDeleteConcurrency?: number; + blobInterceptor?: BlobInterceptor; ++ headsChannel?: string; + } + export declare class SubductionSource implements DocumentSource { + #private; + readonly priority: SourcePriority; +- constructor({ peerId, storage, signer, websocketEndpoints, adapters, onRemoteHeadsChanged, onEphemeral, onHealExhausted, onConnectionChanged, priority, policy, timeouts, compactionDeleteConcurrency, blobInterceptor, }: SubductionSourceOptions); ++ constructor({ peerId, storage, signer, websocketEndpoints, adapters, onRemoteHeadsChanged, onEphemeral, onHealExhausted, onConnectionChanged, priority, policy, timeouts, compactionDeleteConcurrency, blobInterceptor, headsChannel, }: SubductionSourceOptions); + /** True if any connection manager has an established connection. */ + isConnected(): boolean; + getSubduction(): Promise; + attach(query: DocumentQuery): void; +- detach(_documentId: DocumentId): void; ++ detach(documentId: DocumentId): Promise; + shareConfigChanged(): void; + /** Check whether a sedimentree is currently in heal-backoff. */ + isHealing(sedimentreeId: SedimentreeId): boolean; +diff --git a/dist/subduction/source.js b/dist/subduction/source.js +index 4692619444edfae0ef5f111cef6ce29499507afa..7494e3df5893cb2292577008989c7e8a561164ad 100644 +--- a/dist/subduction/source.js ++++ b/dist/subduction/source.js +@@ -1,7 +1,7 @@ + import * as Automerge from "@automerge/automerge/slim"; + import { BlobMeta, CommitId, CommitInput, Fragment, FragmentInput, LooseCommit, Subduction, Topic, setSubductionLogLevel, } from "@automerge/automerge-subduction/slim"; + import { toSedimentreeId, toDocumentId } from "./helpers.js"; +-import { encodeHeads } from "../AutomergeUrl.js"; ++import { encodeHeads, decodeHeads } from "../AutomergeUrl.js"; + import { mergeArrays } from "../helpers/mergeArrays.js"; + import { semaphore } from "../helpers/semaphore.js"; + import { throttle } from "../helpers/throttle.js"; +@@ -27,6 +27,34 @@ const DEFAULT_SYNC_TIMEOUT_MS = 60_000; + * unresponsive or half-torn-down peer. + */ + const SHUTDOWN_SYNC_TIMEOUT_MS = 5_000; ++const MAX_DETACH_SAVE_ROUNDS = 8; ++const MAX_LOCAL_LOAD_RETRIES = 3; ++/** ++ * Cap on how many heal rounds the containment backstop may open for a ++ * single entry before some peer has been seen holding everything the ++ * handle holds. ++ * ++ * A handle can hold a commit this node is structurally unable to push: ++ * a sibling's write, read back out of shared storage, lands in ++ * `knownHashes`, and `#saveNewCommits` filters every future push against ++ * that same set. Containment is then unsatisfiable no matter how many ++ * rounds run, so the backstop has to be counted, not just conditioned. ++ * The counter resets when containment is observed, and on `resync()`. ++ */ ++const MAX_BACKSTOP_SYNCS = 3; ++/** ++ * How long an entry must look diverged before the backstop believes it. ++ * ++ * A sibling's write reaches this handle out of the shared database well ++ * before the tab that made it has finished its own round and said what ++ * the server now holds, so a moment of apparent divergence is the normal ++ * shape of a healthy origin. Waiting, and looking again, is what keeps ++ * the backstop from costing a round every time a sibling types. ++ * ++ * Shares the healer's `healInitialDelayMs` knob and default: both answer ++ * "how long before this is worth a round". ++ */ ++const DEFAULT_BACKSTOP_DELAY_MS = 2_000; + /** + * Deadline for draining in-flight storage-bridge operations at shutdown. + * An adapter promise that never settles (e.g. IndexedDB in a frozen or +@@ -44,6 +72,7 @@ const DEFAULT_COMPACTION_DELETE_CONCURRENCY = 16; + export class SubductionSource { + #subduction; + #storage; ++ #headsChannel = null; + #entries = new Map(); + #log; + #connectionManagers = []; +@@ -52,6 +81,7 @@ export class SubductionSource { + #scheduler; + #syncTimeoutMs; + #syncTimeout; ++ #backstopDelayMs; + #compactionDeleteConcurrency; + #blobInterceptor; + #onRemoteHeadsChanged; +@@ -59,8 +89,13 @@ export class SubductionSource { + /** Last reported aggregate connectedness, to fire onConnectionChanged only + * on transitions. */ + #lastConnected = false; +- /** Last surfaced remote heads per `sid|storageId`, to dedupe repeated +- * notifications (avoids redundant handle events and storage writes). */ ++ /** ++ * Last surfaced remote heads per sedimentree, keyed inside by peer. ++ * Dedupes repeated notifications (avoids redundant handle events and ++ * storage writes), carries the observation timestamp so a sibling ++ * relaying older news cannot walk this node's view backwards, and ++ * keeps the hashes that `#serverHasAllHeads` compares against. ++ */ + #lastRemoteHeads = new Map(); + /** + * Set at the top of `shutdown()`. Makes `#scheduleRecompute` a +@@ -70,12 +105,14 @@ export class SubductionSource { + */ + #shuttingDown = false; + priority; +- constructor({ peerId, storage, signer, websocketEndpoints, adapters, onRemoteHeadsChanged, onEphemeral, onHealExhausted, onConnectionChanged, priority = 2, policy, timeouts, compactionDeleteConcurrency = DEFAULT_COMPACTION_DELETE_CONCURRENCY, blobInterceptor, }) { ++ constructor({ peerId, storage, signer, websocketEndpoints, adapters, onRemoteHeadsChanged, onEphemeral, onHealExhausted, onConnectionChanged, priority = 2, policy, timeouts, compactionDeleteConcurrency = DEFAULT_COMPACTION_DELETE_CONCURRENCY, blobInterceptor, headsChannel, }) { + this.#blobInterceptor = blobInterceptor; + this.#onConnectionChanged = onConnectionChanged; + this.priority = priority; + this.#syncTimeoutMs = timeouts?.syncMs ?? DEFAULT_SYNC_TIMEOUT_MS; + this.#syncTimeout = this.#syncTimeoutMs; ++ this.#backstopDelayMs = ++ timeouts?.healInitialDelayMs ?? DEFAULT_BACKSTOP_DELAY_MS; + this.#compactionDeleteConcurrency = compactionDeleteConcurrency; + // Default roundtrip deadline forwarded to the Subduction constructor. + // The field must be *omitted* (not passed as `undefined`) to get the +@@ -113,7 +150,20 @@ export class SubductionSource { + // raw subduction CommitId hex would never compare equal to a doc's + // own heads). + const urlHeads = encodeHeads(heads.map(h => h.toHexString())); +- this.#surfaceRemoteHeads(sedimentreeId, storageId, urlHeads, Date.now(), true); ++ const timestamp = Date.now(); ++ // Only announce what actually surfaced: the wasm fires repeatedly ++ // with the same heads during a sync burst, and relaying each one ++ // would multiply that burst by the number of tabs on the origin. ++ if (!this.#surfaceRemoteHeads(sedimentreeId, storageId, urlHeads, timestamp, true)) { ++ return; ++ } ++ this.#headsChannel?.postMessage({ ++ kind: "remote", ++ sid: sedimentreeId.toString(), ++ storageId, ++ heads: urlHeads, ++ timestamp, ++ }); + }; + // Construct without hydrating: skip preloading persisted sedimentrees + // from storage at startup. State is loaded lazily on demand instead, +@@ -128,6 +178,27 @@ export class SubductionSource { + onEphemeral, + ...defaultTimeoutOption, + })); ++ if (headsChannel && typeof BroadcastChannel !== "undefined") { ++ this.#headsChannel = new BroadcastChannel(headsChannel); ++ this.#headsChannel.unref?.(); ++ this.#headsChannel.onmessage = ({ data, }) => { ++ const entry = this.#entries.get(data.sid); ++ if (!entry) ++ return; ++ if (data.kind === "remote") { ++ // Never re-announce: the observing tab already told everyone. ++ this.#surfaceRemoteHeads(entry.sedimentreeId, data.storageId, data.heads, data.timestamp, false); ++ return; ++ } ++ if (data.kind !== "saved") ++ return; ++ if (data.heads.every(h => entry.knownHashes.has(h))) ++ return; ++ entry.localLoadPending = true; ++ entry.siblingLoadPending = true; ++ this.#scheduleRecompute(entry); ++ }; ++ } + // ── Connection managers ───────────────────────────────────────── + // Endpoints own *where* their socket lives (in-thread, Worker, ...); + // the shared manager owns the reconnect/backoff loop. Plain URL +@@ -184,16 +255,28 @@ export class SubductionSource { + * when `persist` is set, durably store it so the last-known sync state + * survives reload and replays on the next attach. `heads` must already be + * in automerge-repo `UrlHeads` (bs58check) form. ++ * ++ * Returns whether this was news: `false` means the observation was a ++ * repeat, or older than what is already held for that peer. + */ + #surfaceRemoteHeads(sedimentreeId, storageId, heads, timestamp, persist) { +- // Dedupe: skip if this peer's heads for this sedimentree are unchanged +- // since we last surfaced them (the wasm can fire repeatedly with the +- // same heads during a sync burst). +- const dedupeKey = `${sedimentreeId.toString()}|${storageId}`; ++ const sidStr = sedimentreeId.toString(); + const joined = [...heads].sort().join(","); +- if (this.#lastRemoteHeads.get(dedupeKey) === joined) +- return; +- this.#lastRemoteHeads.set(dedupeKey, joined); ++ const byPeer = this.#lastRemoteHeads.get(sidStr); ++ const previous = byPeer?.get(storageId); ++ // Skip a repeat (the wasm fires repeatedly with the same heads during ++ // a sync burst), and skip stale news: heads carry no order of their ++ // own, so an older observation relayed by a sibling would otherwise ++ // replace a newer one and walk this node's view of the peer back. ++ if (previous !== undefined && ++ (previous.joined === joined || timestamp < previous.timestamp)) { ++ return false; ++ } ++ const surfaced = { joined, timestamp, hashes: decodeHeads(heads) }; ++ if (byPeer) ++ byPeer.set(storageId, surfaced); ++ else ++ this.#lastRemoteHeads.set(sidStr, new Map([[storageId, surfaced]])); + const documentId = toDocumentId(sedimentreeId); + this.#onRemoteHeadsChanged?.(documentId, storageId, heads, timestamp); + if (persist) { +@@ -201,6 +284,21 @@ export class SubductionSource { + .saveRemoteHeads(sedimentreeId, storageId, heads, timestamp) + .catch(e => this.#log.debug("saveRemoteHeads failed: %O", e)); + } ++ return true; ++ } ++ /** ++ * Close the heads channel and forget it. ++ * ++ * Forgetting is the point. `shutdown()` closes the channel after its ++ * flush-then-final-sync pass, but the 100 ms save throttle can still ++ * fire after that, and `postMessage` on a closed BroadcastChannel ++ * throws `InvalidStateError`. `#save` is called as `void this.#save(…)` ++ * with no catch, so the throw would surface as an unhandled rejection ++ * and skip the rest of that save. ++ */ ++ #closeHeadsChannel() { ++ this.#headsChannel?.close(); ++ this.#headsChannel = null; + } + #anyConnectionManagerConnecting() { + return this.#connectionManagers.some(mgr => mgr.isConnecting()); +@@ -244,6 +342,7 @@ export class SubductionSource { + // — recomputes are targeted, so other entries' completions no + // longer walk this one. Mark it dirty so the walk retries the load. + if (entry.syncState === "initializing") { ++ entry.localLoadPending = true; + this.#scheduleRecompute(entry); + return; + } +@@ -364,6 +463,7 @@ export class SubductionSource { + return; + void this.#save(entry); + }, 100); ++ const onHeadsChanged = () => throttledSave(); + this.#entries.set(sidStr, { + syncState: "initializing", + query, +@@ -372,17 +472,23 @@ export class SubductionSource { + syncInFlight: false, + lastSyncResult: null, + lastSyncGeneration: -1, ++ backstopSyncs: 0, ++ backstopTimer: null, + blobLoadInFlight: false, + blobRetries: 0, + needsResync: false, ++ localLoadPending: true, ++ siblingLoadPending: false, + syncSettled: Promise.resolve(), + lastSavedHeads: new Set(), + knownHashes: new Set(), ++ strandedHashes: new Set(), + persistedCommitHashes: new Set(), + untransformedHashes: new Set(), + persistedFragmentHashes: new Set(), + compactionInFlight: null, + flushSave: throttledSave, ++ onHeadsChanged, + saveSettled, + saveInProgress: false, + saveDeltaPending: false, +@@ -394,7 +500,7 @@ export class SubductionSource { + transformRetryScheduled: false, + }); + query.sourcePending("subduction"); +- query.handle.on("heads-changed", () => throttledSave()); ++ query.handle.on("heads-changed", onHeadsChanged); + throttledSave(); + // Subscribe to ephemeral messages for this sedimentree + void (async () => { +@@ -441,6 +547,7 @@ export class SubductionSource { + this.#log.debug(`failed to seed persistedHashes for ${sidStr.slice(0, 8)}: %O`, e); + } + })(); ++ this.#lastRemoteHeads.delete(sidStr); + // Replay the last-known remote heads for this sedimentree so the + // DocHandle reflects the last sync state immediately on (re)attach, + // before any network round-trip completes. Does not re-persist. +@@ -457,7 +564,49 @@ export class SubductionSource { + })(); + this.#scheduleRecompute(this.#entries.get(sidStr)); + } +- detach(_documentId) { } ++ async detach(documentId) { ++ const sid = toSedimentreeId(documentId); ++ const sidStr = sid.toString(); ++ const entry = this.#entries.get(sidStr); ++ if (!entry) ++ return; ++ this.#entries.delete(sidStr); ++ if (entry.backstopTimer !== null) ++ clearTimeout(entry.backstopTimer); ++ entry.handle.off("heads-changed", entry.onHeadsChanged); ++ await entry.saveSettled; ++ for (let round = 0; round < MAX_DETACH_SAVE_ROUNDS; round++) { ++ await this.#save(entry); ++ if (!entry.saveDeltaPending) ++ break; ++ } ++ if (entry.lastSaveError !== null) { ++ this.#log.warn(`detach ${sidStr.slice(0, 8)}: unsaved commits could not be persisted`, entry.lastSaveError); ++ } ++ await entry.syncSettled; ++ if (entry.needsResync || this.#needsOutboundSync(entry)) { ++ await this.#doSync(entry, SHUTDOWN_SYNC_TIMEOUT_MS); ++ } ++ let subduction; ++ try { ++ subduction = await this.#subduction; ++ } ++ catch (e) { ++ this.#log.debug("subduction never initialized, skipping unsubscribe: %O", e); ++ subduction = null; ++ } ++ if (this.#entries.has(sidStr)) ++ return; ++ this.#scheduler.resetHealState(sidStr); ++ if (!subduction) ++ return; ++ try { ++ await subduction.unsubscribeEphemeral([Topic.fromBytes(sid.toBytes())]); ++ } ++ catch (e) { ++ this.#log.debug("ephemeral unsubscribe failed: %O", e); ++ } ++ } + shareConfigChanged() { + for (const entry of this.#entries.values()) { + // A share-config change can flip a previously-denied fetch to +@@ -598,6 +747,208 @@ export class SubductionSource { + entry.lastSyncResult === "no-peers") && + entry.lastSyncGeneration !== this.#connectionGeneration()); + } ++ /** ++ * True when some peer has been seen holding every head this entry's ++ * handle holds — there is nothing left for this node to get across. ++ * Vacuously true for a handle with nothing in it. ++ */ ++ #serverHasAllHeads(entry) { ++ const doc = entry.handle.doc(); ++ if (!doc) ++ return true; ++ const heads = Automerge.getHeads(doc); ++ if (heads.length === 0) ++ return true; ++ const byPeer = this.#lastRemoteHeads.get(entry.sedimentreeId.toString()); ++ if (byPeer === undefined) ++ return false; ++ for (const { hashes } of byPeer.values()) { ++ if (heads.every(h => hashes.includes(h))) ++ return true; ++ } ++ return false; ++ } ++ /** ++ * Backstop for the sync predicate. `#needsOutboundSync` knows only ++ * whether this entry synced since the connection generation moved; it ++ * cannot tell "synced and caught up" from "synced, and the handle has ++ * since advanced by a path that stored nothing" — a sibling's write ++ * read back out of shared storage, say. That entry would otherwise sit ++ * diverged with nothing scheduled to notice. ++ * ++ * Counted, not merely conditioned. See {@link MAX_BACKSTOP_SYNCS}: the ++ * divergence can be permanent, and a bare "sync while not contained" ++ * predicate would then round-trip for the life of the tab. Routed ++ * through the healer so the rounds it does open inherit the healer's ++ * backoff and attempt cap, and left to the healer entirely once a real ++ * failure is being retried. ++ */ ++ #maybeBackstopSync(entry) { ++ if (this.#serverHasAllHeads(entry)) { ++ entry.backstopSyncs = 0; ++ // A peer holding every head holds every ancestor, so nothing the ++ // handle carries is stranded any more. ++ entry.strandedHashes.clear(); ++ return; ++ } ++ if (entry.backstopTimer !== null || this.#backstopBlocked(entry)) ++ return; ++ entry.backstopTimer = setTimeout(() => this.#backstopExpired(entry), this.#backstopDelayMs); ++ } ++ #backstopBlocked(entry) { ++ if (entry.lastSyncResult === "all-failed") ++ return true; ++ // Disconnected there is no round to open, but a stranded commit has ++ // to reach the resident tree *before* the reconnect round reads it, ++ // or that round pushes around it. So the timer still arms offline — ++ // for the re-ingest alone, and only while something is stranded. ++ if (!this.isConnected()) ++ return entry.strandedHashes.size === 0; ++ // The cap counts rounds, not writes: a stranded commit still needs ++ // its one duplicate write after the round budget is spent. ++ if (entry.backstopSyncs >= MAX_BACKSTOP_SYNCS) ++ return entry.strandedHashes.size === 0; ++ return false; ++ } ++ /** Look again once {@link DEFAULT_BACKSTOP_DELAY_MS} is up. */ ++ #backstopExpired(entry) { ++ entry.backstopTimer = null; ++ if (this.#shuttingDown) ++ return; ++ if (this.#serverHasAllHeads(entry)) { ++ entry.backstopSyncs = 0; ++ entry.strandedHashes.clear(); ++ return; ++ } ++ if (this.#backstopBlocked(entry)) ++ return; ++ // Settled, and still diverged: this is the moment the entry has ++ // earned a duplicate write. Offline that is the whole of it — there ++ // is no round to open, and the next reconnect round will find the ++ // tree already holding the sibling's commits. ++ const reingested = this.#reingestStranded(entry); ++ if (!this.isConnected()) ++ return; ++ // The write above is owed to the commit; the round is not. A spent ++ // budget, or a heal chain already retrying this entry, gates the ++ // round alone. ++ if (entry.backstopSyncs >= MAX_BACKSTOP_SYNCS || ++ this.#scheduler.isHealing(entry.sedimentreeId)) { ++ return; ++ } ++ entry.backstopSyncs++; ++ this.#log.debug(`backstop sync ${entry.sedimentreeId.toString().slice(0, 8)} ` + ++ `(${entry.backstopSyncs}/${MAX_BACKSTOP_SYNCS})`); ++ const heal = () => { ++ // The re-ingest can outlive the entry: `detach()` resets this ++ // sedimentree's heal state, and a round for an evicted document is ++ // noise the protocol handler worker pays for in bursts. ++ if (this.#shuttingDown) ++ return; ++ if (this.#entries.get(entry.sedimentreeId.toString()) !== entry) ++ return; ++ this.#scheduler.scheduleHealSync(entry.sedimentreeId); ++ }; ++ // The round has to read a tree that already holds the re-ingested ++ // commits, or it pushes exactly what the last one did. ++ if (reingested) ++ void reingested.then(heal); ++ else ++ heal(); ++ } ++ /** ++ * Hand the next save back the commits this node holds only because it ++ * read a sibling's write out of the shared database. ++ * ++ * Dropping them from `knownHashes` un-filters them in ++ * `#saveNewCommits`, and clearing the save baseline gets them past ++ * `#save`'s heads fast path, which the channel-driven reload already ++ * advanced past them. The save writes them a second time — a ++ * duplicate of a record already on disk — and that write is what puts ++ * them in the resident tree, where a round can finally read them. ++ * ++ * At most one duplicate write per commit, ever: a reload only marks a ++ * hash `knownHashes` did not already hold, and `#saveNewCommits` ++ * unmarks every hash it stores. A server that keeps rejecting the ++ * commit therefore costs re-tries of the round (capped by ++ * {@link MAX_BACKSTOP_SYNCS}), not re-tries of the write. ++ * ++ * Returns null when there was nothing stranded, else a promise for the ++ * save that stores it. ++ */ ++ #reingestStranded(entry) { ++ this.#forgetStrandedOnTheServer(entry); ++ if (entry.strandedHashes.size === 0) ++ return null; ++ const batch = [...entry.strandedHashes]; ++ this.#log.debug(`re-ingesting ${batch.length} stranded commit(s) for ` + ++ entry.sedimentreeId.toString().slice(0, 8)); ++ for (const hash of batch) ++ entry.knownHashes.delete(hash); ++ return (async () => { ++ // A save that started before the drop above filtered on the old ++ // `knownHashes` and would leave the commits stranded; let it go ++ // before arming the one that sees them. It restores the save ++ // baseline as it finishes, so the baseline is cleared after it: ++ // cleared first, the finishing save puts it straight back and the ++ // save below returns on `#save`'s heads fast path. ++ await entry.saveSettled; ++ entry.lastSavedHeads.clear(); ++ entry.flushSave(); ++ entry.flushSave.flush(); ++ await entry.saveSettled; ++ // `#saveNewCommits` has already unmarked what it stored. What is ++ // left of this batch is what it could not store — absorbed into a ++ // fragment since the reload, say — and holding on to that would ++ // arm the backstop for ever. A save that failed outright keeps its ++ // batch, so the next expiry retries it. ++ if (entry.lastSaveError === null) { ++ for (const hash of batch) ++ entry.strandedHashes.delete(hash); ++ } ++ })(); ++ } ++ /** ++ * Forget the stranded commits some peer's heads already descend from. ++ * ++ * `#serverHasAllHeads` asks whether the server holds everything this ++ * handle holds, which is false for the ordinary moment between a local ++ * edit and its confirmation. That is the wrong question here: what the ++ * re-ingest is for is a commit nobody else will push, and a commit the ++ * server's heads descend from has already been pushed by the tab that ++ * wrote it. Asking the narrower question is what keeps a busy origin ++ * from paying a duplicate write per sibling edit. ++ * ++ * Only runs with something stranded, which in normal operation is ++ * nothing at all. ++ */ ++ #forgetStrandedOnTheServer(entry) { ++ if (entry.strandedHashes.size === 0) ++ return; ++ const byPeer = this.#lastRemoteHeads.get(entry.sedimentreeId.toString()); ++ if (byPeer === undefined) ++ return; ++ const doc = entry.handle.doc(); ++ if (!doc) ++ return; ++ for (const { hashes } of byPeer.values()) { ++ let unpushed; ++ try { ++ unpushed = new Set(Automerge.getChangesMetaSince(doc, hashes).map(m => m.hash)); ++ } ++ catch { ++ // Heads this doc does not have; that peer says nothing about ++ // what it holds of ours. ++ continue; ++ } ++ for (const hash of entry.strandedHashes) { ++ if (!unpushed.has(hash)) ++ entry.strandedHashes.delete(hash); ++ } ++ if (entry.strandedHashes.size === 0) ++ return; ++ } ++ } + #recomputeEntry(entry) { + // No new background syncs once shutting down: `shutdown()`'s quiesce + // round is the sole sync initiator from that point. A recompute walk +@@ -625,6 +976,10 @@ export class SubductionSource { + entry.syncInFlight = true; + void this.#doSync(entry); + } ++ else if (!this.isConnected()) { ++ entry.blobLoadInFlight = true; ++ void this.#loadBlobsAndTransition(entry); ++ } + else if (Automerge.getHeads(entry.handle.fullDoc()).length === 0 && + !this.#anyConnectionManagerConnecting()) { + // All connections settled and sync failed — give up. +@@ -632,12 +987,31 @@ export class SubductionSource { + entry.query.sourceUnavailable("subduction"); + } + } ++ if (entry.localLoadPending && !entry.blobLoadInFlight) { ++ entry.localLoadPending = false; ++ const fromSibling = entry.siblingLoadPending; ++ entry.siblingLoadPending = false; ++ entry.blobLoadInFlight = true; ++ void this.#loadBlobsAndTransition(entry, true, fromSibling); ++ } + return; + } + case "running": { +- if (!entry.syncInFlight && this.#needsOutboundSync(entry)) { +- entry.syncInFlight = true; +- void this.#doSync(entry); ++ if (entry.localLoadPending && !entry.blobLoadInFlight) { ++ entry.localLoadPending = false; ++ const fromSibling = entry.siblingLoadPending; ++ entry.siblingLoadPending = false; ++ entry.blobLoadInFlight = true; ++ void this.#loadBlobsAndTransition(entry, true, fromSibling); ++ } ++ if (!entry.syncInFlight) { ++ if (this.#needsOutboundSync(entry)) { ++ entry.syncInFlight = true; ++ void this.#doSync(entry); ++ } ++ else { ++ this.#maybeBackstopSync(entry); ++ } + } + return; + } +@@ -663,6 +1037,8 @@ export class SubductionSource { + await entry.saveSettled; + this.#log.debug(`doSync ${sid} (state=${entry.syncState})`); + const peerResultMap = await subduction.syncWithAllPeers(sedimentreeId, true, timeoutMs !== undefined ? timeoutMs : this.#syncTimeout); ++ if (this.#entries.get(sedimentreeId.toString()) !== entry) ++ return; + const results = peerResultMap.entries(); + const anySuccess = results.some(r => r.success); + this.#log.debug(`doSync ${sid}: ${results.length} peer(s), success=${anySuccess}`); +@@ -676,7 +1052,7 @@ export class SubductionSource { + // arrives, without waiting for further state transitions. + // + // During "initializing", `#loadBlobsAndTransition` performs one +- // atomic getBlobs + loadIncremental after sync succeeds; loading ++ // atomic bulk read + loadIncremental after sync succeeds; loading + // here too would duplicate storage reads on first open. + // + // Also retry when the handle is empty: a previous load may have +@@ -754,42 +1130,63 @@ export class SubductionSource { + return out; + } + /** +- * Load all blobs for a sedimentree from Subduction and apply them to the ++ * Load all blobs for a sedimentree from storage and apply them to the + * handle via `Automerge.loadIncremental`. If new data was loaded, signal + * the query so it can transition to "ready". + * + * This is called reactively after any `syncWithAllPeers` that received + * data, making handle updates immediate. + */ +- async #loadBlobsIntoHandle(entry, subduction) { ++ async #loadBlobsIntoHandle(entry, subduction, fromSibling = false) { + const sid = entry.sedimentreeId.toString().slice(0, 8); ++ // An announcement still waiting for its own reload is answered by ++ // this read too — it reads the same shared database. Honour it ++ // whoever asked for the read, so a peer-fed load landing in that ++ // window cannot file the sibling's commit as known and not stranded. ++ // Left set for `#recomputeEntry` to consume as usual. Only once ++ // running: while initializing the resident tree is hydrated from ++ // this same storage, so nothing read here is stranded. ++ if (entry.syncState === "running" && entry.siblingLoadPending) { ++ fromSibling = true; ++ } + const headCount = () => Automerge.getHeads(entry.handle.fullDoc()).length; +- const allBlobs = await subduction.getBlobs(entry.sedimentreeId); +- const totalBytes = allBlobs +- ? allBlobs.reduce((n, b) => n + b.byteLength, 0) +- : 0; +- this.#log.debug(`loadBlobsIntoHandle ${sid}: ${allBlobs?.length ?? 0} blob(s), ${totalBytes} bytes, heads=${headCount()}`); +- if (!allBlobs || allBlobs.length === 0) +- return false; + // Prefer the id-keyed view of the same blobs. Fall back to the id-less one +- // if it does not cover everything subduction reports. +- let blobsWithIds = null; ++ // if reading it throws. ++ let blobsWithIds; + try { +- const keyed = await this.#idKeyedBlobs(entry); +- if (keyed.length >= allBlobs.length) +- blobsWithIds = keyed; ++ blobsWithIds = await this.#idKeyedBlobs(entry); + } + catch (e) { + this.#log.debug("idKeyedBlobs failed, falling back to id-less load: %O", e); ++ const allBlobs = await subduction.getBlobs(entry.sedimentreeId); ++ blobsWithIds = (allBlobs ?? []).map(blob => ({ idHex: "", blob })); + } +- if (!blobsWithIds) +- blobsWithIds = allBlobs.map(blob => ({ idHex: "", blob })); ++ const totalBytes = blobsWithIds.reduce((n, u) => n + u.blob.byteLength, 0); ++ this.#log.debug(`loadBlobsIntoHandle ${sid}: ${blobsWithIds.length} blob(s), ${totalBytes} bytes, heads=${headCount()}`); ++ if (blobsWithIds.length === 0) ++ return false; + blobsWithIds.sort((a, b) => b.blob.byteLength - a.blob.byteLength); + // Record these IDs as known hashes since nothing in subduction's storage +- // should be re-encrypted and re-saved. ++ // should be re-encrypted and re-saved. A sibling's write read back out ++ // of the shared database is the exception: `knownHashes` then hides a ++ // commit the resident tree has never seen, so mark it stranded too. ++ let newlyStranded = 0; + for (const u of blobsWithIds) { +- if (u.idHex) +- entry.knownHashes.add(u.idHex); ++ if (!u.idHex) ++ continue; ++ if (fromSibling && !entry.knownHashes.has(u.idHex)) { ++ entry.strandedHashes.add(u.idHex); ++ newlyStranded++; ++ } ++ entry.knownHashes.add(u.idHex); ++ } ++ // A commit that has only just arrived is no evidence of divergence: ++ // the tab that wrote it is still mid-push. Start the settling delay ++ // again from here — the recompute in `#loadBlobsAndTransition` re-arms ++ // it — so a stream of sibling edits never expires the backstop. ++ if (newlyStranded > 0 && entry.backstopTimer !== null) { ++ clearTimeout(entry.backstopTimer); ++ entry.backstopTimer = null; + } + let toApply = blobsWithIds.map(u => u.blob); + if (this.#blobInterceptor) { +@@ -835,10 +1232,22 @@ export class SubductionSource { + this.#scheduleCompaction(entry); + return true; + } +- async #loadBlobsAndTransition(entry) { ++ /** ++ * `fromSibling` is consumed by the caller, not read off the entry ++ * here: an announcement that lands while this load is in flight sets ++ * the flag again, and clearing it afterwards would erase it. The ++ * follow-up reload would then file the sibling's commit as known and ++ * not stranded — the one commit the re-ingest exists for. ++ */ ++ async #loadBlobsAndTransition(entry, onlyIfData = false, fromSibling = false) { + try { + const subduction = await this.#subduction; +- const hadBlobs = await this.#loadBlobsIntoHandle(entry, subduction); ++ const hadBlobs = await this.#loadBlobsIntoHandle(entry, subduction, fromSibling); ++ entry.blobRetries = 0; ++ if (onlyIfData && ++ Automerge.getHeads(entry.handle.fullDoc()).length === 0) { ++ return; ++ } + entry.syncState = "running"; + if (Automerge.getHeads(entry.handle.fullDoc()).length === 0) { + if (this.#blobInterceptor && hadBlobs) { +@@ -864,6 +1273,21 @@ export class SubductionSource { + } + } + catch (e) { ++ if (onlyIfData) { ++ if (entry.blobRetries < MAX_LOCAL_LOAD_RETRIES) { ++ entry.blobRetries++; ++ entry.localLoadPending = true; ++ if (fromSibling) ++ entry.siblingLoadPending = true; ++ } ++ else { ++ entry.blobRetries = 0; ++ this.#log.warn(`reload from storage failed ${MAX_LOCAL_LOAD_RETRIES + 1} times for ${entry.sedimentreeId ++ .toString() ++ .slice(0, 8)}; waiting for the next announcement`, e); ++ } ++ return; ++ } + this.#log.debug(`loadBlobsAndTransition threw for ${entry.sedimentreeId + .toString() + .slice(0, 8)}: %O`, e); +@@ -948,6 +1372,7 @@ export class SubductionSource { + entry.saveSettled = new Promise(r => { + resolveSaveSettled = r; + }); ++ let stored = 0; + try { + const doc = entry.handle.fullDoc(); + if (!doc) +@@ -989,6 +1414,14 @@ export class SubductionSource { + // so future mutations to either side stay isolated. + entry.lastSaveError = null; + entry.lastSavedHeads = new Set(currentSet); ++ stored = result.stored; ++ if (stored > 0) { ++ this.#headsChannel?.postMessage({ ++ kind: "saved", ++ sid: entry.sedimentreeId.toString(), ++ heads: currentHeads, ++ }); ++ } + // Detect a post-save delta the `heads-changed` listener didn't + // fire on. `getChangeByHash` calls inside `#saveNewCommits` can + // shift what `getHeads` returns for the same doc reference on +@@ -1014,6 +1447,8 @@ export class SubductionSource { + entry.saveInProgress = false; + resolveSaveSettled(); + } ++ if (stored === 0) ++ return; + // Trigger an immediate sync — the only broadcast path for newly + // saved commits (`#saveNewCommits` is store-only). If a sync is + // already in flight, flag a re-sync for when it completes; otherwise +@@ -1062,7 +1497,7 @@ export class SubductionSource { + const newCommitMetas = commitMetas.filter(m => !entry.knownHashes.has(m.head)); + const newFragmentMetas = fragmentMetas.filter(m => !entry.knownHashes.has(m.head)); + if (newCommitMetas.length === 0 && newFragmentMetas.length === 0) { +- return { prepFailures: 0 }; ++ return { prepFailures: 0, stored: 0 }; + } + const newCommitBytes = newCommitMetas.length === 0 + ? [] +@@ -1143,7 +1578,7 @@ export class SubductionSource { + : new AggregateError(prepErrors, `${prepErrors.length} inputs failed to prepare`); + } + if (commitInputs.length === 0 && fragmentInputs.length === 0) { +- return { prepFailures: prepErrors.length }; ++ return { prepFailures: prepErrors.length, stored: 0 }; + } + // Record hashes BEFORE `storeBuiltBatch` so the synchronous + // `commit-saved` / `fragment-saved` events fired from the storage +@@ -1171,8 +1606,15 @@ export class SubductionSource { + `${fragmentInputs.length} fragments):`, e); + throw e; + } ++ // Stored for real now, so not stranded any more — including anything ++ // a reload marked while this batch was in flight. ++ for (const hash of acceptedHashes) ++ entry.strandedHashes.delete(hash); + this.#scheduleCompaction(entry); +- return { prepFailures: prepErrors.length }; ++ return { ++ prepFailures: prepErrors.length, ++ stored: commitInputs.length + fragmentInputs.length, ++ }; + } + /** + * Kick off a compaction pass without blocking the save loop. +@@ -1321,6 +1763,7 @@ export class SubductionSource { + return; + this.#scheduler.resetHealState(sid); + entry.lastSyncResult = null; ++ entry.backstopSyncs = 0; + this.#scheduleRecompute(entry); + } + // ── Flush ─────────────────────────────────────────────────────────── +@@ -1430,6 +1873,11 @@ export class SubductionSource { + mgr.shutdown(); + } + this.#scheduler.shutdown(); ++ for (const entry of this.#entries.values()) { ++ if (entry.backstopTimer !== null) ++ clearTimeout(entry.backstopTimer); ++ entry.backstopTimer = null; ++ } + for (const entry of this.#entries.values()) + this.#flushInbound(entry); + for (const entry of this.#entries.values()) { +@@ -1468,6 +1916,7 @@ export class SubductionSource { + for (const endpoint of this.#websocketEndpoints) { + endpoint.shutdown?.(); + } ++ this.#closeHeadsChannel(); + this.#storage.close(); + await this.#awaitStorageIdle(); + return; +@@ -1481,6 +1930,7 @@ export class SubductionSource { + for (const endpoint of this.#websocketEndpoints) { + endpoint.shutdown?.(); + } ++ this.#closeHeadsChannel(); + this.#storage.close(); + await this.#awaitStorageIdle(); + } +diff --git a/src/DocumentSource.ts b/src/DocumentSource.ts +index 363cbbad5af3f2db28ed5fa0ca85ca584a60bb49..d9eb3c5d8429607480867bf105d0e772ba1aecfd 100644 +--- a/src/DocumentSource.ts ++++ b/src/DocumentSource.ts +@@ -37,7 +37,7 @@ export interface DocumentSource { + attach(query: DocumentQuery): void + + /** Called when a document is removed from the repo. */ +- detach(documentId: DocumentId): void ++ detach(documentId: DocumentId): void | Promise + + shareConfigChanged(): void + diff --git a/src/Repo.ts b/src/Repo.ts -index 0b408d6bda6582078083949c2fea63cfd0ef5e1b..c3df442b29c5a6ecd59a9464b3a30c4f59e3adc4 100644 +index 0b408d6bda6582078083949c2fea63cfd0ef5e1b..a26a259b4d1c07951bdb75a86cb8b6c297c2e079 100644 --- a/src/Repo.ts +++ b/src/Repo.ts -@@ -1060,8 +1060,11 @@ export interface RepoConfig { +@@ -136,6 +136,7 @@ export class Repo extends EventEmitter { + subductionTimeouts, + subductionCompactionDeleteConcurrency, + subductionBlobInterceptor, ++ headsChannel, + }: RepoConfig = {}) { + super() + this.#peerId = peerId +@@ -213,6 +214,7 @@ export class Repo extends EventEmitter { + timeouts: subductionTimeouts, + compactionDeleteConcurrency: subductionCompactionDeleteConcurrency, + blobInterceptor: subductionBlobInterceptor, ++ headsChannel, + onRemoteHeadsChanged: (documentId, storageId, heads, timestamp) => { + const handle = this.#queries[documentId]?.handle + if (handle) { +@@ -930,11 +932,12 @@ export class Repo extends EventEmitter { + * @param documentId - documentId of the DocHandle to remove from handleCache, if present in cache. + */ + async removeFromCache(documentId: DocumentId): Promise { +- for (const source of this.#sources.values()) { ++ const detached = [...this.#sources.values()].map(source => + source.detach(documentId) +- } ++ ) + delete this.#queries[documentId] + this.#syncStateTracker.delete(documentId) ++ await Promise.all(detached) + } + + async shutdown() { +@@ -1060,8 +1063,11 @@ export interface RepoConfig { adapter: NetworkAdapterInterface serviceName: string /** Whether to initiate ("connect") or accept ("accept") the subduction @@ -82,6 +1033,15 @@ index 0b408d6bda6582078083949c2fea63cfd0ef5e1b..c3df442b29c5a6ecd59a9464b3a30c4f }[] /** +@@ -1088,6 +1094,8 @@ export interface RepoConfig { + * `storage` when only one of the two Repos runs the interceptor. + */ + subductionBlobInterceptor?: BlobInterceptor ++ ++ headsChannel?: string + } + + /** A function that determines whether we should share a document with a peer diff --git a/src/subduction/AdapterConnections.ts b/src/subduction/AdapterConnections.ts index 4426f11291a00b230332c63d77c683bb7ab02d24..e1fb64697b94637cd9dd968a6cebe8f5e06d89b2 100644 --- a/src/subduction/AdapterConnections.ts @@ -135,3 +1095,986 @@ index 0d5510500c1e3a8e0d84b6f9b5f87f42096ee8e3..6b8376be24f932f21ab3ddc4b76bc81f } return false } +diff --git a/src/subduction/source.ts b/src/subduction/source.ts +index 4bd5d5cf682f3792f403efa44914ca88fa8f4006..cc507ab8adfa56f3043beee966157251a7c56ad0 100644 +--- a/src/subduction/source.ts ++++ b/src/subduction/source.ts +@@ -16,7 +16,7 @@ import { DocumentSource } from "../DocumentSource.js" + import { DocumentQuery, SourcePriority } from "../DocumentQuery.js" + import { DocumentId, PeerId } from "../types.js" + import { toSedimentreeId, toDocumentId } from "./helpers.js" +-import { encodeHeads } from "../AutomergeUrl.js" ++import { encodeHeads, decodeHeads } from "../AutomergeUrl.js" + import { DocHandle, NetworkAdapterInterface } from "../index.js" + import { ConnectionManager } from "./ConnectionManager.js" + import type { StorageId } from "../storage/types.js" +@@ -53,6 +53,36 @@ const DEFAULT_SYNC_TIMEOUT_MS = 60_000 + * unresponsive or half-torn-down peer. + */ + const SHUTDOWN_SYNC_TIMEOUT_MS = 5_000 ++const MAX_DETACH_SAVE_ROUNDS = 8 ++const MAX_LOCAL_LOAD_RETRIES = 3 ++ ++/** ++ * Cap on how many heal rounds the containment backstop may open for a ++ * single entry before some peer has been seen holding everything the ++ * handle holds. ++ * ++ * A handle can hold a commit this node is structurally unable to push: ++ * a sibling's write, read back out of shared storage, lands in ++ * `knownHashes`, and `#saveNewCommits` filters every future push against ++ * that same set. Containment is then unsatisfiable no matter how many ++ * rounds run, so the backstop has to be counted, not just conditioned. ++ * The counter resets when containment is observed, and on `resync()`. ++ */ ++const MAX_BACKSTOP_SYNCS = 3 ++ ++/** ++ * How long an entry must look diverged before the backstop believes it. ++ * ++ * A sibling's write reaches this handle out of the shared database well ++ * before the tab that made it has finished its own round and said what ++ * the server now holds, so a moment of apparent divergence is the normal ++ * shape of a healthy origin. Waiting, and looking again, is what keeps ++ * the backstop from costing a round every time a sibling types. ++ * ++ * Shares the healer's `healInitialDelayMs` knob and default: both answer ++ * "how long before this is worth a round". ++ */ ++const DEFAULT_BACKSTOP_DELAY_MS = 2_000 + + /** + * Deadline for draining in-flight storage-bridge operations at shutdown. +@@ -93,8 +123,28 @@ interface SedimentreeEntry { + syncInFlight: boolean + lastSyncResult: SyncResult | null + lastSyncGeneration: number ++ /** ++ * Heal rounds the containment backstop has opened since the last time ++ * a peer was seen holding every head this handle holds. Capped at ++ * {@link MAX_BACKSTOP_SYNCS}. ++ */ ++ backstopSyncs: number ++ /** Pending {@link MAX_BACKSTOP_SYNCS} re-check, or null. */ ++ backstopTimer: ReturnType | null + blobLoadInFlight: boolean + needsResync: boolean ++ localLoadPending: boolean ++ /** ++ * True while a sibling's `saved` announcement is waiting for its ++ * reload. A read of the shared database made while it is set can leave ++ * the handle holding a commit the resident subduction tree has never ++ * seen, so such a read records what it read into ++ * {@link SedimentreeEntry.strandedHashes}. Consumed by ++ * `#recomputeEntry`, which hands it to the reload it starts — clearing ++ * it after the load instead would erase an announcement that arrived ++ * during it. ++ */ ++ siblingLoadPending: boolean + blobRetries: number + /** + * Resolves when the in-flight `#doSync` round (if any) completes. +@@ -123,6 +173,20 @@ interface SedimentreeEntry { + * bounded by automerge's compaction policy, not by total history. + */ + knownHashes: Set ++ /** ++ * The part of `knownHashes` this node has never stored itself: a ++ * sibling's commits, read back out of the shared database by a ++ * channel-driven reload. ++ * ++ * They are in the handle and in `knownHashes`, so `#saveNewCommits` ++ * filters them out of every push — and they are not in this node's ++ * resident subduction tree, so no round can carry them either. If the ++ * tab that wrote them closes before pushing them, nothing else will. ++ * `#reingestStranded` drops them back out of `knownHashes` so the next ++ * save stores them for real; a hash leaves the set for good once some ++ * `storeBuiltBatch` has accepted it. ++ */ ++ strandedHashes: Set + /** + * Hashes of commits we currently believe are persisted to local + * storage as loose-commit records (not as part of a fragment). +@@ -160,6 +224,7 @@ interface SedimentreeEntry { + */ + compactionInFlight: Promise | null + flushSave: ThrottledFunction<() => void> ++ onHeadsChanged: () => void + /** Resolves when any in-progress `#save` completes. */ + saveSettled: Promise + /** +@@ -226,6 +291,37 @@ interface SedimentreeEntry { + inboundFlushScheduled: boolean + } + ++/** ++ * What one tab tells its siblings over the heads channel. ++ * ++ * `saved` says a local write reached the shared database, so a sibling ++ * with the document open should re-read it. `remote` relays an ++ * observation of what a sync server holds: every tab on the origin talks ++ * to the same server as the same identity, so one tab's observation is a ++ * fact for all of them. Only a real wasm `onRemoteHeads` callback ++ * announces `remote`; relaying a relay is what would turn the origin ++ * into an echo chamber. ++ */ ++type HeadsChannelMessage = ++ | { kind: "saved"; sid: string; heads: string[] } ++ | { ++ kind: "remote" ++ sid: string ++ storageId: string ++ heads: string[] ++ timestamp: number ++ } ++ ++/** The last remote heads surfaced for one sedimentree and one peer. */ ++interface SurfacedRemoteHeads { ++ /** Sorted and joined `UrlHeads`, for equality alone. */ ++ joined: string ++ /** When the observation was made, so older news can be dropped. */ ++ timestamp: number ++ /** The same heads as raw hashes, to compare with `Automerge.getHeads`. */ ++ hashes: string[] ++} ++ + /** Callback for remote heads changes from subduction peers. */ + export type OnRemoteHeadsChanged = ( + documentId: DocumentId, +@@ -370,11 +466,13 @@ export interface SubductionSourceOptions { + compactionDeleteConcurrency?: number + + blobInterceptor?: BlobInterceptor ++ headsChannel?: string + } + + export class SubductionSource implements DocumentSource { + #subduction: Promise + #storage: SubductionStorageBridge ++ #headsChannel: (BroadcastChannel & { unref?(): void }) | null = null + #entries = new Map() + #log: Logger + #connectionManagers: ConnectionManager[] = [] +@@ -383,6 +481,7 @@ export class SubductionSource implements DocumentSource { + #scheduler: SyncScheduler + #syncTimeoutMs: number + #syncTimeout: number | null ++ #backstopDelayMs: number + #compactionDeleteConcurrency: number + #blobInterceptor?: BlobInterceptor + #onRemoteHeadsChanged?: OnRemoteHeadsChanged +@@ -390,9 +489,14 @@ export class SubductionSource implements DocumentSource { + /** Last reported aggregate connectedness, to fire onConnectionChanged only + * on transitions. */ + #lastConnected = false +- /** Last surfaced remote heads per `sid|storageId`, to dedupe repeated +- * notifications (avoids redundant handle events and storage writes). */ +- #lastRemoteHeads = new Map() ++ /** ++ * Last surfaced remote heads per sedimentree, keyed inside by peer. ++ * Dedupes repeated notifications (avoids redundant handle events and ++ * storage writes), carries the observation timestamp so a sibling ++ * relaying older news cannot walk this node's view backwards, and ++ * keeps the hashes that `#serverHasAllHeads` compares against. ++ */ ++ #lastRemoteHeads = new Map>() + /** + * Set at the top of `shutdown()`. Makes `#scheduleRecompute` a + * no-op so no new background sync rounds spin up mid-teardown; the +@@ -418,12 +522,15 @@ export class SubductionSource implements DocumentSource { + timeouts, + compactionDeleteConcurrency = DEFAULT_COMPACTION_DELETE_CONCURRENCY, + blobInterceptor, ++ headsChannel, + }: SubductionSourceOptions) { + this.#blobInterceptor = blobInterceptor + this.#onConnectionChanged = onConnectionChanged + this.priority = priority + this.#syncTimeoutMs = timeouts?.syncMs ?? DEFAULT_SYNC_TIMEOUT_MS + this.#syncTimeout = this.#syncTimeoutMs ++ this.#backstopDelayMs = ++ timeouts?.healInitialDelayMs ?? DEFAULT_BACKSTOP_DELAY_MS + this.#compactionDeleteConcurrency = compactionDeleteConcurrency + // Default roundtrip deadline forwarded to the Subduction constructor. + // The field must be *omitted* (not passed as `undefined`) to get the +@@ -467,13 +574,28 @@ export class SubductionSource implements DocumentSource { + // raw subduction CommitId hex would never compare equal to a doc's + // own heads). + const urlHeads = encodeHeads(heads.map(h => h.toHexString())) +- this.#surfaceRemoteHeads( +- sedimentreeId, ++ const timestamp = Date.now() ++ // Only announce what actually surfaced: the wasm fires repeatedly ++ // with the same heads during a sync burst, and relaying each one ++ // would multiply that burst by the number of tabs on the origin. ++ if ( ++ !this.#surfaceRemoteHeads( ++ sedimentreeId, ++ storageId, ++ urlHeads, ++ timestamp, ++ true ++ ) ++ ) { ++ return ++ } ++ this.#headsChannel?.postMessage({ ++ kind: "remote", ++ sid: sedimentreeId.toString(), + storageId, +- urlHeads, +- Date.now(), +- true +- ) ++ heads: urlHeads, ++ timestamp, ++ } satisfies HeadsChannelMessage) + } + + // Construct without hydrating: skip preloading persisted sedimentrees +@@ -491,6 +613,32 @@ export class SubductionSource implements DocumentSource { + ...defaultTimeoutOption, + }) + ) ++ if (headsChannel && typeof BroadcastChannel !== "undefined") { ++ this.#headsChannel = new BroadcastChannel(headsChannel) ++ this.#headsChannel.unref?.() ++ this.#headsChannel.onmessage = ({ ++ data, ++ }: MessageEvent) => { ++ const entry = this.#entries.get(data.sid) ++ if (!entry) return ++ if (data.kind === "remote") { ++ // Never re-announce: the observing tab already told everyone. ++ this.#surfaceRemoteHeads( ++ entry.sedimentreeId, ++ data.storageId, ++ data.heads as UrlHeads, ++ data.timestamp, ++ false ++ ) ++ return ++ } ++ if (data.kind !== "saved") return ++ if (data.heads.every(h => entry.knownHashes.has(h))) return ++ entry.localLoadPending = true ++ entry.siblingLoadPending = true ++ this.#scheduleRecompute(entry) ++ } ++ } + + // ── Connection managers ───────────────────────────────────────── + // Endpoints own *where* their socket lives (in-thread, Worker, ...); +@@ -554,6 +702,9 @@ export class SubductionSource implements DocumentSource { + * when `persist` is set, durably store it so the last-known sync state + * survives reload and replays on the next attach. `heads` must already be + * in automerge-repo `UrlHeads` (bs58check) form. ++ * ++ * Returns whether this was news: `false` means the observation was a ++ * repeat, or older than what is already held for that peer. + */ + #surfaceRemoteHeads( + sedimentreeId: SedimentreeId, +@@ -561,14 +712,24 @@ export class SubductionSource implements DocumentSource { + heads: UrlHeads, + timestamp: number, + persist: boolean +- ) { +- // Dedupe: skip if this peer's heads for this sedimentree are unchanged +- // since we last surfaced them (the wasm can fire repeatedly with the +- // same heads during a sync burst). +- const dedupeKey = `${sedimentreeId.toString()}|${storageId}` ++ ): boolean { ++ const sidStr = sedimentreeId.toString() + const joined = [...heads].sort().join(",") +- if (this.#lastRemoteHeads.get(dedupeKey) === joined) return +- this.#lastRemoteHeads.set(dedupeKey, joined) ++ const byPeer = this.#lastRemoteHeads.get(sidStr) ++ const previous = byPeer?.get(storageId) ++ // Skip a repeat (the wasm fires repeatedly with the same heads during ++ // a sync burst), and skip stale news: heads carry no order of their ++ // own, so an older observation relayed by a sibling would otherwise ++ // replace a newer one and walk this node's view of the peer back. ++ if ( ++ previous !== undefined && ++ (previous.joined === joined || timestamp < previous.timestamp) ++ ) { ++ return false ++ } ++ const surfaced = { joined, timestamp, hashes: decodeHeads(heads) } ++ if (byPeer) byPeer.set(storageId, surfaced) ++ else this.#lastRemoteHeads.set(sidStr, new Map([[storageId, surfaced]])) + + const documentId = toDocumentId(sedimentreeId) + this.#onRemoteHeadsChanged?.( +@@ -582,6 +743,22 @@ export class SubductionSource implements DocumentSource { + .saveRemoteHeads(sedimentreeId, storageId, heads, timestamp) + .catch(e => this.#log.debug("saveRemoteHeads failed: %O", e)) + } ++ return true ++ } ++ ++ /** ++ * Close the heads channel and forget it. ++ * ++ * Forgetting is the point. `shutdown()` closes the channel after its ++ * flush-then-final-sync pass, but the 100 ms save throttle can still ++ * fire after that, and `postMessage` on a closed BroadcastChannel ++ * throws `InvalidStateError`. `#save` is called as `void this.#save(…)` ++ * with no catch, so the throw would surface as an unhandled rejection ++ * and skip the rest of that save. ++ */ ++ #closeHeadsChannel() { ++ this.#headsChannel?.close() ++ this.#headsChannel = null + } + + #anyConnectionManagerConnecting(): boolean { +@@ -631,6 +808,7 @@ export class SubductionSource implements DocumentSource { + // — recomputes are targeted, so other entries' completions no + // longer walk this one. Mark it dirty so the walk retries the load. + if (entry.syncState === "initializing") { ++ entry.localLoadPending = true + this.#scheduleRecompute(entry) + return + } +@@ -766,6 +944,7 @@ export class SubductionSource implements DocumentSource { + if (!entry) return + void this.#save(entry) + }, 100) ++ const onHeadsChanged = () => throttledSave() + + this.#entries.set(sidStr, { + syncState: "initializing", +@@ -775,17 +954,23 @@ export class SubductionSource implements DocumentSource { + syncInFlight: false, + lastSyncResult: null, + lastSyncGeneration: -1, ++ backstopSyncs: 0, ++ backstopTimer: null, + blobLoadInFlight: false, + blobRetries: 0, + needsResync: false, ++ localLoadPending: true, ++ siblingLoadPending: false, + syncSettled: Promise.resolve(), + lastSavedHeads: new Set(), + knownHashes: new Set(), ++ strandedHashes: new Set(), + persistedCommitHashes: new Set(), + untransformedHashes: new Set(), + persistedFragmentHashes: new Set(), + compactionInFlight: null, + flushSave: throttledSave, ++ onHeadsChanged, + saveSettled, + saveInProgress: false, + saveDeltaPending: false, +@@ -799,7 +984,7 @@ export class SubductionSource implements DocumentSource { + + query.sourcePending("subduction") + +- query.handle.on("heads-changed", () => throttledSave()) ++ query.handle.on("heads-changed", onHeadsChanged) + throttledSave() + + // Subscribe to ephemeral messages for this sedimentree +@@ -851,6 +1036,7 @@ export class SubductionSource implements DocumentSource { + } + })() + ++ this.#lastRemoteHeads.delete(sidStr) + // Replay the last-known remote heads for this sedimentree so the + // DocHandle reflects the last sync state immediately on (re)attach, + // before any network round-trip completes. Does not re-persist. +@@ -877,7 +1063,45 @@ export class SubductionSource implements DocumentSource { + this.#scheduleRecompute(this.#entries.get(sidStr)) + } + +- detach(_documentId: DocumentId): void {} ++ async detach(documentId: DocumentId): Promise { ++ const sid = toSedimentreeId(documentId) ++ const sidStr = sid.toString() ++ const entry = this.#entries.get(sidStr) ++ if (!entry) return ++ this.#entries.delete(sidStr) ++ if (entry.backstopTimer !== null) clearTimeout(entry.backstopTimer) ++ entry.handle.off("heads-changed", entry.onHeadsChanged) ++ await entry.saveSettled ++ for (let round = 0; round < MAX_DETACH_SAVE_ROUNDS; round++) { ++ await this.#save(entry) ++ if (!entry.saveDeltaPending) break ++ } ++ if (entry.lastSaveError !== null) { ++ this.#log.warn( ++ `detach ${sidStr.slice(0, 8)}: unsaved commits could not be persisted`, ++ entry.lastSaveError ++ ) ++ } ++ await entry.syncSettled ++ if (entry.needsResync || this.#needsOutboundSync(entry)) { ++ await this.#doSync(entry, SHUTDOWN_SYNC_TIMEOUT_MS) ++ } ++ let subduction: Subduction | null ++ try { ++ subduction = await this.#subduction ++ } catch (e) { ++ this.#log.debug("subduction never initialized, skipping unsubscribe: %O", e) ++ subduction = null ++ } ++ if (this.#entries.has(sidStr)) return ++ this.#scheduler.resetHealState(sidStr) ++ if (!subduction) return ++ try { ++ await subduction.unsubscribeEphemeral([Topic.fromBytes(sid.toBytes())]) ++ } catch (e) { ++ this.#log.debug("ephemeral unsubscribe failed: %O", e) ++ } ++ } + + shareConfigChanged(): void { + for (const entry of this.#entries.values()) { +@@ -1024,6 +1248,205 @@ export class SubductionSource implements DocumentSource { + ) + } + ++ /** ++ * True when some peer has been seen holding every head this entry's ++ * handle holds — there is nothing left for this node to get across. ++ * Vacuously true for a handle with nothing in it. ++ */ ++ #serverHasAllHeads(entry: SedimentreeEntry): boolean { ++ const doc = entry.handle.doc() ++ if (!doc) return true ++ const heads = Automerge.getHeads(doc) ++ if (heads.length === 0) return true ++ const byPeer = this.#lastRemoteHeads.get(entry.sedimentreeId.toString()) ++ if (byPeer === undefined) return false ++ for (const { hashes } of byPeer.values()) { ++ if (heads.every(h => hashes.includes(h))) return true ++ } ++ return false ++ } ++ ++ /** ++ * Backstop for the sync predicate. `#needsOutboundSync` knows only ++ * whether this entry synced since the connection generation moved; it ++ * cannot tell "synced and caught up" from "synced, and the handle has ++ * since advanced by a path that stored nothing" — a sibling's write ++ * read back out of shared storage, say. That entry would otherwise sit ++ * diverged with nothing scheduled to notice. ++ * ++ * Counted, not merely conditioned. See {@link MAX_BACKSTOP_SYNCS}: the ++ * divergence can be permanent, and a bare "sync while not contained" ++ * predicate would then round-trip for the life of the tab. Routed ++ * through the healer so the rounds it does open inherit the healer's ++ * backoff and attempt cap, and left to the healer entirely once a real ++ * failure is being retried. ++ */ ++ #maybeBackstopSync(entry: SedimentreeEntry) { ++ if (this.#serverHasAllHeads(entry)) { ++ entry.backstopSyncs = 0 ++ // A peer holding every head holds every ancestor, so nothing the ++ // handle carries is stranded any more. ++ entry.strandedHashes.clear() ++ return ++ } ++ if (entry.backstopTimer !== null || this.#backstopBlocked(entry)) return ++ entry.backstopTimer = setTimeout( ++ () => this.#backstopExpired(entry), ++ this.#backstopDelayMs ++ ) ++ } ++ ++ #backstopBlocked(entry: SedimentreeEntry): boolean { ++ if (entry.lastSyncResult === "all-failed") return true ++ // Disconnected there is no round to open, but a stranded commit has ++ // to reach the resident tree *before* the reconnect round reads it, ++ // or that round pushes around it. So the timer still arms offline — ++ // for the re-ingest alone, and only while something is stranded. ++ if (!this.isConnected()) return entry.strandedHashes.size === 0 ++ // The cap counts rounds, not writes: a stranded commit still needs ++ // its one duplicate write after the round budget is spent. ++ if (entry.backstopSyncs >= MAX_BACKSTOP_SYNCS) ++ return entry.strandedHashes.size === 0 ++ return false ++ } ++ ++ /** Look again once {@link DEFAULT_BACKSTOP_DELAY_MS} is up. */ ++ #backstopExpired(entry: SedimentreeEntry) { ++ entry.backstopTimer = null ++ if (this.#shuttingDown) return ++ if (this.#serverHasAllHeads(entry)) { ++ entry.backstopSyncs = 0 ++ entry.strandedHashes.clear() ++ return ++ } ++ if (this.#backstopBlocked(entry)) return ++ ++ // Settled, and still diverged: this is the moment the entry has ++ // earned a duplicate write. Offline that is the whole of it — there ++ // is no round to open, and the next reconnect round will find the ++ // tree already holding the sibling's commits. ++ const reingested = this.#reingestStranded(entry) ++ if (!this.isConnected()) return ++ ++ // The write above is owed to the commit; the round is not. A spent ++ // budget, or a heal chain already retrying this entry, gates the ++ // round alone. ++ if ( ++ entry.backstopSyncs >= MAX_BACKSTOP_SYNCS || ++ this.#scheduler.isHealing(entry.sedimentreeId) ++ ) { ++ return ++ } ++ ++ entry.backstopSyncs++ ++ this.#log.debug( ++ `backstop sync ${entry.sedimentreeId.toString().slice(0, 8)} ` + ++ `(${entry.backstopSyncs}/${MAX_BACKSTOP_SYNCS})` ++ ) ++ const heal = () => { ++ // The re-ingest can outlive the entry: `detach()` resets this ++ // sedimentree's heal state, and a round for an evicted document is ++ // noise the protocol handler worker pays for in bursts. ++ if (this.#shuttingDown) return ++ if (this.#entries.get(entry.sedimentreeId.toString()) !== entry) return ++ this.#scheduler.scheduleHealSync(entry.sedimentreeId) ++ } ++ // The round has to read a tree that already holds the re-ingested ++ // commits, or it pushes exactly what the last one did. ++ if (reingested) void reingested.then(heal) ++ else heal() ++ } ++ ++ /** ++ * Hand the next save back the commits this node holds only because it ++ * read a sibling's write out of the shared database. ++ * ++ * Dropping them from `knownHashes` un-filters them in ++ * `#saveNewCommits`, and clearing the save baseline gets them past ++ * `#save`'s heads fast path, which the channel-driven reload already ++ * advanced past them. The save writes them a second time — a ++ * duplicate of a record already on disk — and that write is what puts ++ * them in the resident tree, where a round can finally read them. ++ * ++ * At most one duplicate write per commit, ever: a reload only marks a ++ * hash `knownHashes` did not already hold, and `#saveNewCommits` ++ * unmarks every hash it stores. A server that keeps rejecting the ++ * commit therefore costs re-tries of the round (capped by ++ * {@link MAX_BACKSTOP_SYNCS}), not re-tries of the write. ++ * ++ * Returns null when there was nothing stranded, else a promise for the ++ * save that stores it. ++ */ ++ #reingestStranded(entry: SedimentreeEntry): Promise | null { ++ this.#forgetStrandedOnTheServer(entry) ++ if (entry.strandedHashes.size === 0) return null ++ const batch = [...entry.strandedHashes] ++ this.#log.debug( ++ `re-ingesting ${batch.length} stranded commit(s) for ` + ++ entry.sedimentreeId.toString().slice(0, 8) ++ ) ++ for (const hash of batch) entry.knownHashes.delete(hash) ++ return (async () => { ++ // A save that started before the drop above filtered on the old ++ // `knownHashes` and would leave the commits stranded; let it go ++ // before arming the one that sees them. It restores the save ++ // baseline as it finishes, so the baseline is cleared after it: ++ // cleared first, the finishing save puts it straight back and the ++ // save below returns on `#save`'s heads fast path. ++ await entry.saveSettled ++ entry.lastSavedHeads.clear() ++ entry.flushSave() ++ entry.flushSave.flush() ++ await entry.saveSettled ++ // `#saveNewCommits` has already unmarked what it stored. What is ++ // left of this batch is what it could not store — absorbed into a ++ // fragment since the reload, say — and holding on to that would ++ // arm the backstop for ever. A save that failed outright keeps its ++ // batch, so the next expiry retries it. ++ if (entry.lastSaveError === null) { ++ for (const hash of batch) entry.strandedHashes.delete(hash) ++ } ++ })() ++ } ++ ++ /** ++ * Forget the stranded commits some peer's heads already descend from. ++ * ++ * `#serverHasAllHeads` asks whether the server holds everything this ++ * handle holds, which is false for the ordinary moment between a local ++ * edit and its confirmation. That is the wrong question here: what the ++ * re-ingest is for is a commit nobody else will push, and a commit the ++ * server's heads descend from has already been pushed by the tab that ++ * wrote it. Asking the narrower question is what keeps a busy origin ++ * from paying a duplicate write per sibling edit. ++ * ++ * Only runs with something stranded, which in normal operation is ++ * nothing at all. ++ */ ++ #forgetStrandedOnTheServer(entry: SedimentreeEntry) { ++ if (entry.strandedHashes.size === 0) return ++ const byPeer = this.#lastRemoteHeads.get(entry.sedimentreeId.toString()) ++ if (byPeer === undefined) return ++ const doc = entry.handle.doc() ++ if (!doc) return ++ for (const { hashes } of byPeer.values()) { ++ let unpushed: Set ++ try { ++ unpushed = new Set( ++ Automerge.getChangesMetaSince(doc, hashes).map(m => m.hash) ++ ) ++ } catch { ++ // Heads this doc does not have; that peer says nothing about ++ // what it holds of ours. ++ continue ++ } ++ for (const hash of entry.strandedHashes) { ++ if (!unpushed.has(hash)) entry.strandedHashes.delete(hash) ++ } ++ if (entry.strandedHashes.size === 0) return ++ } ++ } ++ + #recomputeEntry(entry: SedimentreeEntry) { + // No new background syncs once shutting down: `shutdown()`'s quiesce + // round is the sole sync initiator from that point. A recompute walk +@@ -1051,6 +1474,9 @@ export class SubductionSource implements DocumentSource { + entry.query.sourcePending("subduction") + entry.syncInFlight = true + void this.#doSync(entry) ++ } else if (!this.isConnected()) { ++ entry.blobLoadInFlight = true ++ void this.#loadBlobsAndTransition(entry) + } else if ( + Automerge.getHeads(entry.handle.fullDoc()).length === 0 && + !this.#anyConnectionManagerConnecting() +@@ -1060,13 +1486,31 @@ export class SubductionSource implements DocumentSource { + entry.query.sourceUnavailable("subduction") + } + } ++ if (entry.localLoadPending && !entry.blobLoadInFlight) { ++ entry.localLoadPending = false ++ const fromSibling = entry.siblingLoadPending ++ entry.siblingLoadPending = false ++ entry.blobLoadInFlight = true ++ void this.#loadBlobsAndTransition(entry, true, fromSibling) ++ } + return + } + + case "running": { +- if (!entry.syncInFlight && this.#needsOutboundSync(entry)) { +- entry.syncInFlight = true +- void this.#doSync(entry) ++ if (entry.localLoadPending && !entry.blobLoadInFlight) { ++ entry.localLoadPending = false ++ const fromSibling = entry.siblingLoadPending ++ entry.siblingLoadPending = false ++ entry.blobLoadInFlight = true ++ void this.#loadBlobsAndTransition(entry, true, fromSibling) ++ } ++ if (!entry.syncInFlight) { ++ if (this.#needsOutboundSync(entry)) { ++ entry.syncInFlight = true ++ void this.#doSync(entry) ++ } else { ++ this.#maybeBackstopSync(entry) ++ } + } + return + } +@@ -1103,6 +1547,8 @@ export class SubductionSource implements DocumentSource { + timeoutMs !== undefined ? timeoutMs : this.#syncTimeout + ) + ++ if (this.#entries.get(sedimentreeId.toString()) !== entry) return ++ + const results = peerResultMap.entries() + const anySuccess = results.some(r => r.success) + this.#log.debug( +@@ -1120,7 +1566,7 @@ export class SubductionSource implements DocumentSource { + // arrives, without waiting for further state transitions. + // + // During "initializing", `#loadBlobsAndTransition` performs one +- // atomic getBlobs + loadIncremental after sync succeeds; loading ++ // atomic bulk read + loadIncremental after sync succeeds; loading + // here too would duplicate storage reads on first open. + // + // Also retry when the handle is empty: a previous load may have +@@ -1204,7 +1650,7 @@ export class SubductionSource implements DocumentSource { + } + + /** +- * Load all blobs for a sedimentree from Subduction and apply them to the ++ * Load all blobs for a sedimentree from storage and apply them to the + * handle via `Automerge.loadIncremental`. If new data was loaded, signal + * the query so it can transition to "ready". + * +@@ -1213,41 +1659,64 @@ export class SubductionSource implements DocumentSource { + */ + async #loadBlobsIntoHandle( + entry: SedimentreeEntry, +- subduction: Subduction ++ subduction: Subduction, ++ fromSibling = false + ): Promise { + const sid = entry.sedimentreeId.toString().slice(0, 8) ++ // An announcement still waiting for its own reload is answered by ++ // this read too — it reads the same shared database. Honour it ++ // whoever asked for the read, so a peer-fed load landing in that ++ // window cannot file the sibling's commit as known and not stranded. ++ // Left set for `#recomputeEntry` to consume as usual. Only once ++ // running: while initializing the resident tree is hydrated from ++ // this same storage, so nothing read here is stranded. ++ if (entry.syncState === "running" && entry.siblingLoadPending) { ++ fromSibling = true ++ } + const headCount = () => Automerge.getHeads(entry.handle.fullDoc()).length + +- const allBlobs = await subduction.getBlobs(entry.sedimentreeId) +- const totalBytes = allBlobs +- ? allBlobs.reduce((n, b) => n + b.byteLength, 0) +- : 0 +- this.#log.debug( +- `loadBlobsIntoHandle ${sid}: ${ +- allBlobs?.length ?? 0 +- } blob(s), ${totalBytes} bytes, heads=${headCount()}` +- ) +- if (!allBlobs || allBlobs.length === 0) return false +- + // Prefer the id-keyed view of the same blobs. Fall back to the id-less one +- // if it does not cover everything subduction reports. +- let blobsWithIds: Array<{ idHex: string; blob: Uint8Array }> | null = null ++ // if reading it throws. ++ let blobsWithIds: Array<{ idHex: string; blob: Uint8Array }> + try { +- const keyed = await this.#idKeyedBlobs(entry) +- if (keyed.length >= allBlobs.length) blobsWithIds = keyed ++ blobsWithIds = await this.#idKeyedBlobs(entry) + } catch (e) { + this.#log.debug( + "idKeyedBlobs failed, falling back to id-less load: %O", + e + ) ++ const allBlobs = await subduction.getBlobs(entry.sedimentreeId) ++ blobsWithIds = (allBlobs ?? []).map(blob => ({ idHex: "", blob })) + } +- if (!blobsWithIds) blobsWithIds = allBlobs.map(blob => ({ idHex: "", blob })) ++ const totalBytes = blobsWithIds.reduce((n, u) => n + u.blob.byteLength, 0) ++ this.#log.debug( ++ `loadBlobsIntoHandle ${sid}: ${ ++ blobsWithIds.length ++ } blob(s), ${totalBytes} bytes, heads=${headCount()}` ++ ) ++ if (blobsWithIds.length === 0) return false + blobsWithIds.sort((a, b) => b.blob.byteLength - a.blob.byteLength) + + // Record these IDs as known hashes since nothing in subduction's storage +- // should be re-encrypted and re-saved. ++ // should be re-encrypted and re-saved. A sibling's write read back out ++ // of the shared database is the exception: `knownHashes` then hides a ++ // commit the resident tree has never seen, so mark it stranded too. ++ let newlyStranded = 0 + for (const u of blobsWithIds) { +- if (u.idHex) entry.knownHashes.add(u.idHex) ++ if (!u.idHex) continue ++ if (fromSibling && !entry.knownHashes.has(u.idHex)) { ++ entry.strandedHashes.add(u.idHex) ++ newlyStranded++ ++ } ++ entry.knownHashes.add(u.idHex) ++ } ++ // A commit that has only just arrived is no evidence of divergence: ++ // the tab that wrote it is still mid-push. Start the settling delay ++ // again from here — the recompute in `#loadBlobsAndTransition` re-arms ++ // it — so a stream of sibling edits never expires the backstop. ++ if (newlyStranded > 0 && entry.backstopTimer !== null) { ++ clearTimeout(entry.backstopTimer) ++ entry.backstopTimer = null + } + + let toApply = blobsWithIds.map(u => u.blob) +@@ -1300,10 +1769,32 @@ export class SubductionSource implements DocumentSource { + return true + } + +- async #loadBlobsAndTransition(entry: SedimentreeEntry) { ++ /** ++ * `fromSibling` is consumed by the caller, not read off the entry ++ * here: an announcement that lands while this load is in flight sets ++ * the flag again, and clearing it afterwards would erase it. The ++ * follow-up reload would then file the sibling's commit as known and ++ * not stranded — the one commit the re-ingest exists for. ++ */ ++ async #loadBlobsAndTransition( ++ entry: SedimentreeEntry, ++ onlyIfData = false, ++ fromSibling = false ++ ) { + try { + const subduction = await this.#subduction +- const hadBlobs = await this.#loadBlobsIntoHandle(entry, subduction) ++ const hadBlobs = await this.#loadBlobsIntoHandle( ++ entry, ++ subduction, ++ fromSibling ++ ) ++ entry.blobRetries = 0 ++ if ( ++ onlyIfData && ++ Automerge.getHeads(entry.handle.fullDoc()).length === 0 ++ ) { ++ return ++ } + + entry.syncState = "running" + +@@ -1328,6 +1819,24 @@ export class SubductionSource implements DocumentSource { + entry.query.sourcePending("subduction") + } + } catch (e) { ++ if (onlyIfData) { ++ if (entry.blobRetries < MAX_LOCAL_LOAD_RETRIES) { ++ entry.blobRetries++ ++ entry.localLoadPending = true ++ if (fromSibling) entry.siblingLoadPending = true ++ } else { ++ entry.blobRetries = 0 ++ this.#log.warn( ++ `reload from storage failed ${ ++ MAX_LOCAL_LOAD_RETRIES + 1 ++ } times for ${entry.sedimentreeId ++ .toString() ++ .slice(0, 8)}; waiting for the next announcement`, ++ e ++ ) ++ } ++ return ++ } + this.#log.debug( + `loadBlobsAndTransition threw for ${entry.sedimentreeId + .toString() +@@ -1416,6 +1925,7 @@ export class SubductionSource implements DocumentSource { + resolveSaveSettled = r + }) + ++ let stored = 0 + try { + const doc = entry.handle.fullDoc() + if (!doc) return +@@ -1438,7 +1948,7 @@ export class SubductionSource implements DocumentSource { + `syncInFlight=${entry.syncInFlight}` + ) + +- let result: { prepFailures: number } ++ let result: { prepFailures: number; stored: number } + try { + result = await this.#saveNewCommits(entry, doc, subduction) + } catch (e) { +@@ -1467,6 +1977,14 @@ export class SubductionSource implements DocumentSource { + // so future mutations to either side stay isolated. + entry.lastSaveError = null + entry.lastSavedHeads = new Set(currentSet) ++ stored = result.stored ++ if (stored > 0) { ++ this.#headsChannel?.postMessage({ ++ kind: "saved", ++ sid: entry.sedimentreeId.toString(), ++ heads: currentHeads, ++ } satisfies HeadsChannelMessage) ++ } + + // Detect a post-save delta the `heads-changed` listener didn't + // fire on. `getChangeByHash` calls inside `#saveNewCommits` can +@@ -1498,6 +2016,8 @@ export class SubductionSource implements DocumentSource { + resolveSaveSettled() + } + ++ if (stored === 0) return ++ + // Trigger an immediate sync — the only broadcast path for newly + // saved commits (`#saveNewCommits` is store-only). If a sync is + // already in flight, flag a re-sync for when it completes; otherwise +@@ -1521,7 +2041,7 @@ export class SubductionSource implements DocumentSource { + entry: SedimentreeEntry, + doc: Automerge.Doc, + subduction: Subduction +- ): Promise<{ prepFailures: number }> { ++ ): Promise<{ prepFailures: number; stored: number }> { + // Ask the doc directly for its level-0 loose commits and its + // higher-level fragments. Automerge core owns the compaction + // policy (which commits get absorbed into which fragments); we +@@ -1558,7 +2078,7 @@ export class SubductionSource implements DocumentSource { + ) + + if (newCommitMetas.length === 0 && newFragmentMetas.length === 0) { +- return { prepFailures: 0 } ++ return { prepFailures: 0, stored: 0 } + } + + const newCommitBytes = +@@ -1677,7 +2197,7 @@ export class SubductionSource implements DocumentSource { + } + + if (commitInputs.length === 0 && fragmentInputs.length === 0) { +- return { prepFailures: prepErrors.length } ++ return { prepFailures: prepErrors.length, stored: 0 } + } + + // Record hashes BEFORE `storeBuiltBatch` so the synchronous +@@ -1712,9 +2232,16 @@ export class SubductionSource implements DocumentSource { + throw e + } + ++ // Stored for real now, so not stranded any more — including anything ++ // a reload marked while this batch was in flight. ++ for (const hash of acceptedHashes) entry.strandedHashes.delete(hash) ++ + this.#scheduleCompaction(entry) + +- return { prepFailures: prepErrors.length } ++ return { ++ prepFailures: prepErrors.length, ++ stored: commitInputs.length + fragmentInputs.length, ++ } + } + + /** +@@ -1892,6 +2419,7 @@ export class SubductionSource implements DocumentSource { + if (entry === undefined) return + this.#scheduler.resetHealState(sid) + entry.lastSyncResult = null ++ entry.backstopSyncs = 0 + this.#scheduleRecompute(entry) + } + +@@ -2021,6 +2549,11 @@ export class SubductionSource implements DocumentSource { + + this.#scheduler.shutdown() + ++ for (const entry of this.#entries.values()) { ++ if (entry.backstopTimer !== null) clearTimeout(entry.backstopTimer) ++ entry.backstopTimer = null ++ } ++ + for (const entry of this.#entries.values()) this.#flushInbound(entry) + + for (const entry of this.#entries.values()) { +@@ -2076,6 +2609,7 @@ export class SubductionSource implements DocumentSource { + for (const endpoint of this.#websocketEndpoints) { + endpoint.shutdown?.() + } ++ this.#closeHeadsChannel() + this.#storage.close() + await this.#awaitStorageIdle() + return +@@ -2090,6 +2624,7 @@ export class SubductionSource implements DocumentSource { + endpoint.shutdown?.() + } + ++ this.#closeHeadsChannel() + this.#storage.close() + await this.#awaitStorageIdle() + } diff --git a/pnpm-lock.yaml b/pnpm-lock.yaml index ebf6e24d1..8ddc4f611 100644 --- a/pnpm-lock.yaml +++ b/pnpm-lock.yaml @@ -50,7 +50,7 @@ overrides: solid-automerge: ^2.0.1 patchedDependencies: - '@automerge/automerge-repo@2.6.0-subduction.48': 37266639cf8001b55f69d72d72274d4c512898628937a4e0c8b96cb0b8dd67f4 + '@automerge/automerge-repo@2.6.0-subduction.48': 24ddfeb2d5556d98ef58d0620ebec9dbb46f2e35bc70491c7551275cb731f465 importers: @@ -82,7 +82,7 @@ importers: version: 3.5.0 '@automerge/automerge-repo': specifier: 2.6.0-subduction.48 - version: 2.6.0-subduction.48(patch_hash=37266639cf8001b55f69d72d72274d4c512898628937a4e0c8b96cb0b8dd67f4) + version: 2.6.0-subduction.48(patch_hash=24ddfeb2d5556d98ef58d0620ebec9dbb46f2e35bc70491c7551275cb731f465) '@automerge/automerge-repo-keyhive': specifier: 0.5.0-alpha.7 version: 0.5.0-alpha.7(ws@8.21.1) @@ -168,7 +168,7 @@ importers: version: 3.5.0 '@automerge/automerge-repo': specifier: 2.6.0-subduction.48 - version: 2.6.0-subduction.48(patch_hash=37266639cf8001b55f69d72d72274d4c512898628937a4e0c8b96cb0b8dd67f4) + version: 2.6.0-subduction.48(patch_hash=24ddfeb2d5556d98ef58d0620ebec9dbb46f2e35bc70491c7551275cb731f465) '@automerge/automerge-repo-keyhive': specifier: 0.5.0-alpha.7 version: 0.5.0-alpha.7(ws@8.21.1) @@ -213,7 +213,7 @@ importers: version: 3.5.0 '@automerge/automerge-repo': specifier: 2.6.0-subduction.48 - version: 2.6.0-subduction.48(patch_hash=37266639cf8001b55f69d72d72274d4c512898628937a4e0c8b96cb0b8dd67f4) + version: 2.6.0-subduction.48(patch_hash=24ddfeb2d5556d98ef58d0620ebec9dbb46f2e35bc70491c7551275cb731f465) debug: specifier: ^4.4.3 version: 4.4.3 @@ -244,7 +244,7 @@ importers: version: 3.5.0 '@automerge/automerge-repo': specifier: 2.6.0-subduction.48 - version: 2.6.0-subduction.48(patch_hash=37266639cf8001b55f69d72d72274d4c512898628937a4e0c8b96cb0b8dd67f4) + version: 2.6.0-subduction.48(patch_hash=24ddfeb2d5556d98ef58d0620ebec9dbb46f2e35bc70491c7551275cb731f465) '@automerge/automerge-repo-keyhive': specifier: 0.5.0-alpha.7 version: 0.5.0-alpha.7(ws@8.21.1) @@ -318,7 +318,7 @@ importers: version: 3.5.0 '@automerge/automerge-repo': specifier: 2.6.0-subduction.48 - version: 2.6.0-subduction.48(patch_hash=37266639cf8001b55f69d72d72274d4c512898628937a4e0c8b96cb0b8dd67f4) + version: 2.6.0-subduction.48(patch_hash=24ddfeb2d5556d98ef58d0620ebec9dbb46f2e35bc70491c7551275cb731f465) '@automerge/automerge-repo-keyhive': specifier: 0.5.0-alpha.7 version: 0.5.0-alpha.7(ws@8.21.1) @@ -349,7 +349,7 @@ importers: version: 3.5.0 '@automerge/automerge-repo': specifier: 2.6.0-subduction.48 - version: 2.6.0-subduction.48(patch_hash=37266639cf8001b55f69d72d72274d4c512898628937a4e0c8b96cb0b8dd67f4) + version: 2.6.0-subduction.48(patch_hash=24ddfeb2d5556d98ef58d0620ebec9dbb46f2e35bc70491c7551275cb731f465) '@automerge/automerge-subduction': specifier: 0.16.1 version: 0.16.1 @@ -364,7 +364,7 @@ importers: devDependencies: '@automerge/automerge-repo': specifier: 2.6.0-subduction.48 - version: 2.6.0-subduction.48(patch_hash=37266639cf8001b55f69d72d72274d4c512898628937a4e0c8b96cb0b8dd67f4) + version: 2.6.0-subduction.48(patch_hash=24ddfeb2d5556d98ef58d0620ebec9dbb46f2e35bc70491c7551275cb731f465) typescript: specifier: ^5.9.3 version: 5.9.3 @@ -377,7 +377,7 @@ importers: devDependencies: '@automerge/automerge-repo': specifier: 2.6.0-subduction.48 - version: 2.6.0-subduction.48(patch_hash=37266639cf8001b55f69d72d72274d4c512898628937a4e0c8b96cb0b8dd67f4) + version: 2.6.0-subduction.48(patch_hash=24ddfeb2d5556d98ef58d0620ebec9dbb46f2e35bc70491c7551275cb731f465) '@automerge/automerge-repo-react-hooks': specifier: 2.6.0-subduction.48 version: 2.6.0-subduction.48(react-dom@18.3.1(react@18.3.1))(react@18.3.1) @@ -399,10 +399,10 @@ importers: devDependencies: '@automerge/automerge-repo': specifier: 2.6.0-subduction.48 - version: 2.6.0-subduction.48(patch_hash=37266639cf8001b55f69d72d72274d4c512898628937a4e0c8b96cb0b8dd67f4) + version: 2.6.0-subduction.48(patch_hash=24ddfeb2d5556d98ef58d0620ebec9dbb46f2e35bc70491c7551275cb731f465) solid-automerge: specifier: ^2.0.1 - version: 2.0.1(@automerge/automerge-repo@2.6.0-subduction.48(patch_hash=37266639cf8001b55f69d72d72274d4c512898628937a4e0c8b96cb0b8dd67f4))(solid-js@1.9.14) + version: 2.0.1(@automerge/automerge-repo@2.6.0-subduction.48(patch_hash=24ddfeb2d5556d98ef58d0620ebec9dbb46f2e35bc70491c7551275cb731f465))(solid-js@1.9.14) solid-js: specifier: ^1.9.13 version: 1.9.14 @@ -414,7 +414,7 @@ importers: dependencies: '@automerge/automerge-repo': specifier: 2.6.0-subduction.48 - version: 2.6.0-subduction.48(patch_hash=37266639cf8001b55f69d72d72274d4c512898628937a4e0c8b96cb0b8dd67f4) + version: 2.6.0-subduction.48(patch_hash=24ddfeb2d5556d98ef58d0620ebec9dbb46f2e35bc70491c7551275cb731f465) '@automerge/automerge-repo-network-broadcastchannel': specifier: 2.6.0-subduction.48 version: 2.6.0-subduction.48 @@ -2037,7 +2037,7 @@ snapshots: '@automerge/automerge-repo-keyhive@0.5.0-alpha.7(ws@8.21.1)': dependencies: - '@automerge/automerge-repo': 2.6.0-subduction.48(patch_hash=37266639cf8001b55f69d72d72274d4c512898628937a4e0c8b96cb0b8dd67f4) + '@automerge/automerge-repo': 2.6.0-subduction.48(patch_hash=24ddfeb2d5556d98ef58d0620ebec9dbb46f2e35bc70491c7551275cb731f465) '@automerge/automerge-repo-network-websocket': 2.6.0-subduction.48 '@automerge/automerge-subduction': 0.16.1 '@keyhive/keyhive': 0.1.0-alpha.8 @@ -2053,7 +2053,7 @@ snapshots: '@automerge/automerge-repo-network-broadcastchannel@2.6.0-subduction.48': dependencies: - '@automerge/automerge-repo': 2.6.0-subduction.48(patch_hash=37266639cf8001b55f69d72d72274d4c512898628937a4e0c8b96cb0b8dd67f4) + '@automerge/automerge-repo': 2.6.0-subduction.48(patch_hash=24ddfeb2d5556d98ef58d0620ebec9dbb46f2e35bc70491c7551275cb731f465) transitivePeerDependencies: - bufferutil - supports-color @@ -2061,7 +2061,7 @@ snapshots: '@automerge/automerge-repo-network-messagechannel@2.6.0-subduction.48': dependencies: - '@automerge/automerge-repo': 2.6.0-subduction.48(patch_hash=37266639cf8001b55f69d72d72274d4c512898628937a4e0c8b96cb0b8dd67f4) + '@automerge/automerge-repo': 2.6.0-subduction.48(patch_hash=24ddfeb2d5556d98ef58d0620ebec9dbb46f2e35bc70491c7551275cb731f465) eventemitter3: 5.0.4 transitivePeerDependencies: - bufferutil @@ -2070,7 +2070,7 @@ snapshots: '@automerge/automerge-repo-network-websocket@2.6.0-subduction.48': dependencies: - '@automerge/automerge-repo': 2.6.0-subduction.48(patch_hash=37266639cf8001b55f69d72d72274d4c512898628937a4e0c8b96cb0b8dd67f4) + '@automerge/automerge-repo': 2.6.0-subduction.48(patch_hash=24ddfeb2d5556d98ef58d0620ebec9dbb46f2e35bc70491c7551275cb731f465) cbor-x: 1.6.4 debug: 4.4.3 eventemitter3: 5.0.4 @@ -2083,7 +2083,7 @@ snapshots: '@automerge/automerge-repo-react-hooks@2.6.0-subduction.48(react-dom@18.3.1(react@18.3.1))(react@18.3.1)': dependencies: '@automerge/automerge': 3.5.0 - '@automerge/automerge-repo': 2.6.0-subduction.48(patch_hash=37266639cf8001b55f69d72d72274d4c512898628937a4e0c8b96cb0b8dd67f4) + '@automerge/automerge-repo': 2.6.0-subduction.48(patch_hash=24ddfeb2d5556d98ef58d0620ebec9dbb46f2e35bc70491c7551275cb731f465) eventemitter3: 5.0.4 react: 18.3.1 react-dom: 18.3.1(react@18.3.1) @@ -2094,13 +2094,13 @@ snapshots: '@automerge/automerge-repo-storage-indexeddb@2.6.0-subduction.48': dependencies: - '@automerge/automerge-repo': 2.6.0-subduction.48(patch_hash=37266639cf8001b55f69d72d72274d4c512898628937a4e0c8b96cb0b8dd67f4) + '@automerge/automerge-repo': 2.6.0-subduction.48(patch_hash=24ddfeb2d5556d98ef58d0620ebec9dbb46f2e35bc70491c7551275cb731f465) transitivePeerDependencies: - bufferutil - supports-color - utf-8-validate - '@automerge/automerge-repo@2.6.0-subduction.48(patch_hash=37266639cf8001b55f69d72d72274d4c512898628937a4e0c8b96cb0b8dd67f4)': + '@automerge/automerge-repo@2.6.0-subduction.48(patch_hash=24ddfeb2d5556d98ef58d0620ebec9dbb46f2e35bc70491c7551275cb731f465)': dependencies: '@automerge/automerge': 3.5.0 '@automerge/automerge-subduction': 0.16.1 @@ -2124,7 +2124,7 @@ snapshots: '@automerge/vanillajs@2.6.0-subduction.48': dependencies: - '@automerge/automerge-repo': 2.6.0-subduction.48(patch_hash=37266639cf8001b55f69d72d72274d4c512898628937a4e0c8b96cb0b8dd67f4) + '@automerge/automerge-repo': 2.6.0-subduction.48(patch_hash=24ddfeb2d5556d98ef58d0620ebec9dbb46f2e35bc70491c7551275cb731f465) '@automerge/automerge-repo-network-broadcastchannel': 2.6.0-subduction.48 '@automerge/automerge-repo-network-messagechannel': 2.6.0-subduction.48 '@automerge/automerge-repo-network-websocket': 2.6.0-subduction.48 @@ -3348,9 +3348,9 @@ snapshots: slash@3.0.0: {} - solid-automerge@2.0.1(@automerge/automerge-repo@2.6.0-subduction.48(patch_hash=37266639cf8001b55f69d72d72274d4c512898628937a4e0c8b96cb0b8dd67f4))(solid-js@1.9.14): + solid-automerge@2.0.1(@automerge/automerge-repo@2.6.0-subduction.48(patch_hash=24ddfeb2d5556d98ef58d0620ebec9dbb46f2e35bc70491c7551275cb731f465))(solid-js@1.9.14): dependencies: - '@automerge/automerge-repo': 2.6.0-subduction.48(patch_hash=37266639cf8001b55f69d72d72274d4c512898628937a4e0c8b96cb0b8dd67f4) + '@automerge/automerge-repo': 2.6.0-subduction.48(patch_hash=24ddfeb2d5556d98ef58d0620ebec9dbb46f2e35bc70491c7551275cb731f465) '@solid-primitives/utils': 6.4.1(solid-js@1.9.14) cabbages: 0.2.10 solid-js: 1.9.14 diff --git a/scripts/lint-slim-imports.mts b/scripts/lint-slim-imports.mts new file mode 100644 index 000000000..c0290dd05 --- /dev/null +++ b/scripts/lint-slim-imports.mts @@ -0,0 +1,42 @@ +import { readdirSync, readFileSync } from "node:fs"; +import { join, relative } from "node:path"; +import { slim } from "../core/bootloader/src/externals-list.ts"; + +const root = join(import.meta.dirname, ".."); +const roots = ["core", "packages", "sites"]; +const skip = new Set(["node_modules", "dist", "test", "tests"]); +const extensions = /\.(?:[cm]?[jt]s|tsx|jsx)$/; + +const names = Object.keys(slim) + .map((name) => name.replace(/[.*+?^${}()|[\]\\]/g, "\\$&")) + .join("|"); +const bare = new RegExp( + `(?:\\bfrom\\s*|\\bimport\\s*\\(?\\s*)["'](${names})["']`, + "g" +); + +function* files(dir: string): Generator { + for (const entry of readdirSync(dir, { withFileTypes: true })) { + if (skip.has(entry.name)) continue; + const path = join(dir, entry.name); + if (entry.isDirectory()) yield* files(path); + else if (extensions.test(entry.name)) yield path; + } +} + +let failures = 0; +for (const dir of roots) { + for (const file of files(join(root, dir))) { + const source = readFileSync(file, "utf8"); + for (const match of source.matchAll(bare)) { + const line = source.slice(0, match.index).split("\n").length; + const name = match[1]!; + console.error( + `${relative(root, file)}:${line}: "${name}" instantiates its own wasm; import "${slim[name]}"` + ); + failures++; + } + } +} + +if (failures) process.exit(1); diff --git a/sites/bench/README.md b/sites/bench/README.md index 247d56b14..e52240a61 100644 --- a/sites/bench/README.md +++ b/sites/bench/README.md @@ -10,30 +10,71 @@ pnpm --filter patchwork-bench bench:headed Chromium only (`pnpm exec playwright install chromium` once). Talks to the real sync server named in the site build, so the server columns need network. -Results land in `bench-results/results.md` and `results.jsonl`. +Results land in `bench-results/results.md` and `results.jsonl`. Port 5199 has +to be free: a `vite preview` left behind by another checkout would otherwise be +reused and serve its build, so the run refuses one instead. ## Modes The page at `/` builds one Repo and nothing else — no shell, no account, no -package list — in one of five shapes, chosen by `?mode=`: +package list — in one of seven shapes, chosen by `?mode=`: | mode | subduction node | storage | server socket | tabs meet via | | --- | --- | --- | --- | --- | -| `patchwork` | in the tab | each tab, IndexedDB worker | one per tab | the siblings BroadcastChannel (subduction mesh), then the server | -| `pertab` | in the tab | each tab, IndexedDB worker | one per tab | the server (or IndexedDB) | -| `pertab-bc` | in the tab | each tab, IndexedDB worker | one per tab | classic automerge sync over a BroadcastChannel, then the server | +| `patchwork` | in the tab | each tab, in-thread IndexedDB | one per tab | `headsChannel` announcements, as `pertab-heads`, then the server | +| `pertab` | in the tab | each tab, in-thread IndexedDB | one per tab | the server (or IndexedDB) | +| `pertab-bc` | in the tab | each tab, in-thread IndexedDB | one per tab | classic automerge sync over a BroadcastChannel, then the server | +| `pertab-mesh` | in the tab | each tab, in-thread IndexedDB | one per tab | patchwork's old siblings mesh (subduction over a BroadcastChannel), each tab signing as itself | +| `pertab-heads` | in the tab | each tab, in-thread IndexedDB | one per tab | the origin-wide signer, so one subduction peer on N sockets; the Repo's `headsChannel` announces the heads a tab just persisted and the others reload the doc from IndexedDB | | `tab-worker` | a dedicated Worker per tab | the worker, in-thread IndexedDB | one per worker | a BroadcastChannel mesh between the workers, then the server | | `shared-worker` | one SharedWorker | the worker, in-thread IndexedDB | one | the worker | `patchwork` is `createRepo()` as shipped, plus the service worker and the -automerge worker that resolves URLs for it; the other four are built in the -page. In the two worker modes the tab's Repo has no storage and one subduction +automerge worker that resolves URLs for it (that worker builds its Repo on the +first URL it is asked to resolve, so in this bench it is a spawned but idle +process); the other six are built in the page. In the two worker modes the tab's Repo has no storage and one subduction peer, its worker, reached over a MessagePort (`src/worker-link.ts`); the worker runs a Repo of its own as the node (`src/node.ts`) but never opens a document. -`?server=none` runs every non-patchwork mode with no socket. +A node like that stores what a tab pushes and answers what a peer asks, but +doesn't carry one peer's commits to another on its own, so when a tab (or a +mesh peer) announces heads the node runs a sync round for that document with +every peer. `?server=none` runs every non-patchwork mode with no socket. -`?storage=direct` swaps the IndexedDB worker adapter for in-thread IndexedDB in +Two things the worker modes show that aren't bugs in the bench: a `find()` +that races a sibling's `create()` settles unavailable, and a plain second +`find()` returns the same settled query (the entry stays initializing, so data +arriving later is skipped) — the helpers' `find` calls +`repo.resyncSubduction()` between tries, which is what an app would have to +do; and a storageless tab's `flush()` resolves before its commits have reached +the worker, so edits made just before the tab closes can be lost — a dedicated +worker dies with its tab and a SharedWorker with its last tab, mid-write. Closing +right after `flush()` loses the whole doc; 300ms later everything is on disk. +`sync.spec` measures the race with disposable tabs and carries on with fresh +ones. + +A closed tab's mesh peer lingers: the BroadcastChannel adapter only announces a +departure from `disconnect()`, which a closing tab (or its terminated worker) +never calls, so in `pertab-mesh` and `tab-worker` the survivors +keep a phantom peer. `churn.spec` is where that would show. + +`?storage=worker` swaps in-thread IndexedDB for the IndexedDB worker adapter in the bare per-tab modes; `storage-adapter.spec` compares the two in `pertab`. +`?mesh=1` and `?signer=shared` add patchwork's old siblings mesh and its +origin-wide signer (`@inkandswitch/patchwork-bootloader/signer`) to any bare +mode, so the shipped topology can be taken apart one piece at a time +(`pertab-mesh` is `pertab&mesh=1`; `pertab-heads` is `patchwork` +minus the service worker and the idle automerge worker). + +`pertab&mesh=1&signer=shared` and `pertab-mesh` differ in the signer. With the origin-wide signer +every tab presents the same subduction peer id, and with three tabs open the +third stops receiving a sibling's edits over the mesh and the server stops +confirming that sibling's heads — every run, in the old `patchwork`, in +`pertab&mesh=1&signer=shared`, and in a keyhive build of it +(`keyhive: true` in vite.config.ts), whose signer ARK derives from the keypair +every context shares. With per-tab signers (`pertab-mesh`) the same mesh +converges. The mechanism is in subduction-core: connections are keyed by peer +id, a sync round stops at the first connection of a peer that answers, and +relays exclude every connection of the sender's id. The sync server's subduction peer id is learned once per run from a bare tab in a throwaway context and passed to every page as `?serverPeer=`, so "the @@ -47,13 +88,23 @@ server holds our heads" means that peer and not a sibling. once they're all up. - `sync.spec` — find a doc a sibling just created (and whether the first `find()` settled unavailable); edit → seen in the other tabs; edit → server - holds our heads. Medians over 10 edits. -- `storage.spec` — a second tab finds the first's doc through storage alone; - two tabs edit the same doc and close, does a third see everything. -- `offline.spec` — both tabs edit with the network cut, then it returns. - Playwright's offline emulation reaches a tab's own socket but not a worker's, - so the worker modes are also told to drop the server link and hold off - reconnecting; whether the link was actually seen down is recorded. + holds our heads; and whether a client in a fresh browser context can get + the finished doc from the server. Medians over 10 edits. +- `storage.spec` — a tab writes, flushes and closes; a fresh tab finds the doc + through storage alone (no server; patchwork's server is build-time, so its + WebSockets are refused instead). Two tabs edit the same doc + and close, does a third see everything. +- `offline.spec` — both tabs edit with the link to the server cut, then it + returns. Playwright's `setOffline` fails new connections like a dead network + but doesn't close a socket that's already open, so the tabs' sockets are also + proxied (`context.routeWebSocket`) purely to close the live ones; a worker's + socket is out of reach of both, so the worker node is told to drop it and + fails its connect attempts while offline. automerge-repo's reconnect backoff + then applies in every mode. Recorded: whether + the tab modes' link was seen down, whether a stranger could see the offline + edits at the server (it shouldn't), whether the two tabs converged while + offline over a local channel, and the reconnect → converged time only where + they hadn't. - `churn.spec` — close the tab that booted everything, check the rest still sync. - `storage-adapter.spec` — in `pertab`, IndexedDB worker vs in-thread: create 40 docs × 20 edits and flush, edit → flushed latency, cold load of the 40 @@ -63,4 +114,8 @@ server holds our heads" means that peer and not a sibling. Cross-tab timings use epoch milliseconds, since `performance.now()` counts from each page's own navigation start. `find()` in the helpers retries on -"unavailable" and reports how many tries it took. +"unavailable" (with a real re-sync each time) and reports how many tries it +took. + +Each run truncates `bench-results/results.jsonl`; `BENCH_APPEND=1` keeps the +earlier rows, for re-running one spec or one mode into the same table. diff --git a/sites/bench/bench.config.ts b/sites/bench/bench.config.ts index c782f4502..20319662b 100644 --- a/sites/bench/bench.config.ts +++ b/sites/bench/bench.config.ts @@ -27,7 +27,8 @@ export default defineConfig({ command: "pnpm preview", url: `http://localhost:${PORT}`, timeout: 60_000, - reuseExistingServer: true, + // A preview left behind by another checkout would serve that build. + reuseExistingServer: false, env: { PORT: String(PORT) }, }, }); diff --git a/sites/bench/src/main.ts b/sites/bench/src/main.ts index 858458c2c..a5b2e1c04 100644 --- a/sites/bench/src/main.ts +++ b/sites/bench/src/main.ts @@ -4,7 +4,7 @@ // // ?mode=patchwork what patchwork does: createRepo() — the tab's own // subduction node with this origin's IndexedDB, its own -// server socket and the siblings BroadcastChannel — plus +// server socket and `headsChannel` — plus // the automerge worker that resolves URLs for the // service worker // ?mode=pertab the bare node: storage + socket, no siblings channel, @@ -12,6 +12,16 @@ // the database) // ?mode=pertab-bc pertab plus classic automerge sync between tabs over a // BroadcastChannel +// ?mode=pertab-mesh pertab plus patchwork's old siblings mesh (subduction over +// a BroadcastChannel), but each tab signing as itself: +// the old topology minus the shared signer, the +// service worker and the automerge worker +// ?mode=pertab-heads pertab with the origin-wide signer, so every tab is +// the same subduction peer on its own socket, and the +// Repo's `headsChannel`: a tab announces the heads it +// just persisted on a BroadcastChannel and the others +// reload the doc from the shared IndexedDB. No mesh, +// no classic network // ?mode=tab-worker the node in a dedicated Worker the tab spawns: storage, // socket and a mesh to the other tabs' workers live // there; the tab is a storageless Repo on a MessagePort @@ -20,14 +30,22 @@ // // `?server=none` runs everything but patchwork with no socket, so tabs can only // meet through storage (and whatever local channel the mode has). -// `?storage=direct` swaps the IndexedDB worker adapter for in-thread IndexedDB -// in the bare per-tab modes. `?serverPeer=` names the sync server's subduction -// peer id; without it the server is whichever peer isn't this tab. +// `?storage=worker` swaps in-thread IndexedDB for the IndexedDB worker adapter +// in the bare per-tab modes; `?mesh=1` gives them patchwork's old siblings mesh +// (subduction over a BroadcastChannel) and `?signer=shared` patchwork's +// origin-wide signer, so the shipped topology can be taken apart one piece at +// a time. `?serverPeer=` names the sync server's subduction peer id; without +// it the server is whichever peer isn't this tab. import { + parseAutomergeUrl, Repo, + WebSocketTransport, type AutomergeUrl, type DocHandle, + type ManagedTransport, type PeerId, + type RepoConfig, + type WebSocketEndpointInterface, } from "@automerge/automerge-repo/slim"; import { IndexedDBStorageAdapter } from "@automerge/automerge-repo-storage-indexeddb"; import { IndexedDBWorkerStorageAdapter } from "@automerge/automerge-repo-storage-indexeddb/IndexedDBWorkerStorageAdapter"; @@ -35,6 +53,7 @@ import { BroadcastChannelNetworkAdapter } from "@automerge/automerge-repo-networ import { MemorySigner } from "@automerge/automerge-subduction/slim"; import { createRepo, initWasm } from "@inkandswitch/patchwork"; import setupServiceWorker from "@inkandswitch/patchwork-bootloader"; +import { loadOrCreateSigner } from "@inkandswitch/patchwork-bootloader/signer"; import { WorkerSubductionEndpoint } from "./worker-link.js"; import type { ControlPort, NodeMessage, TabMessage } from "./protocol.js"; @@ -44,16 +63,20 @@ type Mode = | "patchwork" | "pertab" | "pertab-bc" + | "pertab-mesh" + | "pertab-heads" | "tab-worker" | "shared-worker"; type Storage = "worker" | "direct"; const params = new URLSearchParams(location.search); const mode = (params.get("mode") ?? "patchwork") as Mode; -const storage = (params.get("storage") ?? "worker") as Storage; +const storage = (params.get("storage") ?? "direct") as Storage; const serverUrl = params.get("server") ?? __SYNC_SERVER__.url; const serverPeer = params.get("serverPeer") ?? undefined; +const stop = params.get("stop"); const workerMode = mode === "tab-worker" || mode === "shared-worker"; +const bareTabMode = !workerMode && mode !== "patchwork"; const marks: Record = {}; const mark = (name: string) => (marks[name] ??= performance.now()); @@ -65,15 +88,51 @@ const online = Promise.withResolvers(); let isOnline: () => Promise = async () => false; let setOffline: (offline: boolean) => void = () => {}; const workerErrors: string[] = []; +const workerLog: string[] = []; +const remoteHeadsSeen: unknown[] = []; +const siblings = params.has("siblings") + ? params.get("siblings") === "1" + : mode === "tab-worker"; +const mesh = params.has("mesh") + ? params.get("mesh") === "1" + : mode === "pertab-mesh"; +const sharedSigner = params.get("signer") === "shared" || mode === "pertab-heads"; + +let serverBytes = 0; +let storageWrites = 0; +let syncRounds = 0; function sameHeads(a: string[], b: string[]): boolean { return a.length === b.length && a.every((head) => b.includes(head)); } function storageAdapter() { - return storage === "direct" - ? new IndexedDBStorageAdapter() - : new IndexedDBWorkerStorageAdapter(); + const adapter = + storage === "direct" + ? new IndexedDBStorageAdapter() + : new IndexedDBWorkerStorageAdapter(); + const saveBatch = adapter.saveBatch.bind(adapter); + adapter.saveBatch = (entries) => { + storageWrites++; + return saveBatch(entries); + }; + return adapter; +} + +function countingEndpoint(url: string): WebSocketEndpointInterface { + return { + url, + async connect(): Promise { + const transport = await WebSocketTransport.connect(url); + const recvBytes = transport.recvBytes.bind(transport); + transport.recvBytes = async () => { + const bytes = await recvBytes(); + serverBytes += bytes.byteLength; + return bytes; + }; + return transport; + }, + }; } // ── Tab modes: the node is this Repo ──────────────────────────────────── @@ -89,25 +148,46 @@ async function buildTabNode(): Promise { repo = tab.repo; ownPeerId = tab.signerIdentity!.peerId; } else { - const signer = new MemorySigner(); + const storage = storageAdapter(); + const signer = sharedSigner + ? await loadOrCreateSigner(storage) + : new MemorySigner(); ownPeerId = signer.peerId().toString(); - repo = new Repo({ + const config: RepoConfig = { signer, - storage: storageAdapter(), + storage, peerId: `bench-tab-${crypto.randomUUID()}` as PeerId, - subductionWebsocketEndpoints: serverUrl === "none" ? [] : [serverUrl], + subductionWebsocketEndpoints: + serverUrl === "none" ? [] : [countingEndpoint(serverUrl)], network: mode === "pertab-bc" ? [new BroadcastChannelNetworkAdapter({ channelName: "bench" })] : [], + subductionAdapters: mesh + ? [ + { + adapter: new BroadcastChannelNetworkAdapter({ + channelName: "bench-mesh", + }), + serviceName: "bench-mesh", + role: "mesh", + }, + ] + : [], enableRemoteHeadsGossiping: true, - }); + headsChannel: mode === "pertab-heads" ? "bench-heads" : undefined, + }; + repo = new Repo(config); } mark("repo"); // Every other node on this origin (siblings, the automerge worker) signs as // this tab does in patchwork mode, and doesn't exist in the bare modes, so - // without a known server peer the server is whoever isn't us. + // without a known server peer the server is whoever isn't us. A mesh with + // distinct signers has no such tell, so it insists on the probe. + if (!serverPeer && serverUrl !== "none" && mesh) { + throw new Error("?serverPeer= is required for a mesh of distinct signers"); + } const isServer = (id: string) => serverPeer ? id === serverPeer : id !== ownPeerId; const connectedServers = async () => @@ -142,24 +222,33 @@ async function buildTabNode(): Promise { // ── Worker modes: the node is in a worker, this Repo is storageless ───── -function spawnControl(): ControlPort { +const PORT_TIMEOUT_MS = 10_000; + +function spawnControl(): { control: ControlPort; errors: EventTarget } { if (mode === "tab-worker") { - return new Worker(new URL("./subduction-worker.ts", import.meta.url), { - type: "module", - }); + const worker = new Worker( + new URL("./subduction-worker.ts", import.meta.url), + { type: "module" } + ); + return { control: worker, errors: worker }; } const shared = new SharedWorker( new URL("./subduction-shared-worker.ts", import.meta.url), { type: "module", name: "bench-subduction" } ); - return shared.port; + return { control: shared.port, errors: shared }; } async function buildWorkerNode(): Promise { - const control = spawnControl(); + const { control, errors } = spawnControl(); const send = (message: TabMessage, transfer?: Transferable[]) => control.postMessage(message, transfer); let connected = false; + errors.addEventListener("error", (event) => { + const message = (event as ErrorEvent).message ?? "worker failed to load"; + workerErrors.push(message); + console.error("[subduction worker]", message); + }); control.addEventListener("message", (event) => { const message = event.data as NodeMessage; @@ -169,6 +258,8 @@ async function buildWorkerNode(): Promise { if (connected) online.resolve(performance.now()); return; case "remote-heads": + remoteHeadsSeen.push(message); + if (message.storageId !== serverPeer) return; serverHeads.set(message.documentId, message.heads); return; case "ready": @@ -179,23 +270,33 @@ async function buildWorkerNode(): Promise { workerErrors.push(message.message); console.error("[subduction worker]", message.message); return; + case "log": + workerLog.push(message.message); + return; } }); control.start?.(); send({ type: "config", - config: { server: serverUrl, serverPeer, siblings: mode === "tab-worker" }, + config: { server: serverUrl, serverPeer, siblings }, }); send({ type: "status" }); + // Bounded so a worker that never answers feeds automerge-repo's reconnect + // loop instead of hanging it. let nextPortId = 0; const openPort = () => new Promise((resolve, reject) => { const id = ++nextPortId; const { port1, port2 } = new MessageChannel(); + const timer = setTimeout(() => { + control.removeEventListener("message", listener); + reject(new Error(`worker didn't accept a port within ${PORT_TIMEOUT_MS}ms`)); + }, PORT_TIMEOUT_MS); const listener = (event: MessageEvent) => { const message = event.data as NodeMessage; if (!("id" in message) || message.id !== id) return; + clearTimeout(timer); control.removeEventListener("message", listener); if (message.type === "port-ready") resolve(port1); else reject(new Error(`worker refused the port: ${message.error}`)); @@ -224,36 +325,56 @@ async function buildWorkerNode(): Promise { return repo; } -async function build(): Promise { +async function build(): Promise { mark("start"); + if (stop === "js") return null; await initWasm(); mark("wasm"); - return workerMode ? buildWorkerNode() : buildTabNode(); + if (stop === "wasm") return null; + const repo = await (workerMode ? buildWorkerNode() : buildTabNode()); + const subduction = await repo.subduction; + const syncWithAllPeers = subduction.syncWithAllPeers.bind(subduction); + subduction.syncWithAllPeers = (...args) => { + syncRounds++; + return syncWithAllPeers(...args); + }; + return repo; } // `find` settles as unavailable when every source has said no, and a sibling -// tab's brand-new doc may not have reached those sources yet. Retrying is what -// an app would have to do; the attempt count is reported so the benches can -// say how often it was needed. +// tab's brand-new doc may not have reached those sources yet. A plain second +// `find` returns the same settled query, so each retry asks the Repo for a +// real re-sync first, which is what an app would have to do; the attempt count +// is reported so the benches can say how often it was needed. The deadline +// also bounds a find that stays pending. async function find( url: string, timeoutMs = 30_000 ): Promise<{ handle: DocHandle>; attempts: number }> { const deadline = performance.now() + timeoutMs; + const { documentId } = parseAutomergeUrl(url as AutomergeUrl); for (let attempts = 1; ; attempts++) { + const remaining = deadline - performance.now(); + if (remaining <= 0) { + throw new Error(`${url} not found within ${timeoutMs}ms (${attempts - 1} retries)`); + } try { const handle = await window.repo.find>( - url as AutomergeUrl + url as AutomergeUrl, + { signal: AbortSignal.timeout(remaining) } ); await handle.whenReady(); return { handle, attempts }; } catch (error) { if (performance.now() > deadline) throw error; + window.repo.resyncSubduction(documentId); await sleep(50); } } } +const ONLINE_TIMEOUT_MS = 60_000; + // How long the main thread was unavailable: a short timer's lateness, plus // whatever the browser reports as long tasks. type Stall = { @@ -273,12 +394,23 @@ window.bench = { storage, marks, find, - online: () => online.promise, + online: (timeoutMs = ONLINE_TIMEOUT_MS) => + Promise.race([ + online.promise, + sleep(timeoutMs).then(() => { + throw new Error(`server not connected within ${timeoutMs}ms`); + }), + ]), isOnline: () => isOnline(), setOffline: (offline) => setOffline(offline), serverPeerIds: () => [...serverPeerIds], serverHeads: (documentId) => serverHeads.get(documentId), workerErrors: () => [...workerErrors], + workerLog: () => [...workerLog], + remoteHeadsSeen: () => [...remoteHeadsSeen], + serverBytes: () => (bareTabMode && serverUrl !== "none" ? serverBytes : null), + storageWrites: () => (bareTabMode ? storageWrites : null), + syncRounds: () => syncRounds, // Resolves with the epoch time the server was seen holding exactly the heads // the document has right now. Polled: a few ms of slop is fine here. async serverConfirmed(url, timeoutMs = 30_000) { @@ -327,8 +459,12 @@ window.bench = { stall = state; }, stallStop() { - if (!stall) throw new Error("stall probe not running"); + if (!stall) return { maxMs: 0, totalMs: 0, longTasks: 0, longTaskMs: 0 }; clearInterval(stall.timer); + for (const entry of stall.observer?.takeRecords() ?? []) { + stall.longTasks++; + stall.longTaskMs += entry.duration; + } stall.observer?.disconnect(); const { maxMs, totalMs, longTasks, longTaskMs } = stall; stall = null; @@ -336,9 +472,11 @@ window.bench = { }, }; -window.repo = await build(); +const repo = await build(); +if (repo) window.repo = repo; mark("ready"); document.body.textContent = `${mode}: ready in ${Math.round(marks.ready - marks.start)}ms`; +if (params.has("ui")) await import("./playground.js"); declare global { interface Window { @@ -348,12 +486,17 @@ declare global { storage: Storage; marks: Record; find: typeof find; - online: () => Promise; + online: (timeoutMs?: number) => Promise; isOnline: () => Promise; setOffline: (offline: boolean) => void; serverPeerIds: () => string[]; serverHeads: (documentId: string) => string[] | undefined; workerErrors: () => string[]; + workerLog: () => string[]; + remoteHeadsSeen: () => unknown[]; + serverBytes: () => number | null; + storageWrites: () => number | null; + syncRounds: () => number; serverConfirmed: (url: string, timeoutMs?: number) => Promise; stallStart: () => void; stallStop: () => { diff --git a/sites/bench/src/node.ts b/sites/bench/src/node.ts index 13b189319..03dca8d6f 100644 --- a/sites/bench/src/node.ts +++ b/sites/bench/src/node.ts @@ -3,14 +3,19 @@ // other nodes on the origin. It never opens a document itself: tabs are // storageless Repos that reach it over MessagePorts and sync through it. import { + documentIdToBinary, initializeWasm, Repo, WebSocketTransport, + type DocumentId, type ManagedTransport, type PeerId, type WebSocketEndpointInterface, } from "@automerge/automerge-repo/slim"; -import { MemorySigner } from "@automerge/automerge-subduction/slim"; +import { + MemorySigner, + SedimentreeId, +} from "@automerge/automerge-subduction/slim"; // eslint-disable-next-line // @ts-ignore — initSync is a wasm-bindgen runtime helper not in the .d.ts import { initSync as initSubductionSync } from "@automerge/automerge-subduction/slim"; @@ -26,15 +31,24 @@ import type { const sleep = (ms: number) => new Promise((resolve) => setTimeout(resolve, ms)); +const RELAY_TIMEOUT_MS = 30_000; + +function toSedimentreeId(documentId: string): SedimentreeId { + const binary = documentIdToBinary(documentId as DocumentId); + if (!binary) throw new Error(`not a document id: ${documentId}`); + const bytes = new Uint8Array(32); + bytes.set(binary); + return SedimentreeId.fromBytes(bytes); +} + /** * The server socket, observable and switchable. automerge-repo's reconnect - * loop calls `connect()`; while offline it waits there, so nothing reconnects - * until told to. + * loop calls `connect()`; while offline that fails like a dead network would, + * so the loop backs off exactly as a tab's does. */ class ServerEndpoint implements WebSocketEndpointInterface { #live: WebSocketTransport | null = null; #offline = false; - #wake: (() => void) | null = null; constructor( readonly url: string, @@ -42,10 +56,12 @@ class ServerEndpoint implements WebSocketEndpointInterface { ) {} async connect(): Promise { - while (this.#offline) { - await new Promise((resolve) => (this.#wake = resolve)); - } + if (this.#offline) throw new Error("offline"); const transport = await WebSocketTransport.connect(this.url); + if (this.#offline) { + void transport.disconnect(); + throw new Error("offline"); + } this.#live = transport; this.events.onOpen(); void transport.closed().then(() => { @@ -58,10 +74,6 @@ class ServerEndpoint implements WebSocketEndpointInterface { setOffline(offline: boolean): void { this.#offline = offline; if (offline) void this.#live?.disconnect(); - else { - this.#wake?.(); - this.#wake = null; - } } } @@ -70,22 +82,27 @@ type Node = { accept(port: MessagePort): void; setOffline(offline: boolean): void; connected(): boolean; - serverHeads(): Map; + remoteHeads(): Map>; }; async function build( config: NodeConfig, broadcast: (message: NodeMessage) => void ): Promise { + if (config.server !== "none" && !config.serverPeer) { + throw new Error( + "a worker node needs the server's peer id (?serverPeer=) to tell it from the tabs" + ); + } const [automergeWasm, subductionWasm] = await Promise.all([ fetch("/automerge.wasm").then((r) => r.bytes()), fetch("/subduction.wasm").then((r) => r.bytes()), ]); await initializeWasm(automergeWasm); - initSubductionSync(subductionWasm); + initSubductionSync({ module: subductionWasm }); const signer = new MemorySigner(); - const serverHeads = new Map(); + const remoteHeads = new Map>(); let socketOpen = false; let connected = false; const setConnected = (value: boolean) => { @@ -93,16 +110,20 @@ async function build( connected = value; broadcast({ type: "connection", connected }); }; + const log = (message: string) => + broadcast({ type: "log", message: `${Math.round(performance.now())}ms ${message}` }); const endpoint = config.server === "none" ? null : new ServerEndpoint(config.server, { onOpen() { + log("server socket open"); socketOpen = true; void awaitHandshake(); }, onClosed() { + log("server socket closed"); socketOpen = false; setConnected(false); }, @@ -127,10 +148,42 @@ async function build( enableRemoteHeadsGossiping: true, }); + // A node that never opens a document has nothing driving sync rounds: it + // stores what a tab pushes and answers what a peer asks, but doesn't carry + // one peer's commits to another on its own. So when a tab (or a mesh peer) + // announces heads, run a round for that document with every peer, which + // pushes the commits to the server and to the other tabs and subscribes to + // updates. One round in flight per document; announcements during it run + // one more. + const relaying = new Map(); + async function relay(documentId: string) { + if (relaying.has(documentId)) { + relaying.set(documentId, true); + return; + } + const subduction = await repo.subduction; + do { + relaying.set(documentId, false); + try { + await subduction.syncWithAllPeers( + toSedimentreeId(documentId), + true, + RELAY_TIMEOUT_MS + ); + } catch (error) { + log(`relay of ${documentId.slice(0, 8)} failed: ${error}`); + } + } while (relaying.get(documentId)); + relaying.delete(documentId); + } + + // Every peer's heads go to the tabs, which know which peer is the server. repo.on("subduction-remote-heads", ({ documentId, storageId, heads }) => { - if (storageId !== config.serverPeer) return; - serverHeads.set(documentId, [...heads]); - broadcast({ type: "remote-heads", documentId, heads: [...heads] }); + let byPeer = remoteHeads.get(documentId); + if (!byPeer) remoteHeads.set(documentId, (byPeer = new Map())); + byPeer.set(storageId, [...heads]); + broadcast({ type: "remote-heads", documentId, storageId, heads: [...heads] }); + if (storageId !== config.serverPeer) void relay(documentId); }); // The socket being open is not the handshake being done; the server counts @@ -139,7 +192,10 @@ async function build( while (socketOpen && !connected) { const peers = await repo.connectedSubductionPeerIds(); if (config.serverPeer && peers.includes(config.serverPeer)) { + log(`server handshake done; peers: ${peers.length}`); setConnected(true); + // Whatever the tabs did while the server was away goes up now. + for (const documentId of remoteHeads.keys()) void relay(documentId); return; } await sleep(10); @@ -160,13 +216,15 @@ async function build( new MessagePortTransport(port), WORKER_SUBDUCTION_SERVICE ) - .catch((error) => - broadcast({ type: "error", message: `accept failed: ${error}` }) + .then( + (peerId) => log(`accepted tab ${peerId.toString().slice(0, 8)}`), + (error) => + broadcast({ type: "error", message: `accept failed: ${error}` }) ); }, setOffline: (offline) => endpoint?.setOffline(offline), connected: () => connected, - serverHeads: () => serverHeads, + remoteHeads: () => remoteHeads, }; } @@ -220,8 +278,10 @@ export function startNode(): { attach(port: ControlPort): void } { case "status": { const built = await node!; post(port, { type: "connection", connected: built.connected() }); - for (const [documentId, heads] of built.serverHeads()) { - post(port, { type: "remote-heads", documentId, heads }); + for (const [documentId, byPeer] of built.remoteHeads()) { + for (const [storageId, heads] of byPeer) { + post(port, { type: "remote-heads", documentId, storageId, heads }); + } } return; } diff --git a/sites/bench/src/playground.ts b/sites/bench/src/playground.ts new file mode 100644 index 000000000..6d39f60e2 --- /dev/null +++ b/sites/bench/src/playground.ts @@ -0,0 +1,147 @@ +import type { AutomergeUrl, DocHandle } from "@automerge/automerge-repo/slim"; + +type Item = { from: string; text: string; at: number }; +type Doc = { items: Item[] }; + +const params = new URLSearchParams(location.search); +const tabName = + params.get("name") ?? + `tab-${Math.random().toString(36).slice(2, 6)}`; + +const css = ` + body { margin: 0; font: 13px/1.5 ui-monospace, SFMono-Regular, Menlo, monospace; + background: #111; color: #eee; } + header { display: flex; gap: 12px; align-items: baseline; flex-wrap: wrap; + padding: 10px 14px; border-bottom: 1px solid #333; } + header b { font-weight: 600; } + .dot { width: 9px; height: 9px; border-radius: 50%; display: inline-block; } + .on { background: #4ade80; } .off { background: #f87171; } + main { display: grid; grid-template-columns: 1fr 320px; gap: 0; height: calc(100vh - 43px); } + #items { overflow: auto; padding: 14px; margin: 0; list-style: none; } + #items li { padding: 2px 0; } + #items .me { color: #93c5fd; } + aside { border-left: 1px solid #333; padding: 14px; overflow: auto; } + aside dl { margin: 0; display: grid; grid-template-columns: max-content 1fr; gap: 2px 10px; } + aside dt { color: #888; } aside dd { margin: 0; word-break: break-all; } + form { display: flex; gap: 8px; padding: 10px 14px; border-top: 1px solid #333; + grid-column: 1 / -1; } + input[type=text] { flex: 1; background: #1c1c1c; border: 1px solid #333; color: #eee; + padding: 6px 8px; font: inherit; } + button { background: #222; border: 1px solid #444; color: #eee; padding: 6px 12px; + font: inherit; cursor: pointer; } + button:hover { background: #2a2a2a; } + a { color: #93c5fd; } +`; + +async function start() { + const bench = window.bench; + const repo = window.repo; + if (!repo) { + document.body.textContent = `${bench.mode}: no repo in this mode`; + return; + } + + let url = params.get("doc") as AutomergeUrl | null; + if (!url) { + const created = repo.create({ items: [] }); + url = created.url; + params.set("doc", url); + history.replaceState(null, "", `?${params}`); + } + + const { handle } = (await bench.find(url)) as unknown as { + handle: DocHandle; + }; + await handle.whenReady(); + + document.head.append(Object.assign(document.createElement("style"), { textContent: css })); + document.body.innerHTML = ` +
+ ${bench.mode} + ${tabName} + … + + open another tab +
+
+
    + +
    + + +
    +
    + `; + + const $ = (id: string) => document.getElementById(id)!; + ($("another") as HTMLAnchorElement).href = `?${params}&name=tab-${Math.random() + .toString(36) + .slice(2, 6)}`; + $("docid").textContent = handle.documentId; + + $("say").addEventListener("submit", (event) => { + event.preventDefault(); + const input = $("text") as HTMLInputElement; + const text = input.value.trim(); + if (!text) return; + input.value = ""; + handle.change((doc) => { + doc.items.push({ from: tabName, text, at: Date.now() }); + }); + }); + + let offline = false; + $("cut").addEventListener("click", () => { + offline = !offline; + bench.setOffline(offline); + $("cut").textContent = offline ? "go online" : "go offline"; + }); + + const renderItems = () => { + const items = handle.doc()?.items ?? []; + $("items").innerHTML = items + .map( + (item) => + `
  • ${item.from}: ${item.text + .replace(/` + ) + .join(""); + $("count").textContent = String(items.length); + }; + handle.on("change", renderItems); + renderItems(); + + const short = (hashes: readonly string[] | undefined) => + hashes?.length ? hashes.map((h) => h.slice(0, 6)).join(" ") : "—"; + + setInterval(async () => { + const heads = handle.heads() ?? []; + const server = bench.serverHeads(handle.documentId); + $("heads").textContent = short(heads); + $("sheads").textContent = short(server); + $("instep").textContent = server + ? heads.every((h) => server.includes(h)) + ? "yes" + : "no" + : "unknown"; + $("rounds").textContent = String(bench.syncRounds()); + $("writes").textContent = String(bench.storageWrites() ?? "—"); + $("bytes").textContent = String(bench.serverBytes() ?? "—"); + const up = await bench.isOnline(); + $("dot").className = `dot ${up ? "on" : "off"}`; + $("net").textContent = up ? "connected" : "offline"; + }, 300); +} + +void start(); diff --git a/sites/bench/src/protocol.ts b/sites/bench/src/protocol.ts index 198dd4d29..05a895be3 100644 --- a/sites/bench/src/protocol.ts +++ b/sites/bench/src/protocol.ts @@ -23,8 +23,9 @@ export type NodeMessage = | { type: "port-ready"; id: number } | { type: "port-failed"; id: number; error: string } | { type: "connection"; connected: boolean } - | { type: "remote-heads"; documentId: string; heads: string[] } - | { type: "error"; message: string }; + | { type: "remote-heads"; documentId: string; storageId: string; heads: string[] } + | { type: "error"; message: string } + | { type: "log"; message: string }; export type ControlPort = { postMessage(message: unknown, transfer?: Transferable[]): void; diff --git a/sites/bench/tests/bench.ts b/sites/bench/tests/bench.ts index d05a6a76f..df7cd15fb 100644 --- a/sites/bench/tests/bench.ts +++ b/sites/bench/tests/bench.ts @@ -6,12 +6,16 @@ export type Mode = | "patchwork" | "pertab" | "pertab-bc" + | "pertab-mesh" + | "pertab-heads" | "tab-worker" | "shared-worker"; export const MODES: Mode[] = [ "patchwork", "pertab", "pertab-bc", + "pertab-mesh", + "pertab-heads", "tab-worker", "shared-worker", ]; @@ -28,7 +32,8 @@ export type Result = { /** Set for the storage-adapter comparison, which is its own table. */ storage?: Storage; tabs?: number; - value: number | boolean; + /** null: no measurement, printed as – */ + value: number | boolean | null; unit: "ms" | "MB" | "n" | "ok"; }; @@ -49,7 +54,7 @@ export function median(values: number[]): number { // worker-hosted node can't tell the server from the tabs it accepts any other // way; the tab modes can, but get it too so every mode measures the same peer. let serverPeer: Promise | undefined; -function probeServerPeer(browser: Browser): Promise { +export function probeServerPeer(browser: Browser): Promise { return (serverPeer ??= (async () => { const context = await browser.newContext(); try { @@ -100,21 +105,72 @@ export function isOnline(page: Page): Promise { return page.evaluate(() => window.bench.isOnline()); } +export type OfflineSwitch = { + /** Cut or restore the link to the sync server in every tab. */ + set(pages: Page[], offline: boolean): Promise; +}; + /** - * Cut the link to the sync server. Playwright's offline emulation reaches a - * page's own socket but not a worker's (it skips worker sessions), so the - * worker modes are also told to drop theirs and hold off reconnecting. + * A real offline switch. Playwright's `setOffline` makes new connections fail + * the way a dead network does but doesn't close a WebSocket that is already + * open, so every WebSocket a page opens is also proxied here, purely so the + * live ones can be closed. (Refusing inside the proxy instead would close the + * page's socket without an `error` event, which automerge-repo's connect never + * recovers from — a shape no real network produces.) A worker's socket is out + * of reach of both, so the worker modes are told to drop theirs and fail + * reconnects while offline. Install before opening any tab. */ -export async function setOffline( - context: BrowserContext, - pages: Page[], - offline: boolean -): Promise { - await context.setOffline(offline); - await Promise.all( - pages.map((page) => - page.evaluate((offline) => window.bench.setOffline(offline), offline) - ) +export async function offlineSwitch( + context: BrowserContext +): Promise { + const live = new Set<{ close(): void }>(); + await context.routeWebSocket( + (url) => url.protocol === "wss:" || url.protocol === "ws:", + (ws) => { + const server = ws.connectToServer(); + const link = { + close() { + live.delete(link); + server.close({ code: 1012, reason: "offline" }); + ws.close({ code: 1012, reason: "offline" }); + }, + }; + live.add(link); + ws.onMessage((message) => server.send(message)); + server.onMessage((message) => ws.send(message)); + ws.onClose((code, reason) => { + live.delete(link); + server.close({ code, reason }); + }); + server.onClose((code, reason) => { + live.delete(link); + ws.close({ code, reason }); + }); + } + ); + return { + async set(pages, offline) { + await context.setOffline(offline); + if (offline) for (const link of [...live]) link.close(); + await Promise.all( + pages.map((page) => + page.evaluate((offline) => window.bench.setOffline(offline), offline) + ) + ); + }, + }; +} + +/** + * No sync server at all, for a build whose server url is baked in: every + * WebSocket a page opens is closed unopened. The connect attempt hangs rather + * than errors (see offlineSwitch), which here is the point — nothing but + * storage can answer. + */ +export function refuseWebSockets(context: BrowserContext): Promise { + return context.routeWebSocket( + (url) => url.protocol === "wss:" || url.protocol === "ws:", + (ws) => ws.close({ code: 1012, reason: "no server" }) ); } @@ -146,13 +202,17 @@ export function createDoc(page: Page, value: object): Promise { */ export function timeFind( page: Page, - url: string + url: string, + timeoutMs = 30_000 ): Promise<{ ms: number; attempts: number }> { - return page.evaluate(async (url) => { - const started = performance.now(); - const { attempts } = await window.bench.find(url); - return { ms: performance.now() - started, attempts }; - }, url); + return page.evaluate( + async ([url, timeoutMs]) => { + const started = performance.now(); + const { attempts } = await window.bench.find(url, timeoutMs); + return { ms: performance.now() - started, attempts }; + }, + [url, timeoutMs] as const + ); } // Cross-page timings use epoch ms: performance.now() counts from each page's @@ -209,8 +269,15 @@ export function awaitField( ); } -export function serverConfirmed(page: Page, url: string): Promise { - return page.evaluate((url) => window.bench.serverConfirmed(url), url); +export function serverConfirmed( + page: Page, + url: string, + timeoutMs = 30_000 +): Promise { + return page.evaluate( + ([url, timeoutMs]) => window.bench.serverConfirmed(url, timeoutMs), + [url, timeoutMs] as const + ); } export function getField(page: Page, url: string, field: string): Promise { @@ -223,6 +290,37 @@ export function getField(page: Page, url: string, field: string): Promise ); } +/** + * The doc as a client that shares nothing with these tabs sees it — a fresh + * browser context, bare per-tab mode, only the server to ask. undefined if it + * can't get the doc at all. + */ +export async function strangerSees( + browser: Browser, + url: string, + timeoutMs = 10_000 +): Promise | undefined> { + const context = await browser.newContext(); + try { + const page = await openTab(context, "pertab"); + await online(page); + return await page + .evaluate( + async ([url, timeoutMs]) => { + const { handle } = await window.bench.find(url, timeoutMs); + return JSON.parse(JSON.stringify(handle.doc())) as Record< + string, + unknown + >; + }, + [url, timeoutMs] as const + ) + .catch(() => undefined); + } finally { + await context.close(); + } +} + export function flush(page: Page): Promise { return page.evaluate(() => window.repo.flush()); } @@ -289,3 +387,31 @@ export async function rendererMemory( .reduce((sum, line) => sum + Number(line.trim()), 0); return { mb: Math.round(kb / 1024), processes: pids.length }; } + +export async function rendererProcesses( + browser: Browser +): Promise> { + const session = await browser.newBrowserCDPSession(); + const { processInfo } = (await session.send("SystemInfo.getProcessInfo")) as { + processInfo: Array<{ type: string; id: number }>; + }; + await session.detach(); + const pids = processInfo + .filter((process) => process.type === "renderer") + .map((process) => process.id); + return pids.map((pid) => { + if (process.platform === "darwin") { + const out = execFileSync( + "footprint", + ["-f", "bytes", "-p", String(pid)], + { encoding: "utf8" } + ); + const bytes = Number(out.match(/phys_footprint:\s+(\d+)/)?.[1] ?? 0); + return { pid, mb: Math.round(bytes / 104857.6) / 10 }; + } + const rss = execFileSync("ps", ["-o", "rss=", "-p", String(pid)], { + encoding: "utf8", + }); + return { pid, mb: Math.round(Number(rss.trim()) / 102.4) / 10 }; + }); +} diff --git a/sites/bench/tests/boot.spec.ts b/sites/bench/tests/boot.spec.ts index 38bff701e..4efff3101 100644 --- a/sites/bench/tests/boot.spec.ts +++ b/sites/bench/tests/boot.spec.ts @@ -47,17 +47,20 @@ for (const mode of MODES) { await Promise.all(pages.map((page) => online(page))); if (WORKER_MODES.includes(mode)) { - // The tab's own link is to its worker; set once the handshake is done. - const linked = (await marks(pages[0])).linked; - if (linked !== undefined) { - record({ - metric: "boot → linked to worker, first tab", - mode, - tabs, - value: linked, - unit: "ms", - }); - } + // The tab's own link is to its worker, handshaken independently of the + // worker's server link, so it may land after `online`. + await pages[0] + .waitForFunction(() => window.bench.marks.linked !== undefined, null, { + timeout: 10_000, + }) + .catch(() => {}); + record({ + metric: "boot → linked to worker, first tab", + mode, + tabs, + value: (await marks(pages[0])).linked ?? null, + unit: "ms", + }); } // Let storage flushes and the first sync rounds settle first. await pages[0].waitForTimeout(2_000); diff --git a/sites/bench/tests/churn.spec.ts b/sites/bench/tests/churn.spec.ts index 54f0b8b62..4cd64e42f 100644 --- a/sites/bench/tests/churn.spec.ts +++ b/sites/bench/tests/churn.spec.ts @@ -1,4 +1,4 @@ -import { expect, test } from "@playwright/test"; +import { test } from "@playwright/test"; import { MODES, awaitField, @@ -11,8 +11,9 @@ import { } from "./bench.js"; // Close the tab that booted everything and check the survivors still sync. -// In patchwork mode that tab spawned the automerge worker; in every mode it -// owned a storage worker mid-write. +// In patchwork mode that tab spawned the automerge worker; in tab-worker mode +// its subduction node dies with it; in shared-worker mode the node it spawned +// outlives it. for (const mode of MODES) { test(`${mode}: closing the first tab doesn't strand the rest`, async ({ context, @@ -22,6 +23,7 @@ for (const mode of MODES) { const c = await openTab(context, mode); await Promise.all([online(first), online(b), online(c)]); const url = await createDoc(first, { n: 0 }); + await serverConfirmed(first, url); await awaitField(b, url, "n", 0); await awaitField(c, url, "n", 0); @@ -46,6 +48,5 @@ for (const mode of MODES) { unit: "ms", }); } - expect(ok).toBe(true); }); } diff --git a/sites/bench/tests/debug.spec.ts b/sites/bench/tests/debug.spec.ts deleted file mode 100644 index 2ea2c0888..000000000 --- a/sites/bench/tests/debug.spec.ts +++ /dev/null @@ -1,49 +0,0 @@ -import { test } from "@playwright/test"; -import { createDoc, flush, getField, openTab, setField, timeFind, online } from "./bench.js"; - -for (const mode of ["shared-worker", "tab-worker"] as const) { - test(`${mode}: debug push`, async ({ context }) => { - const a = await openTab(context, mode, { server: "none" }); - const b = await openTab(context, mode, { server: "none" }); - await Promise.all([online(a), online(b)]); - a.on("console", (m) => console.log("[a]", m.text())); - b.on("console", (m) => console.log("[b]", m.text())); - const url = await createDoc(a, { n: 0 }); - // B asks right away, with the helper's 30s retry cut to 3s. - const first = await b.evaluate(async (url) => { - const started = performance.now(); - try { - const { attempts } = await window.bench.find(url, 3_000); - return { ok: true, attempts, ms: performance.now() - started }; - } catch (e) { - return { ok: false, error: String(e), ms: performance.now() - started }; - } - }, url); - console.log("B find right after create:", JSON.stringify(first)); - await a.waitForTimeout(2_000); - // Same (cached) handle, later. - const second = await b.evaluate(async (url) => { - try { - const { attempts } = await window.bench.find(url, 3_000); - return { ok: true, attempts }; - } catch (e) { - return { ok: false, error: String(e) }; - } - }, url); - console.log("B find 2s later:", JSON.stringify(second)); - // A fresh tab now. - const c = await openTab(context, mode, { server: "none" }); - const fresh = await timeFind(c, url).then((r) => ({ ok: true, ...r }), (e) => ({ ok: false, error: String(e) })); - console.log("C (fresh tab) find:", JSON.stringify(fresh)); - // Now edits without flush, then with flush: does a fresh tab see them? - for (let i = 1; i <= 5; i++) await setField(a, url, "n", i); - await a.waitForTimeout(1_000); - const d = await openTab(context, mode, { server: "none" }); - console.log("D sees n after 5 edits + 1s:", await getField(d, url, "n").catch((e) => String(e))); - for (let i = 6; i <= 10; i++) await setField(a, url, "n", i); - await flush(a); - const e = await openTab(context, mode, { server: "none" }); - console.log("E sees n after 5 more edits + flush:", await getField(e, url, "n").catch((e) => String(e))); - console.log("worker errors A:", JSON.stringify(await a.evaluate(() => window.bench.workerErrors()))); - }); -} diff --git a/sites/bench/tests/evict.spec.ts b/sites/bench/tests/evict.spec.ts new file mode 100644 index 000000000..4cd9c2e4e --- /dev/null +++ b/sites/bench/tests/evict.spec.ts @@ -0,0 +1,142 @@ +import { expect, test, type Page } from "@playwright/test"; +import { awaitField, online, openTab, record, setField, type Mode } from "./bench.js"; + +const mode: Mode = "pertab-mesh"; +const TEXT_BYTES = 1 << 20; +const EDIT_TIMEOUT_MS = 5_000; + +async function jsAndWasmMB(page: Page): Promise { + const cdp = await page.context().newCDPSession(page); + await cdp.send("HeapProfiler.collectGarbage"); + await cdp.detach(); + return page.evaluate(async () => { + const { breakdown } = (await ( + performance as unknown as { + measureUserAgentSpecificMemory(): Promise<{ + breakdown: { bytes: number; types: string[] }[]; + }>; + } + ).measureUserAgentSpecificMemory()); + let bytes = 0; + for (const entry of breakdown) { + if ( + entry.types.includes("JavaScript") || + entry.types.includes("WebAssembly") + ) { + bytes += entry.bytes; + } + } + return bytes / 1_048_576; + }); +} + +test(`${mode}: evict a document, then find it again`, async ({ context }) => { + const a = await openTab(context, mode); + const b = await openTab(context, mode); + await Promise.all([online(a), online(b)]); + + const { url, documentId, heads } = await a.evaluate((bytes) => { + const handle = window.repo.create<{ text: string }>(); + handle.change((d) => { + d.text = "x".repeat(bytes); + }); + return { + url: handle.url, + documentId: handle.documentId, + heads: [...handle.heads()].sort(), + }; + }, TEXT_BYTES); + await a.evaluate(() => window.repo.flush()); + const bHasText = await b.evaluate( + async ([url, bytes]) => { + const { handle } = await window.bench.find(url); + return (handle.doc().text as string | undefined)?.length === bytes; + }, + [url, TEXT_BYTES] as const + ); + expect(bHasText).toBe(true); + + const before = await jsAndWasmMB(a); + const gone = await a.evaluate(async (documentId) => { + await window.repo.removeFromCache(documentId); + return !Object.keys(window.repo.handles).includes(documentId); + }, documentId); + record({ + metric: "evict → documentId gone from repo.handles", + mode, + value: gone, + unit: "ok", + }); + expect.soft(gone).toBe(true); + const after = await jsAndWasmMB(a); + record({ + metric: `evict → JS+wasm memory freed (${TEXT_BYTES >> 20} MB doc)`, + mode, + value: before - after, + unit: "MB", + }); + expect.soft(before - after).toBeGreaterThan(0); + + const refound = await a.evaluate( + async ([url, bytes, heads]) => { + const { handle } = await window.bench.find(url); + return { + sameText: + (handle.doc().text as string | undefined)?.length === bytes, + sameHeads: + JSON.stringify([...handle.heads()].sort()) === JSON.stringify(heads), + }; + }, + [url, TEXT_BYTES, heads] as const + ); + record({ + metric: "evict → re-find has the same text and heads", + mode, + value: refound.sameText && refound.sameHeads, + unit: "ok", + }); + expect.soft(refound.sameText && refound.sameHeads).toBe(true); + + const seen = awaitField(a, url, "counter", 1, EDIT_TIMEOUT_MS).then( + () => true, + () => false + ); + await setField(b, url, "counter", 1); + const siblingEditSeen = await seen; + record({ + metric: "evict → re-find sees a sibling's edit", + mode, + value: siblingEditSeen, + unit: "ok", + }); + expect.soft(siblingEditSeen).toBe(true); + + const reachedSibling = awaitField( + b, + url, + "marker", + "unflushed", + EDIT_TIMEOUT_MS + ).then( + () => true, + () => false + ); + await a.evaluate( + async ([url, documentId]) => { + const { handle } = await window.bench.find(url); + handle.change((d) => { + d.marker = "unflushed"; + }); + await window.repo.removeFromCache(documentId); + }, + [url, documentId] as const + ); + const survived = await reachedSibling; + record({ + metric: "evict with an unflushed edit → a sibling sees it", + mode, + value: survived, + unit: "ok", + }); + expect.soft(survived).toBe(true); +}); diff --git a/sites/bench/tests/global-setup.ts b/sites/bench/tests/global-setup.ts index 1093c6e8a..635bdcebd 100644 --- a/sites/bench/tests/global-setup.ts +++ b/sites/bench/tests/global-setup.ts @@ -2,7 +2,8 @@ import { mkdirSync, writeFileSync } from "node:fs"; import { dirname } from "node:path"; import { RESULTS } from "./bench.js"; +// BENCH_APPEND=1 keeps earlier rows, for re-running one spec or one mode. export default function globalSetup(): void { mkdirSync(dirname(RESULTS), { recursive: true }); - writeFileSync(RESULTS, ""); + if (!process.env.BENCH_APPEND) writeFileSync(RESULTS, ""); } diff --git a/sites/bench/tests/global-teardown.ts b/sites/bench/tests/global-teardown.ts index 5bcf567a1..0c850bcb3 100644 --- a/sites/bench/tests/global-teardown.ts +++ b/sites/bench/tests/global-teardown.ts @@ -4,6 +4,9 @@ import { MODES, RESULTS, STORAGES, type Result } from "./bench.js"; const STORAGE_LABELS = { worker: "idb worker", direct: "idb in-thread" }; function cell({ value, unit }: Result): string { + if (value === null || (typeof value === "number" && !Number.isFinite(value))) { + return "–"; + } if (unit === "ok") return value ? "ok" : "FAIL"; if (unit === "n") return String(value); return `${Math.round(Number(value))} ${unit}`; diff --git a/sites/bench/tests/memory.spec.ts b/sites/bench/tests/memory.spec.ts new file mode 100644 index 000000000..617a7d108 --- /dev/null +++ b/sites/bench/tests/memory.spec.ts @@ -0,0 +1,309 @@ +import { test, type Browser, type BrowserContext, type Page } from "@playwright/test"; +import { appendFileSync, readFileSync, writeFileSync } from "node:fs"; +import { + createDoc, + median, + online, + openTab, + rendererMemory, + rendererProcesses, + timeFind, + type Mode, +} from "./bench.js"; + +const ROWS = "bench-results/memory.jsonl"; +const OUT = "bench-results/memory.md"; +const SETTLE_MS = 3_000; +const GC_MS = 500; +const TABS = [1, 10]; +const TYPE_ORDER = ["JavaScript", "WebAssembly", "DOM", "Shared", "Canvas", "other"]; +const SCOPE_ORDER = [ + "Window", + "DedicatedWorkerGlobalScope", + "SharedWorkerGlobalScope", + "ServiceWorkerGlobalScope", + "none", +]; + +type Scenario = { name: string; url?: string; ready?: "marks" | "repo"; mode?: Mode }; +const SCENARIOS: Scenario[] = [ + { name: "S0 about:blank" }, + { + name: "S1 bundle parsed, no wasm", + url: "/?mode=pertab-mesh&server=none&stop=js", + ready: "marks", + }, + { + name: "S2 wasm instantiated, no Repo", + url: "/?mode=pertab-mesh&server=none&stop=wasm", + ready: "marks", + }, + { + name: "S3 Repo + IndexedDB + mesh, no socket, no doc", + url: "/?mode=pertab-mesh&server=none", + ready: "repo", + }, + { name: "S4 pertab-mesh, server, shared doc", mode: "pertab-mesh" }, + { name: "S5 patchwork, server, shared doc", mode: "patchwork" }, +]; + +type Muasm = { + bytes: number; + breakdown: Array<{ + bytes: number; + types: string[]; + attribution: Array<{ url?: string; scope?: string }>; + }>; +}; +type Grouped = { bytes: number; types: Record; scopes: Record }; +type Row = { + scenario: string; + n: number; + footprintMb?: number; + processes?: Array<{ pid: number; mb: number }>; + heap?: Record; + muasm?: Grouped; + wasmBoot?: [number, number]; + error?: string; +}; + +async function open(context: BrowserContext, scenario: Scenario, n: number): Promise { + const pages: Page[] = []; + if (scenario.mode) { + for (let i = 0; i < n; i++) pages.push(await openTab(context, scenario.mode)); + await Promise.all(pages.map((page) => online(page))); + const url = await createDoc(pages[0], { counter: 0 }); + await Promise.all(pages.slice(1).map((page) => timeFind(page, url))); + return pages; + } + for (let i = 0; i < n; i++) { + const page = await context.newPage(); + page.on("pageerror", (error) => console.error(`[${scenario.name}]`, error.message)); + await page.goto(scenario.url ?? "about:blank"); + if (scenario.ready === "marks") { + await page.waitForFunction(() => window.bench?.marks.ready !== undefined, null, { + timeout: 60_000, + }); + } + if (scenario.ready === "repo") { + await page.waitForFunction(() => window.repo != null, null, { timeout: 60_000 }); + } + pages.push(page); + } + return pages; +} + +function group(muasm: Muasm): Grouped { + const types: Record = {}; + const scopes: Record = {}; + for (const { bytes, types: t, attribution } of muasm.breakdown) { + for (const type of t.length ? t : ["other"]) types[type] = (types[type] ?? 0) + bytes; + const owners = attribution.length ? attribution : [{ scope: "none" }]; + for (const { scope = "none" } of owners) { + scopes[scope] = (scopes[scope] ?? 0) + bytes / owners.length; + } + } + return { bytes: muasm.bytes, types, scopes }; +} + +function medians(records: Record[]): Record { + const keys = [...new Set(records.flatMap((r) => Object.keys(r)))]; + return Object.fromEntries(keys.map((k) => [k, median(records.map((r) => r[k] ?? 0))])); +} + +async function measure(browser: Browser, pages: Page[]): Promise> { + await pages[0].waitForTimeout(SETTLE_MS); + const sessions = await Promise.all( + pages.map((page) => page.context().newCDPSession(page)) + ); + await Promise.all(sessions.map((s) => s.send("HeapProfiler.collectGarbage"))); + await pages[0].waitForTimeout(GC_MS); + const heaps = await Promise.all( + sessions.map(async (s) => { + await s.send("Performance.enable"); + const { metrics } = await s.send("Performance.getMetrics"); + return Object.fromEntries(metrics.map((m) => [m.name, m.value])); + }) + ); + const [footprint, processes] = await Promise.all([ + rendererMemory(browser), + rendererProcesses(browser), + ]); + const isolated = await pages[0].evaluate(() => crossOriginIsolated); + const grouped = isolated + ? await Promise.all( + pages.map((page) => + page + .evaluate(() => + ( + performance as unknown as { + measureUserAgentSpecificMemory(): Promise; + } + ).measureUserAgentSpecificMemory() + ) + .then(group) + ) + ) + : []; + return { + footprintMb: footprint.mb, + processes, + heap: medians(heaps), + muasm: grouped.length + ? { + bytes: median(grouped.map((g) => g.bytes)), + types: medians(grouped.map((g) => g.types)), + scopes: medians(grouped.map((g) => g.scopes)), + } + : undefined, + }; +} + +async function reloadTiming(page: Page): Promise<[number, number]> { + const first = await page.evaluate(() => window.bench.marks.wasm); + await page.reload(); + await page.waitForFunction(() => window.bench?.marks.ready !== undefined, null, { + timeout: 60_000, + }); + const second = await page.evaluate(() => window.bench.marks.wasm); + return [Math.round(first), Math.round(second)]; +} + +const mb = (bytes: number) => (bytes / 1048576).toFixed(1); +const cell = (v: unknown) => (v === undefined || v === null ? "–" : String(v)); + +function scenarioTable(name: string, rows: Row[]): string { + const cols = rows.map((row) => `N=${row.n}`); + const line = (label: string, pick: (row: Row) => unknown) => + `| ${label} | ${rows.map((row) => cell(pick(row))).join(" | ")} |`; + const per = (row: Row) => row.processes!.map((p) => p.mb); + const keys = (pick: (row: Row) => Record | undefined, order: string[]) => { + const all = new Set(rows.flatMap((row) => Object.keys(pick(row) ?? {}))); + return [...order.filter((k) => all.has(k)), ...[...all].filter((k) => !order.includes(k))]; + }; + const out = [ + `### ${name}`, + "", + `| metric | ${cols.join(" | ")} |`, + `| --- | ${cols.map(() => "---").join(" | ")} |`, + line("renderer footprint MB, all processes", (r) => r.footprintMb), + line("renderer processes", (r) => r.processes?.length), + line("helper processes (processes − tabs)", (r) => r.processes && r.processes.length - r.n), + line( + "per-process MB min / median / max", + (r) => + r.processes && + `${Math.min(...per(r))} / ${median(per(r)).toFixed(1)} / ${Math.max(...per(r))}` + ), + line("JSHeapUsedSize MB (median tab)", (r) => r.heap && mb(r.heap.JSHeapUsedSize)), + line("JSHeapTotalSize MB (median tab)", (r) => r.heap && mb(r.heap.JSHeapTotalSize)), + line("Nodes (median tab)", (r) => r.heap?.Nodes), + line("Documents (median tab)", (r) => r.heap?.Documents), + line("MUASM total MB (median tab)", (r) => r.muasm && mb(r.muasm.bytes)), + ...keys((r) => r.muasm?.types, TYPE_ORDER).map((k) => + line(`MUASM type ${k} MB`, (r) => r.muasm && mb(r.muasm.types[k] ?? 0)) + ), + ...keys((r) => r.muasm?.scopes, SCOPE_ORDER).map((k) => + line(`MUASM scope ${k} MB`, (r) => r.muasm && mb(r.muasm.scopes[k] ?? 0)) + ), + ]; + if (rows.some((r) => r.wasmBoot)) { + out.push(line("nav→wasm ms, first load / reload", (r) => r.wasmBoot?.join(" / "))); + } + if (rows.some((r) => r.error)) out.push(line("error", (r) => r.error)); + return out.join("\n"); +} + +function summary(rows: Row[]): string { + const header = [ + "scenario", + "footprint MB N=1", + "footprint MB N=10", + "marginal MB/tab", + "median process MB @N=10", + "JS heap used MB", + "MUASM JavaScript MB", + "MUASM WebAssembly MB", + "MUASM DOM MB", + "medians at N", + ]; + const lines = SCENARIOS.map(({ name }) => { + const one = rows.find((r) => r.scenario === name && r.n === 1); + const ten = rows.find((r) => r.scenario === name && r.n === 10); + const at = ten?.heap ? ten : one?.heap ? one : undefined; + const marginal = + one?.footprintMb !== undefined && ten?.footprintMb !== undefined + ? ((ten.footprintMb - one.footprintMb) / 9).toFixed(1) + : undefined; + return `| ${[ + name, + one?.footprintMb, + ten?.footprintMb, + marginal, + ten?.processes && median(ten.processes.map((p) => p.mb)).toFixed(1), + at?.heap && mb(at.heap.JSHeapUsedSize), + at?.muasm && mb(at.muasm.types.JavaScript ?? 0), + at?.muasm && mb(at.muasm.types.WebAssembly ?? 0), + at?.muasm && mb(at.muasm.types.DOM ?? 0), + at?.n, + ] + .map(cell) + .join(" | ")} |`; + }); + return [ + "## Summary", + "", + `| ${header.join(" | ")} |`, + `| ${header.map(() => "---").join(" | ")} |`, + ...lines, + ].join("\n"); +} + +function save(row: Row): Row[] { + appendFileSync(ROWS, JSON.stringify(row) + "\n"); + const rows = [ + ...new Map( + readFileSync(ROWS, "utf8") + .split("\n") + .filter(Boolean) + .map((line) => JSON.parse(line) as Row) + .map((r) => [`${r.scenario}:${r.n}`, r] as const) + ).values(), + ]; + const tables = SCENARIOS.map(({ name }) => rows.filter((r) => r.scenario === name)) + .filter((group) => group.length) + .map((group) => scenarioTable(group[0].scenario, group)); + writeFileSync(OUT, ["# Renderer memory ladder", ...tables, summary(rows)].join("\n\n") + "\n"); + console.log("\n" + scenarioTable(row.scenario, rows.filter((r) => r.scenario === row.scenario))); + return rows; +} + +test.beforeAll(({}, info) => { + if (info.workerIndex === 0) writeFileSync(ROWS, ""); +}); + +test.afterAll(() => { + console.log("\n" + readFileSync(OUT, "utf8").split("## Summary")[1]); +}); + +for (const scenario of SCENARIOS) { + for (const n of TABS) { + test(`${scenario.name}: ${n} tab(s)`, async ({ browser }) => { + const context = await browser.newContext(); + const row: Row = { scenario: scenario.name, n }; + try { + const pages = await open(context, scenario, n); + Object.assign(row, await measure(browser, pages)); + if (scenario.name.startsWith("S2") && n === 1) { + row.wasmBoot = await reloadTiming(pages[0]); + } + } catch (error) { + row.error = error instanceof Error ? error.message : String(error); + throw error; + } finally { + await context.close(); + save(row); + } + }); + } +} diff --git a/sites/bench/tests/offline.spec.ts b/sites/bench/tests/offline.spec.ts index 9b10bf9d5..8f070e029 100644 --- a/sites/bench/tests/offline.spec.ts +++ b/sites/bench/tests/offline.spec.ts @@ -1,49 +1,78 @@ -import { expect, test } from "@playwright/test"; +import { test } from "@playwright/test"; import { MODES, + WORKER_MODES, awaitField, awaitOnline, createDoc, getField, + offlineSwitch, online, openTab, record, serverConfirmed, setField, - setOffline, + strangerSees, } from "./bench.js"; -// Both tabs edit while the network is cut, then it comes back. The cut is -// Playwright's offline emulation for a tab's own socket, and a control -// message for a worker's (see setOffline); whether the link was actually seen -// down is recorded, so a mode whose socket survived the cut says so. +// Both tabs edit while the link to the server is cut, then it comes back. The +// cut is a WebSocket proxy closing a tab's own socket, and a control message +// for a worker's (see offlineSwitch). Modes with a local channel converge +// while offline; the reconnect timing is only meaningful where they didn't. for (const mode of MODES) { test(`${mode}: concurrent offline edits converge on reconnect`, async ({ + browser, context, }) => { + const network = await offlineSwitch(context); const a = await openTab(context, mode); const b = await openTab(context, mode); await Promise.all([online(a), online(b)]); const url = await createDoc(a, { x: 0, y: 0 }); - await awaitField(b, url, "x", 0); + // Settled at the server before anyone else asks: the racing find is + // sync.spec's business. await serverConfirmed(a, url); + await awaitField(b, url, "x", 0); - await setOffline(context, [a, b], true); + await network.set([a, b], true); const wentDown = (await Promise.all([ awaitOnline(a, false), awaitOnline(b, false), ])).every(Boolean); + // In the worker modes the bench closes the socket itself, so this only + // says something about the tab modes' proxy. + if (!WORKER_MODES.includes(mode)) { + record({ + metric: "server link seen down while offline", + mode, + value: wentDown, + unit: "ok", + }); + } + await setField(a, url, "x", 1); + await setField(b, url, "y", 1); + await a.waitForTimeout(2_000); + + const leaked = await strangerSees(browser, url); record({ - metric: "server link seen down while offline", + metric: "…and the server did not see the offline edits", mode, - value: wentDown, + value: leaked?.x !== 1 && leaked?.y !== 1, + unit: "ok", + }); + const [ay, bx] = await Promise.all([ + getField(a, url, "y"), + getField(b, url, "x"), + ]); + const local = ay === 1 && bx === 1; + record({ + metric: "tabs converged while offline, over a local channel", + mode, + value: local, unit: "ok", }); - await setField(a, url, "x", 1); - await setField(b, url, "y", 1); - await a.waitForTimeout(2_000); - await setOffline(context, [a, b], false); + await network.set([a, b], false); const reconnected = Date.now(); const [seenY, seenX] = await Promise.all([ awaitField(a, url, "y", 1, 60_000).then(() => true, () => false), @@ -58,21 +87,23 @@ for (const mode of MODES) { value: seenY && seenX && confirmed, unit: "ok", }); - if (seenY && seenX) { - record({ metric: "reconnect → tabs converged", mode, value: converged, unit: "ms" }); - } + record({ + metric: "reconnect → tabs converged (only where they hadn't already)", + mode, + value: !local && seenY && seenX ? converged : null, + unit: "ms", + }); const c = await openTab(context, mode); const [x, y] = await Promise.all([ - getField(c, url, "x"), - getField(c, url, "y"), + getField(c, url, "x").catch(() => undefined), + getField(c, url, "y").catch(() => undefined), ]); - expect({ seenY, seenX, confirmed, x, y }).toEqual({ - seenY: true, - seenX: true, - confirmed: true, - x: 1, - y: 1, + record({ + metric: "…and a fresh tab sees both", + mode, + value: x === 1 && y === 1, + unit: "ok", }); }); } diff --git a/sites/bench/tests/scale.spec.ts b/sites/bench/tests/scale.spec.ts new file mode 100644 index 000000000..0bfa50c05 --- /dev/null +++ b/sites/bench/tests/scale.spec.ts @@ -0,0 +1,241 @@ +import { test, type Page } from "@playwright/test"; +import { + type Mode, + awaitField, + createDoc, + marks, + median, + online, + openTab, + probeServerPeer, + record, + rendererMemory, + serverConfirmed, + setField, + timeFind, + withStall, +} from "./bench.js"; + +const MODES = (process.env.SCALE_MODES ?? "patchwork,pertab-heads") + .split(",") + .map((mode) => mode.trim()) + .filter(Boolean) as Mode[]; +const TABS = Number(process.env.SCALE_TABS ?? 40); +const DOC_KB = Number(process.env.SCALE_DOC_KB ?? 0); +const ROUNDS = 5; +const FIND_TIMEOUT_MS = 5_000; +const EDIT_TIMEOUT_MS = 5_000; +const LINKED_WAIT_MS = 3_000; +const WATCHERS_SETTLE_MS = 200; + +function settled(promise: Promise, ms: number): Promise { + return Promise.race([ + promise.catch(() => undefined), + new Promise((resolve) => setTimeout(resolve, ms)), + ]); +} + +function medianOrNull(values: number[]): number | null { + return values.length ? median(values) : null; +} + +function maxOrNull(values: number[]): number | null { + return values.length ? Math.max(...values) : null; +} + +function round(value: number | null, places: number): number | null { + return value === null ? null : Number(value.toFixed(places)); +} + +type Counters = { + rounds: number; + bytes: number | null; + writes: number | null; +}; + +function counters(page: Page): Promise { + return page.evaluate(() => ({ + rounds: window.bench.syncRounds(), + bytes: window.bench.serverBytes(), + writes: window.bench.storageWrites(), + })); +} + +function perEditPerTab( + before: Counters[], + after: Counters[], + key: "bytes" | "writes", + edits: number +): number | null { + let total = 0; + for (const [i, b] of before.entries()) { + const a = after[i][key]; + const v = b[key]; + if (a === null || v === null) return null; + total += a - v; + } + return total / edits / before.length; +} + +for (const mode of MODES) { + test(`${mode}: ${TABS} tabs on one doc`, async ({ browser, context }) => { + test.setTimeout(Math.max(600_000, TABS * 15_000)); + const N = TABS; + await probeServerPeer(browser); + + const opening = Date.now(); + const pages: Page[] = []; + for (let i = 0; i < N; i++) pages.push(await openTab(context, mode)); + const openMs = Date.now() - opening; + const connected = await Promise.all(pages.map((page) => online(page))); + const onlineMs = Date.now() - opening; + + const last = pages[N - 1]; + await last + .waitForFunction(() => window.bench.marks.linked !== undefined, null, { + timeout: LINKED_WAIT_MS, + }) + .catch(() => {}); + const boot = await marks(last); + const sinceStart = (name: string) => + boot[name] === undefined ? null : boot[name] - boot.start; + + const [creator, ...others] = pages; + const url = await createDoc(creator, { + title: "scale", + ...(DOC_KB ? { text: "x".repeat(DOC_KB * 1024) } : {}), + }); + const finds = await Promise.all( + others.map((page) => + timeFind(page, url, FIND_TIMEOUT_MS).then( + (found) => ({ ok: true as const, ...found }), + () => ({ ok: false as const }) + ) + ) + ); + const found = finds.filter((find) => find.ok); + const findMs = found.map((find) => find.ms); + + const before = await Promise.all(pages.map(counters)); + const editors = new Set(); + const seenMs: number[] = []; + const missed = pages.map(() => 0); + const confirmMs: number[] = []; + let unconfirmed = 0; + const { stall } = await withStall(creator, async () => { + for (let r = 0; r < ROUNDS; r++) { + const editor = pages[(r * 7) % N]; + editors.add((r * 7) % N); + const watchers = pages.filter((page) => page !== editor); + const field = String(r); + const seen = watchers.map((page) => + settled( + awaitField(page, url, field, r, EDIT_TIMEOUT_MS), + EDIT_TIMEOUT_MS + 1_000 + ) + ); + await editor.waitForTimeout(WATCHERS_SETTLE_MS); + const edited = await setField(editor, url, field, r); + const [arrived, confirmed] = await Promise.all([ + Promise.all(seen), + settled( + serverConfirmed(editor, url, EDIT_TIMEOUT_MS), + EDIT_TIMEOUT_MS + 1_000 + ), + ]); + for (const [i, at] of arrived.entries()) { + if (at === undefined) missed[pages.indexOf(watchers[i])]++; + else seenMs.push(at - edited); + } + if (confirmed === undefined) unconfirmed++; + else confirmMs.push(confirmed - edited); + } + }); + + const after = await Promise.all(pages.map(counters)); + const roundsPerSiblingEdit = pages + .map((_, i) => i) + .filter((i) => !editors.has(i)) + .map((i) => (after[i].rounds - before[i].rounds) / ROUNDS); + + const memory = await rendererMemory(browser); + const workerErrors = ( + await Promise.all( + pages.map((page) => page.evaluate(() => window.bench.workerErrors())) + ) + ).flat(); + + const result = { + mode, + tabs: N, + openMs, + onlineMs, + bootRepoMs: sinceStart("repo"), + bootLinkedMs: sinceStart("linked"), + bootNodeMs: sinceStart("node"), + bootReadyMs: boot.ready, + bootServerMs: connected[N - 1], + findMedianMs: medianOrNull(findMs), + findSlowestMs: maxOrNull(findMs), + found: found.length, + allFound: found.length === others.length, + retried: found.filter((find) => find.attempts > 1).length, + seenMedianMs: medianOrNull(seenMs), + seenSlowestMs: maxOrNull(seenMs), + tabsSawEveryEdit: missed.slice(1).filter((n) => n === 0).length, + allTabsSawEveryEdit: missed.filter((n) => n === 0).length, + missedPerTab: missed, + confirmMedianMs: medianOrNull(confirmMs), + confirmed: confirmMs.length, + allConfirmed: unconfirmed === 0, + memoryMb: memory.mb, + processes: memory.processes, + stallMaxMs: stall.maxMs, + stallLongTasks: stall.longTasks, + stallLongTaskMs: stall.longTaskMs, + workerErrors: workerErrors.length, + roundsPerSiblingEditInNonEditingTab: round( + medianOrNull(roundsPerSiblingEdit), + 2 + ), + serverBytesPerEditPerTab: round( + perEditPerTab(before, after, "bytes", ROUNDS), + 0 + ), + storageWritesPerEditPerTab: round( + perEditPerTab(before, after, "writes", ROUNDS), + 2 + ), + bootServerLastMs: maxOrNull(connected), + }; + + const rows: Array< + [string, number | boolean | null, "ms" | "MB" | "n" | "ok"] + > = [ + [`open ${N} tabs, wall clock`, result.openMs, "ms"], + [`boot → repo, last of ${N} tabs`, result.bootRepoMs, "ms"], + [`boot → linked, last of ${N} tabs`, result.bootLinkedMs, "ms"], + [`boot → server connected, last of ${N} tabs`, result.bootServerMs, "ms"], + [`renderer memory, ${N} tabs`, result.memoryMb, "MB"], + [`renderer processes, ${N} tabs`, result.processes, "n"], + [`${N - 1} tabs find the doc: median`, result.findMedianMs, "ms"], + [`${N - 1} tabs find the doc: slowest`, result.findSlowestMs, "ms"], + [`…all ${N - 1} found it`, result.allFound, "ok"], + [`…finds that needed a retry (of ${N - 1})`, result.retried, "n"], + [`edit → seen in ${N - 1} tabs: median tab`, result.seenMedianMs, "ms"], + [`edit → seen in ${N - 1} tabs: slowest tab`, result.seenSlowestMs, "ms"], + [`…tabs that saw every edit (of ${N - 1})`, result.tabsSawEveryEdit, "n"], + [`edit → server holds our heads (${N} tabs)`, result.confirmMedianMs, "ms"], + [`…server confirmed every edit (${N} tabs)`, result.allConfirmed, "ok"], + [`…longest main-thread stall in tab 0 during the edits`, result.stallMaxMs, "ms"], + [`sync rounds per sibling edit, non-editing tab`, result.roundsPerSiblingEditInNonEditingTab, "n"], + [`server bytes received per edit per tab`, result.serverBytesPerEditPerTab, "n"], + [`storage writes (saveBatch) per edit per tab`, result.storageWritesPerEditPerTab, "n"], + [`worker errors (${N} tabs)`, result.workerErrors, "n"], + ]; + for (const [metric, value, unit] of rows) { + record({ metric, mode, value, unit }); + } + console.log(`SCALE_RESULT ${JSON.stringify(result)}`); + }); +} diff --git a/sites/bench/tests/storage-adapter.spec.ts b/sites/bench/tests/storage-adapter.spec.ts index 9d6666405..a202d4ac4 100644 --- a/sites/bench/tests/storage-adapter.spec.ts +++ b/sites/bench/tests/storage-adapter.spec.ts @@ -39,8 +39,11 @@ for (const storage of STORAGES) { unit: "ms", }); - // Seed: many docs, each with a run of edits, then one flush for the lot. - // The stall probe says how much of that the main thread felt. + // Seed: many docs, each with a run of edits, yielding between docs so the + // saves those edits trigger interleave with the loop as they would in an + // app, then one flush for the lot. The stall probe says how much of that + // the main thread felt; the changes themselves are the same work under + // either adapter, so the flush is also timed and probed on its own. const { result: seeded, stall: seedStall } = await withStall(a, () => a.evaluate( async ([docs, edits]) => { @@ -66,9 +69,17 @@ for (const storage of STORAGES) { }); } urls.push(handle.url); + await new Promise((resolve) => setTimeout(resolve)); } + const changed = performance.now() - started; + window.bench.stallStop(); + window.bench.stallStart(); await window.repo.flush(); - return { ms: performance.now() - started, urls }; + return { + ms: performance.now() - started, + flushMs: performance.now() - started - changed, + urls, + }; }, [DOCS, EDITS_PER_DOC] as const ) @@ -81,14 +92,21 @@ for (const storage of STORAGES) { unit: "ms", }); record({ - metric: "…longest main-thread stall during that", + metric: "…of which the final flush", + mode, + storage, + value: seeded.flushMs, + unit: "ms", + }); + record({ + metric: "…longest main-thread stall during the flush", mode, storage, value: seedStall.maxMs, unit: "ms", }); record({ - metric: "…main-thread time in long tasks during that", + metric: "…main-thread time in long tasks during the flush", mode, storage, value: seedStall.longTaskMs, @@ -123,7 +141,7 @@ for (const storage of STORAGES) { unit: "ms", }); record({ - metric: "…longest main-thread stall during that", + metric: "…longest main-thread stall over those rounds", mode, storage, value: flushStall.maxMs, @@ -161,18 +179,24 @@ for (const storage of STORAGES) { unit: "ms", }); record({ - metric: "…longest main-thread stall during that", + metric: "…longest main-thread stall during the load", mode, storage, value: loadStall.maxMs, unit: "ms", }); - // Several tabs flushing to the same database at once. + // Several tabs flushing to the same database at once. Each has its doc + // open before the clock starts, so a cold load isn't counted. const writers = [a, b]; while (writers.length < WRITERS) { writers.push(await openTab(context, mode, opts)); } + await Promise.all( + writers.map((page, i) => + page.evaluate((url) => window.bench.find(url), seeded.urls[i + 1]) + ) + ); const started = Date.now(); await Promise.all( writers.map((page, i) => @@ -224,6 +248,7 @@ for (const storage of STORAGES) { const b = await openTab(context, mode, { storage }); await Promise.all([online(a), online(b)]); const url = await createDoc(a, { counter: 0 }); + await serverConfirmed(a, url); await awaitField(b, url, "counter", 0); const propagation: number[] = []; diff --git a/sites/bench/tests/storage.spec.ts b/sites/bench/tests/storage.spec.ts index 4f39be32b..08e00f492 100644 --- a/sites/bench/tests/storage.spec.ts +++ b/sites/bench/tests/storage.spec.ts @@ -1,10 +1,11 @@ -import { expect, test } from "@playwright/test"; +import { test } from "@playwright/test"; import { MODES, createDoc, flush, getField, openTab, + refuseWebSockets, record, setField, type Mode, @@ -12,40 +13,51 @@ import { const EDITS = 20; -// The second-writer question. Every mode but patchwork runs with no server, so -// a tab can only see another's work through storage — its own IndexedDB, or -// the worker's. The patchwork mode's server is build-time, so it keeps its -// socket; the siblings channel is what carries the edits there, and closing -// tabs tests storage all the same. +// The second-writer question, with nothing but storage to answer it: no +// server, and the tab that wrote is closed before the tab that reads opens, +// so a live sibling channel can't answer either. The bare modes take +// `?server=none`; patchwork's server is build-time, so its sockets are +// refused instead. In the worker modes "storage" is the worker's IndexedDB; +// the shared worker also keeps what it relayed in memory. const server = (mode: Mode) => (mode === "patchwork" ? undefined : "none"); +async function storageOnly(context: import("@playwright/test").BrowserContext, mode: Mode) { + if (mode === "patchwork") await refuseWebSockets(context); +} + for (const mode of MODES) { - test(`${mode}: a doc written by one tab is found by the next`, async ({ + test(`${mode}: a doc written by a closed tab is found by the next`, async ({ context, }) => { + await storageOnly(context, mode); const a = await openTab(context, mode, { server: server(mode) }); const url = await createDoc(a, { n: 0 }); for (let i = 1; i <= EDITS; i++) await setField(a, url, "n", i); await flush(a); + await a.close(); const b = await openTab(context, mode, { server: server(mode) }); const started = Date.now(); const seen = await getField(b, url, "n").catch(() => undefined); + const found = seen !== undefined; record({ - metric: "second tab finds first tab's doc (no server)", + metric: "a fresh tab finds a closed tab's doc (no server)", + mode, + value: found, + unit: "ok", + }); + record({ + metric: "…time to find it, from storage", + mode, + value: found ? Date.now() - started : null, + unit: "ms", + }); + record({ + metric: "…and every edit made before the close is in it", mode, value: seen === EDITS, unit: "ok", }); - if (seen === EDITS) { - record({ - metric: "second tab find, from storage", - mode, - value: Date.now() - started, - unit: "ms", - }); - } - expect(seen).toBe(EDITS); }); // Both tabs close right after their last edit, as a user would. This asks @@ -55,6 +67,7 @@ for (const mode of MODES) { test(`${mode}: two tabs write the same doc and close; a third reads it`, async ({ context, }) => { + await storageOnly(context, mode); const a = await openTab(context, mode, { server: server(mode) }); const b = await openTab(context, mode, { server: server(mode) }); const url = await createDoc(a, { a: 0, b: 0 }); @@ -95,6 +108,5 @@ for (const mode of MODES) { value: fromA === EDITS && fromB === EDITS, unit: "ok", }); - expect({ fromA, fromB }).toEqual({ fromA: EDITS, fromB: EDITS }); }); } diff --git a/sites/bench/tests/sync.spec.ts b/sites/bench/tests/sync.spec.ts index 51359752c..c00ea46cc 100644 --- a/sites/bench/tests/sync.spec.ts +++ b/sites/bench/tests/sync.spec.ts @@ -9,61 +9,138 @@ import { record, serverConfirmed, setField, + strangerSees, timeFind, } from "./bench.js"; const TABS = 3; const EDITS = 10; +const RACE_TIMEOUT_MS = 5_000; +const EDIT_TIMEOUT_MS = 5_000; // Tab A creates and edits; tabs B.. find and watch. Each latency is a median // over EDITS rounds. for (const mode of MODES) { - test(`${mode}: cross-tab and server latency`, async ({ context }) => { + test(`${mode}: cross-tab and server latency`, async ({ browser, context }) => { const pages = []; for (let i = 0; i < TABS; i++) pages.push(await openTab(context, mode)); await Promise.all(pages.map((page) => online(page))); const [a, ...others] = pages; const url = await createDoc(a, { counter: 0 }); - // Per-tab modes with no local fan-out only learn of the doc via the - // server, so the first find includes a server round trip by design. - const finds = await Promise.all(others.map((page) => timeFind(page, url))); + // The race: the siblings ask for a doc that may not have left A yet. A + // find that settles unavailable is retried with a real re-sync (see + // `find` in src/main.ts); a tab whose find still failed is replaced below. + const finds = await Promise.all( + others.map((page) => + timeFind(page, url, RACE_TIMEOUT_MS).then( + (found) => ({ ok: true as const, ...found }), + () => ({ ok: false as const }) + ) + ) + ); + const found = finds.filter((find) => find.ok); record({ metric: "find a doc another tab just created", mode, - value: median(finds.map((find) => find.ms)), + value: found.length ? median(found.map((find) => find.ms)) : null, unit: "ms", }); record({ - metric: "…and the first find() didn't settle unavailable", + metric: "…and every such find() succeeded within 5s", mode, - value: finds.every((find) => find.attempts === 1), + value: finds.every((find) => find.ok), + unit: "ok", + }); + record({ + metric: "…and none settled unavailable first", + mode, + value: finds.every((find) => find.ok && find.attempts === 1), unit: "ok", }); + await serverConfirmed(a, url); + const watchers = []; + for (const [i, find] of finds.entries()) { + if (find.ok) { + watchers.push(others[i]); + continue; + } + await others[i].close(); + const fresh = await openTab(context, mode); + await online(fresh); + const { ms } = await timeFind(fresh, url); + record({ + metric: "find that doc from a fresh tab, once the server has it", + mode, + value: ms, + unit: "ms", + }); + watchers.push(fresh); + } + + // Each edit waits a bounded time for the watchers and the server; an edit + // that doesn't arrive is counted rather than aborting the run, so a mode + // that strands a tab still gets a row. const propagation: number[] = []; const confirmation: number[] = []; + let missed = 0; + let unconfirmed = 0; for (let i = 1; i <= EDITS; i++) { - const seen = others.map((page) => awaitField(page, url, "counter", i)); + const seen = watchers.map((page) => + awaitField(page, url, "counter", i, EDIT_TIMEOUT_MS).catch( + () => undefined + ) + ); const edited = await setField(a, url, "counter", i); const [arrived, confirmed] = await Promise.all([ Promise.all(seen), - serverConfirmed(a, url), + serverConfirmed(a, url, EDIT_TIMEOUT_MS).catch(() => undefined), ]); - propagation.push(Math.max(...arrived) - edited); - confirmation.push(confirmed - edited); + if (arrived.every((at) => at !== undefined)) { + propagation.push(Math.max(...(arrived as number[])) - edited); + } else { + missed++; + } + if (confirmed !== undefined) confirmation.push(confirmed - edited); + else unconfirmed++; + } + if (propagation.length) { + record({ + metric: `edit → seen in ${TABS - 1} other tabs`, + mode, + value: median(propagation), + unit: "ms", + }); } record({ - metric: `edit → seen in ${TABS - 1} other tabs`, + metric: `…every edit reached every other tab within ${EDIT_TIMEOUT_MS / 1000}s`, mode, - value: median(propagation), - unit: "ms", + value: missed === 0, + unit: "ok", }); + if (confirmation.length) { + record({ + metric: "edit → server holds our heads", + mode, + value: median(confirmation), + unit: "ms", + }); + } record({ - metric: "edit → server holds our heads", + metric: `…the server confirmed every edit within ${EDIT_TIMEOUT_MS / 1000}s`, mode, - value: median(confirmation), - unit: "ms", + value: unconfirmed === 0, + unit: "ok", + }); + // The server's word, checked: a client that shares nothing with these + // tabs asks the server for the doc. + const stranger = await strangerSees(browser, url); + record({ + metric: "…and a stranger gets the whole doc from the server", + mode, + value: stranger?.counter === EDITS, + unit: "ok", }); }); }