Repository navigation
Conversation
@underlay/core implements the protocol half of the edge redesign: RFC 8785 canonical JSON, the input-rule scanner (duplicate keys, unsafe integer literals, lone surrogates, depth), record and schema hashing with v1 legacy hashes, the key-defined tree (streaming builder, incremental merge, seek/iterate/diff, fsck), and version roots with the salted private commitment. Interior entries carry the child's last key and record leaves carry the record size; see edge-redesign-build.md findings 1-2.
docs/protocol-v2.md specifies format 2 (canonical JSON, input rules, records, trees, access sets, version roots, format 1 aliases). test/vectors/v2.json is generated by scripts/gen-vectors.ts and checked in CI (content comparison, so reformatting is harmless). The chunking experiment kept the fixed-probability rule and lowered the forced leaf split to 8,192 entries.
@underlay/server gets the four ports and their adapters: blob store (S3 API via aws4fetch, filesystem with HMAC-presigned URLs, memory), SQLite via Drizzle (libsql on Node, D1 on Workers; batches only, no interactive transactions), jobs (SQLite table + runner on Node, Queues on Workers) and cache (LRU, Cache API). The object layer stores nodes, bodies (split into parts past 8 MB), out-of-line records, roots, private sets and schemas, and reads them back as a NodeSource. Node and Worker entries serve a skeleton Hono app. Core: TreeSink gains drain() so merges apply write backpressure.
Follows the plan's placements update (primary + mirrors, documented
repository layout, signed version log):
- New @underlay/repo: repository reads and writes through a BlobStore
in the documented layout (nodes/, bodies/<leaf>.ndjson.gz,
records/<hash>.json.gz, roots/, schemas/, private/, files/,
collections/<id>/{collection,head}.json and log/<seq>.json), with
hash verification for untrusted locations. Blob adapters, gzip and
the LRU move here from the server.
- Large leaf bodies are one object of concatenated gzip members
(replacing separate part objects); RepoSink records every key it
writes, which is a commit's mirror sync list.
- Signed, hash-chained version log (Ed25519 via WebCrypto) with
verifyLog for restore and audits.
- Server: storage_locations and placements tables (platform location
seeded; one primary per collection enforced), and Stores replaces
the global blob store: forCollection(id) resolves the primary,
internal holds sessions/uploads/reference log.
- Spec: repository layout and version log (section 11).
commitVersion merges each (set, type) tree with its sorted changes, moves whole trees when a type changes sets, keeps file sets right in O(changes) with per-set reference-count sidecar trees, assembles the root and private set, and publishes with one conditional SQLite batch (insert version / move head / schema_usage, all guarded on the head still being the base). The signed log entry and head.json follow, with a repair job if that write fails. Tests cover CAS conflicts, set moves, type removal, file sets and schema revalidation.
- Sorted runs in the platform's internal area: blocks of ~1 MB, an index per run for per-type reads, k-way merge with later uploads winning, compaction above a fan-in of 16 (property-tested). - Push sessions: inputs object + session row, run bookkeeping, status transitions by compare-and-swap, finalize shared by inline and job-run commits. - Delta push API (/push): upserts and deletes against a base, private flags routed to sets, records validated and hashed at upload, large records stored out of line immediately. - Negotiate API (v1 paths and shapes): presence only from the base version (no global lookup), snapshot diff against the base trees at commit, set moves without re-upload, v1 metadata merge, declared files as a full list, format 1 hashes matched for records with integer-like keys. - Commit now stores schema bodies in the repository. - Collection access resolved once per request; pluggable authenticator (anonymous until better-auth lands).
Port of v1's better-auth setup to the v2 schema (D1/libsql, adapter without transactions): KF Auth OIDC sign-in, organizations with v1's server-controlled fields and slug rules, API keys with v1's scope clamping, personal org on sign-up. The authenticator reads Bearer keys (or ?token= on GET), then the session cookie; org-owned keys act as members of that org only. /api/auth/* is mounted; both entries build auth from v1's env names.
Running the Worker under wrangler dev (workerd, local D1 and Queues, S3 API against the fake server) found that D1 batches only take query builders: the CAS publish is now INSERT ... SELECT from the collection row and conditional UPDATEs, built with Drizzle. scripts/smoke.ts pushes through a running deployment (delta sync and async commits, v1 negotiate with a format 1 hash) and passes on workerd. wrangler.jsonc defines the staging and next environments (placeholder D1 ids; nothing deployed).
Read endpoints with v1's shapes (docs/v1-read-api.md records what the UI and clients rely on): - version views decide privacy once (sets the caller may read); - records page by key or by offset in O(height), across both sets for owners (binary search on ranks) and across types, with a (type, id) cursor; no offset cap; - NDJSON streaming; manifests from nodes only, with ?since= deltas; diff via subtree-skipping tree diff, O(changes); file listings from the visible file trees; - collections list with SQL-side tag/owner filters and facets, detail, account pages and /api/context; - versions carry public file counts and bytes so non-owners never see private sizes.
- Small uploads through the API (32 MB, held in memory to hash), with v1's inert-MIME rule; direct uploads to a staging key via presigned PUT or multipart part URLs, verified by a job that streams the object, checks hash and size, and copies it to the canonical repository key files/<hash> (BlobStore.copy: S3 CopyObject, fs, memory). Staging keys can expire by lifecycle rule. - Proof of possession: an upload always carries bytes. - Downloads: 302 to a presigned attachment URL; non-members only for files in the cumulative public files tree, members also for the head's private files. HEAD reports size and type. - Known gap: copies past 5 GB need UploadPartCopy (marked TODO).
version.published creates one delivery per enabled webhook whose bump filter matches (idempotent per webhook and version) and queues a delivery job each; failures retry as delayed jobs (1 min doubling to 6 h, 5 attempts). Same body, headers and x-underlay-signature HMAC as v1. SSRF: v1's URL and private-address checks at registration; Node resolves and refuses private addresses before fetching (Workers have no private network). Management routes port v1's, for org owners and admins. A maintenance.sweep job expires idle push sessions and purges old deliveries (cron on Workers).
Phase 2: @underlay/core validates with @cfworker/json-schema (no eval or new Function), adjusted to match v1's AJV + ajv-formats: formats ported from ajv-formats, keywords next to $ref applied, later-draft keywords ignored, meta-schema and $ref checks at compile time, AJV's messages, prototype-free data. The differential test (scripts/diff-validators.ts) over all public production data agrees on all 139 schemas, 330,300 records and 85,773 mutated records. checkSchema also refuses field-level privacy at any depth and non-boolean root private. Spec section 5.1 now fixes the dialect. Core: rankOf and entryAt (O(height) seeks by key and by position). Server: create, update, delete, transfer and fork collections, and metadata edits as patch versions reusing every set. A fork is a new root over the source's sets plus a forks row; owners keep the private set and its salt.
Schema ids are hashes. Visibility comes from schema_usage (public set of a public collection, or a collection of the caller's orgs), so a private type's schema stays hidden. q matches labels and type slugs; bodies come from the repository.
After each publish a job diffs the version against the previous one (O(changes)) into events [hash, kind, collection, set, seq, op, type, id], sorts them by hash (spilling to sorted runs past 100k) and writes an immutable tier-0 run of segments: gzip blocks of ~1,024 events, an index with a Bloom filter, at most 256k events per segment. Size- tiered compaction merges FAN_IN runs into the next tier as hash-range parts run as jobs and swaps them in atomically. Segment rows, counters (collections.ref_events/ref_bytes: canonical event bytes) and the indexed flag go in one batch guarded on the version still being unindexed, so retries don't double count. Forks write no events; queries extend a parent's intervals into its forks until the fork removes the record. Routes: GET /api/records/:hash/provenance, POST /api/records/batch, GET /api/collections/files/:hash, all filtered by the caller's access before anything is returned. Format 1 hashes resolve via legacy_hashes.
GET /api/collections/:owner/:slug/export streams manifest.json, README.md (metadata.readme), records/<Type>.ndjson and files/<hash> for what the caller may read. Entry sizes come from tree bytes and counts, so nothing is buffered and there's no record cap (v1 refused over 2M). Long type names use PAX headers. format=tar skips gzip for very large exports (gzip costs Worker CPU per byte).
@underlay/migrate copies accounts, settings, files, webhooks, ARK tables and comments row for row, and replays each collection's ready versions oldest first through the v2 commit engine. A version's changes are the diff between consecutive v1 record sets, computed in Postgres in COLLATE "C" (UTF-8 byte) order, so each version costs O(changes); metadata-patch versions reuse their record sets. Versions keep their v1 semver, time and hashes (as format 1 aliases); records re-hashed by JCS get legacy_hashes aliases; types with field-level privacy become private types and are reported. commitVersion gains a migrated option (semver, createdAt, legacy hashes; indexes the reference log without firing webhooks). @underlay/server gets its package entry (src/index.ts). The test runs v1's own migrations in PGlite and checks the replayed history, privacy, legacy file keys, provenance through an alias, the signed log and lookup by legacy version hash. src/main.ts runs it against a real v1 database into a SQLite file and an S3/R2 bucket.
@underlay/web is the v1 React Router UI rendered on the server with renderToReadableStream. Loaders call the API in-process through the renderPage(request, api) contract (a Worker can't fetch its own zone), the HTML template and asset tags are inlined at build time, and dist/client is served as Workers Static Assets (wrangler assets) or by Node's serveStatic. The SQL query tool and mirror admin are gone; pages for APIs that don't exist yet degrade instead of breaking. Records tables are now in the server-rendered HTML. Verified under workerd (wrangler dev): home, explore, collection, records, versions and schemas pages render; scripts/ssr-smoke.ts passes 25 checks on Node. ARK (ported by a subagent; helpers and tests landed in a4e247e) is mounted; new collections mint an ARK as in v1, and the UI's ARK settings are on.
…e and Worker benchmarks The official JSON Schema Test Suite (draft7/dependencies.json) found that a newline in a property name stopped the dependencies message regex from matching, and the error was dropped with it, so the record passed. Regexes now match across newlines, and an invalid verdict always carries a message. scripts/validator-suite.ts runs the draft-07 suite against @cfworker/json-schema, @hyperjump/json-schema, the v2 wrapper and AJV, and measures throughput over the cached production records; scripts/workerd-bench measures it in workerd.
TreeBuilder gains a leaves-only mode (leafOutput). mergeTree with a range merges only (after, through] and outputs that range's leaves, reporting the last key so a deleted split key can be detected. assembleTree builds the interior levels from the units' segments, reusing base subtrees between them. A property test checks that units plus assembly equal the serial merge (root, counts, change stats), including gaps left to assembly and units merged after a deleted split key.
… compaction during upload A run is now ~64 KB blocks, each gzipped on its own, packed into ~4 MB part objects. The index gives each block's key span, byte range and natural-boundary marks (upsert ids with >= 13 trailing zero bits), so a reader range-GETs only the blocks overlapping a key range and the commit planner reads no data. mergeRuns is a heap over any number of runs; cursors keep one block's lines and a bounded compressed prefetch. Precedence moves from the run to the entry (q), so any group of runs can be compacted without breaking 'later upload wins'. Uploads claim 16 runs of a tier and queue push.compact, which merges them into one run of the next tier (up to tier 2, about 2.5M entries). Commit-time compaction is gone. Migration 0002 adds push_runs.tier and merging_into.
…mmit_units) A delta commit over parallelConfig.above run entries, where every type is plain (no privacy or schema change, none removed), is planned instead of merged in one job. Per (set, type) the planner takes split candidates from the runs' marks and the base tree's upper-level natural keys, leaves intervals no run block touches as gaps, groups the rest into units, and writes each unit a slice of the run blocks overlapping its range. commit.unit runs mergeTree in range mode and stores leaves, change counts and file reference deltas. commit.assemble holds a lease on the session: it merges units whose split key was deleted with the ranges after them up to one that survived, or, once all are done, assembles each tree and hands commitVersion prebuilt trees. The delta change stream moves to push/changes.ts and outcome mapping to push/outcome.ts, shared by the serial and parallel paths. Migration 0003 adds commit_units and push_sessions.commit_plan and assembly_lease. An integration test runs the same pushes in parallel and serially and compares trees, file sets and version stats, with gaps, private records, file refs and deleted split keys.
One package with one entry point (plan decision 18): the format under src/ (the old core barrel is src/format.ts), the repository and version log under src/repo/, and the memory, filesystem and S3 stores under src/stores/. Every consumer imports '@underlay/protocol'. The filesystem store now loads node:fs, node:path and node:stream on first use and signs its presigned URLs with WebCrypto HMAC, so importing the package never pulls in Node-only modules.
Store is the five methods reading, writing, mirroring and restore need: get
(with range), head, put (ifAbsent), list and delete. copy is optional
(copyObject falls back to get and put), and presigning plus multipart are an
optional presigner capability that the S3, file and memory stores have. The
factories are memoryStore(), fileStore(dir[, presign]), s3Store({...}) and the
new r2Store(binding).
One contract suite runs against all four, plain and under a prefix, with R2
through Miniflare. It caught PrefixedStore rewriting list cursors as if they
were keys: S3 and R2 cursors are opaque, so listing past 1,000 keys through a
prefix broke. Cursors now pass through untouched, and the fake S3 pages like S3.
The server gets file bytes from a presigning store (PlatformStorage.files). On
Workers an optional BUCKET R2 binding serves repository and internal objects,
while file URLs stay on the S3 API.
…checks
- SHA-256 goes through process.getBuiltinModule('node:crypto') where it exists
(Node, and Workers with nodejs_compat, checked in a wrangler bundle) and
falls back to @noble/hashes elsewhere; a property test pins the fallback to
the native hash. Salts come from crypto.getRandomValues. The file store gets
Node's modules the same way, so no bundler sees a node: import.
- The validator no longer patches @cfworker/json-schema's shared format table
at import. Its formats are registered under underlay:<name> on first compile
and our schema copies point at them; unknown formats are dropped from the
copy. The differential run over production data is unchanged (330,300
records and 85,773 mutated records agree with AJV).
- The isolate-wide 32 MB LRU is now sharedLru(), created on first use, and the
server asks for it; openRepo(store) defaults to a private 8 MB cache and to
verifying every object (trusted: false).
- PROTOCOL_VERSION and SUPPORTED_PROTOCOL_VERSIONS; Repo.root refuses a root,
or a ulv<n>: hash, of a version it doesn't read (UnsupportedProtocolError).
- sideEffects: false; an oxlint rule keeps the format half free of stores, I/O
and platform imports; a test bundles the package for the browser with esbuild
and runs it in a VM with web globals only (91 KB, 30 KB gzipped).
packVersion yields the objects a version reaches that a base doesn't: new schemas, each tree's new nodes top down (newNodes, the diffTrees walk at node level) with each new leaf's out-of-line records and body, the private set object when asked, and the root last. receiveVersion checks every object against its key before writing it, so a bad pack can't poison a key, then re-derives each tree: the changes from the base merged into the base must give the received root, count and bytes, which checks every boundary in O(changes). The root is written only after that. Tests: full and incremental syncs, public versus all sets, tar round trip, and refusals for corrupt nodes, missing bodies and nodes, swapped bodies, another version's root, foreign keys, and a tree that hashes correctly but is not canonical. A property test pins newNodes on deep trees. The tar writer moves here from the server and becomes pull-based: the old one enqueued the whole archive from start(), so a slow export download buffered it all in memory. A reader is added. Spec: section 11.2 (Sync), and change 9 (tree sync is pull-only).
GET /api/collections/:owner/:slug/versions/:n/pack?base=&sets=public|all streams a tar of packVersion's objects; sets=all needs the private sets, and the base must be a version of the same collection. GET .../log?after=&limit= returns collection.json and the signed entries after a seq. The server now writes collection.json (names and every key that has signed the log) before each log entry, which it didn't before, so mirrors and clients can verify the log. verifyLogEntries checks entries that continue a verified head. Tests pull packs over HTTP into a fresh repository (whole, then incremental), check the private set goes only to members, and verify the log against the published keys.
…ion) The database-free part of commitVersion (per-type merges, privacy moves, removed types, revalidation, file sets and the root) is now buildVersion in the protocol package, with file sizes passed in; file-set bookkeeping (count trees, FileRefDelta) moves with it. The server's commitVersion is buildVersion plus semver, publish and the log, and the CLI will build local versions with the same code, so both make identical trees from identical changes. No behaviour change: every server test passes unchanged.
…pager gets First and Last (staging check F11)
…under the version picker, Versions is a tab, version bar and version rows wrap; footer stays at the bottom (staging check F12)
…ensitive; docs say what recordCount counts for whom (staging check F14)
…ker, — for empty cells, type counts in the sidebar, a message for an unknown type; neutral type icon; one count format and plural everywhere
…ith the header; account settings get a breadcrumb; identifiers in Subscribe boxes wrap at / : and .
…the deployment's host; one-button logo upload; readable topic labels
… operations tools follow a run and say how it went; the private note stays after saving; members page shows placeholders while loading
… and upload limits labelled to match
….txt, API reference, guides and READMEs match v2 - docs/protocol-v2.md: BCP 14 requirements language, terminology, constants table, security considerations, appendices; section numbers kept. Aligned with the reference implementation where it fixes what is hashed or accepted (Appendix B, item 14). - Protocol pages restate the spec section by section. - llms.txt rewritten: v1 claims removed (global record dedup, X-RateLimit headers, offset cap, export parts and cap, fork 403), real limits, missing routes added. - API reference corrected; new pages /docs/api/records and /docs/api/sync-and-integrations. - Concepts, quickstart, integration, README, CLI README, v1-read-api notes corrected. - Docs search anchors: h2s without an id get their slug; nav headings match the pages.
…spec - A key with an invalid escape is syntax, not a thrown SyntaxError (a 500 on the server); \u escapes need exactly four hex digits. - Field-level private and the pattern-length limit reach subschemas named after data keywords (properties.type, definitions.enum, ...). - checkSchemaFull: every spec 5 and 5.1 acceptance check, for the server's session open. - receiveVersion checks the root and private set shape (checkVersionRoot, checkPrivateSetObject); checkSetObject checks slugs and null-root totals. - Vectors: syntax, bad_type, depth-boundary and code-order cases, inputRuleRecipes (record_too_large) and schemaRules; existing vectors unchanged. Spec 11.2 and Appendices A and B updated.
…nder the input rules - Session open runs checkSchemaFull over the effective type set (422); a schema error on upload is a line error, never a 500; commit refuses a session whose schemas fail. - metadata must be an object or null, metadata_patch an object (400). - Deletes pass the input rules, id and slug checks; failures answer validationErrors and totalErrors like records (first 100 lines). - needed_files uses the commit's notion of a file held for the collection (verified uploads included); filesNeeded is bare hex; the commit 409 names currentVersion.
…es across types, privacy moves in manifest deltas
- diff without from= compares with the version before (the first version with empty);
removed entries are {id, type}; the compare page shows their types.
- records.ndjson: ?after_type=&after= resumes after (type, id) through later types, with
an exact X-Underlay-Record-Count; after alone is a 400; members' private lines carry
"private": true.
- manifest ?since=: members get private flags, and a set move with the same hash is
listed under updated with previousPrivate.
- Export uses the version's time for tar entries, so the same version exports to the
same bytes.
…n management, agent links - DELETE /api/schemas/:id/labels/:label is steward-only (any user could mint an admin key). - File reads check access before the denylist, so a 451 never confirms a hash exists. - Collection-scoped keys act as members: no visibility change, deletion, transfer, webhooks or mirrors. Org settings, the NAAN and account deletion need a session or an admin key (capRole). - Members read any file held for the collection, not only the head's. - /_blob/ responses (Node, dev) are sandboxed attachments with nosniff. - Multipart complete without parts is a 400, not a 500. - GET /agent/:token, ported from v1 for the share panel's agent links, describes delta push; 'agent' is a reserved org slug.
…ines and metadata_patch
…c secrets, v1 production conversion scripts
…sions only, account rows mirrored); d1-data --since writes just the rows that changed
… to R2), backing off
…keeps the session open; an empty Bearer token is refused - versions.push_session_id (migration 0023) records the session that committed a version. A commit that finds the head moved, or loses it, by its own session's version reports 201 with that version, so a push.commit job delivered twice no longer settles failed with "Version conflict". finalizeSession only commits a committing session, and when another run settled it first returns what that run recorded. - A commit refused for missing files moves the session back to open, with the refusal in error, a fresh idle timeout and any parallel plan dropped, so the client uploads the files and commits the same session again (as llms.txt said). Docs and spec say so. - "Authorization: Bearer" with no token is 401 "Invalid or expired API key", not anonymous.
…s FormData body stream throws an unhandled error when the route cancels an oversize body (failed CI)
…nd ARK; section 11.3 specifies the read API (routes, response shapes, collection URL forms)
…RK in collection.json, rewritten (with mirrors' copies) on any change and backfilled once; the _/<collectionId> URL form and x-underlay-collection header; the collection response's standard members; a first version's diff counts its files; docs and llms.txt
… required, no reader support for the earlier string owner; the converter rewrites each collection's collection.json at the end of a run (ARK settings and sync changes come after its commits)
…NS record stays), next kept, APP_URL www
…oudflare), not a route
…angler and docs)
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Underlay v2: edge-native rebuild
Rebuilds Underlay to run on edge Workers (and on Node from the same core), replacing the v1 Hono/Postgres/Docker app.