Repository navigation
Narrow AUTO_INCREMENT ids from cluster range claims; log-cursor anti-entropy; per-row last-writer-wins (fixes #170) - #175
Merged
Merged
Conversation
…very client path
Step 1 of the narrow-id work: close the defects that would have made the
allocator's behaviour unobservable or wrong before it exists.
D1 Transpile swallowed every ApplyAST error (`if err == nil && applied`).
Rule errors now propagate verbatim; transform.CodedError carries a MySQL
code, UnsupportedStatementError maps coded -> own code, other rule
errors -> 1105, parser failures -> 1064. No string matching.
D2 NeedsIDInjection and ApplyAST decided independently and could disagree.
One analyzeInsert drives both; column-less positional inserts substitute
at the auto-increment ordinal (index into non-generated columns, not
cid); INSERT ... SELECT and INSERT ... VALUES ROW(...) are refused with
1235 in MySQL's message template instead of letting SQLite assign a key.
D3 The OK packet's insert id came from SQLite's connection-wide last-rowid
register, shared by every client on the single hookDB connection. It
now comes from the statement's own CDC entries by column name: first
generated id, 0 when none. ExecuteLocalWithHooks takes one request.
ConnectionSession.RecordInsertId is the single home of the rule that 0
never overwrites LAST_INSERT_ID(); the replica path stored zeros
unconditionally until now.
D4 Dead code: int_to_bigint rule (never registered), MergeCDCEntries,
PendingLocalExecution, MustFromWireType, HasPrimaryKey, 37 orphaned
regex vars and extractTableName in protocol/parser.go, transpile_test.
parser_vitess.go renamed to vitess_parser_instance.go for its 21 lines.
D5 An unknown wire StatementType panicked the receiving peer through
MustFromWireType with no recover. protocolStatementFromProto returns an
error and both replication call sites answer Success:false.
Replica error path (found by review): ForwardQueryResponse now carries
error_code and sql_state (fields 7, 8) so a rule's 1235/42000 reaches a
replica client intact; the leader fills both from the protocol mapper, the
replica raises the typed error. Code 0 keeps 1105/HY000 for an older leader.
The old string-formatted "ERROR 1105 (HY000): ..." also re-classified any
leader message containing "syntax error" as 1064. Replica-raised link
failures report 1105/HY000 (master's observable behaviour); 2003/2013 are
client-library codes and are never put in an ERR packet.
Generated files regenerated with protoc 33.6 (protobuf@33) and
protoc-gen-go v1.36.10; equivalence from the unchanged proto is header-only.
Evidence: red-before-green on master-identical files for D3; 20 mutation
arms fired at named assertions across three review rounds; unit suite
24 ok, race clean on replica/grpc/protocol/coordinator, integration 30/30,
benchmarks within a measured 2-4% noise floor with identical allocs/op.
Not in this change (pre-existing, recorded): three db data races on master
(transaction records shared by pointer); ./db/ no-tag stub build broken on
master; go vet findings db/vector_index_manager_test.go:315 and
grpc/forward_session.go:356; cmd/vec-bench/main_test.go is gitignored;
protocol/handlers/metadata.go:49 still formats an error code into a string.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014u4bCrv6aSsNKiY6oWuvrt
Quorum is a majority of TOTAL membership, but the membership registry
started with only the local node and was never persisted. A node restarted
while its peers were unreachable, or before gossip converged, computed a
membership of one, was its own majority, and committed alone. Prerequisite
F2 of the narrow-id design.
Persisted membership. grpc/membership_snapshot.go stores {node id,
address, incarnation, status} for every known node as msgpack in the node's
data directory, written atomically (temp file, fsync, rename, directory
fsync). The write hangs off updateClusterMetricsLocked, the choke point
every membership mutator already calls and the per-heartbeat path does
not; a fingerprint over the recorded fields skips the fsync when nothing
the snapshot stores changed. An accepted higher-incarnation update that
changes only address or incarnation now marks the record replaced so it is
persisted too (found in review).
Restore at construction. NewNodeRegistryWithDataDir loads the snapshot
before the gRPC server serves 2PC. Restored peers enter as SUSPECT: they
count toward the denominator and are excluded from replication targets
until gossip hears from them. REMOVED stays REMOVED. Incarnations are
restored so SWIM refutation keeps working. A missing or corrupt snapshot
yields a self-only registry with a warning, never a refusal to boot.
Belt. NodeProvider gains HasSeedNodes; GetClusterState refuses QUORUM,
ALL and unknown levels with a retryable 1205 while a seed-configured node
knows only itself. LOCAL_ONE and ONE are exempt: they make no denominator
claim, and refusing them made a booting node fail SELECT 1 and look dead
to liveness probes (caught by the integration suite). Single-node
deployments without seeds are unchanged.
Tests: snapshot round-trip and restore; a seven-step persist walk; fifty
heartbeats write nothing; corrupt/absent/no-data-dir cases; address and
incarnation change persisted and restored; belt table across all levels;
cluster tests TestLoneRestartedNodeRefusesToBeItsOwnQuorum (kill -9,
peers stopped, restart alone: "need 2 (majority of 3 total members, 1
alive)") and TestFreshNodeWithUnreachableSeedsRefusesWrites (1205).
Five mutation arms fired at named assertions. Unit 24 ok, race clean on
grpc/coordinator/cluster, integration 32/32.
Known and accepted: the snapshot fsync (~6 ms measured) runs under the
registry write lock, only on a real membership change; a peer removed
while this node was down inflates the denominator until gossip delivers
REMOVED (fail-closed).
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014u4bCrv6aSsNKiY6oWuvrt
…TEGER Step 3a of the narrow AUTO_INCREMENT id design. Marmot collapses every MySQL integer type to SQLite's INTEGER so INTEGER PRIMARY KEY keeps aliasing rowid, which is what LAST_INSERT_ID() depends on. That collapse lost the declared width, so a 53-bit id could be minted into an INT column. This change keeps the width where SQLite preserves it verbatim: in the CREATE TABLE text held by sqlite_master. Marker grammar. protocol/query/transform/intmarker defines `/*M:<bits>[u][a]*/` (bits 8/16/24/32, u = UNSIGNED, a = explicit AUTO_INCREMENT), Encode, a column-scoped Decode over sqlite_master text that skips string literals, quoted identifiers and table constraints, Strip, and the signed/unsigned widthMax table. BIGINT emits no marker and its DDL text is byte-identical to before. INTEGER spelled out is a MySQL synonym for INT and is marked as 32 bits. Emit. CREATE TABLE and ALTER TABLE ADD/MODIFY/CHANGE share one collapse, collapseIntegerTypeWithMarker, emitting `INTEGER /*M:..*/`. The ALTER path previously never collapsed integer types at all: ADD COLUMN age INT AUTO_INCREMENT reached SQLite as an unparseable keyword. Read. The schema cache decodes markers from CreateSQL, never from PRAGMA table_info, which normalises the comment away; ColumnSchema carries DeclaredWidth/Unsigned/ExplicitAutoInc and the transpiler SchemaInfo carries the auto-increment column's width and signedness for the allocator that follows on this branch. Client surface. SHOW CREATE TABLE strips the marker; SHOW COLUMNS and information_schema.COLUMNS read PRAGMA and are unchanged. Known limitation: during a rolling upgrade an old binary's SHOW CREATE TABLE shows the marker verbatim for a table created by a new node. Error codes. protocol/mysqlcode is a leaf package holding the MySQL error codes and SQLSTATEs so rule packages can name a code without importing protocol; protocol and transform alias it, and the bare 1064/HY000/08S01 literals in protocol/server.go, mysql_errors.go and error_mapper.go are replaced. Tests: encode/decode across width x signedness x autoinc; decode against hand-written DDL with DEFAULT (expr), string literals, quoted identifiers and table constraints; rowid aliasing with a marker present; ALTER shapes; schema-cache read from CreateSQL vs PRAGMA; SHOW CREATE strip; transpiler schema width/unsigned selection with a decoy narrow column first. Four mutation arms fired at named assertions. Unit 26 ok, race clean across protocol/..., integration 32/32, benchmarks with identical allocs/op. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014u4bCrv6aSsNKiY6oWuvrt
…lumns Step 3b of the narrow AUTO_INCREMENT id design. A node that must mint ids for an INT (or narrower) AUTO_INCREMENT column first claims a range of ids through a quorum, so no two nodes can hand out the same id. Nothing calls the allocator yet; Step 3c wires it into the insert path. On its own this commit changes no insert behaviour, and it is not meant to ship without 3c. Claim protocol. Claim rows live in the existing system database, keyed by (db, tbl), in a hidden table the client cannot write (1142/42000) and CDC never replicates. A claim rides an existing statement type with a flag and a msgpack payload (proto Statement fields 9 and 10), so an older binary rejects it at the CDC-row gate. A participant votes under the claim key's intent lock, reading its own stored base and deriving the width ceiling from its own sqlite_master: accept only when storedBase <= prevBase, newBase == prevBase, size >= 1 and the range fits the column. A rejection returns the participant's base (TransactionResponse field 7) and the coordinator retries above it, pinned to QUORUM, with jittered backoff and a retry on declined PREPARE rounds, then 1205. COMMIT applies the claim only if the stored base has not moved, and refuses (never ACKs) when the claim intent is gone. The system database syncs every commit, with fullfsync on darwin, so an ACKed base survives an OS crash. Seeding. DDL seeds the base to max(MAX(col), declared AUTO_INCREMENT=N-1), raise-only, atomically with the DDL. The width marker gains an optional floor, /*M:<bits>[u][a][:<floor>]*/. A column whose existing ids already exceed its width is refused with 1264/22003, and that code now survives the participant-to-coordinator hop (TransactionResponse field 8). An ALTER that sets a floor without tagging a column, or sets only AUTO_INCREMENT=N, is refused with 1235 instead of being silently dropped. Pre-existing defects fixed on the way, each with a regression test: - Row locks: a stale GC marker let a later writer evict a live intent holder, and emptying a table's row-lock map could drop a lock another transaction had just taken. Markers are gone and maps are kept. - A writer could take over a prepared transaction's lock after the heartbeat timeout, so two conflicting DML commits were both ACKed and one write was silently lost. Heartbeat eviction is removed; a dead coordinator's prepared rows are freed by the stale-transaction GC, including transactions recovered from Pebble after a restart. A DML COMMIT whose prepared rows are gone is refused. - Anti-entropy catch-up on a running node swapped a database's files under open connections and could replace the system database. It now detaches only the database being restored: a SQLite commit hook refuses writes from the moment of the detach (reaching clients as 1213/40001), in-flight writes are drained, lifecycle operations no longer hold the manager lock while stopping GC, and startup restores keep the higher claim base. - Snapshot exports are reference-counted, so overlapping catch-ups no longer leak full database copies; orphans are swept at startup. Benchmarks (same session, alone, medians): claim apply 3.1 ms with fullfsync on darwin, 42 us without; a user write commit costs about 1-4% more for the gate; PREPARE and row-lock paths are unchanged or faster. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014u4bCrv6aSsNKiY6oWuvrt
… claims Step 3c of the narrow AUTO_INCREMENT id design (issue #170). Inserts into an INT (or narrower) AUTO_INCREMENT column now take their ids from ranges the node claimed through a quorum (Step 3b), so ids fit the declared width and no two nodes hand out the same id while the membership is stable. Insert path. A per-(db, table) range allocator hands out ids from the node's claimed range and claims a new one when it is exhausted. An explicit id above the cluster base raises the base; one at or below it is refused with 1235 unless it falls in this node's own claimed range. Exhaustion returns 1062. A bound NULL or 0 is generated. Qualified table names are handled, prepared statements keep the client SQL, and a read-only replica forwards the client SQL rather than the transpiled one. Claim floors. Each claim row keeps three floors: committed (written only when this node applies a claim), seed (local data: DDL seeding, restore, rename) and merged (bases learned from peers). PREPARE votes against their max; COMMIT applies only if committed and seed are at or below the claim's start. Disjointness of the ranges a node ACKs rests on committed alone, so a base learned from a peer between PREPARE and COMMIT never refuses the COMMIT, while a DDL seed in that window still does. Existing system databases are migrated on open. A refused claim apply is reported as a typed error, not as node divergence. Membership. A node whose system database is new or lost holds its vote until it has merged bases from enough peers; cluster.standalone skips the hold for a single node. Every unheld node pulls the maximum bases from all alive peers every 10s, and an authenticated admin sync command does the same on demand and reports whether each node reached every member. The guarantee covers a stable membership with this procedure: change one node at a time, wait until it is ALIVE and unheld, run the sync until it reports complete, then make the next change. Claims carry the schema version they need, and every DDL path forgets cached cursors. Also fixed: - The schema version was written without sync, so a SIGKILL could lose it and leave the node unable to vote after restart. - INFORMATION_SCHEMA queries read through the read pool, so they no longer deadlock the write connection, and prepared IS queries match text ones. - A gossip reconnect now fires node-alive handling, so a node held while its seed was down releases when the seed returns. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014u4bCrv6aSsNKiY6oWuvrt
…rsisted cursors Delta sync could miss ACKed rows on a node that lagged or restarted: it compared watermarks, not the transactions each peer actually committed, and replay did not advance the schema version. Anti-entropy now pulls, per (peer, database), from each alive peer's local commit log (ListCommittedLog / FetchTransactions), with a durable cursor keyed (seq, txnID) that only advances over a contiguous prefix of applied entries. An exact in-flight tracker bounds what a peer lists. Replay is idempotent: an __marmot_applied_txn marker is written in the same SQLite transaction as the effect on every commit path, and the schema version lives in __marmot_schema_version in the DDL's own transaction. Startup catch-up uses the same pull, and a node turns ALIVE only when every database is caught up. A transaction a peer's log holds is decided COMMITTED: a locally prepared copy is committed through the local path, a begun-only record is discarded and replayed, and a late PREPARE for it is refused. A commit whose captured rows or intents no longer match its durable prepare record is refused, and the stale-transaction GC aborts only under the per-transaction commit guard. The log is garbage-collected by the minimum position every member has consumed, bounded by a maximum retention; a pair whose cursor falls behind the truncation point restores from a snapshot and re-applies the node's own log. The database set is reconciled each round with a legacy marker for pre-upgrade registry rows. DDL and CREATE/DROP DATABASE are refused while any member does not serve the log pull protocol; upgrade every node, then run DDL. Delta sync and the snapshot schema-version trailer are removed. Cluster tests gain a log-based stall watchdog, poll readiness instead of sleeping, and each finishes within a minute. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014u4bCrv6aSsNKiY6oWuvrt
…it timestamp A node that caught up from several peers applied each transaction once but not in commit order, so an older row image could overwrite a newer one, an older INSERT could resurrect a deleted row, and nodes diverged for good. No commit timestamp was shared: every node computed its own. The coordinator now merges every participant's PREPARE-time clock into its own, decides one commit HLC, and sends it on COMMIT (unary and streamed) and in the committed log listing. Every node commits with it and merges it into its clock, including on replay, and the clock is seeded at open from the highest commit recorded, so a later conflicting commit always gets a larger version. Each user database keeps __marmot_row_version (table, intent key, wall, logical, origin node, deleted). Every commit path stamps each row in the same SQLite transaction as the change and applies the change only if its version is not older than the stored one. A DELETE leaves a tombstone, an UPDATE of an absent row inserts its after-image, and an UPDATE that moves the primary key tombstones the old key independently of the new one. LOAD DATA rows are stamped per row the same way, in batches under one savepoint. DROP TABLE removes a table's versions, RENAME re-keys them, and a primary key change clears them. Tombstones older than every known member's log retention are purged in bounded chunks through a partial index. The stamp costs one extra indexed upsert per row, with no lock and no round trip. Its statement is prepared once per transaction. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014u4bCrv6aSsNKiY6oWuvrt
maxpert
marked this pull request as ready for review
September 29, 2026 16:10
…narrow AUTO_INCREMENT ids A node restarted after a graceful stop shut itself down again about half a second later. Its peers kept the LEAVING record of its own departure, the node came back at incarnation 0, and it took that record for an admin decommission. The cluster harness never saw it because it only killed nodes. A restarting node now resumes one incarnation above the one it persisted, so its ALIVE supersedes its own earlier LEAVING. Incarnation bumps from refutation and from turning ALIVE are persisted, so a restart never comes back at or below an incarnation it already announced. A REMOVED view of self is never refuted, and an admin decommission of a live node still takes effect, including while the node is still catching up. A node restarted on a wiped data directory after a graceful stop keeps the old behaviour; the operations guide says to treat it as a replaced node. Docs: compatibility.mdx describes narrow AUTO_INCREMENT columns (per-node ranges, the startup vote hold and its retryable 1205, explicit-id rules with 1235/1264/1062, no reset after DROP/CREATE); sql-reference.mdx no longer says INT AUTO_INCREMENT behaves as BIGINT; operations.mdx covers graceful restart, decommission versus removal, and changing membership one node at a time; a new application page (LLDAP 0.6.3) records the tested version, the DATABASE_URL form and the verified 1-node and 3-node scenarios. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014u4bCrv6aSsNKiY6oWuvrt
…d-back explicit transaction Defects found while rewriting the cluster suite, each fixed test first: - A local prepare whose captured rows were lost to SIGKILL wedged catch-up: the puller retried CommitLocallyPrepared until the txn was marked stuck and the node stayed JOINING. It now discards the local prepare under the commit guard and replays the peer's committed copy. - A txn committed just before SIGKILL had its commit record rewritten at open without releasing the row locks restored for it, so every later change to those rows was refused on replay and the node kept a stale value. repairAppliedTxnMetadata now cleans up after the commit it rewrites. - A narrow AUTO_INCREMENT claim inside an explicit transaction deadlocked on the batch committer flush, which waited on the writer the same transaction's pinned session holds. A claim-only commit has no SQLite work and no longer flushes. - A pinned statement whose SQLite transaction the lock wait rolled back now answers 1213 and ends the transaction, so a later COMMIT commits nothing. - XsyncTransactionStore mutated shared TxnState in place while the stale transaction GC read it. Updates are now copy-on-write. The AUTO_INCREMENT merge and base-sync intervals become config knobs with unchanged defaults. The operations guide says to rebuild wiped nodes one at a time. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014u4bCrv6aSsNKiY6oWuvrt
The cluster suite took about 11.5 minutes and had flaky tests. It is rebuilt on one harness: - The binaries are built once per run in TestMain. Each test gets its own ports and directory, and cleanup kills its nodes and prints log tails on failure. - Readiness means MySQL answers, every node sees the others ALIVE, and no node holds its claim votes; the votes are read over gRPC, including from the old-release node in the rolling upgrade. This removes that test's flake. - Fixed sleeps are replaced by polled conditions with deadlines. The only time.Sleep left is inside the poll helper. - Crash tests still use real SIGKILL and SIGTERM, with short intervals. - Assertions check exact row sets, values, ids and MySQL error codes. Retries are typed, use jittered backoff, and never mask the defect a test guards. The 23 tests that took more than 10s are rewritten from scratch. Each test that guards a fix fails when that fix is reverted. 76 tests pass in about 108s, and none takes more than 10s. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014u4bCrv6aSsNKiY6oWuvrt
… of failing at once A client that wrote to a narrow AUTO_INCREMENT table through a node that had just started, before it released its claim votes, got 1205 "lock wait timeout" at once: the node's own prepare declined the claim and the claim path gave up without waiting. An application started together with the cluster failed on its first insert. - A held node's own prepare decline is now typed (ErrLocalVotesHeld). The claim waits for the release, signalled through common.Broadcast, bounded by transaction.lock_wait_timeout_seconds, and returns 1205 only when that expires. - A peer turning ALIVE triggers an immediate merge round, so the hold ends as soon as enough members answer instead of on the next interval. - A narrow insert inside a transaction that already holds a pinned writer still fails fast, as before. Waiting there deadlocks with a joiner's snapshot checkpoint, which needs that writer. An insert that is the first write of its transaction, the common ORM shape, holds no writer and waits. - The in-process test harness now injects ids as a node does. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014u4bCrv6aSsNKiY6oWuvrt
…pplication Test names, fixtures and comments said "LLDAP shape" where they test general behaviour: a narrow insert as the first write of an ORM-style transaction, and a client that starts while a node still holds its claim votes. They now say so. Statements and assertions are unchanged. The LLDAP page no longer says to wait for the vote release before starting: inserts wait for it. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014u4bCrv6aSsNKiY6oWuvrt
maxpert
force-pushed
the
feat/narrow-autoinc-ids
branch
from
September 30, 2026 00:34
66493c2 to
276e96f
Compare
… whole log" The config validation and the tombstone horizon treat a zero max retention as unlimited, but CleanupOldTransactionRecords computed maxCutoff = now, so every GC pass deleted every committed entry whatever the members had consumed. Lagging peers then fell back to snapshots instead of pulling the log. The same behaviour exists on master. With max <= 0 the max-retention branch never fires; only the safe position plus min retention delete entries. Tests that relied on 0 meaning "delete now" use an explicit 1ms. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014u4bCrv6aSsNKiY6oWuvrt
README and the docs still described delta sync as the catch-up mechanism. They now describe the per-peer commit-log pull, applied-txn markers, snapshot fallback, GC by consumed position, the single commit timestamp and per-row versions, persisted membership, and narrow AUTO_INCREMENT range claims beside the BIGINT compact/extended ids. - reference.mdx: what the legacy delta_sync_* keys still control, the two autoinc interval keys, [metastore] strict_prepare_sync, and gc_max_retention_hours = 0 meaning no maximum. - metrics.mdx: the log-pull metrics; the delta sync counter never existed. - replication.mdx also corrects older errors: a nonexistent /intent/ key, node states and transitions, gossip timeouts, and the TxnID, LWW and HLC examples. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014u4bCrv6aSsNKiY6oWuvrt
… merge A joiner starts ALIVE at incarnation 0 and refutes the seed's JOINING record before it marks itself JOINING, and MarkJoining does not bump the incarnation. Its later promotion (ALIVE at a higher incarnation) therefore read as ALIVE -> ALIVE on the seed, AliveChanged never fired, and a held seed waited a full merge interval before releasing its claim votes. With a long interval a narrow insert stalled; this made TestNarrowAutoInc_HeldInsertInTransactionDoesNotStallJoin flaky. An ALIVE node announcing ALIVE at a higher incarnation now signals AliveChanged (the ALIVE callback still fires only on a status change). The cluster harness sets batch_commit.max_wait_ms = 1: its clients write one statement at a time, so every commit waited out the 10ms flush window. The suite drops from about 130s to about 106s. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014u4bCrv6aSsNKiY6oWuvrt
marmot_v2_replication_lag_txns was registered but never set. It now holds, per peer, the committed transactions of that peer's log past this node's pull cursor, summed over databases, as of the last pull round. Log sequence numbers are not dense (a restart skips a lease), so the value counts entries rather than subtracting positions. - Each round counts the listed entries that stay unapplied. When a round stops at the page cap, the last page asks the peer for the committed entries remaining (optional remaining_committed); the count reads index keys only, honours ctx, and a count failure leaves the value unknown without failing the listing. - The gauge is set only when every database's count for the peer is known, never to a false 0, and a removed peer's series is deleted. - The truncateThrough test helper truncates again, so the snapshot fallback test exercises the fallback. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014u4bCrv6aSsNKiY6oWuvrt
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014u4bCrv6aSsNKiY6oWuvrt
This was referenced Sep 30, 2026
This was referenced Sep 30, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes #170.
Summary
Marmot stores every MySQL integer type as SQLite
INTEGER, and generated ids came from a 53-bit HLC-derived space. Any application that declaresINT AUTO_INCREMENT(or a narrower type) and reads ids into a 32-bit type therefore broke; the issue was reported with LLDAP, but the class is general.This PR does three things:
Commits
c063286fix(protocol): deliver AUTO_INCREMENT insert ids and rule errors on every client pathINSERT ... SELECTandVALUES ROW(...)into an injected column are refused with 1235, instead of letting SQLite pick a key.1d9d887fix(cluster): a lone restarted node must not be its own write quorumef6d388feat(schema): record declared MySQL integer widths beside SQLite's INTEGER/*M:<bits>[u][a]*/marker in the storedCREATE TABLEtext keeps the declared width and signedness of TINYINT, SMALLINT, MEDIUMINT and INT.INTEGER PRIMARY KEYstill aliases rowid. BIGINT DDL is byte-identical to before.541fc0ffeat(autoinc): cluster-wide range claims for narrow AUTO_INCREMENT columns7a66596feat(autoinc): mint narrow AUTO_INCREMENT ids from cluster-wide range claimsd0d47f0feat(replication): anti-entropy pulls every peer's commit log from persisted cursorsDelta sync compared watermarks rather than the transactions each peer actually committed, so a lagging or restarted node could miss ACKed rows. It is replaced as follows:
Pulling
ListCommittedLogandFetchTransactions.Idempotent replay
__marmot_applied_txnmarker is written in the same SQLite transaction as the effect, on every commit path.__marmot_schema_version, in the DDL's own transaction.Unfinished local transactions
Log GC and catch-up
Rolling upgrade
CREATE/DROP DATABASEare refused cluster-wide (1213, retryable) while any member does not yet serve the log-pull protocol.Test harness
8914a41feat(replication): per-row last-writer-wins on the coordinator's commit timestampA node catching up from several peers applied each transaction once, but not in commit order. An older row image could overwrite a newer one, or an older INSERT could resurrect a deleted row.
Commit timestamp
Row versions
__marmot_row_version(table, intent key, wall, logical, origin node, deleted) is stamped in the same SQLite transaction as each row change, on every commit path.DDL
Tombstone purge
5ebafa0fix(cluster): a gracefully stopped node rejoins on restart; document narrow AUTO_INCREMENT idsGraceful restart rejoins. On master, a node restarted after a graceful stop shuts itself down again about half a second later. Its peers keep the
LEAVINGrecord from its own departure, and the node reads that as a decommission.REMOVEDis never refuted.Docs
compatibility.mdx: narrow AUTO_INCREMENT behaviour, including every error code above.operations.mdx: the restart note and the procedure for changing membership one node at a time.sql-reference.mdx: corrects theINT AUTO_INCREMENTrow, which said BIGINT.lldap.mdx, for a tested application (see the end-to-end check below).98eafbcfix(replication): recover cleanly from a kill mid-commit; end a rolled-back explicit transactionThe rewrite below exposed these defects. Each was fixed test first.
node stayed JOINING. The puller now discards the local prepare under the commit guard and replays the peer's
committed copy.
releasing its restored row locks. Every later change to those rows was refused on replay, so the node kept a
stale value for an ACKed row. The repair now cleans up like any commit.
waited on the writer that the same transaction's pinned session holds. A claim-only commit no longer flushes.
answers 1213 and ends the transaction. A later COMMIT commits nothing. Before, it answered 1105 and COMMIT
applied the earlier statements.
XsyncTransactionStoreupdates are copy-on-write. This fixes a data race with the stale-transaction GCthat is also on master.
ca2ca4atest(cluster): rewrite the integration suite to run in under two minutesThe cluster suite took about 11.5 minutes and had flaky tests. It is rebuilt on one harness:
prints log tails on failure.
votes. The votes are read over gRPC, including from the old-release node.
time.Sleepleft is inside the poll helper, and every wait has adeadline. SIGKILL and SIGTERM stay real.
The 23 tests over 10s are rewritten from scratch, and a coverage ledger maps each guarded behaviour to its
replacement.
unit test that fails under the same mutation pins it.
Flaky tests, root-caused:
TestGarbageCollectionIntegration -race. A product race, fixed above.718e334fix(autoinc): a narrow insert waits out the startup vote hold instead of failing at onceA narrow insert through a node that had just started, before it released its claim votes, got 1205 at once: the
node's own prepare declined the claim and the claim path gave up without waiting. Any application started together
with the cluster could fail on its first insert.
ErrLocalVotesHeld. The claim waits for the release, which is signalled,not polled. The wait is bounded by
transaction.lock_wait_timeout_seconds, and the claim returns 1205 only whenthat expires.
there would deadlock with a joiner's snapshot checkpoint. An insert that is the first write of its transaction,
the common ORM shape, holds no writer and waits.
276e96ftest: describe narrow AUTO_INCREMENT tests by behaviour, not by one applicationTest names, fixtures and comments now describe the general behaviour they check (a narrow insert as the first write
of an ORM-style transaction; a client starting during the vote hold). Statements and assertions are unchanged.
e5be232fix(gc): gc_max_retention_hours = 0 means no maximumThe config validation and the tombstone horizon treat 0 as unlimited, but the log GC computed its cutoff as "now",
so every pass deleted every committed log entry whatever peers had consumed, forcing snapshot fallbacks. The same
behaviour is on master. With max <= 0 only the safe position and min retention now delete entries; a new test failed
before the fix.
ea2bae8docs: describe the log-pull anti-entropy, row versions and narrow idsreplication.mdxdescribed delta sync as the catch-up mechanism. They now describe the per-peer logpull, applied markers, snapshot fallback, GC by consumed position, the single commit timestamp, per-row versions,
persisted membership and narrow AUTO_INCREMENT beside BIGINT's compact/extended ids.
reference.mdxdocuments what the legacydelta_sync_*keys still control (read-replica bootstrap and GCvalidation only), the two autoinc interval keys,
[metastore] strict_prepare_syncand retention 0.metrics.mdxdocuments the log-pull metrics and drops a delta sync counter that never existed.replication.mdxalso corrects older errors: a nonexistent/intent/key, node states and gossip timeouts, and theTxnID, LWW and HLC examples.
a02b9e6fix(cluster): a joiner's promotion to ALIVE wakes a held seed's claim mergeA joiner refutes the seed's JOINING record as ALIVE before it marks itself JOINING, and its later promotion then read
as ALIVE -> ALIVE on the seed, so the immediate merge never fired and a held seed waited a full merge interval. An ALIVE
node announcing ALIVE at a higher incarnation now signals the merge. This was the root cause of a flaky hold-wait
cluster test (now 20/20). The harness also sets
batch_commit.max_wait_ms = 1, since its clients write one statementat a time; the suite dropped from about 130s to about 106s.
91bf5cffeat(metrics): report replication lag per peer from the log pullmarmot_v2_replication_lag_txnswas registered but never set. It now holds, per peer, the committed transactions pastthis node's pull cursor, summed over databases, as of the last pull round (an upper bound on remaining catch-up).
It counts entries rather than subtracting sequence numbers, which skip on restart. A round stopped at the page cap asks
the peer for its remainder through a new optional field; the count reads keys only, honours cancellation and never
fails the listing. The gauge is never set to a false 0, and a removed peer's series is deleted. Reviewed: PASS.
End-to-end check with a real application
The fixes above are general. As an end-to-end check, LLDAP 0.6.3 (the application in #170, which reads group ids
into an
i32) was run against Marmot in Docker on the final build. Every run passed:Every id fits its declared width, no id is issued twice across nodes, and no application log contains an error.
Compatibility and operations
BIGINTand 64-bit id columns are unchanged.CREATE/DROP DATABASEare refused with a retryable 1213 until every member runs this release. DML keeps working.up to the lock wait. Inside a transaction that already wrote, it returns a retryable 1205 at once.
Testing
All results below are from local runs on this branch.
Unit and static checks
./...with the required build tags, run twice at every freeze: all 28 packages pass.-raceon cfg, coordinator, db, grpc and protocol: clean. The xsync transaction-store race, which also occurs on master, is fixed in98eafbc.go vetreports only 2 findings, both also on master:grpc/forward_session.go:356anddb/vector_index_manager_test.go:315.Cluster suite (after
ca2ca4a)go test -p 1 ./test/three times in a row atca2ca4a: 76 tests pass each time; twice more at718e334(117–119s). One is skipped; it is a subprocess helper.Mutation testing
Regression tests (each fails on the code before its fix)
Benchmarks
Local, alternated A/B runs against the parent commit.
BenchmarkUserDatabaseWriteCommitBenchmarkAutoIncrementInsertPrepared-statement INSERT throughput was not measured; there is no benchmark for it yet.
Known limits and follow-ups (not in this PR)
Retryable errors
such lifetime limit.
Failure model
release their claim votes before hearing from the surviving member.
Documented limits
CREATE new; INSERT ... SELECT; DROP; RENAME) drops that table's tombstones.Test coverage
Pre-existing, unrelated to this PR
REPLACEsemantics, for replicated inserts and for the versioned update path, delete rows that conflict on a secondary unique index without checking their versions.pebble.NoSync, so a SIGKILLed node can lose its unsynced tail.MarkJoiningdoes not bump the incarnation, so peers that already saw a nodeALIVE keep it ALIVE while it catches up.
go test -count>1reuses per-test directories; repeated runs use a loop of-count=1.refreshes
lastSeen(grpc/node_registry.go, "Rule 3").LOAD DATA LOCALreports RowsAffected = 1 for any number of rows (coordinator/handler.go,the
handleMutationdefault). Narrow AUTO_INCREMENT tables report the exact count.on any open pinned writer from a client transaction.
with 1213, with no winner. Clients that retry at a fixed interval can livelock; jittered backoff resolves it.
The same code path exists on master; the runtime behaviour there was not verified.
🤖 Generated with Claude Code