Skip to content

Narrow AUTO_INCREMENT ids from cluster range claims; log-cursor anti-entropy; per-row last-writer-wins (fixes #170) - #175

Merged
maxpert merged 17 commits into
masterfrom
feat/narrow-autoinc-ids
Sep 30, 2026
Merged

maxpert merged 17 commits into
masterfrom
feat/narrow-autoinc-ids

Conversation

@maxpert

@maxpert maxpert commented Sep 29, 2026 •

Copy link
Copy Markdown
Owner

Fixes #170.

Summary

Marmot stores every MySQL integer type as SQLite INTEGER, and generated ids came from a 53-bit HLC-derived space. Any application that declares INT AUTO_INCREMENT (or a narrower type) and reads ids into a 32-bit type therefore broke; the issue was reported with LLDAP, but the class is general.

This PR does three things:

  • It keeps each column's declared width.
  • It mints ids for narrow AUTO_INCREMENT columns from cluster-wide range claims, so ids fit the column and no two nodes hand out the same id.
  • It fixes two pre-existing replication bugs that surfaced while testing this under node outages:
    • a lagging or restarted node could miss committed rows;
    • catch-up could let an older row image overwrite a newer one.

Commits

c063286 fix(protocol): deliver AUTO_INCREMENT insert ids and rule errors on every client path

  • Rule errors from query transforms now reach the client with their MySQL code. Before, they were swallowed.
  • One analysis drives both id injection and the AST rewrite.
  • INSERT ... SELECT and VALUES ROW(...) into an injected column are refused with 1235, instead of letting SQLite pick a key.
  • The OK packet's insert id is now per-statement, where before it was read from the shared connection's last-rowid.

1d9d887 fix(cluster): a lone restarted node must not be its own write quorum

  • Membership (node id, address, incarnation and status) is persisted atomically in the data directory.
  • A node restarted while its peers were unreachable no longer computes a membership of one and commits alone.

ef6d388 feat(schema): record declared MySQL integer widths beside SQLite's INTEGER

  • A /*M:<bits>[u][a]*/ marker in the stored CREATE TABLE text keeps the declared width and signedness of TINYINT, SMALLINT, MEDIUMINT and INT.
  • INTEGER PRIMARY KEY still aliases rowid. BIGINT DDL is byte-identical to before.

541fc0f feat(autoinc): cluster-wide range claims for narrow AUTO_INCREMENT columns

  • A node claims a range of ids through a quorum, keyed by (database, table), in a hidden system table that clients cannot write and CDC never replicates.
  • Participants vote under the claim key's intent lock and check that the range fits the column width from their own schema.

7a66596 feat(autoinc): mint narrow AUTO_INCREMENT ids from cluster-wide range claims

  • Inserts into narrow AUTO_INCREMENT columns take ids from the node's claimed range, and a new range is claimed when it runs out.
  • Explicit ids:
    • Above the cluster base: raises the base.
    • At or below it: refused with 1235, unless the id is in the node's own range.
    • Beyond the column width: 1264.
    • When the column is exhausted: 1062.
  • Each claim keeps three floors: committed, seed and merged. Disjointness rests on the committed floor alone.
  • A node that has just started holds its claim votes until its claim bases are merged with its peers. Inserts during that window get a retryable 1205.
  • Ids are never reset by DROP/CREATE of a table or database.

d0d47f0 feat(replication): anti-entropy pulls every peer's commit log from persisted cursors

Delta sync compared watermarks rather than the transactions each peer actually committed, so a lagging or restarted node could miss ACKed rows. It is replaced as follows:

Pulling

  • Each node pulls from every alive peer's local commit log. There are two new RPCs, ListCommittedLog and FetchTransactions.
  • A durable per-(peer, database) cursor, keyed (seq, txnID), advances only over a contiguous prefix of applied entries.

Idempotent replay

  • An __marmot_applied_txn marker is written in the same SQLite transaction as the effect, on every commit path.
  • The schema version lives in __marmot_schema_version, in the DDL's own transaction.

Unfinished local transactions

  • A transaction that a peer's log holds is treated as committed.
  • A locally prepared copy is committed through the local path.
  • A begin-only record is discarded and replayed.
  • A late PREPARE for it is refused.
  • A commit whose captured rows no longer match its durable prepare record is refused.
  • The stale-transaction GC aborts only under the per-transaction commit guard.

Log GC and catch-up

  • The log is garbage-collected by the minimum position every member has consumed, bounded by a maximum retention.
  • A cursor behind the truncation point falls back to a snapshot, then re-applies the node's own log.
  • A node turns ALIVE only once every database is caught up.

Rolling upgrade

  • DDL and CREATE/DROP DATABASE are refused cluster-wide (1213, retryable) while any member does not yet serve the log-pull protocol.
  • Upgrade every node, then run DDL.

Test harness

  • Cluster tests gain a log-based stall watchdog and poll for readiness instead of sleeping.

8914a41 feat(replication): per-row last-writer-wins on the coordinator's commit timestamp

A node catching up from several peers applied each transaction once, but not in commit order. An older row image could overwrite a newer one, or an older INSERT could resurrect a deleted row.

Commit timestamp

  • The coordinator merges every participant's PREPARE-time clock, decides one commit HLC, and sends it on COMMIT (unary and streamed) and in the log listing.
  • Every node commits with that HLC and merges it into its clock.
  • The clock is seeded at startup from the highest recorded commit.

Row versions

  • __marmot_row_version (table, intent key, wall, logical, origin node, deleted) is stamped in the same SQLite transaction as each row change, on every commit path.
  • A change applies only if its version is not older than the stored one.
  • DELETE leaves a tombstone.
  • An UPDATE that moves the primary key tombstones the old key independently of the new one.
  • LOAD DATA rows are stamped per row, in batches under one savepoint.

DDL

  • DROP TABLE removes a table's versions.
  • RENAME re-keys them.
  • A primary-key change clears them.

Tombstone purge

  • Tombstones are purged in bounded chunks through a partial index once they are older than every known member's log retention.

5ebafa0 fix(cluster): a gracefully stopped node rejoins on restart; document narrow AUTO_INCREMENT ids

Graceful restart rejoins. On master, a node restarted after a graceful stop shuts itself down again about half a second later. Its peers keep the LEAVING record from its own departure, and the node reads that as a decommission.

  • The fix: a restarted node resumes one incarnation above its persisted one.
  • Incarnation bumps from refutation and from turning ALIVE are persisted, so a restart never comes back at or below an incarnation it already announced.
  • REMOVED is never refuted.
  • An admin decommission of a live node still takes effect, including while that node is still catching up.
  • A node restarted on a wiped data directory after a graceful stop keeps the old behaviour. The operations guide says to treat it as a replaced node.
  • There are unit tests and a graceful stop/restart cluster test.

Docs

  • compatibility.mdx: narrow AUTO_INCREMENT behaviour, including every error code above.
  • operations.mdx: the restart note and the procedure for changing membership one node at a time.
  • sql-reference.mdx: corrects the INT AUTO_INCREMENT row, which said BIGINT.
  • A new application compatibility page, lldap.mdx, for a tested application (see the end-to-end check below).

98eafbc fix(replication): recover cleanly from a kill mid-commit; end a rolled-back explicit transaction

The rewrite below exposed these defects. Each was fixed test first.

  • Stuck prepared transaction. A local prepare whose captured rows were lost to SIGKILL wedged catch-up, and the
    node stayed JOINING. The puller now discards the local prepare under the commit guard and replays the peer's
    committed copy.
  • Leaked row locks. A transaction committed just before SIGKILL had its commit record rewritten at open without
    releasing its restored row locks. Every later change to those rows was refused on replay, so the node kept a
    stale value for an ACKed row. The repair now cleans up like any commit.
  • Claim deadlock. A narrow claim inside an explicit transaction deadlocked on the batch-committer flush, which
    waited on the writer that the same transaction's pinned session holds. A claim-only commit no longer flushes.
  • Rolled-back explicit transaction. A pinned statement whose SQLite transaction the lock wait rolled back now
    answers 1213 and ends the transaction. A later COMMIT commits nothing. Before, it answered 1105 and COMMIT
    applied the earlier statements.
  • Race. XsyncTransactionStore updates are copy-on-write. This fixes a data race with the stale-transaction GC
    that is also on master.
  • Config. The AUTO_INCREMENT merge and base-sync intervals become config knobs; the defaults are unchanged.
  • Docs. The operations guide says to rebuild wiped nodes one at a time.

ca2ca4a test(cluster): rewrite the integration suite to run in under two minutes

The cluster suite took about 11.5 minutes and had flaky tests. It is rebuilt on one harness:

  • Binaries are built once per run. Each test gets its own ports and directory, and a cleanup kills its nodes and
    prints log tails on failure.
  • Readiness is checked, not assumed: MySQL answers, every node sees the others ALIVE, and no node holds its claim
    votes. The votes are read over gRPC, including from the old-release node.
  • Fixed sleeps go from 71 to none; the only time.Sleep left is inside the poll helper, and every wait has a
    deadline. SIGKILL and SIGTERM stay real.
  • Assertions check exact rows, values, ids and error codes.
  • Retries are typed and use jittered backoff. None masks the defect a test guards.

The 23 tests over 10s are rewritten from scratch, and a coverage ledger maps each guarded behaviour to its
replacement.

  • Each test that guards a fix fails when that fix is reverted.
  • Where an end-to-end test cannot force the interleaving (per-row LWW, the GC safe position, idempotent replay), a
    unit test that fails under the same mutation pins it.

Flaky tests, root-caused:

  • Rolling upgrade. The harness never waited for the old node's claim votes. It now passes 20/20.
  • TestGarbageCollectionIntegration -race. A product race, fixed above.

718e334 fix(autoinc): a narrow insert waits out the startup vote hold instead of failing at once

A narrow insert through a node that had just started, before it released its claim votes, got 1205 at once: the
node's own prepare declined the claim and the claim path gave up without waiting. Any application started together
with the cluster could fail on its first insert.

  • The local hold is now a typed decline, ErrLocalVotesHeld. The claim waits for the release, which is signalled,
    not polled. The wait is bounded by transaction.lock_wait_timeout_seconds, and the claim returns 1205 only when
    that expires.
  • A peer turning ALIVE triggers an immediate merge round, which shortens the hold.
  • A narrow insert inside a transaction that already holds a pinned writer still fails fast, as before. Waiting
    there would deadlock with a joiner's snapshot checkpoint. An insert that is the first write of its transaction,
    the common ORM shape, holds no writer and waits.

276e96f test: describe narrow AUTO_INCREMENT tests by behaviour, not by one application

Test names, fixtures and comments now describe the general behaviour they check (a narrow insert as the first write
of an ORM-style transaction; a client starting during the vote hold). Statements and assertions are unchanged.

e5be232 fix(gc): gc_max_retention_hours = 0 means no maximum

The config validation and the tombstone horizon treat 0 as unlimited, but the log GC computed its cutoff as "now",
so every pass deleted every committed log entry whatever peers had consumed, forcing snapshot fallbacks. The same
behaviour is on master. With max <= 0 only the safe position and min retention now delete entries; a new test failed
before the fix.

ea2bae8 docs: describe the log-pull anti-entropy, row versions and narrow ids

  • The README and replication.mdx described delta sync as the catch-up mechanism. They now describe the per-peer log
    pull, applied markers, snapshot fallback, GC by consumed position, the single commit timestamp, per-row versions,
    persisted membership and narrow AUTO_INCREMENT beside BIGINT's compact/extended ids.
  • reference.mdx documents what the legacy delta_sync_* keys still control (read-replica bootstrap and GC
    validation only), the two autoinc interval keys, [metastore] strict_prepare_sync and retention 0.
  • metrics.mdx documents the log-pull metrics and drops a delta sync counter that never existed.
  • replication.mdx also corrects older errors: a nonexistent /intent/ key, node states and gossip timeouts, and the
    TxnID, LWW and HLC examples.
  • Each claim was checked against the code by the writers, and 18 were spot-checked by an independent reviewer.

a02b9e6 fix(cluster): a joiner's promotion to ALIVE wakes a held seed's claim merge

A joiner refutes the seed's JOINING record as ALIVE before it marks itself JOINING, and its later promotion then read
as ALIVE -> ALIVE on the seed, so the immediate merge never fired and a held seed waited a full merge interval. An ALIVE
node announcing ALIVE at a higher incarnation now signals the merge. This was the root cause of a flaky hold-wait
cluster test (now 20/20). The harness also sets batch_commit.max_wait_ms = 1, since its clients write one statement
at a time; the suite dropped from about 130s to about 106s.

91bf5cf feat(metrics): report replication lag per peer from the log pull

marmot_v2_replication_lag_txns was registered but never set. It now holds, per peer, the committed transactions past
this node's pull cursor, summed over databases, as of the last pull round (an upper bound on remaining catch-up).
It counts entries rather than subtracting sequence numbers, which skip on restart. A round stopped at the page cap asks
the peer for its remainder through a new optional field; the count reads keys only, honours cancellation and never
fails the listing. The gauge is never set to a false 0, and a removed peer's series is deleted. Reviewed: PASS.

End-to-end check with a real application

The fixes above are general. As an end-to-end check, LLDAP 0.6.3 (the application in #170, which reads group ids
into an i32) was run against Marmot in Docker on the final build. Every run passed:

Scenario Runs
1 node: migrations, users, groups, memberships, password, LDAP bind, restarts 2/2
3 nodes: two app instances on different nodes, one node gracefully restarted, checksums equal on every node 5/5
Same, with the node SIGKILLed 2/2
3 nodes, cold start on a fresh database: migrations replicate, schema equal everywhere 3/3

Every id fits its declared width, no id is issued twice across nodes, and no application log contains an error.

Compatibility and operations

  • Existing BIGINT and 64-bit id columns are unchanged.
  • Rolling upgrade: upgrade every node first. DDL and CREATE/DROP DATABASE are refused with a retryable 1213 until every member runs this release. DML keeps working.
  • Mixed versions: while nodes run different releases, a commit coordinated by an older node has no shared commit timestamp. The row-version guarantee holds once every member is upgraded.
  • Membership: range claims are disjoint while membership is stable. Change membership one node at a time using the documented procedure.
  • Startup: until a node that has just started releases its claim votes, a narrow insert waits for the release,
    up to the lock wait. Inside a transaction that already wrote, it returns a retryable 1205 at once.

Testing

All results below are from local runs on this branch.

Unit and static checks

  • ./... with the required build tags, run twice at every freeze: all 28 packages pass.
  • -race on cfg, coordinator, db, grpc and protocol: clean. The xsync transaction-store race, which also occurs on master, is fixed in 98eafbc.
  • go vet reports only 2 findings, both also on master: grpc/forward_session.go:356 and db/vector_index_manager_test.go:315.

Cluster suite (after ca2ca4a)

  • go test -p 1 ./test/ three times in a row at ca2ca4a: 76 tests pass each time; twice more at 718e334 (117–119s). One is skipped; it is a subprocess helper.
  • Each run took 105–111s, and no test took more than 10s; the slowest took about 6s.
  • Former flaky or failing tests, run 20 times each: 20/20.
    • The rolling upgrade.
    • Explicit transaction across a node outage.
    • Prepared transaction on a killed node.
    • Same rows overwritten across an outage.
  • Before the rewrite, the suite took about 689s and the slowest test 43s.

Mutation testing

  • Every new rule was disabled in turn in a scratch copy, 17 arms for per-row LWW and 18 for its review fixes. Each arm fails at least one test.

Regression tests (each fails on the code before its fix)

  • Missing rows:
    • a node down under load across two databases;
    • an explicit transaction held across an outage;
    • database operations while a node is down;
    • kill -9 during catch-up;
    • a prepared transaction on a killed node.
  • A stale-transaction GC racing a local commit.
  • Rolling upgrade, run against a binary built from the pre-log-pull commit.
  • Stale rows:
    • out-of-order update, delete, and update-before-insert;
    • a primary-key move;
    • LOAD DATA;
    • two peers listing the same transactions in different orders.

Benchmarks

Local, alternated A/B runs against the parent commit.

Benchmark Result
BenchmarkUserDatabaseWriteCommit No measurable difference
BenchmarkAutoIncrementInsert About 10 ms/op on both sides (dominated by the commit batching delay); about 15–23 more allocs/op
Single-row apply and commit (new micro-benchmark) About 18 µs to 33 µs. The row-version stamp is one extra indexed upsert in the same transaction, with no lock and no round trip.
100-row transaction About 0.6 ms, with the stamp statement prepared once per transaction
10,000-row LOAD DATA replay About 9.8 ms before, 23.1 ms after. About 9 ms of the difference is writing 10,000 version rows.

Prepared-statement INSERT throughput was not measured; there is no benchmark for it yet.

Known limits and follow-ups (not in this PR)

Retryable errors

  • An explicit transaction can get a retryable 1213 at COMMIT when its narrow claim goes stale under concurrent claimants on other nodes.
  • An explicit transaction open longer than the lock wait is rolled back and gets 1213. A MySQL transaction has no
    such lifetime limit.

Failure model

  • If two of three nodes lose their data at once, which is more failures than a 3-node cluster tolerates, they can
    release their claim votes before hearing from the surviving member.
    • A narrow insert then fails with 1062. A duplicate id is never stored.
    • The operations guide says to rebuild wiped nodes one at a time.

Documented limits

  • A table rebuild (CREATE new; INSERT ... SELECT; DROP; RENAME) drops that table's tombstones.
  • The read-replica stream applies rows unversioned; it replays one source in order.

Test coverage

  • Admin decommission and removal are covered by unit tests only; there is no cluster test for them yet.

Pre-existing, unrelated to this PR

  • REPLACE semantics, for replicated inserts and for the versioned update path, delete rows that conflict on a secondary unique index without checking their versions.
  • Commit log after SIGKILL.
    • The commit log is written with pebble.NoSync, so a SIGKILLed node can lose its unsynced tail.
    • Its SQLite still holds those rows, but its log can no longer serve them to peers.
    • This is only observable with more than f simultaneous failures.
  • Convergence after more than f failures (not root-caused).
    • Setup: nodes 2 and 3 were SIGKILLed, node 1 was wiped, and all three were restarted.
    • In 2 of 5 runs, a row that node 1 committed with node 3 after the restart never reached node 2.
    • Node 2's pull cursor for node 1 stayed at 0 while it reported itself caught up.
    • With the same scenario inside the failure model (graceful stops, one wipe), every node converges.
  • Joining nodes stay ALIVE to peers. MarkJoining does not bump the incarnation, so peers that already saw a node
    ALIVE keep it ALIVE while it catches up.
  • Harness. go test -count>1 reuses per-test directories; repeated runs use a loop of -count=1.
  • Failure detection. A SIGKILLed node is never marked SUSPECT while its survivors gossip, because every rumour
    refreshes lastSeen (grpc/node_registry.go, "Rule 3").
  • LOAD DATA row count. LOAD DATA LOCAL reports RowsAffected = 1 for any number of rows (coordinator/handler.go,
    the handleMutation default). Narrow AUTO_INCREMENT tables report the exact count.
  • Snapshot and open transactions. A snapshot holds the database manager lock while its TRUNCATE checkpoint waits
    on any open pinned writer from a client transaction.
  • Conflicts with a node down. With 2 of 3 nodes up, two transactions that conflict on one row can both abort
    with 1213, with no winner. Clients that retry at a fixed interval can livelock; jittered backoff resolves it.
    The same code path exists on master; the runtime behaviour there was not verified.

🤖 Generated with Claude Code

Zohaib Sibte Hassan and others added 7 commits September 17, 2026 12:27
…very client path

Step 1 of the narrow-id work: close the defects that would have made the
allocator's behaviour unobservable or wrong before it exists.

D1  Transpile swallowed every ApplyAST error (`if err == nil && applied`).
    Rule errors now propagate verbatim; transform.CodedError carries a MySQL
    code, UnsupportedStatementError maps coded -> own code, other rule
    errors -> 1105, parser failures -> 1064. No string matching.
D2  NeedsIDInjection and ApplyAST decided independently and could disagree.
    One analyzeInsert drives both; column-less positional inserts substitute
    at the auto-increment ordinal (index into non-generated columns, not
    cid); INSERT ... SELECT and INSERT ... VALUES ROW(...) are refused with
    1235 in MySQL's message template instead of letting SQLite assign a key.
D3  The OK packet's insert id came from SQLite's connection-wide last-rowid
    register, shared by every client on the single hookDB connection. It
    now comes from the statement's own CDC entries by column name: first
    generated id, 0 when none. ExecuteLocalWithHooks takes one request.
    ConnectionSession.RecordInsertId is the single home of the rule that 0
    never overwrites LAST_INSERT_ID(); the replica path stored zeros
    unconditionally until now.
D4  Dead code: int_to_bigint rule (never registered), MergeCDCEntries,
    PendingLocalExecution, MustFromWireType, HasPrimaryKey, 37 orphaned
    regex vars and extractTableName in protocol/parser.go, transpile_test.
    parser_vitess.go renamed to vitess_parser_instance.go for its 21 lines.
D5  An unknown wire StatementType panicked the receiving peer through
    MustFromWireType with no recover. protocolStatementFromProto returns an
    error and both replication call sites answer Success:false.

Replica error path (found by review): ForwardQueryResponse now carries
error_code and sql_state (fields 7, 8) so a rule's 1235/42000 reaches a
replica client intact; the leader fills both from the protocol mapper, the
replica raises the typed error. Code 0 keeps 1105/HY000 for an older leader.
The old string-formatted "ERROR 1105 (HY000): ..." also re-classified any
leader message containing "syntax error" as 1064. Replica-raised link
failures report 1105/HY000 (master's observable behaviour); 2003/2013 are
client-library codes and are never put in an ERR packet.

Generated files regenerated with protoc 33.6 (protobuf@33) and
protoc-gen-go v1.36.10; equivalence from the unchanged proto is header-only.

Evidence: red-before-green on master-identical files for D3; 20 mutation
arms fired at named assertions across three review rounds; unit suite
24 ok, race clean on replica/grpc/protocol/coordinator, integration 30/30,
benchmarks within a measured 2-4% noise floor with identical allocs/op.

Not in this change (pre-existing, recorded): three db data races on master
(transaction records shared by pointer); ./db/ no-tag stub build broken on
master; go vet findings db/vector_index_manager_test.go:315 and
grpc/forward_session.go:356; cmd/vec-bench/main_test.go is gitignored;
protocol/handlers/metadata.go:49 still formats an error code into a string.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014u4bCrv6aSsNKiY6oWuvrt
Quorum is a majority of TOTAL membership, but the membership registry
started with only the local node and was never persisted. A node restarted
while its peers were unreachable, or before gossip converged, computed a
membership of one, was its own majority, and committed alone. Prerequisite
F2 of the narrow-id design.

Persisted membership. grpc/membership_snapshot.go stores {node id,
address, incarnation, status} for every known node as msgpack in the node's
data directory, written atomically (temp file, fsync, rename, directory
fsync). The write hangs off updateClusterMetricsLocked, the choke point
every membership mutator already calls and the per-heartbeat path does
not; a fingerprint over the recorded fields skips the fsync when nothing
the snapshot stores changed. An accepted higher-incarnation update that
changes only address or incarnation now marks the record replaced so it is
persisted too (found in review).

Restore at construction. NewNodeRegistryWithDataDir loads the snapshot
before the gRPC server serves 2PC. Restored peers enter as SUSPECT: they
count toward the denominator and are excluded from replication targets
until gossip hears from them. REMOVED stays REMOVED. Incarnations are
restored so SWIM refutation keeps working. A missing or corrupt snapshot
yields a self-only registry with a warning, never a refusal to boot.

Belt. NodeProvider gains HasSeedNodes; GetClusterState refuses QUORUM,
ALL and unknown levels with a retryable 1205 while a seed-configured node
knows only itself. LOCAL_ONE and ONE are exempt: they make no denominator
claim, and refusing them made a booting node fail SELECT 1 and look dead
to liveness probes (caught by the integration suite). Single-node
deployments without seeds are unchanged.

Tests: snapshot round-trip and restore; a seven-step persist walk; fifty
heartbeats write nothing; corrupt/absent/no-data-dir cases; address and
incarnation change persisted and restored; belt table across all levels;
cluster tests TestLoneRestartedNodeRefusesToBeItsOwnQuorum (kill -9,
peers stopped, restart alone: "need 2 (majority of 3 total members, 1
alive)") and TestFreshNodeWithUnreachableSeedsRefusesWrites (1205).
Five mutation arms fired at named assertions. Unit 24 ok, race clean on
grpc/coordinator/cluster, integration 32/32.

Known and accepted: the snapshot fsync (~6 ms measured) runs under the
registry write lock, only on a real membership change; a peer removed
while this node was down inflates the denominator until gossip delivers
REMOVED (fail-closed).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014u4bCrv6aSsNKiY6oWuvrt
…TEGER

Step 3a of the narrow AUTO_INCREMENT id design. Marmot collapses every
MySQL integer type to SQLite's INTEGER so INTEGER PRIMARY KEY keeps
aliasing rowid, which is what LAST_INSERT_ID() depends on. That collapse
lost the declared width, so a 53-bit id could be minted into an INT
column. This change keeps the width where SQLite preserves it verbatim:
in the CREATE TABLE text held by sqlite_master.

Marker grammar. protocol/query/transform/intmarker defines
`/*M:<bits>[u][a]*/` (bits 8/16/24/32, u = UNSIGNED, a = explicit
AUTO_INCREMENT), Encode, a column-scoped Decode over sqlite_master text
that skips string literals, quoted identifiers and table constraints,
Strip, and the signed/unsigned widthMax table. BIGINT emits no marker and
its DDL text is byte-identical to before. INTEGER spelled out is a MySQL
synonym for INT and is marked as 32 bits.

Emit. CREATE TABLE and ALTER TABLE ADD/MODIFY/CHANGE share one collapse,
collapseIntegerTypeWithMarker, emitting `INTEGER /*M:..*/`. The ALTER
path previously never collapsed integer types at all: ADD COLUMN age INT
AUTO_INCREMENT reached SQLite as an unparseable keyword.

Read. The schema cache decodes markers from CreateSQL, never from PRAGMA
table_info, which normalises the comment away; ColumnSchema carries
DeclaredWidth/Unsigned/ExplicitAutoInc and the transpiler SchemaInfo
carries the auto-increment column's width and signedness for the
allocator that follows on this branch.

Client surface. SHOW CREATE TABLE strips the marker; SHOW COLUMNS and
information_schema.COLUMNS read PRAGMA and are unchanged. Known
limitation: during a rolling upgrade an old binary's SHOW CREATE TABLE
shows the marker verbatim for a table created by a new node.

Error codes. protocol/mysqlcode is a leaf package holding the MySQL error
codes and SQLSTATEs so rule packages can name a code without importing
protocol; protocol and transform alias it, and the bare 1064/HY000/08S01
literals in protocol/server.go, mysql_errors.go and error_mapper.go are
replaced.

Tests: encode/decode across width x signedness x autoinc; decode against
hand-written DDL with DEFAULT (expr), string literals, quoted identifiers
and table constraints; rowid aliasing with a marker present; ALTER
shapes; schema-cache read from CreateSQL vs PRAGMA; SHOW CREATE strip;
transpiler schema width/unsigned selection with a decoy narrow column
first. Four mutation arms fired at named assertions. Unit 26 ok, race
clean across protocol/..., integration 32/32, benchmarks with identical
allocs/op.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014u4bCrv6aSsNKiY6oWuvrt
…lumns

Step 3b of the narrow AUTO_INCREMENT id design. A node that must mint ids
for an INT (or narrower) AUTO_INCREMENT column first claims a range of ids
through a quorum, so no two nodes can hand out the same id. Nothing calls
the allocator yet; Step 3c wires it into the insert path. On its own this
commit changes no insert behaviour, and it is not meant to ship without 3c.

Claim protocol. Claim rows live in the existing system database, keyed by
(db, tbl), in a hidden table the client cannot write (1142/42000) and CDC
never replicates. A claim rides an existing statement type with a flag and
a msgpack payload (proto Statement fields 9 and 10), so an older binary
rejects it at the CDC-row gate. A participant votes under the claim key's
intent lock, reading its own stored base and deriving the width ceiling
from its own sqlite_master: accept only when storedBase <= prevBase,
newBase == prevBase, size >= 1 and the range fits the column. A rejection
returns the participant's base (TransactionResponse field 7) and the
coordinator retries above it, pinned to QUORUM, with jittered backoff and
a retry on declined PREPARE rounds, then 1205. COMMIT applies the claim
only if the stored base has not moved, and refuses (never ACKs) when the
claim intent is gone. The system database syncs every commit, with
fullfsync on darwin, so an ACKed base survives an OS crash.

Seeding. DDL seeds the base to max(MAX(col), declared AUTO_INCREMENT=N-1),
raise-only, atomically with the DDL. The width marker gains an optional
floor, /*M:<bits>[u][a][:<floor>]*/. A column whose existing ids already
exceed its width is refused with 1264/22003, and that code now survives
the participant-to-coordinator hop (TransactionResponse field 8). An ALTER
that sets a floor without tagging a column, or sets only AUTO_INCREMENT=N,
is refused with 1235 instead of being silently dropped.

Pre-existing defects fixed on the way, each with a regression test:
- Row locks: a stale GC marker let a later writer evict a live intent
  holder, and emptying a table's row-lock map could drop a lock another
  transaction had just taken. Markers are gone and maps are kept.
- A writer could take over a prepared transaction's lock after the
  heartbeat timeout, so two conflicting DML commits were both ACKed and
  one write was silently lost. Heartbeat eviction is removed; a dead
  coordinator's prepared rows are freed by the stale-transaction GC,
  including transactions recovered from Pebble after a restart. A DML
  COMMIT whose prepared rows are gone is refused.
- Anti-entropy catch-up on a running node swapped a database's files under
  open connections and could replace the system database. It now detaches
  only the database being restored: a SQLite commit hook refuses writes
  from the moment of the detach (reaching clients as 1213/40001), in-flight
  writes are drained, lifecycle operations no longer hold the manager lock
  while stopping GC, and startup restores keep the higher claim base.
- Snapshot exports are reference-counted, so overlapping catch-ups no
  longer leak full database copies; orphans are swept at startup.

Benchmarks (same session, alone, medians): claim apply 3.1 ms with
fullfsync on darwin, 42 us without; a user write commit costs about 1-4%
more for the gate; PREPARE and row-lock paths are unchanged or faster.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014u4bCrv6aSsNKiY6oWuvrt
… claims

Step 3c of the narrow AUTO_INCREMENT id design (issue #170). Inserts into an
INT (or narrower) AUTO_INCREMENT column now take their ids from ranges the
node claimed through a quorum (Step 3b), so ids fit the declared width and
no two nodes hand out the same id while the membership is stable.

Insert path. A per-(db, table) range allocator hands out ids from the
node's claimed range and claims a new one when it is exhausted. An explicit
id above the cluster base raises the base; one at or below it is refused
with 1235 unless it falls in this node's own claimed range. Exhaustion
returns 1062. A bound NULL or 0 is generated. Qualified table names are
handled, prepared statements keep the client SQL, and a read-only replica
forwards the client SQL rather than the transpiled one.

Claim floors. Each claim row keeps three floors: committed (written only
when this node applies a claim), seed (local data: DDL seeding, restore,
rename) and merged (bases learned from peers). PREPARE votes against their
max; COMMIT applies only if committed and seed are at or below the claim's
start. Disjointness of the ranges a node ACKs rests on committed alone, so
a base learned from a peer between PREPARE and COMMIT never refuses the
COMMIT, while a DDL seed in that window still does. Existing system
databases are migrated on open. A refused claim apply is reported as a
typed error, not as node divergence.

Membership. A node whose system database is new or lost holds its vote
until it has merged bases from enough peers; cluster.standalone skips the
hold for a single node. Every unheld node pulls the maximum bases from all
alive peers every 10s, and an authenticated admin sync command does the
same on demand and reports whether each node reached every member. The
guarantee covers a stable membership with this procedure: change one node
at a time, wait until it is ALIVE and unheld, run the sync until it
reports complete, then make the next change. Claims carry the schema
version they need, and every DDL path forgets cached cursors.

Also fixed:
- The schema version was written without sync, so a SIGKILL could lose it
  and leave the node unable to vote after restart.
- INFORMATION_SCHEMA queries read through the read pool, so they no longer
  deadlock the write connection, and prepared IS queries match text ones.
- A gossip reconnect now fires node-alive handling, so a node held while
  its seed was down releases when the seed returns.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014u4bCrv6aSsNKiY6oWuvrt
…rsisted cursors

Delta sync could miss ACKed rows on a node that lagged or restarted: it
compared watermarks, not the transactions each peer actually committed, and
replay did not advance the schema version.

Anti-entropy now pulls, per (peer, database), from each alive peer's local
commit log (ListCommittedLog / FetchTransactions), with a durable cursor
keyed (seq, txnID) that only advances over a contiguous prefix of applied
entries. An exact in-flight tracker bounds what a peer lists. Replay is
idempotent: an __marmot_applied_txn marker is written in the same SQLite
transaction as the effect on every commit path, and the schema version
lives in __marmot_schema_version in the DDL's own transaction. Startup
catch-up uses the same pull, and a node turns ALIVE only when every
database is caught up.

A transaction a peer's log holds is decided COMMITTED: a locally prepared
copy is committed through the local path, a begun-only record is discarded
and replayed, and a late PREPARE for it is refused. A commit whose captured
rows or intents no longer match its durable prepare record is refused, and
the stale-transaction GC aborts only under the per-transaction commit guard.

The log is garbage-collected by the minimum position every member has
consumed, bounded by a maximum retention; a pair whose cursor falls behind
the truncation point restores from a snapshot and re-applies the node's own
log. The database set is reconciled each round with a legacy marker for
pre-upgrade registry rows. DDL and CREATE/DROP DATABASE are refused while
any member does not serve the log pull protocol; upgrade every node, then
run DDL. Delta sync and the snapshot schema-version trailer are removed.

Cluster tests gain a log-based stall watchdog, poll readiness instead of
sleeping, and each finishes within a minute.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014u4bCrv6aSsNKiY6oWuvrt
…it timestamp

A node that caught up from several peers applied each transaction once but
not in commit order, so an older row image could overwrite a newer one, an
older INSERT could resurrect a deleted row, and nodes diverged for good.
No commit timestamp was shared: every node computed its own.

The coordinator now merges every participant's PREPARE-time clock into its
own, decides one commit HLC, and sends it on COMMIT (unary and streamed) and
in the committed log listing. Every node commits with it and merges it into
its clock, including on replay, and the clock is seeded at open from the
highest commit recorded, so a later conflicting commit always gets a larger
version.

Each user database keeps __marmot_row_version (table, intent key, wall,
logical, origin node, deleted). Every commit path stamps each row in the
same SQLite transaction as the change and applies the change only if its
version is not older than the stored one. A DELETE leaves a tombstone, an
UPDATE of an absent row inserts its after-image, and an UPDATE that moves
the primary key tombstones the old key independently of the new one. LOAD
DATA rows are stamped per row the same way, in batches under one savepoint.
DROP TABLE removes a table's versions, RENAME re-keys them, and a primary
key change clears them. Tombstones older than every known member's log
retention are purged in bounded chunks through a partial index.

The stamp costs one extra indexed upsert per row, with no lock and no round
trip. Its statement is prepared once per transaction.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014u4bCrv6aSsNKiY6oWuvrt
@maxpert
maxpert marked this pull request as ready for review September 29, 2026 16:10
Zohaib Sibte Hassan and others added 5 commits September 29, 2026 19:33
…narrow AUTO_INCREMENT ids

A node restarted after a graceful stop shut itself down again about half a
second later. Its peers kept the LEAVING record of its own departure, the
node came back at incarnation 0, and it took that record for an admin
decommission. The cluster harness never saw it because it only killed
nodes.

A restarting node now resumes one incarnation above the one it persisted,
so its ALIVE supersedes its own earlier LEAVING. Incarnation bumps from
refutation and from turning ALIVE are persisted, so a restart never comes
back at or below an incarnation it already announced. A REMOVED view of
self is never refuted, and an admin decommission of a live node still takes
effect, including while the node is still catching up. A node restarted on
a wiped data directory after a graceful stop keeps the old behaviour; the
operations guide says to treat it as a replaced node.

Docs: compatibility.mdx describes narrow AUTO_INCREMENT columns (per-node
ranges, the startup vote hold and its retryable 1205, explicit-id rules
with 1235/1264/1062, no reset after DROP/CREATE); sql-reference.mdx no
longer says INT AUTO_INCREMENT behaves as BIGINT; operations.mdx covers
graceful restart, decommission versus removal, and changing membership one
node at a time; a new application page (LLDAP 0.6.3) records the tested version, the
DATABASE_URL form and the verified 1-node and 3-node scenarios.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014u4bCrv6aSsNKiY6oWuvrt
…d-back explicit transaction

Defects found while rewriting the cluster suite, each fixed test first:

- A local prepare whose captured rows were lost to SIGKILL wedged catch-up:
  the puller retried CommitLocallyPrepared until the txn was marked stuck and
  the node stayed JOINING. It now discards the local prepare under the commit
  guard and replays the peer's committed copy.
- A txn committed just before SIGKILL had its commit record rewritten at open
  without releasing the row locks restored for it, so every later change to
  those rows was refused on replay and the node kept a stale value.
  repairAppliedTxnMetadata now cleans up after the commit it rewrites.
- A narrow AUTO_INCREMENT claim inside an explicit transaction deadlocked on
  the batch committer flush, which waited on the writer the same
  transaction's pinned session holds. A claim-only commit has no SQLite work
  and no longer flushes.
- A pinned statement whose SQLite transaction the lock wait rolled back now
  answers 1213 and ends the transaction, so a later COMMIT commits nothing.
- XsyncTransactionStore mutated shared TxnState in place while the stale
  transaction GC read it. Updates are now copy-on-write.

The AUTO_INCREMENT merge and base-sync intervals become config knobs with
unchanged defaults. The operations guide says to rebuild wiped nodes one at a
time.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014u4bCrv6aSsNKiY6oWuvrt
The cluster suite took about 11.5 minutes and had flaky tests. It is rebuilt
on one harness:

- The binaries are built once per run in TestMain. Each test gets its own
  ports and directory, and cleanup kills its nodes and prints log tails on
  failure.
- Readiness means MySQL answers, every node sees the others ALIVE, and no
  node holds its claim votes; the votes are read over gRPC, including from
  the old-release node in the rolling upgrade. This removes that test's
  flake.
- Fixed sleeps are replaced by polled conditions with deadlines. The only
  time.Sleep left is inside the poll helper.
- Crash tests still use real SIGKILL and SIGTERM, with short intervals.
- Assertions check exact row sets, values, ids and MySQL error codes. Retries
  are typed, use jittered backoff, and never mask the defect a test guards.

The 23 tests that took more than 10s are rewritten from scratch. Each test
that guards a fix fails when that fix is reverted. 76 tests pass in about
108s, and none takes more than 10s.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014u4bCrv6aSsNKiY6oWuvrt
… of failing at once

A client that wrote to a narrow AUTO_INCREMENT table through a node that
had just started, before it released its claim votes, got 1205 "lock wait
timeout" at once: the node's own prepare declined the claim and the claim
path gave up without waiting. An application started together with the
cluster failed on its first insert.

- A held node's own prepare decline is now typed (ErrLocalVotesHeld). The
  claim waits for the release, signalled through common.Broadcast, bounded
  by transaction.lock_wait_timeout_seconds, and returns 1205 only when that
  expires.
- A peer turning ALIVE triggers an immediate merge round, so the hold ends
  as soon as enough members answer instead of on the next interval.
- A narrow insert inside a transaction that already holds a pinned writer
  still fails fast, as before. Waiting there deadlocks with a joiner's
  snapshot checkpoint, which needs that writer. An insert that is the first
  write of its transaction, the common ORM shape, holds no writer and waits.
- The in-process test harness now injects ids as a node does.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014u4bCrv6aSsNKiY6oWuvrt
…pplication

Test names, fixtures and comments said "LLDAP shape" where they test
general behaviour: a narrow insert as the first write of an ORM-style
transaction, and a client that starts while a node still holds its claim
votes. They now say so. Statements and assertions are unchanged.

The LLDAP page no longer says to wait for the vote release before
starting: inserts wait for it.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014u4bCrv6aSsNKiY6oWuvrt
@maxpert
maxpert force-pushed the feat/narrow-autoinc-ids branch from 66493c2 to 276e96f Compare September 30, 2026 00:34
Zohaib Sibte Hassan and others added 5 commits September 29, 2026 20:33
… whole log"

The config validation and the tombstone horizon treat a zero max retention
as unlimited, but CleanupOldTransactionRecords computed maxCutoff = now, so
every GC pass deleted every committed entry whatever the members had
consumed. Lagging peers then fell back to snapshots instead of pulling the
log. The same behaviour exists on master.

With max <= 0 the max-retention branch never fires; only the safe position
plus min retention delete entries. Tests that relied on 0 meaning "delete
now" use an explicit 1ms.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014u4bCrv6aSsNKiY6oWuvrt
README and the docs still described delta sync as the catch-up mechanism.
They now describe the per-peer commit-log pull, applied-txn markers,
snapshot fallback, GC by consumed position, the single commit timestamp and
per-row versions, persisted membership, and narrow AUTO_INCREMENT range
claims beside the BIGINT compact/extended ids.

- reference.mdx: what the legacy delta_sync_* keys still control, the two
  autoinc interval keys, [metastore] strict_prepare_sync, and
  gc_max_retention_hours = 0 meaning no maximum.
- metrics.mdx: the log-pull metrics; the delta sync counter never existed.
- replication.mdx also corrects older errors: a nonexistent /intent/ key,
  node states and transitions, gossip timeouts, and the TxnID, LWW and HLC
  examples.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014u4bCrv6aSsNKiY6oWuvrt
… merge

A joiner starts ALIVE at incarnation 0 and refutes the seed's JOINING
record before it marks itself JOINING, and MarkJoining does not bump the
incarnation. Its later promotion (ALIVE at a higher incarnation) therefore
read as ALIVE -> ALIVE on the seed, AliveChanged never fired, and a held
seed waited a full merge interval before releasing its claim votes. With a
long interval a narrow insert stalled; this made
TestNarrowAutoInc_HeldInsertInTransactionDoesNotStallJoin flaky.

An ALIVE node announcing ALIVE at a higher incarnation now signals
AliveChanged (the ALIVE callback still fires only on a status change).

The cluster harness sets batch_commit.max_wait_ms = 1: its clients write
one statement at a time, so every commit waited out the 10ms flush window.
The suite drops from about 130s to about 106s.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014u4bCrv6aSsNKiY6oWuvrt
marmot_v2_replication_lag_txns was registered but never set. It now holds,
per peer, the committed transactions of that peer's log past this node's
pull cursor, summed over databases, as of the last pull round. Log
sequence numbers are not dense (a restart skips a lease), so the value
counts entries rather than subtracting positions.

- Each round counts the listed entries that stay unapplied. When a round
  stops at the page cap, the last page asks the peer for the committed
  entries remaining (optional remaining_committed); the count reads index
  keys only, honours ctx, and a count failure leaves the value unknown
  without failing the listing.
- The gauge is set only when every database's count for the peer is
  known, never to a false 0, and a removed peer's series is deleted.
- The truncateThrough test helper truncates again, so the snapshot
  fallback test exercises the fallback.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014u4bCrv6aSsNKiY6oWuvrt
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014u4bCrv6aSsNKiY6oWuvrt
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Generated AUTO_INCREMENT ids overflow 32-bit application id types (blocks LLDAP and similar apps)

1 participant