diff --git a/.github/copilot-instructions.md b/.github/copilot-instructions.md
index c617efb26..423eefcdc 100644
--- a/.github/copilot-instructions.md
+++ b/.github/copilot-instructions.md
@@ -115,6 +115,52 @@ baseline truly unaffordable, that is a **decision for the engineer driving the w
raised explicitly per the PRIME DIRECTIVE's blocker protocol — never a shortcut an
assistant takes unilaterally in the name of efficiency.
+## OPTION INTEGRITY — enable choices; foreclose only on analytic grounds
+
+**Our job is to enable options.** Unless an option can be shown to have *no possible
+value*, expose it, and give clients the tools to choose it when it applies and to gather
+the data that makes the choice well-founded. The default posture toward a design
+alternative is to keep it and instrument it, not to rank it.
+
+**The reason this repository needs the rule more than most is recorded separately**, in
+[DESIGN-NOTES.md](../DESIGN-NOTES.md) → [The adoption thesis](../DESIGN-NOTES.md#the-adoption-thesis):
+the hardware that would make these tradeoffs measurable is not available here, and the
+people best placed to judge the options are application authors we have not met. Read it
+before arguing that a particular option is safe to drop — the two conditions it names are
+temporary in principle and are not temporary in practice.
+
+**A measurement that failed to realise an option's value is not a finding against the
+option.** It may mean the hardware, the workload, the software configuration, or the
+apparatus could not reach the conditions where the value appears. Saying "we measured it
+and it did not help" is a statement about the *measurement*; converting it into "it does
+not help" is a category error, and it is the one this rule exists to stop.
+
+**This binds hardest where there is literature or prior art suggesting conditional
+applicability.** Queue and ring topologies, cache and NUMA placement, batching strategies,
+and scheduling disciplines all have regimes where each choice wins. An in-repo measurement
+on one machine cannot settle a question the field treats as workload-dependent, and this
+repository's own decisions (for example that one ring per thread is userspace's proxy for
+one ring per CPU, with NVMe queue pairs as the hardware reason) frequently *are* that prior
+art. Contradicting a recorded decision on the strength of one sample's configuration is a
+defect, not a finding.
+
+**What does justify removing or narrowing an option** is a clear analytic result, of the
+kind that can be argued from the code rather than from a run:
+
+- fewer instructions, fewer allocations, fewer I/Os issued;
+- better locality of reference, argued structurally;
+- a smaller support burden or a clearer programming model;
+- a demonstrated *wrong answer* — an option that binds to incidental behaviour, produces
+ overlapping domains, or silently reports success while doing the wrong thing.
+
+The last is the honest ground for most removals here: not "it measured slower" but "it
+computes the wrong thing."
+
+**When an option stays but cannot be shown to pay, say exactly that**, and say what would
+change the answer: which conditions the apparatus could not reach, and what a consumer
+would need to measure on their own hardware. Foreclosing costs a client a choice they may
+have needed; keeping an unproven option costs a paragraph.
+
## Line endings in tool parameters
All text content passed to tpu tools (`content`, `replacement`, `data` in edit ops) is
@@ -661,7 +707,7 @@ When executing checklist items (CHECKLIST.md files):
- **If items must be done together, say so and do it; don't tease apart.** Once you have decided (and recorded in the checklist if the structure is wrong) that two items must land together, commit them together in one commit citing both IDs. Do **not** try to "unthread" a coupled implementation into per-item commits after the fact — that is fiction, not history.
- **Commit immediately after each item.** In mode b (implementing forward), the commit must happen before moving to the next item. In mode a (recording already-finished work), a single commit citing all the item IDs satisfies this.
- **Commit message format: a Conventional Commits subject line, with the checklist trailer in the body.**
- `release-please` (see "Release process" in [DEVELOPMENT.md](DEVELOPMENT.md)) drives every crate's version
+ `release-please` (see "Release process" in [DEVELOPMENT.md](../DEVELOPMENT.md)) drives every crate's version
bump and CHANGELOG **only** from Conventional Commits subject lines (`type(scope)!: summary`); a subject
that doesn't match that grammar is invisible to it, no matter how much checklist work the commit records.
The mandatory `Completed item:` provenance is therefore never the subject line — it moves to the body, and
@@ -1142,6 +1188,65 @@ If a plan exceeds roughly 10 work items or 3 levels of grouping/nesting, checkpo
into a CHECKLIST.md file in the repository before continuing. The goal is that the plan
survives a lost session — if the plan only exists in the chat, it will be lost.
+## RESOLUTION GRADIENT — sharp at the front, deliberately coarse behind, and never manufacture certainty
+
+**A plan is written at decreasing resolution with distance from the present.** The current
+milestone has great resolution. Later milestones are progressively coarser, and that
+coarseness is **correct** — it is not an omission to be closed, and an audit or review pass
+must not treat it as one.
+
+**Why it cannot be otherwise here.** Some work has the shape *build the blocks → build the
+measurement tools → experiment with those tools to infer things*. On such a project the
+later milestones are not merely unwritten, they are **unwritable**: the experiments that
+would resolve them have not happened. The clarity is an **output** of the work, not an
+input being withheld from it. A project small enough to plan end-to-end before
+implementing is a different case, and the distinction is worth making explicitly before
+planning begins.
+
+**The failure mode this exists to stop.** An assistant asks a *very specific* question of
+someone who holds a *general sense* of the direction. The specificity of the question
+implies an answer of matching precision is available, so one is produced — at low
+confidence. It is then recorded as a decision, and it lands in the wrong milestone, or in
+the wrong order, and later work binds to it. **A low-confidence answer recorded as a
+decision is worse than no answer**, because the uncertainty that surrounded it is now
+invisible to everyone downstream.
+
+Four rules follow:
+
+1. **Calibrate the question to the resolution actually available.** Ask whether the general
+ direction is right before asking which of five options to take. If a question would only
+ be answerable *after* work that has not been done, it is not yet a question — it is a
+ description of that work.
+2. **Make "too early to say" a first-class, explicitly offered answer.** When presenting
+ options, say plainly that leaving it coarse is among them. A question posed without that
+ option is a question that forces a choice, and the person answering may not notice they
+ have been forced.
+3. **When an answer arrives hedged, record the hedge.** A direction that is not settled is a
+ **working position, not a decision**: it gets no decision ID, it lives in Tier 2 or Tier 3
+ or a heading that says so, and nothing binds to it. The worked example already in the tree
+ is "Working position on domain counts (not a decision)" in
+ [DESIGN-SESSION-2026-08-30-numa-sharded-io-execution-domains.md](../design-sessions/DESIGN-SESSION-2026-08-30-numa-sharded-io-execution-domains.md).
+4. **Prefer questions that unblock the current milestone.** If the answer would not change
+ what happens next, asking now mostly converts uncertainty into a record of false
+ precision.
+
+**A deferral is productive, not merely protective.** Naming a deferral is usually read as
+"we avoided building on a guess", which is true and is the smaller half. The larger half is
+that it **buys the interval in which the answer becomes derivable** — the blocks get built,
+the instruments get written, the experiments get run, and the answer that was unavailable
+becomes obvious. So when a deferral discharges, do **not** write it up as though the answer
+existed all along and was waiting to be stated. Say what in the interval produced it. The
+difference matters because the first framing quietly teaches that asking earlier and harder
+would have worked, which is exactly the behaviour rule 1 forbids.
+
+**This does not soften the PRIME DIRECTIVE, and the two must not be confused.** They govern
+different objects. The PRIME DIRECTIVE forbids deferring **work** because no consumer for it
+is currently visible; this rule forbids manufacturing **decisions** the work has not yet made
+available. Building a capability nothing calls yet is required; inventing a specific answer to
+a question the experiments have not reached is not. When they appear to collide, the test is
+whether the thing being deferred is *work you could do now* — if it is, do it, and the
+gradient has nothing to say about it.
+
## Design notes are not a work queue
Design notes (DESIGN-NOTES.md, DESIGN-RATIONALE.md, and related files) record *decisions*
@@ -1304,7 +1409,7 @@ sites in three wordings.
**This is the data-side twin of rule 1.** Rule 1 says define a fact once in code and have everything
ask. This says the same of measurements: hold the number once, and have prose point rather than
-paraphrase.
+paraphrase. Rule 6 extends it once more, to facts that are *derived* rather than measured.
### 5. Present what was observed; never write the conclusion
@@ -1351,6 +1456,57 @@ crate that happens to publish measurements. Every instance found so far has been
existing decision rather than a gap in it. Apply it while writing: no checker can find these,
because nothing is inconsistent.
+### 6. Never store a fact another artifact already owns
+
+Rule 4 governs *measured* numbers. This governs every **derived** fact -- anything a reader could get
+from an artifact that is already authoritative for it. Release or publication status, version numbers,
+which milestones are done, whether a branch has landed, how many crates or tests or files there are.
+Writing one into prose creates a second copy whose only maintenance mechanism is somebody remembering,
+and remembering is what fails.
+
+The tell is that **the copy cannot be wrong at the moment it is written.** It is accurate -- that is why
+it gets written -- and nothing will ever say when it stopped being. A wrong decision gets argued with; a
+stale derived fact is simply believed.
+
+- **Delete rather than update.** When you find a stale derived fact, correcting it is almost never the
+ fix: it re-arms the identical hazard with a fresh date on it. Remove the claim and link the artifact
+ that owns the answer.
+- **Removing the digits is not enough.** "Published at 0.3.1" and "is published" are both copies of the
+ release state; only the first is obviously one. Rule 4's "write the claim, not the digits" shrinks the
+ drift surface of a *measurement whose claim is itself the finding*. It does not license storing a
+ derived fact in words.
+- **An absence may be worth one sentence, once.** Where a reader would expect a status section and find
+ none, say the omission is deliberate and name the artifact that answers it -- otherwise somebody
+ helpfully adds it back.
+- **A characterisation of a sibling item is a derived fact too, and this is the clause that was
+ missing.** "`M22` is a testing-heavy milestone", "`M23.1` touches the crate's contract surface",
+ "those tests only use public API" -- each summarises an artifact that already says what it is, and
+ each is wrong the moment that artifact changes or was misread in the first place. **Link the item;
+ do not describe it.** Measured cost of the omission: both examples above are real, both were
+ written into a milestone's rationale in one session, and both were false when checked -- `M22` is
+ example-only and `M23.1` names the *sample's* `contract.rs`, not the crate's.
+ This is the harder half of the rule to apply, because such a claim arrives as a *subordinate
+ clause supporting an argument* rather than as a statement of fact. "X, because Y is Z" reads as
+ connective tissue; `Y is Z` is nonetheless an assertion about the tree, and the reflex that fires
+ on "I am about to write a version number" does not fire on it. Treat the word **because**,
+ followed by anything about another file, item or milestone, as the tell.
+- **This does not reach the primary record.** Decisions, measurements, rationale, design intent, and a
+ checklist's own contents are owned here and belong here. The test is simply whether some other
+ artifact is already authoritative: if yes, point at it; if no, this *is* the artifact.
+
+**FAIL FAST rule 6 is the sibling, not a contradiction.** That rule says a claim that counts or
+enumerates repository artifacts must come from a command rather than from recollection. This is the
+prior question -- prefer not to state it at all. Bind it to a command only when the claim must exist
+anyway, such as a test asserting a property of the tree.
+
+Worked example, and the reason this is written down: `windows-ioring-sys`' design notes opened with
+"This crate does not exist yet as compiled code", and its published rustdoc said "Under construction",
+several releases after the first one shipped. The first attempt at a fix replaced both with a carefully
+drift-minimised status paragraph -- no version number, linking `CHANGELOG.md` and the checklists -- and
+that was still wrong, because "is published" is itself a copy of the release state. What the crate's
+status is, is a question `CHANGELOG.md` and the git tags answer. The notes now record that they
+deliberately do not answer it.
+
## FAIL FAST — push every rule to the earliest rung that can enforce it
CONTRACT INTEGRITY above tells you to keep restatements in step. This tells you where to put the
diff --git a/.github/workflows/ci.yml b/.github/workflows/ci.yml
index 3d399375e..b9d4943e2 100644
--- a/.github/workflows/ci.yml
+++ b/.github/workflows/ci.yml
@@ -153,6 +153,23 @@ jobs:
shell: pwsh
run: ./tools/check-borrow-surface.ps1
+ # The same mechanism, for a different population that also grew unnoticed:
+ # lib tests that open a real kernel ring. D-49 found 63 of them, which made
+ # `cargo test --lib` an integration suite wearing a unit suite's name. M24 took
+ # it to 41 and none of those is movable without M26.2, so a zero-check would
+ # fail on day one -- hence an inventory, which fails when the SET changes. An
+ # addition obliges the question "does this need the kernel, or only a
+ # ring-shaped thing?"; a removal is progress and needs only regeneration.
+ # Needs no toolchain -- it only reads files.
+ ring-test-population:
+ name: ioring ring-opening lib tests
+ runs-on: windows-latest
+ steps:
+ - uses: actions/checkout@v7
+ - name: Run check-ring-tests.ps1
+ shell: pwsh
+ run: ./tools/check-ring-tests.ps1
+
build-test:
name: build + test
runs-on: windows-latest
@@ -512,6 +529,35 @@ jobs:
RUST_BACKTRACE: 1
RUST_LIB_BACKTRACE: 1
run: cargo test -p windows-ioring-sys --locked --no-fail-fast
+ # `kernel-seam` (M26.2) is orthogonal to `threadpool`, so the two gates
+ # multiply: `--all-features` above builds the seam only alongside the
+ # threadpool, and every step in this job so far builds the no-threadpool
+ # path only with the seam off. The combination is a published
+ # configuration nothing else selects, which is the same argument that
+ # bought this job -- paid here rather than assumed.
+ - name: cargo clippy (kernel-seam, no threadpool)
+ run: cargo clippy -p windows-ioring-sys --all-targets --no-default-features --features kernel-seam --locked -- -D warnings
+ - name: cargo test (kernel-seam, no threadpool)
+ env:
+ RUST_BACKTRACE: 1
+ RUST_LIB_BACKTRACE: 1
+ run: cargo test -p windows-ioring-sys --no-default-features --features kernel-seam --locked --no-fail-fast
+ # `cargo test` compiles the epoch-log sample as a test harness and never
+ # calls `main`, so until M25.1b nothing ran the program itself. That was
+ # not theoretical: M25.1 converted the log's writer to a strided layout
+ # without the reader, which left the log unreadable while every one of
+ # the example's tests passed.
+ #
+ # M25.1b put the contract checks in tests, which is the rung that runs on
+ # every developer's machine. What a test cannot reach is `main` itself --
+ # its path setup, its error plumbing, and its exit code -- and that is a
+ # published example a consumer runs, so a panic on startup is exactly the
+ # failure worth catching. Release, because the sample takes about a second
+ # there against several in debug.
+ - name: cargo run (epoch-log sample, end to end)
+ env:
+ RUST_BACKTRACE: 1
+ run: cargo run -p windows-ioring-sys --example epoch_log --release --locked
placement-probe-no-serde:
name: windows-placement-probe (no serde feature)
diff --git a/.github/workflows/publish-crate.yml b/.github/workflows/publish-crate.yml
index ed918a497..ade7254e9 100644
--- a/.github/workflows/publish-crate.yml
+++ b/.github/workflows/publish-crate.yml
@@ -4,6 +4,7 @@ name: publish-crate
on:
push:
tags:
+ - 'win-numa-sys-v*'
- 'windows-file-enumeration-sys-v*'
- 'windows-file-watcher-v*'
- 'windows-file-watcher-example-test-harness-v*'
@@ -29,6 +30,7 @@ on:
required: true
type: choice
options:
+ - win-numa-sys
- windows-file-enumeration-sys
- windows-file-watcher
- windows-file-watcher-example-test-harness
@@ -109,7 +111,7 @@ jobs:
- name: Wait for workspace-sibling dependencies on crates.io
shell: bash
run: |
- workspace_crates="windows-file-enumeration-sys windows-file-watcher windows-file-watcher-example-test-harness windows-impersonation-token-sys windows-ioring-sys windows-namespace-request-sys windows-overlapped-io-sys windows-thread-ambient-sys windows-threadpool-sys windows-topology-sys windows-waitable-queues wtf-string"
+ workspace_crates="win-numa-sys windows-file-enumeration-sys windows-file-watcher windows-file-watcher-example-test-harness windows-impersonation-token-sys windows-ioring-sys windows-namespace-request-sys windows-overlapped-io-sys windows-thread-ambient-sys windows-threadpool-sys windows-topology-sys windows-waitable-queues wtf-string"
metadata="$(cargo metadata --no-deps --format-version 1)"
# `tr -d '\r'` is load-bearing on the Windows runner: jq.exe writes
# CRLF, and `read` splits on LF alone, so without this the last field
diff --git a/.release-please-manifest.json b/.release-please-manifest.json
index 2f768ec86..3bec1095c 100644
--- a/.release-please-manifest.json
+++ b/.release-please-manifest.json
@@ -5,6 +5,7 @@
"crates/windows-file-watcher-example-test-harness": "0.1.3",
"crates/windows-impersonation-token-sys": "0.1.1",
"crates/windows-ioring-sys": "0.3.1",
+ "crates/win-numa-sys": "0.0.0",
"crates/windows-overlapped-io-sys": "0.1.3",
"crates/windows-namespace-request-sys": "0.2.1",
"crates/windows-thread-ambient-sys": "0.2.0",
diff --git a/CHECKLIST-io-domains.md b/CHECKLIST-io-domains.md
index 74d7948a3..1f925a8df 100644
--- a/CHECKLIST-io-domains.md
+++ b/CHECKLIST-io-domains.md
@@ -466,6 +466,45 @@ Parked, not pending. Shape recorded so it is not lost, per the `M{n}+` conventio
Carry one constraint from the start: the flush barrier stops at the ring's edge, so **an epoch is
per-domain** and a client spanning two domains needs two flushes and an explicit join.
+ **Musing, recorded not prioritized (the engineer, 2026-09-23).** If a consumer could *tell* this
+ layer which of its files share a **flush regime** -- that is, which commits will contend -- that
+ might be useful. It is the "declare rather than discover" shape `S-3` proposed for the storage node,
+ applied to the co-flush group instead. Note the wording: *not* "which files share a device", because
+ the inference from one to the other is exactly what the handover below calls unsound. Fit it in if
+ it falls out naturally; do not build toward it. It belongs here rather than in `windows-ioring-sys`
+ because co-flush *grouping* is reasoning about durability groups, and that crate has none -- see
+ [windows-ioring-sys/DESIGN-NOTES.md](crates/windows-ioring-sys/DESIGN-NOTES.md#d-54).
+
+ **Handed over from `windows-ioring-sys` M23.2(b) on 2026-09-23 under D-54, and sharpened on the way.**
+ The concept to keep is **flush equivalence**, not device identity: the set of files whose flushes are
+ not independent of one another. What a durability group actually needs to know is whether two of its
+ commits land in the same such set, because that is what makes them contend.
+
+ **Device number is a proxy for that set, and the proxy is not known to be sound.** The obvious
+ model -- a device cache flush is per-device, so two logs on one device contend and a group spanning
+ two devices pays the slower flush -- is the *starting* model, not the finding. The engineer's
+ refinement: a flush group may be **larger than the one physical device of interest**, with Storage
+ Spaces the candidate case, since a virtual disk over a pool need not have per-physical-device flush
+ independence. That is the same shape as the already-recorded `Q6` hazard -- "whether a Storage Space
+ reports honestly or reports a fiction" -- reaching the flush question rather than the placement one.
+ Whether the class can also be *smaller* than a device is open and unexamined.
+
+ **This is the argument for declaring rather than discovering, and it is stronger than convenience.**
+ If the equivalence class cannot be soundly derived from a device number, then a consumer stating it
+ is not a stopgap until discovery is implemented -- it may be the only sound mechanism, with
+ discovery serving as a default that must be overridable. That upgrades the musing above from
+ "might be useful" to "might be the answer", without settling it.
+
+ The instruments exist and answer the *proxy* question today:
+ `IOCTL_STORAGE_GET_DEVICE_NUMBER` and `IOCTL_VOLUME_GET_VOLUME_DISK_EXTENTS`, both written and
+ smoke-tested in
+ [file-handle-numa-spike.rs](crates/windows-ioring-sys/design-sessions/spikes/file-handle-numa-spike.rs),
+ which counts distinct `DiskNumber` rather than extents because a volume extended twice onto one disk
+ is still one device. **All of it unmeasured.** The engineer's working position, hedged: the
+ FUA-to-Flush conversion has pushed devices toward better flush behaviour, so several flushes in a
+ row is suboptimal rather than pathological. **Low priority, and explicitly not a blocker** -- record
+ the concept, do not let it gate progress.
+
## M-inf -- Ungated
- [ ] **M-inf.1** -- The linked and sharded MPSC shapes, if and only if M31.5 shows the array queue's tail
diff --git a/CHECKLIST-ship-topology-and-queues.md b/CHECKLIST-ship-topology-and-queues.md
index 7060a66e7..713a9f8c7 100644
--- a/CHECKLIST-ship-topology-and-queues.md
+++ b/CHECKLIST-ship-topology-and-queues.md
@@ -663,7 +663,7 @@ that previously stood in the way are gone:
every case observed so far. It is weaker than its own comment claims, and the comment must be
corrected even if the check is not.
-- [ ] **SH-4.12** -- **`ring_copy`'s `ByL3` policy restates the partition rule instead of asking for
+- [x] **SH-4.12** -- **`ring_copy`'s `ByL3` policy restates the partition rule instead of asking for
it.** Raised by Copilot at reviews `5116772196` and `5116886015`.
[policy.rs](crates/windows-ioring-sys/examples/ring_copy/policy.rs) selects domains with
`matches!(domain.kind, DomainKind::Cache { level: 3, .. })`. The reshaped topology model makes
@@ -675,6 +675,39 @@ that previously stood in the way are gone:
This is the consumer-side twin of the platform-integrity rule: bind to the specified primitive, not
to the level number that happens to be L3 on today's hardware. The fix renames the policy as well
as changing it, since `byl3` is a user-facing CLI value that would no longer describe what it does.
+ > **DONE 2026-09-22, together with `M20.1`** -- the two were one change: the prose rule and the code
+ > that implements it could not land separately without the doc describing something the code did not do.
+ > Measuring the old filter while replacing it found a shape neither item anticipated: this workspace's
+ > own machine reports an L3 spanning all 16 processors above a real 8-way L2 partition, so `level: 3`
+ > matched, did **not** degrade, and returned one whole-machine domain as a successful cache-aware
+ > partition. Previously coupled to `M20.1` in
+ > [crates/windows-ioring-sys/CHECKLIST.md](crates/windows-ioring-sys/CHECKLIST.md) -- **do this item
+ > first**, then that one. `M20.1` sweeps the L3 rule's prose, which reaches `policy.rs`'s doc comments.
+ > Recorded 2026-09-19: M20's header had asserted that no defect was found in `ring_copy`, which this
+ > item superseded, and neither file said so.
+ >
+ > **`M20.3` is no longer coupled, and landed first (2026-09-22).** That coupling read "it rewrites the
+ > selection arm this test would assert against", which is true only of a test asserting through
+ > `ByL3`. The degraded-fallback tail is shared by all five policies and is not what this item changes,
+ > so `M20.3`'s tests exercise it through `ByNode` and `ByPackage` and pin nothing here. **What this
+ > item still owes is `ByL3`'s own degradation condition** -- currently "no `level: 3` domain", after
+ > this "no `outermost_partitioning_cache()`" -- which belongs in this item's verification, where the
+ > rule being degraded on is the new one. `examples/ring_copy/policy/tests.rs` is where it goes.
+ >
+ > **That follow-up is now `SH-4.12.1` below, rather than a note under a checked item.** Raised by
+ > Copilot review on PR #108: leaving it here left scheduled work marked done, which the
+ > checked-means-done rule exists to prevent. `SH-4.12` stays checked for the change it did make.
+
+- [ ] **SH-4.12.1** -- **Give `ByL3` a degradation test on the new rule.** Spawned from `SH-4.12`,
+ which converted the policy from `matches!(domain.kind, DomainKind::Cache { level: 3, .. })` to
+ asking `outermost_partitioning_cache()`, and whose own completion note recorded that the
+ degradation condition still owed a test.
+
+ **What to assert.** The condition being degraded on is now "no `outermost_partitioning_cache()`",
+ not "no `level: 3` domain", so a test written against the old condition would pass while checking
+ the wrong rule. Assert both directions: a machine that reports a partitioning cache selects
+ domains from it, and one that reports none degrades rather than selecting nothing or panicking.
+ [policy/tests.rs](crates/windows-ioring-sys/examples/ring_copy/policy/tests.rs) is where it goes.
- [ ] **SH-4.13** -- **`ProcessorSet` cannot represent every `u8` processor id, and the public API
cannot uphold both "every processor" and "no abort".** Raised by Copilot across three unresolved
diff --git a/CHECKLIST.md b/CHECKLIST.md
index 1e0d64913..88bec1a49 100644
--- a/CHECKLIST.md
+++ b/CHECKLIST.md
@@ -443,3 +443,25 @@ Ungated work with no identified predecessor deliverable.
unexplained result is not mistaken for a tested one.
- [x] **M-inf.2** -- Archived the eight completed milestone groups in [CHECKLIST-thread-ambient.md](CHECKLIST-thread-ambient.md), leaving only the parked `M26+`. -> [completed 2026-09-17](COMPLETED-CHECKLIST.md#m-inf2)
+
+- [ ] **M-inf.3** -- Migrate the existing fourteen `windows-*` crates to the `win-` prefix that
+ [DESIGN-NOTES.md](DESIGN-NOTES.md#new-crates-take-the-win-prefix) makes the go-forward convention.
+ **Horizon work, deliberately unscheduled**, and the cost is not uniform -- so this item is a
+ decision before it is a rename.
+
+ **Three are nearly free**: `windows-guard-alloc`, `windows-placement-probe` and
+ `windows-platform-probes` carry `publish = false`, so they are a directory move plus path
+ dependencies.
+
+ **Eleven are published, and a published name cannot be renamed.** crates.io has no rename: a
+ move is a *new* crate, a final release of the old name pointing at it, and the old name occupying
+ the namespace permanently -- which is a weaker version of the very collision this convention
+ avoids. Each also touches `release-please-config.json`, the publish workflow's tag patterns,
+ CHANGELOG continuity, and every dependent.
+
+ **And one is the repository's own name.** `windows-threadpool-sys` names both a crate and this
+ repository, so renaming the crate either diverges the two or pulls a repository rename along with
+ it, breaking remotes and every inbound link.
+
+ Decide the shape first -- all at once, unpublished-only, or never for the published ones -- rather
+ than starting with the easy three and discovering the policy afterwards.
diff --git a/Cargo.lock b/Cargo.lock
index 703ffad60..cc399b9db 100644
--- a/Cargo.lock
+++ b/Cargo.lock
@@ -104,6 +104,13 @@ version = "1.0.24"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "e6e4313cd5fcd3dad5cafa179702e2b244f760991f45397d14d4ebf38247da75"
+[[package]]
+name = "win-numa-sys"
+version = "0.1.0"
+dependencies = [
+ "windows-sys",
+]
+
[[package]]
name = "windows-core"
version = "0.100.0"
@@ -166,6 +173,7 @@ name = "windows-ioring-sys"
version = "0.3.1"
dependencies = [
"serde_json",
+ "win-numa-sys",
"windows-guard-alloc",
"windows-overlapped-io-sys",
"windows-sys",
diff --git a/Cargo.toml b/Cargo.toml
index d0007510f..823016ee0 100644
--- a/Cargo.toml
+++ b/Cargo.toml
@@ -2,6 +2,7 @@
[workspace]
members = [
+ "crates/win-numa-sys",
"crates/windows-file-enumeration-sys",
"crates/windows-file-watcher",
"crates/windows-file-watcher-example-test-harness",
diff --git a/DESIGN-NOTES.md b/DESIGN-NOTES.md
index 375aa6611..9633db084 100644
--- a/DESIGN-NOTES.md
+++ b/DESIGN-NOTES.md
@@ -70,6 +70,155 @@ wants to avoid contributing to it. The Windows threadpool types are inherently m
choices up to the developer. The `windows-sys` crate published by Microsoft helps with the basics of the FFI
to the APIs, but does little to help turn the alphabet and phrasebook into a useful programming model.
+## The adoption thesis: why locality and queues are built together, and why nothing is foreclosed early
+
+This is the engineer's strategic intent for the ring, queue, topology and durability
+work. It is a **thesis, not a measurement** -- it is stated here because it decides
+questions that would otherwise be decided by an instinct to prune, and because a reader
+who does not have it will mistake deliberate breadth for indecision. The faithful record
+of how it was framed is
+[DESIGN-SESSION-2026-09-23-adoption-thesis.md](design-sessions/DESIGN-SESSION-2026-09-23-adoption-thesis.md).
+
+### The hardware is moving and the programming model is not
+
+Non-uniform memory has been relegated to very expensive machines. Meanwhile Moore's law
+has plateaued, so the way to more performance is wider multiprocessors, and wider
+multiprocessors are arriving much closer to consumers. **A uniform memory architecture
+only scales so far.** Consumer AMD parts are already non-uniform with a uniform facade in
+front of them -- chiplets, core complexes, and a fabric between them -- and the same
+argument is available for Intel. The non-uniformity is present; what is absent is any
+obligation, or any convenient means, to program for it.
+
+This repository has already met the facade twice, and both observations are measured
+rather than argued:
+
+- A shipping ARM consumer laptop reports **no L3 at all** and **zero** NUMA nodes, with
+ its natural cluster boundary at L2
+ ([D-48](crates/windows-ioring-sys/DESIGN-NOTES.md#d-48)).
+- The machine this workspace is developed on reports an **L3 spanning all 16 processors**
+ above a real 8-way L2 partition. `examples/cache_domains.rs` prints it: `L1: 8`,
+ `L2: 8`, `L3: 1`.
+
+In both cases the coarse, advertised boundary is the one that tells you least.
+
+### Why the concept stayed in the datacenter
+
+NUMA is difficult to program for, and its benefits are difficult to measure against
+whatever you would have written otherwise -- you are comparing against a program you did
+not write. But the barrier that matters most is neither of those. **Today, taking
+advantage of non-uniformity is a significant architectural decision made at the very
+beginning of a system's design.** Who makes that commitment unless they are already
+targeting datacenter-class hardware? The commitment is the gate, and it is placed at the
+moment when the least is known.
+
+Queues and rings have a related but distinct problem. The techniques are of general
+purpose utility and the building blocks exist in quantity, but outside a few small
+domains they are not readily graspable *as a way to structure a system from the
+beginning*. That is the same failure
+[The value is existence, not cleverness](#the-value-is-existence-not-cleverness)
+describes: the correct construction is not within reach, so capable people reach for what
+is.
+
+### The thesis: couple them, and a cycle may start
+
+The proposition is that there is a **virtuous cycle** available here, and that it is
+opened by coupling the two problems rather than solving either alone.
+
+Make rings and queues -- and the Windows `IoRing` -- a reachable way to structure an
+application, and let locality benefits arrive **adaptively out of that structure** rather
+than out of a separate up-front architectural bet. Then, in order:
+
+1. Ordinary application writers adopt the structure because it is a good way to organise
+ work, not because they set out to be NUMA-aware.
+2. I/O-bound writers get what `IoRing` and the epoch durability idiom are worth, which
+ they cannot easily get today.
+3. Designers who genuinely do target NUMA systems get a substrate to build on instead of
+ starting from primitives.
+4. If adoption follows, systems designers have a reason to expose more locality facts at
+ the consumer hardware level.
+5. Which closes the loop: lower-capability hardware becomes able to deliver the same kind
+ of benefit.
+
+Step 4 is the payoff and also the part furthest outside our control. It is named because
+it is the reason the earlier steps are worth doing in this order.
+
+### The honest position today
+
+**Expect low benefit to anything but server-class hardware, and say so.** This is not
+pessimism to be edited out later; it is the current state of the evidence.
+
+The machine this is developed on is logically server-class and is nonetheless **just a
+slice**, which does not exhibit non-uniform memory characteristics at all. So the central
+claim of the thesis is not one we can currently measure. The consequences of that gap --
+what is blocked, what is not, and what to run when a real multi-node machine is available
+-- are worked out in
+[DESIGN-SESSION-2026-08-30-numa-sharded-io-execution-domains.md](design-sessions/DESIGN-SESSION-2026-08-30-numa-sharded-io-execution-domains.md)
+under "Working under a hardware gap", and that analysis is unchanged by this section.
+
+### The mechanism: describe the application, not the machine
+
+**The component that removes the up-front commitment is
+[topology-planner](crates/topology-planner/COMPONENT.md)**, and its shape is the thesis made
+concrete. The developer supplies a sufficiently abstract definition of the application's **input,
+output, and processing code paths** -- a dataflow description, which is a statement about their own
+program and something they must know anyway. From it the planner infers the connectivity and
+directed flow needed to realize that graph, with no machine in hand; then, given a physical machine
+model, it returns **one or more suggested realizations** as specific threads pinned to specific
+processor groups, with a stated number of queues of stated types
+([EP-D-6](crates/topology-planner/DESIGN-NOTES.md#ep-d-6)).
+
+Three properties of that arrangement are what make it answer the barrier named above, rather than
+relocating it:
+
+- **The developer never makes a topology decision.** They describe an application; the locality
+ reasoning happens against a machine model, at a point where the machine is actually known, rather
+ than as a bet taken at design time.
+- **The first stage does not involve a machine at all**, so the application's own structure is
+ stable across every machine it will ever run on, and only the second stage is redone when the
+ machine changes.
+- **The answer is plural.** Several arrangements are usually defensible and they differ in ways the
+ planner cannot rank without knowing what the developer values, so it presents candidates and
+ supplies the means to tell them apart. That is OPTION INTEGRITY at component scale -- the same
+ refusal to convert an absence of evidence into a verdict.
+
+### What follows: no early foreclosure
+
+**Avoid all early foreclosure of techniques that may yield benefits to application
+authors.** Two reasons, each independently sufficient:
+
+1. **We are not the application authors.** We do not know what they will want to do, and
+ a design option removed here is one they cannot reach no matter how well it would have
+ suited them.
+2. **We do not have the hardware.** Significant analysis of performance tradeoffs on real
+ NUMA hardware is not available to us, so a tradeoff we "settle" is settled on a machine
+ that cannot exhibit the phenomenon.
+
+Either reason alone forbids pruning an option on the strength of a local measurement. Both
+together make it the repository's standing posture, stated operationally as **OPTION
+INTEGRITY** in [copilot-instructions.md](.github/copilot-instructions.md) -- which is the
+rule, where this section is the reason for it. The corollary that a consumer must be
+handed *data* rather than a verdict is the same thesis seen from the client's side: an
+application author on hardware we have never seen is exactly the person best placed to
+decide, and they can only do it if we give them the means to measure.
+
+**This section does schedule work**, and so differs from
+[The value is existence, not cleverness](#the-value-is-existence-not-cleverness), which
+deliberately schedules none. The mechanism above is the bulk of it, and it is queued in
+that component's own [CHECKLIST.md](crates/topology-planner/CHECKLIST.md) -- `EP-1+.1` for
+the dataflow description's vocabulary, `EP-1+.5` for the connectivity graph's type, and
+`EP-1+.6` for the plural answer.
+
+What the runtime crates owe is the other end: being **realizable from** a plan they did
+not choose. That is `M27` in
+[windows-ioring-sys/CHECKLIST.md](crates/windows-ioring-sys/CHECKLIST.md), which was first
+written as an adaptivity question for that crate and **re-planned the same day it was
+authored**, because the adaptivity has an owner and it is not there. Answering it in the
+ring crate would have grown a second policy surface beside the planner's -- the
+`outermost_partitioning_cache` defect again, a policy answer landing in a crate whose job
+is something else. [D-8](crates/windows-ioring-sys/DESIGN-NOTES.md#d-8) is untouched by any
+of this: being constructible from a policy decision made elsewhere is the opposite of
+taking one.
+
## The value is existence, not cleverness: "it is only a SMOP" is why it is missing, not a reason to skip it
A governing principle for the whole repository, stated because it decides
@@ -205,6 +354,30 @@ and close routines. The new crate inherits an established concept rather than in
**"Ring" was considered and is wrong for the family.** It is accurate for the array shapes and
false for the intrusive-linked one, which is genuinely not a ring. `queues` covers both.
+## New crates take the `win-` prefix, not `windows-`
+
+**The engineer's decision, 2026-09-23, taken when `win-numa-sys` was proposed.** Crates created
+from now on use a `win-` prefix. The reason is namespace collision: `windows` is Microsoft's, and
+a crate published as `windows-numa-sys` today is a name Microsoft may reasonably want tomorrow.
+Abdicating the prefix costs nothing and removes the risk entirely.
+
+**The `-sys` half is unchanged and is still earned rather than assumed.** It means thin-over-Win32:
+memory-safe over an existing API, adding no policy, per
+[the waitable-queues naming decision](#the-waitable-queues-crate-is-named-plural-and-carries-no-sys-suffix).
+A `win-*` crate that decides something on a consumer's behalf drops the suffix exactly as a
+`windows-*` one would.
+
+**The existing fourteen migrate eventually, and the cost is not uniform.** Eleven of them are
+published to crates.io, and a published name cannot be renamed -- a rename is a *new* crate plus a
+final release of the old name, and the old name persists forever. Three are unpublished
+(`windows-guard-alloc`, `windows-placement-probe`, `windows-platform-probes`) and are nearly free to
+move. One further wrinkle: `windows-threadpool-sys` is also the **repository's** name, so renaming
+that crate either diverges the two or drags the repository rename along with it.
+
+The migration is therefore queued at the horizon rather than scheduled, as `M-inf.3` in
+[CHECKLIST.md](CHECKLIST.md). **This decision schedules no rename now**; what it settles is the
+prefix every *new* crate uses, so the divergence stops growing while the question of the existing
+ones stays open.
## Windows SDK model and constraints
This crate targets the object-based thread pool API (introduced in Windows Vista) rather than the legacy
diff --git a/DESIGN-RATIONALE.md b/DESIGN-RATIONALE.md
index 3b38d8de8..be44c288d 100644
--- a/DESIGN-RATIONALE.md
+++ b/DESIGN-RATIONALE.md
@@ -233,6 +233,60 @@ following the rule that a binding which cannot be shown to fail is cosmetic. Fiv
mutations -- three manifest values, a deleted claim, and a stale version planted in prose --
each produce a distinct, located failure.
+## Why no option is foreclosed while the hardware gap lasts
+
+[DESIGN-NOTES.md](DESIGN-NOTES.md#the-adoption-thesis) records the thesis and
+[copilot-instructions.md](.github/copilot-instructions.md) records the operational rule
+(OPTION INTEGRITY). This is how the rule was reached and what was rejected on the way. The faithful
+record of the engineer's framing is
+[DESIGN-SESSION-2026-09-23-adoption-thesis.md](design-sessions/DESIGN-SESSION-2026-09-23-adoption-thesis.md).
+
+The evidence was a specific over-reach, not an argument in the abstract. A harness comparing three
+commit strategies in the epoch-log sample found no blast-radius difference between one ring and two,
+and the conclusion recorded was that the two-ring strategy's justification was "dead on structural
+grounds". Two things were wrong with it, and they fail differently:
+
+- The harness **could not have shown the difference**. Each lane registers its own arena, so the
+ arena is the limiter rather than the ring topology. The finding was a fact about the apparatus
+ presented as a fact about the design.
+- It contradicted [D-27](crates/windows-ioring-sys/DESIGN-NOTES.md#d-27), which had already committed
+ the crate to multiple rings on the strength of per-CPU NVMe queue pairs. One sample's arena sizing
+ was allowed to overrule a decision made on stronger grounds, and nothing flagged the collision.
+
+The second is the more instructive failure. A repository whose decisions are well measured builds an
+instinct to trust a measurement over a recorded position, and that instinct is right often enough to
+be dangerous: it does not ask whether the measurement's configuration could reach the regime the
+recorded position was about.
+
+Three candidate rules were considered.
+
+**"Prefer the measurement"** is what had been happening, and it is the failure above.
+
+**"Prefer the recorded decision"** inverts the bug without fixing it -- a decision that a measurement
+genuinely falsifies should fall, and this repository has correctly retired decisions that way
+(D-47 withdrew half of D-24 on measured grounds).
+
+What survived distinguishes the two cases by **what the measurement was capable of showing**. A
+measurement that reached the regime and found nothing is evidence; a measurement whose apparatus
+excluded the regime is evidence about the apparatus. The rule then enumerates the grounds that do
+justify foreclosing -- fewer instructions, fewer I/Os, better locality argued structurally, smaller
+support burden, clearer model, or a demonstrated wrong answer -- because "measured slower" is absent
+from that list on purpose, while "computes the wrong thing" is on it and is the honest ground for
+most removals here.
+
+The reason this repository needs the rule more than most is in the thesis: the hardware that would
+make these tradeoffs measurable is not available, and the people best placed to judge the options are
+application authors we have not met. Both conditions are temporary in principle and neither is
+temporary in practice, so the posture has to be encoded rather than remembered.
+
+A corollary was adopted with the rule and is worth separating, because it is the part that changes
+code rather than judgement: when an option is narrowed, **say where the choice still lives**. The
+audit that followed found the rule's own author had withdrawn a cache-level policy on sound grounds
+and then failed to say that selecting a level remained available through the topology API. The
+repair was to make the sample print every level beside the heuristic's pick, which turns the
+justification for the withdrawal into something a reader can see rather than something they are
+asked to accept.
+
## Why a measured figure is asked to have one home
[DESIGN-NOTES.md](DESIGN-NOTES.md#prose-volume-and-error-surface) records the rule; this is how it
@@ -269,7 +323,7 @@ moved here from that file, where it had been written inline: Tier 1 is the curre
section carrying its own motivating question, census procedure and superseded drafts had made the
decision harder to find inside it.
-[Restatement drift](#restatement-drift) explains the mechanism and gives the remedy. This note
+[Restatement drift](DESIGN-NOTES.md#restatement-drift) explains the mechanism and gives the remedy. This note
records something that section does not: a measurement of **where** the drift actually lives, taken
after PR #90's eighteenth review round, and what follows from it about formal specification.
@@ -374,7 +428,7 @@ uniformly to hit a volume target would remove the only prose that has never been
leaving the prose that keeps being wrong in proportion.
**A formal spec's most useful property here is not proof -- it is that prose can point at it instead
-of paraphrasing it.** That is [restatement drift](#restatement-drift)'s first remedy applied one
+of paraphrasing it.** That is [restatement drift](DESIGN-NOTES.md#restatement-drift)'s first remedy applied one
level up: define the protocol once in a form that can be checked, and let every document cite it.
This is the real connection between the two ideas, and it is why they belong in the same
conversation despite fixing different things.
diff --git a/PLANS.md b/PLANS.md
index 2e17f06eb..f4bd0ab7e 100644
--- a/PLANS.md
+++ b/PLANS.md
@@ -20,7 +20,7 @@ plans tracker: [crates/windows-file-enumeration-sys/PLANS.md](crates/windows-fil
| Path to CHECKLIST.md | Status | Brief description | Design Notes |
|---|---|---|---|
| [CHECKLIST-mutation-survivors.md](CHECKLIST-mutation-survivors.md) | not started | Work queued from the workspace-wide cargo-mutants sweep of 2026-09-02, whose findings are kept in [mutation-sweeps/2026-09-02/](mutation-sweeps/2026-09-02/README.md) rather than re-derived -- the run took roughly fourteen hours. 2,792 caught, 1,112 survived, 198 timed out. **The headline numbers mislead in three ways and the README says how**: a timeout in a blocking-API crate is usually a detection that lost its name rather than a gap (measured: one of `windows-waitable-queues`' 120 timeouts fails four tests in 0.00s when re-injected alone), a low score on an executable probe crate is measuring the wrong thing, and three kinds of survivor -- equivalent mutants, unreachable code, and constants that want a `const` assertion -- are not missing tests at all. M1 covers the shipping crates; M2 holds the two crates that are not libraries and whose scope is an engineer's decision; M3 re-runs and prunes rather than hand-editing the tool's output into a second source of truth. | [mutation-sweeps/2026-09-02/README.md](mutation-sweeps/2026-09-02/README.md) |
-| [crates/topology-planner/CHECKLIST.md](crates/topology-planner/CHECKLIST.md) | in progress | **Planned, not built** -- the directory holds a plan and no code, and becomes a crate when M2 begins. Owns the mapping from a stated **goal** plus an abstracted idealized machine description to a set of execution domains: which processors host a domain, where each thread pins, which memory node it allocates from, what channel connects each pair, and where each channel's buffer lives. Filed because that mapping was **unowned**: [CHECKLIST-io-domains.md](CHECKLIST-io-domains.md) M32 lists the contracts "the runtime cannot be written without" and all of them concern the queue, while M33+.1 opens with "one pinned thread, its `IoRing`, its node-local registered pool, its shard" -- presupposing a plan nothing computed. Separate from `windows-topology-sys` because that crate states **facts** and this one applies **policy**; fusing them is what produced `outermost_partitioning_cache`, a policy answer sitting in the facts crate that three consumers then re-derived differently (SH-16.9). M1 was a *requirements* milestone -- it states what the topology must answer, and it fed the locality-model session, which has since concluded as `D-13`..`D-21`. **The component was deferred past PR #56 by direction**, contributing only planning documents there; #56 then closed unmerged on 2026-09-15 and its content is landing in peeled pieces instead, so the component is still unlanded and goes in its own pull request. Per `D-21` the topology reshape lands without it, since `windows-topology-sys` publishes a refined view of what the platform publishes and an adapter absorbs the rest. M2+ and M3+ are parked on that session concluding, and are additionally **awaiting a re-cut**: EP-D-4 and EP-D-5 re-scoped the component into four parts (`topology-model` holding the abstract machine description, the planner's traits and the plan type; `topology-planner`; an inward Windows adapter; an outward realizer), and only M1 has been reconciled with that. EP-1.1 is done and already earned its keep: checking the shard-set query against the model found `Processor::capacity` using `0` as both a valid efficiency class and a "not known" sentinel, which collide on every non-hybrid machine (filed as SH-16.12). | [crates/topology-planner/DESIGN-NOTES.md](crates/topology-planner/DESIGN-NOTES.md), [design-sessions/DESIGN-SESSION-2026-09-02-cache-locality-model.md](design-sessions/DESIGN-SESSION-2026-09-02-cache-locality-model.md) |
+| [crates/topology-planner/CHECKLIST.md](crates/topology-planner/CHECKLIST.md) | in progress | **Planned, not built** -- the directory holds a plan and no code, and becomes a crate when M2 begins. Owns the mapping from a **dataflow description of an application** -- its input, output and processing code paths, the shape settled 2026-09-23 as `EP-D-6` and discharging the deferral `EP-D-4` named -- plus an abstracted idealized machine description, to **one or more suggested** sets of execution domains: which processors host a domain, where each thread pins, which memory node it allocates from, what channel connects each pair, and where each channel's buffer lives. `EP-D-6` also split planning into two stages -- connectivity inferred with no machine in hand, then realization against a machine model -- which adds `EP-1+.5` (the connectivity graph is a value and needs a type and a home) and `EP-1+.6` (the answer is a set, so what a caller receives and what `M2+.2` renders are sets). Filed because that mapping was **unowned**: [CHECKLIST-io-domains.md](CHECKLIST-io-domains.md) M32 lists the contracts "the runtime cannot be written without" and all of them concern the queue, while M33+.1 opens with "one pinned thread, its `IoRing`, its node-local registered pool, its shard" -- presupposing a plan nothing computed. Separate from `windows-topology-sys` because that crate states **facts** and this one applies **policy**; fusing them is what produced `outermost_partitioning_cache`, a policy answer sitting in the facts crate that three consumers then re-derived differently (SH-16.9). M1 was a *requirements* milestone -- it states what the topology must answer, and it fed the locality-model session, which has since concluded as `D-13`..`D-21`. **The component was deferred past PR #56 by direction**, contributing only planning documents there; #56 then closed unmerged on 2026-09-15 and its content is landing in peeled pieces instead, so the component is still unlanded and goes in its own pull request. Per `D-21` the topology reshape lands without it, since `windows-topology-sys` publishes a refined view of what the platform publishes and an adapter absorbs the rest. M2+ and M3+ are parked on that session concluding, and are additionally **awaiting a re-cut**: EP-D-4 and EP-D-5 re-scoped the component into four parts (`topology-model` holding the abstract machine description, the planner's traits and the plan type; `topology-planner`; an inward Windows adapter; an outward realizer), and only M1 has been reconciled with that. EP-1.1 is done and already earned its keep: checking the shard-set query against the model found `Processor::capacity` using `0` as both a valid efficiency class and a "not known" sentinel, which collide on every non-hybrid machine (filed as SH-16.12). | [crates/topology-planner/DESIGN-NOTES.md](crates/topology-planner/DESIGN-NOTES.md), [design-sessions/DESIGN-SESSION-2026-09-02-cache-locality-model.md](design-sessions/DESIGN-SESSION-2026-09-02-cache-locality-model.md) |
| [CHECKLIST-io-domains.md](CHECKLIST-io-domains.md) | in progress | **M30 is complete and archived** in [COMPLETED-CHECKLIST.md](COMPLETED-CHECKLIST.md): the queue crate's name, skeleton and SPSC shape. M31 built the bounded-array MPSC (with a lazily created manual-reset doorbell whose reset cannot be separated from the observation that there is nothing to take -- achieved by ordering plus a re-check rather than by a lock, per D-9 and D-15) and is done but for `M31.6`, the `loom` verification, which is re-homed as `M30.4` in [CHECKLIST.md](CHECKLIST.md). M32 remains: the contract decisions -- ordering, correlation, backpressure among them -- the domain runtime cannot be written without. M33+ parks the runtime itself, the creation-time-affinity thread builder, the namespace `Outcome` extension, the client-side `ThreadpoolWait` fan-in helper, and the durability crate. M-inf holds items each gated on a specific measurement rather than on taste. The N=1 path is the whole first deliverable and depends on no NUMA hardware. | [DESIGN-NOTES.md](DESIGN-NOTES.md), [DESIGN-SESSION-2026-08-30-numa-sharded-io-execution-domains.md](design-sessions/DESIGN-SESSION-2026-08-30-numa-sharded-io-execution-domains.md) |
| [CHECKLIST.md](CHECKLIST.md) | in progress | M19: propagate the 2026-08-27 platform measurements (IoRing registration replaces the table; the completion-port/`IoRing` fork; `runs_long` as the growth mechanism; the measured 512 default maximum) into the crates whose code or documentation currently assumes otherwise. M20: decide the session-independent path form, now that path resolution is measured to follow the impersonated token's logon session. M21: reconcile with the impersonation and enumeration crates that landed during the session. M34 carries the review-driven repairs raised while shipping the placement tool: M34.1 (the reusable sabotage harness) is done, and M34.2 (route the placement tool's output through a sink rather than writing to stdout from many sites) and M34.3 (archive the completed item bodies still carried by the three root checklists) are open. M37: discharge the failable-call standard across the workspace. M30: find out how much of this workspace's algorithm correctness can be machine-checked -- a survey matching each argued-but-unchecked algorithm to a class of tool (TLA+/PlusCal, loom, bounded proof, `const` assertions), one pilot chosen because parameter shrinking makes an untestable property exhaustive, and a named list of what the pilot could not reach, which is the deliverable. Scoped as an instrument for narrowing hand-inspection rather than replacing it, and explicitly not a reversal of [D-31](crates/windows-waitable-queues/DESIGN-NOTES.md#d-31). Also re-homes `M31.6`, the `loom` verification the queue crate promises adopters before 1.0: it was previously untracked, referenced from that crate's design notes, a source file and its sabotage manifest with no live checklist item anywhere, and is now queued as M30.4. M30's rationale is in [DESIGN-RATIONALE.md](DESIGN-RATIONALE.md#machine-checking-what-is-argued) -- Tier 2, because no decision is taken yet; M30.5 is what produces one. | [DESIGN-NOTES.md](DESIGN-NOTES.md#remoting-synchronous-namespace-operations) for M19-M21 and M37; N/A for M30 |
| [CHECKLIST-ship-topology-and-queues.md](CHECKLIST-ship-topology-and-queues.md) | in progress | Release `windows-topology-sys` 0.2.0 and `windows-waitable-queues` 0.1.0, which everything shorter-term depends on. **Both reached crates.io on 2026-09-05**, by a route this file does not describe, since PR #56 closed unmerged; `SH-4.15` owns reconciling M4 with what shipped. Deliberately redundant with [CHECKLIST-io-domains.md](CHECKLIST-io-domains.md): that file plans the design, this one plans the release, and a release has failure modes a design checklist does not surface. Two were found while writing it -- `windows-waitable-queues-v*` is missing from the publish workflow's tag list, so release-please would tag it and nothing would publish it, silently; and `windows-ioring-sys` is published against `windows-topology-sys = "0.1.0"`. (That second finding was later **corrected at SH-2.2**: the pin is a *dev*-dependency, which consumers never resolve, so it obliges a pin update but no release.) M1 settled the public surface before it was public (done, archived); M2 repairs the plumbing; M3 lands the branch; M4 releases; M5 verifies from outside the workspace; M6 is long-running validation and gates the queue crate's release specifically. M7-M13 were seven PR #56 review rounds (done, archived); M14, M15 and M16 are the three later rounds and carry the file's open work -- M15 owns the fix for an ABA hole that ships **disclosed rather than fixed**, so it does not block the release. M16 is the SH-3.1.1 diff review, the first to read the branch as a diff rather than react to a comment: seven findings, six fixed, including a publish-workflow regression this branch had introduced two commits earlier and a soundness hole in the crate about to freeze its API. Its remaining four are blocked on [design-sessions/DESIGN-SESSION-2026-09-02-cache-locality-model.md](design-sessions/DESIGN-SESSION-2026-09-02-cache-locality-model.md), which began by asking whether collapsing a seven-kind, any-depth topology onto a single cache boundary is the right projection and has since settled that presence and observation must be modeled rather than collapsed into an `Option`. **That work gated the merge, and has since discharged**: unlike M14 and M15, which concern a defect in an implementation that can ship disclosed, M16 concerned the shape of the public model `windows-topology-sys` 0.2.0 would publish, and a published model cannot be reshaped without another break. It became the `MMT-*` plan, which has landed; the session has concluded and 0.2.0 shipped the new model, so M3 no longer waits on M16. The file opens with a status table. | [crates/windows-topology-sys/DESIGN-NOTES.md](crates/windows-topology-sys/DESIGN-NOTES.md), [crates/windows-waitable-queues/DESIGN-NOTES.md](crates/windows-waitable-queues/DESIGN-NOTES.md) |
@@ -28,6 +28,7 @@ plans tracker: [crates/windows-file-enumeration-sys/PLANS.md](crates/windows-fil
| [CHECKLIST-thread-ambient.md](CHECKLIST-thread-ambient.md) | in progress | **M22-M29 are complete and archived** in [COMPLETED-CHECKLIST.md](COMPLETED-CHECKLIST.md): `windows-thread-ambient-sys` (a standalone layer that captures a thread's ambient state and applies it on another thread), `windows-namespace-request-sys` (marshalable Win32 namespace call parameter sets, over an entry list audited from three real consumers rather than guessed), `windows-platform-probes` (a durable home for the measurements this workspace's designs rest on), and the defects the audit of those three found. What remains is **pending**: the three `M26+` items were each gated on the namespace-facility design branch reaching `main`, and it has, so the gate has lifted -- reconciling the imported design background, applying the M22.2 narrowing to M21.2, and making the merge-or-delete decision on the duplicated path preparation. The file is deleted outright once those land. | [crates/windows-thread-ambient-sys/DESIGN-NOTES.md](crates/windows-thread-ambient-sys/DESIGN-NOTES.md) |
| [crates/windows-overlapped-io-sys/CHECKLIST.md](crates/windows-overlapped-io-sys/CHECKLIST.md) | not started | M14: finish the contract audit -- categories 1, 2, 6, 8, 9 were not examined -- and sweep `outstanding()` for the advisory-predicate hazard. | [crates/windows-overlapped-io-sys/DESIGN-NOTES.md](crates/windows-overlapped-io-sys/DESIGN-NOTES.md) |
| [crates/windows-platform-probes/CHECKLIST.md](crates/windows-platform-probes/CHECKLIST.md) | in progress | M1 (streaming reports) is done and archived: every probe now writes into the sink as it measures, through a `fmt::Write` adapter that left every `writeln!` call site untouched, and the `catch_unwind`/`resume_unwind` pair is gone because there is no longer a buffer to rescue. Measured with a control: a probe killed part-way through a run keeps its banner and heading on every run of the new build, where the previous build kept nothing -- see [crates/windows-platform-probes/DESIGN-NOTES.md](crates/windows-platform-probes/DESIGN-NOTES.md#d-streaming-report) for the figures. M2 built the report oracle -- one executable definition of the correspondences between a report's prose and NDJSON halves, bound inside the renderers so every test that renders inherits it -- along with a derived fact set and a corpus of report shapes. It is complete; its ten unrelated leftovers -- CI hygiene, a doc repair, probe-prose corrections -- were re-sequenced into M4 (gated on M3) and M5 (gated on nothing). M3 then supersedes its central rule. Re-reading M2's own evidence showed that both defects which motivated the oracle were defects in the ENCODED ROW, not in the relation between two renderings, and that the row published its three diagnostic lists as bare counts -- so a survey reading `"parse_incomplete":1` could not tell a probe self-bug from host flakiness. The row is the machine contract and gets the facts and the invariants; the prose is for a reader and gets review. M3 is complete and archived (ten items): those three fields publish arrays of OBJECTS, each carrying a stable `code` plus the values its variant holds -- `{"code":"partitioning_summary_missing","level":9}` rather than the bare `"partitioning_summary_missing"` of the superseded M3.1 form; the surviving correspondences became invariants over the observation rather than over two renderings, so a rule that reads the diagnostic lists -- which would be a restatement of `verdict()` and blind to a deleted push site -- was rewritten to read the observation; the row is emitted from a typed value through one writer with total escaping, which is the crate's only defence against caller text reaching the mined artifact; the prose oracle and every parser serving it were deleted, and no test extracts structured data from prose anywhere in the crate. Four later items came from reviews and are the more instructive half: three instruments were found asserting less than their names claimed, `BlockingState::ALL` was found to be a census the compiler did not check despite a doc comment claiming it did, and the row's hand-written JSON well-formedness check was measured against a real parser over 1807 generated corruptions -- 159 disagreements, every one a FALSE accept -- and replaced by `serde_json`, after which the remaining hand-written string scanners were deleted too. What remains is M4 (four M2 leftovers M3 gated, now unblocked and re-scoped) and M5 (six ungated hygiene items). | [crates/windows-platform-probes/DESIGN-NOTES.md](crates/windows-platform-probes/DESIGN-NOTES.md#d-streaming-report), [#d-encoded-row-is-the-contract](crates/windows-platform-probes/DESIGN-NOTES.md#d-encoded-row-is-the-contract) |
-| [crates/windows-ioring-sys/CHECKLIST.md](crates/windows-ioring-sys/CHECKLIST.md) | in progress | Memory-safe Rust over the Windows `IoRing` submission/completion ring, as a new crate. M1-M19 are complete (0.2.0 shipped 2026-08-30, restoring availability after all three 0.1.x versions were yanked); M1-M19 are archived. **M20** queues documentation and policy-test repairs from the 2026-08-30 NUMA-sharding measurement, and the pinned-thread `M6+` work stays parked. | [crates/windows-ioring-sys/DESIGN-NOTES.md](crates/windows-ioring-sys/DESIGN-NOTES.md) |
+| [crates/win-numa-sys/CHECKLIST.md](crates/win-numa-sys/CHECKLIST.md) | in progress | Memory-safe Rust over the Windows NUMA APIs, and the first crate to take the `win-` prefix ([DESIGN-NOTES.md](DESIGN-NOTES.md#new-crates-take-the-win-prefix)). Created because `VirtualAllocExNuma` had been written twice in shipping library code, in `windows-ioring-sys` and `windows-placement-probe`, which do not depend on each other -- and a third time before that, in a sample, which `M22.3` had hoisted. `NumaBuffer` moved here; `windows-ioring-sys` re-exports it and supplies the `IoBuf`/`IoBufMut` impls via the orphan rule, which keeps `M6+.6`'s deferred trait-merge decision untouched. **The duplication is not yet removed, only relocated**: `N-1.1` collapses the probe's copy, and `N-1.2` decides whether `QueryWorkingSetEx` observation moves with it. | [crates/win-numa-sys/README.md](crates/win-numa-sys/README.md) |
+| [crates/windows-ioring-sys/CHECKLIST.md](crates/windows-ioring-sys/CHECKLIST.md) | in progress | Memory-safe Rust over the Windows `IoRing` submission/completion ring, as a new crate. M1-M19 are complete (0.2.0 shipped 2026-08-30, restoring availability after all three 0.1.x versions were yanked); M1-M19 are archived. **M20** queues documentation and policy-test repairs from the 2026-08-30 NUMA-sharding measurement, and the pinned-thread `M6+` work stays parked. M21-M24 (epoch-log review, arena and submission, and making the unit suite hermetic) are complete; M23, M25 and M26 are open, and **M27 was added 2026-09-23 as the first milestone queued by intent, and re-planned the same day**: the adaptivity the [adoption thesis](DESIGN-NOTES.md#the-adoption-thesis) asks for is owned by [topology-planner](crates/topology-planner/COMPONENT.md), so M27 is now what this crate owes a realizer rather than a policy surface of its own. | [crates/windows-ioring-sys/DESIGN-NOTES.md](crates/windows-ioring-sys/DESIGN-NOTES.md) |
Add a row here when new work is planned, against [CHECKLIST.md](CHECKLIST.md) or any crate's.
diff --git a/crates/topology-planner/CHECKLIST.md b/crates/topology-planner/CHECKLIST.md
index 4ca82654f..0f229daa1 100644
--- a/crates/topology-planner/CHECKLIST.md
+++ b/crates/topology-planner/CHECKLIST.md
@@ -1,9 +1,10 @@
# Checklist: the topology planner
-Plans an arrangement of execution domains from a stated **goal** plus an **abstracted idealized**
-description of a machine. See [COMPONENT.md](COMPONENT.md) for what this crate is and why it is
-separate from both the topology crate and the runtime, and
-[EP-D-4](DESIGN-NOTES.md#ep-d-4) for the architecture it now sits in.
+Plans arrangements of execution domains from a **dataflow description of an application** -- its
+input, output, and processing code paths -- plus an **abstracted idealized** description of a
+machine. See [COMPONENT.md](COMPONENT.md) for what this crate is and why it is separate from both
+the topology crate and the runtime, [EP-D-4](DESIGN-NOTES.md#ep-d-4) for the architecture it now
+sits in, and [EP-D-6](DESIGN-NOTES.md#ep-d-6) for the two-stage shape and the plural answer.
**The component has been re-scoped**, per [EP-D-4](DESIGN-NOTES.md#ep-d-4) and
[EP-D-5](DESIGN-NOTES.md#ep-d-5). It is named `topology-planner` and the directory now matches; it
@@ -33,7 +34,7 @@ prerequisites rather than on someone else's decision.
| Milestone | State | What it is waiting on |
|---|---|---|
| M1 the input contract | 3 done, 2 open | `EP-1.4` and `EP-1.5`'s coverage half, which want a settled model |
-| M1+ scenario and naming | **partly answered** | the name is settled (EP-D-4); the goal input is deferred for litigation, by direction |
+| M1+ scenario and naming | **partly answered** | the name is settled (EP-D-4); the goal input's *shape* is settled as a dataflow description (EP-D-6), and its concrete vocabulary is `EP-1+.1`'s remaining work |
| M2+ the plan as a value | parked, **and needs re-cutting** | re-cut against EP-D-4/EP-D-5, then the topology reshape landing |
| M3+ the policies | parked | M2+ |
| M-inf parked | ungated | not scheduled, deliberately |
@@ -141,32 +142,64 @@ Raised when the engineer described this component's function, which turned out t
"takes a topology, applies policy". Both are gated on the locality-model session, but neither is a
model question -- they are this component's own.
-- [ ] **EP-1+.1** -- **Describe the scenario input.** The synthesizer takes *two* inputs and only one
- is described anywhere. The scenario says what the caller intends to run, and it is what makes a
- measurement meaningful: [EP-D-3](DESIGN-NOTES.md#ep-d-3) established that a measured number means
- nothing without knowing what it measured, so at minimum the scenario must distinguish small-message
- handoff from large-buffer streaming. Its absence is why "what is most useful for consumers" was
- hard to answer in the abstract for so long.
+**Re-planned 2026-09-23 by [EP-D-6](DESIGN-NOTES.md#ep-d-6)**, which settled the input's shape and
+in doing so added work this milestone did not anticipate: the two-stage split means stage 1 has an
+output that is a value in its own right, and the plural answer means stage 2 returns a set rather
+than a plan. `EP-1+.1` is narrowed accordingly and `EP-1+.5` / `EP-1+.6` are new.
+
+- [ ] **EP-1+.1** -- **Give the dataflow description a concrete vocabulary.** Its *shape* is settled
+ by [EP-D-6](DESIGN-NOTES.md#ep-d-6) -- a sufficiently abstract definition of the application's
+ input, output, and processing code paths -- so what remains is the types: how a caller names a
+ source, a sink and a processing step, and how they express an edge's characteristics. That last
+ part is what makes a measurement meaningful, since [EP-D-3](DESIGN-NOTES.md#ep-d-3) established
+ that a measured number means nothing without knowing what it measured; at minimum an edge must
+ distinguish small-message handoff from large-buffer streaming. **Note the correction EP-D-6
+ carried:** that distinction is an attribute *of an edge*, not the whole scenario, which is how
+ this item originally framed it.
- [ ] **EP-1+.2** -- **Decide what the caller-callback traits ask.** Planning is a negotiation: the
component may call back for clarification the scenario did not settle. Enumerating those questions
is what decides whether this is one trait or several, and it cannot be done before EP-1+.1 says
what the scenario already answers.
-- [ ] **EP-1+.3** -- **Settle the naming, before any type is written.** Both inputs and the output
- are graphs of processors and their relations, so "topology" fits all of them and distinguishes
- none -- and a reader seeing the word twice will eventually take one for the other. Decide whether
- the observed machine keeps the bare name (qualified only by its crate), gains a qualifier, or is
- renamed outright; what the synthesized arrangement is called; and whether the inward/outward
- adapters keep those role names or gain more specific crate/type names. Cheap now; expensive once
- any of those names are public. This one blocks nothing but should not be settled by whoever writes
- the first type.
+- [ ] **EP-1+.3** -- **Settle the naming, before any type is written.** There are now **four** graphs
+ in play and "topology" fits several of them while distinguishing none -- a reader meeting the word
+ twice will eventually take one for the other. They are: the **machine** (processors and their
+ relations), the **application's dataflow description** (sources, sinks and processing steps), the
+ **connectivity graph** stage 1 infers from it (the same nodes, with directed flow), and each
+ **realization** stage 2 emits (threads on processor groups, with queues between them). Note that
+ only two of the four are graphs of *processors*, which is itself a correction:
+ [EP-D-6](DESIGN-NOTES.md#ep-d-6) made the input an application-shaped graph, where this item
+ previously assumed every graph was machine-shaped. Decide whether the observed machine keeps the
+ bare name (qualified only by its crate), gains a qualifier, or is renamed outright; what each of
+ the other three is called; and whether the inward/outward adapters keep those role names or gain
+ more specific crate/type names. Cheap now; expensive once any of those names are public. This one
+ blocks nothing but should not be settled by whoever writes the first type. **The connectivity
+ graph is the one with no name at all today.**
- [ ] **EP-1+.4** -- **Assign measurement ownership for directed residency cost in the four-part
architecture.** [EP-D-3](DESIGN-NOTES.md#ep-d-3) requires directed cross-domain cost input with
measurement context. Record which layer owns collecting, validating, and supplying that measurement
context to `topology-model` through the inward adapter/synthesizer path.
+- [ ] **EP-1+.5** -- **Give stage 1's output a type, and decide which crate holds it.**
+ [EP-D-6](DESIGN-NOTES.md#ep-d-6) establishes that the connectivity graph with its directed flow is
+ derived before any machine exists, which makes it a value that can be inspected and reviewed
+ without even a synthetic machine -- a stronger version of the argument
+ [EP-D-5](DESIGN-NOTES.md#ep-d-5) used to make the plan a value. EP-D-5's placement rule points at
+ `topology-model` for both this type and the dataflow description, on the grounds that a component
+ which only describes or realizes must not depend on planning policy. **EP-D-6 recorded that as an
+ argument and deliberately did not take it as a decision**; this item takes it, either way, and says
+ what depends on the answer.
+
+- [ ] **EP-1+.6** -- **Make the plural answer explicit in the plan vocabulary.** The planner returns
+ *one or more* suggested realizations ([EP-D-6](DESIGN-NOTES.md#ep-d-6)), so the type a caller
+ receives is a set and the thing `M2+.2` renders is a set. Decide what a candidate carries beyond
+ the arrangement itself -- at minimum, enough for a developer to tell two candidates apart and say
+ why they would pick one, which is the whole purpose of returning more than one. Ranking is
+ explicitly *not* in scope here: the planner declines to rank because it cannot know what the
+ developer values, and `M3+` owns whether that ever changes.
+
## M2+: the plan as a value
Parked, not pending. Gated on the topology model landing. Shape recorded so it is not lost, per the
diff --git a/crates/topology-planner/COMPONENT.md b/crates/topology-planner/COMPONENT.md
index 65541d650..c57b5f3d4 100644
--- a/crates/topology-planner/COMPONENT.md
+++ b/crates/topology-planner/COMPONENT.md
@@ -10,22 +10,37 @@ emits a platform-neutral plan, so nothing in it is Windows-specific. See
## What it is
-A **planner**. It takes two inputs and produces a third thing:
+A **planner**. It takes a description of an application and produces arrangements for running it.
-- **a stated goal** -- what the caller intends the arrangement to achieve. Its shape is deliberately
- **deferred for litigation**; that is a named deferral, not an omission.
-- **an abstracted idealized description of a machine** -- processors, memory, storage, interconnects,
- distances and bottlenecks. Not Windows-shaped, and richer than any single platform reports. It is
- **mockable by construction**: a description of a machine nobody has is an ordinary input, which is
- what makes this component testable without the hardware it plans for.
+**The input is a dataflow description** -- a sufficiently abstract definition of the application's
+**input, output, and processing code paths**. It says what the application *is*, not how it should
+be arranged: no domain counts, no queue selections, no topology preferences. See
+[DESIGN-NOTES.md](DESIGN-NOTES.md) -> `EP-D-6`.
-From those it produces **a plan**: which processors host domains, where each thread pins, which
-memory node each allocates from, what channel connects each pair, and where each channel's buffer
-lives. The plan **serializes to JSON** and stays abstracted from Windows.
+**Planning happens in two stages**, and the split is load-bearing rather than incidental:
+
+1. **Connectivity, with no machine in hand.** From the dataflow description, infer the general
+ connectivity and the **directed flow of data** needed to realize that graph. This stage depends
+ only on the application, so its result holds for every machine the application will ever run on.
+2. **Realization, given an abstracted idealized description of a machine** -- processors, memory,
+ storage, interconnects, distances and bottlenecks. Not Windows-shaped, and richer than any single
+ platform reports. It is **mockable by construction**: a description of a machine nobody has is an
+ ordinary input, which is what makes this component testable without the hardware it plans for.
+
+The second stage produces **one or more suggested realizations**: which processors host domains,
+where each thread pins, which memory node each allocates from, how many queues of which types, what
+channel connects each pair, and where each channel's buffer lives. Realizations **serialize to
+JSON** and stay abstracted from Windows.
+
+**The answer is plural on purpose.** Several arrangements are usually defensible on a given machine,
+they differ in ways the planner cannot rank without knowing what the developer values, and
+presenting them as candidates is what lets the developer choose. The planner proposes; it does not
+return the one true arrangement.
**It may ask.** Planning is a negotiation, not a pure function: the component may call back to its
-caller through traits for clarifying information the goal did not settle. Which questions those are
-is not yet known, and knowing them is what decides whether that is one trait or several.
+caller through traits for clarifying information the dataflow description did not settle. Which
+questions those are is not yet known, and knowing them is what decides whether that is one trait or
+several.
## The four components, and which way the arrows point
diff --git a/crates/topology-planner/DESIGN-NOTES.md b/crates/topology-planner/DESIGN-NOTES.md
index 448444bfc..1d3b426b8 100644
--- a/crates/topology-planner/DESIGN-NOTES.md
+++ b/crates/topology-planner/DESIGN-NOTES.md
@@ -22,8 +22,9 @@ renamed to match.
| EP-D-1 | **The shard-set query**: what the planner must know to choose which processors host a domain, and what today's model cannot tell it. |
| EP-D-2 | **The proximity query**: how close two processors are, which selects the channel between their domains. Takes an **unordered** pair; the model has no answer today. |
| EP-D-3 | **The residency query**: where a domain's pool lives, and which side of a cross-domain pair should host a shared ring. **Ordered**, with directed cost entering through the abstract model/adapter path under [D-20](../windows-topology-sys/DESIGN-NOTES.md#d-20). |
-| EP-D-4 | **The four-part architecture, and the planner's name.** The engineer's position: the planner is **`topology-planner`** (no `windows-` prefix); it takes a **goal** description (shape deferred for litigation), queries an **abstracted idealized** model covering processors, memory, storage, interconnects, distances and bottlenecks, and emits a **JSON-serializable, platform-neutral** plan. Two kinds of **adapter** bracket it: one exposing the planner's traits over the Windows topology objects, one **realizing** a plan as buffers, rings and threads with the user's code inserted at the right steps. Settles `MMT-1.5` (the facts crate keeps its `-sys` name), the "two graphs, one word" ambiguity, and where distance lives -- the attributed interconnect shape D-9 sketched goes in the abstract model, so D-9's deferral in the facts crate stands unreopened. |
+| EP-D-4 | **The four-part architecture, and the planner's name.** The engineer's position: the planner is **`topology-planner`** (no `windows-` prefix); it takes a **goal** description (shape settled by [EP-D-6](#ep-d-6)), queries an **abstracted idealized** model covering processors, memory, storage, interconnects, distances and bottlenecks, and emits a **JSON-serializable, platform-neutral** plan. Two kinds of **adapter** bracket it: one exposing the planner's traits over the Windows topology objects, one **realizing** a plan as buffers, rings and threads with the user's code inserted at the right steps. Settles `MMT-1.5` (the facts crate keeps its `-sys` name), the "two graphs, one word" ambiguity (widened to four graphs by [EP-D-6](#ep-d-6)), and where distance lives -- the attributed interconnect shape D-9 sketched goes in the abstract model, so D-9's deferral in the facts crate stands unreopened. |
| EP-D-5 | **The component layout: `topology-model` is its own crate, and dependencies point one way.** The abstract model and the traits the planner queries live in `topology-model`, which the planner and both adapters depend on; non-planner components do not depend on `topology-planner`. Putting the traits in the planner would make a crate whose job is to *describe a machine* depend on one that applies *policy* -- the same defect as `outermost_partitioning_cache`, arriving as a dependency edge instead of an API. Two consequences derived from the same rule rather than decided separately: **the plan type also lives in `topology-model`** (otherwise the realizer depends on the planner), and the inward adapter and the realizer are **separate crates** (their dependency sets barely overlap, and fusing them would make reading a topology pull in the whole runtime). |
+| EP-D-6 | **The goal's shape, and planning as two stages with a plural answer.** Discharges the deferral [EP-D-4](#ep-d-4) named. The caller supplies **a sufficiently abstract definition of the application's input, output, and processing code paths** -- a dataflow description, not a topology preference. From it the planner performs **two inferences in sequence**: first, with no machine in hand, the **general connectivity and directed flow of data** needed to realize that graph; then, given a physical machine model, **one or more suggested realizations** of the graph as specific execution threads pinned to specific processor groups, with a stated number of queues of stated types. The answer is **plural by construction**: the planner proposes candidates for the developer to choose between, and does not return the one true arrangement. |
## EP-D-1: the shard-set query
@@ -264,6 +265,8 @@ Detailed trigger analysis and prior framing are recorded in
## EP-D-4: the four-part architecture, and the planner's name
+**The goal's deferred shape is now settled by [EP-D-6](#ep-d-6).** The rest of this decision stands.
+
*The engineer's position, 2026-09-03. This is a **choice**, not one of M1's queries, and it
re-scopes the component that records it.*
@@ -272,8 +275,9 @@ re-scopes the component that records it.*
**The planner is `topology-planner`** -- deliberately with no `windows-` prefix.
- **Input**: a description of the **goal** of the topology -- what the caller intends the
- arrangement to achieve. Its shape is **explicitly deferred for litigation**, which is a named
- deferral rather than an omission.
+ arrangement to achieve. Its shape was **explicitly deferred for litigation** when this decision
+ was written, which was a named deferral rather than an omission; it has since been settled as a
+ dataflow description by [EP-D-6](#ep-d-6).
- **What it queries**: an **abstracted, idealized** description of the machine, covering
**processors, memory, storage (NVMe), interconnects, distances, and bottlenecks**. Not
Windows-shaped, and materially richer than what any one platform reports.
@@ -391,3 +395,85 @@ the caller did not ask for" rule that decided the layout in the first place.
- **Whether `topology-model` is one crate or eventually two.** The machine description and the plan
vocabulary are different enough that they might separate later. They are together now because
splitting on speculation costs more than merging on evidence.
+
+## EP-D-6: the goal's shape, and planning as two stages with a plural answer
+
+*Recorded 2026-09-23, from the engineer's statement of the component's purpose. Discharges the
+deferral [EP-D-4](#ep-d-4) named as "shape deferred for litigation". Supplies the input half that
+[CHECKLIST.md](CHECKLIST.md) `EP-1+.1` was opened to describe.*
+
+### The decision
+
+The caller supplies **a sufficiently abstract definition of the application's input, output, and
+processing code paths**. That is a **dataflow description**: what comes in, what goes out, and what
+the code does between them. It is deliberately not a topology preference, not a domain count, and
+not a queue selection -- the caller states what their application *is*, not how it should be
+arranged.
+
+From that the planner performs **two inferences, in sequence**:
+
+1. **Connectivity, with no machine in hand.** Infer the general connectivity and the **directed flow
+ of data** needed to realize the described graph. This stage answers what must connect to what,
+ and in which direction, and it is complete before any machine is considered.
+
+2. **Realization, given a physical machine model.** Respond with **one or more suggested
+ realizations** of that graph: specific execution threads pinned to specific processor groups, a
+ stated number of queues of stated types, and the placements the plan already covers under
+ [EP-D-5](#ep-d-5).
+
+### What it settles that was previously open
+
+**The goal is a dataflow description.** [EP-D-4](#ep-d-4) recorded the goal as an input whose shape
+was deferred, and [EP-D-3](#ep-d-3) established that a measured number means nothing without knowing
+what it measured. The dataflow description is what makes a measurement meaningful, because an edge
+in the graph carries what is flowing along it. The small-message-handoff versus large-buffer-
+streaming distinction that `EP-1+.1` named as a minimum bar is therefore **an attribute of an edge**,
+not the scenario in its entirety.
+
+**There is an intermediate artifact, and it is machine-independent.** Stage 1's output -- the
+connectivity graph with its directed flow -- is a thing in its own right, derived before any machine
+exists. [EP-D-5](#ep-d-5) argued the plan should be a *value* so it can be inspected, compared and
+reviewed before anything is pinned or allocated; the same argument applies one stage earlier and
+more strongly, because this artifact can be examined without even a synthetic machine. It follows
+that the graph has a type rather than being an internal step.
+
+**The answer is plural.** The planner returns *one or more* suggested realizations, not the
+arrangement. This is not hedging: on a given machine several arrangements are defensible, they
+differ in ways the planner cannot rank without knowing what the developer values, and presenting
+them as candidates is what lets the developer choose. It is the component-level form of OPTION
+INTEGRITY in [copilot-instructions.md](../../.github/copilot-instructions.md) -- propose options and
+supply the data to choose between them, rather than returning a verdict. A consequence to hold onto:
+whatever renders a plan for a human (`M2+.2`) is rendering a *set*, and the difference between
+candidates is the part a reader needs most.
+
+### Why this is the mechanism the adoption thesis needs
+
+[The adoption thesis](../../DESIGN-NOTES.md#the-adoption-thesis) holds that the principal barrier to
+non-uniformity being exploited outside the datacenter is that exploiting it is an architectural
+commitment demanded at the beginning of a design, when the least is known. **This component is how
+that commitment is removed.** The developer describes their own dataflow -- which they must know
+anyway, and which is a statement about their application rather than about any machine -- and never
+makes a topology decision at all. The locality reasoning happens here, against a machine model, at a
+point where the machine is actually known.
+
+That is also why the two stages are separated rather than fused. Stage 1 depends only on the
+application, so it is stable across every machine the application will ever run on; stage 2 is where
+a machine enters, and is the only part that must be redone when the machine changes. Fusing them
+would make the application's own structure re-derivable only in the presence of a machine, which is
+precisely the coupling the thesis objects to.
+
+### What this does not settle
+
+- **The concrete vocabulary of the dataflow description.** "Input, output, and processing code
+ paths" states the *shape*; the types, and how a caller expresses an edge's characteristics, are
+ `EP-1+.1`'s remaining work.
+- **Where the two new types live.** [EP-D-5](#ep-d-5)'s rule -- a component that only describes or
+ realizes must not depend on planning policy -- applies to both the dataflow description and the
+ connectivity graph, and points at `topology-model` for the same reason it placed the plan type
+ there. That is an argument, not yet a decision, and it is queued rather than taken here.
+- **How many candidates, and how they are ordered.** "One or more" is the contract; whether the
+ planner bounds the set, and whether it orders candidates at all given that ranking is what it
+ declines to do, is `M3+` policy work.
+- **Whether stage 1 can fail.** A description whose connectivity cannot be realized is possible, and
+ nothing here says what happens then. Related to `EP-1.4`, which asks the same question for an
+ unanswered model query.
\ No newline at end of file
diff --git a/crates/topology-planner/DESIGN-RATIONALE.md b/crates/topology-planner/DESIGN-RATIONALE.md
index b81e1f0ce..f5f78a963 100644
--- a/crates/topology-planner/DESIGN-RATIONALE.md
+++ b/crates/topology-planner/DESIGN-RATIONALE.md
@@ -35,3 +35,66 @@ EP-D-4 intentionally left several follow-ups unresolved:
These are tracked as checklist work in [CHECKLIST.md](CHECKLIST.md) as `EP-1.4` (not-observed
behavior), `EP-1+.3` (planner/model/adapter naming), and `EP-1+.4` (measurement ownership), rather
than as canonical decisions.
+
+## Why the goal turned out to be a dataflow description, and the answer plural
+
+[DESIGN-NOTES.md](DESIGN-NOTES.md#ep-d-6) records `EP-D-6`. This is how it was reached.
+
+The goal's shape had been deferred since `EP-D-4` (2026-09-03) and the deferral was **named**, which
+is what made it survivable: `COMPONENT.md`, the checklist status table and the decision body all said
+"deferred for litigation" rather than quietly omitting the input, so nothing was built on a guess in
+the meantime.
+
+**That is the smaller half of what the deferral was worth, and the larger half is easy to state
+backwards.** The answer was not sitting formed on 2026-09-03 waiting to be asked for. It did not
+exist. What produced it was the work done in the interval -- the building blocks, then the
+measurement tools, then the experiments that used those tools to infer things -- and the clarity
+arrived as an *output* of that sequence. The deferral's real value was buying the interval, not
+merely guarding it. Writing this up as "the shape was withheld until 2026-09-23" would invert the
+causality and quietly teach that asking earlier and harder would have worked; it would not have, and
+a specific question put to a general sense would have manufactured a lower-confidence answer that
+then got recorded as a decision. (That failure mode is the subject of RESOLUTION GRADIENT in
+[copilot-instructions.md](../../.github/copilot-instructions.md).)
+
+It was stated on 2026-09-23 by the engineer describing the component's purpose, in the course of
+correcting a milestone that had been written in the wrong crate.
+
+**Two candidate shapes had been implicitly in play, and neither was what was chosen.** `EP-1+.1`
+framed the input as a *scenario* -- "what the caller intends to run" -- with a minimum bar of
+distinguishing small-message handoff from large-buffer streaming. That framing is a set of workload
+*characteristics*, and it would have made the planner's input a bag of tuning hints. The other
+implicit shape was a *goal* in the literal sense, some statement of what to optimize (latency,
+throughput, footprint), which would have made the planner a solver over an objective function.
+
+What was chosen is neither: the input describes **the application's own structure** -- its inputs,
+its outputs, and the processing paths between them. The characteristics `EP-1+.1` named do not
+disappear, but they demote from being the scenario to being **attributes of an edge** in that
+structure, which is a strictly more informative place for them: "large buffers" is not a property of
+a workload, it is a property of a particular flow within it, and a real application has several
+flows that differ.
+
+**The two-stage split was not stated as a separate decision and follows from the input's shape.**
+Once the input is the application's structure rather than a set of hints, the connectivity implied by
+that structure can be derived with no machine present at all -- and a derivation that does not need a
+machine should not be entangled with one. That yields an intermediate artifact that is stable for the
+life of the application, where only the second stage is redone per machine. The alternative, deriving
+connectivity and placement together, would make the application's own shape re-derivable only in the
+presence of a machine, which is the coupling the
+[adoption thesis](../../DESIGN-NOTES.md#the-adoption-thesis) exists to object to.
+
+**The plural answer is the part most likely to be eroded later, so the reason is recorded here.** A
+single returned plan is easier to consume, easier to test, and easier to document, and every one of
+those pressures argues for collapsing the set at some future convenient moment. The reason not to is
+that ranking candidates requires knowing what the developer values, which is the one thing this
+component structurally does not know -- it was given a description of an application, not a statement
+of preference. A planner that returns one arrangement has either acquired a preference it was not
+given or hidden a choice it was not entitled to make. That is the same argument as OPTION INTEGRITY
+in [copilot-instructions.md](../../.github/copilot-instructions.md), arriving at component scale
+rather than at documentation scale.
+
+**What was deliberately not decided**, and is queued instead: where the two new types live.
+`EP-D-5`'s placement rule points at `topology-model` for both, and the argument is recorded in
+`EP-D-6` as an argument. Taking it in the same breath as the decision it follows from would have made
+one decision carry two, and the placement question has a consequence -- whether a caller can hold a
+dataflow description without depending on planning policy -- that deserves to be litigated on its
+own. It is `EP-1+.5`.
\ No newline at end of file
diff --git a/crates/topology-planner/PLANS.md b/crates/topology-planner/PLANS.md
index d515fe976..20aba3f79 100644
--- a/crates/topology-planner/PLANS.md
+++ b/crates/topology-planner/PLANS.md
@@ -2,4 +2,4 @@
| Path to CHECKLIST.md | Status | Brief description | Design Notes |
|---|---|---|---|
-| [CHECKLIST.md](CHECKLIST.md) | in progress | Active checklist maintenance for requirements/planning milestones; code implementation milestones are parked. | [DESIGN-NOTES.md](DESIGN-NOTES.md), [DESIGN-RATIONALE.md](DESIGN-RATIONALE.md) |
+| [CHECKLIST.md](CHECKLIST.md) | in progress | Active checklist maintenance for requirements/planning milestones; code implementation milestones are parked. **M1+ was re-planned 2026-09-23** by [EP-D-6](DESIGN-NOTES.md#ep-d-6), which settled the input's shape as a dataflow description of the application -- discharging the deferral EP-D-4 named -- and established two-stage planning with a plural answer. That narrowed `EP-1+.1` to vocabulary, widened `EP-1+.3` from two graphs to four, only two of which are graphs of processors, and added `EP-1+.5` and `EP-1+.6`. | [DESIGN-NOTES.md](DESIGN-NOTES.md), [DESIGN-RATIONALE.md](DESIGN-RATIONALE.md) |
diff --git a/crates/win-numa-sys/CHECKLIST.md b/crates/win-numa-sys/CHECKLIST.md
new file mode 100644
index 000000000..c5180dd13
--- /dev/null
+++ b/crates/win-numa-sys/CHECKLIST.md
@@ -0,0 +1,44 @@
+# Checklist: win-numa-sys
+
+Memory-safe Rust over the Windows NUMA APIs. See [README.md](README.md) for what the crate is.
+
+## M1 -- Finish removing the duplication the crate was created to remove
+
+The crate exists because two independent `VirtualAllocExNuma` implementations had appeared in
+crates that do not depend on each other. Creating it moved one of them; this milestone removes the
+other, and until it does the duplication is **relocated rather than removed**, which is worth saying
+plainly rather than counting the crate as done.
+
+- [ ] **N-1.1** -- **Collapse `windows-placement-probe`'s allocator onto this crate.**
+ `peer_index_cache.rs` has its own `VirtualAllocExNuma` + `VirtualFree` pair, and that crate does
+ not depend on `windows-ioring-sys`, so it never saw the one that was hoisted in `M22.3`.
+
+ **It is not a drop-in, and the differences are the work.** That allocator also faults every page
+ in -- committed pages are demand-zero, so until something writes to them no physical page has been
+ drawn from the preferred node -- and then *observes* which node the pages actually landed on. It
+ is typed over its own `Slot` rather than bytes, and it has a second origin (an ordinary heap
+ allocation) that shares the same `Drop`.
+
+ Decide what of that is general before moving any of it. Page-faulting looks general: any consumer
+ who cares where pages landed has to do it, and this crate's own documentation already tells them
+ so. The typed element and the dual origin look specific to the probe.
+
+- [ ] **N-1.2** -- **Decide whether `QueryWorkingSetEx` observation moves here**, which `N-1.1` will
+ force a view on. `observed_node_of_region`, `observed_node`, `working_set_flags` and the
+ `working_set` bit constants live in `windows-placement-probe` today. "Which node is this page
+ actually on" is a Windows NUMA concept and passes this crate's bar; against that, those helpers
+ carry their own tests for the bit layout, and `working_set_flags` exists *separately from*
+ `observed_node` precisely so a test can check the layout against a field whose value it already
+ knows. Moving them means moving that care too, not just the code.
+
+ Note what this would make possible, since it is the reason to consider it at all: a caller could
+ then ask this crate whether a placement request was honoured, rather than being told by its
+ documentation that success proves nothing and left to write `QueryWorkingSetEx` themselves.
+
+## M2+ -- Parked
+
+- [ ] **M2+.1** -- **Publish.** The crate is a path dependency of `windows-ioring-sys`, which *is*
+ published, so it has to reach crates.io before that crate's next release or the release fails. It
+ is already registered in `release-please-config.json` and in the publish workflow's tag patterns
+ -- both of which this repository has previously been bitten by omitting, silently -- so what
+ remains is the decision to cut `0.1.0`, not the plumbing.
diff --git a/crates/win-numa-sys/Cargo.toml b/crates/win-numa-sys/Cargo.toml
new file mode 100644
index 000000000..8bae51985
--- /dev/null
+++ b/crates/win-numa-sys/Cargo.toml
@@ -0,0 +1,41 @@
+# Copyright (c) 2026 Mike Grier
+
+[package]
+name = "win-numa-sys"
+version = "0.1.0"
+authors.workspace = true
+edition.workspace = true
+rust-version.workspace = true
+license.workspace = true
+repository.workspace = true
+homepage.workspace = true
+description = "Memory-safe Rust over the Windows NUMA APIs: a VirtualAllocExNuma-backed buffer, a NumaNode newtype, and the node queries a caller needs to decide what to pass. Thin over Win32, with no policy about which node anything should use."
+
+# The first crate to take the `win-` prefix rather than `windows-`; see
+# DESIGN-NOTES.md at the repository root. `windows` is Microsoft's namespace,
+# and a crate published as `windows-numa-sys` today is a name they may
+# reasonably want tomorrow.
+
+[dependencies]
+# `default-features = false` matches every other crate in the workspace.
+#
+# `Win32_System_Memory` is `VirtualAllocExNuma` and `VirtualFree`.
+# `Win32_System_Threading` is `GetCurrentProcess` and
+# `GetNumaHighestNodeNumber` -- the latter lives there rather than under
+# `SystemInformation`, which is where its subject matter would suggest.
+# `Win32_System_IO` and `Win32_System_Ioctl` are `DeviceIoControl` and
+# `FSCTL_QUERY_VOLUME_NUMA_INFO`, for asking a handle's volume which node it
+# reports.
+windows-sys = { version = "0.61.2", default-features = false, features = [
+ "Win32_Foundation",
+ "Win32_System_IO",
+ "Win32_System_Ioctl",
+ "Win32_System_Memory",
+ "Win32_System_Threading",
+] }
+
+[package.metadata.docs.rs]
+# The crate is Windows-only, so docs.rs must build on a Windows target or it
+# would render an almost-empty crate.
+default-target = "x86_64-pc-windows-msvc"
+targets = ["x86_64-pc-windows-msvc", "aarch64-pc-windows-msvc"]
diff --git a/crates/win-numa-sys/DESIGN-NOTES.md b/crates/win-numa-sys/DESIGN-NOTES.md
new file mode 100644
index 000000000..95d9bca80
--- /dev/null
+++ b/crates/win-numa-sys/DESIGN-NOTES.md
@@ -0,0 +1,79 @@
+# Design notes: win-numa-sys (Tier 1)
+
+Current canonical decisions for this crate. See [README.md](README.md) for what it is and
+[CHECKLIST.md](CHECKLIST.md) for what is planned.
+
+Repository-level decisions this crate sits under: the `win-` prefix and what `-sys` promises, both in
+the workspace [DESIGN-NOTES.md](../../DESIGN-NOTES.md#new-crates-take-the-win-prefix).
+
+## Decision index
+
+| ID | Decision |
+|---|---|
+| N-D-1 | **Both ways of arriving at a node are offered; choosing between them is not.** A caller may *declare* a node (`NumaBuffer::new`) or *discover* one (`volume_numa_node`), and may qualify what a discovered answer is worth (`highest_numa_node`). What the crate refuses is a constructor that does both in one step -- there is no `NumaBuffer::for_file`. |
+
+## N-D-1: both ways of arriving at a node, and no shortcut between them
+
+*Recorded 2026-09-23 by `M23.2` in
+[windows-ioring-sys/CHECKLIST.md](../windows-ioring-sys/CHECKLIST.md), which asked this crate's
+predecessor whether it should accept a declared storage node as an input.*
+
+### What is offered
+
+- **Declare.** `NumaBuffer::new(len, Some(node))` allocates on a node the caller names. The caller
+ may have got that node from anywhere -- a topology walk, a configuration file, a measurement, a
+ coin toss. This crate does not ask.
+- **Discover.** `volume_numa_node(handle)` asks a handle's volume which node it reports, and returns
+ it. It does not allocate anything.
+- **Qualify.** `highest_numa_node()` answers how many nodes exist, so a caller can tell a real choice
+ from the only choice available.
+
+Both paths have consumers in this workspace already, which is the evidence that neither is
+speculative: `examples/epoch_log` discovers from its log file's volume, and `examples/ring_copy`
+declares a node it computed from the processor topology.
+
+### What is refused, and why
+
+**There is no `NumaBuffer::for_file(handle, len)`** -- no call that queries a node and allocates on
+it in one step. It is the obvious convenience and it is the wrong shape:
+
+- **It hides the answer.** On a single-node machine, "placed on the node the volume named" and "no
+ preference" are the *same allocation*, so a caller could not tell whether the query had found
+ anything. That is the failure the epoch-log sample's report line exists to prevent, and it is the
+ 2026-08-30 session's *report, do not route* position applied one layer down.
+- **It fuses two failure domains.** Allocation can fail; the query can fail. A combined call has to
+ decide what happens when only the query fails -- allocate unplaced, which silently does something
+ other than asked, or return an error, which fails an allocation that would have succeeded. Neither
+ is right for every caller, so neither should be baked in. Split, the caller decides.
+- **The answer is worth more than one buffer.** A node can inform a thread's affinity, a second
+ allocation, a report, or a plan. Tying the query to one constructor means the second consumer
+ writes it again -- which is how this crate came to exist.
+- **And the convenience is already there.** `NumaBuffer::new(len, volume_numa_node(h).ok())` type-
+ checks as written, because the query returns exactly what the constructor takes. The shortcut would
+ save one line.
+
+**An argument that no longer applies, recorded so it is not re-made:** when this was first argued,
+the buffer lived in `windows-ioring-sys` and a `for_file` constructor would have dragged
+`Win32_System_Ioctl` into a crate that otherwise touched only memory. That was a real cost then. It
+is not one now -- `volume_numa_node` lives here and the feature is already declared -- so the
+argument is void and the four above are what the decision rests on.
+
+### Why discovery is offered at all
+
+The discovered answer is weak, and `volume_numa_node`'s own documentation says so: it answers for a
+*volume*, a volume may span devices, and then a single reported node is a fiction rather than an
+answer. A defensible reading is that a query this weak should not be offered.
+
+It is offered anyway, because the alternative is worse. Refusing it does not stop a consumer needing
+the answer; it makes each one write the `DeviceIoControl` themselves. That is not hypothetical --
+this crate exists because `VirtualAllocExNuma` had been written three times in this workspace on
+exactly that logic, and the epoch-log sample had hand-rolled this very FSCTL. The honest move is to
+provide it **and say plainly what it is worth**, which puts the caveat where the caller will read it
+rather than leaving them to discover the limitation themselves.
+
+### What this preserves
+
+[D-8](../windows-ioring-sys/DESIGN-NOTES.md#d-8) -- locality is the consumer's decision -- is intact
+and is the reason the shape is what it is. This crate supplies a fact and an allocator. It does not
+map a file to a node, does not shard anything, and does not choose on a caller's behalf. A consumer
+who wants those things composes them from what is here, on hardware this workspace has never seen.
diff --git a/crates/win-numa-sys/PLANS.md b/crates/win-numa-sys/PLANS.md
new file mode 100644
index 000000000..f2ab79504
--- /dev/null
+++ b/crates/win-numa-sys/PLANS.md
@@ -0,0 +1,5 @@
+# Plans: win-numa-sys
+
+| Path to CHECKLIST.md | Status | Brief description | Design Notes |
+|---|---|---|---|
+| [CHECKLIST.md](CHECKLIST.md) | in progress | M1 finishes what creating the crate started: `windows-placement-probe` still has its own `VirtualAllocExNuma`, so the duplication is currently **relocated rather than removed**, and `N-1.1` is where that is closed. `N-1.2` decides whether `QueryWorkingSetEx` observation moves here with it -- the capability that would let a caller ask whether a placement request was honoured, rather than being told by the documentation that success proves nothing. `M2+.1` is publishing, which is not optional indefinitely: the crate is a path dependency of the published `windows-ioring-sys`. | [DESIGN-NOTES.md](DESIGN-NOTES.md) |
diff --git a/crates/win-numa-sys/README.md b/crates/win-numa-sys/README.md
new file mode 100644
index 000000000..924ecf47b
--- /dev/null
+++ b/crates/win-numa-sys/README.md
@@ -0,0 +1,60 @@
+# win-numa-sys
+
+Memory-safe Rust over the Windows NUMA APIs.
+
+Windows provides some NUMA concepts; this crate builds library notions on top of them. Whether a
+client uses them is up to the client.
+
+## What is here
+
+- **`NumaNode`** -- a node number, as Windows numbers them. A newtype because a node *number* and a
+ *count* of nodes are both small integers and are trivially swapped at a call site.
+- **`NumaBuffer`** -- an owned allocation made with `VirtualAllocExNuma` and released with
+ `VirtualFree`, so a buffer can be placed on a chosen node rather than wherever the default
+ allocator lands it.
+- **`highest_numa_node()`** and **`volume_numa_node(handle)`** -- the two questions Windows will
+ answer about nodes, wrapped so a caller can find a node to pass without writing the FFI.
+
+## What is not here, and will not be
+
+**Any opinion about which node anything should use.** This crate reports what Windows says and
+allocates where it is told. It does not map a file to a node, does not shard anything, and does not
+choose a node on a caller's behalf.
+
+The `-sys` suffix is that promise. In this workspace the suffix means thin over Win32, memory-safe,
+adding no policy -- see the repository's
+[DESIGN-NOTES.md](../../DESIGN-NOTES.md#the-waitable-queues-crate-is-named-plural-and-carries-no-sys-suffix),
+where a sibling crate drops the suffix for failing exactly that test.
+
+## The name
+
+The first crate here to take `win-` rather than `windows-`. `windows` is Microsoft's namespace, and
+a crate published as `windows-numa-sys` today is a name they may reasonably want tomorrow; see
+[DESIGN-NOTES.md](../../DESIGN-NOTES.md#new-crates-take-the-win-prefix). Existing crates keep their
+names for now.
+
+## A node argument is a preference, not an instruction
+
+`VirtualAllocExNuma`'s parameter is `nndPreferred`, and the name is the contract. A successful
+allocation is **not** evidence that the pages landed on the node that was asked for. Two things
+follow, both of which have already caught someone in this workspace:
+
+- Committed pages are demand-zero, so until something writes to them no physical page has been drawn
+ from the preferred node at all. Measuring placement means faulting the pages in first.
+- An *invalid* node is refused -- but `u32::MAX` is not a test of that, because it is the API's own
+ no-preference sentinel and is accepted by design. A measurement that asks for `u32::MAX` and sees
+ it succeed has measured the sentinel, not a range check.
+
+Observing where pages actually landed needs `QueryWorkingSetEx`, which lives in
+`windows-placement-probe` today; whether it moves here is [CHECKLIST.md](CHECKLIST.md) `N-1.2`.
+
+## Buffer traits live with whoever owns them
+
+`NumaBuffer` implements no I/O buffer trait, because this crate defines none. `windows-ioring-sys`
+and `windows-overlapped-io-sys` each already have their own `IoBuf`/`IoBufMut` pair, and a third
+copy here would have made that duplication harder to resolve rather than easier.
+
+A consumer that owns such a trait implements it for `NumaBuffer` -- the orphan rule permits exactly
+that, since the trait is theirs -- over the inherent `as_ptr`, `as_mut_ptr` and `len`.
+`windows-ioring-sys` does so, and keeps re-exporting `NumaBuffer` so `windows_ioring_sys::NumaBuffer`
+still resolves.
diff --git a/crates/win-numa-sys/src/buffer.rs b/crates/win-numa-sys/src/buffer.rs
new file mode 100644
index 000000000..743a538d1
--- /dev/null
+++ b/crates/win-numa-sys/src/buffer.rs
@@ -0,0 +1,182 @@
+// Copyright (c) 2026 Mike Grier
+// Moved from windows-ioring-sys/src/numa_buffer.rs at 6101ca65, which had in
+// turn moved it from examples/ring_copy/buffer.rs at 834c7afa.
+//! A `VirtualAllocExNuma`-backed buffer, so a buffer can be placed on a chosen
+//! NUMA node rather than wherever the default allocator's own heuristics land
+//! it.
+//!
+//! # Why this is its own crate
+//!
+//! It was in `windows-ioring-sys`, which recommends placing a registered pool
+//! near the device and so had a reason to provide the allocator rather than
+//! leave every caller to write it. But the buffer has nothing to do with a
+//! ring: its only connection was one doc line saying it *can* be registered
+//! into one. Meanwhile `windows-placement-probe` -- which does not depend on
+//! the ring crate -- had written the same `VirtualAllocExNuma` call for itself,
+//! making two independent allocators in a workspace that had already hoisted
+//! this code once to stop exactly that.
+//!
+//! # What this does not decide
+//!
+//! **Which node.** A caller who knows their storage or thread layout passes
+//! `Some(node)`; one who does not passes `None` and gets the system's own
+//! choice, which is what the default allocator would have given anyway.
+//! [`volume_numa_node`](crate::volume_numa_node) is available for finding a
+//! candidate, and says plainly what its answer is and is not worth.
+//!
+//! A node argument is a *preference*, which the underlying parameter says in
+//! its own name (`nndPreferred`). So a successful allocation is not by itself
+//! evidence that the pages landed on the node that was asked for, and code
+//! that needs to know must measure rather than assume -- see the crate root.
+
+use std::io;
+use std::ptr;
+
+use windows_sys::Win32::System::Memory::{
+ MEM_COMMIT, MEM_RELEASE, MEM_RESERVE, PAGE_READWRITE, VirtualAllocExNuma, VirtualFree,
+};
+use windows_sys::Win32::System::Threading::GetCurrentProcess;
+
+use crate::NumaNode;
+
+/// `VirtualAllocExNuma`'s documented sentinel for "no NUMA preference" --
+/// windows-sys does not name this constant, so it is named here rather than
+/// written as a bare literal at the call site.
+const NUMA_NO_PREFERRED_NODE: u32 = u32::MAX;
+
+/// An owned buffer allocated with `VirtualAllocExNuma`, freed with
+/// `VirtualFree` on drop.
+///
+/// The allocation is page-granular: `VirtualAllocExNuma` rounds a request up
+/// to the system page size, so a caller asking for a small buffer gets at
+/// least a page. [`NumaBuffer::len`] reports the length that was *requested*,
+/// which is the length a kernel call should be told about.
+///
+/// # Buffer traits live with whoever owns them
+///
+/// This type implements no I/O buffer trait, because this crate defines none
+/// and should not: `windows-ioring-sys` and `windows-overlapped-io-sys` each
+/// have their own `IoBuf`/`IoBufMut`, and a third copy here would be one more
+/// of a thing the workspace is already deciding what to do about. A consumer
+/// that owns such a trait implements it for this type -- the orphan rule
+/// permits exactly that, since the trait is theirs -- over
+/// [`NumaBuffer::as_ptr`], [`NumaBuffer::as_mut_ptr`] and [`NumaBuffer::len`].
+/// `windows-ioring-sys` does so.
+pub struct NumaBuffer {
+ ptr: *mut u8,
+ len: usize,
+}
+
+// SAFETY: the allocation is exclusively owned by this value; sending it
+// across threads only moves that ownership, never aliases it.
+unsafe impl Send for NumaBuffer {}
+
+impl std::fmt::Debug for NumaBuffer {
+ /// Shows the address and requested length, never the contents: a
+ /// registered buffer routinely holds someone's data, and a `Debug` that
+ /// printed it would put that data anywhere a caller logs.
+ fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
+ f.debug_struct("NumaBuffer")
+ .field("ptr", &self.ptr)
+ .field("len", &self.len)
+ .finish()
+ }
+}
+
+impl NumaBuffer {
+ /// Allocate `len` bytes, preferring `node` if given.
+ ///
+ /// Passing `None` requests no NUMA preference, which is the same choice
+ /// the default allocator makes implicitly.
+ ///
+ /// # What a successful return does and does not mean
+ ///
+ /// A valid node is a **preference**, not an instruction -- `nndPreferred`
+ /// says so in its name -- so success means the request was accepted and
+ /// memory was obtained, never that the pages came from the node asked for.
+ ///
+ /// An *invalid* node is a different matter and is genuinely refused: this
+ /// crate's tests assert that `u32::MAX - 1` fails. Note that `u32::MAX`
+ /// itself is **not** a test of that, because it is the API's own
+ /// no-preference sentinel and is accepted by design -- a measurement that
+ /// reads it as an out-of-range node being tolerated has measured the
+ /// sentinel instead.
+ ///
+ /// # Errors
+ ///
+ /// The error from `VirtualAllocExNuma`, which includes a `len` of zero and
+ /// a node number the machine does not have.
+ pub fn new(len: usize, node: Option) -> io::Result {
+ // SAFETY: no pointer arguments; the returned value is a pseudo-handle
+ // that needs no closing.
+ let process = unsafe { GetCurrentProcess() };
+ // SAFETY: `process` is a valid pseudo-handle for the duration of this
+ // call; a null `lpAddress` lets the system choose the address.
+ let ptr = unsafe {
+ VirtualAllocExNuma(
+ process,
+ ptr::null(),
+ len,
+ MEM_COMMIT | MEM_RESERVE,
+ PAGE_READWRITE,
+ node.map_or(NUMA_NO_PREFERRED_NODE, NumaNode::get),
+ )
+ };
+ if ptr.is_null() {
+ return Err(io::Error::last_os_error());
+ }
+ Ok(Self {
+ ptr: ptr.cast(),
+ len,
+ })
+ }
+
+ /// The allocation's base address, fixed for this value's life.
+ #[must_use]
+ pub fn as_ptr(&self) -> *const u8 {
+ self.ptr
+ }
+
+ /// The allocation's base address for writing.
+ ///
+ /// Takes `&mut self` because this value uniquely owns the allocation, so
+ /// exclusive access to the value is exclusive access to the bytes.
+ #[must_use]
+ pub fn as_mut_ptr(&mut self) -> *mut u8 {
+ self.ptr
+ }
+
+ /// The length that was requested.
+ ///
+ /// Not the length that was reserved: `VirtualAllocExNuma` rounds up to a
+ /// page, so the mapping is at least this large and usually larger. This is
+ /// the number a kernel call should be told, because it is the number the
+ /// caller asked to use.
+ #[must_use]
+ pub fn len(&self) -> usize {
+ self.len
+ }
+
+ /// Whether the requested length was zero.
+ ///
+ /// Present because clippy asks for it beside [`Self::len`], and because a
+ /// zero-length request is rejected by `VirtualAllocExNuma` rather than
+ /// producing an empty buffer -- so on any value that exists, this is false.
+ #[must_use]
+ pub fn is_empty(&self) -> bool {
+ self.len == 0
+ }
+}
+
+impl Drop for NumaBuffer {
+ fn drop(&mut self) {
+ // SAFETY: `self.ptr` was returned by `VirtualAllocExNuma` above and
+ // is freed exactly once, here.
+ unsafe {
+ VirtualFree(self.ptr.cast(), 0, MEM_RELEASE);
+ }
+ }
+}
+
+#[cfg(test)]
+mod tests;
diff --git a/crates/win-numa-sys/src/buffer/tests.rs b/crates/win-numa-sys/src/buffer/tests.rs
new file mode 100644
index 000000000..d17627ad5
--- /dev/null
+++ b/crates/win-numa-sys/src/buffer/tests.rs
@@ -0,0 +1,165 @@
+// Copyright (c) 2026 Mike Grier
+//! Tests for [`NumaBuffer`] (M22.3).
+//!
+//! These allocate through `VirtualAllocExNuma` and free on drop. That is an
+//! operating-system call, but this repository targets one operating system, so
+//! by its own Quality rule that alone does not make these integration tests:
+//! they cross no process, device, or network boundary, hold no handle, and
+//! each completes in microseconds.
+//!
+//! # What these cannot establish
+//!
+//! **That the pages landed on the requested node.** The parameter is
+//! `nndPreferred`, so a success proves the request was accepted, not that it
+//! was honoured -- and a single-node host, which is what this is developed on,
+//! could not tell the difference either way. The type's own documentation says
+//! this; these tests do not pretend otherwise by asserting a node back.
+
+use super::NumaBuffer;
+use crate::NumaNode;
+
+/// A node number no machine has, for the rejection cases. `u32::MAX` is not
+/// usable here -- it is `VirtualAllocExNuma`'s own "no preference" sentinel,
+/// so it would be accepted rather than refused.
+const ABSURD_NODE: u32 = u32::MAX - 1;
+
+#[test]
+fn an_unplaced_buffer_allocates() {
+ let buffer = NumaBuffer::new(4096, None).expect("an unplaced allocation");
+ assert!(!buffer.as_ptr().is_null());
+}
+
+#[test]
+fn node_zero_allocates() {
+ // Every machine reports a node 0, including one with NUMA disabled, where
+ // `GetNumaHighestNodeNumber` answers 0.
+ let buffer = NumaBuffer::new(4096, Some(NumaNode::new(0))).expect("node 0 exists everywhere");
+ assert!(!buffer.as_ptr().is_null());
+}
+
+#[test]
+fn the_reported_length_is_the_requested_one_not_the_rounded_one() {
+ // The allocation is page-granular, so the *mapping* is at least a page.
+ // What a kernel call is told about an operation is `len`, and that must
+ // be what the caller asked for -- reporting the rounded-up size would
+ // invite a read or write past the caller's intent.
+ for len in [1_usize, 100, 4095, 4096, 4097, 65536] {
+ let buffer = NumaBuffer::new(len, None).expect("a valid allocation");
+ assert_eq!(buffer.len(), len, "for a request of {len} bytes");
+ }
+}
+
+#[test]
+fn a_fresh_allocation_is_zeroed() {
+ // `VirtualAlloc`-family pages arrive zeroed, which is what lets a caller
+ // register an arena without filling it first.
+ let mut buffer = NumaBuffer::new(8192, None).expect("a valid allocation");
+ // SAFETY: the pointer is this buffer's own allocation of `len`
+ // bytes, and `&mut` makes the borrow exclusive.
+ let bytes = unsafe { std::slice::from_raw_parts(buffer.as_mut_ptr(), buffer.len()) };
+ assert!(bytes.iter().all(|&b| b == 0));
+}
+
+#[test]
+fn bytes_written_read_back() {
+ let mut buffer = NumaBuffer::new(4096, None).expect("a valid allocation");
+ let len = buffer.len();
+ // SAFETY: as above -- this buffer's own allocation, exclusively borrowed.
+ let bytes = unsafe { std::slice::from_raw_parts_mut(buffer.as_mut_ptr(), len) };
+ for (index, byte) in bytes.iter_mut().enumerate() {
+ *byte = (index % 251) as u8;
+ }
+ // SAFETY: as above.
+ let read = unsafe { std::slice::from_raw_parts(buffer.as_ptr(), len) };
+ assert!(read.iter().enumerate().all(|(i, &b)| b == (i % 251) as u8));
+}
+
+#[test]
+fn the_address_is_stable_across_reads() {
+ // The whole reason this type can be registered: the address it reports
+ // does not move for its life.
+ let mut buffer = NumaBuffer::new(4096, None).expect("a valid allocation");
+ let first = buffer.as_ptr();
+ assert_eq!(buffer.as_ptr(), first);
+ assert_eq!(buffer.as_mut_ptr().cast_const(), first);
+ assert_eq!(buffer.as_ptr(), first);
+}
+
+#[test]
+fn an_address_survives_a_move() {
+ // The address must be stable for the value's life,
+ // which includes being moved -- the allocation is behind a pointer, so
+ // moving the handle does not move the bytes.
+ let buffer = NumaBuffer::new(4096, None).expect("a valid allocation");
+ let before = buffer.as_ptr();
+ let moved = buffer;
+ assert_eq!(moved.as_ptr(), before);
+}
+
+#[test]
+fn separate_buffers_do_not_alias() {
+ let buffers: Vec = (0..8)
+ .map(|_| NumaBuffer::new(4096, None).expect("a valid allocation"))
+ .collect();
+ let mut addresses: Vec<*const u8> = buffers.iter().map(NumaBuffer::as_ptr).collect();
+ addresses.sort_unstable();
+ let before = addresses.len();
+ addresses.dedup();
+ assert_eq!(addresses.len(), before, "every slot must be its own memory");
+}
+
+#[test]
+fn a_zero_length_request_is_refused() {
+ let error = NumaBuffer::new(0, None).expect_err("a zero-length mapping is not allocatable");
+ assert!(error.raw_os_error().is_some(), "and it is an OS error");
+}
+
+#[test]
+fn a_node_the_machine_does_not_have_is_refused() {
+ // The other half of the guard: the accepting cases above would all still
+ // pass if `new` ignored its node argument entirely.
+ let error = NumaBuffer::new(4096, Some(NumaNode::new(ABSURD_NODE)))
+ .expect_err("no machine has this node number");
+ assert!(error.raw_os_error().is_some(), "and it is an OS error");
+}
+
+#[test]
+fn a_request_the_address_space_cannot_hold_is_refused() {
+ let error = NumaBuffer::new(usize::MAX, None).expect_err("no process has this much space");
+ assert!(error.raw_os_error().is_some(), "and it is an OS error");
+}
+
+#[test]
+fn a_refused_allocation_leaves_the_allocator_usable() {
+ // A failed `VirtualAllocExNuma` must not leave anything behind that stops
+ // the next one -- the sample arenas allocate in a loop, so one bad request
+ // in the middle would otherwise take the rest with it.
+ let _ = NumaBuffer::new(4096, Some(NumaNode::new(ABSURD_NODE))).expect_err("refused");
+ let buffer = NumaBuffer::new(4096, None).expect("the next allocation still works");
+ assert!(!buffer.as_ptr().is_null());
+}
+
+#[test]
+fn many_buffers_allocate_and_free() {
+ // Drop is what returns the mapping; leaking it would show up here as an
+ // address-space exhaustion long before the loop ends.
+ for _ in 0..512 {
+ let buffer = NumaBuffer::new(65536, None).expect("a valid allocation");
+ assert!(!buffer.as_ptr().is_null());
+ }
+}
+
+#[test]
+fn a_buffer_is_send() {
+ // The `unsafe impl Send` is load-bearing: a domain runtime allocates on
+ // one thread and uses the buffer on the pinned thread that owns the ring.
+ fn assert_send() {}
+ assert_send::();
+
+ let buffer = NumaBuffer::new(4096, None).expect("a valid allocation");
+ let address = buffer.as_ptr() as usize;
+ let moved = std::thread::spawn(move || buffer.as_ptr() as usize)
+ .join()
+ .expect("the thread completes");
+ assert_eq!(moved, address, "and the address travels with it");
+}
diff --git a/crates/win-numa-sys/src/lib.rs b/crates/win-numa-sys/src/lib.rs
new file mode 100644
index 000000000..c4bf991e7
--- /dev/null
+++ b/crates/win-numa-sys/src/lib.rs
@@ -0,0 +1,55 @@
+// Copyright (c) 2026 Mike Grier
+//! Memory-safe Rust over the Windows NUMA APIs.
+//!
+//! Windows provides some NUMA concepts; this crate builds library notions on
+//! top of them. Whether a client uses them is up to the client.
+//!
+//! # What is here
+//!
+//! - [`NumaNode`] -- a node number, as Windows numbers them.
+//! - [`NumaBuffer`] -- an owned allocation made with `VirtualAllocExNuma` and
+//! released with `VirtualFree`, so a buffer can be placed on a chosen node
+//! rather than wherever the default allocator lands it.
+//! - [`highest_numa_node`] and [`volume_numa_node`] -- the two questions
+//! Windows will answer about nodes, wrapped so a caller can find a node to
+//! pass without writing the FFI themselves.
+//!
+//! # What is not here, and will not be
+//!
+//! **Any opinion about which node anything should use.** This crate reports
+//! what Windows says and allocates where it is told. It does not map a file to
+//! a node, does not shard anything, and does not choose a node on a caller's
+//! behalf. Those are workload decisions, and a consumer on hardware this
+//! workspace has never seen is better placed to take them.
+//!
+//! The `-sys` suffix is that promise, and it is the repository's meaning of the
+//! suffix rather than a convention borrowed from elsewhere: thin over Win32,
+//! memory-safe, adding no policy.
+//!
+//! # A node argument is a preference, not an instruction
+//!
+//! The underlying parameter says so in its own name -- `nndPreferred`. A
+//! successful allocation is therefore **not** evidence that the pages landed on
+//! the node that was asked for, and code that needs to know must observe rather
+//! than assume. Two measured facts about that, both from this workspace:
+//!
+//! - Committed pages are demand-zero, so until something writes to them no
+//! physical page has been drawn from the preferred node at all. A caller
+//! measuring placement must fault the pages in first.
+//! - An *invalid* node is refused, and `u32::MAX` is not a test of that: it is
+//! the API's own no-preference sentinel, accepted by design. A measurement
+//! that asks for `u32::MAX` and sees it succeed has measured the sentinel,
+//! not a range check. [`NumaBuffer`]'s tests use `u32::MAX - 1`.
+//!
+//! Observing where pages actually landed needs `QueryWorkingSetEx`, which lives
+//! in `windows-placement-probe` today and is a candidate to move here; see that
+//! crate's `peer_index_cache`.
+
+#![cfg(windows)]
+#![deny(missing_docs)]
+
+mod buffer;
+mod node;
+
+pub use buffer::NumaBuffer;
+pub use node::{NumaNode, highest_numa_node, volume_numa_node};
diff --git a/crates/win-numa-sys/src/node.rs b/crates/win-numa-sys/src/node.rs
new file mode 100644
index 000000000..f08a29428
--- /dev/null
+++ b/crates/win-numa-sys/src/node.rs
@@ -0,0 +1,140 @@
+// Copyright (c) 2026 Mike Grier
+//! [`NumaNode`], and the two questions Windows will answer about nodes.
+
+use std::io;
+use std::os::windows::io::RawHandle;
+
+use windows_sys::Win32::Foundation::HANDLE;
+use windows_sys::Win32::System::IO::DeviceIoControl;
+use windows_sys::Win32::System::Ioctl::FSCTL_QUERY_VOLUME_NUMA_INFO;
+use windows_sys::Win32::System::Threading::GetNumaHighestNodeNumber;
+
+/// A NUMA node number, as Windows numbers them.
+///
+/// A newtype rather than a bare `u32` because the two `u32`s in this area mean
+/// different things and are trivially swapped at a call site: a node *number*
+/// and a *count* of nodes are both small integers, and
+/// [`highest_numa_node`] returns the former while reading like the latter.
+///
+/// **Node numbers are not an index.** Windows does not promise they run
+/// `0..n`, so a caller deriving a node from a position in some list is making
+/// an assumption the platform never offered -- see `windows-placement-probe`,
+/// where a positional index reaching `VirtualAllocExNuma` was a real defect and
+/// is called out at the site that fixed it.
+#[derive(Clone, Copy, Debug, PartialEq, Eq, PartialOrd, Ord, Hash)]
+pub struct NumaNode(u32);
+
+impl NumaNode {
+ /// Wrap a raw node number.
+ ///
+ /// Not validated against the machine: `VirtualAllocExNuma` is where an
+ /// impossible node is rejected, and validating here would mean a second,
+ /// weaker opinion about what exists. See [`NumaBuffer::new`] for what the
+ /// platform actually does with an out-of-range node, which is not what its
+ /// documentation says.
+ ///
+ /// [`NumaBuffer::new`]: crate::NumaBuffer::new
+ #[must_use]
+ pub const fn new(node: u32) -> Self {
+ Self(node)
+ }
+
+ /// The raw node number, for handing to a Win32 call.
+ #[must_use]
+ pub const fn get(self) -> u32 {
+ self.0
+ }
+}
+
+impl std::fmt::Display for NumaNode {
+ fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
+ write!(f, "node {}", self.0)
+ }
+}
+
+/// The highest node number this machine reports, if it answers.
+///
+/// `GetNumaHighestNodeNumber`. The value is a *node number*, not a count: a
+/// machine with one node answers `NumaNode(0)`, not one.
+///
+/// # Why a caller wants this
+///
+/// Mostly to qualify a report rather than to drive a choice. On a machine that
+/// answers `NumaNode(0)` there is exactly one node, so every placement decision
+/// is the same decision, and a line saying "placed on node 0" is true and
+/// misleading. Asking this is how a caller can say which of those it is.
+///
+/// # Errors
+///
+/// `None` when the call fails, which is reported as "unknown" rather than as a
+/// node count of zero -- a machine that will not say how many nodes it has is a
+/// different thing from a machine with none.
+#[must_use]
+pub fn highest_numa_node() -> Option {
+ let mut highest: u32 = 0;
+ // SAFETY: the out parameter is a live local for the call's duration.
+ let ok = unsafe { GetNumaHighestNodeNumber(&raw mut highest) };
+ (ok != 0).then_some(NumaNode::new(highest))
+}
+
+/// The NUMA node `handle`'s **volume** reports, if it reports one.
+///
+/// `FSCTL_QUERY_VOLUME_NUMA_INFO`, issued against the handle directly -- the
+/// documented control code accepts a file or directory handle, so this needs no
+/// device-tree walk and no second open.
+///
+/// # This answers a question about a volume, not about a file
+///
+/// The documented meaning is the node the *volume* resides on. A volume may
+/// span several devices -- an ordinary spanned volume or a Storage Spaces set
+/// does -- and then a single reported node is a fiction rather than an answer,
+/// because the file's extents may live anywhere across the set. **It therefore
+/// cannot tell a caller which node a particular file's I/O is closest to**, even
+/// when it succeeds.
+///
+/// That is a limit of the question, not of this wrapper, and it is why this
+/// function reports rather than acts: a caller who knows their storage layout
+/// can decide what the answer is worth, and one who does not should not have a
+/// placement chosen for them on the strength of it.
+///
+/// # Errors
+///
+/// Any error from `DeviceIoControl`, and
+/// [`io::ErrorKind::InvalidData`] if the control code returns a payload that is
+/// not the documented single `u32`.
+pub fn volume_numa_node(handle: RawHandle) -> io::Result {
+ let mut node: u32 = u32::MAX;
+ let mut returned: u32 = 0;
+ // SAFETY: `handle` is the caller's, and outlives this call by the contract
+ // on this function. The FSCTL takes no input buffer, and its output is
+ // documented as `FSCTL_QUERY_VOLUME_NUMA_INFO_OUTPUT { ULONG NumaNode }` --
+ // one `u32`, which is what `node` provides and what the size argument says.
+ let ok = unsafe {
+ DeviceIoControl(
+ handle as HANDLE,
+ FSCTL_QUERY_VOLUME_NUMA_INFO,
+ std::ptr::null(),
+ 0,
+ (&raw mut node).cast::(),
+ u32::try_from(size_of::()).expect("four fits in a u32"),
+ &raw mut returned,
+ std::ptr::null_mut(),
+ )
+ };
+ if ok == 0 {
+ return Err(io::Error::last_os_error());
+ }
+ if returned as usize != size_of::() {
+ return Err(io::Error::new(
+ io::ErrorKind::InvalidData,
+ format!(
+ "FSCTL_QUERY_VOLUME_NUMA_INFO returned {returned} bytes, not {}",
+ size_of::()
+ ),
+ ));
+ }
+ Ok(NumaNode::new(node))
+}
+
+#[cfg(test)]
+mod tests;
diff --git a/crates/win-numa-sys/src/node/tests.rs b/crates/win-numa-sys/src/node/tests.rs
new file mode 100644
index 000000000..0d855f575
--- /dev/null
+++ b/crates/win-numa-sys/src/node/tests.rs
@@ -0,0 +1,143 @@
+// Copyright (c) 2026 Mike Grier
+//! Tests for [`NumaNode`] and the two node queries.
+//!
+//! # What these can and cannot establish
+//!
+//! The newtype's tests are total: it is a wrapper over a `u32` and every claim
+//! about it holds on any host.
+//!
+//! The query tests are **host-dependent by nature**, and are written to assert
+//! only what is true on every machine rather than what happens to be true on
+//! this one. This workspace is developed on a single-node host, so a test that
+//! asserted a particular node back would be asserting the only answer available
+//! here and would fail on the hardware the crate exists for. What is asserted
+//! instead is shape: that a query either answers or says it could not, that it
+//! never invents a node, and that asking twice agrees.
+
+use std::os::windows::io::AsRawHandle;
+
+use super::{NumaNode, highest_numa_node, volume_numa_node};
+
+/// A scratch file to ask about, named per test so tests running as threads in
+/// one process cannot collide on it.
+fn scratch(tag: &str) -> (std::path::PathBuf, std::fs::File) {
+ let path = std::env::temp_dir().join(format!(
+ "win-numa-sys-node-{}-{tag}.tmp",
+ std::process::id()
+ ));
+ let file = std::fs::OpenOptions::new()
+ .create(true)
+ .truncate(true)
+ .write(true)
+ .open(&path)
+ .expect("a scratch file in the temp directory");
+ (path, file)
+}
+
+#[test]
+fn a_node_round_trips_through_the_newtype() {
+ for raw in [0_u32, 1, 7, 63, u32::MAX - 1, u32::MAX] {
+ assert_eq!(NumaNode::new(raw).get(), raw, "for {raw}");
+ }
+}
+
+#[test]
+fn nodes_compare_and_order_by_their_number() {
+ assert_eq!(NumaNode::new(3), NumaNode::new(3));
+ assert_ne!(NumaNode::new(3), NumaNode::new(4));
+ assert!(NumaNode::new(3) < NumaNode::new(4));
+ let mut nodes = [NumaNode::new(9), NumaNode::new(2), NumaNode::new(5)];
+ nodes.sort_unstable();
+ assert_eq!(
+ nodes,
+ [NumaNode::new(2), NumaNode::new(5), NumaNode::new(9)]
+ );
+}
+
+/// The `Display` form says what the number *is*, because a bare integer in a
+/// report is ambiguous between a node number and a node count -- the very
+/// confusion the newtype exists to stop.
+#[test]
+fn display_names_the_thing_rather_than_printing_a_bare_integer() {
+ assert_eq!(NumaNode::new(0).to_string(), "node 0");
+ assert_eq!(NumaNode::new(12).to_string(), "node 12");
+}
+
+/// Whatever the highest node is, asking twice agrees.
+///
+/// Deliberately not an assertion about the value: on this workspace's host it
+/// is `node 0`, and pinning that would encode the development machine into the
+/// suite. Hot-add could in principle change it between calls, which would make
+/// this flaky rather than wrong -- it has never been observed, and a failure
+/// here would be a genuine finding about the platform rather than a bad test.
+#[test]
+fn the_highest_node_is_stable_across_calls() {
+ assert_eq!(highest_numa_node(), highest_numa_node());
+}
+
+/// A machine that answers reports a node number, not a count.
+///
+/// The distinction is the reason the newtype exists: a single-node machine
+/// answers `node 0`, and code reading that as "zero nodes" would conclude the
+/// machine has no NUMA at all.
+#[test]
+fn the_highest_node_is_a_number_not_a_count() {
+ if let Some(highest) = highest_numa_node() {
+ // Every machine has at least one node, so the highest number is
+ // reachable as a node. This holds on a 1-node host (0) and on a
+ // 64-node one (63).
+ assert!(
+ highest.get() < u32::MAX,
+ "a real machine's highest node cannot be the no-preference sentinel"
+ );
+ }
+}
+
+/// The volume query either answers or reports why not; it never invents a node.
+///
+/// On a local NTFS volume this workspace's host answers. A host whose storage
+/// stack declines is equally valid and must produce an error rather than a
+/// fabricated zero -- which is the failure this asserts against, because zero
+/// is a real node number and so a plausible-looking fabrication.
+#[test]
+fn the_volume_query_answers_or_errors_but_never_fabricates() {
+ let (path, file) = scratch("answers-or-errors");
+ match volume_numa_node(file.as_raw_handle()) {
+ Ok(node) => assert!(
+ node.get() < u32::MAX,
+ "a reported node cannot be the no-preference sentinel"
+ ),
+ Err(error) => assert_ne!(
+ error.kind(),
+ std::io::ErrorKind::Other,
+ "a failure should carry the OS error, not a placeholder"
+ ),
+ }
+ drop(file);
+ let _ = std::fs::remove_file(path);
+}
+
+/// Asking the same handle twice agrees.
+#[test]
+fn the_volume_query_is_stable_for_one_handle() {
+ let (path, file) = scratch("stable");
+ let first = volume_numa_node(file.as_raw_handle()).ok();
+ let second = volume_numa_node(file.as_raw_handle()).ok();
+ assert_eq!(first, second);
+ drop(file);
+ let _ = std::fs::remove_file(path);
+}
+
+/// An invalid handle is an error, not a node.
+///
+/// The cheapest reachable failure path, and worth having because the success
+/// path cannot be forced to fail on a host whose volume does answer.
+#[test]
+fn an_invalid_handle_errors() {
+ let error = volume_numa_node(std::ptr::null_mut())
+ .expect_err("the null handle is not a file or directory handle");
+ assert!(
+ error.raw_os_error().is_some(),
+ "the failure should carry the OS error code, got {error:?}"
+ );
+}
diff --git a/crates/windows-ioring-sys/BORROW-SURFACE.txt b/crates/windows-ioring-sys/BORROW-SURFACE.txt
index 8ed1b15f7..e548286c3 100644
--- a/crates/windows-ioring-sys/BORROW-SURFACE.txt
+++ b/crates/windows-ioring-sys/BORROW-SURFACE.txt
@@ -7,8 +7,13 @@
# forgotten: CI regenerates it and fails when it disagrees with the source.
crates/windows-ioring-sys/src/batch.rs :: get -> io::Result<&[u8]>
crates/windows-ioring-sys/src/batch.rs :: get_mut -> io::Result<&mut [u8]>
+crates/windows-ioring-sys/src/batch.rs :: new(ring: &'ring mut IoRing) -> Self
crates/windows-ioring-sys/src/contract.rs :: violations -> &[Violation]
+crates/windows-ioring-sys/src/error.rs :: as_ioring_error -> Option<&IoRingError>
crates/windows-ioring-sys/src/error.rs :: name -> &'static str
crates/windows-ioring-sys/src/error.rs :: name -> Option<&'static str>
crates/windows-ioring-sys/src/event_delivery.rs :: batch -> Batch<'_>
+crates/windows-ioring-sys/src/event_delivery.rs :: new(mut ring: IoRing, on_completion: F, env: Option<&mut CallbackEnviron<'_>>,) -> io::Result
crates/windows-ioring-sys/src/event_delivery.rs :: scope -> RingScope<'_>
+crates/windows-ioring-sys/src/pending.rs :: contract -> Option<&RingContract>
+crates/windows-ioring-sys/src/ring.rs :: wait(ring: &mut RingWait<'_>, timeout_ms: u32) -> io::Result<()>
diff --git a/crates/windows-ioring-sys/CHECKLIST.md b/crates/windows-ioring-sys/CHECKLIST.md
index 507677f63..b19466f40 100644
--- a/crates/windows-ioring-sys/CHECKLIST.md
+++ b/crates/windows-ioring-sys/CHECKLIST.md
@@ -12,106 +12,359 @@ and M15-M18
M19 is archived [here](COMPLETED-CHECKLIST.md#m19).
-**`M20` is pending; `M6+` is parked rather than pending** -- see the `M{n}+` convention: it is gated work
-with no current obligation, not an unfinished milestone.
+**`M20` through `M23` are pending; `M6+` is parked rather than pending** -- see the `M{n}+` convention: it
+is gated work with no current obligation, not an unfinished milestone.
## M20 -- Repairs from the 2026-08-30 NUMA-sharding measurement
Queued from
[DESIGN-SESSION-2026-08-30-numa-sharded-io-execution-domains.md](../../design-sessions/DESIGN-SESSION-2026-08-30-numa-sharded-io-execution-domains.md),
which measured a shipping ARM laptop and found the L3 heuristic's justification does not hold there. These
-are documentation and policy repairs only; **no defect was found in `ring_copy`** -- `Policy::select`
-already degrades to a whole-machine domain and reports it, which an initial reading of the session got
-wrong and the code corrected.
+were queued as documentation and policy repairs only, on the basis that **no defect was found in
+`ring_copy`** -- `Policy::select` already degrades to a whole-machine domain and reports it, which an
+initial reading of the session got wrong and the code corrected.
+
+**Corrected 2026-09-19: that basis no longer holds, and it changes the order.** `SH-4.12` in
+[CHECKLIST-ship-topology-and-queues.md](../../CHECKLIST-ship-topology-and-queues.md) later found two
+defects in that same function: it selects on `DomainKind::Cache { level: 3, .. }` rather than asking
+`outermost_partitioning_cache()` -- the one definition of which cache level partitions a machine, shipped
+in `windows-topology-sys` 0.2.0 -- so it can produce **overlapping** ring domains where two cache kinds
+report at level 3, and degrades silently on a host whose outermost partition sits at another level.
+`M20.1` and `M20.3` both land on that function and that rule, so both are **coupled to `SH-4.12`** and
+must follow it. `M20.2` and `M20.4` are done. `M20.6` is gated the other way, on `M22.1`. That leaves
+`SH-4.12` as the only thing standing between M20 and completion.
The design questions the session opened are deliberately **not** queued here. It is still open, and its
conclusions belong to it until it converges.
-- [ ] **M20.1** -- Correct the L3 heuristic's justification in
- [DESIGN-NOTES.md](DESIGN-NOTES.md). It currently says the last-level-cache domain "is meaningful on Intel
- and ARM too, where the NUMA node often is not." **Measured counter-example:** a Snapdragon X2 Elite
- (X2E80100, Qualcomm Oryon; 12 cores, no SMT) reports **zero** L3 cache domains -- `L3CacheSize = 0` from
- WMI, and `GetLogicalProcessorInformationEx` yields L1 and L2 only, with L2 forming two domains of six
- processors that agree with the two `Module` domains. The claim that L3 is meaningful on ARM is false on a
- shipping part. Keep the finding that L3 beats the NUMA node; restate the rule as **the outermost cache
- level that actually partitions the machine**, and say what happens when no such level is reported. Sweep
- every restatement of the L3 rule per the repository's blast-radius convention, including the README and
- `ring_copy`'s `policy.rs` doc comments, not only the one sentence quoted above.
-
-- [ ] **M20.2** -- Record the measurement itself as a decision in
- [DESIGN-NOTES.md](DESIGN-NOTES.md), so the next reader inherits the datapoint rather than re-measuring:
- an ARM Windows laptop with no L3 at all, and zero `Win32_NumaNode` instances, is the *common* consumer
- shape now rather than an exotic one. This is the ARM sibling of the existing zero-NUMA-node VM
- observation and belongs beside it.
-
-- [ ] **M20.3** -- Make `ring_copy`'s degraded-fallback path observable in a test. The whole-machine
- fallback in `Policy::select` is the branch every zero-relation machine takes, and this session was the
- first time anyone confirmed it runs. Assert both halves on a synthetic topology: that a policy whose
- relation is absent returns one whole-machine domain with `degraded = true`, and that a policy whose
- relation is present is **not** flagged degraded -- the second half matters because a test of the first
- alone would pass against a function that always degrades.
-
-- [ ] **M20.4** -- Correct "What is not reachable" in [DESIGN-NOTES.md](DESIGN-NOTES.md). It says mapping a
- file handle to its backing device's NUMA node "has no clean user-mode path" and "means walking volume to
- disk to device instance and reading `DEVPKEY_Device_Numa_Node`". **That is wrong on mechanism.**
- `FSCTL_QUERY_VOLUME_NUMA_INFO` is documented in the IFS docs, takes a handle to a **file or directory**
- directly, and returns `FSCTL_QUERY_VOLUME_NUMA_INFO_OUTPUT { ULONG NumaNode }`. No walking required.
- The **conclusion survives for a better reason**, and that is the point of the rewrite: the documented
- meaning is the node the *volume* resides on, not where the file's extents live, so it cannot answer
- "which ring should this file's I/O go to" even when it succeeds; and it is absent whenever the device
- advertised no proximity domain. Record `GetNumaNodeNumberFromHandle` as the other path -- a wrapper over
- `NtQueryInformationFile` with `FileNumaNodeInformation` (class 53) -- and that PHNT and the WDK mark that
- class **reserved for system use**, so this crate must not build on it. State plainly that no published
- measurement of either call succeeding on an ordinary NTFS data file could be found, and cite
- [file-handle-numa-spike.rs](design-sessions/spikes/file-handle-numa-spike.rs) as the unrun instrument.
- **Blocked on hardware, not on a decision:** settling it needs a multi-node machine with storage whose
- PDO advertises a proximity domain. Write the correction now (the documentation defect is independent of
- the measurement) and leave the empirical question open.
+- [x] **M20.1** -- Restate the cache heuristic as "the outermost cache level that actually partitions
+ the machine", sweep every restatement, and replace the consumer that bound to the level number.
+ Done together with `SH-4.12`, which is the code half of the same change.
+ -> [completed 2026-09-22](COMPLETED-CHECKLIST.md#m201)
+
+- [x] **M20.2** -- Record the 2026-08-30 ARM measurement as a decision, beside the zero-NUMA-node
+ observation it is the sibling of.
+ -> [completed 2026-09-19](COMPLETED-CHECKLIST.md#m202)
+
+- [x] **M20.3** -- Make `ring_copy`'s degraded-fallback path observable in a test, asserting both
+ that an absent relation degrades and that a present one does not. Done without waiting on
+ `SH-4.12`: the fallback tail is shared by every policy, so exercising it through `ByNode` and
+ `ByPackage` pins nothing that item rewrites.
+ -> [completed 2026-09-22](COMPLETED-CHECKLIST.md#m203)
+
+- [x] **M20.4** -- Correct "What is not reachable" in [DESIGN-NOTES.md](DESIGN-NOTES.md): the
+ file-handle-to-storage-node mapping is reachable on mechanism, and the conclusion it supported now rests
+ on volume granularity, absence, and spanned volumes instead.
+ -> [completed 2026-09-19](COMPLETED-CHECKLIST.md#m204)
- [x] **M20.5** -- Dissolved by [D-47](DESIGN-NOTES.md#d-47-detail) rather than decided: the
`flush_barrier` assertion was measuring a claim the platform does not honour, so it was never a
flaky test. -> [completed 2026-09-07](COMPLETED-CHECKLIST.md#m205)
-- [ ] **M20.6** -- Re-evaluate `CommitStrategy::AlternatingRings` and the epoch-log benchmark's conclusion
- against [D-47](DESIGN-NOTES.md#d-47-detail). The strategy comparison in
- [strategy.rs](examples/epoch_log/strategy.rs) was designed around D-24's claim that a covering flush holds
- back operations queued behind it: the harness deliberately keeps appending while a commit is outstanding so
- that the stall would be visible in the numbers. D-47 established there is no such stall, so **the rationale
- the benchmark rests on is withdrawn even though the measurements themselves stand**. Two things to settle,
- and they are independent: whether alternating rings still earns its cost now that its stated benefit
- (keeping appends off a stalled ring) does not exist -- the remaining benefit is that epoch *N+1*'s appends
- are provably outside epoch *N*, which is a correctness property rather than a throughput one -- and whether
- the published numbers should be re-read, re-run, or annotated. **Not a documentation-only fix:** if the
- answer is that the strategy no longer earns its place, that is an API change to a published example.
- The corrected prose in [strategy.rs](examples/epoch_log/strategy.rs) and
- [DESIGN-NOTES.md](DESIGN-NOTES.md) both point here.
- *(Numbered M20.6 rather than M20.5 because M20.5 was in flight on a separate branch when this was
- written. That branch was closed unmerged; M20.5 arrives here instead, dissolved -- see above.)*
-
-
-## M6+ -- Model B: explicit-thread delivery and affinity
-
-Parked, not pending. Deferred by the engineer's explicit direction during the 2026-08-22 design session,
-with the plan scoped now so the shape is not lost. This is **not** a fallback for a missing capability
-(D-3) -- it is the high-performance architecture, and M4's thread-pool path is the convenient one.
-
-- [ ] **M6+.1** -- `DeliveryMode::{ThreadpoolWait, PinnedThread}` as an explicit consumer choice, never an
- automatic degradation.
-
-- [ ] **M6+.2** -- Resolve the contention between a thread parked in `SubmitIoRing(ring, n, INFINITE, ..)`
- and callers wanting to build SQEs. This is the hard part and the reason this is its own milestone: it
- directly contradicts M3.1's `&mut`-enforced serialization, and needs either a submit-ownership handoff
- or an internal lock. Neither is obviously right.
-
-- [ ] **M6+.3** -- Shutdown: waking a thread parked on `INFINITE`. `IORING_OP_NOP` is supported and is the
- wake mechanism.
-
-- [ ] **M6+.4** -- Affinity: binding a ring's thread with `SetThreadGroupAffinity`, and documenting the
- execution-domain pattern (one pinned thread, its ring, its node-local registered pool, its shard).
-
-- [ ] **M6+.5** -- A test seam forcing the pinned-thread path even where the completion event is available,
- so it stays testable on every machine rather than only on hardware that lacks the feature.
-
-- [ ] **M6+.6** -- Decide `IoBuf`: extract to a shared crate, re-export from
- `windows-overlapped-io-sys`, or leave duplicated (D-1). The merge-or-delete decision that duplicate-then-decide
- defers to the point where the new path is proven -- which is here, not earlier.
+- [x] **M20.6** -- Re-evaluate `CommitStrategy::AlternatingRings` and the benchmark's conclusion
+ against [D-47](DESIGN-NOTES.md#d-47-detail). **The harness cannot exhibit a blast-radius
+ difference, which is a fact about the harness and not a finding against the strategy** -- each
+ lane's own arena is the limiter there. The strategy stays, with the conditions under which it
+ would pay written down; `M25.5` re-runs the comparison where operations genuinely pend. The
+ sample's output and prose are corrected so they stop claiming to measure a commit.
+ -> [completed 2026-09-23](COMPLETED-CHECKLIST.md#m206)
+
+
+## M21 -- Epoch-log review: correctness repairs
+
+Queued from
+[DESIGN-SESSION-2026-09-19-epoch-log-review.md](design-sessions/DESIGN-SESSION-2026-09-19-epoch-log-review.md)
+(findings `C-1` through `C-5`). Independent of each other; listed in ascending cost. Nothing in this
+milestone was observed failing at the sample's current constants -- these are a withdrawn justification,
+two hang shapes, a mis-keyed trigger, and a specification gap. (`M21.3` predicted that its trigger was
+merely unreachable *today*; measuring it while implementing showed it is unreachable at any constants, so
+what it corrected was the coupling rather than a latent bug. The archived entry has the numbers.)
+
+- [x] **M21.1** -- Correct the last site that still asserts [D-24](DESIGN-NOTES.md#d-24)'s withdrawn
+ half: the epoch-order assertion in the epoch-log committer, whose justification cited the hold-back
+ claim [D-47](DESIGN-NOTES.md#d-47) removed.
+ -> [completed 2026-09-20](COMPLETED-CHECKLIST.md#m211)
+
+- [x] **M21.2** -- Publish a bounded pop and the wait it is generic over, then remove the two unbounded
+ spins. `IoRing::pop_within` / `pop_within_with`, over a `CompletionWait` the caller supplies, because
+ [D-21](DESIGN-NOTES.md#d-21) means the crate cannot choose the wait for them.
+ -> [completed 2026-09-21](COMPLETED-CHECKLIST.md#m212)
+
+- [x] **M21.3** -- Key the epoch commit off a completed append rather than off the counter, so the
+ trigger cannot fire on a pass that appended nothing. The predicted latent bug turned out to be
+ unreachable at any constants -- measured, not re-reasoned -- so this is a coupling change rather than
+ a fix.
+ -> [completed 2026-09-21](COMPLETED-CHECKLIST.md#m213)
+
+- [x] **M21.4** -- State what a *failed* commit does to `durable_through`, and bind it with tests in both
+ directions. Required making the sample a test target at all (`test = true`), and gating the
+ failure-path tests on `fault-injection`, since a healthy flush cannot be made to fail.
+ -> [completed 2026-09-21](COMPLETED-CHECKLIST.md#m214)
+
+- [x] **M21.5** -- Give the harness's wait loops a bound, and collapse the hand-written waits onto the
+ bounded pop. The item named two loops; a census found four, plus two flaky single-`try_pop` sites.
+ -> [completed 2026-09-21](COMPLETED-CHECKLIST.md#m215)
+
+- [x] **M21.6** -- Fix the four defects an independent review of the `M21.2` surface found: the timeout
+ mapping, its victim in `run_down`, the `INFINITE` collision, and the test hole that hid all of them.
+ -> [completed 2026-09-21](COMPLETED-CHECKLIST.md#m216)
+
+## M21+ -- Queued by the 2026-09-21 API review
+
+Queued from the review of the `M21.2` surface, recorded in
+[DESIGN-SESSION-2026-09-21-m21-remediation-findings.md](design-sessions/DESIGN-SESSION-2026-09-21-m21-remediation-findings.md).
+Its other four findings were fixed in `M21.6`.
+
+- [x] **M21+.1** -- Teach [check-borrow-surface.ps1](../../tools/check-borrow-surface.ps1) the two shapes
+ it was blind to: methods of a `pub trait`, and borrows in parameter position. Four entries appeared, one
+ of them predating the widening; the probes also found a latent bug in the checker itself.
+ -> [completed 2026-09-21](COMPLETED-CHECKLIST.md#m21plus1)
+
+## M22 -- Epoch-log review: submission and arena
+
+Queued from the same session (findings `E-1` through `E-3`). `M22.1` is sequenced first because `M20.6`
+re-reads numbers that its change moves.
+
+- [x] **M22.1** -- Batch an epoch's appends into one submission in both append paths, and measure
+ whether the per-record submission cost was flattening the strategy comparison. It was not:
+ throughput did not move out of the noise, though commit p50 did.
+ -> [completed 2026-09-22](COMPLETED-CHECKLIST.md#m221)
+
+- [x] **M22.2** -- Collapse the two free-slot implementations to one, derived from the arena's own
+ outstanding counts rather than tracked beside them. The item called both correct; one was not --
+ the tracked free list leaked a slot on every refused append.
+ -> [completed 2026-09-22](COMPLETED-CHECKLIST.md#m222)
+
+- [x] **M22.3** -- Give the registered arena a stated placement: the epoch-log arena is placed on the
+ NUMA node its own log file's volume reports, and the allocator moved into the library as
+ `NumaBuffer` rather than being copied a second time. The sample says plainly that the placement
+ cannot pay at this workload.
+ -> [completed 2026-09-22](COMPLETED-CHECKLIST.md#m223)
+
+
+## M22+ -- Queued by what the M21 work left behind
+
+- [x] **M22+.1** -- Make [bounded_pop.rs](tests/bounded_pop.rs) independent of how fast a device is, by
+ reading from an overlapped pipe nobody has written to. Filed and completed the same hour; the deferral
+ was a scheduling preference rather than a blocker.
+ -> [completed 2026-09-21](COMPLETED-CHECKLIST.md#m22plus1)
+
+
+
+## M24 -- Make the unit suite hermetic
+
+The defect and its classification are [D-49](DESIGN-NOTES.md#d-49); the remedies and their costs are
+[DESIGN-SESSION-2026-09-21-hermetic-unit-tests.md](design-sessions/DESIGN-SESSION-2026-09-21-hermetic-unit-tests.md).
+It began at **63 of 131 lib tests opening a real kernel ring**, so `cargo test --lib` did not mean
+what its name implies, and the repository's own Quality rule already classifies an operating-system
+API as an external boundary.
+
+**Where it ended: 41 of 151 open a ring, 110 do not.** The remainder is not movable without
+`M26.2`'s FFI seam, and [D-53](DESIGN-NOTES.md#d-53) records the rung that keeps it from climbing
+back -- an inventory of *which* tests open a ring, since a zero-check would fail on day one and
+could only be satisfied by deleting coverage.
+
+**Sequencing is the open question, not whether. Corrected 2026-09-22: the cost of waiting is close
+to zero, which is the opposite of what this paragraph first said.** It claimed that waiting
+compounds, "because every milestone that adds tests adds to the pile to be migrated, and `M22` is a
+testing-heavy milestone". The mechanism is real but the instance was not checked, and it is false:
+**all three `M22` items touch only `examples/epoch_log/`**, and none adds a lib test.
+
+**Unconditional as of 2026-09-22.** `M24.1` concluded and `M24.4` is withdrawn, so nothing in this
+milestone waits on an evaluation any more. The hermetic goal is reached by relocation and by the
+accounting extraction alone; the technique `M24.1` went looking for turned out to be a different
+and larger thing, and is `M26`.
+
+Checked across the whole queue rather than for `M22` alone, since the first claim was wrong for
+want of exactly that: **no pending item outside this milestone modifies `src/**/tests.rs`.** `M22`
+is example-only; `M23.1` is the *sample's* `contract.rs`, not the crate's; `M20.1` and `M20.6` are
+documentation and the `ring_copy` sample; `M23.2` is a decision that may imply API later. The 63
+therefore do not grow while this waits.
+
+So sequencing turns on other things, and they point the other way:
+
+- **`M24.2` is an internals refactor of a published crate**, and the branch carrying this work is
+ already 19 commits with one `feat` and three `fix` commits on it. Stacking a field-layout change
+ on top makes one review cover both a new public API and that refactor.
+- **`M24.1` is an evaluation whose answer could invalidate `M24.4`**, so beginning the build before
+ it concludes risks building something the evaluation rejects.
+- **`M22.1` unblocks `M20.6`**, an open question since 2026-09-07 about whether a strategy still
+ earns its place in a published sample -- which is a decision waiting on a measurement `M22.1`
+ produces.
+
+**`M24.1` concluded (2026-09-22) and nothing here waits on it.** `M24.4` is withdrawn; `M24.2`,
+`M24.3` and `M24.7` are the path to a hermetic suite and are independent of each other.
+
+- [x] **M24.1** -- Settle whether a co-tested fake escapes the mock objection. **Answered: the fake
+ was the wrong instrument.** A shared suite is strong over what we specify and blind to the
+ platform's incidental behaviour, and an assertion about the latter is a frozen observation rather
+ than a contract. Superseded by the resolver in `M26`.
+ -> [completed 2026-09-22](COMPLETED-CHECKLIST.md#m241)
+
+- [x] **M24.2** -- Extract the handle-free accounting into its own type, composed by `IoRing`. The
+ item's field split was verified exactly: five fields carry no kernel state, five do.
+ `Accounting` now owns them with 19 hermetic tests, and `IoRing` delegates nine methods.
+ -> [completed 2026-09-22](COMPLETED-CHECKLIST.md#m242)
+
+- [x] **M24.7** -- Convert the lib tests that construct a ring only to exercise bookkeeping.
+ **61 -> 52**, by narrowing `Token::new` to take the ring's ledger rather than the ring. The
+ remaining 52 are not convertible and the reason is structural, not effort -- see the archive.
+ -> [completed 2026-09-22](COMPLETED-CHECKLIST.md#m247)
+
+- [x] **M24.4** -- **Withdrawn by `M24.1` (2026-09-22).** A shared conformance suite over a
+ hand-written fake is superseded by the response-space resolver in `M26`, which serves the same
+ purpose without encoding a belief about the platform at all. Nothing is deferred by this: `M24`'s
+ goal is a hermetic lib suite, and `M24.2` plus `M24.3` achieve that without it.
+
+- [x] **M24.3** -- Relocate the lib tests that open a ring but use only public API into `tests/`.
+ **52 -> 41.** Eleven moved; the "25" the item predicted was never achievable, and the reason is
+ the same structural one `M24.7` found.
+ -> [completed 2026-09-22](COMPLETED-CHECKLIST.md#m243)
+
+- [x] **M24.5** -- Put the rule on a rung. An **inventory** of which lib tests open a ring
+ ([D-53](DESIGN-NOTES.md#d-53)), not the zero-check the item assumed -- that rule is false and
+ could only be satisfied by deleting coverage. The guard's own bidirectional check found a defect
+ in the guard.
+ -> [completed 2026-09-22](COMPLETED-CHECKLIST.md#m245)
+
+- [x] **M24.6** -- Sweep what this milestone makes false. Two of the three sites the item named
+ were false alarms; the third was false for a different and larger reason than the item gave, and
+ the sweep found two more it did not name.
+ -> [completed 2026-09-23](COMPLETED-CHECKLIST.md#m246)
+
+## M23 -- The ring as a durability domain, and storage affinity
+
+Queued from the 2026-09-19 epoch-log review (findings `S-1` and `S-3`). `S-2` is an addendum to `M20.6`
+rather than an item here. `M23.1` and `M23.2` are done; `M23.3` is the remaining question, and it is
+about this crate's own surface rather than about storage at all.
+
+- [x] **M23.1** -- State in the epoch-log contract that the barrier is ring-wide while the flush names a file, so one ring per log is a precondition of the cost model. -> [completed 2026-09-23](COMPLETED-CHECKLIST.md#m231)
+
+- [x] **M23.2** -- Decide how a caller arrives at a NUMA node: `win-numa-sys` offers declaring and discovering, and refuses the shortcut that does both at once. -> [completed 2026-09-23](COMPLETED-CHECKLIST.md#m232)
+
+- [x] **M23.3** -- Decide what this crate offers for holding a token between push and completion: the ring owns the inventory, `IoRing` becomes generic, and the break is accepted. -> [completed 2026-09-23](COMPLETED-CHECKLIST.md#m233)
+
+- [x] **M23.4** -- Drop guards that panicked during unwind aborted the process instead of reporting; they now stay silent while `std::thread::panicking()`. -> [completed 2026-09-23](COMPLETED-CHECKLIST.md#m234)
+
+- [x] **M23.5** -- Both asserts in `IoRing::drop` are now reached by tests; the raw-HRESULT seam the item priced turned out not to be needed, because the kernel refuses a null ring handle cleanly. -> [completed 2026-09-23](COMPLETED-CHECKLIST.md#m235)
+
+
+## M26+ -- The wakeup window review opened
+
+- [ ] **M26.12** -- **Find why a signal raised just after `wait.arm` can be lost, and fix it.**
+ Raised as a narrower finding by Copilot review on PR #108 -- that `EventDelivery::new` signals
+ only when it attached the event itself -- and the investigation found something wider.
+
+ **What is measured.** The `#[ignore]`d reproducer
+ `a_backlog_is_delivered_even_when_the_caller_attached_the_event_first` in
+ [event_delivery.rs](tests/event_delivery.rs) fails 6 of 6. Signalling unconditionally, which is
+ what the review suggested, does not change that. A 50 ms sleep between `wait.arm` and the signal
+ makes it pass 3 of 3, and so does `--features trace`, which is the same perturbation by another
+ route. Full figures in [UNRESOLVED-TEST-FAILURES.md](UNRESOLVED-TEST-FAILURES.md).
+
+ **Why this is not a small follow-up.** [D-68](DESIGN-NOTES.md#d-68) fixed `M26.9`'s stall by
+ ordering the arm before the signal, measured at 0 failures in 3600 runs. This says that ordering
+ narrows the window rather than closing it, so the decision's reasoning needs revisiting once the
+ mechanism is known -- not before, because the mechanism is currently a guess.
+
+ **Do not apply the sleep.** It is a diagnostic that identified a window, not a fix, and shipping
+ it would convert a reproducible defect into a rare one.
+## M27 -- What this crate owes the topology planner
+
+**Re-planned 2026-09-23, the same day it was written.** M27 was originally "Adaptivity: the benefit
+without the architectural commitment", and asked whether *this crate* should derive a partition for a
+consumer who expresses no preference. That was the wrong owner, and the checklist rules require
+saying so rather than quietly rewriting it. The adaptivity the
+[adoption thesis](../../DESIGN-NOTES.md#the-adoption-thesis) asks for is delivered by
+[topology-planner](../topology-planner/COMPONENT.md), which takes a dataflow description of the
+application and returns one or more suggested realizations
+([EP-D-6](../topology-planner/DESIGN-NOTES.md#ep-d-6)). Had the original M27.1 been answered here it
+would have grown a second, weaker policy surface beside the one that component exists to provide --
+the `outermost_partitioning_cache` defect again, where a policy answer lands in a crate whose job is
+something else.
+
+> **-> CROSS-COMPONENT PREREQUISITE:** `M27.1` and `M27.2` are gated on component
+> `crates/topology-planner` -> `M1+` -> `EP-1+.5` and `EP-1+.6`, which decide the plan vocabulary
+> this crate would be realized from. See [CHECKLIST.md](../topology-planner/CHECKLIST.md).
+
+**What survives here is the realization end, not the policy end.** The planner emits a plan; the
+outward adapter realizes it as buffers, rings and threads
+([EP-D-5](../topology-planner/DESIGN-NOTES.md#ep-d-5)). That adapter is a separate crate, but it can
+only build what this crate exposes, and nothing has ever checked that what it exposes is sufficient.
+[D-8](DESIGN-NOTES.md#d-8) is untouched by all of this: policy stays out of this crate, and being
+*constructible from* a policy decision made elsewhere is the opposite of taking one.
+
+- [ ] **M27.1** -- **Census what a realizer would need from this crate, against the plan vocabulary,
+ and name what is missing.** A plan states which processor a domain pins to, which memory node its
+ pool allocates from, how many queues of which types, and where each channel's buffer lives. Walk
+ each of those to the public API that would realize it and record the gaps. `NumaBuffer`
+ ([D-51](DESIGN-NOTES.md#d-51)) is one half of the pool answer and arrived this month; the ring's
+ own construction takes no placement input at all. **The output is a gap list, not an API** --
+ proposing surface before the plan vocabulary is settled would be binding to a draft.
+
+- [ ] **M27.2** -- **Gated on `M27.1` and on the planner's `EP-1+.6`.** Close the gaps the census
+ names, as ordinary capability on this crate with no policy attached. Each gap is an input a caller
+ supplies, never a choice this crate makes. Verify the way the thesis demands rather than the
+ convenient way: construct from a plan built against a *synthetic* machine, since the planner is
+ mockable by construction and this crate should be realizable without the hardware the plan
+ describes.
+
+- [ ] **M27.3** -- Give a consumer the means to answer placement questions on their own hardware.
+ **Not gated on the planner** -- it is the client-side half of the thesis, and it is what lets a
+ developer disagree with any plan they are handed. `cache_domains.rs` now prints every cache level
+ beside the heuristic's pick; the equivalent for placement is a sample that reports what a chosen
+ arrangement costs and what the alternatives would have cost, on the machine in hand.
+ [ring_copy](examples/ring_copy) is the natural host, being already policy-selectable. **Do not ship
+ a verdict** -- report the observation and let the consumer conclude, per OPTION INTEGRITY.
+
+## M28 -- The ring owns the pending inventory
+
+Queued by [D-55](DESIGN-NOTES.md#d-55), taken 2026-09-23 after the `M23.3` exploration. The break
+is accepted deliberately: `IoRing` becomes generic so the inventory cannot drift from the ring,
+because a consumer never holds a token to lose. The exploration and everything it falsified is in
+[DESIGN-SESSION-2026-09-23-pending-inventory.md](design-sessions/DESIGN-SESSION-2026-09-23-pending-inventory.md);
+`src/pending.rs` is the working spike and is the shape the internal map starts from.
+
+**Sequenced so each step compiles.** The published crate is at 0.3.1, so this is a major bump and
+every consumer names the type -- which means the migration order matters more than usual.
+
+- [ ] **M28.1** -- **Decide what a caller receives, before writing any of it.** If the ring owns
+ the token then `Batch::write` can no longer hand one back, and the shape of what replaces it is
+ the whole design: an identity the caller matches later, or a claim that returns `(T, X)`
+ directly from the ring. The second makes drift impossible and is the point of the break; the
+ first is a smaller change that may not be worth breaking for. Settle it with the
+ `Token::claim_if` safety argument in hand, since that is what currently makes a mismatched
+ completion unclaimable.
+
+- [ ] **M28.2** -- **Bound `RingContract` before anything depends on it more heavily.**
+ `operations: HashMap` is never pruned -- `observe_claim` marks an entry
+ `Completed` and keeps it -- so the oracle retains one entry per operation for the process's
+ life. Undocumented, and not visible in the sample because it appends 24 records. A long-running
+ consumer following the crate's own recommendation leaks. This blocks any design that checks by
+ default, which is why it is here rather than filed separately: `M23.3` reached for always-on
+ checking and this is what ruled it out. Decide whether completed entries are dropped, whether
+ `check_quiescent` needs them, and document the answer either way.
+
+- [ ] **M28.3** -- **Gated on `M28.1`.** Make `IoRing` generic and move the inventory inside.
+ Carry the sidecar: the census found two thirds of consumers keep per-operation data beside the
+ token, so an inventory that holds only tokens serves a minority. Mixed-shape consumers use a
+ closed `enum` -- `tests/generated_sequences.rs` is the worked example and needs no change to
+ keep working.
+
+- [ ] **M28.4** -- **Gated on `M28.3`.** Migrate the ~12 consumers, and delete `Pending` or
+ demote it to the internal map. **Convert all of them or none**: converting a few relocates the
+ duplication rather than removing it, which is the lesson `win-numa-sys` recorded the same day
+ when it moved one `VirtualAllocExNuma` and left the other.
+
+- [ ] **M28.5** -- **Answer the tokenless push.** `flush_raw` returns a bare `usize` and
+ `epoch_log`'s commit path depends on it, because a flush has no buffer and a *borrowed*
+ `RawHandle` gives its token nothing to guard. An inventory the ring owns has to say what it
+ does with operations that have no token -- `RingContract` already models them separately with
+ `observe_tokenless_push`. Note this may dissolve rather than need solving: if the sample owned
+ a `SharedFile` instead of passing a `RawHandle` it could use the safe `flush` and get a token,
+ which `M25.3` reopens anyway by changing how the log is opened.
+
+- [ ] **M28.6** -- **Sweep what the break makes false**, including the README's ring examples, the
+ `D-4` detail section, and every rustdoc that tells a caller to match a completion against a
+ held token -- `Completion::user_data` and `IoRing::push_raw` both do, and they are the evidence
+ D-55 rests on, so they are the first things the change invalidates.
diff --git a/crates/windows-ioring-sys/COMPLETED-CHECKLIST.md b/crates/windows-ioring-sys/COMPLETED-CHECKLIST.md
index 84dbdb328..077e85696 100644
--- a/crates/windows-ioring-sys/COMPLETED-CHECKLIST.md
+++ b/crates/windows-ioring-sys/COMPLETED-CHECKLIST.md
@@ -1565,3 +1565,2259 @@ the API whose breaking change 0.2.0 is being cut for, and it is reachable with n
and D-45 is added to its table of shipped defects of this shape.
**Swept the count restatements too:** that file said "three defects" in four places and is now four, which
is the restatement drift the repository's own conventions warn about.
+
+## Moved 2026-09-19 22:40:58 -07:00 -- M20.4: the file-handle NUMA mechanism correction
+
+### M20.4 -- Correct "What is not reachable" in [DESIGN-NOTES.md](DESIGN-NOTES.md): the file-handle-to-storage-node mapping is reachable on mechanism, and the conclusion it supported now rests on volume granularity, absence, and spanned volumes instead. *(completed 2026-09-19 22:40:58 -07:00)*
+
+The authoritative text is the rewritten "What is not reachable" section in
+[DESIGN-NOTES.md](DESIGN-NOTES.md); the research behind it is `F-1` in
+[DESIGN-SESSION-2026-08-30-numa-sharded-io-execution-domains.md](../../design-sessions/DESIGN-SESSION-2026-08-30-numa-sharded-io-execution-domains.md).
+
+Two things the item's own text had wrong, corrected while doing it rather than copied forward:
+
+- It said to cite [file-handle-numa-spike.rs](design-sessions/spikes/file-handle-numa-spike.rs) as the
+ **unrun** instrument, and to state that no measurement of either call succeeding on an ordinary NTFS
+ data file could be found. `F-1a` of the same session had already smoke-run it: both calls succeed on an
+ ordinary NTFS data file and on a directory handle, and agree. The item was written from `F-1` without
+ `F-1a`. What remains unmeasured is narrower -- whether either call ever names a node that distinguishes
+ one device from another, which needs a multi-node host with storage whose PDO advertises a proximity
+ domain.
+
+- The spikes [README.md](design-sessions/spikes/README.md) carried the same staleness, each instance
+ contradicted by its own body a few paragraphs later. Swept: 3 phrasings in 1 file, plus the sentence
+ promising that a multi-node run "would correct a claim DESIGN-NOTES.md currently makes", which this item
+ has now made false -- it settles an open question instead.
+
+The heading stays "What is not reachable". What is not reachable is the *answer* a ring consumer wants,
+which is still true; renaming it would dangle the pointers in the 2026-09-19 review session.
+
+## Moved 2026-09-19 22:49:09 -07:00 -- M20.2: the ARM no-L3 measurement recorded as D-48
+
+### M20.2 -- Record the 2026-08-30 ARM measurement as a decision, beside the zero-NUMA-node observation it is the sibling of. *(completed 2026-09-19 22:49:09 -07:00)*
+
+Landed as [D-48](DESIGN-NOTES.md#d-48) in the decision index, plus a sibling paragraph in
+"Why the NUMA node is the wrong key" where the existing zero-node observation lives, which is where the
+item asked for it.
+
+Two choices worth recording, because both were places this could have gone wrong:
+
+- **The measurement is cited, not re-transcribed.** The capture is Measurement M-1 in
+ [DESIGN-SESSION-2026-08-30-numa-sharded-io-execution-domains.md](../../design-sessions/DESIGN-SESSION-2026-08-30-numa-sharded-io-execution-domains.md),
+ and D-48 links it rather than copying the probe output into a third place. Pasting the block would have
+ created exactly the restatement the repository conventions warn about.
+
+- **The false clause it falsifies is marked, but not rewritten.** D-48 sits two paragraphs from the
+ sentence saying the last-level-cache domain "is meaningful on Intel and ARM too", which this
+ measurement shows is false on a shipping part -- so leaving it unmarked would have made the document
+ contradict itself. A one-line adjacent marker says so and points at `M20.1`. The restatement of the
+ rule and the sweep across the README, `lib.rs` and `policy.rs` remain `M20.1`, which is coupled to
+ `SH-4.12` and must follow it.
+
+## Moved 2026-09-20 23:16:19 -04:00 -- M21.1: the last site asserting D-24's withdrawn half
+
+### M21.1 -- Correct the last site that still asserts [D-24](DESIGN-NOTES.md#d-24)'s withdrawn half: the epoch-order assertion in the epoch-log committer, whose justification cited the hold-back claim [D-47](DESIGN-NOTES.md#d-47) removed. *(completed 2026-09-20 23:16:19 -04:00)*
+
+The assertion is unchanged, because it was always sound -- just for the other reason. Every commit
+carries `FlushCoverage::CoversPrecedingOperations`, so commit *N* is outstanding when commit *N+1* is
+reached, and D-47's *surviving* half -- no operation queued before a drained flush was ever observed
+completing after it -- is what orders them. The comment now says that, and says explicitly what it does
+not rest on, so a later reader does not "restore" the withdrawn reasoning.
+
+Sweep re-run before committing, as the item required. **21 matches across 14 files**, disposed as:
+
+- **1 violation**, fixed: the justification in [commit.rs](examples/epoch_log/commit.rs).
+- **1 historical site**, marked rather than rewritten:
+ [DESIGN-SESSION-2026-08-28-external-consumer-correspondence.md](design-sessions/DESIGN-SESSION-2026-08-28-external-consumer-correspondence.md)
+ records what the findings became on that date, and glossed D-24 as a stall that "holds operations
+ against unrelated files". A session is a faithful record of its moment, so the gloss stays and a
+ one-line note beside it says which half was later withdrawn.
+- **2 false positives**, left alone: the spikes README ("checked in deliberately rather than held back",
+ about an instrument) and [fault_injection.rs](tests/fault_injection.rs) ("holds it until the completion
+ is claimed", about a token).
+- **17 already correct** -- the decision index, both sides of the crate docs, the README, the stress
+ tests, strategy.rs, and the item's own text.
+
+The item quoted the original sweep as "17 matches across 10 files". Re-running it found more of both,
+and the file count needed care: `rg` groups the two design-session files under a single header, so the
+first reading of its output undercounted by one. Counted with a command rather than by eye.
+
+Verified by running the example in a debug build, where the `debug_assert` is live: it completed, the
+negative control still caught the corrupted record, and all four epochs reported durable in order.
+
+## Moved 2026-09-21 00:23:49 -04:00 -- M21.2: a bounded pop, generic over the wait
+
+### M21.2 -- Publish a bounded pop and the wait it is generic over, then remove the two unbounded spins. *(completed 2026-09-21 00:23:49 -04:00)*
+
+**Decided: the trait, plus a default impl and a convenience.** The crate cannot choose the wait for a
+caller, and that is a contract rather than a shrug -- the completion event is auto-reset with exactly one
+waiter per ring ([D-21](DESIGN-NOTES.md#d-21)), so a wait this crate picked could consume an edge the
+caller's own loop was entitled to. Making it a parameter moves the obligation to the only party who can
+discharge it.
+
+The surface added to [ring.rs](src/ring.rs):
+
+- `IoRing::pop_within(timeout)` -- the convenience, using `SubmitWait`.
+- `IoRing::pop_within_with(wait, timeout)` -- the same loop over a caller-chosen wait; `?Sized`, so a
+ trait object works.
+- `CompletionWait` -- one method, whose contract is deliberately weak: returning early or spuriously is
+ always permitted, because [D-19](DESIGN-NOTES.md#d-19) already makes a wake with nothing to pop normal.
+ An implementation cannot be subtly wrong about *when* to return, only wasteful.
+- `RingWait` -- the ring narrowed to `block` and `outstanding`, so a waiter cannot pop the completion
+ its own caller is waiting for, nor submit work nobody asked for. Same narrowing, same reason, as
+ `RingScope` under [D-43](DESIGN-NOTES.md#d-43).
+- `SubmitWait` -- blocks inside `SubmitIoRing` with nothing queued, touching no event, so it composes
+ with a caller who owns the completion event.
+
+**The borrow question, answered even though the check says the surface is unchanged.** `RingWait` is only
+ever passed *in*, never returned, so [BORROW-SURFACE.txt](BORROW-SURFACE.txt) is untouched (verified: 7
+entries, unchanged). Asked anyway, since it is a lifetime-carrying wrapper: safe code reaching it can call
+`block` and `outstanding` and nothing else; the ring it borrows is held exclusively for the call, so
+nothing it could invalidate is live elsewhere; and the narrow type is the point rather than an accident,
+because a bare mutable reference to the ring would have permitted exactly the two things a waiter must
+not do.
+
+**Measured while building it, and now documented on `RingWait::block`:** `SubmitIoRing` answers
+`E_INVALIDARG` (`0x80070057`) -- not a timeout -- when asked to wait for a completion the kernel has no
+pending operation for. Found by a test that drove the loop with a reservation having no real SQE behind
+it. The precondition holds structurally: `pop_within_with` checks `outstanding()` before consulting the
+wait, so `block` is unreachable with nothing pending.
+
+**An early return that is not an optimisation.** With nothing outstanding and an empty queue no completion
+is possible, so the loop answers immediately rather than sleeping out the bound. That turns "you forgot to
+submit" from a timeout into an instant answer, and it is what the sabotage below pins.
+
+**A panic path found and closed before it shipped.** The first draft computed an instant plus the caller
+timeout directly, which panics on overflow -- so `Duration::MAX`, a reasonable spelling of "no deadline",
+would have taken down the process. Now `checked_add`, with an unrepresentable deadline treated as one that
+never arrives. Two tests cover it: one where the early return answers first, one where an operation is
+outstanding so the overflow branch is actually reached.
+
+**Sabotage-verified**, because a test that cannot go red proves nothing:
+
+- Removing the clamp that keeps the wait timeout above zero turns `the_wait_is_never_handed_a_zero_timeout`
+ red.
+- Removing the nothing-can-arrive early return turns two tests red, and the run takes the full 30-second
+ bound instead of finishing instantly -- which is the behaviour the early return exists to remove.
+
+**Call sites converted:** the two unbounded spins in
+[append.rs](examples/epoch_log/append.rs) and [fault_injection.rs](tests/fault_injection.rs), and the
+test-only `pop_within` helper, which is now a thin panicking wrapper over the public API rather than a
+fourth copy of the loop. The panic is the only test-specific part left: a test wants the name of what it
+waited for, a consumer wants an `Option` it can act on.
+
+## Moved 2026-09-21 02:23:34 -04:00 -- M21.3: the epoch trigger, and a review claim the measurement disproved
+
+### M21.3 -- Key the epoch commit off a completed append rather than off the counter, so the trigger cannot fire on a pass that appended nothing. *(completed 2026-09-21 02:23:34 -04:00)*
+
+The match in [main.rs](examples/epoch_log/main.rs) now yields a `closed_an_epoch` value that every arm
+must produce, rather than a `appended % EPOCH_SIZE == 0` test written after it. Closing an epoch is a
+fact about an append that landed, and making each arm answer the question keeps that local -- a new arm
+added later cannot fall through into a commit.
+
+**The item predicted this was a latent bug. It is not, and the measurement is what settled it.**
+The retry path was instrumented to report when the old shape would have committed, and run at
+`EPOCH_SIZE` of 6, 8, 12, 16 and 24. It fired **zero** times -- including at every value past `SLOTS`,
+which both the item and finding `C-3` predicted would arm it.
+
+The reason is an invariant three blocks from the trigger: `appended % EPOCH_SIZE == 0` is true at exactly
+two moments -- before the first append, and immediately after a commit -- and the arena is empty at both,
+because the commit waits for a covering flush that retires every outstanding write. An append is
+therefore never refused at an epoch boundary, at any constants.
+
+**So this is a coupling change, not a bug fix**, and the distinction is the useful part. The old trigger
+was safe because of something nothing stated, three blocks away; the new one cannot fire because of where
+it is written. The second survives a reader who changes the commit path. The first is what made the
+question take a measurement to answer at all.
+
+Swept the claim rather than only fixing the code: finding `C-3` in
+[DESIGN-SESSION-2026-09-19-epoch-log-review.md](design-sessions/DESIGN-SESSION-2026-09-19-epoch-log-review.md)
+carried the same wrong prediction and now carries the correction beside it. The review lesson recorded
+there is the narrow one: "unreachable today, armed tomorrow" is a claim about a program's reachable
+states, and reading the code is not how to settle one.
+
+## Moved 2026-09-21 16:06:19 -04:00 -- M21.4: what a failed commit means, and the sample's first tests
+
+### M21.4 -- State what a failed commit does to durable_through, and bind it with tests in both directions. *(completed 2026-09-21 16:06:19 -04:00)*
+
+The doc on `Committer::claim` said "A failed commit advances nothing", which reads as a permanent verdict.
+It is not. Every commit here is a **covering** flush, so commit *N+1* reaches epoch *N*'s writes -- queued
+before it -- and observing *N+1* makes *N* durable after all. What makes a record durable is a flush that
+covered it, not the identity of the flush named for its epoch. The doc now says that, and says what a
+caller must not read into a failure: not "epoch *N* is lost", but "not yet".
+
+**The sample had no tests at all.** Examples are not test targets by default, so `cargo test` compiled
+this one and ran nothing. Binding the claim meant adding `test = true` to the `[[example]]` entry first;
+that is the change that makes any of the sample's policy testable, not just this item's part of it.
+
+**The failure path is unreachable by running the sample**, because a flush against a healthy temp file
+does not fail. The crate's fault-injection seam is the only way in, so four of the six tests are gated on
+`fault-injection` and the other two run by default. That follows the precedent and the reasoning already
+written down in [fault_injection.rs](tests/fault_injection.rs), including that CI's
+`cargo test --workspace --all-features` job is what stops gated tests from being tests that never run.
+
+**An assumption caught by asserting it.** The test first asserted that an injected `ERROR_ACCESS_DENIED`
+would surface as `io::ErrorKind::PermissionDenied`. It surfaces as `Other`: the crate preserves the
+HRESULT in an `IoRingError` rather than classifying it. Corrected to assert the Win32 code, which is the
+assertion `tests/fault_injection.rs` already makes one layer down.
+
+**Sabotage-verified in both directions, and the two produce different failure sets** -- which is what
+shows the tests discriminate rather than all keying on one fact:
+
+- `is_durable` returning `true` unconditionally: **4 of 6 fail**, caught by the assertions that an epoch
+ is *not* yet durable.
+- `is_durable` requiring an exact match with the watermark: **2 of 6 fail**, caught by the covering and
+ monotonicity tests.
+
+**Leverage from earlier in this milestone:** `commit_and_pop` is three lines because `M21.2` published
+`IoRing::pop_within`. Without it every test here would have carried its own bounded wait, which is the
+duplication `M21.2` existed to remove.
+
+## Moved 2026-09-21 18:18:52 -04:00 -- M21.5: every unbounded wait in the crate, not the two that were named
+
+### M21.5 -- Give the harness's wait loops a bound, and collapse the hand-written waits onto the bounded pop. *(completed 2026-09-21 18:18:52 -04:00)*
+
+`await_flush` and `await_writes` in [strategy.rs](examples/epoch_log/strategy.rs) blocked in
+`submit_and_wait` for `WAIT_MS`, ignored that it had returned without a completion, and went round
+again forever. Both now carry a deadline and raise `TimedOut`, which is the policy
+[event_loop.rs](examples/epoch_log/event_loop.rs) already documented: a measurement harness that hangs
+reports nothing, which is strictly worse than one that fails.
+
+**The item named two loops. A census found four, and two more of a related shape.** Counted by command
+over every `.rs` outside `target/` and the spikes:
+
+- [strategy.rs](examples/epoch_log/strategy.rs) `await_flush` and `await_writes` -- the two named.
+- [failure_paths.rs](tests/failure_paths.rs) and [kernel_span.rs](tests/kernel_span.rs), each with a
+ helper **called `await_one`**, byte-identical to the other, neither named by the item.
+- [batch/tests.rs](src/batch/tests.rs), where two registration waits were a single `try_pop` -- the
+ flake shape `pop_within` documents -- in a file whose *third* such wait already used the helper.
+ One predicate, three sites, half-converted, which is FAIL FAST rule 1 exactly.
+
+**Sabotage-verified, and it revealed the milestone compounding.** Making `classify` stop filing flush
+results leaves `await_flush` looking for a completion that is never recorded -- an unbounded loop would
+hang forever. It failed with `timed out after 30s waiting for a commit's flush`, and it did so in **two
+seconds**, not thirty: `pop_within`'s nothing-can-arrive early return (`M21.2`) answers immediately once
+the ring is quiesced. The bound is what makes the failure possible; the early return is what makes it
+quick.
+
+A `remaining` helper and a single `timed_out` constructor keep the two waits from describing the same
+condition two ways, and `WAIT` is derived from `WAIT_MS` rather than written twice.
+
+Factoring `Lane::classify` out of `Lane::drain` is what let the bounded waits file a completion they
+blocked for without a second copy of the claim-then-check logic.
+
+## Moved 2026-09-21 21:15:12 -04:00 -- M21.6: the API review's four defects
+
+### M21.6 -- Fix the four defects an independent review of the M21.2 surface found: the timeout mapping, its victim in `run_down`, the INFINITE collision, and the test hole that hid all of them. *(completed 2026-09-21 21:15:12 -04:00)*
+
+Queued and closed in one item because the first two share a root cause and the fourth is the reason
+neither was caught. The surface is unreleased and on a branch, so none of this is a breaking change.
+
+**The defect.** `SubmitIoRing` reports an expired wait as `ERROR_TIMEOUT` (`0x800705B4`), a *failure*
+HRESULT. `RingWait::block` passed it through `check`, so `pop_within` returned `Err` on every real
+timeout and never the `Ok(None)` it documents. All six converted call sites treat `Err` as fatal and
+`Ok(None)` as the timeout signal, so the `ErrorKind::TimedOut` mapping they document was dead code on
+the only path that produces it.
+
+**Its second victim, pre-existing.** `run_down` polls in 50 ms steps with the same `check`, so any
+operation slower than 50 ms made rundown return `Err` with work still outstanding -- and `Drop` then
+asserted and called `CloseIoRing` anyway, which is precisely the "the kernel may still be writing
+through a token's buffer" hazard the function exists to prevent. One `wait_outcome` helper now
+classifies a wait-only `SubmitIoRing` result for both callers, so they cannot disagree again.
+
+The rundown loop is deliberately left unbounded. Every SQE that queues produces exactly one completion
+(M10.2), so it terminates; blocking until that holds is the safe failure, and closing the ring early is
+not.
+
+**The `INFINITE` collision.** A finite bound above ~49.7 days saturated onto `u32::MAX`, which is Win32's
+`INFINITE` -- and `CompletionWait` explicitly invites implementations built on `WaitForMultipleObjects`.
+Clamped to `MAX_WAIT_MS` (one below), and the trait now states that `timeout_ms` is never zero and never
+`INFINITE`, so it can be passed straight to a Win32 wait.
+
+**The contract gap that would have propagated it.** `CompletionWait` never said how to report an expired
+wait, and every Win32 wait reports one as an error -- so any third-party implementation forwarding its
+underlying result would have reproduced the defect exactly. The trait now says an expired bound is
+`Ok(())`, and says why.
+
+**Why no test caught it, which is the finding worth keeping.** The M21.2 tests drive the loop with a wait
+that never enters the kernel. Deterministic, and it leaves the Win32 interaction untested: replacing
+`RingWait::block`'s entire body with an unconditional error left **the whole suite green**.
+
+Closing that needed an operation still pending when the bound expires, and three attempts failed before
+one worked -- each measured, not assumed:
+
+| Attempt | Result |
+|---|---|
+| Buffered read, up to 256 MiB | completion already poppable, 3-5 us |
+| Flush over 512 MiB of dirty cache | 3 us -- the lazy writer had already written it back |
+| Unbuffered read, 256 MiB | 3 us -- **a handle without `FILE_FLAG_OVERLAPPED` is synchronous**, so the read completes inline during submit |
+| Unbuffered **and** overlapped, 64 MiB+ | genuinely pending; the bound expires |
+
+That third row is the general finding: **the crate's existing tests all use synchronous handles**, so
+ring operations complete inline during submit and asynchronous completion is never exercised. That is why
+a counting waiter over the existing flush pattern was reached in 0 of 50 trials.
+
+[tests/bounded_pop.rs](tests/bounded_pop.rs) now covers it with five tests over a 128 MiB unbuffered,
+overlapped read. Each asserts `outstanding() > 0` alongside the expected answer, so a machine fast enough
+to finish the read early fails loudly instead of passing vacuously. Nothing asserts an upper bound on
+elapsed time: Windows' default timer resolution is ~15.6 ms, so a 5 ms bound routinely takes 14-19 ms.
+
+**Sabotage-verified, twice.** Reverting the timeout mapping turns **all five** red. Reproducing the
+review's original mutation -- `block` always failing -- now turns three red, where before it turned none.
+
+The review's fifth finding, that `check-borrow-surface.ps1` is blind to trait methods and to borrows in
+parameter position, is queued as `M21+.1` rather than fixed here: it is a process gap, not a runtime
+defect, and widening the check obliges a borrow-question answer for every entry it newly reports.
+
+## Moved 2026-09-21 21:23:24 -04:00 -- M21+.1: the borrow-surface check learns two shapes
+
+### M21+.1 -- Teach [check-borrow-surface.ps1](../../tools/check-borrow-surface.ps1) the two shapes it was blind to: methods of a \pub trait\, and borrows in parameter position. *(completed 2026-09-21 21:23:24 -04:00)*
+
+The check inspected only the text after the last `->` on lines matching `pub fn`. Trait items are
+declared `fn`, not `pub fn`, so nothing inside any `pub trait` was ever examined; and a borrow-carrying
+type in *parameter* position was invisible in any function. `CompletionWait::wait` is both at once.
+
+**Four entries appeared, and one of them predates the widening by months.**
+`IoRingErrorExt::as_ioring_error` returns `Option<&IoRingError>` and had simply never been inventoried --
+the blind spot made concrete rather than a new risk. The other three are `CompletionWait::wait`,
+`Batch::new` and `EventDelivery::new`, the last two carrying an explicit lifetime in a parameter.
+The borrow question is answered for all four in
+[DESIGN-NOTES.md](DESIGN-NOTES.md#borrow-surface-audit-m21plus1), before the inventory was regenerated,
+as [DESIGN-INSTRUCTIONS.md](DESIGN-INSTRUCTIONS.md) requires. None is a hole.
+
+**A plain `&T` parameter is deliberately not reported.** Reporting every method that borrows something
+would list the whole crate and mean nothing. Only an *explicit lifetime* in parameter position counts --
+the borrow-carrying wrapper whose lifetime the crate chose, not a reference the caller lent us. That is a
+heuristic, and the archived audit says so: it would miss a `&dyn Trait` parameter carrying no named
+lifetime.
+
+**Verified by five probes, restored afterwards** -- and the negative control is the one that matters,
+because a check that fires on everything is as useless as one that fires on nothing:
+
+| Probe | Expected | Result |
+|---|---|---|
+| `pub fn` returning `&[u8]` (control, the old rule) | caught | caught |
+| `pub trait` method returning `&[u8]` | **now caught** | caught |
+| `pub fn` taking `&mut RingWait<->` | **now caught** | caught |
+| `pub fn` taking a plain `&Completion` | **silent** | silent, exit 0 |
+| `pub trait` with a default body plus a borrow-returning sibling | only the sibling | only the sibling |
+
+**The probes found a latent bug in the checker itself**, which is the argument for running them rather
+than reasoning about the regex. A one-line body -- `pub fn f() -> &[u8] { &[] }` -- never satisfied the
+"line ends with `{`" test, so the signature accumulator ran past the end of the file. The old script did
+not crash on it only because it never indexed the lines again afterwards; it silently swallowed the
+following lines instead. Signature termination is now "the accumulated text contains a `{`", and the
+return type is truncated at that brace.
+
+Also swept while here: the script header and its failure message both said **three** shipped defects of
+this shape and listed D-35, D-36, D-43. It is four, and has been since D-45.
+[M19.3](COMPLETED-CHECKLIST.md) swept that count through `DESIGN-INSTRUCTIONS.md` and missed this file,
+which is the restatement-drift pattern landing on the very tool built to stop a different one.
+
+## Moved 2026-09-21 22:08:52 -04:00 -- M22+.1: a pending operation that owes nothing to a device
+
+### M22+.1 -- Make [bounded_pop.rs](tests/bounded_pop.rs) independent of how fast a device is, by reading from an overlapped pipe nobody has written to. *(completed 2026-09-21 22:08:52 -04:00)*
+
+**Queued and completed within the hour, and the queueing was the error.** It was filed as `M22+.1` with
+an entry in `UNRESOLVED-TEST-FAILURES.md` on the grounds that the fix did not belong in a push of
+finished milestones. That is a scheduling preference, not a blocker, and the repository's PRIME
+DIRECTIVE is explicit that only a genuine blocking factor justifies deferral. The mechanism was
+understood when it was filed; the two open questions were each one probe away.
+
+**Both probes answered, and neither was safe to assume.** `IoRing` does accept a pipe handle for
+`read_raw`; and `pop_within(20ms)` against an unwritten overlapped pipe returns `Ok(None)` with
+`outstanding == 1`. `Win32_System_Pipes` was added to the dev-dependency feature set; `PIPE_ACCESS_INBOUND`
+is not re-exported where the module name suggests, so it is a named local constant, as
+`FILE_FLAG_NO_BUFFERING` already was in this file.
+
+**The substance of the change is the question the test asks.** A 128 MiB unbuffered read asks "will this
+device take longer than 5 ms?" -- a question about someone else's hardware, which may answer differently
+on two runs of the same machine. A pipe with no writer asks nothing: the read is pending because no byte
+exists to satisfy it, and it completes exactly when the test writes one.
+
+Where a delay is still needed -- `run_down` polls in 50 ms steps, so forcing it to observe an expired poll
+means releasing the read later than that -- it comes from a `thread::sleep`, whose guarantee runs the safe
+way round: a sleep may overshoot, never undershoot. No assertion depends on an operation *finishing*
+within any bound.
+
+**Verified:** 25 consecutive runs of the target and 3 full `--all-features` suite runs, all green; both
+feature configurations green. **Both sabotages still bite exactly as before the rewrite** -- reverting the
+timeout mapping turns all 5 red, making `RingWait::block` always fail turns 3 red -- which is the
+assertion that matters, because a deterministic test that had lost its discriminating power would be a
+worse outcome than the flake.
+
+Also 11x faster (0.20s against 2.26s), with no 128 MiB fixtures.
+
+The `UNRESOLVED-TEST-FAILURES.md` entry moved to
+[RESOLVED-TEST-FAILURES.md](RESOLVED-TEST-FAILURES.md) in this commit, per the append-only rule.
+
+## Moved 2026-09-22 15:54:53 -04:00 -- M22.1: batching the appends, and the confound it was meant to test
+
+### M22.1 -- Batch an epoch's appends into one submission in both append paths, and measure whether the per-record submission cost was flattening the strategy comparison. *(completed 2026-09-22 15:54:53 -04:00)*
+
+`Appender::append_batch` and `Lane::append_batch` replace the single-record pushes. Both compose as
+many records as there are free arena slots and submit once, so the arena rather than the caller's
+list decides the batch size -- which is what keeps the two halves of an append honest, since a slot
+is composed into only while the kernel is not reading it. With eight slots, an epoch of 64 records
+goes from 64 submissions to 8.
+
+**The teaching defect is the smaller half.** A sample whose job is to teach `Batch` was paying one
+`SubmitIoRing` per record, which is the one thing `Batch` exists to avoid.
+
+**The measurement was the point, and it returned a negative result.** Finding `E-1` raised the
+possibility that the per-record cost was a term every strategy paid equally, and therefore a shared
+constant capable of flattening the three-way comparison into "indistinguishable" without that being
+true. Twenty runs -- ten each side, taken in one sitting by stashing the change so both sets came
+from the same machine and build -- are kept in
+[measurements/2026-09-22-append-batching/](measurements/2026-09-22-append-batching/).
+
+Throughput did not move in a way that can be distinguished from noise: median records/sec shifted by
+1-5% while a single strategy's run-to-run range spans 1.17x to 1.57x. The cross-strategy spread did
+not shrink, and stayed at or below one strategy's own range -- which is the sample's own stated test.
+**So `E-1`'s hypothesis is not supported, and the existing conclusion survives a confound raised
+specifically against it.**
+
+What did move is commit p50, and it is the one figure here that separates: for `covering-flush` the
+ten before-values and ten after-values barely overlap. That is the expected shape rather than a
+surprise -- batching removes seven of every eight submissions from the append path, shortening the
+interval between the last append and the flush being reached. Throughput is bound by the device
+flush and does not move; latency is not, and does.
+
+The figures live in the capture and are **linked** from `strategy.rs` and from `M20.6` rather than
+pasted into either, per the rule that a measurement has one home.
+
+**`M21.3`'s property was preserved deliberately.** Batching changes the append loop, and the naive
+rewrite would have reintroduced a commit trigger reachable on a pass that appended nothing. The loop
+`continue`s when zero records are accepted, so the epoch check is reachable only after progress --
+the same structural guarantee, stated in the same place.
+
+**Unblocks `M20.6`**, whose remaining question is the part no number speaks to: whether alternating
+rings earns its cost on correctness and blast-radius grounds.
+
+### M22.2 -- Collapse the two free-slot implementations to one, derived from the arena's own outstanding counts rather than tracked beside them. The item called both correct; one was not -- the tracked free list leaked a slot on every refused append. *(completed 2026-09-22 16:21:23 -04:00)*
+
+The sample had two answers to "which arena slots are free". `Appender` asked the arena, filtering on
+`outstanding(slot) == Some(0)`. `Lane` kept a `Vec` and maintained it by hand. Both now call one
+`free_slots` in [append.rs](examples/epoch_log/append.rs); `Lane`'s field, its initialiser, and its
+push-on-claim are gone.
+
+**The item's premise was wrong in the direction that mattered.** It said "both are correct and the cost
+difference is nil ... the duplication is the defect, because the two *can* drift". They had already
+drifted. The free list took its slot *before* composing into it, so an append refused between the two --
+a record too long for a slot is the reachable path -- returned with the slot popped and no operation ever
+issued. The arena considered that slot quiet forever; the list never offered it again. `SLOTS` such
+refusals and the harness reports a full arena while the kernel holds nothing.
+
+So this was not a tidying exercise with a correctness footnote: **deriving the fact removed a live bug**,
+and it removed it by construction rather than by fixing the copy -- a slot nothing was pushed against
+never stopped being free.
+
+**Verified by sabotage, in both directions.**
+
+- *Does each caller really bind to the one definition?* Breaking `free_slots` to offer busy slots failed
+ the `Lane` path with the arena's own refusal (`buffer 0 still has 1 operation(s) outstanding`), while
+ the appender ran clean -- so that run proved only half of it. A second sabotage (`take(0)`) starved the
+ appender, which then failed before the strategy section was reached. Both halves bind; one sabotage was
+ not enough to show it, which is the point of running the second.
+- *Was the leak real, or argued from the source?* Re-injecting the free list and failing eight appends
+ left the lane reporting **0** of 8 slots free with the arena entirely idle.
+
+**What the new tests do not catch, stated in the tests.** [strategy/tests.rs](examples/epoch_log/strategy/tests.rs)
+asserts the property from the arena's side, so a re-introduced free list would *not* fail it -- under the
+re-injection above it passed, and only a temporary assertion against the list itself went red. What keeps
+a second definition from returning is that there is one function and both callers call it. Claiming the
+test covers that would be the cosmetic binding the repository's own rules warn about, so the module says
+so plainly instead.
+
+Two harness defects were themselves caught by sabotage discipline and are worth recording, because both
+produce a *false green*: a `.Replace` that matched nothing reported success and ran an unmodified tree,
+and a PowerShell helper that logged with `Write-Output` returned its log line into the patched text. Per-site
+match-count assertions caught the first; the second surfaced as a run with no output at all. The repository
+already requires the match-count check for exactly this reason.
+
+**Raised a layer-placement question rather than answering it silently:** `outstanding()`'s rustdoc names
+this use case, and both in-repo consumers hand-rolled it anyway. Queued as `M22+.2` with its blocker named
+(a public API addition needs a lib test, and that test opens a real ring -- the `M24` pile).
+
+### M22.3 -- Give the registered arena a stated placement: the epoch-log arena is placed on the NUMA node its own log file's volume reports, and the allocator moved into the library as `NumaBuffer` rather than being copied a second time. The sample says plainly that the placement cannot pay at this workload. *(completed 2026-09-22 17:41:01 -04:00)*
+
+The item offered a choice -- adopt `ring_copy`'s allocator, or write down why a durability sample makes no
+locality decision. Both halves turned out to be needed, and a third thing fell out of doing them.
+
+**The decision.** [placement.rs](examples/epoch_log/placement.rs) asks the log's own handle through
+`FSCTL_QUERY_VOLUME_NUMA_INFO` -- the documented call [What is not reachable](DESIGN-NOTES.md) already
+established, which takes a file or directory handle directly and needs no device-tree walk -- and the arena
+is allocated preferring whatever comes back. A volume that names no node yields no preference and the log
+runs on; the report line says which happened and why.
+
+**No benefit is claimed, and the sample says so in its own output.** The arena is eight slots of four
+kilobytes against a workload bound by a per-epoch device flush costing hundreds of microseconds, which
+`M22.1` measured directly. So this demonstrates how the decision is made and reported, not that it was
+worth making; `examples/ring_copy` remains where placement meets a load that could show it. The report is
+also qualified by `GetNumaHighestNodeNumber`, because on a one-node machine "placed on node 0" is true and
+misleading -- the line instead says the choice was never available. Two unit tests hold that in **both**
+directions: the disclaimer must appear for a single-node machine and must **not** appear for a multi-node
+one, since a disclaimer that shows up everywhere trains a reader to ignore it.
+
+**The allocator moved into the library ([D-51](DESIGN-NOTES.md#d-51)), by the engineer's call on a
+question raised before any code was written.** The crate's front page names this allocation as the
+highest-leverage locality decision available and then supplied nothing, so the first consumer wrote it in a
+sample and this item was about to write the second. `git mv` carried the history; `ring_copy` lost its
+local module and binds to the library type.
+
+**That move surfaced a packaging defect that would have reached a consumer.** `cargo check --all-targets`
+unifies dev-dependency features into the build, so the library's newly-required `Win32_System_Memory` was
+being supplied by the dev-dependency list and the lib target compiled clean. `cargo doc` -- which does not
+get dev-dependencies -- failed immediately, and `cargo check --lib` confirmed it: **anyone depending on
+this crate alone would not have compiled it.** The manifest now carries the feature on the library
+dependency, and the stale comment asserting "the library itself needs none of them" is corrected rather
+than left to mislead.
+
+**Verified by sabotage, in both directions.** Forcing the FSCTL to report a node the machine does not have
+failed the run at arena allocation with `ERROR_INVALID_PARAMETER`, which is what proves the queried node
+actually reaches the allocator rather than being reported decoratively. Forcing the FSCTL to fail produced
+the `Unplaced` line, the stated reason, and a log that still kept its contract -- the path that will not
+otherwise execute on a machine where the query succeeds.
+
+**Swept the claim rather than the one site.** `VirtualAllocExNuma` and the arena's old `Vec` allocation
+were stated in six places; [README.md](README.md), [lib.rs](src/lib.rs), two DESIGN-NOTES sections and the
+manifest comment were updated, and the design-session and archive copies were left alone as historical
+record. `M23.2` was **narrowed** in the same pass: its option (a) is no longer "should a sample do this"
+but the residual library question, because the sample half is now done.
+
+Gate: fmt, clippy, the full suite (14 new `NumaBuffer` tests, 9 new placement tests), `cargo doc` clean,
+lib-only and `--no-default-features` builds, the borrow-surface, encoding and publishable checks, and the
+example end to end.
+
+### M20.3 -- Make `ring_copy`'s degraded-fallback path observable in a test, asserting both that an absent relation degrades and that a present one does not. *(completed 2026-09-22 20:11:18 -04:00)*
+
+The whole-machine fallback in `Policy::select` is the branch every zero-relation machine takes -- the
+shape [D-48](DESIGN-NOTES.md#d-48) records as ordinary rather than exotic -- and it cannot be reached
+by *running* the sample on a machine that reports its relations. A synthetic topology reaches it.
+Fifteen tests in [examples/ring_copy/policy/tests.rs](examples/ring_copy/policy/tests.rs).
+
+**The item's reason for demanding both halves was verified rather than trusted.** It argued that a
+test of the absent case alone "would pass against a function that always degrades". Sabotaging
+`select` to degrade unconditionally showed exactly that: `a_policy_whose_relation_is_absent_...`
+**still passed**, while four present-case tests failed. The reverse sabotage -- never degrade -- failed
+five absent-case tests. Neither half is redundant, and that is now a measured statement.
+
+One test, `degrading_unconditionally_would_fail_a_test_here`, exists to put that dependency in code
+rather than in a comment, so a future edit that deletes the present-case coverage has something named
+to delete.
+
+**Done without waiting on `SH-4.12`, and the coupling was narrowed rather than ignored.** The recorded
+callout said that item "rewrites the selection arm this test would assert against". That is true only
+of a test asserting through `ByL3`. The fallback tail is shared by all five policies and is not what
+`SH-4.12` changes -- it changes which domains `ByL3` matches -- so exercising it through `ByNode` and
+`ByPackage` pins nothing. Both checklists now say so, and `SH-4.12` inherited the one assertion that is
+genuinely its own: `ByL3`'s degradation condition, under whichever rule replaces the `level: 3` match.
+
+Beyond the two halves, the cases cover what the fallback must get right and what it must not claim: a
+memory domain with no processors is not a usable node; degradation is per-policy rather than a property
+of the machine; `Single` returns the whole machine **undegraded**, because degrading is a statement
+about not getting what was asked for and `Single` asked for exactly this; the fallback covers every
+online processor, excludes reserved-but-offline slots, and spans processor groups; and it carries no
+observations, because nothing observed it.
+
+The synthetic memory domain uses `Observed::NotObserved` for its size rather than `Known(0)`, which the
+type's own documentation calls the variant "a hand-written description leaves behind". `Known(0)` would
+have asserted a measurement nobody made -- in a test whose subject is honest reporting.
+
+`ring_copy` was auto-discovered and therefore **not a test target**, so `cargo test` would have compiled
+these and run nothing. It now has an explicit `[[example]]` entry with `test = true`, the same reason
+`epoch_log` has one.
+
+### M20.1 -- Restate the cache heuristic as "the outermost cache level that actually partitions the machine", sweep every restatement, and replace the consumer that bound to the level number. *(completed 2026-09-22 20:27:15 -04:00)*
+
+**Done together with `SH-4.12`, because they are one change.** That coupling was real, unlike `M20.3`'s:
+this item's sweep reaches `policy.rs`'s doc comments, and rewriting those to describe the new rule while
+the code still filtered `level: 3` is exactly the contradiction the blast-radius convention exists to
+prevent. Splitting them would have produced a commit whose documentation lied.
+
+**The item's evidence was a shipping ARM part with no L3. Measuring the consumer found a second shape,
+on this workspace's own development machine, that nobody had anticipated.** It reports an L3 spanning
+**all 16 processors** above a real **8-way L2** partition. So the old filter did not fail the way the
+item assumed:
+
+| | domains selected | reported degraded? |
+|---|---|---|
+| old `level: 3` filter | **1**, mask `0xffff` | **no** |
+| `outermost_partitioning_cache()` | **8**, at L2 | no |
+
+The old code *matched something*, so it did not degrade -- it reported success while collapsing an
+eight-domain machine to a single ring. A silent wrong answer, not a visible fallback, and live on the
+machine this repository is developed on rather than on hardware nobody here owns.
+
+Three consumers restated the rule, not the two the items named. `ring_copy`'s `policy.rs` and the prose
+were known; `examples/l3_domains.rs` also hardcoded `cache.level == 3` and was named after the
+assumption. It is now `examples/cache_domains.rs`, asks the same primitive, and reports the level it
+found -- `git mv` kept its history.
+
+**`byl3` and `l3` are rejected rather than aliased.** They named a rule the sample no longer implements;
+mapping them onto `ByCache` would let a script keep asking for L3 and keep believing it got L3, which on
+an L2-partitioned machine is a wrong answer delivered quietly. An unknown policy prints the usage line,
+which is a question rather than a wrong answer.
+
+Five tests were added to the file `M20.3` created two hours earlier -- the assertion `SH-4.12` had
+inherited. Re-injecting the `level: 3` filter fails three of them, including the one that pins the
+measured shape above. The swept sites: `DESIGN-NOTES.md` (the heuristic section, the sizing note, the
+policy list, D-27's pointer, and D-48's own "restating the rule is M20.1" reference, which was itself a
+restatement that would have gone stale), `README.md`, `src/lib.rs`, both examples, and both checklists.
+
+### M24.1 -- Settle whether a co-tested fake escapes the mock objection. Answered: the fake was the wrong instrument. *(completed 2026-09-22 21:18:30 -04:00)*
+
+The item predicted two outcomes and both held -- a wrong **accounting** model was caught by the shared
+suite, a wrong **Windows belief** slipped through. Two further cases changed the answer.
+
+**Case 3, which the item did not predict, is the argument *for* co-testing.** Run an assertion written
+from the *wrong* belief against both peers: the kernel goes red and refutes us, the fake goes green and
+confirms us. That is the manufactured-evidence mechanism made visible, and also the escape -- a
+mock-only world never performs that experiment.
+
+**Case 4 invalidated the line the first three cases suggested.** Raised in review: kernel behaviour is an
+observation at a point in time, not objective truth, and we must not over-index on a record of how it
+runs. Demonstrated: "after submitting, the completion is already queued" reads like a contract and gave
+**opposite answers on two handles of the same API**. So the axis is not "accounting versus Windows
+behaviour" -- it is **our specified contract versus the platform's incidental behaviour**.
+
+**Case 5 replaced the technique.** Also from review: model what the platform is *permitted* to do and
+let a seed pick a resolution, so the assertions are about this crate rather than about the kernel. That
+dissolves the mock objection instead of working around it, because there is no belief to be wrong about.
+A minimal resolver broke a FIFO-assuming consumer under 189 of 200 seeds -- and passed under the other
+11, which is the point: a fixed fake reports whichever single answer it encoded.
+
+**Two apparatus failures in one session, both caught, both the same shape.** The first spike draft ran
+each condition once and printed a verdict -- exactly `D-47`'s error; rewritten to 500 trials it
+immediately found a condition that pends about 1% of the time and would have been called "never". The
+case-4 harness used a bare flush, which completes inline on every handle, so it could not discriminate
+until it was rebuilt around sector-aligned writes. Both are recorded because the repository's rule is
+that an instrument nobody has shown can go red is not evidence.
+
+Outcome: the rejection **stands** with its scope sharpened ([D-52](DESIGN-NOTES.md#d-52)), `M24.4` is
+**withdrawn**, `M24` becomes unconditional, and the technique that actually answers the question is
+`M26`. The apparatus is kept as
+[kernel-response-space-probe.rs](design-sessions/kernel-response-space-probe.rs); the reasoning is
+[DESIGN-SESSION-2026-09-22-kernel-response-space.md](design-sessions/DESIGN-SESSION-2026-09-22-kernel-response-space.md).
+
+### M24.2 -- Extract the handle-free accounting into its own type, composed by `IoRing`. *(completed 2026-09-22 21:29:07 -04:00)*
+
+The item named five fields as handle-free and five as carrying kernel state. **Checked before acting,
+and it was exactly right** -- `ring_id`, `next_user_data`, `outstanding`, `registered_files` and
+`registered_buffers` against `handle`, `version`, `supported_ops`, `registered_buffer_infos` and
+`completion_event`. Nine methods touch only the first five; they moved with the fields, and `IoRing`
+delegates. `RingId` moved too, since the identity counter is part of the ledger rather than of the
+handle. The public surface did not move.
+
+**19 hermetic tests, and they are the first in `src/` for which [D-49](DESIGN-NOTES.md#d-49)'s
+complaint does not apply.** They open no ring, because every rule they check is this crate's own
+specification: an identity is never reused, a refused reservation costs nothing, the counters saturate
+rather than wrap, the two registration indices are independent, and two ledgers never share an
+identity. The last of those used to need two live kernel objects.
+
+**The tests found an off-by-one in their own author's assumptions.** Two of them asserted that the
+last identity handed out is `usize::MAX`, and failed: `checked_add` runs *before* the value is
+returned, so a reservation made at `usize::MAX` fails rather than handing it out, and the identity
+space is `0..=usize::MAX - 1`. Not a defect -- one value out of 2^64, and "fails rather than wraps" is
+the property that matters -- but invisible from the source, so
+`the_last_identity_is_max_minus_one_not_max` records it rather than leaving the next reader to make
+the same wrong assumption.
+
+**Sabotage, four ways, all caught:** a no-op `record_completion` (2 red), a recycled identity (6 red),
+`wrapping_sub` in place of `saturating_sub` (1 red), and the buffer-registration count advancing the
+file counter (4 red). The recycled-identity sabotage first produced *no output at all* rather than a
+red suite -- an ambiguous-integer compile error -- which is the silent-failure shape the repository's
+rules warn about, and was rerun with an explicit type before being believed.
+
+**The payoff was deliberately not taken here.** The item's motivation said most of the ring-opening
+lib tests "become hermetic *in place*". That is 71 tests across four files, which is a different item
+wearing this one's name; it is queued as `M24.7` with the measured per-file census, and with a warning
+that the census must be recounted because the first attempt at it produced false positives by matching
+`to_string()`.
+
+### M24.7 -- Convert the lib tests that construct a ring only to exercise bookkeeping. *(completed 2026-09-22 21:43:29 -04:00)*
+
+**The recount the item demanded was right to demand.** Its figure of 71 was a per-*file*
+`IoRing::new` count. A per-*test* census gives **61**, and the difference is not rounding -- a file
+with 40 tests and 35 constructions has five tests that never touch a ring.
+
+**61 -> 52.** Two changes, one structural and one an excision.
+
+**`Token::new` now takes the ring's ledger rather than the ring.** It only ever used
+`reserve_user_data()` and `ring_id()`, both bookkeeping, so the wide parameter was the only reason
+`token`'s tests opened a ring at all -- **all seven of them, to mint a token and nothing else.** The
+ring was actively a liability there: none of those tests ever submitted, so `Drop`'s run-down would
+wait for completions that were never coming, and a `settle` helper existed purely to stop teardown
+hanging. The helper is gone with the hazard it worked around. `token` is now 7 hermetic, 0 opening.
+
+**Two `ring` tests were removed rather than converted**, being duplicates of what `M24.2`'s hermetic
+tests now cover: `reserve_user_data_increments_outstanding_and_never_repeats_an_id` and
+`record_completion_saturates_rather_than_underflowing`. Deleting them loses no delegation coverage --
+`run_down_returns_once_a_recorded_completion_zeroes_the_count` already drives reserve, `outstanding`
+and `record_completion` through `IoRing`, and must keep a ring for its own sake.
+
+**The remaining 52 are not convertible, and the reason is structural rather than effort.** Recorded
+here so the next reader does not re-derive it:
+
+- `event_delivery` (6) needs a real ring and the thread pool. There is nothing to narrow.
+- `ring`'s injected-failure cluster **looks** convertible by name and is not. It uses a real
+ completion on purpose -- one test says so in an assertion message, "the flush really did succeed,
+ or this test proves nothing" -- because `with_injected_failure` *transforms* a real completion, and
+ fabricating one is precisely the unsoundness the seam exists to avoid.
+- `batch` (13) needs a `Batch`, which needs the handle for its `Build*` calls. Narrowing that
+ parameter is not possible the way `Token`'s was; it becomes reachable only under `M26.2`'s FFI
+ seam, which is a far larger change.
+
+So **relocation, not conversion, is the remedy for the rest**, which is `M24.3` -- updated with this
+finding and with a warning to recount its own stale figure of 25.
+
+### M24.3 -- Relocate the lib tests that open a ring but use only public API into `tests/`. *(completed 2026-09-22 22:43:39 -04:00)*
+
+**52 -> 41.** Eleven tests moved into `tests/ring_lifecycle.rs` (new), `bounded_pop.rs`,
+`event_delivery.rs`, `registration.rs` and `submission_lifecycle.rs`. Totals conserved: 162 lib + 68
+integration before, 151 + 79 after.
+
+**The item predicted 25 and that was never achievable.** Its figure predated two milestones of test
+growth, and more importantly it assumed the constraint was *which* tests had been looked at rather
+than what they reach. `M24.7` had already found the structural version of this; relocation hits the
+same wall from the other side.
+
+**A regex census got the classification wrong, and the compiler caught it.** The scan excluded
+`pop_within` from the blocking list because it is a public method on `IoRing` -- but `ring.rs` also
+has a `#[cfg(test)] pub(crate) fn pop_within(ring, what)` free function, and
+`windows_refuses_an_empty_buffer_registration` uses *that*. It was moved, failed to compile, and was
+returned. A second miss was structural rather than nominal: the scan only looked at `fn` definitions,
+so it did not see that `a_supplied_wait_is_not_consulted_when_nothing_can_arrive` depends on a
+`RecordingWait` **struct** shared with eight other call sites. That one was caught before moving, by
+a second pass that looked for local `struct`/`const` definitions too.
+
+The method that worked was **moving the candidates and letting the compiler rule**, rather than
+trusting the census. Two milestones running, a crude census has produced false classifications here;
+the compiler produced none.
+
+`HugeBuffer` and `NULL_FILE` travelled with the two tests that used them, having no other call sites.
+
+**`M24.5` needed re-planning as a result, and that is the more consequential outcome.** It assumed
+that after `M24.2` and `M24.3` the lib tests would "construct no ring at all", so a zero-check would
+do. Forty-one remain and none is movable, so a zero-check would fail on day one and could only be
+satisfied by deleting real coverage. The item now asks what the rung should actually assert -- a
+ratchet, an allow-list, or nothing until `M26.2` -- rather than presuming the answer.
+
+### M24.5 -- Put the rule on a rung, so it cannot regress. *(completed 2026-09-22 23:44:37 -04:00)*
+
+**The rule the item assumed was false, and that is the decision this item really made**
+([D-53](DESIGN-NOTES.md#d-53)). It expected a zero-check -- "the lib tests should construct no ring at
+all" -- which would have failed on day one and could only ever be satisfied by deleting real coverage.
+Forty-one remain and none is movable without `M26.2`.
+
+So the rung is an **inventory**: [RING-OPENING-LIB-TESTS.txt](RING-OPENING-LIB-TESTS.txt), regenerated
+from source by [check-ring-tests.ps1](../../tools/check-ring-tests.ps1), failing when the two
+disagree. Deliberately the same mechanism as the borrow-surface check, so there is nothing new to
+explain. Per-test rather than per-file, because two thirds of the 41 live in `ring/tests.rs` and a
+file-level allow-list would let exactly that file grow. An inventory rather than a count, because
+add-one-remove-one nets to zero and a bare number is derived data nobody can check by reading.
+
+**The bidirectional verification found a defect in the guard, which is the whole reason the rule
+demands it.** Direction 3 -- a test that reaches a ring only through a helper -- reported the expected
+test *and an innocent one*. The body extraction ended at the next `#[`, so a plain helper defined
+after the last test in a file was swallowed into that test's body, and a helper containing
+`IoRing::new` made the test above it look ring-opening. The body now ends at a column-0 `}`, which
+`cargo fmt` guarantees is a function end. Had only the "must fire" direction been run, the check would
+have shipped with a false positive that fires on innocent changes -- the fastest way to train people
+to ignore it.
+
+Four directions verified after the fix: fires on a direct `IoRing::new`, fires on a ring reached only
+through a helper, stays silent on a new hermetic test, and reports removals as progress needing only
+regeneration. Wired into CI as its own job beside `borrow-surface`; needs no toolchain.
+
+### M24.6 -- Sweep what this milestone makes false. *(completed 2026-09-23 11:49:58 -04:00)*
+
+The item named three sites. **Two were false alarms and the third was false for a different and more
+serious reason than the item gave** -- which is the argument for running the census rather than
+editing the named list.
+
+**Named, and genuinely stale: [D-49](DESIGN-NOTES.md#d-49).** Its "63 of 131" is the figure `M24`
+*started* from; measured after, **41 of 151 open a ring and 110 do not**. It also still said the mock
+rejection "stands until `M24.1` settles it" and that the remedy choice was "gated on one unresolved
+question", both of which `M24.1` closed. Corrected, with pointers to [D-52](DESIGN-NOTES.md#d-52) and
+[D-53](DESIGN-NOTES.md#d-53).
+
+**Named, false alarm: the testing-strategy section.** The item expected it to be stale because it was
+"written when every lib test opened a ring". Reading it, nothing in it turns on hermeticity -- it
+classifies *defect populations* and *techniques*, and `M24` added or removed neither. Its "all five
+techniques" framing is also correctly left alone: the resolver is a sixth *when `M26` builds it*, and
+`M26.6` already owns that edit. Claiming six today would be the opposite error.
+
+**Named, false alarm: `M21.6`'s archive entry.** The item said its "a wait that never enters the
+kernel" clause "stops being the notable exception once the suite is hermetic". Hermeticity does not
+bear on that sentence, and the archive is append-only history describing what was true when written.
+
+**Named, and false -- but not because of `M24`: `F-13`.** Its headline says "Every fixture in this
+crate's tests, examples and samples opens its handle that way [synchronous]". Three do not:
+`flush_barrier.rs`, `handover.rs` and `flush_barrier_stress.rs` open
+`FILE_FLAG_OVERLAPPED | FILE_FLAG_NO_BUFFERING`, dated 2026-08-28, 08-29 and 09-06 -- **weeks before
+F-13 was recorded on 09-21**. So it was false when written, not made false by this milestone.
+
+That matters because the entry's carry-forward escalated from the false half: it says every claim
+about ordering, draining, the completion event and the barrier "was measured against operations that
+may have completed inline" and names D-19, D-23, D-24 and D-47 for re-reading. But D-23, D-24 and
+D-47 were measured by `flush_barrier.rs` -- one of the three overlapped fixtures. The entry even
+hedged correctly ("the drain-ordering spike used `NO_BUFFERING` ... so it is probably fine") and then
+checked only the spike, not the tests sharing its shape. A dated correction was added rather than a
+rewrite, so the record of what was believed survives.
+
+**Unnamed, and found by the sweep: [bounded_pop.rs](tests/bounded_pop.rs) named the wrong gap.** It
+said "every other test of `pop_within` drives the loop with a wait that never enters the kernel".
+Several do enter it through `SubmitWait`. The real gap is narrower and more interesting: the tests
+using the kernel wait drive operations that *complete*, and the one test that lets a bound expire
+fakes both halves -- a `RecordingWait` instead of the kernel and a bare `reserve_user_data` instead of
+a pending operation. **No test had a real operation pending when a real bound expired**, which is
+exactly the state the `ERROR_TIMEOUT` path needs. Corrected in place.
+
+**Unnamed, and found by the sweep: this milestone's own header** still carried the 63-of-131 opening
+figure and a "recount before starting" caution that had been acted on.
+
+**The transferable part.** Three of the five corrections were over-generalisations from a single
+observation -- one fixture becoming "every fixture", one wait shape becoming "every other test". Each
+was a census away from being right, and each then had an alarm built on top of it. That is the same
+shape as the spike that ran one trial per condition, in the same crate, two days earlier.
+
+### M20.6 -- Re-evaluate `CommitStrategy::AlternatingRings` and the epoch-log benchmark's conclusion against D-47. *(completed 2026-09-23 12:06:40 -04:00)*
+
+**Question 1 -- does alternating rings still earn its cost? Answered on its stated grounds: no.**
+`S-2` argued that because a covering flush reaches every operation outstanding on its ring, two rings
+bound what a commit's barrier can be dragged into. That is structurally false for this sample and
+needs no measurement: `RegisteredBuffers::get_mut` refuses a busy slot and there are `SLOTS` slots, so
+at most `SLOTS` appends are outstanding on a ring **by construction** -- and each alternating lane
+registers its own arena of the same size. The arena bounds the blast radius, not the ring topology.
+Probing agreed (8 and 8); the argument does not rest on it and holds whatever the platform does about
+pending. The argument survives against genuinely unrelated traffic from another component; this sample
+has none.
+
+**The strategy is deliberately NOT removed.** The item said that if it no longer earns its place, that
+is an API change to a published example -- and the temptation was to make it. What two rings could
+*also* buy is **overlap**, and overlap is precisely what this harness cannot exhibit: its handle is
+synchronous, so a ring operation completes inline during submit and nothing is ever outstanding across
+a submit boundary. Deleting a strategy on the strength of a measurement that could not have shown it
+working would be the same error as the measurements this item exists to correct. `M25.5` answers it on
+a harness where operations genuinely pend.
+
+**Question 2 -- re-read, re-run, or annotate the numbers?** The investigation concluded "none of
+those": the column measures deferral rather than a commit, and prose cannot fix a measurement.
+**That conclusion was right about the measurement and wrong about the output.** Leaving a column
+labelled `commit p50` in a published sample until `M25` lands is shipping a false claim for the sake
+of a purist position on annotation. The column is now `ack lag`, with a caveat line naming the p99 = 0
+blocking measurement, and pointing at `M25`.
+
+**The rustdoc already knew, which is the finding worth keeping.** `Outcome::commit_latencies` already
+said the figure is "not device flush time", that deferral inflates it, and that alternating rings
+"reports the highest latency of the three while matching them on throughput". A previous pass had
+diagnosed the artifact correctly **and only in the rustdoc** -- the printed output never got the same
+treatment, and the M20.6 investigation re-derived from scratch what was already written one file away.
+What the investigation genuinely added is the extent: blocking is not merely a component of the figure,
+it is **0 us at p99**, so the number is entirely deferral; and underneath that, no pipeline exists at
+all.
+
+**Swept the mechanism, not just the label.** The explanation "the strategies differ about how long the
+flush itself waits and the extra host round trip, and those land in the tens" appears in the module
+docs, in a `main.rs` code comment, and in the printed summary line. It is wrong in the same way at all
+three: those differences cannot occur on a synchronous handle. The three are indistinguishable because
+they do the same serialized work -- both readings give the same ranking and only one is true. All three
+corrected, plus the "a real log keeps appending while a commit is outstanding" claim, which describes
+a state this program has never reached.
+
+### M20.6 -- correction, same day *(recorded 2026-09-23 13:03:35 -04:00)*
+
+**The entry above declared `AlternatingRings`' blast-radius justification "dead on structural grounds".
+That over-reached, and the over-reach is the kind this repository now has a rule against** -- see
+OPTION INTEGRITY in the repository instructions, added by this correction.
+
+What the structural argument actually establishes is that **this harness** cannot exhibit a
+blast-radius difference, because each lane registers its own arena of `SLOTS` slots and the arena is
+the limiter rather than the ring topology. That is a statement about the apparatus. Generalising it to
+"the justification is dead" converted a fact about one sample's configuration into a verdict on a
+design option.
+
+**It also contradicted the crate's own recorded position.** [D-27](DESIGN-NOTES.md#d-27) is this
+crate's decision that one ring per thread is userspace's proxy for one ring per CPU, and records the
+hardware reason: NVMe queue pairs are per-CPU with each pair's completion interrupt routed by its own
+vector. Two rings on two pinned threads *is* that architecture. Declaring a multi-ring strategy's
+justification dead on the strength of one sample's arena sizing sits directly against a decision the
+crate already made on stronger grounds.
+
+The conditions under which alternating rings would pay are now written at
+`CommitStrategy::AlternatingRings`, and they are ordinary rather than exotic: a ring shared with any
+other component, arenas sized asymmetrically from the lanes, real overlap (where the same covered
+count is not the same wait), and per-CPU queue affinity. The sample's job is restated as giving a
+consumer the means to answer this on their own hardware, not handing them a verdict from ours.
+
+Nothing about the measurement corrections in the entry above changes: the ack-lag relabel, the p99 = 0
+blocking finding, and the swept mechanism claim all stand. What changed is the conclusion drawn from
+them.
+
+### M23.1 -- State in the epoch-log contract that the barrier is ring-wide while the flush names a file, so one ring per log is a precondition of the cost model. *(completed 2026-09-23 17:03:35 -04:00)*
+
+*Queued from finding `S-1` of the 2026-09-19 epoch-log review.*
+
+**The item was right that the contract was silent, and wrong about what it was silent on.**
+`S-1` reasoned from [D-47](DESIGN-NOTES.md#d-47)'s surviving half -- the barrier reaches every
+operation outstanding on the ring, not only the current submission batch -- and concluded that
+"one ring per log" is a precondition of the sample's *durability contract*.
+[contract.rs](examples/epoch_log/contract.rs), written before the code precisely so it would state
+preconditions, said nothing about it.
+
+**Writing it found two scopes conflated, and the first draft shipped the conflation.**
+`IOSQE_FLAGS_DRAIN_PRECEDING_OPS` is a flag on the **ring**; `BuildIoRingFlushFile` names a
+**file**. So the barrier bounds what a commit *waits for* and the flush bounds what it *makes
+durable*, and completion is not durability -- a claim the same file already made three paragraphs
+earlier, about a record's own write. The corrected reading: a shared ring does **not** endanger the
+guarantee, which the flush's own file target secures. It endangers the **cost model**, because the
+barrier waits for unrelated traffic unconditionally.
+
+**Three commits, because the first two were not right.** `d845bb28` added the section, an
+assumption, a non-guarantee, `Clause::ALL` -- replacing a hand-written variant list in `main.rs`
+that would have printed one section short had a fourth clause ever been added -- five tests over the
+report's own properties, and the crate's first [sabotage.json](sabotage.json): five injected defects
+caught, plus a control that rewords a statement and survives, so the guards are sensitive to the
+report degrading without being bound to the contract's wording. `4adea675` corrected the
+conflation. `48dc99d3` right-sized what the correction had grown into -- a title giving the device
+equal billing with the ring, plus a bullet and a milestone pointer about multi-device reach, in a
+sample that runs one log file on one ring.
+
+**The conflation was caught by a question, not by the gate**, which stayed green across all three:
+every test passed, the sabotage sweep reported all six cases as declared, and the two contradictory
+sentences sat a screen apart in one file. The tests check that the report *prints* correctly and
+deliberately not what it *says*, so nothing built here could have found it.
+
+**Two things this work left elsewhere.** The primitive-level half moved to the library under
+[D-54](DESIGN-NOTES.md#d-54): the barrier/flush scope distinction is a fact about one flush, so it
+belongs on `FlushCoverage` rather than only in a sample's contract. And
+[checkpoint.rs](examples/epoch_log/checkpoint.rs) gained the reciprocal note -- it already took its
+own ring, for an unrelated *delivery* reason (D-21), so the structure was right twice over with only
+one reason written down. Both sites now point at each other, so collapsing the rings cannot look
+harmless from either end.
+
+### M23.2 -- Decide how a caller arrives at a NUMA node: `win-numa-sys` offers declaring and discovering, and refuses the shortcut that does both at once. *(completed 2026-09-23 19:57:34 -04:00)*
+
+*Queued from finding `S-3` of the 2026-09-19 epoch-log review.*
+
+**The item was narrowed three times before it was answered, and the last narrowing moved it out of
+this crate entirely.** `M22.3` settled its sample half by making the allocator library surface.
+[D-54](DESIGN-NOTES.md#d-54) removed its other half -- sharding by backing device needs the concept
+of a set of operations that commit together, which this crate does not have -- and handed that to
+the durability layer, where it was sharpened from device identity to *flush equivalence*. Then
+`win-numa-sys` was created and `NumaBuffer` moved into it, so "this crate" in the item text stopped
+naming the crate that had to answer.
+
+**What remained was answered by building that crate, so this item's deliverable was the recorded
+decision rather than code.** It is
+[N-D-1](../win-numa-sys/DESIGN-NOTES.md#n-d-1): a caller may *declare* a node
+(`NumaBuffer::new`), *discover* one (`volume_numa_node`), or *qualify* what a discovered answer is
+worth (`highest_numa_node`); what is refused is a `NumaBuffer::for_file` that would query and
+allocate in one step.
+
+Four reasons for the refusal, of which the first is the one that generalises: such a call **hides
+the answer**, and on a single-node machine "placed on the node the volume named" and "no preference"
+are the same allocation, so a caller could not tell whether the query found anything. It also fuses
+two failure domains, withholds an answer useful beyond one buffer, and saves exactly one line,
+since `NumaBuffer::new(len, volume_numa_node(h).ok())` already type-checks.
+
+**A fifth argument was dropped rather than kept, and the decision says so.** When this was first
+argued, a `for_file` constructor would have dragged `Win32_System_Ioctl` into a crate that
+otherwise touched only memory. That was true of `windows-ioring-sys` and is not true of
+`win-numa-sys`, where the query already lives. Recording a void argument as void is cheaper than
+having someone re-make it.
+
+**Neither path is speculative.** `examples/epoch_log` discovers from its log file's volume;
+`examples/ring_copy` declares a node it computed from the processor topology. Both were already
+written against this shape before the decision recorded it.
+
+[D-8](DESIGN-NOTES.md#d-8) is intact, which was the item's stated constraint: locality stays the
+consumer's decision, and the crate supplies a fact and an allocator rather than a choice.
+
+### M23.3 -- Decide what this crate offers for holding a token between push and completion: the ring owns the inventory, `IoRing` becomes generic, and the break is accepted. *(completed 2026-09-23 23:09:13 -04:00)*
+
+*Recorded as [D-55](DESIGN-NOTES.md#d-55). Implementation is `M28`. The exploration, including
+everything it falsified, is in
+[DESIGN-SESSION-2026-09-23-pending-inventory.md](design-sessions/DESIGN-SESSION-2026-09-23-pending-inventory.md).*
+
+**The item was decided against a different argument than the one it was written on.** It argued
+from duplication -- nine sites keeping the same map. A census found ~12 sites of which only a
+third keep the map described, so duplication was both mis-counted and the wrong frame. The
+mechanism is that this crate **mandates** the construct and does not provide it:
+`Batch::write` returns a `Token`, `IoRing::pop_within` returns a `Completion`, and nothing
+connects them but caller-supplied storage -- which the crate's own rustdoc instructs callers to
+build, twice. That follows from [D-4](DESIGN-NOTES.md#d-4) splitting ring-side counting from
+caller-side identity, which is right; what was missing is the half it left to prose.
+
+**A working spike was built and is why the decision is informed rather than argued.**
+`Pending` fits a real consumer -- converting `append.rs` removed its `InFlight` struct and
+its hand-driven oracle calls -- and sabotage established that removing its drop guard or cutting
+its oracle wiring is caught. A test driving a failed write through the injection seam turned
+`M22.2`'s ordering defect from undetected into caught by assertion.
+
+**But the spike also showed why offering it beside the ring is not enough.** Nothing forces a
+minted token into it, so the ring-to-inventory drift survives -- and `Pending::checked()` owning
+an oracle made the converted consumer's existing `RingContract` a decoy, passing
+`assert_quiescent()` vacuously with nothing in the suite catching it. A generic `IoRing`
+owning the map removes the class, because the consumer never holds a token to lose.
+
+**The objection that had ruled that out was false, and checking it was what settled the item.**
+Per-ring monomorphisation holds for every real consumer; `tests/generated_sequences.rs` already
+carries eight token types on one ring behind a closed `enum Held` with no runtime type check;
+and [D-4](DESIGN-NOTES.md#d-4) rules type erasure out in as many words. The dismissal had
+contradicted a decision already on the books, in the opposite direction from the one it assumed.
+
+**Two findings the decision rests on that were not in the item.** `RingContract` never prunes --
+one retained entry per operation for the process's life, undocumented, invisible in a sample that
+appends 24 records -- which rules out checking by default and is `M28.2`. And `epoch_log`'s
+commits are tokenless because a flush has no buffer and a *borrowed* `RawHandle` leaves its token
+nothing to guard, which is `M28.5` and may dissolve when `M25.3` changes how the log is opened.
+
+**What it refuses**, unchanged by the shape: batching, ordering and which slot to pick stay caller
+questions. The sharper refusal is that the inventory does not decide whether a caller is checked.
+
+### M23.4 -- A failing test that left registered buffers outstanding aborted the process instead of reporting; the drop guards now stay silent during unwind. *(completed 2026-09-23 23:16:48 -04:00)*
+
+`RegisteredBuffers::drop` refused to free while an operation was outstanding (M5.3, correctly) and
+said so with a bare `debug_assert!(false, ...)`. Nothing checked `std::thread::panicking()`, so a
+test that failed *because* a slot leaked panicked, unwound, dropped the arena, panicked a second
+time inside `Drop`, and aborted -- replacing its own assertion message with
+`STATUS_STACK_BUFFER_OVERRUN`. The detection was never weakened; what the abort destroyed was the
+diagnosis.
+
+**The item named one site and there were three.** It said "the fix is presumably the same one line
+here", and a sweep of every `impl Drop` in the crate found the assert it named plus **two more** in
+`IoRing::drop` -- the rundown failure and the `CloseIoRing` failure. That second impl is the worse
+of the two: a ring is dropped on the way out of almost every failing test in this crate, so an
+unguarded assert there converts a readable failure into a crash in the *common* case rather than a
+rare one. The reported site was a sample of the population, which is what CONTRACT INTEGRITY rule 3
+says to expect.
+
+**The sweep also produced two false positives worth naming**, because the pattern that produced
+them is the obvious one to reach for. A grep for `impl.*Drop for` matched cargo-mutants-style
+comment text (`::drop -> ()`) inside test files, which read as two
+further unguarded sites. Anchoring the pattern at line start reduced eight candidate impls to the
+three real asserts. A loose grep over a crate that documents its own mutants will find its
+documentation.
+
+**`Pending::drop` was already correct** and is what the fix copies -- it returns early when
+`std::thread::panicking()`, which is why the M23.3 spike never exhibited this.
+
+**Verified by sabotage, as the item required, and the verdict alone would not have shown it.** The
+`M22.2 regression` case in [sabotage.json](sabotage.json) was `caught` before the fix and `caught`
+after; what changed is that it ended in `exit 101` -- a clean `FAILED` naming
+`a_failed_write_still_releases_its_arena_slot` and its message -- instead of `exit -1073740791`.
+That case's `why` text, which had documented the abort as expected behaviour, now carries the
+post-fix failure mode and says a regression in *either* direction (no longer caught, or caught but
+crashing) is visible there. The full sweep stayed at 9-of-9 as declared with the `CONTROL` still
+surviving.
+
+**Only one of the guards has a test that depends on it firing, and the sabotage that established
+that also falsified the first draft of this entry.** Suppressing both guards unconditionally
+(`true` in place of `std::thread::panicking()`) turned
+`batch::tests::dropping_a_registration_with_work_outstanding_is_refused` red, which is the check
+that the fix *narrowed* the guard rather than removing it -- that test drops deliberately, not
+during an unwind, so `thread::panicking()` is false and the assert still fires. But
+`ring::tests::dropping_a_ring_actually_runs_its_drop_body` stayed green under the same sabotage.
+It is not a `#[should_panic]` test and it does not reach either assert; this entry had claimed it
+did, on the strength of its name, until the sabotage said otherwise.
+
+**`IoRing::drop`'s two asserts are therefore unreachable from any test on a healthy host**, and
+that is recorded at the definition rather than left to be rediscovered. `run_down` fails only when
+`SubmitIoRing` or `PopIoCompletion` returns a kernel error HRESULT, and `CloseIoRing` fails only
+when the kernel refuses the close; the crate's `fault-injection` seam sits at the
+*completion-result* level (`Completion::with_injected_failure`) and cannot produce either. Reaching
+them needs a seam over the raw HRESULTs, which is `M23.5` -- spawned rather than assumed, per the
+move-or-spawn rule, because "no test can reach it" is a blocker to name and not a reason to check
+the box and move on. The fix still lands there on its merits: it is precisely the `IoRing` case
+that turns a readable failure into a crash most often, since a ring is dropped on the way out of
+almost every failing test in this crate.
+
+> **The paragraph above was wrong about the remedy, and [M23.5](#m235) overturned it the same
+> evening.** Both asserts are reachable, no seam was built, and no production line changed: the
+> kernel rejects a *null* ring handle cleanly, and `ring::tests` is a child module that can put one
+> in the field. What the paragraph got right is the finding that prompted it -- the guards were
+> genuinely uncovered. It is left standing rather than rewritten because the archive is history,
+> and because the error in it is instructive: it priced a seam it never checked was necessary.
+### M23.5 -- Both asserts in `IoRing::drop` are now reached by tests, and the seam the item priced turned out not to be needed. *(completed 2026-09-23 23:20:27 -04:00)*
+
+M23.4 left both asserts in `IoRing::drop` uncovered: suppressing them entirely left every test in
+the crate green. This item asked whether to build a fault-injection seam over the raw HRESULTs the
+ring's Win32 calls return, or to accept the asserts as documented-unreachable. **Neither. The item's
+premise was false**, and one probe falsified it.
+
+**What the probe measured.** `CloseIoRing(null)` and `SubmitIoRing(null, ..)` both return
+`0x80070006` -- `HRESULT_FROM_WIN32(ERROR_INVALID_HANDLE)` -- a clean refusal. `CloseIoRing` on a
+plausible-looking `0xDEAD_0000` raises `STATUS_ACCESS_VIOLATION`. So a ring handle is **a pointer
+the kernel dereferences, not an index into a handle table**, and null is the one bad value that is
+refused rather than followed. That asymmetry is the whole finding, and it is recorded at the
+definition and in both tests, because a later cleanup that "tidies" the null into a non-null
+sentinel converts two passing tests into a process crash.
+
+**Why no seam was needed.** `mod tests` is a *child* of `ring`, so it already sees the private
+`handle` and `accounting` fields -- a child module can see its ancestors' private items. A
+`#[cfg(test)]` constructor, `IoRing::refused_by_the_kernel`, assembles a whole `IoRing` around a
+null handle, and the tests let the real `Drop` body run against it. Whether `run_down` submits at
+all is what selects between the two asserts, since it loops only while something is outstanding.
+Production code was not touched: the blast radius the item worried about was zero, because the
+change is entirely in test-only code.
+
+**The D-49 ring-test gate improved the design, which is what it is for.** The first working version
+opened a real ring, closed it by hand, and put a null in the field -- and
+[tools/check-ring-tests.ps1](../../tools/check-ring-tests.ps1) flagged two `ADDED` entries and
+asked its standing question: *does this test need the kernel, or only a ring-shaped thing?* Only
+the latter. `RingVersion::V1` is a public const and `OpSupport` derives `Default`, so every one of
+`IoRing`'s six fields is constructible without opening anything. Answering the gate rather than
+re-baselining it removed the real ring, the hand-close, two `unsafe` blocks and their safety
+arguments from both tests, and left the ring-opening population unchanged at 41. These tests need
+the kernel only to *refuse* them, and refusing costs no ring.
+
+**Three sabotages recorded in [sabotage.json](sabotage.json), not one.** Suppressing the rundown
+guard leaves the close test green and vice versa, so the two asserts are independent conditions and
+a single case would have declared the pair covered while half of it was not. The third sabotage is
+of the *test* rather than the code: removing the reservation makes the ring fall through to the
+close, and the panic message becomes `CloseIoRing failed: 0x80070006` against an expected substring
+of `IoRing rundown failed before close`. That is what shows `should_panic`'s `expected` string is
+load-bearing in selecting the assert rather than decorative. The full sweep is 12-of-12 as declared
+with the `CONTROL` still surviving.
+
+**The methodological point, which is the same one this crate keeps paying for.** The item was
+written an hour earlier, by me, and it reasoned from the shape of the code to "this needs a seam"
+without ever asking the kernel what it does with a bad handle. It then priced that seam's blast
+radius and proposed accepting a permanent coverage gap as the alternative. Both options were
+answers to a question that a single `eprintln!` dissolved. The repository's standing instruction is
+never to report that something cannot be done on the basis of reading it; that applies to "no test
+can reach this" exactly as it applies to "this will not compile".
+
+**A note on release builds, swept but deliberately not changed.** These are `#[should_panic]` tests
+over `debug_assert!`, so they would fail under `cargo test --release`, where the assert compiles
+out. That is a pre-existing property of the crate --
+`batch::tests::dropping_a_registration_with_work_outstanding_is_refused` has the same shape and is
+ungated -- and no CI job runs tests in release. Matching the existing precedent was preferred over
+introducing a `cfg(debug_assertions)` gate on two of the three, which would have left the crate
+inconsistent with itself. If release-mode testing is ever added, all three need the gate together.
+
+### M25.1 + M25.2 -- Records gained a fixed sector stride with a zeroed block tail, and replay learned to walk it. *(completed 2026-09-23 23:36:23 -04:00)*
+
+**These two items could not land separately, and that is a defect in how they were written rather
+than a discovery about the code.** A strided writer and an unstrided reader do not describe the same
+file. The checklist sequenced M25.1 (writer) before M25.2 (reader), so the tree between them holds a
+log nothing can read. They are recorded here as one entry, citing both IDs, per the checklist rule
+for acknowledged coupling; the alternative -- restructuring them into independent items -- is not
+available, because the writer and the reader of one format are not independent.
+
+**What the coupling actually cost was nearly a silent break.** With M25.1 applied alone, the log was
+unreadable: replay advanced by a record's own length, landed in a zeroed block tail, decoded
+`NeverWritten`, and reported every record after the first as a missing durable record. And **all 21
+of the example's tests passed anyway.** The only thing that caught it was
+`cargo run --example epoch_log`, which asserts and exits 101 -- and no CI job runs the sample. That
+is the more important finding of the two, and it is queued as `M25.1b` rather than left in this
+entry, because a finding recorded only in an archive is a finding nobody is obliged to act on.
+
+**Three end-to-end tests were added to close the specific hole**, each binding one writer to the
+real reader through a real file: `records_land_one_per_stride_and_replay_walks_them_back` over the
+log's own `Appender`, `a_run_lays_its_records_out_one_per_stride` over the harness's `Lane`, and
+`a_reused_slot_does_not_write_the_previous_records_tail` over the zeroing. The harness test exists
+because a sabotage said it had to: reverting `Lane` to a packed layout was **caught by nothing**
+until it was written.
+
+**The zeroing is not observable through replay, and the comment says so rather than inventing a
+failure mode for it.** Replay decodes only at block starts and takes a record's extent from its own
+header, so a stale fragment past a short record's end is never read. A first draft of the comment
+claimed a stale fragment "would be decoded as a record", which is false for exactly that reason.
+What zeroing actually prevents is the log carrying fragments of unrelated records -- a hygiene
+defect in a format whose purpose is reconstructing what happened after a crash -- so the test
+asserts on the file's bytes rather than on a replay outcome.
+
+**Two duplications were collapsed rather than converted twice.** The sample has two writers over one
+on-disk format, and both carried a packed layout *and* a verbatim copy of the comment justifying it.
+Converting each in place would have turned one duplicated decision into one duplicated rule, so the
+stride moved to `record`, beside the format it describes, and the whole composition -- encode, then
+zero the remainder -- became `record::encode_block`, which both writers call. That makes half the
+drift unrepresentable rather than merely tested for.
+
+**`Decoded::total_len` became a derived `extent()` rather than a silenced warning.** Once replay
+advanced by the stride, the field's only consumer was a test, and it was exactly
+`HEADER_LEN + payload.len()` -- a stored copy of a fact `payload` already carried. The dead-code
+warning was the signal; the fix was to delete the copy, not to `allow` it.
+
+**Replay now confines each decode to its own block.** Previously `decode` received the rest of the
+file, so a corrupted `payload_len` was bounded only by the file's length and a record could claim
+bytes belonging to its successors -- caught, but by the checksum happening to fail rather than
+structurally. The stride is what makes a block boundary exist to confine it to.
+
+**The cost is reported, not described.** M25.1 asked for the write amplification to be "a real cost
+to state rather than hide", and the first draft stated it as a ratio in a doc comment -- a
+hand-maintained copy of a number the program can compute. The sample now measures and prints it
+(`layout: N bytes of records in M bytes of file`) and the doc comment points at that line instead of
+restating it. No conclusion is drawn about whether the ratio is acceptable, because that depends
+entirely on a caller's record size.
+
+**A `const` assertion carries the sector rule**, verified load-bearing in both directions: a stride
+of 4000 fails the build with `error[E0080]: evaluation panicked: RECORD_STRIDE must be a whole
+number of sectors`, and 4096 builds clean. That is the build rung rather than a test, which matters
+because the failure it prevents -- `ERROR_INVALID_PARAMETER` from a `NO_BUFFERING` write in M25.3 --
+would otherwise appear only on 4K-native storage, on somebody else's machine.
+
+**Four sabotages recorded, not one.** The writers' offset advances are separate facts at separate
+sites: measured, reverting the harness lane leaves every appender test green and vice versa, so a
+single case would have declared the pair covered while half of it was not. The full sweep is
+16-of-16 as declared with the `CONTROL` still surviving. Note what the appender case does *not*
+establish: packing its offsets makes records overlap inside a block, so two guards fire at once for
+two different reasons -- the narrower evidence that the stride itself is what is caught is the
+reader case, which moves only one number.
+
+**All three replay paths report the same numbers as before the change**, which is the check that the
+layout moved and the contract did not: 24 durable records verified, 3 tail records tolerated, the
+torn tail still stopping at `Truncated`, and the negative control still catching a corrupted byte.
+The torn-tail simulation needed its arithmetic rewritten to keep meaning that -- it trimmed a fixed
+count of bytes off the end of the file, which after striding lands in the final record's zero
+padding and tears nothing at all. It now derives the cut from the last record's own block, which
+also survives M25.3's pre-allocation.
+
+### M25.3 -- The log and every strategy file are pre-allocated and opened `NO_BUFFERING | OVERLAPPED`. *(completed 2026-09-24 12:45:57 -04:00)*
+
+Two steps in `logfile::create_preallocated`, neither interchangeable with the other: write the
+extent with an ordinary handle and drop it, then reopen `OPEN_EXISTING` with both flags. This is
+[the spike](design-sessions/spikes/write-pending-spike.rs)'s condition D, the only one of four it
+measured as behaving differently from a buffered handle.
+
+**What this buys is an opportunity, not a guarantee, and nothing here claims otherwise.** Per the
+standing constraint on M25, Windows specifies nothing about when a ring operation completes relative
+to `SubmitIoRing`. The log is correct either way; what changes is whether a commit is separately
+*measurable*, which is M25.4's problem and M25.5's to read.
+
+**The sweep found four live sites, and the item named two of them.** It predicted two statements in
+[strategy.rs](examples/epoch_log/strategy.rs) reasoning from a synchronous handle. There were also
+two in [main.rs](examples/epoch_log/main.rs) -- one an internal comment, one **printed to the user**
+as part of the strategy comparison's own narrative. All four were corrected the same way: the
+historical finding is preserved in the past tense, since it was true when written and is how M20.6
+reached its conclusion, and what follows is that the cause has been removed *without* asserting the
+consequence. Whether the strategies are now distinguishable is not settled by changing a flag.
+
+**The third site the item named no longer exists, and that is the right outcome rather than a
+miss.** `placement.rs`'s `volume_numa_node` documented that it could not use
+`windows-overlapped-io-sys`'s typed `BlockingEndpoint::ioctl` partly because the handle was
+synchronous. Earlier this session, `947b252b` moved that function into `win-numa-sys`, and the
+comment went with the move -- correctly, because `win-numa-sys` depends only on `windows-sys` and
+has no occasion to explain why it is not using a crate it does not reference. The item was written
+before that move; its prediction of "a narrowed comment, not a refactor" was answered by the
+comment's home changing.
+
+**The `NO_BUFFERING` alignment rule could not be tested the obvious way, and finding that out is
+what produced the better test.** The first attempt wrote through [`std::io::Write`] and failed on
+the *aligned* write: `write_all` issues `WriteFile` with a null `OVERLAPPED`, which an asynchronous
+handle refuses however well-aligned the transfer is. That is now its own test -- it is the only
+property of the handle's *mode* reachable from here, since `GetFileInformationByHandleEx` does not
+report it and the ring works on synchronous and asynchronous handles alike. The alignment rule is
+tested through a real ring instead, which is also how production reaches this handle, with both
+directions asserted: an aligned write accepted and an unaligned one refused. Each test has a control
+using an ordinary handle, so the refusals are attributable to the flags rather than to anything else
+about the file.
+
+**A blind spot is recorded in [sabotage.json](sabotage.json) as a declared survivor rather than left
+invisible.** Replacing the zero-fill with `set_len` is a **real regression that nothing here
+detects**: both produce a file of the right size whose bytes read back as zero, because reads past
+the valid data length are answered with zeros the filesystem synthesises without touching the disk.
+Only the zero-fill advances that valid data length -- which is the thing that decides whether a
+later write is extending. A `set_len` extent silently returns the log to the configuration measured
+as behaving like a buffered handle. The only user-mode way to read a valid-data length back is
+`FSCTL_QUERY_FILE_REGIONS`, and adding it to a sample purely to check a property the sample does not
+otherwise use was judged machinery for its own sake. Recording it as `expect: "survives"` means a
+future change that makes it observable will show up as a discrepancy in the sweep.
+
+> **Corrected 2026-09-24, the same day, after review challenged the claim rather than the code.**
+> The paragraph above asserts from documentation that `set_len` is a regression, and the reasoning
+> it gives is the wrong mechanism. Two measurements settled it:
+> [2026-09-24-set-len-zero-fill-cost/](measurements/2026-09-24-set-len-zero-fill-cost/README.md)
+> shows the zeroing cost is **identical** for a sequential writer, so that is not the reason; and
+> [2026-09-24-set-len-vs-zero-fill/](measurements/2026-09-24-set-len-vs-zero-fill/README.md) shows
+> the zero-filled extent pends at a median of 471/500 against `set_len`'s 268/500, which is. The
+> blind spot is real and the conclusion survives; the argument for it did not. The second capture
+> also corrects this entry's own framing of the spike, which repeated "only the pre-written extent
+> pended" from a single run that does not replicate.
+
+**The harness caught a stale case in its own manifest, which is worth more than the case was.** The
+`M25.1: the appender packs its record offsets` sabotage stopped compiling, because M25.1 ended by
+removing the `total` binding its patch referenced in order to clear an unused-variable warning --
+so a sabotage that was correct when written was broken by a later edit in the same session. The
+harness reported `MANIFEST DOES NOT COMPILE (tests never ran)` rather than scoring it as caught or
+survived. **A manifest is a restatement site like any other**, and has to be swept when the code it
+patches moves; nothing else would have noticed.
+
+**One assertion had to change because pre-allocation made it vacuous.** The strategy harness
+compared `outcome.bytes` against the file's length to check its own accounting. A pre-allocated file
+spans its whole extent from the moment it is created, whatever was written into it, so that equality
+would have held just as well for a run that wrote nothing. It now checks the accounting against the
+layout rule -- one block per record -- and separately that the extent covers what was written.
+
+**Deliberately left buffered, and said so at the definition.** The checkpoint file's records are
+sixteen bytes from a `Vec` at offset 0, which satisfies none of `NO_BUFFERING`'s three alignment
+rules; the retired segment is written once with `std::fs::write` and never goes through a ring. The
+control plane's correctness comes from its covering flush, not from how its bytes are cached.
+
+**The log is pre-allocated with slack rather than to its exact record count.** A real write-ahead log
+pre-allocates ahead of its writer, because an append that reaches the end of the extent becomes an
+extending write again. It also means a clean log now ends in zeros rather than at EOF, so replay
+stops with `NeverWritten` -- the path M25.2 taught it to tolerate, now actually exercised by the
+sample rather than left for a reader of a real log to meet first.
+
+### M25.1b -- The sample's own verification now runs under `cargo test`, and `main` itself under CI. *(completed 2026-09-24 15:20:02 -04:00)*
+
+The item asked where this belonged: a CI job running the binary, or a `#[test]` calling the sample's
+functions. **Both, because they carry different facts**, and the split follows the FAIL FAST ladder
+rather than splitting the difference.
+
+**Everything the sample asserts is now a test.** `run_log` and `verify` are ordinary functions over
+a generic `Report`, so `tests.rs` drives the real log against a real ring and a real file and then
+runs the real verifier -- the same code path `main` takes. That is not a proxy for running the
+sample; it *is* running it, minus the strategy comparison. It costs nothing measurable: the example
+suite still finishes in well under a second.
+
+**The comparison's strongest assertion had no test, and now does.** `compare_strategies` requires
+all three strategies to write **byte-identical** logs -- replay checks a log against itself, where
+this checks the three against each other, so a dropped record or a wrong offset in any one shows up
+as a difference from the other two. `M25.1`'s harness test runs `CoveringFlush` alone, so this was
+the one assertion only `cargo run` could reach. `every_strategy_writes_the_same_log` runs all three
+at two epochs of three records instead of thirty-two of sixty-four. The expectation survives the
+ring count because a record's offset is its position in the global sequence times the stride, and a
+strategy decides which *ring* submits a write, never where it lands.
+
+**CI runs the binary for what a test cannot reach**: `main` itself -- its path setup, its error
+plumbing, its exit code. This is a published example a consumer runs, so a panic on startup is
+exactly the failure worth catching, and it is the rung that fits because nothing smaller executes
+`main`. Release, where the sample takes about a second.
+
+**The sabotage found a real hole, and it was in the thing this item exists to protect.** Reverting
+`M25.2`'s torn-tail cut to the old file-length form was **survived** by the new end-to-end test.
+With the extent pre-allocated, trimming a fixed count of bytes off the file lands in the slack, so
+every record stays whole -- and `is_clean()` and the durable count both pass while the case
+demonstrates the opposite of what it claims. A verifier that had quietly stopped verifying.
+
+What makes that worth recording is where the hazard already was: **written out in full, in the
+comment directly above the cut**, since `M25.2`. Describing it caught nothing. Two assertions now
+carry it -- that `tail_stopped` is `Truncated` specifically, not merely present, since a cut past
+the last record reports `NeverWritten` and would mean the case had stopped tearing; and that a tail
+record was actually lost. Prose is not a rung, stated once more by a file that had the prose.
+
+**What each guard is load-bearing for**, measured:
+
+| sabotage | end-to-end test | M25.1's layout tests |
+|---|---|---|
+| replay walks by extent (the `M25.1` defect) | red | red |
+| torn cut taken from the file length | **red** | green |
+| negative control moved out of the durable region | **red** | green |
+
+So the end-to-end test would have caught the defect that motivated the item, and it is the only
+thing covering two failures that make the sample's evidence vacuous rather than wrong. Both are
+recorded in [sabotage.json](sabotage.json).
+
+**What is deliberately still only in CI**: the strategy comparison end to end. Running it in a test
+would multiply the suite's cost to re-check a layout and a replay that `M25.1` already covers, and
+what the full run adds beyond the invariant above is a *measurement* -- which is not a contract this
+crate may assert.
+
+### M25.4 -- The commit is measured as submit / blocking / deferral, so the flush's own cost and the deferral window cannot be confused again. *(completed 2026-09-24 16:32:56 -04:00)*
+
+`CommitTiming` replaces the single `Duration` the harness used to publish. The three parts sum to
+that old number, and `flush()` is `submit + blocking` -- the deferral excluded, which is the whole
+point.
+
+**The split satisfies M25's standing constraint by construction rather than by assumption.** Windows
+specifies nothing about when a ring operation completes relative to `SubmitIoRing`, so the harness
+must be meaningful either way: if the flush completes inline the device round trip lands in `submit`
+and `blocking` is zero; if it pends, `submit` is short and the wait appears in `blocking`. Nothing
+has to know which case it got.
+
+**The strategies were always distinguishable on the commit, and the blended number hid it
+completely.** They now separate by roughly sixfold on the flush, where the old figure ranked them in
+the opposite order -- alternating-rings reported the *worst* commit latency while being the fastest,
+because its deferral is about twice the others'. That is not a new finding so much as `M20.6`'s
+finding finally visible in the program's own output. Figures are not quoted here; the sample prints
+them and `M25.5` is where they are read.
+
+**`blocking` reads zero for all three, and the output says plainly that this proves nothing.** A
+zero beside a large deferral is ambiguous: the operation may have completed inline, or it may have
+pended and finished while the program was busy elsewhere. Those are indistinguishable from here.
+Recording that is the point -- reading `blocking` alone would be the same error as before in the
+opposite direction, and the temptation is real now that `M25.3` has given the handle the shape the
+spike measured as pending.
+
+**The settle logic became one function rather than two copies.** Both sites -- the loop's and the
+drain after it -- applied the same rule about where deferral ends and blocking begins, and a rule
+stated twice can be half-corrected. `settle` states it once.
+
+**A sabotage found the guard that the obvious assertions miss.** Asserting the identity
+`flush() == submit + blocking` catches deferral being folded back in, and is **survived** by a part
+that is never measured at all: replacing the deferral measurement with zero satisfies every identity
+while making the decomposition a rename. The test therefore also requires some sample of `deferral`
+and of `submit` to be non-zero -- phrased as "some sample" rather than a lower bound on a duration,
+because this harness defers by construction, so a run in which nothing deferred means the clock is
+not running rather than that the machine was fast. `blocking` deliberately gets no such guard, since
+zero is a legitimate and frequently observed reading for it.
+
+Three cases in [sabotage.json](sabotage.json): the fold, and the two parts that can silently read
+zero. They are listed separately because `submit` and `deferral` are measured at different sites and
+one can be lost without the other.
+
+**What this does not do** is assert any value. Which part carries the cost is a property of the
+machine and the handle, not of this crate.
+
+### M25.5 -- The comparison was re-run over fifteen runs and `M20.6` answered: no other ground found for `AlternatingRings`, and the accounting defect that would have inverted the reading was fixed first. *(completed 2026-09-24 17:19:59 -04:00)*
+
+The capture is [measurements/2026-09-24-commit-decomposed/](measurements/2026-09-24-commit-decomposed/README.md)
+and the figures live there rather than here.
+
+**The measurement had a defect that had to be found before it could answer anything.** `M25.4`'s
+first numbers showed `HostSequenced` committing roughly **six times cheaper** than the other two.
+That was where the clock started: it waits for every write in userspace before pushing an unordered
+flush, and the commit clock began at the *submit*, so its host round trip fell outside every
+measured part. The cost had not gone anywhere; nothing was looking at it.
+
+A fourth part, `prepare`, now covers whatever a strategy must do before its flush can be pushed.
+With it, `HostSequenced` is within noise of the others rather than six times cheaper. **A reader of
+the uncorrected figures would have drawn the opposite of the right conclusion** -- which is the same
+failure mode `M20.6` was opened to fix, one layer down, found by reading the very numbers the fix
+for it produced.
+
+**What fifteen runs show.** The three are **not distinguishable** on throughput or on total commit
+cost: medians within a few percent, every range overlapping every other, and run-to-run spread
+within a single strategy larger than the spread across them -- which is the condition the sample's
+own output tells a reader to check. What *is* structural, and never inverts across fifteen runs, is
+where each spends its commit: `HostSequenced` in `prepare` and almost nothing in `submit`, the
+covering strategies the reverse. That difference is invisible in any blended number, which is what
+`M25.4` was for.
+
+**The open question, answered as far as this can answer it.** `AlternatingRings` shows no advantage
+this harness can measure -- its throughput and commit medians sit inside the others' ranges, its
+deferral is consistently about twice theirs, and the single worst commit p99 in the capture is its
+outlier -- against a doubled arena registration it pays for the life of the run.
+
+**That is not a finding against the strategy**, and the entry says so where a reader will meet it.
+The ground it was built on is blast radius, and `M20.6` established that this harness **cannot
+exhibit that difference at all**, because each lane registers its own arena of the same size so the
+per-ring bound is identical by construction. What this capture could answer is whether some *other*
+ground appears, and none did. Whether that changes the strategy's status is the engineer's decision;
+the conditions under which it would pay are already written in
+[strategy.rs](examples/epoch_log/strategy.rs).
+
+**`block` reads zero at the median in every run**, and the capture records that this establishes
+nothing: an operation that pended and finished during a deferral of several milliseconds is
+indistinguishable from one that completed inline.
+
+**The harness caught stale manifest patches for the third time today.** Adding `prepare` changed
+both `flush()` and the deferred tuple, invalidating two `M25.4` cases that had been correct when
+written hours earlier. Each time the code moved, the manifest's patches stopped applying and only
+`run-sabotage.ps1` noticed -- `MANIFEST STALE: pattern found 0 times`. The note that a manifest is a
+restatement site like any other is now load-bearing three times over, which is enough to call it a
+standing hazard rather than an incident.
+
+### M25.6 -- Swept what this milestone made false, and recorded the two findings as `D-56` and `D-57`. *(completed 2026-09-24 18:03:24 -04:00)*
+
+**Four named sites, and the sweep found two more.** The item listed the "keeps appending while a
+commit is outstanding" rationale, `strategy.rs`'s "what the measurement found" section, the `M22.1`
+capture's commit-p50 claim, and any DESIGN-NOTES text calling the sample's I/O buffered. Grepping
+the falsified *claims* rather than the listed files also turned up the spike's premise that
+`epoch_log` "writes variable-length records at packed offsets", and a second copy of the "two orders
+of magnitude" mechanism inside the `M22.1` capture's `Settles` paragraph. As usual the reported
+sites were a sample of the population.
+
+**One named site turned out not to exist.** No DESIGN-NOTES text describes the sample's I/O as
+buffered. The three near-matches are about other things -- `D-40` is a cached *read* in the handover
+tests, `D-49` is the unit suite's hermeticity, and the note that a ring handle "does not need
+`FILE_FLAG_OVERLAPPED`" is a fact about rings that `M25` did not touch. Recorded because a sweep
+that quietly finds nothing at a named site is indistinguishable from one that did not look.
+
+**The `M22.1` correction is the sharpest of them, and it is not that the number was wrong.** That
+capture reported a commit-p50 reduction as "the one finding here that separates", beside a
+throughput result reported as unmoved. The reduction is real and its stated mechanism is correct as
+written -- batching shortened the interval between the last append and the flush being reached. What
+is wrong is the label: `M20.6` established that figure was **entirely deferral**, so it measured the
+*append path* getting faster. And that makes it not an independent finding at all. Both lines are
+the same fact seen twice -- the appends got faster, the run is flush-bound, so the change appears in
+the metric that is not flush-bound and not in the one that is. Reporting them as two results
+overstates the evidence by exactly one result.
+
+**One figure is now measured and is not what it said.** The old explanation had the strategies
+differing by amounts "two orders of magnitude below" the flush, "in the tens" of microseconds.
+`M25.5` measured hundreds. The conclusion is unchanged, because they remain smaller than the
+run-to-run spread -- but "below the noise" and "two orders of magnitude below the flush" are
+different claims and only the first held, so both copies of the stronger one were corrected rather
+than left standing beside a note.
+
+**`strategy.rs`'s top section was restructured rather than annotated.** It had accumulated a true
+current claim, a superseded mechanism, and two correction sections underneath, so a reader met the
+false explanation first and the correction several paragraphs later. It now states what is measured,
+links the capture instead of quoting figures, and keeps both superseded explanations compactly below
+under a heading that says they are superseded -- which is what CONTRACT INTEGRITY asks for and what
+the file was violating.
+
+**The spike's prediction is left in the past tense rather than deleted**, because the reasoning is
+the reusable part: it said that if only the pre-written condition pends, "the harness fix is not a
+flag change -- it is a change to the log's on-disk format." That is exactly what happened, and it is
+now `D-57`.
+
+**Two decisions recorded.** `D-56`: a benchmark that defers its await measures the deferral, and the
+number survived three rounds of correction because every round re-read the conclusion instead of the
+instrument -- with the generalisation that "the conclusion still holds" is not evidence that the
+instrument does. `D-57`: a flag whose requirements reach into the caller's data layout is not a flag
+change, and costing it as one underestimates it by the size of a format migration.
+
+### M25.7 -- Replay keeps its slice for a reason about failure vocabulary, recorded as `D-58`; the second multi-megabyte buffer became a digest. *(completed 2026-09-24 19:07:31 -04:00)*
+
+The item offered two acceptable answers -- stream the verifier, or stay legible -- and required only
+that an 8 MiB `fs::read` not sit unremarked in a teaching sample.
+
+**The decision is to keep `replay(&[u8])`, and the reason is not simplicity.** The walk is strictly
+forward one block at a time and never looks back, so it genuinely has no need of the whole file, and
+a real log is larger than memory -- which makes reading the whole file the wrong reflex to teach at
+exactly the point a reader is learning to verify one. That argument is real and it lost to a
+stronger one.
+
+**`replay` returns an `Outcome`, not a `Result`.** Every way it can end is a statement about the
+log: verified, tolerated, or a `Violation`. A streaming reader introduces a third kind of ending --
+`io::Error` -- into the one component whose entire job is to distinguish *the log broke its promise*
+from *the log kept it*. Those want different responses from a caller, and a signature returning both
+through one channel invites precisely the conflation this file exists to prevent: an unreadable file
+reported as a missing durable record. **So the streaming version is a different interface, not a
+smaller allocation** -- which is what the item suspected, and the suspicion is what turned out to
+decide it.
+
+The cost of declining it is stated where it is paid rather than hidden: 140 KiB at the log's own
+`fs::read`, 8 MiB per strategy at the harness's, each with a comment saying so. `D-58` records the
+decision and what a consumer building a real verifier should want instead -- the streaming shape,
+*with* the two failure kinds kept apart inside it.
+
+**What was reducible without touching that interface was reduced.** The cross-strategy comparison
+held a whole reference log in memory for the length of the comparison, so two multi-megabyte buffers
+were alive at once. It now keeps a 32-bit digest, which halves the peak and loses nothing a reader
+had: the assertion could already only say *that* two logs differed, never where.
+
+**The digest is a weaker check than the byte comparison it replaced**, and the weakening is guarded
+rather than assumed away. Two different logs can in principle share a digest where two byte arrays
+cannot share their bytes, so `record/tests.rs` -- a test module `record.rs` did not previously have
+-- pins that a flipped byte, a dropped record, and a trailing zeroed block each change it. The
+sabotage confirms the separation is real: a digest folding only the length still distinguishes logs
+of different sizes, so the two length-based tests stay green and only the flipped-byte one fails.
+What none of them establish, and the definition says so, is that no two logs collide.
+
+**The two 64 KiB sites were left with a note rather than churned.** `RETIRED_LEN` is exactly 64 KiB
+-- at the threshold this repository treats as the point to ask the question, not past it -- so both
+the fill that writes it and the read that checks it are within the rule. The note says what a reader
+growing that segment should do: the write has the same shape as `logfile`'s zero-fill, and the check
+is a fold that never needs the bytes all at once.
+
+## Moved 2026-09-24 19:17:17 -04:00 -- M25: the epoch-log sample's I/O became a shape where a commit is observable
+
+The milestone's eight items are archived individually above; what follows is the context the
+section carried, kept because it records what M20.6 found and the constraint every item was
+held to.
+
+## M25 -- Make the epoch-log sample's I/O a shape where a commit is observable
+
+Queued by the `M20.6` investigation, which found three things the item did not anticipate.
+
+**The harness measures the wrong quantity.** Decomposing its commit latency into *deferral* (flush
+pushed -> harness next looked) and *blocking* (time actually waiting) gave blocking p50 **and p99 of
+0 us for all three strategies**. The published `commit p50/p99/max` column is entirely deferral: it
+reports how long the next epoch's appends took, not anything about the commit.
+
+**There is no pipeline to measure.** The commit's `SubmitIoRing` took 289-555 us and returned with
+all 9 completions already queued. The handle has no `FILE_FLAG_OVERLAPPED`, so the batch ran inline,
+and the comment in [strategy.rs](examples/epoch_log/strategy.rs) reading "a real log keeps appending
+while a commit is outstanding" describes something that cannot happen there.
+
+**`AlternatingRings`' blast-radius claim is answered structurally, and needs no run.**
+`RegisteredBuffers::get_mut` refuses a slot with an operation outstanding and there are `SLOTS`
+slots, so at most `SLOTS` appends are outstanding on a ring **by construction** -- and each
+alternating lane registers its own arena of the same size. The per-ring bound is identical either
+way. Measured at 8 and 8, but the argument does not rest on the measurement, and it holds whatever
+the platform does about pending.
+
+[write-pending-spike.rs](design-sessions/spikes/write-pending-spike.rs) then established which
+configurations pend at all. `FILE_FLAG_OVERLAPPED` alone changed nothing (0/500). Only
+`NO_BUFFERING` over a **pre-written extent** pended reliably, and its submit p50 fell from ~500 us to
+116 us -- the flush's cost leaving the submit path is what makes a commit separately observable for
+the first time.
+
+> **Corrected 2026-09-24, after `M25.3` landed: the paragraph above overstates what replicates.**
+> Sixteen runs with a fifth condition added are in
+> [measurements/2026-09-24-set-len-vs-zero-fill/](measurements/2026-09-24-set-len-vs-zero-fill/README.md).
+> What holds is that a **buffered** handle essentially never pends while every `NO_BUFFERING` one
+> pends in most runs. What does not hold is "only the pre-written extent pended": the extending
+> condition has a median of 268/500 over those runs. The zero-filled extent is still the best of
+> the five -- median 471/500, floor 121 against 1 -- so `M25.3`'s choice stands, but as a
+> difference of degree rather than of kind. The single-run reading came from a pair of numbers the
+> spike's own header already warned was unstable. `M25.4` and `M25.5` must be read with that
+> variance in mind rather than against the original framing.
+
+**A standing constraint on every item below.** That 500/500 is an observation, not a contract:
+Windows specifies nothing about when a ring operation completes relative to `SubmitIoRing`. So the
+sample may *adopt* this shape -- it is what real write-ahead logs do, and it is the only shape where
+the measurement means anything -- but **nothing here may depend on an operation pending.** Every item
+must leave the log correct if the platform completes inline tomorrow.
+
+- [x] **M25.1** -- Records gained a fixed sector stride with a zeroed block tail, in both writers. -> [completed 2026-09-23](COMPLETED-CHECKLIST.md#m251)
+
+- [x] **M25.2** -- Replay walks by the stride and confines each decode to its own block. Landed with `M25.1`: a strided writer and an unstrided reader cannot coexist. -> [completed 2026-09-23](COMPLETED-CHECKLIST.md#m251)
+
+- [x] **M25.1b** -- The sample's own verification now runs under `cargo test`, and `main` itself under CI. -> [completed 2026-09-24](COMPLETED-CHECKLIST.md#m251b)
+
+- [x] **M25.3** -- The log and every strategy file are pre-allocated and opened `NO_BUFFERING | OVERLAPPED`. -> [completed 2026-09-24](COMPLETED-CHECKLIST.md#m253)
+
+- [x] **M25.4** -- The commit is measured as submit / blocking / deferral, so the flush's own cost and the deferral window cannot be confused again. -> [completed 2026-09-24](COMPLETED-CHECKLIST.md#m254)
+
+- [x] **M25.5** -- The comparison was re-run over fifteen runs and `M20.6` answered: no other ground found for `AlternatingRings`, and the accounting defect that would have inverted the reading was fixed first. -> [completed 2026-09-24](COMPLETED-CHECKLIST.md#m255)
+
+- [x] **M25.6** -- Swept what this milestone made false, and recorded the two findings as `D-56` and `D-57`. -> [completed 2026-09-24](COMPLETED-CHECKLIST.md#m256)
+
+- [x] **M25.7** -- Replay keeps its slice for a reason about failure vocabulary, recorded as `D-58`; the second multi-megabyte buffer became a digest. -> [completed 2026-09-24](COMPLETED-CHECKLIST.md#m257)
+
+### M26.1 -- The permitted space is specified in [RESPONSE-SPACE.md](RESPONSE-SPACE.md) as eleven cited clauses, and recorded as `D-59`. *(completed 2026-09-24 19:30:19 -04:00)*
+
+**Eleven clauses: seven permissions and four constraints**, each with an ID, a source, and a
+provenance tag. The IDs exist so `M26.3`'s resolver, `M26.4`'s properties and `M26.6`'s kernel tests
+can cite a clause rather than restate it -- and so a clause no code cites is visible as
+unimplemented.
+
+**The provenance tag is what makes it a specification rather than a recording.** Every clause is
+`Observed`, `Over-provision`, or `Decided`, and where a clause is wider than its own observation the
+two parts are split so they can be argued separately. `RS-P-1` is the clearest case: that an
+operation may complete inline or pend is measured, but that operations within one batch resolve
+**independently** is not, and the space permits it anyway -- because a consumer depending on them
+resolving together depends on something Windows never promised.
+
+**The call the item demanded, made rather than defaulted: `RS-C-4` constrains the resolver to
+honour the drain half of `DRAIN_PRECEDING_OPS`.** [D-47](DESIGN-NOTES.md#d-47) measured roughly
+4,500 trials without a single violation; the drain is what this crate's durability story rests on;
+and a resolver permitted to break it would require every consumer to re-verify durability some other
+way, which is to say it would make the primitive useless. The cost is stated plainly: a Windows that
+broke the drain would not be caught by the resolver at all. That is why `M26.6` gained a line
+requiring at least one kernel test to exercise the clause -- otherwise the one constraint the space
+takes on faith is untested in both halves at once.
+
+**The hold-back half stays unconstrained**, since `D-24` claimed it and `D-47` withdrew it. `RS-P-2`
+applies in full to anything queued after a drained flush, which is the defect class that campaign
+found.
+
+**Three citations were checked and one was wrong.** The draft attributed `M22.2`'s defect to the
+checkpoint control plane; it was on the *append* path, and the checkpoint module merely documents
+the same case. Both now appear, distinguished. The other two -- `pop_within`'s "promises nothing
+about poppability" and `D-47`'s trial count -- were verified against the files rather than recalled.
+
+**A working artifact was found rather than assumed missing.**
+[kernel-response-space-probe.rs](design-sessions/kernel-response-space-probe.rs) already contains a
+seeded `Resolver` exercising `RS-P-2`, which broke a FIFO-assuming consumer under 189 of 200 seeds.
+`M26.3` now points at it as a starting point, with the note that the probe marks itself throwaway --
+so promoting it is a deliberate decision rather than a default.
+
+**Four things are listed as deliberately undecided** -- rates, partial transfers, failure-code sets,
+and timing -- so that a later reader can tell an omission from a choice. Rates in particular are
+excluded on principle: a space carrying observed probabilities would be the recording this milestone
+exists to avoid.
+
+## Moved 2026-09-24 20:36:46 -04:00 -- M26.2: the kernel-call seam
+
+### M26.2 -- Build the seam that makes the `windows-sys` calls indirect, so `M26.3`'s resolver can answer them. *(completed 2026-09-24 20:36:46 -04:00)*
+
+The shape is recorded as [D-60](DESIGN-NOTES.md#d-60); what follows is what the work found.
+
+Eight calls became indirect -- `SubmitIoRing`, `PopIoRingCompletion`, and the six `Build*` entry
+points the crate uses -- across 27 call sites in [batch.rs](src/batch.rs) and [ring.rs](src/ring.rs).
+Each now goes through a `pub(crate) unsafe fn` in [sys.rs](src/sys.rs) that is `#[inline(always)]`
+and dispatches through a `through_seam!` macro. With the `kernel-seam` feature off, the macro
+expands to the bare FFI call and nothing else; with it on, the call first asks the installed
+responder.
+
+**The shape was decided by a constraint already on the books, not by taste.** The obvious
+alternative -- parameterising the ring as `IoRing` over a kernel -- is unavailable because
+[D-55](DESIGN-NOTES.md#d-55) has already spent `IoRing`'s type parameter on `M28.3`'s token
+inventory. A kernel generic would publish `IoRing`, which is a two-parameter public type on a
+shipped crate, and the second parameter exists only so the crate can test itself. Module
+indirection costs the public surface nothing.
+
+**The responder is thread-local, for the reason `DROP_RUNS` is.** `cargo test` runs tests as
+threads in one process ([DESIGN-NOTES.md](DESIGN-NOTES.md) records this as the reason this
+workspace is not on nextest), so a process-global responder would let one test answer another
+test's kernel calls. `with()` uses `try_borrow_mut` rather than `borrow_mut`, so a re-entrant call
+from inside a responder falls through to the kernel instead of panicking; `Installed::drop` uses
+`try_with`, so teardown during TLS destruction cannot abort the process ([M23.4](#m234)).
+
+**Five lifecycle calls were deliberately left direct** -- `CreateIoRing`, `CloseIoRing`,
+`GetIoRingInfo`, `IsIoRingOpSupported`, `SetIoRingCompletionEvent`. `M26` is justified by the
+*response space*: what the kernel may answer to submitted work. Routing ring construction and
+teardown through the seam as well would be hermeticity for its own sake, and hermeticity is
+[M24](#m242)'s subject, not this one. The line is recorded so a later reader can tell a boundary
+from an oversight.
+
+**The seam's transparency is measured, not argued.** The full suite passes with the feature off
+(153 lib tests) and on (159 -- the six new ones), and the crate builds clean with zero warnings in
+five configurations: default, `--all-features`, `--no-default-features`, `--features kernel-seam`,
+and release. The sabotage case *`M26.2: the seam consults the responder but ignores its answer`*
+turns `an_installed_responder_answers_instead_of_the_kernel` red, which is what shows the
+consultation is load-bearing rather than decorative -- a seam that asks and discards would pass
+every other test in the crate. The full sweep is 28-of-28 as declared, with the `CONTROL` and the
+`set_len` blind spot both still surviving.
+
+**Two type signatures were wrong on the first attempt and the compiler caught both**, which is
+worth recording because they are the kind of thing a hand-written trait gets wrong silently if it
+is ever allowed to diverge: `BuildIoRingWriteFile`'s caching flag is `i32`, not `u32`, and
+`BuildIoRingRegisterFileHandles` takes `*const *mut c_void`, not `*const isize`. The `real`
+submodule re-exports the `windows-sys` items so the trait's default methods call them directly,
+which is what keeps the two in step -- there is one spelling of each signature, not two.
+
+**`kernel-seam` crossed with `--no-default-features` is a published configuration nothing built.**
+The gate multiplies with `threadpool`: the workspace `--all-features` steps build the seam only
+alongside the threadpool, and the existing `ioring-no-threadpool` job built the no-threadpool path
+only with the seam off. Two steps were added to that job rather than a new job, on the same
+argument that bought the job in the first place. Verified locally before committing: clippy clean
+and 155 lib tests green in that combination.
+
+**A borrow-surface row was owed and was three items late.** `./tools/check-borrow-surface.ps1`
+failed on `Pending::contract -> Option<&RingContract>`, added by `M23.3`'s spike, because that
+change did not run the gate -- so it reported on the next run instead, against unrelated work. The
+row is now in [DESIGN-NOTES.md](DESIGN-NOTES.md)'s audit table: `RingContract` is a pure
+observation record owning no handle, buffer, or registration index, so there is nothing the kernel
+could invalidate, and the borrow is a plain `&self` borrow that blocks submission through that
+`Pending` for its duration. The lateness is recorded in the row itself.
+
+**A tooling mistake destroyed two source files and is worth the warning.** A PowerShell
+`.Replace()` bound the wrong overload and rewrote [batch.rs](src/batch.rs) and [ring.rs](src/ring.rs)
+one character per line. `git checkout --` recovered both, and the conversion was redone with
+`[regex]::Replace` anchored on `(?M26.3 -- Build the resolver over the space `M26.1` specifies, bound to its clause IDs and seeded on its own axis. *(completed 2026-09-24 21:15:16 -04:00)*
+
+The shape is recorded as [D-61](DESIGN-NOTES.md#d-61); what follows is what the work found.
+
+**The resolver is in [resolver.rs](src/sys/resolver.rs), implementing `M26.2`'s `Responses`.** It
+answers the eight submission-path calls itself, so operations the kernel never received still flow
+through this crate's ordinary accounting. Every freedom cites the `RS-P-n` permitting it and every
+restriction cites the `RS-C-n` requiring it, which is what lets a reader check the resolver against
+[RESPONSE-SPACE.md](RESPONSE-SPACE.md) mechanically rather than by reading both and hoping.
+
+**The asymmetry is the design.** `ResolverConfig` has a switch per permission and none for any
+constraint. Narrowing a freedom is how a test isolates another -- a test about ordering does not
+want arbitrary operation failures on top -- while a knob relaxing a constraint would let a test
+assert against a platform that cannot exist. The default is the widest point, so a test that does
+not choose gets every freedom and fails loudly under one it did not handle. A test asserts the
+count: seven fields, seven `RS-P-n`, and it reads them off `Debug` so a field added without a clause
+fails there rather than passing unnoticed.
+
+**Where a permission and a constraint collide, the constraint wins, and that had to be decided
+rather than discovered.** `RS-C-4` holds a drain-flagged operation back even on a tick where
+`RS-P-1`'s coin said complete it now; `RS-C-1` forces a post an operation's coin kept deferring.
+Two bounds exist solely to make `RS-C-1` finite -- per-operation deferrals, and consecutive declined
+submits -- and both are properties of the resolver rather than of the space, which carries no rates
+deliberately.
+
+**A submit resolves; a pop only rescues.** The split is not tidiness. A pop that flipped coins would
+resolve a consumer's work on its first `try_pop`, so "the operation pended" -- the thing `RS-P-1`
+exists to let a test observe -- would be unobservable to exactly the consumer most likely to care.
+The freedom would have been implemented and untestable.
+
+**`SetIoRingCompletionEvent` moved behind the seam, which `M26.2` had left it outside of.** It looks
+like lifecycle and is not: it is how a completion becomes *observable*, so `RS-P-6` is a clause
+about that call. A resolver unable to make it would not satisfy `RS-P-6` vacuously -- it would never
+signal at all, parking every [`EventDelivery`](src/event_delivery.rs) consumer rather than testing
+one. The checklist item had authorised exactly this ("if a clause turns out to need ... extending
+the seam is part of this item"), and this is the clause that needed it.
+
+**`RS-C-4` is decided by position, and the invariant that makes that sound is now asserted.** The
+pool is held in build order, so the set queued before `pool[i]` is exactly `pool[..i]` and a barrier
+is eligible only as the oldest unresolved operation. That reduction is the whole of the constraint's
+implementation and it holds only while the pool stays sorted -- an edit that sorted or reshuffled it
+would relax `RS-C-4` to nothing while every line around it still read as though it applied. A
+`debug_assert!` now says so at the point of use, rather than a comment saying so nearby.
+
+**The first contact with a real ring found a live defect, queued as `M26.8` rather than fixed
+here.** A submit declined under `RS-P-7` propagates out of `IoRing::run_down` as an error with
+`outstanding() > 0`, after which `Drop` asserts and calls `CloseIoRing` anyway -- `M21.6`'s hazard,
+reachable again through a different `HRESULT`. Measured by narrowing one permission at a time: 32
+seeds pass with `may_fail_submits` off, seed `0x1A` fails at `0x80070008` with it on. It is queued
+rather than corrected because `run_down`'s own documentation argues that blocking is the safe
+failure mode while "no hang" is one of the properties `M26.4` is about to write, and the two pull
+opposite ways -- a decision, not a correction. The narrowing is declared in the test that takes it,
+and the current behaviour is pinned by its own test so that whichever way `M26.8` is settled, a test
+has to change.
+
+**The integration test is where it is for the gate's own reason.** `check-ring-tests.ps1` asks
+whether a test needs the kernel or only a ring-shaped thing; a resolver test needs a real ring
+because [D-60](DESIGN-NOTES.md#d-60) deliberately left lifecycle real, which makes it an
+operating-system boundary and therefore `tests/`. The ring-opening lib population is unchanged at
+41.
+
+**Sabotage found a defect in the tests, which is what it is for.** `RS-C-3`'s case survived: the
+test asserted that nothing pops *right now*, which a resolver that had wrongly made staged
+operations eligible also satisfies, because the starvation rescue posts only what has run out of
+deferrals and a fresh operation has not. The test now polls past the bound and asserts the counters,
+so it checks the claim -- unsubmitted work is never eligible -- rather than the symptom. Eight cases
+added, one of them a declared blind spot; the full sweep is 36-of-36 as declared with both blind
+spots and the `CONTROL` surviving.
+
+**The harness caught stale patches a sixth time, and the cause was new.** Four cases reported
+`pattern found 0 times` while every line of each pattern was present individually. The cause is that
+the built-in file-creation tool writes **CRLF** on Windows, so the multi-line patterns could not
+match a file whose line endings were not LF -- and single-line patterns matched fine, which is why
+three of the eight cases passed and hid it. The repository's own instructions warn about this tool;
+`M26.2`'s files escaped it only because git normalised them on commit before that sweep ran. The
+three new files were converted to LF before proceeding.
+
+## Moved 2026-09-24 21:48:36 -04:00 -- M26.4: the properties that must hold under every resolution
+
+### M26.4 -- Write the properties that must hold under every resolution, with `RingContract` as the definition rather than a second copy. *(completed 2026-09-24 21:48:36 -04:00)*
+
+The shape is recorded as [D-62](DESIGN-NOTES.md#d-62); what follows is what the work found.
+
+**Five properties, in
+[properties_under_every_resolution.rs](tests/properties_under_every_resolution.rs).** Conservation
+(P-1), no hang (P-2), `pop_within` honours its bound (P-3), `outstanding` is accurate (P-4), no
+use-after-free (P-5) -- driven over generated plans, each under its own resolution drawn from
+`M26.3`'s resolver at the widest point in the space.
+
+**Only two of the five needed anything new, and that is the item's main point.** P-1 is already
+this crate's own oracle, so the harness reports to [`RingContract`](src/contract.rs) and asks it for
+the verdict. P-4 follows the same rule rather than counting for itself: the expected outstanding
+count is read back out of the contract through its own `Outstanding` violation, because a counter in
+the harness would be a third party to the disagreement and, when the two disagreed, the harness is
+what would get "fixed". P-5 is [`windows_guard_alloc::GuardAlloc`], already the established
+detector.
+
+**Two weaknesses are declared in the file rather than papered over.** `pop_within`'s upper bound is
+nearly free under an ordinary resolution, because the resolver answers a wait immediately and the
+call rarely approaches its deadline -- so the non-vacuous case needs a resolution in which *nothing
+completes during the window*. That is supplied by a separate degenerate responder, and what it is
+matters: not an `RS-C-1` violation, since no finite observation can distinguish "eventually" from
+"never", but the prefix of a satisfying resolution in which the eventually has not happened yet.
+And P-5 covers this crate's memory handling rather than the kernel's, since under a resolver nothing
+external writes into a buffer at all; the kernel-side half stays with
+[generated_sequences.rs](tests/generated_sequences.rs), against a real ring.
+
+**A third seed axis, kept separate.** This file carries the plan seed, the resolver seed and the
+guard allocator's. They are independent on purpose -- pinning the plan alone reproduces the same
+operations against different resolutions, which is what a suspicious plan calls for -- and a failure
+prints all three, because only all three replay the whole run.
+
+**The harness's own first defect was treating a declined submit as a property failure.** Under
+`RS-P-7` a submit may be declined, and `pop_within` surfaces that as an `Err` -- which is *within
+its documented contract*, since it says it returns any error from `SubmitIoRing`. A correct consumer
+retries; a harness that panicked was asserting a contract the crate never offered. It now retries,
+and recognises a refusal by **asking the resolver** whether its decline counter moved rather than by
+matching an `HRESULT`, because a hard-coded code here would be a second copy of a choice the
+resolver owns. Retrying is bounded by P-2's budget in every caller, so a resolution that declined
+forever is still caught.
+
+**That finding extended `M26.8` rather than creating a second item, and the extension is about the
+document.** `RS-P-7` is written as a *consequence* clause -- "if `SubmitIoRing` fails, operations
+already built remain queued" -- citing `D-5`, which establishes the no-rewind consequence and
+nothing about submits failing spontaneously. `M26.3`'s resolver read it as a permission, and the
+space nowhere states that a submit may fail at all. That is a gap rather than a decision, since
+submits demonstrably can fail, so `M26.8` now has to settle both halves together.
+
+**Coverage counters are asserted, not printed, and that is what makes a green run evidence.** All
+five properties are satisfied trivially by a run that does nothing: an empty plan, a resolution that
+completes everything inside its submit, or a harness that quietly stopped reporting would each pass
+every assertion. The counters are what separate "the properties held" from "nothing reached the
+states they are about", and one sabotage exists purely to show they are load-bearing. Their
+thresholds were set from a measured spread across eight fresh seeds rather than guessed; the
+declined-submit count is the thinnest signal and its threshold is only "more than none" for that
+reason, since tightening it would buy a flake rather than a guarantee.
+
+**The integration test is where it is for the gate's own reason**, as in `M26.3`: it needs a real
+ring, because `D-60` deliberately left lifecycle real. The ring-opening lib population is unchanged
+at 41.
+
+**A sabotage case was written, measured, and removed as unsound.** It disabled the P-1 verdict
+check, and it survived -- correctly, because disabling an assertion that does not fire on a green
+baseline cannot fail anything. What actually establishes that the verdict is read is the
+identity-reuse case: measured, the suite fails carrying the oracle's own wording, so the path from
+the crate through `RingContract` to a red test is traversed end to end. That evidence now lives in
+that case's reasoning, and the removal is recorded here because "sabotage a check" is an appealing
+and empty move worth recognising next time. Three cases added, all scoped to this test target on
+purpose -- these mutations are caught by many tests in the crate, and "something went red" would not
+have shown that *these* properties are the ones watching. Full sweep 39-of-39 as declared.
+
+## Moved 2026-09-24 22:13:02 -04:00 -- M26.5: calibrating the resolver
+
+### M26.5 -- Re-inject the two historical defects and confirm the resolver turns red. *(completed 2026-09-24 22:13:02 -04:00)*
+
+The shape is recorded as [D-63](DESIGN-NOTES.md#d-63); what follows is what the work found.
+
+**The two defects are calibrated differently because they live in different places**, and noticing
+that was most of the item. `M21.6`'s -- an expired wait reported as a failure -- was in this crate,
+so re-injecting it means mutating the crate, which a test cannot do; it is a case in
+[sabotage.json](sabotage.json). [D-47](DESIGN-NOTES.md#d-47)'s -- a consumer believing a covering
+flush holds back what follows -- is in a **consumer**, so there is nothing here to mutate and the
+defective consumer had to be written out. It is, in
+[calibration.rs](tests/calibration.rs), as a `HoldBackBeliever` that records when its assumption
+fails; the test asserts that the resolver breaks it.
+
+**What the calibration file adds on the `M21.6` side is the precondition, not the detection.** A
+sabotage of code the suite never executes is caught for some unrelated reason or not at all, and
+either way measures nothing -- so there is a test asserting that `RS-P-4` actually reaches
+`pop_within`, read off the resolver's own counter rather than inferred from a timing, because "the
+call took a while" is not evidence that a wait expired.
+
+**Both directions were verified by execution before being written into the manifest.** Reverting
+`wait_outcome`'s `WAIT_EXPIRED` arm fails the calibration naming the seed and `0x800705B4`. Making
+the resolver enforce the hold-back fails it with "no seed of 64 broke a consumer that assumes a
+covering flush holds back what follows it".
+
+**The second of those runs the opposite way to every other sabotage in this crate**, and the
+manifest says so: it does not break the code, it makes the **instrument** go narrow. That is the
+`M17.4` failure mode -- a suite sampling the right state while being insensitive to the defect
+living in it -- and nothing detects it except a test that demands the sensitivity.
+
+**The two suites' sensitivities were measured and are not equivalent.** `M21.6`'s defect is caught
+by the calibration *and* by `M26.4`'s property suite; the latter only because that suite
+distinguishes a declined submit from a genuine error, so the detection is deliberate rather than
+lucky. A narrowed resolver is caught by the calibration **alone**, and the property suite correctly
+stays green -- a narrower resolution is still a valid one and conservation holds under it. That
+measured gap is the argument for a calibration file rather than a calibration assertion bolted onto
+the property suite, and it is an argument from data rather than from taste.
+
+**The D-47 calibration breaks its believer on every seed, which needs saying rather than
+celebrating.** `D-47` measured the overtake on real hardware at well under one trial in a hundred;
+the resolver does it constantly. That is not the resolver being unfaithful:
+[RESPONSE-SPACE.md](RESPONSE-SPACE.md) carries no rates deliberately, because a resolver
+reproducing an observed frequency would be a model of Windows and therefore the trap
+[D-52](DESIGN-NOTES.md#d-52) was opened to escape. A defect class that is rare on hardware is
+precisely the one a rate-free resolver earns its keep on. The figures are printed by the tests so a
+reader can judge the instrument's strength rather than take "more than none" on trust.
+
+**Seeds here are fixed rather than clock-derived**, unlike the generated suites. A calibration that
+could sometimes fail to demonstrate its own sensitivity would be the exact failure it exists to
+prevent, arriving as a flake.
+
+## Moved 2026-09-24 23:54:45 -04:00 -- M26.6: the kernel tests' new job
+
+### M26.6 -- Point the kernel tests at confirming reality stays inside the declared space. *(completed 2026-09-24 23:54:45 -04:00)*
+
+The shape is recorded as [D-65](DESIGN-NOTES.md#d-65); what follows is what the work found.
+
+**A hole was found exactly where the split is load-bearing.**
+[flush_barrier.rs](tests/flush_barrier.rs) already asserted `RS-C-4` against a real ring -- but the
+assertion sat behind an early return taken whenever its control could not discriminate. On such a
+machine the constraint was untested on **both** sides at once: the resolver is forbidden to produce
+a violation, and the only test that would notice had skipped. That is precisely the condition the
+item was written to prevent, and it was live.
+
+**The fix separates two questions one gate had been answering together.** *Is the barrier doing
+work* is a comparative claim and genuinely meaningless when the control shows no reordering -- the
+covering case would match a control that did nothing. *Did the kernel stay inside `RS-C-4`* is a
+conformance question, where skipping can only ever hide a violation; observing none is weak
+evidence when nothing could have reordered, but observing one is a finding on any machine. The
+conformance assertion now runs everywhere and only the comparative claim is withheld, with the
+output saying `PARTIAL` rather than `SKIP` so the difference is visible in a log.
+
+**The census was green and useless on its first build, and only trying to make it go red found
+that.** It searched each file for the clause ID anywhere in its text. Removing `RS-C-4`'s check
+from the only test performing it did **not** turn it red, because
+[generated_sequences.rs](tests/generated_sequences.rs) mentioned that clause solely to *disclaim*
+it -- "`RS-C-4` is flush_barrier's" -- and under a substring search a disclaimer reads exactly like
+a claim. This repository had already recorded that trap once, for a probe whose only mention of a
+tag was a comment, and it was walked into again anyway.
+
+**So a claim is now a structured marker.** `CONFIRMS:` for a constraint checked against a real
+kernel, `EXERCISES:` for a permission the resolver takes, each naming a clause and nothing else --
+a hedged `CONFIRMS: RS-C-4 eventually` does not parse as a claim, and that case is asserted. The
+rebuilt census was verified to go red in three directions before being trusted: a removed marker, a
+marker naming a clause the document does not declare, and a prose mention standing in for a marker.
+
+**Two blind spots are declared rather than left invisible.** A census over source proves a clause is
+*claimed*, never that the file's assertion still runs or still means anything -- comparing a
+constant against itself leaves the marker in place and the suite green. And dropping the covering
+flag does **not** fire the `RS-C-4` assertion on every machine: measured here, a device stack that
+orders a flush behind its file's outstanding writes by itself produces no violation to see, so that
+sabotage demonstrates nothing portable. The assertion's conformance value -- reporting a violation
+if one occurs -- is separate from its sensitivity, and only the first is claimed.
+
+**The sweep found a stale site from two milestones back.**
+[D-60](DESIGN-NOTES.md#d-60) still listed `SetIoRingCompletionEvent` among the five lifecycle calls
+left outside the seam, which `M26.3` had moved behind it. The correction is recorded in that row
+rather than only in `D-61`, since a reader arriving there would otherwise take the superseded list
+as current.
+
+**The "five techniques" framing becomes six, and the sixth is different in kind.** The other five
+check this crate against its own stated contract and cannot tell you that contract is wrong -- which
+is what happened in the two most expensive defects. The resolver checks the crate against a written
+specification of what the platform may do, and the kernel tests check the platform against that same
+specification. `M26`'s row also joins defect population A, on the observation that the kernel's
+*response* is a precondition and was never varied either.
+
+## Moved 2026-09-25 10:48:58 -04:00 -- M26.7: auditing the suite for frozen observations
+
+### M26.7 -- Audit the existing suite for assertions that are frozen observations rather than contracts. *(completed 2026-09-25 10:48:58 -04:00)*
+
+The shape is recorded as [D-66](DESIGN-NOTES.md#d-66); what follows is what the work found.
+
+**The census came from a command, and it was not the file anyone guessed.** The item named
+[flush_barrier.rs](tests/flush_barrier.rs) as the obvious candidate and said plainly that the
+candidate was a guess. It was: that file turned out to be one of the better-behaved ones, since its
+transfer assertion is explicitly framed as its own precondition and its ordering counter is
+*reported* rather than asserted. The census started from the whole assertion population, narrowed to
+assertions in tests that touch a real ring -- reusing `RING-OPENING-LIB-TESTS.txt` as the classifier
+rather than inventing a second one -- and then to four candidate shapes.
+
+**One class, 31 assertions wide, across five files.** `try_pop()` straight after `submit_and_wait`
+with the `Option` unwrapped, asserting the kernel had **already** queued the completion. That is
+`D-52`'s demonstrated failure exactly -- an assertion that gives opposite answers on two handles of
+the same API -- and this crate's own documentation denies it: a submit-side wait's return "promises
+nothing about poppability". `RESPONSE-SPACE.md` states it as `RS-P-5`. They passed for the reason
+[D-40](DESIGN-NOTES.md#d-40) measured: a buffered read completes inside the submit in 80 of 80
+attempts, while an unbuffered one genuinely pends.
+
+**Restated as the contract `D-52` prescribed** -- ask for the completion within a bound this crate
+chooses -- which holds on every handle rather than on the one a test happens to open. The same sweep
+found an unbounded `try_pop` spin loop in `registration.rs`, which is the shape `pop_within` was
+introduced to replace, and it went the same way.
+
+**Nothing catches a regression by running, and that is the part worth keeping.** Reverting a site
+leaves its test green on any machine where the observation is true, which is precisely why 31 of
+them survived years of review. So the guard is a **census that refuses the shape at the source**,
+plus a resolver-driven test that makes the pending case reachable on demand and shows `try_pop`
+failing where `pop_within` succeeds. One of the three sabotages exists to make that point: reverting
+a restated site is caught by the census and by nothing else, and a reader who checked by re-running
+the test would find it green and conclude the change was cosmetic.
+
+**A second, narrower class was found and deliberately not settled.** Five assertions require a
+complete transfer. `Completion::result` promises only "the transferred byte count", and the space
+lists partial transfers as *deliberately undecided* with the instruction to decide them before a
+resolver relies on either answer. So the suite silently answers a question the specification leaves
+open -- and the two are not yet in conflict only because the resolver reports `Information: 0` and
+no resolver-driven test reads a transfer count. Queued as `M26.10`. One of the five is already in
+the honest form and is left alone.
+
+**What the audit did not find is worth stating too.** The `outstanding() == 0` assertions after a
+rundown are the crate's own contract, not observations. `flush_barrier_stress.rs`'s assertion that
+reordering *happened* is a control verifying its own precondition, with a message explaining why --
+the correct pattern. And the handover tests rest on `D-40`'s synchronous-completion measurement but
+verify that precondition rather than assuming it, which is what a frozen observation fails to do.
+
+## Moved 2026-09-25 12:34:02 -04:00 -- M26.8: retry policy belongs to the caller
+
+### M26.8 -- Decide what `IoRing::run_down` should do when a submit it makes is refused. *(completed 2026-09-25 12:34:02 -04:00)*
+
+The shape is recorded as [D-67](DESIGN-NOTES.md#d-67); what follows is what the work found.
+
+**The item was framed as a decision and was mostly a reading.** It asked which way to resolve a
+tension between "blocking is the safe failure mode" and "no hang". The tension was real but the
+question had an authority nobody had consulted: `SubmitIoRing`'s reference page. The correction that
+made the difference came from the engineer -- **no amount of measurement constitutes a contract** --
+and it was needed, because the session had spent the previous hour measuring and had drawn a
+confident conclusion from it.
+
+**What the documentation settles.** A return-value row gives `IORING_E_WAIT_TIMEOUT` the meaning
+*"All operations were submitted without error and the subsequent wait timed out"*, and the Remarks
+add *"If this function returns an error other than IORING_E_WAIT_TIMEOUT, then all entries remain in
+the submission queue"* and that a per-entry failure arrives as a completion rather than as a submit
+failure.
+
+**A measurement had been read backwards, and the documentation is what caught it.** A probe showed
+an operation completing after a submit reported `E_INVALIDARG`, which was taken to mean the work had
+gone in despite the error. It had not: the entry stayed in the submission queue exactly as
+documented, and what pushed it through was `pop_within`'s *own* internal submit. The observation was
+right and the attribution was wrong -- which is precisely the failure mode a contract prevents and a
+measurement invites.
+
+**`Batch::submit_and_wait` carried `M21.6`'s defect** at the one site that sweep did not reach.
+`do_submit` passed a timed-out wait to `check`, so a fully successful submission was reported as an
+error. The damage is worse than a wrong sign: because any *other* error means the entries are still
+queued, an `Err` was ambiguous between "your buffers are free" and "the kernel still owns them",
+which is [D-5](DESIGN-NOTES.md#d-5)'s hazard with the sign hidden.
+
+**`run_down` was the only waiting API in this crate shaped wrongly**, and the engineer named the
+principle that identifies it: *never implement a retry policy ourselves -- the caller keeps their
+own backoff, counts before giving up, and whatever else they want.* Waiting in segments inside a
+period the caller supplied is fine; `run_down` had the segments and no such period, which made "how
+long to keep trying" this crate's policy. [`IoRing::run_down_within`](src/ring.rs) is the primitive
+the caller bounds; `run_down` is now that with an unbounded period, which is a choice made by
+calling it. Checked against the rest of the surface rather than assumed: `pop_within` and
+`submit_and_wait` were already correctly shaped, so `run_down` really was the sole outlier.
+
+**An error from rundown is no longer terminal, and did not need to be.** The entries remain queued,
+the ring is resumable, and the documentation is what makes that safe to say. The one thing a caller
+must not do after an error is drop the ring.
+
+**`RS-P-7` and `RS-P-3` moved from inference to citation.** `M26.1` wrote `RS-P-7` as a consequence
+clause -- "if a submit fails, entries remain queued" -- citing a decision that established the
+consequence and nothing about submits failing at all, and `M26.3`'s resolver read it as a permission
+regardless. The space now carries a **`Documented`** tag that explicitly outranks `Observed`, since
+a measurement describes one run of one build.
+
+**The manifest drifted and the harness caught it, for the seventh time this session.** Renaming
+`WAIT_EXPIRED` to `IORING_E_WAIT_TIMEOUT` -- so the constant carries the documented name rather than
+a derivation -- left `M26.5`'s sabotage patching text that no longer existed. The sweep that
+followed found four more live sites; the archive was left alone. A per-case uniqueness check now
+runs before the sweep, which would have caught this in seconds rather than in a five-minute run.
+
+## Moved 2026-09-25 14:19:35 -04:00 -- M26.9, the event_delivery stall
+
+### M26.9 -- The `event_delivery` stall: the wait was armed after the event was signalled, which `SetThreadpoolWait` forbids. *(completed 2026-09-25 14:19:35 -04:00)*
+
+Cause, fix and before/after figures are in [DESIGN-NOTES.md](DESIGN-NOTES.md) -> D-68; the
+investigation as it stood when the cause was found is in
+[RESOLVED-TEST-FAILURES.md](RESOLVED-TEST-FAILURES.md). The item as it read when it closed:
+
+- [x] **M26.9** -- Find and fix the intermittent `Timeout` in
+ [event_delivery.rs](tests/event_delivery.rs)'s two threadpool-delivery tests, recorded in
+ [UNRESOLVED-TEST-FAILURES.md](UNRESOLVED-TEST-FAILURES.md).
+
+ **Why this is not merely a flaky test to re-run.** Measured at 1 failure in 80 runs of the
+ compiled binary. The sabotage harness runs the whole suite once per case and the manifest holds
+ 41, so that rate gives roughly a **40% chance of a corrupted sweep** -- and the corruption falsely
+ reports `caught`, which is a sabotage the suite did not catch being recorded as a clean bill of
+ health. The harness is this repository's mechanism for keeping earlier guarantees checked; a 40%
+ chance of a silent false pass undermines every conclusion drawn from it.
+
+ **What has already been ruled out, so it is not re-tried:** ring-resource pressure (zero failures
+ after roughly 18,000 ring create/close cycles), the widened seeded sweeps (zero after repeated
+ property-suite and calibration runs), and CPU starvation (zero under a concurrent `cargo build`
+ saturating the machine).
+
+ **Narrowed 2026-09-25 to a minimal reproducer, and the investigation now leaves this crate.**
+ Measured: 0 failures in 1000 serial runs against 7 in 1000 parallel; every occurrence identical,
+ with **both** delivery tests failing together and `callbacks run: 0` -- the pool never invokes the
+ callback at all, for either ring, and nothing arrives ten seconds later. The two delivery tests
+ alone do not reproduce it (0 in 1000); a third test has to be co-running, and the two that trigger
+ it both create an `EventDelivery` over a ring with nothing outstanding and drop it promptly.
+ Full figures, the reproducer, and what remains unestablished are in
+ [UNRESOLVED-TEST-FAILURES.md](UNRESOLVED-TEST-FAILURES.md).
+
+ **The next step is in [`windows-threadpool-sys`](../windows-threadpool-sys/src/wait.rs)**, not
+ here: both waits are registered on the default process threadpool, and one object's lifecycle
+ appears to stop other, unrelated armed waits from ever firing. Per the repository's mono-repo bug
+ policy, the fix belongs in that layer. `ThreadpoolWait`'s `Drop` has been read and only touches
+ its own object, so the mechanism is **not yet established** -- do not start from a guess about it.
+
+ **Narrowed further the same day, with a configurable trace.** The default pool is **not** wedged:
+ a probe at the moment of failure runs a plain work item and a brand-new armed wait, and measured,
+ both ran. The stalled waits were created and armed -- the trace shows it -- and the trampoline
+ never fires for any of them. **The stall is permanent by design**: the completion event is edge
+ triggered ([D-19](DESIGN-NOTES.md#d-19)), a stalled ring's queue never returns to empty, so the
+ setup signal is the only wakeup that ring will ever receive and losing it once ends delivery for
+ good. The open question is now narrow -- why an armed wait does not observe a signal raised just
+ before it was armed -- and a plausible mechanism is recorded in
+ [UNRESOLVED-TEST-FAILURES.md](UNRESOLVED-TEST-FAILURES.md) **as a hypothesis with no evidence
+ behind it**, together with the experiment that would settle it.
+
+ **The trace is compiled out unless `--features trace` is on**, and narrowed at run time by
+ `WINDOWS_THREADPOOL_TRACE`, because the instrument for a timing-dependent fault must not change
+ the schedule it measures. The flake still reproduces with it on, which was checked first.
+
+ **A cheaper interim mitigation exists and is a separate decision:** the harness could treat a
+ failure in these two tests as *inconclusive* rather than as `caught`, which would stop the false
+ clean bills without pretending the behaviour is understood.
+
+## Moved 2026-09-25 15:02:00 -04:00 -- M26.10, the partial-transfer decision
+
+### M26.10 -- Decided: a completion may report fewer bytes than requested, so `RS-P-8` permits it and the caller owns the remainder. *(completed 2026-09-25 15:02:00 -04:00)*
+
+The decision and its two-layer shape are in [DESIGN-NOTES.md](DESIGN-NOTES.md) -> D-69; the clause
+itself is `RS-P-8` in [RESPONSE-SPACE.md](RESPONSE-SPACE.md). The durability half it spawned is
+`M26.11`. The item as it read when it closed:
+
+- [x] **M26.10** -- Decide whether a completion may report **fewer bytes than requested**, which
+ [RESPONSE-SPACE.md](RESPONSE-SPACE.md) currently lists as deliberately undecided. Found by
+ `M26.7`'s audit, and queued rather than settled there because the space's own instruction is to
+ "decide it before a resolver relies on either answer" -- which is a call about what this crate
+ tolerates, not a correction.
+
+ **The observation.** Five kernel assertions require a full transfer (`assert_eq!(transferred,
+ LEN)`), so the suite already answers the question by assuming one. `Completion::result` promises
+ only "the transferred byte count" and never a complete one, so nothing in the crate backs that
+ assumption up. The two are not in conflict today only because `M26.3`'s resolver reports
+ `Information: 0` for every operation and no resolver-driven test reads a transfer count -- so a
+ resolver and a kernel test would disagree about the same field and nothing would notice.
+
+ **Note one of the five is already right and should be left alone.**
+ [flush_barrier.rs](tests/flush_barrier.rs) checks the transfer explicitly as its *own
+ precondition* -- a short write would make every count in that test meaningless -- and says so.
+ That is the honest form of the assertion whichever way this is decided.
+
+ **What deciding it costs.** Permitting partial transfers widens what every consumer must handle
+ and would make the resolver able to produce them, which the properties in
+ [properties_under_every_resolution.rs](tests/properties_under_every_resolution.rs) would then
+ have to survive. Requiring complete transfers is a `Decided` constraint of the kind `RS-C-1`
+ already is, and would need a `CONFIRMS:` marker on whichever kernel test carries it -- the census
+ added in `M26.6` will then hold the two halves together.
+
+## Moved 2026-09-25 16:21:00 -04:00 -- M26.11, the epoch log's handle requirements
+
+### M26.11 -- The epoch log states the capabilities it requires of a handle; the transfer requirement is checked, the durability requirement is a caller warranty. *(completed 2026-09-25 16:21:00 -04:00)*
+
+The contract gained a fourth clause, `Requires`, beside `Guarantees`,
+`DoesNotGuarantee` and `Assumes` -- the growth the module's own `Clause::ALL`
+doc had anticipated, and its exhaustive `heading()` match is what forced the
+edit. The decision not to gate the handle is [D-70](DESIGN-NOTES.md#d-70). The
+item as it read when it closed:
+
+- [x] **M26.11** -- Document the capability requirements [examples/epoch_log](examples/epoch_log) places on
+ the handle it is given, and let an unmet one surface at the operation that needs it. **No pre-flight
+ handle check** ([D-70](DESIGN-NOTES.md#d-70)).
+
+ **Why no check, which is the part worth not relitigating.** `M26.10` established that the ring permits a
+ short transfer because it never asks what kind of handle it was given ([RS-P-8](RESPONSE-SPACE.md),
+ [D-69](DESIGN-NOTES.md#d-69)), and the obvious next move is for the log to police what the ring does
+ not. It was examined and rejected: the check points the wrong way. A console handle, a pipe or a closed
+ handle is caught by `GetFileType`, but every one of those already fails loudly at the first positioned
+ write -- the check buys a better message. A RAM disk, a remote share, or a volume whose write cache is
+ not power-protected succeeds at every API call and silently fails to be durable, and no probe catches
+ the last of those at all. So the gate guards the failures that were already loud and misses every
+ failure that is silent, while implying a validation that did not happen.
+
+ **What to write instead.** The log's own contract, stated by the log rather than inherited from the
+ ring, listing what the handle must support: positioned I/O at explicit offsets; `FILE_FLAG_OVERLAPPED`;
+ `FILE_FLAG_NO_BUFFERING` together with the sector-aligned buffer, offset and length it requires; a
+ preallocated extent; that a successful write of `N` bytes transfers `N`; and that a completed flush
+ reaches stable media.
+
+ **Mark the last two for what they are.** The transfer requirement is the `RS-P-8` narrowing this log
+ earns by constraining its input -- it is checkable, and the log already compares transferred against
+ requested, so a mismatch is a contract violation to report loudly rather than a case to absorb. The
+ durability requirement is **a warranty the caller gives**, not a property this code can verify;
+ say so in those words, because a contract that merely sounds confident about it is how a silent
+ failure gets built on.
+
+ **Guard what is guardable, and do not pretend about the rest.** The transferred-against-requested
+ comparison is a real assertion and gets a sabotage case: suppress the comparison and the suite must go
+ red. There is no guard for the durability warranty, and the item is complete with that stated rather
+ than papered over. If the contract carries runnable examples they are compiled as doctests, per the
+ repository's rule that prose containing code must compile.
+
+## Moved 2026-09-25 16:24:00 -04:00 -- M26 complete, all eleven items
+
+Every item was archived individually as it closed, so what migrates here is the milestone's own
+framing -- why a resolver over a permitted space was worth building, and the standing constraint
+that the space must be wider than anything observed. The per-item stubs are dropped, carrying
+nothing the entries above do not already hold.
+
+## M26 -- Test against the space of kernel responses, not one observation of it
+
+Queued by
+[DESIGN-SESSION-2026-09-22-kernel-response-space.md](design-sessions/DESIGN-SESSION-2026-09-22-kernel-response-space.md),
+which set out to answer `M24.1` and found a different technique instead.
+
+**The idea.** A fake that models *what Windows does* freezes one run's testimony. A **resolver**
+models what Windows is *permitted* to do, and a seed picks one resolution out of that space: which
+operations finish inside `SubmitIoRing` and which pend, in what order completions are posted, which
+fail. The assertions are then about **us** -- does this crate behave correctly under that resolution
+-- and never about the kernel. There is no belief to be wrong about, which is why this dissolves the
+mock objection rather than working around it.
+
+**Justified by what it catches, not by hermeticity.** `M24` reaches a hermetic lib suite without it,
+so this milestone has to earn its place on the defect class it detects: code that is brittle to
+platform variation *inside* the permitted space. Nothing in the toolkit that preceded it detected
+that -- [DESIGN-NOTES.md](DESIGN-NOTES.md#what-none-of-them-cover) records that the five techniques
+which existed before `M26` all check this crate against *its own stated contract*. The resolver is
+the sixth, added by this milestone.
+
+**The standing constraint, inherited from the session.** The permitted space must be **wider than
+anything observed**, and must not be derived from observation -- deriving it from what we have seen
+closes the trap again. It is a deliberate specification of what we will tolerate, and therefore a
+reviewable artifact rather than a recording.
diff --git a/crates/windows-ioring-sys/Cargo.toml b/crates/windows-ioring-sys/Cargo.toml
index 6ff8a6613..e75f9b3c1 100644
--- a/crates/windows-ioring-sys/Cargo.toml
+++ b/crates/windows-ioring-sys/Cargo.toml
@@ -21,6 +21,12 @@ path = "src/lib.rs"
[features]
default = ["threadpool"]
+# Propagates windows-threadpool-sys's concurrency trace. Off by default and
+# absent from the compiled output when off: the defects it exists for are
+# timing-dependent, so the instrument must not change the schedule it is
+# measuring. Narrow it with WINDOWS_THREADPOOL_TRACE at run time.
+trace = ["windows-threadpool-sys/trace"]
+
# A test-support seam that makes a *real* completion report a failure it did
# not actually have (M16.3). Off by default: it is for testing failure
# handling, and nothing in production has a use for it.
@@ -31,6 +37,15 @@ default = ["threadpool"]
# whole design and is argued in full on `Completion::with_injected_failure`.
# Fabrication stays `#[cfg(test)] pub(crate)` and is not reachable from here.
fault-injection = []
+
+# The indirection the M26 kernel-response resolver plugs into (M26.2). Off by
+# default: with it off, every wrapper in `sys` is an inline forward to the same
+# `windows-sys` call this crate made before, and the install point does not
+# exist at all.
+#
+# A feature rather than `cfg(test)` because the integration suite in tests/
+# cannot see `cfg(test)` items -- the same reason `fault-injection` is one.
+kernel-seam = []
# Model A delivery (`EventDelivery`) is the only thing in this crate that needs
# a thread pool, so a Model B consumer -- a pinned thread parked in
# `Batch::submit_and_wait`, owning its own ring -- otherwise links a dependency
@@ -59,6 +74,21 @@ required-features = ["threadpool"]
name = "epoch_log"
path = "examples/epoch_log/main.rs"
required-features = ["threadpool"]
+# M21.4 gives the sample's durability accounting unit tests -- the failed-commit
+# case cannot be reached by running the sample, because it needs a completion
+# whose result is an injected failure. Examples are not test targets by default,
+# so `cargo test` would compile this and run nothing.
+test = true
+
+# `ring_copy` was auto-discovered until M20.3, which needed it to be a test
+# target for the same reason `epoch_log` is: the degraded-fallback branch in
+# `Policy::select` is the one every zero-relation machine takes, and it cannot
+# be reached by *running* the sample on a machine that reports its relations.
+# A synthetic topology reaches it; a test target is what lets one run.
+[[example]]
+name = "ring_copy"
+path = "examples/ring_copy/main.rs"
+test = true
# The crate is Windows-only, so docs.rs must build on a Windows target or it
# would render an almost-empty crate.
@@ -72,6 +102,11 @@ targets = ["x86_64-pc-windows-msvc"]
# under `Win32_System_IO` as their subject matter might suggest.
# `Win32_System_Threading` supplies the event primitives the Model A delivery
# path (M4) waits on.
+#
+# `Win32_System_Memory` is gone from this list, and its absence is the check
+# that `NumaBuffer` really left: the allocator was the only thing here that
+# reached for it, so a build that still needed it would mean something had been
+# missed. It now belongs to `win-numa-sys`.
windows-sys = { version = "0.61.2", default-features = false, features = [
"Win32_Foundation",
"Win32_Security",
@@ -83,6 +118,13 @@ windows-sys = { version = "0.61.2", default-features = false, features = [
# event. Optional behind the default-on `threadpool` feature (D-22); nothing
# else in the crate references it.
windows-threadpool-sys = { version = "0.1.3", path = "../windows-threadpool-sys", optional = true }
+# NumaBuffer moved here in M22.3 and out again when win-numa-sys was created:
+# the buffer has nothing to do with a ring, and windows-placement-probe had
+# independently written the same VirtualAllocExNuma call. This crate keeps
+# re-exporting the type, so windows_ioring_sys::NumaBuffer still resolves for
+# anyone who bound to it, and supplies the IoBuf/IoBufMut impls that the
+# allocator crate deliberately does not define.
+win-numa-sys = { version = "0.1.0", path = "../win-numa-sys" }
[dev-dependencies]
# M6.3's topology-guidance example enumerates L3 cache domains with this
@@ -111,13 +153,18 @@ windows-topology-sys = { path = "../windows-topology-sys", features = [
] }
# M7's ring-copy sample deserializes a fed-in topology description (--topology).
serde_json = "1.0"
-# M7's ring-copy sample additionally needs `VirtualAllocExNuma` (buffer
-# placement) and `GROUP_AFFINITY` (pinned-thread affinity); M14.1's epoch-log
-# sample needs `Win32_System_Ioctl` for the `FSCTL_SET_ZERO_DATA` reclamation
-# it orders against ring epochs. The library itself needs none of them.
+# M7's ring-copy sample additionally needs `GROUP_AFFINITY` (pinned-thread
+# affinity); M14.1's epoch-log sample needs `Win32_System_Ioctl` for the
+# `FSCTL_SET_ZERO_DATA` reclamation it orders against ring epochs, and for the
+# `FSCTL_QUERY_VOLUME_NUMA_INFO` its arena placement asks (M22.3).
+#
+# `Win32_System_Memory` is deliberately NOT here: buffer placement moved into
+# the library as `NumaBuffer` (M22.3), so the feature is a *library* one now
+# and listing it again would wrongly suggest a sample reaches for
+# `VirtualAllocExNuma` on its own.
windows-sys = { version = "0.61.2", default-features = false, features = [
"Win32_System_Ioctl",
- "Win32_System_Memory",
+ "Win32_System_Pipes",
"Win32_System_SystemInformation",
] }
# M14.1's epoch-log sample orders a non-ring `FSCTL` against ring epochs, which
diff --git a/crates/windows-ioring-sys/DESIGN-NOTES.md b/crates/windows-ioring-sys/DESIGN-NOTES.md
index 5fee420df..c1b4cc167 100644
--- a/crates/windows-ioring-sys/DESIGN-NOTES.md
+++ b/crates/windows-ioring-sys/DESIGN-NOTES.md
@@ -1,11 +1,22 @@
# Design notes: windows-ioring-sys (Tier 1)
-This crate does not exist yet as compiled code. This file, the checklist beside it, and the design session
-it references are the design record that precedes it. Creating the Cargo skeleton is M1.1 in
-[CHECKLIST.md](CHECKLIST.md).
+This file is the crate's current design record: what was decided, and what forced each choice. It was
+written before the implementation, and the design session it references is where its earliest decisions
+came from.
+
+It deliberately states no release or milestone status. That is derivative of
+[CHANGELOG.md](CHANGELOG.md), the git tags, and [CHECKLIST.md](CHECKLIST.md), and a copy of it here would
+be one more thing to keep true by hand.
## Intent
+**Why this crate exists at all is recorded one level up**, in the workspace's
+[DESIGN-NOTES.md](../../DESIGN-NOTES.md) -> [The adoption thesis](../../DESIGN-NOTES.md#the-adoption-thesis).
+Read it before proposing to remove a design option here: it is the reason this crate keeps
+alternatives alive that no measurement on the development machine can justify, and the reason its
+samples hand a consumer data rather than a verdict. The operational form of that posture is OPTION
+INTEGRITY in [copilot-instructions.md](../../.github/copilot-instructions.md).
+
Windows 11 / Server 2022 added `IoRing`: a submission/completion ring for file I/O, closer in shape to
`io_uring` than to anything else Windows offers. This crate raises those primitives into memory-safe Rust
with the minimum additional CPU and memory cost, in the same spirit as the rest of this repository.
@@ -69,6 +80,30 @@ runs a continuation), this crate exposes the mechanism and documents the trade-o
| D-44 | **A spike against the real kernel is a budgeted, first-class technique for every new Win32 surface this crate wraps -- not something that happens after a test fails mysteriously.** The full argument is in [Testing strategy](#testing-strategy-m185); the decision is that the budget is allocated *before* the wrapper is written. Two of the eight defects behind M15-M18 exist because a Win32 contract was assumed rather than measured: the completion event is edge-triggered ([D-19](#d-19)) and `BuildIoRingRegisterBuffers` reads its array when the operation *runs* ([D-32](#d-32)). No oracle, generator, allocator or mutation run supplies that knowledge, because each of them checks code against **our** stated contract -- and in both cases our stated contract was the thing that was wrong. What they detect is a *consequence*, and only on a path some test already walks: the guard allocator does turn D-32 into a hard `STATUS_ACCESS_VIOLATION`, measured in M17.4's calibration, but that is the crash after the mistake, not the knowledge that would have prevented it. A spike is also the only technique here that can be run *before* there is code to test. Two obligations follow, both learned the hard way and recorded in [design-sessions/spikes/README.md](design-sessions/spikes/README.md): a spike must carry a **control case**, because the first two drain-ordering spikes could not discriminate and would have returned confidently wrong answers; and it must be **kept**, as a standalone single-file program depending only on `windows-sys`, so that what it measures stays the operating system's behaviour rather than ours. |
| D-45 | **A borrow-returning method must be audited on two questions, not one: what the returned value *permits*, and how long the *borrow* lasts. `RegisteredBuffers::get` therefore takes `&mut self`.** [M18.1's audit](#borrow-surface-audit-m181) asked only the first, of all nineteen items, and the second is where [D-36](#d-36)'s fix was still open: `get` checked `kernel_writes` at the instant of the call but returned a slice living as long as the borrow, and `Batch::read_registered` takes the registration by **shared** reference -- so safe code could take the borrow while the buffer was quiet, then submit a read into that same buffer and keep reading. Measured before being believed: a probe watched the bytes change from `0x11` to `0xEE` through the live slice while a fresh `get(0)` at that same instant correctly refused with `WouldBlock`. The guard worked; the borrow outlived it. **`&mut self` costs nothing real**, because no caller needs to read a buffer during the window it is refused -- while a read is in flight the bytes are indeterminate and only become meaningful once the completion is observed, so earlier or later is always available. That is not merely an argument: all ~40 read sites in this crate's tests, examples and the epoch-log sample already read at a quiescent point, and converting them needed nothing but `mut` on a local. The concession D-36 deliberately kept (reading a buffer whose own *write* is in flight, where the kernel only reads) is given up with it, and is likewise unused. The arena pattern survives, because a [`Token`] holds a [`RegisteredUse`] rather than a borrow of the registration, so quiet neighbours stay readable while operations are outstanding. Enforced by a `compile_fail` doctest, itself verified by reverting the signature and watching it fail, and paired with a `no_run` doctest asserting the neighbour case still compiles so the guard cannot become over-constraining unnoticed. `get_mut` never had the defect: `&mut self` already conflicted with the shared borrow. |
| D-47 | **`IOSQE_FLAGS_DRAIN_PRECEDING_OPS` is one-sided, and [D-24](#d-24)'s claim that it holds back subsequent operations is withdrawn. Measured over ~4,500 trials.** The drain half is solid: **not once** did an operation queued *before* a drained flush complete after it. The hold-back half is false: post-flush writes overtake the flush at 0.03%-0.8% depending on conditions, and in the worst observed trial *all 32* did. The rate is why this read as a flaky test for three days rather than as a contract defect -- at roughly one run in a thousand, it surfaces every few days in a full-workspace run and never in isolation. Contention raises the rate but is not required (it reproduces on an idle machine); ring depth does not move it (128, 256 and 512 were indistinguishable). Every violation was confined to a single drain of the completion queue, so it is the queue's own posting order rather than an artifact of sampling it twice. **What a consumer may rely on:** a drained flush's completion means everything outstanding when it was reached is durable. It does **not** mean later work has been held. See [D-47 in detail](#d-47-detail). |
+| D-48 | **A shipping ARM consumer laptop reports no L3 cache domains at all, and zero `Win32_NumaNode` instances. Measured, and it is an ordinary consumer shape rather than an exotic one.** A Snapdragon X2 Elite (X2E80100, Qualcomm Oryon), 12 cores and no SMT, reports L1 and L2 only: L2 forms two domains of six processors, agreeing with the two `Module` domains the same probe returned, and WMI reports an L3 size of zero. The capture is Measurement M-1 in [DESIGN-SESSION-2026-08-30-numa-sharded-io-execution-domains.md](../../design-sessions/DESIGN-SESSION-2026-08-30-numa-sharded-io-execution-domains.md). **This is the sibling of the zero-node observation these notes already carried** in [Why the NUMA node is the wrong key](#why-the-numa-node-is-the-wrong-key), and the pair is the point: on one machine the NUMA node is absent because a hypervisor did not present it, on the other the *cache level this crate's guidance names* is absent because the silicon has none. Neither machine is unusual. **What it falsifies is a justification, not the heuristic.** These notes say the last-level-cache domain "is meaningful on Intel and ARM too, where the NUMA node often is not"; on this part it is not meaningful, because it does not exist. That a cache domain beats the node is untouched. **The rule was restated and every restatement swept by `M20.1`/`SH-4.12` (2026-09-22)**, which also replaced the consumer that bound to the level number. Measuring that consumer while fixing it found a second shape this decision did not anticipate: an L3 that spans *every* processor above a real L2 partition, where a `level == 3` filter matches, does **not** degrade, and reports one whole-machine domain as a successful cache-aware partition -- see [Why the NUMA node is the wrong key](#why-the-numa-node-is-the-wrong-key). **Not decided here:** whether the two-cluster L2 structure this part does report is worth sharding on. |
+| D-49 | **The unit suite is not hermetic, and that is a defect rather than a property of wrapping Win32. 63 of 131 lib tests open a real kernel ring.** The repository's Quality rule already classifies this: it reserves integration tests for work that must cross "a real process, filesystem, network, device, **operating-system API**, or other external boundary", and `CreateIoRing` is an operating-system API. So those 63 are integration tests sitting in the unit-test location, and `cargo test --lib` does not mean what its name implies. **This was surfaced by a load-dependent failure and initially mis-diagnosed.** `M21.6` removed five wall-clock assertions from four lib tests, which made their *outcome* independent of load -- measured at 30x duration variation with zero outcome variation under 2x CPU saturation -- and that was reported as the fix. It was not: outcome-stable and hermetic are different properties, and those tests still open a kernel ring. **What is decided here is the defect and the classification, not the remedy.** Three remedies are costed in [DESIGN-SESSION-2026-09-21-hermetic-unit-tests.md](design-sessions/DESIGN-SESSION-2026-09-21-hermetic-unit-tests.md) and the choice between them is gated on one unresolved question: whether a fake whose assertions are *shared* with the kernel escapes the objection in [Two techniques deliberately rejected](#two-techniques-deliberately-rejected), which refuses a mock that would "manufacture evidence" a kernel-behaviour bug was absent. **`M24.1` settled that (2026-09-22) and the answer was that the instrument was wrong**: a shared suite is strong over what this crate specifies and blind to the platform's incidental behaviour, and an assertion about the latter is a frozen observation rather than a contract -- see [D-52](#d-52). The rejection stands with its scope sharpened; the technique that replaces it is the response-space resolver, scheduled as `M26`. **A bright line holds under every remedy:** a fake never answers a question about Windows. Edge-triggered delivery ([D-19](#d-19)), one waiter per ring ([D-21](#d-21)), flush coverage and drain ordering ([D-23](#d-23), [D-24](#d-24), [D-47](#d-47)), the registration array read at run time ([D-32](#d-32)), `ERROR_TIMEOUT` and `E_INVALIDARG` from `SubmitIoRing`, and inline completion on a synchronous handle all stay kernel-tested forever -- which is precisely the set of findings that produced this crate's defects, and the argument for drawing the line exactly there. **`M24` has since run, and the figures above are the ones it started from.** Measured after it: **41 of 151 lib tests open a ring, 110 do not**. The remainder is not movable without `M26.2`'s FFI seam -- `event_delivery` needs the thread pool, `ring`'s injected-failure cluster transforms a *real* completion by design, `batch` needs the handle, and several reach `#[cfg(test)] pub(crate)` helpers. A zero-check is therefore the wrong rung, and [D-53](#d-53) records what replaced it. |
+| D-50 | **The epoch-log sample places its registered arena on the NUMA node its own log file's volume reports, and says in the same breath that the placement cannot pay at this workload.** The arena was `vec![0_u8; SLOT_LEN]` -- heap, no alignment, no node -- while this crate's front page told every consumer that placing the registered pool near the device "is very likely the highest-leverage locality decision available". That silence read as an oversight. The node is asked of the log's own handle through `FSCTL_QUERY_VOLUME_NUMA_INFO`, which [What is not reachable](#what-is-not-reachable) already established as the documented mechanism; a volume that names no node yields no preference, and the log runs on. **What is deliberately not claimed is any benefit.** The arena is eight slots of four kilobytes and this workload is bound by a per-epoch device flush costing hundreds of microseconds -- `M22.1` measured that directly. So the sample demonstrates *how the decision is made and reported*, and `examples/ring_copy` remains where placement is put under a load that could show it. The report line is qualified by `GetNumaHighestNodeNumber` for the same reason: on a one-node machine "placed on node 0" is true and misleading, so the sample says the choice was never available. **This is a sample's local policy, not a retraction of [D-8](#d-8):** the library still maps no file to a node, because a volume may span devices and its node is not where a file's extents live. |
+| D-51 | **`NumaBuffer` moves from `examples/ring_copy` into the library, because recommending an allocation while making every caller write it is what produced the second copy.** This crate's front page names `VirtualAllocExNuma` on the device's node as the highest-leverage locality decision available, and then supplied nothing; the first consumer wrote the allocator in a sample, and `M22.3` was about to make a second. The type is a thin owned mapping implementing [`IoBuf`]/[`IoBufMut`], so it registers like any other buffer. **It decides no policy** -- which node is still the caller's answer, per [D-8](#d-8) -- and its own documentation records that `nndPreferred` is a preference, so a successful allocation is not evidence the pages landed there. The move surfaced a packaging defect that `cargo check --all-targets` cannot see: dev-dependency features are unified into that build, so the missing `Win32_System_Memory` on the *library* dependency only appears in a lib-only build or `cargo doc`. A consumer would have hit it on first compile. |
+| D-52 | **Test this crate against the *space* of responses the platform is permitted to give, not against one observation of what it gave. A fake that models what Windows does freezes one run's testimony; a seeded resolver over the permitted space has no belief to be wrong about.** `M24.1` set out to ask whether a fake with *shared* assertions escapes the mock rejection, and demonstrated that it does not: a fake built from this crate's own pre-`M21.6` belief passed the shared suite green, and the assertion that catches it could only be written after the kernel had already revealed the answer. The deeper finding is that an assertion about platform behaviour is a **frozen observation** rather than a contract -- "after submitting, the completion is already queued" gave *opposite answers on two handles of the same API*, and [write-pending-spike.rs](design-sessions/spikes/write-pending-spike.rs) reported one condition as 5/500 in one run and 271/500 minutes later. Freezing such an observation into a suite, then building a fake to satisfy it, leaves three artifacts agreeing -- which reads as corroboration but is one observation restated three times, the same failure [D-47](#d-47) already cost this crate once. **What replaces it:** the resolver decides from a seed which operations complete inside `SubmitIoRing` and which pend, and in what order completions are posted; the assertions are then about this crate's behaviour under that resolution. Non-reproducibility stops being a threat and becomes the expected case. **The permitted space must be wider than anything observed and must not be derived from observation**, which makes it a deliberate, reviewable specification of what we tolerate -- and it must state the constraints that *do* hold, or the tests demand code defending against impossible kernels. This is a sixth technique beside the five in [What none of them cover](#what-none-of-them-cover): it still cannot tell you the stated contract is wrong, but it detects brittleness to variation inside the space, which none of the five can. Reasoned, not yet measured, against `D-47` and `M21.6`; scheduled as `M26` in [CHECKLIST.md](CHECKLIST.md), whose calibration item exists because `D-41`'s corollary demands it. Session: [DESIGN-SESSION-2026-09-22-kernel-response-space.md](design-sessions/DESIGN-SESSION-2026-09-22-kernel-response-space.md). |
+| D-53 | **The rung guarding the hermetic lib suite is an inventory of *which* lib tests open a ring, not a zero-check and not a count.** `M24.5` originally assumed that after the extraction and the relocation "the lib tests should construct no ring at all", so a zero-check would do. That rule is false and cannot be made true by effort: `event_delivery` needs the thread pool, `ring`'s injected-failure cluster transforms a **real** completion on purpose (fabricating one is the unsoundness the seam exists to avoid), `batch` needs the handle for its `Build*` calls, and several tests reach `#[cfg(test)] pub(crate)` helpers that exist only inside the crate. A zero-check would fail on day one and could only be satisfied by deleting real coverage. **Per-test rather than per-file**, because two thirds of the remaining 41 live in `ring/tests.rs` and a file-level allow-list would let that file grow without limit -- which is where a new ring-opening test would most naturally land. **An inventory rather than a count**, because add-one-remove-one nets to zero and passes, and a bare number is derived data no reader can check. The mechanism is [check-borrow-surface.ps1](../../tools/check-borrow-surface.ps1)'s, deliberately: a committed list regenerated from source, failing when the two disagree, so an addition obliges the question *does this test need the kernel, or only a ring-shaped thing?* A removal is progress and needs only regeneration. **The guard's own bidirectional check found a defect in it**: a plain helper defined after the last test in a file was being swallowed into that test's body, reporting an innocent test as ring-opening -- the body now ends at a column-0 `}` rather than at the next attribute. Inventory: [RING-OPENING-LIB-TESTS.txt](RING-OPENING-LIB-TESTS.txt); script: [check-ring-tests.ps1](../../tools/check-ring-tests.ps1). |
+| D-54 | **This crate owns what *one* flush means. It owns nothing about durability *groups*, and that is a deferral rather than a gap.** The primitives are here because they are facts about a single operation: [`FlushCoverage`] (the barrier is ring-wide), [`FlushMode`] (which modes sync the device), [`WriteCaching`], and the scope distinction between them -- a barrier over the ring, a flush over one file. **Grouping is not here and is not coming here**: what a set of operations that commit together costs, whether two such sets contend, what a co-flush regime implies, and the `Epoch` concept itself. That layer is [C-3](../../design-sessions/DESIGN-SESSION-2026-08-30-numa-sharded-io-execution-domains.md)'s durability crate, queued as `M33+.5` in [CHECKLIST-io-domains.md](../../CHECKLIST-io-domains.md); today `Epoch` exists only in `examples/epoch_log`, and no grouping concept appears anywhere in `src/`. **The test for a proposal is whether it needs the concept of a set of operations that become durable together** -- if it does not, it may belong here and must be justified on its own merits; if it does, it is deferred. Recorded because the boundary was re-litigated three times in one day and because a conversational aside -- calling `M23.3` a "down-payment" on the durability crate -- left the impression that the layer had been folded into this one. It has not been, and building *toward* it from here is what this decision forbids. |
+| D-55 | **The pending-token inventory becomes the ring's, and `IoRing` becomes generic to hold it. The break is accepted.** This crate hands a caller a `Token` from one call and a `Completion` from another, and connects them with nothing -- its own rustdoc twice instructs a caller to "match it against a held `Token`". Every consumer with more than one outstanding tokened operation must therefore build an identity map, and the *correct* one encodes four rules a `HashMap` cannot express; two measured defects in this repository came from the obvious one. Offering a `Pending` beside the ring closes the duplication but not the mechanism: nothing would force a minted token into it. A generic `IoRing` owning the map does, because the consumer never holds a token to lose. **The type-erasure objection is withdrawn as false**: per-ring monomorphisation holds for every real consumer here, a closed `enum` serves the rest -- `tests/generated_sequences.rs` already carries eight token types on one ring that way -- and [D-4](#d-4) independently rules type erasure out ("no slab entry, no box, no type erasure"), so the objection contradicted a decision already on the books. Implementation and migration are `M28`. |
+| D-56 | **A benchmark that defers its await measures the deferral, not the operation -- and the number survived three rounds of correction because every round corrected the conclusion instead of the instrument.** `epoch_log`'s harness published a commit latency measured from pushing a flush to observing its completion. `M20.6` decomposed it and found **blocking p50 and p99 of 0 us for all three strategies**: the figure was entirely the interval in which the program went on appending, so a design that deferred further reported a worse commit while being no slower. Three separate rounds of work had already re-read the *conclusion* drawn from that number -- the harness's serialisation, the `UserData` collision, the per-record submit -- and none had asked whether the number measured what its name said. **The fix is structural: a commit's cost is now reported in parts** (`prepare`, `submit`, `blocking`) with `deferral` beside them and excluded from the total, so the two cannot be read as each other. `M25.5` then found the same defect one layer down in that fix -- the clock started after `HostSequenced`'s host round trip, making it look six times cheaper -- which is why `prepare` exists. **What generalises:** a measurement whose parts are not separately reported can be wrong in a way that no amount of re-reading its output will reveal, and "the conclusion still holds" is not evidence that the instrument does. See [measurements/2026-09-24-commit-decomposed/](measurements/2026-09-24-commit-decomposed/README.md). |
+| D-57 | **A flag whose requirements the caller's data layout cannot satisfy is not a flag change.** Making `epoch_log`'s commit separately observable needed `FILE_FLAG_NO_BUFFERING`, which constrains the transfer's buffer address, file offset **and length** -- and the log wrote variable-length records at packed offsets, satisfying none of them. The work was therefore a change to the log's **on-disk format** (`M25.1`: one record per sector-sized block, tail zeroed, in both writers) before a single flag could move (`M25.3`). [write-pending-spike.rs](design-sessions/spikes/write-pending-spike.rs) predicted exactly this from its conditions, and the prediction is the reusable part: when a configuration's preconditions reach into a caller's data layout, costing it as a flag underestimates it by the size of a format migration. The stride's price is write amplification, which the sample now measures and prints rather than describing, and the reason a real log pays it anyway is sector atomicity -- a record sharing a sector with its neighbour can be torn by that neighbour's write. |
+| D-58 | **A contract checker's failure vocabulary decides its interface, and `epoch_log`'s replay keeps a `&[u8]` for that reason rather than for simplicity.** The walk is strictly forward one block at a time and never looks back, so it has no need of the whole file -- and a real log is larger than memory, which makes reading the whole file the wrong reflex to teach at precisely the point a reader is learning to verify one. The slice is kept anyway because `replay` returns an `Outcome`, not a `Result`: every way it can end is a statement about the log -- verified, tolerated, or a `Violation`. A streaming reader introduces a third kind of ending, `io::Error`, into the one component whose entire job is to distinguish "the log broke its promise" from "the log kept it", and a signature returning both through one channel invites exactly the conflation the file exists to prevent -- an unreadable file reported as a missing durable record. **So the streaming version is a different interface, not a smaller allocation**, and the cost of declining it is stated where it is paid rather than hidden: 140 KiB for the log, 8 MiB per strategy in the harness. A consumer building a real verifier wants the other shape *and* wants the two failure kinds kept apart inside it. What was reducible without touching that interface was reduced: the cross-strategy comparison kept a whole reference log in memory for the length of the comparison and now keeps a digest, which halves the peak and loses nothing a reader had, since the assertion could already only say *that* two logs differed. |
+| D-59 | **The permitted kernel response space is specified in [RESPONSE-SPACE.md](RESPONSE-SPACE.md), as a statement of what this crate will tolerate rather than a record of what Windows did.** [D-52](#d-52) settled that testing against a *space* dissolves the mock objection; this is that space, written down. Eight clauses say what a resolver **may** do -- an operation may complete inside `SubmitIoRing` or pend (`RS-P-1`), completion order is unconstrained (`RS-P-2`), an operation may fail individually (`RS-P-3`), a wait may expire (`RS-P-4`) or return with nothing poppable (`RS-P-5`), a completion behind another need produce no signal (`RS-P-6`), a failed submit leaves operations queued for a later one (`RS-P-7`), and a successful transfer may report fewer bytes than requested (`RS-P-8`, added by `M26.10`). **Four say what it may not**, because a resolver free to violate everything makes this crate defend against a platform that does not exist: completions are conserved and identified (`RS-C-1`, `RS-C-2`), nothing completes before submission (`RS-C-3`), and **the drain half of `DRAIN_PRECEDING_OPS` holds (`RS-C-4`)**. That last is the call `M26.1` demanded rather than defaulted: [D-47](#d-47) measured roughly 4,500 trials without a single violation, the drain is what this crate's durability story rests on, and a resolver permitted to break it would make the primitive useless -- so a Windows that broke it is caught by the kernel tests instead, which is the division of labour they were repointed to in `M26.6` and is now enforced by a census rather than intended. The hold-back half stays unconstrained, since `D-47` withdrew it. **Every clause is tagged `Observed`, `Over-provision`, or `Decided`**, so a reader can tell a measurement from an extrapolation from a call, and the space is deliberately wider than anything observed -- deriving it from observation would close the trap `D-52` was opened to escape. Rates, partial transfers, failure-code sets and timing are listed as deliberately undecided so an omission cannot be mistaken for a choice. |
+| D-60 | **The kernel seam is module indirection, not a type parameter, because [D-55](#d-55) has already spent `IoRing`'s.** `M26.3`'s resolver has to be able to answer the crate's kernel calls, and the textbook shape for that is a generic `IoRing` over a kernel trait. That shape is unavailable here: `D-55` commits `IoRing`'s type parameter to `M28`'s token inventory, so a kernel generic would publish `IoRing` on a shipped crate -- a second public parameter whose only purpose is letting the crate test itself. **The eight submission-path calls therefore route through [sys.rs](src/sys.rs) instead** -- `SubmitIoRing`, `PopIoRingCompletion`, and the six `Build*` entry points -- each an `#[inline(always)]` wrapper whose `through_seam!` macro expands to the bare FFI call when the `kernel-seam` feature is off. The public surface is unchanged and the type parameter stays free. **The responder is thread-local rather than process-global** for the reason this workspace is not on nextest: `cargo test` runs tests as threads in one process, so a global responder would let one test answer another's calls -- the same hazard that puts `DROP_RUNS` inside its test function. `with()` falls through to the kernel on a re-entrant borrow rather than panicking, and `Installed::drop` uses `try_with` so teardown cannot abort ([M23.4](COMPLETED-CHECKLIST.md#m234)). **Four lifecycle calls are deliberately left direct** -- `CreateIoRing`, `CloseIoRing`, `GetIoRingInfo`, `IsIoRingOpSupported`. (This decision originally listed five; `M26.3` moved `SetIoRingCompletionEvent` behind the seam, because it is how a completion becomes *observable* and `RS-P-6` is therefore a clause about it -- see [D-61](#d-61). The correction is recorded here rather than only there, since a reader arriving at this row would otherwise take the superseded list as current.) `M26` justifies itself on the *response space*: what the kernel may answer to submitted work. Routing construction and teardown through the seam too would be hermeticity for its own sake, which is `M24`'s subject and not this one. The boundary is stated here so a later reader can tell it from an oversight -- and `M26.3` carries the converse, that a clause needing one of those five extends the seam rather than working around it. |
+| D-61 | **The resolver's permissions are configurable and its constraints are not, and that asymmetry is what keeps it a specification rather than a fake.** `M26.3` builds the resolver [RESPONSE-SPACE.md](RESPONSE-SPACE.md) was written for: it answers the seam's calls itself, choosing a point in that space from a seed. Every freedom cites the `RS-P-n` permitting it and every restriction cites the `RS-C-n` requiring it, so a clause no code cites is visibly unimplemented and a behaviour citing no clause is the resolver inventing a platform. `ResolverConfig` therefore has a switch per permission and **none for any constraint**: narrowing a freedom is how a test isolates another, while a knob relaxing a constraint would let a test quietly assert against a platform that cannot exist. The default is the **widest** point, so an unconsidered test fails loudly under a freedom it did not handle rather than passing while exercising nothing. **Where a permission and a constraint collide, the constraint wins** -- `RS-C-4` holds a drain-flagged operation back even when `RS-P-1` chose to complete it now, and `RS-C-1` forces a post an operation's coin kept deferring. **Two mechanisms exist only to make `RS-C-1` finite** (a per-operation deferral bound and a bound on consecutive declined submits); both are properties of the resolver rather than of the space, which carries no rates deliberately. **A submit resolves and a pop only rescues**, because a pop that flipped coins would make `RS-P-1`'s pending case unobservable to a polling consumer -- the freedom would be implemented and untestable. `SetIoRingCompletionEvent` moved behind the seam for this item: it is how a completion becomes *observable*, so `RS-P-6` is a clause about it, and a resolver unable to make that call does not satisfy the clause vacuously but hangs every [`EventDelivery`](src/event_delivery.rs) consumer instead. **The resolver's first contact with a real ring found a live defect**, queued as `M26.8`: a submit declined under `RS-P-7` propagates out of `IoRing::run_down` leaving work outstanding, which is `M21.6`'s hazard reachable through a different `HRESULT`. |
+| D-62 | **The properties that must hold under every resolution are checked by *asking* [`RingContract`](src/contract.rs), not by restating it -- and the same rule decides where the expected outstanding count comes from.** `M26.4` states five properties over the resolver `M26.3` built: conservation, no hang, `pop_within` honours its bound, `outstanding` is accurate, and no use-after-free. Only two needed anything new. Conservation is already this crate's own oracle, so the harness reports to it and asks it for the verdict, on the rule that the layer owning an invariant owns the oracle for it -- a copy in a harness is a second implementation, and when the two disagree it is the harness that gets "fixed". **The accurate-`outstanding` property follows the same rule rather than counting for itself**: the expected value is read back out of the contract through its own `Outstanding` violation, because a counter in the harness would be a third party to the disagreement. **"No hang" is a step budget**, since a resolver answers instantly and a hang here is therefore an unterminating loop rather than a block. **Two weaknesses are declared rather than papered over.** `pop_within`'s upper bound is nearly free under an ordinary resolution -- the resolver rarely makes it approach its deadline -- so the non-vacuous case uses a separate degenerate responder in which nothing completes during the window; that is not an `RS-C-1` violation, because no finite observation can distinguish "eventually" from "never", but the prefix of a satisfying resolution in which the eventually has not happened yet. And **no-use-after-free covers this crate's memory handling, not the kernel's**, since under a resolver nothing external ever writes into a buffer; the kernel-side half stays with [generated_sequences.rs](tests/generated_sequences.rs) against a real ring. **Coverage counters are asserted, not printed**: all five properties are satisfied trivially by a run that does nothing, so a vacuity guard is what separates "the properties held" from "nothing reached the states they are about". |
+| D-63 | **The resolver is calibrated against two defects that really happened, and the two suites' sensitivities were measured rather than assumed -- they differ, and the difference is the reason the calibration is its own file.** [D-41](#d-41)'s corollary is the rule: a green result from an instrument nobody has shown can go red is not evidence, and this repository has already produced one instrument of exactly that shape -- `M17.4` reverted issue #47 as it shipped and the generated suite reported **green**, because it sampled the right state while draining by a poll that recovers completions whether the ring signalled or not. **The two defects are calibrated differently because they live in different places.** `M21.6`'s -- an expired wait treated as a failure -- was in this crate, so it is re-injected by [sabotage.json](sabotage.json) and swept; what [calibration.rs](tests/calibration.rs) adds is the precondition that makes that sabotage mean anything, namely that `RS-P-4` reaches `pop_within` at all. [D-47](#d-47)'s is in a **consumer**, so there is nothing to mutate and the defective consumer is written out: a believer in the hold-back `D-24` claimed and `D-47` withdrew, which the resolver must break. **Measured, and the two suites are not equivalent.** `M21.6`'s defect is caught by both the calibration and `M26.4`'s property suite -- the latter only because that suite distinguishes a declined submit from a genuine error, which makes the detection deliberate rather than lucky. A *narrowed* resolver -- one enforcing the hold-back -- is caught by the calibration alone, and the property suite correctly stays green, because a narrower resolution is still a valid one and conservation still holds under it. **That is why a calibration file exists rather than a calibration assertion inside the property suite**: only a test that demands sensitivity can detect an instrument going blind, and such a test fails when the *instrument* regresses rather than when the crate does. |
+| D-64 | **The standard seeded-sweep size is 2048, and a coverage threshold must be stated against the space a test can reach rather than against the sweep size.** The sweeps were widened from 64 (and the property suite's 240 plans) for breadth, on the measurement that seed count is not what these suites cost: a seed is tens of microseconds -- a whole ring, a batch, a drain and a rundown -- so the resolver's 22 unit tests sweep 2048 seeds each inside a lib suite that runs in well under a tenth of a second, far inside this repository's sub-second budget for a submodule. Only [properties_under_every_resolution.rs](tests/properties_under_every_resolution.rs) moved materially, to a little over a second, and even there roughly an eighth of the original cost was a deliberate sleep in the bound test rather than the plans. **The trap the widening exposed is the part worth keeping.** `different_seeds_reach_different_resolutions` asserted that distinct completion orders exceeded *half the seed count*, which is satisfiable only while the seeds are fewer than the outcomes: six operations admit `6! = 720` orders, so that threshold becomes arithmetically impossible past 1440 seeds and the test would have failed with nothing regressed. Thresholds of that kind are now phrased against the achievable space. **Saturation was measured rather than assumed**, because the obvious response -- cap the sweep where it stops gaining -- turned out not to apply: over the resolver's own mixer, 1024 seeds reach 539 of the 720 orders and 2048 reach 670, so the sweep is still gaining breadth at its current size and only around 8192 exhausts the space. **The property suite's vacuity guards are fractions of the plan count** for the same reason in reverse: an absolute floor chosen for 240 plans is a twelvefold margin at 2048, which would let the suite lose most of its work without complaint. Each is set near half the minimum observed over repeated runs, since that file's seeds are clock-derived and its counts therefore vary; the fixed-seed suites are deterministic and need no such margin. |
+| D-65 | **The kernel tests' job is confirming a real Windows stays *inside* the specified space, and that division of labour is enforced by a census that reads markers rather than mentions.** `M26` split one job in two: the resolver sweeps the permissions (`RS-P-n`), and tests against a real ring confirm the constraints (`RS-C-n`). The split is load-bearing for exactly one clause -- `RS-C-4`, which the resolver is forbidden to violate, so a Windows that broke the drain would be caught by nothing the resolver does. **A hole was found where that mattered most**: [flush_barrier.rs](tests/flush_barrier.rs) asserted the clause but sat behind an early return taken when its control could not discriminate, so on such a machine the constraint was untested on both sides at once. The fix separates two questions that one gate had been answering together -- *is the barrier doing work*, a comparative claim that genuinely needs the control, and *did the kernel stay inside `RS-C-4`*, a conformance question where skipping can only hide a violation. The conformance assertion now runs on every machine and only the comparative claim is withheld. **The census was green and useless on its first build, and trying to make it go red is what found that.** It searched each file for the clause ID anywhere in its text; removing `RS-C-4`'s check from the only test performing it did not turn it red, because a second file mentioned the clause only to *disclaim* it -- and under a substring search a disclaimer is indistinguishable from a claim. This is the same trap already recorded for a probe whose only mention of a tag was a comment. A claim is now a structured `CONFIRMS:` / `EXERCISES:` marker naming a clause and nothing else, verified to go red in three directions. **Two blind spots are declared rather than hidden**: a census over source proves a clause is *claimed*, never that the file's assertion still runs or still means anything; and dropping the covering flag does **not** fire the `RS-C-4` assertion on every machine, because a device stack that orders a flush behind its file's outstanding writes by itself produces no violation to see -- so that sabotage is a portable demonstration of nothing, and the assertion's conformance value (reporting a violation if one occurs) is separate from its sensitivity (proving it would notice). **The toolkit's "five techniques" framing becomes six**, and the sixth is different in kind: the other five check this crate against its own stated contract and cannot tell you that contract is wrong, whereas the resolver checks it against a written specification of what the platform may do, and the kernel tests check the platform against that same specification. |
+| D-66 | **The suite's frozen observations were one class, it was 31 assertions wide, and the restatement is guarded by a source census because no run objects to it.** `M26.7` audited the suite for assertions that pin the platform's incidental behaviour rather than this crate's contract -- the shape [D-52](#d-52) demonstrated, where one assertion gave *opposite answers on two handles of the same API*. **The census came from a command**, as the item required, and the answer was not the file anyone guessed: 31 sites across five kernel tests read `try_pop()` straight after `submit_and_wait` and unwrapped the `Option`, asserting the kernel had **already** queued the completion. `pop_within`'s own documentation denies that in this crate's words -- a submit-side wait's return "promises nothing about poppability" -- and [RESPONSE-SPACE.md](RESPONSE-SPACE.md) states it as `RS-P-5`. They passed for the reason [D-40](#d-40) measured: a buffered read completes inside the submit in 80 of 80 attempts, while an unbuffered one genuinely pends. **The restatement is the one `D-52` prescribed** -- ask for the completion within a bound this crate chooses, which holds on every handle -- and it also removed an unbounded `try_pop` spin found in the same sweep. **Nothing catches a regression by running**, which is the part worth keeping: reverting a site leaves its test green on any machine where the observation is true, so the guard is a census that refuses the shape at the source, plus a resolver-driven test that makes the pending case reachable on demand and shows `try_pop` failing where `pop_within` succeeds. **A second, narrower class was found and deliberately not settled**: five assertions require a complete transfer, which the space lists as *deliberately undecided* and `Completion::result` never promises. That is a specification gap the suite silently answers, queued as `M26.10` rather than decided here, since the space's own instruction is to decide it before a resolver relies on either answer. One of the five is already in the honest form -- [flush_barrier.rs](tests/flush_barrier.rs) checks the transfer as its own precondition and says so -- and is left alone. |
+| D-67 | **This crate does not implement retry policy. It supplies bounded primitives and reports what the platform documented, and the caller owns backoff, attempt counts and when to give up.** `M26.8` began as "what should `run_down` do when its submit is refused" and was settled by reading `SubmitIoRing`'s reference page rather than by measuring, which is the correction worth keeping: **no amount of measurement constitutes a contract.** The documentation gives a return-value row for `IORING_E_WAIT_TIMEOUT` -- *"All operations were submitted without error and the subsequent wait timed out"* -- and Remarks stating *"If this function returns an error other than IORING_E_WAIT_TIMEOUT, then all entries remain in the submission queue"*, plus that a per-entry failure arrives as a completion rather than as a submit failure. Three things follow. **`Batch::submit_and_wait` had `M21.6`'s defect** at the one site that sweep did not reach: `do_submit` passed a timed-out wait to `check`, reporting a fully successful submission as an error -- and because any *other* error means the entries are still queued, an `Err` was ambiguous between "your buffers are free" and "the kernel still owns them", which is [D-5](#d-5)'s hazard with the sign hidden. **[`IoRing::run_down`](src/ring.rs) was the only waiting API in this crate shaped wrongly**: it waited in 50 ms segments with no period to sit inside, which made "how long to keep trying" this crate's policy rather than its caller's. [`IoRing::run_down_within`](src/ring.rs) is the primitive -- the caller supplies the bound, `Ok(true)` means safe to drop, `Ok(false)` means call again -- and `run_down` is now that with an unbounded period, which is a choice a caller makes by calling it. Note the contrast that makes this precise: [`IoRing::pop_within`](src/ring.rs) also waits in segments and is *correct*, because its segments sit inside a deadline the caller supplied. **An error from rundown is not terminal and says so**: the entries remain queued, the ring is resumable, and the one thing a caller must not do is drop it. **`RS-P-7` and `RS-P-3` moved from inference to citation** -- `M26.1` had written `RS-P-7` as a consequence clause with no authority for submits failing at all, and `M26.3`'s resolver had read it as a permission anyway; the space now carries a `Documented` tag that outranks `Observed`, because a measurement describes one run of one build. |
+| D-68 | **A thread-pool wait must be armed before the event it watches is signalled, so the ring's completion event is attached unsignalled and the setup signal is raised after arming.** `M26.9`'s intermittent stall was settled the same way [D-67](#d-67) was -- by reading the reference page rather than by measuring. `SetThreadpoolWait`'s Remarks state *"You must re-register the event with the wait object before signaling it each time to trigger the wait callback."* [`IoRing::completion_event`](src/ring.rs) raises its deliberate setup signal as it attaches, and `EventDelivery::new` then built a `ThreadpoolWait` around the returned duplicate and armed it -- signal first, register second, which is the order the sentence forbids. **Three properties compounded to make a dropped signal permanent rather than late.** The event is auto-reset ([D-21](#d-21)), so a signal is consumed rather than left pending for a later arming to observe. It is edge-triggered on the completion queue going empty to non-empty ([D-19](#d-19)), so a ring whose queue is already non-empty is signalled by nothing else. And the setup signal exists precisely to serve the already-non-empty case, so it is the only wakeup such a ring will ever get. The fix separates the two steps: `attach_completion_event_unsignalled` attaches and reports whether the signal is still owed, `raise_setup_signal` raises it, and `EventDelivery::new` arms in between; `completion_event` is the two composed, unchanged, for a caller doing its own waiting. **The ordering was the defect, so [D-21](#d-21) stands.** A manual-reset event would also have survived the wrong order, by not consuming the signal, but that trades the arming rule for a reset the drain has to get right, and nothing measured here argued for reopening a decision whose own rationale is about the drain. **Measured before and after** on the reproducer recorded in [RESOLVED-TEST-FAILURES.md](RESOLVED-TEST-FAILURES.md): 5 failures in 600 and 2 in 600 for the two co-running triggers, against 0 in 3600 after. The rule itself is now stated where a caller meets it -- on `ThreadpoolWait::arm` and `WaitActivation::rearm` in [windows-threadpool-sys](../windows-threadpool-sys/src/wait.rs), on `IoRing::completion_event` with a worked remedy for a caller who holds the handle, and on `EventDelivery::new` -- because every other arming site in this workspace already had the order right, so what was missing was the statement rather than the practice. |
+| D-69 | **A completion may report fewer bytes than requested, and the permission is open because this crate does not constrain the handle type. A consumer that narrows its handle type earns a stronger guarantee, and states it in its own contract.** `M26.10` asked whether a short transfer is possible; the answer has two halves that belong to different layers, and collapsing them is what the question had been deferred over. **The general ring permits it** (`RS-P-8`). `WriteFile`'s Remarks state that *"when writing to a non-blocking, byte-mode pipe handle with insufficient buffer space, WriteFile returns TRUE with \*lpNumberOfBytesWritten < nNumberOfBytesToWrite"*; sockets report a short send against a full transmit buffer, and a communications handle with a write timeout can report a partial count. So the clause is `Documented` rather than `Observed`, which matters because nothing in this repository had measured it and `M26.7` had flagged five assertions quietly depending on the opposite. The reason the permission cannot be narrowed is not that files behave badly -- for an ordinary file on a local volume a successful completion is expected to carry the full length, and a full volume is `ERROR_DISK_FULL` rather than a short success -- but that **`IoRing` takes a handle and never asks what kind it is**. A pipe, a socket and a serial port are all handles, so a consumer of *this* crate has to read the count. **Continuation stays with the caller**, per [D-67](#d-67): the count is reported and nothing reissues the remainder, because how many times to retry and when to give up are policy. A short count may be zero, so a consumer that loops must tolerate making no progress. **The durability layer is the other half.** [examples/epoch_log](examples/epoch_log) makes guarantees -- epoch commits, flush barriers, FUA -- that are only meaningful on a real file on a real volume: a flush barrier means nothing on a socket, and a short write would break "the whole record landed" silently rather than loudly. It therefore has to *constrain* the handle types it accepts, and thereby earn the completeness its accounting already assumes, rather than inherit it from a ring that does not promise it. That constraint is not yet enforced; it is queued as `M26.11` rather than recorded here, because a decision is not a work queue. **The kernel tests were corrected in the honest direction** -- three sites that asserted a full count bare now say that completeness is a property of the temp file they opened, which is the form [flush_barrier.rs](tests/flush_barrier.rs) already used and `M26.7` had singled out as right. |
+| D-70 | **The epoch log states the capabilities it requires of a handle and does not check for them. A pre-flight check points the wrong way: it catches the failures that were already loud and misses every failure that is silent.** [D-69](#d-69) left the durability layer owing a handle constraint, and the obvious discharge was a gate -- `GetFileType`, then `GetDriveType` or `FileRemoteProtocolInfo` for the distinctions the first is too coarse to make. Working the gate out is what showed it was not worth having. **Sort the failures by how they present.** A console handle, a pipe, a socket or a closed handle is exactly what `GetFileType` names -- and every one of them already fails at the first positioned write, so the check buys a clearer message and nothing else. A RAM disk, a remote share, or a volume whose write cache is not power-protected **succeeds at every API call this log makes** and silently fails to be durable; `GetDriveType` can name the first two and nothing in Win32 settles the third from a handle, because write-cache state is a property of the device rather than of the handle. So the gate guards the loud cases and misses the silent ones, which is the inverse of what a durability layer needs. **The cost is not the call, it is the claim.** A check that cannot establish the property still reads to a later maintainer as though the property was established -- the "decoration that reads like enforcement" the repository's FAIL FAST rules name -- and that is worse than no check, because it discourages the reader from asking the question themselves. **So the requirement is documented and the caller warrants it.** This is [D-67](#d-67)'s shape applied to handles rather than to retries: we state what we need, the caller chooses what to hand us, and the platform enforces what it is able to. Two requirements are singled out in the contract: that a successful write of `N` bytes transfers `N`, which this log *can* check and does, and that a completed flush reaches stable media, which it cannot check and says so in those words. **Two alternatives were considered and declined**, both recorded so they are not re-proposed as new: refusing `FILE_TYPE_CHAR` and `FILE_TYPE_UNKNOWN` as a cheap early error, declined because it improves only the already-loud path; and a caller-declared mask of intended handle types, whose value would be the acknowledgment rather than the validation, declined as ceremony every ordinary caller pays for. The work of writing the contract is queued as `M26.11`, since a decision is not a work queue. |
+
## Durability on the ring
Written for consumers, like the two sections that follow it, and for the same reason: the default
@@ -583,7 +618,7 @@ Two consequences follow, and they are why this matters beyond terminology:
cache and NUMA locality that motivated the whole structure comes from the pinning, not from the
per-thread split. This is a configuration people ship by accident.
- **The interesting count is cores, or LLC domains -- not threads.** Which is exactly what
- [D-8](#d-8) and the L3-domain guidance below already recommend; this is the reason underneath them.
+ [D-8](#d-8) and the cache-domain guidance below already recommend; this is the reason underneath them.
### The two models are Windows' own two completion mechanisms
@@ -631,11 +666,43 @@ It is worse in virtualized deployments, which is where most of this code will ru
investigated on reported **zero** `Win32_NumaNode` instances. Any strategy keyed on node must degrade to
"one ring" when the answer is unknowable, which is the common case.
-A better default heuristic is the **last-level cache domain**: `GetLogicalProcessorInformationEx` with
-`RelationCache` filtered to `CacheLevel == 3`. On EPYC that is the CCX/CCD boundary, which has a real
-latency cliff even inside a single NPS1 node, because crossing it goes out to the IO die over Infinity
-Fabric. It is meaningful on Intel and ARM too, where the NUMA node often is not, and it degrades sanely: a
-VM reporting one L3 domain yields one ring, which is correct.
+**And it is not only virtualization.** A shipping ARM consumer laptop reports zero `Win32_NumaNode`
+instances too, and reports no L3 cache domains at all -- see [D-48](#d-48). The two observations are
+siblings, and together they describe the machines most consumers actually have: one where the node is
+absent because a hypervisor did not present it, one where it is absent on bare metal.
+
+A better default heuristic is the **outermost cache level that actually partitions the machine**, which
+`MachineMemoryTopology::outermost_partitioning_cache` answers -- defined in
+[windows-topology-sys](../windows-topology-sys/README.md), so that every consumer asks rather than
+restating. On EPYC that is the L3/CCX boundary,
+which has a real latency cliff even inside a single NPS1 node, because crossing it goes out to the IO die
+over Infinity Fabric. It degrades sanely: a VM whose caches partition nothing yields one ring, which is
+correct.
+
+**The rule is not "L3", and the level number is not the ordering.** Two measurements forced that wording,
+and each falsifies a different half of the old one:
+
+- A shipping Snapdragon X2 Elite reports **no L3 at all**, with its natural cluster boundary at L2
+ ([D-48](#d-48)). So "the last-level cache is L3" is false on a part that ships today.
+- The machine this workspace is developed on reports an L3 that spans **all 16 processors** over a real
+ 8-way L2 partition. So an L3 can exist and still partition nothing -- and a consumer filtering on
+ `level == 3` there does not degrade, because it matched something. It reports one whole-machine domain
+ as a successful cache-aware partition.
+
+That second case is why the rule is stated as a question to ask rather than a level to match: the failure
+is silent, and it collapses an eight-domain machine to a single ring while reporting success. The finding
+that a cache domain beats the NUMA node is untouched by either.
+
+**Choosing a specific level is still available; it is just not the default.** What was withdrawn is a
+*policy* that matched on a level number, because that policy computed the wrong partition on two of the
+three machines above. The underlying capability is untouched: `MachineMemoryTopology::cache_levels` and
+`cache_partitions_at_level` let a consumer who knows their part ask about any level directly, and
+`examples/cache_domains.rs` prints every level's distinct processor sets beside the heuristic's choice,
+so the comparison the old policy got wrong is visible rather than asserted. On the development host that
+output is `L1: 8`, `L2: 8` (chosen, checked pairwise disjoint), `L3: 1` -- a consumer can see in one
+glance why matching `level == 3` there collapses the machine, and equally that on a part where L3 does
+partition, choosing it is theirs to make. A consumer is given the data and the means to decide; what they
+are not given is a preset that answers wrongly without saying so.
**Processor groups are a hard floor.** A thread's affinity is a `GROUP_AFFINITY` and a ring's waiter lives
in exactly one group, so above 64 logical processors the partition is forced whether or not it is wanted.
@@ -650,14 +717,47 @@ So `VirtualAllocExNuma` for the pool, on the node closest to the device, registe
ring, is very likely the highest-leverage locality decision available -- and it is independent of
everything above about completion routing.
+That allocation is `NumaBuffer`, which this crate provides as of [D-51](#d-51). *Which* node is still the
+caller's answer, per [D-8](#d-8); the type supplies the allocation and decides no policy.
+
### What is not reachable
-Mapping a **file handle to the NUMA node of the device backing it** has no clean user-mode path. It means
-walking volume to disk to device instance and reading `DEVPKEY_Device_Numa_Node`, with real failure modes
-(spanned volumes, Storage Spaces, network paths, VHDs) where the question may have no answer. This crate
+What is not reachable is the **answer** -- "which ring should this file's I/O go to". The **mechanism** is
+reachable, and an earlier version of this section had that wrong: it said the mapping has "no clean
+user-mode path" and "means walking volume to disk to device instance and reading
+`DEVPKEY_Device_Numa_Node`". There is one documented call, and it takes the handle a caller already holds.
+
+- **`FSCTL_QUERY_VOLUME_NUMA_INFO`** is documented in the IFS docs, accepts a handle to a **file or
+ directory** directly, and returns `FSCTL_QUERY_VOLUME_NUMA_INFO_OUTPUT { ULONG NumaNode }`. No device-tree
+ walk.
+- **`GetNumaNodeNumberFromHandle`** is the other path: a Win32 wrapper over `NtQueryInformationFile` with
+ `FileNumaNodeInformation` (class 53, Windows 7 and later), yielding
+ `FILE_NUMA_NODE_INFORMATION { USHORT NodeNumber }`. PHNT and the WDK mark that class **reserved for
+ system use**, so this crate must not build on it. It is named here so the next reader does not rediscover
+ it and assume it is available.
+
+Both were observed to succeed on an ordinary NTFS data file, and on a directory handle, and to agree --
+so an ordinary file is not the no-association case. That run was on a single-node host, so neither call is
+shown to name a node that distinguishes anything; the run and its limits are recorded in the spikes
+[README.md](design-sessions/spikes/README.md), and
+[file-handle-numa-spike.rs](design-sessions/spikes/file-handle-numa-spike.rs) is the instrument. Settling
+what a multi-node host reports needs storage whose PDO advertises a proximity domain, which is a hardware
+gap rather than a deferred decision.
+
+**The conclusion this section has always drawn survives, on different grounds than it used to rest on.**
+What either call returns is the node the *volume* resides on, not where the file's extents live. It is
+absent whenever the device layer advertised no proximity domain -- `IoGetDeviceNumaNode` on the PDO, or
+`DEVPKEY_Numa_Proximity_Domain` with `GetNumaProximityNode` from user mode. And one volume may sit on
+several devices, which is the ordinary case for a spanned volume or a Storage Spaces set. So this crate
will not offer an automatic "put this file's I/O on the right ring." It offers "bind a ring to a domain and
submit from there," and leaves the mapping to whoever knows their storage layout.
+**A sample now exercises the mechanism, which is not the same as the crate adopting it.** The epoch-log
+sample asks `FSCTL_QUERY_VOLUME_NUMA_INFO` of its own log handle and places its arena on whatever comes
+back ([D-50](#d-50)), so the call above has a runnable consumer rather than living only in a spike. That is
+a sample making a local policy choice with the caveats in this section attached to it; the library's
+position is unchanged.
+
### The practical shape
Almost nobody runs pure Model B. What works is hybrid: Model B on the hot data path (pinned threads,
@@ -667,13 +767,13 @@ paths, where the thread pool's quiescence is worth more than locality.
Both paths are therefore first-class in this crate, which is what D-3 records.
-On sizing: one domain per physical core (not per SMT sibling) maximizes isolation; one per L3 domain gives
+On sizing: one domain per physical core (not per SMT sibling) maximizes isolation; one per cache domain gives
a smaller number of domains that can still share cache-resident state cheaply -- eight rather than
sixty-four on a 64-core EPYC. Fewer domains balance load better and duplicate registered buffers less; more
isolate better. That is a workload call, and this crate does not make it.
`examples/ring_copy` (M7) is where that workload call actually gets made, for exactly one workload: it
-implements the `ByL3`/`ByNode`/`ByPackage`/`ByCore`/`Single` policies above as runnable code, over a real
+implements the `ByCache`/`ByNode`/`ByPackage`/`ByCore`/`Single` policies above as runnable code, over a real
file copy, so the guidance here has something executable behind it rather than staying prose. The policy
lives in the sample, not the library (D-8); the library still makes none of these choices for a caller.
@@ -731,6 +831,7 @@ evidence rather than left as silence.
| `PendingBufferRegistration::claim_if` | `io::Result>` | Take the registration, or observe the failure. | **Sound, with a documented sharp edge.** A *failed* completion is treated as proof the kernel did not retain the addresses, so the buffers are dropped. M16.4 recorded that injecting a synthetic failure here would therefore free memory the kernel genuinely holds -- inert only because nothing does so outside the fault-injection seam. |
| `PendingFileRegistration::claim_if` | `io::Result` | Take the registration. | **No hole.** `BuildIoRingRegisterFileHandles` reads its array synchronously ([D-32](#d-32)), so nothing outlives the call that could be invalidated. |
| `EventDelivery::scope` (was `ring`) | `RingScope` | Submit work through [`RingScope::batch`], read the ring's read-only state. **Cannot** obtain a `&mut IoRing`, and so cannot replace the ring. | **WAS THE ONE FINDING -- [D-43](#d-43), fixed in M18.6.** The previous `ring -> &Mutex` let safe code assign a whole new ring through the guard, which compiled and silently stopped delivery: measured at one completion delivered before the swap and none after. [D-35](#d-35)'s shape at a different layer, and fixed the same way -- by narrowing the returned type to exactly what the caller needs. Now enforced by a `compile_fail` doctest. |
+| `Pending::contract` | `Option<&RingContract>` | Read the oracle's record of what it has observed -- counts and per-operation states. **Cannot** mutate it, and cannot obtain one at all from an unchecked map, which is what the `Option` reports. | **No hole.** `RingContract` is a pure record: it owns no handle, no buffer, and no index into a registration, so there is nothing here the kernel or a registration could invalidate. The borrow is a plain `&self` borrow of the `Pending`, so nothing can be submitted through that `Pending` while it is alive -- every push takes `&mut self`. Added in M23.3 and audited in M26.2, which is late: the gate reported it on the next run rather than on the change that introduced it, because that change did not run the gate. |
| `IoRing::completion_event` | `OwnedHandle` | Wait on it, close it, hand it elsewhere. | **No hole.** The returned handle is a *duplicate*; the ring keeps its own, so closing the caller's does not stop the ring signalling ([D-20](#d-20)). Repeat calls duplicate the same event rather than attaching a second. The one real hazard -- two waiters on one ring -- is a documented misuse, not a memory-safety hole. |
| `IoRing::try_pop` | `Option` | Read `user_data`, `code`, `result`; use it to claim a token. | **No hole.** `Completion` is a plain value carrying no borrow of ring state. Its power is that it authorises a claim, and that power is bounded by the id/ring checks in `claim_if`. |
| `Completion::with_injected_failure` | `Completion` | Rewrite the result of a *real* completion. | **Sound because it transforms rather than fabricates.** Same `user_data` and `ring_id`, so the "a completion exists therefore the kernel is done" argument is untouched. `Completion::synthetic`, which *would* fabricate, is `#[cfg(test)] pub(crate)` for exactly this reason. Behind the off-by-default `fault-injection` feature. |
@@ -747,6 +848,35 @@ That is the pattern worth carrying forward into M18.2's recurring rule -- the
question that finds these is not "is this correct?" but "what else does this
type allow?"
+### Four more items, surfaced by widening the check (M21+.1)
+
+The audit above is a point-in-time pass, and nineteen was its count. These four
+are not corrections to it; they are items the *check* could not see, and so
+never put to anyone. A 2026-09-21 review of the `M21.2` surface found that
+[check-borrow-surface.ps1](../../tools/check-borrow-surface.ps1) inspected only
+the text after the last `->` on lines matching `pub fn`, which leaves two shapes
+invisible: **a method of a `pub trait`** (declared `fn`, not `pub fn`) and **a
+borrow-carrying type in parameter position**. Widening it reported exactly four
+entries, answered here before the inventory was regenerated.
+
+Worth stating plainly: the check was not wrong about what it covered, and the
+control still passes -- a `pub fn` returning `&[u8]` was caught throughout. It
+was narrow, and nothing said so.
+
+| Item | Shape the old check missed | What safe code may do with it | Finding |
+|---|---|---|---|
+| `IoRingErrorExt::as_ioring_error` | trait method, borrow **returned** | Read `code`, `name`, `condition` through a shared reference. | **No hole, and it is the honest catch of the four.** This predates the widening by months and was simply never inventoried, which is the blind spot made concrete rather than a new risk. The borrow is of the `io::Error` the caller already owns -- not of ring state, not of anything the kernel holds -- and `&IoRingError` permits reads only. |
+| `CompletionWait::wait` | trait method, borrow **parameter** | Call `RingWait::block` and `RingWait::outstanding`, and nothing else. | **No hole, and the narrowing is the reason.** This is the wider exposure of the two directions: the wrapper goes to arbitrary safe code implementing the trait, not to a known caller. `RingWait` exposes no pop -- one would consume the completion its own caller is waiting for -- and no way to build work. It cannot be retained: the `'ring` lifetime is fresh per call and unconstrained by `Self`. Nothing else can touch the `IoRing` while it is alive, because the pop loop holds `&mut self` across the call. Handing out a bare `&mut IoRing` here would have been [D-43](#d-43) again. |
+| `Batch::new` | borrow **parameter** with an explicit lifetime | Nothing it could not already do: the caller supplied the `&mut IoRing`. | **No hole; this is [D-5](#d-5)'s mechanism, not a leak of one.** The exclusive borrow is what makes two concurrent batches fail to compile, which is the point of taking it. The borrow travels *into* the crate and is released when the `Batch` drops. |
+| `EventDelivery::new` | borrow **parameter** with an explicit lifetime | Nothing; the callback environment is forwarded to `ThreadpoolWait::new` and not retained. | **No hole.** `EventDelivery` stores only `wait` and `ring`, so the `&mut CallbackEnviron<'_>` does not outlive the call. |
+
+**A plain `&T` parameter is deliberately not reported**, or the inventory would
+list every method in the crate and say nothing. Lending a reference *to* a
+callee is the caller's business; what this defect class is about is a
+borrow-carrying wrapper whose lifetime the crate chose. The explicit-lifetime
+test is what separates the two, and it is a heuristic -- it would miss a
+hypothetical `&dyn Trait` parameter carrying no named lifetime.
+
## Testing strategy (M18.5)
Eight defects came out of the 0.1.x line and the M11-M14 branch. M15 through
@@ -760,7 +890,7 @@ them reach at all.
| | The defects | What finds them | Built in |
|---|---|---|---|
-| **A -- preconditions never varied** | [#47](https://github.com/MikeGrier/windows-threadpool-sys/issues/47): every `event_delivery` test handed over a *fresh* ring, so "completion queue non-empty at handover" was never a test input | Generated operation sequences, so the state space is sampled rather than enumerated by hand | M17 |
+| **A -- preconditions never varied** | [#47](https://github.com/MikeGrier/windows-threadpool-sys/issues/47): every `event_delivery` test handed over a *fresh* ring, so "completion queue non-empty at handover" was never a test input | Generated operation sequences, so the state space is sampled rather than enumerated by hand. **`M26` added a second technique for this population**, on the observation that *the kernel's response* is a precondition too and was never varied either: a seeded resolver over a specified space | M17, M26 |
| **B -- failure paths never taken** | The checkpoint path authorising a reclaim after a failed write; [#48](https://github.com/MikeGrier/windows-threadpool-sys/issues/48) surfacing as a *lucky* `ERROR_NOACCESS` rather than corruption | Deterministic memory instrumentation, an executable contract oracle, and a seam that injects failure into a real completion | M15, M16 |
| **C -- permissions, not behaviour** | [D-35](#d-35) (`&mut Vec` permits `reserve`), [D-36](#d-36) (`&B` handed out while the kernel writes), and [D-43](#d-43), found by the audit itself | Review, made recurring by a mechanical trigger; mutation testing for the weaker cousin of the same problem | M18 |
@@ -782,6 +912,7 @@ hoped-for ones.
| Generated sequences (M17) | **No new defects.** Found an API precondition the generator was violating, and -- once calibrated -- rediscovers #47 in 10 of 10 runs |
| Borrow-surface audit (M18.1) | **One defect: [D-43](#d-43)**, in 19 items audited. Same shape as the two that prompted the audit |
| Mutation testing (M18.3/4/7) | **A third vacuous test** four review rounds had read past, plus 36 further weak assertions. 79.7% to 95.8% |
+| Seeded resolution over a specified space (M26) | **No defect in shipping code yet**, and the calibration is why that is reportable rather than reassuring: re-injecting `M21.6`'s expired-wait defect turns it red, and so does narrowing the resolver itself. It did find one live defect on its first contact with a real ring -- a declined submit propagating out of `IoRing::run_down` with work still outstanding, queued as `M26.8` |
Two observations that only appear once the table is read as a whole.
@@ -803,13 +934,34 @@ evidence.** Budget the calibration, not just the instrument.
### What none of them cover
-All five techniques check this crate's code against **this crate's stated
-contract**. None of them can tell you the stated contract is wrong -- and in the
-two most expensive defects, that is exactly what happened. The completion event
-is edge-triggered ([D-19](#d-19)) and `BuildIoRingRegisterBuffers` reads its
-array when the operation runs rather than when `Build*` returns
-([D-32](#d-32)). Both were discovered by a spike against the real kernel, and
-neither could have come from anywhere else in this toolkit.
+**Five of the six check this crate's code against this crate's stated
+contract.** None of those five can tell you the stated contract is wrong -- and
+in the two most expensive defects, that is exactly what happened. The
+completion event is edge-triggered ([D-19](#d-19)) and
+`BuildIoRingRegisterBuffers` reads its array when the operation runs rather
+than when `Build*` returns ([D-32](#d-32)). Both were discovered by a spike
+against the real kernel, and neither could have come from anywhere else in that
+toolkit.
+
+**The sixth is different in kind, and `M26.6` is what makes the difference
+real.** The resolver tests this crate against a *written specification of what
+the platform may do* ([RESPONSE-SPACE.md](RESPONSE-SPACE.md)) rather than
+against our beliefs about what it does -- so it catches code that is brittle to
+platform variation inside the permitted space, which is a class the other five
+cannot reach. It still cannot tell you the specification is wrong. What can is
+the **other half of the pair**: the kernel tests now confirm that a real Windows
+stays *inside* the declared space, and a kernel observed outside it is a finding
+about the platform rather than a regression in this crate. `RS-C-4` is the
+clause that makes that job load-bearing rather than nominal, since the resolver
+is forbidden to break the drain half of `DRAIN_PRECEDING_OPS` and therefore
+cannot be what notices if Windows does. The division is enforced by
+[response_space_census.rs](tests/response_space_census.rs), which fails when a
+clause is claimed by nothing on the side that owes it a check.
+
+So the honest statement is narrower than "a spike is the only way to learn the
+platform is not what we assumed", and it is still true: a spike remains the only
+technique that runs **before there is any code to test**, and the space itself
+was written from what spikes established.
Be precise about the failure mode, because "the allocator would not have caught
D-32" is not quite true and the imprecision matters. The guard allocator *does*
@@ -832,7 +984,11 @@ their shape are in
they established is summarised under
[What the spike established](#what-the-spike-established).
-### Two techniques deliberately rejected
+### Two techniques deliberately rejected
+
+**The first was re-examined by [D-49](#d-49) / `M24.1` and now **stands, with its scope sharpened**
+by [D-52](#d-52). The re-examination did not find an escape; it found that the instrument was the
+wrong one.**
Recorded so they are not re-proposed as obvious wins.
@@ -843,6 +999,21 @@ merely have failed to find them, it would have manufactured evidence they were
absent. A model belongs here as an **oracle over observed sequences**
([`RingContract`](src/contract.rs)), never as a substitute for the kernel.
+> **What `M24.1` demonstrated, rather than argued** (2026-09-22, the apparatus is
+> [kernel-response-space-probe.rs](design-sessions/kernel-response-space-probe.rs)). Sharing a
+> suite's assertions between a fake and the kernel does **not** rescue a mock. A fake built from the
+> belief this crate held before `M21.6` -- that an expired wait is an error -- passed the shared
+> suite green; the assertion that catches it could only be written after the kernel had already
+> revealed the answer. Running the *wrong* assertion against both is the one case that helps: the
+> kernel refutes it while the fake confirms it, which is the manufactured-evidence mechanism made
+> visible.
+>
+> And the line is not "accounting versus Windows behaviour", which was the first answer. It is **our
+> specified contract versus the platform's incidental behaviour**. "After submitting, the completion
+> is already queued" reads like a contract and gave *opposite answers on two handles of the same
+> API*; the same question stated as this crate's own contract -- "arrives within a bound we specify"
+> -- holds everywhere. See [D-52](#d-52) for what replaces the technique.
+
**Application Verifier / PageHeap**, rejected after measuring rather than
assuming ([D-37](#d-37)): it works, and needs no SDK, but it is keyed by *image
file name* and cargo rehashes test binaries on every meaningful rebuild -- so
diff --git a/crates/windows-ioring-sys/PLANS.md b/crates/windows-ioring-sys/PLANS.md
index cd609067d..d81c77194 100644
--- a/crates/windows-ioring-sys/PLANS.md
+++ b/crates/windows-ioring-sys/PLANS.md
@@ -6,4 +6,4 @@ contained are archived in [COMPLETED-CHECKLIST.md](COMPLETED-CHECKLIST.md). Desi
| Path to CHECKLIST.md | Status | Brief description | Design Notes |
|---|---|---|---|
-| [CHECKLIST.md](CHECKLIST.md) | in progress | Memory-safe Rust over the Windows 11 / Server 2022 `IoRing` submission/completion ring, as a separate crate from `windows-overlapped-io-sys` (duplicate-then-decide). Covers ring lifecycle and capability negotiation, zero-allocation token-owned buffers, the batch submission builder, threadless delivery through `ThreadpoolWait`, file/buffer registration, consumer-facing documentation, and the `ring-copy` topology-aligned sample (M1-M7 archived). The pinned-thread (Model B) architecture remains parked as `M6+` by the engineer's explicit direction. `M8` (complete) closed a PR #20 review finding: `FileRef::Raw(HANDLE)`'s lifetime gap, fixed with `unsafe fn` raw entry points plus a safe, `Arc`-backed `SharedFile` wrapper for the common case. `M9` (complete) closed further PR #20 review findings: cross-ring `Token`/`RegisteredFile`/`RegisteredBuffers` confusion (a new per-ring `RingId`, checked at claim/push time), `PendingBufferRegistration` freeing its buffers instead of leaking them on an unclaimed drop, and `Batch::do_submit` letting `Drop` silently retry an already-attempted, already-failed submit. `M10` (active) finished auditing the ring completion contract against all ten specification-gap categories (M10.1-M10.3 complete): category 3 found that `supports` answers for the kernel's op table rather than this crate's push surface and that the registration one-shot is spent by queueing rather than succeeding (D-28); categories 1, 2, 6, 8 and 9 established the load-bearing rule that **every successfully queued SQE produces exactly one completion**, plus that this crate deliberately joins nothing (D-29, D-30); and D-14's registration-index continuity assumption was **dissolved rather than measured** -- the collision it guarded against needs a second registration, which was forbidden the day after D-14 was written, leaving only the reserved-not-confirmed meaning of the public counts to state (D-31). The audit also surfaced two API gaps, now queued as work rather than left in the design notes: `M10.4` (complete) gave `FileRef::Registered` safe entry points, since a registered index carries no lifetime obligation and the `unsafe` guarding it was vacuous (D-29): the safe pushes are now generic over a sealed `FileTarget` trait whose associated `Guard` type carries the one real difference between the two targets, which also made the fully-registered (registered file *and* registered buffer) combination expressible for the first time, non-breakingly (D-33). Investigating M10's own recorded test failures then found a **live use-after-free in shipped 0.1.2** and fixed it as `M10.6`: `BuildIoRingRegisterBuffers` reads its `IORING_BUFFER_INFO` array when the op runs rather than at build time -- the opposite of its file-handle sibling, which the rustdoc had wrongly generalized across -- so the array is now owned by the `IoRing` (D-32). `M10.5` (complete) added named predicates for the conditions a consumer must branch on -- `IORING_E_SUBMISSION_QUEUE_FULL` above all, which every push's rustdoc names as the backpressure signal but which `io::Error::kind()` cannot discriminate (D-30): a complete `RingCondition` enum, predicates for the runtime-actionable conditions, and a sealed `IoRingErrorExt` that puts them on `io::Error` so the downcast is named once rather than hand-rolled per call site (D-34). **M10 is complete.** **`M11` is complete and archived** (2026-08-28): it made the completion event a ring primitive, prompted by an external consumer proposal. `IoRing::completion_event` returns an owned duplicate of the ring's own event so a caller can wait on the ring alongside other handles without surrendering it (D-20); its contract is pinned by eleven sabotage-verified tests; `EventDelivery` is re-expressed on top of it, leaving one `SetIoRingCompletionEvent` call site; `windows-threadpool-sys` moved behind a default-on `threadpool` feature with CI building both combinations (D-22); the wakeup shapes and the barrier's ring-edge limit were swept across every place that states them; and `examples/model_b_multiplexed.rs` works the multiplexed shape end to end. The spike that answered the proposal established that the completion event is **edge-triggered** on the completion queue going empty to non-empty (D-19) -- which also exposed a live bug in shipped 0.1.2, where `EventDelivery` permanently stranded completions queued before handover, fixed in M11.3 by the same change that consolidated `EventDelivery` onto the new primitive. **`M12` is complete and archived** (2026-08-28): it addressed durability, from the same exchange. A spike had established that a flush **without** `DRAIN_PRECEDING_OPS` does not cover preceding writes (D-23) while the barrier that fixes it drains everything already outstanding on the ring (D-24, as corrected by D-47), making `Batch::flush` with default options a silent data-loss bug rather than a missing feature. `Batch::flush`/`flush_raw` now require an explicit `FlushCoverage`, so that spelling no longer exists; a `NO_BUFFERING` integration test proves the barrier's behaviour rather than its flag, and measured that *which direction* the reordering shows in is device-dependent (amending D-23); and the parameters the crate had hardcoded away are exposed as `WriteCaching` (`FILE_WRITE_FLAGS`) and `FlushMode` (`FILE_FLUSH_MODE`, whose `NoSync` is the one mode that commits nothing). Durability had been absent from `lib.rs` and `README.md` entirely, and both now state the three facts. **`M13` is complete and archived** (2026-08-29): the `epoch_log` sample is the vehicle D-26 makes for carrying durability *policy* to consumers without this crate owning it. Its own durability contract is written down first, in its own words (Design Autonomy), then implemented -- records composed into a registered arena and appended, epochs closed by one covering flush whose completion is what makes `is_durable` answer `true`, and a multiplexed wait on the ring's completion event alongside a shutdown latch. The replay pass is what turns it from a demonstration into evidence: it holds the durable region to a strict standard, tolerates a torn tail as the contract requires, and is itself proved able to fail by a negative control. Writing it also found and fixed a gap in the crate (D-35: per-buffer outstanding accounting and `RegisteredBuffers::get_mut`). **`M14` is complete and archived** (2026-08-29): the second half of the sample, covering the two things the ring cannot do for a consumer. A non-ring `FSCTL` (`FSCTL_SET_ZERO_DATA`, reclaiming a retired segment) is ordered against ring epochs by the log itself, since `drain_preceding` orders SQEs against SQEs (D-24) and reaches across neither the ring boundary nor a second ring; a thread-pool control plane runs checkpointing as Model A on a *second* ring while the log thread keeps Model B for the data path, because D-21 forbids a ring given to `EventDelivery` also being waited on directly -- so the ordering chain crosses log thread to pool thread to reclaim worker and back with the log thread blocking for none of it. All three epoch-commit strategies are implemented behind one interface, checked both by replay and by requiring the three to be **byte-identical**, and measured on the running machine. The measurement's finding is that the three are **indistinguishable** here -- the cross-strategy spread is the size of one strategy's run-to-run spread, because every strategy pays one device flush per epoch at hundreds of microseconds while their real differences land in the tens -- and it found two harness bugs nothing else caught, including one where a barrier benchmark that awaits each commit before appending again measures the barrier as free. Both findings are promoted into "Durability on the ring" in [DESIGN-NOTES.md](DESIGN-NOTES.md), since both are about the design rather than the demonstration. **`M15`-`M18` (not started) are the testing-strategy response to the eight defects the 0.1.x line and the M11-M14 branch produced.** They are organised by *defect population* rather than by technique, because the populations need different tools and one of them needs a tool that does not exist: (A) preconditions never varied -- every `event_delivery` test handed over a fresh ring, which is why [#47](https://github.com/MikeGrier/windows-threadpool-sys/issues/47) survived; (B) failure paths never taken -- no test ever ran `completion.result()` returning `Err`, which is why the checkpoint path could authorise a reclaim after a failed write; and (C) *permissions rather than behaviour* -- `&mut Vec` permits `reserve`/`resize`/reassign though no code path performs it, which is [D-35](DESIGN-NOTES.md#d-35) and [D-36](DESIGN-NOTES.md#d-36), the two most severe findings, and **no runtime technique reaches that population at all**. M15 and M16 gate the 0.2.0 release. M15 is deterministic memory instrumentation: a guard-page global allocator, chosen over Application Verifier / PageHeap by measurement rather than assumption ([D-37](DESIGN-NOTES.md#d-37) -- PageHeap works and `reg add` alone is enough, but IFEO is keyed by image file name and one test target produced six distinct hashed names in a day), plus a tracked poison pattern covering the gap guard pages structurally cannot see ([D-38](DESIGN-NOTES.md#d-38) -- a guard page catches access to memory that should not be touched, and is blind to the kernel writing into a live, valid buffer, which is exactly what `write_registered` and `read_registered` promise it will not do). M16 makes the contract executable: a public `RingContract` rendering [the category-2 rule](DESIGN-NOTES.md#one-sqe-one-completion), "one SQE, exactly one completion", checkable rather than merely stated, plus a fault-injection seam that finally takes the failure paths. M17 covers A by generating over the operation space instead of enumerating it by hand, gated on an explicit decision about randomized sampling that this component's conventions require be approved and recorded rather than assumed. M18 covered C, where review is a *primary* technique rather than a backstop, and added `cargo-mutants` against a measured rate of vacuous tests. **M8 through M18 are now complete and archived**; `M20` is pending, and `M6+` is parked rather than pending. The strategy as a whole -- which technique reaches which population, what each one actually found, and what none of them reach -- is recorded in [DESIGN-NOTES.md](DESIGN-NOTES.md#testing-strategy-m185); mutation coverage went from 79.7% to 95.8%, and the borrow-surface audit found one further defect of the same shape as the two that prompted it ([D-43](DESIGN-NOTES.md#d-43)). A mock `IoRing` was considered and **rejected**: both shipped defects were the kernel behaving differently from this crate's assumptions, so a mock would have encoded the same assumptions and passed both bugs green. **M20** queues the documentation and policy-test repairs from the 2026-08-30 NUMA-sharding measurement: a shipping ARM laptop reports no L3 cache domain at all, which falsifies the justification given for the last-level-cache heuristic (though not the heuristic's preference over the NUMA node, and not `ring_copy`, whose degraded fallback already handles it correctly). | [DESIGN-NOTES.md](DESIGN-NOTES.md), [DESIGN-SESSION-2026-08-28-completion-event-multiplexing.md](design-sessions/DESIGN-SESSION-2026-08-28-completion-event-multiplexing.md), [DESIGN-SESSION-2026-08-28-external-consumer-correspondence.md](design-sessions/DESIGN-SESSION-2026-08-28-external-consumer-correspondence.md), [DESIGN-SESSION-2026-08-30-numa-sharded-io-execution-domains.md](../../design-sessions/DESIGN-SESSION-2026-08-30-numa-sharded-io-execution-domains.md) |
+| [CHECKLIST.md](CHECKLIST.md) | in progress | Memory-safe Rust over the Windows 11 / Server 2022 `IoRing` submission/completion ring, as a separate crate from `windows-overlapped-io-sys` (duplicate-then-decide). Covers ring lifecycle and capability negotiation, zero-allocation token-owned buffers, the batch submission builder, threadless delivery through `ThreadpoolWait`, file/buffer registration, consumer-facing documentation, and the `ring-copy` topology-aligned sample (M1-M7 archived). The pinned-thread (Model B) architecture remains parked as `M6+` by the engineer's explicit direction. `M8` (complete) closed a PR #20 review finding: `FileRef::Raw(HANDLE)`'s lifetime gap, fixed with `unsafe fn` raw entry points plus a safe, `Arc`-backed `SharedFile` wrapper for the common case. `M9` (complete) closed further PR #20 review findings: cross-ring `Token`/`RegisteredFile`/`RegisteredBuffers` confusion (a new per-ring `RingId`, checked at claim/push time), `PendingBufferRegistration` freeing its buffers instead of leaking them on an unclaimed drop, and `Batch::do_submit` letting `Drop` silently retry an already-attempted, already-failed submit. `M10` (active) finished auditing the ring completion contract against all ten specification-gap categories (M10.1-M10.3 complete): category 3 found that `supports` answers for the kernel's op table rather than this crate's push surface and that the registration one-shot is spent by queueing rather than succeeding (D-28); categories 1, 2, 6, 8 and 9 established the load-bearing rule that **every successfully queued SQE produces exactly one completion**, plus that this crate deliberately joins nothing (D-29, D-30); and D-14's registration-index continuity assumption was **dissolved rather than measured** -- the collision it guarded against needs a second registration, which was forbidden the day after D-14 was written, leaving only the reserved-not-confirmed meaning of the public counts to state (D-31). The audit also surfaced two API gaps, now queued as work rather than left in the design notes: `M10.4` (complete) gave `FileRef::Registered` safe entry points, since a registered index carries no lifetime obligation and the `unsafe` guarding it was vacuous (D-29): the safe pushes are now generic over a sealed `FileTarget` trait whose associated `Guard` type carries the one real difference between the two targets, which also made the fully-registered (registered file *and* registered buffer) combination expressible for the first time, non-breakingly (D-33). Investigating M10's own recorded test failures then found a **live use-after-free in shipped 0.1.2** and fixed it as `M10.6`: `BuildIoRingRegisterBuffers` reads its `IORING_BUFFER_INFO` array when the op runs rather than at build time -- the opposite of its file-handle sibling, which the rustdoc had wrongly generalized across -- so the array is now owned by the `IoRing` (D-32). `M10.5` (complete) added named predicates for the conditions a consumer must branch on -- `IORING_E_SUBMISSION_QUEUE_FULL` above all, which every push's rustdoc names as the backpressure signal but which `io::Error::kind()` cannot discriminate (D-30): a complete `RingCondition` enum, predicates for the runtime-actionable conditions, and a sealed `IoRingErrorExt` that puts them on `io::Error` so the downcast is named once rather than hand-rolled per call site (D-34). **M10 is complete.** **`M11` is complete and archived** (2026-08-28): it made the completion event a ring primitive, prompted by an external consumer proposal. `IoRing::completion_event` returns an owned duplicate of the ring's own event so a caller can wait on the ring alongside other handles without surrendering it (D-20); its contract is pinned by eleven sabotage-verified tests; `EventDelivery` is re-expressed on top of it, leaving one `SetIoRingCompletionEvent` call site; `windows-threadpool-sys` moved behind a default-on `threadpool` feature with CI building both combinations (D-22); the wakeup shapes and the barrier's ring-edge limit were swept across every place that states them; and `examples/model_b_multiplexed.rs` works the multiplexed shape end to end. The spike that answered the proposal established that the completion event is **edge-triggered** on the completion queue going empty to non-empty (D-19) -- which also exposed a live bug in shipped 0.1.2, where `EventDelivery` permanently stranded completions queued before handover, fixed in M11.3 by the same change that consolidated `EventDelivery` onto the new primitive. **`M12` is complete and archived** (2026-08-28): it addressed durability, from the same exchange. A spike had established that a flush **without** `DRAIN_PRECEDING_OPS` does not cover preceding writes (D-23) while the barrier that fixes it drains everything already outstanding on the ring (D-24, as corrected by D-47), making `Batch::flush` with default options a silent data-loss bug rather than a missing feature. `Batch::flush`/`flush_raw` now require an explicit `FlushCoverage`, so that spelling no longer exists; a `NO_BUFFERING` integration test proves the barrier's behaviour rather than its flag, and measured that *which direction* the reordering shows in is device-dependent (amending D-23); and the parameters the crate had hardcoded away are exposed as `WriteCaching` (`FILE_WRITE_FLAGS`) and `FlushMode` (`FILE_FLUSH_MODE`, whose `NoSync` is the one mode that commits nothing). Durability had been absent from `lib.rs` and `README.md` entirely, and both now state the three facts. **`M13` is complete and archived** (2026-08-29): the `epoch_log` sample is the vehicle D-26 makes for carrying durability *policy* to consumers without this crate owning it. Its own durability contract is written down first, in its own words (Design Autonomy), then implemented -- records composed into a registered arena and appended, epochs closed by one covering flush whose completion is what makes `is_durable` answer `true`, and a multiplexed wait on the ring's completion event alongside a shutdown latch. The replay pass is what turns it from a demonstration into evidence: it holds the durable region to a strict standard, tolerates a torn tail as the contract requires, and is itself proved able to fail by a negative control. Writing it also found and fixed a gap in the crate (D-35: per-buffer outstanding accounting and `RegisteredBuffers::get_mut`). **`M14` is complete and archived** (2026-08-29): the second half of the sample, covering the two things the ring cannot do for a consumer. A non-ring `FSCTL` (`FSCTL_SET_ZERO_DATA`, reclaiming a retired segment) is ordered against ring epochs by the log itself, since `drain_preceding` orders SQEs against SQEs (D-24) and reaches across neither the ring boundary nor a second ring; a thread-pool control plane runs checkpointing as Model A on a *second* ring while the log thread keeps Model B for the data path, because D-21 forbids a ring given to `EventDelivery` also being waited on directly -- so the ordering chain crosses log thread to pool thread to reclaim worker and back with the log thread blocking for none of it. All three epoch-commit strategies are implemented behind one interface, checked both by replay and by requiring the three to be **byte-identical**, and measured on the running machine. The measurement's finding is that the three are **indistinguishable** here -- the cross-strategy spread is the size of one strategy's run-to-run spread, because every strategy pays one device flush per epoch at hundreds of microseconds while their real differences land in the tens -- and it found two harness bugs nothing else caught, including one where a barrier benchmark that awaits each commit before appending again measures the barrier as free. Both findings are promoted into "Durability on the ring" in [DESIGN-NOTES.md](DESIGN-NOTES.md), since both are about the design rather than the demonstration. **`M15`-`M18` (not started) are the testing-strategy response to the eight defects the 0.1.x line and the M11-M14 branch produced.** They are organised by *defect population* rather than by technique, because the populations need different tools and one of them needs a tool that does not exist: (A) preconditions never varied -- every `event_delivery` test handed over a fresh ring, which is why [#47](https://github.com/MikeGrier/windows-threadpool-sys/issues/47) survived; (B) failure paths never taken -- no test ever ran `completion.result()` returning `Err`, which is why the checkpoint path could authorise a reclaim after a failed write; and (C) *permissions rather than behaviour* -- `&mut Vec` permits `reserve`/`resize`/reassign though no code path performs it, which is [D-35](DESIGN-NOTES.md#d-35) and [D-36](DESIGN-NOTES.md#d-36), the two most severe findings, and **no runtime technique reaches that population at all**. M15 and M16 gate the 0.2.0 release. M15 is deterministic memory instrumentation: a guard-page global allocator, chosen over Application Verifier / PageHeap by measurement rather than assumption ([D-37](DESIGN-NOTES.md#d-37) -- PageHeap works and `reg add` alone is enough, but IFEO is keyed by image file name and one test target produced six distinct hashed names in a day), plus a tracked poison pattern covering the gap guard pages structurally cannot see ([D-38](DESIGN-NOTES.md#d-38) -- a guard page catches access to memory that should not be touched, and is blind to the kernel writing into a live, valid buffer, which is exactly what `write_registered` and `read_registered` promise it will not do). M16 makes the contract executable: a public `RingContract` rendering [the category-2 rule](DESIGN-NOTES.md#one-sqe-one-completion), "one SQE, exactly one completion", checkable rather than merely stated, plus a fault-injection seam that finally takes the failure paths. M17 covers A by generating over the operation space instead of enumerating it by hand, gated on an explicit decision about randomized sampling that this component's conventions require be approved and recorded rather than assumed. M18 covered C, where review is a *primary* technique rather than a backstop, and added `cargo-mutants` against a measured rate of vacuous tests. **M8 through M18 are now complete and archived**; `M20` is pending, and `M6+` is parked rather than pending. The strategy as a whole -- which technique reaches which population, what each one actually found, and what none of them reach -- is recorded in [DESIGN-NOTES.md](DESIGN-NOTES.md#testing-strategy-m185); mutation coverage went from 79.7% to 95.8%, and the borrow-surface audit found one further defect of the same shape as the two that prompted it ([D-43](DESIGN-NOTES.md#d-43)). A mock `IoRing` was considered and **rejected**: both shipped defects were the kernel behaving differently from this crate's assumptions, so a mock would have encoded the same assumptions and passed both bugs green. **M20** queues the documentation and policy-test repairs from the 2026-08-30 NUMA-sharding measurement: a shipping ARM laptop reports no L3 cache domain at all, which falsifies the justification given for the last-level-cache heuristic (though not the heuristic's preference over the NUMA node, and not `ring_copy`, whose degraded fallback already handles it correctly). **M27 (added 2026-09-23) was queued by intent rather than by a review or a measurement, and re-planned the same day.** It began as an adaptivity question for this crate; the adaptivity the workspace's [adoption thesis](../../DESIGN-NOTES.md#the-adoption-thesis) asks for is owned by [topology-planner](../topology-planner/COMPONENT.md) instead ([EP-D-6](../topology-planner/DESIGN-NOTES.md#ep-d-6)), and answering it here would have grown a second policy surface beside it. What survives is the realization end: M27.1 censuses what a realizer needs from this crate against the plan vocabulary and names the gaps, M27.2 closes them as capability with no policy attached, and M27.3 -- ungated -- gives a consumer the means to answer placement questions on their own hardware. [D-8](DESIGN-NOTES.md#d-8) is untouched: being constructible from a policy decision made elsewhere is the opposite of taking one. | [DESIGN-NOTES.md](DESIGN-NOTES.md), [DESIGN-SESSION-2026-08-28-completion-event-multiplexing.md](design-sessions/DESIGN-SESSION-2026-08-28-completion-event-multiplexing.md), [DESIGN-SESSION-2026-08-28-external-consumer-correspondence.md](design-sessions/DESIGN-SESSION-2026-08-28-external-consumer-correspondence.md), [DESIGN-SESSION-2026-08-30-numa-sharded-io-execution-domains.md](../../design-sessions/DESIGN-SESSION-2026-08-30-numa-sharded-io-execution-domains.md) |
diff --git a/crates/windows-ioring-sys/README.md b/crates/windows-ioring-sys/README.md
index d20e892b5..ec853f30a 100644
--- a/crates/windows-ioring-sys/README.md
+++ b/crates/windows-ioring-sys/README.md
@@ -218,11 +218,15 @@ This crate does not partition anything for you (D-8): it makes a ring cheap and
correct, makes its affinity explicit, and leaves sizing a Model B execution
domain to the caller.
-- **Size a domain by last-level (L3) cache, not by NUMA node.** Node count is a
+- **Size a domain by the outermost cache level that partitions the machine, not by NUMA node.** Node count is a
firmware setting a process cannot see, and most real deployments are
- virtualized, where NUMA topology is often invisible entirely. See
- [examples/l3_domains.rs](examples/l3_domains.rs) for a runnable enumeration,
- built on the safe `GetLogicalProcessorInformationEx` wrapper in
+ virtualized, where NUMA topology is often invisible entirely. Ask
+ `outermost_partitioning_cache()` rather than filtering on `CacheLevel == 3`:
+ a shipping ARM part reports no L3 at all, and a machine can report an L3
+ spanning every processor above a real L2 partition, where a level filter
+ returns one whole-machine domain and calls it a cache-aware partition. See
+ [examples/cache_domains.rs](examples/cache_domains.rs) for a runnable
+ enumeration, built on the safe `GetLogicalProcessorInformationEx` wrapper in
[`windows-topology-sys`](../windows-topology-sys/README.md).
- **Processor groups are a hard floor.** A thread's affinity is a
`GROUP_AFFINITY` and a ring's waiter lives in exactly one group, so above 64
@@ -231,15 +235,16 @@ domain to the caller.
on the node closest to the device, registered once into that domain's ring
via `Batch::register_buffers`, is very likely the highest-leverage locality
decision available -- independent of everything above about completion
- routing.
+ routing. `NumaBuffer` is that allocation; *which* node is still the caller's
+ answer, and this crate does not guess it.
`examples/ring_copy` is where these three points become runnable policy: it
copies one file to another through per-domain rings, sized by a named
-`ByL3`/`ByNode`/`ByPackage`/`ByCore`/`Single` policy, with buffers placed via
-`VirtualAllocExNuma` and a `--placement local|remote` switch to make the
-placement effect measurable. It is a **sample**, not library surface -- this
-crate itself depends on no partitioning policy and does not depend on
-`windows-topology-sys`; only the sample does.
+`ByCache`/`ByNode`/`ByPackage`/`ByCore`/`Single` policy, with placed buffers and a
+`--placement local|remote` switch to make the placement effect measurable. It
+is a **sample**, not library surface -- this crate itself depends on no
+partitioning policy and does not depend on `windows-topology-sys`; only the
+sample does.
## License
diff --git a/crates/windows-ioring-sys/RESOLVED-TEST-FAILURES.md b/crates/windows-ioring-sys/RESOLVED-TEST-FAILURES.md
index 888c86a65..519b31775 100644
--- a/crates/windows-ioring-sys/RESOLVED-TEST-FAILURES.md
+++ b/crates/windows-ioring-sys/RESOLVED-TEST-FAILURES.md
@@ -89,3 +89,239 @@ since a caller may have relied on it.
characterised before it is stabilised. The two natural repairs -- loosen the assertion, or mark the
test serial -- would both have suppressed the only evidence that shipped documentation was wrong,
and "flaky test" and "the platform does not do what we wrote down" produce the same symptom.
+
+## Resolved 2026-09-21 22:08:22 -04:00 -- the unidentified ests/bounded_pop.rs failure
+
+**Recorded and resolved the same day, which is the honest framing:** it was recorded as an unresolved
+failure on the grounds that fixing it did not belong in a push of finished milestones. That was a
+scheduling preference dressed as a blocker. The mechanism and the fix were both understood at the time of
+recording; nothing was actually blocking.
+
+**The failure.** One `cargo test --all-features` run reported `FAILED: 4 passed; 1 failed` in a 2.12s
+target matching [bounded_pop.rs](tests/bounded_pop.rs) by shape. The test name and panic message were not
+captured. It never reproduced -- 15 isolated runs and 4 full-suite runs were green.
+
+**The cause, and why no amount of re-running would have settled it.** Those tests needed an operation
+still pending when a short bound expired, and got it from a 128 MiB unbuffered, overlapped read. That is
+a *margin*, not a guarantee: `FILE_FLAG_NO_BUFFERING` bypasses the system cache but not the drive\'s own,
+so the test was asking "will this device take longer than 5 ms?" -- a question about someone else\'s
+hardware, whose answer may differ between two runs on the same machine.
+
+**The fix: an operation that cannot complete, rather than one that is merely slow.** The read is now
+issued against an **overlapped named pipe that nobody has written to**. It is pending because no byte
+exists to satisfy it, and it completes exactly when the test writes one. There is no device, no cache and
+no margin in the question.
+
+Confirmed by probe before being adopted, since neither half was safe to assume: `IoRing` does accept a
+pipe handle for `read_raw`, and `pop_within(20ms)` against an unwritten pipe returns `Ok(None)` with
+`outstanding == 1`.
+
+Where a delay is genuinely needed -- `run_down` polls in 50 ms steps, so forcing it to observe an expired
+poll means releasing the read later than that -- it comes from a `thread::sleep`, whose guarantee runs the
+safe way round: a sleep may overshoot, never undershoot. No assertion depends on an operation *finishing*
+within any bound.
+
+**Verified:** 25 consecutive runs of the target and 3 full `--all-features` suite runs, all green. Both
+sabotages still bite exactly as before the rewrite -- reverting the timeout mapping turns all 5 red, and
+making `RingWait::block` always fail turns 3 red -- so the rewrite kept every bit of the discriminating
+power it had.
+
+It is also **11x faster** (0.20s against 2.26s) and allocates no 128 MiB fixtures, which was the larger
+part of what the file cost to run.
+
+## Resolved 2026-09-25 14:19:35 -04:00 -- event_delivery stalled because the wait was armed after the event was signalled
+
+`SetThreadpoolWait` documents that "you must re-register the event with the wait object before
+signaling it each time to trigger the wait callback". `EventDelivery::new` did the reverse: it took an
+already-signalled event from `IoRing::completion_event` and armed its wait afterwards.
+
+Three properties compounded to make the dropped signal permanent rather than late. The event is
+auto-reset, so the signal was consumed rather than left pending for the arming to observe. It is
+edge-triggered on the completion queue going empty to non-empty, so a ring whose queue was already
+non-empty was signalled by nothing else. And the setup signal exists to serve exactly that case, so it
+was the only wakeup such a ring would ever get.
+
+Fixed by attaching the event unsignalled and raising the setup signal after arming; see
+[DESIGN-NOTES.md](DESIGN-NOTES.md) -> D-68. Measured on the reproducer recorded below: 0 failures in
+3600 runs after the change, against 5 in 600 and 2 in 600 for the two co-running triggers before it.
+
+The investigation as it stood when the cause was found follows.
+
+### event_delivery's threadpool tests time out at roughly one run in eighty
+
+**Found 2026-09-24**, while widening the seeded sweeps to 2048.
+
+**What happens.** `completions_are_delivered_on_pool_threads_without_the_submitting_thread_waiting`
+and `completions_queued_before_handover_are_still_delivered` in
+[event_delivery.rs](tests/event_delivery.rs) each wait on a channel with
+`recv_timeout(Duration::from_secs(5))` for a completion delivered through the Windows thread pool.
+Occasionally the completion does not arrive inside that bound and the test panics with `Timeout`.
+
+**Measured rather than estimated**, because the rate is the whole point: **1 failure in 80
+consecutive runs** of the compiled test binary when first found. It resisted every targeted attempt
+to provoke it at that stage -- zero failures after roughly 18,000 ring create/close cycles, after
+repeated property-suite and calibration runs, and under a concurrent `cargo build` saturating the
+machine -- so it was neither ring-resource pressure nor CPU load. The narrowing below found what it
+actually needs.
+
+### Narrowed 2026-09-25: it requires parallel test execution, and a co-running create-and-drop
+
+The instrumentation described further down paid for itself immediately. One captured occurrence plus
+four follow-up experiments moved this from "cause unknown" to a minimal reproducer. Every figure
+here comes from running the compiled `event_delivery` binary directly.
+
+**It does not happen serially.** With `--test-threads 1`: **0 failures in 1000 runs**. In parallel:
+**7 in 1000**. At the parallel rate a thousand serial runs would expect about seven, so zero is
+evidence rather than a quiet stretch.
+
+**Every occurrence is identical**, across all seven captures:
+
+- **both** delivery tests fail in the same process, never just one;
+- `callbacks run: 0` -- the pool never invoked the callback, not once, for either ring;
+- `delivered: 0 of 8` and `outstanding: 8` -- nothing was ever popped;
+- the post-mortem finds nothing after a further ten seconds.
+
+So it is **not** a slow device and **not** a single lost wakeup. No callback runs at all, for both
+rings, from the start, and the delivery never arrives.
+
+**The two delivery tests alone do not cause it**: 0 failures in 1000 runs with a filter selecting
+only those two. A third test has to be running. Adding them one at a time, 600 runs each:
+
+| Co-running test | Failures in 600 |
+|---|---|
+| `dropping_with_nothing_outstanding_does_not_hang` | 5 |
+| `new_succeeds_and_the_ring_stays_reachable_for_pushes` | 2 |
+| `teardown_with_operations_in_flight_neither_hangs_nor_closes_the_ring_early` | 0 |
+
+The two that trigger it both create an `EventDelivery` over a ring with **nothing outstanding** and
+drop it promptly; the one that does not is the one holding operations in flight. That is a
+correlation across three tests, not a mechanism, and it is recorded as such.
+
+**Reproducer**, about half a minute:
+
+```powershell
+$ed = 'target\debug\deps\event_delivery-.exe' # the build with 6 tests; check with --list
+$fail = 0
+for ($i=1; $i -le 600; $i++) {
+ & $ed completions_ dropping_with 2>&1 | Out-Null
+ if ($LASTEXITCODE -ne 0) { $fail++ }
+}
+"$fail failures of 600"
+```
+
+**Where this goes next, and why it left this crate.** Both delivery tests pass `env: None` to
+`EventDelivery::new`, so both register their wait on the **default process threadpool** through
+[`windows_threadpool_sys::wait::ThreadpoolWait`](../windows-threadpool-sys/src/wait.rs).
+
+### Narrowed further 2026-09-25: the pool is alive, and the stall is permanent by design
+
+A configurable trace was added for this (see below) and the flake **still reproduces with it on**,
+which is the first thing to check for a timing-dependent fault.
+
+**The default pool is not wedged.** A probe runs at the moment of failure, before anything else: it
+submits a plain work item and separately creates, arms and signals a **brand-new** wait on a
+brand-new event. Measured at a captured stall: `work item ran: true; a fresh wait ran: true`, both
+within two seconds. So the pool dispatches, and its wait mechanism works. Whatever is broken is
+specific to the waits already registered.
+
+**Those waits were created and armed.** The trace shows, for every ring in a failing run:
+`setup-signalled` -> `event-attached` -> `wait created` -> `wait armed`, all within microseconds --
+and then `trampoline-entered` **never appears at all**, for any of them, for the rest of the process.
+
+**The ring's setup signal is raised on the ring's own handle, before the wait is armed on a
+duplicate of it.** The trace records both handle values, and in the captures examined the failing
+ring's handles were not recycled values of the dropped ring's.
+
+**Why the stall is permanent rather than merely late, which the trace explains.** The completion
+event is edge triggered ([D-19](DESIGN-NOTES.md#d-19)): it fires when the queue goes from empty to
+non-empty. A stalled ring has eight completions sitting in its queue, so the queue never returns to
+empty and **no further signal will ever be raised**. The setup signal -- the one deliberate wakeup
+that exists precisely to cover a backlog -- is therefore the only signal that ring will ever get.
+Lose it once and delivery for that ring is dead for good. That is consistent with every capture:
+zero callbacks, nothing after ten more seconds, and both rings affected together.
+
+**So the open question is narrow: why does an armed wait not observe a signal raised before it was
+armed?** An auto-reset event signalled with no waiter stays signalled, so arming afterwards should
+consume it and fire. It does, on better than 99% of runs.
+
+**What is deliberately not concluded.** A mechanism suggests itself -- the pool's internal wait
+thread multiplexes handles, and a concurrent close could plausibly disturb the set it is watching,
+which would fit a fresh wait working while existing ones do not. That is a hypothesis with no
+evidence behind it yet, and it is recorded here as one so the next person does not mistake it for a
+finding. The experiment that would settle it is whether re-arming a stalled wait recovers it;
+`EventDelivery` does not currently expose its wait, so that needs either a test-only accessor or the
+probe moved into `windows-threadpool-sys`.
+
+**Why it is worth recording despite being rare.** The sabotage harness runs the whole suite once per
+case, and the manifest currently holds 41 cases. At the measured rate that is about a **40% chance
+that any given sweep contains at least one corrupted result** -- and the corruption is the dangerous
+direction: a sabotage the suite did not really catch is reported as `caught`, which reads as a clean
+bill of health. Both instances seen so far landed on cases whose patches **provably cannot** affect
+event delivery -- a failure-code bitmask in the resolver, and a prose reword inside an example's
+contract text -- which is how they were recognised as false rather than believed.
+
+**How to tell a false `caught` from a real one.** Read the per-case transcript under
+`.scratch/sabotage/`; a genuine detection names a test related to the patch, while this one names
+one of the two tests above and prints the stall report described next. Do not conclude a sweep is
+clean or dirty from the summary table alone while this is open.
+
+**Explicitly not caused by the 2048 sweep widening**, though that is when it was noticed. The rate
+was measured on the `event_delivery` binary, which uses neither the resolver nor any seeded sweep,
+so its behaviour is independent of those constants. Three sweeps at the previous sizes had passed
+earlier the same day, which is unsurprising at this rate rather than evidence of a change.
+
+### Turning the trace on, and narrowing it
+
+The trace is **compiled out** unless the `trace` feature is on, because the instrument for a
+timing-dependent fault must not change the schedule it is measuring. When on it is still off at run
+time until `WINDOWS_THREADPOOL_TRACE` names the targets wanted, so a session can record one
+subsystem rather than everything:
+
+```powershell
+$env:WINDOWS_THREADPOOL_TRACE = 'wait,delivery' # or 'wait', or '*'
+cargo test -p windows-ioring-sys --features trace --test event_delivery
+```
+
+Recording does not format and does not allocate: an entry is a timestamp, a thread id, two
+`&'static str` labels and two `u64` slots, formatted only when a dump is asked for. The dump is
+included in the stall report automatically, so a captured failure carries its own trace.
+
+Targets currently emitted: `wait` (create, arm, trampoline entry, the three phases of drop) and
+`delivery` (the ring's setup signal with its outstanding count, event attach, arm, callback entry
+and exit).
+
+**The flake still reproduces with the trace on**, which was checked before drawing anything from it.
+
+### What a stalled run now records
+
+Added 2026-09-25. The original failure said only `Timeout`, which ruled nothing out -- that is why
+the investigation above could only proceed by elimination. Both tests now print a report on the way
+out, to stderr and into the panic message, so `cargo test`'s captured output and the sabotage
+harness's per-case transcript both carry it. It states:
+
+- **delivered, of how many expected** -- whether the stall was immediate or partway through.
+- **callbacks run** -- how many times the pool actually invoked the callback. Equal to delivered
+ means everything the callback received reached the test thread; greater means the gap is between
+ the callback and the channel. This is the first fork in the diagnosis and nothing else supplies
+ it.
+- **outstanding** -- the ring's own count, with the caveat that makes it readable: it decrements on
+ pop and the pop happens *inside* the callback, so on its own it cannot separate "the kernel has
+ not finished" from "the callback never ran".
+- **arrival times and inter-arrival gaps** -- whether deliveries were steady and then stopped, or
+ slow throughout.
+- **a post-mortem** -- after the bound expires the test waits a further ten seconds and says whether
+ the delivery arrived late or never came at all. Those have different causes, and no other datum
+ separates them.
+
+Two properties of that reporting are deliberate. The immediate facts are printed **before** the
+post-mortem wait, so they survive the sabotage harness killing a run that exceeds its hang bound --
+losing the report to the very timeout it exists to explain would be the worst outcome. And the
+post-mortem is ten seconds rather than thirty so a failing run stays inside that bound: measured, a
+forced stall completes in about seventeen seconds against a bound of roughly thirty.
+
+**The report states observations and stops.** An earlier draft ended with a verdict, and a
+forced-failure run showed the verdict was wrong -- it blamed something upstream of the channel when
+the injected fault was in the callback body, which the counters it had just printed already ruled
+out.
+
+**Queued as `M26.9`** in [CHECKLIST.md](CHECKLIST.md).
diff --git a/crates/windows-ioring-sys/RESPONSE-SPACE.md b/crates/windows-ioring-sys/RESPONSE-SPACE.md
new file mode 100644
index 000000000..13d1cdefb
--- /dev/null
+++ b/crates/windows-ioring-sys/RESPONSE-SPACE.md
@@ -0,0 +1,302 @@
+# The permitted kernel response space
+
+What `windows-ioring-sys` will tolerate from the platform, stated as a
+specification rather than recorded from a run.
+
+This document is normative. `M26.3`'s resolver generates resolutions **from
+this space**, `M26.4`'s properties must hold under every one of them, and the
+kernel tests confirm that a real Windows stays **inside** it (`M26.6`). Every
+clause carries an ID so those three can cite the clause rather than restate it
+-- and [response_space_census.rs](tests/response_space_census.rs) fails when a
+clause is claimed by nothing on the side that owes it a check, so the division
+of labour is enforced rather than merely described.
+
+## What this is, and what it deliberately is not
+
+**It is not a model of what Windows does.** A model of observed behaviour
+freezes one run's testimony, which is the trap
+[D-52](DESIGN-NOTES.md#d-52) was opened to escape and the objection that
+kept a fake out of this crate for two milestones. A resolver built on a model
+asserts something about the kernel and can be wrong about it.
+
+**It is a statement of what this crate will tolerate.** The resolver asserts
+nothing about Windows. It picks a point in the space below, and the assertions
+are about *us*: does this crate behave correctly under that resolution. There
+is no belief here to be wrong about -- only a specification that can be too
+narrow, which is a reviewable defect rather than a hidden one.
+
+**It is wider than anything observed, on purpose.** Deriving the space from
+observation would close the trap again. Where a clause goes beyond what any
+spike has seen, it says so, because the reader's first question about a
+permissive clause is whether anyone has watched it happen.
+
+## How to read a clause
+
+Every clause has an ID, a statement, a source, and a provenance tag:
+
+- **Observed** -- a spike or a measurement in this repository saw it happen.
+ The citation says which.
+- **Documented** -- Microsoft states it on the API's reference page. This is
+ the strongest tag, and it outranks the others: a measurement describes one
+ run of one build, while a documented return value is what the platform
+ commits to. Where a clause carries both, the documentation is the reason and
+ the measurement is corroboration.
+- **Over-provision** -- wider than anything observed here, allowed
+ deliberately. These are the clauses that make the space a specification
+ rather than a recording.
+- **Decided** -- a constraint this crate chooses to require of the platform.
+ Not measured, not derived; a call, and reviewable as one.
+
+A clause tagged **Observed** may still be wider than its observation. Where
+that is so it is split, so the measured part and the extrapolated part can be
+argued separately.
+
+## Permitted: what a resolver may do
+
+### RS-P-1 -- An operation may complete inside `SubmitIoRing`, or pend
+
+Each operation in a submitted batch resolves independently as either *already
+complete when `SubmitIoRing` returns* or *outstanding*.
+
+- **Observed.**
+ [write-pending-spike.rs](design-sessions/spikes/write-pending-spike.rs)
+ measured both outcomes, and measured the mix varying with handle flags and
+ with whether the extent was written beforehand -- see
+ [2026-09-24-set-len-vs-zero-fill/](measurements/2026-09-24-set-len-vs-zero-fill/README.md),
+ where a buffered handle pended in almost none of 16,000 trials and an
+ unbuffered one pended in most runs.
+- **Over-provision:** the resolver chooses **per operation**, independently,
+ with no rate and no correlation to handle flags. No measurement established
+ that operations within one batch resolve independently. The space permits it
+ because a consumer that depends on them resolving together is depending on
+ something Windows never promised.
+
+### RS-P-2 -- Completion order is unconstrained
+
+Completions may be posted in any order, and that order need bear no relation to
+submission order.
+
+- **Observed, in part.**
+ [D-47](DESIGN-NOTES.md#d-47) measured operations queued *after* a drained
+ flush completing *before* it, at 0.03%-0.8% depending on conditions, with all
+ 32 overtaking in the worst observed trial.
+ [kernel-response-space-probe.rs](design-sessions/kernel-response-space-probe.rs)
+ then built a seeded resolver that permutes completion order and ran a
+ FIFO-assuming consumer against it: broken under 189 of 200 seeds, first at
+ seed 0, where completions arrived as `[3, 1, 4, 2]`. That probe is the
+ working demonstration this clause and `M26.3` are both built on; the
+ [session](design-sessions/DESIGN-SESSION-2026-09-22-kernel-response-space.md)
+ is its write-up.
+- **Over-provision:** any permutation, not merely the reorderings observed.
+ This is the clause the `D-47` defect class lives in, and the reason it is
+ total rather than bounded is that a bound derived from observed rates is a
+ recording.
+
+### RS-P-3 -- An operation may fail individually
+
+Any single operation may complete with a failure result while others submitted
+beside it succeed.
+
+- **Documented.** `SubmitIoRing`'s Remarks state the mechanism: *"Any errors
+ processing a single submission queue entry results in a synchronous
+ completion of that entry posted to the completion queue with an error status
+ code for that operation."* So a per-entry failure is **not** a submit
+ failure; it arrives as an ordinary completion carrying an error. That is the
+ contract, and the two observations below are consistent with it rather than
+ the basis for it.
+- **Observed.** Two independent places in this repository are built around it.
+ [`checkpoint.rs`](examples/epoch_log/checkpoint.rs) documents the case
+ directly -- a record write failing with `ERROR_DISK_FULL` followed by a flush
+ that "completes perfectly happily" -- which is why that control plane checks
+ its write's result separately from its flush's. And on the *append* path,
+ `M22.2` found an ordering defect in handling exactly this: a failed write's
+ token was not claimed, so its arena slot leaked and the failure surfaced
+ `SLOTS` appends later with no trace of the cause.
+- **Over-provision:** any error code, at any position in the batch, including
+ the case where every operation fails. The failure *codes* the platform
+ actually returns are not enumerated here and a resolver must not depend on
+ the set being small.
+
+### RS-P-4 -- A wait may expire
+
+A submit-and-wait may return `HRESULT_FROM_WIN32(ERROR_TIMEOUT)` having waited
+its full timeout, and this is **not** a failure of the ring.
+
+- **Observed.** `M21.6` fixed a defect in this crate that treated an expired
+ wait as an operation failure; `ring.rs`'s `IORING_E_WAIT_TIMEOUT` handling is
+ the correction, and it maps the code to `Ok(())`. `M26.8` found the same
+ defect surviving in [`Batch::submit_and_wait`](src/batch.rs), which that
+ sweep had not reached.
+- **Documented, and the documentation says more than "not a failure".**
+ `SubmitIoRing` gives this code its own return-value row: *"All operations
+ were submitted without error and the subsequent wait timed out."* So it
+ carries a positive guarantee about the submission half, not merely the
+ absence of a failure -- which is why it is classified separately from every
+ other error rather than folded in with them.
+- **Over-provision:** a wait may expire even when completions are available,
+ and may expire on any call including the first.
+
+### RS-P-5 -- A wait may return successfully with nothing poppable
+
+A wait that returns success does not promise that a subsequent pop yields
+anything.
+
+- **Observed.** This crate's own `pop_within` documentation states it -- a
+ submit-side wait's return "promises nothing about poppability" -- and
+ [D-19](DESIGN-NOTES.md#d-19) measured the completion event as **edge**
+ triggered on the queue going empty to non-empty, so a signal is not a count
+ and the queue can be drained by the time a waiter looks.
+- **Over-provision:** this may happen on any wait, any number of times in
+ succession. A resolver is not required to make progress on any particular
+ call, only to satisfy RS-C-1 eventually.
+
+### RS-P-6 -- A signal may not arrive for a completion posted while the queue was already non-empty
+
+The completion event fires on the empty-to-non-empty edge, so a completion
+arriving behind another need not produce its own signal.
+
+- **Observed.** [D-19](DESIGN-NOTES.md#d-19), measured, and
+ [D-21](DESIGN-NOTES.md#d-21) is the consequence this crate drew from it --
+ auto-reset, exactly one waiter per ring.
+- **Over-provision:** none. This clause is the measurement.
+
+### RS-P-7 -- A submit may fail, leaving already-built operations queued for a later submit
+
+`SubmitIoRing` may fail, and when it does every entry it was asked to submit
+remains in the submission queue. A later, unrelated submit is what runs them.
+
+- **Documented, which `M26.8` established and `M26.1` had not.** This clause
+ was originally written as a *consequence* -- "if a submit fails, entries
+ remain queued" -- citing [D-5](DESIGN-NOTES.md#d-5), which establishes the
+ no-rewind consequence and nothing about submits failing at all. `M26.3`'s
+ resolver read it as a permission to fail submits, and the space had no
+ authority for that. It does now: `SubmitIoRing`'s return-value table lists
+ *"Any other error value: Failure to process the submission queue in its
+ entirety"*, and its Remarks state *"If this function returns an error other
+ than IORING_E_WAIT_TIMEOUT, then all entries remain in the submission
+ queue."* Both halves of this clause are therefore Microsoft's, not an
+ inference from a run.
+- **The `IORING_E_WAIT_TIMEOUT` carve-out is part of the clause**, because it
+ is the case where the entries did *not* remain queued: that code means every
+ operation was submitted and only the wait expired. A consumer that cannot
+ tell the two apart cannot know whether its buffers are still owed to the
+ kernel, which is the defect `M26.8` fixed in
+ [`Batch::submit_and_wait`](src/batch.rs).
+- **Over-provision:** the resolver may defer an operation across any number of
+ submits, not only across a failed one. It currently declines a submit only
+ when something is staged, which is *narrower* than this clause -- nothing
+ says a submit carrying no new work cannot fail.
+
+## Constrained: what a resolver may not do
+
+These exist because a resolver free to violate everything makes this crate
+defend against a platform that does not exist, and code written against an
+impossible kernel is untestable and unreviewable. Each is a call.
+
+### RS-P-8 -- A successful transfer may report fewer bytes than requested
+
+A read or write that completes successfully may report an `Information` below
+the length it was given. The remainder is not transferred, and nothing in this
+crate reissues it.
+
+- **Documented, for the handle types that do it.** `WriteFile`'s Remarks state
+ that *"when writing to a non-blocking, byte-mode pipe handle with
+ insufficient buffer space, WriteFile returns TRUE with
+ \*lpNumberOfBytesWritten < nNumberOfBytesToWrite"*. Sockets report a short
+ send when the transmit buffer cannot take the whole buffer, and a
+ communications handle with a write timeout set by `SetCommTimeouts` can
+ report a partial count when the timeout fires after some data has gone out.
+ Reads are shorter still by nature: end of file, a pipe with less buffered
+ than asked for.
+- **Why it is a permission of this space and not a property of a file.** For an
+ ordinary file on a local volume a successful completion is expected to carry
+ the full requested length, and a full volume is an error
+ (`ERROR_DISK_FULL`) rather than a short success. But **this crate does not
+ constrain what a caller registers** -- `IoRing` takes a handle, and a pipe, a
+ socket and a serial port are all handles. A consumer of this crate therefore
+ has to read the count. A consumer that has *also* narrowed its handle type
+ can rely on more, and the place to say so is that consumer's own contract,
+ not this space.
+- **The continuation is the caller's.** This crate reports the count and stops
+ there. Whether to reissue the remainder, how many times, and when to give up
+ are policy, and [D-67](DESIGN-NOTES.md#d-67) keeps policy with the caller.
+- **A short count may be zero.** A consumer looping on the remainder must
+ tolerate a completion that makes no progress rather than assuming each one
+ advances it.
+- Flush and cancel carry no byte count, so this clause does not reach them.
+### RS-C-1 -- Every submitted operation eventually completes exactly once
+
+No completion is lost, none is duplicated, and every successfully submitted
+operation eventually produces exactly one completion.
+
+- **Decided**, not observed. Nothing here has measured the negative, and it
+ could not be measured in bounded time.
+- **Why:** this crate's accounting is driven by observing a real `IORING_CQE`
+ ([D-4](DESIGN-NOTES.md#d-4)) and its [`RingContract`](src/contract.rs) oracle
+ states conservation directly. A platform that lost completions would make
+ every consumer's outstanding count unbounded and every wait a guess. If this
+ is ever observed to fail, the finding is a contract defect in Windows and not
+ a gap in this space.
+
+### RS-C-2 -- A completion identifies the operation that produced it
+
+A completion's `UserData` is the value supplied when the operation was built.
+
+- **Decided.** The alternative is that operation identity is unusable, which
+ would invalidate [D-4](DESIGN-NOTES.md#d-4)'s whole accounting model and
+ `M28`'s pending inventory with it.
+
+### RS-C-3 -- An operation does not complete before it is submitted
+
+- **Decided**, and stated because a resolver that may post a completion for an
+ operation still being built would make the `Build*`/`Submit` boundary
+ meaningless.
+
+### RS-C-4 -- The drain half of `DRAIN_PRECEDING_OPS` holds
+
+No operation queued **before** a flush carrying
+`IOSQE_FLAGS_DRAIN_PRECEDING_OPS` completes after that flush.
+
+- **This is the explicit call `M26.1` demanded, and it is the one place this
+ space is narrower than "anything may happen".**
+- **Observed, with the strongest evidence in this repository.**
+ [D-47](DESIGN-NOTES.md#d-47) measured roughly 4,500 trials in which **not
+ once** did an operation queued before a drained flush complete after it --
+ the same campaign that falsified the *other* half of `D-24`.
+- **Why constrained:** the drain is the documented guarantee this crate's
+ durability story rests on ([D-23](DESIGN-NOTES.md#d-23)). A resolver
+ permitted to break it would require every consumer to re-verify durability by
+ some other means, which is to say it would make the primitive useless. The
+ cost of this call is that a Windows which broke the drain would not be caught
+ by the resolver at all -- it is caught by the kernel tests instead, which is
+ the division of labour they were repointed to in `M26.6`. That is now a
+ mechanical arrangement rather than an intention:
+ [flush_barrier.rs](tests/flush_barrier.rs) carries a `CONFIRMS: RS-C-4`
+ marker and asserts the clause against a real ring on every machine, and
+ [response_space_census.rs](tests/response_space_census.rs) fails if that
+ marker ever disappears.
+- **Note what is *not* constrained:** the hold-back half.
+ [D-24](DESIGN-NOTES.md#d-24) claimed the flag holds back what follows and
+ [D-47](DESIGN-NOTES.md#d-47) withdrew that claim, so RS-P-2 applies in full
+ to operations queued *after* a drained flush. The one-sidedness is the whole
+ point of the pair.
+
+## What this space deliberately leaves undecided
+
+Stated so that a later reader can tell an omission from a choice:
+
+- **Rates.** No clause carries a probability. A resolver weights its choices by
+ seed, and any weighting is a property of the resolver rather than of this
+ space. Recording observed rates here would make the space a recording.
+- **Failure code sets.** RS-P-3 permits any code and enumerates none.
+- **Timing.** Nothing here constrains how long anything takes. `M25`'s standing
+ constraint already forbids this crate from depending on an operation pending,
+ and a space that specified durations would invite exactly that.
+
+## Changing this document
+
+A clause moving from **Over-provision** to **Observed** is an improvement and
+needs only its citation updated. A clause moving in the other direction, or a
+constraint being relaxed, changes what this crate promises to tolerate and
+must be recorded as a decision in [DESIGN-NOTES.md](DESIGN-NOTES.md) with the
+finding that forced it.
diff --git a/crates/windows-ioring-sys/RING-OPENING-LIB-TESTS.txt b/crates/windows-ioring-sys/RING-OPENING-LIB-TESTS.txt
new file mode 100644
index 000000000..8286e346d
--- /dev/null
+++ b/crates/windows-ioring-sys/RING-OPENING-LIB-TESTS.txt
@@ -0,0 +1,49 @@
+# windows-ioring-sys: lib tests that open a real kernel ring.
+#
+# GENERATED by tools/check-ring-tests.ps1 -Update. Do not hand-edit.
+#
+# These are integration tests living in the unit-test location (D-49).
+# The list exists so the population cannot grow unnoticed, which is how
+# it reached 63 before anyone counted. Adding an entry obliges an
+# answer: does this test need the kernel, or only a ring-shaped thing?
+batch::a_pending_buffer_registration_claims_only_its_own_completion
+batch::a_pending_file_registration_claims_only_its_own_completion
+batch::a_pending_file_registration_reports_the_user_data_it_will_claim
+batch::dropping_a_batch_that_queued_nothing_submits_harmlessly
+batch::dropping_a_registration_with_work_outstanding_is_refused
+batch::registered_files_index_from_the_base_and_stop_at_the_end
+batch::registered_files_report_their_extent
+batch::require_refuses_an_op_the_ring_does_not_support
+batch::submit_reports_how_many_operations_it_queued
+batch::windows_refuses_an_empty_buffer_registration
+event_delivery::a_scope_reflects_a_ring_that_genuinely_lacks_support
+event_delivery::a_scope_reports_outstanding_work
+event_delivery::a_scope_reports_registration_counts_that_change_with_registrations
+event_delivery::a_scope_reports_the_rings_static_properties
+ring::a_bound_the_clock_cannot_represent_reaches_the_wait_rather_than_panicking
+ring::a_reserved_opcode_is_not_supported
+ring::a_supplied_wait_is_consulted_when_the_queue_is_not_ready
+ring::a_supplied_wait_is_not_consulted_when_nothing_can_arrive
+ring::a_wait_that_fails_ends_the_pop_with_its_error
+ring::a_wait_that_never_blocks_is_permitted_and_still_terminates
+ring::a_zero_bound_does_not_block
+ring::an_injected_failure_carries_the_condition_it_names
+ring::an_injected_failure_preserves_the_identity_a_token_claims_against
+ring::an_injected_failure_replaces_a_real_success
+ring::an_injected_failure_zeroes_the_transferred_byte_count
+ring::capability_reporting_never_claims_more_than_is_io_ring_op_supported_reports
+ring::dropping_a_ring_actually_runs_its_drop_body
+ring::each_spelling_of_a_failure_produces_the_condition_it_names
+ring::every_named_condition_injects_a_genuine_failure
+ring::injecting_a_success_code_is_refused
+ring::nop_read_and_write_are_supported_on_any_real_ring
+ring::pop_within_returns_successive_completions_one_at_a_time
+ring::pop_within_returns_the_completion_of_a_real_operation
+ring::ring_wait_reports_the_rings_outstanding_count
+ring::run_down_returns_once_a_recorded_completion_zeroes_the_count
+ring::submit_wait_is_what_the_convenience_uses
+ring::supports_reports_exactly_the_capability_set_it_was_given
+ring::the_deadline_is_honoured_when_an_operation_never_completes
+ring::the_debug_rendering_names_the_ring_and_its_key_fields
+ring::the_wait_can_be_supplied_as_a_trait_object
+ring::the_wait_is_never_handed_a_zero_timeout
diff --git a/crates/windows-ioring-sys/UNRESOLVED-TEST-FAILURES.md b/crates/windows-ioring-sys/UNRESOLVED-TEST-FAILURES.md
index abf45c5f5..c6ccebdfa 100644
--- a/crates/windows-ioring-sys/UNRESOLVED-TEST-FAILURES.md
+++ b/crates/windows-ioring-sys/UNRESOLVED-TEST-FAILURES.md
@@ -4,4 +4,32 @@ Pre-existing failures that do not block an unrelated commit, recorded per the re
checklist-execution rules. When one is resolved, move its entry into a sibling
[RESOLVED-TEST-FAILURES.md](RESOLVED-TEST-FAILURES.md) (append-only) rather than deleting it.
-None currently.
+## a backlog is not delivered when the caller attached the event and consumed its signal
+
+**Found 2026-09-25**, by Copilot review on PR #108, which reported the narrower form: that
+`EventDelivery::new` signals only when it attached the event itself, so a caller who attached it
+earlier arms a wait on an already-non-empty, edge-triggered queue with no wakeup owing.
+
+The reproducer is
+`a_backlog_is_delivered_even_when_the_caller_attached_the_event_first` in
+[event_delivery.rs](tests/event_delivery.rs), `#[ignore]`d because it fails. It attaches the event,
+**consumes** the signal attaching raised, submits work, consumes the signal the completions raise,
+and only then hands the ring over -- leaving a non-empty queue with nothing pending, which is the
+state the guarantee is about. An earlier version of the test omitted the two consuming waits and
+passed against the defect, because the leftover signal fired the wait; that version proved nothing.
+
+**The reported repair does not work, which is why nothing is applied.** Signalling unconditionally
+after arming leaves the test failing 6 of 6. What does make it pass is a 50 ms sleep between
+`wait.arm` and the signal: 3 of 3. Building with `--features trace` also makes it pass, which is the
+same schedule perturbation by another route.
+
+So the wakeup is lost in a window *after* arming, rather than never being raised. That is wider than
+the review finding, and it bears on [D-68](DESIGN-NOTES.md#d-68): `M26.9` fixed the delivery stall
+by ordering the arm before the signal, measured at 0 failures in 3600 runs, and this says that
+ordering alone is not sufficient to close the window -- only to narrow it.
+
+**Not established:** why the window exists. `SetThreadpoolWait` is documented as registering the
+wait, so a signal after a completed `arm` should be observed. Whether the pool's wait thread
+re-issues its `WaitForMultipleObjects` asynchronously, and whether an auto-reset signal can be
+consumed and discarded during that re-issue, is a guess and is recorded here as one. Queued as
+`M26.12`.
diff --git a/crates/windows-ioring-sys/design-sessions/DESIGN-SESSION-2026-08-28-external-consumer-correspondence.md b/crates/windows-ioring-sys/design-sessions/DESIGN-SESSION-2026-08-28-external-consumer-correspondence.md
index cdd385015..dbd76d4e0 100644
--- a/crates/windows-ioring-sys/design-sessions/DESIGN-SESSION-2026-08-28-external-consumer-correspondence.md
+++ b/crates/windows-ioring-sys/design-sessions/DESIGN-SESSION-2026-08-28-external-consumer-correspondence.md
@@ -159,6 +159,11 @@ writes -- observed at 17 and 23 of 32 writes completing after it) and
[D-24](../DESIGN-NOTES.md#d-24) (the barrier is a full, ring-wide stall that spans submissions and
holds operations against unrelated files).
+*Recorded as it was concluded on this date. D-24's "holds operations" half was withdrawn on
+2026-09-06 by [D-47](../DESIGN-NOTES.md#d-47), which measured operations queued behind a drained
+flush completing ahead of it; the drain half -- that nothing queued before it completes after it --
+stands.*
+
## Round 4 -- what belongs where
The closing question was whether this crate should provide the emulation a consumer needs on top of
diff --git a/crates/windows-ioring-sys/design-sessions/DESIGN-SESSION-2026-09-19-epoch-log-review.md b/crates/windows-ioring-sys/design-sessions/DESIGN-SESSION-2026-09-19-epoch-log-review.md
new file mode 100644
index 000000000..724df4e0d
--- /dev/null
+++ b/crates/windows-ioring-sys/design-sessions/DESIGN-SESSION-2026-09-19-epoch-log-review.md
@@ -0,0 +1,232 @@
+# Design session 2026-09-19: review of the epoch-log sample and its durability surface
+
+A read-only review of [examples/epoch_log](../examples/epoch_log) and the parts of the crate it
+composes, prompted by two questions from the engineer: whether the sample is correct and efficient,
+and whether the repository's accumulated learnings about ring structuring and storage affinity
+suggest reworking it.
+
+**Nothing was built, run, or measured during this session.** Every finding below is from reading the
+source and the recorded decisions. Where a finding rests on reasoning rather than on a measurement,
+it says so. No claim here is a compile claim.
+
+## What resulted
+
+New work items [M21](../CHECKLIST.md), [M22](../CHECKLIST.md) and [M23](../CHECKLIST.md) in
+[CHECKLIST.md](../CHECKLIST.md), plus an addendum to the already-queued
+[M20.6](../CHECKLIST.md). No decision in [DESIGN-NOTES.md](../DESIGN-NOTES.md) was changed by this
+session; two of the findings are about decisions whose corrections are queued and not yet landed
+(see "Already queued, still undone" below).
+
+## Scope and method
+
+Read in full or in relevant part:
+
+- every module of [examples/epoch_log](../examples/epoch_log);
+- [src/batch.rs](../src/batch.rs)'s submission and flush surface, [src/ring.rs](../src/ring.rs)'s
+ pop and test helpers, [src/lib.rs](../src/lib.rs)'s durability and topology guidance;
+- [DESIGN-NOTES.md](../DESIGN-NOTES.md) decisions D-3, D-5, D-8, D-19, D-21, D-23, D-24, D-47, the
+ "Why the NUMA node is the wrong key" and "What is not reachable" sections;
+- [CHECKLIST.md](../CHECKLIST.md) M20, and the session it was queued from,
+ [DESIGN-SESSION-2026-08-30-numa-sharded-io-execution-domains.md](../../../design-sessions/DESIGN-SESSION-2026-08-30-numa-sharded-io-execution-domains.md);
+- [design-sessions/spikes](spikes) -- the two unrun instruments and their README;
+- [examples/ring_copy](../examples/ring_copy) for comparison, since it is the crate's other sample
+ and the one that does make a locality decision.
+
+## Findings: correctness
+
+### C-1. A withdrawn D-24 claim survives at one site
+
+[examples/epoch_log/commit.rs](../examples/epoch_log/commit.rs) line 156 justifies its epoch-order
+`debug_assert` with "D-24 holds an operation pushed after a drained one until it completes". That is
+the half of [D-24](../DESIGN-NOTES.md#d-24) that [D-47](../DESIGN-NOTES.md#d-47) withdrew, and the
+same file's own module header (line 24) already carries the correction.
+
+A blast-radius sweep of the hold-back phrasing across the crate found 17 matches in 10 files; every
+other site is corrected. This is the last one.
+
+The assertion it guards is still sound, by a different route: commit *N+1* carries the drain flag
+itself, and D-47's *surviving* half ("not once did an operation queued before a drained flush
+complete after it") is what orders it behind commit *N*. So the conclusion holds and the cited
+reason does not -- a correction that did not propagate, rather than a wrong conclusion.
+
+### C-2. A bare `loop { try_pop }` in the flagship example
+
+[examples/epoch_log/append.rs](../examples/epoch_log/append.rs) line 89 spins unbounded after
+`submit_and_wait(1, 30_000)`. [`Batch::submit_and_wait`](../src/batch.rs) documents that returning
+does not mean a completion is poppable, because the timeout can expire first, and
+[`pop_within`](../src/ring.rs) states the consequence outright: "A bare `loop` around `try_pop` is
+worse, because it converts that flake into a hang."
+
+The same step is written three ways in this crate:
+
+| Site | Shape |
+|---|---|
+| [examples/ring_copy/engine.rs](../examples/ring_copy/engine.rs) line 147 | absence is an error (`TimedOut`) |
+| [examples/epoch_log/strategy.rs](../examples/epoch_log/strategy.rs) `Lane::new` | absence is an error |
+| [src/ring.rs](../src/ring.rs) `pop_within` | bounded wait, panics on deadline |
+| [examples/epoch_log/append.rs](../examples/epoch_log/append.rs) line 89 | unbounded hot spin |
+| [tests/fault_injection.rs](../tests/fault_injection.rs) line 50 | unbounded hot spin |
+
+`pop_within` is `#[cfg(test)] pub(crate)`, so neither an example nor an integration test can reach
+it -- examples and `tests/` are separate crates. That is why the duplication exists, and it means
+the fix is an API question rather than a copy-paste: publish a bounded pop, or keep re-deriving it.
+
+The rule is written down in three places and enforced nowhere, which is the detection-ladder point:
+prose is not a rung.
+
+### C-3. The commit trigger keys off the counter, not off the append
+
+[examples/epoch_log/main.rs](../examples/epoch_log/main.rs) line 279 tests
+`appended % EPOCH_SIZE == 0` on every pass of the append loop, including a pass where `append`
+returned `WouldBlock` and `appended` did not move. On such a pass it commits again: a second
+covering flush closing an epoch with nothing in it, and an epoch number consumed for no records.
+
+Not reachable at the sample's current constants -- `SLOTS` is 8, `EPOCH_SIZE` is 6, and the commit
+wait drains the arena, so the arena cannot be full at a boundary. It is armed by anyone who copies
+the sample and raises `EPOCH_SIZE`, which is what the sample exists to be.
+
+> **Corrected 2026-09-21 while implementing `M21.3`: the second paragraph is wrong.** The retry is
+> not reachable at *any* constants, because the predicate is true at exactly two moments -- before
+> the first append, and immediately after a commit -- and the arena is empty at both, the commit
+> having waited for a covering flush that retires every outstanding write. Measured rather than
+> re-reasoned: the retry path was instrumented to report when the old shape would have committed,
+> and it fired **zero** times at `EPOCH_SIZE` of 6, 8, 12, 16 and 24, including the values past
+> `SLOTS` this finding predicted would arm it.
+>
+> What survives is the coupling complaint in the heading, and it is worth the change on its own: the
+> trigger was safe because of an invariant three blocks away that nothing stated, rather than
+> because of where it was written. The lesson for this review is narrower and sharper -- "unreachable
+> today, armed tomorrow" is a claim about a program's reachable states, and reading the code is not
+> how to settle one.
+
+### C-4. `durable_through` across a failed commit is under-specified
+
+[examples/epoch_log/commit.rs](../examples/epoch_log/commit.rs) says "A failed commit advances
+nothing", which reads as though a failed commit of epoch *N* leaves *N* non-durable permanently.
+
+It does not. Epoch *N*'s writes precede commit *N+1*'s covering flush, so a later successful commit
+makes *N* genuinely durable, and the monotonic reading of `durable_through` stays true. That is the
+correct behaviour; the reasoning appears nowhere, so a reader auditing monotonicity after a failure
+has to re-derive it. This is a specification gap, not a defect.
+
+### C-5. Two wait loops hang where their sibling fails
+
+[examples/epoch_log/strategy.rs](../examples/epoch_log/strategy.rs) lines 372 and 384 discard the
+`submit_and_wait` timeout and loop forever.
+[`EventLoop::pump`](../examples/epoch_log/event_loop.rs) raises `TimedOut` on the same condition and
+documents why: "so a stuck loop fails instead of spinning". One program, opposite policies.
+
+## Findings: efficiency
+
+### E-1. One `SubmitIoRing` per record
+
+Both [`Appender::append`](../examples/epoch_log/append.rs) and
+[`Lane::append`](../examples/epoch_log/strategy.rs) construct a `Batch`, push one write, and submit
+it. `Batch` exists to amortise submission across many SQEs; the sample that teaches `Batch` submits
+one entry at a time.
+
+This is not only a throughput observation. It puts a fixed per-record submission cost into all three
+strategies in [strategy.rs](../examples/epoch_log/strategy.rs), which is a shared term in the
+comparison whose headline result is that the three are indistinguishable. Whether batching moves
+that spread is unmeasured; it is the cheapest experiment available, and it bears on
+[M20.6](../CHECKLIST.md).
+
+### E-2. Two implementations of the free-slot pool, in one program
+
+[`Appender::free_slot`](../examples/epoch_log/append.rs) line 132 scans the arena calling
+`outstanding()` per slot; `Lane` keeps a `Vec` free list. Both are correct and the difference
+does not matter at eight slots. The duplication is what matters, because the two can drift.
+
+### E-3. The arena has no placement story
+
+[examples/epoch_log/append.rs](../examples/epoch_log/append.rs) line 84 allocates the registered
+arena as `vec![0_u8; SLOT_LEN]` -- heap, no alignment, no node. The crate's own front page
+([src/lib.rs](../src/lib.rs)) says buffer placement "is very likely the highest-leverage locality
+decision available" and names `VirtualAllocExNuma`, and
+[examples/ring_copy/buffer.rs](../examples/ring_copy/buffer.rs) already implements exactly that.
+
+The durability sample has no locality story at all: no pinning, no node-local arena, one ring. That
+may be the right call for a sample about durability -- but it is currently a silence rather than a
+stated choice, while the crate's headline guidance says the opposite.
+
+## Findings: ring structuring and storage affinity
+
+### S-1. D-47 left one cost standing that the sample never names
+
+[D-47](../DESIGN-NOTES.md#d-47) withdrew the hold-back claim and explicitly kept the other half: the
+barrier "does still reach every outstanding operation on the ring rather than only the current
+submission batch".
+
+So commit latency is a function of whatever else shares the ring. The ring is therefore part of the
+log's durability unit, and "one ring per log" is a precondition rather than a sample convenience.
+
+**Refined when `M23.1` implemented this (2026-09-23); the finding is left as recorded, per Tier 3.**
+The barrier is a *ring* flag and the flush names a *file*, so the two bound different things: the
+barrier bounds what a commit waits for, the flush bounds what it makes durable. Completion is not
+durability, so a shared ring threatens the **cost model** rather than the guarantee. See
+[contract.rs](../examples/epoch_log/contract.rs) -> "One ring per log, because the barrier is
+ring-wide", which is authoritative over this paragraph.
+[contract.rs](../examples/epoch_log/contract.rs) -- which is where this sample puts its
+preconditions, and which was deliberately written before the code -- does not say so.
+
+### S-2. The re-founding of `AlternatingRings` that M20.6 is asking for
+
+[M20.6](../CHECKLIST.md) asks whether alternating rings still earns its cost now that its stated
+benefit (keeping appends off a stalled ring) is withdrawn, and offers "epoch *N+1*'s appends are
+provably outside epoch *N*" as the remaining benefit, characterising that as a correctness property
+rather than a throughput one.
+
+Read against S-1 it is also a throughput property, sited differently. What alternating rings buys is
+a **bound on what a commit's barrier can be dragged into**: under a shared ring, commit latency is
+unbounded in unrelated traffic on that ring; under alternating rings it is bounded by the epoch.
+That is a stronger answer than the one M20.6 currently records, and it is measurable with the
+harness that already exists.
+
+Unmeasured. It follows from the flush's recorded scope plus D-47's surviving half.
+
+### S-3. The storage-affinity question, and the adjacent one that is answerable
+
+The engineer's framing -- Windows does not really present storage NUMA affinity for NVMe-attached
+storage, but that is no reason not to anticipate it -- matches what the repository has already
+established, and [M20.4](../CHECKLIST.md) holds the mechanism research: `FSCTL_QUERY_VOLUME_NUMA_INFO`
+takes a file or directory handle directly with no device-instance walk; `GetNumaNodeNumberFromHandle`
+bottoms out in `NtQueryInformationFile` with `FileNumaNodeInformation` (class 53), which PHNT and the
+WDK mark reserved for system use, so this crate must not build on it; and no published measurement of
+either succeeding on an ordinary NTFS data file could be found. The conclusion survives for a better
+reason than the one originally recorded: the documented meaning is the node the *volume* resides on,
+not where the file's extents live, so it cannot answer "which ring should this file's I/O go to" even
+when it succeeds.
+
+Two ways to anticipate it without claiming it, both of which keep [D-8](../DESIGN-NOTES.md#d-8)
+intact by leaving the policy with the consumer:
+
+1. **Declare rather than discover.** Let a consumer *state* the storage node for a domain and have
+ the arena allocate there via `VirtualAllocExNuma`. The unanswerable question becomes a declared
+ input; [file-handle-numa-spike.rs](spikes/file-handle-numa-spike.rs) fills it in automatically if
+ hardware ever answers.
+2. **Shard by backing device rather than by node.** `IOCTL_STORAGE_GET_DEVICE_NUMBER` and
+ `IOCTL_VOLUME_GET_VOLUME_DISK_EXTENTS` -- both already named in the spike -- answer *which
+ physical device backs this handle*, and that question is reachable today on ordinary hardware. It
+ is also the question that governs the cost this sample is built around: a device cache flush is
+ per-device, so two logs on one device contend at every commit, and a ring spanning two devices
+ takes the slower device's flush on every covering flush.
+
+The second is the substantive suggestion of this session: the placement question people reach for
+(which node) is unanswerable, while the adjacent question that actually sets commit cost (which
+device) is not. Unmeasured; it follows from the flush's scope, and the instruments to settle it
+exist.
+
+## Already queued, still undone
+
+Two corrections were queued before this session and have not landed. Recorded here so this session
+is not read as discovering them:
+
+- [M20.4](../CHECKLIST.md) -- "What is not reachable" in [DESIGN-NOTES.md](../DESIGN-NOTES.md) still
+ carries the superseded mechanism (walking volume to disk to device instance and reading
+ `DEVPKEY_Device_Numa_Node`). The corrected mechanism is written in the checklist item and not in
+ the decision.
+ **Landed the same day**, in this session's follow-up work -- see
+ [COMPLETED-CHECKLIST.md](../COMPLETED-CHECKLIST.md#m204), which also records two things the item's
+ own text had stale.
+- [M20.6](../CHECKLIST.md) -- the `AlternatingRings` re-evaluation. S-2 above is an addendum to it,
+ not a replacement.
diff --git a/crates/windows-ioring-sys/design-sessions/DESIGN-SESSION-2026-09-21-hermetic-unit-tests.md b/crates/windows-ioring-sys/design-sessions/DESIGN-SESSION-2026-09-21-hermetic-unit-tests.md
new file mode 100644
index 000000000..7ce51a0f2
--- /dev/null
+++ b/crates/windows-ioring-sys/design-sessions/DESIGN-SESSION-2026-09-21-hermetic-unit-tests.md
@@ -0,0 +1,228 @@
+# Design session 2026-09-21: hermetic unit tests without losing unit-level coverage
+
+**Status: concluded 2026-09-21.** The engineer's verdict was that the current design is incorrect
+and the question is only when to fix it, so the defect and its classification are recorded as
+[D-49](../DESIGN-NOTES.md#d-49) and the work is scheduled as `M24` in [CHECKLIST.md](../CHECKLIST.md).
+The *remedy* is still open, gated on `M24.1`. What follows is the record as written before that
+verdict; the proposal below is the input to `M24.1`, not its answer. This records a question, the
+measurements taken to answer it, and a proposal, so that a decision can be made against evidence
+rather than against recollection. If the proposal is adopted it becomes a decision in
+[DESIGN-NOTES.md](../DESIGN-NOTES.md) and checklist items in [CHECKLIST.md](../CHECKLIST.md), in the
+same change.
+
+## The question
+
+Raised by the engineer after the M21 work, in two parts:
+
+> A unit test that is affected by system load is not a unit test. [...] Unit tests should be
+> hermetic. The fact that this failure was discovered due to system load is prima facie evidence
+> that the tests are not hermetic.
+
+and then, on being told the structural fix would move those tests to `tests/`:
+
+> While we structurally *can* [move] those tests to be integration tests, that means we lose their
+> coverage at the unit level which is not something that I want to give up lightly. I would prefer to
+> be able to either have the same test code duplicated or be effectively a library and be able to
+> apply to the hermetic IoRing as well as the actual one.
+
+## What is already settled, and what was got wrong
+
+**The classification is not in doubt.** The repository's own Quality rule reserves integration tests
+for cases that "must cross a real process, filesystem, network, device, **operating-system API**, or
+other external boundary". `CreateIoRing` is an operating-system API. Tests that open a real ring are
+integration tests, wherever they currently live.
+
+**What M21.6 fixed was not hermeticity.** Removing five wall-clock assertions from four unit tests
+made their *outcome* independent of load -- measured, under 2x CPU saturation, at 30x duration
+variation with zero outcome variation. It did not make them hermetic: they still open a real kernel
+ring. Outcome-stable and hermetic are different properties, and conflating them is what let the
+original answer sound complete when it was not.
+
+One correction to the evidence, which cuts in the engineer's favour rather than against: the
+load-affected failure actually observed was in `tests/bounded_pop.rs`, which *is* an integration
+test and is where it belongs. The unit suite's hermeticity problem was real but separate -- it was
+the five clock assertions, which failed nothing and were removed on principle.
+
+## Measurements
+
+Counted by command, not by recollection.
+
+**How much of the unit suite crosses the boundary:**
+
+| Module | Tests | Open a real ring |
+|---|---|---|
+| `buf`, `capability`, `contract`, `error` | 60 | 0 |
+| `batch` | 18 | 13 |
+| `ring` | 40 | 37 |
+| `event_delivery` | 6 | 6 |
+| `token` | 7 | 7 |
+| **total** | **131** | **63** |
+
+Four modules are already perfectly hermetic. The non-hermetic mass is concentrated in four others.
+
+**How movable those 63 are, as they stand:**
+
+- **25** use only public API. A pure relocation to `tests/`.
+- **38** also reach crate-private items -- `reserve_user_data`, `record_completion`,
+ `cancel_reservation`, `Completion::synthetic`, `set_supported_ops_for_test`, `raw_handle`,
+ `ring_id`. An integration test cannot see any of those, and widening them would be the wrong
+ trade: several exist specifically in order *not* to be public.
+
+**What those 38 are actually testing:** identity minting, outstanding accounting, completion
+matching, capability gating. That is bookkeeping, not kernel behaviour. They open a ring only
+because the bookkeeping lives as fields on a struct that also owns a handle.
+
+**Whether that bookkeeping separates.** Both questions flagged as unknowns were investigated and
+both came back clean:
+
+- `RingId::next()` is a process-global `AtomicU64` and never touches a handle. Identity minting is
+ trivially handle-independent.
+- `IoRing`'s ten fields split 5/5. Kernel-coupled: `handle`, `completion_event`,
+ `registered_buffer_infos`, plus `version` and `supported_ops` which are *negotiated* from the
+ kernel and then are pure data (and already have a test seam, `set_supported_ops_for_test`). Pure
+ bookkeeping, no handle: `ring_id`, `next_user_data`, `outstanding`, `registered_files`,
+ `registered_buffers`.
+
+## The constraint this must not break
+
+[DESIGN-NOTES.md](../DESIGN-NOTES.md) already rejects a mock `IoRing`, in
+"Two techniques deliberately rejected", and the argument is a good one:
+
+> Both shipped defects were the kernel behaving differently from this crate's assumptions. A mock
+> *encodes* the assumption, so one written before those discoveries would have passed both bugs
+> green -- it would not merely have failed to find them, it would have manufactured evidence they
+> were absent.
+
+**This pass supplied three more confirmations of exactly that**, all within a few hours:
+
+| What the kernel actually does | What a mock would have been written to do |
+|---|---|
+| `SubmitIoRing` reports an expired wait as `ERROR_TIMEOUT`, a *failure* HRESULT | return success, because a timeout is not an error |
+| `SubmitIoRing` answers `E_INVALIDARG` when asked to wait with nothing pending | time out, or succeed |
+| A handle without `FILE_FLAG_OVERLAPPED` completes **inline during submit** | leave the operation pending |
+
+The first is the High-severity defect an independent review found in `M21.2`. A mock would have kept
+the suite green while every real timeout returned `Err`.
+
+So any proposal here has to survive that objection rather than ignore it.
+
+## The proposal: one suite, two backends, shared assertions
+
+Write the bookkeeping tests **once**, as generic functions over a small trait, and run them twice:
+against a hermetic in-memory implementation and against the real kernel ring.
+
+```text
+ +-- run by src/ unit tests, with FakeRing ..... hermetic
+shared suite -------+
+ (generic) +-- run by tests/ integration, with IoRing .... crosses the boundary
+```
+
+**Why this is not the rejected mock.** The rejected thing is a mock used *instead of* the kernel,
+where nothing ever checks the model against reality. Here the same assertions run against both, so
+the fake is continuously differentially tested against the kernel: the moment the model diverges,
+one side goes red and names the divergence. That is the same discipline this crate already demands
+of its spikes -- "a spike must carry a **control case**, because the first two drain-ordering spikes
+could not discriminate and would have returned confidently wrong answers". The kernel run *is* the
+control.
+
+The existing decision would therefore need **amending, not overriding**: it rejects a mock as a
+substitute, and this is a mock as a co-tested peer. That distinction is the whole proposal, and if it
+does not hold up the proposal fails with it.
+
+### The bright line: what the fake is never allowed to answer
+
+The fake models *this crate's bookkeeping*. It never gets a vote on Windows. Concretely, none of
+these may be asserted against the fake, because the fake would only be agreeing with whoever wrote
+it:
+
+- the completion event being edge-triggered ([D-19](../DESIGN-NOTES.md#d-19));
+- one waiter per ring ([D-21](../DESIGN-NOTES.md#d-21));
+- flush coverage and drain ordering ([D-23](../DESIGN-NOTES.md#d-23),
+ [D-24](../DESIGN-NOTES.md#d-24), [D-47](../DESIGN-NOTES.md#d-47));
+- the registration array being read when the operation *runs* ([D-32](../DESIGN-NOTES.md#d-32));
+- `ERROR_TIMEOUT` and `E_INVALIDARG` from `SubmitIoRing`;
+- inline completion on a synchronous handle.
+
+Those stay kernel-only, forever. They are also precisely the findings this pass produced, which is
+the argument for the line being drawn exactly here.
+
+### What the shared suite would cover
+
+Everything whose truth is decided by this crate rather than by Windows:
+
+- identity minting: monotonic, never repeated, exhaustion refused rather than wrapped;
+- outstanding accounting: minted, completed, cancelled, saturating rather than underflowing;
+- token claim matching on `user_data` *and* `ring_id`, including cross-ring rejection;
+- registration index arithmetic (the base index of a second registration);
+- capability gating -- a push refused because the op is unsupported (`set_supported_ops_for_test`
+ already exists for this and needs no kernel);
+- `pop_within`'s deadline arithmetic and its nothing-can-arrive early return, which with a fake ring
+ *and* a fake `CompletionWait` becomes fully hermetic;
+- `RingContract`'s oracle over observed sequences, which is already hermetic and would simply join
+ the suite.
+
+### Mechanics, with costs
+
+Three shapes, cheapest first.
+
+**1. Generic test functions over a narrow trait.** A trait describing only what the bookkeeping
+tests need -- mint an identity, observe a completion, read `outstanding`, claim a token. The real
+implementation wraps `IoRing`; the fake is a few hundred lines of plain Rust. Test bodies are
+`fn identity_is_never_reused(ring: &mut R)`.
+
+*Cost:* the suite must be reachable from both `src/` and `tests/`, and `tests/` cannot see
+crate-private items. The established answer in this repository is a feature-gated `pub` module --
+`windows-file-watcher` already ships `test-util` and `scenario-tool` features for exactly this. So:
+a `pub mod conformance` behind a non-default `test-util` feature.
+
+*Consequence to accept:* feature-gated code is invisible to a default `cargo test` and to
+`cargo mutants` without `--all-features`, which this repository has already been bitten by (a
+`windows-file-watcher` sweep reported 247 survivors of which 147 were in gated modules). CI would
+need the suite in its `--all-features` job, and mutation runs would need the flag.
+
+**2. Parameterise `IoRing` over a backend.** `IoRing` with a private `RingOps`
+trait over the ~10 Win32 entry points. The public spelling `IoRing` survives via the default type
+parameter.
+
+*Cost:* `Batch<'_>` becomes `Batch<'_, B>`, and every `impl` block gains a parameter. Mechanical but
+it touches the whole crate, and it puts a type parameter into a published API for a testing reason.
+Higher risk, more coverage: it would let the *submission* paths be exercised hermetically too, not
+just the bookkeeping.
+
+**3. Extract the bookkeeping into a handle-free type.** `RingAccounting` holding the five pure
+fields, composed by `IoRing`. Most of the 38 internals-touching tests become hermetic *in place*,
+with no fake, no feature gate, and no widened visibility.
+
+*Cost:* a real refactor of a shipped crate's internals, though not of its public surface. It is the
+smallest conceptual change and the one that most directly matches the observation that those 38
+tests were never about the kernel.
+
+These are not exclusive. **3 then 1** is the combination worth considering: extract the accounting so
+much of the coverage becomes hermetic without any fake at all, then add the shared suite for what
+remains and genuinely benefits from running against both backends.
+
+## Open questions for the engineer
+
+1. **Does the co-tested-peer argument actually survive?** It is the load-bearing claim. If a fake can
+ drift in a way the shared assertions do not catch -- because the assertion is about our
+ bookkeeping and the drift is in our model of Windows -- then the rejection stands and only option
+ 3 is safe.
+2. **Is a type parameter in the published API acceptable for a testing reason?** That is option 2's
+ real cost, and it is a question about the crate's public shape rather than about tests.
+3. **Is the feature-gate consequence acceptable?** Gated tests do not run by default, and this
+ repository has measured what that does to a mutation sweep.
+4. **What is the target?** "`cargo test --lib` is hermetic and means it" is achievable. "Every
+ behaviour has a hermetic test" is not, and should not be -- the six items on the bright line above
+ must stay kernel-only.
+
+## What was deliberately not queued, and what changed
+
+As first written this recorded no decision and queued no work, on the grounds that the proposal
+changes a decision that is currently written down and well argued, and that the evidence for
+amending it -- the co-tested-peer distinction -- is an argument rather than a measurement.
+
+The engineer settled the part that did not depend on that argument: **the current structure is
+incorrect regardless of which remedy is chosen**, so the defect is now [D-49](../DESIGN-NOTES.md#d-49)
+and the work is `M24`. The co-tested-peer question survives untouched as `M24.1`, which is
+required to settle it **by demonstration rather than by argument** -- precisely because an argument
+is what is in doubt.
diff --git a/crates/windows-ioring-sys/design-sessions/DESIGN-SESSION-2026-09-21-m21-remediation-findings.md b/crates/windows-ioring-sys/design-sessions/DESIGN-SESSION-2026-09-21-m21-remediation-findings.md
new file mode 100644
index 000000000..365a2ab67
--- /dev/null
+++ b/crates/windows-ioring-sys/design-sessions/DESIGN-SESSION-2026-09-21-m21-remediation-findings.md
@@ -0,0 +1,280 @@
+# Design session 2026-09-21: what remediating the epoch-log review taught
+
+A running record of findings produced **while implementing** the M21 checklist, as distinct from
+[DESIGN-SESSION-2026-09-19-epoch-log-review.md](DESIGN-SESSION-2026-09-19-epoch-log-review.md),
+which recorded the review that produced the items. It exists because the next review pass should
+start from what this one learned rather than rediscovering it -- including the places where the
+review itself was wrong.
+
+Updated as items complete. Entries are numbered `F-n` and never renumbered. M21 is complete as of
+2026-09-21; F-1 to F-12 are its whole record.
+
+## Headline 2: an independent review found what this pass could not
+
+After M21 closed, a fresh reviewer audited the `M21.2` public surface and found a **High**-severity defect
+in it, plus a pre-existing one of the same root cause, plus the reason neither was caught. All four are
+fixed in `M21.6`; `F-13` to `F-15` are what they taught.
+
+The uncomfortable part is not that the review found defects. It is *which* defect: the author of that API
+had written its tests, sabotage-verified them, and reported the sabotage in the commit message -- and the
+sabotage that mattered (delete the real wait entirely) was never run, because the tests had been
+restructured away from the real wait on purpose. **Self-review could not have found this, and did not.**
+
+## Headline 1: a review that reads code produces claims, not findings
+
+Two of this pass's corrections were to **the review**, not to the code it reviewed. Both were
+reachability claims -- statements about which states a program can enter -- and neither could have
+been settled by reading. That is the single most useful thing to carry into the next pass.
+
+- `F-5` below: finding `C-3` asserted a latent bug that measurement showed was unreachable at any
+ constants.
+- `F-8`: an assertion about how an error would surface, wrong in a way only running it revealed.
+
+A reachability claim in a review should be written as a **question with the experiment attached**,
+not as a finding. "Is this reachable if `EPOCH_SIZE` exceeds `SLOTS`? -- instrument the retry path
+and run it" would have cost the same to write and would not have needed correcting afterwards.
+
+## Findings
+
+### F-1 (M21.1) -- a sweep count taken from tool output, not from a count
+
+The blast-radius sweep was reported as "13 files" because that is what `rg`'s summary line said;
+`rg` groups two files sharing a directory prefix under one header, so the real figure was 14. The
+item's own prediction ("17 matches across 10 files") was also low on both axes.
+
+**Carry forward:** a sweep count is a claim about the tree, so it comes from a command that counts,
+per FAIL FAST rule 6. Reading it off a summary line is the same defect one layer up.
+
+### F-2 (M21.2) -- `SubmitIoRing` rejects a wait with no pending operation
+
+`SubmitIoRing` answers `E_INVALIDARG` (`0x80070057`) -- **not** a timeout -- when asked to wait for a
+completion the kernel has no pending operation for. Found because a test drove the new pop loop with
+a reservation that had no real SQE behind it. Now documented on `RingWait::block`, where the
+precondition holds structurally because `pop_within_with` checks `outstanding()` first.
+
+**Carry forward:** `IoRing::run_down` makes the same call and is reachable from `Drop`. Its
+behaviour when the count is non-zero but nothing is genuinely pending is worth a look in the next
+pass -- see `F-4`, which is that combination actually occurring.
+
+### F-3 (M21.2) -- a panic path introduced and closed in the same item
+
+The first draft of `pop_within` computed `Instant::now() + timeout`, which panics on overflow, so
+`Duration::MAX` -- a reasonable spelling of "no deadline" -- would have aborted the process. Closed
+with `checked_add` before the commit, and covered by two tests, one of which reaches the overflow
+branch rather than being answered by the early return.
+
+**Carry forward:** the next pass should check every other public entry point that accepts a
+`Duration` or a timeout for the same shape.
+
+### F-4 (M21.2) -- a failing test can abort the harness instead of reporting
+
+A test that panics while a reservation is outstanding unwinds into `IoRing::drop`, whose rundown
+then fails (per `F-2`) and panics a second time. Rust aborts on a double panic, so the run ends with
+`STATUS_STACK_BUFFER_OVERRUN` and **no test name**. Worked around here by settling every phantom
+reservation before any assertion.
+
+**Carry forward:** this is a diagnosability defect in the crate's own teardown, not only in the
+tests. A `Drop` that can panic turns any unrelated test failure in the same file into an unnamed
+abort. Worth a decision in the next pass: whether rundown failure should be reported some way other
+than `debug_assert!` while unwinding.
+
+### F-5 (M21.3) -- the review claimed a latent bug that does not exist
+
+Finding `C-3` and item `M21.3` both said the epoch-commit trigger was safe at the sample's constants
+but armed for anyone raising `EPOCH_SIZE` past `SLOTS`. Instrumenting the retry path to report when
+the old shape would have committed produced **zero** firings at `EPOCH_SIZE` of 6, 8, 12, 16 and 24.
+
+It is unreachable at any constants: the predicate is true at exactly two moments -- before the first
+append, and immediately after a commit -- and the arena is empty at both, because the commit waits
+for a covering flush that retires every outstanding write.
+
+The change was still worth making, but as a **coupling** change: the old trigger was safe because of
+an invariant three blocks away that nothing stated. Corrections landed in `C-3`, the M21 header, and
+the item.
+
+### F-6 (M21.4) -- the sample had no tests, and could not have had any
+
+Examples are not test targets by default, so `cargo test` compiled the epoch-log sample and ran
+nothing. Every claim its modules made was unbound. `test = true` on the `[[example]]` entry is what
+changed that, and it applies to the whole sample rather than to the one item that needed it.
+
+**Carry forward:** the sample's other modules -- `record`, `replay`, `reclaim`, `checkpoint`,
+`strategy` -- are now testable and still untested. `replay` is the interesting one: it is the
+verifier the sample's own credibility rests on.
+
+### F-7 (M21.4) -- the failure path is unreachable without the injection seam
+
+A commit only fails if its flush fails, and a flush against a healthy temp file does not. So four of
+the six new tests are gated on `fault-injection`, following the precedent in
+[fault_injection.rs](../tests/fault_injection.rs) -- including its reasoning that CI's
+`--all-features` job is what stops a gated test from being a test that never runs.
+
+### F-8 (M21.4) -- an error-surfacing assumption, wrong
+
+The first version of the test expected an injected `ERROR_ACCESS_DENIED` to arrive as
+`io::ErrorKind::PermissionDenied`. It arrives as `Other`: the crate preserves the HRESULT in an
+`IoRingError` rather than classifying it. Corrected to assert the Win32 code, which is what
+`tests/fault_injection.rs` already asserts.
+
+**Carry forward:** the crate does not map Win32 codes onto `io::ErrorKind`. Whether it should is a
+question for the next pass, not a defect -- but consumers matching on `ErrorKind` will match `Other`
+for everything, and nothing currently says so where a consumer would look.
+
+### F-9 (M21.4) -- the milestone compounded
+
+`commit_and_pop` is three lines because `M21.2` published `IoRing::pop_within`. Every test in the
+file would otherwise have carried its own bounded wait -- the exact duplication `M21.2` existed to
+remove, reappearing immediately in the next item.
+
+**Carry forward:** worth checking in the next pass whether the remaining hand-written waits in the
+sample and in `tests/` can now collapse onto it. `M21.5` covers two of them; there may be more.
+
+### F-10 (M21.5) -- the item named two sites; a census found six
+
+`M21.5` was written as "give strategy.rs's two wait loops a bound". Counting by command over every `.rs`
+outside `target/` and the spikes found **four** unbounded wait loops and **two** more of a related
+shape. Two of the four were helpers *both named `await_one`*, byte-identical, in
+[failure_paths.rs](../tests/failure_paths.rs) and [kernel_span.rs](../tests/kernel_span.rs) -- neither
+mentioned by the item or by the review.
+
+The other two were in [batch/tests.rs](../src/batch/tests.rs): registration waits written as a single
+`try_pop`, which is the flake shape `pop_within` documents, in a file whose **third** such wait already
+used the helper. One predicate, three sites, half-converted -- FAIL FAST rule 1 exactly, inside a single
+file.
+
+**Carry forward:** the review found these by reading one sample, so it found what that sample contained.
+A census by command is cheap and finds the population. Every item in the next pass whose subject is a
+*shape* rather than a specific line should carry its census command.
+
+### F-11 (M21.5) -- the milestone compounded again, and measurably
+
+Sabotaging `Lane::classify` to stop filing flush results leaves `await_flush` waiting for a completion
+that is never recorded. An unbounded loop hangs forever there. The new bound reported
+`timed out after 30s waiting for a commit's flush` -- **in two seconds**, because `pop_within`'s
+nothing-can-arrive early return (`F-9`, `M21.2`) answers immediately once the ring is quiesced.
+
+The bound is what makes the failure possible; `M21.2`'s early return is what makes it quick. Neither was
+designed with the other in mind.
+
+### F-12 (M21.5) -- duplicated helpers do not share a name by accident
+
+Two independently written helpers, in two files, both called `await_one`, both the same eight lines. The
+name being identical is the tell: it is what people call this operation, which is the argument for the
+operation belonging to the library. It now does.
+### F-13 (M21.6) -- the crate's tests never exercise asynchronous completion
+
+> **Corrected 2026-09-23 by `M24.6`'s sweep. The headline overstates, and it did so when written.**
+> "Every fixture in this crate's tests, examples and samples opens its handle that way" is false:
+> [flush_barrier.rs](../tests/flush_barrier.rs), [handover.rs](../tests/handover.rs) and
+> [flush_barrier_stress.rs](../tests/flush_barrier_stress.rs) all open
+> `FILE_FLAG_OVERLAPPED | FILE_FLAG_NO_BUFFERING` handles, and had done since 2026-08-28, 08-29 and
+> 09-06 respectively -- weeks before this was recorded on 09-21. What is true is the narrower claim
+> the measurements below actually support: the fixture *that finding was built against* was
+> synchronous, and so are the epoch-log sample's.
+>
+> **The carry-forward escalated from the false half, and is largely unfounded because of it.** It
+> says every claim about ordering, draining, the completion event and the barrier "was measured
+> against operations that may have completed inline", and names D-19, D-23, D-24 and D-47 for
+> re-reading. But D-23, D-24 and D-47 were measured by `flush_barrier.rs`, which is one of the three
+> overlapped, unbuffered fixtures. The entry hedged in the right direction -- "the drain-ordering
+> spike used `NO_BUFFERING` and pre-written extents deliberately, so it is probably fine" -- and then
+> checked only the spike, not the tests that shared its shape.
+>
+> The shape of the error is worth more than the correction: a measurement of **one** fixture was
+> generalised to **every** fixture without a census, and the alarm that followed inherited the
+> generalisation. A census is one command. See `M20.6`'s findings for the case where the same claim
+> *was* true -- the epoch-log sample really did run entirely on synchronous handles, which is what
+> made its strategy comparison measure a pipeline that did not exist.
+
+Measured while building a test that needed a genuinely pending operation. **A file handle opened without
+`FILE_FLAG_OVERLAPPED` is synchronous, so a ring operation against it completes inline during submit.**
+Every fixture in this crate's tests, examples and samples opens its handle that way.
+
+The consequences are larger than the item that found it:
+
+| Attempt at a slow operation | Measured |
+|---|---|
+| Buffered read, up to 256 MiB | 3-5 us -- already poppable |
+| Flush over 512 MiB of dirty cache | 3 us -- lazy writer got there first |
+| Unbuffered read, 256 MiB, synchronous handle | 3 us -- completes during submit |
+| Unbuffered **and** overlapped, 64 MiB and up | genuinely pending |
+
+So the suite has been testing the *synchronous* completion path almost exclusively. A counting waiter over
+the existing flush pattern was reached in **0 of 50 trials**.
+
+**Carry forward, and this is the big one for the next pass:** every claim this crate makes about ordering,
+draining, the completion event, and the barrier was measured against operations that may have completed
+inline. D-19, D-23, D-24 and D-47 all deserve re-reading with that in mind. The drain-ordering spike used
+`NO_BUFFERING` and pre-written extents deliberately, so it is probably fine -- but *probably* is exactly
+the word that needs replacing with a measurement.
+
+### F-14 (M21.6) -- deterministic tests and honest tests are not the same thing
+
+The M21.2 tests were restructured onto a wait that never enters the kernel, precisely to make the loop's
+deadline behaviour deterministic. That was reported in the commit message as a virtue. It was also what
+let a real defect through: replacing `RingWait::block`'s whole body with an unconditional error left the
+entire suite green, because nothing ever reached it.
+
+The isolation was correct for what it tested. The error was not adding anything that drove the real thing
+alongside it -- and then describing the isolation as coverage.
+
+**Carry forward:** when a test double is introduced to make something deterministic, the same change owes
+a test that exercises the real implementation. A mutation that deletes the real implementation should fail
+something.
+
+### F-15 (M21.6) -- an error-vs-timeout mapping is a contract, and Win32 gets it backwards
+
+Every Win32 wait reports an expired bound as a *failure* code -- `ERROR_TIMEOUT` from `SubmitIoRing`,
+`WAIT_TIMEOUT` from the `WaitFor*` family. Any wrapper that forwards its underlying result verbatim
+therefore turns an ordinary timeout into an error, and any API that documents "`Ok(None)` means the bound
+expired" is wrong the moment it does so.
+
+`CompletionWait` had not said which way to report it, so every third-party implementation would have
+reproduced the defect independently. It says so now.
+
+**Carry forward:** check every other place this crate converts a Win32 wait result. The `WaitFor*` calls in
+`event_loop.rs` and `model_b_multiplexed.rs` already handle `WAIT_TIMEOUT` explicitly; whether anything
+else forwards a wait result blindly is worth a census.
+### F-16 (M21+.1) -- the checker had a latent bug that only a probe could find
+
+Widening `check-borrow-surface.ps1` was verified with five probes rather than by re-reading the regex, and
+one of them crashed it: a one-line body -- `pub fn f() -> &[u8] { &[] }` -- never satisfied the "line ends
+with `{`" test, so the signature accumulator ran off the end of the file. The *old* script did not crash on
+that shape only because it never indexed the lines again afterwards; it silently swallowed the following
+lines instead, which means it could have been skipping real signatures all along.
+
+**Carry forward:** a checker is code, and the argument for testing it is the same as for anything else. The
+negative control matters most -- a check that fires on everything is as useless as one that fires on
+nothing, and only the plain-`&T` probe establishes that this one still discriminates.
+
+### F-17 (M21+.1) -- the blind spot had already swallowed something real
+
+The widened check immediately reported `IoRingErrorExt::as_ioring_error -> Option<&IoRingError>`, a public
+trait method returning a borrow that predates the review by months and had never been inventoried. It is
+not a hole -- the borrow is of the `io::Error` the caller owns -- but it was never *put to anyone*, which
+is the whole function of the inventory.
+
+**Carry forward:** when a check is found to be narrow, assume it has already been narrow for a while and
+look at what it let past, rather than only at the change that exposed it.
+
+### F-18 (M21+.1) -- the count sweep missed the tool built to stop drift
+
+The script's header and its failure message both said **three** shipped defects of this shape, listing
+D-35, D-36 and D-43. It has been four since D-45. `M19.3` explicitly swept that count -- its archive records
+"that file said 'three defects' in four places and is now four" -- and swept `DESIGN-INSTRUCTIONS.md` while
+missing `check-borrow-surface.ps1`.
+
+**Carry forward:** the sweep looked at documentation and not at tooling. A `.ps1` file carrying prose is
+still prose, and the next census of any restated fact should include `tools/`.
+## Open questions this pass raised but did not answer
+
+1. Should `IoRing::drop`'s rundown failure be reported some way that does not abort on unwind
+ (`F-4`)?
+2. Do the crate's other timeout-accepting entry points share `F-3`'s overflow shape?
+3. Should Win32 codes map onto `io::ErrorKind`, given that consumers currently see `Other` for
+ everything (`F-8`)?
+4. Which of the sample's now-testable modules deserve tests, and in what order (`F-6`)?
+
+None of these are queued as checklist items yet. They are inputs to the next review pass, which
+should decide whether each is work or a non-issue -- **by measuring, not by reading**, which is the
+lesson of `F-5`.
diff --git a/crates/windows-ioring-sys/design-sessions/DESIGN-SESSION-2026-09-22-kernel-response-space.md b/crates/windows-ioring-sys/design-sessions/DESIGN-SESSION-2026-09-22-kernel-response-space.md
new file mode 100644
index 000000000..085cab741
--- /dev/null
+++ b/crates/windows-ioring-sys/design-sessions/DESIGN-SESSION-2026-09-22-kernel-response-space.md
@@ -0,0 +1,217 @@
+# Design session 2026-09-22: the kernel response space
+
+**Decisions resulting from this session:** D-52 (the resolver technique and what it is for), and an
+amendment to "Two techniques deliberately rejected" in [DESIGN-NOTES.md](../DESIGN-NOTES.md). The work
+it queues is `M26` in [CHECKLIST.md](../CHECKLIST.md); `M24.1` is answered and `M24.4` is withdrawn.
+
+The apparatus built during the session is kept as
+[kernel-response-space-probe.rs](kernel-response-space-probe.rs). It is a **demonstration**, not a
+test: drop it into `tests/` and run with `--nocapture` to reproduce every figure below.
+
+## What the session set out to do
+
+`M24.1` asked one question: does a fake whose assertions are **shared** with the kernel escape the
+objection that led this crate to reject a mock? That objection is specific rather than generic --
+
+> Both shipped defects were the kernel behaving differently from this crate's assumptions. A mock
+> *encodes* the assumption, so one written before those discoveries would have passed both bugs green
+> -- it would not merely have failed to find them, it would have manufactured evidence they were
+> absent.
+
+The item insisted the question be settled by demonstration, "because the argument is exactly what is
+in doubt", and predicted two outcomes: a wrong **accounting** model would be caught by a shared suite,
+and a wrong **Windows belief** would not.
+
+Both predictions held. Then two further cases changed the answer.
+
+## Case 1 and 2: the predictions, confirmed
+
+A five-method slice (`push_one`, `submit`, `try_pop`, `outstanding`, `pop_within`), one generic
+assertion suite, two implementations -- a real `IoRing` over a temp file, and an in-memory fake.
+
+| fake's flaw | shared suite |
+|---|---|
+| none | GREEN, and the kernel agrees |
+| `outstanding` decremented at push instead of at pop | **RED** -- caught |
+| pops without submitting | GREEN -- slipped |
+| pre-`M21.6`: an expired wait is an error | GREEN -- slipped |
+
+The third flaw is the one worth dwelling on, because it is not invented: it is the belief this crate
+actually held until `M21.6`. `SubmitIoRing` reports an expired wait as `ERROR_TIMEOUT`, a *failure*
+HRESULT, so `pop_within` returned `Err` on every ordinary timeout. A fake written before that
+discovery would have encoded exactly this.
+
+Add a timeout assertion and the fake is caught instantly. But **that assertion could only be written
+after the kernel had already revealed the answer.** The fake could never have produced it.
+
+## Case 3: the argument *for* co-testing, which the item did not predict
+
+The first pass missed the case that decides it. What happens when the shared suite itself encodes the
+wrong belief, and is run against both?
+
+| | timeout assertion written from the pre-`M21.6` belief |
+|---|---|
+| kernel | **RED** -- "an expired wait returned `Ok(None)`, not the error we expected" |
+| fake built from the same belief | GREEN |
+
+The kernel **refutes** us. The fake **confirms** us. That is the manufactured-evidence mechanism made
+visible -- and it is also the escape, because running a shared assertion against the kernel is how a
+wrong belief gets contradicted. A mock-only world never performs that experiment.
+
+So after three cases the conclusion was: the rejection stands for mocks-as-substitutes, a co-tested
+peer is admissible for a narrow category, and the bright line is "accounting versus Windows
+behaviour".
+
+**That line was wrong**, and the next case is why.
+
+## Case 4: an assertion that looks like a contract and is a frozen observation
+
+Raised in review: *the kernel's behaviour is not necessarily reproducible run to run. When we observe
+it, that is an observation at a point in time, not a record of objective truth. We must not
+over-index on a record of how it runs as being "right".*
+
+Take the assertion "after submitting, the completion is already queued". It reads like a contract.
+Run it against two handles of the same API:
+
+| | "completion is already queued after submit" |
+|---|---|
+| kernel, buffered handle | GREEN |
+| kernel, `NO_BUFFERING` + `OVERLAPPED`, pre-written extent | **RED** -- "nothing was queued when submit returned" |
+| fake | GREEN -- it encoded whichever one its author saw |
+
+Opposite answers on the same API. So the assertion was never about the ring; it was about a handle, a
+filesystem and a moment. And the run-to-run half was measured the same day:
+[write-pending-spike.rs](spikes/write-pending-spike.rs)'s `NO_BUFFERING`-extending condition reported
+**5/500 in one run and 271/500 minutes later**, same binary, same machine.
+
+Restate the same question as *our* contract -- "the completion arrives within a bound we specify" --
+and all three go green. That form is robust because it is a statement about what this crate promises
+rather than about when the kernel happens to finish.
+
+### The ratchet
+
+The danger is worse than one bad test, and it compounds:
+
+1. Observe the kernel once.
+2. Freeze the observation into a conformance assertion.
+3. Build the fake to satisfy that assertion.
+4. Three artifacts now agree -- and the agreement reads as corroboration when it is **one observation
+ restated three times**.
+
+That is CONTRACT INTEGRITY rule 1 ("a hand-written second copy of a contract rule is not a check of
+the contract, it is a check of the copy") applied to platform behaviour rather than to our own. This
+crate has already paid for it once: [D-47](../DESIGN-NOTES.md#d-47) is exactly this failure, where a
+handful of runs showed a barrier holding, that was written down as a guarantee, and the real violation
+rate was nearer one in a thousand.
+
+### The corrected line
+
+Not "accounting versus Windows behaviour". The axis is:
+
+**our specified contract, versus the platform's incidental behaviour.**
+
+- `pop_within` returns `Ok(None)` on an expired wait -- **ours**. The kernel says `ERROR_TIMEOUT`; the
+ wrapper translates. Legitimate to assert, and legitimate for a fake to encode.
+- "`SubmitIoRing` reports `ERROR_TIMEOUT`" -- **an observation**. Dated, machine-specific, possibly not
+ reproducible. It belongs in a spike with a rate and provenance, never as a pass/fail assertion.
+- "the completion is queued when submit returns" -- **an observation wearing a contract's clothes**,
+ which case 4 demonstrates.
+
+This is why Design Autonomy is a repository rule: we define our behaviour and choose dependencies that
+satisfy it. The wrapper is the layer that absorbs kernel variation, and the conformance suite asserts
+what the wrapper promises -- which is precisely what a fake can faithfully implement.
+
+## Case 5: the reframing, and the actual answer
+
+Also raised in review, and it is a different technique rather than a refinement:
+
+> Our code needs to work in light of all the possible ways that the platform may respond to the rings
+> we submit. Some seed-derivable selection of a set of resolutions to what a given epoch's IoRing
+> would turn into in terms of synchronously versus asynchronously completed items. Input state plus a
+> seed gives a set of kernel responses, and we verify that `windows-ioring-sys` responds correctly to
+> the kernel stimuli.
+
+The fake stops modelling **what Windows does** and starts modelling **what Windows is permitted to
+do**. A seed picks one resolution out of that space: which operations finish inside `SubmitIoRing`
+and which pend, in what order completions are posted, which fail.
+
+**This dissolves the mock objection rather than working around it**, because there is no belief to be
+wrong about. The resolver asserts nothing about the kernel. The assertions are about *us*: does this
+crate behave correctly under this resolution.
+
+It also inverts the problem case 4 raised. Non-reproducibility stops being a threat and becomes the
+expected case -- Windows exercising a different point in a space the tests already sweep. A run-to-run
+change like 5/500 to 271/500 is two samples from a space covered by construction.
+
+A minimal resolver in the probe -- SplitMix64 over one seed, permuting completion order -- was run
+against a consumer that assumes completions arrive in submission order:
+
+```
+a consumer assuming FIFO completion order:
+ broke under 189 of 200 seeds
+ first at seed 0: completions arrived as [3, 1, 4, 2], not [1, 2, 3, 4]
+```
+
+Some seeds pass and some fail, which is the point. A fixed fake reports whichever single answer it
+encoded.
+
+## What this does and does not add to the toolkit
+
+[DESIGN-NOTES.md](../DESIGN-NOTES.md) already draws the boundary this sits on:
+
+> All five techniques check this crate's code against **this crate's stated contract**. None of them
+> can tell you the stated contract is wrong.
+
+The resolver is a sixth technique, and it does something none of the five do: it checks the code
+against a **space** of platform behaviours rather than against one. It still cannot tell you the
+stated contract is wrong -- only a spike does that. What it can tell you is that the code is brittle
+to variation *inside* the space, which nothing in the toolkit currently detects.
+
+It composes with what exists rather than replacing it.
+[generated_sequences.rs](../tests/generated_sequences.rs) (M17.3, `D-41`) already generates the
+**input** space and runs it against the real kernel; the resolver generates the **response** space.
+Both seeded, both replayable from one number, and together a two-dimensional exploration.
+
+## The hard part, which is the whole design
+
+**Where the permitted space comes from.** Derive it from observation and the trap closes again. It has
+to be a deliberate specification -- "we will tolerate these behaviours" -- written wider than anything
+observed, on purpose. That makes this crate's model of Windows an explicit, reviewable, versioned
+artifact instead of an accident of whichever machine ran the tests last.
+
+Three consequences that should be decided rather than defaulted:
+
+1. **Constraints must be modelled too, or the tests demand over-defensive code.** `D-23`'s
+ covering-flush guarantee held with zero failures in roughly 4,500 trials. If the resolver may
+ violate it, we would write code defending against a kernel that breaks a documented guarantee.
+ Making that an explicit call is the improvement; it is still a call.
+2. **The seam is invasive.** A resolver has to sit under the `windows-sys` calls -- `SubmitIoRing`,
+ `PopIoRingCompletion`, the `Build*` family -- which is substantially more than `M24.2`'s field
+ split, on a published crate.
+3. **Kernel tests do not go away; their job changes and improves.** They stop being "run everything
+ against Windows" and become "confirm reality stays *inside* the declared space". Better defined,
+ and if Windows ever moves outside it, that test is what reports something real.
+
+## What it would have caught, stated as reasoning rather than measurement
+
+The FIFO consumer in case 5 is structurally the same defect as `D-47` -- a consumer assuming an
+ordering the platform does not guarantee -- so that class is covered. `M21.6`'s `ERROR_TIMEOUT` defect
+is covered provided the space includes "a wait may expire and report it as a failure HRESULT".
+
+**Both are analogies from the demonstration, not separate measurements.** `M26` carries a calibration
+item to re-inject both defects against a real resolver, because `D-41`'s corollary is the most
+transferable rule this crate has produced: *a green result from an instrument nobody has shown can go
+red is not evidence.* This session produced two apparatus failures of exactly that kind -- a first
+draft that ran each spike condition once, and a case-4 harness whose bare flush completed inline on
+both handles and so could not discriminate until it was rebuilt around aligned writes.
+
+## Consequences for the plan
+
+- `M24.1` is **answered**: the fake was the wrong instrument. Recorded, and the decision amended.
+- `M24.4` is **withdrawn**. A shared conformance suite over a hand-written fake is superseded by the
+ resolver, and its stated purpose -- hermetic bookkeeping tests -- is served better by one.
+- `M24` **no longer depends on either**. Its goal is a hermetic lib suite, and relocation plus the
+ accounting extraction achieve that on their own. It is now unconditional.
+- `M26` is new, and is justified by **what it catches** rather than by hermeticity. That is a real
+ distinction: hermeticity is achievable without it, so the resolver has to earn its place on the
+ defect class it detects.
diff --git a/crates/windows-ioring-sys/design-sessions/DESIGN-SESSION-2026-09-23-pending-inventory.md b/crates/windows-ioring-sys/design-sessions/DESIGN-SESSION-2026-09-23-pending-inventory.md
new file mode 100644
index 000000000..ca0280a6c
--- /dev/null
+++ b/crates/windows-ioring-sys/design-sessions/DESIGN-SESSION-2026-09-23-pending-inventory.md
@@ -0,0 +1,183 @@
+# Design session -- the pending-token inventory (2026-09-23)
+
+**Summary.** An exploration of `M23.3`, which asks whether this crate should offer a
+pending-operations map and a slot arena over it. The session produced a working spike
+(`src/pending.rs`), three measured findings that falsified earlier claims -- two of them
+mine, stated confidently and wrongly -- and a design alternative I dismissed on a reason
+that turned out not to hold. No decision is taken here; `M23.3` still owns that.
+
+## What the checklist said, and what a census found
+
+`M23.3` recorded that nine sites keep a map from `UserData` to an unclaimed `Token`. A
+fresh census found otherwise:
+
+- **There are about twelve**, not nine. The list missed `model_a_delivery.rs`,
+ `model_b_multiplexed.rs` and `generated_sequences.rs`.
+- **Only about a third keep the bare map described.** The rest carry per-operation
+ sidecar data -- a slot index, an expected length, a phase, a sequence number, a round.
+ `flush_barrier_stress.rs` keeps two sidecar maps beside its tokens.
+- **Exactly one site needs synchronisation**: `model_a_delivery.rs` wraps its map in a
+ `Mutex`, because it is the Model A path where completions land on pool threads.
+
+The count was the item's main evidence, and it was taken before `M24` relocated eleven
+tests. **The duplicated thing is not the map**; it is the claim discipline over a map
+whose value type differs at nearly every site. That is why the spike is `Pending`
+and not `Pending` -- the generic is a finding, not a convenience.
+
+## The two types already existed, related by hand
+
+`RingContract` is public, always on, and not feature-gated. Every consumer that uses it
+drives it *in parallel with* its own map:
+
+```text
+self.contract.observe_push(token.id());
+self.in_flight.insert(token.id(), InFlight { token, slot });
+```
+
+The same event, recorded twice, by hand, at every site -- a restatement in the
+repository's own terms, and one that can drift in both directions. The oracle's value
+depends on being driven correctly by the very code it exists to check.
+
+So the pair the engineer was reaching for -- "if the rule is heavy, can we have two
+related types?" -- already exists. What was missing is that the light one should
+**drive** the heavy one, so a consumer updates one thing and both stay true.
+
+## Synchronisation: no, and the reason is already paid for
+
+`pop_within` takes `&mut self`, and `IoRing` is `Send` but **not `Sync`** -- there is only
+`unsafe impl Send for IoRing`. Whoever pops a completion already holds exclusive access,
+so a map reachable through that same `&mut self` needs no `Arc`, no `Mutex`, and no
+interior mutability. The exclusivity exists already.
+
+## Falsified: the type-erasure objection
+
+**This was my claim and it was wrong.** I argued a map owned by the ring would have to
+store heterogeneous `Token`, forcing `Box` and a downcast at the claim site,
+and that this killed the idea. The engineer asked why, since for any particular ring the
+type is fixed. Checking:
+
+- **Per-ring monomorphisation holds for every real consumer.** The epoch-log's log ring
+ carries `Token` for appends, and its commits are *tokenless*.
+ `checkpoint.rs` looked like a counterexample because it has two maps, but its local
+ `Pending` is plain bookkeeping and only one map holds tokens.
+- **Where it genuinely does not hold, the answer is a closed enum, not `dyn Any`** -- and
+ this tree already has one. `generated_sequences.rs` puts eight token types on a single
+ ring behind `enum Held`, with an exhaustive `match` in `claim` and no runtime type
+ check at all.
+
+So a generic `IoRing` owning the map is viable, which matters because it is the shape
+that would make the ring *notify* rather than be told -- and drift between the ring and
+the inventory structurally impossible rather than merely discouraged. Its real costs are
+different from the one I asserted: `IoRing` is not generic today, so this is a breaking
+change to a published crate; a consumer mixing shapes writes a `Held`-style enum; and
+tokenless pushes still need a story.
+
+## Why `commit.rs` uses `flush_raw`, and what follows
+
+Asked during the session, and the answer explains the tokenless commits above.
+
+The safe `flush(&F, ..) -> Token` requires an **owned, guarded**
+file -- `SharedFile` is `Arc`, and the token holds a clone of that `Arc`,
+which is what keeps the handle alive for the kernel. `Committer::commit` is handed a bare
+`RawHandle` that the log's `File` owns, so it cannot build a `SharedFile` without taking
+ownership and closing the log's handle out from under it. It therefore takes the `unsafe`
+`flush_raw`, whose safety argument is hand-written: *"`file` is the log's own handle and
+outlives every operation pushed here; the log drains to empty before it closes."*
+
+**Why that makes the commit tokenless, precisely.** A write's token guards the *buffer*,
+so `write_registered_raw` still returns one even with a raw file. A flush has no buffer --
+its only possible guard is the file -- so with a raw handle there is nothing to hold, and
+`flush_raw` returns a bare `usize`.
+
+Three implications:
+
+1. **The commit path cannot be in any token inventory as currently plumbed.** Not because
+ flushes are special, but because this sample passes a borrowed handle.
+2. **A compile-time guard was traded for a prose argument**, and the prose is load-bearing:
+ nothing enforces "the log drains to empty before it closes".
+3. **It is the same shape as `placement.rs`'s constraint**, found earlier the same day: a
+ borrowed `RawHandle` and an owned handle make different APIs reachable. Two independent
+ places in one sample reach for a lower-level call for the same plumbing reason, which
+ suggests the plumbing rather than the calls is the thing to look at -- and `M25.3`
+ already reopens how the log is opened.
+
+## What the spike established, and what it cannot
+
+**Established, by test and by sabotage:**
+
+- One call site keeps the map and the oracle in step. Cutting the wiring is caught.
+- An unclaimed token is loud at teardown rather than a silent deliberate leak. Removing
+ the `Drop` guard is caught.
+- `Pending` fits a real consumer: converting `append.rs` removed its `InFlight`
+ struct and its hand-driven `observe_completion`/`observe_claim` pair.
+- The `M22.2` ordering defect is now caught by an assertion, via a test that drives a
+ failed write through the injection seam.
+
+**Not established, and two of these are corrections to things I asserted:**
+
+- **The ring does not notify anyone.** `Pending` is consumer-driven. Nothing forces a
+ minted token into the inventory, so a consumer can take a `Token` from `Batch` and never
+ register it, and no type, test or oracle notices. The spike removed drift between the map
+ and the oracle; it did not remove drift between the ring and the map.
+- **One consumer is converted, not twelve.** This is a worked example, not a property of
+ the crate.
+- **`Pending::checked()` owning the oracle creates a decoy hazard.** A consumer that
+ already had a `RingContract` keeps a field that is never written to again. The
+ conversion did exactly that, and its teardown `assert_quiescent()` passed *vacuously* --
+ compiled, ran, every test green. Sabotage confirms nothing catches it. Found by reading.
+- **Ring teardown interaction is untested.** If a `Pending` and an `IoRing` both drop with
+ operations outstanding, the ordering of their guards is unexamined -- and the session
+ already found one drop-order surprise (see below).
+
+## Two findings about the instruments themselves
+
+**A feature-gated test is invisible to the sabotage harness.** The sweep first reported the
+`M22.2` regression as *survived*. The new tests are behind `fault-injection`, off by
+default, so the harness compiled them out and the sabotage landed in code nothing
+exercised. Fixed with a manifest `testArgs` carrying `--all-features`. This is the trap the
+repository already documents for cargo-mutants, arriving through a different tool: a
+feature-gated guard and an absent guard are indistinguishable to a runner that does not
+enable the feature.
+
+**A failing test that leaves registered buffers outstanding aborts instead of reporting.**
+`RegisteredBuffers::drop` refuses to free while operations are outstanding -- correct, from
+`M5.3` -- via a bare `debug_assert!` that does not check `std::thread::panicking()`. So the
+assertion fires and names the leak, then unwinding drops the arena, the `debug_assert`
+panics during unwind, and the process aborts with `STATUS_STACK_BUFFER_OVERRUN`. Detection
+is not weakened; the report is. Queued as `M23.4`.
+
+## On requiring `finish`, and why it is not failable
+
+Rust has no linear types, so nothing can force a method call on a value the caller owns.
+Three rungs were considered:
+
+- **`#[must_use]`** stops the return value being ignored, not the call being skipped.
+- **The drop bomb** -- `Drop` panics when tokens are still held -- is the real enforcement,
+ and is what the spike implements. It is suppressed while already panicking, because a
+ second panic during unwind aborts and replaces the original failure.
+- **A scoped constructor** (`Pending::scope(|p| ...)`) would genuinely force it, since the
+ consumer never owns the value. Not built: it imposes a control-flow shape that suits the
+ appender but not the tests that thread a map through several helpers.
+
+**`finish` reports rather than fails**, and the engineer's observation is why: it *consumes*
+the map, so an `Err` would leave nothing to retry with. A consuming method in a linear
+discipline is normally total for exactly that reason. `-> Result<(), _>` would buy
+`#[must_use]` from the language rather than from an attribute, at the cost of implying a
+recoverable state that does not exist; the spike keeps `-> Vec` with the
+attribute.
+
+The residual worry -- that a path is found so late that a drop bomb ships -- is real but
+narrower than it looks. A failed completion is not unreachable, only untested, and
+`Completion::with_injected_failure` reaches it on demand. What remains is the set of
+conditions the seam cannot manufacture, which is worth naming rather than treating the
+whole class as unreachable.
+
+## Open, for `M23.3` to decide
+
+1. Public type, documented pattern, or `test-util` module. Six of the twelve sites are
+ tests, and test convenience is a weak reason to grow permanent surface.
+2. Whether `checked()` survives in its current form, given the decoy hazard nothing catches.
+3. Whether the stronger shape -- a generic `IoRing` owning the map, so the ring notifies
+ -- is worth a breaking change to a published crate.
+4. What it refuses to decide. Batching, ordering and slot choice are caller questions; the
+ sharper refusal is that the map does not decide whether you are checked.
diff --git a/crates/windows-ioring-sys/design-sessions/kernel-response-space-probe.rs b/crates/windows-ioring-sys/design-sessions/kernel-response-space-probe.rs
new file mode 100644
index 000000000..eaaa15c5d
--- /dev/null
+++ b/crates/windows-ioring-sys/design-sessions/kernel-response-space-probe.rs
@@ -0,0 +1,645 @@
+// Copyright (c) 2026 Mike Grier
+//! **Throwaway probe for `M24.1`.** Not a permanent test -- it exists to
+//! settle one design question by demonstration, and is deleted once the
+//! decision is recorded.
+//!
+//! # The question
+//!
+//! [`DESIGN-NOTES.md`]'s "Two techniques deliberately rejected" refuses a mock
+//! `IoRing`, because a mock *encodes* an assumption: one written before this
+//! crate's two shipped defects were found "would not merely have failed to
+//! find them, it would have manufactured evidence they were absent".
+//!
+//! `M24` wants a hermetic unit suite, and the remedy it would most like is a
+//! fake whose **assertions are shared** with the kernel -- one conformance
+//! suite, run against both. `M24.1` asks whether sharing escapes the
+//! objection, and insists it be settled by demonstration rather than argument,
+//! because the argument is what is in doubt.
+//!
+//! # The two predictions
+//!
+//! 1. Give the fake a wrong **accounting** model. The shared suite must go red
+//! on the fake and stay green on the kernel.
+//! 2. Give the fake a wrong **Windows belief**. The shared suite must *not*
+//! catch it -- the expected result, and the reason `D-49`'s bright line is
+//! a finding rather than a hedge.
+//!
+//! A co-tested peer is defensible only if both halves behave as predicted.
+//! Run with `--nocapture` to read the verdicts.
+
+use std::io;
+use std::os::windows::fs::OpenOptionsExt;
+
+use windows_ioring_sys::{
+ Batch, FlushCoverage, FlushMode, IoRing, NumaBuffer, PushOptions, SharedFile, WriteCaching,
+};
+
+/// The slice of ring behaviour this probe shares between the two peers.
+///
+/// Deliberately the *handle-free accounting* `M24.2` measured as separable:
+/// push something, submit, pop it, and keep an outstanding count. That is the
+/// part a fake could legitimately own.
+trait RingLike {
+ fn push_one(&mut self) -> io::Result;
+ fn submit(&mut self) -> io::Result<()>;
+ fn try_pop(&mut self) -> io::Result