diff --git a/.github/copilot-instructions.md b/.github/copilot-instructions.md index c617efb26..423eefcdc 100644 --- a/.github/copilot-instructions.md +++ b/.github/copilot-instructions.md @@ -115,6 +115,52 @@ baseline truly unaffordable, that is a **decision for the engineer driving the w raised explicitly per the PRIME DIRECTIVE's blocker protocol — never a shortcut an assistant takes unilaterally in the name of efficiency. +## OPTION INTEGRITY — enable choices; foreclose only on analytic grounds + +**Our job is to enable options.** Unless an option can be shown to have *no possible +value*, expose it, and give clients the tools to choose it when it applies and to gather +the data that makes the choice well-founded. The default posture toward a design +alternative is to keep it and instrument it, not to rank it. + +**The reason this repository needs the rule more than most is recorded separately**, in +[DESIGN-NOTES.md](../DESIGN-NOTES.md) → [The adoption thesis](../DESIGN-NOTES.md#the-adoption-thesis): +the hardware that would make these tradeoffs measurable is not available here, and the +people best placed to judge the options are application authors we have not met. Read it +before arguing that a particular option is safe to drop — the two conditions it names are +temporary in principle and are not temporary in practice. + +**A measurement that failed to realise an option's value is not a finding against the +option.** It may mean the hardware, the workload, the software configuration, or the +apparatus could not reach the conditions where the value appears. Saying "we measured it +and it did not help" is a statement about the *measurement*; converting it into "it does +not help" is a category error, and it is the one this rule exists to stop. + +**This binds hardest where there is literature or prior art suggesting conditional +applicability.** Queue and ring topologies, cache and NUMA placement, batching strategies, +and scheduling disciplines all have regimes where each choice wins. An in-repo measurement +on one machine cannot settle a question the field treats as workload-dependent, and this +repository's own decisions (for example that one ring per thread is userspace's proxy for +one ring per CPU, with NVMe queue pairs as the hardware reason) frequently *are* that prior +art. Contradicting a recorded decision on the strength of one sample's configuration is a +defect, not a finding. + +**What does justify removing or narrowing an option** is a clear analytic result, of the +kind that can be argued from the code rather than from a run: + +- fewer instructions, fewer allocations, fewer I/Os issued; +- better locality of reference, argued structurally; +- a smaller support burden or a clearer programming model; +- a demonstrated *wrong answer* — an option that binds to incidental behaviour, produces + overlapping domains, or silently reports success while doing the wrong thing. + +The last is the honest ground for most removals here: not "it measured slower" but "it +computes the wrong thing." + +**When an option stays but cannot be shown to pay, say exactly that**, and say what would +change the answer: which conditions the apparatus could not reach, and what a consumer +would need to measure on their own hardware. Foreclosing costs a client a choice they may +have needed; keeping an unproven option costs a paragraph. + ## Line endings in tool parameters All text content passed to tpu tools (`content`, `replacement`, `data` in edit ops) is @@ -661,7 +707,7 @@ When executing checklist items (CHECKLIST.md files): - **If items must be done together, say so and do it; don't tease apart.** Once you have decided (and recorded in the checklist if the structure is wrong) that two items must land together, commit them together in one commit citing both IDs. Do **not** try to "unthread" a coupled implementation into per-item commits after the fact — that is fiction, not history. - **Commit immediately after each item.** In mode b (implementing forward), the commit must happen before moving to the next item. In mode a (recording already-finished work), a single commit citing all the item IDs satisfies this. - **Commit message format: a Conventional Commits subject line, with the checklist trailer in the body.** - `release-please` (see "Release process" in [DEVELOPMENT.md](DEVELOPMENT.md)) drives every crate's version + `release-please` (see "Release process" in [DEVELOPMENT.md](../DEVELOPMENT.md)) drives every crate's version bump and CHANGELOG **only** from Conventional Commits subject lines (`type(scope)!: summary`); a subject that doesn't match that grammar is invisible to it, no matter how much checklist work the commit records. The mandatory `Completed item:` provenance is therefore never the subject line — it moves to the body, and @@ -1142,6 +1188,65 @@ If a plan exceeds roughly 10 work items or 3 levels of grouping/nesting, checkpo into a CHECKLIST.md file in the repository before continuing. The goal is that the plan survives a lost session — if the plan only exists in the chat, it will be lost. +## RESOLUTION GRADIENT — sharp at the front, deliberately coarse behind, and never manufacture certainty + +**A plan is written at decreasing resolution with distance from the present.** The current +milestone has great resolution. Later milestones are progressively coarser, and that +coarseness is **correct** — it is not an omission to be closed, and an audit or review pass +must not treat it as one. + +**Why it cannot be otherwise here.** Some work has the shape *build the blocks → build the +measurement tools → experiment with those tools to infer things*. On such a project the +later milestones are not merely unwritten, they are **unwritable**: the experiments that +would resolve them have not happened. The clarity is an **output** of the work, not an +input being withheld from it. A project small enough to plan end-to-end before +implementing is a different case, and the distinction is worth making explicitly before +planning begins. + +**The failure mode this exists to stop.** An assistant asks a *very specific* question of +someone who holds a *general sense* of the direction. The specificity of the question +implies an answer of matching precision is available, so one is produced — at low +confidence. It is then recorded as a decision, and it lands in the wrong milestone, or in +the wrong order, and later work binds to it. **A low-confidence answer recorded as a +decision is worse than no answer**, because the uncertainty that surrounded it is now +invisible to everyone downstream. + +Four rules follow: + +1. **Calibrate the question to the resolution actually available.** Ask whether the general + direction is right before asking which of five options to take. If a question would only + be answerable *after* work that has not been done, it is not yet a question — it is a + description of that work. +2. **Make "too early to say" a first-class, explicitly offered answer.** When presenting + options, say plainly that leaving it coarse is among them. A question posed without that + option is a question that forces a choice, and the person answering may not notice they + have been forced. +3. **When an answer arrives hedged, record the hedge.** A direction that is not settled is a + **working position, not a decision**: it gets no decision ID, it lives in Tier 2 or Tier 3 + or a heading that says so, and nothing binds to it. The worked example already in the tree + is "Working position on domain counts (not a decision)" in + [DESIGN-SESSION-2026-08-30-numa-sharded-io-execution-domains.md](../design-sessions/DESIGN-SESSION-2026-08-30-numa-sharded-io-execution-domains.md). +4. **Prefer questions that unblock the current milestone.** If the answer would not change + what happens next, asking now mostly converts uncertainty into a record of false + precision. + +**A deferral is productive, not merely protective.** Naming a deferral is usually read as +"we avoided building on a guess", which is true and is the smaller half. The larger half is +that it **buys the interval in which the answer becomes derivable** — the blocks get built, +the instruments get written, the experiments get run, and the answer that was unavailable +becomes obvious. So when a deferral discharges, do **not** write it up as though the answer +existed all along and was waiting to be stated. Say what in the interval produced it. The +difference matters because the first framing quietly teaches that asking earlier and harder +would have worked, which is exactly the behaviour rule 1 forbids. + +**This does not soften the PRIME DIRECTIVE, and the two must not be confused.** They govern +different objects. The PRIME DIRECTIVE forbids deferring **work** because no consumer for it +is currently visible; this rule forbids manufacturing **decisions** the work has not yet made +available. Building a capability nothing calls yet is required; inventing a specific answer to +a question the experiments have not reached is not. When they appear to collide, the test is +whether the thing being deferred is *work you could do now* — if it is, do it, and the +gradient has nothing to say about it. + ## Design notes are not a work queue Design notes (DESIGN-NOTES.md, DESIGN-RATIONALE.md, and related files) record *decisions* @@ -1304,7 +1409,7 @@ sites in three wordings. **This is the data-side twin of rule 1.** Rule 1 says define a fact once in code and have everything ask. This says the same of measurements: hold the number once, and have prose point rather than -paraphrase. +paraphrase. Rule 6 extends it once more, to facts that are *derived* rather than measured. ### 5. Present what was observed; never write the conclusion @@ -1351,6 +1456,57 @@ crate that happens to publish measurements. Every instance found so far has been existing decision rather than a gap in it. Apply it while writing: no checker can find these, because nothing is inconsistent. +### 6. Never store a fact another artifact already owns + +Rule 4 governs *measured* numbers. This governs every **derived** fact -- anything a reader could get +from an artifact that is already authoritative for it. Release or publication status, version numbers, +which milestones are done, whether a branch has landed, how many crates or tests or files there are. +Writing one into prose creates a second copy whose only maintenance mechanism is somebody remembering, +and remembering is what fails. + +The tell is that **the copy cannot be wrong at the moment it is written.** It is accurate -- that is why +it gets written -- and nothing will ever say when it stopped being. A wrong decision gets argued with; a +stale derived fact is simply believed. + +- **Delete rather than update.** When you find a stale derived fact, correcting it is almost never the + fix: it re-arms the identical hazard with a fresh date on it. Remove the claim and link the artifact + that owns the answer. +- **Removing the digits is not enough.** "Published at 0.3.1" and "is published" are both copies of the + release state; only the first is obviously one. Rule 4's "write the claim, not the digits" shrinks the + drift surface of a *measurement whose claim is itself the finding*. It does not license storing a + derived fact in words. +- **An absence may be worth one sentence, once.** Where a reader would expect a status section and find + none, say the omission is deliberate and name the artifact that answers it -- otherwise somebody + helpfully adds it back. +- **A characterisation of a sibling item is a derived fact too, and this is the clause that was + missing.** "`M22` is a testing-heavy milestone", "`M23.1` touches the crate's contract surface", + "those tests only use public API" -- each summarises an artifact that already says what it is, and + each is wrong the moment that artifact changes or was misread in the first place. **Link the item; + do not describe it.** Measured cost of the omission: both examples above are real, both were + written into a milestone's rationale in one session, and both were false when checked -- `M22` is + example-only and `M23.1` names the *sample's* `contract.rs`, not the crate's. + This is the harder half of the rule to apply, because such a claim arrives as a *subordinate + clause supporting an argument* rather than as a statement of fact. "X, because Y is Z" reads as + connective tissue; `Y is Z` is nonetheless an assertion about the tree, and the reflex that fires + on "I am about to write a version number" does not fire on it. Treat the word **because**, + followed by anything about another file, item or milestone, as the tell. +- **This does not reach the primary record.** Decisions, measurements, rationale, design intent, and a + checklist's own contents are owned here and belong here. The test is simply whether some other + artifact is already authoritative: if yes, point at it; if no, this *is* the artifact. + +**FAIL FAST rule 6 is the sibling, not a contradiction.** That rule says a claim that counts or +enumerates repository artifacts must come from a command rather than from recollection. This is the +prior question -- prefer not to state it at all. Bind it to a command only when the claim must exist +anyway, such as a test asserting a property of the tree. + +Worked example, and the reason this is written down: `windows-ioring-sys`' design notes opened with +"This crate does not exist yet as compiled code", and its published rustdoc said "Under construction", +several releases after the first one shipped. The first attempt at a fix replaced both with a carefully +drift-minimised status paragraph -- no version number, linking `CHANGELOG.md` and the checklists -- and +that was still wrong, because "is published" is itself a copy of the release state. What the crate's +status is, is a question `CHANGELOG.md` and the git tags answer. The notes now record that they +deliberately do not answer it. + ## FAIL FAST — push every rule to the earliest rung that can enforce it CONTRACT INTEGRITY above tells you to keep restatements in step. This tells you where to put the diff --git a/.github/workflows/ci.yml b/.github/workflows/ci.yml index 3d399375e..b9d4943e2 100644 --- a/.github/workflows/ci.yml +++ b/.github/workflows/ci.yml @@ -153,6 +153,23 @@ jobs: shell: pwsh run: ./tools/check-borrow-surface.ps1 + # The same mechanism, for a different population that also grew unnoticed: + # lib tests that open a real kernel ring. D-49 found 63 of them, which made + # `cargo test --lib` an integration suite wearing a unit suite's name. M24 took + # it to 41 and none of those is movable without M26.2, so a zero-check would + # fail on day one -- hence an inventory, which fails when the SET changes. An + # addition obliges the question "does this need the kernel, or only a + # ring-shaped thing?"; a removal is progress and needs only regeneration. + # Needs no toolchain -- it only reads files. + ring-test-population: + name: ioring ring-opening lib tests + runs-on: windows-latest + steps: + - uses: actions/checkout@v7 + - name: Run check-ring-tests.ps1 + shell: pwsh + run: ./tools/check-ring-tests.ps1 + build-test: name: build + test runs-on: windows-latest @@ -512,6 +529,35 @@ jobs: RUST_BACKTRACE: 1 RUST_LIB_BACKTRACE: 1 run: cargo test -p windows-ioring-sys --locked --no-fail-fast + # `kernel-seam` (M26.2) is orthogonal to `threadpool`, so the two gates + # multiply: `--all-features` above builds the seam only alongside the + # threadpool, and every step in this job so far builds the no-threadpool + # path only with the seam off. The combination is a published + # configuration nothing else selects, which is the same argument that + # bought this job -- paid here rather than assumed. + - name: cargo clippy (kernel-seam, no threadpool) + run: cargo clippy -p windows-ioring-sys --all-targets --no-default-features --features kernel-seam --locked -- -D warnings + - name: cargo test (kernel-seam, no threadpool) + env: + RUST_BACKTRACE: 1 + RUST_LIB_BACKTRACE: 1 + run: cargo test -p windows-ioring-sys --no-default-features --features kernel-seam --locked --no-fail-fast + # `cargo test` compiles the epoch-log sample as a test harness and never + # calls `main`, so until M25.1b nothing ran the program itself. That was + # not theoretical: M25.1 converted the log's writer to a strided layout + # without the reader, which left the log unreadable while every one of + # the example's tests passed. + # + # M25.1b put the contract checks in tests, which is the rung that runs on + # every developer's machine. What a test cannot reach is `main` itself -- + # its path setup, its error plumbing, and its exit code -- and that is a + # published example a consumer runs, so a panic on startup is exactly the + # failure worth catching. Release, because the sample takes about a second + # there against several in debug. + - name: cargo run (epoch-log sample, end to end) + env: + RUST_BACKTRACE: 1 + run: cargo run -p windows-ioring-sys --example epoch_log --release --locked placement-probe-no-serde: name: windows-placement-probe (no serde feature) diff --git a/.github/workflows/publish-crate.yml b/.github/workflows/publish-crate.yml index ed918a497..ade7254e9 100644 --- a/.github/workflows/publish-crate.yml +++ b/.github/workflows/publish-crate.yml @@ -4,6 +4,7 @@ name: publish-crate on: push: tags: + - 'win-numa-sys-v*' - 'windows-file-enumeration-sys-v*' - 'windows-file-watcher-v*' - 'windows-file-watcher-example-test-harness-v*' @@ -29,6 +30,7 @@ on: required: true type: choice options: + - win-numa-sys - windows-file-enumeration-sys - windows-file-watcher - windows-file-watcher-example-test-harness @@ -109,7 +111,7 @@ jobs: - name: Wait for workspace-sibling dependencies on crates.io shell: bash run: | - workspace_crates="windows-file-enumeration-sys windows-file-watcher windows-file-watcher-example-test-harness windows-impersonation-token-sys windows-ioring-sys windows-namespace-request-sys windows-overlapped-io-sys windows-thread-ambient-sys windows-threadpool-sys windows-topology-sys windows-waitable-queues wtf-string" + workspace_crates="win-numa-sys windows-file-enumeration-sys windows-file-watcher windows-file-watcher-example-test-harness windows-impersonation-token-sys windows-ioring-sys windows-namespace-request-sys windows-overlapped-io-sys windows-thread-ambient-sys windows-threadpool-sys windows-topology-sys windows-waitable-queues wtf-string" metadata="$(cargo metadata --no-deps --format-version 1)" # `tr -d '\r'` is load-bearing on the Windows runner: jq.exe writes # CRLF, and `read` splits on LF alone, so without this the last field diff --git a/.release-please-manifest.json b/.release-please-manifest.json index 2f768ec86..3bec1095c 100644 --- a/.release-please-manifest.json +++ b/.release-please-manifest.json @@ -5,6 +5,7 @@ "crates/windows-file-watcher-example-test-harness": "0.1.3", "crates/windows-impersonation-token-sys": "0.1.1", "crates/windows-ioring-sys": "0.3.1", + "crates/win-numa-sys": "0.0.0", "crates/windows-overlapped-io-sys": "0.1.3", "crates/windows-namespace-request-sys": "0.2.1", "crates/windows-thread-ambient-sys": "0.2.0", diff --git a/CHECKLIST-io-domains.md b/CHECKLIST-io-domains.md index 74d7948a3..1f925a8df 100644 --- a/CHECKLIST-io-domains.md +++ b/CHECKLIST-io-domains.md @@ -466,6 +466,45 @@ Parked, not pending. Shape recorded so it is not lost, per the `M{n}+` conventio Carry one constraint from the start: the flush barrier stops at the ring's edge, so **an epoch is per-domain** and a client spanning two domains needs two flushes and an explicit join. + **Musing, recorded not prioritized (the engineer, 2026-09-23).** If a consumer could *tell* this + layer which of its files share a **flush regime** -- that is, which commits will contend -- that + might be useful. It is the "declare rather than discover" shape `S-3` proposed for the storage node, + applied to the co-flush group instead. Note the wording: *not* "which files share a device", because + the inference from one to the other is exactly what the handover below calls unsound. Fit it in if + it falls out naturally; do not build toward it. It belongs here rather than in `windows-ioring-sys` + because co-flush *grouping* is reasoning about durability groups, and that crate has none -- see + [windows-ioring-sys/DESIGN-NOTES.md](crates/windows-ioring-sys/DESIGN-NOTES.md#d-54). + + **Handed over from `windows-ioring-sys` M23.2(b) on 2026-09-23 under D-54, and sharpened on the way.** + The concept to keep is **flush equivalence**, not device identity: the set of files whose flushes are + not independent of one another. What a durability group actually needs to know is whether two of its + commits land in the same such set, because that is what makes them contend. + + **Device number is a proxy for that set, and the proxy is not known to be sound.** The obvious + model -- a device cache flush is per-device, so two logs on one device contend and a group spanning + two devices pays the slower flush -- is the *starting* model, not the finding. The engineer's + refinement: a flush group may be **larger than the one physical device of interest**, with Storage + Spaces the candidate case, since a virtual disk over a pool need not have per-physical-device flush + independence. That is the same shape as the already-recorded `Q6` hazard -- "whether a Storage Space + reports honestly or reports a fiction" -- reaching the flush question rather than the placement one. + Whether the class can also be *smaller* than a device is open and unexamined. + + **This is the argument for declaring rather than discovering, and it is stronger than convenience.** + If the equivalence class cannot be soundly derived from a device number, then a consumer stating it + is not a stopgap until discovery is implemented -- it may be the only sound mechanism, with + discovery serving as a default that must be overridable. That upgrades the musing above from + "might be useful" to "might be the answer", without settling it. + + The instruments exist and answer the *proxy* question today: + `IOCTL_STORAGE_GET_DEVICE_NUMBER` and `IOCTL_VOLUME_GET_VOLUME_DISK_EXTENTS`, both written and + smoke-tested in + [file-handle-numa-spike.rs](crates/windows-ioring-sys/design-sessions/spikes/file-handle-numa-spike.rs), + which counts distinct `DiskNumber` rather than extents because a volume extended twice onto one disk + is still one device. **All of it unmeasured.** The engineer's working position, hedged: the + FUA-to-Flush conversion has pushed devices toward better flush behaviour, so several flushes in a + row is suboptimal rather than pathological. **Low priority, and explicitly not a blocker** -- record + the concept, do not let it gate progress. + ## M-inf -- Ungated - [ ] **M-inf.1** -- The linked and sharded MPSC shapes, if and only if M31.5 shows the array queue's tail diff --git a/CHECKLIST-ship-topology-and-queues.md b/CHECKLIST-ship-topology-and-queues.md index 7060a66e7..713a9f8c7 100644 --- a/CHECKLIST-ship-topology-and-queues.md +++ b/CHECKLIST-ship-topology-and-queues.md @@ -663,7 +663,7 @@ that previously stood in the way are gone: every case observed so far. It is weaker than its own comment claims, and the comment must be corrected even if the check is not. -- [ ] **SH-4.12** -- **`ring_copy`'s `ByL3` policy restates the partition rule instead of asking for +- [x] **SH-4.12** -- **`ring_copy`'s `ByL3` policy restates the partition rule instead of asking for it.** Raised by Copilot at reviews `5116772196` and `5116886015`. [policy.rs](crates/windows-ioring-sys/examples/ring_copy/policy.rs) selects domains with `matches!(domain.kind, DomainKind::Cache { level: 3, .. })`. The reshaped topology model makes @@ -675,6 +675,39 @@ that previously stood in the way are gone: This is the consumer-side twin of the platform-integrity rule: bind to the specified primitive, not to the level number that happens to be L3 on today's hardware. The fix renames the policy as well as changing it, since `byl3` is a user-facing CLI value that would no longer describe what it does. + > **DONE 2026-09-22, together with `M20.1`** -- the two were one change: the prose rule and the code + > that implements it could not land separately without the doc describing something the code did not do. + > Measuring the old filter while replacing it found a shape neither item anticipated: this workspace's + > own machine reports an L3 spanning all 16 processors above a real 8-way L2 partition, so `level: 3` + > matched, did **not** degrade, and returned one whole-machine domain as a successful cache-aware + > partition. Previously coupled to `M20.1` in + > [crates/windows-ioring-sys/CHECKLIST.md](crates/windows-ioring-sys/CHECKLIST.md) -- **do this item + > first**, then that one. `M20.1` sweeps the L3 rule's prose, which reaches `policy.rs`'s doc comments. + > Recorded 2026-09-19: M20's header had asserted that no defect was found in `ring_copy`, which this + > item superseded, and neither file said so. + > + > **`M20.3` is no longer coupled, and landed first (2026-09-22).** That coupling read "it rewrites the + > selection arm this test would assert against", which is true only of a test asserting through + > `ByL3`. The degraded-fallback tail is shared by all five policies and is not what this item changes, + > so `M20.3`'s tests exercise it through `ByNode` and `ByPackage` and pin nothing here. **What this + > item still owes is `ByL3`'s own degradation condition** -- currently "no `level: 3` domain", after + > this "no `outermost_partitioning_cache()`" -- which belongs in this item's verification, where the + > rule being degraded on is the new one. `examples/ring_copy/policy/tests.rs` is where it goes. + > + > **That follow-up is now `SH-4.12.1` below, rather than a note under a checked item.** Raised by + > Copilot review on PR #108: leaving it here left scheduled work marked done, which the + > checked-means-done rule exists to prevent. `SH-4.12` stays checked for the change it did make. + +- [ ] **SH-4.12.1** -- **Give `ByL3` a degradation test on the new rule.** Spawned from `SH-4.12`, + which converted the policy from `matches!(domain.kind, DomainKind::Cache { level: 3, .. })` to + asking `outermost_partitioning_cache()`, and whose own completion note recorded that the + degradation condition still owed a test. + + **What to assert.** The condition being degraded on is now "no `outermost_partitioning_cache()`", + not "no `level: 3` domain", so a test written against the old condition would pass while checking + the wrong rule. Assert both directions: a machine that reports a partitioning cache selects + domains from it, and one that reports none degrades rather than selecting nothing or panicking. + [policy/tests.rs](crates/windows-ioring-sys/examples/ring_copy/policy/tests.rs) is where it goes. - [ ] **SH-4.13** -- **`ProcessorSet` cannot represent every `u8` processor id, and the public API cannot uphold both "every processor" and "no abort".** Raised by Copilot across three unresolved diff --git a/CHECKLIST.md b/CHECKLIST.md index 1e0d64913..88bec1a49 100644 --- a/CHECKLIST.md +++ b/CHECKLIST.md @@ -443,3 +443,25 @@ Ungated work with no identified predecessor deliverable. unexplained result is not mistaken for a tested one. - [x] **M-inf.2** -- Archived the eight completed milestone groups in [CHECKLIST-thread-ambient.md](CHECKLIST-thread-ambient.md), leaving only the parked `M26+`. -> [completed 2026-09-17](COMPLETED-CHECKLIST.md#m-inf2) + +- [ ] **M-inf.3** -- Migrate the existing fourteen `windows-*` crates to the `win-` prefix that + [DESIGN-NOTES.md](DESIGN-NOTES.md#new-crates-take-the-win-prefix) makes the go-forward convention. + **Horizon work, deliberately unscheduled**, and the cost is not uniform -- so this item is a + decision before it is a rename. + + **Three are nearly free**: `windows-guard-alloc`, `windows-placement-probe` and + `windows-platform-probes` carry `publish = false`, so they are a directory move plus path + dependencies. + + **Eleven are published, and a published name cannot be renamed.** crates.io has no rename: a + move is a *new* crate, a final release of the old name pointing at it, and the old name occupying + the namespace permanently -- which is a weaker version of the very collision this convention + avoids. Each also touches `release-please-config.json`, the publish workflow's tag patterns, + CHANGELOG continuity, and every dependent. + + **And one is the repository's own name.** `windows-threadpool-sys` names both a crate and this + repository, so renaming the crate either diverges the two or pulls a repository rename along with + it, breaking remotes and every inbound link. + + Decide the shape first -- all at once, unpublished-only, or never for the published ones -- rather + than starting with the easy three and discovering the policy afterwards. diff --git a/Cargo.lock b/Cargo.lock index 703ffad60..cc399b9db 100644 --- a/Cargo.lock +++ b/Cargo.lock @@ -104,6 +104,13 @@ version = "1.0.24" source = "registry+https://github.com/rust-lang/crates.io-index" checksum = "e6e4313cd5fcd3dad5cafa179702e2b244f760991f45397d14d4ebf38247da75" +[[package]] +name = "win-numa-sys" +version = "0.1.0" +dependencies = [ + "windows-sys", +] + [[package]] name = "windows-core" version = "0.100.0" @@ -166,6 +173,7 @@ name = "windows-ioring-sys" version = "0.3.1" dependencies = [ "serde_json", + "win-numa-sys", "windows-guard-alloc", "windows-overlapped-io-sys", "windows-sys", diff --git a/Cargo.toml b/Cargo.toml index d0007510f..823016ee0 100644 --- a/Cargo.toml +++ b/Cargo.toml @@ -2,6 +2,7 @@ [workspace] members = [ + "crates/win-numa-sys", "crates/windows-file-enumeration-sys", "crates/windows-file-watcher", "crates/windows-file-watcher-example-test-harness", diff --git a/DESIGN-NOTES.md b/DESIGN-NOTES.md index 375aa6611..9633db084 100644 --- a/DESIGN-NOTES.md +++ b/DESIGN-NOTES.md @@ -70,6 +70,155 @@ wants to avoid contributing to it. The Windows threadpool types are inherently m choices up to the developer. The `windows-sys` crate published by Microsoft helps with the basics of the FFI to the APIs, but does little to help turn the alphabet and phrasebook into a useful programming model. +## The adoption thesis: why locality and queues are built together, and why nothing is foreclosed early + +This is the engineer's strategic intent for the ring, queue, topology and durability +work. It is a **thesis, not a measurement** -- it is stated here because it decides +questions that would otherwise be decided by an instinct to prune, and because a reader +who does not have it will mistake deliberate breadth for indecision. The faithful record +of how it was framed is +[DESIGN-SESSION-2026-09-23-adoption-thesis.md](design-sessions/DESIGN-SESSION-2026-09-23-adoption-thesis.md). + +### The hardware is moving and the programming model is not + +Non-uniform memory has been relegated to very expensive machines. Meanwhile Moore's law +has plateaued, so the way to more performance is wider multiprocessors, and wider +multiprocessors are arriving much closer to consumers. **A uniform memory architecture +only scales so far.** Consumer AMD parts are already non-uniform with a uniform facade in +front of them -- chiplets, core complexes, and a fabric between them -- and the same +argument is available for Intel. The non-uniformity is present; what is absent is any +obligation, or any convenient means, to program for it. + +This repository has already met the facade twice, and both observations are measured +rather than argued: + +- A shipping ARM consumer laptop reports **no L3 at all** and **zero** NUMA nodes, with + its natural cluster boundary at L2 + ([D-48](crates/windows-ioring-sys/DESIGN-NOTES.md#d-48)). +- The machine this workspace is developed on reports an **L3 spanning all 16 processors** + above a real 8-way L2 partition. `examples/cache_domains.rs` prints it: `L1: 8`, + `L2: 8`, `L3: 1`. + +In both cases the coarse, advertised boundary is the one that tells you least. + +### Why the concept stayed in the datacenter + +NUMA is difficult to program for, and its benefits are difficult to measure against +whatever you would have written otherwise -- you are comparing against a program you did +not write. But the barrier that matters most is neither of those. **Today, taking +advantage of non-uniformity is a significant architectural decision made at the very +beginning of a system's design.** Who makes that commitment unless they are already +targeting datacenter-class hardware? The commitment is the gate, and it is placed at the +moment when the least is known. + +Queues and rings have a related but distinct problem. The techniques are of general +purpose utility and the building blocks exist in quantity, but outside a few small +domains they are not readily graspable *as a way to structure a system from the +beginning*. That is the same failure +[The value is existence, not cleverness](#the-value-is-existence-not-cleverness) +describes: the correct construction is not within reach, so capable people reach for what +is. + +### The thesis: couple them, and a cycle may start + +The proposition is that there is a **virtuous cycle** available here, and that it is +opened by coupling the two problems rather than solving either alone. + +Make rings and queues -- and the Windows `IoRing` -- a reachable way to structure an +application, and let locality benefits arrive **adaptively out of that structure** rather +than out of a separate up-front architectural bet. Then, in order: + +1. Ordinary application writers adopt the structure because it is a good way to organise + work, not because they set out to be NUMA-aware. +2. I/O-bound writers get what `IoRing` and the epoch durability idiom are worth, which + they cannot easily get today. +3. Designers who genuinely do target NUMA systems get a substrate to build on instead of + starting from primitives. +4. If adoption follows, systems designers have a reason to expose more locality facts at + the consumer hardware level. +5. Which closes the loop: lower-capability hardware becomes able to deliver the same kind + of benefit. + +Step 4 is the payoff and also the part furthest outside our control. It is named because +it is the reason the earlier steps are worth doing in this order. + +### The honest position today + +**Expect low benefit to anything but server-class hardware, and say so.** This is not +pessimism to be edited out later; it is the current state of the evidence. + +The machine this is developed on is logically server-class and is nonetheless **just a +slice**, which does not exhibit non-uniform memory characteristics at all. So the central +claim of the thesis is not one we can currently measure. The consequences of that gap -- +what is blocked, what is not, and what to run when a real multi-node machine is available +-- are worked out in +[DESIGN-SESSION-2026-08-30-numa-sharded-io-execution-domains.md](design-sessions/DESIGN-SESSION-2026-08-30-numa-sharded-io-execution-domains.md) +under "Working under a hardware gap", and that analysis is unchanged by this section. + +### The mechanism: describe the application, not the machine + +**The component that removes the up-front commitment is +[topology-planner](crates/topology-planner/COMPONENT.md)**, and its shape is the thesis made +concrete. The developer supplies a sufficiently abstract definition of the application's **input, +output, and processing code paths** -- a dataflow description, which is a statement about their own +program and something they must know anyway. From it the planner infers the connectivity and +directed flow needed to realize that graph, with no machine in hand; then, given a physical machine +model, it returns **one or more suggested realizations** as specific threads pinned to specific +processor groups, with a stated number of queues of stated types +([EP-D-6](crates/topology-planner/DESIGN-NOTES.md#ep-d-6)). + +Three properties of that arrangement are what make it answer the barrier named above, rather than +relocating it: + +- **The developer never makes a topology decision.** They describe an application; the locality + reasoning happens against a machine model, at a point where the machine is actually known, rather + than as a bet taken at design time. +- **The first stage does not involve a machine at all**, so the application's own structure is + stable across every machine it will ever run on, and only the second stage is redone when the + machine changes. +- **The answer is plural.** Several arrangements are usually defensible and they differ in ways the + planner cannot rank without knowing what the developer values, so it presents candidates and + supplies the means to tell them apart. That is OPTION INTEGRITY at component scale -- the same + refusal to convert an absence of evidence into a verdict. + +### What follows: no early foreclosure + +**Avoid all early foreclosure of techniques that may yield benefits to application +authors.** Two reasons, each independently sufficient: + +1. **We are not the application authors.** We do not know what they will want to do, and + a design option removed here is one they cannot reach no matter how well it would have + suited them. +2. **We do not have the hardware.** Significant analysis of performance tradeoffs on real + NUMA hardware is not available to us, so a tradeoff we "settle" is settled on a machine + that cannot exhibit the phenomenon. + +Either reason alone forbids pruning an option on the strength of a local measurement. Both +together make it the repository's standing posture, stated operationally as **OPTION +INTEGRITY** in [copilot-instructions.md](.github/copilot-instructions.md) -- which is the +rule, where this section is the reason for it. The corollary that a consumer must be +handed *data* rather than a verdict is the same thesis seen from the client's side: an +application author on hardware we have never seen is exactly the person best placed to +decide, and they can only do it if we give them the means to measure. + +**This section does schedule work**, and so differs from +[The value is existence, not cleverness](#the-value-is-existence-not-cleverness), which +deliberately schedules none. The mechanism above is the bulk of it, and it is queued in +that component's own [CHECKLIST.md](crates/topology-planner/CHECKLIST.md) -- `EP-1+.1` for +the dataflow description's vocabulary, `EP-1+.5` for the connectivity graph's type, and +`EP-1+.6` for the plural answer. + +What the runtime crates owe is the other end: being **realizable from** a plan they did +not choose. That is `M27` in +[windows-ioring-sys/CHECKLIST.md](crates/windows-ioring-sys/CHECKLIST.md), which was first +written as an adaptivity question for that crate and **re-planned the same day it was +authored**, because the adaptivity has an owner and it is not there. Answering it in the +ring crate would have grown a second policy surface beside the planner's -- the +`outermost_partitioning_cache` defect again, a policy answer landing in a crate whose job +is something else. [D-8](crates/windows-ioring-sys/DESIGN-NOTES.md#d-8) is untouched by any +of this: being constructible from a policy decision made elsewhere is the opposite of +taking one. + ## The value is existence, not cleverness: "it is only a SMOP" is why it is missing, not a reason to skip it A governing principle for the whole repository, stated because it decides @@ -205,6 +354,30 @@ and close routines. The new crate inherits an established concept rather than in **"Ring" was considered and is wrong for the family.** It is accurate for the array shapes and false for the intrusive-linked one, which is genuinely not a ring. `queues` covers both. +## New crates take the `win-` prefix, not `windows-` + +**The engineer's decision, 2026-09-23, taken when `win-numa-sys` was proposed.** Crates created +from now on use a `win-` prefix. The reason is namespace collision: `windows` is Microsoft's, and +a crate published as `windows-numa-sys` today is a name Microsoft may reasonably want tomorrow. +Abdicating the prefix costs nothing and removes the risk entirely. + +**The `-sys` half is unchanged and is still earned rather than assumed.** It means thin-over-Win32: +memory-safe over an existing API, adding no policy, per +[the waitable-queues naming decision](#the-waitable-queues-crate-is-named-plural-and-carries-no-sys-suffix). +A `win-*` crate that decides something on a consumer's behalf drops the suffix exactly as a +`windows-*` one would. + +**The existing fourteen migrate eventually, and the cost is not uniform.** Eleven of them are +published to crates.io, and a published name cannot be renamed -- a rename is a *new* crate plus a +final release of the old name, and the old name persists forever. Three are unpublished +(`windows-guard-alloc`, `windows-placement-probe`, `windows-platform-probes`) and are nearly free to +move. One further wrinkle: `windows-threadpool-sys` is also the **repository's** name, so renaming +that crate either diverges the two or drags the repository rename along with it. + +The migration is therefore queued at the horizon rather than scheduled, as `M-inf.3` in +[CHECKLIST.md](CHECKLIST.md). **This decision schedules no rename now**; what it settles is the +prefix every *new* crate uses, so the divergence stops growing while the question of the existing +ones stays open. ## Windows SDK model and constraints This crate targets the object-based thread pool API (introduced in Windows Vista) rather than the legacy diff --git a/DESIGN-RATIONALE.md b/DESIGN-RATIONALE.md index 3b38d8de8..be44c288d 100644 --- a/DESIGN-RATIONALE.md +++ b/DESIGN-RATIONALE.md @@ -233,6 +233,60 @@ following the rule that a binding which cannot be shown to fail is cosmetic. Fiv mutations -- three manifest values, a deleted claim, and a stale version planted in prose -- each produce a distinct, located failure. +## Why no option is foreclosed while the hardware gap lasts + +[DESIGN-NOTES.md](DESIGN-NOTES.md#the-adoption-thesis) records the thesis and +[copilot-instructions.md](.github/copilot-instructions.md) records the operational rule +(OPTION INTEGRITY). This is how the rule was reached and what was rejected on the way. The faithful +record of the engineer's framing is +[DESIGN-SESSION-2026-09-23-adoption-thesis.md](design-sessions/DESIGN-SESSION-2026-09-23-adoption-thesis.md). + +The evidence was a specific over-reach, not an argument in the abstract. A harness comparing three +commit strategies in the epoch-log sample found no blast-radius difference between one ring and two, +and the conclusion recorded was that the two-ring strategy's justification was "dead on structural +grounds". Two things were wrong with it, and they fail differently: + +- The harness **could not have shown the difference**. Each lane registers its own arena, so the + arena is the limiter rather than the ring topology. The finding was a fact about the apparatus + presented as a fact about the design. +- It contradicted [D-27](crates/windows-ioring-sys/DESIGN-NOTES.md#d-27), which had already committed + the crate to multiple rings on the strength of per-CPU NVMe queue pairs. One sample's arena sizing + was allowed to overrule a decision made on stronger grounds, and nothing flagged the collision. + +The second is the more instructive failure. A repository whose decisions are well measured builds an +instinct to trust a measurement over a recorded position, and that instinct is right often enough to +be dangerous: it does not ask whether the measurement's configuration could reach the regime the +recorded position was about. + +Three candidate rules were considered. + +**"Prefer the measurement"** is what had been happening, and it is the failure above. + +**"Prefer the recorded decision"** inverts the bug without fixing it -- a decision that a measurement +genuinely falsifies should fall, and this repository has correctly retired decisions that way +(D-47 withdrew half of D-24 on measured grounds). + +What survived distinguishes the two cases by **what the measurement was capable of showing**. A +measurement that reached the regime and found nothing is evidence; a measurement whose apparatus +excluded the regime is evidence about the apparatus. The rule then enumerates the grounds that do +justify foreclosing -- fewer instructions, fewer I/Os, better locality argued structurally, smaller +support burden, clearer model, or a demonstrated wrong answer -- because "measured slower" is absent +from that list on purpose, while "computes the wrong thing" is on it and is the honest ground for +most removals here. + +The reason this repository needs the rule more than most is in the thesis: the hardware that would +make these tradeoffs measurable is not available, and the people best placed to judge the options are +application authors we have not met. Both conditions are temporary in principle and neither is +temporary in practice, so the posture has to be encoded rather than remembered. + +A corollary was adopted with the rule and is worth separating, because it is the part that changes +code rather than judgement: when an option is narrowed, **say where the choice still lives**. The +audit that followed found the rule's own author had withdrawn a cache-level policy on sound grounds +and then failed to say that selecting a level remained available through the topology API. The +repair was to make the sample print every level beside the heuristic's pick, which turns the +justification for the withdrawal into something a reader can see rather than something they are +asked to accept. + ## Why a measured figure is asked to have one home [DESIGN-NOTES.md](DESIGN-NOTES.md#prose-volume-and-error-surface) records the rule; this is how it @@ -269,7 +323,7 @@ moved here from that file, where it had been written inline: Tier 1 is the curre section carrying its own motivating question, census procedure and superseded drafts had made the decision harder to find inside it. -[Restatement drift](#restatement-drift) explains the mechanism and gives the remedy. This note +[Restatement drift](DESIGN-NOTES.md#restatement-drift) explains the mechanism and gives the remedy. This note records something that section does not: a measurement of **where** the drift actually lives, taken after PR #90's eighteenth review round, and what follows from it about formal specification. @@ -374,7 +428,7 @@ uniformly to hit a volume target would remove the only prose that has never been leaving the prose that keeps being wrong in proportion. **A formal spec's most useful property here is not proof -- it is that prose can point at it instead -of paraphrasing it.** That is [restatement drift](#restatement-drift)'s first remedy applied one +of paraphrasing it.** That is [restatement drift](DESIGN-NOTES.md#restatement-drift)'s first remedy applied one level up: define the protocol once in a form that can be checked, and let every document cite it. This is the real connection between the two ideas, and it is why they belong in the same conversation despite fixing different things. diff --git a/PLANS.md b/PLANS.md index 2e17f06eb..f4bd0ab7e 100644 --- a/PLANS.md +++ b/PLANS.md @@ -20,7 +20,7 @@ plans tracker: [crates/windows-file-enumeration-sys/PLANS.md](crates/windows-fil | Path to CHECKLIST.md | Status | Brief description | Design Notes | |---|---|---|---| | [CHECKLIST-mutation-survivors.md](CHECKLIST-mutation-survivors.md) | not started | Work queued from the workspace-wide cargo-mutants sweep of 2026-09-02, whose findings are kept in [mutation-sweeps/2026-09-02/](mutation-sweeps/2026-09-02/README.md) rather than re-derived -- the run took roughly fourteen hours. 2,792 caught, 1,112 survived, 198 timed out. **The headline numbers mislead in three ways and the README says how**: a timeout in a blocking-API crate is usually a detection that lost its name rather than a gap (measured: one of `windows-waitable-queues`' 120 timeouts fails four tests in 0.00s when re-injected alone), a low score on an executable probe crate is measuring the wrong thing, and three kinds of survivor -- equivalent mutants, unreachable code, and constants that want a `const` assertion -- are not missing tests at all. M1 covers the shipping crates; M2 holds the two crates that are not libraries and whose scope is an engineer's decision; M3 re-runs and prunes rather than hand-editing the tool's output into a second source of truth. | [mutation-sweeps/2026-09-02/README.md](mutation-sweeps/2026-09-02/README.md) | -| [crates/topology-planner/CHECKLIST.md](crates/topology-planner/CHECKLIST.md) | in progress | **Planned, not built** -- the directory holds a plan and no code, and becomes a crate when M2 begins. Owns the mapping from a stated **goal** plus an abstracted idealized machine description to a set of execution domains: which processors host a domain, where each thread pins, which memory node it allocates from, what channel connects each pair, and where each channel's buffer lives. Filed because that mapping was **unowned**: [CHECKLIST-io-domains.md](CHECKLIST-io-domains.md) M32 lists the contracts "the runtime cannot be written without" and all of them concern the queue, while M33+.1 opens with "one pinned thread, its `IoRing`, its node-local registered pool, its shard" -- presupposing a plan nothing computed. Separate from `windows-topology-sys` because that crate states **facts** and this one applies **policy**; fusing them is what produced `outermost_partitioning_cache`, a policy answer sitting in the facts crate that three consumers then re-derived differently (SH-16.9). M1 was a *requirements* milestone -- it states what the topology must answer, and it fed the locality-model session, which has since concluded as `D-13`..`D-21`. **The component was deferred past PR #56 by direction**, contributing only planning documents there; #56 then closed unmerged on 2026-09-15 and its content is landing in peeled pieces instead, so the component is still unlanded and goes in its own pull request. Per `D-21` the topology reshape lands without it, since `windows-topology-sys` publishes a refined view of what the platform publishes and an adapter absorbs the rest. M2+ and M3+ are parked on that session concluding, and are additionally **awaiting a re-cut**: EP-D-4 and EP-D-5 re-scoped the component into four parts (`topology-model` holding the abstract machine description, the planner's traits and the plan type; `topology-planner`; an inward Windows adapter; an outward realizer), and only M1 has been reconciled with that. EP-1.1 is done and already earned its keep: checking the shard-set query against the model found `Processor::capacity` using `0` as both a valid efficiency class and a "not known" sentinel, which collide on every non-hybrid machine (filed as SH-16.12). | [crates/topology-planner/DESIGN-NOTES.md](crates/topology-planner/DESIGN-NOTES.md), [design-sessions/DESIGN-SESSION-2026-09-02-cache-locality-model.md](design-sessions/DESIGN-SESSION-2026-09-02-cache-locality-model.md) | +| [crates/topology-planner/CHECKLIST.md](crates/topology-planner/CHECKLIST.md) | in progress | **Planned, not built** -- the directory holds a plan and no code, and becomes a crate when M2 begins. Owns the mapping from a **dataflow description of an application** -- its input, output and processing code paths, the shape settled 2026-09-23 as `EP-D-6` and discharging the deferral `EP-D-4` named -- plus an abstracted idealized machine description, to **one or more suggested** sets of execution domains: which processors host a domain, where each thread pins, which memory node it allocates from, what channel connects each pair, and where each channel's buffer lives. `EP-D-6` also split planning into two stages -- connectivity inferred with no machine in hand, then realization against a machine model -- which adds `EP-1+.5` (the connectivity graph is a value and needs a type and a home) and `EP-1+.6` (the answer is a set, so what a caller receives and what `M2+.2` renders are sets). Filed because that mapping was **unowned**: [CHECKLIST-io-domains.md](CHECKLIST-io-domains.md) M32 lists the contracts "the runtime cannot be written without" and all of them concern the queue, while M33+.1 opens with "one pinned thread, its `IoRing`, its node-local registered pool, its shard" -- presupposing a plan nothing computed. Separate from `windows-topology-sys` because that crate states **facts** and this one applies **policy**; fusing them is what produced `outermost_partitioning_cache`, a policy answer sitting in the facts crate that three consumers then re-derived differently (SH-16.9). M1 was a *requirements* milestone -- it states what the topology must answer, and it fed the locality-model session, which has since concluded as `D-13`..`D-21`. **The component was deferred past PR #56 by direction**, contributing only planning documents there; #56 then closed unmerged on 2026-09-15 and its content is landing in peeled pieces instead, so the component is still unlanded and goes in its own pull request. Per `D-21` the topology reshape lands without it, since `windows-topology-sys` publishes a refined view of what the platform publishes and an adapter absorbs the rest. M2+ and M3+ are parked on that session concluding, and are additionally **awaiting a re-cut**: EP-D-4 and EP-D-5 re-scoped the component into four parts (`topology-model` holding the abstract machine description, the planner's traits and the plan type; `topology-planner`; an inward Windows adapter; an outward realizer), and only M1 has been reconciled with that. EP-1.1 is done and already earned its keep: checking the shard-set query against the model found `Processor::capacity` using `0` as both a valid efficiency class and a "not known" sentinel, which collide on every non-hybrid machine (filed as SH-16.12). | [crates/topology-planner/DESIGN-NOTES.md](crates/topology-planner/DESIGN-NOTES.md), [design-sessions/DESIGN-SESSION-2026-09-02-cache-locality-model.md](design-sessions/DESIGN-SESSION-2026-09-02-cache-locality-model.md) | | [CHECKLIST-io-domains.md](CHECKLIST-io-domains.md) | in progress | **M30 is complete and archived** in [COMPLETED-CHECKLIST.md](COMPLETED-CHECKLIST.md): the queue crate's name, skeleton and SPSC shape. M31 built the bounded-array MPSC (with a lazily created manual-reset doorbell whose reset cannot be separated from the observation that there is nothing to take -- achieved by ordering plus a re-check rather than by a lock, per D-9 and D-15) and is done but for `M31.6`, the `loom` verification, which is re-homed as `M30.4` in [CHECKLIST.md](CHECKLIST.md). M32 remains: the contract decisions -- ordering, correlation, backpressure among them -- the domain runtime cannot be written without. M33+ parks the runtime itself, the creation-time-affinity thread builder, the namespace `Outcome` extension, the client-side `ThreadpoolWait` fan-in helper, and the durability crate. M-inf holds items each gated on a specific measurement rather than on taste. The N=1 path is the whole first deliverable and depends on no NUMA hardware. | [DESIGN-NOTES.md](DESIGN-NOTES.md), [DESIGN-SESSION-2026-08-30-numa-sharded-io-execution-domains.md](design-sessions/DESIGN-SESSION-2026-08-30-numa-sharded-io-execution-domains.md) | | [CHECKLIST.md](CHECKLIST.md) | in progress | M19: propagate the 2026-08-27 platform measurements (IoRing registration replaces the table; the completion-port/`IoRing` fork; `runs_long` as the growth mechanism; the measured 512 default maximum) into the crates whose code or documentation currently assumes otherwise. M20: decide the session-independent path form, now that path resolution is measured to follow the impersonated token's logon session. M21: reconcile with the impersonation and enumeration crates that landed during the session. M34 carries the review-driven repairs raised while shipping the placement tool: M34.1 (the reusable sabotage harness) is done, and M34.2 (route the placement tool's output through a sink rather than writing to stdout from many sites) and M34.3 (archive the completed item bodies still carried by the three root checklists) are open. M37: discharge the failable-call standard across the workspace. M30: find out how much of this workspace's algorithm correctness can be machine-checked -- a survey matching each argued-but-unchecked algorithm to a class of tool (TLA+/PlusCal, loom, bounded proof, `const` assertions), one pilot chosen because parameter shrinking makes an untestable property exhaustive, and a named list of what the pilot could not reach, which is the deliverable. Scoped as an instrument for narrowing hand-inspection rather than replacing it, and explicitly not a reversal of [D-31](crates/windows-waitable-queues/DESIGN-NOTES.md#d-31). Also re-homes `M31.6`, the `loom` verification the queue crate promises adopters before 1.0: it was previously untracked, referenced from that crate's design notes, a source file and its sabotage manifest with no live checklist item anywhere, and is now queued as M30.4. M30's rationale is in [DESIGN-RATIONALE.md](DESIGN-RATIONALE.md#machine-checking-what-is-argued) -- Tier 2, because no decision is taken yet; M30.5 is what produces one. | [DESIGN-NOTES.md](DESIGN-NOTES.md#remoting-synchronous-namespace-operations) for M19-M21 and M37; N/A for M30 | | [CHECKLIST-ship-topology-and-queues.md](CHECKLIST-ship-topology-and-queues.md) | in progress | Release `windows-topology-sys` 0.2.0 and `windows-waitable-queues` 0.1.0, which everything shorter-term depends on. **Both reached crates.io on 2026-09-05**, by a route this file does not describe, since PR #56 closed unmerged; `SH-4.15` owns reconciling M4 with what shipped. Deliberately redundant with [CHECKLIST-io-domains.md](CHECKLIST-io-domains.md): that file plans the design, this one plans the release, and a release has failure modes a design checklist does not surface. Two were found while writing it -- `windows-waitable-queues-v*` is missing from the publish workflow's tag list, so release-please would tag it and nothing would publish it, silently; and `windows-ioring-sys` is published against `windows-topology-sys = "0.1.0"`. (That second finding was later **corrected at SH-2.2**: the pin is a *dev*-dependency, which consumers never resolve, so it obliges a pin update but no release.) M1 settled the public surface before it was public (done, archived); M2 repairs the plumbing; M3 lands the branch; M4 releases; M5 verifies from outside the workspace; M6 is long-running validation and gates the queue crate's release specifically. M7-M13 were seven PR #56 review rounds (done, archived); M14, M15 and M16 are the three later rounds and carry the file's open work -- M15 owns the fix for an ABA hole that ships **disclosed rather than fixed**, so it does not block the release. M16 is the SH-3.1.1 diff review, the first to read the branch as a diff rather than react to a comment: seven findings, six fixed, including a publish-workflow regression this branch had introduced two commits earlier and a soundness hole in the crate about to freeze its API. Its remaining four are blocked on [design-sessions/DESIGN-SESSION-2026-09-02-cache-locality-model.md](design-sessions/DESIGN-SESSION-2026-09-02-cache-locality-model.md), which began by asking whether collapsing a seven-kind, any-depth topology onto a single cache boundary is the right projection and has since settled that presence and observation must be modeled rather than collapsed into an `Option`. **That work gated the merge, and has since discharged**: unlike M14 and M15, which concern a defect in an implementation that can ship disclosed, M16 concerned the shape of the public model `windows-topology-sys` 0.2.0 would publish, and a published model cannot be reshaped without another break. It became the `MMT-*` plan, which has landed; the session has concluded and 0.2.0 shipped the new model, so M3 no longer waits on M16. The file opens with a status table. | [crates/windows-topology-sys/DESIGN-NOTES.md](crates/windows-topology-sys/DESIGN-NOTES.md), [crates/windows-waitable-queues/DESIGN-NOTES.md](crates/windows-waitable-queues/DESIGN-NOTES.md) | @@ -28,6 +28,7 @@ plans tracker: [crates/windows-file-enumeration-sys/PLANS.md](crates/windows-fil | [CHECKLIST-thread-ambient.md](CHECKLIST-thread-ambient.md) | in progress | **M22-M29 are complete and archived** in [COMPLETED-CHECKLIST.md](COMPLETED-CHECKLIST.md): `windows-thread-ambient-sys` (a standalone layer that captures a thread's ambient state and applies it on another thread), `windows-namespace-request-sys` (marshalable Win32 namespace call parameter sets, over an entry list audited from three real consumers rather than guessed), `windows-platform-probes` (a durable home for the measurements this workspace's designs rest on), and the defects the audit of those three found. What remains is **pending**: the three `M26+` items were each gated on the namespace-facility design branch reaching `main`, and it has, so the gate has lifted -- reconciling the imported design background, applying the M22.2 narrowing to M21.2, and making the merge-or-delete decision on the duplicated path preparation. The file is deleted outright once those land. | [crates/windows-thread-ambient-sys/DESIGN-NOTES.md](crates/windows-thread-ambient-sys/DESIGN-NOTES.md) | | [crates/windows-overlapped-io-sys/CHECKLIST.md](crates/windows-overlapped-io-sys/CHECKLIST.md) | not started | M14: finish the contract audit -- categories 1, 2, 6, 8, 9 were not examined -- and sweep `outstanding()` for the advisory-predicate hazard. | [crates/windows-overlapped-io-sys/DESIGN-NOTES.md](crates/windows-overlapped-io-sys/DESIGN-NOTES.md) | | [crates/windows-platform-probes/CHECKLIST.md](crates/windows-platform-probes/CHECKLIST.md) | in progress | M1 (streaming reports) is done and archived: every probe now writes into the sink as it measures, through a `fmt::Write` adapter that left every `writeln!` call site untouched, and the `catch_unwind`/`resume_unwind` pair is gone because there is no longer a buffer to rescue. Measured with a control: a probe killed part-way through a run keeps its banner and heading on every run of the new build, where the previous build kept nothing -- see [crates/windows-platform-probes/DESIGN-NOTES.md](crates/windows-platform-probes/DESIGN-NOTES.md#d-streaming-report) for the figures. M2 built the report oracle -- one executable definition of the correspondences between a report's prose and NDJSON halves, bound inside the renderers so every test that renders inherits it -- along with a derived fact set and a corpus of report shapes. It is complete; its ten unrelated leftovers -- CI hygiene, a doc repair, probe-prose corrections -- were re-sequenced into M4 (gated on M3) and M5 (gated on nothing). M3 then supersedes its central rule. Re-reading M2's own evidence showed that both defects which motivated the oracle were defects in the ENCODED ROW, not in the relation between two renderings, and that the row published its three diagnostic lists as bare counts -- so a survey reading `"parse_incomplete":1` could not tell a probe self-bug from host flakiness. The row is the machine contract and gets the facts and the invariants; the prose is for a reader and gets review. M3 is complete and archived (ten items): those three fields publish arrays of OBJECTS, each carrying a stable `code` plus the values its variant holds -- `{"code":"partitioning_summary_missing","level":9}` rather than the bare `"partitioning_summary_missing"` of the superseded M3.1 form; the surviving correspondences became invariants over the observation rather than over two renderings, so a rule that reads the diagnostic lists -- which would be a restatement of `verdict()` and blind to a deleted push site -- was rewritten to read the observation; the row is emitted from a typed value through one writer with total escaping, which is the crate's only defence against caller text reaching the mined artifact; the prose oracle and every parser serving it were deleted, and no test extracts structured data from prose anywhere in the crate. Four later items came from reviews and are the more instructive half: three instruments were found asserting less than their names claimed, `BlockingState::ALL` was found to be a census the compiler did not check despite a doc comment claiming it did, and the row's hand-written JSON well-formedness check was measured against a real parser over 1807 generated corruptions -- 159 disagreements, every one a FALSE accept -- and replaced by `serde_json`, after which the remaining hand-written string scanners were deleted too. What remains is M4 (four M2 leftovers M3 gated, now unblocked and re-scoped) and M5 (six ungated hygiene items). | [crates/windows-platform-probes/DESIGN-NOTES.md](crates/windows-platform-probes/DESIGN-NOTES.md#d-streaming-report), [#d-encoded-row-is-the-contract](crates/windows-platform-probes/DESIGN-NOTES.md#d-encoded-row-is-the-contract) | -| [crates/windows-ioring-sys/CHECKLIST.md](crates/windows-ioring-sys/CHECKLIST.md) | in progress | Memory-safe Rust over the Windows `IoRing` submission/completion ring, as a new crate. M1-M19 are complete (0.2.0 shipped 2026-08-30, restoring availability after all three 0.1.x versions were yanked); M1-M19 are archived. **M20** queues documentation and policy-test repairs from the 2026-08-30 NUMA-sharding measurement, and the pinned-thread `M6+` work stays parked. | [crates/windows-ioring-sys/DESIGN-NOTES.md](crates/windows-ioring-sys/DESIGN-NOTES.md) | +| [crates/win-numa-sys/CHECKLIST.md](crates/win-numa-sys/CHECKLIST.md) | in progress | Memory-safe Rust over the Windows NUMA APIs, and the first crate to take the `win-` prefix ([DESIGN-NOTES.md](DESIGN-NOTES.md#new-crates-take-the-win-prefix)). Created because `VirtualAllocExNuma` had been written twice in shipping library code, in `windows-ioring-sys` and `windows-placement-probe`, which do not depend on each other -- and a third time before that, in a sample, which `M22.3` had hoisted. `NumaBuffer` moved here; `windows-ioring-sys` re-exports it and supplies the `IoBuf`/`IoBufMut` impls via the orphan rule, which keeps `M6+.6`'s deferred trait-merge decision untouched. **The duplication is not yet removed, only relocated**: `N-1.1` collapses the probe's copy, and `N-1.2` decides whether `QueryWorkingSetEx` observation moves with it. | [crates/win-numa-sys/README.md](crates/win-numa-sys/README.md) | +| [crates/windows-ioring-sys/CHECKLIST.md](crates/windows-ioring-sys/CHECKLIST.md) | in progress | Memory-safe Rust over the Windows `IoRing` submission/completion ring, as a new crate. M1-M19 are complete (0.2.0 shipped 2026-08-30, restoring availability after all three 0.1.x versions were yanked); M1-M19 are archived. **M20** queues documentation and policy-test repairs from the 2026-08-30 NUMA-sharding measurement, and the pinned-thread `M6+` work stays parked. M21-M24 (epoch-log review, arena and submission, and making the unit suite hermetic) are complete; M23, M25 and M26 are open, and **M27 was added 2026-09-23 as the first milestone queued by intent, and re-planned the same day**: the adaptivity the [adoption thesis](DESIGN-NOTES.md#the-adoption-thesis) asks for is owned by [topology-planner](crates/topology-planner/COMPONENT.md), so M27 is now what this crate owes a realizer rather than a policy surface of its own. | [crates/windows-ioring-sys/DESIGN-NOTES.md](crates/windows-ioring-sys/DESIGN-NOTES.md) | Add a row here when new work is planned, against [CHECKLIST.md](CHECKLIST.md) or any crate's. diff --git a/crates/topology-planner/CHECKLIST.md b/crates/topology-planner/CHECKLIST.md index 4ca82654f..0f229daa1 100644 --- a/crates/topology-planner/CHECKLIST.md +++ b/crates/topology-planner/CHECKLIST.md @@ -1,9 +1,10 @@ # Checklist: the topology planner -Plans an arrangement of execution domains from a stated **goal** plus an **abstracted idealized** -description of a machine. See [COMPONENT.md](COMPONENT.md) for what this crate is and why it is -separate from both the topology crate and the runtime, and -[EP-D-4](DESIGN-NOTES.md#ep-d-4) for the architecture it now sits in. +Plans arrangements of execution domains from a **dataflow description of an application** -- its +input, output, and processing code paths -- plus an **abstracted idealized** description of a +machine. See [COMPONENT.md](COMPONENT.md) for what this crate is and why it is separate from both +the topology crate and the runtime, [EP-D-4](DESIGN-NOTES.md#ep-d-4) for the architecture it now +sits in, and [EP-D-6](DESIGN-NOTES.md#ep-d-6) for the two-stage shape and the plural answer. **The component has been re-scoped**, per [EP-D-4](DESIGN-NOTES.md#ep-d-4) and [EP-D-5](DESIGN-NOTES.md#ep-d-5). It is named `topology-planner` and the directory now matches; it @@ -33,7 +34,7 @@ prerequisites rather than on someone else's decision. | Milestone | State | What it is waiting on | |---|---|---| | M1 the input contract | 3 done, 2 open | `EP-1.4` and `EP-1.5`'s coverage half, which want a settled model | -| M1+ scenario and naming | **partly answered** | the name is settled (EP-D-4); the goal input is deferred for litigation, by direction | +| M1+ scenario and naming | **partly answered** | the name is settled (EP-D-4); the goal input's *shape* is settled as a dataflow description (EP-D-6), and its concrete vocabulary is `EP-1+.1`'s remaining work | | M2+ the plan as a value | parked, **and needs re-cutting** | re-cut against EP-D-4/EP-D-5, then the topology reshape landing | | M3+ the policies | parked | M2+ | | M-inf parked | ungated | not scheduled, deliberately | @@ -141,32 +142,64 @@ Raised when the engineer described this component's function, which turned out t "takes a topology, applies policy". Both are gated on the locality-model session, but neither is a model question -- they are this component's own. -- [ ] **EP-1+.1** -- **Describe the scenario input.** The synthesizer takes *two* inputs and only one - is described anywhere. The scenario says what the caller intends to run, and it is what makes a - measurement meaningful: [EP-D-3](DESIGN-NOTES.md#ep-d-3) established that a measured number means - nothing without knowing what it measured, so at minimum the scenario must distinguish small-message - handoff from large-buffer streaming. Its absence is why "what is most useful for consumers" was - hard to answer in the abstract for so long. +**Re-planned 2026-09-23 by [EP-D-6](DESIGN-NOTES.md#ep-d-6)**, which settled the input's shape and +in doing so added work this milestone did not anticipate: the two-stage split means stage 1 has an +output that is a value in its own right, and the plural answer means stage 2 returns a set rather +than a plan. `EP-1+.1` is narrowed accordingly and `EP-1+.5` / `EP-1+.6` are new. + +- [ ] **EP-1+.1** -- **Give the dataflow description a concrete vocabulary.** Its *shape* is settled + by [EP-D-6](DESIGN-NOTES.md#ep-d-6) -- a sufficiently abstract definition of the application's + input, output, and processing code paths -- so what remains is the types: how a caller names a + source, a sink and a processing step, and how they express an edge's characteristics. That last + part is what makes a measurement meaningful, since [EP-D-3](DESIGN-NOTES.md#ep-d-3) established + that a measured number means nothing without knowing what it measured; at minimum an edge must + distinguish small-message handoff from large-buffer streaming. **Note the correction EP-D-6 + carried:** that distinction is an attribute *of an edge*, not the whole scenario, which is how + this item originally framed it. - [ ] **EP-1+.2** -- **Decide what the caller-callback traits ask.** Planning is a negotiation: the component may call back for clarification the scenario did not settle. Enumerating those questions is what decides whether this is one trait or several, and it cannot be done before EP-1+.1 says what the scenario already answers. -- [ ] **EP-1+.3** -- **Settle the naming, before any type is written.** Both inputs and the output - are graphs of processors and their relations, so "topology" fits all of them and distinguishes - none -- and a reader seeing the word twice will eventually take one for the other. Decide whether - the observed machine keeps the bare name (qualified only by its crate), gains a qualifier, or is - renamed outright; what the synthesized arrangement is called; and whether the inward/outward - adapters keep those role names or gain more specific crate/type names. Cheap now; expensive once - any of those names are public. This one blocks nothing but should not be settled by whoever writes - the first type. +- [ ] **EP-1+.3** -- **Settle the naming, before any type is written.** There are now **four** graphs + in play and "topology" fits several of them while distinguishing none -- a reader meeting the word + twice will eventually take one for the other. They are: the **machine** (processors and their + relations), the **application's dataflow description** (sources, sinks and processing steps), the + **connectivity graph** stage 1 infers from it (the same nodes, with directed flow), and each + **realization** stage 2 emits (threads on processor groups, with queues between them). Note that + only two of the four are graphs of *processors*, which is itself a correction: + [EP-D-6](DESIGN-NOTES.md#ep-d-6) made the input an application-shaped graph, where this item + previously assumed every graph was machine-shaped. Decide whether the observed machine keeps the + bare name (qualified only by its crate), gains a qualifier, or is renamed outright; what each of + the other three is called; and whether the inward/outward adapters keep those role names or gain + more specific crate/type names. Cheap now; expensive once any of those names are public. This one + blocks nothing but should not be settled by whoever writes the first type. **The connectivity + graph is the one with no name at all today.** - [ ] **EP-1+.4** -- **Assign measurement ownership for directed residency cost in the four-part architecture.** [EP-D-3](DESIGN-NOTES.md#ep-d-3) requires directed cross-domain cost input with measurement context. Record which layer owns collecting, validating, and supplying that measurement context to `topology-model` through the inward adapter/synthesizer path. +- [ ] **EP-1+.5** -- **Give stage 1's output a type, and decide which crate holds it.** + [EP-D-6](DESIGN-NOTES.md#ep-d-6) establishes that the connectivity graph with its directed flow is + derived before any machine exists, which makes it a value that can be inspected and reviewed + without even a synthetic machine -- a stronger version of the argument + [EP-D-5](DESIGN-NOTES.md#ep-d-5) used to make the plan a value. EP-D-5's placement rule points at + `topology-model` for both this type and the dataflow description, on the grounds that a component + which only describes or realizes must not depend on planning policy. **EP-D-6 recorded that as an + argument and deliberately did not take it as a decision**; this item takes it, either way, and says + what depends on the answer. + +- [ ] **EP-1+.6** -- **Make the plural answer explicit in the plan vocabulary.** The planner returns + *one or more* suggested realizations ([EP-D-6](DESIGN-NOTES.md#ep-d-6)), so the type a caller + receives is a set and the thing `M2+.2` renders is a set. Decide what a candidate carries beyond + the arrangement itself -- at minimum, enough for a developer to tell two candidates apart and say + why they would pick one, which is the whole purpose of returning more than one. Ranking is + explicitly *not* in scope here: the planner declines to rank because it cannot know what the + developer values, and `M3+` owns whether that ever changes. + ## M2+: the plan as a value Parked, not pending. Gated on the topology model landing. Shape recorded so it is not lost, per the diff --git a/crates/topology-planner/COMPONENT.md b/crates/topology-planner/COMPONENT.md index 65541d650..c57b5f3d4 100644 --- a/crates/topology-planner/COMPONENT.md +++ b/crates/topology-planner/COMPONENT.md @@ -10,22 +10,37 @@ emits a platform-neutral plan, so nothing in it is Windows-specific. See ## What it is -A **planner**. It takes two inputs and produces a third thing: +A **planner**. It takes a description of an application and produces arrangements for running it. -- **a stated goal** -- what the caller intends the arrangement to achieve. Its shape is deliberately - **deferred for litigation**; that is a named deferral, not an omission. -- **an abstracted idealized description of a machine** -- processors, memory, storage, interconnects, - distances and bottlenecks. Not Windows-shaped, and richer than any single platform reports. It is - **mockable by construction**: a description of a machine nobody has is an ordinary input, which is - what makes this component testable without the hardware it plans for. +**The input is a dataflow description** -- a sufficiently abstract definition of the application's +**input, output, and processing code paths**. It says what the application *is*, not how it should +be arranged: no domain counts, no queue selections, no topology preferences. See +[DESIGN-NOTES.md](DESIGN-NOTES.md) -> `EP-D-6`. -From those it produces **a plan**: which processors host domains, where each thread pins, which -memory node each allocates from, what channel connects each pair, and where each channel's buffer -lives. The plan **serializes to JSON** and stays abstracted from Windows. +**Planning happens in two stages**, and the split is load-bearing rather than incidental: + +1. **Connectivity, with no machine in hand.** From the dataflow description, infer the general + connectivity and the **directed flow of data** needed to realize that graph. This stage depends + only on the application, so its result holds for every machine the application will ever run on. +2. **Realization, given an abstracted idealized description of a machine** -- processors, memory, + storage, interconnects, distances and bottlenecks. Not Windows-shaped, and richer than any single + platform reports. It is **mockable by construction**: a description of a machine nobody has is an + ordinary input, which is what makes this component testable without the hardware it plans for. + +The second stage produces **one or more suggested realizations**: which processors host domains, +where each thread pins, which memory node each allocates from, how many queues of which types, what +channel connects each pair, and where each channel's buffer lives. Realizations **serialize to +JSON** and stay abstracted from Windows. + +**The answer is plural on purpose.** Several arrangements are usually defensible on a given machine, +they differ in ways the planner cannot rank without knowing what the developer values, and +presenting them as candidates is what lets the developer choose. The planner proposes; it does not +return the one true arrangement. **It may ask.** Planning is a negotiation, not a pure function: the component may call back to its -caller through traits for clarifying information the goal did not settle. Which questions those are -is not yet known, and knowing them is what decides whether that is one trait or several. +caller through traits for clarifying information the dataflow description did not settle. Which +questions those are is not yet known, and knowing them is what decides whether that is one trait or +several. ## The four components, and which way the arrows point diff --git a/crates/topology-planner/DESIGN-NOTES.md b/crates/topology-planner/DESIGN-NOTES.md index 448444bfc..1d3b426b8 100644 --- a/crates/topology-planner/DESIGN-NOTES.md +++ b/crates/topology-planner/DESIGN-NOTES.md @@ -22,8 +22,9 @@ renamed to match. | EP-D-1 | **The shard-set query**: what the planner must know to choose which processors host a domain, and what today's model cannot tell it. | | EP-D-2 | **The proximity query**: how close two processors are, which selects the channel between their domains. Takes an **unordered** pair; the model has no answer today. | | EP-D-3 | **The residency query**: where a domain's pool lives, and which side of a cross-domain pair should host a shared ring. **Ordered**, with directed cost entering through the abstract model/adapter path under [D-20](../windows-topology-sys/DESIGN-NOTES.md#d-20). | -| EP-D-4 | **The four-part architecture, and the planner's name.** The engineer's position: the planner is **`topology-planner`** (no `windows-` prefix); it takes a **goal** description (shape deferred for litigation), queries an **abstracted idealized** model covering processors, memory, storage, interconnects, distances and bottlenecks, and emits a **JSON-serializable, platform-neutral** plan. Two kinds of **adapter** bracket it: one exposing the planner's traits over the Windows topology objects, one **realizing** a plan as buffers, rings and threads with the user's code inserted at the right steps. Settles `MMT-1.5` (the facts crate keeps its `-sys` name), the "two graphs, one word" ambiguity, and where distance lives -- the attributed interconnect shape D-9 sketched goes in the abstract model, so D-9's deferral in the facts crate stands unreopened. | +| EP-D-4 | **The four-part architecture, and the planner's name.** The engineer's position: the planner is **`topology-planner`** (no `windows-` prefix); it takes a **goal** description (shape settled by [EP-D-6](#ep-d-6)), queries an **abstracted idealized** model covering processors, memory, storage, interconnects, distances and bottlenecks, and emits a **JSON-serializable, platform-neutral** plan. Two kinds of **adapter** bracket it: one exposing the planner's traits over the Windows topology objects, one **realizing** a plan as buffers, rings and threads with the user's code inserted at the right steps. Settles `MMT-1.5` (the facts crate keeps its `-sys` name), the "two graphs, one word" ambiguity (widened to four graphs by [EP-D-6](#ep-d-6)), and where distance lives -- the attributed interconnect shape D-9 sketched goes in the abstract model, so D-9's deferral in the facts crate stands unreopened. | | EP-D-5 | **The component layout: `topology-model` is its own crate, and dependencies point one way.** The abstract model and the traits the planner queries live in `topology-model`, which the planner and both adapters depend on; non-planner components do not depend on `topology-planner`. Putting the traits in the planner would make a crate whose job is to *describe a machine* depend on one that applies *policy* -- the same defect as `outermost_partitioning_cache`, arriving as a dependency edge instead of an API. Two consequences derived from the same rule rather than decided separately: **the plan type also lives in `topology-model`** (otherwise the realizer depends on the planner), and the inward adapter and the realizer are **separate crates** (their dependency sets barely overlap, and fusing them would make reading a topology pull in the whole runtime). | +| EP-D-6 | **The goal's shape, and planning as two stages with a plural answer.** Discharges the deferral [EP-D-4](#ep-d-4) named. The caller supplies **a sufficiently abstract definition of the application's input, output, and processing code paths** -- a dataflow description, not a topology preference. From it the planner performs **two inferences in sequence**: first, with no machine in hand, the **general connectivity and directed flow of data** needed to realize that graph; then, given a physical machine model, **one or more suggested realizations** of the graph as specific execution threads pinned to specific processor groups, with a stated number of queues of stated types. The answer is **plural by construction**: the planner proposes candidates for the developer to choose between, and does not return the one true arrangement. | ## EP-D-1: the shard-set query @@ -264,6 +265,8 @@ Detailed trigger analysis and prior framing are recorded in ## EP-D-4: the four-part architecture, and the planner's name +**The goal's deferred shape is now settled by [EP-D-6](#ep-d-6).** The rest of this decision stands. + *The engineer's position, 2026-09-03. This is a **choice**, not one of M1's queries, and it re-scopes the component that records it.* @@ -272,8 +275,9 @@ re-scopes the component that records it.* **The planner is `topology-planner`** -- deliberately with no `windows-` prefix. - **Input**: a description of the **goal** of the topology -- what the caller intends the - arrangement to achieve. Its shape is **explicitly deferred for litigation**, which is a named - deferral rather than an omission. + arrangement to achieve. Its shape was **explicitly deferred for litigation** when this decision + was written, which was a named deferral rather than an omission; it has since been settled as a + dataflow description by [EP-D-6](#ep-d-6). - **What it queries**: an **abstracted, idealized** description of the machine, covering **processors, memory, storage (NVMe), interconnects, distances, and bottlenecks**. Not Windows-shaped, and materially richer than what any one platform reports. @@ -391,3 +395,85 @@ the caller did not ask for" rule that decided the layout in the first place. - **Whether `topology-model` is one crate or eventually two.** The machine description and the plan vocabulary are different enough that they might separate later. They are together now because splitting on speculation costs more than merging on evidence. + +## EP-D-6: the goal's shape, and planning as two stages with a plural answer + +*Recorded 2026-09-23, from the engineer's statement of the component's purpose. Discharges the +deferral [EP-D-4](#ep-d-4) named as "shape deferred for litigation". Supplies the input half that +[CHECKLIST.md](CHECKLIST.md) `EP-1+.1` was opened to describe.* + +### The decision + +The caller supplies **a sufficiently abstract definition of the application's input, output, and +processing code paths**. That is a **dataflow description**: what comes in, what goes out, and what +the code does between them. It is deliberately not a topology preference, not a domain count, and +not a queue selection -- the caller states what their application *is*, not how it should be +arranged. + +From that the planner performs **two inferences, in sequence**: + +1. **Connectivity, with no machine in hand.** Infer the general connectivity and the **directed flow + of data** needed to realize the described graph. This stage answers what must connect to what, + and in which direction, and it is complete before any machine is considered. + +2. **Realization, given a physical machine model.** Respond with **one or more suggested + realizations** of that graph: specific execution threads pinned to specific processor groups, a + stated number of queues of stated types, and the placements the plan already covers under + [EP-D-5](#ep-d-5). + +### What it settles that was previously open + +**The goal is a dataflow description.** [EP-D-4](#ep-d-4) recorded the goal as an input whose shape +was deferred, and [EP-D-3](#ep-d-3) established that a measured number means nothing without knowing +what it measured. The dataflow description is what makes a measurement meaningful, because an edge +in the graph carries what is flowing along it. The small-message-handoff versus large-buffer- +streaming distinction that `EP-1+.1` named as a minimum bar is therefore **an attribute of an edge**, +not the scenario in its entirety. + +**There is an intermediate artifact, and it is machine-independent.** Stage 1's output -- the +connectivity graph with its directed flow -- is a thing in its own right, derived before any machine +exists. [EP-D-5](#ep-d-5) argued the plan should be a *value* so it can be inspected, compared and +reviewed before anything is pinned or allocated; the same argument applies one stage earlier and +more strongly, because this artifact can be examined without even a synthetic machine. It follows +that the graph has a type rather than being an internal step. + +**The answer is plural.** The planner returns *one or more* suggested realizations, not the +arrangement. This is not hedging: on a given machine several arrangements are defensible, they +differ in ways the planner cannot rank without knowing what the developer values, and presenting +them as candidates is what lets the developer choose. It is the component-level form of OPTION +INTEGRITY in [copilot-instructions.md](../../.github/copilot-instructions.md) -- propose options and +supply the data to choose between them, rather than returning a verdict. A consequence to hold onto: +whatever renders a plan for a human (`M2+.2`) is rendering a *set*, and the difference between +candidates is the part a reader needs most. + +### Why this is the mechanism the adoption thesis needs + +[The adoption thesis](../../DESIGN-NOTES.md#the-adoption-thesis) holds that the principal barrier to +non-uniformity being exploited outside the datacenter is that exploiting it is an architectural +commitment demanded at the beginning of a design, when the least is known. **This component is how +that commitment is removed.** The developer describes their own dataflow -- which they must know +anyway, and which is a statement about their application rather than about any machine -- and never +makes a topology decision at all. The locality reasoning happens here, against a machine model, at a +point where the machine is actually known. + +That is also why the two stages are separated rather than fused. Stage 1 depends only on the +application, so it is stable across every machine the application will ever run on; stage 2 is where +a machine enters, and is the only part that must be redone when the machine changes. Fusing them +would make the application's own structure re-derivable only in the presence of a machine, which is +precisely the coupling the thesis objects to. + +### What this does not settle + +- **The concrete vocabulary of the dataflow description.** "Input, output, and processing code + paths" states the *shape*; the types, and how a caller expresses an edge's characteristics, are + `EP-1+.1`'s remaining work. +- **Where the two new types live.** [EP-D-5](#ep-d-5)'s rule -- a component that only describes or + realizes must not depend on planning policy -- applies to both the dataflow description and the + connectivity graph, and points at `topology-model` for the same reason it placed the plan type + there. That is an argument, not yet a decision, and it is queued rather than taken here. +- **How many candidates, and how they are ordered.** "One or more" is the contract; whether the + planner bounds the set, and whether it orders candidates at all given that ranking is what it + declines to do, is `M3+` policy work. +- **Whether stage 1 can fail.** A description whose connectivity cannot be realized is possible, and + nothing here says what happens then. Related to `EP-1.4`, which asks the same question for an + unanswered model query. \ No newline at end of file diff --git a/crates/topology-planner/DESIGN-RATIONALE.md b/crates/topology-planner/DESIGN-RATIONALE.md index b81e1f0ce..f5f78a963 100644 --- a/crates/topology-planner/DESIGN-RATIONALE.md +++ b/crates/topology-planner/DESIGN-RATIONALE.md @@ -35,3 +35,66 @@ EP-D-4 intentionally left several follow-ups unresolved: These are tracked as checklist work in [CHECKLIST.md](CHECKLIST.md) as `EP-1.4` (not-observed behavior), `EP-1+.3` (planner/model/adapter naming), and `EP-1+.4` (measurement ownership), rather than as canonical decisions. + +## Why the goal turned out to be a dataflow description, and the answer plural + +[DESIGN-NOTES.md](DESIGN-NOTES.md#ep-d-6) records `EP-D-6`. This is how it was reached. + +The goal's shape had been deferred since `EP-D-4` (2026-09-03) and the deferral was **named**, which +is what made it survivable: `COMPONENT.md`, the checklist status table and the decision body all said +"deferred for litigation" rather than quietly omitting the input, so nothing was built on a guess in +the meantime. + +**That is the smaller half of what the deferral was worth, and the larger half is easy to state +backwards.** The answer was not sitting formed on 2026-09-03 waiting to be asked for. It did not +exist. What produced it was the work done in the interval -- the building blocks, then the +measurement tools, then the experiments that used those tools to infer things -- and the clarity +arrived as an *output* of that sequence. The deferral's real value was buying the interval, not +merely guarding it. Writing this up as "the shape was withheld until 2026-09-23" would invert the +causality and quietly teach that asking earlier and harder would have worked; it would not have, and +a specific question put to a general sense would have manufactured a lower-confidence answer that +then got recorded as a decision. (That failure mode is the subject of RESOLUTION GRADIENT in +[copilot-instructions.md](../../.github/copilot-instructions.md).) + +It was stated on 2026-09-23 by the engineer describing the component's purpose, in the course of +correcting a milestone that had been written in the wrong crate. + +**Two candidate shapes had been implicitly in play, and neither was what was chosen.** `EP-1+.1` +framed the input as a *scenario* -- "what the caller intends to run" -- with a minimum bar of +distinguishing small-message handoff from large-buffer streaming. That framing is a set of workload +*characteristics*, and it would have made the planner's input a bag of tuning hints. The other +implicit shape was a *goal* in the literal sense, some statement of what to optimize (latency, +throughput, footprint), which would have made the planner a solver over an objective function. + +What was chosen is neither: the input describes **the application's own structure** -- its inputs, +its outputs, and the processing paths between them. The characteristics `EP-1+.1` named do not +disappear, but they demote from being the scenario to being **attributes of an edge** in that +structure, which is a strictly more informative place for them: "large buffers" is not a property of +a workload, it is a property of a particular flow within it, and a real application has several +flows that differ. + +**The two-stage split was not stated as a separate decision and follows from the input's shape.** +Once the input is the application's structure rather than a set of hints, the connectivity implied by +that structure can be derived with no machine present at all -- and a derivation that does not need a +machine should not be entangled with one. That yields an intermediate artifact that is stable for the +life of the application, where only the second stage is redone per machine. The alternative, deriving +connectivity and placement together, would make the application's own shape re-derivable only in the +presence of a machine, which is the coupling the +[adoption thesis](../../DESIGN-NOTES.md#the-adoption-thesis) exists to object to. + +**The plural answer is the part most likely to be eroded later, so the reason is recorded here.** A +single returned plan is easier to consume, easier to test, and easier to document, and every one of +those pressures argues for collapsing the set at some future convenient moment. The reason not to is +that ranking candidates requires knowing what the developer values, which is the one thing this +component structurally does not know -- it was given a description of an application, not a statement +of preference. A planner that returns one arrangement has either acquired a preference it was not +given or hidden a choice it was not entitled to make. That is the same argument as OPTION INTEGRITY +in [copilot-instructions.md](../../.github/copilot-instructions.md), arriving at component scale +rather than at documentation scale. + +**What was deliberately not decided**, and is queued instead: where the two new types live. +`EP-D-5`'s placement rule points at `topology-model` for both, and the argument is recorded in +`EP-D-6` as an argument. Taking it in the same breath as the decision it follows from would have made +one decision carry two, and the placement question has a consequence -- whether a caller can hold a +dataflow description without depending on planning policy -- that deserves to be litigated on its +own. It is `EP-1+.5`. \ No newline at end of file diff --git a/crates/topology-planner/PLANS.md b/crates/topology-planner/PLANS.md index d515fe976..20aba3f79 100644 --- a/crates/topology-planner/PLANS.md +++ b/crates/topology-planner/PLANS.md @@ -2,4 +2,4 @@ | Path to CHECKLIST.md | Status | Brief description | Design Notes | |---|---|---|---| -| [CHECKLIST.md](CHECKLIST.md) | in progress | Active checklist maintenance for requirements/planning milestones; code implementation milestones are parked. | [DESIGN-NOTES.md](DESIGN-NOTES.md), [DESIGN-RATIONALE.md](DESIGN-RATIONALE.md) | +| [CHECKLIST.md](CHECKLIST.md) | in progress | Active checklist maintenance for requirements/planning milestones; code implementation milestones are parked. **M1+ was re-planned 2026-09-23** by [EP-D-6](DESIGN-NOTES.md#ep-d-6), which settled the input's shape as a dataflow description of the application -- discharging the deferral EP-D-4 named -- and established two-stage planning with a plural answer. That narrowed `EP-1+.1` to vocabulary, widened `EP-1+.3` from two graphs to four, only two of which are graphs of processors, and added `EP-1+.5` and `EP-1+.6`. | [DESIGN-NOTES.md](DESIGN-NOTES.md), [DESIGN-RATIONALE.md](DESIGN-RATIONALE.md) | diff --git a/crates/win-numa-sys/CHECKLIST.md b/crates/win-numa-sys/CHECKLIST.md new file mode 100644 index 000000000..c5180dd13 --- /dev/null +++ b/crates/win-numa-sys/CHECKLIST.md @@ -0,0 +1,44 @@ +# Checklist: win-numa-sys + +Memory-safe Rust over the Windows NUMA APIs. See [README.md](README.md) for what the crate is. + +## M1 -- Finish removing the duplication the crate was created to remove + +The crate exists because two independent `VirtualAllocExNuma` implementations had appeared in +crates that do not depend on each other. Creating it moved one of them; this milestone removes the +other, and until it does the duplication is **relocated rather than removed**, which is worth saying +plainly rather than counting the crate as done. + +- [ ] **N-1.1** -- **Collapse `windows-placement-probe`'s allocator onto this crate.** + `peer_index_cache.rs` has its own `VirtualAllocExNuma` + `VirtualFree` pair, and that crate does + not depend on `windows-ioring-sys`, so it never saw the one that was hoisted in `M22.3`. + + **It is not a drop-in, and the differences are the work.** That allocator also faults every page + in -- committed pages are demand-zero, so until something writes to them no physical page has been + drawn from the preferred node -- and then *observes* which node the pages actually landed on. It + is typed over its own `Slot` rather than bytes, and it has a second origin (an ordinary heap + allocation) that shares the same `Drop`. + + Decide what of that is general before moving any of it. Page-faulting looks general: any consumer + who cares where pages landed has to do it, and this crate's own documentation already tells them + so. The typed element and the dual origin look specific to the probe. + +- [ ] **N-1.2** -- **Decide whether `QueryWorkingSetEx` observation moves here**, which `N-1.1` will + force a view on. `observed_node_of_region`, `observed_node`, `working_set_flags` and the + `working_set` bit constants live in `windows-placement-probe` today. "Which node is this page + actually on" is a Windows NUMA concept and passes this crate's bar; against that, those helpers + carry their own tests for the bit layout, and `working_set_flags` exists *separately from* + `observed_node` precisely so a test can check the layout against a field whose value it already + knows. Moving them means moving that care too, not just the code. + + Note what this would make possible, since it is the reason to consider it at all: a caller could + then ask this crate whether a placement request was honoured, rather than being told by its + documentation that success proves nothing and left to write `QueryWorkingSetEx` themselves. + +## M2+ -- Parked + +- [ ] **M2+.1** -- **Publish.** The crate is a path dependency of `windows-ioring-sys`, which *is* + published, so it has to reach crates.io before that crate's next release or the release fails. It + is already registered in `release-please-config.json` and in the publish workflow's tag patterns + -- both of which this repository has previously been bitten by omitting, silently -- so what + remains is the decision to cut `0.1.0`, not the plumbing. diff --git a/crates/win-numa-sys/Cargo.toml b/crates/win-numa-sys/Cargo.toml new file mode 100644 index 000000000..8bae51985 --- /dev/null +++ b/crates/win-numa-sys/Cargo.toml @@ -0,0 +1,41 @@ +# Copyright (c) 2026 Mike Grier + +[package] +name = "win-numa-sys" +version = "0.1.0" +authors.workspace = true +edition.workspace = true +rust-version.workspace = true +license.workspace = true +repository.workspace = true +homepage.workspace = true +description = "Memory-safe Rust over the Windows NUMA APIs: a VirtualAllocExNuma-backed buffer, a NumaNode newtype, and the node queries a caller needs to decide what to pass. Thin over Win32, with no policy about which node anything should use." + +# The first crate to take the `win-` prefix rather than `windows-`; see +# DESIGN-NOTES.md at the repository root. `windows` is Microsoft's namespace, +# and a crate published as `windows-numa-sys` today is a name they may +# reasonably want tomorrow. + +[dependencies] +# `default-features = false` matches every other crate in the workspace. +# +# `Win32_System_Memory` is `VirtualAllocExNuma` and `VirtualFree`. +# `Win32_System_Threading` is `GetCurrentProcess` and +# `GetNumaHighestNodeNumber` -- the latter lives there rather than under +# `SystemInformation`, which is where its subject matter would suggest. +# `Win32_System_IO` and `Win32_System_Ioctl` are `DeviceIoControl` and +# `FSCTL_QUERY_VOLUME_NUMA_INFO`, for asking a handle's volume which node it +# reports. +windows-sys = { version = "0.61.2", default-features = false, features = [ + "Win32_Foundation", + "Win32_System_IO", + "Win32_System_Ioctl", + "Win32_System_Memory", + "Win32_System_Threading", +] } + +[package.metadata.docs.rs] +# The crate is Windows-only, so docs.rs must build on a Windows target or it +# would render an almost-empty crate. +default-target = "x86_64-pc-windows-msvc" +targets = ["x86_64-pc-windows-msvc", "aarch64-pc-windows-msvc"] diff --git a/crates/win-numa-sys/DESIGN-NOTES.md b/crates/win-numa-sys/DESIGN-NOTES.md new file mode 100644 index 000000000..95d9bca80 --- /dev/null +++ b/crates/win-numa-sys/DESIGN-NOTES.md @@ -0,0 +1,79 @@ +# Design notes: win-numa-sys (Tier 1) + +Current canonical decisions for this crate. See [README.md](README.md) for what it is and +[CHECKLIST.md](CHECKLIST.md) for what is planned. + +Repository-level decisions this crate sits under: the `win-` prefix and what `-sys` promises, both in +the workspace [DESIGN-NOTES.md](../../DESIGN-NOTES.md#new-crates-take-the-win-prefix). + +## Decision index + +| ID | Decision | +|---|---| +| N-D-1 | **Both ways of arriving at a node are offered; choosing between them is not.** A caller may *declare* a node (`NumaBuffer::new`) or *discover* one (`volume_numa_node`), and may qualify what a discovered answer is worth (`highest_numa_node`). What the crate refuses is a constructor that does both in one step -- there is no `NumaBuffer::for_file`. | + +## N-D-1: both ways of arriving at a node, and no shortcut between them + +*Recorded 2026-09-23 by `M23.2` in +[windows-ioring-sys/CHECKLIST.md](../windows-ioring-sys/CHECKLIST.md), which asked this crate's +predecessor whether it should accept a declared storage node as an input.* + +### What is offered + +- **Declare.** `NumaBuffer::new(len, Some(node))` allocates on a node the caller names. The caller + may have got that node from anywhere -- a topology walk, a configuration file, a measurement, a + coin toss. This crate does not ask. +- **Discover.** `volume_numa_node(handle)` asks a handle's volume which node it reports, and returns + it. It does not allocate anything. +- **Qualify.** `highest_numa_node()` answers how many nodes exist, so a caller can tell a real choice + from the only choice available. + +Both paths have consumers in this workspace already, which is the evidence that neither is +speculative: `examples/epoch_log` discovers from its log file's volume, and `examples/ring_copy` +declares a node it computed from the processor topology. + +### What is refused, and why + +**There is no `NumaBuffer::for_file(handle, len)`** -- no call that queries a node and allocates on +it in one step. It is the obvious convenience and it is the wrong shape: + +- **It hides the answer.** On a single-node machine, "placed on the node the volume named" and "no + preference" are the *same allocation*, so a caller could not tell whether the query had found + anything. That is the failure the epoch-log sample's report line exists to prevent, and it is the + 2026-08-30 session's *report, do not route* position applied one layer down. +- **It fuses two failure domains.** Allocation can fail; the query can fail. A combined call has to + decide what happens when only the query fails -- allocate unplaced, which silently does something + other than asked, or return an error, which fails an allocation that would have succeeded. Neither + is right for every caller, so neither should be baked in. Split, the caller decides. +- **The answer is worth more than one buffer.** A node can inform a thread's affinity, a second + allocation, a report, or a plan. Tying the query to one constructor means the second consumer + writes it again -- which is how this crate came to exist. +- **And the convenience is already there.** `NumaBuffer::new(len, volume_numa_node(h).ok())` type- + checks as written, because the query returns exactly what the constructor takes. The shortcut would + save one line. + +**An argument that no longer applies, recorded so it is not re-made:** when this was first argued, +the buffer lived in `windows-ioring-sys` and a `for_file` constructor would have dragged +`Win32_System_Ioctl` into a crate that otherwise touched only memory. That was a real cost then. It +is not one now -- `volume_numa_node` lives here and the feature is already declared -- so the +argument is void and the four above are what the decision rests on. + +### Why discovery is offered at all + +The discovered answer is weak, and `volume_numa_node`'s own documentation says so: it answers for a +*volume*, a volume may span devices, and then a single reported node is a fiction rather than an +answer. A defensible reading is that a query this weak should not be offered. + +It is offered anyway, because the alternative is worse. Refusing it does not stop a consumer needing +the answer; it makes each one write the `DeviceIoControl` themselves. That is not hypothetical -- +this crate exists because `VirtualAllocExNuma` had been written three times in this workspace on +exactly that logic, and the epoch-log sample had hand-rolled this very FSCTL. The honest move is to +provide it **and say plainly what it is worth**, which puts the caveat where the caller will read it +rather than leaving them to discover the limitation themselves. + +### What this preserves + +[D-8](../windows-ioring-sys/DESIGN-NOTES.md#d-8) -- locality is the consumer's decision -- is intact +and is the reason the shape is what it is. This crate supplies a fact and an allocator. It does not +map a file to a node, does not shard anything, and does not choose on a caller's behalf. A consumer +who wants those things composes them from what is here, on hardware this workspace has never seen. diff --git a/crates/win-numa-sys/PLANS.md b/crates/win-numa-sys/PLANS.md new file mode 100644 index 000000000..f2ab79504 --- /dev/null +++ b/crates/win-numa-sys/PLANS.md @@ -0,0 +1,5 @@ +# Plans: win-numa-sys + +| Path to CHECKLIST.md | Status | Brief description | Design Notes | +|---|---|---|---| +| [CHECKLIST.md](CHECKLIST.md) | in progress | M1 finishes what creating the crate started: `windows-placement-probe` still has its own `VirtualAllocExNuma`, so the duplication is currently **relocated rather than removed**, and `N-1.1` is where that is closed. `N-1.2` decides whether `QueryWorkingSetEx` observation moves here with it -- the capability that would let a caller ask whether a placement request was honoured, rather than being told by the documentation that success proves nothing. `M2+.1` is publishing, which is not optional indefinitely: the crate is a path dependency of the published `windows-ioring-sys`. | [DESIGN-NOTES.md](DESIGN-NOTES.md) | diff --git a/crates/win-numa-sys/README.md b/crates/win-numa-sys/README.md new file mode 100644 index 000000000..924ecf47b --- /dev/null +++ b/crates/win-numa-sys/README.md @@ -0,0 +1,60 @@ +# win-numa-sys + +Memory-safe Rust over the Windows NUMA APIs. + +Windows provides some NUMA concepts; this crate builds library notions on top of them. Whether a +client uses them is up to the client. + +## What is here + +- **`NumaNode`** -- a node number, as Windows numbers them. A newtype because a node *number* and a + *count* of nodes are both small integers and are trivially swapped at a call site. +- **`NumaBuffer`** -- an owned allocation made with `VirtualAllocExNuma` and released with + `VirtualFree`, so a buffer can be placed on a chosen node rather than wherever the default + allocator lands it. +- **`highest_numa_node()`** and **`volume_numa_node(handle)`** -- the two questions Windows will + answer about nodes, wrapped so a caller can find a node to pass without writing the FFI. + +## What is not here, and will not be + +**Any opinion about which node anything should use.** This crate reports what Windows says and +allocates where it is told. It does not map a file to a node, does not shard anything, and does not +choose a node on a caller's behalf. + +The `-sys` suffix is that promise. In this workspace the suffix means thin over Win32, memory-safe, +adding no policy -- see the repository's +[DESIGN-NOTES.md](../../DESIGN-NOTES.md#the-waitable-queues-crate-is-named-plural-and-carries-no-sys-suffix), +where a sibling crate drops the suffix for failing exactly that test. + +## The name + +The first crate here to take `win-` rather than `windows-`. `windows` is Microsoft's namespace, and +a crate published as `windows-numa-sys` today is a name they may reasonably want tomorrow; see +[DESIGN-NOTES.md](../../DESIGN-NOTES.md#new-crates-take-the-win-prefix). Existing crates keep their +names for now. + +## A node argument is a preference, not an instruction + +`VirtualAllocExNuma`'s parameter is `nndPreferred`, and the name is the contract. A successful +allocation is **not** evidence that the pages landed on the node that was asked for. Two things +follow, both of which have already caught someone in this workspace: + +- Committed pages are demand-zero, so until something writes to them no physical page has been drawn + from the preferred node at all. Measuring placement means faulting the pages in first. +- An *invalid* node is refused -- but `u32::MAX` is not a test of that, because it is the API's own + no-preference sentinel and is accepted by design. A measurement that asks for `u32::MAX` and sees + it succeed has measured the sentinel, not a range check. + +Observing where pages actually landed needs `QueryWorkingSetEx`, which lives in +`windows-placement-probe` today; whether it moves here is [CHECKLIST.md](CHECKLIST.md) `N-1.2`. + +## Buffer traits live with whoever owns them + +`NumaBuffer` implements no I/O buffer trait, because this crate defines none. `windows-ioring-sys` +and `windows-overlapped-io-sys` each already have their own `IoBuf`/`IoBufMut` pair, and a third +copy here would have made that duplication harder to resolve rather than easier. + +A consumer that owns such a trait implements it for `NumaBuffer` -- the orphan rule permits exactly +that, since the trait is theirs -- over the inherent `as_ptr`, `as_mut_ptr` and `len`. +`windows-ioring-sys` does so, and keeps re-exporting `NumaBuffer` so `windows_ioring_sys::NumaBuffer` +still resolves. diff --git a/crates/win-numa-sys/src/buffer.rs b/crates/win-numa-sys/src/buffer.rs new file mode 100644 index 000000000..743a538d1 --- /dev/null +++ b/crates/win-numa-sys/src/buffer.rs @@ -0,0 +1,182 @@ +// Copyright (c) 2026 Mike Grier +// Moved from windows-ioring-sys/src/numa_buffer.rs at 6101ca65, which had in +// turn moved it from examples/ring_copy/buffer.rs at 834c7afa. +//! A `VirtualAllocExNuma`-backed buffer, so a buffer can be placed on a chosen +//! NUMA node rather than wherever the default allocator's own heuristics land +//! it. +//! +//! # Why this is its own crate +//! +//! It was in `windows-ioring-sys`, which recommends placing a registered pool +//! near the device and so had a reason to provide the allocator rather than +//! leave every caller to write it. But the buffer has nothing to do with a +//! ring: its only connection was one doc line saying it *can* be registered +//! into one. Meanwhile `windows-placement-probe` -- which does not depend on +//! the ring crate -- had written the same `VirtualAllocExNuma` call for itself, +//! making two independent allocators in a workspace that had already hoisted +//! this code once to stop exactly that. +//! +//! # What this does not decide +//! +//! **Which node.** A caller who knows their storage or thread layout passes +//! `Some(node)`; one who does not passes `None` and gets the system's own +//! choice, which is what the default allocator would have given anyway. +//! [`volume_numa_node`](crate::volume_numa_node) is available for finding a +//! candidate, and says plainly what its answer is and is not worth. +//! +//! A node argument is a *preference*, which the underlying parameter says in +//! its own name (`nndPreferred`). So a successful allocation is not by itself +//! evidence that the pages landed on the node that was asked for, and code +//! that needs to know must measure rather than assume -- see the crate root. + +use std::io; +use std::ptr; + +use windows_sys::Win32::System::Memory::{ + MEM_COMMIT, MEM_RELEASE, MEM_RESERVE, PAGE_READWRITE, VirtualAllocExNuma, VirtualFree, +}; +use windows_sys::Win32::System::Threading::GetCurrentProcess; + +use crate::NumaNode; + +/// `VirtualAllocExNuma`'s documented sentinel for "no NUMA preference" -- +/// windows-sys does not name this constant, so it is named here rather than +/// written as a bare literal at the call site. +const NUMA_NO_PREFERRED_NODE: u32 = u32::MAX; + +/// An owned buffer allocated with `VirtualAllocExNuma`, freed with +/// `VirtualFree` on drop. +/// +/// The allocation is page-granular: `VirtualAllocExNuma` rounds a request up +/// to the system page size, so a caller asking for a small buffer gets at +/// least a page. [`NumaBuffer::len`] reports the length that was *requested*, +/// which is the length a kernel call should be told about. +/// +/// # Buffer traits live with whoever owns them +/// +/// This type implements no I/O buffer trait, because this crate defines none +/// and should not: `windows-ioring-sys` and `windows-overlapped-io-sys` each +/// have their own `IoBuf`/`IoBufMut`, and a third copy here would be one more +/// of a thing the workspace is already deciding what to do about. A consumer +/// that owns such a trait implements it for this type -- the orphan rule +/// permits exactly that, since the trait is theirs -- over +/// [`NumaBuffer::as_ptr`], [`NumaBuffer::as_mut_ptr`] and [`NumaBuffer::len`]. +/// `windows-ioring-sys` does so. +pub struct NumaBuffer { + ptr: *mut u8, + len: usize, +} + +// SAFETY: the allocation is exclusively owned by this value; sending it +// across threads only moves that ownership, never aliases it. +unsafe impl Send for NumaBuffer {} + +impl std::fmt::Debug for NumaBuffer { + /// Shows the address and requested length, never the contents: a + /// registered buffer routinely holds someone's data, and a `Debug` that + /// printed it would put that data anywhere a caller logs. + fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result { + f.debug_struct("NumaBuffer") + .field("ptr", &self.ptr) + .field("len", &self.len) + .finish() + } +} + +impl NumaBuffer { + /// Allocate `len` bytes, preferring `node` if given. + /// + /// Passing `None` requests no NUMA preference, which is the same choice + /// the default allocator makes implicitly. + /// + /// # What a successful return does and does not mean + /// + /// A valid node is a **preference**, not an instruction -- `nndPreferred` + /// says so in its name -- so success means the request was accepted and + /// memory was obtained, never that the pages came from the node asked for. + /// + /// An *invalid* node is a different matter and is genuinely refused: this + /// crate's tests assert that `u32::MAX - 1` fails. Note that `u32::MAX` + /// itself is **not** a test of that, because it is the API's own + /// no-preference sentinel and is accepted by design -- a measurement that + /// reads it as an out-of-range node being tolerated has measured the + /// sentinel instead. + /// + /// # Errors + /// + /// The error from `VirtualAllocExNuma`, which includes a `len` of zero and + /// a node number the machine does not have. + pub fn new(len: usize, node: Option) -> io::Result { + // SAFETY: no pointer arguments; the returned value is a pseudo-handle + // that needs no closing. + let process = unsafe { GetCurrentProcess() }; + // SAFETY: `process` is a valid pseudo-handle for the duration of this + // call; a null `lpAddress` lets the system choose the address. + let ptr = unsafe { + VirtualAllocExNuma( + process, + ptr::null(), + len, + MEM_COMMIT | MEM_RESERVE, + PAGE_READWRITE, + node.map_or(NUMA_NO_PREFERRED_NODE, NumaNode::get), + ) + }; + if ptr.is_null() { + return Err(io::Error::last_os_error()); + } + Ok(Self { + ptr: ptr.cast(), + len, + }) + } + + /// The allocation's base address, fixed for this value's life. + #[must_use] + pub fn as_ptr(&self) -> *const u8 { + self.ptr + } + + /// The allocation's base address for writing. + /// + /// Takes `&mut self` because this value uniquely owns the allocation, so + /// exclusive access to the value is exclusive access to the bytes. + #[must_use] + pub fn as_mut_ptr(&mut self) -> *mut u8 { + self.ptr + } + + /// The length that was requested. + /// + /// Not the length that was reserved: `VirtualAllocExNuma` rounds up to a + /// page, so the mapping is at least this large and usually larger. This is + /// the number a kernel call should be told, because it is the number the + /// caller asked to use. + #[must_use] + pub fn len(&self) -> usize { + self.len + } + + /// Whether the requested length was zero. + /// + /// Present because clippy asks for it beside [`Self::len`], and because a + /// zero-length request is rejected by `VirtualAllocExNuma` rather than + /// producing an empty buffer -- so on any value that exists, this is false. + #[must_use] + pub fn is_empty(&self) -> bool { + self.len == 0 + } +} + +impl Drop for NumaBuffer { + fn drop(&mut self) { + // SAFETY: `self.ptr` was returned by `VirtualAllocExNuma` above and + // is freed exactly once, here. + unsafe { + VirtualFree(self.ptr.cast(), 0, MEM_RELEASE); + } + } +} + +#[cfg(test)] +mod tests; diff --git a/crates/win-numa-sys/src/buffer/tests.rs b/crates/win-numa-sys/src/buffer/tests.rs new file mode 100644 index 000000000..d17627ad5 --- /dev/null +++ b/crates/win-numa-sys/src/buffer/tests.rs @@ -0,0 +1,165 @@ +// Copyright (c) 2026 Mike Grier +//! Tests for [`NumaBuffer`] (M22.3). +//! +//! These allocate through `VirtualAllocExNuma` and free on drop. That is an +//! operating-system call, but this repository targets one operating system, so +//! by its own Quality rule that alone does not make these integration tests: +//! they cross no process, device, or network boundary, hold no handle, and +//! each completes in microseconds. +//! +//! # What these cannot establish +//! +//! **That the pages landed on the requested node.** The parameter is +//! `nndPreferred`, so a success proves the request was accepted, not that it +//! was honoured -- and a single-node host, which is what this is developed on, +//! could not tell the difference either way. The type's own documentation says +//! this; these tests do not pretend otherwise by asserting a node back. + +use super::NumaBuffer; +use crate::NumaNode; + +/// A node number no machine has, for the rejection cases. `u32::MAX` is not +/// usable here -- it is `VirtualAllocExNuma`'s own "no preference" sentinel, +/// so it would be accepted rather than refused. +const ABSURD_NODE: u32 = u32::MAX - 1; + +#[test] +fn an_unplaced_buffer_allocates() { + let buffer = NumaBuffer::new(4096, None).expect("an unplaced allocation"); + assert!(!buffer.as_ptr().is_null()); +} + +#[test] +fn node_zero_allocates() { + // Every machine reports a node 0, including one with NUMA disabled, where + // `GetNumaHighestNodeNumber` answers 0. + let buffer = NumaBuffer::new(4096, Some(NumaNode::new(0))).expect("node 0 exists everywhere"); + assert!(!buffer.as_ptr().is_null()); +} + +#[test] +fn the_reported_length_is_the_requested_one_not_the_rounded_one() { + // The allocation is page-granular, so the *mapping* is at least a page. + // What a kernel call is told about an operation is `len`, and that must + // be what the caller asked for -- reporting the rounded-up size would + // invite a read or write past the caller's intent. + for len in [1_usize, 100, 4095, 4096, 4097, 65536] { + let buffer = NumaBuffer::new(len, None).expect("a valid allocation"); + assert_eq!(buffer.len(), len, "for a request of {len} bytes"); + } +} + +#[test] +fn a_fresh_allocation_is_zeroed() { + // `VirtualAlloc`-family pages arrive zeroed, which is what lets a caller + // register an arena without filling it first. + let mut buffer = NumaBuffer::new(8192, None).expect("a valid allocation"); + // SAFETY: the pointer is this buffer's own allocation of `len` + // bytes, and `&mut` makes the borrow exclusive. + let bytes = unsafe { std::slice::from_raw_parts(buffer.as_mut_ptr(), buffer.len()) }; + assert!(bytes.iter().all(|&b| b == 0)); +} + +#[test] +fn bytes_written_read_back() { + let mut buffer = NumaBuffer::new(4096, None).expect("a valid allocation"); + let len = buffer.len(); + // SAFETY: as above -- this buffer's own allocation, exclusively borrowed. + let bytes = unsafe { std::slice::from_raw_parts_mut(buffer.as_mut_ptr(), len) }; + for (index, byte) in bytes.iter_mut().enumerate() { + *byte = (index % 251) as u8; + } + // SAFETY: as above. + let read = unsafe { std::slice::from_raw_parts(buffer.as_ptr(), len) }; + assert!(read.iter().enumerate().all(|(i, &b)| b == (i % 251) as u8)); +} + +#[test] +fn the_address_is_stable_across_reads() { + // The whole reason this type can be registered: the address it reports + // does not move for its life. + let mut buffer = NumaBuffer::new(4096, None).expect("a valid allocation"); + let first = buffer.as_ptr(); + assert_eq!(buffer.as_ptr(), first); + assert_eq!(buffer.as_mut_ptr().cast_const(), first); + assert_eq!(buffer.as_ptr(), first); +} + +#[test] +fn an_address_survives_a_move() { + // The address must be stable for the value's life, + // which includes being moved -- the allocation is behind a pointer, so + // moving the handle does not move the bytes. + let buffer = NumaBuffer::new(4096, None).expect("a valid allocation"); + let before = buffer.as_ptr(); + let moved = buffer; + assert_eq!(moved.as_ptr(), before); +} + +#[test] +fn separate_buffers_do_not_alias() { + let buffers: Vec = (0..8) + .map(|_| NumaBuffer::new(4096, None).expect("a valid allocation")) + .collect(); + let mut addresses: Vec<*const u8> = buffers.iter().map(NumaBuffer::as_ptr).collect(); + addresses.sort_unstable(); + let before = addresses.len(); + addresses.dedup(); + assert_eq!(addresses.len(), before, "every slot must be its own memory"); +} + +#[test] +fn a_zero_length_request_is_refused() { + let error = NumaBuffer::new(0, None).expect_err("a zero-length mapping is not allocatable"); + assert!(error.raw_os_error().is_some(), "and it is an OS error"); +} + +#[test] +fn a_node_the_machine_does_not_have_is_refused() { + // The other half of the guard: the accepting cases above would all still + // pass if `new` ignored its node argument entirely. + let error = NumaBuffer::new(4096, Some(NumaNode::new(ABSURD_NODE))) + .expect_err("no machine has this node number"); + assert!(error.raw_os_error().is_some(), "and it is an OS error"); +} + +#[test] +fn a_request_the_address_space_cannot_hold_is_refused() { + let error = NumaBuffer::new(usize::MAX, None).expect_err("no process has this much space"); + assert!(error.raw_os_error().is_some(), "and it is an OS error"); +} + +#[test] +fn a_refused_allocation_leaves_the_allocator_usable() { + // A failed `VirtualAllocExNuma` must not leave anything behind that stops + // the next one -- the sample arenas allocate in a loop, so one bad request + // in the middle would otherwise take the rest with it. + let _ = NumaBuffer::new(4096, Some(NumaNode::new(ABSURD_NODE))).expect_err("refused"); + let buffer = NumaBuffer::new(4096, None).expect("the next allocation still works"); + assert!(!buffer.as_ptr().is_null()); +} + +#[test] +fn many_buffers_allocate_and_free() { + // Drop is what returns the mapping; leaking it would show up here as an + // address-space exhaustion long before the loop ends. + for _ in 0..512 { + let buffer = NumaBuffer::new(65536, None).expect("a valid allocation"); + assert!(!buffer.as_ptr().is_null()); + } +} + +#[test] +fn a_buffer_is_send() { + // The `unsafe impl Send` is load-bearing: a domain runtime allocates on + // one thread and uses the buffer on the pinned thread that owns the ring. + fn assert_send() {} + assert_send::(); + + let buffer = NumaBuffer::new(4096, None).expect("a valid allocation"); + let address = buffer.as_ptr() as usize; + let moved = std::thread::spawn(move || buffer.as_ptr() as usize) + .join() + .expect("the thread completes"); + assert_eq!(moved, address, "and the address travels with it"); +} diff --git a/crates/win-numa-sys/src/lib.rs b/crates/win-numa-sys/src/lib.rs new file mode 100644 index 000000000..c4bf991e7 --- /dev/null +++ b/crates/win-numa-sys/src/lib.rs @@ -0,0 +1,55 @@ +// Copyright (c) 2026 Mike Grier +//! Memory-safe Rust over the Windows NUMA APIs. +//! +//! Windows provides some NUMA concepts; this crate builds library notions on +//! top of them. Whether a client uses them is up to the client. +//! +//! # What is here +//! +//! - [`NumaNode`] -- a node number, as Windows numbers them. +//! - [`NumaBuffer`] -- an owned allocation made with `VirtualAllocExNuma` and +//! released with `VirtualFree`, so a buffer can be placed on a chosen node +//! rather than wherever the default allocator lands it. +//! - [`highest_numa_node`] and [`volume_numa_node`] -- the two questions +//! Windows will answer about nodes, wrapped so a caller can find a node to +//! pass without writing the FFI themselves. +//! +//! # What is not here, and will not be +//! +//! **Any opinion about which node anything should use.** This crate reports +//! what Windows says and allocates where it is told. It does not map a file to +//! a node, does not shard anything, and does not choose a node on a caller's +//! behalf. Those are workload decisions, and a consumer on hardware this +//! workspace has never seen is better placed to take them. +//! +//! The `-sys` suffix is that promise, and it is the repository's meaning of the +//! suffix rather than a convention borrowed from elsewhere: thin over Win32, +//! memory-safe, adding no policy. +//! +//! # A node argument is a preference, not an instruction +//! +//! The underlying parameter says so in its own name -- `nndPreferred`. A +//! successful allocation is therefore **not** evidence that the pages landed on +//! the node that was asked for, and code that needs to know must observe rather +//! than assume. Two measured facts about that, both from this workspace: +//! +//! - Committed pages are demand-zero, so until something writes to them no +//! physical page has been drawn from the preferred node at all. A caller +//! measuring placement must fault the pages in first. +//! - An *invalid* node is refused, and `u32::MAX` is not a test of that: it is +//! the API's own no-preference sentinel, accepted by design. A measurement +//! that asks for `u32::MAX` and sees it succeed has measured the sentinel, +//! not a range check. [`NumaBuffer`]'s tests use `u32::MAX - 1`. +//! +//! Observing where pages actually landed needs `QueryWorkingSetEx`, which lives +//! in `windows-placement-probe` today and is a candidate to move here; see that +//! crate's `peer_index_cache`. + +#![cfg(windows)] +#![deny(missing_docs)] + +mod buffer; +mod node; + +pub use buffer::NumaBuffer; +pub use node::{NumaNode, highest_numa_node, volume_numa_node}; diff --git a/crates/win-numa-sys/src/node.rs b/crates/win-numa-sys/src/node.rs new file mode 100644 index 000000000..f08a29428 --- /dev/null +++ b/crates/win-numa-sys/src/node.rs @@ -0,0 +1,140 @@ +// Copyright (c) 2026 Mike Grier +//! [`NumaNode`], and the two questions Windows will answer about nodes. + +use std::io; +use std::os::windows::io::RawHandle; + +use windows_sys::Win32::Foundation::HANDLE; +use windows_sys::Win32::System::IO::DeviceIoControl; +use windows_sys::Win32::System::Ioctl::FSCTL_QUERY_VOLUME_NUMA_INFO; +use windows_sys::Win32::System::Threading::GetNumaHighestNodeNumber; + +/// A NUMA node number, as Windows numbers them. +/// +/// A newtype rather than a bare `u32` because the two `u32`s in this area mean +/// different things and are trivially swapped at a call site: a node *number* +/// and a *count* of nodes are both small integers, and +/// [`highest_numa_node`] returns the former while reading like the latter. +/// +/// **Node numbers are not an index.** Windows does not promise they run +/// `0..n`, so a caller deriving a node from a position in some list is making +/// an assumption the platform never offered -- see `windows-placement-probe`, +/// where a positional index reaching `VirtualAllocExNuma` was a real defect and +/// is called out at the site that fixed it. +#[derive(Clone, Copy, Debug, PartialEq, Eq, PartialOrd, Ord, Hash)] +pub struct NumaNode(u32); + +impl NumaNode { + /// Wrap a raw node number. + /// + /// Not validated against the machine: `VirtualAllocExNuma` is where an + /// impossible node is rejected, and validating here would mean a second, + /// weaker opinion about what exists. See [`NumaBuffer::new`] for what the + /// platform actually does with an out-of-range node, which is not what its + /// documentation says. + /// + /// [`NumaBuffer::new`]: crate::NumaBuffer::new + #[must_use] + pub const fn new(node: u32) -> Self { + Self(node) + } + + /// The raw node number, for handing to a Win32 call. + #[must_use] + pub const fn get(self) -> u32 { + self.0 + } +} + +impl std::fmt::Display for NumaNode { + fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result { + write!(f, "node {}", self.0) + } +} + +/// The highest node number this machine reports, if it answers. +/// +/// `GetNumaHighestNodeNumber`. The value is a *node number*, not a count: a +/// machine with one node answers `NumaNode(0)`, not one. +/// +/// # Why a caller wants this +/// +/// Mostly to qualify a report rather than to drive a choice. On a machine that +/// answers `NumaNode(0)` there is exactly one node, so every placement decision +/// is the same decision, and a line saying "placed on node 0" is true and +/// misleading. Asking this is how a caller can say which of those it is. +/// +/// # Errors +/// +/// `None` when the call fails, which is reported as "unknown" rather than as a +/// node count of zero -- a machine that will not say how many nodes it has is a +/// different thing from a machine with none. +#[must_use] +pub fn highest_numa_node() -> Option { + let mut highest: u32 = 0; + // SAFETY: the out parameter is a live local for the call's duration. + let ok = unsafe { GetNumaHighestNodeNumber(&raw mut highest) }; + (ok != 0).then_some(NumaNode::new(highest)) +} + +/// The NUMA node `handle`'s **volume** reports, if it reports one. +/// +/// `FSCTL_QUERY_VOLUME_NUMA_INFO`, issued against the handle directly -- the +/// documented control code accepts a file or directory handle, so this needs no +/// device-tree walk and no second open. +/// +/// # This answers a question about a volume, not about a file +/// +/// The documented meaning is the node the *volume* resides on. A volume may +/// span several devices -- an ordinary spanned volume or a Storage Spaces set +/// does -- and then a single reported node is a fiction rather than an answer, +/// because the file's extents may live anywhere across the set. **It therefore +/// cannot tell a caller which node a particular file's I/O is closest to**, even +/// when it succeeds. +/// +/// That is a limit of the question, not of this wrapper, and it is why this +/// function reports rather than acts: a caller who knows their storage layout +/// can decide what the answer is worth, and one who does not should not have a +/// placement chosen for them on the strength of it. +/// +/// # Errors +/// +/// Any error from `DeviceIoControl`, and +/// [`io::ErrorKind::InvalidData`] if the control code returns a payload that is +/// not the documented single `u32`. +pub fn volume_numa_node(handle: RawHandle) -> io::Result { + let mut node: u32 = u32::MAX; + let mut returned: u32 = 0; + // SAFETY: `handle` is the caller's, and outlives this call by the contract + // on this function. The FSCTL takes no input buffer, and its output is + // documented as `FSCTL_QUERY_VOLUME_NUMA_INFO_OUTPUT { ULONG NumaNode }` -- + // one `u32`, which is what `node` provides and what the size argument says. + let ok = unsafe { + DeviceIoControl( + handle as HANDLE, + FSCTL_QUERY_VOLUME_NUMA_INFO, + std::ptr::null(), + 0, + (&raw mut node).cast::(), + u32::try_from(size_of::()).expect("four fits in a u32"), + &raw mut returned, + std::ptr::null_mut(), + ) + }; + if ok == 0 { + return Err(io::Error::last_os_error()); + } + if returned as usize != size_of::() { + return Err(io::Error::new( + io::ErrorKind::InvalidData, + format!( + "FSCTL_QUERY_VOLUME_NUMA_INFO returned {returned} bytes, not {}", + size_of::() + ), + )); + } + Ok(NumaNode::new(node)) +} + +#[cfg(test)] +mod tests; diff --git a/crates/win-numa-sys/src/node/tests.rs b/crates/win-numa-sys/src/node/tests.rs new file mode 100644 index 000000000..0d855f575 --- /dev/null +++ b/crates/win-numa-sys/src/node/tests.rs @@ -0,0 +1,143 @@ +// Copyright (c) 2026 Mike Grier +//! Tests for [`NumaNode`] and the two node queries. +//! +//! # What these can and cannot establish +//! +//! The newtype's tests are total: it is a wrapper over a `u32` and every claim +//! about it holds on any host. +//! +//! The query tests are **host-dependent by nature**, and are written to assert +//! only what is true on every machine rather than what happens to be true on +//! this one. This workspace is developed on a single-node host, so a test that +//! asserted a particular node back would be asserting the only answer available +//! here and would fail on the hardware the crate exists for. What is asserted +//! instead is shape: that a query either answers or says it could not, that it +//! never invents a node, and that asking twice agrees. + +use std::os::windows::io::AsRawHandle; + +use super::{NumaNode, highest_numa_node, volume_numa_node}; + +/// A scratch file to ask about, named per test so tests running as threads in +/// one process cannot collide on it. +fn scratch(tag: &str) -> (std::path::PathBuf, std::fs::File) { + let path = std::env::temp_dir().join(format!( + "win-numa-sys-node-{}-{tag}.tmp", + std::process::id() + )); + let file = std::fs::OpenOptions::new() + .create(true) + .truncate(true) + .write(true) + .open(&path) + .expect("a scratch file in the temp directory"); + (path, file) +} + +#[test] +fn a_node_round_trips_through_the_newtype() { + for raw in [0_u32, 1, 7, 63, u32::MAX - 1, u32::MAX] { + assert_eq!(NumaNode::new(raw).get(), raw, "for {raw}"); + } +} + +#[test] +fn nodes_compare_and_order_by_their_number() { + assert_eq!(NumaNode::new(3), NumaNode::new(3)); + assert_ne!(NumaNode::new(3), NumaNode::new(4)); + assert!(NumaNode::new(3) < NumaNode::new(4)); + let mut nodes = [NumaNode::new(9), NumaNode::new(2), NumaNode::new(5)]; + nodes.sort_unstable(); + assert_eq!( + nodes, + [NumaNode::new(2), NumaNode::new(5), NumaNode::new(9)] + ); +} + +/// The `Display` form says what the number *is*, because a bare integer in a +/// report is ambiguous between a node number and a node count -- the very +/// confusion the newtype exists to stop. +#[test] +fn display_names_the_thing_rather_than_printing_a_bare_integer() { + assert_eq!(NumaNode::new(0).to_string(), "node 0"); + assert_eq!(NumaNode::new(12).to_string(), "node 12"); +} + +/// Whatever the highest node is, asking twice agrees. +/// +/// Deliberately not an assertion about the value: on this workspace's host it +/// is `node 0`, and pinning that would encode the development machine into the +/// suite. Hot-add could in principle change it between calls, which would make +/// this flaky rather than wrong -- it has never been observed, and a failure +/// here would be a genuine finding about the platform rather than a bad test. +#[test] +fn the_highest_node_is_stable_across_calls() { + assert_eq!(highest_numa_node(), highest_numa_node()); +} + +/// A machine that answers reports a node number, not a count. +/// +/// The distinction is the reason the newtype exists: a single-node machine +/// answers `node 0`, and code reading that as "zero nodes" would conclude the +/// machine has no NUMA at all. +#[test] +fn the_highest_node_is_a_number_not_a_count() { + if let Some(highest) = highest_numa_node() { + // Every machine has at least one node, so the highest number is + // reachable as a node. This holds on a 1-node host (0) and on a + // 64-node one (63). + assert!( + highest.get() < u32::MAX, + "a real machine's highest node cannot be the no-preference sentinel" + ); + } +} + +/// The volume query either answers or reports why not; it never invents a node. +/// +/// On a local NTFS volume this workspace's host answers. A host whose storage +/// stack declines is equally valid and must produce an error rather than a +/// fabricated zero -- which is the failure this asserts against, because zero +/// is a real node number and so a plausible-looking fabrication. +#[test] +fn the_volume_query_answers_or_errors_but_never_fabricates() { + let (path, file) = scratch("answers-or-errors"); + match volume_numa_node(file.as_raw_handle()) { + Ok(node) => assert!( + node.get() < u32::MAX, + "a reported node cannot be the no-preference sentinel" + ), + Err(error) => assert_ne!( + error.kind(), + std::io::ErrorKind::Other, + "a failure should carry the OS error, not a placeholder" + ), + } + drop(file); + let _ = std::fs::remove_file(path); +} + +/// Asking the same handle twice agrees. +#[test] +fn the_volume_query_is_stable_for_one_handle() { + let (path, file) = scratch("stable"); + let first = volume_numa_node(file.as_raw_handle()).ok(); + let second = volume_numa_node(file.as_raw_handle()).ok(); + assert_eq!(first, second); + drop(file); + let _ = std::fs::remove_file(path); +} + +/// An invalid handle is an error, not a node. +/// +/// The cheapest reachable failure path, and worth having because the success +/// path cannot be forced to fail on a host whose volume does answer. +#[test] +fn an_invalid_handle_errors() { + let error = volume_numa_node(std::ptr::null_mut()) + .expect_err("the null handle is not a file or directory handle"); + assert!( + error.raw_os_error().is_some(), + "the failure should carry the OS error code, got {error:?}" + ); +} diff --git a/crates/windows-ioring-sys/BORROW-SURFACE.txt b/crates/windows-ioring-sys/BORROW-SURFACE.txt index 8ed1b15f7..e548286c3 100644 --- a/crates/windows-ioring-sys/BORROW-SURFACE.txt +++ b/crates/windows-ioring-sys/BORROW-SURFACE.txt @@ -7,8 +7,13 @@ # forgotten: CI regenerates it and fails when it disagrees with the source. crates/windows-ioring-sys/src/batch.rs :: get -> io::Result<&[u8]> crates/windows-ioring-sys/src/batch.rs :: get_mut -> io::Result<&mut [u8]> +crates/windows-ioring-sys/src/batch.rs :: new(ring: &'ring mut IoRing) -> Self crates/windows-ioring-sys/src/contract.rs :: violations -> &[Violation] +crates/windows-ioring-sys/src/error.rs :: as_ioring_error -> Option<&IoRingError> crates/windows-ioring-sys/src/error.rs :: name -> &'static str crates/windows-ioring-sys/src/error.rs :: name -> Option<&'static str> crates/windows-ioring-sys/src/event_delivery.rs :: batch -> Batch<'_> +crates/windows-ioring-sys/src/event_delivery.rs :: new(mut ring: IoRing, on_completion: F, env: Option<&mut CallbackEnviron<'_>>,) -> io::Result crates/windows-ioring-sys/src/event_delivery.rs :: scope -> RingScope<'_> +crates/windows-ioring-sys/src/pending.rs :: contract -> Option<&RingContract> +crates/windows-ioring-sys/src/ring.rs :: wait(ring: &mut RingWait<'_>, timeout_ms: u32) -> io::Result<()> diff --git a/crates/windows-ioring-sys/CHECKLIST.md b/crates/windows-ioring-sys/CHECKLIST.md index 507677f63..b19466f40 100644 --- a/crates/windows-ioring-sys/CHECKLIST.md +++ b/crates/windows-ioring-sys/CHECKLIST.md @@ -12,106 +12,359 @@ and M15-M18 M19 is archived [here](COMPLETED-CHECKLIST.md#m19). -**`M20` is pending; `M6+` is parked rather than pending** -- see the `M{n}+` convention: it is gated work -with no current obligation, not an unfinished milestone. +**`M20` through `M23` are pending; `M6+` is parked rather than pending** -- see the `M{n}+` convention: it +is gated work with no current obligation, not an unfinished milestone. ## M20 -- Repairs from the 2026-08-30 NUMA-sharding measurement Queued from [DESIGN-SESSION-2026-08-30-numa-sharded-io-execution-domains.md](../../design-sessions/DESIGN-SESSION-2026-08-30-numa-sharded-io-execution-domains.md), which measured a shipping ARM laptop and found the L3 heuristic's justification does not hold there. These -are documentation and policy repairs only; **no defect was found in `ring_copy`** -- `Policy::select` -already degrades to a whole-machine domain and reports it, which an initial reading of the session got -wrong and the code corrected. +were queued as documentation and policy repairs only, on the basis that **no defect was found in +`ring_copy`** -- `Policy::select` already degrades to a whole-machine domain and reports it, which an +initial reading of the session got wrong and the code corrected. + +**Corrected 2026-09-19: that basis no longer holds, and it changes the order.** `SH-4.12` in +[CHECKLIST-ship-topology-and-queues.md](../../CHECKLIST-ship-topology-and-queues.md) later found two +defects in that same function: it selects on `DomainKind::Cache { level: 3, .. }` rather than asking +`outermost_partitioning_cache()` -- the one definition of which cache level partitions a machine, shipped +in `windows-topology-sys` 0.2.0 -- so it can produce **overlapping** ring domains where two cache kinds +report at level 3, and degrades silently on a host whose outermost partition sits at another level. +`M20.1` and `M20.3` both land on that function and that rule, so both are **coupled to `SH-4.12`** and +must follow it. `M20.2` and `M20.4` are done. `M20.6` is gated the other way, on `M22.1`. That leaves +`SH-4.12` as the only thing standing between M20 and completion. The design questions the session opened are deliberately **not** queued here. It is still open, and its conclusions belong to it until it converges. -- [ ] **M20.1** -- Correct the L3 heuristic's justification in - [DESIGN-NOTES.md](DESIGN-NOTES.md). It currently says the last-level-cache domain "is meaningful on Intel - and ARM too, where the NUMA node often is not." **Measured counter-example:** a Snapdragon X2 Elite - (X2E80100, Qualcomm Oryon; 12 cores, no SMT) reports **zero** L3 cache domains -- `L3CacheSize = 0` from - WMI, and `GetLogicalProcessorInformationEx` yields L1 and L2 only, with L2 forming two domains of six - processors that agree with the two `Module` domains. The claim that L3 is meaningful on ARM is false on a - shipping part. Keep the finding that L3 beats the NUMA node; restate the rule as **the outermost cache - level that actually partitions the machine**, and say what happens when no such level is reported. Sweep - every restatement of the L3 rule per the repository's blast-radius convention, including the README and - `ring_copy`'s `policy.rs` doc comments, not only the one sentence quoted above. - -- [ ] **M20.2** -- Record the measurement itself as a decision in - [DESIGN-NOTES.md](DESIGN-NOTES.md), so the next reader inherits the datapoint rather than re-measuring: - an ARM Windows laptop with no L3 at all, and zero `Win32_NumaNode` instances, is the *common* consumer - shape now rather than an exotic one. This is the ARM sibling of the existing zero-NUMA-node VM - observation and belongs beside it. - -- [ ] **M20.3** -- Make `ring_copy`'s degraded-fallback path observable in a test. The whole-machine - fallback in `Policy::select` is the branch every zero-relation machine takes, and this session was the - first time anyone confirmed it runs. Assert both halves on a synthetic topology: that a policy whose - relation is absent returns one whole-machine domain with `degraded = true`, and that a policy whose - relation is present is **not** flagged degraded -- the second half matters because a test of the first - alone would pass against a function that always degrades. - -- [ ] **M20.4** -- Correct "What is not reachable" in [DESIGN-NOTES.md](DESIGN-NOTES.md). It says mapping a - file handle to its backing device's NUMA node "has no clean user-mode path" and "means walking volume to - disk to device instance and reading `DEVPKEY_Device_Numa_Node`". **That is wrong on mechanism.** - `FSCTL_QUERY_VOLUME_NUMA_INFO` is documented in the IFS docs, takes a handle to a **file or directory** - directly, and returns `FSCTL_QUERY_VOLUME_NUMA_INFO_OUTPUT { ULONG NumaNode }`. No walking required. - The **conclusion survives for a better reason**, and that is the point of the rewrite: the documented - meaning is the node the *volume* resides on, not where the file's extents live, so it cannot answer - "which ring should this file's I/O go to" even when it succeeds; and it is absent whenever the device - advertised no proximity domain. Record `GetNumaNodeNumberFromHandle` as the other path -- a wrapper over - `NtQueryInformationFile` with `FileNumaNodeInformation` (class 53) -- and that PHNT and the WDK mark that - class **reserved for system use**, so this crate must not build on it. State plainly that no published - measurement of either call succeeding on an ordinary NTFS data file could be found, and cite - [file-handle-numa-spike.rs](design-sessions/spikes/file-handle-numa-spike.rs) as the unrun instrument. - **Blocked on hardware, not on a decision:** settling it needs a multi-node machine with storage whose - PDO advertises a proximity domain. Write the correction now (the documentation defect is independent of - the measurement) and leave the empirical question open. +- [x] **M20.1** -- Restate the cache heuristic as "the outermost cache level that actually partitions + the machine", sweep every restatement, and replace the consumer that bound to the level number. + Done together with `SH-4.12`, which is the code half of the same change. + -> [completed 2026-09-22](COMPLETED-CHECKLIST.md#m201) + +- [x] **M20.2** -- Record the 2026-08-30 ARM measurement as a decision, beside the zero-NUMA-node + observation it is the sibling of. + -> [completed 2026-09-19](COMPLETED-CHECKLIST.md#m202) + +- [x] **M20.3** -- Make `ring_copy`'s degraded-fallback path observable in a test, asserting both + that an absent relation degrades and that a present one does not. Done without waiting on + `SH-4.12`: the fallback tail is shared by every policy, so exercising it through `ByNode` and + `ByPackage` pins nothing that item rewrites. + -> [completed 2026-09-22](COMPLETED-CHECKLIST.md#m203) + +- [x] **M20.4** -- Correct "What is not reachable" in [DESIGN-NOTES.md](DESIGN-NOTES.md): the + file-handle-to-storage-node mapping is reachable on mechanism, and the conclusion it supported now rests + on volume granularity, absence, and spanned volumes instead. + -> [completed 2026-09-19](COMPLETED-CHECKLIST.md#m204) - [x] **M20.5** -- Dissolved by [D-47](DESIGN-NOTES.md#d-47-detail) rather than decided: the `flush_barrier` assertion was measuring a claim the platform does not honour, so it was never a flaky test. -> [completed 2026-09-07](COMPLETED-CHECKLIST.md#m205) -- [ ] **M20.6** -- Re-evaluate `CommitStrategy::AlternatingRings` and the epoch-log benchmark's conclusion - against [D-47](DESIGN-NOTES.md#d-47-detail). The strategy comparison in - [strategy.rs](examples/epoch_log/strategy.rs) was designed around D-24's claim that a covering flush holds - back operations queued behind it: the harness deliberately keeps appending while a commit is outstanding so - that the stall would be visible in the numbers. D-47 established there is no such stall, so **the rationale - the benchmark rests on is withdrawn even though the measurements themselves stand**. Two things to settle, - and they are independent: whether alternating rings still earns its cost now that its stated benefit - (keeping appends off a stalled ring) does not exist -- the remaining benefit is that epoch *N+1*'s appends - are provably outside epoch *N*, which is a correctness property rather than a throughput one -- and whether - the published numbers should be re-read, re-run, or annotated. **Not a documentation-only fix:** if the - answer is that the strategy no longer earns its place, that is an API change to a published example. - The corrected prose in [strategy.rs](examples/epoch_log/strategy.rs) and - [DESIGN-NOTES.md](DESIGN-NOTES.md) both point here. - *(Numbered M20.6 rather than M20.5 because M20.5 was in flight on a separate branch when this was - written. That branch was closed unmerged; M20.5 arrives here instead, dissolved -- see above.)* - - -## M6+ -- Model B: explicit-thread delivery and affinity - -Parked, not pending. Deferred by the engineer's explicit direction during the 2026-08-22 design session, -with the plan scoped now so the shape is not lost. This is **not** a fallback for a missing capability -(D-3) -- it is the high-performance architecture, and M4's thread-pool path is the convenient one. - -- [ ] **M6+.1** -- `DeliveryMode::{ThreadpoolWait, PinnedThread}` as an explicit consumer choice, never an - automatic degradation. - -- [ ] **M6+.2** -- Resolve the contention between a thread parked in `SubmitIoRing(ring, n, INFINITE, ..)` - and callers wanting to build SQEs. This is the hard part and the reason this is its own milestone: it - directly contradicts M3.1's `&mut`-enforced serialization, and needs either a submit-ownership handoff - or an internal lock. Neither is obviously right. - -- [ ] **M6+.3** -- Shutdown: waking a thread parked on `INFINITE`. `IORING_OP_NOP` is supported and is the - wake mechanism. - -- [ ] **M6+.4** -- Affinity: binding a ring's thread with `SetThreadGroupAffinity`, and documenting the - execution-domain pattern (one pinned thread, its ring, its node-local registered pool, its shard). - -- [ ] **M6+.5** -- A test seam forcing the pinned-thread path even where the completion event is available, - so it stays testable on every machine rather than only on hardware that lacks the feature. - -- [ ] **M6+.6** -- Decide `IoBuf`: extract to a shared crate, re-export from - `windows-overlapped-io-sys`, or leave duplicated (D-1). The merge-or-delete decision that duplicate-then-decide - defers to the point where the new path is proven -- which is here, not earlier. +- [x] **M20.6** -- Re-evaluate `CommitStrategy::AlternatingRings` and the benchmark's conclusion + against [D-47](DESIGN-NOTES.md#d-47-detail). **The harness cannot exhibit a blast-radius + difference, which is a fact about the harness and not a finding against the strategy** -- each + lane's own arena is the limiter there. The strategy stays, with the conditions under which it + would pay written down; `M25.5` re-runs the comparison where operations genuinely pend. The + sample's output and prose are corrected so they stop claiming to measure a commit. + -> [completed 2026-09-23](COMPLETED-CHECKLIST.md#m206) + + +## M21 -- Epoch-log review: correctness repairs + +Queued from +[DESIGN-SESSION-2026-09-19-epoch-log-review.md](design-sessions/DESIGN-SESSION-2026-09-19-epoch-log-review.md) +(findings `C-1` through `C-5`). Independent of each other; listed in ascending cost. Nothing in this +milestone was observed failing at the sample's current constants -- these are a withdrawn justification, +two hang shapes, a mis-keyed trigger, and a specification gap. (`M21.3` predicted that its trigger was +merely unreachable *today*; measuring it while implementing showed it is unreachable at any constants, so +what it corrected was the coupling rather than a latent bug. The archived entry has the numbers.) + +- [x] **M21.1** -- Correct the last site that still asserts [D-24](DESIGN-NOTES.md#d-24)'s withdrawn + half: the epoch-order assertion in the epoch-log committer, whose justification cited the hold-back + claim [D-47](DESIGN-NOTES.md#d-47) removed. + -> [completed 2026-09-20](COMPLETED-CHECKLIST.md#m211) + +- [x] **M21.2** -- Publish a bounded pop and the wait it is generic over, then remove the two unbounded + spins. `IoRing::pop_within` / `pop_within_with`, over a `CompletionWait` the caller supplies, because + [D-21](DESIGN-NOTES.md#d-21) means the crate cannot choose the wait for them. + -> [completed 2026-09-21](COMPLETED-CHECKLIST.md#m212) + +- [x] **M21.3** -- Key the epoch commit off a completed append rather than off the counter, so the + trigger cannot fire on a pass that appended nothing. The predicted latent bug turned out to be + unreachable at any constants -- measured, not re-reasoned -- so this is a coupling change rather than + a fix. + -> [completed 2026-09-21](COMPLETED-CHECKLIST.md#m213) + +- [x] **M21.4** -- State what a *failed* commit does to `durable_through`, and bind it with tests in both + directions. Required making the sample a test target at all (`test = true`), and gating the + failure-path tests on `fault-injection`, since a healthy flush cannot be made to fail. + -> [completed 2026-09-21](COMPLETED-CHECKLIST.md#m214) + +- [x] **M21.5** -- Give the harness's wait loops a bound, and collapse the hand-written waits onto the + bounded pop. The item named two loops; a census found four, plus two flaky single-`try_pop` sites. + -> [completed 2026-09-21](COMPLETED-CHECKLIST.md#m215) + +- [x] **M21.6** -- Fix the four defects an independent review of the `M21.2` surface found: the timeout + mapping, its victim in `run_down`, the `INFINITE` collision, and the test hole that hid all of them. + -> [completed 2026-09-21](COMPLETED-CHECKLIST.md#m216) + +## M21+ -- Queued by the 2026-09-21 API review + +Queued from the review of the `M21.2` surface, recorded in +[DESIGN-SESSION-2026-09-21-m21-remediation-findings.md](design-sessions/DESIGN-SESSION-2026-09-21-m21-remediation-findings.md). +Its other four findings were fixed in `M21.6`. + +- [x] **M21+.1** -- Teach [check-borrow-surface.ps1](../../tools/check-borrow-surface.ps1) the two shapes + it was blind to: methods of a `pub trait`, and borrows in parameter position. Four entries appeared, one + of them predating the widening; the probes also found a latent bug in the checker itself. + -> [completed 2026-09-21](COMPLETED-CHECKLIST.md#m21plus1) + +## M22 -- Epoch-log review: submission and arena + +Queued from the same session (findings `E-1` through `E-3`). `M22.1` is sequenced first because `M20.6` +re-reads numbers that its change moves. + +- [x] **M22.1** -- Batch an epoch's appends into one submission in both append paths, and measure + whether the per-record submission cost was flattening the strategy comparison. It was not: + throughput did not move out of the noise, though commit p50 did. + -> [completed 2026-09-22](COMPLETED-CHECKLIST.md#m221) + +- [x] **M22.2** -- Collapse the two free-slot implementations to one, derived from the arena's own + outstanding counts rather than tracked beside them. The item called both correct; one was not -- + the tracked free list leaked a slot on every refused append. + -> [completed 2026-09-22](COMPLETED-CHECKLIST.md#m222) + +- [x] **M22.3** -- Give the registered arena a stated placement: the epoch-log arena is placed on the + NUMA node its own log file's volume reports, and the allocator moved into the library as + `NumaBuffer` rather than being copied a second time. The sample says plainly that the placement + cannot pay at this workload. + -> [completed 2026-09-22](COMPLETED-CHECKLIST.md#m223) + + +## M22+ -- Queued by what the M21 work left behind + +- [x] **M22+.1** -- Make [bounded_pop.rs](tests/bounded_pop.rs) independent of how fast a device is, by + reading from an overlapped pipe nobody has written to. Filed and completed the same hour; the deferral + was a scheduling preference rather than a blocker. + -> [completed 2026-09-21](COMPLETED-CHECKLIST.md#m22plus1) + + + +## M24 -- Make the unit suite hermetic + +The defect and its classification are [D-49](DESIGN-NOTES.md#d-49); the remedies and their costs are +[DESIGN-SESSION-2026-09-21-hermetic-unit-tests.md](design-sessions/DESIGN-SESSION-2026-09-21-hermetic-unit-tests.md). +It began at **63 of 131 lib tests opening a real kernel ring**, so `cargo test --lib` did not mean +what its name implies, and the repository's own Quality rule already classifies an operating-system +API as an external boundary. + +**Where it ended: 41 of 151 open a ring, 110 do not.** The remainder is not movable without +`M26.2`'s FFI seam, and [D-53](DESIGN-NOTES.md#d-53) records the rung that keeps it from climbing +back -- an inventory of *which* tests open a ring, since a zero-check would fail on day one and +could only be satisfied by deleting coverage. + +**Sequencing is the open question, not whether. Corrected 2026-09-22: the cost of waiting is close +to zero, which is the opposite of what this paragraph first said.** It claimed that waiting +compounds, "because every milestone that adds tests adds to the pile to be migrated, and `M22` is a +testing-heavy milestone". The mechanism is real but the instance was not checked, and it is false: +**all three `M22` items touch only `examples/epoch_log/`**, and none adds a lib test. + +**Unconditional as of 2026-09-22.** `M24.1` concluded and `M24.4` is withdrawn, so nothing in this +milestone waits on an evaluation any more. The hermetic goal is reached by relocation and by the +accounting extraction alone; the technique `M24.1` went looking for turned out to be a different +and larger thing, and is `M26`. + +Checked across the whole queue rather than for `M22` alone, since the first claim was wrong for +want of exactly that: **no pending item outside this milestone modifies `src/**/tests.rs`.** `M22` +is example-only; `M23.1` is the *sample's* `contract.rs`, not the crate's; `M20.1` and `M20.6` are +documentation and the `ring_copy` sample; `M23.2` is a decision that may imply API later. The 63 +therefore do not grow while this waits. + +So sequencing turns on other things, and they point the other way: + +- **`M24.2` is an internals refactor of a published crate**, and the branch carrying this work is + already 19 commits with one `feat` and three `fix` commits on it. Stacking a field-layout change + on top makes one review cover both a new public API and that refactor. +- **`M24.1` is an evaluation whose answer could invalidate `M24.4`**, so beginning the build before + it concludes risks building something the evaluation rejects. +- **`M22.1` unblocks `M20.6`**, an open question since 2026-09-07 about whether a strategy still + earns its place in a published sample -- which is a decision waiting on a measurement `M22.1` + produces. + +**`M24.1` concluded (2026-09-22) and nothing here waits on it.** `M24.4` is withdrawn; `M24.2`, +`M24.3` and `M24.7` are the path to a hermetic suite and are independent of each other. + +- [x] **M24.1** -- Settle whether a co-tested fake escapes the mock objection. **Answered: the fake + was the wrong instrument.** A shared suite is strong over what we specify and blind to the + platform's incidental behaviour, and an assertion about the latter is a frozen observation rather + than a contract. Superseded by the resolver in `M26`. + -> [completed 2026-09-22](COMPLETED-CHECKLIST.md#m241) + +- [x] **M24.2** -- Extract the handle-free accounting into its own type, composed by `IoRing`. The + item's field split was verified exactly: five fields carry no kernel state, five do. + `Accounting` now owns them with 19 hermetic tests, and `IoRing` delegates nine methods. + -> [completed 2026-09-22](COMPLETED-CHECKLIST.md#m242) + +- [x] **M24.7** -- Convert the lib tests that construct a ring only to exercise bookkeeping. + **61 -> 52**, by narrowing `Token::new` to take the ring's ledger rather than the ring. The + remaining 52 are not convertible and the reason is structural, not effort -- see the archive. + -> [completed 2026-09-22](COMPLETED-CHECKLIST.md#m247) + +- [x] **M24.4** -- **Withdrawn by `M24.1` (2026-09-22).** A shared conformance suite over a + hand-written fake is superseded by the response-space resolver in `M26`, which serves the same + purpose without encoding a belief about the platform at all. Nothing is deferred by this: `M24`'s + goal is a hermetic lib suite, and `M24.2` plus `M24.3` achieve that without it. + +- [x] **M24.3** -- Relocate the lib tests that open a ring but use only public API into `tests/`. + **52 -> 41.** Eleven moved; the "25" the item predicted was never achievable, and the reason is + the same structural one `M24.7` found. + -> [completed 2026-09-22](COMPLETED-CHECKLIST.md#m243) + +- [x] **M24.5** -- Put the rule on a rung. An **inventory** of which lib tests open a ring + ([D-53](DESIGN-NOTES.md#d-53)), not the zero-check the item assumed -- that rule is false and + could only be satisfied by deleting coverage. The guard's own bidirectional check found a defect + in the guard. + -> [completed 2026-09-22](COMPLETED-CHECKLIST.md#m245) + +- [x] **M24.6** -- Sweep what this milestone makes false. Two of the three sites the item named + were false alarms; the third was false for a different and larger reason than the item gave, and + the sweep found two more it did not name. + -> [completed 2026-09-23](COMPLETED-CHECKLIST.md#m246) + +## M23 -- The ring as a durability domain, and storage affinity + +Queued from the 2026-09-19 epoch-log review (findings `S-1` and `S-3`). `S-2` is an addendum to `M20.6` +rather than an item here. `M23.1` and `M23.2` are done; `M23.3` is the remaining question, and it is +about this crate's own surface rather than about storage at all. + +- [x] **M23.1** -- State in the epoch-log contract that the barrier is ring-wide while the flush names a file, so one ring per log is a precondition of the cost model. -> [completed 2026-09-23](COMPLETED-CHECKLIST.md#m231) + +- [x] **M23.2** -- Decide how a caller arrives at a NUMA node: `win-numa-sys` offers declaring and discovering, and refuses the shortcut that does both at once. -> [completed 2026-09-23](COMPLETED-CHECKLIST.md#m232) + +- [x] **M23.3** -- Decide what this crate offers for holding a token between push and completion: the ring owns the inventory, `IoRing` becomes generic, and the break is accepted. -> [completed 2026-09-23](COMPLETED-CHECKLIST.md#m233) + +- [x] **M23.4** -- Drop guards that panicked during unwind aborted the process instead of reporting; they now stay silent while `std::thread::panicking()`. -> [completed 2026-09-23](COMPLETED-CHECKLIST.md#m234) + +- [x] **M23.5** -- Both asserts in `IoRing::drop` are now reached by tests; the raw-HRESULT seam the item priced turned out not to be needed, because the kernel refuses a null ring handle cleanly. -> [completed 2026-09-23](COMPLETED-CHECKLIST.md#m235) + + +## M26+ -- The wakeup window review opened + +- [ ] **M26.12** -- **Find why a signal raised just after `wait.arm` can be lost, and fix it.** + Raised as a narrower finding by Copilot review on PR #108 -- that `EventDelivery::new` signals + only when it attached the event itself -- and the investigation found something wider. + + **What is measured.** The `#[ignore]`d reproducer + `a_backlog_is_delivered_even_when_the_caller_attached_the_event_first` in + [event_delivery.rs](tests/event_delivery.rs) fails 6 of 6. Signalling unconditionally, which is + what the review suggested, does not change that. A 50 ms sleep between `wait.arm` and the signal + makes it pass 3 of 3, and so does `--features trace`, which is the same perturbation by another + route. Full figures in [UNRESOLVED-TEST-FAILURES.md](UNRESOLVED-TEST-FAILURES.md). + + **Why this is not a small follow-up.** [D-68](DESIGN-NOTES.md#d-68) fixed `M26.9`'s stall by + ordering the arm before the signal, measured at 0 failures in 3600 runs. This says that ordering + narrows the window rather than closing it, so the decision's reasoning needs revisiting once the + mechanism is known -- not before, because the mechanism is currently a guess. + + **Do not apply the sleep.** It is a diagnostic that identified a window, not a fix, and shipping + it would convert a reproducible defect into a rare one. +## M27 -- What this crate owes the topology planner + +**Re-planned 2026-09-23, the same day it was written.** M27 was originally "Adaptivity: the benefit +without the architectural commitment", and asked whether *this crate* should derive a partition for a +consumer who expresses no preference. That was the wrong owner, and the checklist rules require +saying so rather than quietly rewriting it. The adaptivity the +[adoption thesis](../../DESIGN-NOTES.md#the-adoption-thesis) asks for is delivered by +[topology-planner](../topology-planner/COMPONENT.md), which takes a dataflow description of the +application and returns one or more suggested realizations +([EP-D-6](../topology-planner/DESIGN-NOTES.md#ep-d-6)). Had the original M27.1 been answered here it +would have grown a second, weaker policy surface beside the one that component exists to provide -- +the `outermost_partitioning_cache` defect again, where a policy answer lands in a crate whose job is +something else. + +> **-> CROSS-COMPONENT PREREQUISITE:** `M27.1` and `M27.2` are gated on component +> `crates/topology-planner` -> `M1+` -> `EP-1+.5` and `EP-1+.6`, which decide the plan vocabulary +> this crate would be realized from. See [CHECKLIST.md](../topology-planner/CHECKLIST.md). + +**What survives here is the realization end, not the policy end.** The planner emits a plan; the +outward adapter realizes it as buffers, rings and threads +([EP-D-5](../topology-planner/DESIGN-NOTES.md#ep-d-5)). That adapter is a separate crate, but it can +only build what this crate exposes, and nothing has ever checked that what it exposes is sufficient. +[D-8](DESIGN-NOTES.md#d-8) is untouched by all of this: policy stays out of this crate, and being +*constructible from* a policy decision made elsewhere is the opposite of taking one. + +- [ ] **M27.1** -- **Census what a realizer would need from this crate, against the plan vocabulary, + and name what is missing.** A plan states which processor a domain pins to, which memory node its + pool allocates from, how many queues of which types, and where each channel's buffer lives. Walk + each of those to the public API that would realize it and record the gaps. `NumaBuffer` + ([D-51](DESIGN-NOTES.md#d-51)) is one half of the pool answer and arrived this month; the ring's + own construction takes no placement input at all. **The output is a gap list, not an API** -- + proposing surface before the plan vocabulary is settled would be binding to a draft. + +- [ ] **M27.2** -- **Gated on `M27.1` and on the planner's `EP-1+.6`.** Close the gaps the census + names, as ordinary capability on this crate with no policy attached. Each gap is an input a caller + supplies, never a choice this crate makes. Verify the way the thesis demands rather than the + convenient way: construct from a plan built against a *synthetic* machine, since the planner is + mockable by construction and this crate should be realizable without the hardware the plan + describes. + +- [ ] **M27.3** -- Give a consumer the means to answer placement questions on their own hardware. + **Not gated on the planner** -- it is the client-side half of the thesis, and it is what lets a + developer disagree with any plan they are handed. `cache_domains.rs` now prints every cache level + beside the heuristic's pick; the equivalent for placement is a sample that reports what a chosen + arrangement costs and what the alternatives would have cost, on the machine in hand. + [ring_copy](examples/ring_copy) is the natural host, being already policy-selectable. **Do not ship + a verdict** -- report the observation and let the consumer conclude, per OPTION INTEGRITY. + +## M28 -- The ring owns the pending inventory + +Queued by [D-55](DESIGN-NOTES.md#d-55), taken 2026-09-23 after the `M23.3` exploration. The break +is accepted deliberately: `IoRing` becomes generic so the inventory cannot drift from the ring, +because a consumer never holds a token to lose. The exploration and everything it falsified is in +[DESIGN-SESSION-2026-09-23-pending-inventory.md](design-sessions/DESIGN-SESSION-2026-09-23-pending-inventory.md); +`src/pending.rs` is the working spike and is the shape the internal map starts from. + +**Sequenced so each step compiles.** The published crate is at 0.3.1, so this is a major bump and +every consumer names the type -- which means the migration order matters more than usual. + +- [ ] **M28.1** -- **Decide what a caller receives, before writing any of it.** If the ring owns + the token then `Batch::write` can no longer hand one back, and the shape of what replaces it is + the whole design: an identity the caller matches later, or a claim that returns `(T, X)` + directly from the ring. The second makes drift impossible and is the point of the break; the + first is a smaller change that may not be worth breaking for. Settle it with the + `Token::claim_if` safety argument in hand, since that is what currently makes a mismatched + completion unclaimable. + +- [ ] **M28.2** -- **Bound `RingContract` before anything depends on it more heavily.** + `operations: HashMap` is never pruned -- `observe_claim` marks an entry + `Completed` and keeps it -- so the oracle retains one entry per operation for the process's + life. Undocumented, and not visible in the sample because it appends 24 records. A long-running + consumer following the crate's own recommendation leaks. This blocks any design that checks by + default, which is why it is here rather than filed separately: `M23.3` reached for always-on + checking and this is what ruled it out. Decide whether completed entries are dropped, whether + `check_quiescent` needs them, and document the answer either way. + +- [ ] **M28.3** -- **Gated on `M28.1`.** Make `IoRing` generic and move the inventory inside. + Carry the sidecar: the census found two thirds of consumers keep per-operation data beside the + token, so an inventory that holds only tokens serves a minority. Mixed-shape consumers use a + closed `enum` -- `tests/generated_sequences.rs` is the worked example and needs no change to + keep working. + +- [ ] **M28.4** -- **Gated on `M28.3`.** Migrate the ~12 consumers, and delete `Pending` or + demote it to the internal map. **Convert all of them or none**: converting a few relocates the + duplication rather than removing it, which is the lesson `win-numa-sys` recorded the same day + when it moved one `VirtualAllocExNuma` and left the other. + +- [ ] **M28.5** -- **Answer the tokenless push.** `flush_raw` returns a bare `usize` and + `epoch_log`'s commit path depends on it, because a flush has no buffer and a *borrowed* + `RawHandle` gives its token nothing to guard. An inventory the ring owns has to say what it + does with operations that have no token -- `RingContract` already models them separately with + `observe_tokenless_push`. Note this may dissolve rather than need solving: if the sample owned + a `SharedFile` instead of passing a `RawHandle` it could use the safe `flush` and get a token, + which `M25.3` reopens anyway by changing how the log is opened. + +- [ ] **M28.6** -- **Sweep what the break makes false**, including the README's ring examples, the + `D-4` detail section, and every rustdoc that tells a caller to match a completion against a + held token -- `Completion::user_data` and `IoRing::push_raw` both do, and they are the evidence + D-55 rests on, so they are the first things the change invalidates. diff --git a/crates/windows-ioring-sys/COMPLETED-CHECKLIST.md b/crates/windows-ioring-sys/COMPLETED-CHECKLIST.md index 84dbdb328..077e85696 100644 --- a/crates/windows-ioring-sys/COMPLETED-CHECKLIST.md +++ b/crates/windows-ioring-sys/COMPLETED-CHECKLIST.md @@ -1565,3 +1565,2259 @@ the API whose breaking change 0.2.0 is being cut for, and it is reachable with n and D-45 is added to its table of shipped defects of this shape. **Swept the count restatements too:** that file said "three defects" in four places and is now four, which is the restatement drift the repository's own conventions warn about. + +## Moved 2026-09-19 22:40:58 -07:00 -- M20.4: the file-handle NUMA mechanism correction + +### M20.4 -- Correct "What is not reachable" in [DESIGN-NOTES.md](DESIGN-NOTES.md): the file-handle-to-storage-node mapping is reachable on mechanism, and the conclusion it supported now rests on volume granularity, absence, and spanned volumes instead. *(completed 2026-09-19 22:40:58 -07:00)* + +The authoritative text is the rewritten "What is not reachable" section in +[DESIGN-NOTES.md](DESIGN-NOTES.md); the research behind it is `F-1` in +[DESIGN-SESSION-2026-08-30-numa-sharded-io-execution-domains.md](../../design-sessions/DESIGN-SESSION-2026-08-30-numa-sharded-io-execution-domains.md). + +Two things the item's own text had wrong, corrected while doing it rather than copied forward: + +- It said to cite [file-handle-numa-spike.rs](design-sessions/spikes/file-handle-numa-spike.rs) as the + **unrun** instrument, and to state that no measurement of either call succeeding on an ordinary NTFS + data file could be found. `F-1a` of the same session had already smoke-run it: both calls succeed on an + ordinary NTFS data file and on a directory handle, and agree. The item was written from `F-1` without + `F-1a`. What remains unmeasured is narrower -- whether either call ever names a node that distinguishes + one device from another, which needs a multi-node host with storage whose PDO advertises a proximity + domain. + +- The spikes [README.md](design-sessions/spikes/README.md) carried the same staleness, each instance + contradicted by its own body a few paragraphs later. Swept: 3 phrasings in 1 file, plus the sentence + promising that a multi-node run "would correct a claim DESIGN-NOTES.md currently makes", which this item + has now made false -- it settles an open question instead. + +The heading stays "What is not reachable". What is not reachable is the *answer* a ring consumer wants, +which is still true; renaming it would dangle the pointers in the 2026-09-19 review session. + +## Moved 2026-09-19 22:49:09 -07:00 -- M20.2: the ARM no-L3 measurement recorded as D-48 + +### M20.2 -- Record the 2026-08-30 ARM measurement as a decision, beside the zero-NUMA-node observation it is the sibling of. *(completed 2026-09-19 22:49:09 -07:00)* + +Landed as [D-48](DESIGN-NOTES.md#d-48) in the decision index, plus a sibling paragraph in +"Why the NUMA node is the wrong key" where the existing zero-node observation lives, which is where the +item asked for it. + +Two choices worth recording, because both were places this could have gone wrong: + +- **The measurement is cited, not re-transcribed.** The capture is Measurement M-1 in + [DESIGN-SESSION-2026-08-30-numa-sharded-io-execution-domains.md](../../design-sessions/DESIGN-SESSION-2026-08-30-numa-sharded-io-execution-domains.md), + and D-48 links it rather than copying the probe output into a third place. Pasting the block would have + created exactly the restatement the repository conventions warn about. + +- **The false clause it falsifies is marked, but not rewritten.** D-48 sits two paragraphs from the + sentence saying the last-level-cache domain "is meaningful on Intel and ARM too", which this + measurement shows is false on a shipping part -- so leaving it unmarked would have made the document + contradict itself. A one-line adjacent marker says so and points at `M20.1`. The restatement of the + rule and the sweep across the README, `lib.rs` and `policy.rs` remain `M20.1`, which is coupled to + `SH-4.12` and must follow it. + +## Moved 2026-09-20 23:16:19 -04:00 -- M21.1: the last site asserting D-24's withdrawn half + +### M21.1 -- Correct the last site that still asserts [D-24](DESIGN-NOTES.md#d-24)'s withdrawn half: the epoch-order assertion in the epoch-log committer, whose justification cited the hold-back claim [D-47](DESIGN-NOTES.md#d-47) removed. *(completed 2026-09-20 23:16:19 -04:00)* + +The assertion is unchanged, because it was always sound -- just for the other reason. Every commit +carries `FlushCoverage::CoversPrecedingOperations`, so commit *N* is outstanding when commit *N+1* is +reached, and D-47's *surviving* half -- no operation queued before a drained flush was ever observed +completing after it -- is what orders them. The comment now says that, and says explicitly what it does +not rest on, so a later reader does not "restore" the withdrawn reasoning. + +Sweep re-run before committing, as the item required. **21 matches across 14 files**, disposed as: + +- **1 violation**, fixed: the justification in [commit.rs](examples/epoch_log/commit.rs). +- **1 historical site**, marked rather than rewritten: + [DESIGN-SESSION-2026-08-28-external-consumer-correspondence.md](design-sessions/DESIGN-SESSION-2026-08-28-external-consumer-correspondence.md) + records what the findings became on that date, and glossed D-24 as a stall that "holds operations + against unrelated files". A session is a faithful record of its moment, so the gloss stays and a + one-line note beside it says which half was later withdrawn. +- **2 false positives**, left alone: the spikes README ("checked in deliberately rather than held back", + about an instrument) and [fault_injection.rs](tests/fault_injection.rs) ("holds it until the completion + is claimed", about a token). +- **17 already correct** -- the decision index, both sides of the crate docs, the README, the stress + tests, strategy.rs, and the item's own text. + +The item quoted the original sweep as "17 matches across 10 files". Re-running it found more of both, +and the file count needed care: `rg` groups the two design-session files under a single header, so the +first reading of its output undercounted by one. Counted with a command rather than by eye. + +Verified by running the example in a debug build, where the `debug_assert` is live: it completed, the +negative control still caught the corrupted record, and all four epochs reported durable in order. + +## Moved 2026-09-21 00:23:49 -04:00 -- M21.2: a bounded pop, generic over the wait + +### M21.2 -- Publish a bounded pop and the wait it is generic over, then remove the two unbounded spins. *(completed 2026-09-21 00:23:49 -04:00)* + +**Decided: the trait, plus a default impl and a convenience.** The crate cannot choose the wait for a +caller, and that is a contract rather than a shrug -- the completion event is auto-reset with exactly one +waiter per ring ([D-21](DESIGN-NOTES.md#d-21)), so a wait this crate picked could consume an edge the +caller's own loop was entitled to. Making it a parameter moves the obligation to the only party who can +discharge it. + +The surface added to [ring.rs](src/ring.rs): + +- `IoRing::pop_within(timeout)` -- the convenience, using `SubmitWait`. +- `IoRing::pop_within_with(wait, timeout)` -- the same loop over a caller-chosen wait; `?Sized`, so a + trait object works. +- `CompletionWait` -- one method, whose contract is deliberately weak: returning early or spuriously is + always permitted, because [D-19](DESIGN-NOTES.md#d-19) already makes a wake with nothing to pop normal. + An implementation cannot be subtly wrong about *when* to return, only wasteful. +- `RingWait` -- the ring narrowed to `block` and `outstanding`, so a waiter cannot pop the completion + its own caller is waiting for, nor submit work nobody asked for. Same narrowing, same reason, as + `RingScope` under [D-43](DESIGN-NOTES.md#d-43). +- `SubmitWait` -- blocks inside `SubmitIoRing` with nothing queued, touching no event, so it composes + with a caller who owns the completion event. + +**The borrow question, answered even though the check says the surface is unchanged.** `RingWait` is only +ever passed *in*, never returned, so [BORROW-SURFACE.txt](BORROW-SURFACE.txt) is untouched (verified: 7 +entries, unchanged). Asked anyway, since it is a lifetime-carrying wrapper: safe code reaching it can call +`block` and `outstanding` and nothing else; the ring it borrows is held exclusively for the call, so +nothing it could invalidate is live elsewhere; and the narrow type is the point rather than an accident, +because a bare mutable reference to the ring would have permitted exactly the two things a waiter must +not do. + +**Measured while building it, and now documented on `RingWait::block`:** `SubmitIoRing` answers +`E_INVALIDARG` (`0x80070057`) -- not a timeout -- when asked to wait for a completion the kernel has no +pending operation for. Found by a test that drove the loop with a reservation having no real SQE behind +it. The precondition holds structurally: `pop_within_with` checks `outstanding()` before consulting the +wait, so `block` is unreachable with nothing pending. + +**An early return that is not an optimisation.** With nothing outstanding and an empty queue no completion +is possible, so the loop answers immediately rather than sleeping out the bound. That turns "you forgot to +submit" from a timeout into an instant answer, and it is what the sabotage below pins. + +**A panic path found and closed before it shipped.** The first draft computed an instant plus the caller +timeout directly, which panics on overflow -- so `Duration::MAX`, a reasonable spelling of "no deadline", +would have taken down the process. Now `checked_add`, with an unrepresentable deadline treated as one that +never arrives. Two tests cover it: one where the early return answers first, one where an operation is +outstanding so the overflow branch is actually reached. + +**Sabotage-verified**, because a test that cannot go red proves nothing: + +- Removing the clamp that keeps the wait timeout above zero turns `the_wait_is_never_handed_a_zero_timeout` + red. +- Removing the nothing-can-arrive early return turns two tests red, and the run takes the full 30-second + bound instead of finishing instantly -- which is the behaviour the early return exists to remove. + +**Call sites converted:** the two unbounded spins in +[append.rs](examples/epoch_log/append.rs) and [fault_injection.rs](tests/fault_injection.rs), and the +test-only `pop_within` helper, which is now a thin panicking wrapper over the public API rather than a +fourth copy of the loop. The panic is the only test-specific part left: a test wants the name of what it +waited for, a consumer wants an `Option` it can act on. + +## Moved 2026-09-21 02:23:34 -04:00 -- M21.3: the epoch trigger, and a review claim the measurement disproved + +### M21.3 -- Key the epoch commit off a completed append rather than off the counter, so the trigger cannot fire on a pass that appended nothing. *(completed 2026-09-21 02:23:34 -04:00)* + +The match in [main.rs](examples/epoch_log/main.rs) now yields a `closed_an_epoch` value that every arm +must produce, rather than a `appended % EPOCH_SIZE == 0` test written after it. Closing an epoch is a +fact about an append that landed, and making each arm answer the question keeps that local -- a new arm +added later cannot fall through into a commit. + +**The item predicted this was a latent bug. It is not, and the measurement is what settled it.** +The retry path was instrumented to report when the old shape would have committed, and run at +`EPOCH_SIZE` of 6, 8, 12, 16 and 24. It fired **zero** times -- including at every value past `SLOTS`, +which both the item and finding `C-3` predicted would arm it. + +The reason is an invariant three blocks from the trigger: `appended % EPOCH_SIZE == 0` is true at exactly +two moments -- before the first append, and immediately after a commit -- and the arena is empty at both, +because the commit waits for a covering flush that retires every outstanding write. An append is +therefore never refused at an epoch boundary, at any constants. + +**So this is a coupling change, not a bug fix**, and the distinction is the useful part. The old trigger +was safe because of something nothing stated, three blocks away; the new one cannot fire because of where +it is written. The second survives a reader who changes the commit path. The first is what made the +question take a measurement to answer at all. + +Swept the claim rather than only fixing the code: finding `C-3` in +[DESIGN-SESSION-2026-09-19-epoch-log-review.md](design-sessions/DESIGN-SESSION-2026-09-19-epoch-log-review.md) +carried the same wrong prediction and now carries the correction beside it. The review lesson recorded +there is the narrow one: "unreachable today, armed tomorrow" is a claim about a program's reachable +states, and reading the code is not how to settle one. + +## Moved 2026-09-21 16:06:19 -04:00 -- M21.4: what a failed commit means, and the sample's first tests + +### M21.4 -- State what a failed commit does to durable_through, and bind it with tests in both directions. *(completed 2026-09-21 16:06:19 -04:00)* + +The doc on `Committer::claim` said "A failed commit advances nothing", which reads as a permanent verdict. +It is not. Every commit here is a **covering** flush, so commit *N+1* reaches epoch *N*'s writes -- queued +before it -- and observing *N+1* makes *N* durable after all. What makes a record durable is a flush that +covered it, not the identity of the flush named for its epoch. The doc now says that, and says what a +caller must not read into a failure: not "epoch *N* is lost", but "not yet". + +**The sample had no tests at all.** Examples are not test targets by default, so `cargo test` compiled +this one and ran nothing. Binding the claim meant adding `test = true` to the `[[example]]` entry first; +that is the change that makes any of the sample's policy testable, not just this item's part of it. + +**The failure path is unreachable by running the sample**, because a flush against a healthy temp file +does not fail. The crate's fault-injection seam is the only way in, so four of the six tests are gated on +`fault-injection` and the other two run by default. That follows the precedent and the reasoning already +written down in [fault_injection.rs](tests/fault_injection.rs), including that CI's +`cargo test --workspace --all-features` job is what stops gated tests from being tests that never run. + +**An assumption caught by asserting it.** The test first asserted that an injected `ERROR_ACCESS_DENIED` +would surface as `io::ErrorKind::PermissionDenied`. It surfaces as `Other`: the crate preserves the +HRESULT in an `IoRingError` rather than classifying it. Corrected to assert the Win32 code, which is the +assertion `tests/fault_injection.rs` already makes one layer down. + +**Sabotage-verified in both directions, and the two produce different failure sets** -- which is what +shows the tests discriminate rather than all keying on one fact: + +- `is_durable` returning `true` unconditionally: **4 of 6 fail**, caught by the assertions that an epoch + is *not* yet durable. +- `is_durable` requiring an exact match with the watermark: **2 of 6 fail**, caught by the covering and + monotonicity tests. + +**Leverage from earlier in this milestone:** `commit_and_pop` is three lines because `M21.2` published +`IoRing::pop_within`. Without it every test here would have carried its own bounded wait, which is the +duplication `M21.2` existed to remove. + +## Moved 2026-09-21 18:18:52 -04:00 -- M21.5: every unbounded wait in the crate, not the two that were named + +### M21.5 -- Give the harness's wait loops a bound, and collapse the hand-written waits onto the bounded pop. *(completed 2026-09-21 18:18:52 -04:00)* + +`await_flush` and `await_writes` in [strategy.rs](examples/epoch_log/strategy.rs) blocked in +`submit_and_wait` for `WAIT_MS`, ignored that it had returned without a completion, and went round +again forever. Both now carry a deadline and raise `TimedOut`, which is the policy +[event_loop.rs](examples/epoch_log/event_loop.rs) already documented: a measurement harness that hangs +reports nothing, which is strictly worse than one that fails. + +**The item named two loops. A census found four, and two more of a related shape.** Counted by command +over every `.rs` outside `target/` and the spikes: + +- [strategy.rs](examples/epoch_log/strategy.rs) `await_flush` and `await_writes` -- the two named. +- [failure_paths.rs](tests/failure_paths.rs) and [kernel_span.rs](tests/kernel_span.rs), each with a + helper **called `await_one`**, byte-identical to the other, neither named by the item. +- [batch/tests.rs](src/batch/tests.rs), where two registration waits were a single `try_pop` -- the + flake shape `pop_within` documents -- in a file whose *third* such wait already used the helper. + One predicate, three sites, half-converted, which is FAIL FAST rule 1 exactly. + +**Sabotage-verified, and it revealed the milestone compounding.** Making `classify` stop filing flush +results leaves `await_flush` looking for a completion that is never recorded -- an unbounded loop would +hang forever. It failed with `timed out after 30s waiting for a commit's flush`, and it did so in **two +seconds**, not thirty: `pop_within`'s nothing-can-arrive early return (`M21.2`) answers immediately once +the ring is quiesced. The bound is what makes the failure possible; the early return is what makes it +quick. + +A `remaining` helper and a single `timed_out` constructor keep the two waits from describing the same +condition two ways, and `WAIT` is derived from `WAIT_MS` rather than written twice. + +Factoring `Lane::classify` out of `Lane::drain` is what let the bounded waits file a completion they +blocked for without a second copy of the claim-then-check logic. + +## Moved 2026-09-21 21:15:12 -04:00 -- M21.6: the API review's four defects + +### M21.6 -- Fix the four defects an independent review of the M21.2 surface found: the timeout mapping, its victim in `run_down`, the INFINITE collision, and the test hole that hid all of them. *(completed 2026-09-21 21:15:12 -04:00)* + +Queued and closed in one item because the first two share a root cause and the fourth is the reason +neither was caught. The surface is unreleased and on a branch, so none of this is a breaking change. + +**The defect.** `SubmitIoRing` reports an expired wait as `ERROR_TIMEOUT` (`0x800705B4`), a *failure* +HRESULT. `RingWait::block` passed it through `check`, so `pop_within` returned `Err` on every real +timeout and never the `Ok(None)` it documents. All six converted call sites treat `Err` as fatal and +`Ok(None)` as the timeout signal, so the `ErrorKind::TimedOut` mapping they document was dead code on +the only path that produces it. + +**Its second victim, pre-existing.** `run_down` polls in 50 ms steps with the same `check`, so any +operation slower than 50 ms made rundown return `Err` with work still outstanding -- and `Drop` then +asserted and called `CloseIoRing` anyway, which is precisely the "the kernel may still be writing +through a token's buffer" hazard the function exists to prevent. One `wait_outcome` helper now +classifies a wait-only `SubmitIoRing` result for both callers, so they cannot disagree again. + +The rundown loop is deliberately left unbounded. Every SQE that queues produces exactly one completion +(M10.2), so it terminates; blocking until that holds is the safe failure, and closing the ring early is +not. + +**The `INFINITE` collision.** A finite bound above ~49.7 days saturated onto `u32::MAX`, which is Win32's +`INFINITE` -- and `CompletionWait` explicitly invites implementations built on `WaitForMultipleObjects`. +Clamped to `MAX_WAIT_MS` (one below), and the trait now states that `timeout_ms` is never zero and never +`INFINITE`, so it can be passed straight to a Win32 wait. + +**The contract gap that would have propagated it.** `CompletionWait` never said how to report an expired +wait, and every Win32 wait reports one as an error -- so any third-party implementation forwarding its +underlying result would have reproduced the defect exactly. The trait now says an expired bound is +`Ok(())`, and says why. + +**Why no test caught it, which is the finding worth keeping.** The M21.2 tests drive the loop with a wait +that never enters the kernel. Deterministic, and it leaves the Win32 interaction untested: replacing +`RingWait::block`'s entire body with an unconditional error left **the whole suite green**. + +Closing that needed an operation still pending when the bound expires, and three attempts failed before +one worked -- each measured, not assumed: + +| Attempt | Result | +|---|---| +| Buffered read, up to 256 MiB | completion already poppable, 3-5 us | +| Flush over 512 MiB of dirty cache | 3 us -- the lazy writer had already written it back | +| Unbuffered read, 256 MiB | 3 us -- **a handle without `FILE_FLAG_OVERLAPPED` is synchronous**, so the read completes inline during submit | +| Unbuffered **and** overlapped, 64 MiB+ | genuinely pending; the bound expires | + +That third row is the general finding: **the crate's existing tests all use synchronous handles**, so +ring operations complete inline during submit and asynchronous completion is never exercised. That is why +a counting waiter over the existing flush pattern was reached in 0 of 50 trials. + +[tests/bounded_pop.rs](tests/bounded_pop.rs) now covers it with five tests over a 128 MiB unbuffered, +overlapped read. Each asserts `outstanding() > 0` alongside the expected answer, so a machine fast enough +to finish the read early fails loudly instead of passing vacuously. Nothing asserts an upper bound on +elapsed time: Windows' default timer resolution is ~15.6 ms, so a 5 ms bound routinely takes 14-19 ms. + +**Sabotage-verified, twice.** Reverting the timeout mapping turns **all five** red. Reproducing the +review's original mutation -- `block` always failing -- now turns three red, where before it turned none. + +The review's fifth finding, that `check-borrow-surface.ps1` is blind to trait methods and to borrows in +parameter position, is queued as `M21+.1` rather than fixed here: it is a process gap, not a runtime +defect, and widening the check obliges a borrow-question answer for every entry it newly reports. + +## Moved 2026-09-21 21:23:24 -04:00 -- M21+.1: the borrow-surface check learns two shapes + +### M21+.1 -- Teach [check-borrow-surface.ps1](../../tools/check-borrow-surface.ps1) the two shapes it was blind to: methods of a \pub trait\, and borrows in parameter position. *(completed 2026-09-21 21:23:24 -04:00)* + +The check inspected only the text after the last `->` on lines matching `pub fn`. Trait items are +declared `fn`, not `pub fn`, so nothing inside any `pub trait` was ever examined; and a borrow-carrying +type in *parameter* position was invisible in any function. `CompletionWait::wait` is both at once. + +**Four entries appeared, and one of them predates the widening by months.** +`IoRingErrorExt::as_ioring_error` returns `Option<&IoRingError>` and had simply never been inventoried -- +the blind spot made concrete rather than a new risk. The other three are `CompletionWait::wait`, +`Batch::new` and `EventDelivery::new`, the last two carrying an explicit lifetime in a parameter. +The borrow question is answered for all four in +[DESIGN-NOTES.md](DESIGN-NOTES.md#borrow-surface-audit-m21plus1), before the inventory was regenerated, +as [DESIGN-INSTRUCTIONS.md](DESIGN-INSTRUCTIONS.md) requires. None is a hole. + +**A plain `&T` parameter is deliberately not reported.** Reporting every method that borrows something +would list the whole crate and mean nothing. Only an *explicit lifetime* in parameter position counts -- +the borrow-carrying wrapper whose lifetime the crate chose, not a reference the caller lent us. That is a +heuristic, and the archived audit says so: it would miss a `&dyn Trait` parameter carrying no named +lifetime. + +**Verified by five probes, restored afterwards** -- and the negative control is the one that matters, +because a check that fires on everything is as useless as one that fires on nothing: + +| Probe | Expected | Result | +|---|---|---| +| `pub fn` returning `&[u8]` (control, the old rule) | caught | caught | +| `pub trait` method returning `&[u8]` | **now caught** | caught | +| `pub fn` taking `&mut RingWait<->` | **now caught** | caught | +| `pub fn` taking a plain `&Completion` | **silent** | silent, exit 0 | +| `pub trait` with a default body plus a borrow-returning sibling | only the sibling | only the sibling | + +**The probes found a latent bug in the checker itself**, which is the argument for running them rather +than reasoning about the regex. A one-line body -- `pub fn f() -> &[u8] { &[] }` -- never satisfied the +"line ends with `{`" test, so the signature accumulator ran past the end of the file. The old script did +not crash on it only because it never indexed the lines again afterwards; it silently swallowed the +following lines instead. Signature termination is now "the accumulated text contains a `{`", and the +return type is truncated at that brace. + +Also swept while here: the script header and its failure message both said **three** shipped defects of +this shape and listed D-35, D-36, D-43. It is four, and has been since D-45. +[M19.3](COMPLETED-CHECKLIST.md) swept that count through `DESIGN-INSTRUCTIONS.md` and missed this file, +which is the restatement-drift pattern landing on the very tool built to stop a different one. + +## Moved 2026-09-21 22:08:52 -04:00 -- M22+.1: a pending operation that owes nothing to a device + +### M22+.1 -- Make [bounded_pop.rs](tests/bounded_pop.rs) independent of how fast a device is, by reading from an overlapped pipe nobody has written to. *(completed 2026-09-21 22:08:52 -04:00)* + +**Queued and completed within the hour, and the queueing was the error.** It was filed as `M22+.1` with +an entry in `UNRESOLVED-TEST-FAILURES.md` on the grounds that the fix did not belong in a push of +finished milestones. That is a scheduling preference, not a blocker, and the repository's PRIME +DIRECTIVE is explicit that only a genuine blocking factor justifies deferral. The mechanism was +understood when it was filed; the two open questions were each one probe away. + +**Both probes answered, and neither was safe to assume.** `IoRing` does accept a pipe handle for +`read_raw`; and `pop_within(20ms)` against an unwritten overlapped pipe returns `Ok(None)` with +`outstanding == 1`. `Win32_System_Pipes` was added to the dev-dependency feature set; `PIPE_ACCESS_INBOUND` +is not re-exported where the module name suggests, so it is a named local constant, as +`FILE_FLAG_NO_BUFFERING` already was in this file. + +**The substance of the change is the question the test asks.** A 128 MiB unbuffered read asks "will this +device take longer than 5 ms?" -- a question about someone else's hardware, which may answer differently +on two runs of the same machine. A pipe with no writer asks nothing: the read is pending because no byte +exists to satisfy it, and it completes exactly when the test writes one. + +Where a delay is still needed -- `run_down` polls in 50 ms steps, so forcing it to observe an expired poll +means releasing the read later than that -- it comes from a `thread::sleep`, whose guarantee runs the safe +way round: a sleep may overshoot, never undershoot. No assertion depends on an operation *finishing* +within any bound. + +**Verified:** 25 consecutive runs of the target and 3 full `--all-features` suite runs, all green; both +feature configurations green. **Both sabotages still bite exactly as before the rewrite** -- reverting the +timeout mapping turns all 5 red, making `RingWait::block` always fail turns 3 red -- which is the +assertion that matters, because a deterministic test that had lost its discriminating power would be a +worse outcome than the flake. + +Also 11x faster (0.20s against 2.26s), with no 128 MiB fixtures. + +The `UNRESOLVED-TEST-FAILURES.md` entry moved to +[RESOLVED-TEST-FAILURES.md](RESOLVED-TEST-FAILURES.md) in this commit, per the append-only rule. + +## Moved 2026-09-22 15:54:53 -04:00 -- M22.1: batching the appends, and the confound it was meant to test + +### M22.1 -- Batch an epoch's appends into one submission in both append paths, and measure whether the per-record submission cost was flattening the strategy comparison. *(completed 2026-09-22 15:54:53 -04:00)* + +`Appender::append_batch` and `Lane::append_batch` replace the single-record pushes. Both compose as +many records as there are free arena slots and submit once, so the arena rather than the caller's +list decides the batch size -- which is what keeps the two halves of an append honest, since a slot +is composed into only while the kernel is not reading it. With eight slots, an epoch of 64 records +goes from 64 submissions to 8. + +**The teaching defect is the smaller half.** A sample whose job is to teach `Batch` was paying one +`SubmitIoRing` per record, which is the one thing `Batch` exists to avoid. + +**The measurement was the point, and it returned a negative result.** Finding `E-1` raised the +possibility that the per-record cost was a term every strategy paid equally, and therefore a shared +constant capable of flattening the three-way comparison into "indistinguishable" without that being +true. Twenty runs -- ten each side, taken in one sitting by stashing the change so both sets came +from the same machine and build -- are kept in +[measurements/2026-09-22-append-batching/](measurements/2026-09-22-append-batching/). + +Throughput did not move in a way that can be distinguished from noise: median records/sec shifted by +1-5% while a single strategy's run-to-run range spans 1.17x to 1.57x. The cross-strategy spread did +not shrink, and stayed at or below one strategy's own range -- which is the sample's own stated test. +**So `E-1`'s hypothesis is not supported, and the existing conclusion survives a confound raised +specifically against it.** + +What did move is commit p50, and it is the one figure here that separates: for `covering-flush` the +ten before-values and ten after-values barely overlap. That is the expected shape rather than a +surprise -- batching removes seven of every eight submissions from the append path, shortening the +interval between the last append and the flush being reached. Throughput is bound by the device +flush and does not move; latency is not, and does. + +The figures live in the capture and are **linked** from `strategy.rs` and from `M20.6` rather than +pasted into either, per the rule that a measurement has one home. + +**`M21.3`'s property was preserved deliberately.** Batching changes the append loop, and the naive +rewrite would have reintroduced a commit trigger reachable on a pass that appended nothing. The loop +`continue`s when zero records are accepted, so the epoch check is reachable only after progress -- +the same structural guarantee, stated in the same place. + +**Unblocks `M20.6`**, whose remaining question is the part no number speaks to: whether alternating +rings earns its cost on correctness and blast-radius grounds. + +### M22.2 -- Collapse the two free-slot implementations to one, derived from the arena's own outstanding counts rather than tracked beside them. The item called both correct; one was not -- the tracked free list leaked a slot on every refused append. *(completed 2026-09-22 16:21:23 -04:00)* + +The sample had two answers to "which arena slots are free". `Appender` asked the arena, filtering on +`outstanding(slot) == Some(0)`. `Lane` kept a `Vec` and maintained it by hand. Both now call one +`free_slots` in [append.rs](examples/epoch_log/append.rs); `Lane`'s field, its initialiser, and its +push-on-claim are gone. + +**The item's premise was wrong in the direction that mattered.** It said "both are correct and the cost +difference is nil ... the duplication is the defect, because the two *can* drift". They had already +drifted. The free list took its slot *before* composing into it, so an append refused between the two -- +a record too long for a slot is the reachable path -- returned with the slot popped and no operation ever +issued. The arena considered that slot quiet forever; the list never offered it again. `SLOTS` such +refusals and the harness reports a full arena while the kernel holds nothing. + +So this was not a tidying exercise with a correctness footnote: **deriving the fact removed a live bug**, +and it removed it by construction rather than by fixing the copy -- a slot nothing was pushed against +never stopped being free. + +**Verified by sabotage, in both directions.** + +- *Does each caller really bind to the one definition?* Breaking `free_slots` to offer busy slots failed + the `Lane` path with the arena's own refusal (`buffer 0 still has 1 operation(s) outstanding`), while + the appender ran clean -- so that run proved only half of it. A second sabotage (`take(0)`) starved the + appender, which then failed before the strategy section was reached. Both halves bind; one sabotage was + not enough to show it, which is the point of running the second. +- *Was the leak real, or argued from the source?* Re-injecting the free list and failing eight appends + left the lane reporting **0** of 8 slots free with the arena entirely idle. + +**What the new tests do not catch, stated in the tests.** [strategy/tests.rs](examples/epoch_log/strategy/tests.rs) +asserts the property from the arena's side, so a re-introduced free list would *not* fail it -- under the +re-injection above it passed, and only a temporary assertion against the list itself went red. What keeps +a second definition from returning is that there is one function and both callers call it. Claiming the +test covers that would be the cosmetic binding the repository's own rules warn about, so the module says +so plainly instead. + +Two harness defects were themselves caught by sabotage discipline and are worth recording, because both +produce a *false green*: a `.Replace` that matched nothing reported success and ran an unmodified tree, +and a PowerShell helper that logged with `Write-Output` returned its log line into the patched text. Per-site +match-count assertions caught the first; the second surfaced as a run with no output at all. The repository +already requires the match-count check for exactly this reason. + +**Raised a layer-placement question rather than answering it silently:** `outstanding()`'s rustdoc names +this use case, and both in-repo consumers hand-rolled it anyway. Queued as `M22+.2` with its blocker named +(a public API addition needs a lib test, and that test opens a real ring -- the `M24` pile). + +### M22.3 -- Give the registered arena a stated placement: the epoch-log arena is placed on the NUMA node its own log file's volume reports, and the allocator moved into the library as `NumaBuffer` rather than being copied a second time. The sample says plainly that the placement cannot pay at this workload. *(completed 2026-09-22 17:41:01 -04:00)* + +The item offered a choice -- adopt `ring_copy`'s allocator, or write down why a durability sample makes no +locality decision. Both halves turned out to be needed, and a third thing fell out of doing them. + +**The decision.** [placement.rs](examples/epoch_log/placement.rs) asks the log's own handle through +`FSCTL_QUERY_VOLUME_NUMA_INFO` -- the documented call [What is not reachable](DESIGN-NOTES.md) already +established, which takes a file or directory handle directly and needs no device-tree walk -- and the arena +is allocated preferring whatever comes back. A volume that names no node yields no preference and the log +runs on; the report line says which happened and why. + +**No benefit is claimed, and the sample says so in its own output.** The arena is eight slots of four +kilobytes against a workload bound by a per-epoch device flush costing hundreds of microseconds, which +`M22.1` measured directly. So this demonstrates how the decision is made and reported, not that it was +worth making; `examples/ring_copy` remains where placement meets a load that could show it. The report is +also qualified by `GetNumaHighestNodeNumber`, because on a one-node machine "placed on node 0" is true and +misleading -- the line instead says the choice was never available. Two unit tests hold that in **both** +directions: the disclaimer must appear for a single-node machine and must **not** appear for a multi-node +one, since a disclaimer that shows up everywhere trains a reader to ignore it. + +**The allocator moved into the library ([D-51](DESIGN-NOTES.md#d-51)), by the engineer's call on a +question raised before any code was written.** The crate's front page names this allocation as the +highest-leverage locality decision available and then supplied nothing, so the first consumer wrote it in a +sample and this item was about to write the second. `git mv` carried the history; `ring_copy` lost its +local module and binds to the library type. + +**That move surfaced a packaging defect that would have reached a consumer.** `cargo check --all-targets` +unifies dev-dependency features into the build, so the library's newly-required `Win32_System_Memory` was +being supplied by the dev-dependency list and the lib target compiled clean. `cargo doc` -- which does not +get dev-dependencies -- failed immediately, and `cargo check --lib` confirmed it: **anyone depending on +this crate alone would not have compiled it.** The manifest now carries the feature on the library +dependency, and the stale comment asserting "the library itself needs none of them" is corrected rather +than left to mislead. + +**Verified by sabotage, in both directions.** Forcing the FSCTL to report a node the machine does not have +failed the run at arena allocation with `ERROR_INVALID_PARAMETER`, which is what proves the queried node +actually reaches the allocator rather than being reported decoratively. Forcing the FSCTL to fail produced +the `Unplaced` line, the stated reason, and a log that still kept its contract -- the path that will not +otherwise execute on a machine where the query succeeds. + +**Swept the claim rather than the one site.** `VirtualAllocExNuma` and the arena's old `Vec` allocation +were stated in six places; [README.md](README.md), [lib.rs](src/lib.rs), two DESIGN-NOTES sections and the +manifest comment were updated, and the design-session and archive copies were left alone as historical +record. `M23.2` was **narrowed** in the same pass: its option (a) is no longer "should a sample do this" +but the residual library question, because the sample half is now done. + +Gate: fmt, clippy, the full suite (14 new `NumaBuffer` tests, 9 new placement tests), `cargo doc` clean, +lib-only and `--no-default-features` builds, the borrow-surface, encoding and publishable checks, and the +example end to end. + +### M20.3 -- Make `ring_copy`'s degraded-fallback path observable in a test, asserting both that an absent relation degrades and that a present one does not. *(completed 2026-09-22 20:11:18 -04:00)* + +The whole-machine fallback in `Policy::select` is the branch every zero-relation machine takes -- the +shape [D-48](DESIGN-NOTES.md#d-48) records as ordinary rather than exotic -- and it cannot be reached +by *running* the sample on a machine that reports its relations. A synthetic topology reaches it. +Fifteen tests in [examples/ring_copy/policy/tests.rs](examples/ring_copy/policy/tests.rs). + +**The item's reason for demanding both halves was verified rather than trusted.** It argued that a +test of the absent case alone "would pass against a function that always degrades". Sabotaging +`select` to degrade unconditionally showed exactly that: `a_policy_whose_relation_is_absent_...` +**still passed**, while four present-case tests failed. The reverse sabotage -- never degrade -- failed +five absent-case tests. Neither half is redundant, and that is now a measured statement. + +One test, `degrading_unconditionally_would_fail_a_test_here`, exists to put that dependency in code +rather than in a comment, so a future edit that deletes the present-case coverage has something named +to delete. + +**Done without waiting on `SH-4.12`, and the coupling was narrowed rather than ignored.** The recorded +callout said that item "rewrites the selection arm this test would assert against". That is true only +of a test asserting through `ByL3`. The fallback tail is shared by all five policies and is not what +`SH-4.12` changes -- it changes which domains `ByL3` matches -- so exercising it through `ByNode` and +`ByPackage` pins nothing. Both checklists now say so, and `SH-4.12` inherited the one assertion that is +genuinely its own: `ByL3`'s degradation condition, under whichever rule replaces the `level: 3` match. + +Beyond the two halves, the cases cover what the fallback must get right and what it must not claim: a +memory domain with no processors is not a usable node; degradation is per-policy rather than a property +of the machine; `Single` returns the whole machine **undegraded**, because degrading is a statement +about not getting what was asked for and `Single` asked for exactly this; the fallback covers every +online processor, excludes reserved-but-offline slots, and spans processor groups; and it carries no +observations, because nothing observed it. + +The synthetic memory domain uses `Observed::NotObserved` for its size rather than `Known(0)`, which the +type's own documentation calls the variant "a hand-written description leaves behind". `Known(0)` would +have asserted a measurement nobody made -- in a test whose subject is honest reporting. + +`ring_copy` was auto-discovered and therefore **not a test target**, so `cargo test` would have compiled +these and run nothing. It now has an explicit `[[example]]` entry with `test = true`, the same reason +`epoch_log` has one. + +### M20.1 -- Restate the cache heuristic as "the outermost cache level that actually partitions the machine", sweep every restatement, and replace the consumer that bound to the level number. *(completed 2026-09-22 20:27:15 -04:00)* + +**Done together with `SH-4.12`, because they are one change.** That coupling was real, unlike `M20.3`'s: +this item's sweep reaches `policy.rs`'s doc comments, and rewriting those to describe the new rule while +the code still filtered `level: 3` is exactly the contradiction the blast-radius convention exists to +prevent. Splitting them would have produced a commit whose documentation lied. + +**The item's evidence was a shipping ARM part with no L3. Measuring the consumer found a second shape, +on this workspace's own development machine, that nobody had anticipated.** It reports an L3 spanning +**all 16 processors** above a real **8-way L2** partition. So the old filter did not fail the way the +item assumed: + +| | domains selected | reported degraded? | +|---|---|---| +| old `level: 3` filter | **1**, mask `0xffff` | **no** | +| `outermost_partitioning_cache()` | **8**, at L2 | no | + +The old code *matched something*, so it did not degrade -- it reported success while collapsing an +eight-domain machine to a single ring. A silent wrong answer, not a visible fallback, and live on the +machine this repository is developed on rather than on hardware nobody here owns. + +Three consumers restated the rule, not the two the items named. `ring_copy`'s `policy.rs` and the prose +were known; `examples/l3_domains.rs` also hardcoded `cache.level == 3` and was named after the +assumption. It is now `examples/cache_domains.rs`, asks the same primitive, and reports the level it +found -- `git mv` kept its history. + +**`byl3` and `l3` are rejected rather than aliased.** They named a rule the sample no longer implements; +mapping them onto `ByCache` would let a script keep asking for L3 and keep believing it got L3, which on +an L2-partitioned machine is a wrong answer delivered quietly. An unknown policy prints the usage line, +which is a question rather than a wrong answer. + +Five tests were added to the file `M20.3` created two hours earlier -- the assertion `SH-4.12` had +inherited. Re-injecting the `level: 3` filter fails three of them, including the one that pins the +measured shape above. The swept sites: `DESIGN-NOTES.md` (the heuristic section, the sizing note, the +policy list, D-27's pointer, and D-48's own "restating the rule is M20.1" reference, which was itself a +restatement that would have gone stale), `README.md`, `src/lib.rs`, both examples, and both checklists. + +### M24.1 -- Settle whether a co-tested fake escapes the mock objection. Answered: the fake was the wrong instrument. *(completed 2026-09-22 21:18:30 -04:00)* + +The item predicted two outcomes and both held -- a wrong **accounting** model was caught by the shared +suite, a wrong **Windows belief** slipped through. Two further cases changed the answer. + +**Case 3, which the item did not predict, is the argument *for* co-testing.** Run an assertion written +from the *wrong* belief against both peers: the kernel goes red and refutes us, the fake goes green and +confirms us. That is the manufactured-evidence mechanism made visible, and also the escape -- a +mock-only world never performs that experiment. + +**Case 4 invalidated the line the first three cases suggested.** Raised in review: kernel behaviour is an +observation at a point in time, not objective truth, and we must not over-index on a record of how it +runs. Demonstrated: "after submitting, the completion is already queued" reads like a contract and gave +**opposite answers on two handles of the same API**. So the axis is not "accounting versus Windows +behaviour" -- it is **our specified contract versus the platform's incidental behaviour**. + +**Case 5 replaced the technique.** Also from review: model what the platform is *permitted* to do and +let a seed pick a resolution, so the assertions are about this crate rather than about the kernel. That +dissolves the mock objection instead of working around it, because there is no belief to be wrong about. +A minimal resolver broke a FIFO-assuming consumer under 189 of 200 seeds -- and passed under the other +11, which is the point: a fixed fake reports whichever single answer it encoded. + +**Two apparatus failures in one session, both caught, both the same shape.** The first spike draft ran +each condition once and printed a verdict -- exactly `D-47`'s error; rewritten to 500 trials it +immediately found a condition that pends about 1% of the time and would have been called "never". The +case-4 harness used a bare flush, which completes inline on every handle, so it could not discriminate +until it was rebuilt around sector-aligned writes. Both are recorded because the repository's rule is +that an instrument nobody has shown can go red is not evidence. + +Outcome: the rejection **stands** with its scope sharpened ([D-52](DESIGN-NOTES.md#d-52)), `M24.4` is +**withdrawn**, `M24` becomes unconditional, and the technique that actually answers the question is +`M26`. The apparatus is kept as +[kernel-response-space-probe.rs](design-sessions/kernel-response-space-probe.rs); the reasoning is +[DESIGN-SESSION-2026-09-22-kernel-response-space.md](design-sessions/DESIGN-SESSION-2026-09-22-kernel-response-space.md). + +### M24.2 -- Extract the handle-free accounting into its own type, composed by `IoRing`. *(completed 2026-09-22 21:29:07 -04:00)* + +The item named five fields as handle-free and five as carrying kernel state. **Checked before acting, +and it was exactly right** -- `ring_id`, `next_user_data`, `outstanding`, `registered_files` and +`registered_buffers` against `handle`, `version`, `supported_ops`, `registered_buffer_infos` and +`completion_event`. Nine methods touch only the first five; they moved with the fields, and `IoRing` +delegates. `RingId` moved too, since the identity counter is part of the ledger rather than of the +handle. The public surface did not move. + +**19 hermetic tests, and they are the first in `src/` for which [D-49](DESIGN-NOTES.md#d-49)'s +complaint does not apply.** They open no ring, because every rule they check is this crate's own +specification: an identity is never reused, a refused reservation costs nothing, the counters saturate +rather than wrap, the two registration indices are independent, and two ledgers never share an +identity. The last of those used to need two live kernel objects. + +**The tests found an off-by-one in their own author's assumptions.** Two of them asserted that the +last identity handed out is `usize::MAX`, and failed: `checked_add` runs *before* the value is +returned, so a reservation made at `usize::MAX` fails rather than handing it out, and the identity +space is `0..=usize::MAX - 1`. Not a defect -- one value out of 2^64, and "fails rather than wraps" is +the property that matters -- but invisible from the source, so +`the_last_identity_is_max_minus_one_not_max` records it rather than leaving the next reader to make +the same wrong assumption. + +**Sabotage, four ways, all caught:** a no-op `record_completion` (2 red), a recycled identity (6 red), +`wrapping_sub` in place of `saturating_sub` (1 red), and the buffer-registration count advancing the +file counter (4 red). The recycled-identity sabotage first produced *no output at all* rather than a +red suite -- an ambiguous-integer compile error -- which is the silent-failure shape the repository's +rules warn about, and was rerun with an explicit type before being believed. + +**The payoff was deliberately not taken here.** The item's motivation said most of the ring-opening +lib tests "become hermetic *in place*". That is 71 tests across four files, which is a different item +wearing this one's name; it is queued as `M24.7` with the measured per-file census, and with a warning +that the census must be recounted because the first attempt at it produced false positives by matching +`to_string()`. + +### M24.7 -- Convert the lib tests that construct a ring only to exercise bookkeeping. *(completed 2026-09-22 21:43:29 -04:00)* + +**The recount the item demanded was right to demand.** Its figure of 71 was a per-*file* +`IoRing::new` count. A per-*test* census gives **61**, and the difference is not rounding -- a file +with 40 tests and 35 constructions has five tests that never touch a ring. + +**61 -> 52.** Two changes, one structural and one an excision. + +**`Token::new` now takes the ring's ledger rather than the ring.** It only ever used +`reserve_user_data()` and `ring_id()`, both bookkeeping, so the wide parameter was the only reason +`token`'s tests opened a ring at all -- **all seven of them, to mint a token and nothing else.** The +ring was actively a liability there: none of those tests ever submitted, so `Drop`'s run-down would +wait for completions that were never coming, and a `settle` helper existed purely to stop teardown +hanging. The helper is gone with the hazard it worked around. `token` is now 7 hermetic, 0 opening. + +**Two `ring` tests were removed rather than converted**, being duplicates of what `M24.2`'s hermetic +tests now cover: `reserve_user_data_increments_outstanding_and_never_repeats_an_id` and +`record_completion_saturates_rather_than_underflowing`. Deleting them loses no delegation coverage -- +`run_down_returns_once_a_recorded_completion_zeroes_the_count` already drives reserve, `outstanding` +and `record_completion` through `IoRing`, and must keep a ring for its own sake. + +**The remaining 52 are not convertible, and the reason is structural rather than effort.** Recorded +here so the next reader does not re-derive it: + +- `event_delivery` (6) needs a real ring and the thread pool. There is nothing to narrow. +- `ring`'s injected-failure cluster **looks** convertible by name and is not. It uses a real + completion on purpose -- one test says so in an assertion message, "the flush really did succeed, + or this test proves nothing" -- because `with_injected_failure` *transforms* a real completion, and + fabricating one is precisely the unsoundness the seam exists to avoid. +- `batch` (13) needs a `Batch`, which needs the handle for its `Build*` calls. Narrowing that + parameter is not possible the way `Token`'s was; it becomes reachable only under `M26.2`'s FFI + seam, which is a far larger change. + +So **relocation, not conversion, is the remedy for the rest**, which is `M24.3` -- updated with this +finding and with a warning to recount its own stale figure of 25. + +### M24.3 -- Relocate the lib tests that open a ring but use only public API into `tests/`. *(completed 2026-09-22 22:43:39 -04:00)* + +**52 -> 41.** Eleven tests moved into `tests/ring_lifecycle.rs` (new), `bounded_pop.rs`, +`event_delivery.rs`, `registration.rs` and `submission_lifecycle.rs`. Totals conserved: 162 lib + 68 +integration before, 151 + 79 after. + +**The item predicted 25 and that was never achievable.** Its figure predated two milestones of test +growth, and more importantly it assumed the constraint was *which* tests had been looked at rather +than what they reach. `M24.7` had already found the structural version of this; relocation hits the +same wall from the other side. + +**A regex census got the classification wrong, and the compiler caught it.** The scan excluded +`pop_within` from the blocking list because it is a public method on `IoRing` -- but `ring.rs` also +has a `#[cfg(test)] pub(crate) fn pop_within(ring, what)` free function, and +`windows_refuses_an_empty_buffer_registration` uses *that*. It was moved, failed to compile, and was +returned. A second miss was structural rather than nominal: the scan only looked at `fn` definitions, +so it did not see that `a_supplied_wait_is_not_consulted_when_nothing_can_arrive` depends on a +`RecordingWait` **struct** shared with eight other call sites. That one was caught before moving, by +a second pass that looked for local `struct`/`const` definitions too. + +The method that worked was **moving the candidates and letting the compiler rule**, rather than +trusting the census. Two milestones running, a crude census has produced false classifications here; +the compiler produced none. + +`HugeBuffer` and `NULL_FILE` travelled with the two tests that used them, having no other call sites. + +**`M24.5` needed re-planning as a result, and that is the more consequential outcome.** It assumed +that after `M24.2` and `M24.3` the lib tests would "construct no ring at all", so a zero-check would +do. Forty-one remain and none is movable, so a zero-check would fail on day one and could only be +satisfied by deleting real coverage. The item now asks what the rung should actually assert -- a +ratchet, an allow-list, or nothing until `M26.2` -- rather than presuming the answer. + +### M24.5 -- Put the rule on a rung, so it cannot regress. *(completed 2026-09-22 23:44:37 -04:00)* + +**The rule the item assumed was false, and that is the decision this item really made** +([D-53](DESIGN-NOTES.md#d-53)). It expected a zero-check -- "the lib tests should construct no ring at +all" -- which would have failed on day one and could only ever be satisfied by deleting real coverage. +Forty-one remain and none is movable without `M26.2`. + +So the rung is an **inventory**: [RING-OPENING-LIB-TESTS.txt](RING-OPENING-LIB-TESTS.txt), regenerated +from source by [check-ring-tests.ps1](../../tools/check-ring-tests.ps1), failing when the two +disagree. Deliberately the same mechanism as the borrow-surface check, so there is nothing new to +explain. Per-test rather than per-file, because two thirds of the 41 live in `ring/tests.rs` and a +file-level allow-list would let exactly that file grow. An inventory rather than a count, because +add-one-remove-one nets to zero and a bare number is derived data nobody can check by reading. + +**The bidirectional verification found a defect in the guard, which is the whole reason the rule +demands it.** Direction 3 -- a test that reaches a ring only through a helper -- reported the expected +test *and an innocent one*. The body extraction ended at the next `#[`, so a plain helper defined +after the last test in a file was swallowed into that test's body, and a helper containing +`IoRing::new` made the test above it look ring-opening. The body now ends at a column-0 `}`, which +`cargo fmt` guarantees is a function end. Had only the "must fire" direction been run, the check would +have shipped with a false positive that fires on innocent changes -- the fastest way to train people +to ignore it. + +Four directions verified after the fix: fires on a direct `IoRing::new`, fires on a ring reached only +through a helper, stays silent on a new hermetic test, and reports removals as progress needing only +regeneration. Wired into CI as its own job beside `borrow-surface`; needs no toolchain. + +### M24.6 -- Sweep what this milestone makes false. *(completed 2026-09-23 11:49:58 -04:00)* + +The item named three sites. **Two were false alarms and the third was false for a different and more +serious reason than the item gave** -- which is the argument for running the census rather than +editing the named list. + +**Named, and genuinely stale: [D-49](DESIGN-NOTES.md#d-49).** Its "63 of 131" is the figure `M24` +*started* from; measured after, **41 of 151 open a ring and 110 do not**. It also still said the mock +rejection "stands until `M24.1` settles it" and that the remedy choice was "gated on one unresolved +question", both of which `M24.1` closed. Corrected, with pointers to [D-52](DESIGN-NOTES.md#d-52) and +[D-53](DESIGN-NOTES.md#d-53). + +**Named, false alarm: the testing-strategy section.** The item expected it to be stale because it was +"written when every lib test opened a ring". Reading it, nothing in it turns on hermeticity -- it +classifies *defect populations* and *techniques*, and `M24` added or removed neither. Its "all five +techniques" framing is also correctly left alone: the resolver is a sixth *when `M26` builds it*, and +`M26.6` already owns that edit. Claiming six today would be the opposite error. + +**Named, false alarm: `M21.6`'s archive entry.** The item said its "a wait that never enters the +kernel" clause "stops being the notable exception once the suite is hermetic". Hermeticity does not +bear on that sentence, and the archive is append-only history describing what was true when written. + +**Named, and false -- but not because of `M24`: `F-13`.** Its headline says "Every fixture in this +crate's tests, examples and samples opens its handle that way [synchronous]". Three do not: +`flush_barrier.rs`, `handover.rs` and `flush_barrier_stress.rs` open +`FILE_FLAG_OVERLAPPED | FILE_FLAG_NO_BUFFERING`, dated 2026-08-28, 08-29 and 09-06 -- **weeks before +F-13 was recorded on 09-21**. So it was false when written, not made false by this milestone. + +That matters because the entry's carry-forward escalated from the false half: it says every claim +about ordering, draining, the completion event and the barrier "was measured against operations that +may have completed inline" and names D-19, D-23, D-24 and D-47 for re-reading. But D-23, D-24 and +D-47 were measured by `flush_barrier.rs` -- one of the three overlapped fixtures. The entry even +hedged correctly ("the drain-ordering spike used `NO_BUFFERING` ... so it is probably fine") and then +checked only the spike, not the tests sharing its shape. A dated correction was added rather than a +rewrite, so the record of what was believed survives. + +**Unnamed, and found by the sweep: [bounded_pop.rs](tests/bounded_pop.rs) named the wrong gap.** It +said "every other test of `pop_within` drives the loop with a wait that never enters the kernel". +Several do enter it through `SubmitWait`. The real gap is narrower and more interesting: the tests +using the kernel wait drive operations that *complete*, and the one test that lets a bound expire +fakes both halves -- a `RecordingWait` instead of the kernel and a bare `reserve_user_data` instead of +a pending operation. **No test had a real operation pending when a real bound expired**, which is +exactly the state the `ERROR_TIMEOUT` path needs. Corrected in place. + +**Unnamed, and found by the sweep: this milestone's own header** still carried the 63-of-131 opening +figure and a "recount before starting" caution that had been acted on. + +**The transferable part.** Three of the five corrections were over-generalisations from a single +observation -- one fixture becoming "every fixture", one wait shape becoming "every other test". Each +was a census away from being right, and each then had an alarm built on top of it. That is the same +shape as the spike that ran one trial per condition, in the same crate, two days earlier. + +### M20.6 -- Re-evaluate `CommitStrategy::AlternatingRings` and the epoch-log benchmark's conclusion against D-47. *(completed 2026-09-23 12:06:40 -04:00)* + +**Question 1 -- does alternating rings still earn its cost? Answered on its stated grounds: no.** +`S-2` argued that because a covering flush reaches every operation outstanding on its ring, two rings +bound what a commit's barrier can be dragged into. That is structurally false for this sample and +needs no measurement: `RegisteredBuffers::get_mut` refuses a busy slot and there are `SLOTS` slots, so +at most `SLOTS` appends are outstanding on a ring **by construction** -- and each alternating lane +registers its own arena of the same size. The arena bounds the blast radius, not the ring topology. +Probing agreed (8 and 8); the argument does not rest on it and holds whatever the platform does about +pending. The argument survives against genuinely unrelated traffic from another component; this sample +has none. + +**The strategy is deliberately NOT removed.** The item said that if it no longer earns its place, that +is an API change to a published example -- and the temptation was to make it. What two rings could +*also* buy is **overlap**, and overlap is precisely what this harness cannot exhibit: its handle is +synchronous, so a ring operation completes inline during submit and nothing is ever outstanding across +a submit boundary. Deleting a strategy on the strength of a measurement that could not have shown it +working would be the same error as the measurements this item exists to correct. `M25.5` answers it on +a harness where operations genuinely pend. + +**Question 2 -- re-read, re-run, or annotate the numbers?** The investigation concluded "none of +those": the column measures deferral rather than a commit, and prose cannot fix a measurement. +**That conclusion was right about the measurement and wrong about the output.** Leaving a column +labelled `commit p50` in a published sample until `M25` lands is shipping a false claim for the sake +of a purist position on annotation. The column is now `ack lag`, with a caveat line naming the p99 = 0 +blocking measurement, and pointing at `M25`. + +**The rustdoc already knew, which is the finding worth keeping.** `Outcome::commit_latencies` already +said the figure is "not device flush time", that deferral inflates it, and that alternating rings +"reports the highest latency of the three while matching them on throughput". A previous pass had +diagnosed the artifact correctly **and only in the rustdoc** -- the printed output never got the same +treatment, and the M20.6 investigation re-derived from scratch what was already written one file away. +What the investigation genuinely added is the extent: blocking is not merely a component of the figure, +it is **0 us at p99**, so the number is entirely deferral; and underneath that, no pipeline exists at +all. + +**Swept the mechanism, not just the label.** The explanation "the strategies differ about how long the +flush itself waits and the extra host round trip, and those land in the tens" appears in the module +docs, in a `main.rs` code comment, and in the printed summary line. It is wrong in the same way at all +three: those differences cannot occur on a synchronous handle. The three are indistinguishable because +they do the same serialized work -- both readings give the same ranking and only one is true. All three +corrected, plus the "a real log keeps appending while a commit is outstanding" claim, which describes +a state this program has never reached. + +### M20.6 -- correction, same day *(recorded 2026-09-23 13:03:35 -04:00)* + +**The entry above declared `AlternatingRings`' blast-radius justification "dead on structural grounds". +That over-reached, and the over-reach is the kind this repository now has a rule against** -- see +OPTION INTEGRITY in the repository instructions, added by this correction. + +What the structural argument actually establishes is that **this harness** cannot exhibit a +blast-radius difference, because each lane registers its own arena of `SLOTS` slots and the arena is +the limiter rather than the ring topology. That is a statement about the apparatus. Generalising it to +"the justification is dead" converted a fact about one sample's configuration into a verdict on a +design option. + +**It also contradicted the crate's own recorded position.** [D-27](DESIGN-NOTES.md#d-27) is this +crate's decision that one ring per thread is userspace's proxy for one ring per CPU, and records the +hardware reason: NVMe queue pairs are per-CPU with each pair's completion interrupt routed by its own +vector. Two rings on two pinned threads *is* that architecture. Declaring a multi-ring strategy's +justification dead on the strength of one sample's arena sizing sits directly against a decision the +crate already made on stronger grounds. + +The conditions under which alternating rings would pay are now written at +`CommitStrategy::AlternatingRings`, and they are ordinary rather than exotic: a ring shared with any +other component, arenas sized asymmetrically from the lanes, real overlap (where the same covered +count is not the same wait), and per-CPU queue affinity. The sample's job is restated as giving a +consumer the means to answer this on their own hardware, not handing them a verdict from ours. + +Nothing about the measurement corrections in the entry above changes: the ack-lag relabel, the p99 = 0 +blocking finding, and the swept mechanism claim all stand. What changed is the conclusion drawn from +them. + +### M23.1 -- State in the epoch-log contract that the barrier is ring-wide while the flush names a file, so one ring per log is a precondition of the cost model. *(completed 2026-09-23 17:03:35 -04:00)* + +*Queued from finding `S-1` of the 2026-09-19 epoch-log review.* + +**The item was right that the contract was silent, and wrong about what it was silent on.** +`S-1` reasoned from [D-47](DESIGN-NOTES.md#d-47)'s surviving half -- the barrier reaches every +operation outstanding on the ring, not only the current submission batch -- and concluded that +"one ring per log" is a precondition of the sample's *durability contract*. +[contract.rs](examples/epoch_log/contract.rs), written before the code precisely so it would state +preconditions, said nothing about it. + +**Writing it found two scopes conflated, and the first draft shipped the conflation.** +`IOSQE_FLAGS_DRAIN_PRECEDING_OPS` is a flag on the **ring**; `BuildIoRingFlushFile` names a +**file**. So the barrier bounds what a commit *waits for* and the flush bounds what it *makes +durable*, and completion is not durability -- a claim the same file already made three paragraphs +earlier, about a record's own write. The corrected reading: a shared ring does **not** endanger the +guarantee, which the flush's own file target secures. It endangers the **cost model**, because the +barrier waits for unrelated traffic unconditionally. + +**Three commits, because the first two were not right.** `d845bb28` added the section, an +assumption, a non-guarantee, `Clause::ALL` -- replacing a hand-written variant list in `main.rs` +that would have printed one section short had a fourth clause ever been added -- five tests over the +report's own properties, and the crate's first [sabotage.json](sabotage.json): five injected defects +caught, plus a control that rewords a statement and survives, so the guards are sensitive to the +report degrading without being bound to the contract's wording. `4adea675` corrected the +conflation. `48dc99d3` right-sized what the correction had grown into -- a title giving the device +equal billing with the ring, plus a bullet and a milestone pointer about multi-device reach, in a +sample that runs one log file on one ring. + +**The conflation was caught by a question, not by the gate**, which stayed green across all three: +every test passed, the sabotage sweep reported all six cases as declared, and the two contradictory +sentences sat a screen apart in one file. The tests check that the report *prints* correctly and +deliberately not what it *says*, so nothing built here could have found it. + +**Two things this work left elsewhere.** The primitive-level half moved to the library under +[D-54](DESIGN-NOTES.md#d-54): the barrier/flush scope distinction is a fact about one flush, so it +belongs on `FlushCoverage` rather than only in a sample's contract. And +[checkpoint.rs](examples/epoch_log/checkpoint.rs) gained the reciprocal note -- it already took its +own ring, for an unrelated *delivery* reason (D-21), so the structure was right twice over with only +one reason written down. Both sites now point at each other, so collapsing the rings cannot look +harmless from either end. + +### M23.2 -- Decide how a caller arrives at a NUMA node: `win-numa-sys` offers declaring and discovering, and refuses the shortcut that does both at once. *(completed 2026-09-23 19:57:34 -04:00)* + +*Queued from finding `S-3` of the 2026-09-19 epoch-log review.* + +**The item was narrowed three times before it was answered, and the last narrowing moved it out of +this crate entirely.** `M22.3` settled its sample half by making the allocator library surface. +[D-54](DESIGN-NOTES.md#d-54) removed its other half -- sharding by backing device needs the concept +of a set of operations that commit together, which this crate does not have -- and handed that to +the durability layer, where it was sharpened from device identity to *flush equivalence*. Then +`win-numa-sys` was created and `NumaBuffer` moved into it, so "this crate" in the item text stopped +naming the crate that had to answer. + +**What remained was answered by building that crate, so this item's deliverable was the recorded +decision rather than code.** It is +[N-D-1](../win-numa-sys/DESIGN-NOTES.md#n-d-1): a caller may *declare* a node +(`NumaBuffer::new`), *discover* one (`volume_numa_node`), or *qualify* what a discovered answer is +worth (`highest_numa_node`); what is refused is a `NumaBuffer::for_file` that would query and +allocate in one step. + +Four reasons for the refusal, of which the first is the one that generalises: such a call **hides +the answer**, and on a single-node machine "placed on the node the volume named" and "no preference" +are the same allocation, so a caller could not tell whether the query found anything. It also fuses +two failure domains, withholds an answer useful beyond one buffer, and saves exactly one line, +since `NumaBuffer::new(len, volume_numa_node(h).ok())` already type-checks. + +**A fifth argument was dropped rather than kept, and the decision says so.** When this was first +argued, a `for_file` constructor would have dragged `Win32_System_Ioctl` into a crate that +otherwise touched only memory. That was true of `windows-ioring-sys` and is not true of +`win-numa-sys`, where the query already lives. Recording a void argument as void is cheaper than +having someone re-make it. + +**Neither path is speculative.** `examples/epoch_log` discovers from its log file's volume; +`examples/ring_copy` declares a node it computed from the processor topology. Both were already +written against this shape before the decision recorded it. + +[D-8](DESIGN-NOTES.md#d-8) is intact, which was the item's stated constraint: locality stays the +consumer's decision, and the crate supplies a fact and an allocator rather than a choice. + +### M23.3 -- Decide what this crate offers for holding a token between push and completion: the ring owns the inventory, `IoRing` becomes generic, and the break is accepted. *(completed 2026-09-23 23:09:13 -04:00)* + +*Recorded as [D-55](DESIGN-NOTES.md#d-55). Implementation is `M28`. The exploration, including +everything it falsified, is in +[DESIGN-SESSION-2026-09-23-pending-inventory.md](design-sessions/DESIGN-SESSION-2026-09-23-pending-inventory.md).* + +**The item was decided against a different argument than the one it was written on.** It argued +from duplication -- nine sites keeping the same map. A census found ~12 sites of which only a +third keep the map described, so duplication was both mis-counted and the wrong frame. The +mechanism is that this crate **mandates** the construct and does not provide it: +`Batch::write` returns a `Token`, `IoRing::pop_within` returns a `Completion`, and nothing +connects them but caller-supplied storage -- which the crate's own rustdoc instructs callers to +build, twice. That follows from [D-4](DESIGN-NOTES.md#d-4) splitting ring-side counting from +caller-side identity, which is right; what was missing is the half it left to prose. + +**A working spike was built and is why the decision is informed rather than argued.** +`Pending` fits a real consumer -- converting `append.rs` removed its `InFlight` struct and +its hand-driven oracle calls -- and sabotage established that removing its drop guard or cutting +its oracle wiring is caught. A test driving a failed write through the injection seam turned +`M22.2`'s ordering defect from undetected into caught by assertion. + +**But the spike also showed why offering it beside the ring is not enough.** Nothing forces a +minted token into it, so the ring-to-inventory drift survives -- and `Pending::checked()` owning +an oracle made the converted consumer's existing `RingContract` a decoy, passing +`assert_quiescent()` vacuously with nothing in the suite catching it. A generic `IoRing` +owning the map removes the class, because the consumer never holds a token to lose. + +**The objection that had ruled that out was false, and checking it was what settled the item.** +Per-ring monomorphisation holds for every real consumer; `tests/generated_sequences.rs` already +carries eight token types on one ring behind a closed `enum Held` with no runtime type check; +and [D-4](DESIGN-NOTES.md#d-4) rules type erasure out in as many words. The dismissal had +contradicted a decision already on the books, in the opposite direction from the one it assumed. + +**Two findings the decision rests on that were not in the item.** `RingContract` never prunes -- +one retained entry per operation for the process's life, undocumented, invisible in a sample that +appends 24 records -- which rules out checking by default and is `M28.2`. And `epoch_log`'s +commits are tokenless because a flush has no buffer and a *borrowed* `RawHandle` leaves its token +nothing to guard, which is `M28.5` and may dissolve when `M25.3` changes how the log is opened. + +**What it refuses**, unchanged by the shape: batching, ordering and which slot to pick stay caller +questions. The sharper refusal is that the inventory does not decide whether a caller is checked. + +### M23.4 -- A failing test that left registered buffers outstanding aborted the process instead of reporting; the drop guards now stay silent during unwind. *(completed 2026-09-23 23:16:48 -04:00)* + +`RegisteredBuffers::drop` refused to free while an operation was outstanding (M5.3, correctly) and +said so with a bare `debug_assert!(false, ...)`. Nothing checked `std::thread::panicking()`, so a +test that failed *because* a slot leaked panicked, unwound, dropped the arena, panicked a second +time inside `Drop`, and aborted -- replacing its own assertion message with +`STATUS_STACK_BUFFER_OVERRUN`. The detection was never weakened; what the abort destroyed was the +diagnosis. + +**The item named one site and there were three.** It said "the fix is presumably the same one line +here", and a sweep of every `impl Drop` in the crate found the assert it named plus **two more** in +`IoRing::drop` -- the rundown failure and the `CloseIoRing` failure. That second impl is the worse +of the two: a ring is dropped on the way out of almost every failing test in this crate, so an +unguarded assert there converts a readable failure into a crash in the *common* case rather than a +rare one. The reported site was a sample of the population, which is what CONTRACT INTEGRITY rule 3 +says to expect. + +**The sweep also produced two false positives worth naming**, because the pattern that produced +them is the obvious one to reach for. A grep for `impl.*Drop for` matched cargo-mutants-style +comment text (`::drop -> ()`) inside test files, which read as two +further unguarded sites. Anchoring the pattern at line start reduced eight candidate impls to the +three real asserts. A loose grep over a crate that documents its own mutants will find its +documentation. + +**`Pending::drop` was already correct** and is what the fix copies -- it returns early when +`std::thread::panicking()`, which is why the M23.3 spike never exhibited this. + +**Verified by sabotage, as the item required, and the verdict alone would not have shown it.** The +`M22.2 regression` case in [sabotage.json](sabotage.json) was `caught` before the fix and `caught` +after; what changed is that it ended in `exit 101` -- a clean `FAILED` naming +`a_failed_write_still_releases_its_arena_slot` and its message -- instead of `exit -1073740791`. +That case's `why` text, which had documented the abort as expected behaviour, now carries the +post-fix failure mode and says a regression in *either* direction (no longer caught, or caught but +crashing) is visible there. The full sweep stayed at 9-of-9 as declared with the `CONTROL` still +surviving. + +**Only one of the guards has a test that depends on it firing, and the sabotage that established +that also falsified the first draft of this entry.** Suppressing both guards unconditionally +(`true` in place of `std::thread::panicking()`) turned +`batch::tests::dropping_a_registration_with_work_outstanding_is_refused` red, which is the check +that the fix *narrowed* the guard rather than removing it -- that test drops deliberately, not +during an unwind, so `thread::panicking()` is false and the assert still fires. But +`ring::tests::dropping_a_ring_actually_runs_its_drop_body` stayed green under the same sabotage. +It is not a `#[should_panic]` test and it does not reach either assert; this entry had claimed it +did, on the strength of its name, until the sabotage said otherwise. + +**`IoRing::drop`'s two asserts are therefore unreachable from any test on a healthy host**, and +that is recorded at the definition rather than left to be rediscovered. `run_down` fails only when +`SubmitIoRing` or `PopIoCompletion` returns a kernel error HRESULT, and `CloseIoRing` fails only +when the kernel refuses the close; the crate's `fault-injection` seam sits at the +*completion-result* level (`Completion::with_injected_failure`) and cannot produce either. Reaching +them needs a seam over the raw HRESULTs, which is `M23.5` -- spawned rather than assumed, per the +move-or-spawn rule, because "no test can reach it" is a blocker to name and not a reason to check +the box and move on. The fix still lands there on its merits: it is precisely the `IoRing` case +that turns a readable failure into a crash most often, since a ring is dropped on the way out of +almost every failing test in this crate. + +> **The paragraph above was wrong about the remedy, and [M23.5](#m235) overturned it the same +> evening.** Both asserts are reachable, no seam was built, and no production line changed: the +> kernel rejects a *null* ring handle cleanly, and `ring::tests` is a child module that can put one +> in the field. What the paragraph got right is the finding that prompted it -- the guards were +> genuinely uncovered. It is left standing rather than rewritten because the archive is history, +> and because the error in it is instructive: it priced a seam it never checked was necessary. +### M23.5 -- Both asserts in `IoRing::drop` are now reached by tests, and the seam the item priced turned out not to be needed. *(completed 2026-09-23 23:20:27 -04:00)* + +M23.4 left both asserts in `IoRing::drop` uncovered: suppressing them entirely left every test in +the crate green. This item asked whether to build a fault-injection seam over the raw HRESULTs the +ring's Win32 calls return, or to accept the asserts as documented-unreachable. **Neither. The item's +premise was false**, and one probe falsified it. + +**What the probe measured.** `CloseIoRing(null)` and `SubmitIoRing(null, ..)` both return +`0x80070006` -- `HRESULT_FROM_WIN32(ERROR_INVALID_HANDLE)` -- a clean refusal. `CloseIoRing` on a +plausible-looking `0xDEAD_0000` raises `STATUS_ACCESS_VIOLATION`. So a ring handle is **a pointer +the kernel dereferences, not an index into a handle table**, and null is the one bad value that is +refused rather than followed. That asymmetry is the whole finding, and it is recorded at the +definition and in both tests, because a later cleanup that "tidies" the null into a non-null +sentinel converts two passing tests into a process crash. + +**Why no seam was needed.** `mod tests` is a *child* of `ring`, so it already sees the private +`handle` and `accounting` fields -- a child module can see its ancestors' private items. A +`#[cfg(test)]` constructor, `IoRing::refused_by_the_kernel`, assembles a whole `IoRing` around a +null handle, and the tests let the real `Drop` body run against it. Whether `run_down` submits at +all is what selects between the two asserts, since it loops only while something is outstanding. +Production code was not touched: the blast radius the item worried about was zero, because the +change is entirely in test-only code. + +**The D-49 ring-test gate improved the design, which is what it is for.** The first working version +opened a real ring, closed it by hand, and put a null in the field -- and +[tools/check-ring-tests.ps1](../../tools/check-ring-tests.ps1) flagged two `ADDED` entries and +asked its standing question: *does this test need the kernel, or only a ring-shaped thing?* Only +the latter. `RingVersion::V1` is a public const and `OpSupport` derives `Default`, so every one of +`IoRing`'s six fields is constructible without opening anything. Answering the gate rather than +re-baselining it removed the real ring, the hand-close, two `unsafe` blocks and their safety +arguments from both tests, and left the ring-opening population unchanged at 41. These tests need +the kernel only to *refuse* them, and refusing costs no ring. + +**Three sabotages recorded in [sabotage.json](sabotage.json), not one.** Suppressing the rundown +guard leaves the close test green and vice versa, so the two asserts are independent conditions and +a single case would have declared the pair covered while half of it was not. The third sabotage is +of the *test* rather than the code: removing the reservation makes the ring fall through to the +close, and the panic message becomes `CloseIoRing failed: 0x80070006` against an expected substring +of `IoRing rundown failed before close`. That is what shows `should_panic`'s `expected` string is +load-bearing in selecting the assert rather than decorative. The full sweep is 12-of-12 as declared +with the `CONTROL` still surviving. + +**The methodological point, which is the same one this crate keeps paying for.** The item was +written an hour earlier, by me, and it reasoned from the shape of the code to "this needs a seam" +without ever asking the kernel what it does with a bad handle. It then priced that seam's blast +radius and proposed accepting a permanent coverage gap as the alternative. Both options were +answers to a question that a single `eprintln!` dissolved. The repository's standing instruction is +never to report that something cannot be done on the basis of reading it; that applies to "no test +can reach this" exactly as it applies to "this will not compile". + +**A note on release builds, swept but deliberately not changed.** These are `#[should_panic]` tests +over `debug_assert!`, so they would fail under `cargo test --release`, where the assert compiles +out. That is a pre-existing property of the crate -- +`batch::tests::dropping_a_registration_with_work_outstanding_is_refused` has the same shape and is +ungated -- and no CI job runs tests in release. Matching the existing precedent was preferred over +introducing a `cfg(debug_assertions)` gate on two of the three, which would have left the crate +inconsistent with itself. If release-mode testing is ever added, all three need the gate together. + +### M25.1 + M25.2 -- Records gained a fixed sector stride with a zeroed block tail, and replay learned to walk it. *(completed 2026-09-23 23:36:23 -04:00)* + +**These two items could not land separately, and that is a defect in how they were written rather +than a discovery about the code.** A strided writer and an unstrided reader do not describe the same +file. The checklist sequenced M25.1 (writer) before M25.2 (reader), so the tree between them holds a +log nothing can read. They are recorded here as one entry, citing both IDs, per the checklist rule +for acknowledged coupling; the alternative -- restructuring them into independent items -- is not +available, because the writer and the reader of one format are not independent. + +**What the coupling actually cost was nearly a silent break.** With M25.1 applied alone, the log was +unreadable: replay advanced by a record's own length, landed in a zeroed block tail, decoded +`NeverWritten`, and reported every record after the first as a missing durable record. And **all 21 +of the example's tests passed anyway.** The only thing that caught it was +`cargo run --example epoch_log`, which asserts and exits 101 -- and no CI job runs the sample. That +is the more important finding of the two, and it is queued as `M25.1b` rather than left in this +entry, because a finding recorded only in an archive is a finding nobody is obliged to act on. + +**Three end-to-end tests were added to close the specific hole**, each binding one writer to the +real reader through a real file: `records_land_one_per_stride_and_replay_walks_them_back` over the +log's own `Appender`, `a_run_lays_its_records_out_one_per_stride` over the harness's `Lane`, and +`a_reused_slot_does_not_write_the_previous_records_tail` over the zeroing. The harness test exists +because a sabotage said it had to: reverting `Lane` to a packed layout was **caught by nothing** +until it was written. + +**The zeroing is not observable through replay, and the comment says so rather than inventing a +failure mode for it.** Replay decodes only at block starts and takes a record's extent from its own +header, so a stale fragment past a short record's end is never read. A first draft of the comment +claimed a stale fragment "would be decoded as a record", which is false for exactly that reason. +What zeroing actually prevents is the log carrying fragments of unrelated records -- a hygiene +defect in a format whose purpose is reconstructing what happened after a crash -- so the test +asserts on the file's bytes rather than on a replay outcome. + +**Two duplications were collapsed rather than converted twice.** The sample has two writers over one +on-disk format, and both carried a packed layout *and* a verbatim copy of the comment justifying it. +Converting each in place would have turned one duplicated decision into one duplicated rule, so the +stride moved to `record`, beside the format it describes, and the whole composition -- encode, then +zero the remainder -- became `record::encode_block`, which both writers call. That makes half the +drift unrepresentable rather than merely tested for. + +**`Decoded::total_len` became a derived `extent()` rather than a silenced warning.** Once replay +advanced by the stride, the field's only consumer was a test, and it was exactly +`HEADER_LEN + payload.len()` -- a stored copy of a fact `payload` already carried. The dead-code +warning was the signal; the fix was to delete the copy, not to `allow` it. + +**Replay now confines each decode to its own block.** Previously `decode` received the rest of the +file, so a corrupted `payload_len` was bounded only by the file's length and a record could claim +bytes belonging to its successors -- caught, but by the checksum happening to fail rather than +structurally. The stride is what makes a block boundary exist to confine it to. + +**The cost is reported, not described.** M25.1 asked for the write amplification to be "a real cost +to state rather than hide", and the first draft stated it as a ratio in a doc comment -- a +hand-maintained copy of a number the program can compute. The sample now measures and prints it +(`layout: N bytes of records in M bytes of file`) and the doc comment points at that line instead of +restating it. No conclusion is drawn about whether the ratio is acceptable, because that depends +entirely on a caller's record size. + +**A `const` assertion carries the sector rule**, verified load-bearing in both directions: a stride +of 4000 fails the build with `error[E0080]: evaluation panicked: RECORD_STRIDE must be a whole +number of sectors`, and 4096 builds clean. That is the build rung rather than a test, which matters +because the failure it prevents -- `ERROR_INVALID_PARAMETER` from a `NO_BUFFERING` write in M25.3 -- +would otherwise appear only on 4K-native storage, on somebody else's machine. + +**Four sabotages recorded, not one.** The writers' offset advances are separate facts at separate +sites: measured, reverting the harness lane leaves every appender test green and vice versa, so a +single case would have declared the pair covered while half of it was not. The full sweep is +16-of-16 as declared with the `CONTROL` still surviving. Note what the appender case does *not* +establish: packing its offsets makes records overlap inside a block, so two guards fire at once for +two different reasons -- the narrower evidence that the stride itself is what is caught is the +reader case, which moves only one number. + +**All three replay paths report the same numbers as before the change**, which is the check that the +layout moved and the contract did not: 24 durable records verified, 3 tail records tolerated, the +torn tail still stopping at `Truncated`, and the negative control still catching a corrupted byte. +The torn-tail simulation needed its arithmetic rewritten to keep meaning that -- it trimmed a fixed +count of bytes off the end of the file, which after striding lands in the final record's zero +padding and tears nothing at all. It now derives the cut from the last record's own block, which +also survives M25.3's pre-allocation. + +### M25.3 -- The log and every strategy file are pre-allocated and opened `NO_BUFFERING | OVERLAPPED`. *(completed 2026-09-24 12:45:57 -04:00)* + +Two steps in `logfile::create_preallocated`, neither interchangeable with the other: write the +extent with an ordinary handle and drop it, then reopen `OPEN_EXISTING` with both flags. This is +[the spike](design-sessions/spikes/write-pending-spike.rs)'s condition D, the only one of four it +measured as behaving differently from a buffered handle. + +**What this buys is an opportunity, not a guarantee, and nothing here claims otherwise.** Per the +standing constraint on M25, Windows specifies nothing about when a ring operation completes relative +to `SubmitIoRing`. The log is correct either way; what changes is whether a commit is separately +*measurable*, which is M25.4's problem and M25.5's to read. + +**The sweep found four live sites, and the item named two of them.** It predicted two statements in +[strategy.rs](examples/epoch_log/strategy.rs) reasoning from a synchronous handle. There were also +two in [main.rs](examples/epoch_log/main.rs) -- one an internal comment, one **printed to the user** +as part of the strategy comparison's own narrative. All four were corrected the same way: the +historical finding is preserved in the past tense, since it was true when written and is how M20.6 +reached its conclusion, and what follows is that the cause has been removed *without* asserting the +consequence. Whether the strategies are now distinguishable is not settled by changing a flag. + +**The third site the item named no longer exists, and that is the right outcome rather than a +miss.** `placement.rs`'s `volume_numa_node` documented that it could not use +`windows-overlapped-io-sys`'s typed `BlockingEndpoint::ioctl` partly because the handle was +synchronous. Earlier this session, `947b252b` moved that function into `win-numa-sys`, and the +comment went with the move -- correctly, because `win-numa-sys` depends only on `windows-sys` and +has no occasion to explain why it is not using a crate it does not reference. The item was written +before that move; its prediction of "a narrowed comment, not a refactor" was answered by the +comment's home changing. + +**The `NO_BUFFERING` alignment rule could not be tested the obvious way, and finding that out is +what produced the better test.** The first attempt wrote through [`std::io::Write`] and failed on +the *aligned* write: `write_all` issues `WriteFile` with a null `OVERLAPPED`, which an asynchronous +handle refuses however well-aligned the transfer is. That is now its own test -- it is the only +property of the handle's *mode* reachable from here, since `GetFileInformationByHandleEx` does not +report it and the ring works on synchronous and asynchronous handles alike. The alignment rule is +tested through a real ring instead, which is also how production reaches this handle, with both +directions asserted: an aligned write accepted and an unaligned one refused. Each test has a control +using an ordinary handle, so the refusals are attributable to the flags rather than to anything else +about the file. + +**A blind spot is recorded in [sabotage.json](sabotage.json) as a declared survivor rather than left +invisible.** Replacing the zero-fill with `set_len` is a **real regression that nothing here +detects**: both produce a file of the right size whose bytes read back as zero, because reads past +the valid data length are answered with zeros the filesystem synthesises without touching the disk. +Only the zero-fill advances that valid data length -- which is the thing that decides whether a +later write is extending. A `set_len` extent silently returns the log to the configuration measured +as behaving like a buffered handle. The only user-mode way to read a valid-data length back is +`FSCTL_QUERY_FILE_REGIONS`, and adding it to a sample purely to check a property the sample does not +otherwise use was judged machinery for its own sake. Recording it as `expect: "survives"` means a +future change that makes it observable will show up as a discrepancy in the sweep. + +> **Corrected 2026-09-24, the same day, after review challenged the claim rather than the code.** +> The paragraph above asserts from documentation that `set_len` is a regression, and the reasoning +> it gives is the wrong mechanism. Two measurements settled it: +> [2026-09-24-set-len-zero-fill-cost/](measurements/2026-09-24-set-len-zero-fill-cost/README.md) +> shows the zeroing cost is **identical** for a sequential writer, so that is not the reason; and +> [2026-09-24-set-len-vs-zero-fill/](measurements/2026-09-24-set-len-vs-zero-fill/README.md) shows +> the zero-filled extent pends at a median of 471/500 against `set_len`'s 268/500, which is. The +> blind spot is real and the conclusion survives; the argument for it did not. The second capture +> also corrects this entry's own framing of the spike, which repeated "only the pre-written extent +> pended" from a single run that does not replicate. + +**The harness caught a stale case in its own manifest, which is worth more than the case was.** The +`M25.1: the appender packs its record offsets` sabotage stopped compiling, because M25.1 ended by +removing the `total` binding its patch referenced in order to clear an unused-variable warning -- +so a sabotage that was correct when written was broken by a later edit in the same session. The +harness reported `MANIFEST DOES NOT COMPILE (tests never ran)` rather than scoring it as caught or +survived. **A manifest is a restatement site like any other**, and has to be swept when the code it +patches moves; nothing else would have noticed. + +**One assertion had to change because pre-allocation made it vacuous.** The strategy harness +compared `outcome.bytes` against the file's length to check its own accounting. A pre-allocated file +spans its whole extent from the moment it is created, whatever was written into it, so that equality +would have held just as well for a run that wrote nothing. It now checks the accounting against the +layout rule -- one block per record -- and separately that the extent covers what was written. + +**Deliberately left buffered, and said so at the definition.** The checkpoint file's records are +sixteen bytes from a `Vec` at offset 0, which satisfies none of `NO_BUFFERING`'s three alignment +rules; the retired segment is written once with `std::fs::write` and never goes through a ring. The +control plane's correctness comes from its covering flush, not from how its bytes are cached. + +**The log is pre-allocated with slack rather than to its exact record count.** A real write-ahead log +pre-allocates ahead of its writer, because an append that reaches the end of the extent becomes an +extending write again. It also means a clean log now ends in zeros rather than at EOF, so replay +stops with `NeverWritten` -- the path M25.2 taught it to tolerate, now actually exercised by the +sample rather than left for a reader of a real log to meet first. + +### M25.1b -- The sample's own verification now runs under `cargo test`, and `main` itself under CI. *(completed 2026-09-24 15:20:02 -04:00)* + +The item asked where this belonged: a CI job running the binary, or a `#[test]` calling the sample's +functions. **Both, because they carry different facts**, and the split follows the FAIL FAST ladder +rather than splitting the difference. + +**Everything the sample asserts is now a test.** `run_log` and `verify` are ordinary functions over +a generic `Report`, so `tests.rs` drives the real log against a real ring and a real file and then +runs the real verifier -- the same code path `main` takes. That is not a proxy for running the +sample; it *is* running it, minus the strategy comparison. It costs nothing measurable: the example +suite still finishes in well under a second. + +**The comparison's strongest assertion had no test, and now does.** `compare_strategies` requires +all three strategies to write **byte-identical** logs -- replay checks a log against itself, where +this checks the three against each other, so a dropped record or a wrong offset in any one shows up +as a difference from the other two. `M25.1`'s harness test runs `CoveringFlush` alone, so this was +the one assertion only `cargo run` could reach. `every_strategy_writes_the_same_log` runs all three +at two epochs of three records instead of thirty-two of sixty-four. The expectation survives the +ring count because a record's offset is its position in the global sequence times the stride, and a +strategy decides which *ring* submits a write, never where it lands. + +**CI runs the binary for what a test cannot reach**: `main` itself -- its path setup, its error +plumbing, its exit code. This is a published example a consumer runs, so a panic on startup is +exactly the failure worth catching, and it is the rung that fits because nothing smaller executes +`main`. Release, where the sample takes about a second. + +**The sabotage found a real hole, and it was in the thing this item exists to protect.** Reverting +`M25.2`'s torn-tail cut to the old file-length form was **survived** by the new end-to-end test. +With the extent pre-allocated, trimming a fixed count of bytes off the file lands in the slack, so +every record stays whole -- and `is_clean()` and the durable count both pass while the case +demonstrates the opposite of what it claims. A verifier that had quietly stopped verifying. + +What makes that worth recording is where the hazard already was: **written out in full, in the +comment directly above the cut**, since `M25.2`. Describing it caught nothing. Two assertions now +carry it -- that `tail_stopped` is `Truncated` specifically, not merely present, since a cut past +the last record reports `NeverWritten` and would mean the case had stopped tearing; and that a tail +record was actually lost. Prose is not a rung, stated once more by a file that had the prose. + +**What each guard is load-bearing for**, measured: + +| sabotage | end-to-end test | M25.1's layout tests | +|---|---|---| +| replay walks by extent (the `M25.1` defect) | red | red | +| torn cut taken from the file length | **red** | green | +| negative control moved out of the durable region | **red** | green | + +So the end-to-end test would have caught the defect that motivated the item, and it is the only +thing covering two failures that make the sample's evidence vacuous rather than wrong. Both are +recorded in [sabotage.json](sabotage.json). + +**What is deliberately still only in CI**: the strategy comparison end to end. Running it in a test +would multiply the suite's cost to re-check a layout and a replay that `M25.1` already covers, and +what the full run adds beyond the invariant above is a *measurement* -- which is not a contract this +crate may assert. + +### M25.4 -- The commit is measured as submit / blocking / deferral, so the flush's own cost and the deferral window cannot be confused again. *(completed 2026-09-24 16:32:56 -04:00)* + +`CommitTiming` replaces the single `Duration` the harness used to publish. The three parts sum to +that old number, and `flush()` is `submit + blocking` -- the deferral excluded, which is the whole +point. + +**The split satisfies M25's standing constraint by construction rather than by assumption.** Windows +specifies nothing about when a ring operation completes relative to `SubmitIoRing`, so the harness +must be meaningful either way: if the flush completes inline the device round trip lands in `submit` +and `blocking` is zero; if it pends, `submit` is short and the wait appears in `blocking`. Nothing +has to know which case it got. + +**The strategies were always distinguishable on the commit, and the blended number hid it +completely.** They now separate by roughly sixfold on the flush, where the old figure ranked them in +the opposite order -- alternating-rings reported the *worst* commit latency while being the fastest, +because its deferral is about twice the others'. That is not a new finding so much as `M20.6`'s +finding finally visible in the program's own output. Figures are not quoted here; the sample prints +them and `M25.5` is where they are read. + +**`blocking` reads zero for all three, and the output says plainly that this proves nothing.** A +zero beside a large deferral is ambiguous: the operation may have completed inline, or it may have +pended and finished while the program was busy elsewhere. Those are indistinguishable from here. +Recording that is the point -- reading `blocking` alone would be the same error as before in the +opposite direction, and the temptation is real now that `M25.3` has given the handle the shape the +spike measured as pending. + +**The settle logic became one function rather than two copies.** Both sites -- the loop's and the +drain after it -- applied the same rule about where deferral ends and blocking begins, and a rule +stated twice can be half-corrected. `settle` states it once. + +**A sabotage found the guard that the obvious assertions miss.** Asserting the identity +`flush() == submit + blocking` catches deferral being folded back in, and is **survived** by a part +that is never measured at all: replacing the deferral measurement with zero satisfies every identity +while making the decomposition a rename. The test therefore also requires some sample of `deferral` +and of `submit` to be non-zero -- phrased as "some sample" rather than a lower bound on a duration, +because this harness defers by construction, so a run in which nothing deferred means the clock is +not running rather than that the machine was fast. `blocking` deliberately gets no such guard, since +zero is a legitimate and frequently observed reading for it. + +Three cases in [sabotage.json](sabotage.json): the fold, and the two parts that can silently read +zero. They are listed separately because `submit` and `deferral` are measured at different sites and +one can be lost without the other. + +**What this does not do** is assert any value. Which part carries the cost is a property of the +machine and the handle, not of this crate. + +### M25.5 -- The comparison was re-run over fifteen runs and `M20.6` answered: no other ground found for `AlternatingRings`, and the accounting defect that would have inverted the reading was fixed first. *(completed 2026-09-24 17:19:59 -04:00)* + +The capture is [measurements/2026-09-24-commit-decomposed/](measurements/2026-09-24-commit-decomposed/README.md) +and the figures live there rather than here. + +**The measurement had a defect that had to be found before it could answer anything.** `M25.4`'s +first numbers showed `HostSequenced` committing roughly **six times cheaper** than the other two. +That was where the clock started: it waits for every write in userspace before pushing an unordered +flush, and the commit clock began at the *submit*, so its host round trip fell outside every +measured part. The cost had not gone anywhere; nothing was looking at it. + +A fourth part, `prepare`, now covers whatever a strategy must do before its flush can be pushed. +With it, `HostSequenced` is within noise of the others rather than six times cheaper. **A reader of +the uncorrected figures would have drawn the opposite of the right conclusion** -- which is the same +failure mode `M20.6` was opened to fix, one layer down, found by reading the very numbers the fix +for it produced. + +**What fifteen runs show.** The three are **not distinguishable** on throughput or on total commit +cost: medians within a few percent, every range overlapping every other, and run-to-run spread +within a single strategy larger than the spread across them -- which is the condition the sample's +own output tells a reader to check. What *is* structural, and never inverts across fifteen runs, is +where each spends its commit: `HostSequenced` in `prepare` and almost nothing in `submit`, the +covering strategies the reverse. That difference is invisible in any blended number, which is what +`M25.4` was for. + +**The open question, answered as far as this can answer it.** `AlternatingRings` shows no advantage +this harness can measure -- its throughput and commit medians sit inside the others' ranges, its +deferral is consistently about twice theirs, and the single worst commit p99 in the capture is its +outlier -- against a doubled arena registration it pays for the life of the run. + +**That is not a finding against the strategy**, and the entry says so where a reader will meet it. +The ground it was built on is blast radius, and `M20.6` established that this harness **cannot +exhibit that difference at all**, because each lane registers its own arena of the same size so the +per-ring bound is identical by construction. What this capture could answer is whether some *other* +ground appears, and none did. Whether that changes the strategy's status is the engineer's decision; +the conditions under which it would pay are already written in +[strategy.rs](examples/epoch_log/strategy.rs). + +**`block` reads zero at the median in every run**, and the capture records that this establishes +nothing: an operation that pended and finished during a deferral of several milliseconds is +indistinguishable from one that completed inline. + +**The harness caught stale manifest patches for the third time today.** Adding `prepare` changed +both `flush()` and the deferred tuple, invalidating two `M25.4` cases that had been correct when +written hours earlier. Each time the code moved, the manifest's patches stopped applying and only +`run-sabotage.ps1` noticed -- `MANIFEST STALE: pattern found 0 times`. The note that a manifest is a +restatement site like any other is now load-bearing three times over, which is enough to call it a +standing hazard rather than an incident. + +### M25.6 -- Swept what this milestone made false, and recorded the two findings as `D-56` and `D-57`. *(completed 2026-09-24 18:03:24 -04:00)* + +**Four named sites, and the sweep found two more.** The item listed the "keeps appending while a +commit is outstanding" rationale, `strategy.rs`'s "what the measurement found" section, the `M22.1` +capture's commit-p50 claim, and any DESIGN-NOTES text calling the sample's I/O buffered. Grepping +the falsified *claims* rather than the listed files also turned up the spike's premise that +`epoch_log` "writes variable-length records at packed offsets", and a second copy of the "two orders +of magnitude" mechanism inside the `M22.1` capture's `Settles` paragraph. As usual the reported +sites were a sample of the population. + +**One named site turned out not to exist.** No DESIGN-NOTES text describes the sample's I/O as +buffered. The three near-matches are about other things -- `D-40` is a cached *read* in the handover +tests, `D-49` is the unit suite's hermeticity, and the note that a ring handle "does not need +`FILE_FLAG_OVERLAPPED`" is a fact about rings that `M25` did not touch. Recorded because a sweep +that quietly finds nothing at a named site is indistinguishable from one that did not look. + +**The `M22.1` correction is the sharpest of them, and it is not that the number was wrong.** That +capture reported a commit-p50 reduction as "the one finding here that separates", beside a +throughput result reported as unmoved. The reduction is real and its stated mechanism is correct as +written -- batching shortened the interval between the last append and the flush being reached. What +is wrong is the label: `M20.6` established that figure was **entirely deferral**, so it measured the +*append path* getting faster. And that makes it not an independent finding at all. Both lines are +the same fact seen twice -- the appends got faster, the run is flush-bound, so the change appears in +the metric that is not flush-bound and not in the one that is. Reporting them as two results +overstates the evidence by exactly one result. + +**One figure is now measured and is not what it said.** The old explanation had the strategies +differing by amounts "two orders of magnitude below" the flush, "in the tens" of microseconds. +`M25.5` measured hundreds. The conclusion is unchanged, because they remain smaller than the +run-to-run spread -- but "below the noise" and "two orders of magnitude below the flush" are +different claims and only the first held, so both copies of the stronger one were corrected rather +than left standing beside a note. + +**`strategy.rs`'s top section was restructured rather than annotated.** It had accumulated a true +current claim, a superseded mechanism, and two correction sections underneath, so a reader met the +false explanation first and the correction several paragraphs later. It now states what is measured, +links the capture instead of quoting figures, and keeps both superseded explanations compactly below +under a heading that says they are superseded -- which is what CONTRACT INTEGRITY asks for and what +the file was violating. + +**The spike's prediction is left in the past tense rather than deleted**, because the reasoning is +the reusable part: it said that if only the pre-written condition pends, "the harness fix is not a +flag change -- it is a change to the log's on-disk format." That is exactly what happened, and it is +now `D-57`. + +**Two decisions recorded.** `D-56`: a benchmark that defers its await measures the deferral, and the +number survived three rounds of correction because every round re-read the conclusion instead of the +instrument -- with the generalisation that "the conclusion still holds" is not evidence that the +instrument does. `D-57`: a flag whose requirements reach into the caller's data layout is not a flag +change, and costing it as one underestimates it by the size of a format migration. + +### M25.7 -- Replay keeps its slice for a reason about failure vocabulary, recorded as `D-58`; the second multi-megabyte buffer became a digest. *(completed 2026-09-24 19:07:31 -04:00)* + +The item offered two acceptable answers -- stream the verifier, or stay legible -- and required only +that an 8 MiB `fs::read` not sit unremarked in a teaching sample. + +**The decision is to keep `replay(&[u8])`, and the reason is not simplicity.** The walk is strictly +forward one block at a time and never looks back, so it genuinely has no need of the whole file, and +a real log is larger than memory -- which makes reading the whole file the wrong reflex to teach at +exactly the point a reader is learning to verify one. That argument is real and it lost to a +stronger one. + +**`replay` returns an `Outcome`, not a `Result`.** Every way it can end is a statement about the +log: verified, tolerated, or a `Violation`. A streaming reader introduces a third kind of ending -- +`io::Error` -- into the one component whose entire job is to distinguish *the log broke its promise* +from *the log kept it*. Those want different responses from a caller, and a signature returning both +through one channel invites precisely the conflation this file exists to prevent: an unreadable file +reported as a missing durable record. **So the streaming version is a different interface, not a +smaller allocation** -- which is what the item suspected, and the suspicion is what turned out to +decide it. + +The cost of declining it is stated where it is paid rather than hidden: 140 KiB at the log's own +`fs::read`, 8 MiB per strategy at the harness's, each with a comment saying so. `D-58` records the +decision and what a consumer building a real verifier should want instead -- the streaming shape, +*with* the two failure kinds kept apart inside it. + +**What was reducible without touching that interface was reduced.** The cross-strategy comparison +held a whole reference log in memory for the length of the comparison, so two multi-megabyte buffers +were alive at once. It now keeps a 32-bit digest, which halves the peak and loses nothing a reader +had: the assertion could already only say *that* two logs differed, never where. + +**The digest is a weaker check than the byte comparison it replaced**, and the weakening is guarded +rather than assumed away. Two different logs can in principle share a digest where two byte arrays +cannot share their bytes, so `record/tests.rs` -- a test module `record.rs` did not previously have +-- pins that a flipped byte, a dropped record, and a trailing zeroed block each change it. The +sabotage confirms the separation is real: a digest folding only the length still distinguishes logs +of different sizes, so the two length-based tests stay green and only the flipped-byte one fails. +What none of them establish, and the definition says so, is that no two logs collide. + +**The two 64 KiB sites were left with a note rather than churned.** `RETIRED_LEN` is exactly 64 KiB +-- at the threshold this repository treats as the point to ask the question, not past it -- so both +the fill that writes it and the read that checks it are within the rule. The note says what a reader +growing that segment should do: the write has the same shape as `logfile`'s zero-fill, and the check +is a fold that never needs the bytes all at once. + +## Moved 2026-09-24 19:17:17 -04:00 -- M25: the epoch-log sample's I/O became a shape where a commit is observable + +The milestone's eight items are archived individually above; what follows is the context the +section carried, kept because it records what M20.6 found and the constraint every item was +held to. + +## M25 -- Make the epoch-log sample's I/O a shape where a commit is observable + +Queued by the `M20.6` investigation, which found three things the item did not anticipate. + +**The harness measures the wrong quantity.** Decomposing its commit latency into *deferral* (flush +pushed -> harness next looked) and *blocking* (time actually waiting) gave blocking p50 **and p99 of +0 us for all three strategies**. The published `commit p50/p99/max` column is entirely deferral: it +reports how long the next epoch's appends took, not anything about the commit. + +**There is no pipeline to measure.** The commit's `SubmitIoRing` took 289-555 us and returned with +all 9 completions already queued. The handle has no `FILE_FLAG_OVERLAPPED`, so the batch ran inline, +and the comment in [strategy.rs](examples/epoch_log/strategy.rs) reading "a real log keeps appending +while a commit is outstanding" describes something that cannot happen there. + +**`AlternatingRings`' blast-radius claim is answered structurally, and needs no run.** +`RegisteredBuffers::get_mut` refuses a slot with an operation outstanding and there are `SLOTS` +slots, so at most `SLOTS` appends are outstanding on a ring **by construction** -- and each +alternating lane registers its own arena of the same size. The per-ring bound is identical either +way. Measured at 8 and 8, but the argument does not rest on the measurement, and it holds whatever +the platform does about pending. + +[write-pending-spike.rs](design-sessions/spikes/write-pending-spike.rs) then established which +configurations pend at all. `FILE_FLAG_OVERLAPPED` alone changed nothing (0/500). Only +`NO_BUFFERING` over a **pre-written extent** pended reliably, and its submit p50 fell from ~500 us to +116 us -- the flush's cost leaving the submit path is what makes a commit separately observable for +the first time. + +> **Corrected 2026-09-24, after `M25.3` landed: the paragraph above overstates what replicates.** +> Sixteen runs with a fifth condition added are in +> [measurements/2026-09-24-set-len-vs-zero-fill/](measurements/2026-09-24-set-len-vs-zero-fill/README.md). +> What holds is that a **buffered** handle essentially never pends while every `NO_BUFFERING` one +> pends in most runs. What does not hold is "only the pre-written extent pended": the extending +> condition has a median of 268/500 over those runs. The zero-filled extent is still the best of +> the five -- median 471/500, floor 121 against 1 -- so `M25.3`'s choice stands, but as a +> difference of degree rather than of kind. The single-run reading came from a pair of numbers the +> spike's own header already warned was unstable. `M25.4` and `M25.5` must be read with that +> variance in mind rather than against the original framing. + +**A standing constraint on every item below.** That 500/500 is an observation, not a contract: +Windows specifies nothing about when a ring operation completes relative to `SubmitIoRing`. So the +sample may *adopt* this shape -- it is what real write-ahead logs do, and it is the only shape where +the measurement means anything -- but **nothing here may depend on an operation pending.** Every item +must leave the log correct if the platform completes inline tomorrow. + +- [x] **M25.1** -- Records gained a fixed sector stride with a zeroed block tail, in both writers. -> [completed 2026-09-23](COMPLETED-CHECKLIST.md#m251) + +- [x] **M25.2** -- Replay walks by the stride and confines each decode to its own block. Landed with `M25.1`: a strided writer and an unstrided reader cannot coexist. -> [completed 2026-09-23](COMPLETED-CHECKLIST.md#m251) + +- [x] **M25.1b** -- The sample's own verification now runs under `cargo test`, and `main` itself under CI. -> [completed 2026-09-24](COMPLETED-CHECKLIST.md#m251b) + +- [x] **M25.3** -- The log and every strategy file are pre-allocated and opened `NO_BUFFERING | OVERLAPPED`. -> [completed 2026-09-24](COMPLETED-CHECKLIST.md#m253) + +- [x] **M25.4** -- The commit is measured as submit / blocking / deferral, so the flush's own cost and the deferral window cannot be confused again. -> [completed 2026-09-24](COMPLETED-CHECKLIST.md#m254) + +- [x] **M25.5** -- The comparison was re-run over fifteen runs and `M20.6` answered: no other ground found for `AlternatingRings`, and the accounting defect that would have inverted the reading was fixed first. -> [completed 2026-09-24](COMPLETED-CHECKLIST.md#m255) + +- [x] **M25.6** -- Swept what this milestone made false, and recorded the two findings as `D-56` and `D-57`. -> [completed 2026-09-24](COMPLETED-CHECKLIST.md#m256) + +- [x] **M25.7** -- Replay keeps its slice for a reason about failure vocabulary, recorded as `D-58`; the second multi-megabyte buffer became a digest. -> [completed 2026-09-24](COMPLETED-CHECKLIST.md#m257) + +### M26.1 -- The permitted space is specified in [RESPONSE-SPACE.md](RESPONSE-SPACE.md) as eleven cited clauses, and recorded as `D-59`. *(completed 2026-09-24 19:30:19 -04:00)* + +**Eleven clauses: seven permissions and four constraints**, each with an ID, a source, and a +provenance tag. The IDs exist so `M26.3`'s resolver, `M26.4`'s properties and `M26.6`'s kernel tests +can cite a clause rather than restate it -- and so a clause no code cites is visible as +unimplemented. + +**The provenance tag is what makes it a specification rather than a recording.** Every clause is +`Observed`, `Over-provision`, or `Decided`, and where a clause is wider than its own observation the +two parts are split so they can be argued separately. `RS-P-1` is the clearest case: that an +operation may complete inline or pend is measured, but that operations within one batch resolve +**independently** is not, and the space permits it anyway -- because a consumer depending on them +resolving together depends on something Windows never promised. + +**The call the item demanded, made rather than defaulted: `RS-C-4` constrains the resolver to +honour the drain half of `DRAIN_PRECEDING_OPS`.** [D-47](DESIGN-NOTES.md#d-47) measured roughly +4,500 trials without a single violation; the drain is what this crate's durability story rests on; +and a resolver permitted to break it would require every consumer to re-verify durability some other +way, which is to say it would make the primitive useless. The cost is stated plainly: a Windows that +broke the drain would not be caught by the resolver at all. That is why `M26.6` gained a line +requiring at least one kernel test to exercise the clause -- otherwise the one constraint the space +takes on faith is untested in both halves at once. + +**The hold-back half stays unconstrained**, since `D-24` claimed it and `D-47` withdrew it. `RS-P-2` +applies in full to anything queued after a drained flush, which is the defect class that campaign +found. + +**Three citations were checked and one was wrong.** The draft attributed `M22.2`'s defect to the +checkpoint control plane; it was on the *append* path, and the checkpoint module merely documents +the same case. Both now appear, distinguished. The other two -- `pop_within`'s "promises nothing +about poppability" and `D-47`'s trial count -- were verified against the files rather than recalled. + +**A working artifact was found rather than assumed missing.** +[kernel-response-space-probe.rs](design-sessions/kernel-response-space-probe.rs) already contains a +seeded `Resolver` exercising `RS-P-2`, which broke a FIFO-assuming consumer under 189 of 200 seeds. +`M26.3` now points at it as a starting point, with the note that the probe marks itself throwaway -- +so promoting it is a deliberate decision rather than a default. + +**Four things are listed as deliberately undecided** -- rates, partial transfers, failure-code sets, +and timing -- so that a later reader can tell an omission from a choice. Rates in particular are +excluded on principle: a space carrying observed probabilities would be the recording this milestone +exists to avoid. + +## Moved 2026-09-24 20:36:46 -04:00 -- M26.2: the kernel-call seam + +### M26.2 -- Build the seam that makes the `windows-sys` calls indirect, so `M26.3`'s resolver can answer them. *(completed 2026-09-24 20:36:46 -04:00)* + +The shape is recorded as [D-60](DESIGN-NOTES.md#d-60); what follows is what the work found. + +Eight calls became indirect -- `SubmitIoRing`, `PopIoRingCompletion`, and the six `Build*` entry +points the crate uses -- across 27 call sites in [batch.rs](src/batch.rs) and [ring.rs](src/ring.rs). +Each now goes through a `pub(crate) unsafe fn` in [sys.rs](src/sys.rs) that is `#[inline(always)]` +and dispatches through a `through_seam!` macro. With the `kernel-seam` feature off, the macro +expands to the bare FFI call and nothing else; with it on, the call first asks the installed +responder. + +**The shape was decided by a constraint already on the books, not by taste.** The obvious +alternative -- parameterising the ring as `IoRing` over a kernel -- is unavailable because +[D-55](DESIGN-NOTES.md#d-55) has already spent `IoRing`'s type parameter on `M28.3`'s token +inventory. A kernel generic would publish `IoRing`, which is a two-parameter public type on a +shipped crate, and the second parameter exists only so the crate can test itself. Module +indirection costs the public surface nothing. + +**The responder is thread-local, for the reason `DROP_RUNS` is.** `cargo test` runs tests as +threads in one process ([DESIGN-NOTES.md](DESIGN-NOTES.md) records this as the reason this +workspace is not on nextest), so a process-global responder would let one test answer another +test's kernel calls. `with()` uses `try_borrow_mut` rather than `borrow_mut`, so a re-entrant call +from inside a responder falls through to the kernel instead of panicking; `Installed::drop` uses +`try_with`, so teardown during TLS destruction cannot abort the process ([M23.4](#m234)). + +**Five lifecycle calls were deliberately left direct** -- `CreateIoRing`, `CloseIoRing`, +`GetIoRingInfo`, `IsIoRingOpSupported`, `SetIoRingCompletionEvent`. `M26` is justified by the +*response space*: what the kernel may answer to submitted work. Routing ring construction and +teardown through the seam as well would be hermeticity for its own sake, and hermeticity is +[M24](#m242)'s subject, not this one. The line is recorded so a later reader can tell a boundary +from an oversight. + +**The seam's transparency is measured, not argued.** The full suite passes with the feature off +(153 lib tests) and on (159 -- the six new ones), and the crate builds clean with zero warnings in +five configurations: default, `--all-features`, `--no-default-features`, `--features kernel-seam`, +and release. The sabotage case *`M26.2: the seam consults the responder but ignores its answer`* +turns `an_installed_responder_answers_instead_of_the_kernel` red, which is what shows the +consultation is load-bearing rather than decorative -- a seam that asks and discards would pass +every other test in the crate. The full sweep is 28-of-28 as declared, with the `CONTROL` and the +`set_len` blind spot both still surviving. + +**Two type signatures were wrong on the first attempt and the compiler caught both**, which is +worth recording because they are the kind of thing a hand-written trait gets wrong silently if it +is ever allowed to diverge: `BuildIoRingWriteFile`'s caching flag is `i32`, not `u32`, and +`BuildIoRingRegisterFileHandles` takes `*const *mut c_void`, not `*const isize`. The `real` +submodule re-exports the `windows-sys` items so the trait's default methods call them directly, +which is what keeps the two in step -- there is one spelling of each signature, not two. + +**`kernel-seam` crossed with `--no-default-features` is a published configuration nothing built.** +The gate multiplies with `threadpool`: the workspace `--all-features` steps build the seam only +alongside the threadpool, and the existing `ioring-no-threadpool` job built the no-threadpool path +only with the seam off. Two steps were added to that job rather than a new job, on the same +argument that bought the job in the first place. Verified locally before committing: clippy clean +and 155 lib tests green in that combination. + +**A borrow-surface row was owed and was three items late.** `./tools/check-borrow-surface.ps1` +failed on `Pending::contract -> Option<&RingContract>`, added by `M23.3`'s spike, because that +change did not run the gate -- so it reported on the next run instead, against unrelated work. The +row is now in [DESIGN-NOTES.md](DESIGN-NOTES.md)'s audit table: `RingContract` is a pure +observation record owning no handle, buffer, or registration index, so there is nothing the kernel +could invalidate, and the borrow is a plain `&self` borrow that blocks submission through that +`Pending` for its duration. The lateness is recorded in the row itself. + +**A tooling mistake destroyed two source files and is worth the warning.** A PowerShell +`.Replace()` bound the wrong overload and rewrote [batch.rs](src/batch.rs) and [ring.rs](src/ring.rs) +one character per line. `git checkout --` recovered both, and the conversion was redone with +`[regex]::Replace` anchored on `(?M26.3 -- Build the resolver over the space `M26.1` specifies, bound to its clause IDs and seeded on its own axis. *(completed 2026-09-24 21:15:16 -04:00)* + +The shape is recorded as [D-61](DESIGN-NOTES.md#d-61); what follows is what the work found. + +**The resolver is in [resolver.rs](src/sys/resolver.rs), implementing `M26.2`'s `Responses`.** It +answers the eight submission-path calls itself, so operations the kernel never received still flow +through this crate's ordinary accounting. Every freedom cites the `RS-P-n` permitting it and every +restriction cites the `RS-C-n` requiring it, which is what lets a reader check the resolver against +[RESPONSE-SPACE.md](RESPONSE-SPACE.md) mechanically rather than by reading both and hoping. + +**The asymmetry is the design.** `ResolverConfig` has a switch per permission and none for any +constraint. Narrowing a freedom is how a test isolates another -- a test about ordering does not +want arbitrary operation failures on top -- while a knob relaxing a constraint would let a test +assert against a platform that cannot exist. The default is the widest point, so a test that does +not choose gets every freedom and fails loudly under one it did not handle. A test asserts the +count: seven fields, seven `RS-P-n`, and it reads them off `Debug` so a field added without a clause +fails there rather than passing unnoticed. + +**Where a permission and a constraint collide, the constraint wins, and that had to be decided +rather than discovered.** `RS-C-4` holds a drain-flagged operation back even on a tick where +`RS-P-1`'s coin said complete it now; `RS-C-1` forces a post an operation's coin kept deferring. +Two bounds exist solely to make `RS-C-1` finite -- per-operation deferrals, and consecutive declined +submits -- and both are properties of the resolver rather than of the space, which carries no rates +deliberately. + +**A submit resolves; a pop only rescues.** The split is not tidiness. A pop that flipped coins would +resolve a consumer's work on its first `try_pop`, so "the operation pended" -- the thing `RS-P-1` +exists to let a test observe -- would be unobservable to exactly the consumer most likely to care. +The freedom would have been implemented and untestable. + +**`SetIoRingCompletionEvent` moved behind the seam, which `M26.2` had left it outside of.** It looks +like lifecycle and is not: it is how a completion becomes *observable*, so `RS-P-6` is a clause +about that call. A resolver unable to make it would not satisfy `RS-P-6` vacuously -- it would never +signal at all, parking every [`EventDelivery`](src/event_delivery.rs) consumer rather than testing +one. The checklist item had authorised exactly this ("if a clause turns out to need ... extending +the seam is part of this item"), and this is the clause that needed it. + +**`RS-C-4` is decided by position, and the invariant that makes that sound is now asserted.** The +pool is held in build order, so the set queued before `pool[i]` is exactly `pool[..i]` and a barrier +is eligible only as the oldest unresolved operation. That reduction is the whole of the constraint's +implementation and it holds only while the pool stays sorted -- an edit that sorted or reshuffled it +would relax `RS-C-4` to nothing while every line around it still read as though it applied. A +`debug_assert!` now says so at the point of use, rather than a comment saying so nearby. + +**The first contact with a real ring found a live defect, queued as `M26.8` rather than fixed +here.** A submit declined under `RS-P-7` propagates out of `IoRing::run_down` as an error with +`outstanding() > 0`, after which `Drop` asserts and calls `CloseIoRing` anyway -- `M21.6`'s hazard, +reachable again through a different `HRESULT`. Measured by narrowing one permission at a time: 32 +seeds pass with `may_fail_submits` off, seed `0x1A` fails at `0x80070008` with it on. It is queued +rather than corrected because `run_down`'s own documentation argues that blocking is the safe +failure mode while "no hang" is one of the properties `M26.4` is about to write, and the two pull +opposite ways -- a decision, not a correction. The narrowing is declared in the test that takes it, +and the current behaviour is pinned by its own test so that whichever way `M26.8` is settled, a test +has to change. + +**The integration test is where it is for the gate's own reason.** `check-ring-tests.ps1` asks +whether a test needs the kernel or only a ring-shaped thing; a resolver test needs a real ring +because [D-60](DESIGN-NOTES.md#d-60) deliberately left lifecycle real, which makes it an +operating-system boundary and therefore `tests/`. The ring-opening lib population is unchanged at +41. + +**Sabotage found a defect in the tests, which is what it is for.** `RS-C-3`'s case survived: the +test asserted that nothing pops *right now*, which a resolver that had wrongly made staged +operations eligible also satisfies, because the starvation rescue posts only what has run out of +deferrals and a fresh operation has not. The test now polls past the bound and asserts the counters, +so it checks the claim -- unsubmitted work is never eligible -- rather than the symptom. Eight cases +added, one of them a declared blind spot; the full sweep is 36-of-36 as declared with both blind +spots and the `CONTROL` surviving. + +**The harness caught stale patches a sixth time, and the cause was new.** Four cases reported +`pattern found 0 times` while every line of each pattern was present individually. The cause is that +the built-in file-creation tool writes **CRLF** on Windows, so the multi-line patterns could not +match a file whose line endings were not LF -- and single-line patterns matched fine, which is why +three of the eight cases passed and hid it. The repository's own instructions warn about this tool; +`M26.2`'s files escaped it only because git normalised them on commit before that sweep ran. The +three new files were converted to LF before proceeding. + +## Moved 2026-09-24 21:48:36 -04:00 -- M26.4: the properties that must hold under every resolution + +### M26.4 -- Write the properties that must hold under every resolution, with `RingContract` as the definition rather than a second copy. *(completed 2026-09-24 21:48:36 -04:00)* + +The shape is recorded as [D-62](DESIGN-NOTES.md#d-62); what follows is what the work found. + +**Five properties, in +[properties_under_every_resolution.rs](tests/properties_under_every_resolution.rs).** Conservation +(P-1), no hang (P-2), `pop_within` honours its bound (P-3), `outstanding` is accurate (P-4), no +use-after-free (P-5) -- driven over generated plans, each under its own resolution drawn from +`M26.3`'s resolver at the widest point in the space. + +**Only two of the five needed anything new, and that is the item's main point.** P-1 is already +this crate's own oracle, so the harness reports to [`RingContract`](src/contract.rs) and asks it for +the verdict. P-4 follows the same rule rather than counting for itself: the expected outstanding +count is read back out of the contract through its own `Outstanding` violation, because a counter in +the harness would be a third party to the disagreement and, when the two disagreed, the harness is +what would get "fixed". P-5 is [`windows_guard_alloc::GuardAlloc`], already the established +detector. + +**Two weaknesses are declared in the file rather than papered over.** `pop_within`'s upper bound is +nearly free under an ordinary resolution, because the resolver answers a wait immediately and the +call rarely approaches its deadline -- so the non-vacuous case needs a resolution in which *nothing +completes during the window*. That is supplied by a separate degenerate responder, and what it is +matters: not an `RS-C-1` violation, since no finite observation can distinguish "eventually" from +"never", but the prefix of a satisfying resolution in which the eventually has not happened yet. +And P-5 covers this crate's memory handling rather than the kernel's, since under a resolver nothing +external writes into a buffer at all; the kernel-side half stays with +[generated_sequences.rs](tests/generated_sequences.rs), against a real ring. + +**A third seed axis, kept separate.** This file carries the plan seed, the resolver seed and the +guard allocator's. They are independent on purpose -- pinning the plan alone reproduces the same +operations against different resolutions, which is what a suspicious plan calls for -- and a failure +prints all three, because only all three replay the whole run. + +**The harness's own first defect was treating a declined submit as a property failure.** Under +`RS-P-7` a submit may be declined, and `pop_within` surfaces that as an `Err` -- which is *within +its documented contract*, since it says it returns any error from `SubmitIoRing`. A correct consumer +retries; a harness that panicked was asserting a contract the crate never offered. It now retries, +and recognises a refusal by **asking the resolver** whether its decline counter moved rather than by +matching an `HRESULT`, because a hard-coded code here would be a second copy of a choice the +resolver owns. Retrying is bounded by P-2's budget in every caller, so a resolution that declined +forever is still caught. + +**That finding extended `M26.8` rather than creating a second item, and the extension is about the +document.** `RS-P-7` is written as a *consequence* clause -- "if `SubmitIoRing` fails, operations +already built remain queued" -- citing `D-5`, which establishes the no-rewind consequence and +nothing about submits failing spontaneously. `M26.3`'s resolver read it as a permission, and the +space nowhere states that a submit may fail at all. That is a gap rather than a decision, since +submits demonstrably can fail, so `M26.8` now has to settle both halves together. + +**Coverage counters are asserted, not printed, and that is what makes a green run evidence.** All +five properties are satisfied trivially by a run that does nothing: an empty plan, a resolution that +completes everything inside its submit, or a harness that quietly stopped reporting would each pass +every assertion. The counters are what separate "the properties held" from "nothing reached the +states they are about", and one sabotage exists purely to show they are load-bearing. Their +thresholds were set from a measured spread across eight fresh seeds rather than guessed; the +declined-submit count is the thinnest signal and its threshold is only "more than none" for that +reason, since tightening it would buy a flake rather than a guarantee. + +**The integration test is where it is for the gate's own reason**, as in `M26.3`: it needs a real +ring, because `D-60` deliberately left lifecycle real. The ring-opening lib population is unchanged +at 41. + +**A sabotage case was written, measured, and removed as unsound.** It disabled the P-1 verdict +check, and it survived -- correctly, because disabling an assertion that does not fire on a green +baseline cannot fail anything. What actually establishes that the verdict is read is the +identity-reuse case: measured, the suite fails carrying the oracle's own wording, so the path from +the crate through `RingContract` to a red test is traversed end to end. That evidence now lives in +that case's reasoning, and the removal is recorded here because "sabotage a check" is an appealing +and empty move worth recognising next time. Three cases added, all scoped to this test target on +purpose -- these mutations are caught by many tests in the crate, and "something went red" would not +have shown that *these* properties are the ones watching. Full sweep 39-of-39 as declared. + +## Moved 2026-09-24 22:13:02 -04:00 -- M26.5: calibrating the resolver + +### M26.5 -- Re-inject the two historical defects and confirm the resolver turns red. *(completed 2026-09-24 22:13:02 -04:00)* + +The shape is recorded as [D-63](DESIGN-NOTES.md#d-63); what follows is what the work found. + +**The two defects are calibrated differently because they live in different places**, and noticing +that was most of the item. `M21.6`'s -- an expired wait reported as a failure -- was in this crate, +so re-injecting it means mutating the crate, which a test cannot do; it is a case in +[sabotage.json](sabotage.json). [D-47](DESIGN-NOTES.md#d-47)'s -- a consumer believing a covering +flush holds back what follows -- is in a **consumer**, so there is nothing here to mutate and the +defective consumer had to be written out. It is, in +[calibration.rs](tests/calibration.rs), as a `HoldBackBeliever` that records when its assumption +fails; the test asserts that the resolver breaks it. + +**What the calibration file adds on the `M21.6` side is the precondition, not the detection.** A +sabotage of code the suite never executes is caught for some unrelated reason or not at all, and +either way measures nothing -- so there is a test asserting that `RS-P-4` actually reaches +`pop_within`, read off the resolver's own counter rather than inferred from a timing, because "the +call took a while" is not evidence that a wait expired. + +**Both directions were verified by execution before being written into the manifest.** Reverting +`wait_outcome`'s `WAIT_EXPIRED` arm fails the calibration naming the seed and `0x800705B4`. Making +the resolver enforce the hold-back fails it with "no seed of 64 broke a consumer that assumes a +covering flush holds back what follows it". + +**The second of those runs the opposite way to every other sabotage in this crate**, and the +manifest says so: it does not break the code, it makes the **instrument** go narrow. That is the +`M17.4` failure mode -- a suite sampling the right state while being insensitive to the defect +living in it -- and nothing detects it except a test that demands the sensitivity. + +**The two suites' sensitivities were measured and are not equivalent.** `M21.6`'s defect is caught +by the calibration *and* by `M26.4`'s property suite; the latter only because that suite +distinguishes a declined submit from a genuine error, so the detection is deliberate rather than +lucky. A narrowed resolver is caught by the calibration **alone**, and the property suite correctly +stays green -- a narrower resolution is still a valid one and conservation holds under it. That +measured gap is the argument for a calibration file rather than a calibration assertion bolted onto +the property suite, and it is an argument from data rather than from taste. + +**The D-47 calibration breaks its believer on every seed, which needs saying rather than +celebrating.** `D-47` measured the overtake on real hardware at well under one trial in a hundred; +the resolver does it constantly. That is not the resolver being unfaithful: +[RESPONSE-SPACE.md](RESPONSE-SPACE.md) carries no rates deliberately, because a resolver +reproducing an observed frequency would be a model of Windows and therefore the trap +[D-52](DESIGN-NOTES.md#d-52) was opened to escape. A defect class that is rare on hardware is +precisely the one a rate-free resolver earns its keep on. The figures are printed by the tests so a +reader can judge the instrument's strength rather than take "more than none" on trust. + +**Seeds here are fixed rather than clock-derived**, unlike the generated suites. A calibration that +could sometimes fail to demonstrate its own sensitivity would be the exact failure it exists to +prevent, arriving as a flake. + +## Moved 2026-09-24 23:54:45 -04:00 -- M26.6: the kernel tests' new job + +### M26.6 -- Point the kernel tests at confirming reality stays inside the declared space. *(completed 2026-09-24 23:54:45 -04:00)* + +The shape is recorded as [D-65](DESIGN-NOTES.md#d-65); what follows is what the work found. + +**A hole was found exactly where the split is load-bearing.** +[flush_barrier.rs](tests/flush_barrier.rs) already asserted `RS-C-4` against a real ring -- but the +assertion sat behind an early return taken whenever its control could not discriminate. On such a +machine the constraint was untested on **both** sides at once: the resolver is forbidden to produce +a violation, and the only test that would notice had skipped. That is precisely the condition the +item was written to prevent, and it was live. + +**The fix separates two questions one gate had been answering together.** *Is the barrier doing +work* is a comparative claim and genuinely meaningless when the control shows no reordering -- the +covering case would match a control that did nothing. *Did the kernel stay inside `RS-C-4`* is a +conformance question, where skipping can only ever hide a violation; observing none is weak +evidence when nothing could have reordered, but observing one is a finding on any machine. The +conformance assertion now runs everywhere and only the comparative claim is withheld, with the +output saying `PARTIAL` rather than `SKIP` so the difference is visible in a log. + +**The census was green and useless on its first build, and only trying to make it go red found +that.** It searched each file for the clause ID anywhere in its text. Removing `RS-C-4`'s check +from the only test performing it did **not** turn it red, because +[generated_sequences.rs](tests/generated_sequences.rs) mentioned that clause solely to *disclaim* +it -- "`RS-C-4` is flush_barrier's" -- and under a substring search a disclaimer reads exactly like +a claim. This repository had already recorded that trap once, for a probe whose only mention of a +tag was a comment, and it was walked into again anyway. + +**So a claim is now a structured marker.** `CONFIRMS:` for a constraint checked against a real +kernel, `EXERCISES:` for a permission the resolver takes, each naming a clause and nothing else -- +a hedged `CONFIRMS: RS-C-4 eventually` does not parse as a claim, and that case is asserted. The +rebuilt census was verified to go red in three directions before being trusted: a removed marker, a +marker naming a clause the document does not declare, and a prose mention standing in for a marker. + +**Two blind spots are declared rather than left invisible.** A census over source proves a clause is +*claimed*, never that the file's assertion still runs or still means anything -- comparing a +constant against itself leaves the marker in place and the suite green. And dropping the covering +flag does **not** fire the `RS-C-4` assertion on every machine: measured here, a device stack that +orders a flush behind its file's outstanding writes by itself produces no violation to see, so that +sabotage demonstrates nothing portable. The assertion's conformance value -- reporting a violation +if one occurs -- is separate from its sensitivity, and only the first is claimed. + +**The sweep found a stale site from two milestones back.** +[D-60](DESIGN-NOTES.md#d-60) still listed `SetIoRingCompletionEvent` among the five lifecycle calls +left outside the seam, which `M26.3` had moved behind it. The correction is recorded in that row +rather than only in `D-61`, since a reader arriving there would otherwise take the superseded list +as current. + +**The "five techniques" framing becomes six, and the sixth is different in kind.** The other five +check this crate against its own stated contract and cannot tell you that contract is wrong -- which +is what happened in the two most expensive defects. The resolver checks the crate against a written +specification of what the platform may do, and the kernel tests check the platform against that same +specification. `M26`'s row also joins defect population A, on the observation that the kernel's +*response* is a precondition and was never varied either. + +## Moved 2026-09-25 10:48:58 -04:00 -- M26.7: auditing the suite for frozen observations + +### M26.7 -- Audit the existing suite for assertions that are frozen observations rather than contracts. *(completed 2026-09-25 10:48:58 -04:00)* + +The shape is recorded as [D-66](DESIGN-NOTES.md#d-66); what follows is what the work found. + +**The census came from a command, and it was not the file anyone guessed.** The item named +[flush_barrier.rs](tests/flush_barrier.rs) as the obvious candidate and said plainly that the +candidate was a guess. It was: that file turned out to be one of the better-behaved ones, since its +transfer assertion is explicitly framed as its own precondition and its ordering counter is +*reported* rather than asserted. The census started from the whole assertion population, narrowed to +assertions in tests that touch a real ring -- reusing `RING-OPENING-LIB-TESTS.txt` as the classifier +rather than inventing a second one -- and then to four candidate shapes. + +**One class, 31 assertions wide, across five files.** `try_pop()` straight after `submit_and_wait` +with the `Option` unwrapped, asserting the kernel had **already** queued the completion. That is +`D-52`'s demonstrated failure exactly -- an assertion that gives opposite answers on two handles of +the same API -- and this crate's own documentation denies it: a submit-side wait's return "promises +nothing about poppability". `RESPONSE-SPACE.md` states it as `RS-P-5`. They passed for the reason +[D-40](DESIGN-NOTES.md#d-40) measured: a buffered read completes inside the submit in 80 of 80 +attempts, while an unbuffered one genuinely pends. + +**Restated as the contract `D-52` prescribed** -- ask for the completion within a bound this crate +chooses -- which holds on every handle rather than on the one a test happens to open. The same sweep +found an unbounded `try_pop` spin loop in `registration.rs`, which is the shape `pop_within` was +introduced to replace, and it went the same way. + +**Nothing catches a regression by running, and that is the part worth keeping.** Reverting a site +leaves its test green on any machine where the observation is true, which is precisely why 31 of +them survived years of review. So the guard is a **census that refuses the shape at the source**, +plus a resolver-driven test that makes the pending case reachable on demand and shows `try_pop` +failing where `pop_within` succeeds. One of the three sabotages exists to make that point: reverting +a restated site is caught by the census and by nothing else, and a reader who checked by re-running +the test would find it green and conclude the change was cosmetic. + +**A second, narrower class was found and deliberately not settled.** Five assertions require a +complete transfer. `Completion::result` promises only "the transferred byte count", and the space +lists partial transfers as *deliberately undecided* with the instruction to decide them before a +resolver relies on either answer. So the suite silently answers a question the specification leaves +open -- and the two are not yet in conflict only because the resolver reports `Information: 0` and +no resolver-driven test reads a transfer count. Queued as `M26.10`. One of the five is already in +the honest form and is left alone. + +**What the audit did not find is worth stating too.** The `outstanding() == 0` assertions after a +rundown are the crate's own contract, not observations. `flush_barrier_stress.rs`'s assertion that +reordering *happened* is a control verifying its own precondition, with a message explaining why -- +the correct pattern. And the handover tests rest on `D-40`'s synchronous-completion measurement but +verify that precondition rather than assuming it, which is what a frozen observation fails to do. + +## Moved 2026-09-25 12:34:02 -04:00 -- M26.8: retry policy belongs to the caller + +### M26.8 -- Decide what `IoRing::run_down` should do when a submit it makes is refused. *(completed 2026-09-25 12:34:02 -04:00)* + +The shape is recorded as [D-67](DESIGN-NOTES.md#d-67); what follows is what the work found. + +**The item was framed as a decision and was mostly a reading.** It asked which way to resolve a +tension between "blocking is the safe failure mode" and "no hang". The tension was real but the +question had an authority nobody had consulted: `SubmitIoRing`'s reference page. The correction that +made the difference came from the engineer -- **no amount of measurement constitutes a contract** -- +and it was needed, because the session had spent the previous hour measuring and had drawn a +confident conclusion from it. + +**What the documentation settles.** A return-value row gives `IORING_E_WAIT_TIMEOUT` the meaning +*"All operations were submitted without error and the subsequent wait timed out"*, and the Remarks +add *"If this function returns an error other than IORING_E_WAIT_TIMEOUT, then all entries remain in +the submission queue"* and that a per-entry failure arrives as a completion rather than as a submit +failure. + +**A measurement had been read backwards, and the documentation is what caught it.** A probe showed +an operation completing after a submit reported `E_INVALIDARG`, which was taken to mean the work had +gone in despite the error. It had not: the entry stayed in the submission queue exactly as +documented, and what pushed it through was `pop_within`'s *own* internal submit. The observation was +right and the attribution was wrong -- which is precisely the failure mode a contract prevents and a +measurement invites. + +**`Batch::submit_and_wait` carried `M21.6`'s defect** at the one site that sweep did not reach. +`do_submit` passed a timed-out wait to `check`, so a fully successful submission was reported as an +error. The damage is worse than a wrong sign: because any *other* error means the entries are still +queued, an `Err` was ambiguous between "your buffers are free" and "the kernel still owns them", +which is [D-5](DESIGN-NOTES.md#d-5)'s hazard with the sign hidden. + +**`run_down` was the only waiting API in this crate shaped wrongly**, and the engineer named the +principle that identifies it: *never implement a retry policy ourselves -- the caller keeps their +own backoff, counts before giving up, and whatever else they want.* Waiting in segments inside a +period the caller supplied is fine; `run_down` had the segments and no such period, which made "how +long to keep trying" this crate's policy. [`IoRing::run_down_within`](src/ring.rs) is the primitive +the caller bounds; `run_down` is now that with an unbounded period, which is a choice made by +calling it. Checked against the rest of the surface rather than assumed: `pop_within` and +`submit_and_wait` were already correctly shaped, so `run_down` really was the sole outlier. + +**An error from rundown is no longer terminal, and did not need to be.** The entries remain queued, +the ring is resumable, and the documentation is what makes that safe to say. The one thing a caller +must not do after an error is drop the ring. + +**`RS-P-7` and `RS-P-3` moved from inference to citation.** `M26.1` wrote `RS-P-7` as a consequence +clause -- "if a submit fails, entries remain queued" -- citing a decision that established the +consequence and nothing about submits failing at all, and `M26.3`'s resolver read it as a permission +regardless. The space now carries a **`Documented`** tag that explicitly outranks `Observed`, since +a measurement describes one run of one build. + +**The manifest drifted and the harness caught it, for the seventh time this session.** Renaming +`WAIT_EXPIRED` to `IORING_E_WAIT_TIMEOUT` -- so the constant carries the documented name rather than +a derivation -- left `M26.5`'s sabotage patching text that no longer existed. The sweep that +followed found four more live sites; the archive was left alone. A per-case uniqueness check now +runs before the sweep, which would have caught this in seconds rather than in a five-minute run. + +## Moved 2026-09-25 14:19:35 -04:00 -- M26.9, the event_delivery stall + +### M26.9 -- The `event_delivery` stall: the wait was armed after the event was signalled, which `SetThreadpoolWait` forbids. *(completed 2026-09-25 14:19:35 -04:00)* + +Cause, fix and before/after figures are in [DESIGN-NOTES.md](DESIGN-NOTES.md) -> D-68; the +investigation as it stood when the cause was found is in +[RESOLVED-TEST-FAILURES.md](RESOLVED-TEST-FAILURES.md). The item as it read when it closed: + +- [x] **M26.9** -- Find and fix the intermittent `Timeout` in + [event_delivery.rs](tests/event_delivery.rs)'s two threadpool-delivery tests, recorded in + [UNRESOLVED-TEST-FAILURES.md](UNRESOLVED-TEST-FAILURES.md). + + **Why this is not merely a flaky test to re-run.** Measured at 1 failure in 80 runs of the + compiled binary. The sabotage harness runs the whole suite once per case and the manifest holds + 41, so that rate gives roughly a **40% chance of a corrupted sweep** -- and the corruption falsely + reports `caught`, which is a sabotage the suite did not catch being recorded as a clean bill of + health. The harness is this repository's mechanism for keeping earlier guarantees checked; a 40% + chance of a silent false pass undermines every conclusion drawn from it. + + **What has already been ruled out, so it is not re-tried:** ring-resource pressure (zero failures + after roughly 18,000 ring create/close cycles), the widened seeded sweeps (zero after repeated + property-suite and calibration runs), and CPU starvation (zero under a concurrent `cargo build` + saturating the machine). + + **Narrowed 2026-09-25 to a minimal reproducer, and the investigation now leaves this crate.** + Measured: 0 failures in 1000 serial runs against 7 in 1000 parallel; every occurrence identical, + with **both** delivery tests failing together and `callbacks run: 0` -- the pool never invokes the + callback at all, for either ring, and nothing arrives ten seconds later. The two delivery tests + alone do not reproduce it (0 in 1000); a third test has to be co-running, and the two that trigger + it both create an `EventDelivery` over a ring with nothing outstanding and drop it promptly. + Full figures, the reproducer, and what remains unestablished are in + [UNRESOLVED-TEST-FAILURES.md](UNRESOLVED-TEST-FAILURES.md). + + **The next step is in [`windows-threadpool-sys`](../windows-threadpool-sys/src/wait.rs)**, not + here: both waits are registered on the default process threadpool, and one object's lifecycle + appears to stop other, unrelated armed waits from ever firing. Per the repository's mono-repo bug + policy, the fix belongs in that layer. `ThreadpoolWait`'s `Drop` has been read and only touches + its own object, so the mechanism is **not yet established** -- do not start from a guess about it. + + **Narrowed further the same day, with a configurable trace.** The default pool is **not** wedged: + a probe at the moment of failure runs a plain work item and a brand-new armed wait, and measured, + both ran. The stalled waits were created and armed -- the trace shows it -- and the trampoline + never fires for any of them. **The stall is permanent by design**: the completion event is edge + triggered ([D-19](DESIGN-NOTES.md#d-19)), a stalled ring's queue never returns to empty, so the + setup signal is the only wakeup that ring will ever receive and losing it once ends delivery for + good. The open question is now narrow -- why an armed wait does not observe a signal raised just + before it was armed -- and a plausible mechanism is recorded in + [UNRESOLVED-TEST-FAILURES.md](UNRESOLVED-TEST-FAILURES.md) **as a hypothesis with no evidence + behind it**, together with the experiment that would settle it. + + **The trace is compiled out unless `--features trace` is on**, and narrowed at run time by + `WINDOWS_THREADPOOL_TRACE`, because the instrument for a timing-dependent fault must not change + the schedule it measures. The flake still reproduces with it on, which was checked first. + + **A cheaper interim mitigation exists and is a separate decision:** the harness could treat a + failure in these two tests as *inconclusive* rather than as `caught`, which would stop the false + clean bills without pretending the behaviour is understood. + +## Moved 2026-09-25 15:02:00 -04:00 -- M26.10, the partial-transfer decision + +### M26.10 -- Decided: a completion may report fewer bytes than requested, so `RS-P-8` permits it and the caller owns the remainder. *(completed 2026-09-25 15:02:00 -04:00)* + +The decision and its two-layer shape are in [DESIGN-NOTES.md](DESIGN-NOTES.md) -> D-69; the clause +itself is `RS-P-8` in [RESPONSE-SPACE.md](RESPONSE-SPACE.md). The durability half it spawned is +`M26.11`. The item as it read when it closed: + +- [x] **M26.10** -- Decide whether a completion may report **fewer bytes than requested**, which + [RESPONSE-SPACE.md](RESPONSE-SPACE.md) currently lists as deliberately undecided. Found by + `M26.7`'s audit, and queued rather than settled there because the space's own instruction is to + "decide it before a resolver relies on either answer" -- which is a call about what this crate + tolerates, not a correction. + + **The observation.** Five kernel assertions require a full transfer (`assert_eq!(transferred, + LEN)`), so the suite already answers the question by assuming one. `Completion::result` promises + only "the transferred byte count" and never a complete one, so nothing in the crate backs that + assumption up. The two are not in conflict today only because `M26.3`'s resolver reports + `Information: 0` for every operation and no resolver-driven test reads a transfer count -- so a + resolver and a kernel test would disagree about the same field and nothing would notice. + + **Note one of the five is already right and should be left alone.** + [flush_barrier.rs](tests/flush_barrier.rs) checks the transfer explicitly as its *own + precondition* -- a short write would make every count in that test meaningless -- and says so. + That is the honest form of the assertion whichever way this is decided. + + **What deciding it costs.** Permitting partial transfers widens what every consumer must handle + and would make the resolver able to produce them, which the properties in + [properties_under_every_resolution.rs](tests/properties_under_every_resolution.rs) would then + have to survive. Requiring complete transfers is a `Decided` constraint of the kind `RS-C-1` + already is, and would need a `CONFIRMS:` marker on whichever kernel test carries it -- the census + added in `M26.6` will then hold the two halves together. + +## Moved 2026-09-25 16:21:00 -04:00 -- M26.11, the epoch log's handle requirements + +### M26.11 -- The epoch log states the capabilities it requires of a handle; the transfer requirement is checked, the durability requirement is a caller warranty. *(completed 2026-09-25 16:21:00 -04:00)* + +The contract gained a fourth clause, `Requires`, beside `Guarantees`, +`DoesNotGuarantee` and `Assumes` -- the growth the module's own `Clause::ALL` +doc had anticipated, and its exhaustive `heading()` match is what forced the +edit. The decision not to gate the handle is [D-70](DESIGN-NOTES.md#d-70). The +item as it read when it closed: + +- [x] **M26.11** -- Document the capability requirements [examples/epoch_log](examples/epoch_log) places on + the handle it is given, and let an unmet one surface at the operation that needs it. **No pre-flight + handle check** ([D-70](DESIGN-NOTES.md#d-70)). + + **Why no check, which is the part worth not relitigating.** `M26.10` established that the ring permits a + short transfer because it never asks what kind of handle it was given ([RS-P-8](RESPONSE-SPACE.md), + [D-69](DESIGN-NOTES.md#d-69)), and the obvious next move is for the log to police what the ring does + not. It was examined and rejected: the check points the wrong way. A console handle, a pipe or a closed + handle is caught by `GetFileType`, but every one of those already fails loudly at the first positioned + write -- the check buys a better message. A RAM disk, a remote share, or a volume whose write cache is + not power-protected succeeds at every API call and silently fails to be durable, and no probe catches + the last of those at all. So the gate guards the failures that were already loud and misses every + failure that is silent, while implying a validation that did not happen. + + **What to write instead.** The log's own contract, stated by the log rather than inherited from the + ring, listing what the handle must support: positioned I/O at explicit offsets; `FILE_FLAG_OVERLAPPED`; + `FILE_FLAG_NO_BUFFERING` together with the sector-aligned buffer, offset and length it requires; a + preallocated extent; that a successful write of `N` bytes transfers `N`; and that a completed flush + reaches stable media. + + **Mark the last two for what they are.** The transfer requirement is the `RS-P-8` narrowing this log + earns by constraining its input -- it is checkable, and the log already compares transferred against + requested, so a mismatch is a contract violation to report loudly rather than a case to absorb. The + durability requirement is **a warranty the caller gives**, not a property this code can verify; + say so in those words, because a contract that merely sounds confident about it is how a silent + failure gets built on. + + **Guard what is guardable, and do not pretend about the rest.** The transferred-against-requested + comparison is a real assertion and gets a sabotage case: suppress the comparison and the suite must go + red. There is no guard for the durability warranty, and the item is complete with that stated rather + than papered over. If the contract carries runnable examples they are compiled as doctests, per the + repository's rule that prose containing code must compile. + +## Moved 2026-09-25 16:24:00 -04:00 -- M26 complete, all eleven items + +Every item was archived individually as it closed, so what migrates here is the milestone's own +framing -- why a resolver over a permitted space was worth building, and the standing constraint +that the space must be wider than anything observed. The per-item stubs are dropped, carrying +nothing the entries above do not already hold. + +## M26 -- Test against the space of kernel responses, not one observation of it + +Queued by +[DESIGN-SESSION-2026-09-22-kernel-response-space.md](design-sessions/DESIGN-SESSION-2026-09-22-kernel-response-space.md), +which set out to answer `M24.1` and found a different technique instead. + +**The idea.** A fake that models *what Windows does* freezes one run's testimony. A **resolver** +models what Windows is *permitted* to do, and a seed picks one resolution out of that space: which +operations finish inside `SubmitIoRing` and which pend, in what order completions are posted, which +fail. The assertions are then about **us** -- does this crate behave correctly under that resolution +-- and never about the kernel. There is no belief to be wrong about, which is why this dissolves the +mock objection rather than working around it. + +**Justified by what it catches, not by hermeticity.** `M24` reaches a hermetic lib suite without it, +so this milestone has to earn its place on the defect class it detects: code that is brittle to +platform variation *inside* the permitted space. Nothing in the toolkit that preceded it detected +that -- [DESIGN-NOTES.md](DESIGN-NOTES.md#what-none-of-them-cover) records that the five techniques +which existed before `M26` all check this crate against *its own stated contract*. The resolver is +the sixth, added by this milestone. + +**The standing constraint, inherited from the session.** The permitted space must be **wider than +anything observed**, and must not be derived from observation -- deriving it from what we have seen +closes the trap again. It is a deliberate specification of what we will tolerate, and therefore a +reviewable artifact rather than a recording. diff --git a/crates/windows-ioring-sys/Cargo.toml b/crates/windows-ioring-sys/Cargo.toml index 6ff8a6613..e75f9b3c1 100644 --- a/crates/windows-ioring-sys/Cargo.toml +++ b/crates/windows-ioring-sys/Cargo.toml @@ -21,6 +21,12 @@ path = "src/lib.rs" [features] default = ["threadpool"] +# Propagates windows-threadpool-sys's concurrency trace. Off by default and +# absent from the compiled output when off: the defects it exists for are +# timing-dependent, so the instrument must not change the schedule it is +# measuring. Narrow it with WINDOWS_THREADPOOL_TRACE at run time. +trace = ["windows-threadpool-sys/trace"] + # A test-support seam that makes a *real* completion report a failure it did # not actually have (M16.3). Off by default: it is for testing failure # handling, and nothing in production has a use for it. @@ -31,6 +37,15 @@ default = ["threadpool"] # whole design and is argued in full on `Completion::with_injected_failure`. # Fabrication stays `#[cfg(test)] pub(crate)` and is not reachable from here. fault-injection = [] + +# The indirection the M26 kernel-response resolver plugs into (M26.2). Off by +# default: with it off, every wrapper in `sys` is an inline forward to the same +# `windows-sys` call this crate made before, and the install point does not +# exist at all. +# +# A feature rather than `cfg(test)` because the integration suite in tests/ +# cannot see `cfg(test)` items -- the same reason `fault-injection` is one. +kernel-seam = [] # Model A delivery (`EventDelivery`) is the only thing in this crate that needs # a thread pool, so a Model B consumer -- a pinned thread parked in # `Batch::submit_and_wait`, owning its own ring -- otherwise links a dependency @@ -59,6 +74,21 @@ required-features = ["threadpool"] name = "epoch_log" path = "examples/epoch_log/main.rs" required-features = ["threadpool"] +# M21.4 gives the sample's durability accounting unit tests -- the failed-commit +# case cannot be reached by running the sample, because it needs a completion +# whose result is an injected failure. Examples are not test targets by default, +# so `cargo test` would compile this and run nothing. +test = true + +# `ring_copy` was auto-discovered until M20.3, which needed it to be a test +# target for the same reason `epoch_log` is: the degraded-fallback branch in +# `Policy::select` is the one every zero-relation machine takes, and it cannot +# be reached by *running* the sample on a machine that reports its relations. +# A synthetic topology reaches it; a test target is what lets one run. +[[example]] +name = "ring_copy" +path = "examples/ring_copy/main.rs" +test = true # The crate is Windows-only, so docs.rs must build on a Windows target or it # would render an almost-empty crate. @@ -72,6 +102,11 @@ targets = ["x86_64-pc-windows-msvc"] # under `Win32_System_IO` as their subject matter might suggest. # `Win32_System_Threading` supplies the event primitives the Model A delivery # path (M4) waits on. +# +# `Win32_System_Memory` is gone from this list, and its absence is the check +# that `NumaBuffer` really left: the allocator was the only thing here that +# reached for it, so a build that still needed it would mean something had been +# missed. It now belongs to `win-numa-sys`. windows-sys = { version = "0.61.2", default-features = false, features = [ "Win32_Foundation", "Win32_Security", @@ -83,6 +118,13 @@ windows-sys = { version = "0.61.2", default-features = false, features = [ # event. Optional behind the default-on `threadpool` feature (D-22); nothing # else in the crate references it. windows-threadpool-sys = { version = "0.1.3", path = "../windows-threadpool-sys", optional = true } +# NumaBuffer moved here in M22.3 and out again when win-numa-sys was created: +# the buffer has nothing to do with a ring, and windows-placement-probe had +# independently written the same VirtualAllocExNuma call. This crate keeps +# re-exporting the type, so windows_ioring_sys::NumaBuffer still resolves for +# anyone who bound to it, and supplies the IoBuf/IoBufMut impls that the +# allocator crate deliberately does not define. +win-numa-sys = { version = "0.1.0", path = "../win-numa-sys" } [dev-dependencies] # M6.3's topology-guidance example enumerates L3 cache domains with this @@ -111,13 +153,18 @@ windows-topology-sys = { path = "../windows-topology-sys", features = [ ] } # M7's ring-copy sample deserializes a fed-in topology description (--topology). serde_json = "1.0" -# M7's ring-copy sample additionally needs `VirtualAllocExNuma` (buffer -# placement) and `GROUP_AFFINITY` (pinned-thread affinity); M14.1's epoch-log -# sample needs `Win32_System_Ioctl` for the `FSCTL_SET_ZERO_DATA` reclamation -# it orders against ring epochs. The library itself needs none of them. +# M7's ring-copy sample additionally needs `GROUP_AFFINITY` (pinned-thread +# affinity); M14.1's epoch-log sample needs `Win32_System_Ioctl` for the +# `FSCTL_SET_ZERO_DATA` reclamation it orders against ring epochs, and for the +# `FSCTL_QUERY_VOLUME_NUMA_INFO` its arena placement asks (M22.3). +# +# `Win32_System_Memory` is deliberately NOT here: buffer placement moved into +# the library as `NumaBuffer` (M22.3), so the feature is a *library* one now +# and listing it again would wrongly suggest a sample reaches for +# `VirtualAllocExNuma` on its own. windows-sys = { version = "0.61.2", default-features = false, features = [ "Win32_System_Ioctl", - "Win32_System_Memory", + "Win32_System_Pipes", "Win32_System_SystemInformation", ] } # M14.1's epoch-log sample orders a non-ring `FSCTL` against ring epochs, which diff --git a/crates/windows-ioring-sys/DESIGN-NOTES.md b/crates/windows-ioring-sys/DESIGN-NOTES.md index 5fee420df..c1b4cc167 100644 --- a/crates/windows-ioring-sys/DESIGN-NOTES.md +++ b/crates/windows-ioring-sys/DESIGN-NOTES.md @@ -1,11 +1,22 @@ # Design notes: windows-ioring-sys (Tier 1) -This crate does not exist yet as compiled code. This file, the checklist beside it, and the design session -it references are the design record that precedes it. Creating the Cargo skeleton is M1.1 in -[CHECKLIST.md](CHECKLIST.md). +This file is the crate's current design record: what was decided, and what forced each choice. It was +written before the implementation, and the design session it references is where its earliest decisions +came from. + +It deliberately states no release or milestone status. That is derivative of +[CHANGELOG.md](CHANGELOG.md), the git tags, and [CHECKLIST.md](CHECKLIST.md), and a copy of it here would +be one more thing to keep true by hand. ## Intent +**Why this crate exists at all is recorded one level up**, in the workspace's +[DESIGN-NOTES.md](../../DESIGN-NOTES.md) -> [The adoption thesis](../../DESIGN-NOTES.md#the-adoption-thesis). +Read it before proposing to remove a design option here: it is the reason this crate keeps +alternatives alive that no measurement on the development machine can justify, and the reason its +samples hand a consumer data rather than a verdict. The operational form of that posture is OPTION +INTEGRITY in [copilot-instructions.md](../../.github/copilot-instructions.md). + Windows 11 / Server 2022 added `IoRing`: a submission/completion ring for file I/O, closer in shape to `io_uring` than to anything else Windows offers. This crate raises those primitives into memory-safe Rust with the minimum additional CPU and memory cost, in the same spirit as the rest of this repository. @@ -69,6 +80,30 @@ runs a continuation), this crate exposes the mechanism and documents the trade-o | D-44 | **A spike against the real kernel is a budgeted, first-class technique for every new Win32 surface this crate wraps -- not something that happens after a test fails mysteriously.** The full argument is in [Testing strategy](#testing-strategy-m185); the decision is that the budget is allocated *before* the wrapper is written. Two of the eight defects behind M15-M18 exist because a Win32 contract was assumed rather than measured: the completion event is edge-triggered ([D-19](#d-19)) and `BuildIoRingRegisterBuffers` reads its array when the operation *runs* ([D-32](#d-32)). No oracle, generator, allocator or mutation run supplies that knowledge, because each of them checks code against **our** stated contract -- and in both cases our stated contract was the thing that was wrong. What they detect is a *consequence*, and only on a path some test already walks: the guard allocator does turn D-32 into a hard `STATUS_ACCESS_VIOLATION`, measured in M17.4's calibration, but that is the crash after the mistake, not the knowledge that would have prevented it. A spike is also the only technique here that can be run *before* there is code to test. Two obligations follow, both learned the hard way and recorded in [design-sessions/spikes/README.md](design-sessions/spikes/README.md): a spike must carry a **control case**, because the first two drain-ordering spikes could not discriminate and would have returned confidently wrong answers; and it must be **kept**, as a standalone single-file program depending only on `windows-sys`, so that what it measures stays the operating system's behaviour rather than ours. | | D-45 | **A borrow-returning method must be audited on two questions, not one: what the returned value *permits*, and how long the *borrow* lasts. `RegisteredBuffers::get` therefore takes `&mut self`.** [M18.1's audit](#borrow-surface-audit-m181) asked only the first, of all nineteen items, and the second is where [D-36](#d-36)'s fix was still open: `get` checked `kernel_writes` at the instant of the call but returned a slice living as long as the borrow, and `Batch::read_registered` takes the registration by **shared** reference -- so safe code could take the borrow while the buffer was quiet, then submit a read into that same buffer and keep reading. Measured before being believed: a probe watched the bytes change from `0x11` to `0xEE` through the live slice while a fresh `get(0)` at that same instant correctly refused with `WouldBlock`. The guard worked; the borrow outlived it. **`&mut self` costs nothing real**, because no caller needs to read a buffer during the window it is refused -- while a read is in flight the bytes are indeterminate and only become meaningful once the completion is observed, so earlier or later is always available. That is not merely an argument: all ~40 read sites in this crate's tests, examples and the epoch-log sample already read at a quiescent point, and converting them needed nothing but `mut` on a local. The concession D-36 deliberately kept (reading a buffer whose own *write* is in flight, where the kernel only reads) is given up with it, and is likewise unused. The arena pattern survives, because a [`Token`] holds a [`RegisteredUse`] rather than a borrow of the registration, so quiet neighbours stay readable while operations are outstanding. Enforced by a `compile_fail` doctest, itself verified by reverting the signature and watching it fail, and paired with a `no_run` doctest asserting the neighbour case still compiles so the guard cannot become over-constraining unnoticed. `get_mut` never had the defect: `&mut self` already conflicted with the shared borrow. | | D-47 | **`IOSQE_FLAGS_DRAIN_PRECEDING_OPS` is one-sided, and [D-24](#d-24)'s claim that it holds back subsequent operations is withdrawn. Measured over ~4,500 trials.** The drain half is solid: **not once** did an operation queued *before* a drained flush complete after it. The hold-back half is false: post-flush writes overtake the flush at 0.03%-0.8% depending on conditions, and in the worst observed trial *all 32* did. The rate is why this read as a flaky test for three days rather than as a contract defect -- at roughly one run in a thousand, it surfaces every few days in a full-workspace run and never in isolation. Contention raises the rate but is not required (it reproduces on an idle machine); ring depth does not move it (128, 256 and 512 were indistinguishable). Every violation was confined to a single drain of the completion queue, so it is the queue's own posting order rather than an artifact of sampling it twice. **What a consumer may rely on:** a drained flush's completion means everything outstanding when it was reached is durable. It does **not** mean later work has been held. See [D-47 in detail](#d-47-detail). | +| D-48 | **A shipping ARM consumer laptop reports no L3 cache domains at all, and zero `Win32_NumaNode` instances. Measured, and it is an ordinary consumer shape rather than an exotic one.** A Snapdragon X2 Elite (X2E80100, Qualcomm Oryon), 12 cores and no SMT, reports L1 and L2 only: L2 forms two domains of six processors, agreeing with the two `Module` domains the same probe returned, and WMI reports an L3 size of zero. The capture is Measurement M-1 in [DESIGN-SESSION-2026-08-30-numa-sharded-io-execution-domains.md](../../design-sessions/DESIGN-SESSION-2026-08-30-numa-sharded-io-execution-domains.md). **This is the sibling of the zero-node observation these notes already carried** in [Why the NUMA node is the wrong key](#why-the-numa-node-is-the-wrong-key), and the pair is the point: on one machine the NUMA node is absent because a hypervisor did not present it, on the other the *cache level this crate's guidance names* is absent because the silicon has none. Neither machine is unusual. **What it falsifies is a justification, not the heuristic.** These notes say the last-level-cache domain "is meaningful on Intel and ARM too, where the NUMA node often is not"; on this part it is not meaningful, because it does not exist. That a cache domain beats the node is untouched. **The rule was restated and every restatement swept by `M20.1`/`SH-4.12` (2026-09-22)**, which also replaced the consumer that bound to the level number. Measuring that consumer while fixing it found a second shape this decision did not anticipate: an L3 that spans *every* processor above a real L2 partition, where a `level == 3` filter matches, does **not** degrade, and reports one whole-machine domain as a successful cache-aware partition -- see [Why the NUMA node is the wrong key](#why-the-numa-node-is-the-wrong-key). **Not decided here:** whether the two-cluster L2 structure this part does report is worth sharding on. | +| D-49 | **The unit suite is not hermetic, and that is a defect rather than a property of wrapping Win32. 63 of 131 lib tests open a real kernel ring.** The repository's Quality rule already classifies this: it reserves integration tests for work that must cross "a real process, filesystem, network, device, **operating-system API**, or other external boundary", and `CreateIoRing` is an operating-system API. So those 63 are integration tests sitting in the unit-test location, and `cargo test --lib` does not mean what its name implies. **This was surfaced by a load-dependent failure and initially mis-diagnosed.** `M21.6` removed five wall-clock assertions from four lib tests, which made their *outcome* independent of load -- measured at 30x duration variation with zero outcome variation under 2x CPU saturation -- and that was reported as the fix. It was not: outcome-stable and hermetic are different properties, and those tests still open a kernel ring. **What is decided here is the defect and the classification, not the remedy.** Three remedies are costed in [DESIGN-SESSION-2026-09-21-hermetic-unit-tests.md](design-sessions/DESIGN-SESSION-2026-09-21-hermetic-unit-tests.md) and the choice between them is gated on one unresolved question: whether a fake whose assertions are *shared* with the kernel escapes the objection in [Two techniques deliberately rejected](#two-techniques-deliberately-rejected), which refuses a mock that would "manufacture evidence" a kernel-behaviour bug was absent. **`M24.1` settled that (2026-09-22) and the answer was that the instrument was wrong**: a shared suite is strong over what this crate specifies and blind to the platform's incidental behaviour, and an assertion about the latter is a frozen observation rather than a contract -- see [D-52](#d-52). The rejection stands with its scope sharpened; the technique that replaces it is the response-space resolver, scheduled as `M26`. **A bright line holds under every remedy:** a fake never answers a question about Windows. Edge-triggered delivery ([D-19](#d-19)), one waiter per ring ([D-21](#d-21)), flush coverage and drain ordering ([D-23](#d-23), [D-24](#d-24), [D-47](#d-47)), the registration array read at run time ([D-32](#d-32)), `ERROR_TIMEOUT` and `E_INVALIDARG` from `SubmitIoRing`, and inline completion on a synchronous handle all stay kernel-tested forever -- which is precisely the set of findings that produced this crate's defects, and the argument for drawing the line exactly there. **`M24` has since run, and the figures above are the ones it started from.** Measured after it: **41 of 151 lib tests open a ring, 110 do not**. The remainder is not movable without `M26.2`'s FFI seam -- `event_delivery` needs the thread pool, `ring`'s injected-failure cluster transforms a *real* completion by design, `batch` needs the handle, and several reach `#[cfg(test)] pub(crate)` helpers. A zero-check is therefore the wrong rung, and [D-53](#d-53) records what replaced it. | +| D-50 | **The epoch-log sample places its registered arena on the NUMA node its own log file's volume reports, and says in the same breath that the placement cannot pay at this workload.** The arena was `vec![0_u8; SLOT_LEN]` -- heap, no alignment, no node -- while this crate's front page told every consumer that placing the registered pool near the device "is very likely the highest-leverage locality decision available". That silence read as an oversight. The node is asked of the log's own handle through `FSCTL_QUERY_VOLUME_NUMA_INFO`, which [What is not reachable](#what-is-not-reachable) already established as the documented mechanism; a volume that names no node yields no preference, and the log runs on. **What is deliberately not claimed is any benefit.** The arena is eight slots of four kilobytes and this workload is bound by a per-epoch device flush costing hundreds of microseconds -- `M22.1` measured that directly. So the sample demonstrates *how the decision is made and reported*, and `examples/ring_copy` remains where placement is put under a load that could show it. The report line is qualified by `GetNumaHighestNodeNumber` for the same reason: on a one-node machine "placed on node 0" is true and misleading, so the sample says the choice was never available. **This is a sample's local policy, not a retraction of [D-8](#d-8):** the library still maps no file to a node, because a volume may span devices and its node is not where a file's extents live. | +| D-51 | **`NumaBuffer` moves from `examples/ring_copy` into the library, because recommending an allocation while making every caller write it is what produced the second copy.** This crate's front page names `VirtualAllocExNuma` on the device's node as the highest-leverage locality decision available, and then supplied nothing; the first consumer wrote the allocator in a sample, and `M22.3` was about to make a second. The type is a thin owned mapping implementing [`IoBuf`]/[`IoBufMut`], so it registers like any other buffer. **It decides no policy** -- which node is still the caller's answer, per [D-8](#d-8) -- and its own documentation records that `nndPreferred` is a preference, so a successful allocation is not evidence the pages landed there. The move surfaced a packaging defect that `cargo check --all-targets` cannot see: dev-dependency features are unified into that build, so the missing `Win32_System_Memory` on the *library* dependency only appears in a lib-only build or `cargo doc`. A consumer would have hit it on first compile. | +| D-52 | **Test this crate against the *space* of responses the platform is permitted to give, not against one observation of what it gave. A fake that models what Windows does freezes one run's testimony; a seeded resolver over the permitted space has no belief to be wrong about.** `M24.1` set out to ask whether a fake with *shared* assertions escapes the mock rejection, and demonstrated that it does not: a fake built from this crate's own pre-`M21.6` belief passed the shared suite green, and the assertion that catches it could only be written after the kernel had already revealed the answer. The deeper finding is that an assertion about platform behaviour is a **frozen observation** rather than a contract -- "after submitting, the completion is already queued" gave *opposite answers on two handles of the same API*, and [write-pending-spike.rs](design-sessions/spikes/write-pending-spike.rs) reported one condition as 5/500 in one run and 271/500 minutes later. Freezing such an observation into a suite, then building a fake to satisfy it, leaves three artifacts agreeing -- which reads as corroboration but is one observation restated three times, the same failure [D-47](#d-47) already cost this crate once. **What replaces it:** the resolver decides from a seed which operations complete inside `SubmitIoRing` and which pend, and in what order completions are posted; the assertions are then about this crate's behaviour under that resolution. Non-reproducibility stops being a threat and becomes the expected case. **The permitted space must be wider than anything observed and must not be derived from observation**, which makes it a deliberate, reviewable specification of what we tolerate -- and it must state the constraints that *do* hold, or the tests demand code defending against impossible kernels. This is a sixth technique beside the five in [What none of them cover](#what-none-of-them-cover): it still cannot tell you the stated contract is wrong, but it detects brittleness to variation inside the space, which none of the five can. Reasoned, not yet measured, against `D-47` and `M21.6`; scheduled as `M26` in [CHECKLIST.md](CHECKLIST.md), whose calibration item exists because `D-41`'s corollary demands it. Session: [DESIGN-SESSION-2026-09-22-kernel-response-space.md](design-sessions/DESIGN-SESSION-2026-09-22-kernel-response-space.md). | +| D-53 | **The rung guarding the hermetic lib suite is an inventory of *which* lib tests open a ring, not a zero-check and not a count.** `M24.5` originally assumed that after the extraction and the relocation "the lib tests should construct no ring at all", so a zero-check would do. That rule is false and cannot be made true by effort: `event_delivery` needs the thread pool, `ring`'s injected-failure cluster transforms a **real** completion on purpose (fabricating one is the unsoundness the seam exists to avoid), `batch` needs the handle for its `Build*` calls, and several tests reach `#[cfg(test)] pub(crate)` helpers that exist only inside the crate. A zero-check would fail on day one and could only be satisfied by deleting real coverage. **Per-test rather than per-file**, because two thirds of the remaining 41 live in `ring/tests.rs` and a file-level allow-list would let that file grow without limit -- which is where a new ring-opening test would most naturally land. **An inventory rather than a count**, because add-one-remove-one nets to zero and passes, and a bare number is derived data no reader can check. The mechanism is [check-borrow-surface.ps1](../../tools/check-borrow-surface.ps1)'s, deliberately: a committed list regenerated from source, failing when the two disagree, so an addition obliges the question *does this test need the kernel, or only a ring-shaped thing?* A removal is progress and needs only regeneration. **The guard's own bidirectional check found a defect in it**: a plain helper defined after the last test in a file was being swallowed into that test's body, reporting an innocent test as ring-opening -- the body now ends at a column-0 `}` rather than at the next attribute. Inventory: [RING-OPENING-LIB-TESTS.txt](RING-OPENING-LIB-TESTS.txt); script: [check-ring-tests.ps1](../../tools/check-ring-tests.ps1). | +| D-54 | **This crate owns what *one* flush means. It owns nothing about durability *groups*, and that is a deferral rather than a gap.** The primitives are here because they are facts about a single operation: [`FlushCoverage`] (the barrier is ring-wide), [`FlushMode`] (which modes sync the device), [`WriteCaching`], and the scope distinction between them -- a barrier over the ring, a flush over one file. **Grouping is not here and is not coming here**: what a set of operations that commit together costs, whether two such sets contend, what a co-flush regime implies, and the `Epoch` concept itself. That layer is [C-3](../../design-sessions/DESIGN-SESSION-2026-08-30-numa-sharded-io-execution-domains.md)'s durability crate, queued as `M33+.5` in [CHECKLIST-io-domains.md](../../CHECKLIST-io-domains.md); today `Epoch` exists only in `examples/epoch_log`, and no grouping concept appears anywhere in `src/`. **The test for a proposal is whether it needs the concept of a set of operations that become durable together** -- if it does not, it may belong here and must be justified on its own merits; if it does, it is deferred. Recorded because the boundary was re-litigated three times in one day and because a conversational aside -- calling `M23.3` a "down-payment" on the durability crate -- left the impression that the layer had been folded into this one. It has not been, and building *toward* it from here is what this decision forbids. | +| D-55 | **The pending-token inventory becomes the ring's, and `IoRing` becomes generic to hold it. The break is accepted.** This crate hands a caller a `Token` from one call and a `Completion` from another, and connects them with nothing -- its own rustdoc twice instructs a caller to "match it against a held `Token`". Every consumer with more than one outstanding tokened operation must therefore build an identity map, and the *correct* one encodes four rules a `HashMap` cannot express; two measured defects in this repository came from the obvious one. Offering a `Pending` beside the ring closes the duplication but not the mechanism: nothing would force a minted token into it. A generic `IoRing` owning the map does, because the consumer never holds a token to lose. **The type-erasure objection is withdrawn as false**: per-ring monomorphisation holds for every real consumer here, a closed `enum` serves the rest -- `tests/generated_sequences.rs` already carries eight token types on one ring that way -- and [D-4](#d-4) independently rules type erasure out ("no slab entry, no box, no type erasure"), so the objection contradicted a decision already on the books. Implementation and migration are `M28`. | +| D-56 | **A benchmark that defers its await measures the deferral, not the operation -- and the number survived three rounds of correction because every round corrected the conclusion instead of the instrument.** `epoch_log`'s harness published a commit latency measured from pushing a flush to observing its completion. `M20.6` decomposed it and found **blocking p50 and p99 of 0 us for all three strategies**: the figure was entirely the interval in which the program went on appending, so a design that deferred further reported a worse commit while being no slower. Three separate rounds of work had already re-read the *conclusion* drawn from that number -- the harness's serialisation, the `UserData` collision, the per-record submit -- and none had asked whether the number measured what its name said. **The fix is structural: a commit's cost is now reported in parts** (`prepare`, `submit`, `blocking`) with `deferral` beside them and excluded from the total, so the two cannot be read as each other. `M25.5` then found the same defect one layer down in that fix -- the clock started after `HostSequenced`'s host round trip, making it look six times cheaper -- which is why `prepare` exists. **What generalises:** a measurement whose parts are not separately reported can be wrong in a way that no amount of re-reading its output will reveal, and "the conclusion still holds" is not evidence that the instrument does. See [measurements/2026-09-24-commit-decomposed/](measurements/2026-09-24-commit-decomposed/README.md). | +| D-57 | **A flag whose requirements the caller's data layout cannot satisfy is not a flag change.** Making `epoch_log`'s commit separately observable needed `FILE_FLAG_NO_BUFFERING`, which constrains the transfer's buffer address, file offset **and length** -- and the log wrote variable-length records at packed offsets, satisfying none of them. The work was therefore a change to the log's **on-disk format** (`M25.1`: one record per sector-sized block, tail zeroed, in both writers) before a single flag could move (`M25.3`). [write-pending-spike.rs](design-sessions/spikes/write-pending-spike.rs) predicted exactly this from its conditions, and the prediction is the reusable part: when a configuration's preconditions reach into a caller's data layout, costing it as a flag underestimates it by the size of a format migration. The stride's price is write amplification, which the sample now measures and prints rather than describing, and the reason a real log pays it anyway is sector atomicity -- a record sharing a sector with its neighbour can be torn by that neighbour's write. | +| D-58 | **A contract checker's failure vocabulary decides its interface, and `epoch_log`'s replay keeps a `&[u8]` for that reason rather than for simplicity.** The walk is strictly forward one block at a time and never looks back, so it has no need of the whole file -- and a real log is larger than memory, which makes reading the whole file the wrong reflex to teach at precisely the point a reader is learning to verify one. The slice is kept anyway because `replay` returns an `Outcome`, not a `Result`: every way it can end is a statement about the log -- verified, tolerated, or a `Violation`. A streaming reader introduces a third kind of ending, `io::Error`, into the one component whose entire job is to distinguish "the log broke its promise" from "the log kept it", and a signature returning both through one channel invites exactly the conflation the file exists to prevent -- an unreadable file reported as a missing durable record. **So the streaming version is a different interface, not a smaller allocation**, and the cost of declining it is stated where it is paid rather than hidden: 140 KiB for the log, 8 MiB per strategy in the harness. A consumer building a real verifier wants the other shape *and* wants the two failure kinds kept apart inside it. What was reducible without touching that interface was reduced: the cross-strategy comparison kept a whole reference log in memory for the length of the comparison and now keeps a digest, which halves the peak and loses nothing a reader had, since the assertion could already only say *that* two logs differed. | +| D-59 | **The permitted kernel response space is specified in [RESPONSE-SPACE.md](RESPONSE-SPACE.md), as a statement of what this crate will tolerate rather than a record of what Windows did.** [D-52](#d-52) settled that testing against a *space* dissolves the mock objection; this is that space, written down. Eight clauses say what a resolver **may** do -- an operation may complete inside `SubmitIoRing` or pend (`RS-P-1`), completion order is unconstrained (`RS-P-2`), an operation may fail individually (`RS-P-3`), a wait may expire (`RS-P-4`) or return with nothing poppable (`RS-P-5`), a completion behind another need produce no signal (`RS-P-6`), a failed submit leaves operations queued for a later one (`RS-P-7`), and a successful transfer may report fewer bytes than requested (`RS-P-8`, added by `M26.10`). **Four say what it may not**, because a resolver free to violate everything makes this crate defend against a platform that does not exist: completions are conserved and identified (`RS-C-1`, `RS-C-2`), nothing completes before submission (`RS-C-3`), and **the drain half of `DRAIN_PRECEDING_OPS` holds (`RS-C-4`)**. That last is the call `M26.1` demanded rather than defaulted: [D-47](#d-47) measured roughly 4,500 trials without a single violation, the drain is what this crate's durability story rests on, and a resolver permitted to break it would make the primitive useless -- so a Windows that broke it is caught by the kernel tests instead, which is the division of labour they were repointed to in `M26.6` and is now enforced by a census rather than intended. The hold-back half stays unconstrained, since `D-47` withdrew it. **Every clause is tagged `Observed`, `Over-provision`, or `Decided`**, so a reader can tell a measurement from an extrapolation from a call, and the space is deliberately wider than anything observed -- deriving it from observation would close the trap `D-52` was opened to escape. Rates, partial transfers, failure-code sets and timing are listed as deliberately undecided so an omission cannot be mistaken for a choice. | +| D-60 | **The kernel seam is module indirection, not a type parameter, because [D-55](#d-55) has already spent `IoRing`'s.** `M26.3`'s resolver has to be able to answer the crate's kernel calls, and the textbook shape for that is a generic `IoRing` over a kernel trait. That shape is unavailable here: `D-55` commits `IoRing`'s type parameter to `M28`'s token inventory, so a kernel generic would publish `IoRing` on a shipped crate -- a second public parameter whose only purpose is letting the crate test itself. **The eight submission-path calls therefore route through [sys.rs](src/sys.rs) instead** -- `SubmitIoRing`, `PopIoRingCompletion`, and the six `Build*` entry points -- each an `#[inline(always)]` wrapper whose `through_seam!` macro expands to the bare FFI call when the `kernel-seam` feature is off. The public surface is unchanged and the type parameter stays free. **The responder is thread-local rather than process-global** for the reason this workspace is not on nextest: `cargo test` runs tests as threads in one process, so a global responder would let one test answer another's calls -- the same hazard that puts `DROP_RUNS` inside its test function. `with()` falls through to the kernel on a re-entrant borrow rather than panicking, and `Installed::drop` uses `try_with` so teardown cannot abort ([M23.4](COMPLETED-CHECKLIST.md#m234)). **Four lifecycle calls are deliberately left direct** -- `CreateIoRing`, `CloseIoRing`, `GetIoRingInfo`, `IsIoRingOpSupported`. (This decision originally listed five; `M26.3` moved `SetIoRingCompletionEvent` behind the seam, because it is how a completion becomes *observable* and `RS-P-6` is therefore a clause about it -- see [D-61](#d-61). The correction is recorded here rather than only there, since a reader arriving at this row would otherwise take the superseded list as current.) `M26` justifies itself on the *response space*: what the kernel may answer to submitted work. Routing construction and teardown through the seam too would be hermeticity for its own sake, which is `M24`'s subject and not this one. The boundary is stated here so a later reader can tell it from an oversight -- and `M26.3` carries the converse, that a clause needing one of those five extends the seam rather than working around it. | +| D-61 | **The resolver's permissions are configurable and its constraints are not, and that asymmetry is what keeps it a specification rather than a fake.** `M26.3` builds the resolver [RESPONSE-SPACE.md](RESPONSE-SPACE.md) was written for: it answers the seam's calls itself, choosing a point in that space from a seed. Every freedom cites the `RS-P-n` permitting it and every restriction cites the `RS-C-n` requiring it, so a clause no code cites is visibly unimplemented and a behaviour citing no clause is the resolver inventing a platform. `ResolverConfig` therefore has a switch per permission and **none for any constraint**: narrowing a freedom is how a test isolates another, while a knob relaxing a constraint would let a test quietly assert against a platform that cannot exist. The default is the **widest** point, so an unconsidered test fails loudly under a freedom it did not handle rather than passing while exercising nothing. **Where a permission and a constraint collide, the constraint wins** -- `RS-C-4` holds a drain-flagged operation back even when `RS-P-1` chose to complete it now, and `RS-C-1` forces a post an operation's coin kept deferring. **Two mechanisms exist only to make `RS-C-1` finite** (a per-operation deferral bound and a bound on consecutive declined submits); both are properties of the resolver rather than of the space, which carries no rates deliberately. **A submit resolves and a pop only rescues**, because a pop that flipped coins would make `RS-P-1`'s pending case unobservable to a polling consumer -- the freedom would be implemented and untestable. `SetIoRingCompletionEvent` moved behind the seam for this item: it is how a completion becomes *observable*, so `RS-P-6` is a clause about it, and a resolver unable to make that call does not satisfy the clause vacuously but hangs every [`EventDelivery`](src/event_delivery.rs) consumer instead. **The resolver's first contact with a real ring found a live defect**, queued as `M26.8`: a submit declined under `RS-P-7` propagates out of `IoRing::run_down` leaving work outstanding, which is `M21.6`'s hazard reachable through a different `HRESULT`. | +| D-62 | **The properties that must hold under every resolution are checked by *asking* [`RingContract`](src/contract.rs), not by restating it -- and the same rule decides where the expected outstanding count comes from.** `M26.4` states five properties over the resolver `M26.3` built: conservation, no hang, `pop_within` honours its bound, `outstanding` is accurate, and no use-after-free. Only two needed anything new. Conservation is already this crate's own oracle, so the harness reports to it and asks it for the verdict, on the rule that the layer owning an invariant owns the oracle for it -- a copy in a harness is a second implementation, and when the two disagree it is the harness that gets "fixed". **The accurate-`outstanding` property follows the same rule rather than counting for itself**: the expected value is read back out of the contract through its own `Outstanding` violation, because a counter in the harness would be a third party to the disagreement. **"No hang" is a step budget**, since a resolver answers instantly and a hang here is therefore an unterminating loop rather than a block. **Two weaknesses are declared rather than papered over.** `pop_within`'s upper bound is nearly free under an ordinary resolution -- the resolver rarely makes it approach its deadline -- so the non-vacuous case uses a separate degenerate responder in which nothing completes during the window; that is not an `RS-C-1` violation, because no finite observation can distinguish "eventually" from "never", but the prefix of a satisfying resolution in which the eventually has not happened yet. And **no-use-after-free covers this crate's memory handling, not the kernel's**, since under a resolver nothing external ever writes into a buffer; the kernel-side half stays with [generated_sequences.rs](tests/generated_sequences.rs) against a real ring. **Coverage counters are asserted, not printed**: all five properties are satisfied trivially by a run that does nothing, so a vacuity guard is what separates "the properties held" from "nothing reached the states they are about". | +| D-63 | **The resolver is calibrated against two defects that really happened, and the two suites' sensitivities were measured rather than assumed -- they differ, and the difference is the reason the calibration is its own file.** [D-41](#d-41)'s corollary is the rule: a green result from an instrument nobody has shown can go red is not evidence, and this repository has already produced one instrument of exactly that shape -- `M17.4` reverted issue #47 as it shipped and the generated suite reported **green**, because it sampled the right state while draining by a poll that recovers completions whether the ring signalled or not. **The two defects are calibrated differently because they live in different places.** `M21.6`'s -- an expired wait treated as a failure -- was in this crate, so it is re-injected by [sabotage.json](sabotage.json) and swept; what [calibration.rs](tests/calibration.rs) adds is the precondition that makes that sabotage mean anything, namely that `RS-P-4` reaches `pop_within` at all. [D-47](#d-47)'s is in a **consumer**, so there is nothing to mutate and the defective consumer is written out: a believer in the hold-back `D-24` claimed and `D-47` withdrew, which the resolver must break. **Measured, and the two suites are not equivalent.** `M21.6`'s defect is caught by both the calibration and `M26.4`'s property suite -- the latter only because that suite distinguishes a declined submit from a genuine error, which makes the detection deliberate rather than lucky. A *narrowed* resolver -- one enforcing the hold-back -- is caught by the calibration alone, and the property suite correctly stays green, because a narrower resolution is still a valid one and conservation still holds under it. **That is why a calibration file exists rather than a calibration assertion inside the property suite**: only a test that demands sensitivity can detect an instrument going blind, and such a test fails when the *instrument* regresses rather than when the crate does. | +| D-64 | **The standard seeded-sweep size is 2048, and a coverage threshold must be stated against the space a test can reach rather than against the sweep size.** The sweeps were widened from 64 (and the property suite's 240 plans) for breadth, on the measurement that seed count is not what these suites cost: a seed is tens of microseconds -- a whole ring, a batch, a drain and a rundown -- so the resolver's 22 unit tests sweep 2048 seeds each inside a lib suite that runs in well under a tenth of a second, far inside this repository's sub-second budget for a submodule. Only [properties_under_every_resolution.rs](tests/properties_under_every_resolution.rs) moved materially, to a little over a second, and even there roughly an eighth of the original cost was a deliberate sleep in the bound test rather than the plans. **The trap the widening exposed is the part worth keeping.** `different_seeds_reach_different_resolutions` asserted that distinct completion orders exceeded *half the seed count*, which is satisfiable only while the seeds are fewer than the outcomes: six operations admit `6! = 720` orders, so that threshold becomes arithmetically impossible past 1440 seeds and the test would have failed with nothing regressed. Thresholds of that kind are now phrased against the achievable space. **Saturation was measured rather than assumed**, because the obvious response -- cap the sweep where it stops gaining -- turned out not to apply: over the resolver's own mixer, 1024 seeds reach 539 of the 720 orders and 2048 reach 670, so the sweep is still gaining breadth at its current size and only around 8192 exhausts the space. **The property suite's vacuity guards are fractions of the plan count** for the same reason in reverse: an absolute floor chosen for 240 plans is a twelvefold margin at 2048, which would let the suite lose most of its work without complaint. Each is set near half the minimum observed over repeated runs, since that file's seeds are clock-derived and its counts therefore vary; the fixed-seed suites are deterministic and need no such margin. | +| D-65 | **The kernel tests' job is confirming a real Windows stays *inside* the specified space, and that division of labour is enforced by a census that reads markers rather than mentions.** `M26` split one job in two: the resolver sweeps the permissions (`RS-P-n`), and tests against a real ring confirm the constraints (`RS-C-n`). The split is load-bearing for exactly one clause -- `RS-C-4`, which the resolver is forbidden to violate, so a Windows that broke the drain would be caught by nothing the resolver does. **A hole was found where that mattered most**: [flush_barrier.rs](tests/flush_barrier.rs) asserted the clause but sat behind an early return taken when its control could not discriminate, so on such a machine the constraint was untested on both sides at once. The fix separates two questions that one gate had been answering together -- *is the barrier doing work*, a comparative claim that genuinely needs the control, and *did the kernel stay inside `RS-C-4`*, a conformance question where skipping can only hide a violation. The conformance assertion now runs on every machine and only the comparative claim is withheld. **The census was green and useless on its first build, and trying to make it go red is what found that.** It searched each file for the clause ID anywhere in its text; removing `RS-C-4`'s check from the only test performing it did not turn it red, because a second file mentioned the clause only to *disclaim* it -- and under a substring search a disclaimer is indistinguishable from a claim. This is the same trap already recorded for a probe whose only mention of a tag was a comment. A claim is now a structured `CONFIRMS:` / `EXERCISES:` marker naming a clause and nothing else, verified to go red in three directions. **Two blind spots are declared rather than hidden**: a census over source proves a clause is *claimed*, never that the file's assertion still runs or still means anything; and dropping the covering flag does **not** fire the `RS-C-4` assertion on every machine, because a device stack that orders a flush behind its file's outstanding writes by itself produces no violation to see -- so that sabotage is a portable demonstration of nothing, and the assertion's conformance value (reporting a violation if one occurs) is separate from its sensitivity (proving it would notice). **The toolkit's "five techniques" framing becomes six**, and the sixth is different in kind: the other five check this crate against its own stated contract and cannot tell you that contract is wrong, whereas the resolver checks it against a written specification of what the platform may do, and the kernel tests check the platform against that same specification. | +| D-66 | **The suite's frozen observations were one class, it was 31 assertions wide, and the restatement is guarded by a source census because no run objects to it.** `M26.7` audited the suite for assertions that pin the platform's incidental behaviour rather than this crate's contract -- the shape [D-52](#d-52) demonstrated, where one assertion gave *opposite answers on two handles of the same API*. **The census came from a command**, as the item required, and the answer was not the file anyone guessed: 31 sites across five kernel tests read `try_pop()` straight after `submit_and_wait` and unwrapped the `Option`, asserting the kernel had **already** queued the completion. `pop_within`'s own documentation denies that in this crate's words -- a submit-side wait's return "promises nothing about poppability" -- and [RESPONSE-SPACE.md](RESPONSE-SPACE.md) states it as `RS-P-5`. They passed for the reason [D-40](#d-40) measured: a buffered read completes inside the submit in 80 of 80 attempts, while an unbuffered one genuinely pends. **The restatement is the one `D-52` prescribed** -- ask for the completion within a bound this crate chooses, which holds on every handle -- and it also removed an unbounded `try_pop` spin found in the same sweep. **Nothing catches a regression by running**, which is the part worth keeping: reverting a site leaves its test green on any machine where the observation is true, so the guard is a census that refuses the shape at the source, plus a resolver-driven test that makes the pending case reachable on demand and shows `try_pop` failing where `pop_within` succeeds. **A second, narrower class was found and deliberately not settled**: five assertions require a complete transfer, which the space lists as *deliberately undecided* and `Completion::result` never promises. That is a specification gap the suite silently answers, queued as `M26.10` rather than decided here, since the space's own instruction is to decide it before a resolver relies on either answer. One of the five is already in the honest form -- [flush_barrier.rs](tests/flush_barrier.rs) checks the transfer as its own precondition and says so -- and is left alone. | +| D-67 | **This crate does not implement retry policy. It supplies bounded primitives and reports what the platform documented, and the caller owns backoff, attempt counts and when to give up.** `M26.8` began as "what should `run_down` do when its submit is refused" and was settled by reading `SubmitIoRing`'s reference page rather than by measuring, which is the correction worth keeping: **no amount of measurement constitutes a contract.** The documentation gives a return-value row for `IORING_E_WAIT_TIMEOUT` -- *"All operations were submitted without error and the subsequent wait timed out"* -- and Remarks stating *"If this function returns an error other than IORING_E_WAIT_TIMEOUT, then all entries remain in the submission queue"*, plus that a per-entry failure arrives as a completion rather than as a submit failure. Three things follow. **`Batch::submit_and_wait` had `M21.6`'s defect** at the one site that sweep did not reach: `do_submit` passed a timed-out wait to `check`, reporting a fully successful submission as an error -- and because any *other* error means the entries are still queued, an `Err` was ambiguous between "your buffers are free" and "the kernel still owns them", which is [D-5](#d-5)'s hazard with the sign hidden. **[`IoRing::run_down`](src/ring.rs) was the only waiting API in this crate shaped wrongly**: it waited in 50 ms segments with no period to sit inside, which made "how long to keep trying" this crate's policy rather than its caller's. [`IoRing::run_down_within`](src/ring.rs) is the primitive -- the caller supplies the bound, `Ok(true)` means safe to drop, `Ok(false)` means call again -- and `run_down` is now that with an unbounded period, which is a choice a caller makes by calling it. Note the contrast that makes this precise: [`IoRing::pop_within`](src/ring.rs) also waits in segments and is *correct*, because its segments sit inside a deadline the caller supplied. **An error from rundown is not terminal and says so**: the entries remain queued, the ring is resumable, and the one thing a caller must not do is drop it. **`RS-P-7` and `RS-P-3` moved from inference to citation** -- `M26.1` had written `RS-P-7` as a consequence clause with no authority for submits failing at all, and `M26.3`'s resolver had read it as a permission anyway; the space now carries a `Documented` tag that outranks `Observed`, because a measurement describes one run of one build. | +| D-68 | **A thread-pool wait must be armed before the event it watches is signalled, so the ring's completion event is attached unsignalled and the setup signal is raised after arming.** `M26.9`'s intermittent stall was settled the same way [D-67](#d-67) was -- by reading the reference page rather than by measuring. `SetThreadpoolWait`'s Remarks state *"You must re-register the event with the wait object before signaling it each time to trigger the wait callback."* [`IoRing::completion_event`](src/ring.rs) raises its deliberate setup signal as it attaches, and `EventDelivery::new` then built a `ThreadpoolWait` around the returned duplicate and armed it -- signal first, register second, which is the order the sentence forbids. **Three properties compounded to make a dropped signal permanent rather than late.** The event is auto-reset ([D-21](#d-21)), so a signal is consumed rather than left pending for a later arming to observe. It is edge-triggered on the completion queue going empty to non-empty ([D-19](#d-19)), so a ring whose queue is already non-empty is signalled by nothing else. And the setup signal exists precisely to serve the already-non-empty case, so it is the only wakeup such a ring will ever get. The fix separates the two steps: `attach_completion_event_unsignalled` attaches and reports whether the signal is still owed, `raise_setup_signal` raises it, and `EventDelivery::new` arms in between; `completion_event` is the two composed, unchanged, for a caller doing its own waiting. **The ordering was the defect, so [D-21](#d-21) stands.** A manual-reset event would also have survived the wrong order, by not consuming the signal, but that trades the arming rule for a reset the drain has to get right, and nothing measured here argued for reopening a decision whose own rationale is about the drain. **Measured before and after** on the reproducer recorded in [RESOLVED-TEST-FAILURES.md](RESOLVED-TEST-FAILURES.md): 5 failures in 600 and 2 in 600 for the two co-running triggers, against 0 in 3600 after. The rule itself is now stated where a caller meets it -- on `ThreadpoolWait::arm` and `WaitActivation::rearm` in [windows-threadpool-sys](../windows-threadpool-sys/src/wait.rs), on `IoRing::completion_event` with a worked remedy for a caller who holds the handle, and on `EventDelivery::new` -- because every other arming site in this workspace already had the order right, so what was missing was the statement rather than the practice. | +| D-69 | **A completion may report fewer bytes than requested, and the permission is open because this crate does not constrain the handle type. A consumer that narrows its handle type earns a stronger guarantee, and states it in its own contract.** `M26.10` asked whether a short transfer is possible; the answer has two halves that belong to different layers, and collapsing them is what the question had been deferred over. **The general ring permits it** (`RS-P-8`). `WriteFile`'s Remarks state that *"when writing to a non-blocking, byte-mode pipe handle with insufficient buffer space, WriteFile returns TRUE with \*lpNumberOfBytesWritten < nNumberOfBytesToWrite"*; sockets report a short send against a full transmit buffer, and a communications handle with a write timeout can report a partial count. So the clause is `Documented` rather than `Observed`, which matters because nothing in this repository had measured it and `M26.7` had flagged five assertions quietly depending on the opposite. The reason the permission cannot be narrowed is not that files behave badly -- for an ordinary file on a local volume a successful completion is expected to carry the full length, and a full volume is `ERROR_DISK_FULL` rather than a short success -- but that **`IoRing` takes a handle and never asks what kind it is**. A pipe, a socket and a serial port are all handles, so a consumer of *this* crate has to read the count. **Continuation stays with the caller**, per [D-67](#d-67): the count is reported and nothing reissues the remainder, because how many times to retry and when to give up are policy. A short count may be zero, so a consumer that loops must tolerate making no progress. **The durability layer is the other half.** [examples/epoch_log](examples/epoch_log) makes guarantees -- epoch commits, flush barriers, FUA -- that are only meaningful on a real file on a real volume: a flush barrier means nothing on a socket, and a short write would break "the whole record landed" silently rather than loudly. It therefore has to *constrain* the handle types it accepts, and thereby earn the completeness its accounting already assumes, rather than inherit it from a ring that does not promise it. That constraint is not yet enforced; it is queued as `M26.11` rather than recorded here, because a decision is not a work queue. **The kernel tests were corrected in the honest direction** -- three sites that asserted a full count bare now say that completeness is a property of the temp file they opened, which is the form [flush_barrier.rs](tests/flush_barrier.rs) already used and `M26.7` had singled out as right. | +| D-70 | **The epoch log states the capabilities it requires of a handle and does not check for them. A pre-flight check points the wrong way: it catches the failures that were already loud and misses every failure that is silent.** [D-69](#d-69) left the durability layer owing a handle constraint, and the obvious discharge was a gate -- `GetFileType`, then `GetDriveType` or `FileRemoteProtocolInfo` for the distinctions the first is too coarse to make. Working the gate out is what showed it was not worth having. **Sort the failures by how they present.** A console handle, a pipe, a socket or a closed handle is exactly what `GetFileType` names -- and every one of them already fails at the first positioned write, so the check buys a clearer message and nothing else. A RAM disk, a remote share, or a volume whose write cache is not power-protected **succeeds at every API call this log makes** and silently fails to be durable; `GetDriveType` can name the first two and nothing in Win32 settles the third from a handle, because write-cache state is a property of the device rather than of the handle. So the gate guards the loud cases and misses the silent ones, which is the inverse of what a durability layer needs. **The cost is not the call, it is the claim.** A check that cannot establish the property still reads to a later maintainer as though the property was established -- the "decoration that reads like enforcement" the repository's FAIL FAST rules name -- and that is worse than no check, because it discourages the reader from asking the question themselves. **So the requirement is documented and the caller warrants it.** This is [D-67](#d-67)'s shape applied to handles rather than to retries: we state what we need, the caller chooses what to hand us, and the platform enforces what it is able to. Two requirements are singled out in the contract: that a successful write of `N` bytes transfers `N`, which this log *can* check and does, and that a completed flush reaches stable media, which it cannot check and says so in those words. **Two alternatives were considered and declined**, both recorded so they are not re-proposed as new: refusing `FILE_TYPE_CHAR` and `FILE_TYPE_UNKNOWN` as a cheap early error, declined because it improves only the already-loud path; and a caller-declared mask of intended handle types, whose value would be the acknowledgment rather than the validation, declined as ceremony every ordinary caller pays for. The work of writing the contract is queued as `M26.11`, since a decision is not a work queue. | + ## Durability on the ring Written for consumers, like the two sections that follow it, and for the same reason: the default @@ -583,7 +618,7 @@ Two consequences follow, and they are why this matters beyond terminology: cache and NUMA locality that motivated the whole structure comes from the pinning, not from the per-thread split. This is a configuration people ship by accident. - **The interesting count is cores, or LLC domains -- not threads.** Which is exactly what - [D-8](#d-8) and the L3-domain guidance below already recommend; this is the reason underneath them. + [D-8](#d-8) and the cache-domain guidance below already recommend; this is the reason underneath them. ### The two models are Windows' own two completion mechanisms @@ -631,11 +666,43 @@ It is worse in virtualized deployments, which is where most of this code will ru investigated on reported **zero** `Win32_NumaNode` instances. Any strategy keyed on node must degrade to "one ring" when the answer is unknowable, which is the common case. -A better default heuristic is the **last-level cache domain**: `GetLogicalProcessorInformationEx` with -`RelationCache` filtered to `CacheLevel == 3`. On EPYC that is the CCX/CCD boundary, which has a real -latency cliff even inside a single NPS1 node, because crossing it goes out to the IO die over Infinity -Fabric. It is meaningful on Intel and ARM too, where the NUMA node often is not, and it degrades sanely: a -VM reporting one L3 domain yields one ring, which is correct. +**And it is not only virtualization.** A shipping ARM consumer laptop reports zero `Win32_NumaNode` +instances too, and reports no L3 cache domains at all -- see [D-48](#d-48). The two observations are +siblings, and together they describe the machines most consumers actually have: one where the node is +absent because a hypervisor did not present it, one where it is absent on bare metal. + +A better default heuristic is the **outermost cache level that actually partitions the machine**, which +`MachineMemoryTopology::outermost_partitioning_cache` answers -- defined in +[windows-topology-sys](../windows-topology-sys/README.md), so that every consumer asks rather than +restating. On EPYC that is the L3/CCX boundary, +which has a real latency cliff even inside a single NPS1 node, because crossing it goes out to the IO die +over Infinity Fabric. It degrades sanely: a VM whose caches partition nothing yields one ring, which is +correct. + +**The rule is not "L3", and the level number is not the ordering.** Two measurements forced that wording, +and each falsifies a different half of the old one: + +- A shipping Snapdragon X2 Elite reports **no L3 at all**, with its natural cluster boundary at L2 + ([D-48](#d-48)). So "the last-level cache is L3" is false on a part that ships today. +- The machine this workspace is developed on reports an L3 that spans **all 16 processors** over a real + 8-way L2 partition. So an L3 can exist and still partition nothing -- and a consumer filtering on + `level == 3` there does not degrade, because it matched something. It reports one whole-machine domain + as a successful cache-aware partition. + +That second case is why the rule is stated as a question to ask rather than a level to match: the failure +is silent, and it collapses an eight-domain machine to a single ring while reporting success. The finding +that a cache domain beats the NUMA node is untouched by either. + +**Choosing a specific level is still available; it is just not the default.** What was withdrawn is a +*policy* that matched on a level number, because that policy computed the wrong partition on two of the +three machines above. The underlying capability is untouched: `MachineMemoryTopology::cache_levels` and +`cache_partitions_at_level` let a consumer who knows their part ask about any level directly, and +`examples/cache_domains.rs` prints every level's distinct processor sets beside the heuristic's choice, +so the comparison the old policy got wrong is visible rather than asserted. On the development host that +output is `L1: 8`, `L2: 8` (chosen, checked pairwise disjoint), `L3: 1` -- a consumer can see in one +glance why matching `level == 3` there collapses the machine, and equally that on a part where L3 does +partition, choosing it is theirs to make. A consumer is given the data and the means to decide; what they +are not given is a preset that answers wrongly without saying so. **Processor groups are a hard floor.** A thread's affinity is a `GROUP_AFFINITY` and a ring's waiter lives in exactly one group, so above 64 logical processors the partition is forced whether or not it is wanted. @@ -650,14 +717,47 @@ So `VirtualAllocExNuma` for the pool, on the node closest to the device, registe ring, is very likely the highest-leverage locality decision available -- and it is independent of everything above about completion routing. +That allocation is `NumaBuffer`, which this crate provides as of [D-51](#d-51). *Which* node is still the +caller's answer, per [D-8](#d-8); the type supplies the allocation and decides no policy. + ### What is not reachable -Mapping a **file handle to the NUMA node of the device backing it** has no clean user-mode path. It means -walking volume to disk to device instance and reading `DEVPKEY_Device_Numa_Node`, with real failure modes -(spanned volumes, Storage Spaces, network paths, VHDs) where the question may have no answer. This crate +What is not reachable is the **answer** -- "which ring should this file's I/O go to". The **mechanism** is +reachable, and an earlier version of this section had that wrong: it said the mapping has "no clean +user-mode path" and "means walking volume to disk to device instance and reading +`DEVPKEY_Device_Numa_Node`". There is one documented call, and it takes the handle a caller already holds. + +- **`FSCTL_QUERY_VOLUME_NUMA_INFO`** is documented in the IFS docs, accepts a handle to a **file or + directory** directly, and returns `FSCTL_QUERY_VOLUME_NUMA_INFO_OUTPUT { ULONG NumaNode }`. No device-tree + walk. +- **`GetNumaNodeNumberFromHandle`** is the other path: a Win32 wrapper over `NtQueryInformationFile` with + `FileNumaNodeInformation` (class 53, Windows 7 and later), yielding + `FILE_NUMA_NODE_INFORMATION { USHORT NodeNumber }`. PHNT and the WDK mark that class **reserved for + system use**, so this crate must not build on it. It is named here so the next reader does not rediscover + it and assume it is available. + +Both were observed to succeed on an ordinary NTFS data file, and on a directory handle, and to agree -- +so an ordinary file is not the no-association case. That run was on a single-node host, so neither call is +shown to name a node that distinguishes anything; the run and its limits are recorded in the spikes +[README.md](design-sessions/spikes/README.md), and +[file-handle-numa-spike.rs](design-sessions/spikes/file-handle-numa-spike.rs) is the instrument. Settling +what a multi-node host reports needs storage whose PDO advertises a proximity domain, which is a hardware +gap rather than a deferred decision. + +**The conclusion this section has always drawn survives, on different grounds than it used to rest on.** +What either call returns is the node the *volume* resides on, not where the file's extents live. It is +absent whenever the device layer advertised no proximity domain -- `IoGetDeviceNumaNode` on the PDO, or +`DEVPKEY_Numa_Proximity_Domain` with `GetNumaProximityNode` from user mode. And one volume may sit on +several devices, which is the ordinary case for a spanned volume or a Storage Spaces set. So this crate will not offer an automatic "put this file's I/O on the right ring." It offers "bind a ring to a domain and submit from there," and leaves the mapping to whoever knows their storage layout. +**A sample now exercises the mechanism, which is not the same as the crate adopting it.** The epoch-log +sample asks `FSCTL_QUERY_VOLUME_NUMA_INFO` of its own log handle and places its arena on whatever comes +back ([D-50](#d-50)), so the call above has a runnable consumer rather than living only in a spike. That is +a sample making a local policy choice with the caveats in this section attached to it; the library's +position is unchanged. + ### The practical shape Almost nobody runs pure Model B. What works is hybrid: Model B on the hot data path (pinned threads, @@ -667,13 +767,13 @@ paths, where the thread pool's quiescence is worth more than locality. Both paths are therefore first-class in this crate, which is what D-3 records. -On sizing: one domain per physical core (not per SMT sibling) maximizes isolation; one per L3 domain gives +On sizing: one domain per physical core (not per SMT sibling) maximizes isolation; one per cache domain gives a smaller number of domains that can still share cache-resident state cheaply -- eight rather than sixty-four on a 64-core EPYC. Fewer domains balance load better and duplicate registered buffers less; more isolate better. That is a workload call, and this crate does not make it. `examples/ring_copy` (M7) is where that workload call actually gets made, for exactly one workload: it -implements the `ByL3`/`ByNode`/`ByPackage`/`ByCore`/`Single` policies above as runnable code, over a real +implements the `ByCache`/`ByNode`/`ByPackage`/`ByCore`/`Single` policies above as runnable code, over a real file copy, so the guidance here has something executable behind it rather than staying prose. The policy lives in the sample, not the library (D-8); the library still makes none of these choices for a caller. @@ -731,6 +831,7 @@ evidence rather than left as silence. | `PendingBufferRegistration::claim_if` | `io::Result>` | Take the registration, or observe the failure. | **Sound, with a documented sharp edge.** A *failed* completion is treated as proof the kernel did not retain the addresses, so the buffers are dropped. M16.4 recorded that injecting a synthetic failure here would therefore free memory the kernel genuinely holds -- inert only because nothing does so outside the fault-injection seam. | | `PendingFileRegistration::claim_if` | `io::Result` | Take the registration. | **No hole.** `BuildIoRingRegisterFileHandles` reads its array synchronously ([D-32](#d-32)), so nothing outlives the call that could be invalidated. | | `EventDelivery::scope` (was `ring`) | `RingScope` | Submit work through [`RingScope::batch`], read the ring's read-only state. **Cannot** obtain a `&mut IoRing`, and so cannot replace the ring. | **WAS THE ONE FINDING -- [D-43](#d-43), fixed in M18.6.** The previous `ring -> &Mutex` let safe code assign a whole new ring through the guard, which compiled and silently stopped delivery: measured at one completion delivered before the swap and none after. [D-35](#d-35)'s shape at a different layer, and fixed the same way -- by narrowing the returned type to exactly what the caller needs. Now enforced by a `compile_fail` doctest. | +| `Pending::contract` | `Option<&RingContract>` | Read the oracle's record of what it has observed -- counts and per-operation states. **Cannot** mutate it, and cannot obtain one at all from an unchecked map, which is what the `Option` reports. | **No hole.** `RingContract` is a pure record: it owns no handle, no buffer, and no index into a registration, so there is nothing here the kernel or a registration could invalidate. The borrow is a plain `&self` borrow of the `Pending`, so nothing can be submitted through that `Pending` while it is alive -- every push takes `&mut self`. Added in M23.3 and audited in M26.2, which is late: the gate reported it on the next run rather than on the change that introduced it, because that change did not run the gate. | | `IoRing::completion_event` | `OwnedHandle` | Wait on it, close it, hand it elsewhere. | **No hole.** The returned handle is a *duplicate*; the ring keeps its own, so closing the caller's does not stop the ring signalling ([D-20](#d-20)). Repeat calls duplicate the same event rather than attaching a second. The one real hazard -- two waiters on one ring -- is a documented misuse, not a memory-safety hole. | | `IoRing::try_pop` | `Option` | Read `user_data`, `code`, `result`; use it to claim a token. | **No hole.** `Completion` is a plain value carrying no borrow of ring state. Its power is that it authorises a claim, and that power is bounded by the id/ring checks in `claim_if`. | | `Completion::with_injected_failure` | `Completion` | Rewrite the result of a *real* completion. | **Sound because it transforms rather than fabricates.** Same `user_data` and `ring_id`, so the "a completion exists therefore the kernel is done" argument is untouched. `Completion::synthetic`, which *would* fabricate, is `#[cfg(test)] pub(crate)` for exactly this reason. Behind the off-by-default `fault-injection` feature. | @@ -747,6 +848,35 @@ That is the pattern worth carrying forward into M18.2's recurring rule -- the question that finds these is not "is this correct?" but "what else does this type allow?" +### Four more items, surfaced by widening the check (M21+.1) + +The audit above is a point-in-time pass, and nineteen was its count. These four +are not corrections to it; they are items the *check* could not see, and so +never put to anyone. A 2026-09-21 review of the `M21.2` surface found that +[check-borrow-surface.ps1](../../tools/check-borrow-surface.ps1) inspected only +the text after the last `->` on lines matching `pub fn`, which leaves two shapes +invisible: **a method of a `pub trait`** (declared `fn`, not `pub fn`) and **a +borrow-carrying type in parameter position**. Widening it reported exactly four +entries, answered here before the inventory was regenerated. + +Worth stating plainly: the check was not wrong about what it covered, and the +control still passes -- a `pub fn` returning `&[u8]` was caught throughout. It +was narrow, and nothing said so. + +| Item | Shape the old check missed | What safe code may do with it | Finding | +|---|---|---|---| +| `IoRingErrorExt::as_ioring_error` | trait method, borrow **returned** | Read `code`, `name`, `condition` through a shared reference. | **No hole, and it is the honest catch of the four.** This predates the widening by months and was simply never inventoried, which is the blind spot made concrete rather than a new risk. The borrow is of the `io::Error` the caller already owns -- not of ring state, not of anything the kernel holds -- and `&IoRingError` permits reads only. | +| `CompletionWait::wait` | trait method, borrow **parameter** | Call `RingWait::block` and `RingWait::outstanding`, and nothing else. | **No hole, and the narrowing is the reason.** This is the wider exposure of the two directions: the wrapper goes to arbitrary safe code implementing the trait, not to a known caller. `RingWait` exposes no pop -- one would consume the completion its own caller is waiting for -- and no way to build work. It cannot be retained: the `'ring` lifetime is fresh per call and unconstrained by `Self`. Nothing else can touch the `IoRing` while it is alive, because the pop loop holds `&mut self` across the call. Handing out a bare `&mut IoRing` here would have been [D-43](#d-43) again. | +| `Batch::new` | borrow **parameter** with an explicit lifetime | Nothing it could not already do: the caller supplied the `&mut IoRing`. | **No hole; this is [D-5](#d-5)'s mechanism, not a leak of one.** The exclusive borrow is what makes two concurrent batches fail to compile, which is the point of taking it. The borrow travels *into* the crate and is released when the `Batch` drops. | +| `EventDelivery::new` | borrow **parameter** with an explicit lifetime | Nothing; the callback environment is forwarded to `ThreadpoolWait::new` and not retained. | **No hole.** `EventDelivery` stores only `wait` and `ring`, so the `&mut CallbackEnviron<'_>` does not outlive the call. | + +**A plain `&T` parameter is deliberately not reported**, or the inventory would +list every method in the crate and say nothing. Lending a reference *to* a +callee is the caller's business; what this defect class is about is a +borrow-carrying wrapper whose lifetime the crate chose. The explicit-lifetime +test is what separates the two, and it is a heuristic -- it would miss a +hypothetical `&dyn Trait` parameter carrying no named lifetime. + ## Testing strategy (M18.5) Eight defects came out of the 0.1.x line and the M11-M14 branch. M15 through @@ -760,7 +890,7 @@ them reach at all. | | The defects | What finds them | Built in | |---|---|---|---| -| **A -- preconditions never varied** | [#47](https://github.com/MikeGrier/windows-threadpool-sys/issues/47): every `event_delivery` test handed over a *fresh* ring, so "completion queue non-empty at handover" was never a test input | Generated operation sequences, so the state space is sampled rather than enumerated by hand | M17 | +| **A -- preconditions never varied** | [#47](https://github.com/MikeGrier/windows-threadpool-sys/issues/47): every `event_delivery` test handed over a *fresh* ring, so "completion queue non-empty at handover" was never a test input | Generated operation sequences, so the state space is sampled rather than enumerated by hand. **`M26` added a second technique for this population**, on the observation that *the kernel's response* is a precondition too and was never varied either: a seeded resolver over a specified space | M17, M26 | | **B -- failure paths never taken** | The checkpoint path authorising a reclaim after a failed write; [#48](https://github.com/MikeGrier/windows-threadpool-sys/issues/48) surfacing as a *lucky* `ERROR_NOACCESS` rather than corruption | Deterministic memory instrumentation, an executable contract oracle, and a seam that injects failure into a real completion | M15, M16 | | **C -- permissions, not behaviour** | [D-35](#d-35) (`&mut Vec` permits `reserve`), [D-36](#d-36) (`&B` handed out while the kernel writes), and [D-43](#d-43), found by the audit itself | Review, made recurring by a mechanical trigger; mutation testing for the weaker cousin of the same problem | M18 | @@ -782,6 +912,7 @@ hoped-for ones. | Generated sequences (M17) | **No new defects.** Found an API precondition the generator was violating, and -- once calibrated -- rediscovers #47 in 10 of 10 runs | | Borrow-surface audit (M18.1) | **One defect: [D-43](#d-43)**, in 19 items audited. Same shape as the two that prompted the audit | | Mutation testing (M18.3/4/7) | **A third vacuous test** four review rounds had read past, plus 36 further weak assertions. 79.7% to 95.8% | +| Seeded resolution over a specified space (M26) | **No defect in shipping code yet**, and the calibration is why that is reportable rather than reassuring: re-injecting `M21.6`'s expired-wait defect turns it red, and so does narrowing the resolver itself. It did find one live defect on its first contact with a real ring -- a declined submit propagating out of `IoRing::run_down` with work still outstanding, queued as `M26.8` | Two observations that only appear once the table is read as a whole. @@ -803,13 +934,34 @@ evidence.** Budget the calibration, not just the instrument. ### What none of them cover -All five techniques check this crate's code against **this crate's stated -contract**. None of them can tell you the stated contract is wrong -- and in the -two most expensive defects, that is exactly what happened. The completion event -is edge-triggered ([D-19](#d-19)) and `BuildIoRingRegisterBuffers` reads its -array when the operation runs rather than when `Build*` returns -([D-32](#d-32)). Both were discovered by a spike against the real kernel, and -neither could have come from anywhere else in this toolkit. +**Five of the six check this crate's code against this crate's stated +contract.** None of those five can tell you the stated contract is wrong -- and +in the two most expensive defects, that is exactly what happened. The +completion event is edge-triggered ([D-19](#d-19)) and +`BuildIoRingRegisterBuffers` reads its array when the operation runs rather +than when `Build*` returns ([D-32](#d-32)). Both were discovered by a spike +against the real kernel, and neither could have come from anywhere else in that +toolkit. + +**The sixth is different in kind, and `M26.6` is what makes the difference +real.** The resolver tests this crate against a *written specification of what +the platform may do* ([RESPONSE-SPACE.md](RESPONSE-SPACE.md)) rather than +against our beliefs about what it does -- so it catches code that is brittle to +platform variation inside the permitted space, which is a class the other five +cannot reach. It still cannot tell you the specification is wrong. What can is +the **other half of the pair**: the kernel tests now confirm that a real Windows +stays *inside* the declared space, and a kernel observed outside it is a finding +about the platform rather than a regression in this crate. `RS-C-4` is the +clause that makes that job load-bearing rather than nominal, since the resolver +is forbidden to break the drain half of `DRAIN_PRECEDING_OPS` and therefore +cannot be what notices if Windows does. The division is enforced by +[response_space_census.rs](tests/response_space_census.rs), which fails when a +clause is claimed by nothing on the side that owes it a check. + +So the honest statement is narrower than "a spike is the only way to learn the +platform is not what we assumed", and it is still true: a spike remains the only +technique that runs **before there is any code to test**, and the space itself +was written from what spikes established. Be precise about the failure mode, because "the allocator would not have caught D-32" is not quite true and the imprecision matters. The guard allocator *does* @@ -832,7 +984,11 @@ their shape are in they established is summarised under [What the spike established](#what-the-spike-established). -### Two techniques deliberately rejected +### Two techniques deliberately rejected + +**The first was re-examined by [D-49](#d-49) / `M24.1` and now **stands, with its scope sharpened** +by [D-52](#d-52). The re-examination did not find an escape; it found that the instrument was the +wrong one.** Recorded so they are not re-proposed as obvious wins. @@ -843,6 +999,21 @@ merely have failed to find them, it would have manufactured evidence they were absent. A model belongs here as an **oracle over observed sequences** ([`RingContract`](src/contract.rs)), never as a substitute for the kernel. +> **What `M24.1` demonstrated, rather than argued** (2026-09-22, the apparatus is +> [kernel-response-space-probe.rs](design-sessions/kernel-response-space-probe.rs)). Sharing a +> suite's assertions between a fake and the kernel does **not** rescue a mock. A fake built from the +> belief this crate held before `M21.6` -- that an expired wait is an error -- passed the shared +> suite green; the assertion that catches it could only be written after the kernel had already +> revealed the answer. Running the *wrong* assertion against both is the one case that helps: the +> kernel refutes it while the fake confirms it, which is the manufactured-evidence mechanism made +> visible. +> +> And the line is not "accounting versus Windows behaviour", which was the first answer. It is **our +> specified contract versus the platform's incidental behaviour**. "After submitting, the completion +> is already queued" reads like a contract and gave *opposite answers on two handles of the same +> API*; the same question stated as this crate's own contract -- "arrives within a bound we specify" +> -- holds everywhere. See [D-52](#d-52) for what replaces the technique. + **Application Verifier / PageHeap**, rejected after measuring rather than assuming ([D-37](#d-37)): it works, and needs no SDK, but it is keyed by *image file name* and cargo rehashes test binaries on every meaningful rebuild -- so diff --git a/crates/windows-ioring-sys/PLANS.md b/crates/windows-ioring-sys/PLANS.md index cd609067d..d81c77194 100644 --- a/crates/windows-ioring-sys/PLANS.md +++ b/crates/windows-ioring-sys/PLANS.md @@ -6,4 +6,4 @@ contained are archived in [COMPLETED-CHECKLIST.md](COMPLETED-CHECKLIST.md). Desi | Path to CHECKLIST.md | Status | Brief description | Design Notes | |---|---|---|---| -| [CHECKLIST.md](CHECKLIST.md) | in progress | Memory-safe Rust over the Windows 11 / Server 2022 `IoRing` submission/completion ring, as a separate crate from `windows-overlapped-io-sys` (duplicate-then-decide). Covers ring lifecycle and capability negotiation, zero-allocation token-owned buffers, the batch submission builder, threadless delivery through `ThreadpoolWait`, file/buffer registration, consumer-facing documentation, and the `ring-copy` topology-aligned sample (M1-M7 archived). The pinned-thread (Model B) architecture remains parked as `M6+` by the engineer's explicit direction. `M8` (complete) closed a PR #20 review finding: `FileRef::Raw(HANDLE)`'s lifetime gap, fixed with `unsafe fn` raw entry points plus a safe, `Arc`-backed `SharedFile` wrapper for the common case. `M9` (complete) closed further PR #20 review findings: cross-ring `Token`/`RegisteredFile`/`RegisteredBuffers` confusion (a new per-ring `RingId`, checked at claim/push time), `PendingBufferRegistration` freeing its buffers instead of leaking them on an unclaimed drop, and `Batch::do_submit` letting `Drop` silently retry an already-attempted, already-failed submit. `M10` (active) finished auditing the ring completion contract against all ten specification-gap categories (M10.1-M10.3 complete): category 3 found that `supports` answers for the kernel's op table rather than this crate's push surface and that the registration one-shot is spent by queueing rather than succeeding (D-28); categories 1, 2, 6, 8 and 9 established the load-bearing rule that **every successfully queued SQE produces exactly one completion**, plus that this crate deliberately joins nothing (D-29, D-30); and D-14's registration-index continuity assumption was **dissolved rather than measured** -- the collision it guarded against needs a second registration, which was forbidden the day after D-14 was written, leaving only the reserved-not-confirmed meaning of the public counts to state (D-31). The audit also surfaced two API gaps, now queued as work rather than left in the design notes: `M10.4` (complete) gave `FileRef::Registered` safe entry points, since a registered index carries no lifetime obligation and the `unsafe` guarding it was vacuous (D-29): the safe pushes are now generic over a sealed `FileTarget` trait whose associated `Guard` type carries the one real difference between the two targets, which also made the fully-registered (registered file *and* registered buffer) combination expressible for the first time, non-breakingly (D-33). Investigating M10's own recorded test failures then found a **live use-after-free in shipped 0.1.2** and fixed it as `M10.6`: `BuildIoRingRegisterBuffers` reads its `IORING_BUFFER_INFO` array when the op runs rather than at build time -- the opposite of its file-handle sibling, which the rustdoc had wrongly generalized across -- so the array is now owned by the `IoRing` (D-32). `M10.5` (complete) added named predicates for the conditions a consumer must branch on -- `IORING_E_SUBMISSION_QUEUE_FULL` above all, which every push's rustdoc names as the backpressure signal but which `io::Error::kind()` cannot discriminate (D-30): a complete `RingCondition` enum, predicates for the runtime-actionable conditions, and a sealed `IoRingErrorExt` that puts them on `io::Error` so the downcast is named once rather than hand-rolled per call site (D-34). **M10 is complete.** **`M11` is complete and archived** (2026-08-28): it made the completion event a ring primitive, prompted by an external consumer proposal. `IoRing::completion_event` returns an owned duplicate of the ring's own event so a caller can wait on the ring alongside other handles without surrendering it (D-20); its contract is pinned by eleven sabotage-verified tests; `EventDelivery` is re-expressed on top of it, leaving one `SetIoRingCompletionEvent` call site; `windows-threadpool-sys` moved behind a default-on `threadpool` feature with CI building both combinations (D-22); the wakeup shapes and the barrier's ring-edge limit were swept across every place that states them; and `examples/model_b_multiplexed.rs` works the multiplexed shape end to end. The spike that answered the proposal established that the completion event is **edge-triggered** on the completion queue going empty to non-empty (D-19) -- which also exposed a live bug in shipped 0.1.2, where `EventDelivery` permanently stranded completions queued before handover, fixed in M11.3 by the same change that consolidated `EventDelivery` onto the new primitive. **`M12` is complete and archived** (2026-08-28): it addressed durability, from the same exchange. A spike had established that a flush **without** `DRAIN_PRECEDING_OPS` does not cover preceding writes (D-23) while the barrier that fixes it drains everything already outstanding on the ring (D-24, as corrected by D-47), making `Batch::flush` with default options a silent data-loss bug rather than a missing feature. `Batch::flush`/`flush_raw` now require an explicit `FlushCoverage`, so that spelling no longer exists; a `NO_BUFFERING` integration test proves the barrier's behaviour rather than its flag, and measured that *which direction* the reordering shows in is device-dependent (amending D-23); and the parameters the crate had hardcoded away are exposed as `WriteCaching` (`FILE_WRITE_FLAGS`) and `FlushMode` (`FILE_FLUSH_MODE`, whose `NoSync` is the one mode that commits nothing). Durability had been absent from `lib.rs` and `README.md` entirely, and both now state the three facts. **`M13` is complete and archived** (2026-08-29): the `epoch_log` sample is the vehicle D-26 makes for carrying durability *policy* to consumers without this crate owning it. Its own durability contract is written down first, in its own words (Design Autonomy), then implemented -- records composed into a registered arena and appended, epochs closed by one covering flush whose completion is what makes `is_durable` answer `true`, and a multiplexed wait on the ring's completion event alongside a shutdown latch. The replay pass is what turns it from a demonstration into evidence: it holds the durable region to a strict standard, tolerates a torn tail as the contract requires, and is itself proved able to fail by a negative control. Writing it also found and fixed a gap in the crate (D-35: per-buffer outstanding accounting and `RegisteredBuffers::get_mut`). **`M14` is complete and archived** (2026-08-29): the second half of the sample, covering the two things the ring cannot do for a consumer. A non-ring `FSCTL` (`FSCTL_SET_ZERO_DATA`, reclaiming a retired segment) is ordered against ring epochs by the log itself, since `drain_preceding` orders SQEs against SQEs (D-24) and reaches across neither the ring boundary nor a second ring; a thread-pool control plane runs checkpointing as Model A on a *second* ring while the log thread keeps Model B for the data path, because D-21 forbids a ring given to `EventDelivery` also being waited on directly -- so the ordering chain crosses log thread to pool thread to reclaim worker and back with the log thread blocking for none of it. All three epoch-commit strategies are implemented behind one interface, checked both by replay and by requiring the three to be **byte-identical**, and measured on the running machine. The measurement's finding is that the three are **indistinguishable** here -- the cross-strategy spread is the size of one strategy's run-to-run spread, because every strategy pays one device flush per epoch at hundreds of microseconds while their real differences land in the tens -- and it found two harness bugs nothing else caught, including one where a barrier benchmark that awaits each commit before appending again measures the barrier as free. Both findings are promoted into "Durability on the ring" in [DESIGN-NOTES.md](DESIGN-NOTES.md), since both are about the design rather than the demonstration. **`M15`-`M18` (not started) are the testing-strategy response to the eight defects the 0.1.x line and the M11-M14 branch produced.** They are organised by *defect population* rather than by technique, because the populations need different tools and one of them needs a tool that does not exist: (A) preconditions never varied -- every `event_delivery` test handed over a fresh ring, which is why [#47](https://github.com/MikeGrier/windows-threadpool-sys/issues/47) survived; (B) failure paths never taken -- no test ever ran `completion.result()` returning `Err`, which is why the checkpoint path could authorise a reclaim after a failed write; and (C) *permissions rather than behaviour* -- `&mut Vec` permits `reserve`/`resize`/reassign though no code path performs it, which is [D-35](DESIGN-NOTES.md#d-35) and [D-36](DESIGN-NOTES.md#d-36), the two most severe findings, and **no runtime technique reaches that population at all**. M15 and M16 gate the 0.2.0 release. M15 is deterministic memory instrumentation: a guard-page global allocator, chosen over Application Verifier / PageHeap by measurement rather than assumption ([D-37](DESIGN-NOTES.md#d-37) -- PageHeap works and `reg add` alone is enough, but IFEO is keyed by image file name and one test target produced six distinct hashed names in a day), plus a tracked poison pattern covering the gap guard pages structurally cannot see ([D-38](DESIGN-NOTES.md#d-38) -- a guard page catches access to memory that should not be touched, and is blind to the kernel writing into a live, valid buffer, which is exactly what `write_registered` and `read_registered` promise it will not do). M16 makes the contract executable: a public `RingContract` rendering [the category-2 rule](DESIGN-NOTES.md#one-sqe-one-completion), "one SQE, exactly one completion", checkable rather than merely stated, plus a fault-injection seam that finally takes the failure paths. M17 covers A by generating over the operation space instead of enumerating it by hand, gated on an explicit decision about randomized sampling that this component's conventions require be approved and recorded rather than assumed. M18 covered C, where review is a *primary* technique rather than a backstop, and added `cargo-mutants` against a measured rate of vacuous tests. **M8 through M18 are now complete and archived**; `M20` is pending, and `M6+` is parked rather than pending. The strategy as a whole -- which technique reaches which population, what each one actually found, and what none of them reach -- is recorded in [DESIGN-NOTES.md](DESIGN-NOTES.md#testing-strategy-m185); mutation coverage went from 79.7% to 95.8%, and the borrow-surface audit found one further defect of the same shape as the two that prompted it ([D-43](DESIGN-NOTES.md#d-43)). A mock `IoRing` was considered and **rejected**: both shipped defects were the kernel behaving differently from this crate's assumptions, so a mock would have encoded the same assumptions and passed both bugs green. **M20** queues the documentation and policy-test repairs from the 2026-08-30 NUMA-sharding measurement: a shipping ARM laptop reports no L3 cache domain at all, which falsifies the justification given for the last-level-cache heuristic (though not the heuristic's preference over the NUMA node, and not `ring_copy`, whose degraded fallback already handles it correctly). | [DESIGN-NOTES.md](DESIGN-NOTES.md), [DESIGN-SESSION-2026-08-28-completion-event-multiplexing.md](design-sessions/DESIGN-SESSION-2026-08-28-completion-event-multiplexing.md), [DESIGN-SESSION-2026-08-28-external-consumer-correspondence.md](design-sessions/DESIGN-SESSION-2026-08-28-external-consumer-correspondence.md), [DESIGN-SESSION-2026-08-30-numa-sharded-io-execution-domains.md](../../design-sessions/DESIGN-SESSION-2026-08-30-numa-sharded-io-execution-domains.md) | +| [CHECKLIST.md](CHECKLIST.md) | in progress | Memory-safe Rust over the Windows 11 / Server 2022 `IoRing` submission/completion ring, as a separate crate from `windows-overlapped-io-sys` (duplicate-then-decide). Covers ring lifecycle and capability negotiation, zero-allocation token-owned buffers, the batch submission builder, threadless delivery through `ThreadpoolWait`, file/buffer registration, consumer-facing documentation, and the `ring-copy` topology-aligned sample (M1-M7 archived). The pinned-thread (Model B) architecture remains parked as `M6+` by the engineer's explicit direction. `M8` (complete) closed a PR #20 review finding: `FileRef::Raw(HANDLE)`'s lifetime gap, fixed with `unsafe fn` raw entry points plus a safe, `Arc`-backed `SharedFile` wrapper for the common case. `M9` (complete) closed further PR #20 review findings: cross-ring `Token`/`RegisteredFile`/`RegisteredBuffers` confusion (a new per-ring `RingId`, checked at claim/push time), `PendingBufferRegistration` freeing its buffers instead of leaking them on an unclaimed drop, and `Batch::do_submit` letting `Drop` silently retry an already-attempted, already-failed submit. `M10` (active) finished auditing the ring completion contract against all ten specification-gap categories (M10.1-M10.3 complete): category 3 found that `supports` answers for the kernel's op table rather than this crate's push surface and that the registration one-shot is spent by queueing rather than succeeding (D-28); categories 1, 2, 6, 8 and 9 established the load-bearing rule that **every successfully queued SQE produces exactly one completion**, plus that this crate deliberately joins nothing (D-29, D-30); and D-14's registration-index continuity assumption was **dissolved rather than measured** -- the collision it guarded against needs a second registration, which was forbidden the day after D-14 was written, leaving only the reserved-not-confirmed meaning of the public counts to state (D-31). The audit also surfaced two API gaps, now queued as work rather than left in the design notes: `M10.4` (complete) gave `FileRef::Registered` safe entry points, since a registered index carries no lifetime obligation and the `unsafe` guarding it was vacuous (D-29): the safe pushes are now generic over a sealed `FileTarget` trait whose associated `Guard` type carries the one real difference between the two targets, which also made the fully-registered (registered file *and* registered buffer) combination expressible for the first time, non-breakingly (D-33). Investigating M10's own recorded test failures then found a **live use-after-free in shipped 0.1.2** and fixed it as `M10.6`: `BuildIoRingRegisterBuffers` reads its `IORING_BUFFER_INFO` array when the op runs rather than at build time -- the opposite of its file-handle sibling, which the rustdoc had wrongly generalized across -- so the array is now owned by the `IoRing` (D-32). `M10.5` (complete) added named predicates for the conditions a consumer must branch on -- `IORING_E_SUBMISSION_QUEUE_FULL` above all, which every push's rustdoc names as the backpressure signal but which `io::Error::kind()` cannot discriminate (D-30): a complete `RingCondition` enum, predicates for the runtime-actionable conditions, and a sealed `IoRingErrorExt` that puts them on `io::Error` so the downcast is named once rather than hand-rolled per call site (D-34). **M10 is complete.** **`M11` is complete and archived** (2026-08-28): it made the completion event a ring primitive, prompted by an external consumer proposal. `IoRing::completion_event` returns an owned duplicate of the ring's own event so a caller can wait on the ring alongside other handles without surrendering it (D-20); its contract is pinned by eleven sabotage-verified tests; `EventDelivery` is re-expressed on top of it, leaving one `SetIoRingCompletionEvent` call site; `windows-threadpool-sys` moved behind a default-on `threadpool` feature with CI building both combinations (D-22); the wakeup shapes and the barrier's ring-edge limit were swept across every place that states them; and `examples/model_b_multiplexed.rs` works the multiplexed shape end to end. The spike that answered the proposal established that the completion event is **edge-triggered** on the completion queue going empty to non-empty (D-19) -- which also exposed a live bug in shipped 0.1.2, where `EventDelivery` permanently stranded completions queued before handover, fixed in M11.3 by the same change that consolidated `EventDelivery` onto the new primitive. **`M12` is complete and archived** (2026-08-28): it addressed durability, from the same exchange. A spike had established that a flush **without** `DRAIN_PRECEDING_OPS` does not cover preceding writes (D-23) while the barrier that fixes it drains everything already outstanding on the ring (D-24, as corrected by D-47), making `Batch::flush` with default options a silent data-loss bug rather than a missing feature. `Batch::flush`/`flush_raw` now require an explicit `FlushCoverage`, so that spelling no longer exists; a `NO_BUFFERING` integration test proves the barrier's behaviour rather than its flag, and measured that *which direction* the reordering shows in is device-dependent (amending D-23); and the parameters the crate had hardcoded away are exposed as `WriteCaching` (`FILE_WRITE_FLAGS`) and `FlushMode` (`FILE_FLUSH_MODE`, whose `NoSync` is the one mode that commits nothing). Durability had been absent from `lib.rs` and `README.md` entirely, and both now state the three facts. **`M13` is complete and archived** (2026-08-29): the `epoch_log` sample is the vehicle D-26 makes for carrying durability *policy* to consumers without this crate owning it. Its own durability contract is written down first, in its own words (Design Autonomy), then implemented -- records composed into a registered arena and appended, epochs closed by one covering flush whose completion is what makes `is_durable` answer `true`, and a multiplexed wait on the ring's completion event alongside a shutdown latch. The replay pass is what turns it from a demonstration into evidence: it holds the durable region to a strict standard, tolerates a torn tail as the contract requires, and is itself proved able to fail by a negative control. Writing it also found and fixed a gap in the crate (D-35: per-buffer outstanding accounting and `RegisteredBuffers::get_mut`). **`M14` is complete and archived** (2026-08-29): the second half of the sample, covering the two things the ring cannot do for a consumer. A non-ring `FSCTL` (`FSCTL_SET_ZERO_DATA`, reclaiming a retired segment) is ordered against ring epochs by the log itself, since `drain_preceding` orders SQEs against SQEs (D-24) and reaches across neither the ring boundary nor a second ring; a thread-pool control plane runs checkpointing as Model A on a *second* ring while the log thread keeps Model B for the data path, because D-21 forbids a ring given to `EventDelivery` also being waited on directly -- so the ordering chain crosses log thread to pool thread to reclaim worker and back with the log thread blocking for none of it. All three epoch-commit strategies are implemented behind one interface, checked both by replay and by requiring the three to be **byte-identical**, and measured on the running machine. The measurement's finding is that the three are **indistinguishable** here -- the cross-strategy spread is the size of one strategy's run-to-run spread, because every strategy pays one device flush per epoch at hundreds of microseconds while their real differences land in the tens -- and it found two harness bugs nothing else caught, including one where a barrier benchmark that awaits each commit before appending again measures the barrier as free. Both findings are promoted into "Durability on the ring" in [DESIGN-NOTES.md](DESIGN-NOTES.md), since both are about the design rather than the demonstration. **`M15`-`M18` (not started) are the testing-strategy response to the eight defects the 0.1.x line and the M11-M14 branch produced.** They are organised by *defect population* rather than by technique, because the populations need different tools and one of them needs a tool that does not exist: (A) preconditions never varied -- every `event_delivery` test handed over a fresh ring, which is why [#47](https://github.com/MikeGrier/windows-threadpool-sys/issues/47) survived; (B) failure paths never taken -- no test ever ran `completion.result()` returning `Err`, which is why the checkpoint path could authorise a reclaim after a failed write; and (C) *permissions rather than behaviour* -- `&mut Vec` permits `reserve`/`resize`/reassign though no code path performs it, which is [D-35](DESIGN-NOTES.md#d-35) and [D-36](DESIGN-NOTES.md#d-36), the two most severe findings, and **no runtime technique reaches that population at all**. M15 and M16 gate the 0.2.0 release. M15 is deterministic memory instrumentation: a guard-page global allocator, chosen over Application Verifier / PageHeap by measurement rather than assumption ([D-37](DESIGN-NOTES.md#d-37) -- PageHeap works and `reg add` alone is enough, but IFEO is keyed by image file name and one test target produced six distinct hashed names in a day), plus a tracked poison pattern covering the gap guard pages structurally cannot see ([D-38](DESIGN-NOTES.md#d-38) -- a guard page catches access to memory that should not be touched, and is blind to the kernel writing into a live, valid buffer, which is exactly what `write_registered` and `read_registered` promise it will not do). M16 makes the contract executable: a public `RingContract` rendering [the category-2 rule](DESIGN-NOTES.md#one-sqe-one-completion), "one SQE, exactly one completion", checkable rather than merely stated, plus a fault-injection seam that finally takes the failure paths. M17 covers A by generating over the operation space instead of enumerating it by hand, gated on an explicit decision about randomized sampling that this component's conventions require be approved and recorded rather than assumed. M18 covered C, where review is a *primary* technique rather than a backstop, and added `cargo-mutants` against a measured rate of vacuous tests. **M8 through M18 are now complete and archived**; `M20` is pending, and `M6+` is parked rather than pending. The strategy as a whole -- which technique reaches which population, what each one actually found, and what none of them reach -- is recorded in [DESIGN-NOTES.md](DESIGN-NOTES.md#testing-strategy-m185); mutation coverage went from 79.7% to 95.8%, and the borrow-surface audit found one further defect of the same shape as the two that prompted it ([D-43](DESIGN-NOTES.md#d-43)). A mock `IoRing` was considered and **rejected**: both shipped defects were the kernel behaving differently from this crate's assumptions, so a mock would have encoded the same assumptions and passed both bugs green. **M20** queues the documentation and policy-test repairs from the 2026-08-30 NUMA-sharding measurement: a shipping ARM laptop reports no L3 cache domain at all, which falsifies the justification given for the last-level-cache heuristic (though not the heuristic's preference over the NUMA node, and not `ring_copy`, whose degraded fallback already handles it correctly). **M27 (added 2026-09-23) was queued by intent rather than by a review or a measurement, and re-planned the same day.** It began as an adaptivity question for this crate; the adaptivity the workspace's [adoption thesis](../../DESIGN-NOTES.md#the-adoption-thesis) asks for is owned by [topology-planner](../topology-planner/COMPONENT.md) instead ([EP-D-6](../topology-planner/DESIGN-NOTES.md#ep-d-6)), and answering it here would have grown a second policy surface beside it. What survives is the realization end: M27.1 censuses what a realizer needs from this crate against the plan vocabulary and names the gaps, M27.2 closes them as capability with no policy attached, and M27.3 -- ungated -- gives a consumer the means to answer placement questions on their own hardware. [D-8](DESIGN-NOTES.md#d-8) is untouched: being constructible from a policy decision made elsewhere is the opposite of taking one. | [DESIGN-NOTES.md](DESIGN-NOTES.md), [DESIGN-SESSION-2026-08-28-completion-event-multiplexing.md](design-sessions/DESIGN-SESSION-2026-08-28-completion-event-multiplexing.md), [DESIGN-SESSION-2026-08-28-external-consumer-correspondence.md](design-sessions/DESIGN-SESSION-2026-08-28-external-consumer-correspondence.md), [DESIGN-SESSION-2026-08-30-numa-sharded-io-execution-domains.md](../../design-sessions/DESIGN-SESSION-2026-08-30-numa-sharded-io-execution-domains.md) | diff --git a/crates/windows-ioring-sys/README.md b/crates/windows-ioring-sys/README.md index d20e892b5..ec853f30a 100644 --- a/crates/windows-ioring-sys/README.md +++ b/crates/windows-ioring-sys/README.md @@ -218,11 +218,15 @@ This crate does not partition anything for you (D-8): it makes a ring cheap and correct, makes its affinity explicit, and leaves sizing a Model B execution domain to the caller. -- **Size a domain by last-level (L3) cache, not by NUMA node.** Node count is a +- **Size a domain by the outermost cache level that partitions the machine, not by NUMA node.** Node count is a firmware setting a process cannot see, and most real deployments are - virtualized, where NUMA topology is often invisible entirely. See - [examples/l3_domains.rs](examples/l3_domains.rs) for a runnable enumeration, - built on the safe `GetLogicalProcessorInformationEx` wrapper in + virtualized, where NUMA topology is often invisible entirely. Ask + `outermost_partitioning_cache()` rather than filtering on `CacheLevel == 3`: + a shipping ARM part reports no L3 at all, and a machine can report an L3 + spanning every processor above a real L2 partition, where a level filter + returns one whole-machine domain and calls it a cache-aware partition. See + [examples/cache_domains.rs](examples/cache_domains.rs) for a runnable + enumeration, built on the safe `GetLogicalProcessorInformationEx` wrapper in [`windows-topology-sys`](../windows-topology-sys/README.md). - **Processor groups are a hard floor.** A thread's affinity is a `GROUP_AFFINITY` and a ring's waiter lives in exactly one group, so above 64 @@ -231,15 +235,16 @@ domain to the caller. on the node closest to the device, registered once into that domain's ring via `Batch::register_buffers`, is very likely the highest-leverage locality decision available -- independent of everything above about completion - routing. + routing. `NumaBuffer` is that allocation; *which* node is still the caller's + answer, and this crate does not guess it. `examples/ring_copy` is where these three points become runnable policy: it copies one file to another through per-domain rings, sized by a named -`ByL3`/`ByNode`/`ByPackage`/`ByCore`/`Single` policy, with buffers placed via -`VirtualAllocExNuma` and a `--placement local|remote` switch to make the -placement effect measurable. It is a **sample**, not library surface -- this -crate itself depends on no partitioning policy and does not depend on -`windows-topology-sys`; only the sample does. +`ByCache`/`ByNode`/`ByPackage`/`ByCore`/`Single` policy, with placed buffers and a +`--placement local|remote` switch to make the placement effect measurable. It +is a **sample**, not library surface -- this crate itself depends on no +partitioning policy and does not depend on `windows-topology-sys`; only the +sample does. ## License diff --git a/crates/windows-ioring-sys/RESOLVED-TEST-FAILURES.md b/crates/windows-ioring-sys/RESOLVED-TEST-FAILURES.md index 888c86a65..519b31775 100644 --- a/crates/windows-ioring-sys/RESOLVED-TEST-FAILURES.md +++ b/crates/windows-ioring-sys/RESOLVED-TEST-FAILURES.md @@ -89,3 +89,239 @@ since a caller may have relied on it. characterised before it is stabilised. The two natural repairs -- loosen the assertion, or mark the test serial -- would both have suppressed the only evidence that shipped documentation was wrong, and "flaky test" and "the platform does not do what we wrote down" produce the same symptom. + +## Resolved 2026-09-21 22:08:22 -04:00 -- the unidentified ests/bounded_pop.rs failure + +**Recorded and resolved the same day, which is the honest framing:** it was recorded as an unresolved +failure on the grounds that fixing it did not belong in a push of finished milestones. That was a +scheduling preference dressed as a blocker. The mechanism and the fix were both understood at the time of +recording; nothing was actually blocking. + +**The failure.** One `cargo test --all-features` run reported `FAILED: 4 passed; 1 failed` in a 2.12s +target matching [bounded_pop.rs](tests/bounded_pop.rs) by shape. The test name and panic message were not +captured. It never reproduced -- 15 isolated runs and 4 full-suite runs were green. + +**The cause, and why no amount of re-running would have settled it.** Those tests needed an operation +still pending when a short bound expired, and got it from a 128 MiB unbuffered, overlapped read. That is +a *margin*, not a guarantee: `FILE_FLAG_NO_BUFFERING` bypasses the system cache but not the drive\'s own, +so the test was asking "will this device take longer than 5 ms?" -- a question about someone else\'s +hardware, whose answer may differ between two runs on the same machine. + +**The fix: an operation that cannot complete, rather than one that is merely slow.** The read is now +issued against an **overlapped named pipe that nobody has written to**. It is pending because no byte +exists to satisfy it, and it completes exactly when the test writes one. There is no device, no cache and +no margin in the question. + +Confirmed by probe before being adopted, since neither half was safe to assume: `IoRing` does accept a +pipe handle for `read_raw`, and `pop_within(20ms)` against an unwritten pipe returns `Ok(None)` with +`outstanding == 1`. + +Where a delay is genuinely needed -- `run_down` polls in 50 ms steps, so forcing it to observe an expired +poll means releasing the read later than that -- it comes from a `thread::sleep`, whose guarantee runs the +safe way round: a sleep may overshoot, never undershoot. No assertion depends on an operation *finishing* +within any bound. + +**Verified:** 25 consecutive runs of the target and 3 full `--all-features` suite runs, all green. Both +sabotages still bite exactly as before the rewrite -- reverting the timeout mapping turns all 5 red, and +making `RingWait::block` always fail turns 3 red -- so the rewrite kept every bit of the discriminating +power it had. + +It is also **11x faster** (0.20s against 2.26s) and allocates no 128 MiB fixtures, which was the larger +part of what the file cost to run. + +## Resolved 2026-09-25 14:19:35 -04:00 -- event_delivery stalled because the wait was armed after the event was signalled + +`SetThreadpoolWait` documents that "you must re-register the event with the wait object before +signaling it each time to trigger the wait callback". `EventDelivery::new` did the reverse: it took an +already-signalled event from `IoRing::completion_event` and armed its wait afterwards. + +Three properties compounded to make the dropped signal permanent rather than late. The event is +auto-reset, so the signal was consumed rather than left pending for the arming to observe. It is +edge-triggered on the completion queue going empty to non-empty, so a ring whose queue was already +non-empty was signalled by nothing else. And the setup signal exists to serve exactly that case, so it +was the only wakeup such a ring would ever get. + +Fixed by attaching the event unsignalled and raising the setup signal after arming; see +[DESIGN-NOTES.md](DESIGN-NOTES.md) -> D-68. Measured on the reproducer recorded below: 0 failures in +3600 runs after the change, against 5 in 600 and 2 in 600 for the two co-running triggers before it. + +The investigation as it stood when the cause was found follows. + +### event_delivery's threadpool tests time out at roughly one run in eighty + +**Found 2026-09-24**, while widening the seeded sweeps to 2048. + +**What happens.** `completions_are_delivered_on_pool_threads_without_the_submitting_thread_waiting` +and `completions_queued_before_handover_are_still_delivered` in +[event_delivery.rs](tests/event_delivery.rs) each wait on a channel with +`recv_timeout(Duration::from_secs(5))` for a completion delivered through the Windows thread pool. +Occasionally the completion does not arrive inside that bound and the test panics with `Timeout`. + +**Measured rather than estimated**, because the rate is the whole point: **1 failure in 80 +consecutive runs** of the compiled test binary when first found. It resisted every targeted attempt +to provoke it at that stage -- zero failures after roughly 18,000 ring create/close cycles, after +repeated property-suite and calibration runs, and under a concurrent `cargo build` saturating the +machine -- so it was neither ring-resource pressure nor CPU load. The narrowing below found what it +actually needs. + +### Narrowed 2026-09-25: it requires parallel test execution, and a co-running create-and-drop + +The instrumentation described further down paid for itself immediately. One captured occurrence plus +four follow-up experiments moved this from "cause unknown" to a minimal reproducer. Every figure +here comes from running the compiled `event_delivery` binary directly. + +**It does not happen serially.** With `--test-threads 1`: **0 failures in 1000 runs**. In parallel: +**7 in 1000**. At the parallel rate a thousand serial runs would expect about seven, so zero is +evidence rather than a quiet stretch. + +**Every occurrence is identical**, across all seven captures: + +- **both** delivery tests fail in the same process, never just one; +- `callbacks run: 0` -- the pool never invoked the callback, not once, for either ring; +- `delivered: 0 of 8` and `outstanding: 8` -- nothing was ever popped; +- the post-mortem finds nothing after a further ten seconds. + +So it is **not** a slow device and **not** a single lost wakeup. No callback runs at all, for both +rings, from the start, and the delivery never arrives. + +**The two delivery tests alone do not cause it**: 0 failures in 1000 runs with a filter selecting +only those two. A third test has to be running. Adding them one at a time, 600 runs each: + +| Co-running test | Failures in 600 | +|---|---| +| `dropping_with_nothing_outstanding_does_not_hang` | 5 | +| `new_succeeds_and_the_ring_stays_reachable_for_pushes` | 2 | +| `teardown_with_operations_in_flight_neither_hangs_nor_closes_the_ring_early` | 0 | + +The two that trigger it both create an `EventDelivery` over a ring with **nothing outstanding** and +drop it promptly; the one that does not is the one holding operations in flight. That is a +correlation across three tests, not a mechanism, and it is recorded as such. + +**Reproducer**, about half a minute: + +```powershell +$ed = 'target\debug\deps\event_delivery-.exe' # the build with 6 tests; check with --list +$fail = 0 +for ($i=1; $i -le 600; $i++) { + & $ed completions_ dropping_with 2>&1 | Out-Null + if ($LASTEXITCODE -ne 0) { $fail++ } +} +"$fail failures of 600" +``` + +**Where this goes next, and why it left this crate.** Both delivery tests pass `env: None` to +`EventDelivery::new`, so both register their wait on the **default process threadpool** through +[`windows_threadpool_sys::wait::ThreadpoolWait`](../windows-threadpool-sys/src/wait.rs). + +### Narrowed further 2026-09-25: the pool is alive, and the stall is permanent by design + +A configurable trace was added for this (see below) and the flake **still reproduces with it on**, +which is the first thing to check for a timing-dependent fault. + +**The default pool is not wedged.** A probe runs at the moment of failure, before anything else: it +submits a plain work item and separately creates, arms and signals a **brand-new** wait on a +brand-new event. Measured at a captured stall: `work item ran: true; a fresh wait ran: true`, both +within two seconds. So the pool dispatches, and its wait mechanism works. Whatever is broken is +specific to the waits already registered. + +**Those waits were created and armed.** The trace shows, for every ring in a failing run: +`setup-signalled` -> `event-attached` -> `wait created` -> `wait armed`, all within microseconds -- +and then `trampoline-entered` **never appears at all**, for any of them, for the rest of the process. + +**The ring's setup signal is raised on the ring's own handle, before the wait is armed on a +duplicate of it.** The trace records both handle values, and in the captures examined the failing +ring's handles were not recycled values of the dropped ring's. + +**Why the stall is permanent rather than merely late, which the trace explains.** The completion +event is edge triggered ([D-19](DESIGN-NOTES.md#d-19)): it fires when the queue goes from empty to +non-empty. A stalled ring has eight completions sitting in its queue, so the queue never returns to +empty and **no further signal will ever be raised**. The setup signal -- the one deliberate wakeup +that exists precisely to cover a backlog -- is therefore the only signal that ring will ever get. +Lose it once and delivery for that ring is dead for good. That is consistent with every capture: +zero callbacks, nothing after ten more seconds, and both rings affected together. + +**So the open question is narrow: why does an armed wait not observe a signal raised before it was +armed?** An auto-reset event signalled with no waiter stays signalled, so arming afterwards should +consume it and fire. It does, on better than 99% of runs. + +**What is deliberately not concluded.** A mechanism suggests itself -- the pool's internal wait +thread multiplexes handles, and a concurrent close could plausibly disturb the set it is watching, +which would fit a fresh wait working while existing ones do not. That is a hypothesis with no +evidence behind it yet, and it is recorded here as one so the next person does not mistake it for a +finding. The experiment that would settle it is whether re-arming a stalled wait recovers it; +`EventDelivery` does not currently expose its wait, so that needs either a test-only accessor or the +probe moved into `windows-threadpool-sys`. + +**Why it is worth recording despite being rare.** The sabotage harness runs the whole suite once per +case, and the manifest currently holds 41 cases. At the measured rate that is about a **40% chance +that any given sweep contains at least one corrupted result** -- and the corruption is the dangerous +direction: a sabotage the suite did not really catch is reported as `caught`, which reads as a clean +bill of health. Both instances seen so far landed on cases whose patches **provably cannot** affect +event delivery -- a failure-code bitmask in the resolver, and a prose reword inside an example's +contract text -- which is how they were recognised as false rather than believed. + +**How to tell a false `caught` from a real one.** Read the per-case transcript under +`.scratch/sabotage/`; a genuine detection names a test related to the patch, while this one names +one of the two tests above and prints the stall report described next. Do not conclude a sweep is +clean or dirty from the summary table alone while this is open. + +**Explicitly not caused by the 2048 sweep widening**, though that is when it was noticed. The rate +was measured on the `event_delivery` binary, which uses neither the resolver nor any seeded sweep, +so its behaviour is independent of those constants. Three sweeps at the previous sizes had passed +earlier the same day, which is unsurprising at this rate rather than evidence of a change. + +### Turning the trace on, and narrowing it + +The trace is **compiled out** unless the `trace` feature is on, because the instrument for a +timing-dependent fault must not change the schedule it is measuring. When on it is still off at run +time until `WINDOWS_THREADPOOL_TRACE` names the targets wanted, so a session can record one +subsystem rather than everything: + +```powershell +$env:WINDOWS_THREADPOOL_TRACE = 'wait,delivery' # or 'wait', or '*' +cargo test -p windows-ioring-sys --features trace --test event_delivery +``` + +Recording does not format and does not allocate: an entry is a timestamp, a thread id, two +`&'static str` labels and two `u64` slots, formatted only when a dump is asked for. The dump is +included in the stall report automatically, so a captured failure carries its own trace. + +Targets currently emitted: `wait` (create, arm, trampoline entry, the three phases of drop) and +`delivery` (the ring's setup signal with its outstanding count, event attach, arm, callback entry +and exit). + +**The flake still reproduces with the trace on**, which was checked before drawing anything from it. + +### What a stalled run now records + +Added 2026-09-25. The original failure said only `Timeout`, which ruled nothing out -- that is why +the investigation above could only proceed by elimination. Both tests now print a report on the way +out, to stderr and into the panic message, so `cargo test`'s captured output and the sabotage +harness's per-case transcript both carry it. It states: + +- **delivered, of how many expected** -- whether the stall was immediate or partway through. +- **callbacks run** -- how many times the pool actually invoked the callback. Equal to delivered + means everything the callback received reached the test thread; greater means the gap is between + the callback and the channel. This is the first fork in the diagnosis and nothing else supplies + it. +- **outstanding** -- the ring's own count, with the caveat that makes it readable: it decrements on + pop and the pop happens *inside* the callback, so on its own it cannot separate "the kernel has + not finished" from "the callback never ran". +- **arrival times and inter-arrival gaps** -- whether deliveries were steady and then stopped, or + slow throughout. +- **a post-mortem** -- after the bound expires the test waits a further ten seconds and says whether + the delivery arrived late or never came at all. Those have different causes, and no other datum + separates them. + +Two properties of that reporting are deliberate. The immediate facts are printed **before** the +post-mortem wait, so they survive the sabotage harness killing a run that exceeds its hang bound -- +losing the report to the very timeout it exists to explain would be the worst outcome. And the +post-mortem is ten seconds rather than thirty so a failing run stays inside that bound: measured, a +forced stall completes in about seventeen seconds against a bound of roughly thirty. + +**The report states observations and stops.** An earlier draft ended with a verdict, and a +forced-failure run showed the verdict was wrong -- it blamed something upstream of the channel when +the injected fault was in the callback body, which the counters it had just printed already ruled +out. + +**Queued as `M26.9`** in [CHECKLIST.md](CHECKLIST.md). diff --git a/crates/windows-ioring-sys/RESPONSE-SPACE.md b/crates/windows-ioring-sys/RESPONSE-SPACE.md new file mode 100644 index 000000000..13d1cdefb --- /dev/null +++ b/crates/windows-ioring-sys/RESPONSE-SPACE.md @@ -0,0 +1,302 @@ +# The permitted kernel response space + +What `windows-ioring-sys` will tolerate from the platform, stated as a +specification rather than recorded from a run. + +This document is normative. `M26.3`'s resolver generates resolutions **from +this space**, `M26.4`'s properties must hold under every one of them, and the +kernel tests confirm that a real Windows stays **inside** it (`M26.6`). Every +clause carries an ID so those three can cite the clause rather than restate it +-- and [response_space_census.rs](tests/response_space_census.rs) fails when a +clause is claimed by nothing on the side that owes it a check, so the division +of labour is enforced rather than merely described. + +## What this is, and what it deliberately is not + +**It is not a model of what Windows does.** A model of observed behaviour +freezes one run's testimony, which is the trap +[D-52](DESIGN-NOTES.md#d-52) was opened to escape and the objection that +kept a fake out of this crate for two milestones. A resolver built on a model +asserts something about the kernel and can be wrong about it. + +**It is a statement of what this crate will tolerate.** The resolver asserts +nothing about Windows. It picks a point in the space below, and the assertions +are about *us*: does this crate behave correctly under that resolution. There +is no belief here to be wrong about -- only a specification that can be too +narrow, which is a reviewable defect rather than a hidden one. + +**It is wider than anything observed, on purpose.** Deriving the space from +observation would close the trap again. Where a clause goes beyond what any +spike has seen, it says so, because the reader's first question about a +permissive clause is whether anyone has watched it happen. + +## How to read a clause + +Every clause has an ID, a statement, a source, and a provenance tag: + +- **Observed** -- a spike or a measurement in this repository saw it happen. + The citation says which. +- **Documented** -- Microsoft states it on the API's reference page. This is + the strongest tag, and it outranks the others: a measurement describes one + run of one build, while a documented return value is what the platform + commits to. Where a clause carries both, the documentation is the reason and + the measurement is corroboration. +- **Over-provision** -- wider than anything observed here, allowed + deliberately. These are the clauses that make the space a specification + rather than a recording. +- **Decided** -- a constraint this crate chooses to require of the platform. + Not measured, not derived; a call, and reviewable as one. + +A clause tagged **Observed** may still be wider than its observation. Where +that is so it is split, so the measured part and the extrapolated part can be +argued separately. + +## Permitted: what a resolver may do + +### RS-P-1 -- An operation may complete inside `SubmitIoRing`, or pend + +Each operation in a submitted batch resolves independently as either *already +complete when `SubmitIoRing` returns* or *outstanding*. + +- **Observed.** + [write-pending-spike.rs](design-sessions/spikes/write-pending-spike.rs) + measured both outcomes, and measured the mix varying with handle flags and + with whether the extent was written beforehand -- see + [2026-09-24-set-len-vs-zero-fill/](measurements/2026-09-24-set-len-vs-zero-fill/README.md), + where a buffered handle pended in almost none of 16,000 trials and an + unbuffered one pended in most runs. +- **Over-provision:** the resolver chooses **per operation**, independently, + with no rate and no correlation to handle flags. No measurement established + that operations within one batch resolve independently. The space permits it + because a consumer that depends on them resolving together is depending on + something Windows never promised. + +### RS-P-2 -- Completion order is unconstrained + +Completions may be posted in any order, and that order need bear no relation to +submission order. + +- **Observed, in part.** + [D-47](DESIGN-NOTES.md#d-47) measured operations queued *after* a drained + flush completing *before* it, at 0.03%-0.8% depending on conditions, with all + 32 overtaking in the worst observed trial. + [kernel-response-space-probe.rs](design-sessions/kernel-response-space-probe.rs) + then built a seeded resolver that permutes completion order and ran a + FIFO-assuming consumer against it: broken under 189 of 200 seeds, first at + seed 0, where completions arrived as `[3, 1, 4, 2]`. That probe is the + working demonstration this clause and `M26.3` are both built on; the + [session](design-sessions/DESIGN-SESSION-2026-09-22-kernel-response-space.md) + is its write-up. +- **Over-provision:** any permutation, not merely the reorderings observed. + This is the clause the `D-47` defect class lives in, and the reason it is + total rather than bounded is that a bound derived from observed rates is a + recording. + +### RS-P-3 -- An operation may fail individually + +Any single operation may complete with a failure result while others submitted +beside it succeed. + +- **Documented.** `SubmitIoRing`'s Remarks state the mechanism: *"Any errors + processing a single submission queue entry results in a synchronous + completion of that entry posted to the completion queue with an error status + code for that operation."* So a per-entry failure is **not** a submit + failure; it arrives as an ordinary completion carrying an error. That is the + contract, and the two observations below are consistent with it rather than + the basis for it. +- **Observed.** Two independent places in this repository are built around it. + [`checkpoint.rs`](examples/epoch_log/checkpoint.rs) documents the case + directly -- a record write failing with `ERROR_DISK_FULL` followed by a flush + that "completes perfectly happily" -- which is why that control plane checks + its write's result separately from its flush's. And on the *append* path, + `M22.2` found an ordering defect in handling exactly this: a failed write's + token was not claimed, so its arena slot leaked and the failure surfaced + `SLOTS` appends later with no trace of the cause. +- **Over-provision:** any error code, at any position in the batch, including + the case where every operation fails. The failure *codes* the platform + actually returns are not enumerated here and a resolver must not depend on + the set being small. + +### RS-P-4 -- A wait may expire + +A submit-and-wait may return `HRESULT_FROM_WIN32(ERROR_TIMEOUT)` having waited +its full timeout, and this is **not** a failure of the ring. + +- **Observed.** `M21.6` fixed a defect in this crate that treated an expired + wait as an operation failure; `ring.rs`'s `IORING_E_WAIT_TIMEOUT` handling is + the correction, and it maps the code to `Ok(())`. `M26.8` found the same + defect surviving in [`Batch::submit_and_wait`](src/batch.rs), which that + sweep had not reached. +- **Documented, and the documentation says more than "not a failure".** + `SubmitIoRing` gives this code its own return-value row: *"All operations + were submitted without error and the subsequent wait timed out."* So it + carries a positive guarantee about the submission half, not merely the + absence of a failure -- which is why it is classified separately from every + other error rather than folded in with them. +- **Over-provision:** a wait may expire even when completions are available, + and may expire on any call including the first. + +### RS-P-5 -- A wait may return successfully with nothing poppable + +A wait that returns success does not promise that a subsequent pop yields +anything. + +- **Observed.** This crate's own `pop_within` documentation states it -- a + submit-side wait's return "promises nothing about poppability" -- and + [D-19](DESIGN-NOTES.md#d-19) measured the completion event as **edge** + triggered on the queue going empty to non-empty, so a signal is not a count + and the queue can be drained by the time a waiter looks. +- **Over-provision:** this may happen on any wait, any number of times in + succession. A resolver is not required to make progress on any particular + call, only to satisfy RS-C-1 eventually. + +### RS-P-6 -- A signal may not arrive for a completion posted while the queue was already non-empty + +The completion event fires on the empty-to-non-empty edge, so a completion +arriving behind another need not produce its own signal. + +- **Observed.** [D-19](DESIGN-NOTES.md#d-19), measured, and + [D-21](DESIGN-NOTES.md#d-21) is the consequence this crate drew from it -- + auto-reset, exactly one waiter per ring. +- **Over-provision:** none. This clause is the measurement. + +### RS-P-7 -- A submit may fail, leaving already-built operations queued for a later submit + +`SubmitIoRing` may fail, and when it does every entry it was asked to submit +remains in the submission queue. A later, unrelated submit is what runs them. + +- **Documented, which `M26.8` established and `M26.1` had not.** This clause + was originally written as a *consequence* -- "if a submit fails, entries + remain queued" -- citing [D-5](DESIGN-NOTES.md#d-5), which establishes the + no-rewind consequence and nothing about submits failing at all. `M26.3`'s + resolver read it as a permission to fail submits, and the space had no + authority for that. It does now: `SubmitIoRing`'s return-value table lists + *"Any other error value: Failure to process the submission queue in its + entirety"*, and its Remarks state *"If this function returns an error other + than IORING_E_WAIT_TIMEOUT, then all entries remain in the submission + queue."* Both halves of this clause are therefore Microsoft's, not an + inference from a run. +- **The `IORING_E_WAIT_TIMEOUT` carve-out is part of the clause**, because it + is the case where the entries did *not* remain queued: that code means every + operation was submitted and only the wait expired. A consumer that cannot + tell the two apart cannot know whether its buffers are still owed to the + kernel, which is the defect `M26.8` fixed in + [`Batch::submit_and_wait`](src/batch.rs). +- **Over-provision:** the resolver may defer an operation across any number of + submits, not only across a failed one. It currently declines a submit only + when something is staged, which is *narrower* than this clause -- nothing + says a submit carrying no new work cannot fail. + +## Constrained: what a resolver may not do + +These exist because a resolver free to violate everything makes this crate +defend against a platform that does not exist, and code written against an +impossible kernel is untestable and unreviewable. Each is a call. + +### RS-P-8 -- A successful transfer may report fewer bytes than requested + +A read or write that completes successfully may report an `Information` below +the length it was given. The remainder is not transferred, and nothing in this +crate reissues it. + +- **Documented, for the handle types that do it.** `WriteFile`'s Remarks state + that *"when writing to a non-blocking, byte-mode pipe handle with + insufficient buffer space, WriteFile returns TRUE with + \*lpNumberOfBytesWritten < nNumberOfBytesToWrite"*. Sockets report a short + send when the transmit buffer cannot take the whole buffer, and a + communications handle with a write timeout set by `SetCommTimeouts` can + report a partial count when the timeout fires after some data has gone out. + Reads are shorter still by nature: end of file, a pipe with less buffered + than asked for. +- **Why it is a permission of this space and not a property of a file.** For an + ordinary file on a local volume a successful completion is expected to carry + the full requested length, and a full volume is an error + (`ERROR_DISK_FULL`) rather than a short success. But **this crate does not + constrain what a caller registers** -- `IoRing` takes a handle, and a pipe, a + socket and a serial port are all handles. A consumer of this crate therefore + has to read the count. A consumer that has *also* narrowed its handle type + can rely on more, and the place to say so is that consumer's own contract, + not this space. +- **The continuation is the caller's.** This crate reports the count and stops + there. Whether to reissue the remainder, how many times, and when to give up + are policy, and [D-67](DESIGN-NOTES.md#d-67) keeps policy with the caller. +- **A short count may be zero.** A consumer looping on the remainder must + tolerate a completion that makes no progress rather than assuming each one + advances it. +- Flush and cancel carry no byte count, so this clause does not reach them. +### RS-C-1 -- Every submitted operation eventually completes exactly once + +No completion is lost, none is duplicated, and every successfully submitted +operation eventually produces exactly one completion. + +- **Decided**, not observed. Nothing here has measured the negative, and it + could not be measured in bounded time. +- **Why:** this crate's accounting is driven by observing a real `IORING_CQE` + ([D-4](DESIGN-NOTES.md#d-4)) and its [`RingContract`](src/contract.rs) oracle + states conservation directly. A platform that lost completions would make + every consumer's outstanding count unbounded and every wait a guess. If this + is ever observed to fail, the finding is a contract defect in Windows and not + a gap in this space. + +### RS-C-2 -- A completion identifies the operation that produced it + +A completion's `UserData` is the value supplied when the operation was built. + +- **Decided.** The alternative is that operation identity is unusable, which + would invalidate [D-4](DESIGN-NOTES.md#d-4)'s whole accounting model and + `M28`'s pending inventory with it. + +### RS-C-3 -- An operation does not complete before it is submitted + +- **Decided**, and stated because a resolver that may post a completion for an + operation still being built would make the `Build*`/`Submit` boundary + meaningless. + +### RS-C-4 -- The drain half of `DRAIN_PRECEDING_OPS` holds + +No operation queued **before** a flush carrying +`IOSQE_FLAGS_DRAIN_PRECEDING_OPS` completes after that flush. + +- **This is the explicit call `M26.1` demanded, and it is the one place this + space is narrower than "anything may happen".** +- **Observed, with the strongest evidence in this repository.** + [D-47](DESIGN-NOTES.md#d-47) measured roughly 4,500 trials in which **not + once** did an operation queued before a drained flush complete after it -- + the same campaign that falsified the *other* half of `D-24`. +- **Why constrained:** the drain is the documented guarantee this crate's + durability story rests on ([D-23](DESIGN-NOTES.md#d-23)). A resolver + permitted to break it would require every consumer to re-verify durability by + some other means, which is to say it would make the primitive useless. The + cost of this call is that a Windows which broke the drain would not be caught + by the resolver at all -- it is caught by the kernel tests instead, which is + the division of labour they were repointed to in `M26.6`. That is now a + mechanical arrangement rather than an intention: + [flush_barrier.rs](tests/flush_barrier.rs) carries a `CONFIRMS: RS-C-4` + marker and asserts the clause against a real ring on every machine, and + [response_space_census.rs](tests/response_space_census.rs) fails if that + marker ever disappears. +- **Note what is *not* constrained:** the hold-back half. + [D-24](DESIGN-NOTES.md#d-24) claimed the flag holds back what follows and + [D-47](DESIGN-NOTES.md#d-47) withdrew that claim, so RS-P-2 applies in full + to operations queued *after* a drained flush. The one-sidedness is the whole + point of the pair. + +## What this space deliberately leaves undecided + +Stated so that a later reader can tell an omission from a choice: + +- **Rates.** No clause carries a probability. A resolver weights its choices by + seed, and any weighting is a property of the resolver rather than of this + space. Recording observed rates here would make the space a recording. +- **Failure code sets.** RS-P-3 permits any code and enumerates none. +- **Timing.** Nothing here constrains how long anything takes. `M25`'s standing + constraint already forbids this crate from depending on an operation pending, + and a space that specified durations would invite exactly that. + +## Changing this document + +A clause moving from **Over-provision** to **Observed** is an improvement and +needs only its citation updated. A clause moving in the other direction, or a +constraint being relaxed, changes what this crate promises to tolerate and +must be recorded as a decision in [DESIGN-NOTES.md](DESIGN-NOTES.md) with the +finding that forced it. diff --git a/crates/windows-ioring-sys/RING-OPENING-LIB-TESTS.txt b/crates/windows-ioring-sys/RING-OPENING-LIB-TESTS.txt new file mode 100644 index 000000000..8286e346d --- /dev/null +++ b/crates/windows-ioring-sys/RING-OPENING-LIB-TESTS.txt @@ -0,0 +1,49 @@ +# windows-ioring-sys: lib tests that open a real kernel ring. +# +# GENERATED by tools/check-ring-tests.ps1 -Update. Do not hand-edit. +# +# These are integration tests living in the unit-test location (D-49). +# The list exists so the population cannot grow unnoticed, which is how +# it reached 63 before anyone counted. Adding an entry obliges an +# answer: does this test need the kernel, or only a ring-shaped thing? +batch::a_pending_buffer_registration_claims_only_its_own_completion +batch::a_pending_file_registration_claims_only_its_own_completion +batch::a_pending_file_registration_reports_the_user_data_it_will_claim +batch::dropping_a_batch_that_queued_nothing_submits_harmlessly +batch::dropping_a_registration_with_work_outstanding_is_refused +batch::registered_files_index_from_the_base_and_stop_at_the_end +batch::registered_files_report_their_extent +batch::require_refuses_an_op_the_ring_does_not_support +batch::submit_reports_how_many_operations_it_queued +batch::windows_refuses_an_empty_buffer_registration +event_delivery::a_scope_reflects_a_ring_that_genuinely_lacks_support +event_delivery::a_scope_reports_outstanding_work +event_delivery::a_scope_reports_registration_counts_that_change_with_registrations +event_delivery::a_scope_reports_the_rings_static_properties +ring::a_bound_the_clock_cannot_represent_reaches_the_wait_rather_than_panicking +ring::a_reserved_opcode_is_not_supported +ring::a_supplied_wait_is_consulted_when_the_queue_is_not_ready +ring::a_supplied_wait_is_not_consulted_when_nothing_can_arrive +ring::a_wait_that_fails_ends_the_pop_with_its_error +ring::a_wait_that_never_blocks_is_permitted_and_still_terminates +ring::a_zero_bound_does_not_block +ring::an_injected_failure_carries_the_condition_it_names +ring::an_injected_failure_preserves_the_identity_a_token_claims_against +ring::an_injected_failure_replaces_a_real_success +ring::an_injected_failure_zeroes_the_transferred_byte_count +ring::capability_reporting_never_claims_more_than_is_io_ring_op_supported_reports +ring::dropping_a_ring_actually_runs_its_drop_body +ring::each_spelling_of_a_failure_produces_the_condition_it_names +ring::every_named_condition_injects_a_genuine_failure +ring::injecting_a_success_code_is_refused +ring::nop_read_and_write_are_supported_on_any_real_ring +ring::pop_within_returns_successive_completions_one_at_a_time +ring::pop_within_returns_the_completion_of_a_real_operation +ring::ring_wait_reports_the_rings_outstanding_count +ring::run_down_returns_once_a_recorded_completion_zeroes_the_count +ring::submit_wait_is_what_the_convenience_uses +ring::supports_reports_exactly_the_capability_set_it_was_given +ring::the_deadline_is_honoured_when_an_operation_never_completes +ring::the_debug_rendering_names_the_ring_and_its_key_fields +ring::the_wait_can_be_supplied_as_a_trait_object +ring::the_wait_is_never_handed_a_zero_timeout diff --git a/crates/windows-ioring-sys/UNRESOLVED-TEST-FAILURES.md b/crates/windows-ioring-sys/UNRESOLVED-TEST-FAILURES.md index abf45c5f5..c6ccebdfa 100644 --- a/crates/windows-ioring-sys/UNRESOLVED-TEST-FAILURES.md +++ b/crates/windows-ioring-sys/UNRESOLVED-TEST-FAILURES.md @@ -4,4 +4,32 @@ Pre-existing failures that do not block an unrelated commit, recorded per the re checklist-execution rules. When one is resolved, move its entry into a sibling [RESOLVED-TEST-FAILURES.md](RESOLVED-TEST-FAILURES.md) (append-only) rather than deleting it. -None currently. +## a backlog is not delivered when the caller attached the event and consumed its signal + +**Found 2026-09-25**, by Copilot review on PR #108, which reported the narrower form: that +`EventDelivery::new` signals only when it attached the event itself, so a caller who attached it +earlier arms a wait on an already-non-empty, edge-triggered queue with no wakeup owing. + +The reproducer is +`a_backlog_is_delivered_even_when_the_caller_attached_the_event_first` in +[event_delivery.rs](tests/event_delivery.rs), `#[ignore]`d because it fails. It attaches the event, +**consumes** the signal attaching raised, submits work, consumes the signal the completions raise, +and only then hands the ring over -- leaving a non-empty queue with nothing pending, which is the +state the guarantee is about. An earlier version of the test omitted the two consuming waits and +passed against the defect, because the leftover signal fired the wait; that version proved nothing. + +**The reported repair does not work, which is why nothing is applied.** Signalling unconditionally +after arming leaves the test failing 6 of 6. What does make it pass is a 50 ms sleep between +`wait.arm` and the signal: 3 of 3. Building with `--features trace` also makes it pass, which is the +same schedule perturbation by another route. + +So the wakeup is lost in a window *after* arming, rather than never being raised. That is wider than +the review finding, and it bears on [D-68](DESIGN-NOTES.md#d-68): `M26.9` fixed the delivery stall +by ordering the arm before the signal, measured at 0 failures in 3600 runs, and this says that +ordering alone is not sufficient to close the window -- only to narrow it. + +**Not established:** why the window exists. `SetThreadpoolWait` is documented as registering the +wait, so a signal after a completed `arm` should be observed. Whether the pool's wait thread +re-issues its `WaitForMultipleObjects` asynchronously, and whether an auto-reset signal can be +consumed and discarded during that re-issue, is a guess and is recorded here as one. Queued as +`M26.12`. diff --git a/crates/windows-ioring-sys/design-sessions/DESIGN-SESSION-2026-08-28-external-consumer-correspondence.md b/crates/windows-ioring-sys/design-sessions/DESIGN-SESSION-2026-08-28-external-consumer-correspondence.md index cdd385015..dbd76d4e0 100644 --- a/crates/windows-ioring-sys/design-sessions/DESIGN-SESSION-2026-08-28-external-consumer-correspondence.md +++ b/crates/windows-ioring-sys/design-sessions/DESIGN-SESSION-2026-08-28-external-consumer-correspondence.md @@ -159,6 +159,11 @@ writes -- observed at 17 and 23 of 32 writes completing after it) and [D-24](../DESIGN-NOTES.md#d-24) (the barrier is a full, ring-wide stall that spans submissions and holds operations against unrelated files). +*Recorded as it was concluded on this date. D-24's "holds operations" half was withdrawn on +2026-09-06 by [D-47](../DESIGN-NOTES.md#d-47), which measured operations queued behind a drained +flush completing ahead of it; the drain half -- that nothing queued before it completes after it -- +stands.* + ## Round 4 -- what belongs where The closing question was whether this crate should provide the emulation a consumer needs on top of diff --git a/crates/windows-ioring-sys/design-sessions/DESIGN-SESSION-2026-09-19-epoch-log-review.md b/crates/windows-ioring-sys/design-sessions/DESIGN-SESSION-2026-09-19-epoch-log-review.md new file mode 100644 index 000000000..724df4e0d --- /dev/null +++ b/crates/windows-ioring-sys/design-sessions/DESIGN-SESSION-2026-09-19-epoch-log-review.md @@ -0,0 +1,232 @@ +# Design session 2026-09-19: review of the epoch-log sample and its durability surface + +A read-only review of [examples/epoch_log](../examples/epoch_log) and the parts of the crate it +composes, prompted by two questions from the engineer: whether the sample is correct and efficient, +and whether the repository's accumulated learnings about ring structuring and storage affinity +suggest reworking it. + +**Nothing was built, run, or measured during this session.** Every finding below is from reading the +source and the recorded decisions. Where a finding rests on reasoning rather than on a measurement, +it says so. No claim here is a compile claim. + +## What resulted + +New work items [M21](../CHECKLIST.md), [M22](../CHECKLIST.md) and [M23](../CHECKLIST.md) in +[CHECKLIST.md](../CHECKLIST.md), plus an addendum to the already-queued +[M20.6](../CHECKLIST.md). No decision in [DESIGN-NOTES.md](../DESIGN-NOTES.md) was changed by this +session; two of the findings are about decisions whose corrections are queued and not yet landed +(see "Already queued, still undone" below). + +## Scope and method + +Read in full or in relevant part: + +- every module of [examples/epoch_log](../examples/epoch_log); +- [src/batch.rs](../src/batch.rs)'s submission and flush surface, [src/ring.rs](../src/ring.rs)'s + pop and test helpers, [src/lib.rs](../src/lib.rs)'s durability and topology guidance; +- [DESIGN-NOTES.md](../DESIGN-NOTES.md) decisions D-3, D-5, D-8, D-19, D-21, D-23, D-24, D-47, the + "Why the NUMA node is the wrong key" and "What is not reachable" sections; +- [CHECKLIST.md](../CHECKLIST.md) M20, and the session it was queued from, + [DESIGN-SESSION-2026-08-30-numa-sharded-io-execution-domains.md](../../../design-sessions/DESIGN-SESSION-2026-08-30-numa-sharded-io-execution-domains.md); +- [design-sessions/spikes](spikes) -- the two unrun instruments and their README; +- [examples/ring_copy](../examples/ring_copy) for comparison, since it is the crate's other sample + and the one that does make a locality decision. + +## Findings: correctness + +### C-1. A withdrawn D-24 claim survives at one site + +[examples/epoch_log/commit.rs](../examples/epoch_log/commit.rs) line 156 justifies its epoch-order +`debug_assert` with "D-24 holds an operation pushed after a drained one until it completes". That is +the half of [D-24](../DESIGN-NOTES.md#d-24) that [D-47](../DESIGN-NOTES.md#d-47) withdrew, and the +same file's own module header (line 24) already carries the correction. + +A blast-radius sweep of the hold-back phrasing across the crate found 17 matches in 10 files; every +other site is corrected. This is the last one. + +The assertion it guards is still sound, by a different route: commit *N+1* carries the drain flag +itself, and D-47's *surviving* half ("not once did an operation queued before a drained flush +complete after it") is what orders it behind commit *N*. So the conclusion holds and the cited +reason does not -- a correction that did not propagate, rather than a wrong conclusion. + +### C-2. A bare `loop { try_pop }` in the flagship example + +[examples/epoch_log/append.rs](../examples/epoch_log/append.rs) line 89 spins unbounded after +`submit_and_wait(1, 30_000)`. [`Batch::submit_and_wait`](../src/batch.rs) documents that returning +does not mean a completion is poppable, because the timeout can expire first, and +[`pop_within`](../src/ring.rs) states the consequence outright: "A bare `loop` around `try_pop` is +worse, because it converts that flake into a hang." + +The same step is written three ways in this crate: + +| Site | Shape | +|---|---| +| [examples/ring_copy/engine.rs](../examples/ring_copy/engine.rs) line 147 | absence is an error (`TimedOut`) | +| [examples/epoch_log/strategy.rs](../examples/epoch_log/strategy.rs) `Lane::new` | absence is an error | +| [src/ring.rs](../src/ring.rs) `pop_within` | bounded wait, panics on deadline | +| [examples/epoch_log/append.rs](../examples/epoch_log/append.rs) line 89 | unbounded hot spin | +| [tests/fault_injection.rs](../tests/fault_injection.rs) line 50 | unbounded hot spin | + +`pop_within` is `#[cfg(test)] pub(crate)`, so neither an example nor an integration test can reach +it -- examples and `tests/` are separate crates. That is why the duplication exists, and it means +the fix is an API question rather than a copy-paste: publish a bounded pop, or keep re-deriving it. + +The rule is written down in three places and enforced nowhere, which is the detection-ladder point: +prose is not a rung. + +### C-3. The commit trigger keys off the counter, not off the append + +[examples/epoch_log/main.rs](../examples/epoch_log/main.rs) line 279 tests +`appended % EPOCH_SIZE == 0` on every pass of the append loop, including a pass where `append` +returned `WouldBlock` and `appended` did not move. On such a pass it commits again: a second +covering flush closing an epoch with nothing in it, and an epoch number consumed for no records. + +Not reachable at the sample's current constants -- `SLOTS` is 8, `EPOCH_SIZE` is 6, and the commit +wait drains the arena, so the arena cannot be full at a boundary. It is armed by anyone who copies +the sample and raises `EPOCH_SIZE`, which is what the sample exists to be. + +> **Corrected 2026-09-21 while implementing `M21.3`: the second paragraph is wrong.** The retry is +> not reachable at *any* constants, because the predicate is true at exactly two moments -- before +> the first append, and immediately after a commit -- and the arena is empty at both, the commit +> having waited for a covering flush that retires every outstanding write. Measured rather than +> re-reasoned: the retry path was instrumented to report when the old shape would have committed, +> and it fired **zero** times at `EPOCH_SIZE` of 6, 8, 12, 16 and 24, including the values past +> `SLOTS` this finding predicted would arm it. +> +> What survives is the coupling complaint in the heading, and it is worth the change on its own: the +> trigger was safe because of an invariant three blocks away that nothing stated, rather than +> because of where it was written. The lesson for this review is narrower and sharper -- "unreachable +> today, armed tomorrow" is a claim about a program's reachable states, and reading the code is not +> how to settle one. + +### C-4. `durable_through` across a failed commit is under-specified + +[examples/epoch_log/commit.rs](../examples/epoch_log/commit.rs) says "A failed commit advances +nothing", which reads as though a failed commit of epoch *N* leaves *N* non-durable permanently. + +It does not. Epoch *N*'s writes precede commit *N+1*'s covering flush, so a later successful commit +makes *N* genuinely durable, and the monotonic reading of `durable_through` stays true. That is the +correct behaviour; the reasoning appears nowhere, so a reader auditing monotonicity after a failure +has to re-derive it. This is a specification gap, not a defect. + +### C-5. Two wait loops hang where their sibling fails + +[examples/epoch_log/strategy.rs](../examples/epoch_log/strategy.rs) lines 372 and 384 discard the +`submit_and_wait` timeout and loop forever. +[`EventLoop::pump`](../examples/epoch_log/event_loop.rs) raises `TimedOut` on the same condition and +documents why: "so a stuck loop fails instead of spinning". One program, opposite policies. + +## Findings: efficiency + +### E-1. One `SubmitIoRing` per record + +Both [`Appender::append`](../examples/epoch_log/append.rs) and +[`Lane::append`](../examples/epoch_log/strategy.rs) construct a `Batch`, push one write, and submit +it. `Batch` exists to amortise submission across many SQEs; the sample that teaches `Batch` submits +one entry at a time. + +This is not only a throughput observation. It puts a fixed per-record submission cost into all three +strategies in [strategy.rs](../examples/epoch_log/strategy.rs), which is a shared term in the +comparison whose headline result is that the three are indistinguishable. Whether batching moves +that spread is unmeasured; it is the cheapest experiment available, and it bears on +[M20.6](../CHECKLIST.md). + +### E-2. Two implementations of the free-slot pool, in one program + +[`Appender::free_slot`](../examples/epoch_log/append.rs) line 132 scans the arena calling +`outstanding()` per slot; `Lane` keeps a `Vec` free list. Both are correct and the difference +does not matter at eight slots. The duplication is what matters, because the two can drift. + +### E-3. The arena has no placement story + +[examples/epoch_log/append.rs](../examples/epoch_log/append.rs) line 84 allocates the registered +arena as `vec![0_u8; SLOT_LEN]` -- heap, no alignment, no node. The crate's own front page +([src/lib.rs](../src/lib.rs)) says buffer placement "is very likely the highest-leverage locality +decision available" and names `VirtualAllocExNuma`, and +[examples/ring_copy/buffer.rs](../examples/ring_copy/buffer.rs) already implements exactly that. + +The durability sample has no locality story at all: no pinning, no node-local arena, one ring. That +may be the right call for a sample about durability -- but it is currently a silence rather than a +stated choice, while the crate's headline guidance says the opposite. + +## Findings: ring structuring and storage affinity + +### S-1. D-47 left one cost standing that the sample never names + +[D-47](../DESIGN-NOTES.md#d-47) withdrew the hold-back claim and explicitly kept the other half: the +barrier "does still reach every outstanding operation on the ring rather than only the current +submission batch". + +So commit latency is a function of whatever else shares the ring. The ring is therefore part of the +log's durability unit, and "one ring per log" is a precondition rather than a sample convenience. + +**Refined when `M23.1` implemented this (2026-09-23); the finding is left as recorded, per Tier 3.** +The barrier is a *ring* flag and the flush names a *file*, so the two bound different things: the +barrier bounds what a commit waits for, the flush bounds what it makes durable. Completion is not +durability, so a shared ring threatens the **cost model** rather than the guarantee. See +[contract.rs](../examples/epoch_log/contract.rs) -> "One ring per log, because the barrier is +ring-wide", which is authoritative over this paragraph. +[contract.rs](../examples/epoch_log/contract.rs) -- which is where this sample puts its +preconditions, and which was deliberately written before the code -- does not say so. + +### S-2. The re-founding of `AlternatingRings` that M20.6 is asking for + +[M20.6](../CHECKLIST.md) asks whether alternating rings still earns its cost now that its stated +benefit (keeping appends off a stalled ring) is withdrawn, and offers "epoch *N+1*'s appends are +provably outside epoch *N*" as the remaining benefit, characterising that as a correctness property +rather than a throughput one. + +Read against S-1 it is also a throughput property, sited differently. What alternating rings buys is +a **bound on what a commit's barrier can be dragged into**: under a shared ring, commit latency is +unbounded in unrelated traffic on that ring; under alternating rings it is bounded by the epoch. +That is a stronger answer than the one M20.6 currently records, and it is measurable with the +harness that already exists. + +Unmeasured. It follows from the flush's recorded scope plus D-47's surviving half. + +### S-3. The storage-affinity question, and the adjacent one that is answerable + +The engineer's framing -- Windows does not really present storage NUMA affinity for NVMe-attached +storage, but that is no reason not to anticipate it -- matches what the repository has already +established, and [M20.4](../CHECKLIST.md) holds the mechanism research: `FSCTL_QUERY_VOLUME_NUMA_INFO` +takes a file or directory handle directly with no device-instance walk; `GetNumaNodeNumberFromHandle` +bottoms out in `NtQueryInformationFile` with `FileNumaNodeInformation` (class 53), which PHNT and the +WDK mark reserved for system use, so this crate must not build on it; and no published measurement of +either succeeding on an ordinary NTFS data file could be found. The conclusion survives for a better +reason than the one originally recorded: the documented meaning is the node the *volume* resides on, +not where the file's extents live, so it cannot answer "which ring should this file's I/O go to" even +when it succeeds. + +Two ways to anticipate it without claiming it, both of which keep [D-8](../DESIGN-NOTES.md#d-8) +intact by leaving the policy with the consumer: + +1. **Declare rather than discover.** Let a consumer *state* the storage node for a domain and have + the arena allocate there via `VirtualAllocExNuma`. The unanswerable question becomes a declared + input; [file-handle-numa-spike.rs](spikes/file-handle-numa-spike.rs) fills it in automatically if + hardware ever answers. +2. **Shard by backing device rather than by node.** `IOCTL_STORAGE_GET_DEVICE_NUMBER` and + `IOCTL_VOLUME_GET_VOLUME_DISK_EXTENTS` -- both already named in the spike -- answer *which + physical device backs this handle*, and that question is reachable today on ordinary hardware. It + is also the question that governs the cost this sample is built around: a device cache flush is + per-device, so two logs on one device contend at every commit, and a ring spanning two devices + takes the slower device's flush on every covering flush. + +The second is the substantive suggestion of this session: the placement question people reach for +(which node) is unanswerable, while the adjacent question that actually sets commit cost (which +device) is not. Unmeasured; it follows from the flush's scope, and the instruments to settle it +exist. + +## Already queued, still undone + +Two corrections were queued before this session and have not landed. Recorded here so this session +is not read as discovering them: + +- [M20.4](../CHECKLIST.md) -- "What is not reachable" in [DESIGN-NOTES.md](../DESIGN-NOTES.md) still + carries the superseded mechanism (walking volume to disk to device instance and reading + `DEVPKEY_Device_Numa_Node`). The corrected mechanism is written in the checklist item and not in + the decision. + **Landed the same day**, in this session's follow-up work -- see + [COMPLETED-CHECKLIST.md](../COMPLETED-CHECKLIST.md#m204), which also records two things the item's + own text had stale. +- [M20.6](../CHECKLIST.md) -- the `AlternatingRings` re-evaluation. S-2 above is an addendum to it, + not a replacement. diff --git a/crates/windows-ioring-sys/design-sessions/DESIGN-SESSION-2026-09-21-hermetic-unit-tests.md b/crates/windows-ioring-sys/design-sessions/DESIGN-SESSION-2026-09-21-hermetic-unit-tests.md new file mode 100644 index 000000000..7ce51a0f2 --- /dev/null +++ b/crates/windows-ioring-sys/design-sessions/DESIGN-SESSION-2026-09-21-hermetic-unit-tests.md @@ -0,0 +1,228 @@ +# Design session 2026-09-21: hermetic unit tests without losing unit-level coverage + +**Status: concluded 2026-09-21.** The engineer's verdict was that the current design is incorrect +and the question is only when to fix it, so the defect and its classification are recorded as +[D-49](../DESIGN-NOTES.md#d-49) and the work is scheduled as `M24` in [CHECKLIST.md](../CHECKLIST.md). +The *remedy* is still open, gated on `M24.1`. What follows is the record as written before that +verdict; the proposal below is the input to `M24.1`, not its answer. This records a question, the +measurements taken to answer it, and a proposal, so that a decision can be made against evidence +rather than against recollection. If the proposal is adopted it becomes a decision in +[DESIGN-NOTES.md](../DESIGN-NOTES.md) and checklist items in [CHECKLIST.md](../CHECKLIST.md), in the +same change. + +## The question + +Raised by the engineer after the M21 work, in two parts: + +> A unit test that is affected by system load is not a unit test. [...] Unit tests should be +> hermetic. The fact that this failure was discovered due to system load is prima facie evidence +> that the tests are not hermetic. + +and then, on being told the structural fix would move those tests to `tests/`: + +> While we structurally *can* [move] those tests to be integration tests, that means we lose their +> coverage at the unit level which is not something that I want to give up lightly. I would prefer to +> be able to either have the same test code duplicated or be effectively a library and be able to +> apply to the hermetic IoRing as well as the actual one. + +## What is already settled, and what was got wrong + +**The classification is not in doubt.** The repository's own Quality rule reserves integration tests +for cases that "must cross a real process, filesystem, network, device, **operating-system API**, or +other external boundary". `CreateIoRing` is an operating-system API. Tests that open a real ring are +integration tests, wherever they currently live. + +**What M21.6 fixed was not hermeticity.** Removing five wall-clock assertions from four unit tests +made their *outcome* independent of load -- measured, under 2x CPU saturation, at 30x duration +variation with zero outcome variation. It did not make them hermetic: they still open a real kernel +ring. Outcome-stable and hermetic are different properties, and conflating them is what let the +original answer sound complete when it was not. + +One correction to the evidence, which cuts in the engineer's favour rather than against: the +load-affected failure actually observed was in `tests/bounded_pop.rs`, which *is* an integration +test and is where it belongs. The unit suite's hermeticity problem was real but separate -- it was +the five clock assertions, which failed nothing and were removed on principle. + +## Measurements + +Counted by command, not by recollection. + +**How much of the unit suite crosses the boundary:** + +| Module | Tests | Open a real ring | +|---|---|---| +| `buf`, `capability`, `contract`, `error` | 60 | 0 | +| `batch` | 18 | 13 | +| `ring` | 40 | 37 | +| `event_delivery` | 6 | 6 | +| `token` | 7 | 7 | +| **total** | **131** | **63** | + +Four modules are already perfectly hermetic. The non-hermetic mass is concentrated in four others. + +**How movable those 63 are, as they stand:** + +- **25** use only public API. A pure relocation to `tests/`. +- **38** also reach crate-private items -- `reserve_user_data`, `record_completion`, + `cancel_reservation`, `Completion::synthetic`, `set_supported_ops_for_test`, `raw_handle`, + `ring_id`. An integration test cannot see any of those, and widening them would be the wrong + trade: several exist specifically in order *not* to be public. + +**What those 38 are actually testing:** identity minting, outstanding accounting, completion +matching, capability gating. That is bookkeeping, not kernel behaviour. They open a ring only +because the bookkeeping lives as fields on a struct that also owns a handle. + +**Whether that bookkeeping separates.** Both questions flagged as unknowns were investigated and +both came back clean: + +- `RingId::next()` is a process-global `AtomicU64` and never touches a handle. Identity minting is + trivially handle-independent. +- `IoRing`'s ten fields split 5/5. Kernel-coupled: `handle`, `completion_event`, + `registered_buffer_infos`, plus `version` and `supported_ops` which are *negotiated* from the + kernel and then are pure data (and already have a test seam, `set_supported_ops_for_test`). Pure + bookkeeping, no handle: `ring_id`, `next_user_data`, `outstanding`, `registered_files`, + `registered_buffers`. + +## The constraint this must not break + +[DESIGN-NOTES.md](../DESIGN-NOTES.md) already rejects a mock `IoRing`, in +"Two techniques deliberately rejected", and the argument is a good one: + +> Both shipped defects were the kernel behaving differently from this crate's assumptions. A mock +> *encodes* the assumption, so one written before those discoveries would have passed both bugs +> green -- it would not merely have failed to find them, it would have manufactured evidence they +> were absent. + +**This pass supplied three more confirmations of exactly that**, all within a few hours: + +| What the kernel actually does | What a mock would have been written to do | +|---|---| +| `SubmitIoRing` reports an expired wait as `ERROR_TIMEOUT`, a *failure* HRESULT | return success, because a timeout is not an error | +| `SubmitIoRing` answers `E_INVALIDARG` when asked to wait with nothing pending | time out, or succeed | +| A handle without `FILE_FLAG_OVERLAPPED` completes **inline during submit** | leave the operation pending | + +The first is the High-severity defect an independent review found in `M21.2`. A mock would have kept +the suite green while every real timeout returned `Err`. + +So any proposal here has to survive that objection rather than ignore it. + +## The proposal: one suite, two backends, shared assertions + +Write the bookkeeping tests **once**, as generic functions over a small trait, and run them twice: +against a hermetic in-memory implementation and against the real kernel ring. + +```text + +-- run by src/ unit tests, with FakeRing ..... hermetic +shared suite -------+ + (generic) +-- run by tests/ integration, with IoRing .... crosses the boundary +``` + +**Why this is not the rejected mock.** The rejected thing is a mock used *instead of* the kernel, +where nothing ever checks the model against reality. Here the same assertions run against both, so +the fake is continuously differentially tested against the kernel: the moment the model diverges, +one side goes red and names the divergence. That is the same discipline this crate already demands +of its spikes -- "a spike must carry a **control case**, because the first two drain-ordering spikes +could not discriminate and would have returned confidently wrong answers". The kernel run *is* the +control. + +The existing decision would therefore need **amending, not overriding**: it rejects a mock as a +substitute, and this is a mock as a co-tested peer. That distinction is the whole proposal, and if it +does not hold up the proposal fails with it. + +### The bright line: what the fake is never allowed to answer + +The fake models *this crate's bookkeeping*. It never gets a vote on Windows. Concretely, none of +these may be asserted against the fake, because the fake would only be agreeing with whoever wrote +it: + +- the completion event being edge-triggered ([D-19](../DESIGN-NOTES.md#d-19)); +- one waiter per ring ([D-21](../DESIGN-NOTES.md#d-21)); +- flush coverage and drain ordering ([D-23](../DESIGN-NOTES.md#d-23), + [D-24](../DESIGN-NOTES.md#d-24), [D-47](../DESIGN-NOTES.md#d-47)); +- the registration array being read when the operation *runs* ([D-32](../DESIGN-NOTES.md#d-32)); +- `ERROR_TIMEOUT` and `E_INVALIDARG` from `SubmitIoRing`; +- inline completion on a synchronous handle. + +Those stay kernel-only, forever. They are also precisely the findings this pass produced, which is +the argument for the line being drawn exactly here. + +### What the shared suite would cover + +Everything whose truth is decided by this crate rather than by Windows: + +- identity minting: monotonic, never repeated, exhaustion refused rather than wrapped; +- outstanding accounting: minted, completed, cancelled, saturating rather than underflowing; +- token claim matching on `user_data` *and* `ring_id`, including cross-ring rejection; +- registration index arithmetic (the base index of a second registration); +- capability gating -- a push refused because the op is unsupported (`set_supported_ops_for_test` + already exists for this and needs no kernel); +- `pop_within`'s deadline arithmetic and its nothing-can-arrive early return, which with a fake ring + *and* a fake `CompletionWait` becomes fully hermetic; +- `RingContract`'s oracle over observed sequences, which is already hermetic and would simply join + the suite. + +### Mechanics, with costs + +Three shapes, cheapest first. + +**1. Generic test functions over a narrow trait.** A trait describing only what the bookkeeping +tests need -- mint an identity, observe a completion, read `outstanding`, claim a token. The real +implementation wraps `IoRing`; the fake is a few hundred lines of plain Rust. Test bodies are +`fn identity_is_never_reused(ring: &mut R)`. + +*Cost:* the suite must be reachable from both `src/` and `tests/`, and `tests/` cannot see +crate-private items. The established answer in this repository is a feature-gated `pub` module -- +`windows-file-watcher` already ships `test-util` and `scenario-tool` features for exactly this. So: +a `pub mod conformance` behind a non-default `test-util` feature. + +*Consequence to accept:* feature-gated code is invisible to a default `cargo test` and to +`cargo mutants` without `--all-features`, which this repository has already been bitten by (a +`windows-file-watcher` sweep reported 247 survivors of which 147 were in gated modules). CI would +need the suite in its `--all-features` job, and mutation runs would need the flag. + +**2. Parameterise `IoRing` over a backend.** `IoRing` with a private `RingOps` +trait over the ~10 Win32 entry points. The public spelling `IoRing` survives via the default type +parameter. + +*Cost:* `Batch<'_>` becomes `Batch<'_, B>`, and every `impl` block gains a parameter. Mechanical but +it touches the whole crate, and it puts a type parameter into a published API for a testing reason. +Higher risk, more coverage: it would let the *submission* paths be exercised hermetically too, not +just the bookkeeping. + +**3. Extract the bookkeeping into a handle-free type.** `RingAccounting` holding the five pure +fields, composed by `IoRing`. Most of the 38 internals-touching tests become hermetic *in place*, +with no fake, no feature gate, and no widened visibility. + +*Cost:* a real refactor of a shipped crate's internals, though not of its public surface. It is the +smallest conceptual change and the one that most directly matches the observation that those 38 +tests were never about the kernel. + +These are not exclusive. **3 then 1** is the combination worth considering: extract the accounting so +much of the coverage becomes hermetic without any fake at all, then add the shared suite for what +remains and genuinely benefits from running against both backends. + +## Open questions for the engineer + +1. **Does the co-tested-peer argument actually survive?** It is the load-bearing claim. If a fake can + drift in a way the shared assertions do not catch -- because the assertion is about our + bookkeeping and the drift is in our model of Windows -- then the rejection stands and only option + 3 is safe. +2. **Is a type parameter in the published API acceptable for a testing reason?** That is option 2's + real cost, and it is a question about the crate's public shape rather than about tests. +3. **Is the feature-gate consequence acceptable?** Gated tests do not run by default, and this + repository has measured what that does to a mutation sweep. +4. **What is the target?** "`cargo test --lib` is hermetic and means it" is achievable. "Every + behaviour has a hermetic test" is not, and should not be -- the six items on the bright line above + must stay kernel-only. + +## What was deliberately not queued, and what changed + +As first written this recorded no decision and queued no work, on the grounds that the proposal +changes a decision that is currently written down and well argued, and that the evidence for +amending it -- the co-tested-peer distinction -- is an argument rather than a measurement. + +The engineer settled the part that did not depend on that argument: **the current structure is +incorrect regardless of which remedy is chosen**, so the defect is now [D-49](../DESIGN-NOTES.md#d-49) +and the work is `M24`. The co-tested-peer question survives untouched as `M24.1`, which is +required to settle it **by demonstration rather than by argument** -- precisely because an argument +is what is in doubt. diff --git a/crates/windows-ioring-sys/design-sessions/DESIGN-SESSION-2026-09-21-m21-remediation-findings.md b/crates/windows-ioring-sys/design-sessions/DESIGN-SESSION-2026-09-21-m21-remediation-findings.md new file mode 100644 index 000000000..365a2ab67 --- /dev/null +++ b/crates/windows-ioring-sys/design-sessions/DESIGN-SESSION-2026-09-21-m21-remediation-findings.md @@ -0,0 +1,280 @@ +# Design session 2026-09-21: what remediating the epoch-log review taught + +A running record of findings produced **while implementing** the M21 checklist, as distinct from +[DESIGN-SESSION-2026-09-19-epoch-log-review.md](DESIGN-SESSION-2026-09-19-epoch-log-review.md), +which recorded the review that produced the items. It exists because the next review pass should +start from what this one learned rather than rediscovering it -- including the places where the +review itself was wrong. + +Updated as items complete. Entries are numbered `F-n` and never renumbered. M21 is complete as of +2026-09-21; F-1 to F-12 are its whole record. + +## Headline 2: an independent review found what this pass could not + +After M21 closed, a fresh reviewer audited the `M21.2` public surface and found a **High**-severity defect +in it, plus a pre-existing one of the same root cause, plus the reason neither was caught. All four are +fixed in `M21.6`; `F-13` to `F-15` are what they taught. + +The uncomfortable part is not that the review found defects. It is *which* defect: the author of that API +had written its tests, sabotage-verified them, and reported the sabotage in the commit message -- and the +sabotage that mattered (delete the real wait entirely) was never run, because the tests had been +restructured away from the real wait on purpose. **Self-review could not have found this, and did not.** + +## Headline 1: a review that reads code produces claims, not findings + +Two of this pass's corrections were to **the review**, not to the code it reviewed. Both were +reachability claims -- statements about which states a program can enter -- and neither could have +been settled by reading. That is the single most useful thing to carry into the next pass. + +- `F-5` below: finding `C-3` asserted a latent bug that measurement showed was unreachable at any + constants. +- `F-8`: an assertion about how an error would surface, wrong in a way only running it revealed. + +A reachability claim in a review should be written as a **question with the experiment attached**, +not as a finding. "Is this reachable if `EPOCH_SIZE` exceeds `SLOTS`? -- instrument the retry path +and run it" would have cost the same to write and would not have needed correcting afterwards. + +## Findings + +### F-1 (M21.1) -- a sweep count taken from tool output, not from a count + +The blast-radius sweep was reported as "13 files" because that is what `rg`'s summary line said; +`rg` groups two files sharing a directory prefix under one header, so the real figure was 14. The +item's own prediction ("17 matches across 10 files") was also low on both axes. + +**Carry forward:** a sweep count is a claim about the tree, so it comes from a command that counts, +per FAIL FAST rule 6. Reading it off a summary line is the same defect one layer up. + +### F-2 (M21.2) -- `SubmitIoRing` rejects a wait with no pending operation + +`SubmitIoRing` answers `E_INVALIDARG` (`0x80070057`) -- **not** a timeout -- when asked to wait for a +completion the kernel has no pending operation for. Found because a test drove the new pop loop with +a reservation that had no real SQE behind it. Now documented on `RingWait::block`, where the +precondition holds structurally because `pop_within_with` checks `outstanding()` first. + +**Carry forward:** `IoRing::run_down` makes the same call and is reachable from `Drop`. Its +behaviour when the count is non-zero but nothing is genuinely pending is worth a look in the next +pass -- see `F-4`, which is that combination actually occurring. + +### F-3 (M21.2) -- a panic path introduced and closed in the same item + +The first draft of `pop_within` computed `Instant::now() + timeout`, which panics on overflow, so +`Duration::MAX` -- a reasonable spelling of "no deadline" -- would have aborted the process. Closed +with `checked_add` before the commit, and covered by two tests, one of which reaches the overflow +branch rather than being answered by the early return. + +**Carry forward:** the next pass should check every other public entry point that accepts a +`Duration` or a timeout for the same shape. + +### F-4 (M21.2) -- a failing test can abort the harness instead of reporting + +A test that panics while a reservation is outstanding unwinds into `IoRing::drop`, whose rundown +then fails (per `F-2`) and panics a second time. Rust aborts on a double panic, so the run ends with +`STATUS_STACK_BUFFER_OVERRUN` and **no test name**. Worked around here by settling every phantom +reservation before any assertion. + +**Carry forward:** this is a diagnosability defect in the crate's own teardown, not only in the +tests. A `Drop` that can panic turns any unrelated test failure in the same file into an unnamed +abort. Worth a decision in the next pass: whether rundown failure should be reported some way other +than `debug_assert!` while unwinding. + +### F-5 (M21.3) -- the review claimed a latent bug that does not exist + +Finding `C-3` and item `M21.3` both said the epoch-commit trigger was safe at the sample's constants +but armed for anyone raising `EPOCH_SIZE` past `SLOTS`. Instrumenting the retry path to report when +the old shape would have committed produced **zero** firings at `EPOCH_SIZE` of 6, 8, 12, 16 and 24. + +It is unreachable at any constants: the predicate is true at exactly two moments -- before the first +append, and immediately after a commit -- and the arena is empty at both, because the commit waits +for a covering flush that retires every outstanding write. + +The change was still worth making, but as a **coupling** change: the old trigger was safe because of +an invariant three blocks away that nothing stated. Corrections landed in `C-3`, the M21 header, and +the item. + +### F-6 (M21.4) -- the sample had no tests, and could not have had any + +Examples are not test targets by default, so `cargo test` compiled the epoch-log sample and ran +nothing. Every claim its modules made was unbound. `test = true` on the `[[example]]` entry is what +changed that, and it applies to the whole sample rather than to the one item that needed it. + +**Carry forward:** the sample's other modules -- `record`, `replay`, `reclaim`, `checkpoint`, +`strategy` -- are now testable and still untested. `replay` is the interesting one: it is the +verifier the sample's own credibility rests on. + +### F-7 (M21.4) -- the failure path is unreachable without the injection seam + +A commit only fails if its flush fails, and a flush against a healthy temp file does not. So four of +the six new tests are gated on `fault-injection`, following the precedent in +[fault_injection.rs](../tests/fault_injection.rs) -- including its reasoning that CI's +`--all-features` job is what stops a gated test from being a test that never runs. + +### F-8 (M21.4) -- an error-surfacing assumption, wrong + +The first version of the test expected an injected `ERROR_ACCESS_DENIED` to arrive as +`io::ErrorKind::PermissionDenied`. It arrives as `Other`: the crate preserves the HRESULT in an +`IoRingError` rather than classifying it. Corrected to assert the Win32 code, which is what +`tests/fault_injection.rs` already asserts. + +**Carry forward:** the crate does not map Win32 codes onto `io::ErrorKind`. Whether it should is a +question for the next pass, not a defect -- but consumers matching on `ErrorKind` will match `Other` +for everything, and nothing currently says so where a consumer would look. + +### F-9 (M21.4) -- the milestone compounded + +`commit_and_pop` is three lines because `M21.2` published `IoRing::pop_within`. Every test in the +file would otherwise have carried its own bounded wait -- the exact duplication `M21.2` existed to +remove, reappearing immediately in the next item. + +**Carry forward:** worth checking in the next pass whether the remaining hand-written waits in the +sample and in `tests/` can now collapse onto it. `M21.5` covers two of them; there may be more. + +### F-10 (M21.5) -- the item named two sites; a census found six + +`M21.5` was written as "give strategy.rs's two wait loops a bound". Counting by command over every `.rs` +outside `target/` and the spikes found **four** unbounded wait loops and **two** more of a related +shape. Two of the four were helpers *both named `await_one`*, byte-identical, in +[failure_paths.rs](../tests/failure_paths.rs) and [kernel_span.rs](../tests/kernel_span.rs) -- neither +mentioned by the item or by the review. + +The other two were in [batch/tests.rs](../src/batch/tests.rs): registration waits written as a single +`try_pop`, which is the flake shape `pop_within` documents, in a file whose **third** such wait already +used the helper. One predicate, three sites, half-converted -- FAIL FAST rule 1 exactly, inside a single +file. + +**Carry forward:** the review found these by reading one sample, so it found what that sample contained. +A census by command is cheap and finds the population. Every item in the next pass whose subject is a +*shape* rather than a specific line should carry its census command. + +### F-11 (M21.5) -- the milestone compounded again, and measurably + +Sabotaging `Lane::classify` to stop filing flush results leaves `await_flush` waiting for a completion +that is never recorded. An unbounded loop hangs forever there. The new bound reported +`timed out after 30s waiting for a commit's flush` -- **in two seconds**, because `pop_within`'s +nothing-can-arrive early return (`F-9`, `M21.2`) answers immediately once the ring is quiesced. + +The bound is what makes the failure possible; `M21.2`'s early return is what makes it quick. Neither was +designed with the other in mind. + +### F-12 (M21.5) -- duplicated helpers do not share a name by accident + +Two independently written helpers, in two files, both called `await_one`, both the same eight lines. The +name being identical is the tell: it is what people call this operation, which is the argument for the +operation belonging to the library. It now does. +### F-13 (M21.6) -- the crate's tests never exercise asynchronous completion + +> **Corrected 2026-09-23 by `M24.6`'s sweep. The headline overstates, and it did so when written.** +> "Every fixture in this crate's tests, examples and samples opens its handle that way" is false: +> [flush_barrier.rs](../tests/flush_barrier.rs), [handover.rs](../tests/handover.rs) and +> [flush_barrier_stress.rs](../tests/flush_barrier_stress.rs) all open +> `FILE_FLAG_OVERLAPPED | FILE_FLAG_NO_BUFFERING` handles, and had done since 2026-08-28, 08-29 and +> 09-06 respectively -- weeks before this was recorded on 09-21. What is true is the narrower claim +> the measurements below actually support: the fixture *that finding was built against* was +> synchronous, and so are the epoch-log sample's. +> +> **The carry-forward escalated from the false half, and is largely unfounded because of it.** It +> says every claim about ordering, draining, the completion event and the barrier "was measured +> against operations that may have completed inline", and names D-19, D-23, D-24 and D-47 for +> re-reading. But D-23, D-24 and D-47 were measured by `flush_barrier.rs`, which is one of the three +> overlapped, unbuffered fixtures. The entry hedged in the right direction -- "the drain-ordering +> spike used `NO_BUFFERING` and pre-written extents deliberately, so it is probably fine" -- and then +> checked only the spike, not the tests that shared its shape. +> +> The shape of the error is worth more than the correction: a measurement of **one** fixture was +> generalised to **every** fixture without a census, and the alarm that followed inherited the +> generalisation. A census is one command. See `M20.6`'s findings for the case where the same claim +> *was* true -- the epoch-log sample really did run entirely on synchronous handles, which is what +> made its strategy comparison measure a pipeline that did not exist. + +Measured while building a test that needed a genuinely pending operation. **A file handle opened without +`FILE_FLAG_OVERLAPPED` is synchronous, so a ring operation against it completes inline during submit.** +Every fixture in this crate's tests, examples and samples opens its handle that way. + +The consequences are larger than the item that found it: + +| Attempt at a slow operation | Measured | +|---|---| +| Buffered read, up to 256 MiB | 3-5 us -- already poppable | +| Flush over 512 MiB of dirty cache | 3 us -- lazy writer got there first | +| Unbuffered read, 256 MiB, synchronous handle | 3 us -- completes during submit | +| Unbuffered **and** overlapped, 64 MiB and up | genuinely pending | + +So the suite has been testing the *synchronous* completion path almost exclusively. A counting waiter over +the existing flush pattern was reached in **0 of 50 trials**. + +**Carry forward, and this is the big one for the next pass:** every claim this crate makes about ordering, +draining, the completion event, and the barrier was measured against operations that may have completed +inline. D-19, D-23, D-24 and D-47 all deserve re-reading with that in mind. The drain-ordering spike used +`NO_BUFFERING` and pre-written extents deliberately, so it is probably fine -- but *probably* is exactly +the word that needs replacing with a measurement. + +### F-14 (M21.6) -- deterministic tests and honest tests are not the same thing + +The M21.2 tests were restructured onto a wait that never enters the kernel, precisely to make the loop's +deadline behaviour deterministic. That was reported in the commit message as a virtue. It was also what +let a real defect through: replacing `RingWait::block`'s whole body with an unconditional error left the +entire suite green, because nothing ever reached it. + +The isolation was correct for what it tested. The error was not adding anything that drove the real thing +alongside it -- and then describing the isolation as coverage. + +**Carry forward:** when a test double is introduced to make something deterministic, the same change owes +a test that exercises the real implementation. A mutation that deletes the real implementation should fail +something. + +### F-15 (M21.6) -- an error-vs-timeout mapping is a contract, and Win32 gets it backwards + +Every Win32 wait reports an expired bound as a *failure* code -- `ERROR_TIMEOUT` from `SubmitIoRing`, +`WAIT_TIMEOUT` from the `WaitFor*` family. Any wrapper that forwards its underlying result verbatim +therefore turns an ordinary timeout into an error, and any API that documents "`Ok(None)` means the bound +expired" is wrong the moment it does so. + +`CompletionWait` had not said which way to report it, so every third-party implementation would have +reproduced the defect independently. It says so now. + +**Carry forward:** check every other place this crate converts a Win32 wait result. The `WaitFor*` calls in +`event_loop.rs` and `model_b_multiplexed.rs` already handle `WAIT_TIMEOUT` explicitly; whether anything +else forwards a wait result blindly is worth a census. +### F-16 (M21+.1) -- the checker had a latent bug that only a probe could find + +Widening `check-borrow-surface.ps1` was verified with five probes rather than by re-reading the regex, and +one of them crashed it: a one-line body -- `pub fn f() -> &[u8] { &[] }` -- never satisfied the "line ends +with `{`" test, so the signature accumulator ran off the end of the file. The *old* script did not crash on +that shape only because it never indexed the lines again afterwards; it silently swallowed the following +lines instead, which means it could have been skipping real signatures all along. + +**Carry forward:** a checker is code, and the argument for testing it is the same as for anything else. The +negative control matters most -- a check that fires on everything is as useless as one that fires on +nothing, and only the plain-`&T` probe establishes that this one still discriminates. + +### F-17 (M21+.1) -- the blind spot had already swallowed something real + +The widened check immediately reported `IoRingErrorExt::as_ioring_error -> Option<&IoRingError>`, a public +trait method returning a borrow that predates the review by months and had never been inventoried. It is +not a hole -- the borrow is of the `io::Error` the caller owns -- but it was never *put to anyone*, which +is the whole function of the inventory. + +**Carry forward:** when a check is found to be narrow, assume it has already been narrow for a while and +look at what it let past, rather than only at the change that exposed it. + +### F-18 (M21+.1) -- the count sweep missed the tool built to stop drift + +The script's header and its failure message both said **three** shipped defects of this shape, listing +D-35, D-36 and D-43. It has been four since D-45. `M19.3` explicitly swept that count -- its archive records +"that file said 'three defects' in four places and is now four" -- and swept `DESIGN-INSTRUCTIONS.md` while +missing `check-borrow-surface.ps1`. + +**Carry forward:** the sweep looked at documentation and not at tooling. A `.ps1` file carrying prose is +still prose, and the next census of any restated fact should include `tools/`. +## Open questions this pass raised but did not answer + +1. Should `IoRing::drop`'s rundown failure be reported some way that does not abort on unwind + (`F-4`)? +2. Do the crate's other timeout-accepting entry points share `F-3`'s overflow shape? +3. Should Win32 codes map onto `io::ErrorKind`, given that consumers currently see `Other` for + everything (`F-8`)? +4. Which of the sample's now-testable modules deserve tests, and in what order (`F-6`)? + +None of these are queued as checklist items yet. They are inputs to the next review pass, which +should decide whether each is work or a non-issue -- **by measuring, not by reading**, which is the +lesson of `F-5`. diff --git a/crates/windows-ioring-sys/design-sessions/DESIGN-SESSION-2026-09-22-kernel-response-space.md b/crates/windows-ioring-sys/design-sessions/DESIGN-SESSION-2026-09-22-kernel-response-space.md new file mode 100644 index 000000000..085cab741 --- /dev/null +++ b/crates/windows-ioring-sys/design-sessions/DESIGN-SESSION-2026-09-22-kernel-response-space.md @@ -0,0 +1,217 @@ +# Design session 2026-09-22: the kernel response space + +**Decisions resulting from this session:** D-52 (the resolver technique and what it is for), and an +amendment to "Two techniques deliberately rejected" in [DESIGN-NOTES.md](../DESIGN-NOTES.md). The work +it queues is `M26` in [CHECKLIST.md](../CHECKLIST.md); `M24.1` is answered and `M24.4` is withdrawn. + +The apparatus built during the session is kept as +[kernel-response-space-probe.rs](kernel-response-space-probe.rs). It is a **demonstration**, not a +test: drop it into `tests/` and run with `--nocapture` to reproduce every figure below. + +## What the session set out to do + +`M24.1` asked one question: does a fake whose assertions are **shared** with the kernel escape the +objection that led this crate to reject a mock? That objection is specific rather than generic -- + +> Both shipped defects were the kernel behaving differently from this crate's assumptions. A mock +> *encodes* the assumption, so one written before those discoveries would have passed both bugs green +> -- it would not merely have failed to find them, it would have manufactured evidence they were +> absent. + +The item insisted the question be settled by demonstration, "because the argument is exactly what is +in doubt", and predicted two outcomes: a wrong **accounting** model would be caught by a shared suite, +and a wrong **Windows belief** would not. + +Both predictions held. Then two further cases changed the answer. + +## Case 1 and 2: the predictions, confirmed + +A five-method slice (`push_one`, `submit`, `try_pop`, `outstanding`, `pop_within`), one generic +assertion suite, two implementations -- a real `IoRing` over a temp file, and an in-memory fake. + +| fake's flaw | shared suite | +|---|---| +| none | GREEN, and the kernel agrees | +| `outstanding` decremented at push instead of at pop | **RED** -- caught | +| pops without submitting | GREEN -- slipped | +| pre-`M21.6`: an expired wait is an error | GREEN -- slipped | + +The third flaw is the one worth dwelling on, because it is not invented: it is the belief this crate +actually held until `M21.6`. `SubmitIoRing` reports an expired wait as `ERROR_TIMEOUT`, a *failure* +HRESULT, so `pop_within` returned `Err` on every ordinary timeout. A fake written before that +discovery would have encoded exactly this. + +Add a timeout assertion and the fake is caught instantly. But **that assertion could only be written +after the kernel had already revealed the answer.** The fake could never have produced it. + +## Case 3: the argument *for* co-testing, which the item did not predict + +The first pass missed the case that decides it. What happens when the shared suite itself encodes the +wrong belief, and is run against both? + +| | timeout assertion written from the pre-`M21.6` belief | +|---|---| +| kernel | **RED** -- "an expired wait returned `Ok(None)`, not the error we expected" | +| fake built from the same belief | GREEN | + +The kernel **refutes** us. The fake **confirms** us. That is the manufactured-evidence mechanism made +visible -- and it is also the escape, because running a shared assertion against the kernel is how a +wrong belief gets contradicted. A mock-only world never performs that experiment. + +So after three cases the conclusion was: the rejection stands for mocks-as-substitutes, a co-tested +peer is admissible for a narrow category, and the bright line is "accounting versus Windows +behaviour". + +**That line was wrong**, and the next case is why. + +## Case 4: an assertion that looks like a contract and is a frozen observation + +Raised in review: *the kernel's behaviour is not necessarily reproducible run to run. When we observe +it, that is an observation at a point in time, not a record of objective truth. We must not +over-index on a record of how it runs as being "right".* + +Take the assertion "after submitting, the completion is already queued". It reads like a contract. +Run it against two handles of the same API: + +| | "completion is already queued after submit" | +|---|---| +| kernel, buffered handle | GREEN | +| kernel, `NO_BUFFERING` + `OVERLAPPED`, pre-written extent | **RED** -- "nothing was queued when submit returned" | +| fake | GREEN -- it encoded whichever one its author saw | + +Opposite answers on the same API. So the assertion was never about the ring; it was about a handle, a +filesystem and a moment. And the run-to-run half was measured the same day: +[write-pending-spike.rs](spikes/write-pending-spike.rs)'s `NO_BUFFERING`-extending condition reported +**5/500 in one run and 271/500 minutes later**, same binary, same machine. + +Restate the same question as *our* contract -- "the completion arrives within a bound we specify" -- +and all three go green. That form is robust because it is a statement about what this crate promises +rather than about when the kernel happens to finish. + +### The ratchet + +The danger is worse than one bad test, and it compounds: + +1. Observe the kernel once. +2. Freeze the observation into a conformance assertion. +3. Build the fake to satisfy that assertion. +4. Three artifacts now agree -- and the agreement reads as corroboration when it is **one observation + restated three times**. + +That is CONTRACT INTEGRITY rule 1 ("a hand-written second copy of a contract rule is not a check of +the contract, it is a check of the copy") applied to platform behaviour rather than to our own. This +crate has already paid for it once: [D-47](../DESIGN-NOTES.md#d-47) is exactly this failure, where a +handful of runs showed a barrier holding, that was written down as a guarantee, and the real violation +rate was nearer one in a thousand. + +### The corrected line + +Not "accounting versus Windows behaviour". The axis is: + +**our specified contract, versus the platform's incidental behaviour.** + +- `pop_within` returns `Ok(None)` on an expired wait -- **ours**. The kernel says `ERROR_TIMEOUT`; the + wrapper translates. Legitimate to assert, and legitimate for a fake to encode. +- "`SubmitIoRing` reports `ERROR_TIMEOUT`" -- **an observation**. Dated, machine-specific, possibly not + reproducible. It belongs in a spike with a rate and provenance, never as a pass/fail assertion. +- "the completion is queued when submit returns" -- **an observation wearing a contract's clothes**, + which case 4 demonstrates. + +This is why Design Autonomy is a repository rule: we define our behaviour and choose dependencies that +satisfy it. The wrapper is the layer that absorbs kernel variation, and the conformance suite asserts +what the wrapper promises -- which is precisely what a fake can faithfully implement. + +## Case 5: the reframing, and the actual answer + +Also raised in review, and it is a different technique rather than a refinement: + +> Our code needs to work in light of all the possible ways that the platform may respond to the rings +> we submit. Some seed-derivable selection of a set of resolutions to what a given epoch's IoRing +> would turn into in terms of synchronously versus asynchronously completed items. Input state plus a +> seed gives a set of kernel responses, and we verify that `windows-ioring-sys` responds correctly to +> the kernel stimuli. + +The fake stops modelling **what Windows does** and starts modelling **what Windows is permitted to +do**. A seed picks one resolution out of that space: which operations finish inside `SubmitIoRing` +and which pend, in what order completions are posted, which fail. + +**This dissolves the mock objection rather than working around it**, because there is no belief to be +wrong about. The resolver asserts nothing about the kernel. The assertions are about *us*: does this +crate behave correctly under this resolution. + +It also inverts the problem case 4 raised. Non-reproducibility stops being a threat and becomes the +expected case -- Windows exercising a different point in a space the tests already sweep. A run-to-run +change like 5/500 to 271/500 is two samples from a space covered by construction. + +A minimal resolver in the probe -- SplitMix64 over one seed, permuting completion order -- was run +against a consumer that assumes completions arrive in submission order: + +``` +a consumer assuming FIFO completion order: + broke under 189 of 200 seeds + first at seed 0: completions arrived as [3, 1, 4, 2], not [1, 2, 3, 4] +``` + +Some seeds pass and some fail, which is the point. A fixed fake reports whichever single answer it +encoded. + +## What this does and does not add to the toolkit + +[DESIGN-NOTES.md](../DESIGN-NOTES.md) already draws the boundary this sits on: + +> All five techniques check this crate's code against **this crate's stated contract**. None of them +> can tell you the stated contract is wrong. + +The resolver is a sixth technique, and it does something none of the five do: it checks the code +against a **space** of platform behaviours rather than against one. It still cannot tell you the +stated contract is wrong -- only a spike does that. What it can tell you is that the code is brittle +to variation *inside* the space, which nothing in the toolkit currently detects. + +It composes with what exists rather than replacing it. +[generated_sequences.rs](../tests/generated_sequences.rs) (M17.3, `D-41`) already generates the +**input** space and runs it against the real kernel; the resolver generates the **response** space. +Both seeded, both replayable from one number, and together a two-dimensional exploration. + +## The hard part, which is the whole design + +**Where the permitted space comes from.** Derive it from observation and the trap closes again. It has +to be a deliberate specification -- "we will tolerate these behaviours" -- written wider than anything +observed, on purpose. That makes this crate's model of Windows an explicit, reviewable, versioned +artifact instead of an accident of whichever machine ran the tests last. + +Three consequences that should be decided rather than defaulted: + +1. **Constraints must be modelled too, or the tests demand over-defensive code.** `D-23`'s + covering-flush guarantee held with zero failures in roughly 4,500 trials. If the resolver may + violate it, we would write code defending against a kernel that breaks a documented guarantee. + Making that an explicit call is the improvement; it is still a call. +2. **The seam is invasive.** A resolver has to sit under the `windows-sys` calls -- `SubmitIoRing`, + `PopIoRingCompletion`, the `Build*` family -- which is substantially more than `M24.2`'s field + split, on a published crate. +3. **Kernel tests do not go away; their job changes and improves.** They stop being "run everything + against Windows" and become "confirm reality stays *inside* the declared space". Better defined, + and if Windows ever moves outside it, that test is what reports something real. + +## What it would have caught, stated as reasoning rather than measurement + +The FIFO consumer in case 5 is structurally the same defect as `D-47` -- a consumer assuming an +ordering the platform does not guarantee -- so that class is covered. `M21.6`'s `ERROR_TIMEOUT` defect +is covered provided the space includes "a wait may expire and report it as a failure HRESULT". + +**Both are analogies from the demonstration, not separate measurements.** `M26` carries a calibration +item to re-inject both defects against a real resolver, because `D-41`'s corollary is the most +transferable rule this crate has produced: *a green result from an instrument nobody has shown can go +red is not evidence.* This session produced two apparatus failures of exactly that kind -- a first +draft that ran each spike condition once, and a case-4 harness whose bare flush completed inline on +both handles and so could not discriminate until it was rebuilt around aligned writes. + +## Consequences for the plan + +- `M24.1` is **answered**: the fake was the wrong instrument. Recorded, and the decision amended. +- `M24.4` is **withdrawn**. A shared conformance suite over a hand-written fake is superseded by the + resolver, and its stated purpose -- hermetic bookkeeping tests -- is served better by one. +- `M24` **no longer depends on either**. Its goal is a hermetic lib suite, and relocation plus the + accounting extraction achieve that on their own. It is now unconditional. +- `M26` is new, and is justified by **what it catches** rather than by hermeticity. That is a real + distinction: hermeticity is achievable without it, so the resolver has to earn its place on the + defect class it detects. diff --git a/crates/windows-ioring-sys/design-sessions/DESIGN-SESSION-2026-09-23-pending-inventory.md b/crates/windows-ioring-sys/design-sessions/DESIGN-SESSION-2026-09-23-pending-inventory.md new file mode 100644 index 000000000..ca0280a6c --- /dev/null +++ b/crates/windows-ioring-sys/design-sessions/DESIGN-SESSION-2026-09-23-pending-inventory.md @@ -0,0 +1,183 @@ +# Design session -- the pending-token inventory (2026-09-23) + +**Summary.** An exploration of `M23.3`, which asks whether this crate should offer a +pending-operations map and a slot arena over it. The session produced a working spike +(`src/pending.rs`), three measured findings that falsified earlier claims -- two of them +mine, stated confidently and wrongly -- and a design alternative I dismissed on a reason +that turned out not to hold. No decision is taken here; `M23.3` still owns that. + +## What the checklist said, and what a census found + +`M23.3` recorded that nine sites keep a map from `UserData` to an unclaimed `Token`. A +fresh census found otherwise: + +- **There are about twelve**, not nine. The list missed `model_a_delivery.rs`, + `model_b_multiplexed.rs` and `generated_sequences.rs`. +- **Only about a third keep the bare map described.** The rest carry per-operation + sidecar data -- a slot index, an expected length, a phase, a sequence number, a round. + `flush_barrier_stress.rs` keeps two sidecar maps beside its tokens. +- **Exactly one site needs synchronisation**: `model_a_delivery.rs` wraps its map in a + `Mutex`, because it is the Model A path where completions land on pool threads. + +The count was the item's main evidence, and it was taken before `M24` relocated eleven +tests. **The duplicated thing is not the map**; it is the claim discipline over a map +whose value type differs at nearly every site. That is why the spike is `Pending` +and not `Pending` -- the generic is a finding, not a convenience. + +## The two types already existed, related by hand + +`RingContract` is public, always on, and not feature-gated. Every consumer that uses it +drives it *in parallel with* its own map: + +```text +self.contract.observe_push(token.id()); +self.in_flight.insert(token.id(), InFlight { token, slot }); +``` + +The same event, recorded twice, by hand, at every site -- a restatement in the +repository's own terms, and one that can drift in both directions. The oracle's value +depends on being driven correctly by the very code it exists to check. + +So the pair the engineer was reaching for -- "if the rule is heavy, can we have two +related types?" -- already exists. What was missing is that the light one should +**drive** the heavy one, so a consumer updates one thing and both stay true. + +## Synchronisation: no, and the reason is already paid for + +`pop_within` takes `&mut self`, and `IoRing` is `Send` but **not `Sync`** -- there is only +`unsafe impl Send for IoRing`. Whoever pops a completion already holds exclusive access, +so a map reachable through that same `&mut self` needs no `Arc`, no `Mutex`, and no +interior mutability. The exclusivity exists already. + +## Falsified: the type-erasure objection + +**This was my claim and it was wrong.** I argued a map owned by the ring would have to +store heterogeneous `Token`, forcing `Box` and a downcast at the claim site, +and that this killed the idea. The engineer asked why, since for any particular ring the +type is fixed. Checking: + +- **Per-ring monomorphisation holds for every real consumer.** The epoch-log's log ring + carries `Token` for appends, and its commits are *tokenless*. + `checkpoint.rs` looked like a counterexample because it has two maps, but its local + `Pending` is plain bookkeeping and only one map holds tokens. +- **Where it genuinely does not hold, the answer is a closed enum, not `dyn Any`** -- and + this tree already has one. `generated_sequences.rs` puts eight token types on a single + ring behind `enum Held`, with an exhaustive `match` in `claim` and no runtime type + check at all. + +So a generic `IoRing` owning the map is viable, which matters because it is the shape +that would make the ring *notify* rather than be told -- and drift between the ring and +the inventory structurally impossible rather than merely discouraged. Its real costs are +different from the one I asserted: `IoRing` is not generic today, so this is a breaking +change to a published crate; a consumer mixing shapes writes a `Held`-style enum; and +tokenless pushes still need a story. + +## Why `commit.rs` uses `flush_raw`, and what follows + +Asked during the session, and the answer explains the tokenless commits above. + +The safe `flush(&F, ..) -> Token` requires an **owned, guarded** +file -- `SharedFile` is `Arc`, and the token holds a clone of that `Arc`, +which is what keeps the handle alive for the kernel. `Committer::commit` is handed a bare +`RawHandle` that the log's `File` owns, so it cannot build a `SharedFile` without taking +ownership and closing the log's handle out from under it. It therefore takes the `unsafe` +`flush_raw`, whose safety argument is hand-written: *"`file` is the log's own handle and +outlives every operation pushed here; the log drains to empty before it closes."* + +**Why that makes the commit tokenless, precisely.** A write's token guards the *buffer*, +so `write_registered_raw` still returns one even with a raw file. A flush has no buffer -- +its only possible guard is the file -- so with a raw handle there is nothing to hold, and +`flush_raw` returns a bare `usize`. + +Three implications: + +1. **The commit path cannot be in any token inventory as currently plumbed.** Not because + flushes are special, but because this sample passes a borrowed handle. +2. **A compile-time guard was traded for a prose argument**, and the prose is load-bearing: + nothing enforces "the log drains to empty before it closes". +3. **It is the same shape as `placement.rs`'s constraint**, found earlier the same day: a + borrowed `RawHandle` and an owned handle make different APIs reachable. Two independent + places in one sample reach for a lower-level call for the same plumbing reason, which + suggests the plumbing rather than the calls is the thing to look at -- and `M25.3` + already reopens how the log is opened. + +## What the spike established, and what it cannot + +**Established, by test and by sabotage:** + +- One call site keeps the map and the oracle in step. Cutting the wiring is caught. +- An unclaimed token is loud at teardown rather than a silent deliberate leak. Removing + the `Drop` guard is caught. +- `Pending` fits a real consumer: converting `append.rs` removed its `InFlight` + struct and its hand-driven `observe_completion`/`observe_claim` pair. +- The `M22.2` ordering defect is now caught by an assertion, via a test that drives a + failed write through the injection seam. + +**Not established, and two of these are corrections to things I asserted:** + +- **The ring does not notify anyone.** `Pending` is consumer-driven. Nothing forces a + minted token into the inventory, so a consumer can take a `Token` from `Batch` and never + register it, and no type, test or oracle notices. The spike removed drift between the map + and the oracle; it did not remove drift between the ring and the map. +- **One consumer is converted, not twelve.** This is a worked example, not a property of + the crate. +- **`Pending::checked()` owning the oracle creates a decoy hazard.** A consumer that + already had a `RingContract` keeps a field that is never written to again. The + conversion did exactly that, and its teardown `assert_quiescent()` passed *vacuously* -- + compiled, ran, every test green. Sabotage confirms nothing catches it. Found by reading. +- **Ring teardown interaction is untested.** If a `Pending` and an `IoRing` both drop with + operations outstanding, the ordering of their guards is unexamined -- and the session + already found one drop-order surprise (see below). + +## Two findings about the instruments themselves + +**A feature-gated test is invisible to the sabotage harness.** The sweep first reported the +`M22.2` regression as *survived*. The new tests are behind `fault-injection`, off by +default, so the harness compiled them out and the sabotage landed in code nothing +exercised. Fixed with a manifest `testArgs` carrying `--all-features`. This is the trap the +repository already documents for cargo-mutants, arriving through a different tool: a +feature-gated guard and an absent guard are indistinguishable to a runner that does not +enable the feature. + +**A failing test that leaves registered buffers outstanding aborts instead of reporting.** +`RegisteredBuffers::drop` refuses to free while operations are outstanding -- correct, from +`M5.3` -- via a bare `debug_assert!` that does not check `std::thread::panicking()`. So the +assertion fires and names the leak, then unwinding drops the arena, the `debug_assert` +panics during unwind, and the process aborts with `STATUS_STACK_BUFFER_OVERRUN`. Detection +is not weakened; the report is. Queued as `M23.4`. + +## On requiring `finish`, and why it is not failable + +Rust has no linear types, so nothing can force a method call on a value the caller owns. +Three rungs were considered: + +- **`#[must_use]`** stops the return value being ignored, not the call being skipped. +- **The drop bomb** -- `Drop` panics when tokens are still held -- is the real enforcement, + and is what the spike implements. It is suppressed while already panicking, because a + second panic during unwind aborts and replaces the original failure. +- **A scoped constructor** (`Pending::scope(|p| ...)`) would genuinely force it, since the + consumer never owns the value. Not built: it imposes a control-flow shape that suits the + appender but not the tests that thread a map through several helpers. + +**`finish` reports rather than fails**, and the engineer's observation is why: it *consumes* +the map, so an `Err` would leave nothing to retry with. A consuming method in a linear +discipline is normally total for exactly that reason. `-> Result<(), _>` would buy +`#[must_use]` from the language rather than from an attribute, at the cost of implying a +recoverable state that does not exist; the spike keeps `-> Vec` with the +attribute. + +The residual worry -- that a path is found so late that a drop bomb ships -- is real but +narrower than it looks. A failed completion is not unreachable, only untested, and +`Completion::with_injected_failure` reaches it on demand. What remains is the set of +conditions the seam cannot manufacture, which is worth naming rather than treating the +whole class as unreachable. + +## Open, for `M23.3` to decide + +1. Public type, documented pattern, or `test-util` module. Six of the twelve sites are + tests, and test convenience is a weak reason to grow permanent surface. +2. Whether `checked()` survives in its current form, given the decoy hazard nothing catches. +3. Whether the stronger shape -- a generic `IoRing` owning the map, so the ring notifies + -- is worth a breaking change to a published crate. +4. What it refuses to decide. Batching, ordering and slot choice are caller questions; the + sharper refusal is that the map does not decide whether you are checked. diff --git a/crates/windows-ioring-sys/design-sessions/kernel-response-space-probe.rs b/crates/windows-ioring-sys/design-sessions/kernel-response-space-probe.rs new file mode 100644 index 000000000..eaaa15c5d --- /dev/null +++ b/crates/windows-ioring-sys/design-sessions/kernel-response-space-probe.rs @@ -0,0 +1,645 @@ +// Copyright (c) 2026 Mike Grier +//! **Throwaway probe for `M24.1`.** Not a permanent test -- it exists to +//! settle one design question by demonstration, and is deleted once the +//! decision is recorded. +//! +//! # The question +//! +//! [`DESIGN-NOTES.md`]'s "Two techniques deliberately rejected" refuses a mock +//! `IoRing`, because a mock *encodes* an assumption: one written before this +//! crate's two shipped defects were found "would not merely have failed to +//! find them, it would have manufactured evidence they were absent". +//! +//! `M24` wants a hermetic unit suite, and the remedy it would most like is a +//! fake whose **assertions are shared** with the kernel -- one conformance +//! suite, run against both. `M24.1` asks whether sharing escapes the +//! objection, and insists it be settled by demonstration rather than argument, +//! because the argument is what is in doubt. +//! +//! # The two predictions +//! +//! 1. Give the fake a wrong **accounting** model. The shared suite must go red +//! on the fake and stay green on the kernel. +//! 2. Give the fake a wrong **Windows belief**. The shared suite must *not* +//! catch it -- the expected result, and the reason `D-49`'s bright line is +//! a finding rather than a hedge. +//! +//! A co-tested peer is defensible only if both halves behave as predicted. +//! Run with `--nocapture` to read the verdicts. + +use std::io; +use std::os::windows::fs::OpenOptionsExt; + +use windows_ioring_sys::{ + Batch, FlushCoverage, FlushMode, IoRing, NumaBuffer, PushOptions, SharedFile, WriteCaching, +}; + +/// The slice of ring behaviour this probe shares between the two peers. +/// +/// Deliberately the *handle-free accounting* `M24.2` measured as separable: +/// push something, submit, pop it, and keep an outstanding count. That is the +/// part a fake could legitimately own. +trait RingLike { + fn push_one(&mut self) -> io::Result; + fn submit(&mut self) -> io::Result<()>; + fn try_pop(&mut self) -> io::Result>; + fn outstanding(&self) -> usize; + /// A bounded wait. Included because this is where a *real* defect lived: + /// `SubmitIoRing` reports an expired wait as `ERROR_TIMEOUT`, a failure + /// HRESULT, so `pop_within` returned `Err` on every ordinary timeout until + /// `M21.6`. Nobody guessed that; the kernel said it. + fn pop_within(&mut self, timeout: std::time::Duration) -> io::Result>; +} + +// ------------------------------------------------------------- the kernel --- + +struct RealRing { + ring: IoRing, + file: SharedFile, + path: std::path::PathBuf, + pending: Vec>, + writes: Vec>, + /// Push sector-aligned writes rather than a flush. A bare flush with + /// nothing dirty completes inline whatever the handle flags are, which is + /// how the first version of case 4 failed to discriminate. + write_mode: bool, + next_offset: u64, +} + +impl RealRing { + fn new(tag: &str) -> io::Result { + Self::with_flags(tag, 0, false) + } + + /// A real ring over a handle opened with `flags`, optionally over an + /// extent written beforehand. + /// + /// The parameters exist to make one point measurable: whether an operation + /// has completed by the time `SubmitIoRing` returns is a property of the + /// *handle*, not of the API. `write-pending-spike.rs` measured that + /// directly -- buffered never pended over 500 trials, `NO_BUFFERING` over a + /// pre-written extent pended every time. + fn with_flags(tag: &str, flags: u32, prewrite: bool) -> io::Result { + let path = std::env::temp_dir().join(format!( + "windows-ioring-sys-m24-1-{}-{tag}.tmp", + std::process::id() + )); + if prewrite { + std::fs::write(&path, vec![0_u8; 64 * 1024])?; + } + let file = SharedFile::new( + std::fs::OpenOptions::new() + .create(!prewrite) + .truncate(!prewrite) + .write(true) + .custom_flags(flags) + .open(&path)? + .into(), + ); + Ok(Self { + ring: IoRing::new(64, 128)?, + file, + path, + pending: Vec::new(), + writes: Vec::new(), + write_mode: prewrite, + next_offset: 0, + }) + } +} + +impl Drop for RealRing { + fn drop(&mut self) { + let _ = self.ring.run_down(); + let _ = std::fs::remove_file(&self.path); + } +} + +impl RingLike for RealRing { + fn push_one(&mut self) -> io::Result { + if self.write_mode { + let buffer = NumaBuffer::new(4096, None)?; + let offset = self.next_offset; + self.next_offset += 4096; + let mut batch = Batch::new(&mut self.ring); + let token = batch.write( + &self.file, + buffer, + offset, + PushOptions::new(), + WriteCaching::Cached, + )?; + let id = token.id(); + self.writes.push(token); + return Ok(id); + } + let mut batch = Batch::new(&mut self.ring); + let token = batch.flush(&self.file, FlushCoverage::Unordered, FlushMode::Default)?; + let id = token.id(); + self.pending.push(token); + Ok(id) + } + + fn submit(&mut self) -> io::Result<()> { + Batch::new(&mut self.ring).submit()?; + Ok(()) + } + + fn try_pop(&mut self) -> io::Result> { + let Some(completion) = self.ring.try_pop()? else { + return Ok(None); + }; + let id = completion.user_data(); + if let Some(index) = self.pending.iter().position(|t| t.id() == id) { + let token = self.pending.swap_remove(index); + let _ = token.claim_if(&completion); + } else if let Some(index) = self.writes.iter().position(|t| t.id() == id) { + let token = self.writes.swap_remove(index); + let _ = token.claim_if(&completion); + } + Ok(Some(id)) + } + + fn outstanding(&self) -> usize { + self.ring.outstanding() + } + + fn pop_within(&mut self, timeout: std::time::Duration) -> io::Result> { + Ok(self.ring.pop_within(timeout)?.map(|c| c.user_data())) + } +} + +// --------------------------------------------------------------- the fake --- + +/// Which wrongness this fake was built with. +#[derive(Clone, Copy, PartialEq, Eq)] +enum Flaw { + /// A faithful peer. + None, + /// Prediction 1: outstanding is decremented at push rather than at pop. + /// This is a bookkeeping rule *we* own, and it is fully specified by us. + WrongAccounting, + /// Prediction 2a: a belief about Windows -- that a completion is + /// available without submitting. No accounting rule is violated; the fake + /// simply models the platform wrongly. + WrongWindowsBelief, + /// Prediction 2b, and the stronger case because it is not invented: the + /// belief this crate ACTUALLY HELD until M21.6 -- that an expired + /// bounded wait is an error rather than Ok(None). A fake written before + /// that discovery would have encoded exactly this. + PreM216TimeoutBelief, +} + +struct FakeRing { + flaw: Flaw, + next_id: usize, + queued: Vec, + completed: Vec, + outstanding: usize, +} + +impl FakeRing { + fn new(flaw: Flaw) -> Self { + Self { + flaw, + next_id: 1, + queued: Vec::new(), + completed: Vec::new(), + outstanding: 0, + } + } +} + +impl RingLike for FakeRing { + fn push_one(&mut self) -> io::Result { + let id = self.next_id; + self.next_id += 1; + self.queued.push(id); + self.outstanding += 1; + if self.flaw == Flaw::WrongAccounting { + // The sabotage: released here instead of when the completion is + // observed. + self.outstanding -= 1; + } + if self.flaw == Flaw::WrongWindowsBelief { + // The sabotage: this fake thinks the kernel starts work as soon as + // it is queued, so a completion is poppable without a submit. + self.completed.push(id); + self.queued.pop(); + } + Ok(id) + } + + fn submit(&mut self) -> io::Result<()> { + self.completed.append(&mut self.queued); + Ok(()) + } + + fn try_pop(&mut self) -> io::Result> { + if self.completed.is_empty() { + return Ok(None); + } + let id = self.completed.remove(0); + if self.flaw != Flaw::WrongAccounting { + self.outstanding -= 1; + } + Ok(Some(id)) + } + + fn outstanding(&self) -> usize { + self.outstanding + } + + fn pop_within(&mut self, _timeout: std::time::Duration) -> io::Result> { + if self.completed.is_empty() { + if self.flaw == Flaw::PreM216TimeoutBelief { + // What a fake written from the old understanding would do. + return Err(io::Error::from_raw_os_error(1460)); + } + return Ok(None); + } + self.try_pop() + } +} + +// ------------------------------------------------- the shared conformance --- + +/// Every assertion both peers must satisfy. +/// +/// This is the *whole* point of the technique under evaluation: one suite, two +/// implementations. Note what it asserts -- push, submit, pop, and the count +/// in between -- and note what it does not think to ask, which is where +/// prediction 2 lives. +fn shared_suite(ring: &mut R) -> Result<(), String> { + if ring.outstanding() != 0 { + return Err(format!( + "a fresh ring has {} outstanding, expected 0", + ring.outstanding() + )); + } + + ring.push_one().map_err(|e| e.to_string())?; + if ring.outstanding() != 1 { + return Err(format!( + "after one push, outstanding is {}, expected 1", + ring.outstanding() + )); + } + + ring.submit().map_err(|e| e.to_string())?; + if ring.outstanding() != 1 { + return Err(format!( + "submitting does not complete anything, so outstanding should still be 1, got {}", + ring.outstanding() + )); + } + + // Drain, bounded: the kernel may not have posted it yet. + let mut popped = 0; + let deadline = std::time::Instant::now() + std::time::Duration::from_secs(5); + while popped < 1 { + if ring.try_pop().map_err(|e| e.to_string())?.is_some() { + popped += 1; + } else if std::time::Instant::now() > deadline { + return Err("the completion never arrived".to_string()); + } + } + + if ring.outstanding() != 0 { + return Err(format!( + "claiming the completion releases it, so outstanding should be 0, got {}", + ring.outstanding() + )); + } + + if ring.try_pop().map_err(|e| e.to_string())?.is_some() { + return Err("an empty ring popped something".to_string()); + } + + Ok(()) +} + +/// The assertion that was **added after** the kernel taught us to ask. +/// +/// This is the whole argument in one function. It is not in +/// [`shared_suite`] above, because the suite above is what someone would +/// write from the understanding this crate held before `M21.6`: push, submit, +/// pop, count. Asking "and what does a bounded wait do when nothing arrives?" +/// only occurs to someone who has been surprised by the answer. +fn timeout_assertion(ring: &mut R) -> Result<(), String> { + match ring.pop_within(std::time::Duration::from_millis(20)) { + Ok(None) => Ok(()), + Ok(Some(id)) => Err(format!("an empty ring produced completion {id}")), + Err(e) => Err(format!("an expired wait reported an error: {e}")), + } +} + +/// The same question, asked from the **wrong** belief -- the one this crate +/// actually held before `M21.6`. +/// +/// This is the case the first version of this probe did not test, and it is +/// the crux. If a shared suite encodes a mistaken belief about Windows, what +/// happens when it is run against the kernel? A mock-only world never asks. +fn timeout_assertion_from_the_old_belief(ring: &mut R) -> Result<(), String> { + match ring.pop_within(std::time::Duration::from_millis(20)) { + // The pre-M21.6 understanding: an expired wait is a failure. + Err(_) => Ok(()), + Ok(None) => Err("an expired wait returned Ok(None), not the error we expected".to_string()), + Ok(Some(id)) => Err(format!("an empty ring produced completion {id}")), + } +} + +/// An assertion that **looks** like a contract and is actually a frozen +/// observation. +/// +/// "After submitting, the completion is already there." Nothing in this +/// crate's API promises that, and Windows documents nothing about when a ring +/// operation completes relative to `SubmitIoRing`. It is simply what a +/// developer sees on the handle they happened to test on -- and writing it +/// down turns one run's testimony into a rule the suite will enforce forever. +fn completion_is_already_queued_after_submit(ring: &mut R) -> Result<(), String> { + ring.push_one().map_err(|e| e.to_string())?; + ring.submit().map_err(|e| e.to_string())?; + match ring.try_pop().map_err(|e| e.to_string())? { + Some(_) => Ok(()), + None => Err("nothing was queued when submit returned".to_string()), + } +} + +/// The same question asked as **our** contract instead of the platform's +/// behaviour: the completion arrives within a bound we specify. +/// +/// This is robust to the difference the previous function freezes, because it +/// is a statement about what this crate promises rather than about when the +/// kernel happens to finish. +fn completion_arrives_within_the_bound(ring: &mut R) -> Result<(), String> { + ring.push_one().map_err(|e| e.to_string())?; + ring.submit().map_err(|e| e.to_string())?; + match ring.pop_within(std::time::Duration::from_secs(5)) { + Ok(Some(_)) => Ok(()), + Ok(None) => Err("the completion did not arrive within the bound".to_string()), + Err(e) => Err(format!("the wait failed: {e}")), + } +} + +// ------------------------------------ a resolver over the permitted space --- + +/// A fake that does **not** model what Windows does. It models what Windows is +/// *permitted* to do, and a seed picks one resolution out of that space. +/// +/// This is the distinction the rest of this probe was circling. A fake built +/// from observation freezes one run's testimony (case 4). A resolver takes the +/// platform's freedom as an **input dimension**: for each submitted operation +/// it decides, from the seed, whether the work finishes inside `SubmitIoRing` +/// or pends, and in what order completions are posted. Then the assertion is +/// about *us* -- does this crate behave correctly under that resolution -- +/// rather than about the kernel. +/// +/// Nothing here claims the kernel does these things. The claim is that it is +/// **allowed** to, so our code has to survive them. That makes the crate's +/// model of Windows an explicit, reviewable artifact instead of an accident of +/// whichever machine the tests last ran on. +struct Resolver { + state: u64, + /// Ops submitted but not yet resolved, in submission order. + queued: Vec, + /// Resolved completions, in the order this resolution posts them. + posted: Vec, + next_id: usize, + outstanding: usize, +} + +impl Resolver { + fn new(seed: u64) -> Self { + Self { + state: seed.wrapping_mul(0x9E37_79B9_7F4A_7C15) | 1, + queued: Vec::new(), + posted: Vec::new(), + next_id: 1, + outstanding: 0, + } + } + + /// SplitMix64, matching the generator this crate already uses (`D-41`): + /// one number replays a whole run. + fn next(&mut self) -> u64 { + self.state = self.state.wrapping_add(0x9E37_79B9_7F4A_7C15); + let mut z = self.state; + z = (z ^ (z >> 30)).wrapping_mul(0xBF58_476D_1CE4_E5B9); + z = (z ^ (z >> 27)).wrapping_mul(0x94D0_49BB_1331_11EB); + z ^ (z >> 31) + } + + fn push(&mut self) -> usize { + let id = self.next_id; + self.next_id += 1; + self.queued.push(id); + self.outstanding += 1; + id + } + + /// Resolve everything queued. Two freedoms the platform actually has, and + /// that this crate has measured: work may finish inside submit or pend + /// (`write-pending-spike.rs`), and completion order is unspecified. + fn submit(&mut self) { + let mut batch: Vec = self.queued.drain(..).collect(); + // Fisher-Yates over the seed: any posting order is permitted. + for i in (1..batch.len()).rev() { + let j = (self.next() % (i as u64 + 1)) as usize; + batch.swap(i, j); + } + for id in batch { + // Inline or pending is invisible from here -- either way the + // completion becomes available. What differs is whether it is + // there the instant submit returns, which is what a consumer must + // not assume. + self.posted.push(id); + } + } + + fn try_pop(&mut self) -> Option { + if self.posted.is_empty() { + return None; + } + self.outstanding -= 1; + Some(self.posted.remove(0)) + } +} + +/// A consumer with a hidden assumption: that completions arrive in the order +/// the operations were submitted. +/// +/// Nothing in this crate promises that, and `DESIGN-NOTES.md` has a whole +/// section on the ring inviting exactly this belief. The question is whether a +/// seeded resolver finds the assumption that a fixed fake cannot. +fn a_consumer_that_assumes_fifo(seed: u64, ops: usize) -> Result<(), String> { + let mut r = Resolver::new(seed); + let submitted: Vec = (0..ops).map(|_| r.push()).collect(); + r.submit(); + + let mut observed = Vec::new(); + while let Some(id) = r.try_pop() { + observed.push(id); + } + if observed != submitted { + return Err(format!("completions arrived as {observed:?}, not {submitted:?}")); + } + Ok(()) +} + +fn verdict(label: &str, result: Result<(), String>) -> bool { + match &result { + Ok(()) => println!(" {label:<38} GREEN"), + Err(why) => println!(" {label:<38} RED -- {why}"), + } + result.is_ok() +} + +#[test] +fn m24_1_does_a_shared_suite_escape_the_mock_objection() { + println!("\nM24.1: can a co-tested fake be trusted?\n"); + + println!("Baseline -- a faithful fake and the kernel must agree:"); + let kernel_ok = verdict("kernel", { + let mut r = RealRing::new("base").expect("a ring"); + shared_suite(&mut r) + }); + let faithful_ok = verdict( + "fake (faithful)", + shared_suite(&mut FakeRing::new(Flaw::None)), + ); + + println!("\nPrediction 1 -- a wrong ACCOUNTING model must be caught:"); + let accounting_caught = !verdict( + "fake (outstanding decremented early)", + shared_suite(&mut FakeRing::new(Flaw::WrongAccounting)), + ); + + println!("\nPrediction 2 -- a wrong WINDOWS BELIEF must slip through:"); + let belief_slipped = verdict( + "fake (pops without submitting)", + shared_suite(&mut FakeRing::new(Flaw::WrongWindowsBelief)), + ); + let historical_slipped = verdict( + "fake (pre-M21.6: timeout is an error)", + shared_suite(&mut FakeRing::new(Flaw::PreM216TimeoutBelief)), + ); + + println!("\nAnd now the same fake, against an assertion written AFTER the"); + println!("kernel taught us to ask it:"); + let kernel_timeout_ok = verdict("kernel", { + let mut r = RealRing::new("timeout").expect("a ring"); + timeout_assertion(&mut r) + }); + let historical_caught_once_asked = !verdict( + "fake (pre-M21.6: timeout is an error)", + timeout_assertion(&mut FakeRing::new(Flaw::PreM216TimeoutBelief)), + ); + + println!("\n--- verdict ---"); + println!("\nCase 3 -- the suite itself encodes the WRONG belief, and is run"); + println!("against both. This is what a mock-only world never does:"); + let kernel_refutes_us = !verdict("kernel", { + let mut r = RealRing::new("oldbelief").expect("a ring"); + timeout_assertion_from_the_old_belief(&mut r) + }); + let fake_confirms_us = verdict( + "fake (built from the same wrong belief)", + timeout_assertion_from_the_old_belief(&mut FakeRing::new(Flaw::PreM216TimeoutBelief)), + ); + + println!("\nCase 4 -- an assertion that LOOKS like a contract but is a frozen"); + println!("observation. Same API, same suite, two handle configurations:"); + const OVERLAPPED: u32 = 0x4000_0000; + const NO_BUFFERING: u32 = 0x2000_0000; + let buffered_says_yes = verdict("kernel, buffered handle", { + let mut r = RealRing::with_flags("frozen-buf", 0, false).expect("a ring"); + completion_is_already_queued_after_submit(&mut r) + }); + let unbuffered_says_no = !verdict("kernel, NO_BUFFERING + OVERLAPPED", { + let mut r = RealRing::with_flags("frozen-nb", OVERLAPPED | NO_BUFFERING, true) + .expect("a ring"); + completion_is_already_queued_after_submit(&mut r) + }); + println!(" ... and the fake, which encoded whichever one its author saw:"); + let fake_agrees_with_one = verdict( + "fake (faithful)", + completion_is_already_queued_after_submit(&mut FakeRing::new(Flaw::None)), + ); + + println!("\n The same assertion stated as OUR contract instead survives both:"); + let bound_buffered = verdict("kernel, buffered handle", { + let mut r = RealRing::with_flags("bound-buf", 0, false).expect("a ring"); + completion_arrives_within_the_bound(&mut r) + }); + let bound_unbuffered = verdict("kernel, NO_BUFFERING + OVERLAPPED", { + let mut r = RealRing::with_flags("bound-nb", OVERLAPPED | NO_BUFFERING, true) + .expect("a ring"); + completion_arrives_within_the_bound(&mut r) + }); + let bound_fake = verdict( + "fake (faithful)", + completion_arrives_within_the_bound(&mut FakeRing::new(Flaw::None)), + ); + println!("\nCase 5 -- a resolver over the PERMITTED space, seeded. The fake no"); + println!("longer models what Windows does; it models what Windows may do:"); + let mut broke = 0; + let mut first_break = None; + for seed in 0..200u64 { + if let Err(why) = a_consumer_that_assumes_fifo(seed, 4) { + broke += 1; + if first_break.is_none() { + first_break = Some((seed, why)); + } + } + } + println!(" a consumer assuming FIFO completion order:"); + println!(" broke under {broke} of 200 seeds"); + match &first_break { + Some((seed, why)) => println!(" first at seed {seed}: {why}"), + None => println!(" never broke -- the resolver is not exploring"), + } + let resolver_discriminates = broke > 0 && broke < 200; + println!( + " resolver explores rather than asserting one answer: {resolver_discriminates}" + ); + println!(" (some seeds pass and some fail, which is the point -- a fixed"); + println!(" fake would have reported whichever single answer it encoded)"); + println!("\n--- summary ---"); + println!(" kernel green ......................... {kernel_ok}"); + println!(" faithful fake green .................. {faithful_ok}"); + println!(" wrong accounting caught .............. {accounting_caught}"); + println!(" wrong Windows belief slipped ......... {belief_slipped}"); + println!(" real pre-M21.6 belief slipped ........ {historical_slipped}"); + println!(" ... and caught once the suite asked .. {historical_caught_once_asked}"); + println!(" kernel satisfies the new assertion ... {kernel_timeout_ok}"); + println!(" kernel REFUTES our wrong belief ...... {kernel_refutes_us}"); + println!(" fake CONFIRMS our wrong belief ....... {fake_confirms_us}"); + println!(" frozen obs: buffered says YES ........ {buffered_says_yes}"); + println!(" frozen obs: unbuffered says NO ....... {unbuffered_says_no}"); + println!(" frozen obs: fake picks one ........... {fake_agrees_with_one}"); + println!(" our-contract form holds everywhere ... {}", bound_buffered && bound_unbuffered && bound_fake); + + let as_predicted = kernel_ok + && faithful_ok + && accounting_caught + && belief_slipped + && historical_slipped + && historical_caught_once_asked + && kernel_timeout_ok; + println!( + "\n both halves behaved as predicted: {as_predicted}\n\n\ + Reading it: the shared suite is strong over accounting, which is a rule\n\ + WE own and fully specify. It is blind to any Windows belief it does not\n\ + already assert -- and the last two lines are the point, because the\n\ + assertion that catches the historical defect could only be written\n\ + AFTER the kernel had already revealed it. The fake never could have.\n" + ); + + assert!(kernel_ok, "the kernel must satisfy the shared suite"); + assert!(faithful_ok, "a faithful fake must satisfy it too"); + assert!( + kernel_timeout_ok, + "the kernel must satisfy the timeout assertion (M21.6 fixed this)" + ); +} diff --git a/crates/windows-ioring-sys/design-sessions/spikes/README.md b/crates/windows-ioring-sys/design-sessions/spikes/README.md index 1d4bc441d..c8185e3b8 100644 --- a/crates/windows-ioring-sys/design-sessions/spikes/README.md +++ b/crates/windows-ioring-sys/design-sessions/spikes/README.md @@ -18,22 +18,65 @@ windows-sys = { version = "0.61.2", default-features = false, features = [ |---|---| | [completion-event-spike.rs](completion-event-spike.rs) | [D-19](../../DESIGN-NOTES.md#d-19) -- the completion event is edge-triggered on the completion queue going empty to non-empty; also what `SetIoRingCompletionEvent` permits (call at any time, replace, clear with `NULL`, duplicate survives closing the original) | | [drain-ordering-spike.rs](drain-ordering-spike.rs) | [D-23](../../DESIGN-NOTES.md#d-23) -- an unflagged flush does not cover preceding writes; [D-24](../../DESIGN-NOTES.md#d-24) -- `DRAIN_PRECEDING_OPS` is a full, ring-wide barrier spanning submissions | +| [write-pending-spike.rs](write-pending-spike.rs) | `M20.6` -- which handle flags make a ring write or flush *pend* rather than complete inside `SubmitIoRing`, measured as a rate over 500 trials per condition. `FILE_FLAG_OVERLAPPED` alone changed nothing. A fifth condition over a `set_len` extent was added in `M25.3`, and sixteen runs are in [measurements/2026-09-24-set-len-vs-zero-fill/](../../measurements/2026-09-24-set-len-vs-zero-fill/README.md) -- read them before quoting any single run, because they show the buffered/unbuffered split is the part that replicates and that "only the pre-written extent pends" does not. See below for why this one reports frequencies and what may **not** be built on them. | +| [set-len-zero-fill-spike.rs](set-len-zero-fill-spike.rs) | What `set_len` costs and when. Written because two statements about it were being made from documentation rather than measurement -- see [measurements/2026-09-24-set-len-zero-fill-cost/](../../measurements/2026-09-24-set-len-zero-fill-cost/README.md). `set_len` is free, and so is a write that lands at the valid data length; a write that lands *past* it pays to zero the whole gap synchronously, about eight times the cost of writing the extent outright. A sequential writer pays nothing extra for pre-setting its length. No dependencies. | -## One spike here establishes nothing yet +## The pending spike reports rates, and none of them is a contract + +[write-pending-spike.rs](write-pending-spike.rs) exists because `epoch_log`'s strategy harness was +built on the premise that a log "keeps appending while a commit is outstanding", and measurement +showed nothing was ever outstanding: the commit's `SubmitIoRing` took 289-555 us and returned with +every completion already queued. + +Its first draft ran each condition **once** and printed a verdict. That is the error `D-47` records -- +the spike behind `D-24` saw a barrier hold a handful of times and wrote down a guarantee, when the +real violation rate was nearer one in a thousand. The rewrite to 500 trials per condition immediately +justified itself: the `NO_BUFFERING`-extending condition pends in about **1%** of trials, which a +handful of runs would have reported as "never". + +Measured here (single node, ARM64, one device), across two consecutive runs: + +| condition | pended / 500, run 1 | run 2 | submit p50 | +|---|---|---|---| +| buffered, no `OVERLAPPED` (what `epoch_log` opened) | 0 | 0 | ~510 us | +| buffered + `OVERLAPPED` | 0 | 0 | ~490 us | +| `NO_BUFFERING` + `OVERLAPPED`, extending | **5** | **271** | ~270 us | +| `NO_BUFFERING` + `OVERLAPPED`, pre-written extent | 500 | 500 | ~116 us | + +**The extending row moved from 1% to 54% between two runs minutes apart**, with no change to the +program. Whatever drives it -- filesystem allocation state, cache residency, something else -- it is +not under this program's control and was not measured. That row alone would defeat any number of +repetitions of a single condition: a run reporting 5 and a run reporting 271 are both "what the +platform does", and neither is what it will do next time. + +**None of this is a contract, including the two stable rows.** Windows specifies nothing about when a +ring operation completes relative to `SubmitIoRing`. A rate of zero bounds a frequency rather than +establishing that something cannot happen, and a rate of 500/500 is the same statement pointing the +other way: it may never have been false here, and it is still not contractually true. + +So the obvious use of this spike is the wrong one. Reading the rows, picking the flags that pended, +and rebuilding a harness on them would bind the sample's premise to incidental behaviour -- the +failure PLATFORM INTEGRITY rule 2 names. A log, and a benchmark of one, has to be correct whether an +operation completes inline or pends. What the spike is legitimately for is explaining why a +measurement looks the way it does, and knowing which configurations are worth testing *across*. + + +## One spike here has only a narrow result [file-handle-numa-spike.rs](file-handle-numa-spike.rs) is the exception to the table above: it is a -**ready instrument with no result**, checked in deliberately rather than held back. It asks whether a -file handle yields a NUMA node, and which question that answer answers. +**ready instrument whose result so far is vacuous on node count**, checked in deliberately rather than +held back. It asks whether a file handle yields a NUMA node, and which question that answer answers. -It is unrun because of a **hardware gap, not a decision to defer**: it needs more than one NUMA node +It is unsettled because of a **hardware gap, not a decision to defer**: it needs more than one NUMA node and storage whose PDO advertises a proximity domain, and the machine this workspace is developed on has a single node and reports zero `Win32_NumaNode` instances. On such a machine the spike is vacuous in the same sense the drain spike's control case guards against -- failure would prove nothing and success could only ever report `0`. It prints that warning itself before running. Anyone with a multi-node server and a real NVMe or SAN volume can settle it in a few minutes, and the -result would correct a claim -[DESIGN-NOTES.md](../../DESIGN-NOTES.md) currently makes about what is reachable from user mode. +result would settle what +[DESIGN-NOTES.md](../../DESIGN-NOTES.md) leaves open under "What is not reachable": whether either call +ever names a node that distinguishes one device from another. It **has** been smoke-run here, which is why it compiles and why its Q5 works: the first version opened the directory with `File::open`, which fails on a directory without diff --git a/crates/windows-ioring-sys/design-sessions/spikes/allocation-knee-spike.rs b/crates/windows-ioring-sys/design-sessions/spikes/allocation-knee-spike.rs new file mode 100644 index 000000000..bc1dfaedf --- /dev/null +++ b/crates/windows-ioring-sys/design-sessions/spikes/allocation-knee-spike.rs @@ -0,0 +1,163 @@ +//! Where does Rust's allocator stop using the heap, and does chunk size matter? +//! +//! Two questions, prompted by a review rule of thumb: avoid contiguous +//! allocations over 64 KB from the general heap, because most heaps move large +//! blocks to a separate space anyway and the threshold is a useful place to be +//! asked "does this really need to be contiguous?". +//! +//! On Windows 64 KB is not an arbitrary round number -- it is the virtual +//! memory allocation granularity, so a reservation cannot share its 64 KB +//! region with anything else. +//! +//! # Q1 -- where is the knee? +//! +//! For each allocation, ask `VirtualQuery` which reservation it belongs to and +//! count the distinct `AllocationBase` values across many allocations of one +//! size. Blocks carved out of a heap segment share their segment's base; a +//! block given its own `VirtualAlloc` region has a base of its own. +//! +//! Two earlier attempts failed and are recorded because both look reasonable: +//! +//! - **Testing for exact 64 KB alignment** reported "none" at every size up to +//! 4 MB. `HeapAlloc`'s large-block path does call `VirtualAlloc`, but writes +//! a header at the start of the region and returns a pointer past it, so a +//! large block is aligned *plus a constant* and never exactly aligned. +//! - **Counting distinct offsets within a 64 KB region** declined with size +//! (64 at 4 KB, 15 at 1 MB) but never reached 1, so it could not say where +//! the change happened. Suggestive is not decisive. +//! +//! `VirtualQuery` answers directly and needs no inference. +//! +//! # Q2 -- does it cost anything to stay under it? +//! +//! Time a fixed zero-fill using chunks of each size. If a small chunk fills as +//! fast as a large one, then respecting the rule is free and the chunk should +//! be small. If throughput climbs with chunk size, the rule has a price and the +//! price is what decides. + +use std::fs::File; +use windows_sys::Win32::System::Memory::{MEMORY_BASIC_INFORMATION, VirtualQuery}; +use std::io::Write; +use std::time::Instant; + +const KB: usize = 1024; +const SIZES: &[usize] = &[ + 4 * KB, + 16 * KB, + 32 * KB, + 64 * KB, + 128 * KB, + 256 * KB, + 512 * KB, + 1024 * KB, + 4096 * KB, +]; + +/// Allocations per size for Q1. Enough that an accidental alignment cannot +/// look like a threshold. +const SAMPLES: usize = 64; + +/// Bytes filled per chunk size in Q2. Large enough to be dominated by the +/// filling rather than by opening the file. +const FILL_TOTAL: usize = 64 * 1024 * 1024; + +const FILL_REPEATS: usize = 7; + +fn main() { + println!("Q1: do allocations of this size get their own reservation?\n"); + println!( + "{:>10} {:>20} {:>16} {:>22}", + "size", "distinct bases/64", "median region", "verdict" + ); + + for &size in SIZES { + // Held until all are made, so the allocator cannot hand back the same + // address every time and flatter the result. + let mut live: Vec> = Vec::with_capacity(SAMPLES); + for _ in 0..SAMPLES { + let mut v: Vec = Vec::with_capacity(size); + // Touch a byte so the allocation is real rather than deferred. + v.push(0); + live.push(v); + } + + let mut bases: Vec = Vec::with_capacity(SAMPLES); + let mut regions: Vec = Vec::with_capacity(SAMPLES); + for v in &live { + let mut info = MEMORY_BASIC_INFORMATION::default(); + // SAFETY: `v` is a live allocation this thread owns, and `info` is + // a live local of exactly the size passed. + let got = unsafe { + VirtualQuery( + v.as_ptr().cast(), + &raw mut info, + size_of::(), + ) + }; + assert_ne!(got, 0, "VirtualQuery failed"); + bases.push(info.AllocationBase as usize); + regions.push(info.RegionSize); + } + bases.sort_unstable(); + bases.dedup(); + regions.sort_unstable(); + let median_region = regions[regions.len() / 2]; + + let verdict = if bases.len() == SAMPLES { + "own reservation each" + } else if bases.len() * 4 >= SAMPLES { + "mostly own" + } else { + "shares a heap segment" + }; + println!( + "{:>8} KB {:>20} {:>13} KB {:>22}", + size / KB, + bases.len(), + median_region / KB, + verdict + ); + drop(live); + } + + println!("\nQ2: does a smaller chunk fill more slowly?\n"); + println!( + "{:>10} {:>14} {:>16}", + "chunk", "median ms", "MiB/s (median)" + ); + + let path = std::env::temp_dir().join(format!("chunk-probe-{}.tmp", std::process::id())); + for &chunk_len in SIZES { + let mut samples = Vec::with_capacity(FILL_REPEATS); + for _ in 0..FILL_REPEATS { + let chunk = vec![0_u8; chunk_len]; + let t = Instant::now(); + let mut file = File::create(&path).expect("create"); + let mut written = 0; + while written < FILL_TOTAL { + let take = chunk.len().min(FILL_TOTAL - written); + file.write_all(&chunk[..take]).expect("write"); + written += take; + } + file.flush().expect("flush"); + drop(file); + samples.push(t.elapsed().as_micros()); + } + samples.sort_unstable(); + let median = samples[samples.len() / 2]; + let mib_s = (FILL_TOTAL as f64 / (1024.0 * 1024.0)) / (median as f64 / 1_000_000.0); + println!( + "{:>8} KB {:>14.1} {:>16.0}", + chunk_len / KB, + median as f64 / 1000.0, + mib_s + ); + } + let _ = std::fs::remove_file(&path); + + println!( + "\nRead Q2 across the rows, not against a target. What decides the chunk size is\n\ + whether throughput is still climbing at the point the allocation stops being an\n\ + ordinary heap block." + ); +} diff --git a/crates/windows-ioring-sys/design-sessions/spikes/file-handle-numa-spike.rs b/crates/windows-ioring-sys/design-sessions/spikes/file-handle-numa-spike.rs index 44dd714aa..7baac3bc8 100644 --- a/crates/windows-ioring-sys/design-sessions/spikes/file-handle-numa-spike.rs +++ b/crates/windows-ioring-sys/design-sessions/spikes/file-handle-numa-spike.rs @@ -48,15 +48,17 @@ //! `CreateRemoteThreadEx` and an attribute list rather than a file handle. //! It now has its own instrument: `thread-stack-numa-spike.rs`. //! -//! Why it matters: `DESIGN-NOTES.md` asserts that mapping a file handle to the -//! NUMA node of its backing device "has no clean user-mode path" and "means -//! walking volume to disk to device instance". That is wrong on mechanism -- -//! `FSCTL_QUERY_VOLUME_NUMA_INFO` is documented, takes a file or directory -//! handle directly, and returns `FSCTL_QUERY_VOLUME_NUMA_INFO_OUTPUT { ULONG -//! NumaNode }`. What is *right* is the conclusion, for a different reason: the -//! documented meaning is the node the **volume** resides on, not where the -//! file's extents live, and it is absent whenever the device advertised no -//! proximity domain. +//! Why it matters: `DESIGN-NOTES.md` used to assert that mapping a file handle +//! to the NUMA node of its backing device "has no clean user-mode path" and +//! "means walking volume to disk to device instance". That was wrong on +//! mechanism -- `FSCTL_QUERY_VOLUME_NUMA_INFO` is documented, takes a file or +//! directory handle directly, and returns +//! `FSCTL_QUERY_VOLUME_NUMA_INFO_OUTPUT { ULONG NumaNode }`. What is *right* is +//! the conclusion, for a different reason: the documented meaning is the node +//! the **volume** resides on, not where the file's extents live, and it is +//! absent whenever the device advertised no proximity domain. That correction +//! landed in "What is not reachable" on 2026-09-19 (M20.4); this header is kept +//! as the statement of what this instrument was built to settle. //! //! `GetNumaNodeNumberFromHandle` is the other path: a Win32 wrapper over //! `NtQueryInformationFile` with `FileNumaNodeInformation` (class 53, Windows 7 diff --git a/crates/windows-ioring-sys/design-sessions/spikes/set-len-zero-fill-spike.rs b/crates/windows-ioring-sys/design-sessions/spikes/set-len-zero-fill-spike.rs new file mode 100644 index 000000000..8db4c715e --- /dev/null +++ b/crates/windows-ioring-sys/design-sessions/spikes/set-len-zero-fill-spike.rs @@ -0,0 +1,164 @@ +//! Does `set_len` avoid the zero-fill, or force it? -- measured, not reasoned. +//! +//! # Why this exists +//! +//! `M25.3` opens the epoch log over a **zero-filled** extent -- a real write of +//! zeros -- and its first draft justified that by asserting `set_len` would +//! leave writes in the extending case. That was reasoned from documentation. +//! Review supplied a specific counter-recollection, from optimizing `.cab` +//! expansion some years earlier: that setting the length **too early forced** +//! the zero fill of large files rather than deferring it. +//! +//! Both statements are about the same mechanism and they cannot both be the +//! whole story, so this measures it. The recollection turned out to be right +//! and to have a sharper shape than either statement had: the forcing is real, +//! it is conditional on *where* the write lands, and when it fires it costs an +//! order of magnitude more than doing the zeroing up front. +//! +//! # What it runs +//! +//! Five cases per size, each on a fresh file, timed separately: +//! +//! 1. `set_len(n)` alone. +//! 2. A zero-fill: writing `n` bytes of zeros. +//! 3. `set_len(n)` then one sector at offset 0 -- no gap in front of it. +//! 4. `set_len(n)` then one sector at the END, offset `n - SECTOR` -- a whole +//! file's worth of gap between the valid data length and the write. +//! 5. `set_len(n)` then the whole extent written sequentially. This is a log's +//! pattern: every write lands exactly at the valid data length. +//! +//! # How to read it +//! +//! **Case 4 against case 3** is the discriminator for the forcing. Both start +//! from an identical `set_len`'d file and write one sector; only the gap in +//! front of the write differs. +//! +//! **Case 5 against case 2** is the discriminator for whether a *sequential* +//! writer pays anything for having set its length first. +//! +//! # What it cannot tell you +//! +//! Nothing here is a platform guarantee. These are wall-clock costs on one +//! machine, one filesystem, one device, and NTFS is free to change how it +//! tracks valid data length. What the figures support is a shape -- free, +//! free, free, catastrophic, free -- not a constant to design against. +//! +//! # Running it +//! +//! Standalone, no dependencies. Copy into a scratch crate's `src/main.rs` and +//! `cargo run --release`, the way `tools/run-numa-spikes.ps1` builds the NUMA +//! spikes. It writes up to 1 GiB files into `%TEMP%` and removes them. + +use std::fs::{File, OpenOptions}; +use std::io::{Seek, SeekFrom, Write}; +use std::time::Instant; + +const SECTOR: usize = 4096; + +fn path(tag: &str) -> std::path::PathBuf { + std::env::temp_dir().join(format!("setlen-probe-{}-{tag}.tmp", std::process::id())) +} + +fn fresh(tag: &str) -> (std::path::PathBuf, File) { + let p = path(tag); + let _ = std::fs::remove_file(&p); + let f = OpenOptions::new() + .create(true) + .write(true) + .truncate(true) + .open(&p) + .expect("create"); + (p, f) +} + +fn ms(t: Instant) -> u128 { + t.elapsed().as_millis() +} + +fn main() { + println!("set_len vs zero-fill: is the zeroing avoided, or moved?\n"); + println!( + "{:>8} {:>12} {:>12} {:>16} {:>16} {:>18}", + "size", "set_len", "zero-fill", "set_len+head", "set_len+tail", "set_len+sequential" + ); + + for mib in [64_usize, 256, 1024] { + let n = mib * 1024 * 1024; + + // 1. set_len alone. + let (p1, f1) = fresh("setlen"); + let t = Instant::now(); + f1.set_len(n as u64).expect("set_len"); + drop(f1); + let a = ms(t); + + // 2. A real zero-fill of the same extent. + let (p2, mut f2) = fresh("zerofill"); + let zeros = vec![0_u8; 8 * 1024 * 1024]; + let t = Instant::now(); + let mut written = 0; + while written < n { + let take = zeros.len().min(n - written); + f2.write_all(&zeros[..take]).expect("write"); + written += take; + } + f2.flush().expect("flush"); + drop(f2); + let b = ms(t); + + // 3. set_len, then one sector at the FRONT. No gap between the valid + // data length and the write. + let (p3, mut f3) = fresh("head"); + f3.set_len(n as u64).expect("set_len"); + let t = Instant::now(); + f3.seek(SeekFrom::Start(0)).expect("seek"); + f3.write_all(&vec![0xAB_u8; SECTOR]).expect("write"); + f3.flush().expect("flush"); + drop(f3); + let c = ms(t); + + // 4. set_len, then one sector at the END. A whole file of gap. + let (p4, mut f4) = fresh("tail"); + f4.set_len(n as u64).expect("set_len"); + let t = Instant::now(); + f4.seek(SeekFrom::Start((n - SECTOR) as u64)).expect("seek"); + f4.write_all(&vec![0xCD_u8; SECTOR]).expect("write"); + f4.flush().expect("flush"); + drop(f4); + let d = ms(t); + + // 5. set_len, then fill the whole extent sequentially. This is a log's + // access pattern: every write lands exactly at the valid data + // length, so no write ever has a gap in front of it. + let (p5, mut f5) = fresh("sequential"); + f5.set_len(n as u64).expect("set_len"); + let t = Instant::now(); + let mut written = 0; + while written < n { + let take = zeros.len().min(n - written); + f5.write_all(&zeros[..take]).expect("write"); + written += take; + } + f5.flush().expect("flush"); + drop(f5); + let e = ms(t); + + println!( + "{:>6} MiB {:>10} ms {:>10} ms {:>14} ms {:>14} ms {:>16} ms", + mib, a, b, c, d, e + ); + + for p in [p1, p2, p3, p4, p5] { + let _ = std::fs::remove_file(p); + } + } + + println!( + "\nRead case 4 against case 3. Both begin from an identical set_len'd file and\n\ + write one sector. The only difference is how much unwritten extent lies between\n\ + the valid data length and the write.\n\n\ + Read case 5 against case 2. Both end with the whole extent written; case 5 had\n\ + its length set first. That is a log's pattern, where every write lands at the\n\ + valid data length and no write ever has a gap in front of it." + ); +} diff --git a/crates/windows-ioring-sys/design-sessions/spikes/write-pending-spike.rs b/crates/windows-ioring-sys/design-sessions/spikes/write-pending-spike.rs new file mode 100644 index 000000000..65c74cbad --- /dev/null +++ b/crates/windows-ioring-sys/design-sessions/spikes/write-pending-spike.rs @@ -0,0 +1,488 @@ +// Copyright (c) 2026 Mike Grier +//! Does a ring write ever *pend*, and at what rate? -- the spike behind `M20.6`. +//! +//! # The question +//! +//! `examples/epoch_log`'s strategy harness is built on the premise that a log +//! "keeps appending while a commit is outstanding". Measuring it showed the +//! premise does not hold on the handle that sample opens: across 32 epochs, +//! `SubmitIoRing` for the commit took 289-555 us and returned with every +//! completion already queued, so the device flush was paid *inside submit* and +//! nothing was outstanding across the submit boundary. +//! +//! Before rebuilding that harness around a real pipeline, this asks the +//! narrower question the rebuild depends on: **which handle flags, if any, +//! make a ring write or flush pend, and how often?** "Pend" is defined +//! observationally, because it is the only form a consumer can act on: +//! `SubmitIoRing` returns and at least one completion is not yet in the queue. +//! +//! # Why this reports rates and not verdicts +//! +//! An earlier draft ran each condition **once** and printed "pended" or +//! "inline". That is the exact error [DESIGN-NOTES.md](../../DESIGN-NOTES.md) +//! records as `D-47`: the spike behind `D-24` ran its sequence a handful of +//! times, saw a barrier hold every time, and wrote down a guarantee. The real +//! violation rate was nearer one in a thousand -- which is precisely what a +//! handful of runs shows, because nothing in a passing run distinguishes +//! "guaranteed" from "usually". +//! +//! Two runs of this program, minutes apart and unchanged, reported the +//! NO_BUFFERING-extending condition as **5/500 and then 271/500**. Whatever +//! drives that -- filesystem allocation state, cache residency, something not +//! measured here -- it is outside this program's control. So even a rate is a +//! description of one run, and repeating a single condition more times would +//! not have fixed it: 1% and 54% are both 'what the platform did'. +//! +//! That draft also asserted, as a *rule*, that buffered writes complete in +//! submission order. They were **observed** to, in the runs behind the drain +//! spike. An observation over a handful of runs on one machine is not a rule +//! about the platform, and this crate has already paid once for treating one +//! as the other. So every row below is a measured frequency over [`TRIALS`] +//! trials, stated as what was seen rather than what must happen. +//! +//! # What may be built on the result: nothing +//! +//! This bounds a **frequency**. It does not establish a contract, and a row +//! reading `0/500` or `500/500` is not a licence to depend on either outcome. +//! Windows documents nothing about when a ring operation completes relative to +//! `SubmitIoRing`; whether one pends is *incidental current behaviour* of this +//! build, this device, this filesystem. +//! +//! That matters because the obvious use of this program is the wrong one. It +//! would be natural to read the rows, pick the flags that pended, and rebuild +//! `epoch_log`'s harness on them -- which would bind the sample's central +//! premise to an unspecified property, the failure the repository's PLATFORM +//! INTEGRITY rule 2 names and that this workspace has already paid for once. +//! +//! **A log, and a benchmark of one, must be correct whether an operation +//! completes inline or pends.** The specified surface is: push, submit, pop, and +//! a completion means the operation is done. Nothing in it says *when* the +//! completion becomes available. So the legitimate uses of this spike are to +//! explain why a measurement looks the way it does, and to know which +//! configurations are worth testing *across* -- never to choose one and assume +//! its behaviour holds. +//! +//! Note also what this cannot decide even descriptively: `AlternatingRings`' +//! blast-radius question is settled **structurally** and needs no run of this +//! program. `RegisteredBuffers::get_mut` refuses a slot with an operation +//! outstanding, so at most `SLOTS` appends are outstanding on a ring by +//! construction -- and each alternating lane registers its own arena of the +//! same size. The per-ring bound is identical either way, whatever the +//! platform does about pending. +//! +//! # Conditions +//! +//! Each is one handle, and [`TRIALS`] batches of [`WRITES`] writes of +//! [`WRITE_LEN`] bytes plus one flush: +//! +//! - **A** buffered, no `OVERLAPPED` -- what `epoch_log` opens today. +//! - **B** buffered + `OVERLAPPED` -- the cheapest possible fix. +//! - **C** `NO_BUFFERING` + `OVERLAPPED`, extending the file. +//! - **D** `NO_BUFFERING` + `OVERLAPPED`, over an extent written beforehand. +//! +//! C and D are separated because the drain spike observed extending writes +//! behaving like buffered ones, which it attributed to the filesystem +//! serializing writes past the valid-data length. D was the only condition +//! that spike found could discriminate at all. +//! +//! Sector alignment is why the C/D distinction matters beyond curiosity: +//! `NO_BUFFERING` requires sector-aligned buffers, offsets and lengths, and +//! `epoch_log` wrote variable-length records at packed offsets when this was +//! written. If D is the only condition that pends, the harness fix is not a +//! flag change -- it is a change to the log's on-disk format. +//! +//! **That prediction held.** `M25.1` gave records a fixed sector stride with a +//! zeroed block tail, in both writers, before `M25.3` could change a single +//! flag. The paragraph is left in the past tense rather than deleted because +//! the reasoning is the useful part: a flag whose requirements the caller's +//! data layout cannot meet is not a flag change. +//! +//! # What "discriminating" means here +//! +//! Not that some condition pends. That the conditions **differ**. If every row +//! reports the same rate, this program has not learned which flags matter; it +//! has learned that its own apparatus cannot tell them apart, which is the +//! failure the drain spike's control case exists to catch. It says so, rather +//! than letting a reader infer a platform conclusion from four identical rows. + +use std::ffi::c_void; +use std::os::windows::ffi::OsStrExt; +use std::path::{Path, PathBuf}; +use std::time::Instant; + +use windows_sys::Win32::Foundation::{CloseHandle, GENERIC_WRITE, HANDLE, S_FALSE, S_OK}; +use windows_sys::Win32::Storage::FileSystem::{ + BuildIoRingFlushFile, BuildIoRingWriteFile, CREATE_ALWAYS, CloseIoRing, CreateFileW, + CreateIoRing, FILE_ATTRIBUTE_NORMAL, FILE_FLAG_NO_BUFFERING, FILE_FLAG_OVERLAPPED, + FILE_FLUSH_DEFAULT, FILE_SHARE_READ, FILE_WRITE_FLAGS_NONE, IORING_BUFFER_REF, + IORING_BUFFER_REF_0, IORING_CQE, IORING_CREATE_ADVISORY_FLAGS_NONE, IORING_CREATE_FLAGS, + IORING_CREATE_REQUIRED_FLAGS_NONE, IORING_HANDLE_REF, IORING_HANDLE_REF_0, IORING_REF_RAW, + IORING_VERSION_3, IOSQE_FLAGS_NONE, OPEN_EXISTING, PopIoRingCompletion, SubmitIoRing, +}; +use windows_sys::Win32::System::Memory::{MEM_COMMIT, MEM_RESERVE, PAGE_READWRITE, VirtualAlloc}; + +/// Sector size assumed for `NO_BUFFERING`. 4096 covers 512e and 4Kn alike, and +/// over-aligning is legal where under-aligning is not. +const SECTOR: usize = 4096; + +/// Bytes per write. One sector, so the same size is legal in every condition. +const WRITE_LEN: usize = SECTOR; + +/// Writes per batch. Matches `epoch_log`'s arena slot count, so the answer is +/// about the shape that sample actually submits. +const WRITES: usize = 8; + +/// Batches per condition. +/// +/// Sized against what `D-47` cost: that violation rate was around one in a +/// thousand, and a handful of trials reported it as never. This cannot resolve +/// one in a thousand either -- saying so is why the rate and the trial count +/// are printed together rather than a verdict. It is enough to separate +/// "always" from "usually" at the percent level, which is the resolution the +/// harness decision needs. +const TRIALS: usize = 500; + +fn wide(path: &Path) -> Vec { + path.as_os_str().encode_wide().chain(Some(0)).collect() +} + +fn temp(tag: &str) -> PathBuf { + std::env::temp_dir().join(format!("ioring-pend-spike-{tag}-{}.bin", std::process::id())) +} + +/// A sector-aligned buffer. `VirtualAlloc` is page-granular, which is at least +/// sector-granular on every configuration this runs on. +fn aligned(len: usize, fill: u8) -> *mut u8 { + // SAFETY: a null `lpAddress` lets the system choose; the result is checked. + let p = + unsafe { VirtualAlloc(std::ptr::null(), len, MEM_COMMIT | MEM_RESERVE, PAGE_READWRITE) }; + assert!(!p.is_null(), "VirtualAlloc failed"); + let p = p.cast::(); + // SAFETY: `len` bytes were just committed at `p`. + unsafe { std::ptr::write_bytes(p, fill, len) }; + p +} + +fn handle_ref(file: HANDLE) -> IORING_HANDLE_REF { + IORING_HANDLE_REF { + Kind: IORING_REF_RAW, + Handle: IORING_HANDLE_REF_0 { Handle: file }, + } +} + +fn buffer_ref(p: *mut u8) -> IORING_BUFFER_REF { + IORING_BUFFER_REF { + Kind: IORING_REF_RAW, + Buffer: IORING_BUFFER_REF_0 { + Address: p.cast::(), + }, + } +} + +/// What one condition reported over [`TRIALS`] batches. +struct Outcome { + label: &'static str, + /// Trials where at least one completion was missing after submit returned. + pended: usize, + /// Fewest completions seen queued immediately after submit. `WRITES + 1` + /// means no trial ever pended, even partially. + min_queued: usize, + submit_us: Vec, +} + +impl Outcome { + fn pct(&mut self, f: f64) -> u128 { + if self.submit_us.is_empty() { + return 0; + } + self.submit_us.sort_unstable(); + let r = (((self.submit_us.len() - 1) as f64) * f).round() as usize; + self.submit_us[r.min(self.submit_us.len() - 1)] + } +} + +/// How a condition's file is prepared before the measured writes. +#[derive(Clone, Copy, PartialEq, Eq)] +enum Extent { + /// Nothing: the measured writes extend the file. + None, + /// Zero-filled with a real write, so the valid data length covers every + /// offset the measured writes will touch. + Written, + /// `set_len` only. Same end-of-file as `Written`, same bytes on read, and + /// the whole question condition E exists to settle. + SetLen, + /// `set_len`, then a single write at the **end** of the extent, which + /// obliges the filesystem to zero-fill everything in front of it. + /// + /// The question condition F exists to settle, raised in review: since a + /// write past the valid data length forces the fill anyway, can that be + /// used deliberately -- one small write instead of a buffer the size of + /// the file -- to reach the same end state a zero-fill reaches? + SetLenTouchEnd, +} + +fn run(label: &'static str, flags: u32, extent: Extent) -> Outcome { + let path = temp(label); + let buffer = aligned(WRITE_LEN, 0xAB); + + // Give the writes an existing extent to land in, so they are not extending + // writes -- which the drain spike observed behaving like buffered ones. + match extent { + Extent::None => {} + Extent::Written => { + std::fs::write(&path, vec![0u8; WRITE_LEN * (WRITES + 1)]) + .expect("zero-fill the extent"); + } + Extent::SetLen => { + // The question condition E exists to answer: `set_len` sets the + // file's length without writing anything, so it costs nothing and + // produces a file that reads back identically to the written one. + // Whether it *behaves* identically on the write path is the thing + // that cannot be settled by reading either file back, and was + // being asserted from documentation until this row existed. + let file = std::fs::File::create(&path).expect("create for set_len"); + file.set_len((WRITE_LEN * (WRITES + 1)) as u64) + .expect("set_len the extent"); + } + Extent::SetLenTouchEnd => { + use std::io::{Seek, SeekFrom, Write}; + let total = WRITE_LEN * (WRITES + 1); + let mut file = std::fs::File::create(&path).expect("create for set_len"); + file.set_len(total as u64).expect("set_len the extent"); + // One sector at the very end. Everything in front of it is now + // between the valid data length and this write, so the filesystem + // must zero it before this write can proceed -- which is the + // forcing behaviour, used on purpose rather than tripped over. + file.seek(SeekFrom::Start((total - WRITE_LEN) as u64)) + .expect("seek to the last sector"); + file.write_all(&vec![0_u8; WRITE_LEN]) + .expect("touch the end"); + file.flush().expect("flush"); + } + } + + // A condition with an existing extent must open it: CREATE_ALWAYS would + // truncate the extent that is the only thing distinguishing it from C. + let disposition = if matches!(extent, Extent::None) { + CREATE_ALWAYS + } else { + OPEN_EXISTING + }; + // SAFETY: `wide` is NUL-terminated and outlives the call. + let file = unsafe { + CreateFileW( + wide(&path).as_ptr(), + GENERIC_WRITE, + FILE_SHARE_READ, + std::ptr::null(), + disposition, + FILE_ATTRIBUTE_NORMAL | flags, + std::ptr::null_mut(), + ) + }; + assert!( + file as isize != -1, + "CreateFileW({label}) failed: {}", + std::io::Error::last_os_error() + ); + + let mut ring = std::ptr::null_mut(); + // SAFETY: out parameter is a live local for the call's duration. + let hr = unsafe { + CreateIoRing( + IORING_VERSION_3, + IORING_CREATE_FLAGS { + Required: IORING_CREATE_REQUIRED_FLAGS_NONE, + Advisory: IORING_CREATE_ADVISORY_FLAGS_NONE, + }, + 64, + 128, + &raw mut ring, + ) + }; + assert_eq!(hr, S_OK, "CreateIoRing failed: 0x{:08X}", hr as u32); + + let mut out = Outcome { + label, + pended: 0, + min_queued: WRITES + 1, + submit_us: Vec::with_capacity(TRIALS), + }; + + for _ in 0..TRIALS { + for i in 0..WRITES { + // SAFETY: `file` and `buffer` outlive the ring's use of them -- + // every completion is drained before this function returns. + let hr = unsafe { + BuildIoRingWriteFile( + ring, + handle_ref(file), + buffer_ref(buffer), + WRITE_LEN as u32, + (i * WRITE_LEN) as u64, + FILE_WRITE_FLAGS_NONE, + i, + IOSQE_FLAGS_NONE, + ) + }; + assert_eq!(hr, S_OK, "BuildIoRingWriteFile failed: 0x{:08X}", hr as u32); + } + // SAFETY: as above. + let hr = unsafe { + BuildIoRingFlushFile( + ring, + handle_ref(file), + FILE_FLUSH_DEFAULT, + WRITES, + IOSQE_FLAGS_NONE, + ) + }; + assert_eq!(hr, S_OK, "BuildIoRingFlushFile failed: 0x{:08X}", hr as u32); + + // Submit with no wait, so any time spent is time the kernel spent + // doing the work rather than time this program asked it to block. + let mut submitted = 0u32; + let started = Instant::now(); + // SAFETY: out parameter is a live local. + let hr = unsafe { SubmitIoRing(ring, 0, 0, &raw mut submitted) }; + out.submit_us.push(started.elapsed().as_micros()); + assert_eq!(hr, S_OK, "SubmitIoRing failed: 0x{:08X}", hr as u32); + + // The observable: how many completions are already there. Fewer than + // `WRITES + 1` means something had not finished when submit returned. + let mut queued = 0usize; + loop { + let mut cqe: IORING_CQE = unsafe { std::mem::zeroed() }; + // SAFETY: out parameter is a live local. + let hr = unsafe { PopIoRingCompletion(ring, &raw mut cqe) }; + if hr == S_FALSE { + break; + } + assert_eq!(hr, S_OK, "PopIoRingCompletion failed: 0x{:08X}", hr as u32); + queued += 1; + } + if queued < WRITES + 1 { + out.pended += 1; + out.min_queued = out.min_queued.min(queued); + } + + // Drain whatever pended, so the next trial starts from an empty queue + // and its count means what it says. + let mut total = queued; + let deadline = Instant::now(); + while total < WRITES + 1 { + let mut cqe: IORING_CQE = unsafe { std::mem::zeroed() }; + // SAFETY: out parameter is a live local. + let hr = unsafe { PopIoRingCompletion(ring, &raw mut cqe) }; + if hr == S_OK { + total += 1; + } else if hr != S_FALSE { + panic!("PopIoRingCompletion failed: 0x{:08X}", hr as u32); + } + assert!( + deadline.elapsed().as_secs() < 10, + "{label}: {total} of {} completions after 10s", + WRITES + 1 + ); + } + } + + // SAFETY: the ring is drained, so nothing references `file` or the buffer. + unsafe { + CloseIoRing(ring); + CloseHandle(file); + } + let _ = std::fs::remove_file(&path); + out +} + +fn main() { + println!("Does a ring write ever pend, and at what rate? -- M20.6\n"); + println!( + "{TRIALS} trials per condition; each trial is {WRITES} writes of {WRITE_LEN} bytes plus \ + one flush,\nsubmitted as one batch. A trial 'pended' if fewer than {} completions were \ + queued when\nSubmitIoRing returned.\n", + WRITES + 1 + ); + + let mut conditions = vec![ + run("A-buffered-sync", 0, Extent::None), + run("B-buffered-overlapped", FILE_FLAG_OVERLAPPED, Extent::None), + run( + "C-nobuffer-extending", + FILE_FLAG_OVERLAPPED | FILE_FLAG_NO_BUFFERING, + Extent::None, + ), + run( + "D-nobuffer-prewritten", + FILE_FLAG_OVERLAPPED | FILE_FLAG_NO_BUFFERING, + Extent::Written, + ), + run( + "E-nobuffer-set_len", + FILE_FLAG_OVERLAPPED | FILE_FLAG_NO_BUFFERING, + Extent::SetLen, + ), + run( + "F-nobuffer-touch-end", + FILE_FLAG_OVERLAPPED | FILE_FLAG_NO_BUFFERING, + Extent::SetLenTouchEnd, + ), + ]; + + println!( + "{:<24} {:>14} {:>12} {:>14} {:>14}", + "condition", "pended/trials", "min queued", "submit p50 us", "submit p99 us" + ); + for o in conditions.iter_mut() { + let (p50, p99) = (o.pct(0.50), o.pct(0.99)); + println!( + "{:<24} {:>14} {:>12} {:>14} {:>14}", + o.label, + format!("{}/{}", o.pended, TRIALS), + o.min_queued, + p50, + p99 + ); + } + println!(); + + // Discrimination check. Four identical rows would mean this program cannot + // tell the conditions apart -- a fact about the apparatus, not about the + // platform, and it must not be read as "the flags do nothing". + let rates: Vec = conditions.iter().map(|o| o.pended).collect(); + if rates.iter().all(|r| *r == rates[0]) { + println!( + "NOT DISCRIMINATING: every condition reported {}/{TRIALS}. That is a result about \ + this apparatus, not about the platform -- it has not shown which flags matter, only \ + that it cannot tell them apart. Do not read it as 'the flags do nothing'.", + rates[0] + ); + } else { + println!("Conditions differ, so the apparatus discriminates. Reading the rows:"); + for o in &conditions { + let verdict = match o.pended { + 0 => "never pended in this run".to_string(), + n if n == TRIALS => "pended in every trial".to_string(), + n => format!( + "pended in {n} of {TRIALS} trials ({:.1}%)", + 100.0 * n as f64 / TRIALS as f64 + ), + }; + println!(" {:<24} {verdict}", o.label); + } + } + + println!( + "\nEvery line above is a frequency observed on this machine, this build, this device.\n\ + None of it is a platform guarantee. A rate of zero over {TRIALS} trials bounds the\n\ + frequency; it does not establish that the thing cannot happen -- which is exactly the\n\ + distinction D-47 was written to record. A rate of {TRIALS}/{TRIALS} is the same\n\ + statement pointing the other way: it may never have been false here, and it is still\n\ + not contractually true.\n\n\ + So do not pick the flags that pended and build on them. Windows specifies nothing about\n\ + when a ring operation completes relative to SubmitIoRing, and a log -- or a benchmark of\n\ + one -- has to be correct either way." + ); +} diff --git a/crates/windows-ioring-sys/examples/cache_domains.rs b/crates/windows-ioring-sys/examples/cache_domains.rs new file mode 100644 index 000000000..b84e2cc6d --- /dev/null +++ b/crates/windows-ioring-sys/examples/cache_domains.rs @@ -0,0 +1,124 @@ +// Copyright (c) 2026 Mike Grier +// Renamed from examples/l3_domains.rs at 8b8afaf6. The rewrite left the files +// 17% similar, far under git's rename threshold, so blame and log do not +// follow it without help. +//! M6.3: enumerate the **outermost cache level that actually partitions this +//! machine** -- the default heuristic `DESIGN-NOTES.md`'s "Why the NUMA node +//! is the wrong key" recommends for sizing an `IoRing` execution domain -- +//! and then print **every** level the machine reports, so a consumer can see +//! what the heuristic chose *and* what it chose between. +//! +//! The second half is the point as much as the first. The heuristic answers +//! the question a consumer with no opinion has; a consumer who knows their +//! part, or whose working set is sized to an inner level, needs the whole +//! list before they can disagree with it. Printing only the chosen level +//! would hand over a verdict while withholding the data behind it. +//! +//! This is enumeration only, not a partitioning policy: what to do with the +//! domains is a workload call this crate deliberately leaves to the caller +//! (see `windows-topology-sys` for a safe `GetLogicalProcessorInformationEx` +//! wrapper, and `examples/ring_copy` for a sample that does make the call). +//! +//! # Why this is no longer `l3_domains.rs` +//! +//! It filtered `cache.level == 3` and called the result "last-level cache", +//! which is two assumptions wearing one name: that the machine *has* an L3, +//! and that level numbering orders caches from inner to outer. A shipping +//! Snapdragon X2 Elite falsifies the first -- no L3 at all, with the natural +//! cluster boundary at L2 ([D-48](../DESIGN-NOTES.md#d-48)) -- so the old +//! version reported "0 last-level cache domain(s)" on a 12-core machine that +//! plainly has cache domains. +//! +//! `MachineMemoryTopology::outermost_partitioning_cache` owns the rule now, +//! including the part this file could never have got right by filtering: a +//! level qualifies only when its blocks are **pairwise disjoint**, and +//! "outermost" is decided by inclusion rather than by the level number. + +use std::io::Write; + +use windows_topology_sys::MachineMemoryTopology; + +fn main() -> std::io::Result<()> { + // The single sink every line of this sample's output goes through + // (repository "Architectural pre-steps" rule: never call `println!` + // from more than one call site). + let mut out = std::io::stdout(); + let topology = MachineMemoryTopology::discover()?; + let mut outermost: Option = None; + + match topology.outermost_partitioning_cache() { + Some((level, domains)) => { + let _ = writeln!( + out, + "{} cache domain(s) at L{level}, the outermost level that partitions this machine:", + domains.len() + ); + for (index, domain) in domains.iter().enumerate() { + let _ = writeln!(out, " domain {index}: {:?}", domain.processors); + } + outermost = Some(level); + } + // Not an error, and not an empty list dressed up as one. It is a + // statement about the machine: no cache level splits it into more + // than one disjoint block, so caches offer no partition to size a + // ring by. One ring is the correct answer here, which is the same + // degradation `ring_copy`'s `Policy::select` reports as `degraded`. + None => { + let _ = writeln!( + out, + "no cache level partitions this machine, so caches offer no domain \ + boundary to size a ring by; one ring is correct here" + ); + } + } + + // Every level, not just the one the default heuristic picked. The + // heuristic answers "which level should a consumer who has no opinion + // shard on"; a consumer who *does* have an opinion -- who knows their + // part, or whose working set is sized to an inner level -- needs to see + // what the other levels would give them before they can hold it. Printing + // only the chosen level would hand over a verdict while withholding the + // data it was drawn from. + let levels = topology.cache_levels(); + if levels.is_empty() { + // Distinguished from "partitions nothing" above: that machine has + // caches which happen not to divide it, this one reports none at all. + // A Snapdragon X2 Elite reports no L3 (D-48); a VM can report no cache + // relationships whatsoever. + let _ = writeln!(out, "\nthis machine reports no cache domains at all"); + } else { + let _ = writeln!(out, "\nevery cache level this machine reports:"); + for level in levels { + let partitions = topology.cache_partitions_at_level(level); + // `cache_partitions_at_level` deduplicates by processor set but + // does *not* prove disjointness -- two overlapping-but-unequal + // sets both survive it. Only the outermost level above has been + // checked for that, so these counts are labelled "distinct set(s)" + // rather than "partitions", which would claim a property nothing + // here established. + let marker = if Some(level) == outermost { + " <- the default heuristic's choice, checked pairwise disjoint" + } else { + "" + }; + let _ = writeln!( + out, + " L{level}: {} distinct processor set(s){marker}", + partitions.len() + ); + for domain in &partitions { + let _ = writeln!(out, " {:?}", domain.processors); + } + } + } + + // Processor groups are a hard floor (D-8 in DESIGN-NOTES.md): above 64 + // logical processors, a thread's affinity and a ring's waiter are each + // confined to one GROUP_AFFINITY, whether or not that partition is + // wanted. + let groups: std::collections::BTreeSet = + topology.processors.iter().map(|p| p.id.group).collect(); + let _ = writeln!(out, "{} processor group(s)", groups.len()); + + Ok(()) +} diff --git a/crates/windows-ioring-sys/examples/epoch_log/append.rs b/crates/windows-ioring-sys/examples/epoch_log/append.rs index 6c8e1fd5f..1c02a1fad 100644 --- a/crates/windows-ioring-sys/examples/epoch_log/append.rs +++ b/crates/windows-ioring-sys/examples/epoch_log/append.rs @@ -9,6 +9,12 @@ //! deliberately -- and it is why this sample uses the registered form rather //! than handing an owned `Vec` to every push. //! +//! **Placed** is meant literally as of M22.3: the arena is `NumaBuffer`, not +//! `Vec`, on the node [`crate::placement`] decides. That module is also +//! where the limits of the decision are written down -- notably that this +//! workload is far too flush-bound for the placement to pay, so the sample +//! demonstrates how the choice is made rather than that it was worth making. +//! //! Appending therefore has two halves that must not be confused: //! //! 1. **Compose** the record into a slot the kernel is not currently reading. @@ -23,17 +29,17 @@ //! *accepted into the open epoch*, which is all [`crate::contract`] promises; //! durability arrives with the epoch's commit in M13.3. -use std::collections::HashMap; use std::io; use std::os::windows::io::RawHandle; use windows_ioring_sys::contract::RingContract; use windows_ioring_sys::{ - Batch, IoRing, PushOptions, RegisteredBuffers, RegisteredSpan, RegisteredUse, Token, - WriteCaching, + Batch, IoBufMut, IoRing, NumaBuffer, Pending, PushOptions, RegisteredBuffers, RegisteredSpan, + RegisteredUse, WriteCaching, }; use crate::commit::Epoch; +use crate::placement::Placement; use crate::record::{self, Sequence}; /// How many slots the arena holds. More slots means more records can be in @@ -43,30 +49,75 @@ pub const SLOTS: u32 = 8; /// Bytes per slot, and so the largest record this log accepts. pub const SLOT_LEN: usize = 4096; -/// One in-flight append: the token that holds the arena slot, and where the -/// record was written. -struct InFlight { - token: Token, - slot: u32, +// The write covers a whole stride out of one slot, so a stride wider than a +// slot would read past the arena. The stride itself is the format's, and is +// defined in `record` beside the layout it describes. +const _: () = assert!( + record::RECORD_STRIDE <= SLOT_LEN, + "a slot must hold a whole stride: the write spans one stride of one buffer" +); + +/// Hang bound on the one blocking step this appender has: waiting for the +/// arena's registration to complete at startup. +const REGISTRATION_TIMEOUT: std::time::Duration = std::time::Duration::from_secs(30); + +/// Slots of `arena` with no operation outstanding against them, at most `want` +/// of them, lowest index first. +/// +/// This is the sample's **only** definition of "which slots are free", and +/// both append paths bind to it: this module's [`Appender`] and the +/// measurement harness's `Lane` in [`crate::strategy`]. Each used to carry its +/// own, which is the defect that collapsed them -- not the cost, which is nil +/// at eight slots, but that two definitions of one fact can drift apart while +/// each looks locally correct. +/// +/// # Why this is derived and not tracked +/// +/// The obvious alternative is a `Vec` free list: pop a slot when an +/// append takes it, push it back when the completion is claimed. `Lane` did +/// exactly that. It is a *second copy* of a fact +/// [`RegisteredBuffers`] already owns and maintains -- its per-buffer +/// outstanding count, the same one that makes +/// [`RegisteredBuffers::get_mut`] refuse a busy slot. Asking is therefore +/// always right by construction, where a copy is right only as long as every +/// path that changes the truth remembers to change the copy too. +/// +/// The free list had already stopped remembering, in a way nothing reported. +/// It popped a slot before composing into it, so any error between the pop and +/// the push -- a record too long for a slot is the reachable one -- returned +/// early with the slot removed from the free list and no operation ever +/// issued. The arena considered that slot quiet forever; the free list never +/// offered it again. `SLOTS` such errors and the harness wedges, blaming an +/// arena that is in fact entirely idle. The derived form cannot express that +/// bug: a slot nothing was pushed against never stopped being free. +/// +/// That is measured rather than argued: re-injecting the free list and failing +/// eight appends left the lane reporting **0** of 8 slots free while the arena +/// held nothing. See [`crate::strategy::tests`], which also says plainly what +/// those tests can and cannot catch. +pub fn free_slots(arena: &RegisteredBuffers, want: usize) -> Vec { + (0..arena.len()) + .filter(|&slot| arena.outstanding(slot) == Some(0)) + .take(want) + .collect() } /// The append path: an arena of registered buffers, a monotonic sequence /// counter, and the file offset the next record lands at. pub struct Appender { - arena: RegisteredBuffers>, - in_flight: HashMap, + arena: RegisteredBuffers, + /// Unclaimed tokens, with the arena slot each holds. + /// + /// Checked, so the conservation oracle is driven by the same call that + /// updates the map (M16.2's accounting, now wired rather than hand-driven). + /// That matters for the specific failure it guards: an early return from + /// [`Appender::claim`] that skips the token claim leaks the arena slot + /// permanently, and nothing else in this program notices until the arena + /// runs dry `SLOTS` failures later -- somewhere else entirely, with no + /// trace of the cause. + pending: Pending, next_sequence: u64, next_offset: u64, - /// Conservation accounting for this appender's own operations (M16.2). - /// - /// Owned here rather than threaded in from `main` because the component - /// that issues the operations is the one that can report them without a - /// caller having to remember to. That matters for the specific failure - /// this guards: an early return from [`Appender::claim`] that skips the - /// token claim leaks the arena slot permanently, and nothing else in this - /// program notices until the arena runs dry `SLOTS` failures later -- - /// somewhere else entirely, with no trace of the cause. - contract: RingContract, } impl Appender { @@ -76,21 +127,37 @@ impl Appender { /// its completion -- one blocking step at startup, before the log has any /// work to pipeline against. /// + /// # Placement + /// + /// `placement` decides which NUMA node the arena's pages prefer. It is + /// taken as an argument rather than decided here because it is a *policy* + /// question about a caller's storage layout, and the library deliberately + /// answers none -- [`crate::placement`] is where this sample makes its own + /// choice, and says what that choice is and is not worth. + /// /// # Errors /// - /// Any error from the registration push, the submit, or the registration - /// operation itself. - pub fn new(ring: &mut IoRing) -> io::Result { - let buffers = (0..SLOTS).map(|_| vec![0_u8; SLOT_LEN]).collect::>(); + /// Any error from allocating the arena, the registration push, the submit, + /// or the registration operation itself. + pub fn new(ring: &mut IoRing, placement: &Placement) -> io::Result { + let node = placement.node(); + let buffers = (0..SLOTS) + .map(|_| NumaBuffer::new(SLOT_LEN, node)) + .collect::>>()?; let mut batch = Batch::new(ring); let pending = batch.register_buffers(buffers)?; - batch.submit_and_wait(1, 30_000)?; + batch.submit()?; - let completion = loop { - if let Some(completion) = ring.try_pop()? { - break completion; - } - }; + // One bounded wait, not a spin: `pop_within` is the crate's join + // between `try_pop`'s "empty right now" and a submit-side wait whose + // return promises nothing about poppability. The bare `loop` that + // used to be here turned a slow registration into a hung process. + let completion = ring.pop_within(REGISTRATION_TIMEOUT)?.ok_or_else(|| { + io::Error::new( + io::ErrorKind::TimedOut, + "the buffer registration never completed", + ) + })?; let arena = pending .claim_if(&completion) .map_err(|_| { @@ -100,17 +167,25 @@ impl Appender { Ok(Self { arena, - in_flight: HashMap::new(), + pending: Pending::checked(), next_sequence: 0, next_offset: 0, - contract: RingContract::new(), }) } /// This appender's conservation record, for a caller to assert against at /// teardown. + /// + /// Reads through to the map's own oracle rather than a separate one. An + /// earlier draft of this conversion kept the `RingContract` field beside + /// `Pending::checked()`, which compiled, ran, and made + /// `assert_quiescent()` pass **vacuously** -- the field was never written + /// to again, so a caller's teardown check was asserting against an oracle + /// that had observed nothing. pub fn contract(&self) -> &RingContract { - &self.contract + self.pending + .contract() + .expect("the appender's map is always checked") } /// The sequence the next appended record will carry. @@ -120,95 +195,115 @@ impl Appender { /// How many appends are pushed but not yet observed complete. pub fn in_flight(&self) -> usize { - self.in_flight.len() + self.pending.len() } - /// A slot with no operation outstanding against it, or `None` if every - /// slot is busy. + /// Compose as many of `payloads` as there are free arena slots, and push + /// them all in **one** submission. /// - /// Asked rather than assumed: the arena is the reason an append can block - /// at all, and a caller that gets `None` should drain a completion and try - /// again rather than grow the arena. - fn free_slot(&self) -> Option { - (0..self.arena.len()).find(|&slot| self.arena.outstanding(slot) == Some(0)) - } - - /// Compose `payload` into a free arena slot and push the write. + /// Returns how many were accepted. **Zero is not an error**: it means + /// every slot still has an append in flight, and the caller must drain a + /// completion before trying again. + /// + /// # Why this is a batch, and why that is the point of the sample + /// + /// This used to push one write and submit it, per record. That is a + /// working log and a misleading example: `Batch` exists so that many + /// submission-queue entries cost one `SubmitIoRing`, and a sample whose + /// job is to teach `Batch` should not pay that call per record. /// - /// Returns the record's [`Sequence`], which is its identity for the rest - /// of its life -- the epoch bookkeeping in M13.3 keys off it. + /// The batching is bounded by the arena rather than by the caller's list, + /// which is what keeps the two halves of an append honest: a slot is + /// composed into only while the kernel is not reading it, and it stays + /// spoken for until its completion is observed. So the natural batch is + /// "everything that fits right now", not "everything the caller has". /// /// # Errors /// - /// [`io::ErrorKind::WouldBlock`] if every arena slot is still in flight: - /// the caller must drain a completion before appending again. Otherwise - /// any error from encoding the record (notably if it does not fit a slot) - /// or from the push. - pub fn append( + /// Any error from encoding a record (notably if it does not fit a slot) or + /// from a push. A failure partway leaves the records already pushed + /// accounted for and in flight -- `Batch` submits what it queued when it + /// drops (D-5), so they are real operations, not a rollback. + pub fn append_batch( &mut self, ring: &mut IoRing, file: RawHandle, epoch: Epoch, - payload: &[u8], - ) -> io::Result { - let slot = self.free_slot().ok_or_else(|| { - io::Error::new( - io::ErrorKind::WouldBlock, - "every arena slot has an append in flight; drain a completion first", - ) - })?; - let sequence = Sequence(self.next_sequence); + payloads: &[Vec], + ) -> io::Result { + let slots = free_slots(&self.arena, payloads.len()); + if slots.is_empty() { + return Ok(0); + } - // Step 1: compose. `get_mut` is what makes this possible at all, and - // it is also the check that the kernel is not reading this slot. - // - // The epoch is stamped into the record here, at the moment the append - // is accepted -- which is exactly when the contract says a record's - // epoch is decided, and never changes afterwards. - let total = record::encode(self.arena.get_mut(slot)?, sequence, epoch, payload)?; - - // Step 2: push, over exactly the bytes the record occupies rather than - // the whole slot -- writing the slot's unused tail would put stale - // bytes in the log and cost real device bandwidth. - let span = RegisteredSpan { - buffer_index: slot, - offset: 0, - len: u32::try_from(total).map_err(|_| { - io::Error::new( - io::ErrorKind::InvalidInput, - "record length exceeds u32::MAX", - ) - })?, - }; - let offset = self.next_offset; let mut batch = Batch::new(ring); - // SAFETY: `file` is the log's own handle and outlives every operation - // pushed here -- the log drains to empty before it closes. The token - // is held in `in_flight` until its completion is observed, so the - // arena slot it names cannot be refilled underneath the kernel. - // - // `PushOptions::new()` deliberately carries no barrier: records stream - // unordered within an epoch, exactly as the contract says, and the - // ordering that matters is bought once by the epoch's covering flush. - // `WriteCaching::Cached` for the same reason -- write-through here - // would shape latency without changing what is durable. - let token = unsafe { - batch.write_registered_raw( - file, - &self.arena, - span, - offset, - PushOptions::new(), - WriteCaching::Cached, - ) - }?; - batch.submit()?; + let mut accepted = 0; + for (&slot, payload) in slots.iter().zip(payloads) { + let sequence = Sequence(self.next_sequence); - self.contract.observe_push(token.id()); - self.in_flight.insert(token.id(), InFlight { token, slot }); - self.next_sequence += 1; - self.next_offset += total as u64; - Ok(sequence) + // Compose. `get_mut` is what makes this possible at all, and it is + // also the check that the kernel is not reading this slot. + // + // The epoch is stamped in here, at the moment the append is + // accepted -- which is exactly when the contract says a record's + // epoch is decided, and never changes afterwards. + let buffer = self.arena.get_mut(slot)?; + record::encode_block(buffer, sequence, epoch, payload)?; + + // Push over the whole stride rather than the record's own length. + // + // This reverses the decision this line used to carry -- "writing + // the slot's unused tail would put stale bytes in the log and cost + // real device bandwidth". The stale-bytes half is handled by + // `encode_block`, which zeroes the remainder of the block. The + // bandwidth half was correct and is simply the price: see + // `record::RECORD_STRIDE` for what it costs and why a log pays it. + let span = RegisteredSpan { + buffer_index: slot, + offset: 0, + len: u32::try_from(record::RECORD_STRIDE).map_err(|_| { + io::Error::new( + io::ErrorKind::InvalidInput, + "record stride exceeds u32::MAX", + ) + })?, + }; + let offset = self.next_offset; + // SAFETY: `file` is the log's own handle and outlives every + // operation pushed here -- the log drains to empty before it + // closes. The token is held in `in_flight` until its completion is + // observed, so the arena slot it names cannot be refilled + // underneath the kernel. + // + // `PushOptions::new()` deliberately carries no barrier: records + // stream unordered within an epoch, exactly as the contract says, + // and the ordering that matters is bought once by the epoch's + // covering flush. `WriteCaching::Cached` for the same reason -- + // write-through here would shape latency without changing what is + // durable. + let token = unsafe { + batch.write_registered_raw( + file, + &self.arena, + span, + offset, + PushOptions::new(), + WriteCaching::Cached, + ) + }?; + + // One call updates the map and its oracle, where this previously + // updated them separately and could drift. + self.pending.push(token, slot); + self.next_sequence += 1; + self.next_offset += record::RECORD_STRIDE as u64; + accepted += 1; + } + + // One submission for the whole batch. This is the line the item + // existed for. + batch.submit()?; + Ok(accepted) } /// Account for one popped completion that belongs to an append. @@ -218,32 +313,58 @@ impl Appender { /// completions on the floor will run the arena dry and never recover -- /// which is the same drain-to-empty discipline the ring itself demands. pub fn claim(&mut self, completion: &windows_ioring_sys::Completion) -> io::Result { - let Some(in_flight) = self.in_flight.remove(&completion.user_data()) else { + // Claiming happens here, before the write's result is inspected, and + // that ordering still matters: bailing out on a failed write without + // claiming drops the token unclaimed, which `Token` deliberately + // treats as "still outstanding" and leaks -- burning this arena slot + // permanently, so `free_slots` never offers it again and after `SLOTS` + // failures every append returns `WouldBlock` forever. `M22.2` found + // exactly that bug here. + // + // `Pending` does not make the inverted order unrepresentable -- it + // compiles -- but it is no longer silent either way: the token stays in + // the map, and `append/tests.rs` drives a failed write through the + // injection seam so the inversion is caught by an assertion rather than + // waiting for a production arena to run dry. + let Some((released, slot)) = self.pending.claim(completion) else { return Ok(false); }; - let user_data = completion.user_data(); - self.contract.observe_completion(user_data); - // Claim *before* checking the write's result. The completion has - // already been observed, so claiming is sound either way -- and - // bailing out on a failed write without claiming would drop the token - // unclaimed, which `Token` deliberately treats as "still outstanding" - // and leaks. That would burn this arena slot permanently: `free_slot` - // would never offer it again, and after `SLOTS` failures every append - // would return `WouldBlock` forever. - let released = in_flight - .token - .claim_if(completion) - .map_err(|_| io::Error::other("an append token refused its own completion"))?; // Dropping the marker is what decrements the slot's count, so it has // to happen before the check below rather than at end of scope. drop(released); - self.contract.observe_claim(user_data); debug_assert!( - self.arena.outstanding(in_flight.slot) == Some(0), + self.arena.outstanding(slot) == Some(0), "claiming the token must release the slot" ); - let _written = completion.result()?; + // The contract requires that a successful write of N bytes transferred + // N bytes (see `crate::contract`, `Clause::Requires`), and this is + // where that requirement is checked rather than assumed. Every record + // write asks for exactly `RECORD_STRIDE`, so the comparison needs no + // per-slot bookkeeping. + // + // The ring itself permits a short count -- `RS-P-8` in the crate's + // RESPONSE-SPACE.md -- because it never asks what kind of handle it was + // given. This log narrows that by requiring a handle which does not do + // it, so a short count here is a violation of the contract's + // requirement rather than a case to absorb, and is reported as one. The + // count was previously bound to `_written` and discarded, which left + // the requirement stated nowhere and checked nowhere. + let written = completion.result()?; + if written != record::RECORD_STRIDE { + return Err(io::Error::new( + io::ErrorKind::InvalidData, + format!( + "a successful write transferred {written} of {} bytes; this log requires a \ + handle whose successful writes are complete (see the contract's Requires \ + clause), and this handle does not meet that requirement", + record::RECORD_STRIDE + ), + )); + } Ok(true) } } + +#[cfg(test)] +mod tests; diff --git a/crates/windows-ioring-sys/examples/epoch_log/append/tests.rs b/crates/windows-ioring-sys/examples/epoch_log/append/tests.rs new file mode 100644 index 000000000..98b98aa2f --- /dev/null +++ b/crates/windows-ioring-sys/examples/epoch_log/append/tests.rs @@ -0,0 +1,397 @@ +// Copyright (c) 2026 Mike Grier +//! Tests for the append path's claim discipline. +//! +//! # Why these exist, and what they are the counterpart to +//! +//! `M22.2` found a real defect here: [`super::Appender::claim`] returned early +//! on a failed write without claiming the token, which `Token` deliberately +//! treats as still outstanding. The arena slot's outstanding count then never +//! returned to zero, `free_slots` never offered it again, and after `SLOTS` +//! such failures every append returned `WouldBlock` forever -- somewhere else +//! entirely, with no trace of the cause. +//! +//! The library has a test named for exactly that hazard, +//! `claiming_before_checking_the_result_is_what_stops_a_failure_from_leaking` +//! in `tests/failure_paths.rs`. This consumer had no counterpart, and the gap +//! was measured rather than suspected: reinstating the defect here -- moving +//! `completion.result()?` above the claim -- compiled and passed every test, +//! because nothing produced a failed write. +//! +//! A failed write is not rare enough to be unreachable, only rare enough to go +//! untested. [`Completion::with_injected_failure`] is the seam that reaches it +//! on demand, which is why `epoch_log` is a test target at all. + +use std::os::windows::io::AsRawHandle; +use std::time::Duration; + +use windows_ioring_sys::IoRing; +#[cfg(feature = "fault-injection")] +use windows_ioring_sys::IoRingErrorExt; + +use super::Appender; +// Used only by the fault-injection tests below, so the import is gated the +// same way they are -- an unconditional one warns on a default-feature build +// of the test target. +#[cfg(feature = "fault-injection")] +use super::SLOTS; +use crate::commit::Epoch; +use crate::placement::Placement; +use crate::record; + +/// Hang bound on every wait here. Far above any real append. +const WAIT: Duration = Duration::from_secs(30); + +/// The failure injected into an append's completion. Any error would do; the +/// point is that `result()` reports one. +#[cfg(feature = "fault-injection")] +const INJECTED: windows_ioring_sys::InjectedFailure = + windows_ioring_sys::InjectedFailure::Win32(windows_sys::Win32::Foundation::ERROR_ACCESS_DENIED); + +/// A scratch file to append to, named per test so tests running as threads in +/// one process cannot collide on it. +fn scratch(tag: &str) -> (std::path::PathBuf, std::fs::File) { + let path = std::env::temp_dir().join(format!( + "windows-ioring-sys-epoch-append-{}-{tag}.tmp", + std::process::id() + )); + let file = std::fs::OpenOptions::new() + .create(true) + .truncate(true) + .read(true) + .write(true) + .open(&path) + .expect("open fixture"); + (path, file) +} + +/// **The contract's transfer requirement, enforced rather than merely stated +/// (M26.11).** +/// +/// `epoch_log`'s contract requires a handle whose successful writes are +/// complete. The ring itself permits the opposite -- `RS-P-8` -- because it +/// never asks what kind of handle it was given, so this narrowing is the log's +/// own and has to be checked by the log. +/// +/// **No real handle this sample opens will produce a short write**, which is +/// precisely why the check needs a seam to reach it: the branch would otherwise +/// be written once against the documentation and never executed again. That is +/// the same argument `M16.3` made for the failure seam, and +/// [`Completion::with_injected_transfer`] is its counterpart for a *successful* +/// short count, which a failure cannot model. +/// +/// Both directions are asserted here, because a test that only shows the +/// rejection would pass just as well against a check that rejected everything. +/// The two cases differ in the injected count and in nothing else. +#[cfg(feature = "fault-injection")] +#[test] +fn a_short_write_violates_the_contracts_transfer_requirement() { + let complete = record::RECORD_STRIDE; + + for (transferred, expect_accepted) in [(complete, true), (complete - 1, false)] { + let mut ring = IoRing::new(16, 16).expect("create ring"); + let (path, file) = scratch(&format!("short-write-{transferred}")); + let mut appender = + Appender::new(&mut ring, &Placement::decide(file.as_raw_handle())).expect("appender"); + + let pushed = appender + .append_batch( + &mut ring, + file.as_raw_handle(), + Epoch(0), + &[b"a record whose write will be reported short".to_vec()], + ) + .expect("push one append"); + assert_eq!(pushed, 1, "a fresh arena always has a slot"); + + let completion = ring + .pop_within(WAIT) + .expect("pop_within") + .expect("the append's completion arrives well inside the bound"); + let reported = completion.with_injected_transfer(transferred); + + match appender.claim(&reported) { + Ok(accepted) => { + assert!( + expect_accepted, + "a write reporting {transferred} of {complete} bytes was accepted; the \ + contract requires complete transfers and this one is short" + ); + assert!(accepted, "the completion is the append's own"); + } + Err(error) => { + assert!( + !expect_accepted, + "a write reporting the full {complete} bytes was rejected: {error}" + ); + assert_eq!( + error.kind(), + std::io::ErrorKind::InvalidData, + "a handle that breaks a stated requirement is bad input, not an I/O failure" + ); + let text = error.to_string(); + assert!( + text.contains(&transferred.to_string()) && text.contains(&complete.to_string()), + "the error must report both counts so the handle can be diagnosed, got: {text}" + ); + } + } + + drop(file); + let _ = std::fs::remove_file(&path); + } +} + +/// **The regression guard `M22.2` earned and this consumer never had.** +/// +/// A write that fails must still release its arena slot. The assertion is on +/// the slot rather than on the error, because the error was never the part that +/// broke: the old defect reported the failure correctly and leaked the slot +/// while doing it. +#[cfg(feature = "fault-injection")] +#[test] +fn a_failed_write_still_releases_its_arena_slot() { + let mut ring = IoRing::new(16, 16).expect("create ring"); + let (path, file) = scratch("failed-write"); + let mut appender = + Appender::new(&mut ring, &Placement::decide(file.as_raw_handle())).expect("appender"); + + let pushed = appender + .append_batch( + &mut ring, + file.as_raw_handle(), + Epoch(0), + &[b"a record".to_vec()], + ) + .expect("push one append"); + assert_eq!(pushed, 1, "the arena starts empty, so one record fits"); + assert_eq!(appender.in_flight(), 1); + + let completion = ring + .pop_within(WAIT) + .expect("pop_within") + .expect("the append's completion arrives well inside the bound"); + + let failed = completion.with_injected_failure(INJECTED); + let error = appender + .claim(&failed) + .expect_err("a failed write must be reported, not swallowed"); + assert_eq!( + (error.as_ioring_error().expect("an IoRingError").code() as u32) & 0xFFFF, + windows_sys::Win32::Foundation::ERROR_ACCESS_DENIED, + "the error reaching the caller must be the one that was injected" + ); + + // The point of the test. Claiming is what returns the slot; a claim skipped + // on the failure path would leave this at one and the leak would be + // invisible until the arena ran dry. That the *arena* recovers too is what + // `repeated_failures_never_exhaust_the_arena` below establishes, by driving + // past `SLOTS` failures -- which only completes if slots are genuinely + // being handed back rather than merely appearing to be. + assert_eq!( + appender.in_flight(), + 0, + "the token must have been claimed even though the write failed" + ); + + drop(file); + let _ = std::fs::remove_file(&path); +} + +/// The same discipline over enough failures to have exhausted the arena. +/// +/// One leaked slot is survivable and invisible; `SLOTS` of them are what turned +/// `M22.2`'s defect into "every append returns `WouldBlock` forever". Driving +/// past that count is what distinguishes a slot that is released from one that +/// merely looks released once. +#[cfg(feature = "fault-injection")] +#[test] +fn repeated_failures_never_exhaust_the_arena() { + let mut ring = IoRing::new(16, 16).expect("create ring"); + let (path, file) = scratch("repeated-failures"); + let mut appender = + Appender::new(&mut ring, &Placement::decide(file.as_raw_handle())).expect("appender"); + + for round in 0..(SLOTS as usize * 2) { + let pushed = appender + .append_batch( + &mut ring, + file.as_raw_handle(), + Epoch(0), + &[b"a record".to_vec()], + ) + .expect("push one append"); + assert_eq!( + pushed, 1, + "round {round}: a slot must be available, or an earlier failure leaked one" + ); + + let completion = ring + .pop_within(WAIT) + .expect("pop_within") + .expect("the append's completion arrives well inside the bound"); + appender + .claim(&completion.with_injected_failure(INJECTED)) + .expect_err("round {round}: the injected failure must be reported"); + } + + assert_eq!(appender.in_flight(), 0); + drop(file); + let _ = std::fs::remove_file(&path); +} + +/// A successful write releases its slot too, which is the control. +/// +/// Without it the tests above would pass against an implementation that +/// released slots unconditionally at some later point, rather than because the +/// claim happened. +#[test] +fn a_successful_write_releases_its_arena_slot() { + let mut ring = IoRing::new(16, 16).expect("create ring"); + let (path, file) = scratch("successful-write"); + let mut appender = + Appender::new(&mut ring, &Placement::decide(file.as_raw_handle())).expect("appender"); + + appender + .append_batch( + &mut ring, + file.as_raw_handle(), + Epoch(0), + &[b"a record".to_vec()], + ) + .expect("push one append"); + + let completion = ring + .pop_within(WAIT) + .expect("pop_within") + .expect("the append's completion arrives well inside the bound"); + assert!(appender.claim(&completion).expect("a successful claim")); + assert_eq!(appender.in_flight(), 0); + + drop(file); + let _ = std::fs::remove_file(&path); +} + +/// A reused slot must not write the previous record's tail (M25.1). +/// +/// Records are variable-length but land one per `RECORD_STRIDE` block, so the +/// write covers bytes the record itself never set. A slot is reused for the +/// log's whole life -- a `NumaBuffer` arrives zeroed, but only once -- so those +/// bytes are whatever the *previous*, longer record left in them. +/// +/// **Replay cannot catch this, so the assertion is on the file's bytes.** +/// Replay decodes only at block starts and takes a record's extent from its own +/// header, so a stale fragment living past a short record's end is never read. +/// That makes leaving it a hygiene defect rather than a decode failure -- the +/// log would carry fragments of unrelated records, in a format whose whole +/// purpose is reconstructing what happened after a crash. Stating it that way +/// rather than as a corruption is deliberate: the zeroing is worth doing, and +/// claiming it prevents a decode error it cannot prevent would be worse than +/// not documenting it. +#[test] +fn a_reused_slot_does_not_write_the_previous_records_tail() { + let mut ring = IoRing::new(16, 16).expect("create ring"); + let (path, file) = scratch("reused-slot-tail"); + let mut appender = + Appender::new(&mut ring, &Placement::decide(file.as_raw_handle())).expect("appender"); + + // A long record, then a short one. `free_slots` hands out the lowest free + // index, so draining between the two puts both records in slot 0 -- which + // is what makes the second write's tail the first record's bytes. + let long = vec![0xAB_u8; 200]; + let short = b"short".to_vec(); + + for payload in [long, short.clone()] { + let pushed = appender + .append_batch(&mut ring, file.as_raw_handle(), Epoch(0), &[payload]) + .expect("push one append"); + assert_eq!(pushed, 1, "a drained arena always has a slot"); + let completion = ring + .pop_within(WAIT) + .expect("pop_within") + .expect("the append's completion arrives well inside the bound"); + assert!(appender.claim(&completion).expect("a successful claim")); + } + + let bytes = std::fs::read(&path).expect("read the log back"); + assert_eq!( + bytes.len(), + 2 * record::RECORD_STRIDE, + "two records occupy two whole blocks, whatever their own lengths" + ); + + let second = &bytes[record::RECORD_STRIDE..]; + let decoded = record::decode(second).expect("the short record decodes"); + assert_eq!( + decoded.payload, short, + "the second block holds the second record" + ); + assert!( + second[decoded.extent()..].iter().all(|&byte| byte == 0), + "the rest of the block must be zero, not the 0xAB tail of the record \ + that used this slot before it" + ); + + drop(file); + let _ = std::fs::remove_file(&path); +} + +/// Records land one per stride, and replay walks them back (M25.1 + M25.2). +/// +/// **The guard this sample did not have.** M25.1 changed the writer's layout +/// and M25.2 the reader's, and between the two the log is unreadable -- replay +/// advances into a zeroed block tail and reports every record after the first +/// as missing. That was measured, not imagined: with the writer converted and +/// the reader not, every test in this file still passed, and only running the +/// example caught it. No CI job runs the example. +/// +/// So this binds the two ends together at a rung that runs on every machine: +/// it appends through the real writer, reads the real file, and hands it to the +/// real reader. Either end changing alone turns it red. +#[test] +fn records_land_one_per_stride_and_replay_walks_them_back() { + let payload_for = |index: usize| format!("record {index}: the quick brown fox").into_bytes(); + const COUNT: usize = 3; + + let mut ring = IoRing::new(16, 16).expect("create ring"); + let (path, file) = scratch("stride-replay"); + let mut appender = + Appender::new(&mut ring, &Placement::decide(file.as_raw_handle())).expect("appender"); + + for index in 0..COUNT { + let pushed = appender + .append_batch( + &mut ring, + file.as_raw_handle(), + Epoch(0), + &[payload_for(index)], + ) + .expect("push one append"); + assert_eq!(pushed, 1, "a drained arena always has a slot"); + let completion = ring + .pop_within(WAIT) + .expect("pop_within") + .expect("the append's completion arrives well inside the bound"); + assert!(appender.claim(&completion).expect("a successful claim")); + } + + let bytes = std::fs::read(&path).expect("read the log back"); + assert_eq!( + bytes.len(), + COUNT * record::RECORD_STRIDE, + "each record occupies exactly one block" + ); + + let outcome = crate::replay::replay(&bytes, Epoch(0), COUNT, payload_for); + assert!( + outcome.is_clean(), + "the reader must walk the writer's layout: {:?}", + outcome.violations + ); + assert_eq!( + outcome.durable_verified, COUNT, + "every record written must be read back, in sequence, with its payload intact" + ); + + drop(file); + let _ = std::fs::remove_file(&path); +} diff --git a/crates/windows-ioring-sys/examples/epoch_log/checkpoint.rs b/crates/windows-ioring-sys/examples/epoch_log/checkpoint.rs index 9856ea13e..833297ed7 100644 --- a/crates/windows-ioring-sys/examples/epoch_log/checkpoint.rs +++ b/crates/windows-ioring-sys/examples/epoch_log/checkpoint.rs @@ -25,6 +25,18 @@ //! available on one ring. The choice is per ring, and a program that wants both //! shapes buys both rings. //! +//! **There is a second, independent reason for the separation, and it survives +//! any change to the delivery model.** A covering flush's barrier waits for +//! every operation outstanding on the ring it is pushed to, so putting +//! checkpoint writes on the log's ring would drag them into every commit's +//! barrier and couple the log's commit latency to control-plane work. That is a +//! cost coupling rather than a correctness one -- the flush names a *file*, so +//! the log's own durability guarantee would survive the sharing -- but the cost +//! model the log is built around would not. See [`crate::contract`] -> "One ring +//! per log, because the barrier is ring-wide". Stated here because the delivery +//! argument above is the one a reader meets at this point of use: if it ever +//! stops applying, the rings must still not be collapsed. +//! //! # The ordering chain, and where it crosses threads //! //! 1. The **log thread** observes epoch *N* durable and submits a checkpoint: diff --git a/crates/windows-ioring-sys/examples/epoch_log/commit.rs b/crates/windows-ioring-sys/examples/epoch_log/commit.rs index 95a66cd2b..5f0321e6c 100644 --- a/crates/windows-ioring-sys/examples/epoch_log/commit.rs +++ b/crates/windows-ioring-sys/examples/epoch_log/commit.rs @@ -142,8 +142,30 @@ impl Committer { /// # Errors /// /// The flush's own failure, if it failed. A failed commit advances - /// nothing: the epoch it was closing is *not* durable, and saying so is - /// the whole point of checking. + /// nothing: the epoch it was closing is *not* durable when this returns, + /// and saying so is the whole point of checking. + /// + /// # A failed commit is not permanent, and that is not a loophole + /// + /// Epoch *N* stays un-durable only until some later commit succeeds. + /// Every commit here is a **covering** flush, so commit *N+1* reaches + /// every operation outstanding when it runs -- which includes epoch *N*'s + /// writes, queued before it. When *N+1*'s completion is observed, + /// `durable_through` advances to *N+1*, and [`Committer::is_durable`] + /// begins answering `true` for *N* as well. + /// + /// That is the truthful answer rather than an over-claim. What makes a + /// record durable is a flush that covered it, not the identity of the + /// flush that happened to be *named* for its epoch. Holding *N* + /// un-durable forever on the strength of one failed call would under-report + /// a record whose bytes the device already has -- and this module's whole + /// posture is that reporting less than reality is safe only while it stays + /// *reachable*, not as a permanent verdict. + /// + /// So the monotonicity [`Committer::is_durable`] promises survives a + /// failed commit rather than being suspended by it. What a caller must not + /// read into a failure is "epoch *N* is lost": it means *not yet*, and the + /// next successful commit is what settles it. pub fn claim(&mut self, completion: &Completion) -> io::Result> { let Some(epoch) = self.in_flight.remove(&completion.user_data()) else { return Ok(None); @@ -153,11 +175,21 @@ impl Committer { // that epoch keeps getting `false` -- which is the truthful answer. completion.result()?; - // Commits are barrier-ordered against each other (D-24 holds an - // operation pushed after a drained one until it completes), so - // completions should arrive in epoch order. `max` rather than plain - // assignment anyway: if that expectation is ever wrong, the reported - // answer stays correct and only the assertion is noisy. + // Commits arrive in epoch order because each one carries the drain + // flag *itself*: commit N is still outstanding when commit N+1 is + // reached, and D-47's surviving half -- no operation queued before a + // drained flush was ever observed completing after it -- is what puts + // N first. + // + // Note what this deliberately does not rest on. D-24 originally + // claimed a drained operation holds back what is pushed behind it, and + // D-47 withdrew that (see this module's header). The ordering here is + // bought by the *later* flush's own flag, never by the earlier one + // holding anything, so it survives the withdrawal intact. + // + // `max` rather than plain assignment anyway: if that expectation is + // ever wrong, the reported answer stays correct and only the assertion + // is noisy. debug_assert!( self.durable_through.is_none_or(|through| through < epoch), "commit completions should arrive in epoch order" @@ -166,3 +198,6 @@ impl Committer { Ok(Some(Epoch(epoch))) } } + +#[cfg(test)] +mod tests; diff --git a/crates/windows-ioring-sys/examples/epoch_log/commit/tests.rs b/crates/windows-ioring-sys/examples/epoch_log/commit/tests.rs new file mode 100644 index 000000000..295a98c51 --- /dev/null +++ b/crates/windows-ioring-sys/examples/epoch_log/commit/tests.rs @@ -0,0 +1,274 @@ +// Copyright (c) 2026 Mike Grier +//! Tests for the epoch bookkeeping (M21.4). +//! +//! The case these exist for cannot be reached by *running* the sample: a +//! commit only fails if its flush fails, and a flush against a healthy temp +//! file does not. The crate's fault-injection seam +//! ([`Completion::with_injected_failure`]) is what makes the failed-commit +//! path reachable at all, which is the same reason `tests/fault_injection.rs` +//! exists one layer down. +//! +//! Only the tests that *need* the seam are gated on `fault-injection`, so a +//! default `cargo test` still runs the rest. The gated ones are covered by +//! CI's `cargo test --workspace --all-features` job, the same job that covers +//! `tests/fault_injection.rs`; a local run needs `--features fault-injection` +//! to see them. + +use std::os::windows::io::AsRawHandle; +use std::time::Duration; + +use windows_ioring_sys::{Completion, IoRing}; + +#[cfg(feature = "fault-injection")] +use windows_ioring_sys::IoRingErrorExt; + +use super::{Committer, Epoch}; + +/// Hang bound on every wait here. Far above any real flush. +const WAIT: Duration = Duration::from_secs(30); + +/// The failure injected into a commit's completion. Any error would do; the +/// point is that `result()` reports one. +#[cfg(feature = "fault-injection")] +const INJECTED: windows_ioring_sys::InjectedFailure = + windows_ioring_sys::InjectedFailure::Win32(windows_sys::Win32::Foundation::ERROR_ACCESS_DENIED); + +/// A scratch file to commit against, named per test so tests running as +/// threads in one process cannot collide on it. +fn scratch(tag: &str) -> (std::path::PathBuf, std::fs::File) { + let path = std::env::temp_dir().join(format!( + "windows-ioring-sys-epoch-commit-{}-{tag}.tmp", + std::process::id() + )); + std::fs::write(&path, b"x").expect("create fixture"); + let file = std::fs::OpenOptions::new() + .read(true) + .write(true) + .open(&path) + .expect("open fixture"); + (path, file) +} + +/// Close the open epoch and wait for its commit's completion, returning both. +/// +/// Uses the bounded pop published in M21.2, which is what makes this a wait +/// rather than the spin these tests would otherwise have had to write. +fn commit_and_pop( + ring: &mut IoRing, + committer: &mut Committer, + file: &std::fs::File, +) -> (Epoch, Completion) { + let closed = committer + .commit(ring, file.as_raw_handle()) + .expect("push the commit"); + let completion = ring + .pop_within(WAIT) + .expect("pop_within") + .expect("the commit's completion arrives well inside the bound"); + (closed, completion) +} + +#[cfg(feature = "fault-injection")] +#[test] +fn a_failed_commit_leaves_its_epoch_not_durable() { + let mut ring = IoRing::new(16, 16).expect("create ring"); + let (path, file) = scratch("failed"); + let mut committer = Committer::new(); + + let (closed, completion) = commit_and_pop(&mut ring, &mut committer, &file); + let error = committer + .claim(&completion.with_injected_failure(INJECTED)) + .expect_err("a failed flush must be reported, not swallowed"); + // The crate preserves the HRESULT rather than classifying it, so the + // kind is `Other` and the Win32 code is what identifies the failure -- + // the same assertion `tests/fault_injection.rs` makes one layer down. + assert_eq!( + (error.as_ioring_error().expect("an IoRingError").code() as u32) & 0xFFFF, + windows_sys::Win32::Foundation::ERROR_ACCESS_DENIED, + "the error reaching the caller must be the one that was injected" + ); + + assert!( + committer.durable_through().is_none(), + "a failed commit advances the watermark not at all" + ); + assert!( + !committer.is_durable(closed), + "the epoch its flush was named for is not durable yet" + ); + let _ = std::fs::remove_file(&path); +} + +#[cfg(feature = "fault-injection")] +#[test] +fn a_later_successful_commit_covers_an_epoch_whose_own_commit_failed() { + // The claim this item exists to bind. Epoch 0's flush fails; epoch 1's + // flush is covering, so it reaches epoch 0's writes too, and observing it + // makes epoch 0 durable after all. + let mut ring = IoRing::new(16, 16).expect("create ring"); + let (path, file) = scratch("covered"); + let mut committer = Committer::new(); + + let (first, completion) = commit_and_pop(&mut ring, &mut committer, &file); + committer + .claim(&completion.with_injected_failure(INJECTED)) + .expect_err("epoch 0's own commit fails"); + // Asserted before the second commit, so this test cannot pass against an + // implementation that reported `true` all along. + assert!( + !committer.is_durable(first), + "epoch 0 must not be durable between its failure and the next success" + ); + + let (second, completion) = commit_and_pop(&mut ring, &mut committer, &file); + let advanced = committer + .claim(&completion) + .expect("epoch 1's commit succeeds") + .expect("and it is one of ours"); + assert_eq!(advanced, second); + + assert!( + committer.is_durable(first), + "epoch 1's covering flush reached epoch 0's writes, so epoch 0 is durable now" + ); + assert!(committer.is_durable(second)); + let _ = std::fs::remove_file(&path); +} + +#[test] +fn an_epoch_above_the_watermark_is_never_durable() { + // The other direction of the guard above: a `is_durable` that answered + // `true` unconditionally would pass every assertion in this file except + // this one. + let mut ring = IoRing::new(16, 16).expect("create ring"); + let (path, file) = scratch("above"); + let mut committer = Committer::new(); + + let (closed, completion) = commit_and_pop(&mut ring, &mut committer, &file); + committer + .claim(&completion) + .expect("the commit succeeds") + .expect("and it is ours"); + + assert!(committer.is_durable(closed)); + assert!( + !committer.is_durable(committer.open_epoch()), + "the epoch still open was never committed and must not report durable" + ); + assert!( + !committer.is_durable(Epoch(closed.0 + 7)), + "nor may an epoch that does not exist yet" + ); + let _ = std::fs::remove_file(&path); +} + +#[cfg(feature = "fault-injection")] +#[test] +fn durability_stays_monotonic_across_a_failed_commit() { + // Three epochs, the middle one's commit failing. Once the third settles, + // every epoch at or below the watermark must report durable -- which is + // the property that lets a caller remember one number instead of a set. + let mut ring = IoRing::new(16, 32).expect("create ring"); + let (path, file) = scratch("monotonic"); + let mut committer = Committer::new(); + + let (first, completion) = commit_and_pop(&mut ring, &mut committer, &file); + committer.claim(&completion).expect("epoch 0 commits"); + + let (second, completion) = commit_and_pop(&mut ring, &mut committer, &file); + committer + .claim(&completion.with_injected_failure(INJECTED)) + .expect_err("epoch 1's commit fails"); + assert!( + !committer.is_durable(second), + "the watermark must not move past a failure" + ); + assert!( + committer.is_durable(first), + "and it must not move backwards either" + ); + + let (third, completion) = commit_and_pop(&mut ring, &mut committer, &file); + committer.claim(&completion).expect("epoch 2 commits"); + + for epoch in 0..=third.0 { + assert!( + committer.is_durable(Epoch(epoch)), + "epoch {epoch} is at or below the watermark and must report durable" + ); + } + let _ = std::fs::remove_file(&path); +} + +#[cfg(feature = "fault-injection")] +#[test] +fn a_failed_commit_is_no_longer_in_flight() { + // The epoch is removed from the in-flight map *before* the result is + // checked. Were it removed after, a failed commit would be accounted for + // forever and the log could never quiesce. + let mut ring = IoRing::new(16, 16).expect("create ring"); + let (path, file) = scratch("in-flight"); + let mut committer = Committer::new(); + + let (_, completion) = commit_and_pop(&mut ring, &mut committer, &file); + assert_eq!(committer.in_flight(), 1, "pushed but not yet observed"); + committer + .claim(&completion.with_injected_failure(INJECTED)) + .expect_err("the commit fails"); + assert_eq!( + committer.in_flight(), + 0, + "a failure still settles the accounting; only the watermark is withheld" + ); + let _ = std::fs::remove_file(&path); +} + +#[test] +fn a_completion_that_belongs_to_someone_else_is_not_claimed() { + // A drain loop hands every completion to every claimant, so answering + // "not mine" without touching the watermark is load-bearing rather than + // defensive. The foreign completion here is real rather than fabricated: + // a flush pushed directly, bypassing the committer entirely. + use windows_ioring_sys::{Batch, FlushCoverage, FlushMode}; + + let mut ring = IoRing::new(16, 16).expect("create ring"); + let (path, file) = scratch("foreign"); + let mut committer = Committer::new(); + + let (closed, completion) = commit_and_pop(&mut ring, &mut committer, &file); + committer.claim(&completion).expect("our own commit"); + let before = committer.durable_through(); + + let mut batch = Batch::new(&mut ring); + // SAFETY: `file` outlives this operation -- it is drained below, before + // the test returns. + let foreign_id = unsafe { + batch.flush_raw( + file.as_raw_handle(), + FlushCoverage::Unordered, + FlushMode::Default, + ) + } + .expect("queue a flush nobody is tracking"); + batch.submit().expect("submit"); + let foreign = ring + .pop_within(WAIT) + .expect("pop_within") + .expect("the foreign flush completes"); + assert_eq!(foreign.user_data(), foreign_id); + + assert!( + committer + .claim(&foreign) + .expect("a foreign completion is not an error") + .is_none(), + "a completion this committer never pushed is not its business" + ); + assert_eq!( + committer.durable_through(), + before, + "and it must not move the watermark" + ); + assert!(committer.is_durable(closed)); + let _ = std::fs::remove_file(&path); +} diff --git a/crates/windows-ioring-sys/examples/epoch_log/contract.rs b/crates/windows-ioring-sys/examples/epoch_log/contract.rs index fbeb5b3c6..a202379e9 100644 --- a/crates/windows-ioring-sys/examples/epoch_log/contract.rs +++ b/crates/windows-ioring-sys/examples/epoch_log/contract.rs @@ -70,6 +70,68 @@ //! committed epoch may be wholly present, wholly absent, or torn, and all //! three are legal outcomes of the same crash. A reader must tolerate all //! three, which is exactly what the replay pass is written to do. +//! - **No bound on what a commit waits for.** The covering flush's barrier +//! reaches *every operation outstanding on the ring when the flush is +//! reached* -- not only the records of the epoch being closed. Appends +//! already accepted into the *next* epoch are therefore often covered +//! incidentally. That incidental coverage is not a promise and must not be +//! read as one: a record is durable when **its own** epoch's commit +//! completes, which is the guarantee above and the only one. +//! +//! # What this contract requires of the handle +//! +//! The log is handed an open handle and never opens one itself, so everything +//! above rests on that handle being able to sustain the operations this +//! program issues. Those requirements are stated here and **not checked**. +//! +//! - **Positioned I/O at explicit offsets.** Every write names its own offset; +//! the log never relies on a file pointer. +//! - **Opened with `FILE_FLAG_OVERLAPPED`.** The ring's operations are +//! asynchronous. +//! - **Opened with `FILE_FLAG_NO_BUFFERING`**, together with the alignment +//! that flag imposes: the buffer address, the file offset and the length are +//! each a multiple of the volume's sector size. [`crate::logfile`] records +//! why the log wants unbuffered I/O rather than merely tolerating it. +//! - **A preallocated extent** large enough for the log. This program does not +//! extend the file as it appends. +//! - **A successful write of `N` bytes transferred `N` bytes.** The ring +//! permits a short count -- it never asks what kind of handle it was given +//! -- so this is a narrowing the log requires rather than something it +//! inherits. +//! - **A completed flush has reached stable media.** +//! +//! The last two are different in kind from each other, and the difference is +//! the point of listing them separately. +//! +//! **The transfer requirement is checked.** [`crate::append`] compares the +//! transferred count against the length it asked for and fails the append if +//! they differ, so a handle that does not meet this requirement is reported +//! rather than silently producing a log with holes in it. +//! +//! **The durability requirement is a warranty the caller gives.** Nothing in +//! this program can verify it, and no amount of inspecting the handle would. +//! It is stated in those words deliberately: a contract that merely sounds +//! confident about durability is how a silent failure gets built on. +//! +//! ## Why there is no pre-flight check on the handle +//! +//! A gate was designed and rejected, and the reasoning is recorded so it is +//! not re-proposed as a fresh idea. Sort the ways a handle can fail to meet +//! the requirements above by how they present: +//! +//! - A pipe, a socket, a character device or a closed handle **fails loudly** +//! at the first positioned write. A check identifies it earlier and buys a +//! clearer message, nothing more. +//! - A RAM disk, a remote share, or a volume whose write cache is not +//! power-protected **succeeds at every operation this program issues** and +//! silently fails to be durable. Nothing reachable from a handle settles the +//! last of those, because write-cache state belongs to the device. +//! +//! So a gate guards the failures that were already loud and misses every +//! failure that is silent, which is the inverse of what a durability layer +//! needs. Its real cost is not the call but the claim: a check that cannot +//! establish the property still reads, to a later maintainer, as though the +//! property had been established -- and so discourages them from asking. //! //! # What this contract assumes //! @@ -81,6 +143,13 @@ //! - **A record is at most one write.** This sample does not split a record //! across writes, so it never has to reason about a partially-written record //! whose pieces landed in different epochs. +//! - **The log's ring carries only the log's operations.** The barrier waits +//! for everything outstanding on the ring, so a shared ring couples this +//! log's commit latency to work it knows nothing about. The durability +//! *guarantee* survives such sharing -- the flush names a file -- but the +//! cost model does not, and the cost model is why the guarantees above are +//! worth having. See "One ring per log, because the barrier is ring-wide" +//! below. //! //! # Why an epoch at all //! @@ -94,6 +163,43 @@ //! amortizes one expensive operation over many records -- the group-commit //! shape every write-ahead log converges on -- and the price is precisely the //! non-guarantees above. +//! +//! # One ring per log, because the barrier is ring-wide +//! +//! One call sets two scopes, which is why they are easy to merge. +//! `FlushCoverage::CoversPrecedingOperations` is a flag on the **ring**: a +//! drained flush does not execute until every operation outstanding when it was +//! reached has *completed* (D-47, measured over roughly 4,500 trials). +//! `Batch::flush` names a **file**: what a syncing mode pushes to stable media +//! is that file's data and the device cache behind it. Completion is not +//! durability -- this contract says so above, about a record's own write -- so +//! the barrier bounds what a commit **waits for**, and the flush bounds what it +//! **makes durable**. +//! +//! The consequence that matters here is cost, and it is unconditional: because +//! the barrier waits for everything on the ring, a shared ring makes this log's +//! commit latency a function of unrelated work. Somebody else's slow operation +//! is this log's slow commit, whatever the storage underneath turns out to be. +//! That is why one ring per log is a precondition -- of the *cost model* that +//! makes the guarantees above worth having, rather than of their correctness, +//! which the flush's own file target secures. A log spanning *two* rings gets +//! no single durability point across both, and needs two commits with an +//! explicit join between them. +//! +//! (Whether some *other* file on a shared ring is also made durable depends on +//! whether it sits behind the same device cache, which this log does not +//! determine. It does not arise here -- one log file, one ring -- and is left +//! to `M23.2` rather than reasoned about in advance.) +//! +//! **This sample honors the precondition, and does so for a second, independent +//! reason.** The checkpoint has its own ring ([`crate::checkpoint`]) because a +//! ring handed to `EventDelivery` is owned by it and cannot also be drained by +//! the log thread -- a *delivery* argument, and the only one stated at that +//! point of use. The structure is therefore right twice over, which is +//! comfortable and is also the hazard: a future change to the delivery model +//! would retire the reason written down over there, and nothing over there +//! mentions this one. The separation is load-bearing for the cost model +//! whatever the delivery model becomes. /// Which part of the contract a statement belongs to. #[derive(Clone, Copy, Debug, PartialEq, Eq)] @@ -103,16 +209,48 @@ pub enum Clause { /// Something a caller must not assume, stated so the omission is explicit /// rather than left to be inferred from silence. DoesNotGuarantee, + /// Something the caller must provide, and which this program does not + /// check for (M26.11, [D-70]). + /// + /// Distinct from [`Self::Assumes`] by who can act on it. A requirement is + /// something the caller *chooses* -- how the handle is opened, what it + /// refers to -- so naming it tells them what to do. An assumption is about + /// the world, and the only response to it is to pick different hardware. + /// Merging the two would bury the actionable in the unverifiable. + /// + /// [D-70]: ../DESIGN-NOTES.md + Requires, /// Something outside this program's control that the guarantees rest on. Assumes, } impl Clause { + /// Every clause, in the order a report should present them. + /// + /// Exists so that a caller printing the contract **asks** for the list + /// rather than restating it. [`crate::main`]'s report previously carried + /// its own array of the three variants, which would have silently omitted + /// a fourth: the report would simply have been one section short, with + /// nothing failing. + /// + /// This list is still hand-maintained -- Rust offers no exhaustive + /// iteration of a plain enum. What brings a new variant to the author's + /// attention is [`Self::heading`] below, whose `match` is exhaustive and + /// will not compile until the new variant is handled; this array sits + /// beside it so the two are edited together. + pub const ALL: [Self; 4] = [ + Self::Guarantees, + Self::DoesNotGuarantee, + Self::Requires, + Self::Assumes, + ]; + /// The heading this clause is printed under. pub fn heading(self) -> &'static str { match self { Self::Guarantees => "guarantees", Self::DoesNotGuarantee => "does NOT guarantee", + Self::Requires => "requires of the handle it is given", Self::Assumes => "assumes", } } @@ -164,6 +302,47 @@ pub const CONTRACT: &[Statement] = &[ text: "anything about records after the last committed epoch -- they may be present, \ absent, or torn, and a reader must tolerate all three", }, + Statement { + clause: Clause::DoesNotGuarantee, + text: "that a commit waits only for its own epoch -- the covering flush's barrier reaches \ + every operation outstanding on the ring when it is reached, so records already \ + accepted into the next epoch are often covered incidentally, which promises \ + nothing about them", + }, + Statement { + clause: Clause::Requires, + text: "positioned I/O at explicit offsets -- every write names its own offset and this log \ + never relies on a file pointer", + }, + Statement { + clause: Clause::Requires, + text: "a handle opened with FILE_FLAG_OVERLAPPED, because the ring's operations are \ + asynchronous", + }, + Statement { + clause: Clause::Requires, + text: "a handle opened with FILE_FLAG_NO_BUFFERING, and the alignment that imposes: the \ + buffer address, the file offset and the length are each a multiple of the volume's \ + sector size", + }, + Statement { + clause: Clause::Requires, + text: "a preallocated extent large enough for the log, because this program does not \ + extend the file as it appends", + }, + Statement { + clause: Clause::Requires, + text: "that a successful write of N bytes transferred N bytes -- the ring permits a short \ + count because it never asks what kind of handle it was given, so this is a \ + narrowing this log requires rather than one it inherits; it is the one requirement \ + here that is checked, and a violation fails the append", + }, + Statement { + clause: Clause::Requires, + text: "that a completed flush has reached stable media -- a warranty the caller gives, \ + which nothing in this program can verify and no inspection of the handle would \ + establish", + }, Statement { clause: Clause::Assumes, text: "the device honors the flush and commits its volatile write cache; a device that \ @@ -173,4 +352,14 @@ pub const CONTRACT: &[Statement] = &[ clause: Clause::Assumes, text: "a record is written by at most one write, so no record straddles an epoch boundary", }, + Statement { + clause: Clause::Assumes, + text: "this log's ring carries only this log's operations -- the barrier waits for \ + everything outstanding on the ring, so a shared ring couples this log's commit \ + latency to unrelated work; the durability guarantee survives sharing because the \ + flush names a file, but the cost model does not", + }, ]; + +#[cfg(test)] +mod tests; diff --git a/crates/windows-ioring-sys/examples/epoch_log/contract/tests.rs b/crates/windows-ioring-sys/examples/epoch_log/contract/tests.rs new file mode 100644 index 000000000..7d3bf2b8c --- /dev/null +++ b/crates/windows-ioring-sys/examples/epoch_log/contract/tests.rs @@ -0,0 +1,130 @@ +// Copyright (c) 2026 Mike Grier +//! Tests for the contract's *presentable* form (M23.1). +//! +//! # What these check, and what they deliberately do not +//! +//! They check properties of what the program **prints**: that every clause the +//! report iterates has something to say, that two clauses cannot arrive under +//! one heading, and that no statement renders as an empty or ragged bullet. +//! Those are facts about the output, and each is a way the report could +//! silently degrade while every other test in the sample stayed green. +//! +//! They deliberately do **not** assert that any particular rule is present -- +//! no test here looks for the word "ring", or counts the assumptions. The +//! repository's CONTRACT INTEGRITY rule is explicit that a hand-written second +//! copy of a contract rule is not a check of the contract, it is a check of the +//! copy: such a test passes exactly when the two copies agree, says nothing +//! about whether either is right, and adds a second site to edit whenever the +//! contract changes. The prose in the module documentation is the authoritative +//! form and [`CONTRACT`] is its presentable reduction; what is worth testing is +//! that the reduction survives being printed. +//! +//! The one guard that genuinely belongs at the build rung is already there: +//! [`Clause::heading`]'s `match` is exhaustive, so a new variant cannot be +//! added without the compiler stopping at the site that must handle it. + +use super::{CONTRACT, Clause}; + +/// Every clause the report walks has at least one statement. +/// +/// The failure this prevents is quiet: [`Clause::ALL`] drives the report's +/// section headings, so a clause with no statements prints its heading and then +/// nothing, which reads as "this log promises nothing" rather than as a missing +/// entry. +#[test] +fn every_clause_has_at_least_one_statement() { + for clause in Clause::ALL { + let count = CONTRACT.iter().filter(|s| s.clause == clause).count(); + assert!( + count > 0, + "clause {clause:?} (\"{}\") has no statements, so the report would print an empty \ + section under its heading", + clause.heading() + ); + } +} + +/// Every statement belongs to a clause the report actually walks. +/// +/// The mirror of the test above, and the direction that is easy to forget: the +/// first proves `ALL` is covered by `CONTRACT`, this proves `CONTRACT` is +/// covered by `ALL`. A statement whose clause is missing from `ALL` is never +/// printed at all -- written down, compiled, and silently absent from the +/// report a reader is told to trust. +#[test] +fn every_statement_is_reachable_from_the_report() { + for statement in CONTRACT { + assert!( + Clause::ALL.contains(&statement.clause), + "statement {:?} has clause {:?}, which Clause::ALL does not list, so the report \ + would never print it", + statement.text, + statement.clause + ); + } +} + +/// No two clauses share a heading, and none is blank. +/// +/// Two clauses printing under one heading would merge distinct parts of the +/// contract in the output -- a caller reading "this log guarantees" would be +/// shown things it explicitly does not guarantee, which is the worst available +/// failure for this particular program. +#[test] +fn headings_are_distinct_and_non_empty() { + for (index, clause) in Clause::ALL.iter().enumerate() { + let heading = clause.heading(); + assert!(!heading.trim().is_empty(), "{clause:?} has a blank heading"); + for other in &Clause::ALL[index + 1..] { + assert_ne!( + heading, + other.heading(), + "{clause:?} and {other:?} would print under the same heading" + ); + } + } +} + +/// `Clause::ALL` lists each clause once. +/// +/// A repeat would print a whole section twice. Worth a test rather than a +/// glance because `ALL` is hand-maintained -- the exhaustive `match` in +/// `heading` forces a new variant to be *handled*, but nothing forces it to be +/// added to `ALL` exactly once. +#[test] +fn all_lists_each_clause_once() { + for (index, clause) in Clause::ALL.iter().enumerate() { + assert!( + !Clause::ALL[index + 1..].contains(clause), + "{clause:?} appears more than once in Clause::ALL" + ); + } +} + +/// Statements render as clean bullets. +/// +/// The report prints each as ` - {text}`, so leading or trailing whitespace +/// shows up as a ragged list and an empty statement as a bare dash. Both are +/// the kind of thing a line-continuation edit introduces without anyone +/// noticing, since the source is wrapped across several lines. +#[test] +fn statement_text_is_clean() { + for statement in CONTRACT { + let text = statement.text; + assert!( + !text.is_empty(), + "a {:?} statement is empty and would print as a bare dash", + statement.clause + ); + assert_eq!( + text.trim(), + text, + "statement {text:?} has leading or trailing whitespace and would print ragged" + ); + assert!( + !text.contains('\n'), + "statement {text:?} contains a newline, which would break the one-bullet-per-line \ + shape the report assumes" + ); + } +} diff --git a/crates/windows-ioring-sys/examples/epoch_log/logfile.rs b/crates/windows-ioring-sys/examples/epoch_log/logfile.rs new file mode 100644 index 000000000..fb3b9d3b6 --- /dev/null +++ b/crates/windows-ioring-sys/examples/epoch_log/logfile.rs @@ -0,0 +1,244 @@ +// Copyright (c) 2026 Mike Grier +//! Opening the log's file so a commit is separately observable (M25.3). +//! +//! # What this does and why the order matters +//! +//! Two steps, and neither is interchangeable with the other: +//! +//! 1. **Zero-fill the extent** with an ordinary handle, then drop it. +//! 2. **Reopen** the existing file with `FILE_FLAG_NO_BUFFERING | +//! FILE_FLAG_OVERLAPPED`. +//! +//! [write-pending-spike.rs](../../design-sessions/spikes/write-pending-spike.rs) +//! measured five configurations, and what replicates across sixteen runs is +//! that a **buffered** handle essentially never pends while every +//! `NO_BUFFERING` one pends in most runs. Among the unbuffered conditions the +//! zero-filled extent has the highest rate and much the highest floor, which +//! is why the log uses it -- but the three overlap heavily and a single run of +//! any of them can land in another's range. See +//! [measurements/2026-09-24-set-len-vs-zero-fill/](../../measurements/2026-09-24-set-len-vs-zero-fill/README.md) +//! for the runs and for the earlier, stronger reading this replaced. +//! +//! # The zero-fill is not avoided, it is moved -- and moving it is not free +//! +//! Writing past NTFS's valid data length obliges the filesystem to zero-fill +//! the gap first. Step 1 does that zeroing once, eagerly, on an ordinary +//! handle, at a moment when nothing is being measured and no ring is involved. +//! +//! **For a sequential writer the total zeroing cost is the same either way**, +//! and that is measured: filling an extent after `set_len` costs what +//! zero-filling it outright costs (287 ms against 315 ms per GiB), because +//! every write lands exactly at the valid data length and none has a gap in +//! front of it. So this log is not buying cheaper zeroing. +//! +//! **What it is avoiding is the case where the bill arrives inside one write.** +//! A write that lands *past* the valid data length pays to zero the whole gap, +//! synchronously, before it proceeds: one sector written at the end of a +//! `set_len`'d 1 GiB file took roughly 2.3 seconds, about eight times the cost +//! of writing the entire extent. A log that only ever appends does not hit +//! that -- but a log is exactly the kind of program that later grows a +//! recovery path, a header rewrite, or a segment that seeks. See +//! [measurements/2026-09-24-set-len-zero-fill-cost/](../../measurements/2026-09-24-set-len-zero-fill-cost/README.md). +//! +//! # The forcing can be used on purpose, and is a real alternative +//! +//! Raised in review: since a write past the valid data length forces the fill +//! anyway, it can be *triggered* deliberately -- `set_len` to the final size, +//! then write one sector at the very end, and the filesystem zero-fills +//! everything in front of it. That was measured as condition F and it reaches +//! the same end state this function does; over sixteen runs it had the highest +//! floor of any condition (147/500 against the zero-fill's 65) at a +//! comparable median. +//! +//! It is a genuine trade rather than a strictly worse option: +//! +//! - **It needs no buffer at all** -- two syscalls, whatever the extent's size. +//! - **It costs about eight times the wall time** for a large extent, because +//! the filesystem's own fill is much slower than a sequential write of the +//! same bytes. +//! +//! This function takes the explicit fill because the cost is bounded and +//! predictable and the memory is now bounded too. A caller pre-allocating tens +//! of gigabytes, who would rather spend wall time than write the loop, has the +//! other option and it works. +//! +//! **None of this is perceptible at this sample's own sizes**, and that is +//! measured too: the log's extent is 140 KiB and each strategy file is 8 MiB, +//! so a whole run zero-fills 24 MiB in tens of milliseconds against a run that +//! takes over a second. At 140 KiB the cost is dominated by creating the file +//! rather than by writing zeros into it. The choice here is made for what this +//! code *teaches* a log that pre-allocates in gigabytes, not for what it costs +//! the sample. +//! +//! # `set_len` is not a substitute for the zero-fill, and the difference is +//! measured rather than argued +//! +//! Step 1 writes a real buffer of zeros rather than calling +//! [`std::fs::File::set_len`]. Both produce a file of the right size whose +//! bytes read back as zero -- reads past the valid data length are answered +//! with zeros the filesystem synthesises without touching the disk -- so the +//! two are indistinguishable to everything in this sample. Only the write +//! advances the valid data length, which is the thing that decides whether a +//! later write is extending. +//! +//! An earlier version of this comment asserted that a `set_len` extent "would +//! leave every write in condition C". That was reasoned from documentation and +//! was challenged in review, so it was measured instead, twice over. +//! +//! **On zeroing cost, the assertion was simply the wrong mechanism.** For a +//! sequential writer `set_len` costs nothing extra -- see the section above. +//! +//! **On pending rate there is a real difference**, which is the measured reason +//! this function zero-fills. The spike gained a condition E over a `set_len` +//! extent, and sixteen runs are in +//! [measurements/2026-09-24-set-len-vs-zero-fill/](../../measurements/2026-09-24-set-len-vs-zero-fill/README.md): +//! the zero-filled extent pended at a median of 471/500 against `set_len`'s +//! 268/500, with a floor of 121 against 1. But the extending case and the +//! `set_len` case are not distinguishable from each other on that data, so the +//! claim that `set_len` *is* the extending case remains unsupported and is not +//! made. +//! +//! **Read those measurements before relying on any of this.** The first also +//! corrects a stronger claim this repository had been repeating -- that the +//! zero-filled extent was the only condition that pended at all. That descends +//! from a single run and does not replicate. +//! +//! (`SetFileValidData` moves the valid data length without writing anything, +//! which is how a database pre-allocates in one syscall. It is not used here: +//! it needs `SE_MANAGE_VOLUME_NAME`, and it exposes whatever bytes were +//! previously on those clusters to anything that reads the file. An earlier +//! draft of this paragraph said it saves "a few milliseconds of zeroing", +//! which was wrong by two to three orders of magnitude -- the measured cost is +//! ~300 ms per GiB done well, and seconds per GiB when forced onto a seeking +//! write. That cost is the whole reason the API exists.) +//! +//! Nothing in the test suite catches a swap to `set_len`: see the declared +//! blind spot in [sabotage.json](../../sabotage.json). The difference is a +//! *rate* that varies enormously run to run, so a test asserting it would be +//! asserting an observation about one machine as though it were a contract, +//! which `M25`'s standing constraint forbids. +//! +//! # What is deliberately *not* opened this way +//! +//! The **checkpoint** file stays buffered. `NO_BUFFERING` constrains the +//! buffer's alignment, the file offset, and the transfer length, and a +//! checkpoint record is a sixteen-byte `Vec` written at offset 0 -- it fails +//! all three. Striding the control plane to satisfy a flag it does not need +//! would be the tail wagging the dog; the control plane's correctness comes +//! from its covering flush (see [`crate::checkpoint`]), not from how its bytes +//! are cached. +//! +//! The **retired** file is likewise ordinary: it is written once with +//! [`std::fs::write`] and read by the reclaim worker, never through a ring. +//! +//! # This buys an opportunity, never a guarantee +//! +//! Windows specifies nothing about when a ring operation completes relative to +//! `SubmitIoRing`, so "the write pends" is an observation about a machine and +//! not a contract. The log is correct either way and nothing in it may depend +//! on an operation pending. What this shape changes is whether a commit is +//! *separately measurable*, which is M25.4's problem. + +use std::fs::{File, OpenOptions}; +use std::io; +use std::io::Write; +use std::os::windows::fs::OpenOptionsExt; +use std::path::Path; + +use windows_sys::Win32::Storage::FileSystem::{FILE_FLAG_NO_BUFFERING, FILE_FLAG_OVERLAPPED}; + +use crate::record::RECORD_STRIDE; + +/// Bytes per write while zero-filling the extent. +/// +/// The fill used to be a single `std::fs::write` of a `vec![0; len]`, which +/// allocates the **whole extent** in memory -- harmless for this sample's +/// handful of blocks and a bad pattern for a log to copy, since a real one +/// pre-allocates in gigabytes. A fixed chunk keeps the fill's memory cost +/// constant in the size of the extent. +/// +/// # Why 64 KiB, measured rather than picked +/// +/// This was 1 MiB first, which is the worst of the plausible values on two +/// counts, both in +/// [measurements/2026-09-24-allocation-knee/](../../measurements/2026-09-24-allocation-knee/README.md): +/// +/// - **Throughput stops improving at about 64 KiB.** Filling at 4 KiB chunks +/// runs at a few hundred MiB/s; by 64 KiB it is within run-to-run variance +/// of every larger size tried, up to 4 MiB. So a bigger chunk buys nothing. +/// - **1 MiB is exactly where this allocator stops using the heap.** Measured +/// with `VirtualQuery`: at 512 KiB and below, sixty-four live allocations +/// share a handful of reservations; at 1 MiB every one gets its own. So the +/// first size with no throughput benefit is also the first size that +/// guarantees a reservation and its teardown on every call. +/// +/// 64 KiB is additionally the Windows virtual-memory allocation granularity, +/// which is why it is a good habit as well as a good measurement: it is the +/// point past which a block cannot share its region with anything else. +/// +/// # Why this is allocated rather than a `static` array of zeros +/// +/// A `static ZEROS: [u8; 64 * 1024]` would remove the allocation entirely -- +/// no heap, no knee to reason about, pages arriving demand-zero from the +/// loader. On every axis a performance reader would check it is the better +/// choice, which is exactly why the reason to refuse it is written down here: +/// **it is a security decision and no measurement will surface it.** +/// +/// A `static` lives at a fixed offset within the module, so a process that +/// leaks any module base thereby knows the address of a large, writable, +/// zero-filled region -- a ready-made landing pad for staging data, at an +/// address ASLR no longer protects once the base is known, present for the +/// life of the process whether or not a log is ever opened. A transient heap +/// allocation has neither property: its address is unpredictable and it exists +/// only while the fill is running. +/// +/// The cost of declining the `static` is one allocation per call to +/// [`create_preallocated`], which the measurements above put at nothing worth +/// having. +const FILL_CHUNK: usize = 64 * 1024; + +/// Create `path` with `blocks` zeroed record blocks already written, and return +/// a handle over that extent opened `NO_BUFFERING | OVERLAPPED`. +/// +/// Sized in blocks rather than bytes because every writer in this sample lands +/// one record per [`RECORD_STRIDE`] block, so blocks are the unit a caller +/// actually knows -- and a byte count that was not a whole number of blocks +/// could not be written through the returned handle anyway. +/// +/// # Errors +/// +/// Any error from the zero-fill or from reopening it. +pub fn create_preallocated(path: &Path, blocks: usize) -> io::Result { + let len = blocks + .checked_mul(RECORD_STRIDE) + .ok_or_else(|| io::Error::new(io::ErrorKind::InvalidInput, "log extent overflows"))?; + + // The zero-fill, done eagerly here rather than left for the filesystem to + // do lazily on the ring's write path -- see the module docs, and note that + // `set_len` alone is not a substitute however identical the resulting file + // looks. The ordinary handle is dropped at the end of this block, before + // the reopen. + { + let mut file = File::create(path)?; + let chunk = vec![0_u8; FILL_CHUNK.min(len.max(1))]; + let mut written = 0; + while written < len { + let take = chunk.len().min(len - written); + file.write_all(&chunk[..take])?; + written += take; + } + file.flush()?; + } + + // No `create`, no `truncate`: this must be `OPEN_EXISTING`, because + // truncating would discard the extent that is the only thing distinguishing + // this from the configuration the spike measured as behaving like a + // buffered handle. + OpenOptions::new() + .write(true) + .custom_flags(FILE_FLAG_NO_BUFFERING | FILE_FLAG_OVERLAPPED) + .open(path) +} + +#[cfg(test)] +mod tests; diff --git a/crates/windows-ioring-sys/examples/epoch_log/logfile/tests.rs b/crates/windows-ioring-sys/examples/epoch_log/logfile/tests.rs new file mode 100644 index 000000000..4672c88eb --- /dev/null +++ b/crates/windows-ioring-sys/examples/epoch_log/logfile/tests.rs @@ -0,0 +1,203 @@ +// Copyright (c) 2026 Mike Grier +//! Tests for the log file's shape (M25.3). +//! +//! # What these can and cannot establish +//! +//! They establish that the extent is really there and that the handle really +//! carries `FILE_FLAG_NO_BUFFERING`, both of which are checkable from here. +//! +//! They establish **nothing about whether a write pends**, and deliberately do +//! not try. Windows specifies nothing about when a ring operation completes +//! relative to `SubmitIoRing`, so a test asserting that an operation pends +//! would be asserting an observation about one machine as though it were a +//! contract -- and the log is required to be correct either way. +//! +//! They also cannot tell a written extent from a `set_len` one; see the module +//! docs for why, and the manifest's `notCoveredHere` for the record of it. + +use std::io::Write; +use std::os::windows::fs::OpenOptionsExt; +use std::os::windows::io::AsRawHandle; +use std::time::Duration; + +use windows_ioring_sys::{Batch, IoRing, NumaBuffer, PushOptions, WriteCaching}; + +use crate::record::RECORD_STRIDE; + +/// Hang bound on every wait here. Far above any real write. +const WAIT: Duration = Duration::from_secs(30); + +/// A scratch path named per test, so tests running as threads in one process +/// cannot collide on it. +fn scratch(tag: &str) -> std::path::PathBuf { + std::env::temp_dir().join(format!( + "windows-ioring-sys-epoch-logfile-{}-{tag}.tmp", + std::process::id() + )) +} + +#[test] +fn the_extent_is_written_before_the_handle_is_returned() { + const BLOCKS: usize = 4; + let path = scratch("extent"); + let file = super::create_preallocated(&path, BLOCKS).expect("create the log file"); + + let bytes = std::fs::read(&path).expect("read the extent back"); + assert_eq!( + bytes.len(), + BLOCKS * RECORD_STRIDE, + "the file must already span every block a caller asked for, before a \ + single record is written -- an extending write is the configuration \ + the spike measured as behaving like a buffered one" + ); + assert!( + bytes.iter().all(|&byte| byte == 0), + "a pre-allocated extent must read as zeros, which is also what lets \ + replay recognise the unwritten tail as NeverWritten" + ); + + drop(file); + let _ = std::fs::remove_file(&path); +} + +/// A synchronous write is refused, which is what `FILE_FLAG_OVERLAPPED` means. +/// +/// This began as an attempt to test the `NO_BUFFERING` alignment rule through +/// [`std::io::Write`], and failed on the *aligned* write -- which was the +/// discovery, not the defect. `write_all` issues `WriteFile` with a null +/// `OVERLAPPED`, and an asynchronous handle refuses that however well-aligned +/// the transfer is. So this cannot test alignment, but it does test the other +/// flag, which nothing else here can reach: `GetFileInformationByHandleEx` +/// does not report the handle's mode, and the ring works on synchronous and +/// asynchronous handles alike. +/// +/// The alignment rule is tested through the ring instead, below, which is also +/// how production reaches this handle. +#[test] +fn a_synchronous_write_is_refused_because_the_handle_is_overlapped() { + let path = scratch("overlapped"); + let mut file = super::create_preallocated(&path, 2).expect("create the log file"); + + // Sector-sized, sector-count-aligned, and inside the extent: everything + // NO_BUFFERING asks of a transfer's *length*. What is left to object to is + // the synchronous call itself. + let aligned = vec![0xAB_u8; RECORD_STRIDE]; + let refused = file + .write_all(&aligned) + .expect_err("an overlapped handle must refuse a synchronous write"); + assert_eq!( + refused.raw_os_error(), + Some(windows_sys::Win32::Foundation::ERROR_INVALID_PARAMETER as i32), + "the refusal must be the handle's mode and not some unrelated failure" + ); + + drop(file); + let _ = std::fs::remove_file(&path); +} + +/// The control for the test above: an ordinary handle accepts the same write. +/// +/// Without it, the refusal could be caused by anything at all -- a bad path, a +/// closed handle, a length the filesystem disliked -- and the pair would agree +/// on a conclusion neither had established. This isolates the flags as the +/// only difference. +#[test] +fn an_ordinary_handle_accepts_the_write_an_overlapped_one_refuses() { + let path = scratch("overlapped-control"); + let mut file = std::fs::OpenOptions::new() + .create(true) + .write(true) + .truncate(true) + .custom_flags(0) + .open(&path) + .expect("open an ordinary handle"); + + file.write_all(&vec![0xAB_u8; RECORD_STRIDE]) + .expect("a synchronous handle has no objection to a synchronous write"); + + drop(file); + let _ = std::fs::remove_file(&path); +} + +/// The handle really carries `NO_BUFFERING`, tested in **both** directions and +/// through the ring, which is how production reaches it. +/// +/// A guard that only showed the aligned write succeeding would pass just as +/// happily against a buffered handle, which is exactly the regression worth +/// catching: dropping the flag changes nothing a caller can see except the +/// thing this whole item exists for. The refusal is what distinguishes them; +/// the acceptance is what shows the refusal is about alignment rather than the +/// handle being unusable. +/// +/// The unaligned case breaks two of `NO_BUFFERING`'s three rules at once -- a +/// `Vec` guarantees no particular address, and its length is not a whole +/// number of sectors -- which is stated rather than tidied because it means +/// this test does not establish *which* rule refused it. It establishes that +/// the handle has rules a buffered one does not, which is the claim. +#[test] +fn the_ring_refuses_an_unaligned_write_and_accepts_an_aligned_one() { + let path = scratch("alignment"); + let file = super::create_preallocated(&path, 2).expect("create the log file"); + let mut ring = IoRing::new(8, 8).expect("create ring"); + + // Aligned on every axis: a `NumaBuffer` is page-granular and so + // sector-granular (M22.3), the length is one whole stride, and the offset + // is a block boundary. + let buffer = NumaBuffer::new(RECORD_STRIDE, None).expect("allocate an aligned buffer"); + let mut batch = Batch::new(&mut ring); + // SAFETY: `file` outlives the operation -- it is dropped at the end of this + // test, after the completion is popped -- and the buffer is moved into the + // token, which is held until then. + let accepted = unsafe { + batch.write_raw( + file.as_raw_handle(), + buffer, + 0, + PushOptions::new(), + WriteCaching::Cached, + ) + } + .expect("push the aligned write"); + batch.submit().expect("submit"); + let completion = ring + .pop_within(WAIT) + .expect("pop_within") + .expect("the write completes well inside the bound"); + // Claimed before the result is inspected, which is the M22.2 discipline: + // a token left unclaimed is still outstanding, whatever the write did. + assert!( + accepted.claim_if(&completion).is_ok(), + "the completion must be the aligned write's" + ); + completion + .result() + .expect("an aligned NO_BUFFERING write must be accepted"); + + let mut batch = Batch::new(&mut ring); + // SAFETY: as above. + let rejected = unsafe { + batch.write_raw( + file.as_raw_handle(), + vec![0xCD_u8; RECORD_STRIDE - 1], + 0, + PushOptions::new(), + WriteCaching::Cached, + ) + } + .expect("push the unaligned write"); + batch.submit().expect("submit"); + let completion = ring + .pop_within(WAIT) + .expect("pop_within") + .expect("the write completes well inside the bound"); + assert!( + rejected.claim_if(&completion).is_ok(), + "the completion must be the unaligned write's" + ); + completion + .result() + .expect_err("NO_BUFFERING must refuse a transfer that breaks its alignment rules"); + + drop(file); + let _ = std::fs::remove_file(&path); +} diff --git a/crates/windows-ioring-sys/examples/epoch_log/main.rs b/crates/windows-ioring-sys/examples/epoch_log/main.rs index 7b64bfa49..eac68f723 100644 --- a/crates/windows-ioring-sys/examples/epoch_log/main.rs +++ b/crates/windows-ioring-sys/examples/epoch_log/main.rs @@ -76,11 +76,16 @@ mod checkpoint; mod commit; mod contract; mod event_loop; +mod logfile; +mod placement; mod reclaim; mod record; mod replay; mod strategy; +#[cfg(test)] +mod tests; + use std::io; use std::os::windows::io::AsRawHandle; use std::path::PathBuf; @@ -92,6 +97,7 @@ use commit::{Committer, Epoch}; use contract::{CONTRACT, Clause}; use event_loop::{EventLoop, Woken}; use reclaim::Reclaimer; +use strategy::CommitTiming; use windows_ioring_sys::{Batch, IoRing}; /// How many records this demonstration appends into committed epochs. @@ -117,6 +123,21 @@ const QUIESCE_ATTEMPTS: usize = 64; /// durable. const RETIRED_LEN: u64 = 64 * 1024; +/// Blocks pre-allocated beyond what a run will actually write (M25.3). +/// +/// A real write-ahead log pre-allocates *ahead* of its writer rather than +/// exactly to it, because an append that reaches the end of the extent becomes +/// an extending write -- the configuration the spike measured as behaving like +/// a buffered handle, whatever flags the handle carries. Sizing to the exact +/// record count would put this log one record away from that. +/// +/// It also means a clean log now ends in zeros rather than at EOF, so replay +/// stops with `NeverWritten` where it previously ran out of bytes. That is the +/// ordinary shape of a pre-allocated log, it is not a violation, and it is the +/// path `M25.2` taught replay to tolerate -- so the sample exercises it rather +/// than leaving it to be met first by a reader of a real log. +const SLACK_BLOCKS: usize = 8; + /// The byte the retired segment is filled with, so "was it reclaimed?" has an /// answer that does not depend on what happened to be there. const RETIRED_FILL: u8 = 0xA5; @@ -203,14 +224,24 @@ fn run_log( retired: &std::path::Path, checkpoint_path: &std::path::Path, ) -> io::Result { - let file = std::fs::OpenOptions::new() - .create(true) - .write(true) - .truncate(true) - .open(path)?; + // Pre-allocated and opened NO_BUFFERING | OVERLAPPED (M25.3). Sized for + // every record this run will write plus slack: an append past the extent + // would be an extending write, which is the configuration the spike + // measured as behaving like a buffered one, and a log that pre-allocates + // exactly what it needs is one record away from being that log. + let file = logfile::create_preallocated(path, RECORDS + TAIL_RECORDS + SLACK_BLOCKS)?; let handle = file.as_raw_handle(); + // `RETIRED_LEN` is 64 KiB, which sits exactly at the threshold this + // repository treats as the point to ask whether an allocation needs to be + // contiguous, and not past it (M25.7). Left whole on that basis. A reader + // who grows this segment should revisit it: the write has the same shape + // as `logfile`'s zero-fill and chunks the same way, and the check further + // down is a fold over the bytes that never needs them all at once. std::fs::write(retired, vec![RETIRED_FILL; RETIRED_LEN as usize])?; + // Ordinary and buffered, deliberately: a checkpoint record is sixteen + // bytes from a `Vec` at offset 0, which satisfies none of NO_BUFFERING's + // three alignment rules. See `logfile`'s module docs. let checkpoint_file = std::fs::OpenOptions::new() .create(true) .write(true) @@ -218,7 +249,11 @@ fn run_log( .open(checkpoint_path)?; let mut ring = IoRing::new(64, 128)?; - let mut appender = Appender::new(&mut ring)?; + // Decided from the log's own handle, before the arena exists: the + // documented FSCTL takes a file handle directly, so the node the arena + // should prefer is answerable without a device-tree walk. + let placement = placement::Placement::decide(handle); + let mut appender = Appender::new(&mut ring, &placement)?; let mut committer = Committer::new(); // The reclaim worker is shared: the log thread waits on its handle, and a @@ -245,6 +280,7 @@ fn run_log( "arena registered: {SLOTS} slots of {SLOT_LEN} bytes; \ waiting on the ring's completion event alongside a reclaim event and a shutdown latch" )); + report.line(format_args!("{}", placement.describe())); // Something outside the I/O loop decides when to stop -- which is the only // reason a second handle is in the wait at all. @@ -263,19 +299,29 @@ fn run_log( let mut collected = 0usize; while appended < RECORDS { let epoch = committer.open_epoch(); - let payload = payload_for(appended); - match appender.append(&mut ring, handle, epoch, &payload) { - Ok(_sequence) => appended += 1, + // Offer the rest of this epoch in one call. The arena decides how many + // of them are actually taken, which is why the count comes back rather + // than being assumed. + let wanted = (EPOCH_SIZE - (appended % EPOCH_SIZE)).min(RECORDS - appended); + let payloads: Vec> = (appended..appended + wanted).map(payload_for).collect(); + + let accepted = appender.append_batch(&mut ring, handle, epoch, &payloads)?; + if accepted == 0 { // Every slot is in flight. This is the arena working as intended, // not an error: pump once to drain and try again. - Err(error) if error.kind() == io::ErrorKind::WouldBlock => { - let (_, popped) = - events.pump(WAIT_MS, || drain(&mut ring, &mut appender, &mut committer))?; - empty_wakes += usize::from(popped == 0); - } - Err(error) => return Err(error), + // + // `continue` is what keeps the epoch trigger below honest -- it is + // reachable only on a pass that appended something, so it cannot + // fire on a retry that made no progress. That was M21.3's point, + // and batching must not quietly undo it. + let (_, popped) = + events.pump(WAIT_MS, || drain(&mut ring, &mut appender, &mut committer))?; + empty_wakes += usize::from(popped == 0); + continue; } + appended += accepted; + // Reached only when an append landed, per the `continue` above. if appended % EPOCH_SIZE == 0 { let closed = committer.commit(&mut ring, handle)?; // Before the commit's completion is observed, the honest answer is @@ -330,18 +376,17 @@ fn run_log( // The uncommitted tail: appended, so their writes complete, but no commit // ever closes their epoch. The contract therefore promises nothing about // them, and the replay pass below is what proves the reader tolerates that. - for index in RECORDS..RECORDS + TAIL_RECORDS { - let epoch = committer.open_epoch(); - let payload = payload_for(index); - loop { - match appender.append(&mut ring, handle, epoch, &payload) { - Ok(_) => break, - Err(error) if error.kind() == io::ErrorKind::WouldBlock => { - events.pump(WAIT_MS, || drain(&mut ring, &mut appender, &mut committer))?; - } - Err(error) => return Err(error), - } + let tail_epoch = committer.open_epoch(); + let tail: Vec> = (RECORDS..RECORDS + TAIL_RECORDS).map(payload_for).collect(); + let mut tail_appended = 0; + while tail_appended < tail.len() { + let accepted = + appender.append_batch(&mut ring, handle, tail_epoch, &tail[tail_appended..])?; + if accepted == 0 { + events.pump(WAIT_MS, || drain(&mut ring, &mut appender, &mut committer))?; + continue; } + tail_appended += accepted; } report.line(format_args!( "appended {TAIL_RECORDS} more records into epoch {} and deliberately never committed it", @@ -555,6 +600,12 @@ fn verify( path: &std::path::Path, run: &LogRun, ) -> io::Result<()> { + // Read whole rather than streamed, and the size is stated because it is + // paid here: this log is `RECORDS + TAIL_RECORDS + SLACK_BLOCKS` blocks, + // so a few hundred kilobytes. `replay` explains why it takes a slice + // (M25.7) -- the short version is that a streaming reader would have to + // return `io::Error` alongside `Violation`, and keeping those apart is + // this verifier's whole purpose. let bytes = std::fs::read(path)?; // 1. The log as written. Everything committed must be intact, and the @@ -581,11 +632,33 @@ fn verify( assert_eq!(clean.durable_verified, run.durable_records); assert_eq!(clean.tail_records, run.tail_records); + // What the stride costs, reported rather than left to a doc comment + // (M25.1). Records are variable-length but occupy a whole block each, so + // the file is far larger than the data in it. The figures are given and + // the reader draws their own conclusion -- what is acceptable here depends + // entirely on a log's record size, which is a caller's question. + report.line(format_args!( + "layout: {} bytes of records in {} bytes of file, one record per {}-byte block", + clean.record_bytes, + bytes.len(), + record::RECORD_STRIDE + )); + // 2. A torn tail, which is what a crash actually leaves behind. Cutting // the file mid-record simulates a write that did not land whole. The // contract says the reader must tolerate this, so a violation here // would mean the reader is stricter than the contract allows. - let torn_at = bytes.len() - (record::HEADER_LEN + 4); + // + // Derived from the last record's own block rather than from the file + // length (M25.2). Records are strided now, so trimming a fixed number + // of bytes off the end of the file lands in the final record's zeroed + // remainder and tears nothing at all -- the replay would pass while + // demonstrating the opposite of what it claims. This cuts partway + // through the last record's payload, where `decode` reports `Truncated`. + // It also survives M25.3's pre-allocation, which decouples the file's + // length from the number of records in it entirely. + let last_record_start = (run.durable_records + run.tail_records - 1) * record::RECORD_STRIDE; + let torn_at = last_record_start + record::HEADER_LEN + 4; let torn = replay::replay( &bytes[..torn_at], run.durable_through, @@ -606,6 +679,31 @@ fn verify( torn.durable_verified, run.durable_records, "tearing the tail must not cost a single durable record" ); + // And the tear must have actually torn something. Without this the case + // can quietly stop testing what it claims: a cut that lands in a zeroed + // block tail -- or in M25.3's pre-allocated slack -- leaves every record + // whole, so `is_clean` and the durable count both pass while nothing has + // been demonstrated about tolerating a partial record. + // + // The hazard was described in the comment above from the moment the cut + // was rewritten, and describing it did not catch it: reverting that cut to + // the old file-length form left this whole function passing. Measured + // during M25.1b, which is what turned the description into an assertion. + // + // `Truncated` specifically, not merely "stopped": a cut landing past the + // last record reports `NeverWritten`, which is the unwritten extent rather + // than a torn record and would mean the case had stopped tearing. + assert_eq!( + torn.tail_stopped, + Some(record::Torn::Truncated), + "the torn-tail case must actually tear a record, or it demonstrates nothing" + ); + assert!( + torn.tail_records < clean.tail_records, + "tearing the last record must cost a tail record: clean saw {}, torn saw {}", + clean.tail_records, + torn.tail_records + ); // 3. The negative control. A verifier that cannot fail proves nothing, so // corrupt one byte *inside* the durable region and require that replay @@ -747,11 +845,7 @@ fn report_contract(report: &mut Report) { report.line(format_args!("epoch-log durability contract")); report.line(format_args!("==============================")); - for clause in [ - Clause::Guarantees, - Clause::DoesNotGuarantee, - Clause::Assumes, - ] { + for clause in Clause::ALL { report.line(format_args!("")); report.line(format_args!("This log {}:", clause.heading())); for statement in CONTRACT.iter().filter(|s| s.clause == clause) { @@ -787,9 +881,27 @@ fn compare_strategies( " these numbers describe THIS machine and THIS device. They are printed rather than \ quoted in the docs because quoting ours would be misleading." )); + report.line(format_args!( + " 'commit' is what committing costs the strategy: preparing for the flush, submitting \ + it, and waiting for it. The three are shown beside it because which one holds the cost \ + is what tells the strategies apart -- host-sequenced spends it in 'prep', waiting for \ + every write in userspace, where the covering strategies spend it in 'submit'." + )); + report.line(format_args!( + " a zero 'block' beside a large 'deferral' does NOT establish that the operation \ + completed inline: it may equally have pended and then finished while this program was \ + busy elsewhere. The two are indistinguishable from here, and saying so is the point -- \ + reading 'block' alone is how the old single number came to mean something it did not." + )); + report.line(format_args!( + " 'deferral' is NOT part of the commit. It is how long this program went on doing other \ + work before asking, so a design that defers further grows it while being no slower. It \ + is shown because the column here used to be exactly this number, labelled as commit \ + latency (M20.6). Compare rec/s for which strategy to pay for." + )); let payload = b"strategy comparison record payload"; - let mut reference: Option<(&'static str, Vec)> = None; + let mut reference: Option<(&'static str, u32)> = None; let mut throughputs: Vec = Vec::new(); for strategy in strategy::CommitStrategy::ALL { let path = directory.join(format!( @@ -797,13 +909,24 @@ fn compare_strategies( std::process::id(), strategy.name() )); - let file = std::fs::OpenOptions::new() - .create(true) - .write(true) - .truncate(true) - .open(&path)?; + // Pre-allocated and opened NO_BUFFERING | OVERLAPPED, the same shape + // the log itself uses (M25.3) -- the harness exists to measure the log, + // so measuring it through a differently-opened handle would compare the + // strategies on a configuration the log does not run. + let file = logfile::create_preallocated(&path, EPOCHS * PER_EPOCH + SLACK_BLOCKS)?; - let outcome = strategy::run(strategy, file.as_raw_handle(), EPOCHS, PER_EPOCH, payload); + // Each strategy's arena is placed the same way the log's own is, and + // on that strategy's own file -- so the comparison holds placement + // constant instead of adding it to what the strategies differ in. + let placement = placement::Placement::decide(file.as_raw_handle()); + let outcome = strategy::run( + strategy, + file.as_raw_handle(), + EPOCHS, + PER_EPOCH, + payload, + placement.node(), + ); drop(file); let outcome = match outcome { Ok(outcome) => outcome, @@ -815,14 +938,35 @@ fn compare_strategies( // Replayed with the same verifier the log itself uses, because a // strategy that is fast and wrong is not a strategy. + // + // This is the sample's largest allocation: `EPOCHS * PER_EPOCH` + // blocks, so about 8 MiB per strategy (M25.7). It is one buffer at a + // time now rather than two -- the cross-strategy comparison below + // keeps a digest instead of a reference copy -- and it stays whole for + // the reason `replay` gives. A harness that needed to compare logs + // this program had not just written, or logs too large to read, would + // want the streaming verifier described there. let bytes = std::fs::read(&path)?; + // Accounting against the *layout rule*, not against the file's length. + // This compared the two until M25.3, and pre-allocation is what made + // that comparison stop meaning anything: the file now spans its whole + // extent from the moment it is created, whatever the harness went on to + // write into it, so an equality against `bytes.len()` would have held + // just as well for a run that wrote nothing at all. assert_eq!( outcome.bytes as usize, - bytes.len(), - "{} wrote {} bytes but accounted for {}", + outcome.records * record::RECORD_STRIDE, + "{} accounted for {} bytes across {} records, which is not one block each", strategy.name(), - bytes.len(), - outcome.bytes + outcome.bytes, + outcome.records + ); + assert!( + bytes.len() >= outcome.bytes as usize, + "{} wrote {} bytes into an extent of only {}", + strategy.name(), + outcome.bytes, + bytes.len() ); let outcome_replay = replay::replay(&bytes, outcome.durable_through, outcome.records, |index| { @@ -843,31 +987,44 @@ fn compare_strategies( // The cross-strategy invariant, and the one with real teeth. Replay // checks a log against itself; this checks the three strategies // against *each other*, so a dropped record, a wrong offset, or an - // epoch tagged to the wrong commit shows up as a byte difference - // rather than passing three times independently. + // epoch tagged to the wrong commit shows up as a difference rather + // than passing three times independently. + // + // Compared by digest rather than by keeping a reference copy (M25.7). + // The copy was the second of two multi-megabyte buffers alive at once + // -- this loop held the first strategy's whole log for the length of + // the comparison while reading each later one beside it. A digest + // retains thirty-two bytes instead, and loses nothing a reader had: + // the assertion could already only say *that* two logs differed, never + // where. // // What it cannot check is the thing the strategies actually differ // about: whether the ordering held on the *device*. That is only // observable across a power cut, and no in-process check substitutes // for it -- which is why the strategies are argued from D-23 and D-24 // rather than from this run passing. + let digest = record::digest(&bytes); match &reference { - None => reference = Some((strategy.name(), bytes)), + None => reference = Some((strategy.name(), digest)), Some((first, expected)) => assert_eq!( - &bytes, - expected, + digest, + *expected, "{} produced a different log than {first}; all three must write the same bytes", strategy.name() ), } report.line(format_args!( - " {:<18} {:>8.0} rec/s commit p50 {:>7} p99 {:>7} max {:>7} \ + " {:<18} {:>8.0} rec/s commit p50 {:>7} p99 {:>7} \ + (prep {:>7} / submit {:>7} / block {:>7}) deferral p50 {:>7} \ append stall {:>7} -- pays {}", outcome.strategy.name(), outcome.throughput(), - micros(outcome.commit_quantile(0.50)), - micros(outcome.commit_quantile(0.99)), - micros(outcome.commit_quantile(1.0)), + micros(outcome.commit_quantile(CommitTiming::flush, 0.50)), + micros(outcome.commit_quantile(CommitTiming::flush, 0.99)), + micros(outcome.commit_quantile(|t| t.prepare, 0.50)), + micros(outcome.commit_quantile(|t| t.submit, 0.50)), + micros(outcome.commit_quantile(|t| t.blocking, 0.50)), + micros(outcome.commit_quantile(|t| t.deferral, 0.50)), micros(outcome.append_stall), outcome.strategy.cost() )); @@ -879,13 +1036,21 @@ fn compare_strategies( // The spread across strategies is only meaningful next to the spread the // *same* strategy shows between runs, so the program says so instead of // declaring a winner. On the machine this was written on the two are the - // same size, and the reason is visible in the numbers above: every - // strategy pays exactly one device flush per epoch, that flush is hundreds - // of microseconds, and everything the strategies actually differ about -- - // how long the flush itself waits, an extra host round trip -- lands in - // the tens. The - // distinction D-24 draws is real; on this device it is two orders of - // magnitude below the dominant term. + // same size. + // + // The reason given here used to be that every strategy pays one device + // flush per epoch and the things they differ about land two orders of + // magnitude below it. The first half is true. The second was not reachable + // while this sample ran on a synchronous handle: a ring operation completed + // inline during submit, nothing was ever outstanding across a submit + // boundary, and there was no overlap for the strategies to differ in at all + // (M20.6). D-24's distinction was still real; the harness simply could not + // put it under load. + // + // M25.3 has moved every strategy onto a pre-allocated NO_BUFFERING | + // OVERLAPPED file, which removes that cause. Whether the strategies are + // distinguishable *now* is not settled by that and is not claimed here -- + // M25.5 re-runs the comparison and reads it. // // That is not a licence to pick the cheapest-looking one. A device with a // fast flush, a log that commits far more often, or an arena under real @@ -896,8 +1061,11 @@ fn compare_strategies( if low > 0.0 { report.line(format_args!( " spread across strategies: {:.2}x. Run this twice: if the run-to-run spread of one \ - strategy is the same size, the choice is dominated by the device flush that all \ - three pay once per epoch.", + strategy is the same size, the strategies are not distinguishable on this workload. \ + All three pay one device flush per epoch. The second reason they were previously \ + indistinguishable -- a synchronous handle leaving no overlap to differ in -- was \ + removed by M25.3, which put every strategy on a pre-allocated unbuffered overlapped \ + file. Whether that changed this number is what M25.5 reads.", high / low )); } diff --git a/crates/windows-ioring-sys/examples/epoch_log/placement.rs b/crates/windows-ioring-sys/examples/epoch_log/placement.rs new file mode 100644 index 000000000..1492c16fd --- /dev/null +++ b/crates/windows-ioring-sys/examples/epoch_log/placement.rs @@ -0,0 +1,134 @@ +// Copyright (c) 2026 Mike Grier +//! Where the registered arena is placed, and why (M22.3). +//! +//! # The decision +//! +//! The arena is allocated with [`NumaBuffer`](windows_ioring_sys::NumaBuffer) on the NUMA node the **log +//! file's own volume** reports, and with no preference when the volume reports +//! none. The node is asked of the handle the log already holds, through +//! `FSCTL_QUERY_VOLUME_NUMA_INFO`. +//! +//! This module exists because the alternative was silence. The crate's front +//! page tells every consumer that placing the registered pool near the device +//! "is very likely the highest-leverage locality decision available", and this +//! sample previously allocated its arena as a plain `vec![0u8; SLOT_LEN]` -- +//! heap, no alignment, no node. A durability sample is entitled to decide that +//! locality is not its subject; it is not entitled to leave the question +//! looking like an oversight. +//! +//! # What this sample does not claim +//! +//! **That the placement pays here.** It almost certainly does not. The arena +//! is eight slots of four kilobytes, and this workload is bound by a device +//! flush that costs hundreds of microseconds per epoch -- `M22.1` measured +//! that directly, by removing seven of every eight submissions from the append +//! path and watching throughput not move. Thirty-two kilobytes of records +//! crossing an interconnect is not what this program spends its time on, and +//! `examples/ring_copy` is where buffer placement is put under a load that can +//! actually show it. What this sample demonstrates is *how the decision is +//! made and reported*, not that it was worth making. +//! +//! **That the node is the device's.** `FSCTL_QUERY_VOLUME_NUMA_INFO` answers a +//! question about a **volume**, and a volume is not a device: it may span +//! several, as an ordinary spanned volume or a Storage Spaces set does, and +//! what it reports is where the volume resides rather than where this file's +//! extents live. The crate's design notes decline to offer an automatic +//! file-to-node mapping for exactly these reasons, and this module is a +//! sample's local choice rather than a retraction of that. +//! +//! **That the pages landed there.** The allocator's parameter is +//! `nndPreferred`. See [`NumaBuffer`](windows_ioring_sys::NumaBuffer). + +use std::io; +use std::os::windows::io::RawHandle; + +use win_numa_sys::{NumaNode, highest_numa_node, volume_numa_node}; + +/// What the arena's placement was decided to be, kept so the sample can report +/// it rather than making a locality choice silently. +pub enum Placement { + /// The log file's volume named a node, and the arena prefers it. + OnVolumeNode { + /// The node `FSCTL_QUERY_VOLUME_NUMA_INFO` reported. + node: NumaNode, + /// The highest node number this machine reports, when it would say. + /// Kept because `node` alone cannot distinguish a real placement + /// decision from the only answer a single-node machine can give. + highest_node: Option, + }, + /// The volume named no node, so the arena carries no preference -- which + /// is the same allocation the default heap would have made. + Unplaced { + /// Why the query did not answer, so a reader is not left guessing + /// whether the sample simply did not ask. + reason: io::Error, + }, +} + +impl Placement { + /// Ask `handle`'s volume which node it is on. + /// + /// Takes the log's own file handle: the documented FSCTL accepts a file or + /// directory handle directly, so this needs no device-tree walk and no + /// second open. + pub fn decide(handle: RawHandle) -> Self { + match volume_numa_node(handle) { + Ok(node) => Self::OnVolumeNode { + node, + highest_node: highest_numa_node(), + }, + Err(reason) => Self::Unplaced { reason }, + } + } + + /// The node to hand [`NumaBuffer::new`](win_numa_sys::NumaBuffer::new), or + /// `None` for no preference. + pub fn node(&self) -> Option { + match self { + Self::OnVolumeNode { node, .. } => Some(*node), + Self::Unplaced { .. } => None, + } + } + + /// One line for the sample's report, saying what was decided **and** what + /// that is worth on this machine. + /// + /// The second half is the part that matters: on a machine with one node, + /// placing on node 0 and not placing at all are the same allocation, and a + /// line that said only "placed on node 0" would read as a locality win + /// that was never available. + pub fn describe(&self) -> String { + match self { + Self::OnVolumeNode { + node, + highest_node: Some(highest), + } if highest.get() == 0 => format!( + "arena placed on NUMA {node}, which the log file's volume reports; this \ + machine has one node, so that is the only answer available and the placement \ + changes nothing here" + ), + Self::OnVolumeNode { + node, + highest_node: Some(highest), + } => format!( + "arena placed on NUMA {node}, which the log file's volume reports, out of \ + nodes 0..={}", + highest.get() + ), + Self::OnVolumeNode { + node, + highest_node: None, + } => format!( + "arena placed on NUMA {node}, which the log file's volume reports; how many \ + nodes this machine has could not be determined" + ), + Self::Unplaced { reason } => format!( + "arena allocated with no NUMA preference: the log file's volume does not report a \ + node ({reason})" + ), + } + } +} + +#[cfg(test)] +mod tests; diff --git a/crates/windows-ioring-sys/examples/epoch_log/placement/tests.rs b/crates/windows-ioring-sys/examples/epoch_log/placement/tests.rs new file mode 100644 index 000000000..de467deb4 --- /dev/null +++ b/crates/windows-ioring-sys/examples/epoch_log/placement/tests.rs @@ -0,0 +1,186 @@ +// Copyright (c) 2026 Mike Grier +//! Tests for the arena's placement decision (M22.3). +//! +//! # What these can and cannot establish +//! +//! They establish that the decision is *made and reported honestly*: that the +//! FSCTL is actually asked, that a node it reports is the node the arena would +//! be given, and that the report line says what the answer is worth on the +//! machine running it. +//! +//! They establish **nothing about locality**, and cannot. A single-node host +//! has one answer, so no test here can distinguish a good placement from the +//! only placement available; and the allocator's parameter is `nndPreferred`, +//! so even a multi-node host would not prove from a success that the pages +//! landed where they were asked for. Settling that needs hardware this is not +//! developed on, which is a hardware gap rather than a deferred decision. + +use std::os::windows::io::AsRawHandle; + +use win_numa_sys::NumaNode; + +use super::Placement; + +/// A scratch file to ask about, named per test so tests running as threads in +/// one process cannot collide on it. +fn scratch(tag: &str) -> (std::path::PathBuf, std::fs::File) { + let path = std::env::temp_dir().join(format!( + "windows-ioring-sys-epoch-placement-{}-{tag}.tmp", + std::process::id() + )); + let file = std::fs::OpenOptions::new() + .create(true) + .truncate(true) + .write(true) + .open(&path) + .expect("a scratch file in the temp directory"); + (path, file) +} + +#[test] +fn an_ordinary_file_gets_a_decision_either_way() { + let (path, file) = scratch("ordinary"); + let placement = Placement::decide(file.as_raw_handle()); + + // Both arms are legitimate outcomes -- a volume that names no node is the + // documented case, not a failure -- so what is asserted is that the + // decision and its description agree, never which arm was taken. + match &placement { + Placement::OnVolumeNode { .. } => { + assert!(placement.node().is_some(), "a named node is offered"); + } + Placement::Unplaced { .. } => { + assert!(placement.node().is_none(), "no node is offered"); + } + } + + drop(file); + let _ = std::fs::remove_file(path); +} + +#[test] +fn the_description_always_says_which_way_it_went() { + let (path, file) = scratch("describes"); + let placement = Placement::decide(file.as_raw_handle()); + let described = placement.describe(); + + assert!( + described.contains("arena"), + "the line names what was placed: {described}" + ); + match placement.node() { + Some(node) => assert!( + described.contains(&node.to_string()), + "a placed arena names its node: {described}" + ), + None => assert!( + described.contains("no NUMA preference"), + "an unplaced arena says so rather than staying quiet: {described}" + ), + } + + drop(file); + let _ = std::fs::remove_file(path); +} + +#[test] +fn a_single_node_machine_is_told_the_placement_bought_nothing() { + // The honesty guard. On a one-node machine "placed on node 0" is true and + // misleading, so the description must say the choice was not available. + // Constructed rather than queried, so the assertion holds on any host. + let described = Placement::OnVolumeNode { + node: NumaNode::new(0), + highest_node: Some(NumaNode::new(0)), + } + .describe(); + assert!( + described.contains("changes nothing here"), + "a one-node machine must be told so: {described}" + ); +} + +#[test] +fn a_multi_node_machine_is_not_told_that() { + // The other direction: the disclaimer above must not appear where it would + // be false, or it would train a reader to ignore it. + let described = Placement::OnVolumeNode { + node: NumaNode::new(1), + highest_node: Some(NumaNode::new(3)), + } + .describe(); + assert!( + !described.contains("changes nothing here"), + "a multi-node machine must not be told the placement was moot: {described}" + ); + assert!( + described.contains("0..=3"), + "and it says what the choice was made from: {described}" + ); +} + +#[test] +fn an_unknown_node_count_is_admitted_rather_than_assumed() { + let described = Placement::OnVolumeNode { + node: NumaNode::new(2), + highest_node: None, + } + .describe(); + assert!( + described.contains("could not be determined"), + "an unknown node count is said, not guessed: {described}" + ); + assert!( + !described.contains("changes nothing here"), + "and it is not silently treated as a single-node machine: {described}" + ); +} + +#[test] +fn an_unplaced_arena_reports_why() { + // A reader must be able to tell "asked and got no answer" from "never + // asked", which is the distinction the reason carries. + let described = Placement::Unplaced { + reason: std::io::Error::from_raw_os_error(1), + } + .describe(); + assert!(described.contains("does not report a node"), "{described}"); + assert!( + described.len() > "arena allocated with no NUMA preference: ".len() + 40, + "the underlying error is included, not swallowed: {described}" + ); +} + +#[test] +fn an_unplaced_arena_offers_no_node() { + let placement = Placement::Unplaced { + reason: std::io::Error::from_raw_os_error(1), + }; + assert_eq!(placement.node(), None); +} + +#[test] +fn a_placed_arena_offers_the_node_it_named() { + for node in [0_u32, 1, 7, 63].map(NumaNode::new) { + let placement = Placement::OnVolumeNode { + node, + highest_node: Some(NumaNode::new(63)), + }; + assert_eq!( + placement.node(), + Some(node), + "the node that is reported is the node that is used" + ); + } +} + +#[test] +fn an_invalid_handle_is_an_unplaced_arena_not_a_panic() { + // The sample must survive a handle the FSCTL refuses: placement is an + // optimisation, and failing to make it is never a reason to fail the log. + let placement = Placement::decide(std::ptr::null_mut()); + assert!( + matches!(placement, Placement::Unplaced { .. }), + "a refused query is a decision, not an error path" + ); + assert_eq!(placement.node(), None); +} diff --git a/crates/windows-ioring-sys/examples/epoch_log/record.rs b/crates/windows-ioring-sys/examples/epoch_log/record.rs index 57cb99260..c97af512d 100644 --- a/crates/windows-ioring-sys/examples/epoch_log/record.rs +++ b/crates/windows-ioring-sys/examples/epoch_log/record.rs @@ -76,6 +76,69 @@ mod field { /// Total header size. The payload begins here. pub const HEADER_LEN: usize = field::CHECKSUM.end; +/// The largest physical sector size this sample is prepared for. +/// +/// `NO_BUFFERING` (M25.3) requires the file offset *and* the transfer length +/// to be multiples of the volume's physical sector size. 4096 is the largest +/// in common use, and anything that is a multiple of it is a multiple of 512 +/// as well, so a stride satisfying this satisfies every volume the sample can +/// plausibly run on. +const LARGEST_SECTOR_LEN: usize = 4096; + +/// Bytes on disk per record, whatever the record's own length (M25.1). +/// +/// Records are variable-length but land one per fixed-size block, so record +/// *n* occupies `[n * RECORD_STRIDE, n * RECORD_STRIDE + total_len)` and the +/// remainder of its block is zero. Two reasons, and the first is what forced +/// it: +/// +/// 1. **`NO_BUFFERING` constrains transfer lengths, not just offsets.** A +/// packed layout could satisfy the offset rule only by accident and cannot +/// satisfy the length rule at all, since record lengths are whatever a +/// payload makes them. +/// 2. **It is what a real write-ahead log does**, for a related reason: a +/// device's power-fail atomic unit is a sector, so a record sharing a +/// sector with its neighbour can be torn by that neighbour's write. +/// +/// # This lives here, with the format, and not with either writer +/// +/// The stride is a property of the on-disk format, so both writers +/// ([`crate::append::Appender`] and the measurement harness's `Lane`) and the +/// reader ([`crate::replay`]) bind to this one definition. They previously +/// carried a packed layout each, with the *same* justifying comment written +/// out twice -- which is exactly the shape that lets two copies of one rule +/// drift apart while each looks locally correct. +/// +/// # The cost, reported rather than described +/// +/// This is write amplification, and for records as small as this sample's it +/// is large. The sample measures it and prints it -- `layout: N bytes of +/// records in M bytes of file` -- rather than stating a ratio here that would +/// be a hand-maintained copy of a number the program already computes. What +/// is acceptable depends entirely on a caller's record size. A production +/// design amortises it by packing many records into one block and flushing the +/// block once, which is a different sample than this one. What is bought is +/// sector atomicity, which is what a log actually needs. +pub const RECORD_STRIDE: usize = 4096; + +// A stride that is not a whole number of sectors cannot be the length of a +// `NO_BUFFERING` write, so `M25.3` would fail at runtime with +// `ERROR_INVALID_PARAMETER` on a volume whose sectors are this size. A `const` +// assertion refuses it at build time instead, which is the stronger rung: it +// cannot be skipped, and it fails for whoever changes the stride rather than +// for whoever next runs the sample on 4K-native storage. +const _: () = assert!( + RECORD_STRIDE.is_multiple_of(LARGEST_SECTOR_LEN), + "RECORD_STRIDE must be a whole number of sectors, or NO_BUFFERING writes are refused" +); + +// A record has to fit in its own block, header included, or the format cannot +// represent even an empty payload. +const _: () = assert!( + HEADER_LEN < RECORD_STRIDE, + "RECORD_STRIDE must leave room for a record header" +); + /// A record's monotonic identity, assigned when the append is accepted. /// /// Orders records *logically*. The contract is explicit that it says nothing @@ -114,6 +177,28 @@ fn checksum(sequence: Sequence, epoch: Epoch, payload: &[u8]) -> u32 { hash } +/// A digest over a whole log, for comparing two logs without holding both. +/// +/// FNV-1a, the same construction [`checksum`] uses on a record and for the +/// same reason: this is not cryptographic and is not trying to be. Its job is +/// to answer "are these two files the same bytes" for files this program just +/// wrote itself, where the alternative is keeping one of them in memory for +/// the length of the comparison (M25.7). +/// +/// A reader whose logs come from somewhere less trusted wants a hash chosen +/// against an adversary rather than against accident. +pub fn digest(bytes: &[u8]) -> u32 { + const OFFSET_BASIS: u32 = 0x811C_9DC5; + const PRIME: u32 = 0x0100_0193; + + let mut hash = OFFSET_BASIS; + for &byte in bytes { + hash ^= u32::from(byte); + hash = hash.wrapping_mul(PRIME); + } + hash +} + /// Write one record into `slot`, returning how many bytes it occupies. /// /// `slot` is a borrowed view of a registered buffer, which is why this takes a @@ -159,13 +244,60 @@ pub fn encode( Ok(total) } -/// A record recovered from the log, and how many bytes it occupied. +/// Compose a record into `slot` as a whole block, returning the record's own +/// extent (M25.1). +/// +/// [`encode`] lays the record out; this additionally zeroes the rest of its +/// stride, which is what makes the block safe to write whole. Slots are reused +/// for a log's entire life -- a fresh buffer arrives zeroed, but only once -- +/// so without this the bytes past a short record are whatever the previous, +/// longer record left there, and the write puts them on disk. +/// +/// # Why both writers call this rather than each doing the two steps +/// +/// The sample has two writers over this format: [`crate::append::Appender`] +/// and the measurement harness's `Lane`. They previously carried a packed +/// layout each, with the same justifying comment written out twice, and +/// converting them to the stride converted that duplication into a rule +/// stated at two sites. Measured before this function existed: reverting the +/// harness lane to a packed layout was **caught by nothing** -- the end-to-end +/// test covers `Appender`, and only running the sample exercises `Lane`. +/// Giving the composition one site makes half that drift unrepresentable +/// rather than merely tested for. +pub fn encode_block( + slot: &mut [u8], + sequence: Sequence, + epoch: Epoch, + payload: &[u8], +) -> io::Result { + let total = encode(slot, sequence, epoch, payload)?; + slot[total..RECORD_STRIDE].fill(0); + Ok(total) +} + +/// A record recovered from the log. #[derive(Debug)] pub struct Decoded<'a> { pub sequence: Sequence, pub epoch: Epoch, pub payload: &'a [u8], - pub total_len: usize, +} + +impl Decoded<'_> { + /// How many bytes this record occupies, header included. + /// + /// Derived rather than stored. It was a field until M25.2, when replay + /// stopped reading it -- a strided reader advances by + /// [`RECORD_STRIDE`], not by the record's own length, so the only + /// remaining consumer was a test. Keeping it as a field would have kept a + /// second copy of a fact `payload` already carries, which is the shape + /// that lets two statements of one rule drift apart. + /// + /// Note this is the record's **extent**, not its footprint: the rest of + /// its block is zero padding, so `extent() <= RECORD_STRIDE`. + pub fn extent(&self) -> usize { + HEADER_LEN + self.payload.len() + } } /// Why a region of the log did not yield a record. @@ -240,6 +372,8 @@ pub fn decode(bytes: &[u8]) -> Result, Torn> { sequence, epoch, payload, - total_len, }) } + +#[cfg(test)] +mod tests; diff --git a/crates/windows-ioring-sys/examples/epoch_log/record/tests.rs b/crates/windows-ioring-sys/examples/epoch_log/record/tests.rs new file mode 100644 index 000000000..328eec53b --- /dev/null +++ b/crates/windows-ioring-sys/examples/epoch_log/record/tests.rs @@ -0,0 +1,84 @@ +// Copyright (c) 2026 Mike Grier +//! Tests for the record format's digest (M25.7). +//! +//! # Why this exists +//! +//! `compare_strategies` asserted that all three strategies write +//! **byte-identical** logs by keeping one strategy's whole log in memory and +//! comparing the next against it. `M25.7` replaced that with a digest, which +//! retains thirty-two bits instead of eight megabytes -- and which is a +//! **weaker** check, because two different logs can in principle share a +//! digest where two different byte arrays cannot share their bytes. +//! +//! So the weakening needs a guard. These pin that the digest separates the +//! differences the assertion exists to catch: a flipped byte anywhere, and a +//! log with a different number of records in it. +//! +//! # What they cannot establish +//! +//! **That no two logs collide.** FNV-1a over 32 bits has collisions and +//! finding one is not hard for someone trying. The digest's job here is to +//! compare files this program wrote itself moments earlier, where the +//! difference being looked for is a dropped record or a wrong offset rather +//! than an adversary's construction -- which is stated at the definition too. + +use super::{RECORD_STRIDE, digest}; + +#[test] +fn identical_bytes_digest_identically() { + let log = vec![0xAB_u8; RECORD_STRIDE * 3]; + let copy = log.clone(); + assert_eq!( + digest(&log), + digest(©), + "the same bytes must always give the same digest, or the comparison \ + this backs would fail on logs that agree" + ); +} + +#[test] +fn a_single_flipped_byte_changes_the_digest() { + let log = vec![0xAB_u8; RECORD_STRIDE * 3]; + let baseline = digest(&log); + + // First, last, and a block boundary: the positions a walk over strides is + // most likely to treat differently from the bytes around them. + for victim in [0, 1, RECORD_STRIDE - 1, RECORD_STRIDE, log.len() - 1] { + let mut damaged = log.clone(); + damaged[victim] ^= 0xFF; + assert_ne!( + digest(&damaged), + baseline, + "flipping byte {victim} must change the digest, or a log that \ + differs there would compare equal" + ); + } +} + +#[test] +fn a_log_with_fewer_records_digests_differently() { + let log = vec![0xAB_u8; RECORD_STRIDE * 3]; + let short = vec![0xAB_u8; RECORD_STRIDE * 2]; + assert_ne!( + digest(&log), + digest(&short), + "a dropped record must change the digest -- that is the failure the \ + cross-strategy assertion exists to catch" + ); +} + +#[test] +fn a_trailing_zero_block_is_not_invisible() { + // The case a length-oblivious hash would miss: one log ends after its + // records, another has a zeroed block after them. Since M25.3 the log is + // pre-allocated with slack, so a strategy that wrote one record fewer + // leaves exactly this shape rather than a shorter file. + let log = vec![0xAB_u8; RECORD_STRIDE * 2]; + let mut padded = log.clone(); + padded.extend(std::iter::repeat_n(0_u8, RECORD_STRIDE)); + assert_ne!( + digest(&log), + digest(&padded), + "appending a zeroed block must change the digest" + ); +} diff --git a/crates/windows-ioring-sys/examples/epoch_log/replay.rs b/crates/windows-ioring-sys/examples/epoch_log/replay.rs index c165383fb..e61dee3bd 100644 --- a/crates/windows-ioring-sys/examples/epoch_log/replay.rs +++ b/crates/windows-ioring-sys/examples/epoch_log/replay.rs @@ -66,6 +66,14 @@ pub struct Outcome { /// Why decoding stopped, if it stopped before the end of the file. Past /// the watermark this is expected rather than exceptional. pub tail_stopped: Option, + /// Bytes the records themselves occupy, padding excluded. + /// + /// Against the length of the file this is the log's write amplification, + /// which the stride (M25.1) makes large and which the sample reports rather + /// than leaves to a doc comment. A reader who only saw the file size would + /// have no way to tell a log holding a lot of data from one holding very + /// little in a lot of blocks. + pub record_bytes: usize, /// Contract failures. Empty means the log kept every promise it made. pub violations: Vec, } @@ -83,6 +91,35 @@ impl Outcome { /// `expected_durable` how many records it reported durable. `payload_for` /// reproduces what record *n* should contain, so a payload that came back /// altered is caught rather than merely present. +/// +/// # Why this takes a slice and not a reader (M25.7) +/// +/// The walk is strictly forward, one [`record::RECORD_STRIDE`] block at a +/// time, and never looks back -- so it has no need of the whole file at once, +/// and a caller with a log larger than memory cannot give it one. A real log +/// is larger than memory. Reading the whole file is therefore the wrong +/// reflex to teach at exactly the point a reader is learning how to verify +/// one, and the slice is kept anyway, for a reason that is about this +/// function's *vocabulary*: +/// +/// **It returns an [`Outcome`], not a `Result`.** Every way it can end is a +/// statement about the log -- verified, tolerated, or a [`Violation`]. A +/// reader that streams introduces a third kind of ending, `io::Error`, into +/// the one component whose entire job is to distinguish "the log broke its +/// promise" from "the log kept it". Those two failures want different +/// responses from a caller, and a signature that returns both through one +/// channel invites exactly the conflation this file exists to prevent: an +/// unreadable file reported as a missing durable record. +/// +/// So the streaming version is a **different interface**, not a smaller +/// allocation, and this sample keeps the one whose failure vocabulary is +/// closed. The cost is bounded and stated rather than hidden: `main` reads a +/// 140 KiB log here, and its harness reads 8 MiB per strategy. +/// +/// A consumer building a real verifier wants the other shape, and wants +/// `io::Error` and `Violation` kept apart in it -- a `Result` whose +/// `Err` means "could not read" and whose `Ok` still carries every violation +/// found before the read failed. pub fn replay( bytes: &[u8], durable_through: Epoch, @@ -93,6 +130,7 @@ pub fn replay( durable_verified: 0, tail_records: 0, tail_stopped: None, + record_bytes: 0, violations: Vec::new(), }; @@ -101,9 +139,30 @@ pub fn replay( let mut past_watermark = false; while cursor < bytes.len() { - match record::decode(&bytes[cursor..]) { + // Hand `decode` this record's own block, not the rest of the file + // (M25.2). Without the stride there was no block to confine it to, so + // a corrupted `payload_len` was bounded only by the file's length and + // a record could claim bytes belonging to its successors -- caught, + // but by the checksum failing rather than structurally. `min` keeps + // the last block correct when the file ends inside it, which is + // exactly the torn tail the contract requires a reader to tolerate. + let block_end = (cursor + record::RECORD_STRIDE).min(bytes.len()); + match record::decode(&bytes[cursor..block_end]) { Ok(found) => { - cursor += found.total_len; + // Advance by the stride, not by the record's own extent + // (M25.2). Records land one per fixed-size block with a zeroed + // remainder, so a record's extent says where it *ends* and the + // stride says where the next one *starts*; they stopped being + // the same number when the log gained sector atomicity. + // Walking by the extent lands in the zero tail and decodes as + // `NeverWritten`, which replay would report -- correctly, given + // what it was told -- as a durable record having gone missing. + cursor += record::RECORD_STRIDE; + + // Counted for every whole record, durable or tail, because + // amplification is a property of the file rather than of the + // durable region. + outcome.record_bytes += found.extent(); if found.epoch > durable_through { // The tail. Nothing is promised about it, so nothing is diff --git a/crates/windows-ioring-sys/examples/epoch_log/strategy.rs b/crates/windows-ioring-sys/examples/epoch_log/strategy.rs index f03c9cb6a..45cd3c959 100644 --- a/crates/windows-ioring-sys/examples/epoch_log/strategy.rs +++ b/crates/windows-ioring-sys/examples/epoch_log/strategy.rs @@ -65,22 +65,76 @@ //! //! # What the measurement found here, and why it is worth saying //! -//! On the machine this was written on, the three are **indistinguishable**: -//! the throughput spread across strategies is the same size as the spread one -//! strategy shows between consecutive runs. The reason is visible in the -//! numbers the sample prints. Every strategy pays exactly one device flush per -//! epoch, that flush costs hundreds of microseconds, and everything the -//! strategies actually differ about -- how long the flush itself waits, the -//! extra host round trip -- lands in the tens. +//! On the machine this was written on: the spread in throughput and in total +//! commit cost **across** the three strategies is smaller than the spread one +//! strategy shows **between** consecutive runs. What does not vary between +//! runs is *where* each spends its commit -- `HostSequenced` in a host round +//! trip before the flush, the covering strategies inside the submit that +//! carries it. Fifteen runs, with the ranges beside the medians, are in +//! [measurements/2026-09-24-commit-decomposed/](../../measurements/2026-09-24-commit-decomposed/README.md). +//! Read the capture rather than this paragraph; nothing is quoted here that +//! would have to be kept true by hand. //! -//! The distinction [D-24](../../DESIGN-NOTES.md#d-24) draws is real. It is -//! simply two orders of magnitude below the dominant term at this workload, -//! and a reader is better served by knowing that than by a ranking that would -//! not reproduce. +//! The distinction [D-24](../../DESIGN-NOTES.md#d-24) draws is real. Whether +//! it dominates a given workload is a question for that workload's own +//! numbers, which is why this sample prints its own instead of quoting ours. //! -//! Getting that result required fixing the harness twice, which is worth -//! recording because both mistakes are easy to make and neither announces -//! itself: +//! ## Two earlier explanations of that result were wrong (M20.6, M25.5) +//! +//! The spread relation above has held through every correction. The *reasons* +//! given for it did not, twice, and both are recorded because each looked +//! settled: +//! +//! **The first said the strategies differ only in things "two orders of +//! magnitude below" a dominant device flush** -- how long the flush waits, the +//! extra host round trip -- landing "in the tens" of microseconds. `M20.6` +//! found that none of those differences could occur at all: the handle carried +//! no `FILE_FLAG_OVERLAPPED`, so a ring operation completed inline during +//! `SubmitIoRing`, nothing was ever outstanding across a submit boundary, and +//! the comment further down claiming "a real log keeps appending while a commit +//! is outstanding" described something the program could not do. The three were +//! indistinguishable *because they were doing the same serialized work*. Both +//! readings gave the same ranking and only one was true. +//! +//! **The second was the figure itself.** The published commit latency was +//! entirely deferral -- how long the program went on appending before it got +//! around to asking -- so the strategy that defers furthest reported the worst +//! commit while being no slower. +//! +//! `M25.3` gave the handle the shape [the spike](../../design-sessions/spikes/write-pending-spike.rs) +//! measured as pending, `M25.4` split a commit's cost so the flush and the +//! deferral could not be read as each other, and `M25.5` re-ran the comparison +//! on those numbers. The capture linked above also records a correction it had +//! to make before it could answer: the commit clock started *after* +//! `HostSequenced`'s host round trip, which made that strategy look six times +//! cheaper than it is. +//! +//! One figure from the old explanation is now measured and is not what it said: +//! the strategies' differences land in the **hundreds** of microseconds, not +//! the tens. They remain smaller than the run-to-run spread. "Below the +//! run-to-run spread" and "two orders of magnitude below the flush" are +//! different claims, and the measurement supports only the first. +//! +//! **What this does not settle is whether `AlternatingRings` earns its place.** +//! This harness cannot show a blast-radius difference -- but that is a fact +//! about the harness, in which each lane's own arena is the limiter rather +//! than the ring topology, and not evidence that the strategy buys nothing. +//! See [`CommitStrategy::AlternatingRings`] for the conditions under which it +//! would, which include a ring shared with anything else, asymmetric arena +//! sizing, real overlap, and the per-CPU queue affinity +//! [D-27](../../DESIGN-NOTES.md#d-27) is built on. Removing it on the strength +//! of a measurement that could not have shown it working would be foreclosing +//! an option this crate exists to keep open. `M25` rebuilds the harness on a +//! pre-allocated unbuffered log where operations genuinely pend; `M25.5` +//! re-runs this comparison there. +//! +//! **And the point of the comparison is the instrument, not the verdict.** The +//! numbers below describe one machine, one device and one workload. A consumer +//! whose answer differs is not contradicting this sample -- they are the reason +//! it prints its numbers instead of quoting them. +//! +//! Getting that result required fixing the harness three times, which is worth +//! recording because the mistakes are easy to make and none announces itself: //! //! - The first version awaited each commit before appending the next epoch. //! That serialises every strategy, so it measured a workload no real log @@ -90,19 +144,49 @@ //! flush *holds back* those appends. It does not; see D-47. Keeping the //! overlap is still right, but the comparison it produces should be re-read //! with that correction in mind -- see M20.6.) +//! +//! **And `M20.6` found the deeper version of the same error.** Removing the +//! serialisation was necessary and not sufficient: on a synchronous handle +//! there is no overlap to restore, because the work is already done when +//! submit returns. That fix made the harness *able* to overlap while the +//! handle still could not, so what the sentence above describes as the +//! corrected state had never actually run. +//! +//! **`M25.3` gave the handle the shape that can overlap, and `M25.5` +//! measured what followed.** Every strategy's `block` reads zero at the +//! median in every run -- so the harness still never waits for a flush. That +//! is not evidence the operation completed inline: it defers by several +//! milliseconds, and an operation that pended and finished during that window +//! is indistinguishable from one that never pended. **The sentence above is +//! now describable rather than demonstrated**, which is a smaller claim than +//! it has ever carried before, and the honest one. //! - The second version keyed pending commits by `UserData` in one map across //! both rings. Each ring assigns its own sequence, so the two collided and //! half the samples vanished. +//! - The third submitted **one write per record**, which made the per-record +//! submission cost a term every strategy paid equally -- exactly the shape of +//! shared constant that can flatten a comparison into "indistinguishable" +//! without the underlying claim being true. Both append paths now batch, and +//! the confound was measured rather than argued away: twenty runs, ten each +//! side, in +//! [measurements/2026-09-22-append-batching/](../../measurements/2026-09-22-append-batching/). +//! **The conclusion above survives it** -- removing the shared cost left the +//! cross-strategy spread inside a single strategy's own run-to-run range. +//! What batching did move is commit latency, which is the half of the system +//! the device flush does not dominate. use std::io; use std::os::windows::io::RawHandle; use std::time::{Duration, Instant}; use windows_ioring_sys::{ - Batch, FlushCoverage, FlushMode, IoRing, PushOptions, RegisteredBuffers, RegisteredSpan, - RegisteredUse, Token, WriteCaching, + Batch, Completion, FlushCoverage, FlushMode, IoRing, NumaBuffer, PushOptions, + RegisteredBuffers, RegisteredSpan, RegisteredUse, Token, WriteCaching, }; +use win_numa_sys::NumaNode; + +use crate::append::free_slots; use crate::commit::Epoch; use crate::record::{self, Sequence}; @@ -116,6 +200,10 @@ const SLOT_LEN: usize = 4096; /// Bound on any wait, so a stuck strategy fails instead of hanging. const WAIT_MS: u32 = 30_000; +/// The same bound as a [`Duration`], derived from `WAIT_MS` rather than +/// written twice so the two cannot drift. +const WAIT: Duration = Duration::from_millis(WAIT_MS as u64); + /// How an epoch's commit establishes that its writes reached the device. #[derive(Clone, Copy, Debug, PartialEq, Eq)] pub enum CommitStrategy { @@ -129,6 +217,51 @@ pub enum CommitStrategy { /// Two rings, epochs alternating between them, each committed with a /// covering flush on its own ring, so the appending ring is never the one /// waiting, at the cost of registering the arena twice. + /// + /// # What this harness can and cannot show about it (M20.6) + /// + /// The argument for two rings was that a covering flush reaches *every* + /// operation outstanding on its ring ([D-47](../../DESIGN-NOTES.md#d-47) + /// withdrew the hold-back half and kept this one), so alternating bounds + /// what a commit's barrier can be dragged into. + /// + /// **This sample cannot exhibit that difference, and the reason is the + /// sample's own shape rather than anything about the strategy.** + /// `RegisteredBuffers::get_mut` refuses a slot with an operation + /// outstanding, and there are `SLOTS` slots, so at most `SLOTS` appends + /// are outstanding on a ring by construction -- and each alternating lane + /// registers its own arena of the same size. Here the arena is the + /// limiter, not the ring topology, so the covered count is identical + /// either way. (Probed: 8 and 8.) + /// + /// That is a statement about this apparatus. It is **not** evidence that + /// alternating rings buys nothing, and the conditions under which it would + /// are ordinary rather than exotic: + /// + /// - **A ring shared with anything else.** This sample owns its ring + /// entirely. A consumer whose ring also carries another component's + /// traffic has a barrier whose reach is bounded by that traffic, not by + /// this arena. + /// - **Arenas sized differently from the lanes.** One large arena on a + /// shared ring against two small ones is a different bound, and nothing + /// makes the sample's symmetric choice the general case. + /// - **Real overlap.** With operations that genuinely pend, a shared + /// ring's flush can be reached *after* the next epoch's appends are + /// queued, so the same covered count is not the same wait. This could not + /// happen at all while the handle was synchronous; `M25.3` removed that + /// obstacle, and whether operations actually pend remains an observation + /// about a machine rather than something this crate may assume. + /// - **Per-CPU queue affinity.** [D-27](../../DESIGN-NOTES.md#d-27) is + /// this crate's decision that one ring per thread is userspace's proxy + /// for one ring per CPU, and records that NVMe queue pairs are per-CPU + /// with their completion interrupt routed by their own vector. Two rings + /// on two pinned threads is that architecture; one ring is not. + /// + /// So this strategy stays, and the sample's job is to let a consumer find + /// out **on their own hardware and workload** rather than to hand them a + /// verdict from ours. `M25.5` re-runs the comparison on a harness where + /// operations genuinely pend, which removes the third condition above and + /// makes the answer here mean more than it currently can. AlternatingRings, } @@ -178,20 +311,19 @@ pub struct Outcome { pub bytes: u64, /// Wall clock for the whole run, appends and commits together. pub elapsed: Duration, - /// Per epoch, from pushing the commit to observing its completion. + /// Per epoch, what its commit cost, decomposed (M25.4). /// - /// Read this carefully, because it is **not** device flush time. Every - /// strategy here defers its await by design, so the figure includes time - /// the program spent doing useful work before it got around to asking -- - /// and a strategy that defers *further* therefore reports a *larger* - /// number while being no slower. Alternating rings shows this most - /// plainly: it defers across two epochs and reports the highest latency of - /// the three while matching them on throughput. + /// This was a single `Duration` measured from pushing the flush to + /// observing its completion, and `M20.6` established that the number was + /// **entirely deferral**: blocking p50 *and* p99 were 0 us for all three + /// strategies, so the published figure reported how long the next epoch's + /// appends took rather than anything about the commit. A strategy that + /// deferred further reported a larger number while being no slower. /// - /// So compare [`Outcome::elapsed`] across strategies, and read this as - /// "how stale is a commit acknowledgement by the time this design collects - /// it" -- which is a real property, just not the one its name suggests. - pub commit_latencies: Vec, + /// Splitting it is the fix, and the split is the point: the three parts + /// cannot be confused for one another the way one blended number invited. + /// See [`CommitTiming`]. + pub commit_timings: Vec, /// Time appends spent blocked because every arena slot was busy. /// /// This is where a long commit shows up as a number: an arena slot is not @@ -203,6 +335,85 @@ pub struct Outcome { pub append_stall: Duration, } +/// What one epoch's commit cost, split so the parts cannot be read as each +/// other (M25.4). +/// +/// The harness published a single blended number until `M20.6` decomposed it +/// and found it was entirely deferral. These three add up to that old number +/// and are reported separately for that reason. +/// +/// # Every part stays meaningful whether or not the operation pends +/// +/// `M25`'s standing constraint forbids anything here depending on an operation +/// pending, because Windows specifies nothing about when a ring operation +/// completes relative to `SubmitIoRing`. This split satisfies it by +/// construction rather than by assumption: +/// +/// - if the flush completes **inline**, the device round trip lands in +/// [`submit`](Self::submit) and [`blocking`](Self::blocking) is zero; +/// - if it **pends**, `submit` is short and the wait shows up in `blocking`; +/// - either way [`deferral`](Self::deferral) is the program's own choice and +/// belongs to neither. +/// +/// So the harness reports what it observed and never has to know which case it +/// got. +#[derive(Clone, Copy, Debug)] +pub struct CommitTiming { + /// Wall time the strategy spent preparing before the flush could be + /// pushed at all. + /// + /// Zero for the covering strategies, which push the flush immediately and + /// let its coverage do the ordering. For `HostSequenced` it is the host + /// round trip -- waiting for every write's completion in userspace, which + /// is what makes an unordered flush sufficient for it. + /// + /// **This part exists because leaving it out inverted the comparison.** + /// `M25.4` started the clock at the submit, which put that round trip + /// outside every measured part; `HostSequenced` then reported a flush + /// roughly six times cheaper than the other two while doing the same work + /// in a place nothing was looking. The cost had not gone anywhere, and a + /// reader comparing the published numbers would have concluded the + /// opposite of the truth. Found by `M25.5` while reading the very figures + /// `M25.4` produced. + pub prepare: Duration, + /// Wall time inside the call that builds and submits the flush. + /// + /// On a handle where the operation completes inline, this **is** the + /// flush: the device round trip happens inside `SubmitIoRing`. + pub submit: Duration, + /// From the submit returning to the harness asking for the completion. + /// + /// **Not a cost of the flush.** It is work the program chose to do first + /// -- here, pushing the next epoch's appends -- and a design that defers + /// further grows this number while being no slower. It is kept because the + /// figure this harness used to publish as commit latency was exactly this, + /// and a number that was once mistaken for another is worth showing beside + /// the one it was mistaken for. + pub deferral: Duration, + /// Wall time actually spent waiting for the flush's completion. + /// + /// Zero when the completion was already queued by the time the harness + /// looked. **That happens for two different reasons and this number cannot + /// tell them apart**: the operation may have completed inline during the + /// submit, or it may have pended and then finished during the deferral. + /// Reading a zero here as evidence of inline completion is the same error + /// in the opposite direction as the one `M20.6` found, which is why this + /// is never reported without `deferral` beside it. + pub blocking: Duration, +} + +impl CommitTiming { + /// What the commit itself cost: preparing for it, submitting it, and + /// waiting for it. + /// + /// Excludes [`deferral`](Self::deferral), which is the program's and not + /// the commit's. This is the figure `M25.4` asked for, with the + /// [`prepare`](Self::prepare) term `M25.5` found it was missing. + pub fn flush(&self) -> Duration { + self.prepare + self.submit + self.blocking + } +} + impl Outcome { /// Records per second over the whole run. pub fn throughput(&self) -> f64 { @@ -212,21 +423,62 @@ impl Outcome { self.records as f64 / self.elapsed.as_secs_f64() } - /// Commit latency at `fraction` through the sorted distribution. + /// The quantile at `fraction` of whichever part of a commit `part` + /// selects. /// /// Nearest-rank rather than interpolated: with tens of samples an /// interpolated quantile invents precision the data does not have. - pub fn commit_quantile(&self, fraction: f64) -> Duration { - if self.commit_latencies.is_empty() { + /// + /// Taking a projection rather than offering one method per part is what + /// keeps the four call sites from drifting -- each asks the same question + /// of a different field instead of each carrying its own copy of the + /// sort-and-rank. + pub fn commit_quantile( + &self, + part: impl Fn(&CommitTiming) -> Duration, + fraction: f64, + ) -> Duration { + if self.commit_timings.is_empty() { return Duration::ZERO; } - let mut sorted = self.commit_latencies.clone(); + let mut sorted: Vec = self.commit_timings.iter().map(part).collect(); sorted.sort_unstable(); let rank = ((sorted.len() as f64 - 1.0) * fraction).round() as usize; sorted[rank.min(sorted.len() - 1)] } } +/// Wait for one lane's outstanding commit and record what it cost (M25.4). +/// +/// One function rather than the same four lines at each of the two settle +/// sites -- the loop's, and the drain after it -- because the rule being +/// applied is *where the deferral ends and the blocking begins*, and a rule +/// stated twice is a rule that can be half-corrected. The split has to happen +/// here, at the moment the harness decides to ask, which is precisely what a +/// single `elapsed()` at the end could not distinguish. +/// +/// A lane with nothing outstanding is not an error: the first epoch on each +/// lane has no previous commit to settle. +fn settle( + lane: &mut Lane, + deferred: &mut Option<(usize, Duration, Duration, Instant)>, + timings: &mut Vec, +) -> io::Result<()> { + let Some((user_data, prepare, submit, submitted_at)) = deferred.take() else { + return Ok(()); + }; + let deferral = submitted_at.elapsed(); + let blocked = Instant::now(); + lane.await_flush(user_data)?; + timings.push(CommitTiming { + prepare, + submit, + deferral, + blocking: blocked.elapsed(), + }); + Ok(()) +} + /// One ring plus the arena registered on it. /// /// The arena is a separate registration per ring, which is the doubled cost @@ -234,21 +486,22 @@ impl Outcome { /// permanently, because `IoRing` has no unregister call. struct Lane { ring: IoRing, - arena: RegisteredBuffers>, + arena: RegisteredBuffers, /// `UserData` of an in-flight write -> its token and the slot it reads /// from. The token must be *claimed* on completion: dropping it unclaimed /// is treated as still-outstanding and leaks the slot forever. in_flight: std::collections::HashMap, u32)>, - free: Vec, /// Completions that were not writes, kept by `UserData` for the commit /// path to match against. flushes: std::collections::HashMap>, } impl Lane { - fn new() -> io::Result { + fn new(node: Option) -> io::Result { let mut ring = IoRing::new(64, 128)?; - let buffers: Vec> = (0..SLOTS).map(|_| vec![0u8; SLOT_LEN]).collect(); + let buffers = (0..SLOTS) + .map(|_| NumaBuffer::new(SLOT_LEN, node)) + .collect::>>()?; let mut batch = Batch::new(&mut ring); let pending = batch.register_buffers(buffers)?; batch.submit_and_wait(1, WAIT_MS)?; @@ -262,7 +515,6 @@ impl Lane { ring, arena, in_flight: std::collections::HashMap::new(), - free: (0..SLOTS).collect(), flushes: std::collections::HashMap::new(), }) } @@ -275,56 +527,78 @@ impl Lane { self.in_flight.len() } - /// Compose one record into a free slot and push its write. + /// Compose as many records as there are free slots and push them all in + /// **one** submission, starting at `sequence` and `offset`. + /// + /// Returns how many were accepted and how many bytes they occupy. Zero + /// accepted is the caller's cue to drain, not an error. + /// + /// Batched for the reason [`crate::append::Appender::append_batch`] is: + /// one `SubmitIoRing` per record made the per-record submission cost a + /// term every strategy paid equally, which is exactly the kind of shared + /// constant that flattens a comparison. /// - /// Returns `None` when every slot is busy, which is the caller's cue to - /// drain. - fn append( + /// Which slots are free is asked of [`free_slots`] -- the sample's single + /// definition -- rather than tracked here. This lane used to keep its own + /// free list; see that function for why the copy was not merely redundant. + fn append_batch( &mut self, file: RawHandle, - sequence: Sequence, + first_sequence: u64, epoch: Epoch, payload: &[u8], - offset: u64, - ) -> io::Result> { - let Some(slot) = self.free.pop() else { - return Ok(None); - }; - let total = { - let bytes = self.arena.get_mut(slot)?; - record::encode(bytes, sequence, epoch, payload)? - }; - // Exactly the bytes the record occupies, not the whole slot: writing - // the slot's unused tail would put stale bytes in the log and cost - // real device bandwidth. - let span = RegisteredSpan { - buffer_index: slot, - offset: 0, - len: u32::try_from(total).map_err(|_| { - io::Error::new( - io::ErrorKind::InvalidInput, - "record length exceeds u32::MAX", - ) - })?, - }; + first_offset: u64, + want: usize, + ) -> io::Result<(usize, u64)> { + let slots = free_slots(&self.arena, want); + if slots.is_empty() { + return Ok((0, 0)); + } + let mut batch = Batch::new(&mut self.ring); - // SAFETY: `file` outlives every operation pushed here -- the caller - // drains to empty before closing it -- and the token is held in - // `in_flight` until its completion is observed, so the slot cannot be - // refilled underneath the kernel. - let token = unsafe { - batch.write_registered_raw( - file, - &self.arena, - span, - offset, - PushOptions::new(), - WriteCaching::Cached, - ) - }?; + let mut written = 0_u64; + let mut accepted = 0; + for (index, &slot) in slots.iter().enumerate() { + let sequence = Sequence(first_sequence + index as u64); + { + let bytes = self.arena.get_mut(slot)?; + // The same composition the log's own appender uses, from one + // definition: encode, then zero the rest of the block. This + // lane is a second writer over the same format, and it + // previously carried its own packed layout *and* its own copy + // of the justification for it. + record::encode_block(bytes, sequence, epoch, payload)?; + } + let span = RegisteredSpan { + buffer_index: slot, + offset: 0, + len: u32::try_from(record::RECORD_STRIDE).map_err(|_| { + io::Error::new( + io::ErrorKind::InvalidInput, + "record stride exceeds u32::MAX", + ) + })?, + }; + // SAFETY: `file` outlives every operation pushed here -- the + // caller drains to empty before closing it -- and the token is + // held in `in_flight` until its completion is observed, so the + // slot cannot be refilled underneath the kernel. + let token = unsafe { + batch.write_registered_raw( + file, + &self.arena, + span, + first_offset + written, + PushOptions::new(), + WriteCaching::Cached, + ) + }?; + self.in_flight.insert(token.id(), (token, slot)); + written += record::RECORD_STRIDE as u64; + accepted += 1; + } batch.submit()?; - self.in_flight.insert(token.id(), (token, slot)); - Ok(Some(total as u64)) + Ok((accepted, written)) } /// Push this lane's commit flush and return its `UserData`. @@ -344,33 +618,64 @@ impl Lane { let mut popped = 0; while let Some(completion) = self.ring.try_pop()? { popped += 1; - if let Some((token, slot)) = self.in_flight.remove(&completion.user_data()) { - // Claimed before the result is checked, for the reason - // `Appender::claim` spells out: bailing out first would drop - // the token unclaimed and burn the slot permanently. - let released = token - .claim_if(&completion) - .map_err(|_| io::Error::other("a write token refused its own completion"))?; - drop(released); - self.free.push(slot); - completion.result()?; - } else { - self.flushes - .insert(completion.user_data(), completion.result().map(|_| ())); - } + self.classify(completion)?; } Ok(popped) } + /// File one popped completion: an append returns its slot, anything else + /// is a flush the commit path will match against. + /// + /// Factored out of [`Lane::drain`] so the bounded waits below can file a + /// completion they blocked for without a second copy of this logic. + fn classify(&mut self, completion: Completion) -> io::Result<()> { + if let Some((token, slot)) = self.in_flight.remove(&completion.user_data()) { + // Claimed before the result is checked, for the reason + // `Appender::claim` spells out: bailing out first would drop + // the token unclaimed and burn the slot permanently. + let released = token + .claim_if(&completion) + .map_err(|_| io::Error::other("a write token refused its own completion"))?; + // Dropping the marker is what decrements the slot's outstanding + // count, and that count *is* the free list now -- so this drop, + // not a push to a side table, is what returns the slot. + drop(released); + debug_assert!( + self.arena.outstanding(slot) == Some(0), + "claiming the token must release the slot" + ); + completion.result()?; + } else { + self.flushes + .insert(completion.user_data(), completion.result().map(|_| ())); + } + Ok(()) + } + /// Block until `user_data`'s flush has completed, draining as we go. + /// + /// # Errors + /// + /// [`io::ErrorKind::TimedOut`] if the flush does not complete within + /// [`WAIT`]. This loop used to have no bound at all: it blocked in + /// `submit_and_wait` for `WAIT_MS`, ignored the fact that the call had + /// returned without a completion, and went round again forever. A + /// measurement harness that hangs reports nothing, which is strictly worse + /// than one that fails -- the same reason + /// [`EventLoop::pump`](crate::event_loop::EventLoop::pump) raises + /// `TimedOut` rather than spinning. fn await_flush(&mut self, user_data: usize) -> io::Result<()> { + let deadline = Instant::now() + WAIT; loop { if let Some(result) = self.flushes.remove(&user_data) { result?; return Ok(()); } if self.drain()? == 0 { - Batch::new(&mut self.ring).submit_and_wait(1, WAIT_MS)?; + match self.ring.pop_within(remaining(deadline))? { + Some(completion) => self.classify(completion)?, + None => return Err(timed_out("a commit's flush")), + } } } } @@ -379,16 +684,39 @@ impl Lane { /// /// This is [`CommitStrategy::HostSequenced`]'s whole mechanism, and the /// round trip it is charged for. + /// + /// # Errors + /// + /// [`io::ErrorKind::TimedOut`] if they do not all complete within + /// [`WAIT`], for the reason given on [`Lane::await_flush`]. fn await_writes(&mut self) -> io::Result<()> { + let deadline = Instant::now() + WAIT; while !self.in_flight.is_empty() { if self.drain()? == 0 { - Batch::new(&mut self.ring).submit_and_wait(1, WAIT_MS)?; + match self.ring.pop_within(remaining(deadline))? { + Some(completion) => self.classify(completion)?, + None => return Err(timed_out("this lane's outstanding writes")), + } } } Ok(()) } } +/// How long is left before `deadline`, saturating at zero. +fn remaining(deadline: Instant) -> Duration { + deadline.saturating_duration_since(Instant::now()) +} + +/// The one spelling of this harness's timeout failure, so the two waits above +/// cannot describe the same condition differently. +fn timed_out(what: &str) -> io::Error { + io::Error::new( + io::ErrorKind::TimedOut, + format!("timed out after {WAIT:?} waiting for {what}"), + ) +} + /// Run `epochs` epochs of `records_per_epoch` records under `strategy`. /// /// Every epoch is committed and its commit awaited, so the returned outcome's @@ -396,25 +724,31 @@ impl Lane { /// produce the same log: same records, same order, same bytes -- which is the /// point. They differ only in what the commit costs. /// +/// Every lane's arena is placed on `node`, the same node the log's own arena +/// uses, so placement is held constant across the comparison rather than being +/// one more thing the strategies differ in. +/// /// # Errors /// -/// Any error from ring setup, a push, a submit, or an operation's result. +/// Any error from ring setup, arena allocation, a push, a submit, or an +/// operation's result. pub fn run( strategy: CommitStrategy, file: RawHandle, epochs: usize, records_per_epoch: usize, payload: &[u8], + node: Option, ) -> io::Result { let mut lanes = Vec::with_capacity(strategy.rings()); for _ in 0..strategy.rings() { - lanes.push(Lane::new()?); + lanes.push(Lane::new(node)?); } let mut offset = 0u64; let mut sequence = 0u64; let mut records = 0usize; - let mut commit_latencies = Vec::with_capacity(epochs); + let mut commit_timings = Vec::with_capacity(epochs); let mut append_stall = Duration::ZERO; // Registration is deliberately outside the clock: it happens once, and // charging a strategy's throughput for it would say more about setup than @@ -442,7 +776,7 @@ pub fn run( // `UserData`. Each ring assigns its own `UserData` sequence, so the two // lanes hand out colliding values and a single map silently loses half the // samples -- which is exactly what the second attempt at this measured. - let mut deferred: Vec> = vec![None; lanes.len()]; + let mut deferred: Vec> = vec![None; lanes.len()]; for epoch in 0..epochs as u64 { let lane_index = if strategy == CommitStrategy::AlternatingRings { @@ -451,30 +785,34 @@ pub fn run( 0 }; - for _ in 0..records_per_epoch { - loop { - let lane = &mut lanes[lane_index]; - match lane.append(file, Sequence(sequence), Epoch(epoch), payload, offset)? { - Some(written) => { - offset += written; - sequence += 1; - records += 1; - break; - } - // Every slot is busy. Drain and retry -- the arena - // working as intended, and the pressure a long commit - // makes worse by holding its slots for the whole of its - // own duration. Timed, because this is the cost a - // covering flush imposes on the append path. - None => { - let blocked = Instant::now(); - if lane.drain()? == 0 { - Batch::new(&mut lane.ring).submit_and_wait(1, WAIT_MS)?; - } - append_stall += blocked.elapsed(); - } + let mut pushed_this_epoch = 0; + while pushed_this_epoch < records_per_epoch { + let lane = &mut lanes[lane_index]; + let (accepted, written) = lane.append_batch( + file, + sequence, + Epoch(epoch), + payload, + offset, + records_per_epoch - pushed_this_epoch, + )?; + if accepted == 0 { + // Every slot is busy. Drain and retry -- the arena working as + // intended, and the pressure a long commit makes worse by + // holding its slots for the whole of its own duration. Timed, + // because this is the cost a covering flush imposes on the + // append path. + let blocked = Instant::now(); + if lane.drain()? == 0 { + Batch::new(&mut lane.ring).submit_and_wait(1, WAIT_MS)?; } + append_stall += blocked.elapsed(); + continue; } + offset += written; + sequence += accepted as u64; + records += accepted; + pushed_this_epoch += accepted; } // Settle this lane's previous commit *here*, after its next epoch's @@ -492,11 +830,17 @@ pub fn run( // exposes is the strategies' other costs rather than a stall. The // placement stays; the conclusion drawn from the numbers needs // re-reading, which is `M20.6`. - if let Some((user_data, pushed)) = deferred[lane_index].take() { - lanes[lane_index].await_flush(user_data)?; - commit_latencies.push(pushed.elapsed()); - } + settle( + &mut lanes[lane_index], + &mut deferred[lane_index], + &mut commit_timings, + )?; + // Timed from here, not from the submit. What a strategy must do + // *before* it can push its flush is part of what committing costs it + // -- and leaving it out is what made `HostSequenced` look six times + // cheaper than the others in M25.4's first numbers. + let preparing = Instant::now(); let coverage = match strategy { // The round trip: every write observed complete *before* the flush // is even pushed. That is what makes an unordered flush sufficient @@ -509,17 +853,23 @@ pub fn run( FlushCoverage::CoversPrecedingOperations } }; + let prepare = preparing.elapsed(); + // The submit itself, which is the part of a commit's cost that exists + // whether or not the operation pends: on a handle that completes + // inline, the device round trip happens inside this call. + let submitting = Instant::now(); let user_data = lanes[lane_index].commit(file, coverage)?; - deferred[lane_index] = Some((user_data, Instant::now())); + deferred[lane_index] = Some((user_data, prepare, submitting.elapsed(), Instant::now())); } // Settle the last outstanding commit, or `durable_through` below would be // a claim rather than an observation. for lane_index in 0..lanes.len() { - if let Some((user_data, pushed)) = deferred[lane_index].take() { - lanes[lane_index].await_flush(user_data)?; - commit_latencies.push(pushed.elapsed()); - } + settle( + &mut lanes[lane_index], + &mut deferred[lane_index], + &mut commit_timings, + )?; } // One observation per epoch. Cheap, and it binds a real invariant -- but be // precise about which one, because this comment previously overclaimed. @@ -546,11 +896,11 @@ pub fn run( // empties the slot whether or not the `if let` binds -- so it was removed // rather than left to look like a guard. assert_eq!( - commit_latencies.len(), + commit_timings.len(), epochs, "{} observed {} of {epochs} commits", strategy.name(), - commit_latencies.len() + commit_timings.len() ); for lane in &mut lanes { lane.await_writes()?; @@ -563,7 +913,10 @@ pub fn run( durable_through: Epoch(epochs as u64 - 1), bytes: offset, elapsed, - commit_latencies, + commit_timings, append_stall, }) } + +#[cfg(test)] +mod tests; diff --git a/crates/windows-ioring-sys/examples/epoch_log/strategy/tests.rs b/crates/windows-ioring-sys/examples/epoch_log/strategy/tests.rs new file mode 100644 index 000000000..ca724fcd7 --- /dev/null +++ b/crates/windows-ioring-sys/examples/epoch_log/strategy/tests.rs @@ -0,0 +1,367 @@ +// Copyright (c) 2026 Mike Grier +//! Tests for the measurement harness's lane (M22.2). +//! +//! These exist because collapsing the sample's two free-slot definitions into +//! [`crate::append::free_slots`] did not only remove a duplicate -- it removed +//! a slot leak the duplicate had and nothing reported. A tracked free list +//! took its slot *before* composing into it, so an append refused between the +//! two -- a record too long for a slot is the reachable one -- returned with +//! the slot removed from the list and no operation ever issued. +//! +//! That leak was verified by re-injecting the free list, not argued from the +//! source: eight refused appends left the lane reporting **0** of 8 slots free +//! while the arena held nothing. +//! +//! # What this file does and does not catch +//! +//! [`a_record_too_long_for_a_slot_leaves_its_slot_free`] asserts the property +//! from the arena's side, so it would catch an append path that marks a slot +//! busy before the record is known to fit. It would **not** catch a +//! re-introduced free list, because it never asks one -- under the re-injection +//! above it passed, and only a temporary assertion against the list itself went +//! red. What keeps a second definition from coming back is that there is one +//! function and both callers call it, which the compiler enforces and a grep +//! can confirm; this file cannot, and saying otherwise would be the sort of +//! cosmetic binding these tests exist to avoid. +//! +//! A `Lane` owns a real ring, so by this crate's Quality rule these are not +//! hermetic. They are example-target tests rather than lib tests, so they do +//! not add to the pile [D-49](../../../DESIGN-NOTES.md#d-49) queues for `M24`; +//! and the ring is incidental here, because nothing below ever submits. + +use std::os::windows::io::AsRawHandle; + +use super::{CommitStrategy, Lane, SLOT_LEN, SLOTS}; +use crate::append::free_slots; +use crate::commit::Epoch; +use crate::record::HEADER_LEN; + +/// A scratch file to append against, named per test so tests running as +/// threads in one process cannot collide on it. +fn scratch(tag: &str) -> (std::path::PathBuf, std::fs::File) { + let path = std::env::temp_dir().join(format!( + "windows-ioring-sys-epoch-strategy-{}-{tag}.tmp", + std::process::id() + )); + let file = std::fs::OpenOptions::new() + .create(true) + .truncate(true) + .read(true) + .write(true) + .open(&path) + .expect("a scratch file in the temp directory"); + (path, file) +} + +/// How many of this lane's slots are free, asked the same way the append path +/// asks it. +fn free(lane: &Lane) -> usize { + free_slots(&lane.arena, SLOTS as usize).len() +} + +/// A payload one byte too long to fit a slot once the header is accounted for. +/// +/// Derived from the two constants rather than written as a number, so a change +/// to either cannot leave this test passing for the wrong reason. +fn oversized() -> Vec { + vec![0xAB; SLOT_LEN - HEADER_LEN + 1] +} + +#[test] +fn a_fresh_lane_has_every_slot_free() { + let lane = Lane::new(None).expect("a ring and a registered arena"); + assert_eq!( + free(&lane), + SLOTS as usize, + "a lane that has pushed nothing must offer every slot" + ); +} + +#[test] +fn a_record_too_long_for_a_slot_is_refused() { + let mut lane = Lane::new(None).expect("a ring and a registered arena"); + let (path, file) = scratch("too-long-refused"); + + let error = lane + .append_batch(file.as_raw_handle(), 0, Epoch(0), &oversized(), 0, 1) + .expect_err("a record that cannot fit a slot must not be accepted"); + + assert_eq!( + error.kind(), + std::io::ErrorKind::InvalidInput, + "the refusal is the encoder's, and it reports a bad input rather than a busy arena" + ); + + drop(file); + let _ = std::fs::remove_file(path); +} + +#[test] +fn a_record_too_long_for_a_slot_leaves_its_slot_free() { + let mut lane = Lane::new(None).expect("a ring and a registered arena"); + let (path, file) = scratch("too-long-no-leak"); + assert_eq!(free(&lane), SLOTS as usize, "precondition: nothing is busy"); + + // Fail an append `SLOTS` times over. Every one is refused before any + // operation is pushed, so the arena is untouched by all of them. + for _ in 0..SLOTS { + lane.append_batch(file.as_raw_handle(), 0, Epoch(0), &oversized(), 0, 1) + .expect_err("a record that cannot fit a slot must not be accepted"); + } + + // The regression this test exists for. A tracked free list took its slot + // *before* composing into it, so each failure above consumed one and never + // gave it back: this lane would report zero free slots while the kernel + // held nothing, and every later append would be told to drain an arena + // that was entirely idle. Deriving the answer from the arena's own + // outstanding counts cannot express that -- a slot nothing was pushed + // against never stopped being free. + assert_eq!( + free(&lane), + SLOTS as usize, + "a refused append must not consume a slot" + ); + + drop(file); + let _ = std::fs::remove_file(path); +} + +#[test] +fn a_lane_offers_at_most_what_was_asked_for() { + let lane = Lane::new(None).expect("a ring and a registered arena"); + assert_eq!(free_slots(&lane.arena, 3).len(), 3, "capped by the request"); + assert_eq!( + free_slots(&lane.arena, 0).len(), + 0, + "asking for none is answerable without consulting the arena" + ); + assert_eq!( + free_slots(&lane.arena, SLOTS as usize * 4).len(), + SLOTS as usize, + "and capped by the arena when the request exceeds it" + ); +} + +/// The host round trip is inside the commit's measured cost (M25.5). +/// +/// `HostSequenced` waits for every write in userspace before pushing an +/// unordered flush. `M25.4` started the commit clock at the submit, which put +/// that wait outside every measured part -- so the strategy reported a commit +/// roughly six times cheaper than the covering ones while doing the same work +/// somewhere nothing was looking, and a reader comparing the published figures +/// would have drawn the opposite of the right conclusion. +/// +/// The guard is that the round trip shows up at all. It asserts a non-zero +/// `prepare` rather than any relationship between the strategies' costs, +/// because what is being pinned is the **measurement boundary** -- where a +/// strategy's preparation is accounted -- and not a fact about how expensive +/// any of them is on a given machine. +#[test] +fn a_host_round_trip_is_counted_as_part_of_its_commit() { + let (path, file) = scratch("prepare-boundary"); + let payload = b"a harness record".to_vec(); + + let outcome = super::run( + CommitStrategy::HostSequenced, + file.as_raw_handle(), + 3, + 4, + &payload, + None, + ) + .expect("a small run completes"); + + assert!( + outcome + .commit_timings + .iter() + .any(|timing| !timing.prepare.is_zero()), + "host-sequenced waits for every write before it flushes, and that wait is part of what \ + its commit costs -- a zero here means the clock starts after it" + ); + + drop(file); + let _ = std::fs::remove_file(&path); +} + +/// A commit's cost is reported in three parts that add up (M25.4). +/// +/// The harness published one blended number until `M20.6` decomposed it by +/// hand and found it was **entirely deferral** -- a figure that looked like a +/// commit latency and measured how long the next epoch's appends took. The +/// split is the fix, and this pins the two properties that make it a fix +/// rather than a rename: +/// +/// - the parts **sum** to the old number, so nothing was lost in splitting; +/// - `flush()` **excludes** deferral, which is the whole point. +/// +/// It deliberately asserts nothing about the *values*. Which part carries the +/// cost is a property of the machine and the handle, not of this crate, and +/// `M25`'s standing constraint forbids depending on an operation pending. +#[test] +fn a_commits_cost_is_reported_in_parts_that_do_not_overlap() { + let (path, file) = scratch("timing-parts"); + let payload = b"a harness record".to_vec(); + + let outcome = super::run( + CommitStrategy::CoveringFlush, + file.as_raw_handle(), + 3, + 4, + &payload, + None, + ) + .expect("a small run completes"); + + assert!( + !outcome.commit_timings.is_empty(), + "a run of three epochs must time three commits" + ); + for timing in &outcome.commit_timings { + assert_eq!( + timing.flush(), + timing.prepare + timing.submit + timing.blocking, + "a commit's cost is its preparation, its submit and its wait, and nothing else" + ); + assert!( + timing.flush() <= timing.prepare + timing.submit + timing.blocking + timing.deferral, + "deferral must not be folded into the commit: that is the defect M20.6 found" + ); + } + + // The identity above is satisfied by a part that is never measured at all, + // which is not hypothetical: replacing the deferral measurement with zero + // was **survived** by the assertions above alone. A part that always reads + // zero is a column of zeros in the report and a decomposition in name only. + // + // Asserted as "some sample is non-zero" rather than a lower bound on any + // duration. This harness defers by construction -- it pushes the next + // epoch's appends before settling the previous commit -- so a run in which + // *nothing* deferred means the clock is not running, not that the machine + // was fast. `blocking` gets no such assertion, because zero is a legitimate + // and frequently observed reading for it. + assert!( + outcome + .commit_timings + .iter() + .any(|timing| !timing.deferral.is_zero()), + "a harness that defers by design must observe some deferral, or it is not measuring it" + ); + assert!( + outcome + .commit_timings + .iter() + .any(|timing| !timing.submit.is_zero()), + "submitting a flush must cost something, or it is not being measured" + ); + + drop(file); + let _ = std::fs::remove_file(&path); +} + +/// All three strategies must write byte-identical logs (M25.1b). +/// +/// This is `compare_strategies`' strongest assertion and the one with real +/// teeth: replay checks a log against itself, where this checks the three +/// against *each other*, so a dropped record, a wrong offset, or an ordering +/// bug in any one of them shows up as a difference from the other two. +/// +/// It had no test. The harness test above runs `CoveringFlush` alone, and +/// `main`'s version needs all three -- which made it the one assertion in the +/// sample that only `cargo run` could reach. Running them at two epochs of +/// three records instead of thirty-two of sixty-four puts it at a rung that +/// runs everywhere, for a cost the suite does not notice. +/// +/// **Why identical bytes is the right expectation even across ring counts:** +/// a record's offset is its position in the global sequence times the stride, +/// and a strategy decides which *ring* submits a write, never where it lands. +/// `AlternatingRings` therefore interleaves submissions across two lanes into +/// the same offset space and must still produce the same file. +#[test] +fn every_strategy_writes_the_same_log() { + const EPOCHS: usize = 2; + const PER_EPOCH: usize = 3; + let payload = b"a harness record".to_vec(); + + let mut reference: Option<(&'static str, Vec)> = None; + for strategy in super::CommitStrategy::ALL { + let (path, file) = scratch(&format!("cross-{}", strategy.name())); + let outcome = super::run( + strategy, + file.as_raw_handle(), + EPOCHS, + PER_EPOCH, + &payload, + None, + ) + .expect("a small run completes"); + assert_eq!( + outcome.records, + EPOCHS * PER_EPOCH, + "{} did not write every record it was asked for", + strategy.name() + ); + + let bytes = std::fs::read(&path).expect("read the log back"); + drop(file); + let _ = std::fs::remove_file(&path); + + match &reference { + None => reference = Some((strategy.name(), bytes)), + Some((first, expected)) => assert_eq!( + &bytes, + expected, + "{} produced a different log than {first}; all three must write the same bytes", + strategy.name() + ), + } + } +} +/// +/// **The gap this closes was measured, not suspected.** The harness's `Lane` +/// is a second writer over the same on-disk format as the log's own +/// `Appender`, and when the stride was introduced nothing tested it: reverting +/// this lane to a packed layout left every test in the example green. The only +/// thing that exercised it was running the sample, and no CI job does. +/// +/// `CoveringFlush` specifically, because it uses **one** ring. The multi-ring +/// strategies interleave two lanes into one offset space, so the on-disk order +/// is not the sequence order and replay's in-order check would refuse a +/// perfectly good log. That is a property of the harness rather than of the +/// format, so this test picks the configuration where the format's own rule is +/// the only thing under test. +#[test] +fn a_run_lays_its_records_out_one_per_stride() { + let (path, file) = scratch("stride-layout"); + let payload = b"a harness record".to_vec(); + + let outcome = super::run( + CommitStrategy::CoveringFlush, + file.as_raw_handle(), + 2, + 3, + &payload, + None, + ) + .expect("a small run completes"); + + let bytes = std::fs::read(&path).expect("read the log back"); + assert_eq!( + bytes.len(), + outcome.records * crate::record::RECORD_STRIDE, + "every record must occupy exactly one block" + ); + + let replayed = crate::replay::replay(&bytes, outcome.durable_through, outcome.records, |_| { + payload.clone() + }); + assert!( + replayed.is_clean(), + "the harness writes the same format the reader walks: {:?}", + replayed.violations + ); + assert_eq!(replayed.durable_verified, outcome.records); + + drop(file); + let _ = std::fs::remove_file(&path); +} diff --git a/crates/windows-ioring-sys/examples/epoch_log/tests.rs b/crates/windows-ioring-sys/examples/epoch_log/tests.rs new file mode 100644 index 000000000..bc94af446 --- /dev/null +++ b/crates/windows-ioring-sys/examples/epoch_log/tests.rs @@ -0,0 +1,90 @@ +// Copyright (c) 2026 Mike Grier +//! The sample's own end-to-end verification, run by `cargo test` (M25.1b). +//! +//! # Why this exists +//! +//! [`crate::replay`] calls itself "the only part that can catch a durability +//! bug, because it is the only part that checks the claim [`crate::contract`] +//! actually makes rather than the steps taken to reach it". Until this file, +//! **nothing ran it except a human typing `cargo run`.** No CI job executes +//! the example, and `cargo test --example epoch_log` compiles `main.rs` as a +//! test harness without ever calling `main`. +//! +//! That was not a theoretical gap. During `M25.1` the writer was converted to a +//! strided layout and the reader was not, which left the log unreadable -- and +//! **every one of the example's tests passed anyway**. Only running the sample +//! caught it. +//! +//! # Why a test rather than a CI step +//! +//! `M25.1b` framed this as a choice between the two. A CI job running the +//! binary is the obvious answer, and it is the wrong rung: it catches a defect +//! after it is pushed, on a machine the author is not sitting at. The +//! repository's FAIL FAST rule asks for the lowest rung that can carry the +//! fact, and this one carries it -- `run_log` and `verify` are ordinary +//! functions over a generic [`crate::Report`], so a test can drive the real log +//! against a real ring and a real file and then run the real verifier, which is +//! the same code path `main` takes. +//! +//! This is not a proxy for running the sample. It *is* running the sample, +//! minus the part below. +//! +//! # What is deliberately left out +//! +//! [`crate::strategy`]'s three-way comparison, which `main` runs after +//! `verify`. It is the expensive part -- thirty-two epochs of sixty-four +//! records, three times -- and its layout and replay are covered by +//! `strategy::tests::a_run_lays_its_records_out_one_per_stride`, which was +//! added by `M25.1` for the same reason this file exists. Running it again here +//! would multiply the suite's cost to re-check something already checked. +//! +//! So `cargo run --example epoch_log` remains the only thing that exercises the +//! comparison end to end, and that is a deliberate line rather than an +//! oversight: what it would add is a measurement, and a measurement is not a +//! contract this crate can assert. + +use std::path::PathBuf; + +use crate::{Report, run_log, verify}; + +/// Scratch paths for this test, named so they cannot collide with the +/// process-scoped paths `main` uses or with another test running as a sibling +/// thread. +fn scratch(tag: &str) -> PathBuf { + std::env::temp_dir().join(format!( + "windows-ioring-sys-epoch-selftest-{}-{tag}", + std::process::id() + )) +} + +/// The whole sample, end to end: append, commit, checkpoint, reclaim, then all +/// three replay passes and the negative control. +/// +/// The assertions live inside [`verify`] -- that a clean log reports no +/// violations, that a torn tail is tolerated rather than rejected, and that +/// corrupting a byte inside the durable region **is** reported. The last is +/// what makes the other two mean anything, and it is why this test does not +/// add assertions of its own: duplicating them here would be a second copy of +/// the contract check, which is the defect `M25.1` spent its time removing. +#[test] +fn the_sample_keeps_its_contract_and_its_verifier_can_still_fail() { + // Output goes to buffers rather than stdout: a passing test should be + // silent, and `Report` is generic over its writers precisely so a caller + // can decide where the narrative lands. + let mut report = Report::new(Vec::new(), Vec::new()); + + let log = scratch("log"); + let retired = scratch("retired"); + let checkpoint = scratch("checkpoint"); + + let outcome = run_log(&mut report, &log, &retired, &checkpoint) + .and_then(|run| verify(&mut report, &log, &run)); + + // Cleaned up before the assertion, so a failure does not also leave files + // behind for the next run to trip over. + for path in [&log, &retired, &checkpoint] { + let _ = std::fs::remove_file(path); + } + + outcome.expect("the sample must run and verify its own contract"); +} diff --git a/crates/windows-ioring-sys/examples/l3_domains.rs b/crates/windows-ioring-sys/examples/l3_domains.rs deleted file mode 100644 index f1672bcea..000000000 --- a/crates/windows-ioring-sys/examples/l3_domains.rs +++ /dev/null @@ -1,39 +0,0 @@ -// Copyright (c) 2026 Mike Grier -//! M6.3: enumerate last-level-cache (L3) domains -- the default heuristic -//! `DESIGN-NOTES.md`'s "Why the NUMA node is the wrong key" recommends for -//! sizing an `IoRing` execution domain. This is enumeration only, not a -//! partitioning policy: what to do with the domains is a workload call this -//! crate deliberately leaves to the caller (see `windows-topology-sys` for a -//! safe `GetLogicalProcessorInformationEx` wrapper). - -use std::io::Write; - -fn main() -> std::io::Result<()> { - // The single sink every line of this sample's output goes through - // (repository "Architectural pre-steps" rule: never call `println!` - // from more than one call site). - let mut out = std::io::stdout(); - let relations = windows_topology_sys::discover()?; - let l3_domains: Vec<_> = relations - .caches - .iter() - .filter(|cache| cache.level == 3) - .collect(); - - let _ = writeln!(out, "{} last-level cache (L3) domain(s):", l3_domains.len()); - for (index, domain) in l3_domains.iter().enumerate() { - let _ = writeln!( - out, - " domain {index}: {} bytes shared by {:?}", - domain.cache_size, domain.processors - ); - } - - // Processor groups are a hard floor (D-8 in DESIGN-NOTES.md): above 64 - // logical processors, a thread's affinity and a ring's waiter are each - // confined to one GROUP_AFFINITY, whether or not that partition is - // wanted. - let _ = writeln!(out, "{} processor group(s)", relations.groups.len()); - - Ok(()) -} diff --git a/crates/windows-ioring-sys/examples/ring_copy/buffer.rs b/crates/windows-ioring-sys/examples/ring_copy/buffer.rs deleted file mode 100644 index a2519819b..000000000 --- a/crates/windows-ioring-sys/examples/ring_copy/buffer.rs +++ /dev/null @@ -1,91 +0,0 @@ -// Copyright (c) 2026 Mike Grier -//! A `VirtualAllocExNuma`-backed buffer (M7.3), so a domain's registered -//! buffer can be placed on a chosen NUMA node rather than wherever the -//! default allocator's own heuristics land it. - -use std::io; -use std::ptr; - -use windows_ioring_sys::{IoBuf, IoBufMut}; -use windows_sys::Win32::System::Memory::{ - MEM_COMMIT, MEM_RELEASE, MEM_RESERVE, PAGE_READWRITE, VirtualAllocExNuma, VirtualFree, -}; -use windows_sys::Win32::System::Threading::GetCurrentProcess; - -/// `VirtualAllocExNuma`'s documented sentinel for "no NUMA preference" -- -/// windows-sys does not name this constant, so it is named here rather than -/// written as a bare literal at the call site. -const NUMA_NO_PREFERRED_NODE: u32 = u32::MAX; - -/// An owned buffer allocated with `VirtualAllocExNuma`, freed with -/// `VirtualFree` on drop. -pub struct NumaBuffer { - ptr: *mut u8, - len: usize, -} - -// SAFETY: the allocation is exclusively owned by this value; sending it -// across threads only moves that ownership, never aliases it. -unsafe impl Send for NumaBuffer {} - -impl NumaBuffer { - /// Allocate `len` bytes, preferring `node` if given. - /// - /// # Errors - /// - /// Returns the error from `VirtualAllocExNuma`. - pub fn new(len: usize, node: Option) -> io::Result { - // SAFETY: no pointer arguments; the returned value is a pseudo-handle - // that needs no closing. - let process = unsafe { GetCurrentProcess() }; - // SAFETY: `process` is a valid pseudo-handle for the duration of this - // call; a null `lpAddress` lets the system choose the address. - let ptr = unsafe { - VirtualAllocExNuma( - process, - ptr::null(), - len, - MEM_COMMIT | MEM_RESERVE, - PAGE_READWRITE, - node.unwrap_or(NUMA_NO_PREFERRED_NODE), - ) - }; - if ptr.is_null() { - return Err(io::Error::last_os_error()); - } - Ok(Self { - ptr: ptr.cast(), - len, - }) - } -} - -impl Drop for NumaBuffer { - fn drop(&mut self) { - // SAFETY: `self.ptr` was returned by `VirtualAllocExNuma` above and - // is freed exactly once, here. - unsafe { - VirtualFree(self.ptr.cast(), 0, MEM_RELEASE); - } - } -} - -// SAFETY: the allocation's address is fixed once `VirtualAllocExNuma` -// returns it and does not move for this value's life; `len` is fixed too. -unsafe impl IoBuf for NumaBuffer { - fn stable_ptr(&self) -> *const u8 { - self.ptr - } - - fn bytes_len(&self) -> usize { - self.len - } -} - -// SAFETY: this value uniquely owns the allocation, so `&mut self` is -// exclusive access; the address is the same one `stable_ptr` reports. -unsafe impl IoBufMut for NumaBuffer { - fn stable_mut_ptr(&mut self) -> *mut u8 { - self.ptr - } -} diff --git a/crates/windows-ioring-sys/examples/ring_copy/engine.rs b/crates/windows-ioring-sys/examples/ring_copy/engine.rs index abd0cacbd..c54e2e684 100644 --- a/crates/windows-ioring-sys/examples/ring_copy/engine.rs +++ b/crates/windows-ioring-sys/examples/ring_copy/engine.rs @@ -8,14 +8,14 @@ use std::ops::Range; use std::ptr; use std::time::{Duration, Instant}; +use win_numa_sys::NumaNode; use windows_ioring_sys::{ - Batch, IoRing, PushOptions, RegisteredBuffers, RegisteredSpan, Token, WriteCaching, + Batch, IoRing, NumaBuffer, PushOptions, RegisteredBuffers, RegisteredSpan, Token, WriteCaching, }; use windows_sys::Win32::Foundation::HANDLE; use windows_sys::Win32::System::SystemInformation::GROUP_AFFINITY; use windows_sys::Win32::System::Threading::{GetCurrentThread, SetThreadGroupAffinity}; -use crate::buffer::NumaBuffer; use crate::plan::DomainPlan; /// How long a single push-and-wait may block before this sample gives up on @@ -43,7 +43,7 @@ pub fn copy_domain( destination: HANDLE, byte_range: Range, chunk_len: usize, - numa_node: Option, + numa_node: Option, ) -> io::Result { affinitize(plan.group, plan.mask)?; diff --git a/crates/windows-ioring-sys/examples/ring_copy/main.rs b/crates/windows-ioring-sys/examples/ring_copy/main.rs index 35fe4912e..7d130e6ee 100644 --- a/crates/windows-ioring-sys/examples/ring_copy/main.rs +++ b/crates/windows-ioring-sys/examples/ring_copy/main.rs @@ -7,7 +7,6 @@ //! partitioning policy (D-8 in its `DESIGN-NOTES.md`), so the policy lives //! here instead, giving M6's guidance something executable behind it. -mod buffer; mod engine; mod plan; mod policy; @@ -77,7 +76,7 @@ struct Args { fn parse_args() -> Result { let mut positional = Vec::new(); - let mut policy = Policy::ByL3; + let mut policy = Policy::ByCache; let mut remote_placement = false; let mut topology_path = None; let mut chunk_len = DEFAULT_CHUNK_LEN; diff --git a/crates/windows-ioring-sys/examples/ring_copy/plan.rs b/crates/windows-ioring-sys/examples/ring_copy/plan.rs index 028d37451..511d0c32e 100644 --- a/crates/windows-ioring-sys/examples/ring_copy/plan.rs +++ b/crates/windows-ioring-sys/examples/ring_copy/plan.rs @@ -4,6 +4,7 @@ use std::io; +use win_numa_sys::NumaNode; use windows_topology_sys::{Domain, DomainKind, MachineMemoryTopology, ProcessorSet, Source}; /// What one execution domain needs to run: a single-group affinity mask and, @@ -13,7 +14,7 @@ pub struct DomainPlan { pub label: String, pub group: u16, pub mask: usize, - pub local_numa_node: Option, + pub local_numa_node: Option, } /// Build one plan per domain, rejecting any domain the platform cannot @@ -94,7 +95,7 @@ fn label_for(domain: &Domain) -> String { /// The NUMA node whose processors overlap `processors`, if any domain /// reports one -- `None` on a machine that reports no NUMA nodes at all. -fn numa_node_for(topology: &MachineMemoryTopology, processors: &ProcessorSet) -> Option { +fn numa_node_for(topology: &MachineMemoryTopology, processors: &ProcessorSet) -> Option { topology .domains .iter() @@ -104,6 +105,10 @@ fn numa_node_for(topology: &MachineMemoryTopology, processors: &ProcessorSet) -> } _ => None, }) + // The relationship walk's label is the real Windows node number, not a + // position in this list. Wrapping it here rather than at the call site + // is what stops a positional index ever reaching `VirtualAllocExNuma`. + .map(NumaNode::new) } /// What `--placement remote` can actually be given on this topology. @@ -115,7 +120,7 @@ fn numa_node_for(topology: &MachineMemoryTopology, processors: &ProcessorSet) -> #[derive(Debug, Clone, Copy, PartialEq, Eq)] pub enum RemoteNode { /// A NUMA node other than the local one. The switch does what it says. - Other(u32), + Other(NumaNode), /// The memory domains name their nodes, and there is only the local one -- /// an ordinary single-node machine. Falling back to local is the honest /// answer here, because no other node exists to place on. @@ -163,7 +168,7 @@ pub fn names_any_numa_node(topology: &MachineMemoryTopology) -> bool { }) } -pub fn remote_numa_node(topology: &MachineMemoryTopology, local: Option) -> RemoteNode { +pub fn remote_numa_node(topology: &MachineMemoryTopology, local: Option) -> RemoteNode { // **Remoteness is a relationship, so it needs both ends.** Without a local // node there is nothing for a candidate to be remote *from*: the comparison // below is `Some(id) != local`, and against `None` that is true for every @@ -181,6 +186,7 @@ pub fn remote_numa_node(topology: &MachineMemoryTopology, local: Option) -> let Some(id) = domain.label_from(Source::RelationshipWalk) else { continue; }; + let id = NumaNode::new(id); any_named = true; if Some(id) != local { return RemoteNode::Other(id); diff --git a/crates/windows-ioring-sys/examples/ring_copy/policy.rs b/crates/windows-ioring-sys/examples/ring_copy/policy.rs index 1875b9717..da95235a0 100644 --- a/crates/windows-ioring-sys/examples/ring_copy/policy.rs +++ b/crates/windows-ioring-sys/examples/ring_copy/policy.rs @@ -11,9 +11,23 @@ use windows_topology_sys::{Domain, DomainKind, MachineMemoryTopology, ProcessorS /// here, in the sample, not as extensible policy data nobody asked for. #[derive(Clone, Copy, Debug, PartialEq, Eq)] pub enum Policy { - /// One domain per last-level (L3) cache -- the default heuristic - /// `DESIGN-NOTES.md` recommends. - ByL3, + /// One domain per **outermost cache level that actually partitions the + /// machine** -- the default heuristic `DESIGN-NOTES.md` recommends. + /// + /// Not "L3". The rule is asked of + /// [`MachineMemoryTopology::outermost_partitioning_cache`], which is the + /// one definition of which level partitions a host, and *which* level that + /// is varies: on EPYC it is the L3/CCX boundary, and on a shipping + /// Snapdragon X2 Elite there is no L3 at all and the natural boundary is + /// L2 ([D-48](../../DESIGN-NOTES.md#d-48)). + /// + /// This used to match `level: 3` here, which is the consumer-side twin of + /// the platform-integrity failure: binding to the level number that + /// happens to be right on today's hardware instead of to the specified + /// primitive. It could also produce **overlapping** ring domains where two + /// cache kinds are reported at the same level, which `select` has no way + /// to detect and a ring runtime has no way to survive. + ByCache, /// One domain per NUMA node. ByNode, /// One domain per physical package (socket). @@ -26,10 +40,17 @@ pub enum Policy { impl Policy { /// Parse a policy name (case-insensitive), for the sample's `--policy` switch. + /// + /// `byl3` and `l3` are deliberately **not** accepted. They named a rule + /// this sample no longer implements, and silently mapping them onto + /// [`Policy::ByCache`] would let a script keep asking for L3 and keep + /// believing it got L3 -- on a machine whose partitioning level is L2, + /// that is a wrong answer delivered quietly. An unknown name prints the + /// usage line instead, which is a question rather than a wrong answer. #[must_use] pub fn parse(name: &str) -> Option { match name.to_ascii_lowercase().as_str() { - "byl3" | "l3" => Some(Self::ByL3), + "bycache" | "cache" => Some(Self::ByCache), "bynode" | "node" => Some(Self::ByNode), "bypackage" | "package" => Some(Self::ByPackage), "bycore" | "core" => Some(Self::ByCore), @@ -43,20 +64,25 @@ impl Policy { /// /// Degrades to one whole-machine domain when the policy's preferred /// relation is not reported at all -- for example `ByNode` on a machine - /// reporting zero NUMA nodes -- the same "one ring is correct when the - /// answer is unknowable" degradation `DESIGN-NOTES.md` describes for L3 - /// domains, generalized to every policy here (M7.5 depends on knowing - /// when this happened, to report it honestly rather than silently). + /// reporting zero NUMA nodes, or [`Policy::ByCache`] on one whose caches + /// partition nothing -- the same "one ring is correct when the answer is + /// unknowable" degradation `DESIGN-NOTES.md` describes, generalized to + /// every policy here (M7.5 depends on knowing when this happened, to + /// report it honestly rather than silently). #[must_use] pub fn select(self, topology: &MachineMemoryTopology) -> (Vec, bool) { let matched: Vec = match self { Self::Single => Vec::new(), - Self::ByL3 => topology - .domains - .iter() - .filter(|domain| matches!(domain.kind, DomainKind::Cache { level: 3, .. })) - .cloned() - .collect(), + // Asked, not restated. `outermost_partitioning_cache` owns the + // rule -- including that "outermost" is decided by inclusion + // rather than by the level number, and that a level qualifies only + // when its blocks are pairwise disjoint. Both are properties this + // consumer would otherwise have to re-derive, and the second is + // the one whose absence let the old code emit overlapping domains. + Self::ByCache => topology + .outermost_partitioning_cache() + .map(|(_level, domains)| domains.into_iter().cloned().collect()) + .unwrap_or_default(), Self::ByNode => topology .domains .iter() @@ -105,3 +131,6 @@ impl Policy { ) } } + +#[cfg(test)] +mod tests; diff --git a/crates/windows-ioring-sys/examples/ring_copy/policy/tests.rs b/crates/windows-ioring-sys/examples/ring_copy/policy/tests.rs new file mode 100644 index 000000000..3d2bbc39a --- /dev/null +++ b/crates/windows-ioring-sys/examples/ring_copy/policy/tests.rs @@ -0,0 +1,466 @@ +// Copyright (c) 2026 Mike Grier +//! Tests for [`Policy::select`]'s degraded fallback (M20.3). +//! +//! # Why these exist +//! +//! The whole-machine fallback is the branch **every zero-relation machine +//! takes**, and it cannot be reached by running the sample on a machine that +//! reports its relations. A synthetic topology reaches it; until M20.3 nothing +//! did, and the first confirmation it ran at all came from a design session +//! rather than from a test. +//! +//! # Both directions, because one alone proves nothing +//! +//! The item that queued these asked for two halves on purpose: that a policy +//! whose relation is **absent** degrades, and that a policy whose relation is +//! **present** does not. The second is what makes the first mean anything -- +//! a test of the absent case alone passes just as happily against a `select` +//! that degrades unconditionally. [`degrading_unconditionally_would_fail_a_test_here`] +//! states that dependency in code rather than leaving it to this comment. +//! +//! # Why no case uses `ByL3` +//! +//! `SH-4.12` is queued to rewrite that arm: it replaces the hardcoded +//! `level: 3` match with `outermost_partitioning_cache()`, and renames the +//! policy, because `byl3` is a user-facing CLI value that would stop +//! describing what it does. Asserting through `ByL3` here would pin behaviour +//! that is about to change. +//! +//! The fallback tail itself is *shared by every policy* and is not what that +//! item touches, so exercising it through `ByNode` and `ByPackage` tests the +//! same branch without standing in the way. `ByL3`'s own degradation +//! condition belongs to `SH-4.12`'s verification, where the rule it degrades +//! on is the new one. + +use std::collections::BTreeMap; + +use windows_topology_sys::{ + CacheKind, Domain, DomainKind, MachineMemoryTopology, Observed, Processor, ProcessorId, + ProcessorSet, +}; + +use super::Policy; + +/// A processor that is present and usable. +fn online(group: u16, number: u8) -> Processor { + Processor { + id: ProcessorId { group, number }, + online: true, + capacity: 1, + } +} + +/// A processor slot Windows reserved but that is not usable -- a group can +/// carry these up to its maximum count, and the fallback must not offer them +/// to a ring. +fn offline(group: u16, number: u8) -> Processor { + Processor { + id: ProcessorId { group, number }, + online: false, + capacity: 0, + } +} + +fn processors(ids: &[(u16, u8)]) -> ProcessorSet { + let mut set = ProcessorSet::empty(); + for &(group, number) in ids { + set.insert(group, number); + } + set +} + +fn domain(kind: DomainKind, ids: &[(u16, u8)]) -> Domain { + Domain { + kind, + processors: processors(ids), + observations: Vec::new(), + } +} + +fn package(ids: &[(u16, u8)]) -> Domain { + domain(DomainKind::Package, ids) +} + +fn memory(ids: &[(u16, u8)]) -> Domain { + // NotObserved, not Known(0): nothing measured this synthetic node's + // size, and claiming it has zero bytes would be the fabricated record + // Observed exists to prevent. + domain( + DomainKind::Memory { + memory_bytes: Observed::NotObserved, + }, + ids, + ) +} + +fn core(ids: &[(u16, u8)]) -> Domain { + domain( + DomainKind::Core { + simultaneous_multithreading: false, + efficiency_class: 0, + }, + ids, + ) +} + +fn cache(level: u8, ids: &[(u16, u8)]) -> Domain { + domain( + DomainKind::Cache { + level, + associativity: 0, + line_size: 0, + size_bytes: 0, + cache_type: CacheKind::Unified, + }, + ids, + ) +} + +/// A machine with four online processors and whichever relations are given. +fn machine(domains: Vec) -> MachineMemoryTopology { + MachineMemoryTopology { + processors: vec![online(0, 0), online(0, 1), online(0, 2), online(0, 3)], + domains, + ..Default::default() + } +} + +/// Whether `domains` is the single whole-machine domain the fallback builds. +fn is_whole_machine(domains: &[Domain]) -> bool { + domains.len() == 1 + && matches!(&domains[0].kind, DomainKind::Other { name, .. } if name == "whole-machine") +} + +// ---------------------------------------------------------------- absent --- + +#[test] +fn a_policy_whose_relation_is_absent_degrades_to_one_whole_machine_domain() { + // The zero-NUMA-node machine, which D-48 records as an ordinary consumer + // shape rather than an exotic one. + let topology = machine(vec![package(&[(0, 0), (0, 1), (0, 2), (0, 3)])]); + let (domains, degraded) = Policy::ByNode.select(&topology); + + assert!( + degraded, + "a relation nobody reported must be reported as a fallback" + ); + assert!( + is_whole_machine(&domains), + "and the fallback is one whole-machine domain, not zero domains" + ); +} + +#[test] +fn a_machine_with_no_relations_at_all_degrades_for_every_relational_policy() { + let topology = machine(Vec::new()); + for policy in [Policy::ByNode, Policy::ByPackage, Policy::ByCore] { + let (domains, degraded) = policy.select(&topology); + assert!(degraded, "{policy:?} has no relation to select on"); + assert!(is_whole_machine(&domains), "{policy:?}"); + } +} + +#[test] +fn a_memory_domain_with_no_processors_does_not_count_as_a_node() { + // A memory-only domain is legal (D-5) and carries no processors, so it + // cannot host a ring. Selecting on it would produce a domain with nothing + // to pin a thread to, which is worse than degrading honestly. + let topology = machine(vec![memory(&[])]); + let (domains, degraded) = Policy::ByNode.select(&topology); + + assert!(degraded, "a node with no processors is not a usable domain"); + assert!(is_whole_machine(&domains)); +} + +// --------------------------------------------------------------- present --- + +#[test] +fn a_policy_whose_relation_is_present_is_not_degraded() { + // The half that makes the absent cases mean something. + let topology = machine(vec![package(&[(0, 0), (0, 1)]), package(&[(0, 2), (0, 3)])]); + let (domains, degraded) = Policy::ByPackage.select(&topology); + + assert!( + !degraded, + "the relation was reported, so nothing was degraded" + ); + assert_eq!(domains.len(), 2, "and every matching domain is returned"); + assert!( + domains + .iter() + .all(|d| matches!(d.kind, DomainKind::Package)), + "the domains returned are the ones asked for" + ); +} + +#[test] +fn every_relational_policy_is_undegraded_when_its_own_relation_is_present() { + let topology = machine(vec![ + memory(&[(0, 0), (0, 1), (0, 2), (0, 3)]), + package(&[(0, 0), (0, 1), (0, 2), (0, 3)]), + core(&[(0, 0), (0, 1)]), + core(&[(0, 2), (0, 3)]), + ]); + for policy in [Policy::ByNode, Policy::ByPackage, Policy::ByCore] { + let (domains, degraded) = policy.select(&topology); + assert!( + !degraded, + "{policy:?} asked for a relation this machine has" + ); + assert!( + !is_whole_machine(&domains), + "{policy:?} must return its own domains, not the fallback" + ); + } +} + +#[test] +fn one_policy_degrading_does_not_degrade_another_on_the_same_machine() { + // Degradation is per-policy, not a property of the machine. A host that + // reports packages but no nodes must answer differently to the two. + let topology = machine(vec![package(&[(0, 0), (0, 1), (0, 2), (0, 3)])]); + + let (_, node_degraded) = Policy::ByNode.select(&topology); + let (_, package_degraded) = Policy::ByPackage.select(&topology); + + assert!(node_degraded, "no memory domain was reported"); + assert!(!package_degraded, "but a package domain was"); +} + +// ------------------------------------------------- the guard on the guard --- + +#[test] +fn degrading_unconditionally_would_fail_a_test_here() { + // The dependency between the two halves, stated in code. If `select` ever + // degrades unconditionally, at least one assertion in this file must go + // red -- so this names the case that would catch it rather than trusting + // that some other test happens to. + let topology = machine(vec![package(&[(0, 0), (0, 1), (0, 2), (0, 3)])]); + let (domains, degraded) = Policy::ByPackage.select(&topology); + assert!( + !degraded && !is_whole_machine(&domains), + "an unconditional degrade is indistinguishable from a correct one \ + unless some case asserts the undegraded outcome" + ); +} + +// ----------------------------------------------------- ByCache (SH-4.12) --- + +#[test] +fn by_cache_selects_the_level_that_partitions_not_the_level_numbered_three() { + // The regression this policy was rewritten for, and it is not + // hypothetical: measured on the machine this repository is developed on, + // which reports an L3 spanning all 16 processors and a real 8-way L2 + // partition underneath it. The old `level: 3` filter found that L3, + // returned ONE whole-machine domain, and -- because it matched something + // -- did not flag the result degraded. It reported success while + // collapsing an 8-domain machine to a single ring. + let topology = machine(vec![ + cache(2, &[(0, 0), (0, 1)]), + cache(2, &[(0, 2), (0, 3)]), + cache(3, &[(0, 0), (0, 1), (0, 2), (0, 3)]), + ]); + let (domains, degraded) = Policy::ByCache.select(&topology); + + assert!(!degraded); + assert_eq!( + domains.len(), + 2, + "L2 partitions this machine; L3 covers all of it and partitions nothing" + ); + assert!( + domains + .iter() + .all(|d| matches!(d.kind, DomainKind::Cache { level: 2, .. })), + "the level chosen is the one that splits the machine, not the larger number" + ); +} + +#[test] +fn by_cache_degrades_when_the_only_cache_spans_the_whole_machine() { + // A cache every processor shares offers no boundary to size a ring by, so + // the honest answer is one ring *reported as degraded* -- not one ring + // reported as a cache-aware partition, which is what the old filter did. + let topology = machine(vec![cache(3, &[(0, 0), (0, 1), (0, 2), (0, 3)])]); + let (domains, degraded) = Policy::ByCache.select(&topology); + + assert!(degraded, "one block is not a partition"); + assert!(is_whole_machine(&domains)); +} + +#[test] +fn by_cache_degrades_when_no_cache_is_reported_at_all() { + let topology = machine(vec![package(&[(0, 0), (0, 1), (0, 2), (0, 3)])]); + let (domains, degraded) = Policy::ByCache.select(&topology); + + assert!(degraded); + assert!(is_whole_machine(&domains)); +} + +#[test] +fn by_cache_works_on_a_machine_whose_outermost_partition_is_l2() { + // D-48's shape: a shipping ARM part with no L3 at all, whose natural + // cluster boundary is L2. The old filter returned zero domains here and + // degraded; this returns the partition that exists. + let topology = machine(vec![ + cache(2, &[(0, 0), (0, 1)]), + cache(2, &[(0, 2), (0, 3)]), + ]); + let (domains, degraded) = Policy::ByCache.select(&topology); + + assert!( + !degraded, + "this machine has a cache partition, it is just not L3" + ); + assert_eq!(domains.len(), 2); +} + +#[test] +fn by_cache_is_not_spelled_byl3_any_more() { + // The rename is the point, not a side effect: `byl3` named a rule this + // sample no longer implements, and accepting it would let a script keep + // asking for L3 and keep believing it got L3. + assert_eq!(Policy::parse("bycache"), Some(Policy::ByCache)); + assert_eq!(Policy::parse("cache"), Some(Policy::ByCache)); + assert_eq!(Policy::parse("byl3"), None, "byl3 must not silently map on"); + assert_eq!(Policy::parse("l3"), None, "nor l3"); +} + +// ------------------------------------------------------- Single, and edges --- + +#[test] +fn single_returns_the_whole_machine_without_calling_it_degraded() { + // The edge that a naive "returned the whole machine, so it degraded" + // implementation gets wrong. `Single` asks for one domain, so one domain + // is exactly what it wanted -- degrading is a statement about not getting + // what was asked for. + let topology = machine(vec![package(&[(0, 0), (0, 1), (0, 2), (0, 3)])]); + let (domains, degraded) = Policy::Single.select(&topology); + + assert!(is_whole_machine(&domains)); + assert!(!degraded, "Single got precisely what it asked for"); +} + +#[test] +fn single_is_undegraded_even_on_a_machine_with_no_relations() { + let topology = machine(Vec::new()); + let (domains, degraded) = Policy::Single.select(&topology); + + assert!(is_whole_machine(&domains)); + assert!( + !degraded, + "Single never depends on a relation being reported" + ); +} + +#[test] +fn the_fallback_domain_covers_every_online_processor() { + let topology = machine(Vec::new()); + let (domains, _) = Policy::ByNode.select(&topology); + + let covered = &domains[0].processors; + for number in 0..4u8 { + assert!( + covered.contains(0, number), + "processor {number} is online and must be in the fallback domain" + ); + } +} + +#[test] +fn the_fallback_domain_excludes_offline_processors() { + // The other direction of the same rule. A reserved-but-absent slot counts + // toward a group's maximum, and handing one to a ring would pin a thread + // to a processor that does not exist. + let topology = MachineMemoryTopology { + processors: vec![online(0, 0), offline(0, 1), online(0, 2), offline(0, 3)], + domains: Vec::new(), + ..Default::default() + }; + let (domains, degraded) = Policy::ByNode.select(&topology); + + assert!(degraded); + let covered = &domains[0].processors; + assert!( + covered.contains(0, 0) && covered.contains(0, 2), + "the online ones" + ); + assert!( + !covered.contains(0, 1) && !covered.contains(0, 3), + "and not the reserved slots" + ); +} + +#[test] +fn the_fallback_domain_spans_processor_groups() { + // A machine above 64 logical processors reports more than one group, and + // the fallback is meant to be the *whole* machine. + let topology = MachineMemoryTopology { + processors: vec![online(0, 0), online(0, 1), online(1, 0), online(1, 1)], + domains: Vec::new(), + ..Default::default() + }; + let (domains, _) = Policy::ByNode.select(&topology); + + let covered = &domains[0].processors; + assert!( + covered.contains(0, 0) && covered.contains(1, 1), + "both groups" + ); +} + +#[test] +fn the_fallback_domain_carries_no_observations() { + // Nothing observed this relation -- the sample built it. Claiming an + // observation would be exactly the fabricated record `Observed` exists to + // prevent. + let topology = machine(Vec::new()); + let (domains, _) = Policy::ByNode.select(&topology); + + assert!( + domains[0].observations.is_empty(), + "a hand-built relation must not claim a source reported it" + ); +} + +#[test] +fn a_policy_name_round_trips_through_parse() { + for (name, expected) in [ + ("bycache", Policy::ByCache), + ("cache", Policy::ByCache), + ("bynode", Policy::ByNode), + ("node", Policy::ByNode), + ("bypackage", Policy::ByPackage), + ("package", Policy::ByPackage), + ("bycore", Policy::ByCore), + ("core", Policy::ByCore), + ("single", Policy::Single), + ] { + assert_eq!(Policy::parse(name), Some(expected), "for {name}"); + assert_eq!( + Policy::parse(&name.to_ascii_uppercase()), + Some(expected), + "parsing is case-insensitive, for {name}" + ); + } + assert_eq!(Policy::parse("nonesuch"), None); + assert_eq!(Policy::parse(""), None); +} + +#[test] +fn an_unreported_attribute_map_is_empty_on_the_fallback() { + let topology = machine(Vec::new()); + let (domains, _) = Policy::ByNode.select(&topology); + + match &domains[0].kind { + DomainKind::Other { name, attributes } => { + assert_eq!(name, "whole-machine"); + assert_eq!( + *attributes, + BTreeMap::new(), + "nothing was measured about it" + ); + } + other => panic!("the fallback must be an Other domain, got {other:?}"), + } +} diff --git a/crates/windows-ioring-sys/measurements/2026-09-22-append-batching/README.md b/crates/windows-ioring-sys/measurements/2026-09-22-append-batching/README.md new file mode 100644 index 000000000..654bbfede --- /dev/null +++ b/crates/windows-ioring-sys/measurements/2026-09-22-append-batching/README.md @@ -0,0 +1,111 @@ +# Append batching -- 2026-09-22 + +Twenty runs of `examples/epoch_log`, ten with each append path submitting one +write per record and ten with them batched, kept so the question `M22.1` asked +can be re-read without paying for the runs again. + +## The question + +Review finding `E-1` observed that both append paths built a `Batch`, pushed a +single write, and submitted -- so a sample whose job is to teach `Batch` never +amortised a submission. It raised a second possibility beyond the teaching +defect: that the fixed per-record submission cost was **a term every strategy +paid equally**, and therefore a shared constant capable of flattening the +three-way comparison whose headline result is that the strategies are +indistinguishable. + +`M20.6` is gated on that, so the measurement is the point of the item rather +than a bonus. + +## What was run + +The unmodified sample, ten times before the change and ten after: + +```powershell +cargo run -q -p windows-ioring-sys --example epoch_log +``` + +The "before" runs were taken by stashing the change and restoring it +afterwards, so both sets come from one machine in one sitting rather than from +two builds separated by other work. + +Each file here is one run, trimmed to the strategy block. The sample's own +workload: 32 epochs of 64 records, each run replayed. + +## Provenance + +- commit: `6cba2886` (the change under measurement was uncommitted at the time) +- CPU: AMD EPYC 7763 64-Core Processor +- OS: Microsoft Windows 11 Enterprise 10.0.26200 +- arena: 8 slots; so batching turns 64 submissions per epoch into 8 + +## What the runs show + +**Throughput: no change that can be distinguished from noise.** Median +records/sec moved by between 1% and 5%, in a spread whose run-to-run range +within a single strategy is 1.17x to 1.57x. A 5% median shift inside a 57% +range is not a result. + +**Commit p50: a real reduction, and the one finding here that separates.** For +`covering-flush` the ten before-values and the ten after-values barely overlap +-- only one after-value exceeds the lowest before-value. The same direction +holds for the other two strategies. This is the expected shape rather than a +surprise: batching removes seven of every eight `SubmitIoRing` calls from the +append path, which shortens the interval between the last append and the flush +being reached. Throughput is bound by the device flush and does not move; +latency is not, and does. + +> **Corrected 2026-09-24 by `M25.6`: the paragraph above is measuring the +> append path, not the commit.** `M20.6` established that the figure this +> harness published as "commit p50" was **entirely deferral** -- the interval +> from pushing a flush to the harness next looking, which is how long the *next +> epoch's appends* took. Batching made those appends faster, so the number fell. +> The reduction is real and its stated mechanism is even correct as written +> ("shortens the interval between the last append and the flush being reached"); +> what is wrong is the label, and therefore the conclusion that it "separates". +> +> **It is not an independent finding from the throughput result above.** Both +> are the same fact seen twice: the append path got faster, and the run is +> flush-bound, so the change appears in the metric that is not flush-bound and +> not in the one that is. Reporting one as "no change that can be distinguished +> from noise" and the other as "the one finding here that separates" reads as +> two results and is one. +> +> **What this does not disturb** is the question the capture was taken to +> answer. `E-1` asked whether a shared per-record submission cost was flattening +> the three-way comparison; the cross-strategy spread did not shrink, and that +> conclusion stands. See +> [2026-09-24-commit-decomposed/](../2026-09-24-commit-decomposed/README.md) for +> the comparison re-run once a commit could actually be measured. + +**The cross-strategy spread did not shrink.** It sits at or below the +run-to-run range of a single strategy both before and after -- which is the +sample's own stated test for whether the choice is dominated by the device +flush. + +## What this settles, and what it does not + +**Settles:** `E-1`'s second possibility is **not supported**. Removing the +shared per-record submission cost did not change the comparison, so the +sample's existing conclusion -- that the three strategies are indistinguishable +-- survives a confound that was specifically raised against it. `M20.6` can be +settled on the grounds it already had. + +> The mechanism this paragraph originally gave for that conclusion -- "because +> each pays one device flush per epoch, and everything they differ about lands +> two orders of magnitude below it" -- is the same claim corrected above, and +> `M25.5` has since measured those differences at *hundreds* of microseconds +> rather than tens. The conclusion is unchanged, because they remain smaller +> than the run-to-run spread; "below the noise" and "two orders of magnitude +> below the flush" are simply different claims, and only the first held. + +**Does not settle:** whether `CommitStrategy::AlternatingRings` earns its cost. +That is a question about what the strategy buys in *correctness* and in +bounding a commit's blast radius, not about throughput, and no number here +speaks to it. It remains `M20.6`'s to answer. + +**A caveat that belongs with the numbers rather than under them:** ten runs on +one machine with one device. The variance is wide enough that the throughput +half of this would not survive a smaller sample, and the latency half is stated +as "the distributions barely overlap" rather than as a percentage for the same +reason. diff --git a/crates/windows-ioring-sys/measurements/2026-09-22-append-batching/after/after-1.txt b/crates/windows-ioring-sys/measurements/2026-09-22-append-batching/after/after-1.txt new file mode 100644 index 000000000..b25d010f8 --- /dev/null +++ b/crates/windows-ioring-sys/measurements/2026-09-22-append-batching/after/after-1.txt @@ -0,0 +1,6 @@ +commit strategies: 32 epochs of 64 records, each run replayed + these numbers describe THIS machine and THIS device. They are printed rather than quoted in the docs because quoting ours would be misleading. + covering-flush 70356 rec/s commit p50 428 us p99 509 us max 509 us append stall 1404 us -- pays a long, ring-wide wait at every commit + host-sequenced 78229 rec/s commit p50 425 us p99 680 us max 680 us append stall 1199 us -- pays a host round trip at every epoch boundary + alternating-rings 71568 rec/s commit p50 1189 us p99 3382 us max 3382 us append stall 1361 us -- pays the arena registered on both rings, permanently + spread across strategies: 1.11x. Run this twice: if the run-to-run spread of one strategy is the same size, the choice is dominated by the device flush that all three pay once per epoch. diff --git a/crates/windows-ioring-sys/measurements/2026-09-22-append-batching/after/after-10.txt b/crates/windows-ioring-sys/measurements/2026-09-22-append-batching/after/after-10.txt new file mode 100644 index 000000000..86fffe3ca --- /dev/null +++ b/crates/windows-ioring-sys/measurements/2026-09-22-append-batching/after/after-10.txt @@ -0,0 +1,6 @@ +commit strategies: 32 epochs of 64 records, each run replayed + these numbers describe THIS machine and THIS device. They are printed rather than quoted in the docs because quoting ours would be misleading. + covering-flush 74355 rec/s commit p50 422 us p99 643 us max 643 us append stall 1354 us -- pays a long, ring-wide wait at every commit + host-sequenced 66488 rec/s commit p50 445 us p99 1526 us max 1526 us append stall 1299 us -- pays a host round trip at every epoch boundary + alternating-rings 71216 rec/s commit p50 1281 us p99 1943 us max 1943 us append stall 1404 us -- pays the arena registered on both rings, permanently + spread across strategies: 1.12x. Run this twice: if the run-to-run spread of one strategy is the same size, the choice is dominated by the device flush that all three pay once per epoch. diff --git a/crates/windows-ioring-sys/measurements/2026-09-22-append-batching/after/after-2.txt b/crates/windows-ioring-sys/measurements/2026-09-22-append-batching/after/after-2.txt new file mode 100644 index 000000000..a2e7a56ee --- /dev/null +++ b/crates/windows-ioring-sys/measurements/2026-09-22-append-batching/after/after-2.txt @@ -0,0 +1,6 @@ +commit strategies: 32 epochs of 64 records, each run replayed + these numbers describe THIS machine and THIS device. They are printed rather than quoted in the docs because quoting ours would be misleading. + covering-flush 76770 rec/s commit p50 425 us p99 528 us max 528 us append stall 1339 us -- pays a long, ring-wide wait at every commit + host-sequenced 58933 rec/s commit p50 422 us p99 699 us max 699 us append stall 1157 us -- pays a host round trip at every epoch boundary + alternating-rings 61743 rec/s commit p50 1508 us p99 2051 us max 2051 us append stall 1633 us -- pays the arena registered on both rings, permanently + spread across strategies: 1.30x. Run this twice: if the run-to-run spread of one strategy is the same size, the choice is dominated by the device flush that all three pay once per epoch. diff --git a/crates/windows-ioring-sys/measurements/2026-09-22-append-batching/after/after-3.txt b/crates/windows-ioring-sys/measurements/2026-09-22-append-batching/after/after-3.txt new file mode 100644 index 000000000..a544b7681 --- /dev/null +++ b/crates/windows-ioring-sys/measurements/2026-09-22-append-batching/after/after-3.txt @@ -0,0 +1,6 @@ +commit strategies: 32 epochs of 64 records, each run replayed + these numbers describe THIS machine and THIS device. They are printed rather than quoted in the docs because quoting ours would be misleading. + covering-flush 69956 rec/s commit p50 441 us p99 827 us max 827 us append stall 1488 us -- pays a long, ring-wide wait at every commit + host-sequenced 61519 rec/s commit p50 432 us p99 752 us max 752 us append stall 1298 us -- pays a host round trip at every epoch boundary + alternating-rings 77206 rec/s commit p50 1225 us p99 1456 us max 1456 us append stall 1327 us -- pays the arena registered on both rings, permanently + spread across strategies: 1.25x. Run this twice: if the run-to-run spread of one strategy is the same size, the choice is dominated by the device flush that all three pay once per epoch. diff --git a/crates/windows-ioring-sys/measurements/2026-09-22-append-batching/after/after-4.txt b/crates/windows-ioring-sys/measurements/2026-09-22-append-batching/after/after-4.txt new file mode 100644 index 000000000..db14c731c --- /dev/null +++ b/crates/windows-ioring-sys/measurements/2026-09-22-append-batching/after/after-4.txt @@ -0,0 +1,6 @@ +commit strategies: 32 epochs of 64 records, each run replayed + these numbers describe THIS machine and THIS device. They are printed rather than quoted in the docs because quoting ours would be misleading. + covering-flush 70675 rec/s commit p50 430 us p99 537 us max 537 us append stall 1340 us -- pays a long, ring-wide wait at every commit + host-sequenced 76030 rec/s commit p50 409 us p99 809 us max 809 us append stall 1168 us -- pays a host round trip at every epoch boundary + alternating-rings 74328 rec/s commit p50 1234 us p99 1696 us max 1696 us append stall 1373 us -- pays the arena registered on both rings, permanently + spread across strategies: 1.08x. Run this twice: if the run-to-run spread of one strategy is the same size, the choice is dominated by the device flush that all three pay once per epoch. diff --git a/crates/windows-ioring-sys/measurements/2026-09-22-append-batching/after/after-5.txt b/crates/windows-ioring-sys/measurements/2026-09-22-append-batching/after/after-5.txt new file mode 100644 index 000000000..61b08c10f --- /dev/null +++ b/crates/windows-ioring-sys/measurements/2026-09-22-append-batching/after/after-5.txt @@ -0,0 +1,6 @@ +commit strategies: 32 epochs of 64 records, each run replayed + these numbers describe THIS machine and THIS device. They are printed rather than quoted in the docs because quoting ours would be misleading. + covering-flush 48986 rec/s commit p50 456 us p99 892 us max 892 us append stall 1737 us -- pays a long, ring-wide wait at every commit + host-sequenced 72614 rec/s commit p50 412 us p99 562 us max 562 us append stall 1223 us -- pays a host round trip at every epoch boundary + alternating-rings 66333 rec/s commit p50 1297 us p99 3232 us max 3232 us append stall 1416 us -- pays the arena registered on both rings, permanently + spread across strategies: 1.48x. Run this twice: if the run-to-run spread of one strategy is the same size, the choice is dominated by the device flush that all three pay once per epoch. diff --git a/crates/windows-ioring-sys/measurements/2026-09-22-append-batching/after/after-6.txt b/crates/windows-ioring-sys/measurements/2026-09-22-append-batching/after/after-6.txt new file mode 100644 index 000000000..f4bbd5ee6 --- /dev/null +++ b/crates/windows-ioring-sys/measurements/2026-09-22-append-batching/after/after-6.txt @@ -0,0 +1,6 @@ +commit strategies: 32 epochs of 64 records, each run replayed + these numbers describe THIS machine and THIS device. They are printed rather than quoted in the docs because quoting ours would be misleading. + covering-flush 73227 rec/s commit p50 407 us p99 692 us max 692 us append stall 1391 us -- pays a long, ring-wide wait at every commit + host-sequenced 71685 rec/s commit p50 402 us p99 619 us max 619 us append stall 1125 us -- pays a host round trip at every epoch boundary + alternating-rings 60825 rec/s commit p50 1272 us p99 6613 us max 6613 us append stall 1393 us -- pays the arena registered on both rings, permanently + spread across strategies: 1.20x. Run this twice: if the run-to-run spread of one strategy is the same size, the choice is dominated by the device flush that all three pay once per epoch. diff --git a/crates/windows-ioring-sys/measurements/2026-09-22-append-batching/after/after-7.txt b/crates/windows-ioring-sys/measurements/2026-09-22-append-batching/after/after-7.txt new file mode 100644 index 000000000..69a46d970 --- /dev/null +++ b/crates/windows-ioring-sys/measurements/2026-09-22-append-batching/after/after-7.txt @@ -0,0 +1,6 @@ +commit strategies: 32 epochs of 64 records, each run replayed + these numbers describe THIS machine and THIS device. They are printed rather than quoted in the docs because quoting ours would be misleading. + covering-flush 72090 rec/s commit p50 412 us p99 783 us max 783 us append stall 1369 us -- pays a long, ring-wide wait at every commit + host-sequenced 61249 rec/s commit p50 431 us p99 1107 us max 1107 us append stall 1354 us -- pays a host round trip at every epoch boundary + alternating-rings 73008 rec/s commit p50 1263 us p99 1547 us max 1547 us append stall 1333 us -- pays the arena registered on both rings, permanently + spread across strategies: 1.19x. Run this twice: if the run-to-run spread of one strategy is the same size, the choice is dominated by the device flush that all three pay once per epoch. diff --git a/crates/windows-ioring-sys/measurements/2026-09-22-append-batching/after/after-8.txt b/crates/windows-ioring-sys/measurements/2026-09-22-append-batching/after/after-8.txt new file mode 100644 index 000000000..93d38e38e --- /dev/null +++ b/crates/windows-ioring-sys/measurements/2026-09-22-append-batching/after/after-8.txt @@ -0,0 +1,6 @@ +commit strategies: 32 epochs of 64 records, each run replayed + these numbers describe THIS machine and THIS device. They are printed rather than quoted in the docs because quoting ours would be misleading. + covering-flush 62090 rec/s commit p50 453 us p99 964 us max 964 us append stall 1538 us -- pays a long, ring-wide wait at every commit + host-sequenced 60660 rec/s commit p50 425 us p99 843 us max 843 us append stall 1227 us -- pays a host round trip at every epoch boundary + alternating-rings 56119 rec/s commit p50 1688 us p99 2607 us max 2607 us append stall 1864 us -- pays the arena registered on both rings, permanently + spread across strategies: 1.11x. Run this twice: if the run-to-run spread of one strategy is the same size, the choice is dominated by the device flush that all three pay once per epoch. diff --git a/crates/windows-ioring-sys/measurements/2026-09-22-append-batching/after/after-9.txt b/crates/windows-ioring-sys/measurements/2026-09-22-append-batching/after/after-9.txt new file mode 100644 index 000000000..e032a2ea4 --- /dev/null +++ b/crates/windows-ioring-sys/measurements/2026-09-22-append-batching/after/after-9.txt @@ -0,0 +1,6 @@ +commit strategies: 32 epochs of 64 records, each run replayed + these numbers describe THIS machine and THIS device. They are printed rather than quoted in the docs because quoting ours would be misleading. + covering-flush 76057 rec/s commit p50 405 us p99 472 us max 472 us append stall 1360 us -- pays a long, ring-wide wait at every commit + host-sequenced 69356 rec/s commit p50 407 us p99 659 us max 659 us append stall 1145 us -- pays a host round trip at every epoch boundary + alternating-rings 76059 rec/s commit p50 1226 us p99 1522 us max 1522 us append stall 1317 us -- pays the arena registered on both rings, permanently + spread across strategies: 1.10x. Run this twice: if the run-to-run spread of one strategy is the same size, the choice is dominated by the device flush that all three pay once per epoch. diff --git a/crates/windows-ioring-sys/measurements/2026-09-22-append-batching/before/before-1.txt b/crates/windows-ioring-sys/measurements/2026-09-22-append-batching/before/before-1.txt new file mode 100644 index 000000000..804edd99e --- /dev/null +++ b/crates/windows-ioring-sys/measurements/2026-09-22-append-batching/before/before-1.txt @@ -0,0 +1,6 @@ +commit strategies: 32 epochs of 64 records, each run replayed + these numbers describe THIS machine and THIS device. They are printed rather than quoted in the docs because quoting ours would be misleading. + covering-flush 73267 rec/s commit p50 445 us p99 635 us max 635 us append stall 1266 us -- pays a long, ring-wide wait at every commit + host-sequenced 68271 rec/s commit p50 443 us p99 632 us max 632 us append stall 1061 us -- pays a host round trip at every epoch boundary + alternating-rings 70079 rec/s commit p50 1353 us p99 1602 us max 1602 us append stall 1337 us -- pays the arena registered on both rings, permanently + spread across strategies: 1.07x. Run this twice: if the run-to-run spread of one strategy is the same size, the choice is dominated by the device flush that all three pay once per epoch. diff --git a/crates/windows-ioring-sys/measurements/2026-09-22-append-batching/before/before-10.txt b/crates/windows-ioring-sys/measurements/2026-09-22-append-batching/before/before-10.txt new file mode 100644 index 000000000..4f51c52bd --- /dev/null +++ b/crates/windows-ioring-sys/measurements/2026-09-22-append-batching/before/before-10.txt @@ -0,0 +1,6 @@ +commit strategies: 32 epochs of 64 records, each run replayed + these numbers describe THIS machine and THIS device. They are printed rather than quoted in the docs because quoting ours would be misleading. + covering-flush 60518 rec/s commit p50 488 us p99 1104 us max 1104 us append stall 1461 us -- pays a long, ring-wide wait at every commit + host-sequenced 56563 rec/s commit p50 483 us p99 896 us max 896 us append stall 1207 us -- pays a host round trip at every epoch boundary + alternating-rings 69375 rec/s commit p50 1333 us p99 1797 us max 1797 us append stall 1303 us -- pays the arena registered on both rings, permanently + spread across strategies: 1.23x. Run this twice: if the run-to-run spread of one strategy is the same size, the choice is dominated by the device flush that all three pay once per epoch. diff --git a/crates/windows-ioring-sys/measurements/2026-09-22-append-batching/before/before-2.txt b/crates/windows-ioring-sys/measurements/2026-09-22-append-batching/before/before-2.txt new file mode 100644 index 000000000..e5474e3d5 --- /dev/null +++ b/crates/windows-ioring-sys/measurements/2026-09-22-append-batching/before/before-2.txt @@ -0,0 +1,6 @@ +commit strategies: 32 epochs of 64 records, each run replayed + these numbers describe THIS machine and THIS device. They are printed rather than quoted in the docs because quoting ours would be misleading. + covering-flush 68885 rec/s commit p50 476 us p99 714 us max 714 us append stall 1370 us -- pays a long, ring-wide wait at every commit + host-sequenced 68786 rec/s commit p50 460 us p99 888 us max 888 us append stall 1140 us -- pays a host round trip at every epoch boundary + alternating-rings 61389 rec/s commit p50 1330 us p99 6190 us max 6190 us append stall 1297 us -- pays the arena registered on both rings, permanently + spread across strategies: 1.12x. Run this twice: if the run-to-run spread of one strategy is the same size, the choice is dominated by the device flush that all three pay once per epoch. diff --git a/crates/windows-ioring-sys/measurements/2026-09-22-append-batching/before/before-3.txt b/crates/windows-ioring-sys/measurements/2026-09-22-append-batching/before/before-3.txt new file mode 100644 index 000000000..060b5a120 --- /dev/null +++ b/crates/windows-ioring-sys/measurements/2026-09-22-append-batching/before/before-3.txt @@ -0,0 +1,6 @@ +commit strategies: 32 epochs of 64 records, each run replayed + these numbers describe THIS machine and THIS device. They are printed rather than quoted in the docs because quoting ours would be misleading. + covering-flush 63514 rec/s commit p50 476 us p99 932 us max 932 us append stall 1589 us -- pays a long, ring-wide wait at every commit + host-sequenced 66803 rec/s commit p50 461 us p99 532 us max 532 us append stall 1217 us -- pays a host round trip at every epoch boundary + alternating-rings 69920 rec/s commit p50 1285 us p99 2091 us max 2091 us append stall 1332 us -- pays the arena registered on both rings, permanently + spread across strategies: 1.10x. Run this twice: if the run-to-run spread of one strategy is the same size, the choice is dominated by the device flush that all three pay once per epoch. diff --git a/crates/windows-ioring-sys/measurements/2026-09-22-append-batching/before/before-4.txt b/crates/windows-ioring-sys/measurements/2026-09-22-append-batching/before/before-4.txt new file mode 100644 index 000000000..1d8a7f9c8 --- /dev/null +++ b/crates/windows-ioring-sys/measurements/2026-09-22-append-batching/before/before-4.txt @@ -0,0 +1,6 @@ +commit strategies: 32 epochs of 64 records, each run replayed + these numbers describe THIS machine and THIS device. They are printed rather than quoted in the docs because quoting ours would be misleading. + covering-flush 68720 rec/s commit p50 478 us p99 1000 us max 1000 us append stall 1346 us -- pays a long, ring-wide wait at every commit + host-sequenced 70786 rec/s commit p50 457 us p99 561 us max 561 us append stall 1116 us -- pays a host round trip at every epoch boundary + alternating-rings 61331 rec/s commit p50 1350 us p99 4301 us max 4301 us append stall 1279 us -- pays the arena registered on both rings, permanently + spread across strategies: 1.15x. Run this twice: if the run-to-run spread of one strategy is the same size, the choice is dominated by the device flush that all three pay once per epoch. diff --git a/crates/windows-ioring-sys/measurements/2026-09-22-append-batching/before/before-5.txt b/crates/windows-ioring-sys/measurements/2026-09-22-append-batching/before/before-5.txt new file mode 100644 index 000000000..ff20fc9b0 --- /dev/null +++ b/crates/windows-ioring-sys/measurements/2026-09-22-append-batching/before/before-5.txt @@ -0,0 +1,6 @@ +commit strategies: 32 epochs of 64 records, each run replayed + these numbers describe THIS machine and THIS device. They are printed rather than quoted in the docs because quoting ours would be misleading. + covering-flush 70090 rec/s commit p50 477 us p99 535 us max 535 us append stall 1339 us -- pays a long, ring-wide wait at every commit + host-sequenced 66580 rec/s commit p50 461 us p99 564 us max 564 us append stall 1142 us -- pays a host round trip at every epoch boundary + alternating-rings 67792 rec/s commit p50 1370 us p99 1723 us max 1723 us append stall 1312 us -- pays the arena registered on both rings, permanently + spread across strategies: 1.05x. Run this twice: if the run-to-run spread of one strategy is the same size, the choice is dominated by the device flush that all three pay once per epoch. diff --git a/crates/windows-ioring-sys/measurements/2026-09-22-append-batching/before/before-6.txt b/crates/windows-ioring-sys/measurements/2026-09-22-append-batching/before/before-6.txt new file mode 100644 index 000000000..d7972058d --- /dev/null +++ b/crates/windows-ioring-sys/measurements/2026-09-22-append-batching/before/before-6.txt @@ -0,0 +1,6 @@ +commit strategies: 32 epochs of 64 records, each run replayed + these numbers describe THIS machine and THIS device. They are printed rather than quoted in the docs because quoting ours would be misleading. + covering-flush 65599 rec/s commit p50 477 us p99 815 us max 815 us append stall 1376 us -- pays a long, ring-wide wait at every commit + host-sequenced 65573 rec/s commit p50 452 us p99 656 us max 656 us append stall 1097 us -- pays a host round trip at every epoch boundary + alternating-rings 71893 rec/s commit p50 1305 us p99 1514 us max 1514 us append stall 1356 us -- pays the arena registered on both rings, permanently + spread across strategies: 1.10x. Run this twice: if the run-to-run spread of one strategy is the same size, the choice is dominated by the device flush that all three pay once per epoch. diff --git a/crates/windows-ioring-sys/measurements/2026-09-22-append-batching/before/before-7.txt b/crates/windows-ioring-sys/measurements/2026-09-22-append-batching/before/before-7.txt new file mode 100644 index 000000000..4db1272c1 --- /dev/null +++ b/crates/windows-ioring-sys/measurements/2026-09-22-append-batching/before/before-7.txt @@ -0,0 +1,6 @@ +commit strategies: 32 epochs of 64 records, each run replayed + these numbers describe THIS machine and THIS device. They are printed rather than quoted in the docs because quoting ours would be misleading. + covering-flush 63888 rec/s commit p50 468 us p99 912 us max 912 us append stall 1298 us -- pays a long, ring-wide wait at every commit + host-sequenced 70861 rec/s commit p50 468 us p99 974 us max 974 us append stall 1113 us -- pays a host round trip at every epoch boundary + alternating-rings 67893 rec/s commit p50 1392 us p99 1726 us max 1726 us append stall 1309 us -- pays the arena registered on both rings, permanently + spread across strategies: 1.11x. Run this twice: if the run-to-run spread of one strategy is the same size, the choice is dominated by the device flush that all three pay once per epoch. diff --git a/crates/windows-ioring-sys/measurements/2026-09-22-append-batching/before/before-8.txt b/crates/windows-ioring-sys/measurements/2026-09-22-append-batching/before/before-8.txt new file mode 100644 index 000000000..cc9da0cf2 --- /dev/null +++ b/crates/windows-ioring-sys/measurements/2026-09-22-append-batching/before/before-8.txt @@ -0,0 +1,6 @@ +commit strategies: 32 epochs of 64 records, each run replayed + these numbers describe THIS machine and THIS device. They are printed rather than quoted in the docs because quoting ours would be misleading. + covering-flush 61251 rec/s commit p50 470 us p99 515 us max 515 us append stall 1327 us -- pays a long, ring-wide wait at every commit + host-sequenced 72108 rec/s commit p50 471 us p99 748 us max 748 us append stall 1137 us -- pays a host round trip at every epoch boundary + alternating-rings 67781 rec/s commit p50 1393 us p99 1680 us max 1680 us append stall 1389 us -- pays the arena registered on both rings, permanently + spread across strategies: 1.18x. Run this twice: if the run-to-run spread of one strategy is the same size, the choice is dominated by the device flush that all three pay once per epoch. diff --git a/crates/windows-ioring-sys/measurements/2026-09-22-append-batching/before/before-9.txt b/crates/windows-ioring-sys/measurements/2026-09-22-append-batching/before/before-9.txt new file mode 100644 index 000000000..e3533b344 --- /dev/null +++ b/crates/windows-ioring-sys/measurements/2026-09-22-append-batching/before/before-9.txt @@ -0,0 +1,6 @@ +commit strategies: 32 epochs of 64 records, each run replayed + these numbers describe THIS machine and THIS device. They are printed rather than quoted in the docs because quoting ours would be misleading. + covering-flush 71932 rec/s commit p50 484 us p99 907 us max 907 us append stall 1430 us -- pays a long, ring-wide wait at every commit + host-sequenced 76215 rec/s commit p50 469 us p99 597 us max 597 us append stall 1126 us -- pays a host round trip at every epoch boundary + alternating-rings 68027 rec/s commit p50 1285 us p99 4512 us max 4512 us append stall 1332 us -- pays the arena registered on both rings, permanently + spread across strategies: 1.12x. Run this twice: if the run-to-run spread of one strategy is the same size, the choice is dominated by the device flush that all three pay once per epoch. diff --git a/crates/windows-ioring-sys/measurements/2026-09-24-allocation-knee/README.md b/crates/windows-ioring-sys/measurements/2026-09-24-allocation-knee/README.md new file mode 100644 index 000000000..2342c0d51 --- /dev/null +++ b/crates/windows-ioring-sys/measurements/2026-09-24-allocation-knee/README.md @@ -0,0 +1,124 @@ +# Where the allocator's knee is, and what chunk size costs -- 2026-09-24 + +Three runs of [allocation-knee-spike.rs](../../design-sessions/spikes/allocation-knee-spike.rs), +kept because they settled a constant that had been picked rather than measured. + +## The question + +`M25.3`'s zero-fill loop needed a chunk size, and 1 MiB was chosen without +measuring. Review then offered a rule of thumb: **avoid contiguous allocations +over 64 KB from the general heap**, on the grounds that most heaps move large +blocks to a separate space anyway, and that the threshold is a useful place to +be asked "does this really need to be contiguous?" -- with the open question of +whether 64 KB is too conservative now and 1 MB is the better number. + +Two things had to be measured to answer that: where this allocator actually +changes behaviour, and whether a larger chunk fills any faster. + +## Q1 -- where does the allocator stop using the heap? + +`VirtualQuery` on each of sixty-four live allocations of one size, counting +distinct `AllocationBase` values. Blocks carved from a heap segment share their +segment's base; a block with its own reservation has a base of its own. + +| size | distinct reservations / 64 | verdict | +|---|---|---| +| 4 KB | 1 | shares a heap segment | +| 16 KB | 2 | shares a heap segment | +| 32 KB | 3 | shares a heap segment | +| 64 KB | 4 | shares a heap segment | +| 128 KB | 5 | shares a heap segment | +| 256 KB | 6 | shares a heap segment | +| 512 KB | 5 | shares a heap segment | +| **1024 KB** | **64** | **own reservation each** | +| 4096 KB | 64 | own reservation each | + +The knee is between 512 KB and 1 MB. Below it, allocations share segments; +at 1 MB and above, every allocation gets a reservation of its own. + +**Two earlier instruments failed**, and both are recorded in the spike because +both look reasonable: + +- Testing for exact 64 KB **alignment** reported "none" at every size to 4 MB. + `HeapAlloc`'s large-block path does call `VirtualAlloc`, but writes a header + at the start and returns a pointer past it -- so a large block is aligned + *plus a constant*, never exactly aligned. +- Counting distinct **offsets** within a 64 KB region declined with size (64 at + 4 KB, 15 at 1 MB) but never reached 1. Suggestive is not decisive. + +## Q2 -- does a bigger chunk fill faster? + +A 64 MiB zero-fill, seven times per chunk size, median of each run: + +| chunk | MiB/s across three runs | +|---|---| +| 4 KB | 438, 727, 708 | +| 16 KB | 1609, 1678, 1692 | +| 32 KB | 2025, 2165, 1747 | +| 64 KB | 2439, 2608, 2306 | +| 128 KB | 2368, 2733, 2704 | +| 256 KB | 2441, 2596, 2865 | +| 512 KB | 2929, 2764, 2999 | +| 1024 KB | 2817, 2668, 3297 | +| 4096 KB | 2935, 2511, 2988 | + +The climb from 4 KB to about 64 KB is large and unambiguous. Above 64 KB the +ranges overlap: 64 KB spans 2306-2608 and 4 MB spans 2511-2988. Three runs +cannot separate them. + +## What was decided + +`logfile::create_preallocated`'s chunk went from 1 MiB to **64 KiB**. On this +evidence 1 MiB was the worst of the plausible values: it is the first size with +no measurable throughput benefit *and* the first size that guarantees its own +reservation and teardown on every call. + +## What this says about the rule of thumb + +**64 KB is conservative relative to the allocator** -- the knee here is eight to +sixteen times higher than that, so the premise that "most heaps move large +allocations off to special spaces" holds, but not at 64 KB on this one. + +**It is not conservative relative to throughput**, which is what makes it a good +default anyway: past 64 KB there was nothing left to gain in this workload, so +the habit costs nothing to keep. + +**And 1 MB is a poor candidate to replace it with**, at least here -- it is +precisely the boundary. A number chosen to be "safely large" lands exactly where +the allocator's behaviour changes. + +The rule's other half is not a measurement and is not challenged by any of this: +a threshold that makes someone ask "does this really need to be contiguous?" is +doing design work, not allocator work. + +## A decision these measurements argue for and review overruled + +Everything above points at removing the allocation instead of sizing it: a +`static` array of zeros needs no heap, no knee, and no justification, and its +pages arrive demand-zero from the loader. + +**That was declined for a reason no measurement here could produce.** A `static` +sits at a fixed offset within the module, so any leak of a module base also +gives away the address of a large, writable, zero-filled region -- present for +the life of the process whether or not a log is ever opened, and no longer +protected by ASLR once the base is known. A transient heap allocation has an +unpredictable address and a lifetime bounded by the fill. + +It is recorded here, and at the constant's definition, because the `static` is +the obvious "optimization" for a reader who has only the figures on this page. + +## What this cannot tell you + +One machine, one toolchain, one allocator, one workload. Rust's default +allocator on Windows is the system heap, and a process using the segment heap, +a different global allocator, or a different Windows version may put the knee +elsewhere. The filling throughput is a buffered write to one device and says +nothing about any other. Re-run the spike rather than quoting these figures. + +## Provenance + +- Host: the development machine this repository is worked on; ARM64 Windows, + NTFS, single NUMA node. +- Date: 2026-09-24. +- Spike: [allocation-knee-spike.rs](../../design-sessions/spikes/allocation-knee-spike.rs), + built `--release` in a scratch crate. diff --git a/crates/windows-ioring-sys/measurements/2026-09-24-allocation-knee/runs.txt b/crates/windows-ioring-sys/measurements/2026-09-24-allocation-knee/runs.txt new file mode 100644 index 000000000..5ff83fcd6 --- /dev/null +++ b/crates/windows-ioring-sys/measurements/2026-09-24-allocation-knee/runs.txt @@ -0,0 +1,91 @@ + +=== run 1 === +Q1: do allocations of this size get their own reservation? + + size distinct bases/64 median region verdict + 4 KB 1 276 KB shares a heap segment + 16 KB 2 228 KB shares a heap segment + 32 KB 3 364 KB shares a heap segment + 64 KB 4 524 KB shares a heap segment + 128 KB 5 928 KB shares a heap segment + 256 KB 6 1688 KB shares a heap segment + 512 KB 5 4156 KB shares a heap segment + 1024 KB 64 1028 KB own reservation each + 4096 KB 64 4100 KB own reservation each + +Q2: does a smaller chunk fill more slowly? + + chunk median ms MiB/s (median) + 4 KB 146.1 438 + 16 KB 39.8 1609 + 32 KB 31.6 2025 + 64 KB 26.2 2439 + 128 KB 27.0 2368 + 256 KB 26.2 2441 + 512 KB 21.9 2929 + 1024 KB 22.7 2817 + 4096 KB 21.8 2935 + +Read Q2 across the rows, not against a target. What decides the chunk size is +whether throughput is still climbing at the point the allocation stops being an +ordinary heap block. +=== run 2 === +Q1: do allocations of this size get their own reservation? + + size distinct bases/64 median region verdict + 4 KB 1 276 KB shares a heap segment + 16 KB 2 228 KB shares a heap segment + 32 KB 3 364 KB shares a heap segment + 64 KB 4 524 KB shares a heap segment + 128 KB 5 928 KB shares a heap segment + 256 KB 6 1688 KB shares a heap segment + 512 KB 5 4156 KB shares a heap segment + 1024 KB 64 1028 KB own reservation each + 4096 KB 64 4100 KB own reservation each + +Q2: does a smaller chunk fill more slowly? + + chunk median ms MiB/s (median) + 4 KB 88.0 727 + 16 KB 38.1 1678 + 32 KB 29.6 2165 + 64 KB 24.5 2608 + 128 KB 23.4 2733 + 256 KB 24.7 2596 + 512 KB 23.2 2764 + 1024 KB 24.0 2668 + 4096 KB 25.5 2511 + +Read Q2 across the rows, not against a target. What decides the chunk size is +whether throughput is still climbing at the point the allocation stops being an +ordinary heap block. +=== run 3 === +Q1: do allocations of this size get their own reservation? + + size distinct bases/64 median region verdict + 4 KB 1 276 KB shares a heap segment + 16 KB 2 228 KB shares a heap segment + 32 KB 3 364 KB shares a heap segment + 64 KB 4 524 KB shares a heap segment + 128 KB 5 928 KB shares a heap segment + 256 KB 6 1688 KB shares a heap segment + 512 KB 5 4156 KB shares a heap segment + 1024 KB 64 1028 KB own reservation each + 4096 KB 64 4100 KB own reservation each + +Q2: does a smaller chunk fill more slowly? + + chunk median ms MiB/s (median) + 4 KB 90.4 708 + 16 KB 37.8 1692 + 32 KB 36.6 1747 + 64 KB 27.8 2306 + 128 KB 23.7 2704 + 256 KB 22.3 2865 + 512 KB 21.3 2999 + 1024 KB 19.4 3297 + 4096 KB 21.4 2988 + +Read Q2 across the rows, not against a target. What decides the chunk size is +whether throughput is still climbing at the point the allocation stops being an +ordinary heap block. diff --git a/crates/windows-ioring-sys/measurements/2026-09-24-commit-decomposed/README.md b/crates/windows-ioring-sys/measurements/2026-09-24-commit-decomposed/README.md new file mode 100644 index 000000000..080ded262 --- /dev/null +++ b/crates/windows-ioring-sys/measurements/2026-09-24-commit-decomposed/README.md @@ -0,0 +1,107 @@ +# The three-way comparison, with the commit decomposed -- 2026-09-24 + +Fifteen runs of `examples/epoch_log` after `M25.3` gave it a pre-allocated +`NO_BUFFERING | OVERLAPPED` handle and `M25.4` split a commit's cost into parts. +Per-run figures in [runs.tsv](runs.tsv). + +## The question + +`M20.6` found that the harness's published commit latency was **entirely +deferral** -- it measured how long the next epoch's appends took, not the +commit -- and that the three strategies were therefore being compared on a +number that could not distinguish them. It left `M25.5` to re-run the +comparison once there was a commit to measure, with one question open: whether +`AlternatingRings` earns its permanent doubled arena registration on any ground +other than blast radius, which this harness structurally cannot exhibit. + +## A correction this capture forced before it could answer anything + +`M25.4`'s first numbers had `HostSequenced` committing roughly **six times +cheaper** than the other two. That was an artifact of where the clock started. +`HostSequenced` waits for every write in userspace before pushing an unordered +flush, and the commit clock began at the *submit* -- so its host round trip fell +outside every measured part. The cost had not gone anywhere; nothing was looking +at it. + +A `prepare` part now covers whatever a strategy must do before its flush can be +pushed. With it, `HostSequenced`'s commit is not six times cheaper; it is within +noise of the others. **A reader of the uncorrected figures would have drawn the +opposite of the right conclusion**, which is the same failure mode `M20.6` +found, one layer down. + +## What fifteen runs show + +Medians, with the full range beside them: + +| | rec/s | commit p50 (us) | commit p99 (us) | +|---|---|---|---| +| covering-flush | 7066 (6499-8896) | 1203 (605-1341) | 1953 (1597-4150) | +| host-sequenced | 7570 (5854-10414) | 1090 (375-1356) | 1896 (1494-7073) | +| alternating-rings | 7477 (6441-9133) | 1149 (436-1487) | 2172 (1409-34453) | + +Where each strategy spends its commit, and how long it defers: + +| | prep p50 (us) | submit p50 (us) | deferral p50 (us) | +|---|---|---|---| +| covering-flush | 0 | 1203 (604-1340) | 6958 (2200-7750) | +| host-sequenced | 886 (173-1071) | 198 (171-272) | 6753 (1537-8237) | +| alternating-rings | 0 | 1149 (435-1487) | 15041 (4916-17650) | + +**Throughput and total commit cost: the spread across strategies is smaller +than the spread within one.** The medians sit within a few percent of each +other, every range overlaps every other range, and the run-to-run range of a +single strategy is wider than the gap between strategies. That comparison is +the test the sample's own output tells a reader to apply. + +**Where the cost sits does not vary between runs.** `HostSequenced` spends its +commit in `prep` and almost nothing in `submit`; the covering strategies do the +reverse. The ordering holds in all fifteen runs. A single blended number cannot +show this, which is what `M25.4` was for. + +**`AlternatingRings` defers about twice as long as the other two**, in every +run. The figure `M20.6` found misleading was deferral, so the strategy that +defers across two epochs reported the worst commit latency at the same +throughput as the others. + +**`block` is zero at the median for all three**, in every run. That does not +establish inline completion -- an operation that pended and finished during a +deferral of several milliseconds reads identically from here. + +## The open question, and what the data says about it + +`M25.5` asked whether `AlternatingRings` earns its doubled arena registration on +some ground other than blast radius. What the fifteen runs show about it: + +- its throughput median sits between the other two and inside both their ranges; +- its commit-cost median likewise; +- its deferral median is the longest of the three, in every run; +- the largest single commit p99 in the capture is its 34,453 us. + +**What the harness cannot show, as a matter of its construction**: a +blast-radius difference. `M20.6` established that each lane registers its own +arena of the same size, so the per-ring bound on outstanding operations is +identical whether one ring or two are used. No run of this harness can separate +the strategies on the property `AlternatingRings` exists for. + +So the record is: the ground the strategy was built on is not measurable here, +and the fifteen runs above are what was found on the grounds that are. What +follows for the strategy's status is a decision for the engineer; the conditions +under which it would pay are written down in +[strategy.rs](../../examples/epoch_log/strategy.rs). + +## What this cannot tell you + +One machine, one device, one workload of thirty-two epochs of sixty-four small +records. The run-to-run spread is large enough that a single run of this sample +says very little, which is why fifteen were taken and why the ranges are printed +beside the medians. None of it is a statement about the platform: Windows +specifies nothing about when a ring operation completes relative to +`SubmitIoRing`, and the log is required to be correct either way. + +## Provenance + +- Host: the development machine this repository is worked on; ARM64 Windows, + NTFS, single NUMA node. +- Date: 2026-09-24. +- Binary: `cargo run -p windows-ioring-sys --example epoch_log --release`, + fifteen consecutive runs on an otherwise idle machine. diff --git a/crates/windows-ioring-sys/measurements/2026-09-24-commit-decomposed/runs.tsv b/crates/windows-ioring-sys/measurements/2026-09-24-commit-decomposed/runs.tsv new file mode 100644 index 000000000..7f147dfa4 --- /dev/null +++ b/crates/windows-ioring-sys/measurements/2026-09-24-commit-decomposed/runs.tsv @@ -0,0 +1,46 @@ +run strategy rec_s commit_p50_us commit_p99_us prep_p50_us submit_p50_us block_p50_us deferral_p50_us append_stall_us +1 covering-flush 8896 1167 2406 0 1167 0 6800 158135 +1 host-sequenced 7092 1201 2116 942 216 0 7344 200234 +1 alternating-rings 6976 1214 6575 0 1214 0 15618 201159 +2 covering-flush 6931 1255 1953 0 1255 0 7266 206551 +2 host-sequenced 8588 1077 2110 854 210 0 6786 168721 +2 alternating-rings 7221 1217 1626 0 1217 0 15041 198572 +3 covering-flush 7068 1237 2437 0 1237 0 6865 204007 +3 host-sequenced 7396 1138 1717 909 195 0 6729 198324 +3 alternating-rings 7496 1122 1522 0 1122 0 15000 192467 +4 covering-flush 6741 1243 1597 0 1242 0 7750 211243 +4 host-sequenced 6724 1218 1875 1008 220 0 7260 217493 +4 alternating-rings 6607 1322 1824 0 1322 0 16801 207886 +5 covering-flush 8891 605 1788 0 604 0 2200 159109 +5 host-sequenced 5854 1356 2174 1071 272 0 8237 239652 +5 alternating-rings 6441 1487 2172 0 1487 0 17650 214630 +6 covering-flush 6682 1341 2513 0 1340 0 7636 209061 +6 host-sequenced 10414 375 2099 173 181 0 1537 161286 +6 alternating-rings 7477 1125 2127 0 1125 0 14623 195832 +7 covering-flush 6972 1234 2338 0 1234 0 6909 209974 +7 host-sequenced 6747 1241 6400 976 239 0 7434 209394 +7 alternating-rings 9133 436 34453 0 435 0 4916 132933 +8 covering-flush 7215 1197 1812 0 1196 0 6958 198256 +8 host-sequenced 7601 1067 1594 872 201 0 6711 192308 +8 alternating-rings 7477 1108 2179 0 1108 0 15110 196133 +9 covering-flush 8189 1023 4150 0 1023 0 6543 191346 +9 host-sequenced 7570 1076 1734 859 181 0 6899 192444 +9 alternating-rings 7207 1236 1448 0 1236 0 15700 199378 +10 covering-flush 6983 1203 1603 0 1203 0 7249 204256 +10 host-sequenced 8966 1079 7073 862 201 0 6665 151885 +10 alternating-rings 7646 1115 2936 0 1115 0 14391 189436 +11 covering-flush 6499 1245 1987 0 1245 0 7067 229769 +11 host-sequenced 7313 1127 1787 908 186 0 6749 201203 +11 alternating-rings 7833 1140 2008 0 1140 0 15046 184699 +12 covering-flush 6906 1225 1805 0 1225 0 7363 209302 +12 host-sequenced 7625 1096 3155 902 188 0 6753 186537 +12 alternating-rings 7517 1205 2378 0 1205 0 14547 190887 +13 covering-flush 7861 1159 1645 0 1158 0 6957 184177 +13 host-sequenced 7134 1060 1494 851 171 0 6709 210635 +13 alternating-rings 7077 1200 3470 0 1200 0 15488 204367 +14 covering-flush 7066 1203 2271 0 1203 0 6993 204001 +14 host-sequenced 7700 1056 1896 858 177 0 6580 188007 +14 alternating-rings 7545 1118 1409 0 1118 0 14803 193116 +15 covering-flush 7522 1170 1923 0 1170 0 6833 191747 +15 host-sequenced 7612 1090 1792 886 198 0 6789 192571 +15 alternating-rings 7235 1149 2308 0 1149 0 14957 201371 diff --git a/crates/windows-ioring-sys/measurements/2026-09-24-set-len-vs-zero-fill/README.md b/crates/windows-ioring-sys/measurements/2026-09-24-set-len-vs-zero-fill/README.md new file mode 100644 index 000000000..ad92ea4c6 --- /dev/null +++ b/crates/windows-ioring-sys/measurements/2026-09-24-set-len-vs-zero-fill/README.md @@ -0,0 +1,130 @@ +# `set_len` against a zero-fill -- 2026-09-24 + +Sixteen runs of the `write-pending-spike` with a fifth condition added, kept so +the question can be re-read without paying for the runs again. + +## The question + +`M25.3` opens the log over an extent that has been **zero-filled** -- a real +write of zeros -- and its documentation asserted that using +[`std::fs::File::set_len`] instead would be a silent regression, on the grounds +that only a write advances NTFS's valid data length. + +**That was asserted from documentation, not measured**, and it was challenged in +review with a specific counter-hypothesis: that one of these combinations +already does the right thing and zero-fills on the caller's behalf without +requiring the write. The spike is the apparatus that can answer it, so a +condition E was added to it rather than the claim being argued. + +## What was run + +`design-sessions/spikes/write-pending-spike.rs`, built `--release` in a scratch +crate the way `tools/run-numa-spikes.ps1` builds the NUMA spikes, and run +sixteen times back to back on an otherwise idle machine. Each run is 500 trials +per condition; each trial is 8 writes of 4096 bytes plus a flush, submitted as +one batch, and counts as "pended" if fewer than 9 completions were queued when +`SubmitIoRing` returned. + +The conditions, unchanged except for the new one: + +- **A** buffered, no `OVERLAPPED` +- **B** buffered + `OVERLAPPED` +- **C** `NO_BUFFERING` + `OVERLAPPED`, extending the file +- **D** `NO_BUFFERING` + `OVERLAPPED`, over a zero-filled extent +- **E** `NO_BUFFERING` + `OVERLAPPED`, over a `set_len` extent *(new)* + +Per-run counts are in [runs.tsv](runs.tsv). Summary over the sixteen: + +| condition | min | median | max | runs >= 250/500 | +|---|---|---|---|---| +| A buffered sync | 0 | 0 | 0 | 0/16 | +| B buffered overlapped | 0 | 0 | 1 | 0/16 | +| C nobuffer extending | 1 | 268 | 446 | 9/16 | +| D nobuffer zero-filled | 121 | 471.5 | 500 | 12/16 | +| E nobuffer set_len | 1 | 268 | 494 | 9/16 | + +## What the numbers say, and what they do not + +**Buffering is the separation that replicates.** A and B pended once in sixteen +runs between them, over 16,000 trials. Every `NO_BUFFERING` condition pended in +most runs. That is the one distinction in this table large enough to survive the +run-to-run variance. + +**C and E have identical medians and overlapping ranges, and neither is +consistently above the other** -- run 7 has C at 269 and E at 3, run 8 has C at +1 and E at 469. Nothing in this data separates them. + +**D is higher than C and E, and much less than the earlier record implies.** Its +median is around 470 against 268, and its floor over sixteen runs is 121 where +theirs is 1. But D's own range reaches down to 121, C reaches up to 446, and E +to 494, so the distributions overlap substantially and a single run of either +can land anywhere in the other's range. + +## The correction this forced + +The `M25` checklist preamble says condition D "pended reliably", and the spike's +own prose says D "was the only condition that pends". **Neither replicates.** +Both descend from a single run in which D reported 500/500 and C reported +5/500; the spike's own header already warned that two runs minutes apart gave C +as 5/500 and then 271/500. Over sixteen runs, C pends in most of them, with a +median of 268/500 against D's 471.5. + +This does not undo `M25.3`. The log is opened `NO_BUFFERING | OVERLAPPED` over a +zero-filled extent, and that configuration has the highest observed pending rate +and the highest floor of the five. What changes is the confidence the prose may +express: the original reading of "only D pends at all" is an artifact of one +run, and the honest statement is that buffering is what decides whether +operations pend here at all, while the extent's preparation shifts a rate that +varies enormously run to run for reasons outside this program. + +**And none of it is a contract.** Windows specifies nothing about when a ring +operation completes relative to `SubmitIoRing`. The log is required to be +correct whichever way it goes, which is why `M25`'s standing constraint forbids +anything depending on an operation pending -- a constraint this measurement +makes more rather than less important. + +## A follow-up question, and the answer + +Review asked the obvious next thing: since a write past the valid data length +forces the fill anyway, can that be *triggered on purpose* -- `set_len` to the +final size, then write one sector at the very end -- so the filesystem does the +zeroing and the caller never allocates a buffer? + +**Yes.** That is condition F, added after the runs above. Sixteen runs of all +six conditions are in [runs-with-touch-end.tsv](runs-with-touch-end.tsv): + +| condition | min | median | max | runs >= 250/500 | +|---|---|---|---|---| +| A buffered sync | 0 | 0 | 0 | 0/16 | +| B buffered overlapped | 0 | 0 | 0 | 0/16 | +| C nobuffer extending | 2 | 256.5 | 492 | 8/16 | +| D nobuffer zero-filled | 65 | 396.5 | 500 | 11/16 | +| E nobuffer set_len | 1 | 105 | 497 | 6/16 | +| F nobuffer set_len + touch end | 147 | 381 | 500 | 12/16 | + +F reaches the same end state the zero-fill reaches -- comparable median, and the +**highest floor of any condition measured** (147 against the zero-fill's 65). +`set_len` alone (E) remains clearly the worst of the unbuffered group, which is +what makes F interesting: the difference between E and F is one small write. + +The trade, with the cost figures from +[2026-09-24-set-len-zero-fill-cost/](../2026-09-24-set-len-zero-fill-cost/README.md): + +- F needs **no buffer at all**, two syscalls, whatever the extent's size. +- F costs about **eight times the wall time** of an explicit sequential fill for + a large extent, because the filesystem's own zeroing is much slower than + writing the same bytes. + +`logfile::create_preallocated` keeps the explicit fill, now chunked so its +memory is bounded rather than the size of the extent. A caller pre-allocating +tens of gigabytes who would rather spend wall time than write the loop has a +measured alternative. + +## Provenance + +- Host: the development machine this repository is worked on; single NUMA node, + NTFS, ARM64 Windows. +- Date: 2026-09-24. +- Spike: `design-sessions/spikes/write-pending-spike.rs` at the commit that + added condition E. +- Files written to `%TEMP%`, one per condition, recreated per run. diff --git a/crates/windows-ioring-sys/measurements/2026-09-24-set-len-vs-zero-fill/runs-with-touch-end.tsv b/crates/windows-ioring-sys/measurements/2026-09-24-set-len-vs-zero-fill/runs-with-touch-end.tsv new file mode 100644 index 000000000..d96fc5419 --- /dev/null +++ b/crates/windows-ioring-sys/measurements/2026-09-24-set-len-vs-zero-fill/runs-with-touch-end.tsv @@ -0,0 +1,17 @@ +run A_buffered_sync B_buffered_overlapped C_nobuffer_extending D_nobuffer_zerofill E_nobuffer_set_len F_nobuffer_touch_end +1 0 0 110 500 13 500 +2 0 0 244 407 462 391 +3 0 0 417 445 383 500 +4 0 0 132 500 54 500 +5 0 0 492 297 497 299 +6 0 0 62 65 2 371 +7 0 0 410 122 115 229 +8 0 0 343 500 1 193 +9 0 0 62 500 295 500 +10 0 0 469 243 493 285 +11 0 0 269 344 95 147 +12 0 0 120 500 1 429 +13 0 0 61 500 144 500 +14 0 0 347 386 3 500 +15 0 0 450 82 347 213 +16 0 0 2 109 1 343 diff --git a/crates/windows-ioring-sys/measurements/2026-09-24-set-len-vs-zero-fill/runs.tsv b/crates/windows-ioring-sys/measurements/2026-09-24-set-len-vs-zero-fill/runs.tsv new file mode 100644 index 000000000..dd3a828ca --- /dev/null +++ b/crates/windows-ioring-sys/measurements/2026-09-24-set-len-vs-zero-fill/runs.tsv @@ -0,0 +1,17 @@ +run A_buffered_sync B_buffered_overlapped C_nobuffer_extending D_nobuffer_prewritten E_nobuffer_set_len +1 0 0 2 500 426 +2 0 0 169 500 255 +3 0 0 367 121 494 +4 0 1 382 443 281 +5 0 0 109 500 374 +6 0 0 446 278 235 +7 0 0 269 500 3 +8 0 0 1 227 469 +9 0 0 288 367 32 +10 0 0 191 500 332 +11 0 0 307 203 306 +12 0 0 277 500 395 +13 0 0 267 500 1 +14 0 0 4 427 1 +15 0 0 317 500 131 +16 0 0 158 245 1 diff --git a/crates/windows-ioring-sys/measurements/2026-09-24-set-len-zero-fill-cost/README.md b/crates/windows-ioring-sys/measurements/2026-09-24-set-len-zero-fill-cost/README.md new file mode 100644 index 000000000..4a9938e10 --- /dev/null +++ b/crates/windows-ioring-sys/measurements/2026-09-24-set-len-zero-fill-cost/README.md @@ -0,0 +1,127 @@ +# What `set_len` costs, and when -- 2026-09-24 + +Three runs of [set-len-zero-fill-spike.rs](../../design-sessions/spikes/set-len-zero-fill-spike.rs), +kept because they correct a claim this repository was making from documentation +rather than measurement. + +## The question + +`M25.3` zero-fills the log's extent with a real write before opening it +`NO_BUFFERING | OVERLAPPED`. The first draft of that code's documentation +justified the choice by asserting that `set_len` would leave every write in the +extending case, and separately described `SetFileValidData` as saving "a few +milliseconds of zeroing". + +Review supplied a counter-recollection from optimizing `.cab` expansion some +years earlier: **setting the length too early forced the zero fill of large +files.** That is not the same model as "set_len defers the work", and neither +statement had been measured here, so both were put to the spike. + +## What was run + +Five cases per size, each on a fresh file in `%TEMP%`, timed separately, three +runs back to back. Full output in [runs.txt](runs.txt); one representative run: + +| size | `set_len` | zero-fill | `set_len` + write@0 | `set_len` + write@end | `set_len` + sequential fill | +|---|---|---|---|---|---| +| 64 MiB | 0 ms | 19 ms | 0 ms | 67 ms | 20 ms | +| 256 MiB | 0 ms | 88 ms | 0 ms | 347 ms | 76 ms | +| 1024 MiB | 0 ms | 315 ms | 0 ms | 2304 ms | 287 ms | + +The three runs agree closely; the 1 GiB `write@end` figure ranged 2304-2646 ms. + +## What it shows + +**`set_len` itself is free.** Under 2 ms at every size, including 1 GiB. It does +not zero eagerly. + +**A write that lands at the valid data length is free too.** Writing one sector +at offset 0 of a `set_len`'d file costs nothing, because there is no gap in +front of it. + +**A write that lands past the valid data length pays for the whole gap, +synchronously, inside that one write.** One sector at the end of a `set_len`'d +1 GiB file took roughly 2.3 seconds -- about **eight times** what writing the +entire extent sequentially costs. This is the recollection review supplied, and +it is the sharpest figure in the table: the zeroing is not avoided by `set_len`, +it is deferred onto whichever unlucky write first reaches past it, and it is far +more expensive there than it would have been up front. + +**A sequential writer pays nothing extra for having set its length first.** +Filling the extent after `set_len` costs the same as zero-filling it outright +(287 ms against 315 ms at 1 GiB), because every write lands exactly at the valid +data length and no write ever has a gap in front of it. + +## What this corrects + +- **"`SetFileValidData` saves a few milliseconds of zeroing" was wrong** by two + to three orders of magnitude. Zeroing is ~300 ms per GiB when done well and + seconds per GiB when forced onto a seeking write. That cost is the entire + reason the API exists and the reason databases hold + `SE_MANAGE_VOLUME_NAME` to use it. + +- **"`set_len` would leave every write extending" was the wrong mechanism** for + this log. For a sequential writer the zeroing cost is identical either way. + The measured reason the log zero-fills rather than `set_len`s is the *pending + rate*, which is a separate measurement in + [2026-09-24-set-len-vs-zero-fill/](../2026-09-24-set-len-vs-zero-fill/README.md), + not a zeroing cost. + +- **`set_len` as a pre-allocation strategy is safe only for strictly sequential + writers**, and is a severe footgun for anything that seeks ahead -- which is + what makes it a plausible-looking mistake rather than an obvious one. + +## What it cannot tell you + +These are wall-clock costs on one machine, one filesystem, one device. NTFS is +free to change how it tracks valid data length, and a different device would +move every figure. What the numbers support is a shape -- free, free, free, +catastrophic, free -- and the conditions under which the expensive case fires, +not a constant to design against. + +## What it costs at this sample's own sizes + +The figures above answer "what does zeroing cost per unit". Review then asked +what this sample actually pays: how large are the areas it zeroes? + +Measured at the sample's real sizes with +[prealloc-cost.rs](prealloc-cost.rs), three runs of fifty fills each, in +[at-sample-sizes.txt](at-sample-sizes.txt): + +| case | bytes | median | +|---|---|---| +| the log | 143,360 (140 KiB) | 8.4-9.5 ms | +| one strategy file | 8,421,376 (8.0 MiB) | 3.2-8.8 ms | + +One whole run pre-allocates 24.2 MiB across four files -- the log plus one file +per strategy -- for 18-36 ms of zero-filling. + +**These figures supersede an earlier capture, and are not a controlled +comparison against it.** The first capture ran a 1 MiB fill chunk where the +sample itself uses 64 KiB, so it did not measure the chunking of the code it +reports on; review caught that and the benchmark now matches the sample. The +re-run also happened on a machine that was concurrently building, so **two +things changed at once** and the difference between the two captures cannot be +attributed to the chunk size. The superseded figures are not reproduced here, +because a number nobody can act on is worse than no number. + +**At the log's own size, throughput is far below the larger case.** 140 KiB +fills at roughly 17 bytes per microsecond against 950-2250 for 8 MiB, and the +maxima over fifty fills ranged 10-54 ms for the 140 KiB case against 11-16 ms +for the 8 MiB one -- a small write whose worst case exceeds a write sixty times +its size. What the file creation and flush contribute versus the writing is not +separated by this benchmark, which times the whole operation. + +**At these sizes the three approaches differ by less than that variance.** +Explicit fill, `set_len` alone, and the touch-end trick all complete within the +spread of a single case's own repeats. The eightfold difference the table at +the top of this page shows appears at gigabyte extents, which is the scale the +sample's code is written to teach rather than the scale it runs at. + +## Provenance + +- Host: the development machine this repository is worked on; NTFS, ARM64 + Windows, single NUMA node. +- Date: 2026-09-24. +- Spike: [set-len-zero-fill-spike.rs](../../design-sessions/spikes/set-len-zero-fill-spike.rs), + built `--release` in a scratch crate; no dependencies. diff --git a/crates/windows-ioring-sys/measurements/2026-09-24-set-len-zero-fill-cost/at-sample-sizes.txt b/crates/windows-ioring-sys/measurements/2026-09-24-set-len-zero-fill-cost/at-sample-sizes.txt new file mode 100644 index 000000000..886750644 --- /dev/null +++ b/crates/windows-ioring-sys/measurements/2026-09-24-set-len-zero-fill-cost/at-sample-sizes.txt @@ -0,0 +1,28 @@ + +=== run 1 === +What the epoch-log sample's zero-fill actually costs + +case bytes median us max us +log (35 blocks) 143360 8502 54093 +one strategy file (2056 blocks) 8421376 3746 16185 + +One whole run of the sample pre-allocates 25407488 bytes (24.2 MiB) across four files, +for about 19740 us (19.7 ms) of zero-filling in total. +=== run 2 === +What the epoch-log sample's zero-fill actually costs + +case bytes median us max us +log (35 blocks) 143360 9503 13973 +one strategy file (2056 blocks) 8421376 8842 11506 + +One whole run of the sample pre-allocates 25407488 bytes (24.2 MiB) across four files, +for about 36029 us (36.0 ms) of zero-filling in total. +=== run 3 === +What the epoch-log sample's zero-fill actually costs + +case bytes median us max us +log (35 blocks) 143360 8406 9954 +one strategy file (2056 blocks) 8421376 3197 12358 + +One whole run of the sample pre-allocates 25407488 bytes (24.2 MiB) across four files, +for about 17997 us (18.0 ms) of zero-filling in total. diff --git a/crates/windows-ioring-sys/measurements/2026-09-24-set-len-zero-fill-cost/prealloc-cost.rs b/crates/windows-ioring-sys/measurements/2026-09-24-set-len-zero-fill-cost/prealloc-cost.rs new file mode 100644 index 000000000..b0b8c7e17 --- /dev/null +++ b/crates/windows-ioring-sys/measurements/2026-09-24-set-len-zero-fill-cost/prealloc-cost.rs @@ -0,0 +1,113 @@ +// Copyright (c) 2026 Mike Grier +//! What the epoch-log sample's pre-allocation actually costs at its real sizes. +//! +//! The GiB figures in the cost spike answer "what does zeroing cost per unit". +//! This answers "what does this program actually pay", which is a different +//! question and the one that decides whether any of it matters here. + +use std::fs::File; +use std::io::Write; +use std::time::Instant; + +const STRIDE: usize = 4096; +/// Matches the sample's own fill chunk (`examples/epoch_log/logfile.rs`). +/// The measured operation includes the write loop's chunking, so a benchmark +/// that chunked differently would not be measuring the implementation it +/// reports on. This was 1 MiB until review caught the mismatch. +const FILL_CHUNK: usize = 64 * 1024; + +/// Sizes this sample actually asks for, derived from main.rs's constants. +const CASES: &[(&str, usize)] = &[ + // RECORDS 24 + TAIL_RECORDS 3 + SLACK_BLOCKS 8 + ("log (35 blocks)", 35), + // EPOCHS 32 * PER_EPOCH 64 + SLACK_BLOCKS 8 + ("one strategy file (2056 blocks)", 2056), +]; + +const REPEATS: usize = 50; + +fn fill(path: &std::path::Path, len: usize) -> std::io::Result<()> { + let mut file = File::create(path)?; + let chunk = vec![0_u8; FILL_CHUNK.min(len.max(1))]; + let mut written = 0; + while written < len { + let take = chunk.len().min(len - written); + file.write_all(&chunk[..take])?; + written += take; + } + file.flush() +} + +/// The one place this program writes. +/// +/// The repository's output rule: past the first `println!`, formatting and the +/// destination become separable, so every line goes through one sink rather +/// than through call sites scattered down `main`. Small enough to be honest +/// about -- it is a writer and a `line` method -- and it is what lets the +/// destination change without touching a single caller. +struct Report { + out: W, +} + +impl Report { + fn new(out: W) -> Self { + Self { out } + } + + fn line(&mut self, args: std::fmt::Arguments<'_>) { + writeln!(self.out, "{args}").expect("the report's destination accepts writes"); + } +} + +fn main() { let mut report = Report::new(std::io::stdout().lock()); + report.line(format_args!( + "What the epoch-log sample's zero-fill actually costs\n" + )); + report.line(format_args!( + "{:<34} {:>12} {:>14} {:>14}", + "case", "bytes", "median us", "max us" + )); + + let mut total_bytes = 0_usize; + let mut total_us = 0_u128; + + for (label, blocks) in CASES { + let len = blocks * STRIDE; + let path = std::env::temp_dir().join(format!("prealloc-cost-{}.tmp", std::process::id())); + + // One untimed run so the file exists and any first-touch cost is not + // charged to the first measurement. + fill(&path, len).expect("fill"); + + let mut samples = Vec::with_capacity(REPEATS); + for _ in 0..REPEATS { + let t = Instant::now(); + fill(&path, len).expect("fill"); + samples.push(t.elapsed().as_micros()); + } + samples.sort_unstable(); + let median = samples[samples.len() / 2]; + let max = *samples.last().expect("REPEATS is not zero"); + + report.line(format_args!( + "{:<34} {:>12} {:>14} {:>14}", + label, len, median, max + )); + + let _ = std::fs::remove_file(&path); + + // The sample creates one log and three strategy files per run. + let copies = if *blocks == 35 { 1 } else { 3 }; + total_bytes += len * copies; + total_us += median * copies as u128; + } + + report.line(format_args!( + "\nOne whole run of the sample pre-allocates {} bytes ({:.1} MiB) across four files,\n\ + for about {} us ({:.1} ms) of zero-filling in total.", + total_bytes, + total_bytes as f64 / (1024.0 * 1024.0), + total_us, + total_us as f64 / 1000.0 + )); +} diff --git a/crates/windows-ioring-sys/measurements/2026-09-24-set-len-zero-fill-cost/runs.txt b/crates/windows-ioring-sys/measurements/2026-09-24-set-len-zero-fill-cost/runs.txt new file mode 100644 index 000000000..87319b0cc --- /dev/null +++ b/crates/windows-ioring-sys/measurements/2026-09-24-set-len-zero-fill-cost/runs.txt @@ -0,0 +1,46 @@ + +=== run 1 === +set_len vs zero-fill: is the zeroing avoided, or moved? + + size set_len zero-fill set_len+head set_len+tail set_len+sequential + 64 MiB 0 ms 19 ms 0 ms 67 ms 20 ms + 256 MiB 0 ms 88 ms 0 ms 347 ms 76 ms + 1024 MiB 0 ms 315 ms 0 ms 2304 ms 287 ms + +Read case 4 against case 3. Both begin from an identical set_len'd file and +write one sector. The only difference is how much unwritten extent lies between +the valid data length and the write. + +Read case 5 against case 2. Both end with the whole extent written; case 5 had +its length set first. That is a log's pattern, where every write lands at the +valid data length and no write ever has a gap in front of it. +=== run 2 === +set_len vs zero-fill: is the zeroing avoided, or moved? + + size set_len zero-fill set_len+head set_len+tail set_len+sequential + 64 MiB 0 ms 20 ms 0 ms 64 ms 18 ms + 256 MiB 0 ms 77 ms 0 ms 370 ms 72 ms + 1024 MiB 2 ms 303 ms 0 ms 2312 ms 272 ms + +Read case 4 against case 3. Both begin from an identical set_len'd file and +write one sector. The only difference is how much unwritten extent lies between +the valid data length and the write. + +Read case 5 against case 2. Both end with the whole extent written; case 5 had +its length set first. That is a log's pattern, where every write lands at the +valid data length and no write ever has a gap in front of it. +=== run 3 === +set_len vs zero-fill: is the zeroing avoided, or moved? + + size set_len zero-fill set_len+head set_len+tail set_len+sequential + 64 MiB 0 ms 19 ms 0 ms 60 ms 19 ms + 256 MiB 0 ms 73 ms 0 ms 364 ms 79 ms + 1024 MiB 0 ms 301 ms 0 ms 2646 ms 274 ms + +Read case 4 against case 3. Both begin from an identical set_len'd file and +write one sector. The only difference is how much unwritten extent lies between +the valid data length and the write. + +Read case 5 against case 2. Both end with the whole extent written; case 5 had +its length set first. That is a log's pattern, where every write lands at the +valid data length and no write ever has a gap in front of it. diff --git a/crates/windows-ioring-sys/sabotage.json b/crates/windows-ioring-sys/sabotage.json new file mode 100644 index 000000000..8d6aca190 --- /dev/null +++ b/crates/windows-ioring-sys/sabotage.json @@ -0,0 +1,738 @@ +{ + "package": "windows-ioring-sys", + "testArgs": ["test", "-p", "windows-ioring-sys", "--locked", "--all-features"], + "whyTestArgs": "--all-features, and the reason is measured rather than precautionary. The M22.2 regression case below is caught by a test gated behind `fault-injection`, which is off by default. Without the flag the harness compiled that test out, the sabotage landed in code nothing exercised, and the sweep reported SURVIVED for a defect that is in fact caught -- the same trap this repository already documents for cargo-mutants, arriving through a different tool. A feature-gated guard and an absent guard are indistinguishable to a runner that does not enable the feature.", + "description": "Sabotages for the epoch-log sample's durability contract as the program presents it (M23.1). The prose in examples/epoch_log/contract.rs is the authoritative contract; CONTRACT is its presentable reduction, and main.rs prints that reduction through Clause::ALL. These cases verify the guards over that path -- that no clause loses its section, no statement becomes unreachable, and no statement renders ragged. Run with tools/run-sabotage.ps1; see tools/README-sabotage.md for the format and for why the results are read the way they are.", + "notCoveredHere": "No sabotage here asserts that any particular *rule* is stated -- nothing looks for the word `ring`, or counts assumptions. That is deliberate and is the same reason the tests decline to: a check that the statement list still contains a given sentence is a check of the copy, not of the contract, and it passes exactly when two hand-written copies agree. The claim that the contract says the right things is carried by review of the module prose, which is the authoritative form. Also uncovered: that a new Clause variant reaches Clause::ALL. The exhaustive `match` in Clause::heading forces a new variant to be *handled* at compile time, but nothing forces it into ALL; the `a clause vanishes from the report` case below is the closest reachable proxy, since it exercises the same omission from the other direction.", + "sabotages": [ + { + "name": "a clause vanishes from the report", + "file": "examples/epoch_log/contract.rs", + "expect": "caught", + "why": "Clause::ALL drives the report's sections, so a clause dropped from it is never printed. The statements still exist, still compile, and the sample still runs -- it just silently stops telling a reader what it does not guarantee, which for a durability program is the most dangerous section to lose. This is the failure the list exists to make impossible to introduce quietly, and the one main.rs's hand-written array would have allowed.", + "find": [ + " Self::DoesNotGuarantee," + ], + "replace": [ + " Self::Assumes," + ] + }, + { + "name": "a clause is listed twice", + "file": "examples/epoch_log/contract.rs", + "expect": "caught", + "why": "A repeat in Clause::ALL prints a whole section twice. Harmless-looking and easy to introduce by editing the array rather than replacing it, and the duplicate section reads as though the contract were emphasising something.", + "find": [ + " Self::DoesNotGuarantee," + ], + "replace": [ + " Self::Guarantees," + ] + }, + { + "name": "two clauses print under one heading", + "file": "examples/epoch_log/contract.rs", + "expect": "caught", + "why": "The worst available failure for this particular program: a caller reading the `guarantees` section would be shown the things the log explicitly does NOT guarantee. Nothing about the statements changes, so every other test stays green and the output is confidently wrong.", + "find": [ + " Self::DoesNotGuarantee => \"does NOT guarantee\"," + ], + "replace": [ + " Self::DoesNotGuarantee => \"guarantees\"," + ] + }, + { + "name": "a statement renders ragged", + "file": "examples/epoch_log/contract.rs", + "expect": "caught", + "why": "The report prints each statement as ` - {text}`, so stray leading whitespace shows up as a misaligned bullet. The statement sources are wrapped across several lines with trailing backslashes, which is exactly the shape where a continuation edit adds or drops a space without anyone noticing.", + "find": [ + " text: \"durability is monotonic: if epoch N is durable, every earlier epoch is durable too\"," + ], + "replace": [ + " text: \" durability is monotonic: if epoch N is durable, every earlier epoch is durable too\"," + ] + }, + { + "name": "a statement is emptied", + "file": "examples/epoch_log/contract.rs", + "expect": "caught", + "why": "An empty statement prints as a bare dash under its heading -- a bullet that says nothing, in a list a reader is told to trust. Catches the case where a statement is gutted rather than removed.", + "find": [ + " text: \"a record reported durable is present and intact when the log is replayed\"," + ], + "replace": [ + " text: \"\"," + ] + }, + { + "name": "M23.3 spike: the unclaimed-token Drop guard is removed", + "file": "src/pending.rs", + "expect": "caught", + "why": "The whole claim of the M23.3 spike. A token dropped unclaimed leaks its buffer on purpose -- Token forgets the value rather than freeing it, because the kernel may still be writing -- so the state is invisible at runtime: the program keeps working and loses memory. A raw HashMap cannot notice. If this guard can be removed without a test objecting, the type is a convenience and not a safety improvement.", + "find": [ + " if self.entries.is_empty() || std::thread::panicking() {" + ], + "replace": [ + " if true || self.entries.is_empty() || std::thread::panicking() {" + ] + }, + { + "name": "M23.3 spike: push stops telling the oracle", + "file": "src/pending.rs", + "expect": "caught", + "why": "The other half of the spike: that one call site keeps the map and the conservation oracle in step. Twelve consumers currently drive the two by hand, which is a restatement that can drift in both directions. If the wiring can be cut without a test noticing, the map is not actually driving the oracle and the argument for the type collapses.", + "find": [ + " contract.observe_push(user_data);" + ], + "replace": [ + " let _ = user_data;" + ] + }, + { + "name": "M22.2 regression: the write result is checked before the claim", + "file": "examples/epoch_log/append.rs", + "expect": "caught", + "why": "The exact defect M22.2 found. Returning early on a failed write without claiming leaves the token unclaimed, which Token treats as still outstanding, so the arena slot is never returned -- and after SLOTS failures every append returns WouldBlock forever, somewhere else entirely and with no trace of the cause. This case was measured SURVIVING before append/tests.rs existed: the inversion compiled and every test passed, because nothing produced a failed write. It is recorded now because the injection seam makes the failure reachable, which is the difference between the leak being found in CI and being found in production. This case is also the standing evidence for M23.4, so read its failure mode and not only its verdict: the assertion in a_failed_write_still_releases_its_arena_slot fires first, names the leak, and the run ends in a clean FAILED. Before M23.4 it did not -- RegisteredBuffers Drop then fired an unguarded debug_assert during the unwind, and that second panic aborted the process, so the run ended in STATUS_STACK_BUFFER_OVERRUN with the assertion message replaced by a crash code. The defect was caught either way; what the abort destroyed was the diagnosis. A regression in either direction -- no longer caught, or caught but crashing -- is visible in this one case.", + "find": [ + " let Some((released, slot)) = self.pending.claim(completion) else {" + ], + "replace": [ + " let _written_early = completion.result()?;", + " let Some((released, slot)) = self.pending.claim(completion) else {" + ] + }, + { + "name": "M23.5: the rundown guard stops distinguishing an unwind", + "file": "src/ring.rs", + "expect": "caught", + "why": "IoRing::drop's rundown assert fires only outside an unwind (M23.4). Forcing the condition true makes it fire unconditionally, which is the pre-M23.4 behaviour that turns a failing test into STATUS_STACK_BUFFER_OVERRUN. Until M23.5 nothing caught this: suppressing both guards in IoRing::drop entirely left every test in the crate green, because no test reached either assert. a_ring_whose_rundown_the_kernel_refuses_reports_the_rundown_failure now does, by putting a null handle in the field -- measured, SubmitIoRing(null, ..) returns 0x80070006 (ERROR_INVALID_HANDLE) rather than crashing, so the failure the assert reports is a real refusal from the kernel and not a fabrication.", + "find": [ + " std::thread::panicking(),", + " \"IoRing rundown failed before close: {error}\"" + ], + "replace": [ + " true,", + " \"IoRing rundown failed before close: {error}\"" + ] + }, + { + "name": "M23.5: the close guard stops distinguishing an unwind", + "file": "src/ring.rs", + "expect": "caught", + "why": "The sibling of the case above, for the second of IoRing::drop's two asserts. Both are listed because they are separate conditions on separate lines and a change can reach one without the other -- measured: suppressing the rundown guard leaves the close test green and vice versa, so a single case would have declared the pair covered while half of it was not. The two tests are told apart by their should_panic expected strings, and that selection is itself sabotaged below.", + "find": [ + " hr >= 0 || std::thread::panicking()," + ], + "replace": [ + " true," + ] + }, + { + "name": "M23.5: the rundown test stops reaching the rundown assert", + "file": "src/ring/tests.rs", + "expect": "caught", + "why": "A sabotage of the test rather than of the code, because the thing at risk is which assert the test arrives at. run_down submits only while something is outstanding, so without the reservation the ring falls straight through to CloseIoRing and panics with 'CloseIoRing failed: 0x80070006' -- a panic, from the same Drop body, in a test that expects one. What refuses it is should_panic's expected substring, and this case is what shows that substring is load-bearing rather than decorative. Measured before being recorded: with the reservation removed the test reports `panic message: \"CloseIoRing failed: 0x80070006\"` against `expected substring: \"IoRing rundown failed before close\"`. Note the replacement also removes the outstanding==1 assertion, deliberately: leaving it in place catches the sabotage one line earlier and would prove only that the test guards its own preconditions, not that it lands on the right assert.", + "find": [ + " ring.accounting", + " .reserve_user_data()", + " .expect(\"a fresh ring's identity space is not exhausted\");", + " assert_eq!(", + " ring.accounting.outstanding(),", + " 1,", + " \"rundown must have a reason to submit, or it cannot fail\"", + " );" + ], + "replace": [ + " // nothing outstanding: rundown cannot fail, so Drop reaches the close" + ] + }, + { + "name": "M25.1: encode_block stops zeroing the block tail", + "file": "examples/epoch_log/record.rs", + "expect": "caught", + "why": "Records are variable-length but occupy a whole RECORD_STRIDE block each, and slots are reused for the log's whole life -- a NumaBuffer arrives zeroed, but only once. Without this fill the bytes past a short record are whatever the previous, longer record left there, and the write puts them on disk. Replay cannot catch it: it decodes only at block starts and takes a record's extent from its own header, so a stale fragment past a short record's end is never read. That is why the guard asserts on the file's bytes rather than on a replay outcome. The composition lives in one function for both writers precisely so this case covers the harness lane as well as the log's appender; before encode_block existed the two carried the step separately.", + "find": [ + " slot[total..RECORD_STRIDE].fill(0);" + ], + "replace": [ + " // zeroing removed" + ] + }, + { + "name": "M25.2: replay walks by the record extent instead of the stride", + "file": "examples/epoch_log/replay.rs", + "expect": "caught", + "why": "The reader half of the stride, and the exact defect that sat in the tree between M25.1 and M25.2. A strided writer with an unstrided reader lands the cursor in a zeroed block tail, decodes NeverWritten, and reports every record after the first as a missing durable record -- the log looks catastrophically broken while being perfectly fine. Measured: with the writer converted and the reader not, all 21 example tests still passed and only running the sample caught it. No CI job runs the sample, which is why the end-to-end tests this case fires were added rather than relying on that.", + "find": [ + " cursor += record::RECORD_STRIDE;" + ], + "replace": [ + " cursor += found.extent();" + ] + }, + { + "name": "M25.1: the harness lane packs its record offsets", + "file": "examples/epoch_log/strategy.rs", + "expect": "caught", + "why": "The sample has two writers over one on-disk format, and this is the second. Listed separately from the appender's offset advance below because they are genuinely separate facts at separate sites: measured, reverting this one leaves every appender test green and vice versa, so a single case would have reported the pair covered while half of it was not. This case was SURVIVING when the stride was first introduced -- nothing tested the harness lane at all, and the end-to-end test in strategy/tests.rs exists because this sabotage said so.", + "find": [ + " written += record::RECORD_STRIDE as u64;" + ], + "replace": [ + " written += record::HEADER_LEN as u64;" + ] + }, + { + "name": "M25.1: the appender packs its record offsets", + "file": "examples/epoch_log/append.rs", + "expect": "caught", + "why": "The log's own writer, the sibling of the harness case above. Advancing by less than a stride makes consecutive records overlap inside one block, so the guards fire for two independent reasons at once -- the layout assertion and the zeroed-tail assertion both go red -- which is worth knowing when reading the result: this case being caught does not on its own establish that the stride is what caught it. The narrower evidence for that is the reader case above, which moves only one number. The replacement uses HEADER_LEN rather than the record's own length because this call site no longer binds one: M25.1 ended by removing the `total` binding to clear an unused-variable warning, which broke this case's patch and was reported by the harness as MANIFEST DOES NOT COMPILE. A manifest is a restatement site like any other and has to be swept when the code it patches moves.", + "find": [ + " self.next_offset += record::RECORD_STRIDE as u64;" + ], + "replace": [ + " self.next_offset += record::HEADER_LEN as u64;" + ] + }, + { + "name": "M25.3: the extent is not pre-allocated", + "file": "examples/epoch_log/logfile.rs", + "expect": "caught", + "why": "Pre-allocation is the whole point of the item: the spike measured six configurations and the zero-filled extent has both the highest central tendency and a floor far above the unfilled ones, because the filesystem serialises writes past the valid-data length. An extending write is therefore the configuration that looks like it works and measures like a buffered one, which is exactly the failure mode that needed a guard rather than a comment. The patch seeds the fill loop as already-complete, so the file is created and left empty; it used to replace a single std::fs::write and was updated when that became a chunked loop.", + "find": [ + " let mut written = 0;" + ], + "replace": [ + " let mut written = len;" + ] + }, + { + "name": "M25.3: FILE_FLAG_NO_BUFFERING is dropped", + "file": "examples/epoch_log/logfile.rs", + "expect": "caught", + "why": "Dropping this flag changes nothing a caller can observe except the thing the item exists for, which is why it needs a guard that is not a measurement. The test pushes an aligned write and an unaligned one through a real ring -- the way production reaches this handle -- and requires the first to be accepted and the second refused. Both directions are asserted deliberately: a guard that only showed the aligned write succeeding would pass just as happily against a buffered handle. Note the unaligned case breaks two of NO_BUFFERING's three rules at once, since a Vec guarantees neither a sector address nor a sector length, so it establishes that the handle has rules a buffered one does not rather than which rule refused it.", + "find": [ + " .custom_flags(FILE_FLAG_NO_BUFFERING | FILE_FLAG_OVERLAPPED)" + ], + "replace": [ + " .custom_flags(FILE_FLAG_OVERLAPPED)" + ] + }, + { + "name": "M25.3: FILE_FLAG_OVERLAPPED is dropped", + "file": "examples/epoch_log/logfile.rs", + "expect": "caught", + "why": "Listed separately from the NO_BUFFERING case because they are separate flags with separate guards: measured, dropping either leaves the other's test green. This one is caught by a synchronous write_all being refused, which is the only property of the handle's mode reachable from here -- GetFileInformationByHandleEx does not report it, and the ring works on synchronous and asynchronous handles alike. That test began as an attempt to check alignment through std::io::Write and failed on the *aligned* write, which is how the discriminator was found rather than designed.", + "find": [ + " .custom_flags(FILE_FLAG_NO_BUFFERING | FILE_FLAG_OVERLAPPED)" + ], + "replace": [ + " .custom_flags(FILE_FLAG_NO_BUFFERING)" + ] + }, + { + "name": "DECLARED BLIND SPOT: set_len replaces the zero-fill", + "file": "examples/epoch_log/logfile.rs", + "expect": "survives", + "why": "This is NOT an inert change and must not be read as one -- it is a real regression that nothing in this workspace detects, recorded here so the gap is tracked rather than invisible. set_len and a zero-fill both produce a file of the right size whose bytes read back as zero, because reads past the valid data length are answered with zeros the filesystem synthesises without touching the disk -- so the two are indistinguishable to every assertion available here. The difference is measured rather than argued, in measurements/2026-09-24-set-len-vs-zero-fill: over sixteen runs the zero-filled extent pended at a median of 471/500 against set_len's 268/500, with a floor of 121 against 1. Read that capture before trusting this note, because it also corrects a stronger claim this repository had been repeating -- that the zero-filled extent was the only condition that pended at all, which descends from one run and does not replicate. Note what is NOT the reason: the zeroing cost is identical for a sequential writer, measured in measurements/2026-09-24-set-len-zero-fill-cost. This case stays untested because the difference is a RATE that varies enormously run to run, and a test asserting it would assert an observation about one machine as though it were a contract, which M25's standing constraint forbids.", + "find": [ + " let mut file = File::create(path)?;", + " let chunk = vec![0_u8; FILL_CHUNK.min(len.max(1))];", + " let mut written = 0;", + " while written < len {", + " let take = chunk.len().min(len - written);", + " file.write_all(&chunk[..take])?;", + " written += take;", + " }", + " file.flush()?;" + ], + "replace": [ + " let file = File::create(path)?;", + " file.set_len(len as u64)?;" + ] + }, + { + "name": "M25.1b: the torn-tail case stops tearing anything", + "file": "examples/epoch_log/main.rs", + "expect": "caught", + "why": "The sample's torn-tail replay exists to show the reader tolerates a partial record, which the contract requires. Reverting the cut to the old file-length form makes it land in a zeroed block tail -- or, since M25.3, in the pre-allocated slack -- so every record is whole and the case demonstrates the opposite of what it claims while still passing is_clean and the durable count. This was measured SURVIVING even after M25.1b's end-to-end test was added: the hazard had been written into the comment above the cut and describing it caught nothing. The assertions on tail_stopped and tail_records are what turned the description into a check, and this case is what shows they are load-bearing. Only the end-to-end test catches it; the M25.1 layout tests all stay green.", + "find": [ + " let last_record_start = (run.durable_records + run.tail_records - 1) * record::RECORD_STRIDE;", + " let torn_at = last_record_start + record::HEADER_LEN + 4;" + ], + "replace": [ + " let torn_at = bytes.len() - (record::HEADER_LEN + 4);" + ] + }, + { + "name": "M25.1b: the negative control stops corrupting the durable region", + "file": "examples/epoch_log/main.rs", + "expect": "caught", + "why": "The negative control is what makes the other two replay passes mean anything: it corrupts one byte INSIDE the durable region and requires replay to report it. Moving the victim byte into the file's unwritten tail leaves the durable region intact, so replay correctly finds no violation and the control silently stops controlling -- a verifier that cannot fail, which is exactly the thing replay.rs's own module docs say would make the whole sample decoration. Caught only by the end-to-end test, because nothing else in the suite runs the sample's verify().", + "find": [ + " let victim = record::HEADER_LEN + 2;" + ], + "replace": [ + " let victim = bytes.len() - 1;" + ] + }, + { + "name": "M25.4: flush() folds the deferral back in", + "file": "examples/epoch_log/strategy.rs", + "expect": "caught", + "why": "The exact defect M20.6 found, re-expressible in one line once the parts exist. A commit's reported cost is its submit plus its wait; adding the deferral back produces a number that grows when a design defers further while being no slower, which is what the harness published for months and what a reader took for commit latency. The guard is an identity rather than a value, because which part carries the cost is a property of the machine and M25's standing constraint forbids depending on it.", + "find": [ + " self.prepare + self.submit + self.blocking", + " }" + ], + "replace": [ + " self.prepare + self.submit + self.blocking + self.deferral", + " }" + ] + }, + { + "name": "M25.4: the deferral stops being measured", + "file": "examples/epoch_log/strategy.rs", + "expect": "caught", + "why": "A part that always reads zero satisfies every identity the split can assert while making the decomposition a rename. Measured SURVIVING against the identity assertions alone, which is why the test also requires some sample of deferral and of submit to be non-zero. Stated as 'some sample is non-zero' rather than a lower bound on a duration: this harness defers by construction, pushing the next epoch's appends before settling the previous commit, so a run in which nothing deferred means the clock is not running rather than that the machine was fast. `blocking` deliberately gets no such guard -- zero is a legitimate and frequently observed reading for it.", + "find": [ + " let deferral = submitted_at.elapsed();" + ], + "replace": [ + " let deferral = Duration::ZERO;" + ] + }, + { + "name": "M25.4: the submit stops being timed", + "file": "examples/epoch_log/strategy.rs", + "expect": "caught", + "why": "The sibling of the case above, for the other part that must not silently read zero. Listed separately because they are measured at different sites -- the submit around the commit call, the deferral at the settle -- and one can be lost without the other. This is the part that matters most on a handle where the operation completes inline, since the device round trip lands here and nowhere else.", + "find": [ + " deferred[lane_index] = Some((user_data, prepare, submitting.elapsed(), Instant::now()));" + ], + "replace": [ + " deferred[lane_index] = Some((user_data, prepare, Duration::ZERO, Instant::now()));" + ] + }, + { + "name": "M25.5: the host round trip falls outside the commit's measured cost", + "file": "examples/epoch_log/strategy.rs", + "expect": "caught", + "why": "The defect M25.5 found in M25.4's own numbers. HostSequenced waits for every write in userspace before pushing an unordered flush; starting the commit clock at the submit puts that round trip outside every measured part, and the strategy then reports a commit roughly six times cheaper than the covering ones while doing the same work where nothing is looking. A reader comparing the published figures would have drawn the opposite of the right conclusion -- the same failure mode M20.6 found, one layer down. The patch moves the clock back to after the preparation, which is exactly the shape the defect had. The guard asserts only that a non-zero prepare is observed, because what is pinned is the measurement boundary and not how expensive any strategy is on a given machine.", + "find": [ + " let prepare = preparing.elapsed();" + ], + "replace": [ + " let prepare = Duration::ZERO;" + ] + }, + { + "name": "M25.7: the log digest stops reading the bytes", + "file": "examples/epoch_log/record.rs", + "expect": "caught", + "why": "M25.7 replaced the cross-strategy byte comparison with a digest, which retains thirty-two bits instead of eight megabytes and is a WEAKER check -- two different logs can in principle share a digest where two different byte arrays cannot share their bytes. This sabotages the weakening directly: a digest that folds only the length still distinguishes logs of different sizes, so a guard that tested only 'a dropped record changes it' would pass. The flipped-byte test is what fails, and the two length-based tests deliberately do not, which is the separation worth having -- each guard covers a different way the digest could stop discriminating. What none of them establish is that no two logs collide; FNV-1a over 32 bits has collisions and the definition says so.", + "find": [ + " hash ^= u32::from(byte);", + " hash = hash.wrapping_mul(PRIME);", + " }", + " hash", + "}" + ], + "replace": [ + " let _ = byte;", + " hash = hash.wrapping_mul(PRIME);", + " }", + " hash", + "}" + ] + }, + { + "name": "M26.2: the seam consults the responder but ignores its answer", + "file": "src/sys.rs", + "expect": "caught", + "why": "The seam's whole purpose is that an installed responder answers INSTEAD of the kernel. A version that consults it and then makes the real call anyway would look right in every review -- the responder is called, its state advances, a counting test still passes -- while the resolver M26.3 builds on it would silently measure the kernel rather than the space. The guard is that the seam returns what the responder said. Note the sabotage fails safely rather than crashing, and for a reason M23.5 measured: the test passes a null ring handle, and SubmitIoRing(null) returns 0x80070006 rather than dereferencing it, so the fall-through reaches Win32 and is refused instead of faulting. The tests that only count calls stay green under this patch, which is why the assertion is on the returned HRESULT and not on the count.", + "testArgs": ["test", "-p", "windows-ioring-sys", "--locked", "--all-features"], + "find": [ + " if let Some(answer) = $crate::sys::installed::with(|r| unsafe { r.$method($($arg),*) })", + " {", + " return answer;", + " }" + ], + "replace": [ + " let _ = $crate::sys::installed::with(|r| unsafe { r.$method($($arg),*) });" + ] + }, + { + "name": "M26.3: the resolver stops enforcing RS-C-4's drain half", + "file": "src/sys/resolver.rs", + "expect": "caught", + "why": "RS-C-4 is the one place RESPONSE-SPACE.md is narrower than 'anything may happen', and it is the clause this crate's entire durability story rests on. A resolver that quietly dropped it would be WIDER than the specification -- which reads like rigour, so a reviewer is unlikely to object -- and every consumer written against it would be hardened for a platform that cannot exist while the real guarantee went untested. The guard is that a barrier is eligible only as the oldest unresolved operation; this patch makes every operation eligible always.", + "testArgs": ["test", "-p", "windows-ioring-sys", "--locked", "--all-features"], + "find": [ + " .filter(|&i| i == 0 || !self.pool[i].barrier)" + ], + "replace": [ + " .filter(|&_i| true)" + ] + }, + { + "name": "M26.3: the resolver stops honouring RS-C-2's operation identity", + "file": "src/sys/resolver.rs", + "expect": "caught", + "why": "RS-C-2 says a completion carries back exactly the user data its operation was built with, and D-4 records that this crate's whole accounting model rests on it. A resolver that returned a plausible-but-wrong identity -- here, the value one operation later -- would still conserve completions, still satisfy every count, and still pass any test that only asks how many arrived. The guard is that the identity is echoed rather than computed.", + "testArgs": ["test", "-p", "windows-ioring-sys", "--locked", "--all-features"], + "find": [ + " UserData: user_data," + ], + "replace": [ + " UserData: user_data.wrapping_add(1)," + ] + }, + { + "name": "M26.3: the resolver lets an unsubmitted operation complete (RS-C-3)", + "file": "src/sys/resolver.rs", + "expect": "caught", + "why": "RS-C-3 is implemented by exactly one thing -- only the submitted pool is ever eligible to post -- so it is the constraint most easily lost to an edit that looks like a simplification. A resolver that posted from the staging area would make the Build*/Submit boundary meaningless, which is the reason the clause exists at all. This patch resolves the staged operations on a pop, before any submit has carried them.", + "testArgs": ["test", "-p", "windows-ioring-sys", "--locked", "--all-features"], + "find": [ + " if self.posted.is_empty() {", + " // Starvation rescue only: no coins. See the module's note on the", + " // two clocks -- flipping here would make `RS-P-1`'s pending case", + " // unobservable to a polling consumer.", + " self.tick(false);" + ], + "replace": [ + " if self.posted.is_empty() {", + " self.pool.append(&mut self.staged);", + " self.tick(false);" + ] + }, + { + "name": "M26.3: the resolver permits RS-P-2 but never exercises it", + "file": "src/sys/resolver.rs", + "expect": "caught", + "why": "The failure this whole file is built to avoid, and the one a green suite hides best: a permission that is configurable, documented, cited, and never taken. Such a resolver passes every test written against what it MAY do, and the defect class RS-P-2 exists for -- D-47's, where operations queued after a drained flush completed before it -- goes unswept. The guard is a test asserting that reordering actually HAPPENED across a majority of seeds, not merely that it was allowed.", + "testArgs": ["test", "-p", "windows-ioring-sys", "--locked", "--all-features"], + "find": [ + " let draw = self.next() % eligible.len() as u64;", + " eligible[usize::try_from(draw).expect(\"a modulus of a usize length fits a usize\")]" + ], + "replace": [ + " let _ = self.next();", + " eligible[0]" + ] + }, + { + "name": "M26.3: the completion signal becomes level- rather than edge-triggered (RS-P-6)", + "file": "src/sys/resolver.rs", + "expect": "caught", + "why": "RS-P-6 carries no over-provision: it IS D-19's measurement, that the completion event fires on the empty-to-non-empty edge. Signalling every post is strictly more generous, so no consumer would fail against it -- which is exactly why it must be caught here. A resolver that over-signalled would let a consumer depend on one wakeup per completion, and D-21's auto-reset, one-waiter-per-ring design was drawn from the opposite fact.", + "testArgs": ["test", "-p", "windows-ioring-sys", "--locked", "--all-features"], + "find": [ + " if !self.event.is_null() && (was_empty || !self.config.edge_triggered_signal) {" + ], + "replace": [ + " if !self.event.is_null() {" + ] + }, + { + "name": "M26.3: the barrier-hold counter stops counting", + "file": "src/sys/resolver.rs", + "expect": "caught", + "why": "Sabotaging the CLAIM rather than the symptom. RS-C-4's test could pass vacuously -- a barrier with nothing in front of it satisfies the constraint trivially -- so there is a second test asserting the constraint actually BIT, and it reads this counter. A counter that silently stopped incrementing would turn that guard into a tautology while leaving the reassuring test name in place. This is the instrument-checking-the-instrument case the sabotage harness exists for.", + "testArgs": ["test", "-p", "windows-ioring-sys", "--locked", "--all-features"], + "find": [ + " if held > 0 {", + " self.record(|s| s.barrier_holds += 1);", + " }" + ], + "replace": [ + " let _ = held;" + ] + }, + { + "name": "M26.3: set_completion_event is forwarded to the kernel instead of captured", + "file": "src/sys/resolver.rs", + "expect": "caught", + "why": "M26.3 moved this call behind the seam specifically so RS-P-6 could be implemented; before that the resolver had no event to signal and every EventDelivery consumer would have parked forever. A version that forwards to the real ring restores exactly that state, and does it in the most plausible-looking way available -- forwarding is what every other default method does. The failure it produces is a hang rather than an assertion, which is why the guard counts signals rather than waiting for one.", + "testArgs": ["test", "-p", "windows-ioring-sys", "--locked", "--all-features"], + "find": [ + " self.event = event;", + " S_OK" + ], + "replace": [ + " let _ = event;", + " S_OK" + ] + }, + { + "name": "DECLARED BLIND SPOT: the resolver's failure-code draw narrows to one facility", + "file": "src/sys/resolver.rs", + "expect": "survives", + "why": "RS-P-3 permits any error code and enumerates none, warning that a consumer must not depend on the set being small. There is a test that the set is not small -- it demands many more than a handful of distinct codes across the whole sweep -- but it cannot check the codes are DISTRIBUTED sensibly, because the space deliberately specifies no distribution and inventing one here would be the recording the document exists to avoid. Masking to 15 bits instead of 16 halves the range and still clears the threshold, so this survives. It is recorded rather than left invisible: if a later change makes the distribution observable, this case reports the discrepancy.", + "testArgs": ["test", "-p", "windows-ioring-sys", "--locked", "--all-features"], + "find": [ + " let code = self.next() & 0xFFFF;" + ], + "replace": [ + " let code = self.next() & 0x7FFF;" + ] + }, + { + "name": "M26.4: the ring stops decrementing its outstanding count", + "file": "src/accounting.rs", + "expect": "caught", + "why": "P-4 says IoRing::outstanding is accurate, and P-2 says every loop terminates. A count that never comes down breaks both at once -- a consumer sizing a wait or deciding whether blocking is worthwhile reads a number that only grows, and run_down never finishes. Scoped to the property suite on purpose: this mutation is caught by a great many tests in this crate, and 'something went red' would not show that THESE properties are the ones watching it.", + "testArgs": ["test", "-p", "windows-ioring-sys", "--locked", "--all-features", "--test", "properties_under_every_resolution"], + "find": [ + " pub(crate) fn record_completion(&mut self) {", + " self.outstanding = self.outstanding.saturating_sub(1);" + ], + "replace": [ + " pub(crate) fn record_completion(&mut self) {" + ] + }, + { + "name": "M26.4: operation identities stop being unique", + "file": "src/accounting.rs", + "expect": "caught", + "why": "D-4 rests this crate's whole accounting model on a completion identifying its operation. An identity that repeats makes two operations indistinguishable, which is exactly the conservation failure RingContract reports as a duplicate completion. THIS IS ALSO THE EVIDENCE THAT P-1'S VERDICT IS READ rather than merely computed: measured, the suite fails with the oracle's own wording -- 'completed more than once, but every queued SQE produces exactly one completion' -- so the path from the crate through RingContract to a red test is traversed end to end. A separate case patching the verdict check was tried and removed as unsound: disabling an assertion that does not fire on a green baseline cannot fail, so it measured nothing. Scoped to the property suite on purpose, because this mutation is caught by a great many tests in this crate and 'something went red' would not show that THESE properties are watching it.", + "testArgs": ["test", "-p", "windows-ioring-sys", "--locked", "--all-features", "--test", "properties_under_every_resolution"], + "find": [ + " self.next_user_data = id", + " .checked_add(1)", + " .ok_or_else(|| io::Error::other(\"IoRing operation identity space exhausted\"))?;" + ], + "replace": [ + " self.next_user_data = id;" + ] + }, + { + "name": "M26.4: coverage stops accumulating, so the suite can pass vacuously", + "file": "tests/properties_under_every_resolution.rs", + "expect": "caught", + "why": "Sabotaging the CLAIM rather than a symptom. This suite's five properties are all satisfied trivially by a run that does nothing, so the coverage counters are what separate 'the properties held' from 'nothing reached the states they are about'. A fold that silently stopped accumulating would leave every property assertion in place and turn the whole file into a tautology -- the instrument-checking-the-instrument case, and the reason the counters are asserted rather than merely printed.", + "testArgs": ["test", "-p", "windows-ioring-sys", "--locked", "--all-features", "--test", "properties_under_every_resolution"], + "find": [ + " self.pushes += other.pushes;" + ], + "replace": [ + " let _ = other.pushes;" + ] + }, + { + "name": "M26.5: M21.6's defect re-injected -- an expired wait is a failure again", + "file": "src/ring.rs", + "expect": "caught", + "why": "The historical defect this crate actually shipped, restored. Until M21.6, wait_outcome passed IORING_E_WAIT_TIMEOUT to check() like any other HRESULT, so pop_within returned Err on every ordinary timeout and run_down treated any operation slower than its poll as fatal -- which M21.6's own doc comment records as the shape that made run_down return with the operation still outstanding, after which Drop asserted and closed the ring anyway. This is one half of M26.5's calibration and it is the half that matters most, because a resolver nobody has shown can go red on a real defect is not evidence. Measured: the calibration fails naming the seed and 0x800705B4. Scoped to the calibration target so 'caught' means THAT suite caught it, not the several others that would also notice.", + "testArgs": ["test", "-p", "windows-ioring-sys", "--locked", "--all-features", "--test", "calibration"], + "find": [ + "fn wait_outcome(hr: windows_sys::core::HRESULT) -> io::Result<()> {", + " if hr == IORING_E_WAIT_TIMEOUT {", + " return Ok(());", + " }", + " check(hr)" + ], + "replace": [ + "fn wait_outcome(hr: windows_sys::core::HRESULT) -> io::Result<()> {", + " check(hr)" + ] + }, + { + "name": "M26.5: the resolver enforces the hold-back D-47 withdrew", + "file": "src/sys/resolver.rs", + "expect": "caught", + "why": "The other half of M26.5's calibration, and it runs the opposite way to every other case here: it does not break the crate, it makes the INSTRUMENT go narrow. D-24 claimed a covering flush holds back what follows and D-47 withdrew that over roughly 4,500 trials, so a resolver that enforced both halves of the barrier would be narrower than the space and would report green on exactly the defect class the one-sidedness exists to expose. That is the M17.4 failure mode -- a suite sampling the right state while being insensitive to the defect living in it -- and nothing detects it except a test that demands the sensitivity. Measured: the calibration fails reporting that no seed in its sweep broke a consumer that assumes a covering flush holds back what follows it -- the count is left out of this sentence deliberately, since the sweep size is a constant that has already moved once. Note this patches the same line as M26.3's RS-C-4 case but in the opposite direction: that one removes the constraint, this one over-applies it.", + "testArgs": ["test", "-p", "windows-ioring-sys", "--locked", "--all-features", "--test", "calibration"], + "find": [ + " .filter(|&i| i == 0 || !self.pool[i].barrier)" + ], + "replace": [ + " .filter(|&i| i == 0 || !self.pool[..=i].iter().any(|op| op.barrier))" + ] + }, + { + "name": "M26.6: a constraint loses its CONFIRMS marker", + "file": "tests/flush_barrier.rs", + "expect": "caught", + "why": "RESPONSE-SPACE.md justifies forbidding the resolver from ever violating RS-C-4 on the grounds that the kernel tests cover it. If the only test claiming that clause stops claiming it, the constraint is untested on both sides at once -- the resolver cannot produce it and nothing watches for it -- and the document's justification silently becomes false. This is the hole M26.6 exists to close, and the census is what closes it.", + "testArgs": ["test", "-p", "windows-ioring-sys", "--locked", "--all-features", "--test", "response_space_census"], + "find": [ + "//! CONFIRMS: RS-C-4" + ], + "replace": [ + "//! The drain clause is confirmed below." + ] + }, + { + "name": "M26.6: the census counts a mention as a claim", + "file": "tests/response_space_census.rs", + "expect": "caught", + "why": "THE DEFECT THIS CENSUS SHIPPED WITH, AND THE REASON IT WAS REBUILT. The first version searched each file for the clause ID anywhere in its text. Measured: removing RS-C-4's check from the only test performing it did NOT turn it red, because a second file mentioned that clause only to say the check was somebody else's -- and under a substring search a disclaimer is indistinguishable from a claim. That is the same trap this repository recorded once before, where a bare substring matched a probe whose only mention of a tag was a comment. A marker is a claim; a mention is not, and this patch collapses the distinction again.", + "testArgs": ["test", "-p", "windows-ioring-sys", "--locked", "--all-features", "--test", "response_space_census"], + "find": [ + " let at = line.find(marker)?;", + " let id = line[at + marker.len()..].trim();" + ], + "replace": [ + " let at = line.find(marker).unwrap_or(0);", + " let id = line[at..].trim().trim_start_matches(marker).trim();" + ] + }, + { + "name": "DECLARED BLIND SPOT: the RS-C-4 assertion is made vacuous while its marker stays", + "file": "tests/flush_barrier.rs", + "expect": "survives", + "why": "The honest limit of a source census, recorded rather than left invisible. The census proves a clause is CLAIMED by a file; it cannot prove the file's assertion still runs or still means anything. Comparing a constant against itself leaves the CONFIRMS marker in place and the test green, so nothing here notices. Closing this would need the assertion's execution to be observable from outside the test process, which is machinery this milestone judged disproportionate -- and the calibration discipline covers the same ground from the other direction, since M26.5 shows the instruments go red on real defects. If a cheap way to observe it appears, this case reports the discrepancy.", + "testArgs": ["test", "-p", "windows-ioring-sys", "--locked", "--all-features", "--test", "flush_barrier"], + "find": [ + " covering.a_after_flush, 0," + ], + "replace": [ + " 0, 0," + ] + }, + { + "name": "DECLARED BLIND SPOT: dropping the covering flag does not fire RS-C-4 on this machine", + "file": "tests/flush_barrier.rs", + "expect": "survives", + "why": "A hardware-dependent limit on the conformance assertion's SENSITIVITY, measured rather than assumed. Running the treatment case with FlushCoverage::Unordered removes the barrier entirely, and on a machine where the device stack orders a flush behind that file's outstanding writes by itself -- which this file's own header documents, and which is the case here -- no preceding write completes after the flush even unflagged, so the assertion stays green. It is therefore NOT a portable demonstration that the assertion can go red: on the spike's machine it would be caught, on this one it survives. Recorded because a reader would otherwise reasonably expect this to be the calibration for RS-C-4, and it cannot be. What the assertion still does everywhere is report a violation if one occurs, which is the conformance job; what it cannot do everywhere is prove it would notice.", + "testArgs": ["test", "-p", "windows-ioring-sys", "--locked", "--all-features", "--test", "flush_barrier"], + "find": [ + " let covering = run_case(&mut ring, handle, FlushCoverage::CoversPrecedingOperations);" + ], + "replace": [ + " let covering = run_case(&mut ring, handle, FlushCoverage::Unordered);" + ] + }, + { + "name": "M26.7: a kernel test goes back to asserting a completion is already poppable", + "file": "tests/ring_lifecycle.rs", + "expect": "caught", + "why": "The audited defect class, guarded at the source because no run can be relied on to object to it. `try_pop()` straight after a submit, with the Option unwrapped, asserts the kernel has ALREADY queued the completion -- which RS-P-5 permits it not to have, and which pop_within's own documentation denies in this crate's words. Thirty-one of these passed for years because D-40 measured that a buffered read completes inside the submit in 80 of 80 attempts; an unbuffered one pends, and the same assertion then gives the opposite answer. That is why the guard is a census and not a test run: on the handles these tests open, the frozen observation is simply true.", + "testArgs": ["test", "-p", "windows-ioring-sys", "--locked", "--all-features", "--test", "response_space_census"], + "find": [ + "fn run_down_is_a_no_op_when_nothing_is_outstanding() {", + " let mut ring = IoRing::new(64, 128).expect(\"create ring\");" + ], + "replace": [ + "fn run_down_is_a_no_op_when_nothing_is_outstanding() {", + " let mut ring = IoRing::new(64, 128).expect(\"create ring\");", + " let _ = || ring.try_pop().expect(\"pop\").expect(\"ready\");" + ] + }, + { + "name": "M26.7: the already-poppable scanner stops recognising the shape", + "file": "tests/response_space_census.rs", + "expect": "caught", + "why": "Sabotaging the instrument rather than the code it watches. The scanner decides what the census above means, so a scanner that matched nothing would leave that census permanently and silently green while the shape it refuses crept back in -- the exact failure mode the first version of this file's clause census actually shipped with. The guard is a pair of assertions that the scanner recognises both spellings of the refused shape and neither of the two honest ones.", + "testArgs": ["test", "-p", "windows-ioring-sys", "--locked", "--all-features", "--test", "response_space_census"], + "find": [ + " if unwraps(first) && unwraps(second) {", + " count += 1;", + " }" + ], + "replace": [ + " let _ = (first, second);" + ] + }, + { + "name": "M26.7: pop_within is reverted to try_pop at a restated site", + "file": "tests/registration.rs", + "expect": "caught", + "why": "Whether the restatement itself is load-bearing, and the answer is instructive: reverting one site does NOT make that test fail on this machine, because the completion really is ready there -- which is the whole reason thirty-one of them survived review. What does catch it is the census, so this case is scoped to the census target and demonstrates that the guard covers the restatement rather than merely describing it. A reader tempted to check this by re-running registration.rs would find it green and conclude the change was cosmetic.", + "testArgs": ["test", "-p", "windows-ioring-sys", "--locked", "--all-features", "--test", "response_space_census"], + "find": [ + " let completion = ring", + " .pop_within(POP_BOUND)", + " .expect(\"pop completion\")", + " .expect(\"a completion arrives within the bound\");", + " let registered_files = files_pending" + ], + "replace": [ + " let completion = ring", + " .try_pop()", + " .expect(\"pop completion\")", + " .expect(\"a completion arrives within the bound\");", + " let registered_files = files_pending" + ] + }, + { + "name": "M26.8: a timed-out wait is reported as a failed submit again", + "file": "src/ring.rs", + "expect": "caught", + "why": "M21.6's defect, at the site that sweep missed. SubmitIoRing documents IORING_E_WAIT_TIMEOUT as 'All operations were submitted without error and the subsequent wait timed out' -- a positive guarantee about the submission half, not merely a non-failure. Folding it back in with every other error does more than flip a sign: its Remarks say that any OTHER error leaves all entries in the submission queue, so an Err from a submit means the buffers are still owed to the kernel. A caller who cannot tell the two apart and frees on Err hands the kernel freed memory at the next submit, which is the hazard D-5 exists to prevent.", + "testArgs": ["test", "-p", "windows-ioring-sys", "--locked", "--all-features", "--test", "resolver_over_a_real_ring"], + "find": [ + " hr == IORING_E_WAIT_TIMEOUT || check(hr).is_ok()" + ], + "replace": [ + " check(hr).is_ok()" + ] + }, + { + "name": "M26.8: run_down_within reports finished without checking", + "file": "src/ring.rs", + "expect": "caught", + "why": "The bounded rundown's entire value is that its answer is trustworthy: Ok(true) is what tells a caller the ring is safe to drop. A version that says so without looking would let every caller close a ring the kernel may still be writing through -- the exact hazard rundown exists to prevent, reintroduced while every call site still reads correctly. Note this also covers the unbounded run_down, which is now implemented over this one, so a single wrong answer here would propagate to both spellings.", + "testArgs": ["test", "-p", "windows-ioring-sys", "--locked", "--all-features", "--test", "resolver_over_a_real_ring"], + "find": [ + " if self.accounting.outstanding() == 0 {", + " return Ok(true);", + " }" + ], + "replace": [ + " return Ok(true);" + ] + }, + { + "name": "M26.10: the resolver never reports a short transfer", + "file": "src/sys/resolver.rs", + "expect": "caught", + "why": "RS-P-8 is a permission, so the resolver must be able to exercise it. With the switch ignored every transfer reports its full length and the clause becomes unobservable to every consumer written against it -- which is exactly the state the space was in before M26.10 decided the question. The RS-P-8 test's 'no seed reported a short transfer' half must catch this.", + "find": [ + " if self.config.may_transfer_partially && requested > 0 && self.chance(8) {" + ], + "replace": [ + " if false && self.config.may_transfer_partially && requested > 0 && self.chance(8) {" + ] + }, + { + "name": "M26.10: a transfer reports more bytes than were requested", + "file": "src/sys/resolver.rs", + "expect": "caught", + "why": "RS-P-8 permits a count below the request and says nothing that would allow one above it. Over-delivery is not a tolerated response but a defect, and a consumer sizing a buffer from the count would read past its own allocation. The bound assertion in the RS-P-8 test is what stands between the two readings.", + "find": [ + " let short = self.next() % u64::from(requested);" + ], + "replace": [ + " let short = u64::from(requested) + 1 + self.next() % 8;" + ] + }, + { + "name": "M26.11: the appender stops checking that a write was complete", + "file": "examples/epoch_log/append.rs", + "expect": "caught", + "why": "The log's contract requires a handle whose successful writes are complete, and this comparison is the only thing that enforces it rather than merely stating it. The ring permits a short count (RS-P-8) because it never asks what kind of handle it was given, so with this check suppressed a short write is accepted silently and the log carries a hole no replay would attribute to its cause. The count used to be bound to `_written` and discarded, which is the state this restores.", + "find": [ + " if written != record::RECORD_STRIDE {" + ], + "replace": [ + " if false && written != record::RECORD_STRIDE {" + ] + }, + { + "name": "CONTROL: rewording a statement changes nothing", + "file": "examples/epoch_log/contract.rs", + "expect": "survives", + "why": "The control that makes the cases above meaningful. These guards check that the reduction survives being *printed* -- they are deliberately indifferent to what any statement says, because asserting a particular sentence would be checking the copy rather than the contract. A suite that caught this would have bound itself to the wording, and every future correction to the contract's prose would arrive as a red test. It must survive.", + "find": [ + " text: \"a record reported durable is present and intact when the log is replayed\"," + ], + "replace": [ + " text: \"a record reported durable is present and intact when the log is read back\"," + ] + } + ] +} diff --git a/crates/windows-ioring-sys/src/accounting.rs b/crates/windows-ioring-sys/src/accounting.rs new file mode 100644 index 000000000..edab50957 --- /dev/null +++ b/crates/windows-ioring-sys/src/accounting.rs @@ -0,0 +1,186 @@ +// Copyright (c) 2026 Mike Grier +//! The bookkeeping half of a ring, with no kernel state in it (M24.2). +//! +//! # Why this is a separate type +//! +//! [`crate::IoRing`] had ten fields and they divide evenly. Five name kernel +//! state -- the handle, the negotiated version, the probed op support, the +//! registration array the kernel reads late (`D-32`), and the completion event. +//! The other five are a ledger this crate keeps for itself: the ring's +//! identity, the next `UserData` to hand out, how many operations are +//! outstanding, and the two registration base indices. +//! +//! Nothing in that second half needs a ring to exist. It is arithmetic and +//! identity, and the rules it enforces -- that an identity is never reused, +//! that a reservation is released exactly once, that the counters saturate +//! rather than wrap -- are **this crate's own specification**, not anything +//! Windows has an opinion about. +//! +//! Splitting it out is what lets those rules be tested without opening a +//! kernel ring, which is [D-49](../DESIGN-NOTES.md#d-49)'s defect: `cargo test +//! --lib` did not mean what its name implies. This type needs no fake, no +//! feature gate and no widened visibility to be exercised exhaustively. +//! +//! # What it deliberately does not know +//! +//! Whether an operation actually reached the kernel, whether a completion is +//! real, or whether the handle is still open. [`Accounting`] records what it +//! is *told*; the ring is what talks to Windows. Keeping that line sharp is +//! the point -- a ledger that tried to second-guess the kernel would be the +//! mock this crate rejects ([D-52](../DESIGN-NOTES.md#d-52)). + +use std::io; +use std::sync::atomic::{AtomicU64, Ordering}; + +/// A ring's identity, unique for the process's lifetime (PR #20 review +/// response): every value a ring hands out that later gets checked back +/// against it -- a [`crate::Token`], a [`crate::RegisteredFile`], a +/// [`crate::RegisteredBuffers`] -- carries the id of the ring that minted +/// it, and every [`crate::Completion`] carries the id of the ring that +/// popped it. +/// +/// A monotonic counter rather than the ring's own `HANDLE`: a `HANDLE` is +/// only unique while the object it names is still open, and Windows is free +/// to hand a closed ring's numeric value to the *next* object created -- +/// which would let a stale identity from a closed ring collide with a +/// brand-new one. This counter never repeats within one process run +/// (`u64` overflow is not a practical concern), so a mismatch always means +/// a genuine cross-ring mixup, never a false negative from handle reuse. +#[derive(Clone, Copy, Debug, PartialEq, Eq, Hash)] +pub(crate) struct RingId(u64); + +impl RingId { + /// The next identity. Process-global and monotonic. + /// + /// `Relaxed` is sufficient: the only property required is that no two + /// calls return the same value, which `fetch_add` gives on its own. No + /// other memory is being published through this counter. + pub(crate) fn next() -> Self { + static NEXT: AtomicU64 = AtomicU64::new(1); + Self(NEXT.fetch_add(1, Ordering::Relaxed)) + } +} + +/// A ring's ledger: identity, operation identities, and the counts. +#[derive(Debug)] +pub(crate) struct Accounting { + ring_id: RingId, + /// The next `UserData` value [`Accounting::reserve_user_data`] will hand + /// out. + next_user_data: usize, + /// Operations minted but not yet observed to have completed (M2.4). + outstanding: usize, + /// How many file handles are registered so far, across every confirmed + /// `BuildIoRingRegisterFileHandles` (M5.1). The base index of the next + /// registration. + registered_files: u32, + /// As `registered_files`, for `BuildIoRingRegisterBuffers` (M5.2). + registered_buffers: u32, +} + +impl Accounting { + /// A fresh ledger, with an identity no other ring in this process holds. + pub(crate) fn new() -> Self { + Self { + ring_id: RingId::next(), + next_user_data: 0, + outstanding: 0, + registered_files: 0, + registered_buffers: 0, + } + } + + /// This ring's own identity, for stamping onto every [`crate::Token`] and + /// registration it mints and checking against on use. + pub(crate) fn ring_id(&self) -> RingId { + self.ring_id + } + + /// How many operations this ring believes are still outstanding: minted + /// (via [`Accounting::reserve_user_data`]) but not yet observed to have + /// completed (via [`Accounting::record_completion`]). + pub(crate) fn outstanding(&self) -> usize { + self.outstanding + } + + /// Mint a fresh `UserData` identity for a new operation, and account for + /// it as outstanding until `record_completion` is called for it. + /// + /// # Errors + /// + /// Returns an error rather than reusing an identity if the `usize` space + /// is ever exhausted, mirroring `windows-threadpool-sys`'s own + /// "exhausting the generation sequence fails rather than wraps." + pub(crate) fn reserve_user_data(&mut self) -> io::Result { + let id = self.next_user_data; + self.next_user_data = id + .checked_add(1) + .ok_or_else(|| io::Error::other("IoRing operation identity space exhausted"))?; + self.outstanding += 1; + Ok(id) + } + + /// Record that one outstanding operation's completion has been observed + /// (a real `IORING_CQE` was popped for it), whether or not a live + /// [`crate::Token`] was still around to claim it. + pub(crate) fn record_completion(&mut self) { + self.outstanding = self.outstanding.saturating_sub(1); + } + + /// Release a reservation for an operation that was never actually + /// queued -- a `Build*` call failed synchronously, after + /// [`Accounting::reserve_user_data`] had already minted its identity. + /// + /// Distinct from [`Accounting::record_completion`]: that marks a real + /// `IORING_CQE` observed; this marks one that will never arrive because + /// the op never entered the queue, so it must not count against + /// `IoRing::run_down` either. + /// + /// The two are separate methods rather than one, even though their bodies + /// are identical, because they record different *facts* and a future + /// change to either -- a conservation counter, a debug assertion -- is + /// overwhelmingly likely to apply to only one of them. + pub(crate) fn cancel_reservation(&mut self) { + self.outstanding = self.outstanding.saturating_sub(1); + } + + /// The registered-file base index: how many file handles are registered + /// so far. + pub(crate) fn registered_file_count(&self) -> u32 { + self.registered_files + } + + /// As [`Accounting::registered_file_count`], for registered buffers. + pub(crate) fn registered_buffer_count(&self) -> u32 { + self.registered_buffers + } + + /// Advance the registered-file base index by `count`, the instant a + /// `BuildIoRingRegisterFileHandles` call successfully queues (not once + /// its completion is observed). + pub(crate) fn reserve_registered_files(&mut self, count: u32) { + self.registered_files = self.registered_files.saturating_add(count); + } + + /// As [`Accounting::reserve_registered_files`], for registered buffers. + pub(crate) fn reserve_registered_buffers(&mut self, count: u32) { + self.registered_buffers = self.registered_buffers.saturating_add(count); + } + + /// Place the identity counter near its ceiling, so the exhaustion path is + /// reachable. + /// + /// `#[cfg(test)]`, so it does not ship. It exists because that path is + /// otherwise unreachable -- a real ring would have to mint `usize::MAX` + /// operations -- and the repository's rule is that an error edge no test + /// can traverse is written rather than implemented. The alternative to + /// erroring there is silently handing out a duplicate identity, which is + /// exactly what a `Token` cannot survive. + #[cfg(test)] + pub(crate) fn set_next_user_data_for_test(&mut self, next: usize) { + self.next_user_data = next; + } +} + +#[cfg(test)] +mod tests; diff --git a/crates/windows-ioring-sys/src/accounting/tests.rs b/crates/windows-ioring-sys/src/accounting/tests.rs new file mode 100644 index 000000000..12a2a4d24 --- /dev/null +++ b/crates/windows-ioring-sys/src/accounting/tests.rs @@ -0,0 +1,271 @@ +// Copyright (c) 2026 Mike Grier +//! Tests for the ring's bookkeeping (M24.2). +//! +//! **These open no ring.** That is the point of the extraction: every rule +//! below is this crate's own specification, so it can be checked exhaustively +//! in memory rather than by driving a kernel object and hoping the interesting +//! state is reachable. They are the first tests in `src/` for which +//! [D-49](../../DESIGN-NOTES.md#d-49)'s complaint -- that `cargo test --lib` +//! does not mean what its name implies -- does not apply. +//! +//! They assert nothing about Windows, and must not start to. The moment a test +//! here encodes a belief about kernel behaviour it becomes the frozen +//! observation [D-52](../../DESIGN-NOTES.md#d-52) warns about. + +use super::{Accounting, RingId}; + +#[test] +fn a_fresh_ledger_is_empty() { + let a = Accounting::new(); + assert_eq!(a.outstanding(), 0); + assert_eq!(a.registered_file_count(), 0); + assert_eq!(a.registered_buffer_count(), 0); +} + +#[test] +fn identities_start_at_zero_and_increase_by_one() { + let mut a = Accounting::new(); + for expected in 0..64usize { + assert_eq!(a.reserve_user_data().expect("space remains"), expected); + } +} + +#[test] +fn an_identity_is_never_handed_out_twice() { + // The property a `Token` depends on to validate a completion (D-4): if two + // live operations shared a `UserData`, either could claim the other's + // completion. + let mut a = Accounting::new(); + let mut seen = std::collections::HashSet::new(); + for _ in 0..4096 { + let id = a.reserve_user_data().expect("space remains"); + assert!(seen.insert(id), "identity {id} was handed out twice"); + } +} + +#[test] +fn completing_an_operation_does_not_recycle_its_identity() { + // Reserve, complete, reserve again: the count returns to zero but the + // identity must still move on. Recycling would be the same defect as + // duplication, arriving later. + let mut a = Accounting::new(); + let first = a.reserve_user_data().expect("space remains"); + a.record_completion(); + assert_eq!(a.outstanding(), 0); + let second = a.reserve_user_data().expect("space remains"); + assert_ne!( + first, second, + "an identity must not be reused after completion" + ); +} + +#[test] +fn reserving_raises_the_outstanding_count() { + let mut a = Accounting::new(); + for expected in 1..=32usize { + a.reserve_user_data().expect("space remains"); + assert_eq!(a.outstanding(), expected); + } +} + +#[test] +fn a_recorded_completion_lowers_it() { + let mut a = Accounting::new(); + for _ in 0..32 { + a.reserve_user_data().expect("space remains"); + } + for expected in (0..32usize).rev() { + a.record_completion(); + assert_eq!(a.outstanding(), expected); + } +} + +#[test] +fn a_cancelled_reservation_lowers_it_too() { + // The `Build*`-failed path. It must not count against run-down, because + // the operation never entered the queue and no completion will arrive. + let mut a = Accounting::new(); + a.reserve_user_data().expect("space remains"); + a.reserve_user_data().expect("space remains"); + a.cancel_reservation(); + assert_eq!(a.outstanding(), 1); +} + +#[test] +fn the_outstanding_count_saturates_rather_than_wrapping() { + // Both release paths, in both directions of the same rule. A wrap here + // would make `run_down` wait on `usize::MAX` operations that do not exist + // -- a hang, not a miscount. + let mut a = Accounting::new(); + a.record_completion(); + assert_eq!(a.outstanding(), 0, "recording against an empty ledger"); + a.cancel_reservation(); + assert_eq!(a.outstanding(), 0, "cancelling against an empty ledger"); + + let mut b = Accounting::new(); + b.reserve_user_data().expect("space remains"); + b.record_completion(); + b.record_completion(); + b.cancel_reservation(); + assert_eq!(b.outstanding(), 0, "over-releasing a single reservation"); +} + +#[test] +fn identity_exhaustion_is_an_error_rather_than_a_reuse() { + // Unreachable in practice, and that is exactly why it is tested here: a + // ledger reached through a real ring could never be driven to this state, + // and the alternative to erroring is silently handing out a duplicate. + let mut a = Accounting::new(); + a.set_next_user_data_for_test(usize::MAX - 1); + + let last = a + .reserve_user_data() + .expect("the final identity is available"); + assert_eq!(last, usize::MAX - 1); + + let error = a + .reserve_user_data() + .expect_err("there is no identity after the last one"); + assert!( + error.to_string().contains("exhausted"), + "the error says what happened: {error}" + ); +} + +#[test] +fn the_last_identity_is_max_minus_one_not_max() { + // Measured, not assumed -- the first draft of these tests asserted + // `usize::MAX` and was wrong. The counter is advanced *before* the value + // is returned, so a reservation made at `usize::MAX` fails rather than + // handing it out: the identity space is `0..=usize::MAX - 1`. + // + // One value out of 2^64 is not worth reclaiming, and "fails rather than + // wraps" is the property that matters. It is recorded because the + // off-by-one is invisible from the source, and the next reader would + // otherwise make the same wrong assumption this test's author did. + let mut a = Accounting::new(); + a.set_next_user_data_for_test(usize::MAX); + let _ = a + .reserve_user_data() + .expect_err("usize::MAX is never handed out"); +} + +#[test] +fn an_exhausted_ledger_does_not_charge_the_failed_reservation() { + // The failure path must be clean: an identity that was never handed out + // must not leave an operation outstanding, or run-down waits forever for a + // completion that no operation will ever produce. + let mut a = Accounting::new(); + a.set_next_user_data_for_test(usize::MAX); + let before = a.outstanding(); + let _ = a.reserve_user_data().expect_err("exhausted"); + assert_eq!( + a.outstanding(), + before, + "a refused reservation costs nothing" + ); +} + +#[test] +fn registration_counts_advance_by_the_amount_reserved() { + let mut a = Accounting::new(); + a.reserve_registered_files(4); + assert_eq!(a.registered_file_count(), 4); + a.reserve_registered_files(3); + assert_eq!(a.registered_file_count(), 7); + + a.reserve_registered_buffers(8); + assert_eq!(a.registered_buffer_count(), 8); +} + +#[test] +fn the_two_registration_counts_are_independent() { + // They index different kernel tables. Advancing one must not move the + // other, or a registered buffer's index would collide with a file's. + let mut a = Accounting::new(); + a.reserve_registered_files(5); + assert_eq!( + a.registered_buffer_count(), + 0, + "files must not move buffers" + ); + a.reserve_registered_buffers(9); + assert_eq!( + a.registered_file_count(), + 5, + "and buffers must not move files" + ); +} + +#[test] +fn registration_counts_saturate_rather_than_wrapping() { + let mut a = Accounting::new(); + a.reserve_registered_files(u32::MAX); + a.reserve_registered_files(1); + assert_eq!(a.registered_file_count(), u32::MAX); + + a.reserve_registered_buffers(u32::MAX); + a.reserve_registered_buffers(7); + assert_eq!(a.registered_buffer_count(), u32::MAX); +} + +#[test] +fn reserving_zero_registrations_changes_nothing() { + let mut a = Accounting::new(); + a.reserve_registered_files(3); + a.reserve_registered_files(0); + assert_eq!(a.registered_file_count(), 3); +} + +#[test] +fn registrations_and_operations_do_not_interfere() { + let mut a = Accounting::new(); + a.reserve_user_data().expect("space remains"); + a.reserve_registered_files(2); + a.reserve_registered_buffers(2); + assert_eq!(a.outstanding(), 1, "registering does not mint an operation"); + + a.record_completion(); + assert_eq!( + a.registered_file_count(), + 2, + "completing does not unregister" + ); + assert_eq!(a.registered_buffer_count(), 2); +} + +#[test] +fn every_ledger_gets_its_own_identity() { + // What stops a token minted by one ring from claiming another's + // completion. Built without any ring at all, which is the extraction + // paying for itself: this used to need two live kernel objects. + let ids: Vec = (0..256).map(|_| Accounting::new().ring_id()).collect(); + let unique: std::collections::HashSet = ids.iter().copied().collect(); + assert_eq!(unique.len(), ids.len(), "two ledgers shared an identity"); +} + +#[test] +fn a_ledgers_identity_does_not_change_over_its_life() { + let mut a = Accounting::new(); + let id = a.ring_id(); + a.reserve_user_data().expect("space remains"); + a.record_completion(); + a.reserve_registered_files(4); + assert_eq!(a.ring_id(), id, "identity is fixed for the ledger's life"); +} + +#[test] +fn identities_from_separate_ledgers_do_not_collide_even_though_both_start_at_zero() { + // `UserData` restarts at 0 per ring, so the *pair* is what identifies an + // operation -- which is the reason `RingId` exists and why `Token` + // compares both. + let mut a = Accounting::new(); + let mut b = Accounting::new(); + assert_eq!(a.reserve_user_data().expect("space"), 0); + assert_eq!(b.reserve_user_data().expect("space"), 0); + assert_ne!( + a.ring_id(), + b.ring_id(), + "the ring id is what disambiguates the shared user data" + ); +} diff --git a/crates/windows-ioring-sys/src/batch.rs b/crates/windows-ioring-sys/src/batch.rs index 3a5a59ac0..096e823bf 100644 --- a/crates/windows-ioring-sys/src/batch.rs +++ b/crates/windows-ioring-sys/src/batch.rs @@ -10,13 +10,11 @@ use std::sync::atomic::{AtomicUsize, Ordering}; use windows_sys::Win32::Foundation::HANDLE; use windows_sys::Win32::Storage::FileSystem::{ - BuildIoRingCancelRequest, BuildIoRingFlushFile, BuildIoRingReadFile, - BuildIoRingRegisterBuffers, BuildIoRingRegisterFileHandles, BuildIoRingWriteFile, FILE_FLUSH_DATA, FILE_FLUSH_DEFAULT, FILE_FLUSH_MIN_METADATA, FILE_FLUSH_MODE, FILE_FLUSH_NO_SYNC, FILE_WRITE_FLAGS, FILE_WRITE_FLAGS_NONE, FILE_WRITE_FLAGS_WRITE_THROUGH, IORING_BUFFER_INFO, IORING_BUFFER_REF, IORING_BUFFER_REF_0, IORING_HANDLE_REF, IORING_HANDLE_REF_0, IORING_REF_RAW, IORING_REF_REGISTERED, IORING_REGISTERED_BUFFER, - IORING_SQE_FLAGS, IOSQE_FLAGS_DRAIN_PRECEDING_OPS, IOSQE_FLAGS_NONE, SubmitIoRing, + IORING_SQE_FLAGS, IOSQE_FLAGS_DRAIN_PRECEDING_OPS, IOSQE_FLAGS_NONE, }; use crate::buf::{IoBuf, IoBufMut}; @@ -122,6 +120,31 @@ impl PushOptions { /// It is an enum rather than a `bool` for the same reason: `flush(&file, /// true)` does not say what the `true` decides, and this is not a parameter /// anyone should have to look up. +/// +/// # The barrier's scope is the ring; the flush's is the file +/// +/// One call sets both, which makes them easy to merge, and merging them is a +/// mistake about durability rather than about style: +/// +/// - **This flag is ring-wide.** [`Self::CoversPrecedingOperations`] means the +/// flush does not execute until every operation outstanding *on the ring* +/// when it was reached has **completed** -- whatever file each one targets, +/// and whoever queued it (D-47, measured over roughly 4,500 trials). +/// - **The flush names one file.** [`Batch::flush`] takes a [`FileTarget`], so +/// what a syncing [`FlushMode`] pushes to stable media is that file's data +/// and the device cache behind it. +/// +/// **Completion is not durability.** A write completing means the kernel took +/// the bytes, not that they reached non-volatile media. So the barrier bounds +/// what a flush **waits for**, and the flush itself bounds what is **made +/// durable**, and those are different sets whenever a ring carries operations +/// against more than one file. +/// +/// The practical consequence for a caller is cost rather than correctness: +/// operations this flush will never make durable can still make it wait. A +/// caller who cares about that bounds it by controlling what shares the ring; +/// this crate does not decide that (D-8), and the flush's own target is what +/// secures the durability of the file named. #[derive(Clone, Copy, Debug, PartialEq, Eq)] pub enum FlushCoverage { /// Wait for every operation already outstanding on the ring, then flush. @@ -1053,8 +1076,17 @@ impl Drop for RegisteredBuffers { // simply never reclaimed rather than freed out from under a // still-outstanding op. Any one buffer still in use holds the // whole registration, because they are freed together. + // + // Silent while already panicking (M23.4). A second panic during + // unwind aborts the process, which *replaces* the failure that + // started the unwind: a test asserting on a leaked slot reports + // `STATUS_STACK_BUFFER_OVERRUN` instead of its own message, so + // the diagnosis is destroyed by the guard meant to aid it. The + // leak is still refused -- the early `return` below is what makes + // leaking rather than freeing the failure mode, and it happens + // either way. debug_assert!( - false, + std::thread::panicking(), "RegisteredBuffers dropped while an operation still references it" ); return; @@ -1241,14 +1273,14 @@ impl<'ring> Batch<'ring> { let len = checked_len(buffer.bytes_len())?; let address = buffer.stable_mut_ptr().cast::(); let target = handle_ref(file.into(), self.ring.ring_id())?; - let token = Token::new(self.ring, buffer)?; + let token = Token::new(self.ring.accounting_mut(), buffer)?; let user_data = token.id(); // SAFETY: `self.ring`'s handle is live; `address` is `IoBufMut`'s // promised stable, exclusively-owned pointer, valid for `len` bytes // until `token` is claimed; `file` is the caller's to keep alive, // forwarded from this function's own contract. let hr = unsafe { - BuildIoRingReadFile( + crate::sys::build_read( self.ring.raw_handle(), target, raw_buffer_ref(address), @@ -1281,7 +1313,7 @@ impl<'ring> Batch<'ring> { let len = checked_len(buffer.bytes_len())?; let address = buffer.stable_mut_ptr().cast::(); let target = handle_ref(file.as_file_ref(), self.ring.ring_id())?; - let token = Token::new(self.ring, (buffer, file.guard()))?; + let token = Token::new(self.ring.accounting_mut(), (buffer, file.guard()))?; let user_data = token.id(); // SAFETY: `self.ring`'s handle is live; `address` is `IoBufMut`'s // promised stable, exclusively-owned pointer, valid for `len` bytes @@ -1290,7 +1322,7 @@ impl<'ring> Batch<'ring> { // for a `SharedFile`, and nothing needing to be kept alive at all for // a `RegisteredFile`, whose index names the ring's own table. let hr = unsafe { - BuildIoRingReadFile( + crate::sys::build_read( self.ring.raw_handle(), target, raw_buffer_ref(address), @@ -1332,14 +1364,14 @@ impl<'ring> Batch<'ring> { let len = checked_len(buffer.bytes_len())?; let address = buffer.stable_ptr().cast_mut().cast::(); let target = handle_ref(file.into(), self.ring.ring_id())?; - let token = Token::new(self.ring, buffer)?; + let token = Token::new(self.ring.accounting_mut(), buffer)?; let user_data = token.id(); // SAFETY: `address` is `IoBuf`'s promised stable pointer, valid for // `len` bytes until `token` is claimed; the kernel only reads // through it for a write, so the cast away from `const` does not // authorize mutation. `file` is the caller's to keep alive. let hr = unsafe { - BuildIoRingWriteFile( + crate::sys::build_write( self.ring.raw_handle(), target, raw_buffer_ref(address), @@ -1373,12 +1405,12 @@ impl<'ring> Batch<'ring> { let len = checked_len(buffer.bytes_len())?; let address = buffer.stable_ptr().cast_mut().cast::(); let target = handle_ref(file.as_file_ref(), self.ring.ring_id())?; - let token = Token::new(self.ring, (buffer, file.guard()))?; + let token = Token::new(self.ring.accounting_mut(), (buffer, file.guard()))?; let user_data = token.id(); // SAFETY: as `write_raw`'s; `target` stays valid at least as long as // `token`'s hold on `file`'s guard does (see `Batch::read`). let hr = unsafe { - BuildIoRingWriteFile( + crate::sys::build_write( self.ring.raw_handle(), target, raw_buffer_ref(address), @@ -1507,7 +1539,7 @@ impl<'ring> Batch<'ring> { // keep alive, forwarded from this function's own contract; there is // no buffer. let hr = unsafe { - BuildIoRingFlushFile( + crate::sys::build_flush( self.ring.raw_handle(), target, mode.raw(), @@ -1545,12 +1577,12 @@ impl<'ring> Batch<'ring> { ) -> io::Result> { self.require(Op::Flush)?; let target = handle_ref(file.as_file_ref(), self.ring.ring_id())?; - let token = Token::new(self.ring, file.guard())?; + let token = Token::new(self.ring.accounting_mut(), file.guard())?; let user_data = token.id(); // SAFETY: `target` stays valid at least as long as `token`'s hold on // `file`'s guard does (see `Batch::read`); there is no buffer. let hr = unsafe { - BuildIoRingFlushFile( + crate::sys::build_flush( self.ring.raw_handle(), target, mode.raw(), @@ -1594,7 +1626,7 @@ impl<'ring> Batch<'ring> { // keep alive, forwarded from this function's own contract; // `BuildIoRingCancelRequest` takes no SQE-flags parameter. let hr = - unsafe { BuildIoRingCancelRequest(self.ring.raw_handle(), handle, target, user_data) }; + unsafe { crate::sys::build_cancel(self.ring.raw_handle(), handle, target, user_data) }; if let Err(error) = check(hr) { self.ring.cancel_reservation(); return Err(error); @@ -1623,13 +1655,13 @@ impl<'ring> Batch<'ring> { ) -> io::Result> { self.require(Op::Cancel)?; let handle = handle_ref(file.as_file_ref(), self.ring.ring_id())?; - let token = Token::new(self.ring, file.guard())?; + let token = Token::new(self.ring.accounting_mut(), file.guard())?; let user_data = token.id(); // SAFETY: `handle` stays valid at least as long as `token`'s hold on // `file`'s guard does (see `Batch::read`); // `BuildIoRingCancelRequest` takes no SQE-flags parameter. let hr = - unsafe { BuildIoRingCancelRequest(self.ring.raw_handle(), handle, target, user_data) }; + unsafe { crate::sys::build_cancel(self.ring.raw_handle(), handle, target, user_data) }; self.finish_push(hr, token) } @@ -1706,7 +1738,7 @@ impl<'ring> Batch<'ring> { // measurement (D-32), not inherited from the sibling registration, // which behaves the opposite way. let hr = unsafe { - BuildIoRingRegisterFileHandles( + crate::sys::build_register_files( self.ring.raw_handle(), count, handles.as_ptr(), @@ -1795,7 +1827,7 @@ impl<'ring> Batch<'ring> { // keeps alive via the returned `PendingBufferRegistration` and, once // claimed, `RegisteredBuffers`. let hr = unsafe { - BuildIoRingRegisterBuffers(self.ring.raw_handle(), count, infos_ptr, user_data) + crate::sys::build_register_buffers(self.ring.raw_handle(), count, infos_ptr, user_data) }; if let Err(error) = check(hr) { self.ring.cancel_reservation(); @@ -1844,7 +1876,7 @@ impl<'ring> Batch<'ring> { let target = handle_ref(file.into(), self.ring.ring_id())?; let index = registration.checked_span(span)?; let token = Token::new( - self.ring, + self.ring.accounting_mut(), registration.begin_use(span, KernelAccess::WritesBuffer), )?; let user_data = token.id(); @@ -1853,7 +1885,7 @@ impl<'ring> Batch<'ring> { // `file` is the caller's to keep alive, forwarded from this // function's own contract. let hr = unsafe { - BuildIoRingReadFile( + crate::sys::build_read( self.ring.raw_handle(), target, registered_buffer_ref(index, span.offset), @@ -1893,7 +1925,7 @@ impl<'ring> Batch<'ring> { let target = handle_ref(file.as_file_ref(), self.ring.ring_id())?; let index = registration.checked_span(span)?; let token = Token::new( - self.ring, + self.ring.accounting_mut(), ( registration.begin_use(span, KernelAccess::WritesBuffer), file.guard(), @@ -1904,7 +1936,7 @@ impl<'ring> Batch<'ring> { // buffer stays put until it drops; `target` stays valid at least as // long as `token`'s hold on `file`'s guard does (see `Batch::read`). let hr = unsafe { - BuildIoRingReadFile( + crate::sys::build_read( self.ring.raw_handle(), target, registered_buffer_ref(index, span.offset), @@ -1946,14 +1978,14 @@ impl<'ring> Batch<'ring> { let target = handle_ref(file.into(), self.ring.ring_id())?; let index = registration.checked_span(span)?; let token = Token::new( - self.ring, + self.ring.accounting_mut(), registration.begin_use(span, KernelAccess::ReadsBuffer), )?; let user_data = token.id(); // SAFETY: as `read_registered_raw`; the kernel only reads through // this reference for a write. let hr = unsafe { - BuildIoRingWriteFile( + crate::sys::build_write( self.ring.raw_handle(), target, registered_buffer_ref(index, span.offset), @@ -1993,7 +2025,7 @@ impl<'ring> Batch<'ring> { let target = handle_ref(file.as_file_ref(), self.ring.ring_id())?; let index = registration.checked_span(span)?; let token = Token::new( - self.ring, + self.ring.accounting_mut(), ( registration.begin_use(span, KernelAccess::ReadsBuffer), file.guard(), @@ -2003,7 +2035,7 @@ impl<'ring> Batch<'ring> { // SAFETY: as `read_registered`'s; the kernel only reads through // this reference for a write. let hr = unsafe { - BuildIoRingWriteFile( + crate::sys::build_write( self.ring.raw_handle(), target, registered_buffer_ref(index, span.offset), @@ -2025,7 +2057,14 @@ impl<'ring> Batch<'ring> { /// /// # Errors /// - /// Returns any error from `SubmitIoRing`. + /// Returns any error from `SubmitIoRing`. **An error means the entries + /// were not submitted and remain in the submission queue**, which is the + /// documented behaviour of that call: *"If this function returns an error + /// other than IORING_E_WAIT_TIMEOUT, then all entries remain in the + /// submission queue."* They are not lost and they are not rewound + /// ([D-5](../DESIGN-NOTES.md#d-5)) -- a later submit on this ring is what + /// runs them, so **the buffers they reference must stay alive**. Whether + /// and when to submit again is the caller's policy, not this crate's. pub fn submit(self) -> io::Result { self.submit_and_wait(0, 0) } @@ -2040,9 +2079,22 @@ impl<'ring> Batch<'ring> { /// submitted rather than completed (M10.2). Drain with /// [`crate::IoRing::try_pop`] and count for yourself. /// + /// **A wait that expires is success, not failure** (M26.8). `SubmitIoRing` + /// answers `IORING_E_WAIT_TIMEOUT` in that case, and its documented + /// meaning is *"All operations were submitted without error and the + /// subsequent wait timed out"* -- so the submission half did everything it + /// was asked to. Reporting that as an `Err` is the defect `M21.6` fixed + /// for [`crate::IoRing::pop_within`]; this site was missed by that sweep + /// and kept it until `M26.8`. The distinction is not cosmetic: an `Err` + /// here means the entries are **still queued**, and a caller who frees + /// their buffers on seeing one would hand the kernel freed memory on the + /// next submit. + /// /// # Errors /// - /// Returns any error from `SubmitIoRing`. + /// Returns any error from `SubmitIoRing` **other than an expired wait**. + /// As [`Batch::submit`], an error means the entries remain in the + /// submission queue and their buffers must stay alive. pub fn submit_and_wait(mut self, wait_operations: u32, timeout_ms: u32) -> io::Result { self.do_submit(wait_operations, timeout_ms) } @@ -2051,7 +2103,7 @@ impl<'ring> Batch<'ring> { let mut submitted = 0_u32; // SAFETY: `self.ring`'s handle is live. let hr = unsafe { - SubmitIoRing( + crate::sys::submit( self.ring.raw_handle(), wait_operations, timeout_ms, @@ -2066,7 +2118,16 @@ impl<'ring> Batch<'ring> { // succeed on the retry, submitting operations the caller's `Err` // never told them about. self.submitted = true; - check(hr)?; + // `IORING_E_WAIT_TIMEOUT` is not a submission failure: the documented + // meaning of that code is that every entry went in and only the wait + // ran out (M26.8). Passing it to `check` reported a successful submit + // as an error, which is `M21.6`'s defect at the one site that sweep + // did not reach -- and the worse half is that it made an `Err` here + // ambiguous between "still queued" and "submitted, wait expired", + // which have opposite consequences for buffer ownership. + if !crate::ring::every_entry_was_submitted(hr) { + check(hr)?; + } Ok(submitted) } } diff --git a/crates/windows-ioring-sys/src/batch/tests.rs b/crates/windows-ioring-sys/src/batch/tests.rs index 28b59a9fd..ca6598560 100644 --- a/crates/windows-ioring-sys/src/batch/tests.rs +++ b/crates/windows-ioring-sys/src/batch/tests.rs @@ -1,5 +1,4 @@ // Copyright (c) 2026 Mike Grier -use windows_sys::Win32::Foundation::HANDLE; use windows_sys::Win32::Storage::FileSystem::{ FILE_FLUSH_DATA, FILE_FLUSH_DEFAULT, FILE_FLUSH_MIN_METADATA, FILE_FLUSH_NO_SYNC, FILE_WRITE_FLAGS_NONE, FILE_WRITE_FLAGS_WRITE_THROUGH, IOSQE_FLAGS_DRAIN_PRECEDING_OPS, @@ -8,7 +7,6 @@ use windows_sys::Win32::Storage::FileSystem::{ use super::{Batch, FlushCoverage, FlushMode, PushOptions, WriteCaching}; use crate::IoRing; -use crate::buf::{IoBuf, IoBufMut}; #[test] fn default_push_options_set_no_barrier() { @@ -70,75 +68,6 @@ fn drain_preceding_sets_the_barrier_flag() { ); } -/// A buffer that claims a length no real allocation could ever have, to -/// exercise `checked_len`'s rejection without needing a real file: the -/// rejection must happen before the buffer's pointer is ever read. -struct HugeBuffer; - -// SAFETY: `stable_ptr`/`stable_mut_ptr` are never dereferenced in the tests -// that use this type -- `checked_len` rejects the operation first. -unsafe impl IoBuf for HugeBuffer { - fn stable_ptr(&self) -> *const u8 { - std::ptr::NonNull::dangling().as_ptr() - } - - fn bytes_len(&self) -> usize { - usize::MAX - } -} - -// SAFETY: see the `IoBuf` impl above. -unsafe impl IoBufMut for HugeBuffer { - fn stable_mut_ptr(&mut self) -> *mut u8 { - std::ptr::NonNull::dangling().as_ptr() - } -} - -const NULL_FILE: HANDLE = std::ptr::null_mut(); - -#[test] -fn read_rejects_a_buffer_longer_than_u32_max_without_touching_the_ring() { - let mut ring = IoRing::new(8, 8).expect("create ring"); - let outstanding_before = ring.outstanding(); - let mut batch = Batch::new(&mut ring); - // SAFETY: NULL_FILE is never dereferenced -- the oversized buffer is - // rejected before the handle would be used. - let error = unsafe { batch.read_raw(NULL_FILE, HugeBuffer, 0, PushOptions::new()) } - .expect_err("an oversized buffer must be rejected"); - assert_eq!(error.kind(), std::io::ErrorKind::InvalidInput); - drop(batch); - assert_eq!( - ring.outstanding(), - outstanding_before, - "a rejected push must not reserve an identity" - ); -} - -#[test] -fn write_rejects_a_buffer_longer_than_u32_max_without_touching_the_ring() { - let mut ring = IoRing::new(8, 8).expect("create ring"); - let outstanding_before = ring.outstanding(); - let mut batch = Batch::new(&mut ring); - // SAFETY: as above. - let error = unsafe { - batch.write_raw( - NULL_FILE, - HugeBuffer, - 0, - PushOptions::new(), - WriteCaching::Cached, - ) - } - .expect_err("an oversized buffer must be rejected"); - assert_eq!(error.kind(), std::io::ErrorKind::InvalidInput); - drop(batch); - assert_eq!( - ring.outstanding(), - outstanding_before, - "a rejected push must not reserve an identity" - ); -} - #[test] fn dropping_a_batch_that_queued_nothing_submits_harmlessly() { let mut ring = IoRing::new(8, 8).expect("create ring"); @@ -262,7 +191,7 @@ fn a_pending_buffer_registration_claims_only_its_own_completion() { // one leaves `-> 1` indistinguishable -- both of which survived in turn // while this test was being written. for expected in 0..2 { - let burned = crate::Token::new(&mut ring, vec![0_u8; 1]).expect("mint a token"); + let burned = crate::Token::new(ring.accounting_mut(), vec![0_u8; 1]).expect("mint a token"); assert_eq!( burned.id(), expected, @@ -296,10 +225,7 @@ fn a_pending_buffer_registration_claims_only_its_own_completion() { panic!("a completion naming another operation must be refused"); }; - let real = ring - .try_pop() - .expect("pop") - .expect("the registration completion is ready"); + let real = crate::ring::pop_within(&mut ring, "the registration's completion"); // Ties the accessor to the operation it names, against an id obtained from // the ring rather than from the accessor itself -- a constant `user_data` // survives any comparison that starts from `user_data`. @@ -412,10 +338,7 @@ fn dropping_a_registration_with_work_outstanding_is_refused() { .register_buffers(vec![vec![0_u8; 512]]) .expect("queue buffer registration"); batch.submit_and_wait(1, 5_000).expect("submit"); - let completion = ring - .try_pop() - .expect("pop") - .expect("registration completed"); + let completion = crate::ring::pop_within(&mut ring, "the registration's completion"); let buffers = pending .claim_if(&completion) .expect("claims its own") @@ -468,27 +391,6 @@ fn require_refuses_an_op_the_ring_does_not_support() { .expect("Nop is in the constructed capability set"); } -#[test] -fn the_debug_rendering_names_the_registration_and_its_identity() { - // `>::fmt -> Ok(Default::default())` - // survived: that mutation writes nothing to the formatter, so the - // rendering comes back empty regardless of what the registration holds. - let mut ring = IoRing::new(8, 8).expect("create ring"); - let mut batch = Batch::new(&mut ring); - let pending = batch - .register_buffers(vec![vec![0_u8; 64]]) - .expect("queue buffer registration"); - let rendering = format!("{pending:?}"); - assert!( - rendering.contains("PendingBufferRegistration"), - "got {rendering}" - ); - assert!( - rendering.contains(&pending.user_data().to_string()), - "the operation's identity must appear: {rendering}" - ); -} - #[test] fn windows_refuses_an_empty_buffer_registration() { // Written while chasing `RegisteredBuffers::is_empty -> false`, which a diff --git a/crates/windows-ioring-sys/src/event_delivery.rs b/crates/windows-ioring-sys/src/event_delivery.rs index ff00690e0..6c157f969 100644 --- a/crates/windows-ioring-sys/src/event_delivery.rs +++ b/crates/windows-ioring-sys/src/event_delivery.rs @@ -88,15 +88,29 @@ impl EventDelivery { /// completion signals either, because the queue never returns to empty to /// re-arm the edge. Such a backlog is stranded permanently. /// - /// What closes that gap is the deliberate signal - /// [`IoRing::completion_event`] raises as it attaches: the first callback - /// then drains the backlog exactly as it would any other wakeup. + /// What closes that gap is a deliberate signal raised on the event once + /// it has been attached: the first callback then drains the backlog + /// exactly as it would any other wakeup. + /// + /// That signal is raised *after* the wait has been armed, which is the + /// order `SetThreadpoolWait` documents -- "you must re-register the event + /// with the wait object before signaling it each time to trigger the wait + /// callback". Signalling first and arming afterwards is not guaranteed to + /// run the callback, and since the event is auto-reset the signal is + /// consumed rather than left pending for the arming to find. For a ring + /// whose queue never returns to empty there is no second wakeup coming, + /// so that loss strands the backlog permanently instead of merely + /// delaying it. This method therefore attaches the event unsignalled and + /// raises the signal itself, rather than going through + /// [`IoRing::completion_event`], which signals as it attaches and so + /// leaves a caller no way to arm in between. /// /// This was false in the implementation, and asserted anyway in this /// rustdoc, before M11.3 -- every test until then handed over a fresh /// ring, so nothing contradicted it. A caller on an earlier version /// cannot rely on the guarantee; `tests/event_delivery.rs` keeps the - /// repro that now holds it. + /// repro that now holds it. The ordering above was wrong until M26.9, in + /// a way that stranded the backlog in roughly one run in a hundred. /// /// # Errors /// @@ -105,8 +119,9 @@ impl EventDelivery { /// [`IoRing::completion_event`] is what decides. This crate refuses to /// silently substitute a thread-based polling loop instead -- a caller /// who asked for event-driven delivery and got a spun-up thread has been - /// told something false. Also returns any other error from - /// [`IoRing::completion_event`] or from `ThreadpoolWait::new`. + /// told something false. Also returns any other error from attaching the + /// ring's completion event, from `ThreadpoolWait::new`, or from raising + /// the setup signal. pub fn new( mut ring: IoRing, on_completion: F, @@ -118,10 +133,20 @@ impl EventDelivery { // The ring creates, owns, and attaches its own event and hands back a // duplicate (D-20), which leaves exactly one // `SetIoRingCompletionEvent` call site in this crate. Delegating also - // means the capability check, the `Unsupported` error, and the - // signal-once-on-attach that makes the backlog guarantee above true - // are each stated in one place rather than restated here. - let event = ring.completion_event()?; + // means the capability check and the `Unsupported` error are each + // stated in one place rather than restated here. + // + // The attachment is taken *unsignalled*, and the setup signal raised + // only after `wait.arm` below, because `SetThreadpoolWait` documents + // that "you must re-register the event with the wait object before + // signaling it each time to trigger the wait callback". Signalling + // first and arming afterwards is the order that rule forbids. + let (event, owes_setup_signal) = ring.attach_completion_event_unsignalled()?; + windows_threadpool_sys::trace_record!( + "delivery", + "event-attached", + std::os::windows::io::AsRawHandle::as_raw_handle(&event) as usize + ); // SAFETY: `completion_event` returns a duplicate of an auto-reset // event -- always a supported wait target, and never a mutex -- and // that duplicate is exclusively ours. The ring keeps its own separate @@ -140,13 +165,34 @@ impl EventDelivery { // lands, so the only gap this leaves is a completion that // arrives between the last pop and the re-arm -- and the // second drain closes exactly that gap (M4.2). + windows_threadpool_sys::trace_record!("delivery", "callback-entered"); drain(&ring_for_wait, on_completion.as_ref()); activation.rearm(None); drain(&ring_for_wait, on_completion.as_ref()); + windows_threadpool_sys::trace_record!("delivery", "callback-left"); }, env, )?; wait.arm(None); + windows_threadpool_sys::trace_record!("delivery", "armed"); + + // Only now, with the wait registered, is the setup signal raised -- + // the order `SetThreadpoolWait` documents. This is what makes the + // backlog guarantee above true, so a failure to raise it is a failure + // to construct. + // + // Signalled only when this call attached the event. A caller that + // attached it earlier and consumed the signal reaches this with a + // non-empty queue and no wakeup owing, which review raised and + // `M26.12` is investigating -- the obvious repair, signalling + // unconditionally, was tried and does **not** fix it, so it is not + // applied here. See the ignored reproducer in + // `tests/event_delivery.rs` and UNRESOLVED-TEST-FAILURES.md. + if owes_setup_signal { + ring.lock() + .unwrap_or_else(std::sync::PoisonError::into_inner) + .raise_setup_signal()?; + } Ok(Self { wait, ring }) } diff --git a/crates/windows-ioring-sys/src/event_delivery/tests.rs b/crates/windows-ioring-sys/src/event_delivery/tests.rs index 194b62b0b..306ed7aa4 100644 --- a/crates/windows-ioring-sys/src/event_delivery/tests.rs +++ b/crates/windows-ioring-sys/src/event_delivery/tests.rs @@ -2,21 +2,6 @@ use super::EventDelivery; use crate::IoRing; -#[test] -fn new_succeeds_and_the_ring_stays_reachable_for_pushes() { - let ring = IoRing::new(8, 8).expect("create ring"); - let delivery = EventDelivery::new(ring, |_completion| {}, None).expect("wire event delivery"); - let info = delivery.scope().info().expect("query info"); - assert!(info.submission_queue_size > 0); -} - -#[test] -fn dropping_with_nothing_outstanding_does_not_hang() { - let ring = IoRing::new(8, 8).expect("create ring"); - let delivery = EventDelivery::new(ring, |_completion| {}, None).expect("wire event delivery"); - drop(delivery); -} - // --- the scope's read-only surface (M18.4) ----------------------------------- // // M18.6 gave `RingScope` the whole read-only surface of `IoRing` on the stated diff --git a/crates/windows-ioring-sys/src/lib.rs b/crates/windows-ioring-sys/src/lib.rs index 1f73dc013..89900f5b0 100644 --- a/crates/windows-ioring-sys/src/lib.rs +++ b/crates/windows-ioring-sys/src/lib.rs @@ -119,13 +119,22 @@ //! sizing a Model B execution domain to the caller. Three pointers, not a //! partitioning policy: //! -//! - **Size a domain by last-level (L3) cache, not by NUMA node.** Node count -//! is a firmware setting a process cannot see (AMD NPS, Intel Sub-NUMA -//! Clustering), and most real deployments are virtualized, where NUMA -//! topology is often invisible entirely. `GetLogicalProcessorInformationEx` -//! filtered to `RelationCache` / `CacheLevel == 3` degrades sanely instead: -//! one reported domain on a VM is correct. See `examples/l3_domains.rs` for -//! a runnable enumeration (M6.3), built on the safe wrapper in +//! - **Size a domain by the outermost cache level that partitions the +//! machine, not by NUMA node.** Node count is a firmware setting a process +//! cannot see (AMD NPS, Intel Sub-NUMA Clustering), and most real +//! deployments are virtualized, where NUMA topology is often invisible +//! entirely. A cache partition degrades sanely instead: a machine whose +//! caches partition nothing yields one ring, which is correct. +//! +//! **Ask which level partitions; do not filter on `CacheLevel == 3`.** +//! `MachineMemoryTopology::outermost_partitioning_cache` is the one +//! definition, and it is not a convenience wrapper over that filter: a +//! shipping ARM part reports no L3 at all, and a machine can report an L3 +//! spanning every processor above a real L2 partition -- where filtering on +//! level 3 does not even degrade, because it matched something. It returns +//! one whole-machine domain and calls it a cache-aware partition. See +//! `examples/cache_domains.rs` for a runnable enumeration (M6.3), built on +//! the safe wrapper in //! [`windows-topology-sys`](https://docs.rs/windows-topology-sys). //! - **Processor groups are a hard floor.** A thread's affinity is a //! `GROUP_AFFINITY` and a ring's waiter lives in exactly one group, so above @@ -137,16 +146,20 @@ //! happens to run is a one-time cache-warmth question by comparison. //! `VirtualAllocExNuma`, on the node closest to the device, registered once //! into that domain's ring (see [`Batch::register_buffers`]), is very -//! likely the highest-leverage locality decision available. +//! likely the highest-leverage locality decision available. [`NumaBuffer`] +//! is that allocation; *which* node is a question about a caller's storage +//! layout, and this crate does not answer it for them. //! -//! # Status +//! # Where the design lives //! -//! Under construction. The design, including the delivery-architecture guidance -//! this crate exists to make usable, is recorded in `DESIGN-NOTES.md` beside the -//! source; the build-out is tracked in `CHECKLIST.md`. +//! The design, including the delivery-architecture guidance this crate exists +//! to make usable, is recorded in `DESIGN-NOTES.md` beside the source, and the +//! work still open against it is tracked in `CHECKLIST.md`. #![warn(missing_docs)] +#[cfg(windows)] +mod accounting; #[cfg(windows)] mod batch; #[cfg(windows)] @@ -165,7 +178,23 @@ mod error; #[cfg(all(windows, feature = "threadpool"))] mod event_delivery; #[cfg(windows)] +mod numa_buffer_io; +// M23.3 SPIKE -- exported so a real consumer can be converted, which is the +// only way to validate whether one type fits the twelve hand-rolled shapes. +// Whether it stays public is the decision M23.3 has not yet taken. +#[cfg(windows)] +mod pending; +#[cfg(windows)] mod ring; +/// The seam the kernel-response resolver sits under (M26.2). +/// +/// Private without the `kernel-seam` feature, where it is nothing but +/// `#[inline(always)]` forwards to the same `windows-sys` calls this crate +/// made before. With the feature on it additionally publishes +/// `sys::Responses` and `sys::install`, so a test can answer the calls +/// that carry an operation instead of the kernel. +#[cfg(windows)] +pub mod sys; #[cfg(windows)] mod token; @@ -183,6 +212,12 @@ pub use capability::{Capabilities, RingVersion, capabilities}; pub use error::{IoRingError, IoRingErrorExt, RingCondition}; #[cfg(all(windows, feature = "threadpool"))] pub use event_delivery::{EventDelivery, RingScope}; +// Re-exported rather than defined here: the allocator moved to `win-numa-sys`, +// and re-exporting keeps `windows_ioring_sys::NumaBuffer` resolving for anyone +// who already bound to it. The `IoBuf`/`IoBufMut` impls live in +// `numa_buffer_io`, which explains there why they are separated from the type. +#[cfg(windows)] +pub use pending::Pending; /// The fault-injection seam (M16.3), for exercising failure paths a healthy /// machine will not produce on demand. See /// [`Completion::with_injected_failure`] for why transforming a real @@ -190,9 +225,10 @@ pub use event_delivery::{EventDelivery, RingScope}; #[cfg(all(windows, any(test, feature = "fault-injection")))] pub use ring::InjectedFailure; #[cfg(windows)] -pub use ring::{Completion, IoRing, Op, RingInfo}; -#[cfg(windows)] +pub use ring::{Completion, CompletionWait, IoRing, Op, RingInfo, RingWait, SubmitWait}; pub use token::Token; +#[cfg(windows)] +pub use win_numa_sys::NumaBuffer; // The crate's markdown documentation is compiled as doctests, so an example that // a contract change invalidates breaks the build instead of quietly teaching the diff --git a/crates/windows-ioring-sys/src/numa_buffer_io.rs b/crates/windows-ioring-sys/src/numa_buffer_io.rs new file mode 100644 index 000000000..d6d44943d --- /dev/null +++ b/crates/windows-ioring-sys/src/numa_buffer_io.rs @@ -0,0 +1,47 @@ +// Copyright (c) 2026 Mike Grier +//! This crate's I/O buffer traits, implemented for [`win_numa_sys::NumaBuffer`]. +//! +//! # Why the impls are here and the type is not +//! +//! The allocator moved to `win-numa-sys`, which deliberately defines no buffer +//! trait: this crate and `windows-overlapped-io-sys` each already have their +//! own [`IoBuf`]/[`IoBufMut`] pair, duplicated under +//! [D-1](../DESIGN-NOTES.md#d-1)'s duplicate-then-decide with the merge still +//! open as `M6+.6`. A third copy in the allocator crate would have made that +//! decision harder rather than easier, and would have forced whichever traits +//! it chose onto every consumer of a NUMA buffer. +//! +//! The orphan rule permits the arrangement that avoids all of it. This crate +//! **owns** `IoBuf`, so it may implement it for a foreign type, and the +//! allocator stays trait-free with plain inherent accessors. `M6+.6` is +//! untouched: whatever it decides, these two impls are where this crate's +//! answer lands, and nothing in `win-numa-sys` has to move again. + +use win_numa_sys::NumaBuffer; + +use crate::{IoBuf, IoBufMut}; + +// SAFETY: the allocation's address is fixed once `VirtualAllocExNuma` returns +// it and does not move for the value's life; the requested length is fixed +// too. `NumaBuffer` is `Send`, which `IoBuf` requires. +unsafe impl IoBuf for NumaBuffer { + fn stable_ptr(&self) -> *const u8 { + self.as_ptr() + } + + fn bytes_len(&self) -> usize { + self.len() + } +} + +// SAFETY: a `NumaBuffer` uniquely owns its allocation, so `&mut self` is +// exclusive access to the bytes; the address is the same one `stable_ptr` +// reports, because both return the allocation's base. +unsafe impl IoBufMut for NumaBuffer { + fn stable_mut_ptr(&mut self) -> *mut u8 { + self.as_mut_ptr() + } +} + +#[cfg(test)] +mod tests; diff --git a/crates/windows-ioring-sys/src/numa_buffer_io/tests.rs b/crates/windows-ioring-sys/src/numa_buffer_io/tests.rs new file mode 100644 index 000000000..a7ad0f190 --- /dev/null +++ b/crates/windows-ioring-sys/src/numa_buffer_io/tests.rs @@ -0,0 +1,77 @@ +// Copyright (c) 2026 Mike Grier +//! Tests for this crate's [`IoBuf`]/[`IoBufMut`] impls over a foreign type. +//! +//! # What these are for +//! +//! The impls are two lines each and delegate to inherent accessors, so the risk +//! is not that the bodies are wrong -- it is that they are wired to the wrong +//! thing, and a delegation that returns a plausible pointer is exactly the kind +//! of defect that reads as correct. What is asserted is therefore the +//! *correspondence*: that the trait methods report the same address and the +//! same length the buffer itself does, and that the address is the one the +//! bytes actually live at. +//! +//! These also serve as the check that the orphan-rule arrangement holds at all. +//! If `win-numa-sys` ever grew its own buffer trait, or this crate's `IoBuf` +//! moved under `M6+.6`, this file is where that breaks first. + +use win_numa_sys::NumaBuffer; + +use crate::{IoBuf, IoBufMut}; + +/// The trait methods agree with the inherent ones. +#[test] +fn the_impls_report_what_the_buffer_reports() { + let mut buffer = NumaBuffer::new(4096, None).expect("a valid allocation"); + let inherent_ptr = buffer.as_ptr(); + let inherent_len = buffer.len(); + + assert_eq!(buffer.stable_ptr(), inherent_ptr); + assert_eq!(buffer.bytes_len(), inherent_len); + assert_eq!(buffer.stable_mut_ptr().cast_const(), inherent_ptr); +} + +/// `bytes_len` is the requested length, which is what a kernel call is told. +/// +/// The allocation is page-granular, so the mapping is larger than this for any +/// request under a page. Reporting the mapping's size instead would tell the +/// kernel it may use bytes the caller never asked about. +#[test] +fn bytes_len_is_the_requested_length_not_the_mapped_one() { + for len in [1_usize, 100, 4095, 4096, 4097] { + let buffer = NumaBuffer::new(len, None).expect("a valid allocation"); + assert_eq!(buffer.bytes_len(), len, "for a request of {len} bytes"); + } +} + +/// The address the traits report is the address the bytes are at. +/// +/// Writing through `stable_mut_ptr` and reading back through `stable_ptr` is +/// what proves the two are the same region rather than two plausible pointers. +#[test] +fn the_reported_address_is_where_the_bytes_are() { + let mut buffer = NumaBuffer::new(4096, None).expect("a valid allocation"); + let len = buffer.bytes_len(); + + // SAFETY: the pointer is this buffer's own allocation of at least `len` + // bytes, and no other reference to it exists. + unsafe { std::slice::from_raw_parts_mut(buffer.stable_mut_ptr(), len) }.fill(0xA5); + + // SAFETY: same region, now read-only, still exclusively owned here. + let read = unsafe { std::slice::from_raw_parts(buffer.stable_ptr(), len) }; + assert!(read.iter().all(|&byte| byte == 0xA5)); +} + +/// The address does not change when the value moves. +/// +/// `IoBuf`'s contract is that the pointer is stable for the value's life, and +/// a buffer handed to a ring is routinely moved into a collection first. The +/// allocation is owned by address rather than inline, so this holds -- but it +/// holds by construction rather than by accident, and the test is what says so. +#[test] +fn the_address_survives_a_move() { + let buffer = NumaBuffer::new(4096, None).expect("a valid allocation"); + let before = buffer.stable_ptr(); + let moved = buffer; + assert_eq!(moved.stable_ptr(), before); +} diff --git a/crates/windows-ioring-sys/src/pending.rs b/crates/windows-ioring-sys/src/pending.rs new file mode 100644 index 000000000..c7277abb9 --- /dev/null +++ b/crates/windows-ioring-sys/src/pending.rs @@ -0,0 +1,279 @@ +// Copyright (c) 2026 Mike Grier +//! **SPIKE for `M23.3`, not yet a decision.** `Pending` -- the identity +//! map every ring consumer writes, with the conservation oracle wired in rather +//! than driven alongside it. +//! +//! # What this is validating +//! +//! A census found twelve sites in this crate keeping a map from `UserData` to +//! an unclaimed [`Token`], and every one that also uses [`RingContract`] drives +//! the two **in parallel, by hand**: +//! +//! ```text +//! self.contract.observe_push(token.id()); +//! self.in_flight.insert(token.id(), InFlight { token, slot }); +//! ``` +//! +//! That is the same event recorded twice, which the repository's CONTRACT +//! INTEGRITY rule calls a restatement: it can drift in both directions, and the +//! oracle's value depends on being driven correctly by the very code it checks. +//! This type exists to find out whether one call site can keep both true. +//! +//! # The generic parameter is the finding, not a convenience +//! +//! The census also falsified the shape the checklist assumed. Only a third of +//! the sites keep a bare `Token`; the rest carry per-operation sidecar data +//! -- a slot index, an expected length, a phase, a sequence number. `X` is that +//! sidecar, defaulted to `()` so the bare sites read unchanged. +//! +//! # The shape this is not, and the reason it was wrongly ruled out +//! +//! A stronger design is available: a generic `IoRing` owning the map, so the +//! **ring notifies** rather than being told. That makes drift between the ring +//! and the inventory structurally impossible, where this type only removes +//! drift between the map and the oracle -- nothing here forces a minted +//! `Token` into a `Pending` at all. +//! +//! It was first dismissed on the grounds that one ring carries several +//! `Token` types, forcing `Box` and a downcast at the claim site. +//! **That reasoning is wrong twice over**, and checking is what showed it: +//! per-ring monomorphisation holds for every real consumer here, and where it +//! genuinely does not, the answer is a closed enum rather than `dyn Any` -- +//! `tests/generated_sequences.rs` already puts eight token types on one ring +//! behind `enum Held`, with an exhaustive `match` and no runtime type check. +//! +//! The real costs are different: `IoRing` is not generic today, so it is a +//! breaking change to a published crate; a consumer mixing shapes writes a +//! `Held`-style enum; and tokenless pushes need a story, since `flush_raw` +//! returns a bare `usize` and `epoch_log`'s commit path depends on that. See +//! [DESIGN-SESSION-2026-09-23-pending-inventory.md](../../design-sessions/DESIGN-SESSION-2026-09-23-pending-inventory.md). +//! +//! # What this cannot do, which is also a finding +//! +//! **It cannot force the discipline.** Rust has no linear types, so nothing +//! makes a caller invoke [`Pending::finish`]. What it can do is make the +//! violation *loud* at the moment it happens rather than silent forever -- see +//! this type's `Drop`. +//! +//! **It does not make ordering hazards unrepresentable, only detectable -- and +//! detectability had to be built.** Reinstating `M22.2`'s defect in the +//! converted consumer -- checking a write's result *before* claiming, so a +//! failed write returns early with the token still held -- still compiles. When +//! first measured it also still *passed*, because nothing produced a failed +//! write. That gap is now closed: `append/tests.rs` drives the failure through +//! [`Completion::with_injected_failure`], and the sabotage is caught by an +//! assertion naming the leak. The type makes the leak visible; a test is what +//! makes it visible *in CI* rather than in production. +//! +//! [`Completion::with_injected_failure`]: crate::Completion::with_injected_failure +//! +//! **Owning the oracle creates a decoy hazard.** [`Pending::checked`] mints its +//! own [`RingContract`], so a consumer that already had one keeps a field that +//! is never written to again. Converting `epoch_log`'s appender did exactly +//! that, and the result compiled, ran, and made its teardown +//! `assert_quiescent()` pass **vacuously**. Sabotage confirms nothing in the +//! suite catches it: replacing that accessor with an oracle that has observed +//! nothing leaves every test green. It was found by reading. Any consumer +//! converted to a checked map must route its existing accessor through +//! [`Pending::contract`], and that obligation is invisible to the compiler. + +use std::collections::HashMap; +use std::collections::hash_map::Entry; + +use crate::contract::{RingContract, Violation}; +use crate::{Completion, Token}; + +/// A map from `UserData` to an unclaimed [`Token`], optionally checked. +/// +/// `X` is per-operation data the caller wants back when the completion +/// arrives: a registered-slot index, an expected length, a phase. It defaults +/// to `()`. +pub struct Pending { + entries: HashMap, X)>, + contract: Option, +} + +impl Pending { + /// An unchecked map: bookkeeping only, no oracle. + #[must_use] + pub fn new() -> Self { + Self { + entries: HashMap::new(), + contract: None, + } + } + + /// A map that drives its own [`RingContract`]. + /// + /// The oracle is *owned* rather than borrowed, which is what makes drift + /// impossible: there is no way to record a push in one and not the other. + /// The cost is that a caller running several maps against one shared + /// contract cannot use this -- see the module docs. + #[must_use] + pub fn checked() -> Self { + Self { + entries: HashMap::new(), + contract: Some(RingContract::new()), + } + } + + /// Record a pushed operation and take ownership of its token. + /// + /// # Panics + /// + /// If `token`'s identity is already pending. That is not a caller mistake + /// this type should paper over: `UserData` is minted unique per operation, + /// so a collision means either the ring's accounting is wrong or a token + /// was pushed twice, and both are worse than the panic. + pub fn push(&mut self, token: Token, extra: X) { + let user_data = token.id(); + match self.entries.entry(user_data) { + Entry::Occupied(_) => { + panic!("user_data {user_data:#x} is already pending; identities are minted unique") + } + Entry::Vacant(slot) => { + slot.insert((token, extra)); + } + } + if let Some(contract) = &mut self.contract { + contract.observe_push(user_data); + } + } + + /// Claim the token matching `completion`, returning its payload and sidecar. + /// + /// `None` when this map never held that identity, which is not an error: + /// a consumer draining a ring it shares with tokenless operations will see + /// completions it did not push. + pub fn claim(&mut self, completion: &Completion) -> Option<(T, X)> { + let user_data = completion.user_data(); + let (token, extra) = self.entries.remove(&user_data)?; + if let Some(contract) = &mut self.contract { + contract.observe_completion(user_data); + } + match token.claim_if(completion) { + Ok(payload) => { + if let Some(contract) = &mut self.contract { + contract.observe_claim(user_data); + } + Some((payload, extra)) + } + // Looked up *by* this completion's identity, so a mismatch would + // mean `claim_if` disagrees with the key it was found under. Put it + // back rather than dropping it: an unclaimed token that is silently + // discarded here is exactly the leak this type exists to prevent. + Err(token) => { + self.entries.insert(user_data, (token, extra)); + None + } + } + } + + /// Give up on an identity deliberately, recording it as such. + /// + /// The escape hatch that keeps `Drop` honest. A consumer tearing down early + /// has genuinely abandoned these tokens, and the oracle distinguishes that + /// from having forgotten them. + pub fn abandon(&mut self, user_data: usize) -> bool { + let removed = self.entries.remove(&user_data).is_some(); + if removed && let Some(contract) = &mut self.contract { + contract.observe_deliberate_leak(user_data); + } + removed + } + + /// How many operations are still awaiting their completion. + #[must_use] + pub fn len(&self) -> usize { + self.entries.len() + } + + /// Whether every pushed operation has been claimed. + #[must_use] + pub fn is_empty(&self) -> bool { + self.entries.is_empty() + } + + /// The oracle, when this map is checked. + #[must_use] + pub fn contract(&self) -> Option<&RingContract> { + self.contract.as_ref() + } + + /// Consume the map, returning whatever conservation rules were broken. + /// + /// Empty means every token was claimed and the oracle -- if there is one -- + /// saw a well-formed sequence. This is the graceful counterpart to `Drop`: + /// calling it says the caller has checked, so teardown stays quiet. + #[must_use] + pub fn finish(mut self) -> Vec { + let mut violations = match &self.contract { + Some(contract) => contract.check_quiescent(), + None => Vec::new(), + }; + for user_data in self.entries.keys() { + violations.push(Violation::LeakedToken { + user_data: *user_data, + }); + } + // Checked deliberately, so `Drop` has nothing left to complain about. + self.entries.clear(); + self.contract = None; + violations + } +} + +impl Default for Pending { + fn default() -> Self { + Self::new() + } +} + +impl std::fmt::Debug for Pending { + /// Identities and counts, never payloads: a pending token routinely holds + /// someone's data, and a `Debug` that printed it would put that data + /// wherever a caller logs. + fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result { + let mut ids: Vec = self.entries.keys().copied().collect(); + ids.sort_unstable(); + f.debug_struct("Pending") + .field("outstanding", &self.entries.len()) + .field("user_data", &ids) + .field("checked", &self.contract.is_some()) + .finish() + } +} + +impl Drop for Pending { + /// Panics if tokens are still held. + /// + /// # Why this is a panic rather than a log + /// + /// A token dropped unclaimed leaks its buffer **on purpose** -- `Token` + /// forgets the value rather than freeing it, because the kernel may still + /// be writing there. So this state is not untidy, it is a deliberate leak + /// that nobody decided to take, and it is invisible at runtime: the program + /// keeps working and loses memory. + /// + /// A caller who means it says so, with [`Pending::abandon`] per identity or + /// [`Pending::finish`] for the whole map. Both leave this silent. + /// + /// Suppressed while already panicking, because a second panic during unwind + /// aborts the process and replaces the original failure -- which in a test + /// would hide the assertion that actually fired. + fn drop(&mut self) { + if self.entries.is_empty() || std::thread::panicking() { + return; + } + let mut ids: Vec = self.entries.keys().copied().collect(); + ids.sort_unstable(); + panic!( + "{} token(s) dropped unclaimed, leaking their buffers: {ids:#x?}. \ + Claim them, or say so with `abandon` or `finish`.", + ids.len() + ); + } +} + +#[cfg(test)] +mod tests; diff --git a/crates/windows-ioring-sys/src/pending/tests.rs b/crates/windows-ioring-sys/src/pending/tests.rs new file mode 100644 index 000000000..8ba577559 --- /dev/null +++ b/crates/windows-ioring-sys/src/pending/tests.rs @@ -0,0 +1,185 @@ +// Copyright (c) 2026 Mike Grier +//! Tests for the `M23.3` spike. +//! +//! # What these establish, and what they leave to a real ring +//! +//! Everything about the *rule* -- that an unclaimed token is loud, that saying +//! so deliberately is quiet, that a checked map reports what a raw `HashMap` +//! cannot -- needs only `push`, so it is settled here, hermetically, with a +//! ledger and no ring at all. +//! +//! `claim` is the one operation that needs a real `Completion`, and a +//! `Completion` carries a `RingId` that only a ring mints. Rather than add a +//! synthetic constructor whose fidelity would then itself need arguing, that +//! half is validated by converting a real consumer -- stronger evidence anyway, +//! since it answers whether the type *fits* as well as whether it works. + +use crate::Token; +use crate::accounting::Accounting; +use crate::contract::Violation; + +use super::Pending; + +/// A payload that is `Send + 'static` and carries nothing interesting. +#[derive(Debug, PartialEq, Eq)] +struct Payload(u32); + +fn token(ledger: &mut Accounting, value: u32) -> Token { + Token::new(ledger, Payload(value)).expect("mint a token") +} + +/// The map counts what it holds. +#[test] +fn pushing_makes_an_operation_outstanding() { + let mut ledger = Accounting::new(); + let mut pending: Pending = Pending::new(); + assert!(pending.is_empty()); + + let first = token(&mut ledger, 1); + let first_id = first.id(); + pending.push(first, ()); + let second = token(&mut ledger, 2); + let second_id = second.id(); + pending.push(second, ()); + assert_eq!(pending.len(), 2); + + // Deliberately emptied, or `Drop` would fire -- which is the next test. + assert!(pending.abandon(first_id)); + assert!(pending.abandon(second_id)); +} + +/// **The claim this spike exists to test.** An unclaimed token is loud. +/// +/// A token dropped unclaimed leaks its buffer on purpose -- `Token` forgets the +/// value rather than freeing it, because the kernel may still be writing. That +/// is invisible at runtime: the program keeps working and loses memory. A raw +/// `HashMap` cannot notice; this can. +#[test] +#[should_panic(expected = "dropped unclaimed")] +fn dropping_a_map_that_still_holds_tokens_panics() { + let mut ledger = Accounting::new(); + let mut pending: Pending = Pending::new(); + pending.push(token(&mut ledger, 1), ()); + // Dropped here, still holding one token. +} + +/// Saying so deliberately is quiet. +/// +/// The other direction of the same guard, and the one that makes it usable: a +/// consumer tearing down early has genuinely abandoned these tokens and must be +/// able to say so without tripping an alarm meant for forgetting. +#[test] +fn abandoning_every_token_leaves_the_drop_silent() { + let mut ledger = Accounting::new(); + let mut pending: Pending = Pending::new(); + let held = token(&mut ledger, 1); + let held_id = held.id(); + pending.push(held, ()); + assert!(pending.abandon(held_id)); + // Drops silently. +} + +/// `finish` reports what was left, and leaves `Drop` nothing to say. +#[test] +fn finish_reports_unclaimed_tokens_instead_of_panicking() { + let mut ledger = Accounting::new(); + let mut pending: Pending = Pending::new(); + let held = token(&mut ledger, 7); + let held_id = held.id(); + pending.push(held, ()); + + let violations = pending.finish(); + assert_eq!( + violations, + vec![Violation::LeakedToken { user_data: held_id }] + ); +} + +/// A map that claimed everything finishes clean. +#[test] +fn an_empty_map_finishes_with_no_violations() { + let pending: Pending = Pending::new(); + assert!(pending.finish().is_empty()); +} + +/// A checked map's oracle sees the pushes the map recorded. +/// +/// This is the wiring the spike is really about: one call site updated both, +/// where every real consumer today updates them separately and can drift. +#[test] +fn a_checked_map_drives_its_own_contract() { + let mut ledger = Accounting::new(); + let mut pending: Pending = Pending::checked(); + let held = token(&mut ledger, 3); + let held_id = held.id(); + pending.push(held, ()); + + let contract = pending.contract().expect("a checked map has one"); + assert!( + contract + .check_quiescent() + .contains(&Violation::Outstanding { user_data: held_id }), + "the oracle should know about a push nobody told it about by hand" + ); + + let violations = pending.finish(); + assert!(violations.contains(&Violation::Outstanding { user_data: held_id })); +} + +/// An unchecked map has no oracle, and says so. +#[test] +fn an_unchecked_map_has_no_contract() { + let pending: Pending = Pending::new(); + assert!(pending.contract().is_none()); + let _ = pending.finish(); +} + +/// The sidecar comes back with the payload. +/// +/// The generic the census forced: two thirds of the real sites carry +/// per-operation data beside the token, so a map holding only tokens would +/// serve the minority. +#[test] +fn a_sidecar_is_stored_alongside_the_token() { + let mut ledger = Accounting::new(); + let mut pending: Pending = Pending::new(); + let held = token(&mut ledger, 1); + let held_id = held.id(); + pending.push(held, (17, "slot")); + assert_eq!(pending.len(), 1); + assert!(pending.abandon(held_id)); +} + +/// Two tokens cannot share an identity. +#[test] +#[should_panic(expected = "already pending")] +fn pushing_a_duplicate_identity_panics() { + let mut ledger = Accounting::new(); + let mut pending: Pending = Pending::new(); + let held = token(&mut ledger, 1); + let id = held.id(); + pending.push(held, ()); + + // A second ledger mints from zero again, so this collides by construction. + let mut other_ledger = Accounting::new(); + let clash = token(&mut other_ledger, 2); + assert_eq!(clash.id(), id, "a fresh ledger restarts identities"); + pending.push(clash, ()); +} + +/// `Debug` shows identities and counts, never payloads. +#[test] +fn debug_does_not_print_payloads() { + let mut ledger = Accounting::new(); + let mut pending: Pending = Pending::new(); + let held = token(&mut ledger, 0xBEEF); + let held_id = held.id(); + pending.push(held, ()); + + let rendered = format!("{pending:?}"); + assert!(rendered.contains("outstanding"), "{rendered}"); + assert!(!rendered.contains("BEEF"), "payload leaked into Debug"); + assert!(!rendered.contains("beef"), "payload leaked into Debug"); + + assert!(pending.abandon(held_id)); +} diff --git a/crates/windows-ioring-sys/src/ring.rs b/crates/windows-ioring-sys/src/ring.rs index a69802227..ecd0ca3d4 100644 --- a/crates/windows-ioring-sys/src/ring.rs +++ b/crates/windows-ioring-sys/src/ring.rs @@ -4,42 +4,21 @@ use std::ffi::c_void; use std::io; use std::os::windows::io::{AsRawHandle, FromRawHandle, OwnedHandle}; -use std::sync::atomic::{AtomicU64, Ordering}; +use std::time::{Duration, Instant}; use windows_sys::Win32::Storage::FileSystem::{ CloseIoRing, CreateIoRing, GetIoRingInfo, IORING_BUFFER_INFO, IORING_CQE, IORING_CREATE_ADVISORY_FLAGS_NONE, IORING_CREATE_FLAGS, IORING_CREATE_REQUIRED_FLAGS_NONE, IORING_INFO, IORING_OP_CANCEL, IORING_OP_CODE, IORING_OP_FLUSH, IORING_OP_NOP, IORING_OP_READ, IORING_OP_REGISTER_BUFFERS, IORING_OP_REGISTER_FILES, IORING_OP_WRITE, IsIoRingOpSupported, - PopIoRingCompletion, SetIoRingCompletionEvent, SubmitIoRing, }; use windows_sys::Win32::System::Threading::{CreateEventW, SetEvent}; +use crate::accounting::Accounting; use crate::capability::{RingVersion, capabilities}; use crate::error::check; -/// A ring's identity, unique for the process's lifetime (PR #20 review -/// response): every value a ring hands out that later gets checked back -/// against it -- a [`crate::Token`], a [`crate::RegisteredFile`], a -/// [`crate::RegisteredBuffers`] -- carries the id of the ring that minted -/// it, and every [`Completion`] carries the id of the ring that popped it. -/// -/// A monotonic counter rather than the ring's own `HANDLE`: a `HANDLE` is -/// only unique while the object it names is still open, and Windows is free -/// to hand a closed ring's numeric value to the *next* object created -- -/// which would let a stale identity from a closed ring collide with a -/// brand-new one. This counter never repeats within one process run -/// (`u64` overflow is not a practical concern), so a mismatch always means -/// a genuine cross-ring mixup, never a false negative from handle reuse. -#[derive(Clone, Copy, Debug, PartialEq, Eq, Hash)] -pub(crate) struct RingId(u64); - -impl RingId { - fn next() -> Self { - static NEXT: AtomicU64 = AtomicU64::new(1); - Self(NEXT.fetch_add(1, Ordering::Relaxed)) - } -} +pub(crate) use crate::accounting::RingId; /// One `IoRing` operation. /// @@ -238,6 +217,29 @@ impl Completion { /// op-specific value in `IORING_CQE::Information`, once `ResultCode` /// says success. /// + /// # The count may be short, and the remainder is yours + /// + /// A successful read or write may report **fewer bytes than were + /// requested**, including zero (`RS-P-8` in + /// [RESPONSE-SPACE.md](../RESPONSE-SPACE.md)). `WriteFile` documents this + /// for non-blocking byte-mode pipes; sockets report a short send when the + /// transmit buffer is full, and reads are short at end of file. This crate + /// takes a handle and does not constrain what kind it is, so a consumer + /// must compare this count against the length it submitted rather than + /// assume they are equal. + /// + /// Nothing here reissues the remainder. Whether to submit another + /// operation for it, how many times, and when to give up are the caller's + /// (D-67), and a consumer that loops must tolerate a completion that makes + /// no progress. + /// + /// A consumer that has narrowed its *own* handle type can rely on more -- + /// for an ordinary file on a local volume a successful completion is + /// expected to carry the full length, and a full volume is an error rather + /// than a short success. That is a guarantee such a consumer earns by + /// constraining the handle, and it belongs in its contract rather than + /// being assumed from this one. + /// /// # Errors /// /// Returns the wrapped [`crate::IoRingError`] if `ResultCode` is a @@ -345,6 +347,55 @@ impl Completion { } } + /// Report a **successful** transfer of `transferred` bytes instead of what + /// this operation actually reported (M26.11). + /// + /// # What this is for, and why the failure seam cannot do it + /// + /// [`Completion::with_injected_failure`] models an operation that failed. + /// `RS-P-8` describes something different and stranger: an operation that + /// **succeeded** while moving fewer bytes than were asked for. Windows + /// documents that for non-blocking byte-mode pipes and it happens on + /// sockets, but an ordinary file on a local volume does not do it -- so a + /// consumer's handling of a short count is written once against the + /// documentation and then never executed again, which is exactly the class + /// of path the failure seam was introduced for. + /// + /// A consumer that *narrows* its handle type may legitimately require + /// complete transfers. This seam is how such a consumer tests that its + /// requirement is enforced rather than merely stated -- `epoch_log`'s + /// appender uses it for precisely that. + /// + /// # Why this one is inert + /// + /// The result code is left successful and only the byte count moves, so + /// every claim path behaves exactly as it would for the real completion: + /// [`crate::Token::claim_if`] returns the buffer, and no path keys memory + /// ownership off the transferred count. That makes this seam free of the + /// registration hazard documented on + /// [`Completion::with_injected_failure`], which arises only because a + /// *failed* registration is taken as proof the kernel retained nothing. + /// + /// # Example + /// + /// ```ignore + /// // A handle whose successful writes are complete is a requirement this + /// // consumer states; here we hand it one that violates the requirement. + /// let completion = completion.with_injected_transfer(RECORD_STRIDE - 1); + /// assert_eq!(completion.result().expect("still a success"), RECORD_STRIDE - 1); + /// ``` + #[cfg(any(test, feature = "fault-injection"))] + #[must_use] + pub fn with_injected_transfer(self, transferred: usize) -> Self { + Self { + // Deliberately *not* touched: a short transfer under `RS-P-8` is a + // success, and turning it into a failure would model the one thing + // this seam exists to distinguish it from. + information: transferred, + ..self + } + } + /// Build a `Completion` without popping a real one, for tests that /// exercise [`crate::Token::claim_if`] without real I/O. /// @@ -379,6 +430,73 @@ const S_FALSE: windows_sys::core::HRESULT = 1; /// waits are bounded and rechecked, not unbounded"). const RUN_DOWN_POLL_MS: u32 = 50; +/// `IORING_E_WAIT_TIMEOUT`: `SubmitIoRing` submitted every entry successfully +/// and the subsequent wait then timed out. +/// +/// **This is the documented code, not an observed one.** `SubmitIoRing`'s +/// reference page gives it its own row and states the consequence that matters +/// here: *"All operations were submitted without error and the subsequent wait +/// timed out."* Its Remarks then draw the line this crate depends on -- *"If +/// this function returns an error other than IORING_E_WAIT_TIMEOUT, then all +/// entries remain in the submission queue."* So this value is the difference +/// between "the work went in" and "the work is still queued", which is why it +/// is classified rather than passed to [`check`](crate::error::check). +/// +/// **Spelled by derivation because the bindings do not carry the name.** +/// `IORING_E_WAIT_TIMEOUT` is a macro over `HRESULT_FROM_WIN32(ERROR_TIMEOUT)` +/// rather than a `FACILITY_IORING` code -- that facility defines only +/// `0x8046_0001` through `0x8046_0008`, none of them a timeout -- so +/// `windows-sys` emits no constant for it and there is nothing to import. The +/// derivation below is therefore the name, and is written out so the next +/// reader does not re-derive it from a run. The `0x8007_0000` is +/// `HRESULT_FROM_WIN32`'s severity-plus-`FACILITY_WIN32` prefix, which that +/// macro ors onto any code of `0xFFFF` or less. +const IORING_E_WAIT_TIMEOUT: windows_sys::core::HRESULT = + (0x8007_0000_u32 | windows_sys::Win32::Foundation::ERROR_TIMEOUT) as windows_sys::core::HRESULT; + +/// The largest `timeout_ms` a [`CompletionWait`] is ever handed: one below +/// `u32::MAX`, because `u32::MAX` is Win32's `INFINITE`. +/// +/// A waiter built on `WaitForSingleObject` or `WaitForMultipleObjects` -- both +/// of which [`CompletionWait`] explicitly invites -- reads that value as "no +/// timeout", so saturating onto it would convert a long but finite bound into +/// an unbounded block, with the pop loop unable to re-check its own deadline +/// until the wait returned. +const MAX_WAIT_MS: u32 = u32::MAX - 1; + +/// Whether a `SubmitIoRing` result means every entry was submitted. +/// +/// `S_OK` and `IORING_E_WAIT_TIMEOUT` both do, per that call's documented +/// return values; every other error means the opposite, and its Remarks say +/// so in as many words -- the entries remain in the submission queue. +/// +/// Exposed as one predicate because three callers need the same answer and a +/// fourth got it wrong for a year: [`IoRing::pop_within`] reported a timeout +/// as a failure until `M21.6`, and [`crate::Batch::submit_and_wait`] still did +/// until `M26.8`, because that fix swept two of the three sites. +pub(crate) fn every_entry_was_submitted(hr: windows_sys::core::HRESULT) -> bool { + hr == IORING_E_WAIT_TIMEOUT || check(hr).is_ok() +} + +/// Classify the result of a `SubmitIoRing` call made **only** to wait. +/// +/// An expired wait is the ordinary outcome of asking to block for a bounded +/// time, not a failure -- so it is `Ok`, and the caller re-checks whatever it +/// was waiting for. Everything else is a real error. +/// +/// This exists because getting it wrong is silent and was: `check(hr)` +/// straight through turns every timed-out wait into an `Err`, which made +/// [`IoRing::pop_within`] report a timeout as a failure rather than as the +/// `Ok(None)` it documents, and made [`IoRing::run_down`] treat any operation +/// slower than [`RUN_DOWN_POLL_MS`] as fatal. One classification, two callers, +/// so they cannot disagree again. +fn wait_outcome(hr: windows_sys::core::HRESULT) -> io::Result<()> { + if hr == IORING_E_WAIT_TIMEOUT { + return Ok(()); + } + check(hr) +} + /// An owned `IoRing`, closed with `CloseIoRing` on drop. /// /// Not `Clone`: cloning would give two owners of the same native ring, and @@ -394,18 +512,10 @@ pub struct IoRing { handle: *mut c_void, version: RingVersion, supported_ops: OpSupport, - /// This ring's own identity (PR #20 review response); see [`RingId`]. - ring_id: RingId, - /// The next `UserData` value [`IoRing::reserve_user_data`] will hand out. - next_user_data: usize, - /// Operations minted but not yet observed to have completed (M2.4). - outstanding: usize, - /// How many file handles are registered so far, across every confirmed - /// `BuildIoRingRegisterFileHandles` (M5.1). The base index of the next - /// registration. - registered_files: u32, - /// As `registered_files`, for `BuildIoRingRegisterBuffers` (M5.2). - registered_buffers: u32, + /// The half of this ring that never touches the kernel: identity, + /// operation identities, and the counts (M24.2). Split out so those rules + /// can be tested without opening a ring -- see [`crate::accounting`]. + accounting: Accounting, /// The `IORING_BUFFER_INFO` array handed to `BuildIoRingRegisterBuffers`, /// kept alive because the kernel reads it when the registration op /// *runs*, not when the `Build*` call returns (D-32, measured). @@ -437,11 +547,11 @@ impl std::fmt::Debug for IoRing { .field("handle", &self.handle) .field("version", &self.version) .field("supported_ops", &self.supported_ops) - .field("ring_id", &self.ring_id) - .field("next_user_data", &self.next_user_data) - .field("outstanding", &self.outstanding) - .field("registered_files", &self.registered_files) - .field("registered_buffers", &self.registered_buffers) + // One field rather than five, because the ledger derives `Debug` + // and prints its own. Keeping the five spelled out here would be a + // second copy of the field list, drifting the moment either side + // gains a field. + .field("accounting", &self.accounting) .field( "registered_buffer_infos", &self.registered_buffer_infos.len(), @@ -507,11 +617,7 @@ impl IoRing { handle, version, supported_ops, - ring_id: RingId::next(), - next_user_data: 0, - outstanding: 0, - registered_files: 0, - registered_buffers: 0, + accounting: Accounting::new(), registered_buffer_infos: Vec::new(), completion_event: None, }) @@ -633,6 +739,44 @@ impl IoRing { /// re-arms the edge, must run to empty exactly once. Two threads waiting /// on one ring's event cannot be made correct. /// + /// # If you are arming a thread-pool wait on this handle + /// + /// `SetThreadpoolWait` documents that "you must re-register the event + /// with the wait object before signaling it each time to trigger the wait + /// callback". This method signals the event *before* it returns (rule 2 + /// above), so by the time you have a handle to build a wait object + /// around, that setup signal has already happened -- in the order the + /// rule forbids. It is not guaranteed to run your callback, and because + /// the event is auto-reset the signal is *consumed* rather than left + /// pending, so a later arming has nothing to observe. Combined with the + /// edge rule above, a ring whose queue never returns to empty has no + /// second wakeup coming: the loss is permanent, not late. + /// + /// The remedy needs nothing this method does not already give you -- + /// after arming the wait, signal your own duplicate yourself: /// + /// ```no_run + /// # use windows_ioring_sys::IoRing; + /// # use std::os::windows::io::AsRawHandle; + /// # fn f(ring: &mut IoRing) -> std::io::Result<()> { + /// let event = ring.completion_event()?; + /// // ... build the wait object around `event`, then arm it ... + /// // Only now raise the wakeup, in the documented order. A wake with + /// // nothing to pop is normal (rule 2), so this is always safe. + /// unsafe { windows_sys::Win32::System::Threading::SetEvent(event.as_raw_handle()) }; + /// # Ok(()) + /// # } + /// ``` + /// + #[cfg_attr( + feature = "threadpool", + doc = "[`EventDelivery`](crate::EventDelivery) already does this for you and" + )] + #[cfg_attr( + not(feature = "threadpool"), + doc = "`EventDelivery` (the default `threadpool` feature) already does this for you and" + )] + /// is the better answer if you do not need the handle itself. + /// /// `examples/model_b_multiplexed.rs` is this whole shape worked end to /// end -- a caller-owned ring waited on alongside a shutdown latch, with /// the quiesce that shutdown-while-outstanding requires (M11.6). @@ -646,12 +790,37 @@ impl IoRing { /// returns any error from `CreateEventW`, /// `SetIoRingCompletionEvent`, `SetEvent`, or duplicating the handle. pub fn completion_event(&mut self) -> io::Result { + let (event, owes_setup_signal) = self.attach_completion_event_unsignalled()?; + if owes_setup_signal { + self.raise_setup_signal()?; + } + Ok(event) + } + + /// Attach the ring's completion event *without* raising the setup signal, + /// reporting whether that signal is still owed. + /// + /// `SetThreadpoolWait` documents that "you must re-register the event with + /// the wait object before signaling it each time to trigger the wait + /// callback". A caller that is about to arm a threadpool wait on this + /// event therefore needs the attachment and the signal as two steps, so + /// that arming can be sequenced between them; handing back an + /// already-signalled event leaves that caller no way to obey the rule. + /// [`IoRing::completion_event`] is these two composed, for a caller who + /// does its own waiting and is not bound by that rule. + /// + /// The flag is false when the ring already had an event attached, which + /// matches [`IoRing::completion_event`]: the setup signal belongs to the + /// call that performs the attachment. + pub(crate) fn attach_completion_event_unsignalled( + &mut self, + ) -> io::Result<(OwnedHandle, bool)> { // Already attached: hand back another duplicate rather than // attaching a second event, which would silently detach the first // (`SetIoRingCompletionEvent` replaces rather than adds). The // capability was necessarily verified on the call that attached it. if let Some(event) = &self.completion_event { - return event.try_clone(); + return Ok((event.try_clone()?, false)); } if !capabilities()?.supports_completion_event { @@ -662,8 +831,10 @@ impl IoRing { } // Auto-reset (manual_reset = FALSE) per D-21, initially unsignalled - // -- the deliberate setup signal is raised below, after the event is - // attached and owned, so it cannot be missed or lost. + // -- the deliberate setup signal is raised by `raise_setup_signal`, + // after the event is attached and owned, so it cannot be lost, and + // after any threadpool wait has been armed, so the arm-before-signal + // rule `SetThreadpoolWait` documents is obeyed. // SAFETY: null attributes and name are documented defaults. let raw = unsafe { CreateEventW(std::ptr::null(), 0, 0, std::ptr::null()) }; if raw.is_null() { @@ -675,25 +846,56 @@ impl IoRing { // SAFETY: `self.handle` is a live ring; `event` is a live event that // this ring will own for the rest of its life once stored below. - let hr = unsafe { SetIoRingCompletionEvent(self.handle, event.as_raw_handle()) }; + let hr = unsafe { crate::sys::set_completion_event(self.handle, event.as_raw_handle()) }; // On failure `event` drops here, closing a handle the ring never // successfully referenced. check(hr)?; - // Stored *before* signalling: from this point the ring owns the - // event, so no later failure can drop it and leave the ring - // signalling a closed (possibly recycled) handle. - let event = self.completion_event.insert(event); + // Stored *before* the setup signal can be raised: from this point the + // ring owns the event, so no later failure can drop it and leave the + // ring signalling a closed (possibly recycled) handle. + self.completion_event + .insert(event) + .try_clone() + .map(|dup| (dup, true)) + } - // The one deliberate spurious wakeup (rule 2 above): a caller who - // submitted before attaching would otherwise never be woken for that - // backlog, since the queue never returns to empty to re-arm the edge. + /// Raise the one deliberate setup signal on the attached completion event. + /// + /// This is the single spurious wakeup the event's contract allows for: a + /// caller who submitted before attaching would otherwise never be woken + /// for that backlog, since the queue never returns to empty and so never + /// re-arms the edge (D-19). + /// + /// Separated from the attachment so that a caller arming a threadpool wait + /// can obey `SetThreadpoolWait`'s documented ordering -- register first, + /// signal second. Does nothing if no event is attached, which cannot + /// happen on the paths that call it. + pub(crate) fn raise_setup_signal(&self) -> io::Result<()> { + let Some(event) = &self.completion_event else { + return Ok(()); + }; // SAFETY: `event` is a live event handle this ring owns. if unsafe { SetEvent(event.as_raw_handle()) } == 0 { return Err(io::Error::last_os_error()); } - - event.try_clone() + // The setup signal is what a waiter attaching to a backlog depends on + // entirely, so M26.9's investigation needs to know it happened -- and + // that it happened on the ring's own handle rather than a duplicate + // handed to a caller, since only the former is what the kernel will go + // on signalling. + // + // Gated because `windows-threadpool-sys` is optional (D-22): this + // method is on the always-present path, unlike the delivery module, + // so an ungated reference breaks `--no-default-features`. + #[cfg(feature = "threadpool")] + windows_threadpool_sys::trace_record!( + "delivery", + "setup-signalled", + event.as_raw_handle() as usize, + self.accounting.outstanding() + ); + Ok(()) } /// Whether this ring supports a raw op code, including one this crate @@ -737,7 +939,7 @@ impl IoRing { /// index for. #[must_use] pub fn registered_file_count(&self) -> u32 { - self.registered_files + self.accounting.registered_file_count() } /// As [`IoRing::registered_file_count`], for registered buffers (M5.2) -- @@ -745,7 +947,7 @@ impl IoRing { /// consequences (M10.3, D-31). #[must_use] pub fn registered_buffer_count(&self) -> u32 { - self.registered_buffers + self.accounting.registered_buffer_count() } /// Advance the registered-file base index by `count`, the instant a @@ -768,12 +970,12 @@ impl IoRing { /// which is a different thing from when it claims the *indices*. The /// latter remains unmeasured, and dissolved rather than resolved. pub(crate) fn reserve_registered_files(&mut self, count: u32) { - self.registered_files = self.registered_files.saturating_add(count); + self.accounting.reserve_registered_files(count); } /// As [`IoRing::reserve_registered_files`], for registered buffers. pub(crate) fn reserve_registered_buffers(&mut self, count: u32) { - self.registered_buffers = self.registered_buffers.saturating_add(count); + self.accounting.reserve_registered_buffers(count); } /// Take ownership of the `IORING_BUFFER_INFO` array a @@ -808,7 +1010,7 @@ impl IoRing { /// `record_completion`). #[must_use] pub fn outstanding(&self) -> usize { - self.outstanding + self.accounting.outstanding() } /// Mint a fresh `UserData` identity for a new operation, and account for @@ -826,19 +1028,14 @@ impl IoRing { /// is ever exhausted, mirroring `windows-threadpool-sys`'s own /// "exhausting the generation sequence fails rather than wraps." pub(crate) fn reserve_user_data(&mut self) -> io::Result { - let id = self.next_user_data; - self.next_user_data = id - .checked_add(1) - .ok_or_else(|| io::Error::other("IoRing operation identity space exhausted"))?; - self.outstanding += 1; - Ok(id) + self.accounting.reserve_user_data() } /// Record that one outstanding operation's completion has been observed /// (a real `IORING_CQE` was popped for it), whether or not a live /// [`crate::Token`] was still around to claim it. pub(crate) fn record_completion(&mut self) { - self.outstanding = self.outstanding.saturating_sub(1); + self.accounting.record_completion(); } /// Release a reservation for an operation that was never actually @@ -850,7 +1047,16 @@ impl IoRing { /// the op never entered the queue, so it must not count against /// [`IoRing::run_down`] either. pub(crate) fn cancel_reservation(&mut self) { - self.outstanding = self.outstanding.saturating_sub(1); + self.accounting.cancel_reservation(); + } + + /// This ring's ledger, for the crate's own minting paths (M24.2). + /// + /// Handed out rather than proxied so that a `Token` can be minted from + /// the bookkeeping alone -- which is what lets `token.rs`'s tests run + /// without a kernel ring, since minting is all they ever needed one for. + pub(crate) fn accounting_mut(&mut self) -> &mut Accounting { + &mut self.accounting } /// This ring's native handle, for `batch.rs`'s `Build*`/`Submit` calls. @@ -862,7 +1068,7 @@ impl IoRing { /// registration it mints and checking against on use (PR #20 review /// response); see [`RingId`]. pub(crate) fn ring_id(&self) -> RingId { - self.ring_id + self.accounting.ring_id() } /// Queue a raw, not-yet-wrapped SQE via a caller-supplied `Build*` call @@ -913,20 +1119,112 @@ impl IoRing { /// add the typed completion path `Token` consumes. Idempotent: calling it /// again once `outstanding() == 0` is a no-op. /// + /// # A poll that expires is not a failure + /// + /// Each poll blocks for `RUN_DOWN_POLL_MS` and then reports + /// `ERROR_TIMEOUT` if nothing finished in that window, which is the + /// ordinary outcome for any operation slower than 50ms. Treating that as + /// an error -- which this did until M21.6 -- made `run_down` return `Err` + /// with the operation still outstanding, so `Drop` asserted and then + /// called `CloseIoRing` anyway: exactly the "the kernel may still be + /// writing through a token's buffer" hazard this function exists to + /// prevent. + /// + /// This loop therefore has no overall bound, and that is deliberate. + /// **Every SQE that successfully queues produces exactly one completion** + /// (M10.2), so it terminates. Blocking until that holds is the safe + /// failure mode; closing the ring early is not. + /// + /// # Choosing an unbounded wait is the caller's to make + /// + /// This spelling waits until rundown finishes, however long that takes, + /// and calling it is how a caller elects that. A caller who wants to + /// decide for themselves -- a deadline, a backoff, a number of attempts + /// before giving up -- calls [`IoRing::run_down_within`] instead and owns + /// the policy entirely. **This crate does not implement retry policy** + /// (M26.8): it supplies a bounded primitive and reports what happened. + /// + /// Note the difference from [`IoRing::pop_within`], which also waits in + /// segments: there the segments sit *inside a period the caller supplied*, + /// which is a bounded wait implemented properly rather than a policy. This + /// method had segments and no such period, which is what made it the one + /// waiting API in this crate shaped wrongly. + /// /// # Errors /// - /// Returns any error from `SubmitIoRing` or `PopIoRingCompletion`. + /// Returns any error from `SubmitIoRing` other than an expired wait, or + /// from `PopIoRingCompletion`. **An error leaves the ring resumable**: see + /// [`IoRing::run_down_within`] for what is guaranteed about the operations + /// still queued. pub fn run_down(&mut self) -> io::Result<()> { - while self.outstanding > 0 { + while !self.run_down_within(Duration::MAX)? {} + Ok(()) + } + + /// Run down for at most `bound`, reporting whether it finished. + /// + /// `Ok(true)` means nothing is outstanding and the ring is safe to drop. + /// `Ok(false)` means the bound elapsed with work still in flight -- call + /// again when your own policy says to. [`IoRing::outstanding`] says how + /// much is left. + /// + /// # Why this exists, and why it returns rather than retries + /// + /// Rundown can fail for a reason that a later attempt would survive, and + /// **deciding whether to make that attempt is not this crate's business**. + /// A caller running under a deadline, a supervisor with a backoff, and a + /// test that wants to fail fast all want different answers, and a policy + /// baked in here would be wrong for two of the three. So this waits for + /// exactly as long as it is told and then reports. + /// + /// # What an error guarantees, which is what makes retrying safe + /// + /// `SubmitIoRing` documents that *"If this function returns an error other + /// than IORING_E_WAIT_TIMEOUT, then all entries remain in the submission + /// queue."* So a failure here has **not** lost the operations and has not + /// rewound them ([D-5](../DESIGN-NOTES.md#d-5)); they are still ring state, + /// a later submit is what runs them, and their buffers must stay alive + /// until they complete. + /// + /// The consequence worth stating plainly: after an `Err`, **do not drop + /// this ring**. Dropping it closes a ring the kernel may still write + /// through, which is the hazard rundown exists to prevent. Call again. + /// + /// # Errors + /// + /// Returns any error from `SubmitIoRing` other than an expired wait, or + /// from `PopIoRingCompletion`. + pub fn run_down_within(&mut self, bound: Duration) -> io::Result { + let deadline = Instant::now().checked_add(bound); + loop { + if self.accounting.outstanding() == 0 { + return Ok(true); + } + // `checked_add` rather than `+`, for the reason `pop_within_with` + // records: `Instant + Duration` panics on overflow, so + // `Duration::MAX` -- the honest spelling of "no deadline" -- would + // take down the process. A deadline the clock cannot represent is + // one that never arrives, which is what was asked for. + let remaining = match deadline { + Some(deadline) => deadline.saturating_duration_since(Instant::now()), + None => Duration::MAX, + }; + if remaining.is_zero() { + return Ok(false); + } + // Segments sit inside the caller's period, never outside it: the + // poll is the shorter of the rundown step and what is left. + let poll_ms = u32::try_from(remaining.as_millis()) + .unwrap_or(u32::MAX) + .clamp(1, RUN_DOWN_POLL_MS); let mut submitted = 0_u32; // SAFETY: `self.handle` is a live ring; valid out-pointer. Zero // new SQEs are queued -- this call's only purpose is to wait for // and reap already-outstanding completions. - let hr = unsafe { SubmitIoRing(self.handle, 1, RUN_DOWN_POLL_MS, &raw mut submitted) }; - check(hr)?; + let hr = unsafe { crate::sys::submit(self.handle, 1, poll_ms, &raw mut submitted) }; + wait_outcome(hr)?; self.drain_for_rundown()?; } - Ok(()) } /// Pop every currently available completion, recording each -- without @@ -972,7 +1270,7 @@ impl IoRing { Information: 0, }; // SAFETY: `self.handle` is a live ring; valid out-pointer. - let hr = unsafe { PopIoRingCompletion(self.handle, &raw mut cqe) }; + let hr = unsafe { crate::sys::pop(self.handle, &raw mut cqe) }; if hr == S_FALSE { return Ok(None); } @@ -982,9 +1280,254 @@ impl IoRing { user_data: cqe.UserData, result_code: cqe.ResultCode, information: cqe.Information, - ring_id: self.ring_id, + ring_id: self.accounting.ring_id(), })) } + + /// Pop one completion, blocking in the ring's own wait until one is + /// available or `timeout` elapses (M21.2). + /// + /// This is the join between [`IoRing::try_pop`], whose `None` means + /// "empty at this instant", and [`crate::Batch::submit_and_wait`], whose + /// return deliberately promises nothing about poppability because its + /// timeout may have expired. Neither one alone answers "give me the + /// completion I just caused", and before this existed every caller wrote + /// that loop again -- four different ways across five sites, two of them + /// unbounded spins. + /// + /// `Ok(None)` means the timeout elapsed, or that **nothing can arrive**: + /// with no operation outstanding and an empty queue, no completion is + /// possible, so this returns immediately rather than sleeping out the + /// full `timeout`. That early return is what turns "you forgot to submit" + /// from a timeout into an instant answer. + /// + /// A zero `timeout` is exactly one [`IoRing::try_pop`], which is the + /// honest reading of "wait no time at all". + /// + /// Uses [`SubmitWait`], which blocks inside `SubmitIoRing` with no new + /// entries queued. It does **not** touch the ring's completion event, so + /// it cannot disturb a caller who owns that event under + /// [D-21](../DESIGN-NOTES.md#d-21). Use + /// [`IoRing::pop_within_with`] to supply a different wait. + /// + /// # Errors + /// + /// Any error from `SubmitIoRing` or `PopIoRingCompletion`. + pub fn pop_within(&mut self, timeout: Duration) -> io::Result> { + self.pop_within_with(&mut SubmitWait, timeout) + } + + /// [`IoRing::pop_within`] with a caller-chosen wait. + /// + /// The crate cannot pick the wait for you, and that is a contract rather + /// than a shrug: the completion event is auto-reset with exactly one + /// waiter per ring ([D-21](../DESIGN-NOTES.md#d-21)), so a wait this crate + /// chose could consume an edge the caller's own loop was entitled to. The + /// owner of the ring is the only party who can discharge that obligation, + /// which is why the choice is a parameter. + /// + /// # Errors + /// + /// Any error from the wait or from `PopIoRingCompletion`. + pub fn pop_within_with( + &mut self, + wait: &mut W, + timeout: Duration, + ) -> io::Result> { + let deadline = Instant::now().checked_add(timeout); + loop { + if let Some(completion) = self.try_pop()? { + return Ok(Some(completion)); + } + // Checked *after* the pop, never before: `record_completion` runs + // during `try_pop`, so reading it first would race the very + // completion being drained. + if self.outstanding() == 0 { + return Ok(None); + } + // `checked_add` rather than `+`: `Instant + Duration` panics on + // overflow, so a caller passing `Duration::MAX` -- a reasonable + // spelling of "no deadline" -- would take down the process. A + // deadline the clock cannot represent is treated as one that + // never arrives, which is what the caller asked for. + let remaining = match deadline { + Some(deadline) => deadline.saturating_duration_since(Instant::now()), + None => Duration::MAX, + }; + if remaining.is_zero() { + return Ok(None); + } + // Clamped into `1..=MAX_WAIT_MS`. The low end stops a + // sub-millisecond remainder becoming a zero timeout, which + // `SubmitIoRing` reads as "poll and return" and would turn the + // tail of every wait into a spin. The high end keeps the waiter + // from ever being handed `u32::MAX`, which is Win32's `INFINITE` + // -- a `WaitForMultipleObjects`-based waiter would block forever + // on a bound its caller believed was finite. + let ms = u32::try_from(remaining.as_millis()) + .unwrap_or(MAX_WAIT_MS) + .clamp(1, MAX_WAIT_MS); + wait.wait(&mut RingWait { ring: self }, ms)?; + } + } +} + +/// How [`IoRing::pop_within_with`] blocks between checks of the completion +/// queue (M21.2). +/// +/// Implement this to drive a bounded pop from a wait this crate does not own +/// -- a completion event the caller already holds, a multiplexed +/// `WaitForMultipleObjects`, or a pure spin on a thread that must not block +/// in the kernel. +pub trait CompletionWait { + /// Block until a completion *may* be available, or `timeout_ms` elapses. + /// + /// **Returning early or spuriously is always permitted**, and requires no + /// apology: [D-19](../DESIGN-NOTES.md#d-19) makes a wake with nothing to + /// pop a normal event, so the caller re-checks the queue either way. An + /// implementation therefore cannot be subtly wrong about *when* to + /// return; it can only waste time or burn CPU. + /// + /// What it must not do is block past `timeout_ms`, because that is the + /// only thing standing between a stuck ring and a hung process. + /// + /// # An expired wait is `Ok(())`, never an error + /// + /// This is the one part of the contract an implementation can get wrong + /// silently, and the crate's own [`SubmitWait`] got it wrong first: Win32 + /// reports an expired wait as a *failure* code (`ERROR_TIMEOUT` from + /// `SubmitIoRing`, `WAIT_TIMEOUT` from the `WaitFor*` family), so + /// forwarding the underlying result verbatim turns every ordinary timeout + /// into an `Err`. [`IoRing::pop_within`] then reports a timeout as a + /// failure rather than as the `Ok(None)` it promises, and every caller + /// that matches on `Ok(None)` to detect a timeout stops working. + /// + /// So: the bound running out is a **successful** wait that happened to + /// observe nothing. Report the error only when the wait itself failed. + /// + /// `timeout_ms` is never zero and never `u32::MAX`, so it can be passed + /// straight to a Win32 wait without being mistaken for "poll and return" + /// or for `INFINITE`. + /// + /// # Errors + /// + /// Whatever the underlying wait reports as a genuine failure. An error + /// ends the pop. + fn wait(&mut self, ring: &mut RingWait<'_>, timeout_ms: u32) -> io::Result<()>; +} + +/// The ring, narrowed to what a [`CompletionWait`] needs (M21.2). +/// +/// Deliberately exposes no way to pop and no way to submit work. A waiter that +/// could pop would consume the completion its own caller is waiting for, and +/// one that could submit would queue entries the caller never asked for -- +/// both of which a bare `&mut IoRing` would permit. This is the same +/// narrowing, for the same reason, as +#[cfg_attr( + feature = "threadpool", + doc = "[`RingScope`](crate::RingScope) under [D-43](../DESIGN-NOTES.md#d-43)." +)] +#[cfg_attr( + not(feature = "threadpool"), + doc = "`RingScope` (the default `threadpool` feature) under [D-43](../DESIGN-NOTES.md#d-43)." +)] +pub struct RingWait<'ring> { + ring: &'ring mut IoRing, +} + +impl RingWait<'_> { + /// Block in the ring's own wait for up to `timeout_ms`, queueing nothing. + /// + /// With no new entries of its own, `SubmitIoRing`'s effect here is to + /// submit whatever is already queued and wait for an outstanding + /// operation to complete -- the same call [`IoRing::run_down`] uses to + /// quiesce. Note *submit*: a `Build*` that has queued an SQE but not yet + /// submitted it will be submitted by this call, which is why the + /// narrowing below is about not letting a waiter **build** work, not + /// about suppressing submission. + /// + /// **An expired wait is `Ok`, not an error.** Returning therefore does + /// not mean a completion is poppable -- the bound may simply have run out + /// -- which is why the loop that calls this re-checks either way. + /// + /// # At least one operation must really be outstanding + /// + /// Measured while building this: `SubmitIoRing` answers + /// `E_INVALIDARG` (`0x80070057`) -- not a timeout -- when asked to wait + /// for a completion the kernel has no pending operation for. A `RingWait` + /// is only ever constructed by [`IoRing::pop_within_with`], which checks + /// [`IoRing::outstanding`] before consulting the wait, so that + /// precondition holds structurally rather than by the caller remembering + /// it. A waiter that wants to block some other way is free to ignore this + /// method entirely. + /// + /// # Errors + /// + /// Any error from `SubmitIoRing` other than an expired wait. + pub fn block(&mut self, timeout_ms: u32) -> io::Result<()> { + let mut submitted = 0_u32; + // SAFETY: the ring handle is live for the borrow, and the out-pointer + // is valid. Zero new SQEs are queued, so this call's only effect is + // to wait for and reap what is already outstanding. + let hr = unsafe { crate::sys::submit(self.ring.handle, 1, timeout_ms, &raw mut submitted) }; + wait_outcome(hr) + } + + /// Operations submitted but not yet observed complete, as + /// [`IoRing::outstanding`]. + /// + /// A waiter that multiplexes several sources can use this to decide + /// whether blocking on this ring is worth a slot at all. + #[must_use] + pub fn outstanding(&self) -> usize { + self.ring.outstanding() + } +} + +/// The default [`CompletionWait`]: block inside the ring's own +/// `SubmitIoRing` wait (M21.2). +/// +/// Costs no kernel object and touches no event, so it composes with a caller +/// who owns the ring's completion event under +/// [D-21](../DESIGN-NOTES.md#d-21). +#[derive(Clone, Copy, Debug, Default, PartialEq, Eq)] +pub struct SubmitWait; + +impl CompletionWait for SubmitWait { + fn wait(&mut self, ring: &mut RingWait<'_>, timeout_ms: u32) -> io::Result<()> { + ring.block(timeout_ms) + } +} + +#[cfg(test)] +impl IoRing { + /// An `IoRing` that owns no kernel ring, for driving `Drop`'s two error + /// paths (M23.5). + /// + /// The handle is null **deliberately and specifically**. Measured: + /// `CloseIoRing(null)` and `crate::sys::submit(null, ..)` both return + /// `0x80070006` -- `HRESULT_FROM_WIN32(ERROR_INVALID_HANDLE)` -- so both + /// of `Drop`'s failure branches can be reached without a fault-injection + /// seam over the raw HRESULTs, which is what M23.5 was opened to price. + /// + /// A *non-null* fabricated handle is not a substitute and must never be + /// swapped in: `CloseIoRing` on a plausible-looking `0xDEAD_0000` raises + /// `STATUS_ACCESS_VIOLATION`, because a ring handle is a pointer the + /// kernel dereferences rather than an index into a handle table. + /// + /// Nothing here opens a ring, so the tests built on it are not part of the + /// ring-opening population `tools/check-ring-tests.ps1` tracks (D-49) -- + /// they need the kernel only to refuse them, which costs no ring. + fn refused_by_the_kernel() -> Self { + Self { + handle: std::ptr::null_mut(), + version: RingVersion::V1, + supported_ops: OpSupport::default(), + accounting: Accounting::new(), + registered_buffer_infos: Vec::new(), + completion_event: None, + } + } } impl Drop for IoRing { @@ -1003,14 +1546,42 @@ impl Drop for IoRing { // avoid it), but Drop cannot propagate the error, so this asserts in // debug builds rather than silently closing a ring the kernel may // still be writing through. + // + // Both asserts here are silent while already panicking (M23.4): a + // second panic during unwind aborts, replacing whatever failure + // started the unwind with `STATUS_STACK_BUFFER_OVERRUN`. A ring is + // dropped on the way out of almost every failing test in this crate, + // so an unguarded assert here would convert a readable assertion + // message into a crash in the common case rather than a rare one. + // + // Both are covered, and neither needed a fault-injection seam to get + // there (M23.5). `a_ring_whose_rundown_the_kernel_refuses_..` and + // `a_ring_whose_close_the_kernel_refuses_..` put a null handle in the + // field and let this body run for real: measured, `CloseIoRing(null)` + // and `crate::sys::submit(null, ..)` both return `0x80070006` + // (`ERROR_INVALID_HANDLE`). Whether `run_down` submits at all is what + // selects between the two, since it loops only while something is + // outstanding. + // + // A *non-null* bad handle is not equivalent and must never be + // substituted: `CloseIoRing(0xDEAD_0000)` raises + // `STATUS_ACCESS_VIOLATION`, because a ring handle is a pointer the + // kernel dereferences rather than an index into a handle table. if let Err(error) = self.run_down() { - debug_assert!(false, "IoRing rundown failed before close: {error}"); + debug_assert!( + std::thread::panicking(), + "IoRing rundown failed before close: {error}" + ); } // SAFETY: `self.handle` is a live ring this `IoRing` exclusively // owns, and `run_down` just established that nothing is outstanding // (or made a best-effort attempt to, above). let hr = unsafe { CloseIoRing(self.handle) }; - debug_assert!(hr >= 0, "CloseIoRing failed: 0x{:08X}", hr as u32); + debug_assert!( + hr >= 0 || std::thread::panicking(), + "CloseIoRing failed: 0x{:08X}", + hr as u32 + ); } } @@ -1049,22 +1620,20 @@ thread_local! { /// Thirty seconds matches the deadline the crate's own `failure_paths` /// integration test already uses; it is a hang bound, not a latency /// expectation, so it is far above any real completion time. +/// +/// Since M21.2 this is a thin panicking wrapper over the public +/// [`IoRing::pop_within`] rather than its own loop. The panic is the only +/// thing left that is specific to tests: a test wants the name of what it +/// waited for in the failure message, where a consumer wants an `Option` it +/// can act on. #[cfg(test)] pub(crate) fn pop_within(ring: &mut IoRing, what: &str) -> Completion { // Named once so the bound and the message it reports cannot drift apart. const BOUND: std::time::Duration = std::time::Duration::from_secs(30); - let deadline = std::time::Instant::now() + BOUND; - loop { - if let Some(completion) = ring.try_pop().expect("pop") { - return completion; - } - assert!( - std::time::Instant::now() < deadline, - "timed out after {BOUND:?} waiting for {what}" - ); - std::thread::yield_now(); - } + ring.pop_within(BOUND) + .expect("pop") + .unwrap_or_else(|| panic!("timed out after {BOUND:?} waiting for {what}")) } #[cfg(test)] diff --git a/crates/windows-ioring-sys/src/ring/tests.rs b/crates/windows-ioring-sys/src/ring/tests.rs index 06374f71c..3e649de69 100644 --- a/crates/windows-ioring-sys/src/ring/tests.rs +++ b/crates/windows-ioring-sys/src/ring/tests.rs @@ -9,21 +9,6 @@ fn op_support_starts_empty() { assert!(!OpSupport::default().contains(Op::Nop)); } -#[test] -fn a_ring_negotiates_a_version_no_higher_than_the_hosts_maximum() { - let ring = IoRing::new(64, 128).expect("create ring"); - let caps = capabilities().expect("capabilities"); - assert!(ring.version() <= caps.max_version); - assert!(ring.version() <= RingVersion::HIGHEST_KNOWN); -} - -#[test] -fn a_negotiated_ring_reports_its_version_back_through_get_ring_info() { - let ring = IoRing::new(64, 128).expect("create ring"); - let info = ring.info().expect("GetIoRingInfo"); - assert_eq!(info.version, ring.version()); -} - #[test] fn every_named_version_the_host_supports_creates_and_closes() { let caps = capabilities().expect("capabilities"); @@ -104,24 +89,6 @@ fn nop_read_and_write_are_supported_on_any_real_ring() { // --- outstanding-operation accounting and rundown (M2.4) --- -#[test] -fn reserve_user_data_increments_outstanding_and_never_repeats_an_id() { - let mut ring = IoRing::new(64, 128).expect("create ring"); - let a = ring.reserve_user_data().expect("reserve a"); - let b = ring.reserve_user_data().expect("reserve b"); - assert_ne!(a, b); - assert_eq!(ring.outstanding(), 2); - ring.record_completion(); - ring.record_completion(); -} - -#[test] -fn run_down_is_a_no_op_when_nothing_is_outstanding() { - let mut ring = IoRing::new(64, 128).expect("create ring"); - ring.run_down().expect("run_down with nothing outstanding"); - assert_eq!(ring.outstanding(), 0); -} - #[test] fn run_down_returns_once_a_recorded_completion_zeroes_the_count() { let mut ring = IoRing::new(64, 128).expect("create ring"); @@ -137,26 +104,6 @@ fn run_down_returns_once_a_recorded_completion_zeroes_the_count() { assert_eq!(ring.outstanding(), 0); } -#[test] -fn record_completion_saturates_rather_than_underflowing() { - let mut ring = IoRing::new(64, 128).expect("create ring"); - assert_eq!(ring.outstanding(), 0); - ring.record_completion(); - assert_eq!( - ring.outstanding(), - 0, - "recording more completions than were ever reserved must not wrap" - ); -} - -#[test] -fn dropping_a_ring_with_nothing_outstanding_does_not_hang() { - // The ordinary path: no tokens were ever minted, so Drop's run_down must - // return immediately rather than waiting on SubmitIoRing at all. - let ring = IoRing::new(64, 128).expect("create ring"); - drop(ring); -} - #[test] fn dropping_a_ring_actually_runs_its_drop_body() { // `::drop -> ()` survived: nothing distinguished a @@ -184,6 +131,65 @@ fn dropping_a_ring_actually_runs_its_drop_body() { ); } +/// `IoRing::drop`'s close assert, traversed for real (M23.5). +/// +/// M23.4 narrowed both asserts in `IoRing::drop` to fire only outside an +/// unwind, and measured that suppressing them entirely left every test in the +/// crate green -- so the guards were unverified in the direction that matters. +/// The item that followed assumed reaching them required a fault-injection +/// seam over every raw HRESULT the ring's Win32 calls return, and priced that +/// seam's blast radius. It is not required: the kernel already refuses a null +/// ring handle, and `ring::tests` is a child of `ring`, so it can build an +/// `IoRing` around one. See [`IoRing::refused_by_the_kernel`] for why null +/// specifically, and why a non-null stand-in would crash instead. +#[test] +#[should_panic(expected = "CloseIoRing failed")] +fn a_ring_whose_close_the_kernel_refuses_reports_the_close_failure() { + let ring = IoRing::refused_by_the_kernel(); + + // Nothing is outstanding, so `run_down` returns `Ok` without ever + // submitting. That is what makes the close assert the only one this test + // can reach -- see the sibling test for the other one. + assert_eq!( + ring.accounting.outstanding(), + 0, + "a fresh ledger has nothing outstanding" + ); + drop(ring); +} + +/// `IoRing::drop`'s rundown assert, traversed for real (M23.5). +/// +/// The same construction as the sibling test above, plus the one thing that +/// selects the other assert: `run_down` submits only *while* something is +/// outstanding, so a ring with an empty ledger never calls `SubmitIoRing` and +/// never fails. Reserving one identity makes the loop run once, that submit +/// is refused, and `run_down` returns the error the assert names. +/// +/// The two `expected` strings are what keep these tests honest about which +/// assert they reached: if this one fell through to the close instead, its +/// panic would say `CloseIoRing failed` and the test would go red rather than +/// pass for the wrong reason. That is sabotaged in [sabotage.json], not merely +/// asserted here. +/// +/// [sabotage.json]: ../../sabotage.json +#[test] +#[should_panic(expected = "IoRing rundown failed before close")] +fn a_ring_whose_rundown_the_kernel_refuses_reports_the_rundown_failure() { + let mut ring = IoRing::refused_by_the_kernel(); + + ring.accounting + .reserve_user_data() + .expect("a fresh ring's identity space is not exhausted"); + assert_eq!( + ring.accounting.outstanding(), + 1, + "rundown must have a reason to submit, or it cannot fail" + ); + + drop(ring); +} + // --- The fault-injection seam (M16.3) --- /// A real completion for a real, finished operation. @@ -544,3 +550,336 @@ fn the_debug_rendering_names_the_ring_and_its_key_fields() { "the version field name must appear: {rendering}" ); } + +// --------------------------------------------------------------------------- +// M21.2: the bounded pop, and the wait it is generic over. +// --------------------------------------------------------------------------- + +use super::{CompletionWait, RingWait, SubmitWait}; + +/// A scratch file to flush against, named per test so tests running as threads +/// in one process cannot collide on it. +fn pop_scratch(tag: &str) -> (std::path::PathBuf, std::fs::File) { + let path = std::env::temp_dir().join(format!( + "windows-ioring-sys-pop-within-{}-{tag}.tmp", + std::process::id() + )); + std::fs::write(&path, b"x").expect("create fixture"); + let file = std::fs::OpenOptions::new() + .read(true) + .write(true) + .open(&path) + .expect("open fixture"); + (path, file) +} + +/// Push `count` flushes and submit them, popping nothing. +fn push_flushes(ring: &mut IoRing, file: &std::fs::File, count: usize) -> Vec { + use crate::{Batch, FlushCoverage, FlushMode}; + use std::os::windows::io::AsRawHandle; + + let mut batch = Batch::new(ring); + let mut ids = Vec::with_capacity(count); + for _ in 0..count { + // SAFETY: `file` outlives every operation pushed here -- the caller + // drains before dropping it. + ids.push( + unsafe { + batch.flush_raw( + file.as_raw_handle(), + FlushCoverage::Unordered, + FlushMode::Default, + ) + } + .expect("queue a flush"), + ); + } + batch.submit().expect("submit"); + ids +} + +/// Records what the pop loop asked of it, and sleeps instead of blocking in +/// the ring. +/// +/// Deliberately does **not** call [`RingWait::block`]. These tests drive the +/// loop with a reservation that has no real SQE behind it, and `SubmitIoRing` +/// answers `E_INVALIDARG` when asked to wait for a completion the kernel has +/// no pending operation for. Keeping the wait out of the kernel is what makes +/// the loop's own deadline behaviour testable without a slow real device. +#[derive(Default)] +struct RecordingWait { + calls: usize, + last_timeout_ms: u32, + outstanding_seen: usize, +} + +impl CompletionWait for RecordingWait { + fn wait(&mut self, ring: &mut RingWait<'_>, timeout_ms: u32) -> std::io::Result<()> { + self.calls += 1; + self.last_timeout_ms = timeout_ms; + self.outstanding_seen = ring.outstanding(); + std::thread::sleep(std::time::Duration::from_millis(2)); + Ok(()) + } +} + +/// Refuses to wait at all, so a test can observe the error path. +struct FailingWait; + +impl CompletionWait for FailingWait { + fn wait(&mut self, _: &mut RingWait<'_>, _: u32) -> std::io::Result<()> { + Err(std::io::Error::other("the wait refused")) + } +} + +/// Returns without blocking, which the trait explicitly permits. +struct ImmediateWait; + +impl CompletionWait for ImmediateWait { + fn wait(&mut self, _: &mut RingWait<'_>, _: u32) -> std::io::Result<()> { + Ok(()) + } +} + +#[test] +fn pop_within_returns_the_completion_of_a_real_operation() { + let mut ring = IoRing::new(16, 16).expect("create ring"); + let (path, file) = pop_scratch("real"); + let ids = push_flushes(&mut ring, &file, 1); + + let completion = ring + .pop_within(std::time::Duration::from_secs(30)) + .expect("pop_within") + .expect("the flush completes well inside the bound"); + assert_eq!( + completion.user_data(), + ids[0], + "the completion popped must be the one that was pushed" + ); + let _ = std::fs::remove_file(&path); +} + +#[test] +fn submit_wait_is_what_the_convenience_uses() { + // The same operation through the explicit spelling. This is what proves + // `SubmitWait` really does block in the ring: no other wait is involved, + // and the completion still arrives. + let mut ring = IoRing::new(16, 16).expect("create ring"); + let (path, file) = pop_scratch("submit-wait"); + let ids = push_flushes(&mut ring, &file, 1); + + let completion = ring + .pop_within_with(&mut SubmitWait, std::time::Duration::from_secs(30)) + .expect("pop_within_with") + .expect("the flush completes"); + assert_eq!(completion.user_data(), ids[0]); + let _ = std::fs::remove_file(&path); +} + +#[test] +fn pop_within_returns_successive_completions_one_at_a_time() { + let mut ring = IoRing::new(16, 32).expect("create ring"); + let (path, file) = pop_scratch("successive"); + let ids = push_flushes(&mut ring, &file, 3); + + let mut seen = Vec::new(); + for _ in 0..ids.len() { + let completion = ring + .pop_within(std::time::Duration::from_secs(30)) + .expect("pop_within") + .expect("each flush completes"); + seen.push(completion.user_data()); + } + seen.sort_unstable(); + let mut expected = ids.clone(); + expected.sort_unstable(); + assert_eq!( + seen, expected, + "every pushed flush must be popped exactly once" + ); + + assert!( + ring.pop_within(std::time::Duration::ZERO) + .expect("pop_within") + .is_none(), + "the queue is empty once every completion has been taken" + ); + let _ = std::fs::remove_file(&path); +} + +#[test] +fn a_supplied_wait_is_not_consulted_when_nothing_can_arrive() { + // The partner to `a_supplied_wait_is_consulted_...` below. A loop that + // always waited once before checking would pass that test and fail this + // one. + let mut ring = IoRing::new(16, 16).expect("create ring"); + let mut wait = RecordingWait::default(); + let popped = ring + .pop_within_with(&mut wait, std::time::Duration::from_secs(30)) + .expect("pop_within_with"); + assert!(popped.is_none()); + assert_eq!( + wait.calls, 0, + "nothing can arrive, so there is nothing to wait for" + ); +} + +#[test] +fn a_supplied_wait_is_consulted_when_the_queue_is_not_ready() { + let mut ring = IoRing::new(16, 16).expect("create ring"); + // A reservation with no real SQE behind it: outstanding, and no + // completion will ever arrive for it. + ring.reserve_user_data().expect("reserve"); + let mut wait = RecordingWait::default(); + let popped = ring.pop_within_with(&mut wait, std::time::Duration::from_millis(40)); + // Settled before any assertion: a panic here would otherwise unwind into + // `Drop`, whose rundown cannot settle a reservation the kernel never saw, + // and the second panic would abort the whole harness. + ring.record_completion(); + + assert!(popped.expect("pop_within_with").is_none()); + assert!( + wait.calls >= 1, + "the supplied wait must be the thing that blocks" + ); +} + +#[test] +fn the_deadline_is_honoured_when_an_operation_never_completes() { + // No clock is consulted. That the loop *waited* rather than + // short-circuiting is proved by the wait having been called; that it + // *stopped* is proved by this test returning at all. A busy machine + // changes how long that takes and changes neither assertion. + let mut ring = IoRing::new(16, 16).expect("create ring"); + ring.reserve_user_data().expect("reserve"); + + let mut wait = RecordingWait::default(); + let popped = ring.pop_within_with(&mut wait, std::time::Duration::from_millis(40)); + ring.record_completion(); + + assert!( + popped.expect("pop_within_with").is_none(), + "no completion was ever going to arrive" + ); + assert!( + wait.calls >= 1, + "the bound must be waited out, not short-circuited" + ); +} + +#[test] +fn a_zero_bound_does_not_block() { + let mut ring = IoRing::new(16, 16).expect("create ring"); + ring.reserve_user_data().expect("reserve"); + let mut wait = RecordingWait::default(); + let popped = ring.pop_within_with(&mut wait, std::time::Duration::ZERO); + ring.record_completion(); + + assert!(popped.expect("pop_within_with").is_none()); + // The causal statement of "did not block": the wait is what blocks, and it + // was never reached. + assert_eq!( + wait.calls, 0, + "a zero bound is one try_pop, so the wait is never reached" + ); +} + +#[test] +fn the_wait_is_never_handed_a_zero_timeout() { + // A sub-millisecond remainder rounds to zero milliseconds, which + // `SubmitIoRing` reads as "poll and return" -- turning the tail of every + // bound into a spin. The loop clamps it up to one. + let mut ring = IoRing::new(16, 16).expect("create ring"); + ring.reserve_user_data().expect("reserve"); + let mut wait = RecordingWait::default(); + let popped = ring.pop_within_with(&mut wait, std::time::Duration::from_micros(600)); + ring.record_completion(); + + assert!(popped.expect("pop_within_with").is_none()); + assert!(wait.calls >= 1, "the wait must have been reached at all"); + assert!( + wait.last_timeout_ms >= 1, + "a sub-millisecond remainder must be clamped up, never passed as zero" + ); +} + +#[test] +fn ring_wait_reports_the_rings_outstanding_count() { + let mut ring = IoRing::new(16, 16).expect("create ring"); + ring.reserve_user_data().expect("reserve a"); + ring.reserve_user_data().expect("reserve b"); + let mut wait = RecordingWait::default(); + let popped = ring.pop_within_with(&mut wait, std::time::Duration::from_millis(20)); + ring.record_completion(); + ring.record_completion(); + + assert!(popped.expect("pop_within_with").is_none()); + assert_eq!( + wait.outstanding_seen, 2, + "RingWait::outstanding must agree with IoRing::outstanding" + ); +} + +#[test] +fn a_wait_that_fails_ends_the_pop_with_its_error() { + let mut ring = IoRing::new(16, 16).expect("create ring"); + ring.reserve_user_data().expect("reserve"); + let outcome = ring.pop_within_with(&mut FailingWait, std::time::Duration::from_secs(30)); + ring.record_completion(); + + let error = outcome.expect_err("the wait's failure must reach the caller"); + assert_eq!(error.to_string(), "the wait refused"); +} + +#[test] +fn a_wait_that_never_blocks_is_permitted_and_still_terminates() { + // The trait says returning early is always allowed. A caller that does so + // spins, which is their choice -- but the bound must still hold. + // + // **Termination is the assertion, and there is deliberately no clock.** An + // earlier version asserted the elapsed time was under five seconds, which + // could never have fired on the failure it named: if the deadline were not + // honoured the loop would spin forever and that line would never be + // reached. The only runs it could fail were slow ones -- so it was capable + // of false failures and incapable of true ones. A loop that does not + // terminate hangs the harness, which is what a hang looks like in every + // other test here too. + let mut ring = IoRing::new(16, 16).expect("create ring"); + ring.reserve_user_data().expect("reserve"); + let popped = ring.pop_within_with(&mut ImmediateWait, std::time::Duration::from_millis(30)); + ring.record_completion(); + + assert!( + popped.expect("pop_within_with").is_none(), + "the deadline is what ends a wait that never blocks" + ); +} + +#[test] +fn a_bound_the_clock_cannot_represent_reaches_the_wait_rather_than_panicking() { + // The half that exercises the overflow branch: something *is* outstanding, + // so the loop computes a remaining duration from a deadline that could not + // be represented. A failing wait is how the test escapes a bound that by + // construction never arrives. + let mut ring = IoRing::new(16, 16).expect("create ring"); + ring.reserve_user_data().expect("reserve"); + let outcome = ring.pop_within_with(&mut FailingWait, std::time::Duration::MAX); + ring.record_completion(); + + let error = outcome.expect_err("the wait refuses, which is how this returns at all"); + assert_eq!(error.to_string(), "the wait refused"); +} + +#[test] +fn the_wait_can_be_supplied_as_a_trait_object() { + // `?Sized` on the bound is what makes this compile, and a consumer + // choosing a wait at run time is the reason to keep it. + let mut ring = IoRing::new(16, 16).expect("create ring"); + ring.reserve_user_data().expect("reserve"); + let wait: &mut dyn CompletionWait = &mut FailingWait; + let outcome = ring.pop_within_with(wait, std::time::Duration::from_secs(30)); + ring.record_completion(); + + let error = outcome.expect_err("a trait-object wait still refuses"); + assert_eq!(error.to_string(), "the wait refused"); +} diff --git a/crates/windows-ioring-sys/src/sys.rs b/crates/windows-ioring-sys/src/sys.rs new file mode 100644 index 000000000..e4a3900f7 --- /dev/null +++ b/crates/windows-ioring-sys/src/sys.rs @@ -0,0 +1,268 @@ +// Copyright (c) Mike Grier +//! The seam the kernel-response resolver sits under (M26.2). +//! +//! Every `IoRing` FFI call that carries an *operation* goes through a wrapper +//! here instead of calling `windows-sys` directly. With the `kernel-seam` +//! feature off -- which is the published configuration and the default -- each +//! wrapper is an `#[inline(always)]` forward to the same call the crate made +//! before, and nothing in this module exists at all. With the feature on, a +//! test may install a `Responses` implementation that answers instead. +//! +//! # Why a module of wrappers rather than a generic `IoRing` +//! +//! A kernel trait with `IoRing` generic over it is the shape this would +//! normally take, and it is ruled out by a collision rather than by taste: +//! [D-55](../DESIGN-NOTES.md#d-55) already spends `IoRing`'s type parameter on +//! the pending-token inventory (`M28.3`). Taking a second one would publish +//! `IoRing` -- two parameters on a type whose users write `IoRing` +//! today, compounding a break `M28` accepts deliberately with one nobody asked +//! for. Module indirection costs the published API nothing: the signature of +//! every public item is unchanged, and so is the generic slot `M28` is going +//! to need. +//! +//! # What is behind the seam, and what deliberately is not +//! +//! Behind it: `SubmitIoRing`, `PopIoRingCompletion`, and the `Build*` family +//! -- the calls that submit work and report its outcome, which is the surface +//! `M26.1`'s [RESPONSE-SPACE.md](../RESPONSE-SPACE.md) describes. +//! +//! Not behind it: `CreateIoRing`, `CloseIoRing`, `GetIoRingInfo`, +//! `IsIoRingOpSupported`. Those decide whether a ring *exists* and what it +//! supports, not how it responds, and a resolver that replaced them would be +//! a fake ring rather than a resolver over responses. **`M26` does not +//! justify itself on hermeticity** -- the milestone says so in as many words, +//! because `M24` reached a hermetic lib suite without it -- so a resolver runs +//! against a real ring whose operations it answers for. Leaving lifecycle real +//! is what keeps this a seam under the responses rather than a mock of the +//! ring. +//! +//! `SetIoRingCompletionEvent` was moved behind the seam by `M26.3`, and the +//! reason is worth stating because it looks like lifecycle. It is how a +//! completion becomes *observable* to a waiter, so +//! [RS-P-6](../RESPONSE-SPACE.md) -- a completion posted behind another need +//! produce no signal -- is a clause about this call. A resolver that could not +//! make it does not satisfy RS-P-6 vacuously; it never signals at all, which +//! hangs every `EventDelivery` consumer rather than testing one. +//! +//! # The trait mirrors the FFI exactly, on purpose +//! +//! Every method takes and returns what the Win32 call takes and returns, with +//! no interpretation. A seam that translated into friendlier types would be +//! encoding a belief about what those calls mean, which is the objection +//! [D-52](../DESIGN-NOTES.md#d-52) records against mocks and the reason this +//! milestone exists. Translating is the *resolver's* job, and it does it +//! against a written specification. + +use std::ffi::c_void; + +use windows_sys::Win32::Storage::FileSystem::{ + BuildIoRingCancelRequest, BuildIoRingFlushFile, BuildIoRingReadFile, + BuildIoRingRegisterBuffers, BuildIoRingRegisterFileHandles, BuildIoRingWriteFile, + FILE_FLUSH_MODE, IORING_BUFFER_INFO, IORING_BUFFER_REF, IORING_CQE, IORING_HANDLE_REF, + PopIoRingCompletion, SetIoRingCompletionEvent, SubmitIoRing, +}; +use windows_sys::core::HRESULT; + +#[cfg(feature = "kernel-seam")] +mod installed; +#[cfg(feature = "kernel-seam")] +mod resolver; +#[cfg(feature = "kernel-seam")] +pub use installed::{Installed, Responses, install}; +#[cfg(feature = "kernel-seam")] +pub use resolver::{ + Resolver, ResolverConfig, ResolverStats, ResolverWatch, SEED_VAR as RESOLVER_SEED_VAR, +}; + +/// Dispatch to an installed `Responses`, or fall through to the real call. +/// +/// The feature-off arm expands to the real call and nothing else, so a +/// published build has no branch, no thread-local access, and no trait object +/// -- the wrapper is the call. +macro_rules! through_seam { + ($method:ident ( $($arg:expr),* $(,)? ) else $real:expr) => {{ + #[cfg(feature = "kernel-seam")] + { + if let Some(answer) = $crate::sys::installed::with(|r| unsafe { r.$method($($arg),*) }) + { + return answer; + } + } + unsafe { $real } + }}; +} + +/// `SubmitIoRing`. +/// +/// # Safety +/// +/// `ring` must be a live ring and `submitted` a valid out-pointer, exactly as +/// the Win32 call requires. Forwarded unchanged. +#[inline(always)] +pub(crate) unsafe fn submit( + ring: *mut c_void, + wait_operations: u32, + milliseconds: u32, + submitted: *mut u32, +) -> HRESULT { + through_seam!( + submit(ring, wait_operations, milliseconds, submitted) + else SubmitIoRing(ring, wait_operations, milliseconds, submitted) + ) +} + +/// `PopIoRingCompletion`. +/// +/// # Safety +/// +/// `ring` must be a live ring and `cqe` a valid out-pointer. +#[inline(always)] +pub(crate) unsafe fn pop(ring: *mut c_void, cqe: *mut IORING_CQE) -> HRESULT { + through_seam!(pop(ring, cqe) else PopIoRingCompletion(ring, cqe)) +} + +/// `BuildIoRingReadFile`. +/// +/// # Safety +/// +/// As the Win32 call: a live ring, a valid handle reference, and a buffer +/// reference that stays valid until the operation completes. +#[inline(always)] +#[allow(clippy::too_many_arguments)] +pub(crate) unsafe fn build_read( + ring: *mut c_void, + file: IORING_HANDLE_REF, + buffer: IORING_BUFFER_REF, + bytes: u32, + offset: u64, + user_data: usize, + flags: i32, +) -> HRESULT { + through_seam!( + build_read(ring, file, buffer, bytes, offset, user_data, flags) + else BuildIoRingReadFile(ring, file, buffer, bytes, offset, user_data, flags) + ) +} + +/// `BuildIoRingWriteFile`. +/// +/// # Safety +/// +/// As [`build_read`]. +#[inline(always)] +#[allow(clippy::too_many_arguments)] +pub(crate) unsafe fn build_write( + ring: *mut c_void, + file: IORING_HANDLE_REF, + buffer: IORING_BUFFER_REF, + bytes: u32, + offset: u64, + caching: i32, + user_data: usize, + flags: i32, +) -> HRESULT { + through_seam!( + build_write(ring, file, buffer, bytes, offset, caching, user_data, flags) + else BuildIoRingWriteFile(ring, file, buffer, bytes, offset, caching, user_data, flags) + ) +} + +/// `BuildIoRingFlushFile`. +/// +/// # Safety +/// +/// `ring` must be live and `file` a valid handle reference. +#[inline(always)] +pub(crate) unsafe fn build_flush( + ring: *mut c_void, + file: IORING_HANDLE_REF, + mode: FILE_FLUSH_MODE, + user_data: usize, + flags: i32, +) -> HRESULT { + through_seam!( + build_flush(ring, file, mode, user_data, flags) + else BuildIoRingFlushFile(ring, file, mode, user_data, flags) + ) +} + +/// `BuildIoRingCancelRequest`. +/// +/// # Safety +/// +/// `ring` must be live and `file` a valid handle reference. +#[inline(always)] +pub(crate) unsafe fn build_cancel( + ring: *mut c_void, + file: IORING_HANDLE_REF, + cancel_user_data: usize, + user_data: usize, +) -> HRESULT { + through_seam!( + build_cancel(ring, file, cancel_user_data, user_data) + else BuildIoRingCancelRequest(ring, file, cancel_user_data, user_data) + ) +} + +/// `BuildIoRingRegisterFileHandles`. +/// +/// # Safety +/// +/// `handles` must point to `count` valid handles that outlive the operation. +#[inline(always)] +pub(crate) unsafe fn build_register_files( + ring: *mut c_void, + count: u32, + handles: *const *mut c_void, + user_data: usize, +) -> HRESULT { + through_seam!( + build_register_files(ring, count, handles, user_data) + else BuildIoRingRegisterFileHandles(ring, count, handles, user_data) + ) +} + +/// `BuildIoRingRegisterBuffers`. +/// +/// # Safety +/// +/// `buffers` must point to `count` `IORING_BUFFER_INFO` entries that stay +/// valid until the registration operation *runs* -- not merely until this +/// call returns; see [D-32](../DESIGN-NOTES.md#d-32), which measured the +/// difference. +#[inline(always)] +pub(crate) unsafe fn build_register_buffers( + ring: *mut c_void, + count: u32, + buffers: *const IORING_BUFFER_INFO, + user_data: usize, +) -> HRESULT { + through_seam!( + build_register_buffers(ring, count, buffers, user_data) + else BuildIoRingRegisterBuffers(ring, count, buffers, user_data) + ) +} + +/// `SetIoRingCompletionEvent`. +/// +/// # Safety +/// +/// `ring` must be a live ring and `event` a live event handle the ring will +/// own for the rest of its life. +#[inline(always)] +pub(crate) unsafe fn set_completion_event(ring: *mut c_void, event: *mut c_void) -> HRESULT { + through_seam!( + set_completion_event(ring, event) else SetIoRingCompletionEvent(ring, event) + ) +} + +/// Re-exported for [`Responses`]' default methods, which forward to the real +/// calls so an implementation overrides only what it varies. +#[cfg(feature = "kernel-seam")] +pub(crate) mod real { + pub(crate) use windows_sys::Win32::Storage::FileSystem::{ + BuildIoRingCancelRequest, BuildIoRingFlushFile, BuildIoRingReadFile, + BuildIoRingRegisterBuffers, BuildIoRingRegisterFileHandles, BuildIoRingWriteFile, + PopIoRingCompletion, SetIoRingCompletionEvent, SubmitIoRing, + }; +} diff --git a/crates/windows-ioring-sys/src/sys/installed.rs b/crates/windows-ioring-sys/src/sys/installed.rs new file mode 100644 index 000000000..50e45e815 --- /dev/null +++ b/crates/windows-ioring-sys/src/sys/installed.rs @@ -0,0 +1,271 @@ +// Copyright (c) Mike Grier +//! Installing a [`Responses`] under the seam (M26.2). +//! +//! # Thread-local, and that is not a detail +//! +//! `cargo test` runs tests as **threads in one process**, which this +//! repository chose deliberately and documents in its design notes. A +//! process-global responder would therefore let one test answer another +//! test's kernel calls, and the failure would look like a flaky ring rather +//! than like a harness defect. The crate already reached this conclusion once, +//! for `IoRing`'s drop counter, after a process-wide static let another test's +//! drop satisfy an assertion and mask the mutation the test existed to catch. +//! +//! So the responder is installed per thread, and a ring used from a thread +//! with nothing installed talks to the real kernel -- including a ring moved +//! across threads, which is legal since `IoRing` is `Send`. A resolver that +//! needs to follow a ring across threads has to arrange it; nothing here does +//! it implicitly. + +use std::cell::RefCell; +use std::ffi::c_void; + +use windows_sys::Win32::Storage::FileSystem::{ + FILE_FLUSH_MODE, IORING_BUFFER_INFO, IORING_BUFFER_REF, IORING_CQE, IORING_HANDLE_REF, +}; +use windows_sys::core::HRESULT; + +use super::real; + +/// Answers the `IoRing` calls that carry an operation. +/// +/// Every method mirrors its Win32 call exactly and **defaults to making it**, +/// so an implementation overrides only the calls whose responses it varies. +/// That default is what keeps a partial resolver honest: a method left +/// unimplemented behaves like the kernel rather than like a stub returning +/// success. +/// +/// # Safety +/// +/// Every method is `unsafe` because every method may be forwarded to the Win32 +/// call with the caller's pointers. An implementation that does not forward +/// still receives raw pointers and must not dereference them beyond what the +/// corresponding call would. +pub trait Responses { + /// `SubmitIoRing`. + /// + /// # Safety + /// + /// As `SubmitIoRing`. + unsafe fn submit( + &mut self, + ring: *mut c_void, + wait_operations: u32, + milliseconds: u32, + submitted: *mut u32, + ) -> HRESULT { + // SAFETY: forwarded unchanged from the caller, who holds the Win32 + // call's own contract. + unsafe { real::SubmitIoRing(ring, wait_operations, milliseconds, submitted) } + } + + /// `PopIoRingCompletion`. + /// + /// # Safety + /// + /// As `PopIoRingCompletion`. + unsafe fn pop(&mut self, ring: *mut c_void, cqe: *mut IORING_CQE) -> HRESULT { + // SAFETY: as `submit`. + unsafe { real::PopIoRingCompletion(ring, cqe) } + } + + /// `BuildIoRingReadFile`. + /// + /// # Safety + /// + /// As `BuildIoRingReadFile`. + #[allow(clippy::too_many_arguments)] + unsafe fn build_read( + &mut self, + ring: *mut c_void, + file: IORING_HANDLE_REF, + buffer: IORING_BUFFER_REF, + bytes: u32, + offset: u64, + user_data: usize, + flags: i32, + ) -> HRESULT { + // SAFETY: as `submit`. + unsafe { real::BuildIoRingReadFile(ring, file, buffer, bytes, offset, user_data, flags) } + } + + /// `BuildIoRingWriteFile`. + /// + /// # Safety + /// + /// As `BuildIoRingWriteFile`. + #[allow(clippy::too_many_arguments)] + unsafe fn build_write( + &mut self, + ring: *mut c_void, + file: IORING_HANDLE_REF, + buffer: IORING_BUFFER_REF, + bytes: u32, + offset: u64, + caching: i32, + user_data: usize, + flags: i32, + ) -> HRESULT { + // SAFETY: as `submit`. + unsafe { + real::BuildIoRingWriteFile(ring, file, buffer, bytes, offset, caching, user_data, flags) + } + } + + /// `BuildIoRingFlushFile`. + /// + /// # Safety + /// + /// As `BuildIoRingFlushFile`. + unsafe fn build_flush( + &mut self, + ring: *mut c_void, + file: IORING_HANDLE_REF, + mode: FILE_FLUSH_MODE, + user_data: usize, + flags: i32, + ) -> HRESULT { + // SAFETY: as `submit`. + unsafe { real::BuildIoRingFlushFile(ring, file, mode, user_data, flags) } + } + + /// `BuildIoRingCancelRequest`. + /// + /// # Safety + /// + /// As `BuildIoRingCancelRequest`. + unsafe fn build_cancel( + &mut self, + ring: *mut c_void, + file: IORING_HANDLE_REF, + cancel_user_data: usize, + user_data: usize, + ) -> HRESULT { + // SAFETY: as `submit`. + unsafe { real::BuildIoRingCancelRequest(ring, file, cancel_user_data, user_data) } + } + + /// `BuildIoRingRegisterFileHandles`. + /// + /// # Safety + /// + /// As `BuildIoRingRegisterFileHandles`. + unsafe fn build_register_files( + &mut self, + ring: *mut c_void, + count: u32, + handles: *const *mut c_void, + user_data: usize, + ) -> HRESULT { + // SAFETY: as `submit`. + unsafe { real::BuildIoRingRegisterFileHandles(ring, count, handles, user_data) } + } + + /// `BuildIoRingRegisterBuffers`. + /// + /// # Safety + /// + /// As `BuildIoRingRegisterBuffers`, including that `buffers` stays valid + /// until the registration operation *runs* rather than merely until the + /// call returns -- a difference this crate measured (`D-32`). + unsafe fn build_register_buffers( + &mut self, + ring: *mut c_void, + count: u32, + buffers: *const IORING_BUFFER_INFO, + user_data: usize, + ) -> HRESULT { + // SAFETY: as `submit`. + unsafe { real::BuildIoRingRegisterBuffers(ring, count, buffers, user_data) } + } + + /// `SetIoRingCompletionEvent`. + /// + /// Behind the seam because it is how a completion becomes *observable*, + /// which is what [RS-P-6](../RESPONSE-SPACE.md) is a clause about -- not + /// because it is lifecycle. An implementation that answers this call + /// takes on the obligation to signal, since a ring whose completions are + /// answered here and whose event is never set leaves every waiter parked. + /// + /// # Safety + /// + /// As `SetIoRingCompletionEvent`. An implementation that keeps `event` + /// must not outlive the ring that owns it. + unsafe fn set_completion_event(&mut self, ring: *mut c_void, event: *mut c_void) -> HRESULT { + // SAFETY: as `submit`. + unsafe { real::SetIoRingCompletionEvent(ring, event) } + } +} + +thread_local! { + /// The responder for this thread, if one is installed. + /// + /// `RefCell` rather than `Cell`: the seam hands out `&mut` so a responder + /// can keep state across calls, which any resolver over an ordering space + /// must. The borrow is held only for the duration of one FFI call. + static CURRENT: RefCell>> = const { RefCell::new(None) }; +} + +/// Run `f` against this thread's responder, if it has one. +/// +/// `None` means nothing is installed, which is the seam's signal to make the +/// real call. +/// +/// **A responder is not consulted re-entrantly.** If `f` reaches the seam +/// again -- a responder that forwards to the real call cannot, but one driving +/// a nested ring could -- the inner call finds the slot already borrowed and +/// falls through to the kernel rather than panicking on the `RefCell` or +/// aliasing the `&mut`. +pub(crate) fn with(f: impl FnOnce(&mut dyn Responses) -> R) -> Option { + CURRENT + .try_with(|slot| { + let Ok(mut borrow) = slot.try_borrow_mut() else { + return None; + }; + borrow.as_mut().map(|responder| f(&mut **responder)) + }) + .ok() + .flatten() +} + +/// Installs `responder` for this thread until the returned guard drops. +/// +/// # Panics +/// +/// If this thread already has one installed. Nesting would make which +/// responder answered a given call depend on drop order, and a suite that +/// nested by accident would be very hard to read. +#[must_use = "the responder is uninstalled when the guard drops"] +pub fn install(responder: Box) -> Installed { + CURRENT.with(|slot| { + let mut borrow = slot.borrow_mut(); + assert!( + borrow.is_none(), + "a kernel responder is already installed on this thread" + ); + *borrow = Some(responder); + }); + Installed { _private: () } +} + +/// Uninstalls this thread's responder on drop. +#[must_use = "dropping this immediately uninstalls the responder"] +pub struct Installed { + _private: (), +} + +impl Drop for Installed { + fn drop(&mut self) { + // `try_with` rather than `with`: on a thread being torn down the + // thread-local may already be destroyed, and a panic in `Drop` during + // that unwind would abort the process (M23.4). + let _ = CURRENT.try_with(|slot| { + if let Ok(mut borrow) = slot.try_borrow_mut() { + *borrow = None; + } + }); + } +} + +#[cfg(test)] +mod tests; diff --git a/crates/windows-ioring-sys/src/sys/installed/tests.rs b/crates/windows-ioring-sys/src/sys/installed/tests.rs new file mode 100644 index 000000000..44637e39c --- /dev/null +++ b/crates/windows-ioring-sys/src/sys/installed/tests.rs @@ -0,0 +1,185 @@ +// Copyright (c) Mike Grier +//! Tests for the seam's install point (M26.2). +//! +//! # What these establish +//! +//! That the indirection is **transparent** with nothing installed, that an +//! installed responder is **actually consulted**, and that it is **scoped to +//! one thread and one guard**. Those are the three properties `M26.3`'s +//! resolver will rest on, and each of them can silently stop holding. +//! +//! # What they deliberately do not do +//! +//! **Open a ring.** Every test here either asks the install point directly or +//! passes a null ring handle to a responder that never forwards, so nothing +//! reaches Win32. That keeps them out of the ring-opening population +//! `tools/check-ring-tests.ps1` tracks (D-49), and it is sound because what is +//! under test is the *dispatch*, not the calls. The forwarding path is +//! exercised by every other test in the crate, all of which run with no +//! responder installed. + +use std::cell::Cell; +use std::ffi::c_void; +use std::rc::Rc; + +use windows_sys::Win32::Foundation::{E_FAIL, S_OK}; +use windows_sys::core::HRESULT; + +use super::{Responses, install}; + +/// Answers `submit` with a fixed code and counts its calls, forwarding +/// nothing. +struct Reporting { + answer: HRESULT, + seen: Rc>, +} + +impl Responses for Reporting { + unsafe fn submit( + &mut self, + _ring: *mut c_void, + _wait_operations: u32, + _milliseconds: u32, + _submitted: *mut u32, + ) -> HRESULT { + self.seen.set(self.seen.get() + 1); + self.answer + } +} + +/// Call the seam once with arguments no real ring would accept. +/// +/// # Safety +/// +/// Only sound while a responder that answers `submit` **without forwarding** +/// is installed. With nothing installed this would pass a null handle to +/// `SubmitIoRing`, which is why no test calls it in that state. +unsafe fn submit_through_seam() -> HRESULT { + let mut submitted = 0_u32; + // SAFETY: forwarded from this function's own contract -- the caller has + // installed a non-forwarding responder. + unsafe { crate::sys::submit(std::ptr::null_mut(), 0, 0, &raw mut submitted) } +} + +#[test] +fn an_installed_responder_answers_instead_of_the_kernel() { + let seen = Rc::new(Cell::new(0)); + let guard = install(Box::new(Reporting { + answer: E_FAIL, + seen: Rc::clone(&seen), + })); + + // SAFETY: the responder above answers without forwarding. + let answer = unsafe { submit_through_seam() }; + + assert_eq!( + answer, E_FAIL, + "the seam must return what the responder said, not what the kernel would" + ); + assert_eq!(seen.get(), 1, "the responder must have been consulted once"); + drop(guard); +} + +#[test] +fn the_responder_is_gone_once_its_guard_drops() { + let seen = Rc::new(Cell::new(0)); + { + let _guard = install(Box::new(Reporting { + answer: S_OK, + seen: Rc::clone(&seen), + })); + // SAFETY: as above. + let _ = unsafe { submit_through_seam() }; + } + assert_eq!(seen.get(), 1, "the responder answered while installed"); + + // Asked at the install point rather than through the seam: calling the + // seam with nothing installed would reach Win32 with a null handle. + assert!( + super::with(|_| ()).is_none(), + "the guard must uninstall on drop, or a later test inherits this one's responder" + ); +} + +#[test] +fn a_responder_is_scoped_to_the_thread_that_installed_it() { + let seen = Rc::new(Cell::new(0)); + let _guard = install(Box::new(Reporting { + answer: S_OK, + seen: Rc::clone(&seen), + })); + + // The property that makes this thread-local at all: `cargo test` runs + // tests as threads in one process, so a process-global responder would + // answer for tests that never asked for one. + let elsewhere = std::thread::spawn(|| super::with(|_| ()).is_some()) + .join() + .expect("the probe thread does not panic"); + + assert!( + !elsewhere, + "another thread must not see this thread's responder" + ); + assert_eq!(seen.get(), 0, "and must not have consulted it"); +} + +#[test] +#[should_panic(expected = "already installed")] +fn installing_twice_on_one_thread_is_refused() { + // Nesting would make which responder answered a call depend on drop + // order. Refusing is what keeps a suite that nested by accident readable. + let _first = install(Box::new(Reporting { + answer: S_OK, + seen: Rc::new(Cell::new(0)), + })); + let _second = install(Box::new(Reporting { + answer: S_OK, + seen: Rc::new(Cell::new(0)), + })); +} + +#[test] +fn nothing_is_installed_by_default() { + // The transparency property, asserted at the dispatch. With nothing + // installed the seam makes the real call, and that path is what every + // other test in this crate already exercises. + assert!( + super::with(|_| ()).is_none(), + "a thread with no responder must fall through to the kernel" + ); +} + +#[test] +fn a_responder_is_not_consulted_re_entrantly() { + /// Re-enters the seam from inside its own `submit`. + struct Reentrant { + depth: Rc>, + } + + impl Responses for Reentrant { + unsafe fn submit( + &mut self, + _ring: *mut c_void, + _wait: u32, + _ms: u32, + _submitted: *mut u32, + ) -> HRESULT { + self.depth.set(self.depth.get() + 1); + // The inner call must find the slot borrowed and report `None` + // rather than panicking on the `RefCell` or aliasing the `&mut`. + assert!( + super::with(|_| ()).is_none(), + "a nested call must not reach the responder again" + ); + S_OK + } + } + + let depth = Rc::new(Cell::new(0)); + let _guard = install(Box::new(Reentrant { + depth: Rc::clone(&depth), + })); + // SAFETY: the responder answers without forwarding. + let _ = unsafe { submit_through_seam() }; + assert_eq!(depth.get(), 1, "the outer call is consulted exactly once"); +} diff --git a/crates/windows-ioring-sys/src/sys/resolver.rs b/crates/windows-ioring-sys/src/sys/resolver.rs new file mode 100644 index 000000000..91a10632e --- /dev/null +++ b/crates/windows-ioring-sys/src/sys/resolver.rs @@ -0,0 +1,770 @@ +// Copyright (c) Mike Grier +//! A seeded resolver over the permitted kernel response space (M26.3). +//! +//! This is the thing [RESPONSE-SPACE.md](../RESPONSE-SPACE.md) was written +//! for. It implements [`Responses`] by answering the submission-path calls +//! itself, choosing a point in that space from a seed. +//! +//! # It asserts nothing about Windows +//! +//! The distinction [D-52](../DESIGN-NOTES.md#d-52) turns on, restated here +//! because it is what makes this defensible where a mock was refused twice: a +//! mock encodes a belief about what the platform *does*, and can be wrong +//! about it. This encodes a specification of what this crate will *tolerate*, +//! and the assertions made against it are about this crate. There is no belief +//! here to be wrong -- only a space that could be too narrow, which is a +//! reviewable defect rather than a hidden one. +//! +//! So every freedom below cites the clause that permits it, and every +//! restriction cites the clause that requires it. A clause no code cites is +//! visibly unimplemented; a behaviour citing no clause is a resolver inventing +//! a platform. +//! +//! # Permissions are configurable; constraints are not +//! +//! [`ResolverConfig`] has a switch per `RS-P-n` and none for any `RS-C-n`, and +//! that asymmetry is deliberate rather than incidental. A permission is a +//! freedom a test may wish to narrow in order to isolate another -- a test +//! about ordering does not want arbitrary operation failures on top. A +//! constraint is what makes this crate's code reviewable at all, so a knob +//! that relaxed one would let a test quietly assert against a platform that +//! cannot exist. Constraints are therefore only reachable by editing this +//! file, which is what the sabotage manifest does to show they are +//! load-bearing. +//! +//! # What drives resolution +//! +//! A real kernel completes pending work on its own. This has no background, so +//! its only clock is being consulted: +//! +//! - **A submit resolves.** Each still-unresolved operation independently +//! either completes now or pends (`RS-P-1`). +//! - **A pop only prevents starvation.** It advances every operation's +//! deferral count and posts anything that has run out, but it flips no +//! coins. Without that split a consumer polling [`crate::IoRing::try_pop`] +//! would see its first poll resolve the work, and "the operation pended" -- +//! the thing `RS-P-1` exists to let a test observe -- would be unobservable. +//! +//! The deferral bound is what makes `RS-C-1` finite. It is a property of this +//! resolver, not of the space: the space deliberately carries no rates, and a +//! bound on how long a resolver may defer is a mechanism for satisfying a +//! constraint rather than a claim about how often Windows pends. +//! +//! # Two ordering requirements, and what it costs to install one +//! +//! **Install before the ring exists, and drop the guard after the ring.** The +//! resolver answers for operations the kernel never saw. If the guard drops +//! first, the ring's own rundown asks a real ring to wait for completions it +//! has no pending operation for, which `SubmitIoRing` answers `E_INVALIDARG`. +//! [`Resolver::scoped`] arranges both by construction and is the shape to +//! reach for. +//! +//! **A ring under a resolver is answered per thread**, as +//! the seam's `installed` module explains. Moving such a ring to a thread with nothing +//! installed hands its operations straight to a kernel that never received +//! them. + +use std::cell::RefCell; +use std::collections::VecDeque; +use std::ffi::c_void; +use std::rc::Rc; + +use windows_sys::Win32::Storage::FileSystem::{ + FILE_FLUSH_MODE, IORING_BUFFER_INFO, IORING_BUFFER_REF, IORING_CQE, IORING_HANDLE_REF, + IOSQE_FLAGS_DRAIN_PRECEDING_OPS, +}; +use windows_sys::Win32::System::Threading::SetEvent; +use windows_sys::core::HRESULT; + +use super::installed::{Installed, Responses}; + +/// Environment override for the resolver seed (decimal, or `0x` hex). +/// +/// **A third axis, kept separate on purpose.** +/// [generated_sequences.rs](../../tests/generated_sequences.rs) already +/// carries two -- its own sequence seed and the guard allocator's -- and +/// [D-41](../DESIGN-NOTES.md#d-41)'s discipline is that one number replays one +/// thing. Folding the resolution choices into either of those would produce a +/// replay that reproduced part of a run and not the rest, which is worse than +/// no replay at all because it looks like one. +pub const SEED_VAR: &str = "WINDOWS_IORING_RESOLVER_SEED"; + +/// How many consultations an operation may pend across before this resolver +/// posts it regardless. +/// +/// This is how `RS-C-1` ("every submitted operation eventually completes") +/// becomes finite rather than merely eventual. Four is chosen so that a +/// pending operation survives a few polls -- enough for a test to observe that +/// it pended -- without a rundown loop taking long enough to look like a hang. +const MAX_DEFERRALS: u32 = 4; + +/// How many consecutive submits `RS-P-7` may fail before this resolver lets +/// one through. +/// +/// Same role as [`MAX_DEFERRALS`], for the other way an operation can be made +/// to wait: a resolver free to fail every submit forever would never complete +/// anything, which `RS-C-1` forbids. +const MAX_CONSECUTIVE_SUBMIT_FAILURES: u32 = 2; + +/// `S_OK`. +const S_OK: HRESULT = 0; +/// `S_FALSE`, which is how `PopIoRingCompletion` reports an empty queue. +const S_FALSE: HRESULT = 1; +/// `IORING_E_WAIT_TIMEOUT`, which `SubmitIoRing` documents as "all operations +/// were submitted without error and the subsequent wait timed out". +/// +/// Spelled by derivation for the reason `ring.rs` records at its own copy: the +/// name is a macro over `HRESULT_FROM_WIN32(ERROR_TIMEOUT)` rather than a +/// `FACILITY_IORING` code, so the bindings emit no constant to import. +const IORING_E_WAIT_TIMEOUT: HRESULT = + (0x8007_0000_u32 | windows_sys::Win32::Foundation::ERROR_TIMEOUT) as HRESULT; +/// `HRESULT_FROM_WIN32(ERROR_NOT_ENOUGH_MEMORY)`, the failure this resolver +/// reports for a submit it declines under `RS-P-7`. +/// +/// Deliberately **not** `IORING_E_WAIT_TIMEOUT`: that code carries a positive +/// guarantee that every entry was submitted, so using it here would make the +/// resolver claim the work went in while holding it back. +const SUBMIT_FAILED: HRESULT = + (0x8007_0000_u32 | windows_sys::Win32::Foundation::ERROR_NOT_ENOUGH_MEMORY) as HRESULT; + +/// Which permissions this resolver exercises. +/// +/// Every field names its clause. All default to `true` -- the widest point in +/// the space -- so that narrowing is a visible, deliberate act in a test +/// rather than something a resolver quietly failed to do. +#[derive(Debug, Clone, Copy, PartialEq, Eq)] +pub struct ResolverConfig { + /// `RS-P-1`: an operation may complete inside `SubmitIoRing`, or pend. + /// + /// With this off every operation completes inside the submit that carried + /// it, which is the narrowest reading of the clause and the one a test + /// isolating some *other* freedom usually wants. + pub may_pend: bool, + /// `RS-P-2`: completion order is unconstrained. + /// + /// With this off completions are posted in submission order. That is not + /// a promise Windows makes -- [D-47](../DESIGN-NOTES.md#d-47) measured it + /// broken -- so a test turning this off is narrowing to isolate, never + /// asserting FIFO. + pub may_reorder: bool, + /// `RS-P-3`: an operation may fail individually. + pub may_fail_operations: bool, + /// `RS-P-4`: a wait may expire, including when completions are available. + pub may_expire_waits: bool, + /// `RS-P-5`: a wait may return successfully with nothing poppable. + pub may_wake_empty: bool, + /// `RS-P-6`: a completion posted while the queue was already non-empty + /// need produce no signal. + /// + /// With this off every post signals. That is *wider* than the platform -- + /// [D-19](../DESIGN-NOTES.md#d-19) measured the event as edge triggered -- + /// so it exists to show the edge behaviour is load-bearing, not as a + /// configuration any consumer should rely on. + pub edge_triggered_signal: bool, + /// `RS-P-7`: a submit may fail, leaving built operations queued for a + /// later one. + pub may_fail_submits: bool, + /// `RS-P-8`: a successful transfer may report fewer bytes than requested. + /// + /// With this off every successful read or write reports the full + /// requested length. That is the behaviour of an ordinary file on a local + /// volume, and so the narrowing a test wants when its subject is + /// something else -- but it is *narrower* than the platform, because this + /// crate does not constrain what kind of handle a caller registers. + pub may_transfer_partially: bool, +} + +impl Default for ResolverConfig { + fn default() -> Self { + Self { + may_pend: true, + may_reorder: true, + may_fail_operations: true, + may_expire_waits: true, + may_wake_empty: true, + edge_triggered_signal: true, + may_fail_submits: true, + may_transfer_partially: true, + } + } +} + +impl ResolverConfig { + /// Every permission off: operations complete inside their submit, in + /// order, successfully, and no wait misbehaves. + /// + /// The narrowest point in the space, and therefore the *weakest* test. It + /// exists so a test can turn on exactly the one freedom it is about, and + /// so a reader can tell that choice from an accident. + #[must_use] + pub const fn narrowest() -> Self { + Self { + may_pend: false, + may_reorder: false, + may_fail_operations: false, + may_expire_waits: false, + may_wake_empty: false, + edge_triggered_signal: true, + may_fail_submits: false, + may_transfer_partially: false, + } + } +} + +/// What a resolution actually did, readable while the resolver is installed. +/// +/// A test that narrows the config to isolate one freedom needs to know the +/// freedom was *exercised*; otherwise a green run means only that the seed +/// never took that branch. These counters are how a test asserts against the +/// run rather than against the configuration. +#[derive(Debug, Default, Clone, Copy, PartialEq, Eq)] +pub struct ResolverStats { + /// Operations built (any `Build*` call that returned success). + pub built: usize, + /// Operations moved from built to submitted. + pub submitted: usize, + /// Completions posted to the completion queue. + pub posted: usize, + /// Completions handed to a `PopIoRingCompletion`. + pub popped: usize, + /// Posts that carried a failure result (`RS-P-3` exercised). + pub failed_operations: usize, + /// Submits declined, leaving built operations queued (`RS-P-7`). + pub failed_submits: usize, + /// Waits answered as expired (`RS-P-4`). + pub expired_waits: usize, + /// Waits answered successfully with nothing poppable (`RS-P-5`). + pub empty_wakes: usize, + /// Event signals raised (`RS-P-6`; zero when no event is attached). + pub signals: usize, + /// Posts that took something other than the oldest unresolved operation + /// (`RS-P-2` exercised -- reordering actually happened, as distinct from + /// having been permitted). + pub reorderings: usize, + /// Times an operation was left unresolved by a consultation (`RS-P-1` + /// exercised -- something actually pended). + pub deferrals: usize, + /// Times a drain-flagged operation was held back because something queued + /// before it was still unresolved (`RS-C-4` actually bit, as distinct from + /// having been vacuously satisfied). + pub barrier_holds: usize, + /// Successful transfers reported short under `RS-P-8`. + pub partial_transfers: usize, +} + +/// A live view of [`ResolverStats`] for an installed resolver. +/// +/// `Rc` rather than `Arc` because a responder is thread-local +/// (the seam's `installed` module), so sharing it across threads is already +/// meaningless. +#[derive(Debug, Clone)] +pub struct ResolverWatch(Rc>); + +impl ResolverWatch { + /// The statistics as of now. + #[must_use] + pub fn stats(&self) -> ResolverStats { + *self.0.borrow() + } +} + +/// One operation the resolver is answering for. +#[derive(Debug, Clone, Copy)] +struct Op { + /// `RS-C-2`: echoed into the completion exactly as supplied. + user_data: usize, + /// Build order, ring-wide. `RS-C-4` is stated over this. + seq: u64, + /// Carries `IOSQE_FLAGS_DRAIN_PRECEDING_OPS`. + barrier: bool, + /// Consultations survived without being posted. + deferrals: u32, + /// Bytes asked for, when this operation is a transfer. `None` for flush + /// and cancel, whose `Information` is not a byte count at all -- which is + /// why `RS-P-8` is stated over transfers rather than over completions. + requested: Option, +} + +/// A resolver over [RESPONSE-SPACE.md](../RESPONSE-SPACE.md). +/// +/// See the module documentation for what drives resolution and for the two +/// ordering requirements installing one imposes. +pub struct Resolver { + seed: u64, + state: u64, + config: ResolverConfig, + /// Built, not yet submitted. `RS-C-3` is exactly the rule that nothing + /// here is eligible to complete. + staged: Vec, + /// Submitted, not yet posted. Held in build order, which is what lets + /// `RS-C-4` be decided by position. + pool: Vec, + /// Posted, not yet popped. The completion queue: identity, result code, + /// and the `Information` the completion will carry (`RS-P-8`). + posted: VecDeque<(usize, HRESULT, usize)>, + next_seq: u64, + consecutive_submit_failures: u32, + /// The ring's completion event, if one has been attached. Borrowed, never + /// owned: the ring closes it, and this resolver must not outlive the ring. + event: *mut c_void, + stats: Rc>, +} + +impl std::fmt::Debug for Resolver { + fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result { + f.debug_struct("Resolver") + .field("seed", &format_args!("0x{:016X}", self.seed)) + .field("config", &self.config) + .field("staged", &self.staged.len()) + .field("pool", &self.pool.len()) + .field("posted", &self.posted.len()) + .field("stats", &*self.stats.borrow()) + .finish() + } +} + +impl Resolver { + /// A resolver at the widest point in the space, from an explicit seed. + #[must_use] + pub fn new(seed: u64) -> Self { + Self::with_config(seed, ResolverConfig::default()) + } + + /// A resolver exercising exactly the permissions `config` allows. + #[must_use] + pub fn with_config(seed: u64, config: ResolverConfig) -> Self { + Self { + seed, + // `| 1` so a zero seed is not a fixed point of the mixer. + state: seed.wrapping_mul(0x9E37_79B9_7F4A_7C15) | 1, + config, + staged: Vec::new(), + pool: Vec::new(), + posted: VecDeque::new(), + next_seq: 0, + consecutive_submit_failures: 0, + event: std::ptr::null_mut(), + stats: Rc::new(RefCell::new(ResolverStats::default())), + } + } + + /// The seed this resolver is replaying. + #[must_use] + pub fn seed(&self) -> u64 { + self.seed + } + + /// A live view of what this resolution has done so far. + #[must_use] + pub fn watch(&self) -> ResolverWatch { + ResolverWatch(Rc::clone(&self.stats)) + } + + /// The seed for this run: pinned from [`SEED_VAR`] when set, otherwise + /// derived from the clock so successive runs explore different points. + /// + /// # Panics + /// + /// If [`SEED_VAR`] is set to something that is neither decimal nor `0x` + /// hex. A seed that silently fell back to the clock would make the + /// variable look like it worked while the run it was pinning drifted. + #[must_use] + pub fn seed_from_env() -> u64 { + match std::env::var(SEED_VAR) { + Ok(text) => { + let trimmed = text.trim(); + trimmed + .strip_prefix("0x") + .or_else(|| trimmed.strip_prefix("0X")) + .map_or_else( + || trimmed.parse::().ok(), + |hex| u64::from_str_radix(hex, 16).ok(), + ) + .unwrap_or_else(|| { + panic!("{SEED_VAR} is set to {text:?}, which is neither decimal nor 0x hex") + }) + } + Err(_) => std::time::SystemTime::now() + .duration_since(std::time::UNIX_EPOCH) + .map_or(0x243F_6A88_85A3_08D3, |elapsed| elapsed.as_nanos() as u64), + } + } + + /// The command that replays this resolver's choices. + /// + /// Print it from any test that installs one. It names **only this axis**: + /// a test that also draws on another seeded source announces that source + /// too, because a partial replay is the failure `SEED_VAR` documents. + #[must_use] + pub fn replay_hint(&self) -> String { + format!("$env:{SEED_VAR}='0x{:016X}'", self.seed) + } + + /// Install this resolver, run `body`, then uninstall it. + /// + /// This is the shape to reach for, because it makes the module's first + /// ordering requirement structural: a ring created inside `body` is also + /// dropped inside it, so its rundown is still answered by the resolver + /// that owns its operations. + /// + /// # Panics + /// + /// If a responder is already installed on this thread. + pub fn scoped(self, body: impl FnOnce(&ResolverWatch) -> R) -> R { + let watch = self.watch(); + let guard: Installed = super::install(Box::new(self)); + let outcome = body(&watch); + drop(guard); + outcome + } + + /// SplitMix64, matching the generator this crate already uses + /// ([D-41](../DESIGN-NOTES.md#d-41)): one number replays a whole run. + fn next(&mut self) -> u64 { + self.state = self.state.wrapping_add(0x9E37_79B9_7F4A_7C15); + let mut z = self.state; + z = (z ^ (z >> 30)).wrapping_mul(0xBF58_476D_1CE4_E5B9); + z = (z ^ (z >> 27)).wrapping_mul(0x94D0_49BB_1331_11EB); + z ^ (z >> 31) + } + + /// True with probability `percent`. + /// + /// The weighting is a property of this resolver, never of the space -- + /// [RESPONSE-SPACE.md](../RESPONSE-SPACE.md) carries no rates, deliberately, + /// because a space with observed probabilities in it is a recording. + fn chance(&mut self, percent: u64) -> bool { + self.next() % 100 < percent + } + + fn record(&self, f: impl FnOnce(&mut ResolverStats)) { + f(&mut self.stats.borrow_mut()); + } + + /// Record a built operation and return the success every `Build*` call + /// returns when it queues an SQE. + fn build(&mut self, user_data: usize, flags: i32, requested: Option) -> HRESULT { + let seq = self.next_seq; + self.next_seq += 1; + self.staged.push(Op { + user_data, + seq, + barrier: flags & IOSQE_FLAGS_DRAIN_PRECEDING_OPS != 0, + deferrals: 0, + requested, + }); + self.record(|s| s.built += 1); + S_OK + } + + /// Post one completion, choosing which under `RS-P-2` and `RS-C-4`. + /// + /// Returns false only when there is nothing to post. + fn post_one(&mut self) -> bool { + if self.pool.is_empty() { + return false; + } + + // `RS-C-4`: no operation queued before a drain-flagged operation may + // complete after it. `pool` is in build order, so the set queued + // before `pool[i]` is exactly `pool[..i]` -- a barrier is therefore + // eligible only when it is the oldest unresolved operation. + // + // That reduction is the whole of this constraint's implementation, and + // it is load-bearing rather than descriptive: it holds only while the + // pool stays sorted by build order. An edit that sorted or reshuffled + // the pool would relax `RS-C-4` to nothing while every line below kept + // reading as though it still applied, so the invariant is asserted + // rather than asserted-in-a-comment. + debug_assert!( + self.pool.windows(2).all(|pair| pair[0].seq < pair[1].seq), + "the pool must stay in build order or RS-C-4's positional test is meaningless" + ); + // + // Note which way this resolves against `RS-P-2`: ordering is + // "unconstrained" as a permission, and a constraint overrides a + // permission. The hold-back half is *not* constrained -- D-24 claimed + // it and D-47 withdrew the claim -- so an operation queued after a + // barrier is eligible at any time, including before the barrier. + let eligible: Vec = (0..self.pool.len()) + .filter(|&i| i == 0 || !self.pool[i].barrier) + .collect(); + let held = self.pool.len() - eligible.len(); + if held > 0 { + self.record(|s| s.barrier_holds += 1); + } + + let pick = if self.config.may_reorder { + // `RS-P-2`: uniform over everything eligible. Repeated to + // exhaustion this is a uniform shuffle of each barrier-delimited + // segment, which is the widest order the constraint leaves. + let draw = self.next() % eligible.len() as u64; + eligible[usize::try_from(draw).expect("a modulus of a usize length fits a usize")] + } else { + eligible[0] + }; + if pick != 0 { + self.record(|s| s.reorderings += 1); + } + + let op = self.pool.remove(pick); + + // `RS-P-3`: any single operation may fail while its neighbours + // succeed. The *set* of codes is a property of this resolver; the + // space enumerates none and says a consumer must not depend on the + // set being small, so this draws across the whole Win32 facility + // rather than from a handful of plausible ones. + let result = if self.config.may_fail_operations && self.chance(20) { + self.record(|s| s.failed_operations += 1); + let code = self.next() & 0xFFFF; + (0x8007_0000_u32 | u32::try_from(code).expect("masked to 16 bits")) as HRESULT + } else { + S_OK + }; + + let was_empty = self.posted.is_empty(); + // `RS-P-8`: a successful transfer may report fewer bytes than were + // asked for. Documented for non-blocking byte-mode pipes -- and this + // crate constrains the handle type not at all -- so a consumer must + // read the count rather than assume it. A failed operation transfers + // nothing, and flush and cancel carry no byte count, so neither is + // eligible. + let information = match (result == S_OK, op.requested) { + (true, Some(requested)) => { + if self.config.may_transfer_partially && requested > 0 && self.chance(8) { + self.record(|s| s.partial_transfers += 1); + // Strictly short, including zero: a caller who loops on + // the remainder must tolerate making no progress rather + // than assuming every completion advances it. + let short = self.next() % u64::from(requested); + usize::try_from(short).expect("short count fits, it is below a u32") + } else { + usize::try_from(requested).expect("a u32 fits a usize on this target") + } + } + _ => 0, + }; + self.posted.push_back((op.user_data, result, information)); + self.record(|s| s.posted += 1); + + // `RS-P-6`: the completion event fires on the empty-to-non-empty + // edge, so a completion arriving behind another produces no signal of + // its own (D-19, measured). With `edge_triggered_signal` off this + // signals every post, which is wider than the platform and exists to + // show the edge is load-bearing. + if !self.event.is_null() && (was_empty || !self.config.edge_triggered_signal) { + self.record(|s| s.signals += 1); + // SAFETY: `event` was handed to `set_completion_event` by the ring + // that owns it, and this resolver is required to be uninstalled + // only after that ring is gone. + unsafe { SetEvent(self.event) }; + } + true + } + + /// Advance resolution by one consultation. + /// + /// `flip` distinguishes the two clocks the module documents: a submit + /// flips coins, a pop only rescues operations that have run out of + /// deferrals. + fn tick(&mut self, flip: bool) { + if self.pool.is_empty() { + return; + } + + let mut want = 0_usize; + for i in 0..self.pool.len() { + let forced = self.pool[i].deferrals >= MAX_DEFERRALS; + // `RS-P-1`: each operation independently completes now or pends. + // A forced post is `RS-C-1` overriding that permission, which is + // the same precedence `RS-C-4` takes over `RS-P-2`. + if forced || (flip && (!self.config.may_pend || self.chance(50))) { + want += 1; + } + } + + let deferred = self.pool.len() - want.min(self.pool.len()); + if deferred > 0 { + self.record(|s| s.deferrals += deferred); + } + for op in &mut self.pool { + op.deferrals += 1; + } + + for _ in 0..want { + if !self.post_one() { + break; + } + } + } +} + +impl Responses for Resolver { + unsafe fn submit( + &mut self, + _ring: *mut c_void, + wait_operations: u32, + _milliseconds: u32, + submitted: *mut u32, + ) -> HRESULT { + // `RS-P-7`: a failed submit leaves already-built operations queued for + // a later submit -- there is no rewind once `Build*` returns (D-5). + // Bounded, because a resolver that failed every submit forever would + // never complete anything and `RS-C-1` forbids that. + if self.config.may_fail_submits + && !self.staged.is_empty() + && self.consecutive_submit_failures < MAX_CONSECUTIVE_SUBMIT_FAILURES + && self.chance(15) + { + self.consecutive_submit_failures += 1; + self.record(|s| s.failed_submits += 1); + // SAFETY: the caller passes a valid out-pointer, as the Win32 call + // requires. + unsafe { submitted.write(0) }; + return SUBMIT_FAILED; + } + self.consecutive_submit_failures = 0; + + // `RS-C-3`: nothing completes before it is submitted. That rule is + // this move and nothing else -- only `pool` is ever eligible to post. + let moved = self.staged.len(); + self.pool.append(&mut self.staged); + self.record(|s| s.submitted += moved); + // SAFETY: as above. + unsafe { + submitted.write(u32::try_from(moved).unwrap_or(u32::MAX)); + } + + self.tick(true); + + if wait_operations == 0 { + return S_OK; + } + + // `RS-P-4`: a wait may expire, and may do so even when completions are + // available and on the very first call. This is not a ring failure -- + // M21.6 fixed a defect in this crate that read it as one. + if self.config.may_expire_waits && self.chance(25) { + self.record(|s| s.expired_waits += 1); + return IORING_E_WAIT_TIMEOUT; + } + + // `RS-P-5`: a wait returning success promises nothing about + // poppability. Reached whenever the queue is empty here -- either + // because resolution deferred everything, or because this clause + // chose to wake with work still pending. + if self.posted.is_empty() { + self.record(|s| s.empty_wakes += 1); + } + S_OK + } + + unsafe fn pop(&mut self, _ring: *mut c_void, cqe: *mut IORING_CQE) -> HRESULT { + if self.posted.is_empty() { + // Starvation rescue only: no coins. See the module's note on the + // two clocks -- flipping here would make `RS-P-1`'s pending case + // unobservable to a polling consumer. + self.tick(false); + } + + let Some((user_data, result, information)) = self.posted.pop_front() else { + return S_FALSE; + }; + self.record(|s| s.popped += 1); + // SAFETY: the caller passes a valid out-pointer, as the Win32 call + // requires. + unsafe { + cqe.write(IORING_CQE { + // `RS-C-2`: a completion identifies the operation that + // produced it, by carrying back exactly the value supplied + // when it was built. + UserData: user_data, + ResultCode: result, + Information: information, + }); + } + S_OK + } + + unsafe fn build_read( + &mut self, + _ring: *mut c_void, + _file: IORING_HANDLE_REF, + _buffer: IORING_BUFFER_REF, + bytes: u32, + _offset: u64, + user_data: usize, + flags: i32, + ) -> HRESULT { + self.build(user_data, flags, Some(bytes)) + } + + unsafe fn build_write( + &mut self, + _ring: *mut c_void, + _file: IORING_HANDLE_REF, + _buffer: IORING_BUFFER_REF, + bytes: u32, + _offset: u64, + _caching: i32, + user_data: usize, + flags: i32, + ) -> HRESULT { + self.build(user_data, flags, Some(bytes)) + } + + unsafe fn build_flush( + &mut self, + _ring: *mut c_void, + _file: IORING_HANDLE_REF, + _mode: FILE_FLUSH_MODE, + user_data: usize, + flags: i32, + ) -> HRESULT { + self.build(user_data, flags, None) + } + + unsafe fn build_cancel( + &mut self, + _ring: *mut c_void, + _file: IORING_HANDLE_REF, + _cancel_user_data: usize, + user_data: usize, + ) -> HRESULT { + // A cancel carries no SQE flags of its own in this crate's surface, so + // it is never a barrier. It is still an operation and still completes + // exactly once under `RS-C-1`. + self.build(user_data, 0, None) + } + + unsafe fn build_register_files( + &mut self, + _ring: *mut c_void, + _count: u32, + _handles: *const *mut c_void, + user_data: usize, + ) -> HRESULT { + self.build(user_data, 0, None) + } + + unsafe fn build_register_buffers( + &mut self, + _ring: *mut c_void, + _count: u32, + _buffers: *const IORING_BUFFER_INFO, + user_data: usize, + ) -> HRESULT { + self.build(user_data, 0, None) + } + + unsafe fn set_completion_event(&mut self, _ring: *mut c_void, event: *mut c_void) -> HRESULT { + // Kept, not forwarded. The real ring would signal for operations it + // never received; this resolver signals for the ones it is answering + // for, on the edge `RS-P-6` describes. + self.event = event; + S_OK + } +} + +#[cfg(test)] +mod tests; diff --git a/crates/windows-ioring-sys/src/sys/resolver/tests.rs b/crates/windows-ioring-sys/src/sys/resolver/tests.rs new file mode 100644 index 000000000..12e19a71b --- /dev/null +++ b/crates/windows-ioring-sys/src/sys/resolver/tests.rs @@ -0,0 +1,927 @@ +// Copyright (c) Mike Grier +//! Tests for the response-space resolver (M26.3). +//! +//! # What these assert, and what they deliberately do not +//! +//! Each test names the clause it is about and asserts two things: that the +//! resolver *may* do what the clause permits, and -- where the clause is a +//! freedom rather than a shape -- that it actually **did**, by reading +//! [`ResolverStats`]. The second half is the one worth having. A resolver that +//! permits reordering and never reorders passes any test written only against +//! the first, and a permission nothing exercises is indistinguishable from one +//! nothing implemented. +//! +//! These drive the resolver through [`Responses`] directly rather than through +//! an `IoRing`. That is not a convenience: this file is about whether the +//! resolver occupies the space it claims to, which is a property of the +//! resolver alone. `M26.4` is where a real ring is driven across it, and +//! `M26.5` is where the resolution is calibrated against defects this crate +//! actually shipped. +//! +//! # The markers below are read by a test +//! +//! `M26.6` added a census ([response_space_census.rs](../../../tests/response_space_census.rs)) +//! that fails when a clause in the space is claimed by nothing. It reads these +//! `EXERCISES:` lines rather than searching for the clause ID anywhere in the +//! text, and the difference is not cosmetic -- the first version of that +//! census matched any mention, so a file saying "that clause is somebody +//! else's job" counted as checking it. A marker is a claim; a mention is not. +//! +//! EXERCISES: RS-P-1 +//! EXERCISES: RS-P-2 +//! EXERCISES: RS-P-3 +//! EXERCISES: RS-P-4 +//! EXERCISES: RS-P-5 +//! EXERCISES: RS-P-6 +//! EXERCISES: RS-P-7 +//! EXERCISES: RS-P-8 + +use std::ffi::c_void; +use std::ptr; + +use windows_sys::Win32::Storage::FileSystem::{ + IORING_BUFFER_REF, IORING_BUFFER_REF_0, IORING_CQE, IORING_HANDLE_REF, IORING_HANDLE_REF_0, + IORING_REF_RAW, IOSQE_FLAGS_DRAIN_PRECEDING_OPS, +}; + +use super::{Resolver, ResolverConfig, ResolverStats}; +use crate::sys::Responses; + +/// Seeds every ordering test sweeps. +/// +/// Fixed rather than drawn, because a unit test must be reproducible without +/// an announcement (the repository's standing rule on random sampling). The +/// *environment* seed exists for the generated suites, which announce it; +/// here a fixed sweep is what makes a failure mean the same thing twice. +/// +/// Raised from 64 to 2048 for breadth. The cost is a property of the sweep and +/// not of the seed count alone -- measured, a seed here is tens of +/// microseconds, so the whole file stays far inside the sub-second budget this +/// repository sets for a submodule's unit tests. +const SEEDS: std::ops::Range = 0..2048; + +/// How many seeds [`SEEDS`] covers. +/// +/// Spelled out because `Range` is not an `ExactSizeIterator` -- 64-bit +/// ranges can exceed `usize` -- so there is no `len()` to ask for. +const SEED_COUNT: usize = 2048; + +/// Distinct orders six independently-resolving operations can be posted in. +/// +/// `6!`. This is the ceiling on what +/// [`different_seeds_reach_different_resolutions`] can observe, and it is +/// *below* [`SEED_COUNT`] -- which is exactly the trap that assertion fell +/// into when the sweep was widened. A threshold phrased as a fraction of the +/// seed count silently becomes unsatisfiable once the seeds outnumber the +/// outcomes, so the bound is stated against the space the test can actually +/// reach. +/// +/// Measured over the resolver's own mixer: 64 seeds reach 64 of these, 1024 +/// reach 539, 2048 reach 670, and 8192 are needed for all 720. So the sweep is +/// still gaining breadth at its current size rather than re-drawing orders it +/// has already seen. +const DISTINCT_ORDERS_OF_SIX: usize = 720; + +/// A handle reference that is never dereferenced. +/// +/// The resolver answers every call it is given without touching the caller's +/// pointers, so a null one is the honest argument: if some future edit did +/// start dereferencing, these tests would fault rather than silently pass on +/// a plausible-looking fake. That is the same reasoning `M23.5` recorded about +/// ring handles -- a fake that looks valid is worse than one that cannot be. +fn handle() -> IORING_HANDLE_REF { + IORING_HANDLE_REF { + Kind: IORING_REF_RAW, + Handle: IORING_HANDLE_REF_0 { + Handle: ptr::null_mut(), + }, + } +} + +/// Drive a read build, the cheapest operation that carries a requested length +/// and so the only shape `RS-P-8` is stated over. +fn build_read(resolver: &mut Resolver, user_data: usize, bytes: u32) { + // SAFETY: the resolver dereferences none of these; see `handle`. + let hr = unsafe { + resolver.build_read( + ptr::null_mut(), + handle(), + IORING_BUFFER_REF { + Kind: IORING_REF_RAW, + Buffer: IORING_BUFFER_REF_0 { + Address: ptr::null_mut(), + }, + }, + bytes, + 0, + user_data, + 0, + ) + }; + assert_eq!(hr, 0, "a build should report success"); +} + +/// Drive a flush build, which is the cheapest operation to synthesise and the +/// only one that can carry a barrier through this crate's public surface. +fn build_flush(resolver: &mut Resolver, user_data: usize, barrier: bool) { + let flags = if barrier { + IOSQE_FLAGS_DRAIN_PRECEDING_OPS + } else { + 0 + }; + // SAFETY: the resolver dereferences none of these; see `handle`. + let hr = unsafe { resolver.build_flush(ptr::null_mut(), handle(), 0, user_data, flags) }; + assert_eq!(hr, 0, "a build should report success"); +} + +/// Submit without waiting. +fn submit(resolver: &mut Resolver) -> i32 { + let mut submitted = 0_u32; + // SAFETY: `submitted` is a valid out-pointer; the ring pointer is unused. + unsafe { resolver.submit(ptr::null_mut(), 0, 0, &raw mut submitted) } +} + +/// Submit asking to wait, which is what `RS-P-4` and `RS-P-5` are about. +fn submit_waiting(resolver: &mut Resolver) -> i32 { + let mut submitted = 0_u32; + // SAFETY: as `submit`. + unsafe { resolver.submit(ptr::null_mut(), 1, 50, &raw mut submitted) } +} + +/// Pop one completion, or `None` when the queue is empty. +/// Pop carrying the transferred count, which `pop` discards. +fn pop_full(resolver: &mut Resolver) -> Option<(usize, i32, usize)> { + let mut cqe = IORING_CQE { + UserData: 0, + ResultCode: 0, + Information: 0, + }; + // SAFETY: `cqe` is a valid out-pointer. + let hr = unsafe { resolver.pop(ptr::null_mut(), &raw mut cqe) }; + if hr == 1 { + return None; + } + assert_eq!( + hr, 0, + "a pop should either succeed or report an empty queue" + ); + Some((cqe.UserData, cqe.ResultCode, cqe.Information)) +} + +fn pop(resolver: &mut Resolver) -> Option<(usize, i32)> { + let mut cqe = IORING_CQE { + UserData: 0, + ResultCode: 0, + Information: 0, + }; + // SAFETY: `cqe` is a valid out-pointer. + let hr = unsafe { resolver.pop(ptr::null_mut(), &raw mut cqe) }; + if hr == 1 { + return None; + } + assert_eq!( + hr, 0, + "a pop should either succeed or report an empty queue" + ); + Some((cqe.UserData, cqe.ResultCode)) +} + +/// Drain everything the resolver will ever produce, submitting to drive +/// resolution, and return the completions in the order they were handed over. +/// +/// Bounded so a resolver that stopped making progress fails as a test rather +/// than as a hang -- which is the whole difference between a caught defect and +/// a CI job somebody cancels. +fn drain(resolver: &mut Resolver, expected: usize) -> Vec<(usize, i32)> { + let mut seen = Vec::new(); + for _ in 0..(expected + 2) * (super::MAX_DEFERRALS as usize + 2) { + while let Some(completion) = pop(resolver) { + seen.push(completion); + } + if seen.len() == expected { + return seen; + } + submit(resolver); + } + panic!( + "resolution did not finish: {} of {expected} completions after the bound, which is \ + RS-C-1 not holding in finite time", + seen.len() + ); +} + +/// Run `body` over every seed in [`SEEDS`] and return the accumulated stats. +fn sweep(config: ResolverConfig, mut body: impl FnMut(&mut Resolver)) -> Vec { + SEEDS + .map(|seed| { + let mut resolver = Resolver::with_config(seed, config); + let watch = resolver.watch(); + body(&mut resolver); + watch.stats() + }) + .collect() +} + +// ------------------------------------------------------- RS-C-1 and RS-C-2 --- + +#[test] +fn every_submitted_operation_completes_exactly_once_carrying_its_own_identity() { + // RS-C-1 and RS-C-2 together, over the widest configuration: whatever else + // the resolver chooses, the multiset of completions equals the multiset of + // submissions. This is the constraint every other test rests on, so it is + // swept rather than sampled. + for seed in SEEDS { + let mut resolver = Resolver::new(seed); + for id in 1..=8_usize { + build_flush(&mut resolver, id, false); + } + submit(&mut resolver); + let mut ids: Vec = drain(&mut resolver, 8) + .into_iter() + .map(|(id, _)| id) + .collect(); + ids.sort_unstable(); + assert_eq!( + ids, + (1..=8).collect::>(), + "seed {seed}: every operation must complete exactly once, carrying the user data \ + it was built with" + ); + } +} + +#[test] +fn an_operation_that_was_never_submitted_does_not_complete() { + // RS-C-3. Built is not submitted: D-5 records that there is no rewind once + // `Build*` returns, but that is about the operation being *queued*, not + // about it having run. + // + // POLLED PAST THE DEFERRAL BOUND ON PURPOSE, and the reason is a defect + // this test had until the sabotage sweep found it. A single pop asserts + // only that nothing completed *yet*, and a resolver that had wrongly made + // the staged operations eligible would still answer that pop with nothing + // -- the starvation rescue posts only what has run out of deferrals, and + // a freshly built operation has not. The claim is that unsubmitted work is + // never ELIGIBLE, so the test has to outlast the bound that would make an + // eligible operation surface. + let mut resolver = Resolver::new(1); + for id in 1..=4_usize { + build_flush(&mut resolver, id, false); + } + for poll in 0..=(super::MAX_DEFERRALS + 2) { + assert_eq!( + pop(&mut resolver), + None, + "poll {poll}: nothing may complete before a submit carries it" + ); + } + let stats = resolver.watch().stats(); + assert_eq!(stats.built, 4, "four operations were built"); + assert_eq!( + stats.submitted, 0, + "none was submitted, so none may have become eligible" + ); + assert_eq!(stats.posted, 0, "and none may have been posted"); +} + +// ----------------------------------------------------------------- RS-C-4 --- + +#[test] +fn nothing_queued_before_a_drained_flush_completes_after_it() { + // RS-C-4, the one place this space is narrower than "anything may happen", + // swept across every seed because it is a claim about *all* resolutions. + for seed in SEEDS { + let mut resolver = Resolver::new(seed); + // Three ordinary operations, then a barrier, then three more. The + // trailing three are deliberately present: the hold-back half is NOT + // constrained (D-24 claimed it, D-47 withdrew it), so they may appear + // anywhere at all -- including before the barrier. + for id in 1..=3_usize { + build_flush(&mut resolver, id, false); + } + build_flush(&mut resolver, 100, true); + for id in 4..=6_usize { + build_flush(&mut resolver, id, false); + } + submit(&mut resolver); + + let order: Vec = drain(&mut resolver, 7) + .into_iter() + .map(|(id, _)| id) + .collect(); + let barrier_at = order + .iter() + .position(|&id| id == 100) + .expect("the barrier must complete"); + for id in 1..=3_usize { + let at = order + .iter() + .position(|&candidate| candidate == id) + .expect("every operation completes"); + assert!( + at < barrier_at, + "seed {seed}: operation {id} was queued before the drained flush and completed \ + after it, which RS-C-4 forbids -- order was {order:?}" + ); + } + } +} + +#[test] +fn the_hold_back_half_of_the_drain_flag_is_not_constrained() { + // The other side of RS-C-4, and the reason the clause is one-sided. D-47 + // withdrew D-24's claim that a drained flush holds back what follows, so a + // resolver that enforced both halves would be narrower than the space and + // would hide exactly the defect class D-47 found. + let overtook = SEEDS + .filter(|&seed| { + let mut resolver = Resolver::new(seed); + build_flush(&mut resolver, 1, false); + build_flush(&mut resolver, 100, true); + build_flush(&mut resolver, 2, false); + submit(&mut resolver); + let order: Vec = drain(&mut resolver, 3) + .into_iter() + .map(|(id, _)| id) + .collect(); + let barrier_at = order.iter().position(|&id| id == 100).unwrap(); + let after_at = order.iter().position(|&id| id == 2).unwrap(); + after_at < barrier_at + }) + .count(); + assert!( + overtook > 0, + "no seed let an operation queued after a drained flush complete before it, so the \ + resolver is enforcing a hold-back D-47 withdrew" + ); +} + +#[test] +fn a_barrier_is_actually_held_back_rather_than_vacuously_satisfied() { + // RS-C-4 could be satisfied by a resolver that never had a barrier with + // anything in front of it. The counter distinguishes "the constraint bit" + // from "the constraint never applied", which is the difference between a + // test and a tautology. + let stats = sweep(ResolverConfig::default(), |resolver| { + build_flush(resolver, 1, false); + build_flush(resolver, 2, false); + build_flush(resolver, 100, true); + submit(resolver); + drain(resolver, 3); + }); + let held: usize = stats.iter().map(|s| s.barrier_holds).sum(); + assert!( + held > 0, + "no seed ever held a barrier back, so RS-C-4 was satisfied vacuously and this suite \ + proves nothing about it" + ); +} + +// ----------------------------------------------------------------- RS-P-1 --- + +#[test] +fn an_operation_may_pend_rather_than_complete_inside_the_submit() { + let stats = sweep(ResolverConfig::default(), |resolver| { + for id in 1..=6_usize { + build_flush(resolver, id, false); + } + submit(resolver); + drain(resolver, 6); + }); + assert!( + stats.iter().any(|s| s.deferrals > 0), + "no seed left an operation unresolved by the submit that carried it, so RS-P-1's \ + pending case is permitted but never exercised" + ); +} + +#[test] +fn narrowing_rs_p_1_completes_everything_inside_the_submit() { + // The converse, which is what makes the permission's switch meaningful: a + // test isolating some other freedom needs `may_pend: false` to actually + // remove this one. + for seed in SEEDS { + let mut resolver = Resolver::with_config( + seed, + ResolverConfig { + may_pend: false, + ..ResolverConfig::narrowest() + }, + ); + for id in 1..=5_usize { + build_flush(&mut resolver, id, false); + } + submit(&mut resolver); + let mut seen = Vec::new(); + while let Some((id, _)) = pop(&mut resolver) { + seen.push(id); + } + assert_eq!( + seen.len(), + 5, + "seed {seed}: with RS-P-1 narrowed, every operation completes inside its submit" + ); + } +} + +// ----------------------------------------------------------------- RS-P-2 --- + +#[test] +fn completion_order_is_unconstrained() { + // RS-P-2, and the demonstration the probe already made: a FIFO-assuming + // consumer breaks. Asserting that *some* seed reorders is the honest form + // -- a specific permutation at a specific seed would be asserting against + // the mixer rather than against the clause. + let reordered = SEEDS + .filter(|&seed| { + let mut resolver = Resolver::with_config( + seed, + ResolverConfig { + may_reorder: true, + ..ResolverConfig::narrowest() + }, + ); + for id in 1..=5_usize { + build_flush(&mut resolver, id, false); + } + submit(&mut resolver); + let order: Vec = drain(&mut resolver, 5) + .into_iter() + .map(|(id, _)| id) + .collect(); + order != (1..=5).collect::>() + }) + .count(); + assert!( + reordered > SEED_COUNT / 2, + "only {reordered} of {SEED_COUNT} seeds reordered, which is too few for a clause that \ + permits any permutation" + ); +} + +#[test] +fn narrowing_rs_p_2_posts_in_submission_order() { + for seed in SEEDS { + let mut resolver = Resolver::with_config(seed, ResolverConfig::narrowest()); + for id in 1..=5_usize { + build_flush(&mut resolver, id, false); + } + submit(&mut resolver); + let order: Vec = drain(&mut resolver, 5) + .into_iter() + .map(|(id, _)| id) + .collect(); + assert_eq!( + order, + (1..=5).collect::>(), + "seed {seed}: with RS-P-2 narrowed the order is submission order" + ); + } +} + +// ----------------------------------------------------------------- RS-P-3 --- + +#[test] +fn an_operation_may_fail_while_its_neighbours_succeed() { + // RS-P-3. The mixed case is the one that matters: M22.2's defect was an + // ordering error in handling a *single* failed write among successes, and + // a resolver that only ever failed all or none would not have reached it. + let mixed = SEEDS + .filter(|&seed| { + let mut resolver = Resolver::with_config( + seed, + ResolverConfig { + may_fail_operations: true, + ..ResolverConfig::narrowest() + }, + ); + for id in 1..=6_usize { + build_flush(&mut resolver, id, false); + } + submit(&mut resolver); + let results = drain(&mut resolver, 6); + let failures = results.iter().filter(|(_, hr)| *hr != 0).count(); + failures > 0 && failures < results.len() + }) + .count(); + assert!( + mixed > 0, + "no seed produced a batch in which some operations failed and others succeeded" + ); +} + +#[test] +fn failure_codes_are_not_drawn_from_a_small_fixed_set() { + // RS-P-3 says a consumer must not depend on the set of codes being small, + // and enumerates none. A resolver returning one hard-coded error would + // satisfy the clause's letter while teaching every consumer built against + // it to match on that one code. + let mut codes = std::collections::BTreeSet::new(); + for seed in SEEDS { + let mut resolver = Resolver::with_config( + seed, + ResolverConfig { + may_fail_operations: true, + ..ResolverConfig::narrowest() + }, + ); + for id in 1..=4_usize { + build_flush(&mut resolver, id, false); + } + submit(&mut resolver); + for (_, hr) in drain(&mut resolver, 4) { + if hr != 0 { + codes.insert(hr); + } + } + } + assert!( + codes.len() > 16, + "only {} distinct failure codes across {SEED_COUNT} seeds, which is a small fixed set \ + in all but name", + codes.len() + ); +} + +// --------------------------------------------------------- RS-P-4, RS-P-5 --- + +#[test] +fn a_wait_may_expire() { + // RS-P-4. `IORING_E_WAIT_TIMEOUT` is not a ring failure -- M21.6 fixed a defect in + // this crate that read it as one -- so a resolver that never produced it + // would leave that correction untested. + let expired = SEEDS + .filter(|&seed| { + let mut resolver = Resolver::with_config( + seed, + ResolverConfig { + may_expire_waits: true, + ..ResolverConfig::narrowest() + }, + ); + build_flush(&mut resolver, 1, false); + let hr = submit_waiting(&mut resolver); + let expired = hr == super::IORING_E_WAIT_TIMEOUT; + drain(&mut resolver, 1); + expired + }) + .count(); + assert!( + expired > 0, + "no seed expired a wait, so RS-P-4 is permitted but never exercised" + ); +} + +#[test] +fn a_wait_may_expire_even_when_a_completion_is_available() { + // The over-provision half of RS-P-4, stated separately because it is the + // part no measurement here established: a consumer that treats a + // successful return as "therefore something is poppable", or an expiry as + // "therefore nothing is", is wrong in both directions. + let expired_with_work = SEEDS + .filter(|&seed| { + let mut resolver = Resolver::with_config( + seed, + ResolverConfig { + may_expire_waits: true, + ..ResolverConfig::narrowest() + }, + ); + build_flush(&mut resolver, 1, false); + let hr = submit_waiting(&mut resolver); + let available = !resolver.posted.is_empty(); + drain(&mut resolver, 1); + hr == super::IORING_E_WAIT_TIMEOUT && available + }) + .count(); + assert!( + expired_with_work > 0, + "no seed expired a wait while a completion sat in the queue, so RS-P-4's \ + over-provision is unexercised and a consumer could still read an expiry as emptiness" + ); +} + +#[test] +fn a_successful_wait_promises_nothing_about_poppability() { + // RS-P-5, which this crate's own `pop_within` documentation already + // states, and which D-19's edge-triggered measurement explains. + let stats = sweep( + ResolverConfig { + may_wake_empty: true, + may_pend: true, + ..ResolverConfig::narrowest() + }, + |resolver| { + for id in 1..=3_usize { + build_flush(resolver, id, false); + } + submit_waiting(resolver); + drain(resolver, 3); + }, + ); + assert!( + stats.iter().any(|s| s.empty_wakes > 0), + "no seed returned a successful wait with an empty queue, so RS-P-5 is unexercised" + ); +} + +// ----------------------------------------------------------------- RS-P-6 --- + +#[test] +fn a_completion_posted_behind_another_raises_no_signal_of_its_own() { + // RS-P-6, which is a measurement rather than an over-provision: D-19 found + // the completion event edge triggered on empty-to-non-empty, and D-21 is + // the consequence this crate drew from it. + // + // A real event handle, because the resolver's signalling is the behaviour + // under test and a null handle would skip it entirely. + // SAFETY: a manual-reset, initially-unsignalled, unnamed event. + let event = unsafe { + windows_sys::Win32::System::Threading::CreateEventW(ptr::null(), 1, 0, ptr::null()) + }; + assert!(!event.is_null(), "CreateEventW failed"); + + let mut resolver = Resolver::with_config( + 7, + ResolverConfig { + may_pend: false, + ..ResolverConfig::narrowest() + }, + ); + // SAFETY: a live event this test owns for the resolver's whole lifetime. + let hr = unsafe { resolver.set_completion_event(ptr::null_mut(), event.cast::()) }; + assert_eq!(hr, 0); + + for id in 1..=4_usize { + build_flush(&mut resolver, id, false); + } + submit(&mut resolver); + assert_eq!( + resolver.watch().stats().signals, + 1, + "four completions posted back to back must raise exactly one signal -- the \ + empty-to-non-empty edge" + ); + + // Drain, then post again: the queue went empty, so the next post is a new + // edge and signals again. Without this half the test would pass for a + // resolver that signalled exactly once ever. + while pop(&mut resolver).is_some() {} + build_flush(&mut resolver, 5, false); + submit(&mut resolver); + assert_eq!( + resolver.watch().stats().signals, + 2, + "a post onto a queue that had drained is a fresh edge and must signal" + ); + + // SAFETY: this test owns the handle and the resolver is done with it. + unsafe { windows_sys::Win32::Foundation::CloseHandle(event) }; +} + +// ----------------------------------------------------------------- RS-P-7 --- + +#[test] +fn a_failed_submit_leaves_operations_queued_for_a_later_one() { + // RS-P-7, which follows from D-5: the submission queue is ring state, not + // batch state, and there is no rewind once `Build*` returns. + let stats = sweep( + ResolverConfig { + may_fail_submits: true, + ..ResolverConfig::narrowest() + }, + |resolver| { + for id in 1..=4_usize { + build_flush(resolver, id, false); + } + submit(resolver); + drain(resolver, 4); + }, + ); + let failed: usize = stats.iter().map(|s| s.failed_submits).sum(); + assert!( + failed > 0, + "no seed declined a submit, so RS-P-7 is unexercised" + ); + for (seed, s) in stats.iter().enumerate() { + assert_eq!( + s.popped, 4, + "seed {seed}: a declined submit must leave the operations queued, not lose them" + ); + } +} + +// ----------------------------------------------------- the seed discipline --- + +#[test] +fn one_seed_replays_a_whole_resolution() { + // D-41's discipline, applied to this axis. Two resolvers on the same seed + // must agree completely; if they did not, the replay command the resolver + // prints would be a lie. + let run = |seed| { + let mut resolver = Resolver::new(seed); + for id in 1..=6_usize { + build_flush(&mut resolver, id, id == 3); + } + submit(&mut resolver); + (drain(&mut resolver, 6), resolver.watch().stats()) + }; + for seed in SEEDS { + assert_eq!( + run(seed), + run(seed), + "seed {seed} did not replay identically" + ); + } +} + +#[test] +fn different_seeds_reach_different_resolutions() { + // The converse, and not a formality: a resolver whose seed did nothing + // would pass every replay test above while exploring one point forever. + // + // The bound is a fraction of [`DISTINCT_ORDERS_OF_SIX`], not of the seed + // count, and the difference is load-bearing rather than pedantic: six + // operations admit 720 orders, so a threshold of `SEED_COUNT / 2` becomes + // arithmetically unsatisfiable the moment the sweep exceeds 1440 seeds -- + // the test would fail without anything having regressed. + let orders: std::collections::BTreeSet> = SEEDS + .map(|seed| { + let mut resolver = Resolver::with_config( + seed, + ResolverConfig { + may_reorder: true, + ..ResolverConfig::narrowest() + }, + ); + for id in 1..=6_usize { + build_flush(&mut resolver, id, false); + } + submit(&mut resolver); + drain(&mut resolver, 6) + .into_iter() + .map(|(id, _)| id) + .collect() + }) + .collect(); + assert!( + orders.len() > DISTINCT_ORDERS_OF_SIX / 2, + "only {} of {DISTINCT_ORDERS_OF_SIX} possible orders across {SEED_COUNT} seeds", + orders.len() + ); +} + +#[test] +fn the_environment_seed_is_parsed_in_both_spellings_and_rejected_when_malformed() { + // The variable is the whole replay mechanism, so the failure that matters + // is a malformed value silently falling back to the clock: the run would + // look pinned and would not be. + // + // `seed_from_env` reads a process-wide variable, so this drives the + // parsing through the same expression rather than mutating the + // environment -- setting one here would race every other test in this + // binary, which runs its tests as threads in one process. + let parse = |text: &str| -> Option { + let trimmed = text.trim(); + trimmed + .strip_prefix("0x") + .or_else(|| trimmed.strip_prefix("0X")) + .map_or_else( + || trimmed.parse::().ok(), + |hex| u64::from_str_radix(hex, 16).ok(), + ) + }; + assert_eq!(parse("42"), Some(42)); + assert_eq!(parse("0x2A"), Some(42)); + assert_eq!(parse("0X2a"), Some(42)); + assert_eq!(parse(" 0x2A "), Some(42)); + assert_eq!( + parse("2A"), + None, + "bare hex is not decimal and must not parse" + ); + assert_eq!(parse("banana"), None); + assert_eq!(parse(""), None); +} + +#[test] +fn the_replay_hint_names_this_axis_and_the_seed_it_pins() { + let resolver = Resolver::new(0x2A); + assert_eq!( + resolver.replay_hint(), + format!("$env:{}='0x000000000000002A'", super::SEED_VAR) + ); + assert_eq!(resolver.seed(), 0x2A); +} + +// ------------------------------------------------------------ the defaults --- + +#[test] +fn the_default_configuration_is_the_widest_point_in_the_space() { + // The default matters because it decides what a test gets when it does not + // choose. Wide-by-default means an unconsidered test fails loudly under a + // freedom it did not handle; narrow-by-default would mean it passes while + // exercising nothing. + let wide = ResolverConfig::default(); + assert!(wide.may_pend); + assert!(wide.may_reorder); + assert!(wide.may_fail_operations); + assert!(wide.may_expire_waits); + assert!(wide.may_wake_empty); + assert!(wide.may_fail_submits); + + let narrow = ResolverConfig::narrowest(); + assert!(!narrow.may_pend); + assert!(!narrow.may_reorder); + assert!(!narrow.may_fail_operations); + assert!(!narrow.may_expire_waits); + assert!(!narrow.may_wake_empty); + assert!(!narrow.may_fail_submits); + + // Both keep the edge-triggered signal, because turning it *off* is the + // widening. RS-P-6 carries no over-provision -- it is the measurement -- + // so the narrowest point in the space still has it. + assert!(wide.edge_triggered_signal); + assert!(narrow.edge_triggered_signal); +} + +#[test] +fn a_successful_transfer_may_report_fewer_bytes_than_requested() { + // RS-P-8. Both directions, because a permission tested one way only says + // half of what it means: with the switch on some seed must report short, + // and with it off none may -- otherwise the switch is not what decides it. + const LEN: u32 = 4096; + + let mut short_seen = 0_usize; + let mut full_seen = 0_usize; + for seed in SEEDS { + let mut resolver = Resolver::with_config( + seed, + ResolverConfig { + may_transfer_partially: true, + ..ResolverConfig::narrowest() + }, + ); + for id in 1..=8_usize { + build_read(&mut resolver, id, LEN); + } + submit(&mut resolver); + while let Some((_, result, information)) = pop_full(&mut resolver) { + assert_eq!(result, 0, "narrowest() forbids a failed operation"); + assert!( + information <= LEN as usize, + "a transfer reported {information} bytes for a {LEN}-byte request, \ + which is over-delivery rather than a short count; seed {seed:#x}" + ); + if information == LEN as usize { + full_seen += 1; + } else { + short_seen += 1; + } + } + } + assert!( + short_seen > 0, + "no seed reported a short transfer, so RS-P-8 is unexercised" + ); + assert!( + full_seen > 0, + "every transfer was short, so the clause is being applied unconditionally \ + rather than as a permission" + ); + + // The other direction: with the permission withdrawn, a short count is a + // defect rather than a tolerated response. + for seed in SEEDS { + let mut resolver = Resolver::with_config(seed, ResolverConfig::narrowest()); + for id in 1..=8_usize { + build_read(&mut resolver, id, LEN); + } + submit(&mut resolver); + while let Some((_, _, information)) = pop_full(&mut resolver) { + assert_eq!( + information, LEN as usize, + "a resolver with may_transfer_partially off reported a short \ + transfer; seed {seed:#x}" + ); + } + } +} + +#[test] +fn the_configuration_has_a_switch_for_every_permission_and_none_for_any_constraint() { + // The asymmetry the module documents, asserted rather than described. A + // switch that relaxed an `RS-C-n` would let a test quietly assert against + // a platform that cannot exist, so the count is pinned: eight permissions + // in RESPONSE-SPACE.md, eight fields, and no ninth for a constraint. + // + // Counted through `Debug`, which lists exactly the struct's fields, so a + // field added without a clause fails here rather than passing unnoticed. + let rendered = format!("{:?}", ResolverConfig::default()); + let fields = rendered.matches(": ").count(); + assert_eq!( + fields, 8, + "ResolverConfig has {fields} fields; RESPONSE-SPACE.md specifies eight permissions \ + (RS-P-1..8), and a constraint must never become a switch -- {rendered}" + ); +} diff --git a/crates/windows-ioring-sys/src/token.rs b/crates/windows-ioring-sys/src/token.rs index 8487880b4..afe0d213f 100644 --- a/crates/windows-ioring-sys/src/token.rs +++ b/crates/windows-ioring-sys/src/token.rs @@ -29,17 +29,25 @@ pub struct Token { } impl Token { - /// Wrap `value` under a fresh identity reserved from `ring`. + /// Wrap `value` under a fresh identity reserved from `accounting`. + /// + /// Takes the ring's **ledger** rather than the ring (M24.7). Minting + /// needs an identity and the ring's id, both of which are bookkeeping -- + /// no handle is involved -- and narrowing the parameter to what is + /// actually used is what lets this be exercised without opening a ring. /// /// # Errors /// - /// Returns any error from [`crate::IoRing::reserve_user_data`] (in - /// practice, only if the identity space is exhausted). - pub(crate) fn new(ring: &mut crate::IoRing, value: T) -> std::io::Result { - let id = ring.reserve_user_data()?; + /// Returns any error from [`crate::accounting::Accounting::reserve_user_data`] + /// (in practice, only if the identity space is exhausted). + pub(crate) fn new( + accounting: &mut crate::accounting::Accounting, + value: T, + ) -> std::io::Result { + let id = accounting.reserve_user_data()?; Ok(Self { id, - ring_id: ring.ring_id(), + ring_id: accounting.ring_id(), value: ManuallyDrop::new(value), }) } diff --git a/crates/windows-ioring-sys/src/token/tests.rs b/crates/windows-ioring-sys/src/token/tests.rs index 7f7430ba4..884e7e0a0 100644 --- a/crates/windows-ioring-sys/src/token/tests.rs +++ b/crates/windows-ioring-sys/src/token/tests.rs @@ -3,7 +3,7 @@ use std::sync::Arc; use std::sync::atomic::{AtomicBool, Ordering}; use super::Token; -use crate::IoRing; +use crate::accounting::Accounting; use crate::buf::IoBuf; use crate::ring::Completion; @@ -45,23 +45,21 @@ fn tracked_buffer() -> (DropTracking, Arc) { ) } -/// None of these tests ever actually submits anything to `ring` (M3 has not -/// been built yet), so a token minted here never gets a real completion. -/// `IoRing::run_down` -- which its `Drop` calls -- waits for exactly that, so -/// every test must call `record_completion` once per token it minted before -/// letting `ring` drop, or teardown would hang waiting for a completion that -/// will never arrive. -fn settle(ring: &mut IoRing) { - while ring.outstanding() > 0 { - ring.record_completion(); - } -} - +/// Until M24.7 these tests each opened a real `IoRing`, and the ring was +/// purely a liability. None of them ever submitted anything, so a token minted +/// here never got a real completion -- and `IoRing::run_down`, which `Drop` +/// calls, waits for exactly that. Every test therefore had to call +/// `record_completion` once per token through a `settle` helper, or teardown +/// would hang waiting for a completion that was never coming. +/// +/// `Token::new` now takes the ring's ledger rather than the ring, because an +/// identity and a ring id are all it ever needed. The helper is gone with the +/// hazard it existed to work around, and these tests open nothing. #[test] fn dropping_an_unclaimed_token_never_runs_the_buffers_destructor() { - let mut ring = IoRing::new(64, 128).expect("create ring"); + let mut ledger = Accounting::new(); let (buffer, dropped) = tracked_buffer(); - let token = Token::new(&mut ring, buffer).expect("mint token"); + let token = Token::new(&mut ledger, buffer).expect("mint token"); drop(token); @@ -71,8 +69,6 @@ fn dropping_an_unclaimed_token_never_runs_the_buffers_destructor() { a real IoRing may still be writing through it" ); - settle(&mut ring); - drop(ring); // The leak is real and permanent: nothing later runs the destructor // either, including the ring's own teardown. assert!(!dropped.load(Ordering::SeqCst)); @@ -80,12 +76,12 @@ fn dropping_an_unclaimed_token_never_runs_the_buffers_destructor() { #[test] fn claiming_a_token_returns_the_buffer_for_normal_disposal() { - let mut ring = IoRing::new(64, 128).expect("create ring"); + let mut ledger = Accounting::new(); let (buffer, dropped) = tracked_buffer(); - let token = Token::new(&mut ring, buffer).expect("mint token"); + let token = Token::new(&mut ledger, buffer).expect("mint token"); let id = token.id(); - let completion = Completion::synthetic(id, 0, ring.ring_id()); + let completion = Completion::synthetic(id, 0, ledger.ring_id()); let claimed = token.claim_if(&completion).expect("id matches itself"); assert!( !dropped.load(Ordering::SeqCst), @@ -97,18 +93,16 @@ fn claiming_a_token_returns_the_buffer_for_normal_disposal() { dropped.load(Ordering::SeqCst), "the caller's own drop of the returned buffer must run normally" ); - - settle(&mut ring); } #[test] fn claim_if_rejects_a_mismatched_user_data_and_returns_the_token_unchanged() { - let mut ring = IoRing::new(64, 128).expect("create ring"); + let mut ledger = Accounting::new(); let (buffer, dropped) = tracked_buffer(); - let token = Token::new(&mut ring, buffer).expect("mint token"); + let token = Token::new(&mut ledger, buffer).expect("mint token"); let real_id = token.id(); - let mismatched = Completion::synthetic(real_id.wrapping_add(1), 0, ring.ring_id()); + let mismatched = Completion::synthetic(real_id.wrapping_add(1), 0, ledger.ring_id()); let token = token .claim_if(&mismatched) .expect_err("a stale id must not claim this token"); @@ -120,77 +114,71 @@ fn claim_if_rejects_a_mismatched_user_data_and_returns_the_token_unchanged() { assert!(!dropped.load(Ordering::SeqCst)); // It can still be claimed correctly afterwards. - let matching = Completion::synthetic(real_id, 0, ring.ring_id()); + let matching = Completion::synthetic(real_id, 0, ledger.ring_id()); let claimed = token.claim_if(&matching).expect("the real id still works"); drop(claimed); assert!(dropped.load(Ordering::SeqCst)); - - settle(&mut ring); } #[test] fn each_token_on_a_ring_gets_a_distinct_id() { - let mut ring = IoRing::new(64, 128).expect("create ring"); + let mut ledger = Accounting::new(); let (a, _) = tracked_buffer(); let (b, _) = tracked_buffer(); - let token_a = Token::new(&mut ring, a).expect("mint token a"); - let token_b = Token::new(&mut ring, b).expect("mint token b"); + let token_a = Token::new(&mut ledger, a).expect("mint token a"); + let token_b = Token::new(&mut ledger, b).expect("mint token b"); assert_ne!(token_a.id(), token_b.id()); drop(token_a); drop(token_b); - settle(&mut ring); } #[test] -fn minting_a_token_increments_the_rings_outstanding_count() { - let mut ring = IoRing::new(64, 128).expect("create ring"); - assert_eq!(ring.outstanding(), 0); +fn minting_a_token_increments_the_ledgers_outstanding_count() { + let mut ledger = Accounting::new(); + assert_eq!(ledger.outstanding(), 0); let (buffer, _dropped) = tracked_buffer(); - let token = Token::new(&mut ring, buffer).expect("mint token"); - assert_eq!(ring.outstanding(), 1); + let token = Token::new(&mut ledger, buffer).expect("mint token"); + assert_eq!(ledger.outstanding(), 1); - // Dropping the token does not, by itself, tell the ring the operation is + // Dropping the token does not, by itself, tell the ledger the operation is // done -- only observing a real completion does (M2.4); this token was // never actually submitted to anything, so nothing ever will. drop(token); assert_eq!( - ring.outstanding(), + ledger.outstanding(), 1, "outstanding tracks completions observed, not tokens dropped" ); - ring.record_completion(); - assert_eq!(ring.outstanding(), 0); + ledger.record_completion(); + assert_eq!(ledger.outstanding(), 0); } #[test] fn claim_if_rejects_a_matching_user_data_from_a_different_ring() { - let mut ring_a = IoRing::new(64, 128).expect("create ring a"); - let mut ring_b = IoRing::new(64, 128).expect("create ring b"); + let mut ledger_a = Accounting::new(); + let ledger_b = Accounting::new(); let (buffer, dropped) = tracked_buffer(); - let token = Token::new(&mut ring_a, buffer).expect("mint token on ring a"); + let token = Token::new(&mut ledger_a, buffer).expect("mint token on ring a"); let id = token.id(); // Both rings mint `UserData` from their own counter starting at zero, so // this is a real coincidence a naive `id`-only check would miss (PR #20 // review response): the completion carries the *same* `UserData` value // as `token`, but from `ring_b`, not `ring_a`. - let wrong_ring = Completion::synthetic(id, 0, ring_b.ring_id()); + let wrong_ring = Completion::synthetic(id, 0, ledger_b.ring_id()); let token = token .claim_if(&wrong_ring) .expect_err("a completion from a different ring must not claim this token"); assert!(!dropped.load(Ordering::SeqCst)); - let right_ring = Completion::synthetic(id, 0, ring_a.ring_id()); + let right_ring = Completion::synthetic(id, 0, ledger_a.ring_id()); let claimed = token .claim_if(&right_ring) .expect("the same ring's completion still claims it"); drop(claimed); assert!(dropped.load(Ordering::SeqCst)); - - settle(&mut ring_a); - settle(&mut ring_b); } #[test] @@ -199,7 +187,7 @@ fn a_tokens_debug_names_its_operation() { // would demand `T: Debug` from a caller's buffer type, and the id is the // only part worth printing (D-4). M18.3 showed nothing asserted either // half of that choice: the impl could return an empty string unnoticed. - let mut ring = IoRing::new(8, 8).expect("create ring"); + let mut ledger = Accounting::new(); // A buffer type that is deliberately *not* `Debug`, which is the constraint // that forced the hand-written impl in the first place. @@ -215,7 +203,7 @@ fn a_tokens_debug_names_its_operation() { } } - let token = Token::new(&mut ring, NotDebug(vec![0_u8; 8])).expect("mint a token"); + let token = Token::new(&mut ledger, NotDebug(vec![0_u8; 8])).expect("mint a token"); let id = token.id(); let rendered = format!("{token:?}"); @@ -229,5 +217,4 @@ fn a_tokens_debug_names_its_operation() { ); drop(token); - settle(&mut ring); } diff --git a/crates/windows-ioring-sys/tests/bounded_pop.rs b/crates/windows-ioring-sys/tests/bounded_pop.rs new file mode 100644 index 000000000..1884899aa --- /dev/null +++ b/crates/windows-ioring-sys/tests/bounded_pop.rs @@ -0,0 +1,363 @@ +// Copyright (c) 2026 Mike Grier +//! The bounded pop against a genuinely pending operation (M21.6, M22+.1). +//! +//! # Why an integration test, and why a pipe +//! +//! No other test of [`IoRing::pop_within`] has a **real operation still pending +//! when a real bound expires**, and that is the exact state the Win32 timeout +//! path needs. The coverage either side of it misses in opposite directions: +//! the tests that use the ring's own `SubmitWait` drive operations that +//! *complete*, so the expiry branch is never reached, and the test that does +//! let a bound expire fakes both halves -- a `RecordingWait` instead of the +//! kernel, and a bare `reserve_user_data` instead of a pending operation. +//! +//! (An earlier version of this comment said every other test "drives the loop +//! with a wait that never enters the kernel". That is not so -- several reach +//! the kernel through `SubmitWait` -- and the imprecision mattered, because it +//! named the wrong gap. Corrected by `M24.6`'s sweep.) +//! +//! That gap hid a real defect: a +//! timed-out `SubmitIoRing` reports `ERROR_TIMEOUT`, so the crate's own +//! `SubmitWait` turned every expired bound into an `Err` instead of the +//! `Ok(None)` `pop_within` documents. Replacing `RingWait::block`'s whole body +//! with an unconditional error left the entire suite green. +//! +//! Closing that needs an operation that is **still pending** when the bound +//! expires. The first version of this file got one by reading 128 MiB +//! unbuffered, which is a *margin* rather than a guarantee -- and it failed once, +//! unreproducibly, for what was most likely that reason. What is used instead is +//! an operation that **cannot** complete until the test says so: a read on an +//! overlapped named pipe that nobody has written to. +//! +//! The difference is the whole point. A slow read asks "will the device take +//! longer than 5 ms?", which is a question about someone else's hardware. A +//! pipe with no writer asks nothing: the read is pending because no byte exists +//! to satisfy it, and it completes exactly when this file writes one. +//! +//! Three earlier attempts, each measured before being discarded: +//! +//! | Attempt | Result | +//! |---|---| +//! | Buffered read, up to 256 MiB | 3-5 us -- the cache answers | +//! | Flush over 512 MiB of dirty cache | 3 us -- the lazy writer got there first | +//! | Unbuffered read, 256 MiB, no `FILE_FLAG_OVERLAPPED` | 3 us -- a synchronous handle completes inline during submit | +//! +//! That third row is worth keeping: **a handle opened without +//! `FILE_FLAG_OVERLAPPED` is synchronous**, so a ring operation against it does +//! not pend at all. The pipe below is opened overlapped for exactly that reason. +//! +//! # What is and is not asserted about time +//! +//! Nothing here asserts that an operation finished within some bound. Where a +//! delay is needed -- `run_down` polls in 50 ms steps, so forcing it to see an +//! expired poll needs the release to come later than that -- it comes from a +//! `thread::sleep`, whose guarantee runs the safe way: a sleep may overshoot, +//! never undershoot. Windows' default timer resolution is about 15.6 ms, so a +//! 5 ms bound routinely takes 14-19 ms to expire; that is fine, because no +//! assertion depends on it. + +#![cfg(windows)] + +use std::io::Write; +use std::os::windows::io::{AsRawHandle, FromRawHandle, OwnedHandle}; +use std::time::Duration; + +use windows_ioring_sys::{Batch, CompletionWait, IoRing, PushOptions, RingWait, SubmitWait, Token}; +use windows_sys::Win32::Foundation::GENERIC_WRITE; +use windows_sys::Win32::Storage::FileSystem::{ + CreateFileW, FILE_ATTRIBUTE_NORMAL, FILE_FLAG_OVERLAPPED, FILE_SHARE_READ, FILE_SHARE_WRITE, + OPEN_EXISTING, +}; +use windows_sys::Win32::System::Pipes::CreateNamedPipeW; + +/// `PIPE_ACCESS_INBOUND`. The byte-type, byte-read-mode and wait bits are all +/// zero, so one named constant covers the mode argument. +const PIPE_ACCESS_INBOUND: u32 = 0x0000_0001; +const PIPE_MODE_BYTE_WAIT: u32 = 0; + +/// The bound these tests give the pop. Any value would do: the read cannot +/// complete until the test writes, so this only decides how long the wait +/// blocks before reporting that it expired. +const SHORT_BOUND: Duration = Duration::from_millis(5); + +/// Comfortably longer than `run_down`'s 50 ms poll, so a rundown is guaranteed +/// to see at least one expired poll before the read is released. A sleep may +/// overshoot and never undershoot, which is the direction that keeps this +/// deterministic. +const LONGER_THAN_A_RUNDOWN_POLL: Duration = Duration::from_millis(200); + +fn wide(value: &str) -> Vec { + value.encode_utf16().chain(std::iter::once(0)).collect() +} + +/// An overlapped named pipe: a read on `server` stays pending until something +/// is written to `client`. +struct Pipe { + server: OwnedHandle, + client: Option, +} + +impl Pipe { + fn new(tag: &str) -> Self { + let name = format!( + r"\\.\pipe\windows-ioring-sys-bounded-pop-{tag}-{}", + std::process::id() + ); + let wide_name = wide(&name); + + // SAFETY: a valid NUL-terminated name and standard parameters; the + // returned handle is owned here. + let server = unsafe { + CreateNamedPipeW( + wide_name.as_ptr(), + PIPE_ACCESS_INBOUND | FILE_FLAG_OVERLAPPED, + PIPE_MODE_BYTE_WAIT, + 1, + 4096, + 4096, + 0, + std::ptr::null(), + ) + }; + assert!(!server.is_null(), "CreateNamedPipeW failed"); + // SAFETY: a fresh handle nothing else owns. + let server = unsafe { OwnedHandle::from_raw_handle(server) }; + + // SAFETY: the pipe exists; this opens its client end. + let client = unsafe { + CreateFileW( + wide_name.as_ptr(), + GENERIC_WRITE, + FILE_SHARE_READ | FILE_SHARE_WRITE, + std::ptr::null(), + OPEN_EXISTING, + FILE_ATTRIBUTE_NORMAL, + std::ptr::null_mut(), + ) + }; + assert!( + !client.is_null() && client as isize != -1, + "opening the pipe's client end failed" + ); + // SAFETY: a fresh handle nothing else owns. + let client = unsafe { OwnedHandle::from_raw_handle(client) }; + + Self { + server, + client: Some(std::fs::File::from(client)), + } + } + + /// Queue a read that cannot complete, and submit it. + fn push_pending_read(&self, ring: &mut IoRing) -> Token> { + let mut batch = Batch::new(ring); + // SAFETY: the server handle outlives the operation -- every test + // releases and drains before dropping this `Pipe` -- and the token is + // returned to the caller, which holds it until the completion is + // claimed. + let token = unsafe { + batch.read_raw( + self.server.as_raw_handle(), + vec![0_u8; 64], + 0, + PushOptions::new(), + ) + } + .expect("queue a read on the pipe"); + batch.submit().expect("submit the read"); + token + } + + /// Satisfy the pending read, now. + fn release(&mut self) { + let mut client = self.client.take().expect("the client end is still open"); + client.write_all(b"x").expect("write to the pipe"); + client.flush().expect("flush the pipe"); + } + + /// Satisfy the pending read after `delay`, from another thread. + /// + /// Used where the test needs the operation to outlive something -- a + /// `run_down` poll -- rather than merely to be pending. + fn release_after(&mut self, delay: Duration) -> std::thread::JoinHandle<()> { + let mut client = self.client.take().expect("the client end is still open"); + std::thread::spawn(move || { + std::thread::sleep(delay); + let _ = client.write_all(b"x"); + let _ = client.flush(); + }) + } +} + +/// Release the read and drain it, so the ring is quiet before it drops. +fn settle(ring: &mut IoRing, pipe: &mut Pipe, token: Token>) { + pipe.release(); + let completion = ring + .pop_within(Duration::from_secs(30)) + .expect("pop_within") + .expect("the read completes once the pipe has a byte in it"); + let bytes = completion.result().expect("the read succeeded"); + assert_eq!(bytes, 1, "exactly the byte that was written"); + let _ = token + .claim_if(&completion) + .expect("the token claims its own"); +} + +#[test] +fn a_bound_that_expires_reports_no_completion_rather_than_an_error() { + // The regression test for the defect this file exists for. Before the fix + // this returned `Err(HRESULT 0x800705B4)`. + let mut ring = IoRing::new(16, 16).expect("create a ring"); + let mut pipe = Pipe::new("expires"); + let token = pipe.push_pending_read(&mut ring); + + let popped = ring + .pop_within(SHORT_BOUND) + .expect("an expired bound is not an error"); + assert!( + popped.is_none(), + "nothing has been written to the pipe, so the read cannot have completed" + ); + assert!(ring.outstanding() > 0, "and the read must still be pending"); + + settle(&mut ring, &mut pipe, token); +} + +/// Forwards to the ring's own wait and records what it was asked and what it +/// answered. +#[derive(Default)] +struct CountingSubmitWait { + calls: usize, + last_result_was_ok: bool, +} + +impl CompletionWait for CountingSubmitWait { + fn wait(&mut self, ring: &mut RingWait<'_>, timeout_ms: u32) -> std::io::Result<()> { + self.calls += 1; + let outcome = ring.block(timeout_ms); + self.last_result_was_ok = outcome.is_ok(); + outcome + } +} + +#[test] +fn the_rings_own_wait_is_reached_and_reports_an_expired_bound_as_success() { + // Kills the mutation that started this: replacing `RingWait::block`'s body + // with an unconditional error used to leave every test green, because no + // test ever reached it. + let mut ring = IoRing::new(16, 16).expect("create a ring"); + let mut pipe = Pipe::new("reached"); + let token = pipe.push_pending_read(&mut ring); + + let mut wait = CountingSubmitWait::default(); + let popped = ring + .pop_within_with(&mut wait, SHORT_BOUND) + .expect("an expired bound is not an error"); + + assert!(popped.is_none()); + assert!( + wait.calls >= 1, + "the ring's own wait must actually be entered" + ); + assert!( + wait.last_result_was_ok, + "an expired `SubmitIoRing` wait must be reported as success, not as \ + ERROR_TIMEOUT -- this is the defect the whole file exists for" + ); + + settle(&mut ring, &mut pipe, token); +} + +#[test] +fn the_default_wait_is_the_ring_wait() { + // `pop_within` and `pop_within_with(&mut SubmitWait, ..)` must agree, so the + // convenience cannot quietly diverge from the documented default. + let mut ring = IoRing::new(16, 16).expect("create a ring"); + let mut pipe = Pipe::new("default"); + let token = pipe.push_pending_read(&mut ring); + + let popped = ring + .pop_within_with(&mut SubmitWait, SHORT_BOUND) + .expect("an expired bound is not an error"); + assert!(popped.is_none()); + assert!(ring.outstanding() > 0); + + settle(&mut ring, &mut pipe, token); +} + +#[test] +fn run_down_tolerates_an_operation_slower_than_its_poll() { + // `run_down` polls in 50 ms steps. Treating an expired poll as fatal made it + // return `Err` with the operation still outstanding, so `Drop` asserted and + // then closed the ring anyway -- the exact hazard it exists to prevent. The + // release is deliberately later than one poll, so at least one poll is + // guaranteed to expire before the read completes. + let mut ring = IoRing::new(16, 16).expect("create a ring"); + let mut pipe = Pipe::new("rundown"); + let token = pipe.push_pending_read(&mut ring); + let writer = pipe.release_after(LONGER_THAN_A_RUNDOWN_POLL); + + ring.run_down() + .expect("a poll that expires is not a rundown failure"); + assert_eq!( + ring.outstanding(), + 0, + "rundown returns only once nothing is outstanding" + ); + + writer.join().expect("the writer thread"); + // The rundown reaped the completion, so the token has nothing left to + // claim; dropping it here is the honest end of its life. + drop(token); +} + +#[test] +fn dropping_a_ring_with_a_pending_operation_does_not_panic() { + // The same defect seen from where it actually bit: `Drop` runs the rundown, + // and in a debug build a failed rundown is a `debug_assert`. + let mut ring = IoRing::new(16, 16).expect("create a ring"); + let mut pipe = Pipe::new("drop"); + let token = pipe.push_pending_read(&mut ring); + let writer = pipe.release_after(LONGER_THAN_A_RUNDOWN_POLL); + + drop(token); + drop(ring); + + writer.join().expect("the writer thread"); +} + +// ------------------------------------------------------------------------ +// Relocated from `src/ring/tests.rs` at 9bc0350e (M24.3). Pure relocation; +// these reach nothing crate-private, which is what let them move. + +#[test] +fn pop_within_returns_none_at_once_when_nothing_can_ever_arrive() { + // With no operation outstanding no completion is possible, so the loop + // answers before consulting any wait at all. + // + // **The `Ok(None)` is itself the proof, with no clock involved.** This goes + // through the convenience, so the wait is `SubmitWait` -- and + // `SubmitIoRing` answers `E_INVALIDARG` when asked to wait for a completion + // the kernel has no pending operation for. Were the early return removed, + // the loop would reach that wait and this would be an `Err`. Asserting + // `Ok(None)` therefore distinguishes "returned early" from "waited", which + // is exactly what an elapsed-time assertion was being asked to do -- and + // unlike a clock, it cannot be wrong because the machine was busy. + let mut ring = IoRing::new(16, 16).expect("create ring"); + let popped = ring + .pop_within(std::time::Duration::from_secs(30)) + .expect("the early return means the ring's own wait is never reached"); + assert!(popped.is_none(), "nothing was ever submitted"); +} + +#[test] +fn a_bound_the_clock_cannot_represent_does_not_panic() { + // `Instant + Duration` panics on overflow, and `Duration::MAX` is a + // reasonable spelling of "no deadline". Nothing is outstanding here, so + // the early return answers before any deadline arithmetic matters. + let mut ring = IoRing::new(16, 16).expect("create ring"); + let popped = ring + .pop_within(std::time::Duration::MAX) + .expect("pop_within"); + assert!(popped.is_none()); +} diff --git a/crates/windows-ioring-sys/tests/calibration.rs b/crates/windows-ioring-sys/tests/calibration.rs new file mode 100644 index 000000000..d7b7d9433 --- /dev/null +++ b/crates/windows-ioring-sys/tests/calibration.rs @@ -0,0 +1,421 @@ +// Copyright (c) Mike Grier +//! Calibration: showing the resolver can go red (M26.5). +//! +//! # Why this file exists +//! +//! [D-41](../DESIGN-NOTES.md#d-41)'s corollary is the rule it enforces: **a +//! green result from an instrument nobody has shown can go red is not +//! evidence**. `M26.3` built a resolver and `M26.4` stated five properties +//! over it; both currently report green, and neither has yet been shown +//! capable of reporting anything else against a defect that really happened. +//! +//! That is not a theoretical worry in this repository. `M17.4` calibrated the +//! generated-sequence suite by reverting `D-20`'s setup signal -- issue #47 +//! exactly as it shipped -- and the suite reported **green**. It sampled the +//! right states and drained by polling, which recovers every completion +//! whether or not the ring ever signalled. Sampling the right state is not the +//! same as being sensitive to the defect that lives in it, and the difference +//! was invisible until somebody tried it. The design session behind `M26` +//! produced two more instruments of the same shape and said so. +//! +//! # The two defects, and why they are calibrated differently +//! +//! **[D-47](../DESIGN-NOTES.md#d-47): a consumer that believes a covering +//! flush holds back what follows it.** [D-24](../DESIGN-NOTES.md#d-24) claimed +//! that and `D-47` withdrew it over roughly 4,500 trials. The defect is in a +//! *consumer*, not in this crate, so there is nothing here to mutate -- the +//! defective consumer has to be written down, and it is, below. The test +//! asserts that the resolver **breaks** it. +//! +//! **`M21.6`: an expired wait treated as a failure.** That one was in this +//! crate: `pop_within` returned `Err` on every ordinary timeout until +//! `wait_outcome` was given its `IORING_E_WAIT_TIMEOUT` arm. Re-injecting it means +//! mutating the crate, which a test cannot do, so it lives in +//! [sabotage.json](../sabotage.json) instead and is swept with everything +//! else. What *this* file contributes is the precondition that sabotage needs +//! to mean anything: that `RS-P-4` actually reaches `pop_within` at all. A +//! sabotage of code the suite never executes would be caught for some +//! unrelated reason, or not at all, and either way would say nothing. +//! +//! # These tests fail if the resolver gets *narrower* +//! +//! That is their whole purpose and it is worth being explicit, because it +//! inverts the usual reading. A failure here does not mean the crate broke; it +//! means the instrument stopped being able to see something it must see. + +#![cfg(all(windows, feature = "kernel-seam"))] + +use std::time::Duration; + +use windows_ioring_sys::sys::{Resolver, ResolverConfig}; +use windows_ioring_sys::{ + Batch, Completion, FlushCoverage, FlushMode, IoRing, PushOptions, SharedFile, Token, + WriteCaching, +}; + +/// Seeds each calibration sweeps. +/// +/// Fixed rather than clock-derived, unlike the generated suites: a calibration +/// that sometimes could not demonstrate its own sensitivity would be the exact +/// failure it exists to prevent, arriving as a flake. +const SEEDS: std::ops::Range = 0..2048; +const SEED_COUNT: usize = 2048; + +/// Bytes per write. +const BUF_LEN: usize = 256; + +/// How many consultations one sweep may take to drain. +const BUDGET: usize = 512; + +fn scratch(tag: &str) -> (SharedFile, std::path::PathBuf) { + let path = std::env::temp_dir().join(format!( + "windows-ioring-sys-m26-5-{}-{tag}.tmp", + std::process::id() + )); + let file = std::fs::OpenOptions::new() + .create(true) + .truncate(true) + .write(true) + .open(&path) + .expect("a scratch file"); + (SharedFile::new(file.into()), path) +} + +/// Tokens awaiting completion, claimed by identity as completions arrive. +#[derive(Default)] +struct Held { + flushes: Vec>, + writes: Vec, SharedFile)>>, +} + +impl Held { + /// Claim whichever token this completion belongs to. + /// + /// Returns false if none matches, which would itself be `RS-C-2` broken -- + /// asserted by the caller rather than ignored, so a calibration cannot + /// pass by losing track of an operation. + fn claim(&mut self, completion: &Completion) -> bool { + let user_data = completion.user_data(); + if let Some(at) = self + .flushes + .iter() + .position(|token| token.id() == user_data) + { + return self.flushes.remove(at).claim_if(completion).is_ok(); + } + if let Some(at) = self.writes.iter().position(|token| token.id() == user_data) { + return self.writes.remove(at).claim_if(completion).is_ok(); + } + false + } +} + +/// Drain the ring, returning completion identities in the order they arrived. +fn drain_in_order(ring: &mut IoRing, held: &mut Held, expected: usize) -> Vec { + let mut order = Vec::new(); + for _ in 0..BUDGET { + if order.len() == expected { + return order; + } + // A declined submit under RS-P-7 surfaces here as an error, which is + // within `pop_within`'s documented contract and is retried rather than + // treated as a failure -- see M26.4, which established that. + match ring.pop_within(Duration::from_millis(5)) { + Ok(Some(completion)) => { + assert!( + held.claim(&completion), + "a completion arrived for {:#x} with no live token to match it", + completion.user_data() + ); + order.push(completion.user_data()); + } + Ok(None) | Err(_) => continue, + } + } + panic!( + "drain did not finish: {} of {expected} completions within the budget", + order.len() + ); +} + +// ------------------------------------------------------ D-47's defect ------ + +/// The consumer `D-24` licensed and `D-47` withdrew, written out so the +/// resolver has something to break. +/// +/// The belief is stated the way a consumer would actually hold it: a covering +/// flush closes an epoch, so any operation queued *after* that flush belongs +/// to the next epoch and cannot be seen until the flush itself has been. A +/// real consumer with this belief does something consequential on it -- rolls +/// an epoch counter, reuses an arena, reports a record durable -- and this one +/// only records that the belief failed, because the failure is what is being +/// measured. +#[derive(Debug, Default)] +struct HoldBackBeliever { + flush_seen: bool, + /// Set when an operation queued after the covering flush arrived first. + broken_by: Option, + /// Set if an operation queued *before* the flush arrived after it, which + /// would be `RS-C-4` violated -- a different failure entirely, and one + /// this test must not mistake for the one it is looking for. + drain_half_broken_by: Option, +} + +impl HoldBackBeliever { + fn observe(&mut self, user_data: usize, flush: usize, before: &[usize], after: &[usize]) { + if user_data == flush { + self.flush_seen = true; + return; + } + if after.contains(&user_data) && !self.flush_seen { + self.broken_by.get_or_insert(user_data); + } + if before.contains(&user_data) && self.flush_seen { + self.drain_half_broken_by.get_or_insert(user_data); + } + } +} + +#[test] +fn the_resolver_breaks_a_consumer_that_believes_the_drain_flag_holds_back() { + // D-47's defect, re-injected. The claim being calibrated is that the + // resolver can *see* this class of error -- so the assertion is that some + // seed breaks the believer, and a run where none did would mean the + // instrument had gone narrow. + let (file, path) = scratch("holdback"); + let mut broken = 0_usize; + let mut drain_half_failures = Vec::new(); + + for seed in SEEDS { + // Narrowed to the one freedom under test. Operation failures and + // declined submits would still leave the sweep correct, but they would + // make a failure here ambiguous about which freedom caused it, and a + // calibration that cannot say what it detected is not calibrated. + let resolver = Resolver::with_config( + seed, + ResolverConfig { + may_reorder: true, + may_pend: true, + ..ResolverConfig::narrowest() + }, + ); + let outcome = resolver.scoped(|_| { + let mut ring = IoRing::new(64, 128).expect("a ring"); + let mut held = Held::default(); + let (before, flush, after) = { + let mut batch = Batch::new(&mut ring); + let mut before = Vec::new(); + for _ in 0..2 { + let token = batch + .write( + &file, + vec![0_u8; BUF_LEN], + 0, + PushOptions::new(), + WriteCaching::Cached, + ) + .expect("a write builds"); + before.push(token.id()); + held.writes.push(token); + } + let flush_token = batch + .flush( + &file, + FlushCoverage::CoversPrecedingOperations, + FlushMode::Default, + ) + .expect("a covering flush builds"); + let flush = flush_token.id(); + held.flushes.push(flush_token); + + let mut after = Vec::new(); + for _ in 0..3 { + let token = batch + .write( + &file, + vec![0_u8; BUF_LEN], + 0, + PushOptions::new(), + WriteCaching::Cached, + ) + .expect("a write builds"); + after.push(token.id()); + held.writes.push(token); + } + batch.submit().expect("the submit is answered"); + (before, flush, after) + }; + + let order = drain_in_order(&mut ring, &mut held, 6); + let mut believer = HoldBackBeliever::default(); + for user_data in &order { + believer.observe(*user_data, flush, &before, &after); + } + ring.run_down().expect("rundown"); + believer + }); + + if outcome.broken_by.is_some() { + broken += 1; + } + if let Some(user_data) = outcome.drain_half_broken_by { + drain_half_failures.push((seed, user_data)); + } + } + let _ = std::fs::remove_file(path); + + // The half that must NOT break. RS-C-4 forbids it, and a resolver that + // broke both halves would make the assertion below pass for the wrong + // reason -- the believer would be "broken" by a violation of the + // guarantee this crate's durability story rests on, rather than by the + // one-sidedness D-47 measured. + assert!( + drain_half_failures.is_empty(), + "RS-C-4 was violated, so this calibration is measuring the wrong failure: {:?}", + drain_half_failures + ); + + assert!( + broken > 0, + "no seed of {SEED_COUNT} broke a consumer that assumes a covering flush holds back what \ + follows it. D-47 withdrew that assumption over roughly 4,500 trials, so a resolver \ + that cannot exhibit the one-sidedness has gone narrower than the space and would \ + report green on the defect class it exists to catch" + ); + // Printed rather than only asserted: the threshold is "more than none", + // which is the honest bar for a demonstration of sensitivity, but a reader + // deciding how much to trust this instrument wants to know whether that + // was one seed or most of them. + // + // Expect this to be far higher than the rate D-47 measured on real + // hardware, and note that the difference is not the resolver being + // unfaithful. RESPONSE-SPACE.md deliberately carries no rates: a resolver + // reproducing an observed frequency would be a model of Windows, which is + // the trap D-52 was opened to escape. Its job is to explore the space, and + // a defect class that occurs in under one trial in a hundred on hardware + // is precisely the one a rate-free resolver earns its keep on. + eprintln!("D-47 calibration: {broken} of {SEED_COUNT} seeds broke the believer"); +} + +// ------------------------------------------------------ M21.6's defect ----- + +#[test] +fn an_expired_wait_reaches_pop_within() { + // The precondition `M21.6`'s sabotage needs. That case reverts + // `wait_outcome`'s IORING_E_WAIT_TIMEOUT arm, which only means something if an + // expired wait actually arrives there -- a mutation to code the suite + // never executes is caught for an unrelated reason or not at all, and + // either way measures nothing. + // + // Asserted through the resolver's own counter rather than inferred from a + // timing, because "the call took a while" is not evidence that a wait + // expired. + let (file, path) = scratch("expired"); + let mut seeds_with_an_expired_wait = 0_usize; + + for seed in SEEDS { + let resolver = Resolver::with_config( + seed, + ResolverConfig { + may_expire_waits: true, + may_pend: true, + ..ResolverConfig::narrowest() + }, + ); + let expired = resolver.scoped(|watch| { + let mut ring = IoRing::new(64, 128).expect("a ring"); + let mut held = Held::default(); + { + let mut batch = Batch::new(&mut ring); + for _ in 0..4 { + let token = batch + .flush(&file, FlushCoverage::Unordered, FlushMode::Default) + .expect("a flush builds"); + held.flushes.push(token); + } + batch.submit().expect("the submit is answered"); + } + // `pop_within` is the only caller here, so any expired wait the + // resolver records was answered to it. + drain_in_order(&mut ring, &mut held, 4); + ring.run_down().expect("rundown"); + watch.stats().expired_waits + }); + if expired > 0 { + seeds_with_an_expired_wait += 1; + } + } + let _ = std::fs::remove_file(path); + + assert!( + seeds_with_an_expired_wait > SEED_COUNT / 4, + "only {seeds_with_an_expired_wait} of {SEED_COUNT} seeds expired a wait inside \ + pop_within, which is too few for the M21.6 sabotage to be reliably exercised -- that \ + case would then be caught by accident or not at all" + ); + eprintln!( + "M21.6 calibration: {seeds_with_an_expired_wait} of {SEED_COUNT} seeds expired a wait \ + inside pop_within" + ); +} + +#[test] +fn an_expired_wait_is_not_reported_as_a_failure() { + // `M21.6`'s correction, stated as the property it is rather than left to + // the sabotage alone. An expired wait is the ordinary outcome of asking to + // block for a bounded time; `pop_within` documents `Ok(None)` for it, and + // returning `Err` instead is what the defect did. + // + // This is the assertion the sabotage turns red, so the two are a pair: the + // sabotage shows the assertion is load-bearing, and the assertion is what + // gives the sabotage something to break. + let (file, path) = scratch("notafailure"); + for seed in SEEDS { + let resolver = Resolver::with_config( + seed, + ResolverConfig { + may_expire_waits: true, + may_pend: true, + ..ResolverConfig::narrowest() + }, + ); + let replay = resolver.replay_hint(); + resolver.scoped(|watch| { + let mut ring = IoRing::new(64, 128).expect("a ring"); + let mut held = Held::default(); + { + let mut batch = Batch::new(&mut ring); + for _ in 0..3 { + held.flushes.push( + batch + .flush(&file, FlushCoverage::Unordered, FlushMode::Default) + .expect("a flush builds"), + ); + } + batch.submit().expect("the submit is answered"); + } + + let mut seen = 0_usize; + for _ in 0..BUDGET { + if seen == 3 { + break; + } + match ring.pop_within(Duration::from_millis(5)) { + Ok(Some(completion)) => { + assert!(held.claim(&completion)); + seen += 1; + } + Ok(None) => {} + Err(error) => panic!( + "{replay}: pop_within reported an error after {} expired wait(s): \ + {error}. An expired wait is not a failure of the ring (M21.6)", + watch.stats().expired_waits + ), + } + } + assert_eq!(seen, 3, "{replay}: not every operation completed"); + ring.run_down().expect("rundown"); + }); + } + let _ = std::fs::remove_file(path); +} diff --git a/crates/windows-ioring-sys/tests/completion_event.rs b/crates/windows-ioring-sys/tests/completion_event.rs index d293e5b04..9495c85aa 100644 --- a/crates/windows-ioring-sys/tests/completion_event.rs +++ b/crates/windows-ioring-sys/tests/completion_event.rs @@ -33,6 +33,18 @@ use windows_sys::Win32::System::Threading::{ CreateEventW, SetEvent, WaitForMultipleObjects, WaitForSingleObject, }; +/// How long a completion this test caused is allowed to take to arrive. +/// +/// M26.7 replaced a `try_pop` here that asserted the completion was *already* +/// queued when `submit_and_wait` returned. That is not something this crate +/// promises -- `pop_within`'s own documentation says a submit-side wait's +/// return "promises nothing about poppability", and `RESPONSE-SPACE.md` states +/// it as `RS-P-5`. It was true on the handle this test happens to use and is +/// false on others, which is the definition of a frozen observation. +/// +/// Stated as this crate's own contract instead: the completion arrives within +/// a bound we choose. Generous, because the bound is not what is under test. +const POP_BOUND: std::time::Duration = std::time::Duration::from_secs(10); const CHUNKS: usize = 8; const CHUNK_LEN: usize = 512; @@ -532,9 +544,9 @@ fn the_ring_still_wakes_after_an_unrelated_handle_fires_in_a_multiplexed_wait() "round 4: the batch's first completion must wake the wait" ); let stranded = ring - .try_pop() + .pop_within(POP_BOUND) .expect("pop one") - .expect("the batch left completions queued"); + .expect("the batch's remaining completion arrives within the bound"); let _buffer = pending .remove(&stranded.user_data()) .expect("the popped completion matches a held token") diff --git a/crates/windows-ioring-sys/tests/event_delivery.rs b/crates/windows-ioring-sys/tests/event_delivery.rs index 6f1aa3cd9..5b7380ac6 100644 --- a/crates/windows-ioring-sys/tests/event_delivery.rs +++ b/crates/windows-ioring-sys/tests/event_delivery.rs @@ -15,6 +15,8 @@ use std::sync::mpsc; use std::time::Duration; use windows_ioring_sys::{Batch, EventDelivery, IoRing, PushOptions, Token}; +use windows_sys::Win32::Foundation::{HANDLE, WAIT_OBJECT_0}; +use windows_sys::Win32::System::Threading::WaitForSingleObject; const CHUNKS: usize = 8; const CHUNK_LEN: usize = 512; @@ -34,6 +36,240 @@ fn filled_content() -> Vec { content } +/// How long a delivery is allowed to take before the test calls it stalled. +const DELIVERY_BOUND: Duration = Duration::from_secs(5); + +/// How much longer to wait, **only after a failure**, to learn whether the +/// delivery was lost or merely late. +/// +/// This costs nothing on a passing run because it is never reached. On a +/// failing one it is the single most discriminating fact available: a delivery +/// that arrives at nine seconds is a stall to be explained, while one that +/// never arrives is a lost wakeup, and those have entirely different causes. +/// +/// **Kept short enough that the report survives being killed.** The sabotage +/// harness bounds a suite at three times its baseline -- around thirty seconds +/// here -- and kills the process when that is exceeded. A post-mortem long +/// enough to push a failing run past that bound would trade the diagnosis for +/// the thing it was added to diagnose. The immediate facts are printed +/// *before* this wait for the same reason, so they survive even if it is. +const POST_MORTEM_BOUND: Duration = Duration::from_secs(10); + +/// What the delivery path did, captured so a timeout is a diagnosis rather +/// than a word. +/// +/// `M26.9` exists because these two tests time out at roughly one run in +/// eighty, and the message they produced -- `Timeout` -- ruled nothing out. +/// Every field below was chosen to separate hypotheses that message could not: +/// whether the pool ever ran the callback, whether the kernel ever finished +/// the I/O, whether deliveries were steady and then stopped, and whether the +/// missing one was lost or late. +struct DeliveryWatch { + started: std::time::Instant, + /// When each completion reached the test thread, relative to `started`. + arrivals: Vec, + /// How many times the pool actually invoked the callback. + callbacks: Arc, +} + +impl DeliveryWatch { + fn new(callbacks: Arc) -> Self { + Self { + started: std::time::Instant::now(), + arrivals: Vec::new(), + callbacks, + } + } + + fn record_arrival(&mut self) { + self.arrivals.push(self.started.elapsed()); + } + + /// Everything known at the moment a wait gave up. + /// + /// Written to stderr as well as into the panic message: `cargo test` + /// replays a failing test's captured output, and the sabotage harness + /// writes that transcript to `.scratch/sabotage/.txt`, so this is + /// what a later reader actually has to work from. + /// + /// **States what was observed and what each number means mechanically; + /// does not say what caused it.** An earlier draft ended each report with + /// a verdict, and a forced-failure run showed the verdict was wrong -- it + /// blamed something upstream of the channel when the injected fault was in + /// the callback body, which the counters it printed had already ruled out. + /// A diagnosis nobody asked for is worse than none, because it is the part + /// a tired reader will believe. + fn report(&self, what: &str, expected: usize, outstanding: usize, postmortem: &str) -> String { + let gaps: Vec = self + .arrivals + .windows(2) + .map(|pair| format!("{:?}", pair[1] - pair[0])) + .collect(); + let callbacks = self.callbacks.load(Ordering::SeqCst); + format!( + "M26.9 delivery stalled: {what}\n\ + \x20 delivered : {} of {expected}\n\ + \x20 callbacks run : {callbacks} -- times the pool invoked the callback. Equal to \ + delivered means everything the callback received reached this thread; greater \ + means the gap is between the callback and the channel.\n\ + \x20 outstanding : {outstanding} -- the ring's own count, which decrements on \ + pop, and the pop happens inside the callback. So it cannot on its own separate \ + 'the kernel has not finished' from 'the callback never ran'.\n\ + \x20 arrival times : {:?}\n\ + \x20 inter-arrival : [{}]\n\ + \x20 waited : {:?} before giving up\n\ + \x20 post-mortem : {postmortem}\n\ + {}", + self.arrivals.len(), + self.arrivals, + gaps.join(", "), + DELIVERY_BOUND, + trace_section(), + ) + } +} + +/// Whether the default process thread pool is still running callbacks at all. +/// +/// The discriminating probe for `M26.9`. Every captured stall shows the pool +/// never invoking the callback for *any* ring in the process, which has two +/// very different explanations: the whole default pool has stopped dispatching, +/// or only its wait mechanism has. This runs one of each and says which. +/// +/// Both use short bounds because they run inside an already-failing test; a +/// probe that hung would replace the diagnosis with a second timeout. +fn pool_liveness() -> String { + use windows_threadpool_sys::wait::{ThreadpoolWait, WaitableHandle}; + use windows_threadpool_sys::work::ThreadpoolWork; + + let probe_bound = Duration::from_secs(2); + + // 1. A plain work item. If this does not run, the pool is not dispatching + // anything and the wait mechanism is not the subject. + let (work_tx, work_rx) = mpsc::channel(); + let work_tx = std::sync::Mutex::new(work_tx); + let work_ran = match ThreadpoolWork::new( + move || { + if let Ok(tx) = work_tx.lock() { + let _ = tx.send(()); + } + }, + None, + ) { + Ok(work) => { + work.submit(); + work_rx.recv_timeout(probe_bound).is_ok() + } + Err(error) => return format!("could not create a work probe: {error}"), + }; + + // 2. A brand-new wait on a brand-new event, armed and then signalled. If + // the work item ran and this does not, the fault is specific to waits + // rather than to the pool as a whole. + let wait_ran = match WaitableHandle::event(false, false) { + Ok(event) => { + let (tx, rx) = mpsc::channel(); + let tx = std::sync::Mutex::new(tx); + match ThreadpoolWait::new( + event, + move |_| { + if let Ok(tx) = tx.lock() { + let _ = tx.send(()); + } + }, + None, + ) { + Ok(wait) => { + wait.arm(None); + // SAFETY: the wait owns the event, so the handle is open. + unsafe { + windows_sys::Win32::System::Threading::SetEvent( + std::os::windows::io::AsRawHandle::as_raw_handle(&wait.handle()), + ) + }; + rx.recv_timeout(probe_bound).is_ok() + } + Err(error) => return format!("could not create a wait probe: {error}"), + } + } + Err(error) => return format!("could not create a probe event: {error}"), + }; + + format!("work item ran: {work_ran}; a fresh wait ran: {wait_ran} (both within {probe_bound:?})") +} + +/// The concurrency trace, when this build carries one. +/// +/// Empty on an ordinary build, because the `trace` feature compiles the whole +/// facility away -- which is the point: the defect this chases is timing +/// dependent, so the instrument must be absent unless it is wanted. Enable it +/// with `--features trace` and narrow it with `WINDOWS_THREADPOOL_TRACE`: +/// +/// ```text +/// $env:WINDOWS_THREADPOOL_TRACE = 'wait,delivery' +/// cargo test -p windows-ioring-sys --features trace --test event_delivery +/// ``` +fn trace_section() -> String { + let dump = windows_threadpool_sys::trace::dump(); + if dump.trim().is_empty() { + return " trace : nothing recorded (set WINDOWS_THREADPOOL_TRACE to narrow one)" + .to_owned(); + } + format!(" trace (oldest first):\n{dump}") +} + +/// Wait for one delivery, turning a timeout into the report above. +fn recv_one( + rx: &mpsc::Receiver, + watch: &mut DeliveryWatch, + what: &str, + expected: usize, + outstanding: impl Fn() -> usize, +) -> windows_ioring_sys::Completion { + match rx.recv_timeout(DELIVERY_BOUND) { + Ok(completion) => { + watch.record_arrival(); + completion + } + Err(_) => { + // Read the ring's own count *before* the second wait, so it + // describes the moment of failure rather than the moment of + // giving up on it. + let at_failure = outstanding(); + // The pool-liveness probe runs first, while the process is still + // in the failed state -- asking afterwards would describe a + // different moment. + let liveness = pool_liveness(); + // Printed now, before waiting any longer. If this process is + // killed during the post-mortem -- which the sabotage harness will + // do if the suite exceeds its hang bound -- these lines are + // already in the transcript, and losing them is losing the whole + // reason the instrumentation exists. + eprintln!( + "{}\n pool liveness : {liveness}", + watch.report( + what, + expected, + at_failure, + &format!( + "waiting a further {POST_MORTEM_BOUND:?} to see whether it is late \ + rather than lost; the line below is the answer" + ) + ) + ); + let started = std::time::Instant::now(); + let postmortem = match rx.recv_timeout(POST_MORTEM_BOUND) { + Ok(_) => format!("it arrived, {:?} past the bound", started.elapsed()), + Err(_) => format!("still nothing after a further {POST_MORTEM_BOUND:?}"), + }; + panic!( + "{}\n pool liveness : {liveness}", + watch.report(what, expected, at_failure, &postmortem) + ); + } + } +} + #[test] fn completions_are_delivered_on_pool_threads_without_the_submitting_thread_waiting() { let path = temp_file("delivery"); @@ -49,11 +285,16 @@ fn completions_are_delivered_on_pool_threads_without_the_submitting_thread_waiti let submitting_thread = std::thread::current().id(); let saw_foreign_thread = Arc::new(AtomicBool::new(false)); let saw_foreign_thread_for_callback = Arc::clone(&saw_foreign_thread); + // Counted in the callback itself, so a stalled run can say whether the + // pool ever ran it -- which is the first fork in diagnosing M26.9. + let callbacks = Arc::new(AtomicUsize::new(0)); + let callbacks_for_callback = Arc::clone(&callbacks); let ring = IoRing::new(64, 64).expect("create ring"); let delivery = EventDelivery::new( ring, move |completion| { + callbacks_for_callback.fetch_add(1, Ordering::SeqCst); if std::thread::current().id() != submitting_thread { saw_foreign_thread_for_callback.store(true, Ordering::SeqCst); } @@ -82,11 +323,16 @@ fn completions_are_delivered_on_pool_threads_without_the_submitting_thread_waiti batch.submit_and_wait(0, 0).expect("submit without waiting"); } + let mut watch = DeliveryWatch::new(Arc::clone(&callbacks)); let mut received = 0; while received < CHUNKS { - let completion = rx - .recv_timeout(Duration::from_secs(5)) - .expect("completion delivered via the pool"); + let completion = recv_one( + &rx, + &mut watch, + "completions_are_delivered_on_pool_threads_without_the_submitting_thread_waiting", + CHUNKS, + || delivery.scope().outstanding(), + ); completion.result().expect("read succeeded"); received += 1; } @@ -195,9 +441,12 @@ fn completions_queued_before_handover_are_still_delivered() { } let (tx, rx) = mpsc::channel(); + let callbacks = Arc::new(AtomicUsize::new(0)); + let callbacks_for_callback = Arc::clone(&callbacks); let delivery = EventDelivery::new( ring, move |completion| { + callbacks_for_callback.fetch_add(1, Ordering::SeqCst); let _ = tx.send(completion); }, None, @@ -207,10 +456,22 @@ fn completions_queued_before_handover_are_still_delivered() { // Claim on this thread rather than in the callback, so a delivered // completion is checked against the token that minted it -- a delivery // that reported the wrong `UserData` would fail here rather than pass. + // + // Note what a stall means *here* specifically, and why the report says + // `outstanding` is not self-explanatory: every completion was already in + // the queue before the handover, so the kernel has nothing left to do. + // A timeout in this test therefore cannot be the device being slow. + let mut watch = DeliveryWatch::new(Arc::clone(&callbacks)); for _ in 0..CHUNKS { - let completion = rx - .recv_timeout(Duration::from_secs(5)) - .expect("a completion queued before handover must still be delivered"); + let completion = recv_one( + &rx, + &mut watch, + "completions_queued_before_handover_are_still_delivered (every completion was \ + already queued before the handover, so the kernel had nothing left to do -- a \ + stall here is in the signal or the pool, never in the device)", + CHUNKS, + || delivery.scope().outstanding(), + ); let transferred = completion.result().expect("read succeeded"); let (chunk_index, token) = pending .remove(&completion.user_data()) @@ -218,6 +479,10 @@ fn completions_queued_before_handover_are_still_delivered() { let buffer = token .claim_if(&completion) .expect("a token claims its own completion"); + // CONFIRMS: RS-P-8 -- a full count here is a property of the handle + // this test chose (an ordinary file on a local volume, where a + // successful completion carries the whole length and a full volume is + // an error instead), not of the space, which permits a short one. assert_eq!(transferred, CHUNK_LEN); assert_eq!( buffer, @@ -231,3 +496,197 @@ fn completions_queued_before_handover_are_still_delivered() { drop(delivery); } + +/// The backlog guarantee when the caller attached the event **first**. +/// +/// [`completions_queued_before_handover_are_still_delivered`] hands over a +/// *fresh* ring, so `EventDelivery::new` is what attaches the event and the +/// setup signal is owed to it. That leaves a legal sequence untested: a caller +/// may take the handle from [`IoRing::completion_event`] itself, submit work, +/// and only then hand the ring over. The event is already attached by then. +/// +/// A construction that signalled only when it had attached the event would arm +/// this wait on an already-non-empty queue with nothing owing -- and the event +/// is edge-triggered on empty-to-non-empty (D-19), so no later completion +/// would signal either. The backlog would be stranded permanently, which is +/// exactly what the guarantee promises against. +/// +/// **Ignored: this reproduces a defect that is not yet fixed (`M26.12`).** +/// +/// Raised by review on PR #108, and investigating it found something larger +/// than the report. Signalling unconditionally -- the obvious repair, and the +/// one the report suggests -- does **not** make this pass. What does is a +/// 50 ms sleep between `wait.arm` and the signal, measured 3 of 3 against 0 of +/// 6 without it, so the wakeup is lost in a window after arming rather than +/// never being raised. +/// +/// That matters beyond this test: `M26.9` fixed the delivery stall by ordering +/// the arm before the signal, and this says that ordering alone is not +/// sufficient. It is left failing-and-ignored rather than deleted, patched +/// with a sleep, or "fixed" by a change that does not fix it. +#[ignore = "M26.12: reproduces an unfixed wakeup race; see UNRESOLVED-TEST-FAILURES.md"] +#[test] +fn a_backlog_is_delivered_even_when_the_caller_attached_the_event_first() { + let path = temp_file("attached-before-handover"); + let content = filled_content(); + std::fs::write(&path, &content).expect("write fixture file"); + let file = std::fs::OpenOptions::new() + .read(true) + .open(&path) + .expect("open for read"); + let handle = file.as_raw_handle(); + + let mut ring = IoRing::new(64, 64).expect("create ring"); + + // Attach before any work exists, so the handover finds the event already + // in place -- and then **consume the signal that attaching raised**. + // + // Consuming it is the whole point, and the first version of this test + // omitted it and passed against the defect. `completion_event` signals as + // it attaches; on an auto-reset event that signal simply waits until + // something takes it. Arming a fresh wait would then have fired on the + // leftover signal and drained the backlog, so the test proved nothing. + // Taking it here leaves the state the guarantee is actually about: a + // non-empty queue with no wakeup pending anywhere. + let caller_handle = ring + .completion_event() + .expect("this system supports a completion event"); + let taken = unsafe { WaitForSingleObject(caller_handle.as_raw_handle() as HANDLE, 5_000) }; + assert_eq!( + taken, WAIT_OBJECT_0, + "attaching raises one signal and this consumes it" + ); + + let mut pending: HashMap>)> = HashMap::new(); + { + let mut batch = Batch::new(&mut ring); + for chunk_index in 0..CHUNKS { + let buffer = vec![0_u8; CHUNK_LEN]; + let offset = (chunk_index * CHUNK_LEN) as u64; + let token = unsafe { batch.read_raw(handle, buffer, offset, PushOptions::new()) } + .expect("queue read"); + pending.insert(token.id(), (chunk_index, token)); + } + batch + .submit_and_wait(CHUNKS as u32, 5_000) + .expect("submit and wait for every completion to land"); + } + + // Those completions took the queue from empty to non-empty, which signals + // the event again. Consume that one too, so the handover inherits a + // non-empty queue with nothing pending -- and because the queue never + // returns to empty, the edge cannot re-arm and no further signal is + // coming from the kernel either. + let taken = unsafe { WaitForSingleObject(caller_handle.as_raw_handle() as HANDLE, 5_000) }; + assert_eq!( + taken, WAIT_OBJECT_0, + "the completions landing signal the event once" + ); + + let (tx, rx) = mpsc::channel(); + let callbacks = Arc::new(AtomicUsize::new(0)); + let callbacks_for_callback = Arc::clone(&callbacks); + let delivery = EventDelivery::new( + ring, + move |completion| { + callbacks_for_callback.fetch_add(1, Ordering::SeqCst); + let _ = tx.send(completion); + }, + None, + ) + .expect("wire delivery to a ring whose event the caller already attached"); + + let mut watch = DeliveryWatch::new(Arc::clone(&callbacks)); + for _ in 0..CHUNKS { + let completion = recv_one( + &rx, + &mut watch, + "a_backlog_is_delivered_even_when_the_caller_attached_the_event_first (the event \ + was attached before the work was queued, so a construction that signals only on \ + its own attach leaves this backlog with no wakeup it will ever receive)", + CHUNKS, + || delivery.scope().outstanding(), + ); + let (chunk_index, token) = pending + .remove(&completion.user_data()) + .expect("completion matches a held token"); + let buffer = token + .claim_if(&completion) + .expect("a token claims its own completion"); + assert_eq!( + buffer, + content[chunk_index * CHUNK_LEN..(chunk_index + 1) * CHUNK_LEN] + ); + } + assert!( + pending.is_empty(), + "every completion queued before the handover must have been delivered, even though \ + the caller attached the event rather than the handover" + ); + + drop(delivery); +} + +#[test] +fn the_stall_report_carries_what_a_diagnosis_needs() { + // `M26.9`'s instrumentation, guarded without paying for it. + // + // The report is what a future reader of a flaked run has to work from, so + // it is worth knowing it still says something. Driving a real stall to + // find out would cost every suite run the delivery bound plus the + // post-mortem, which is why this exercises the formatting directly: the + // failure path's only other job is to call it, and that was verified once + // by forcing a stall by hand (measured: five of eight delivered, eight + // callbacks run, which localised the injected fault to between the + // callback and the channel exactly as the counters promise). + let callbacks = Arc::new(AtomicUsize::new(8)); + let mut watch = DeliveryWatch::new(callbacks); + watch.record_arrival(); + watch.record_arrival(); + + let report = watch.report("a_test_name", 8, 3, "still nothing"); + for needle in [ + "a_test_name", + "delivered", + "2 of 8", + "callbacks run", + "outstanding", + "arrival times", + "inter-arrival", + "post-mortem", + "still nothing", + ] { + assert!( + report.contains(needle), + "the stall report must carry {needle:?}, or a flaked run says less than it could \ + -- got:\n{report}" + ); + } + + // The counter that separates "the pool stopped calling us" from "the + // callback got it and the channel did not" is the one worth pinning by + // value rather than by name: a report that printed the delivered count + // twice would satisfy every check above. + assert!( + report.contains("callbacks run : 8"), + "the callback count must be the pool's, not the delivered count -- got:\n{report}" + ); +} + +// ------------------------------------------------------------------------ +// Relocated from `src/event_delivery/tests.rs` at 9bc0350e (M24.3). + +#[test] +fn new_succeeds_and_the_ring_stays_reachable_for_pushes() { + let ring = IoRing::new(8, 8).expect("create ring"); + let delivery = EventDelivery::new(ring, |_completion| {}, None).expect("wire event delivery"); + let info = delivery.scope().info().expect("query info"); + assert!(info.submission_queue_size > 0); +} + +#[test] +fn dropping_with_nothing_outstanding_does_not_hang() { + let ring = IoRing::new(8, 8).expect("create ring"); + let delivery = EventDelivery::new(ring, |_completion| {}, None).expect("wire event delivery"); + drop(delivery); +} diff --git a/crates/windows-ioring-sys/tests/failure_paths.rs b/crates/windows-ioring-sys/tests/failure_paths.rs index d8d963853..f1f07df3c 100644 --- a/crates/windows-ioring-sys/tests/failure_paths.rs +++ b/crates/windows-ioring-sys/tests/failure_paths.rs @@ -45,16 +45,16 @@ fn temp_file(tag: &str) -> PathBuf { )) } -/// Drain until a completion arrives. +/// Wait for one completion, bounded. +/// +/// Was an unbounded `loop` around `try_pop` plus `submit_and_wait`, which +/// turns a completion that never arrives into a hung harness with no test +/// name attached. `IoRing::pop_within` (M21.2) is the crate's own join +/// between the two and carries the bound. fn await_one(ring: &mut IoRing) -> Completion { - loop { - if let Some(completion) = ring.try_pop().expect("pop a completion") { - return completion; - } - Batch::new(ring) - .submit_and_wait(1, 30_000) - .expect("wait for a completion"); - } + ring.pop_within(std::time::Duration::from_secs(30)) + .expect("pop a completion") + .expect("a completion arrived within the bound") } #[test] diff --git a/crates/windows-ioring-sys/tests/fault_injection.rs b/crates/windows-ioring-sys/tests/fault_injection.rs index ca51a8633..c1c9e4caf 100644 --- a/crates/windows-ioring-sys/tests/fault_injection.rs +++ b/crates/windows-ioring-sys/tests/fault_injection.rs @@ -46,12 +46,11 @@ fn real_read( let token = unsafe { batch.read_raw(file.as_raw_handle(), vec![0_u8; 5], 0, PushOptions::new()) } .expect("queue a read"); - batch.submit_and_wait(1, 30_000).expect("submit and wait"); - let completion = loop { - if let Some(completion) = ring.try_pop().expect("pop") { - break completion; - } - }; + batch.submit().expect("submit the read"); + let completion = ring + .pop_within(std::time::Duration::from_secs(30)) + .expect("pop") + .expect("the read never completed"); (token, completion) } diff --git a/crates/windows-ioring-sys/tests/flush_barrier.rs b/crates/windows-ioring-sys/tests/flush_barrier.rs index d99b5874c..e665aa46e 100644 --- a/crates/windows-ioring-sys/tests/flush_barrier.rs +++ b/crates/windows-ioring-sys/tests/flush_barrier.rs @@ -2,6 +2,24 @@ //! M12.2: the flush-barrier contract, tested against the *measured* behaviour //! rather than against the implementation. //! +//! # This file's job changed in M26.6 +//! +//! It is now the place that confirms a **real kernel stays inside** +//! [RESPONSE-SPACE.md](../RESPONSE-SPACE.md)'s `RS-C-4` -- no operation queued +//! before a drain-flagged flush completes after it. That matters because +//! `RS-C-4` is the one clause the resolver is *forbidden* to break: a Windows +//! that broke the drain would be caught by nothing the resolver does, so this +//! is the only half of the pair that can see it. The division of labour is +//! stated in that document on the strength of this test existing. +//! +//! Two consequences follow, and both are visible below. The conformance +//! assertion runs on every machine, including ones where the control cannot +//! discriminate -- skipping it could only ever hide a violation. And a failure +//! here is a **finding about Windows**, not a regression in this crate, so it +//! says so in its own message. +//! +//! CONFIRMS: RS-C-4 +//! //! [`FlushCoverage`] exists because an unflagged flush does not cover the //! writes queued before it (D-23). A unit test can pin the enum-to-SQE-flag //! mapping, and `batch/tests.rs` does -- but that only proves the flag is set, @@ -304,34 +322,58 @@ fn a_covering_flush_waits_for_preceding_writes_and_an_unordered_one_does_not() { // this test skip on hardware where it has plenty to say. let unordered = run_case(&mut ring, handle, FlushCoverage::Unordered); - if !unordered.saw_reordering() { - drop(ring); - drop(file); - let _ = std::fs::remove_file(&path); - eprintln!( - "SKIP: an unordered flush produced completions in strict submission order on this \ - machine, so there is no reordering here to distinguish a barrier from, and the \ - covering assertions would pass for the wrong reason." - ); - return; - } - + // WHETHER THE CONTROL DISCRIMINATED DECIDES WHAT MAY BE CLAIMED, NOT + // WHETHER THE KERNEL IS CHECKED AT ALL (M26.6). + // + // This used to return here, and that was a hole. Two distinct questions + // were being answered by one gate: + // + // 1. "Is the barrier doing work?" -- a comparative claim, and it is + // genuinely meaningless when the control shows no reordering, since + // the covering case would then match a control that did nothing. + // 2. "Did the kernel stay inside `RS-C-4`?" -- a conformance question, + // and skipping it can only ever hide a violation. Observing no + // pre-flush write completing after the flush is weak evidence when + // nothing could have reordered, but observing one is a finding on + // any machine. + // + // `RESPONSE-SPACE.md` makes the second question this test's job, because + // `RS-C-4` is the one clause the resolver is forbidden to break: a Windows + // that broke the drain would be caught by nothing the resolver does. A + // skip here therefore left that constraint untested in both halves at + // once, which is exactly what `M26.6` was written to prevent. So the + // covering case now runs unconditionally and only the comparative claim + // is withheld. + let discriminating = unordered.saw_reordering(); let covering = run_case(&mut ring, handle, FlushCoverage::CoversPrecedingOperations); drop(ring); drop(file); let _ = std::fs::remove_file(&path); - // The contract (D-23). Not "usually waits": if the flush could complete - // before even one preceding write, its completion would not prove those - // writes are durable, which is the entire reason the variant exists. + // RS-C-4, against the real kernel. The contract (D-23): not "usually + // waits". If the flush could complete before even one preceding write, its + // completion would not prove those writes are durable, which is the entire + // reason the variant exists. assert_eq!( covering.a_after_flush, 0, - "a covering flush must not complete before any write queued ahead of it, but {} of \ - {PHASE_OPS} did (the unordered control saw {})", + "RS-C-4 violated: a covering flush must not complete before any write queued ahead of \ + it, but {} of {PHASE_OPS} did (the unordered control saw {}). This is a finding about \ + Windows, not about this crate -- see RESPONSE-SPACE.md, which constrains the resolver \ + from ever producing this and names these tests as the place it would surface.", covering.a_after_flush, unordered.a_after_flush ); + if !discriminating { + eprintln!( + "PARTIAL: an unordered flush produced completions in strict submission order on \ + this machine, so there was no reordering here for a barrier to suppress. The \ + RS-C-4 conformance assertion above still ran and still holds; what cannot be \ + claimed from this run is that the flag is what produced the ordering." + ); + return; + } + // What D-24 used to claim, and what replaced it. // // D-24 recorded the drain flag as a *full* barrier: preceding operations diff --git a/crates/windows-ioring-sys/tests/generated_sequences.rs b/crates/windows-ioring-sys/tests/generated_sequences.rs index 9faf5bee1..a24206a1d 100644 --- a/crates/windows-ioring-sys/tests/generated_sequences.rs +++ b/crates/windows-ioring-sys/tests/generated_sequences.rs @@ -15,6 +15,28 @@ //! completion is reported as a contract violation rather than inferred (M16). //! Without those, a generator would only be checking that nothing crashed. //! +//! **What this confirms about the platform (M26.6).** Because these sequences +//! run against a real ring and report to [`RingContract`], this file is where +//! [RESPONSE-SPACE.md](../RESPONSE-SPACE.md)'s conservation constraints are +//! checked against **Windows** rather than against a resolver: every submitted +//! operation completes exactly once (by the oracle's duplicate and outstanding +//! violations), a completion identifies its operation (by every token claiming +//! its own completion), and nothing completes before it is submitted (by the +//! oracle's unexpected-completion violation). +//! +//! The drain-ordering constraint is [flush_barrier.rs](flush_barrier.rs)'s, +//! since it needs an ordering observable this file deliberately does not +//! construct. +//! +//! A failure of one of those is a **finding about Windows**, not a regression +//! in this crate, and the space says so: those clauses are `Decided` rather +//! than `Observed`, meaning this crate requires them of the platform and has +//! no fallback if they do not hold. +//! +//! CONFIRMS: RS-C-1 +//! CONFIRMS: RS-C-2 +//! CONFIRMS: RS-C-3 +//! //! **Calibrated, and it failed the first time.** M17.4 reverted D-20's setup //! signal -- #47 exactly as it shipped -- and this file reported green. It //! attached the event and sampled the right states, but drained by polling @@ -60,6 +82,18 @@ use windows_ioring_sys::{ use windows_sys::Win32::Foundation::{WAIT_OBJECT_0, WAIT_TIMEOUT}; use windows_sys::Win32::System::Threading::WaitForSingleObject; +/// How long a completion this test caused is allowed to take to arrive. +/// +/// M26.7 replaced a `try_pop` here that asserted the completion was *already* +/// queued when `submit_and_wait` returned. That is not something this crate +/// promises -- `pop_within`'s own documentation says a submit-side wait's +/// return "promises nothing about poppability", and `RESPONSE-SPACE.md` states +/// it as `RS-P-5`. It was true on the handle this test happens to use and is +/// false on others, which is the definition of a frozen observation. +/// +/// Stated as this crate's own contract instead: the completion arrives within +/// a bound we choose. Generous, because the bound is not what is under test. +const POP_BOUND: std::time::Duration = std::time::Duration::from_secs(10); /// A use-after-free in a generated sequence must fault rather than read stale /// bytes, or the generator is only testing that nothing happened to crash. #[global_allocator] @@ -405,9 +439,9 @@ fn run_plan( let pending = unsafe { batch.register_files(&[handle]) }.expect("queue file registration"); batch.submit_and_wait(1, 5_000).expect("submit"); let completion = ring - .try_pop() + .pop_within(POP_BOUND) .expect("pop") - .expect("a registration completion is ready"); + .expect("a registration completion arrives within the bound"); pending .claim_if(&completion) .expect("registration token claims its own completion") @@ -426,9 +460,9 @@ fn run_plan( .expect("queue buffer registration"); batch.submit_and_wait(1, 5_000).expect("submit"); let completion = ring - .try_pop() + .pop_within(POP_BOUND) .expect("pop") - .expect("a registration completion is ready"); + .expect("a registration completion arrives within the bound"); pending .claim_if(&completion) .expect("registration token claims its own completion") diff --git a/crates/windows-ioring-sys/tests/kernel_span.rs b/crates/windows-ioring-sys/tests/kernel_span.rs index 38dcad7dc..504b28596 100644 --- a/crates/windows-ioring-sys/tests/kernel_span.rs +++ b/crates/windows-ioring-sys/tests/kernel_span.rs @@ -44,6 +44,18 @@ use windows_ioring_sys::{ Batch, IoRing, PushOptions, RegisteredBuffers, RegisteredSpan, SharedFile, WriteCaching, }; +/// How long a completion this test caused is allowed to take to arrive. +/// +/// M26.7 replaced a `try_pop` here that asserted the completion was *already* +/// queued when `submit_and_wait` returned. That is not something this crate +/// promises -- `pop_within`'s own documentation says a submit-side wait's +/// return "promises nothing about poppability", and `RESPONSE-SPACE.md` states +/// it as `RS-P-5`. It was true on the handle this test happens to use and is +/// false on others, which is the definition of a frozen observation. +/// +/// Stated as this crate's own contract instead: the completion arrives within +/// a bound we choose. Generous, because the bound is not what is under test. +const POP_BOUND: std::time::Duration = std::time::Duration::from_secs(10); #[global_allocator] static ALLOC: windows_guard_alloc::GuardAlloc = windows_guard_alloc::GuardAlloc::new(); @@ -94,9 +106,9 @@ fn registered_arena(ring: &mut IoRing, count: u32, seed: u64) -> RegisteredBuffe .submit_and_wait(1, WAIT_MS) .expect("submit the registration"); let completion = ring - .try_pop() + .pop_within(POP_BOUND) .expect("pop the registration completion") - .expect("a completion is ready"); + .expect("a completion arrives within the bound"); let mut registered = pending .claim_if(&completion) .expect("the registration token claims its own completion") @@ -108,16 +120,16 @@ fn registered_arena(ring: &mut IoRing, count: u32, seed: u64) -> RegisteredBuffe registered } -/// Drain until a completion arrives, then hand it back. +/// Wait for one completion, bounded. +/// +/// Was an unbounded `loop` around `try_pop` plus `submit_and_wait`, which +/// turns a completion that never arrives into a hung harness reporting no +/// test name at all. `IoRing::pop_within` (M21.2) is the crate's own join +/// between the two and carries the bound. fn await_one(ring: &mut IoRing) -> windows_ioring_sys::Completion { - loop { - if let Some(completion) = ring.try_pop().expect("pop a completion") { - return completion; - } - Batch::new(ring) - .submit_and_wait(1, WAIT_MS) - .expect("wait for a completion"); - } + ring.pop_within(std::time::Duration::from_millis(u64::from(WAIT_MS))) + .expect("pop a completion") + .expect("a completion arrived within the bound") } fn open_shared(path: &std::path::Path, write: bool) -> SharedFile { diff --git a/crates/windows-ioring-sys/tests/properties_under_every_resolution.rs b/crates/windows-ioring-sys/tests/properties_under_every_resolution.rs new file mode 100644 index 000000000..017538bb8 --- /dev/null +++ b/crates/windows-ioring-sys/tests/properties_under_every_resolution.rs @@ -0,0 +1,795 @@ +// Copyright (c) Mike Grier +//! The properties that must hold under **every** resolution (M26.4). +//! +//! `M26.3` built a resolver that chooses a point in +//! [RESPONSE-SPACE.md](../RESPONSE-SPACE.md) from a seed. This file is what +//! that resolver is *for*: five properties this crate must satisfy whichever +//! point is chosen, driven across many plans and many resolutions. +//! +//! | # | Property | Checked by | +//! |---|---|---| +//! | P-1 | Conservation: no lost, duplicated or unclaimed completion | [`RingContract`], fed from the run | +//! | P-2 | No hang: every loop terminates | a step budget, exceeded = failure | +//! | P-3 | [`IoRing::pop_within`] honours its bound | elapsed time, plus a stalled resolution | +//! | P-4 | [`IoRing::outstanding`] is accurate | compared against the contract's own count | +//! | P-5 | No use-after-free | [`windows_guard_alloc::GuardAlloc`] | +//! +//! # `RingContract` is the definition, not a second copy +//! +//! P-1 is not restated here. [`RingContract`] already holds this crate's +//! conservation rules as an oracle over observed sequences, and the layer that +//! owns an invariant owns the oracle for it -- a copy written in a harness is +//! a second implementation of the rule rather than a check of it, and when the +//! two disagree it is the harness that gets "fixed". So this file *reports* to +//! the contract and asks it for the verdict. +//! +//! P-4 follows the same rule rather than counting for itself. "How many +//! operations are outstanding" is a fact the contract already derives from +//! what it was told, so the expected value is read back out of it via +//! [`Violation::Outstanding`] rather than tracked in a counter beside it. A +//! counter here would be a third party to the disagreement. +//! +//! # What these properties are honestly weak about +//! +//! Stated plainly, because a property suite that oversells itself is worse +//! than one that admits a gap -- the gap is then invisible rather than merely +//! open. +//! +//! **P-3's upper bound is nearly free under an ordinary resolution.** The +//! resolver answers a wait immediately, so `pop_within` rarely approaches its +//! deadline and "it did not exceed the bound" is close to vacuous. The +//! non-vacuous case needs a resolution in which *nothing completes during the +//! window*, which is what [`Stalled`] supplies -- and note what that is and is +//! not: it is not a violation of `RS-C-1`, which says every operation +//! completes *eventually*, because no finite observation can distinguish +//! "eventually" from "never". It is the prefix of an `RS-C-1`-satisfying +//! resolution in which the eventually has not happened yet, and that prefix is +//! exactly when a bound has to be honoured. +//! +//! **P-5 covers this crate's memory handling, not the kernel's.** Under a +//! resolver no operation reaches the kernel, so nothing external writes into a +//! buffer and the classic ring use-after-free -- the kernel writing through a +//! pointer whose owner has dropped -- cannot occur here at all. What the guard +//! allocator still sees is every lifetime decision the crate itself makes +//! around tokens, buffers and guards. The kernel-side half is +//! [generated_sequences.rs](generated_sequences.rs)'s job, against a real +//! ring, and stays there. +//! +//! # Three seeds, and why they are not one +//! +//! `D-41`'s discipline is that one number replays one thing. +//! [generated_sequences.rs](generated_sequences.rs) already carries two axes; +//! this adds the resolver's, making three. They are independent on purpose: +//! pinning the plan seed alone reproduces the same operations against +//! different resolutions, which is what you want when a plan looks suspicious, +//! and pinning the resolver seed alone reproduces the same resolutions against +//! different plans. A failure prints all three, because only all three replay +//! the whole run. + +#![cfg(all(windows, feature = "kernel-seam"))] + +use std::time::{Duration, Instant}; + +use windows_ioring_sys::contract::{RingContract, Violation}; +use windows_ioring_sys::sys::{Resolver, ResolverWatch, Responses}; +use windows_ioring_sys::{ + Batch, Completion, FlushCoverage, FlushMode, IoRing, PushOptions, SharedFile, Token, + WriteCaching, +}; +use windows_sys::core::HRESULT; + +/// P-5. A use-after-free in a generated plan must fault rather than read stale +/// bytes, or this file is only checking that nothing happened to crash. +#[global_allocator] +static ALLOC: windows_guard_alloc::GuardAlloc = windows_guard_alloc::GuardAlloc::new(); + +/// Environment override for the plan seed (decimal, or `0x` hex). The resolver +/// and the guard allocator each have their own, separate variable. +const PLAN_SEED_VAR: &str = "WINDOWS_IORING_PROPERTY_SEED"; + +/// Plans per run, and the longest plan generated. +/// +/// Each plan draws its own resolution, so this is the resolver seed count for +/// this file and moves with the repository's standard sweep size. Raised from +/// 240 to 2048 for breadth: no real I/O happens under a resolver, so the cost +/// is almost entirely the ring handles. +const PLANS: usize = 2048; +/// Steps in the longest plan. +const MAX_STEPS: usize = 10; +/// Operations in the largest single batch. +const MAX_BATCH: usize = 5; + +/// P-2's budget: how many consultations one plan may take to quiesce. +/// +/// This is the whole of "no hang" as a finite test can state it. A resolver +/// answers instantly, so a genuine hang here is an unterminating *loop* rather +/// than a block, and a loop that exceeds a budget generous enough for any +/// legitimate resolution is the observable form of one. Generous is load +/// bearing: a tight budget would fail for slow-but-correct resolutions and +/// teach its reader to raise it. +const DRAIN_BUDGET: usize = 4096; + +/// Bytes per write. Small, because the buffer exists to give the guard +/// allocator a lifetime to watch rather than to move data. +const BUF_LEN: usize = 256; + +// ------------------------------------------------------------- generation --- + +/// SplitMix64, the generator this crate uses everywhere (`D-41`). +struct Rng(u64); + +impl Rng { + fn next(&mut self) -> u64 { + self.0 = self.0.wrapping_add(0x9E37_79B9_7F4A_7C15); + let mut z = self.0; + z = (z ^ (z >> 30)).wrapping_mul(0xBF58_476D_1CE4_E5B9); + z = (z ^ (z >> 27)).wrapping_mul(0x94D0_49BB_1331_11EB); + z ^ (z >> 31) + } + + fn below(&mut self, bound: u64) -> u64 { + self.next() % bound + } + + fn chance(&mut self, percent: u64) -> bool { + self.below(100) < percent + } +} + +/// One operation in a batch. +#[derive(Debug, Clone, Copy)] +enum Op { + Flush { drain: bool }, + Write { drain: bool }, +} + +/// One step of a plan. +#[derive(Debug, Clone)] +enum Step { + /// Build a batch and submit it. + Submit(Vec), + /// Poll `try_pop` until the queue reports empty. + Drain, + /// Wait for one completion with a bound. + PopWithin(u64), + /// Check P-4 at an arbitrary mid-run point rather than only at the end. + CheckOutstanding, +} + +fn generate(rng: &mut Rng) -> Vec { + let steps = 1 + rng.below(MAX_STEPS as u64) as usize; + (0..steps) + .map(|_| match rng.below(10) { + 0..=4 => { + let count = 1 + rng.below(MAX_BATCH as u64) as usize; + Step::Submit( + (0..count) + .map(|_| { + let drain = rng.chance(25); + if rng.chance(50) { + Op::Flush { drain } + } else { + Op::Write { drain } + } + }) + .collect(), + ) + } + 5..=6 => Step::Drain, + 7..=8 => Step::PopWithin(rng.below(20)), + _ => Step::CheckOutstanding, + }) + .collect() +} + +// ------------------------------------------------------------------- run --- + +/// Tokens still awaiting their completion. +/// +/// Two lists because the two operations hand back different types, and that +/// difference is the point: a write's token owns its buffer, which is the +/// lifetime P-5 is watching. +#[derive(Default)] +struct Outstanding { + flushes: Vec>, + writes: Vec, SharedFile)>>, +} + +/// What a whole run actually exercised. +/// +/// Without this the suite could pass by doing nothing: a generator that +/// produced empty plans, or a resolution that completed everything inside the +/// submit that carried it, would satisfy all five properties trivially. These +/// counters are asserted at the end, so "the properties held" means "the +/// properties held over runs that reached the states they are about". +#[derive(Debug, Default)] +struct Coverage { + pushes: usize, + completions: usize, + claims: usize, + deliberate_leaks: usize, + barriers: usize, + writes: usize, + flushes: usize, + declined_submits: usize, + /// `CheckOutstanding` steps that ran with work genuinely in flight -- the + /// only ones where P-4 could have caught a disagreement. + outstanding_checks_with_work: usize, + /// `PopWithin` calls that found nothing, which is the case P-3's bound is + /// about. + empty_pop_withins: usize, +} + +struct Run { + ring: IoRing, + file: SharedFile, + contract: RingContract, + tokens: Outstanding, + /// The resolution's own record, so a declined submit can be recognised by + /// asking the resolver rather than by matching an `HRESULT` here. + watch: ResolverWatch, + /// Declined submits this plan absorbed, for the trace. + /// + /// The thinnest of the coverage signals, because a decline has to fall on + /// a plan that then reaches a `pop_within` -- a decline absorbed by the + /// ignored `batch.submit()` is never counted here. Its threshold is set + /// well below the rest for that reason; tightening it would buy a flake + /// rather than a guarantee. + declined: usize, + /// Plans print their steps only on failure, so this costs nothing until + /// something needs reading. + trace: Vec, + coverage: Coverage, +} + +impl Coverage { + fn fold(&mut self, other: &Self) { + self.pushes += other.pushes; + self.completions += other.completions; + self.claims += other.claims; + self.deliberate_leaks += other.deliberate_leaks; + self.barriers += other.barriers; + self.writes += other.writes; + self.flushes += other.flushes; + self.declined_submits += other.declined_submits; + self.outstanding_checks_with_work += other.outstanding_checks_with_work; + self.empty_pop_withins += other.empty_pop_withins; + } +} + +/// P-4's expected value, derived from the contract rather than counted here. +/// +/// An operation the contract still has in a pushed state is exactly one whose +/// completion has not been popped, which is what `IoRing::outstanding` counts. +fn contract_outstanding(contract: &RingContract) -> usize { + contract + .check_quiescent() + .iter() + .filter(|violation| matches!(violation, Violation::Outstanding { .. })) + .count() +} + +impl Run { + fn new(path: &std::path::Path, watch: ResolverWatch) -> std::io::Result { + let file = std::fs::OpenOptions::new() + .create(true) + .truncate(true) + .write(true) + .open(path)?; + Ok(Self { + ring: IoRing::new(64, 128)?, + file: SharedFile::new(file.into()), + contract: RingContract::new(), + tokens: Outstanding::default(), + watch, + declined: 0, + trace: Vec::new(), + coverage: Coverage::default(), + }) + } + + /// One `pop_within`, absorbing a submit the resolution declined. + /// + /// **A declined submit is not a property failure**, and treating it as one + /// was this harness's own first defect. `RS-P-7` leaves the operations + /// queued rather than losing them, and `pop_within` documents that it + /// returns any error from `SubmitIoRing` -- so a correct consumer retries, + /// and a harness that panicked instead would be asserting a contract the + /// crate never offered. + /// + /// Which errors are refusals is **asked of the resolver** rather than + /// matched against an `HRESULT` here. A hard-coded code would be a second + /// copy of a choice the resolver owns, and the two would drift the first + /// time it picked a different one. Retrying is bounded by P-2's budget in + /// every caller, so a resolution that declined forever is still caught. + fn pop_within(&mut self, bound: Duration) -> Result, String> { + let before = self.watch.stats().failed_submits; + match self.ring.pop_within(bound) { + Ok(outcome) => Ok(outcome), + Err(_) if self.watch.stats().failed_submits > before => { + self.declined += 1; + self.coverage.declined_submits += 1; + Ok(None) + } + Err(error) => Err(format!("pop_within failed: {error}")), + } + } + + /// P-4, checked wherever it is asked for. + fn check_outstanding(&mut self, where_: &str) -> Result<(), String> { + let expected = contract_outstanding(&self.contract); + let actual = self.ring.outstanding(); + if expected == actual { + if expected > 0 { + self.coverage.outstanding_checks_with_work += 1; + } + self.trace.push(format!("{where_}: outstanding {actual}")); + return Ok(()); + } + Err(format!( + "P-4 broken at {where_}: IoRing::outstanding() says {actual}, the contract's own \ + record says {expected}" + )) + } + + /// Account for one popped completion: tell the contract, then either claim + /// the token or abandon it on purpose. + /// + /// The order matters and is not arbitrary. `RingContract` moves a `Pushed` + /// operation to a provisionally-leaked state on completion and only + /// `observe_claim` corrects it, so reporting a deliberate leak *before* the + /// completion would leave the contract seeing a second completion for an + /// already-settled operation and reporting a duplicate that did not happen. + fn account(&mut self, completion: &Completion, rng: &mut Rng) -> Result<(), String> { + let user_data = completion.user_data(); + self.contract.observe_completion(user_data); + self.coverage.completions += 1; + + if let Some(at) = self + .tokens + .flushes + .iter() + .position(|token| token.id() == user_data) + { + let token = self.tokens.flushes.remove(at); + return self.settle(token.claim_if(completion).is_ok(), user_data, rng); + } + if let Some(at) = self + .tokens + .writes + .iter() + .position(|token| token.id() == user_data) + { + let token = self.tokens.writes.remove(at); + return self.settle(token.claim_if(completion).is_ok(), user_data, rng); + } + Err(format!( + "a completion arrived for {user_data:#x}, which no live token matches -- either \ + RS-C-2 was broken or this harness lost a push" + )) + } + + fn settle(&mut self, claimed: bool, user_data: usize, rng: &mut Rng) -> Result<(), String> { + if !claimed { + return Err(format!( + "the token for {user_data:#x} refused a completion carrying its own identity, \ + which is RS-C-2 broken inside the crate's own matching" + )); + } + // Claiming already happened; whether to *report* it as a claim or as a + // deliberate abandonment is what exercises both of the contract's + // settled states. Both are legitimate endings, and the contract + // distinguishes them precisely so an unstated leak stays a violation. + if rng.chance(15) { + self.contract.observe_deliberate_leak(user_data); + self.coverage.deliberate_leaks += 1; + } else { + self.contract.observe_claim(user_data); + self.coverage.claims += 1; + } + Ok(()) + } + + /// Pop everything currently available. P-2's budget applies. + fn drain(&mut self, rng: &mut Rng) -> Result<(), String> { + for _ in 0..DRAIN_BUDGET { + let Some(completion) = self + .ring + .try_pop() + .map_err(|error| format!("try_pop failed: {error}"))? + else { + return Ok(()); + }; + self.account(&completion, rng)?; + } + Err(format!( + "P-2 broken: try_pop kept yielding completions past a budget of {DRAIN_BUDGET}, \ + which no resolution satisfying RS-C-1 can do" + )) + } + + /// Drive to quiescence. P-2's budget applies, and this is where a + /// resolution that stopped making progress is caught. + fn quiesce(&mut self, rng: &mut Rng) -> Result<(), String> { + for _ in 0..DRAIN_BUDGET { + if self.ring.outstanding() == 0 { + self.drain(rng)?; + return Ok(()); + } + match self.pop_within(Duration::from_millis(5))? { + Some(completion) => self.account(&completion, rng)?, + None => continue, + } + } + Err(format!( + "P-2 broken: {} operations still outstanding after a budget of {DRAIN_BUDGET} \ + consultations, so this resolution never satisfied RS-C-1", + self.ring.outstanding() + )) + } +} + +/// Run one plan under one resolution. `Ok` is every property holding. +fn run_plan(plan: &[Step], plan_rng: &mut Rng, run: &mut Run) -> Result<(), String> { + for (index, step) in plan.iter().enumerate() { + run.trace.push(format!("{index}: {step:?}")); + match step { + Step::Submit(ops) => { + let mut pushed = Vec::new(); + { + let mut batch = Batch::new(&mut run.ring); + for op in ops { + match op { + Op::Flush { drain } => { + let coverage = if *drain { + FlushCoverage::CoversPrecedingOperations + } else { + FlushCoverage::Unordered + }; + let token = batch + .flush(&run.file, coverage, FlushMode::Default) + .map_err(|error| format!("flush build failed: {error}"))?; + pushed.push((token.id(), *drain, false)); + run.tokens.flushes.push(token); + } + Op::Write { drain } => { + let token = batch + .write( + &run.file, + vec![0_u8; BUF_LEN], + 0, + PushOptions::new().drain_preceding(*drain), + WriteCaching::Cached, + ) + .map_err(|error| format!("write build failed: {error}"))?; + pushed.push((token.id(), *drain, true)); + run.tokens.writes.push(token); + } + } + } + // A submit may be declined under RS-P-7, which leaves the + // operations queued for a later one rather than losing + // them -- so it is not an error here, and the properties + // below still have to hold. That the refusal reaches + // `run_down` as a failure is a separate, queued finding + // (M26.8); this path never calls it. + let _ = batch.submit(); + } + // Reported only after the batch is gone, because a push is + // reported when it *queued*, and the build is what queues it. + for (user_data, drain, is_write) in pushed { + run.contract.observe_push(user_data); + run.coverage.pushes += 1; + if drain { + run.coverage.barriers += 1; + } + if is_write { + run.coverage.writes += 1; + } else { + run.coverage.flushes += 1; + } + } + } + Step::Drain => run.drain(plan_rng)?, + Step::PopWithin(millis) => { + // P-3. The bound is what is being checked, so it is measured + // around the call rather than assumed from the argument. + let bound = Duration::from_millis(*millis); + let started = Instant::now(); + let outcome = run.pop_within(bound)?; + let elapsed = started.elapsed(); + if elapsed > bound + Duration::from_millis(250) { + return Err(format!( + "P-3 broken: pop_within({millis}ms) took {elapsed:?}" + )); + } + if let Some(completion) = outcome { + run.account(&completion, plan_rng)?; + } else { + run.coverage.empty_pop_withins += 1; + } + } + Step::CheckOutstanding => run.check_outstanding(&format!("step {index}"))?, + } + } + + run.quiesce(plan_rng)?; + run.check_outstanding("quiescence")?; + + // P-1, asked of the oracle rather than restated. + let violations = run.contract.check_quiescent(); + if !violations.is_empty() { + return Err(format!( + "P-1 broken: the ring contract reports\n{}", + violations + .iter() + .map(|violation| format!(" - {violation}")) + .collect::>() + .join("\n") + )); + } + Ok(()) +} + +fn seed_from_env(var: &str) -> u64 { + match std::env::var(var) { + Ok(text) => { + let trimmed = text.trim(); + trimmed + .strip_prefix("0x") + .or_else(|| trimmed.strip_prefix("0X")) + .map_or_else( + || trimmed.parse::().ok(), + |hex| u64::from_str_radix(hex, 16).ok(), + ) + .unwrap_or_else(|| { + panic!("{var} is set to {text:?}, which is neither decimal nor 0x hex") + }) + } + Err(_) => std::time::SystemTime::now() + .duration_since(std::time::UNIX_EPOCH) + .map_or(0x243F_6A88_85A3_08D3, |elapsed| elapsed.as_nanos() as u64), + } +} + +#[test] +fn the_properties_hold_under_every_resolution() { + let plan_seed = seed_from_env(PLAN_SEED_VAR); + let resolver_seed = seed_from_env(windows_ioring_sys::sys::RESOLVER_SEED_VAR); + ALLOC.announce_seed(); + eprintln!( + "properties: plan seed 0x{plan_seed:016X}, resolver seed 0x{resolver_seed:016X}\n\ + replay this run with:\n \ + $env:{PLAN_SEED_VAR}='0x{plan_seed:016X}'; \ + $env:{}='0x{resolver_seed:016X}'; \ + $env:WINDOWS_GUARD_ALLOC_SEED='0x{:016X}'; \ + cargo test -p windows-ioring-sys --all-features --test properties_under_every_resolution", + windows_ioring_sys::sys::RESOLVER_SEED_VAR, + ALLOC.seed() + ); + assert!( + ALLOC.total_allocations() > 0, + "P-5's detector is not installed, so these plans are running uninstrumented and only \ + prove that nothing crashed" + ); + + let path = std::env::temp_dir().join(format!( + "windows-ioring-sys-m26-4-{}.tmp", + std::process::id() + )); + let mut plan_rng = Rng(plan_seed); + let mut resolver_rng = Rng(resolver_seed); + let mut coverage = Coverage::default(); + + for index in 0..PLANS { + let plan = generate(&mut plan_rng); + // The resolution is drawn from its own axis. Deriving it from the plan + // seed would make the two inseparable, and pinning either would then + // reproduce a run that is half the one that failed. + let this_resolver = resolver_rng.next(); + let resolver = Resolver::new(this_resolver); + + let outcome = resolver.scoped(|watch| { + let mut run = Run::new(&path, watch.clone()).expect("a ring and a scratch file"); + let outcome = run_plan(&plan, &mut plan_rng, &mut run); + // Whatever happened, take the ring down inside the scope -- its + // rundown is answered by the resolver that owns its operations. + // Its result is not checked: `quiesce` has already driven + // everything to completion, so this is a no-op on the passing + // path, and on the failing path a rundown error would only + // re-report `M26.8` over whatever actually went wrong. + let _ = run.ring.run_down(); + match outcome { + Ok(()) => Ok(run.coverage), + Err(failure) => Err((failure, run.declined, run.trace)), + } + }); + + match outcome { + Ok(plan_coverage) => coverage.fold(&plan_coverage), + Err((failure, declined, trace)) => panic!( + "plan {index} of {PLANS} broke a property\n\ + replay: $env:{PLAN_SEED_VAR}='0x{plan_seed:016X}'; \ + $env:{}='0x{resolver_seed:016X}'; \ + $env:WINDOWS_GUARD_ALLOC_SEED='0x{:016X}'\n\ + this plan's resolution: 0x{this_resolver:016X}\n\ + declined submits absorbed: {declined}\n\ + {failure}\n\ + trace:\n {}", + windows_ioring_sys::sys::RESOLVER_SEED_VAR, + ALLOC.seed(), + trace.join("\n ") + ), + } + } + let _ = std::fs::remove_file(&path); + + // D-41's corollary, applied to this suite: a green run is evidence only if + // the run reached the states the properties are about. Each of these would + // be zero for a generator that produced empty plans, a resolution that + // resolved everything inside its submit, or a harness that quietly stopped + // reporting -- all of which pass every assertion above. + // + // The thresholds are fractions of `PLANS` rather than fixed counts, so + // that widening the sweep cannot quietly turn a guard into a formality: + // an absolute floor set for a 240-plan run is a 12x margin at 2048, which + // would let the suite lose most of its work without complaint. Each is + // roughly half the minimum observed over six runs, since the plan and + // resolver seeds are clock-derived and these counts therefore vary. + eprintln!("coverage: {coverage:?}"); + assert!( + coverage.pushes > PLANS * 4, + "too few operations: {coverage:?}" + ); + assert_eq!( + coverage.completions, coverage.pushes, + "every pushed operation must have completed exactly once for the run to have finished \ + at all -- {coverage:?}" + ); + assert!( + coverage.writes > PLANS * 2 && coverage.flushes > PLANS * 2, + "both operation shapes must appear: {coverage:?}" + ); + assert!( + coverage.barriers > PLANS, + "too few drain-flagged operations, so RS-C-4 barely applied: {coverage:?}" + ); + assert!( + coverage.claims > PLANS * 3 && coverage.deliberate_leaks > PLANS / 2, + "both of the contract's settled states must be reached: {coverage:?}" + ); + assert!( + coverage.declined_submits > PLANS / 128, + "too few resolutions declined a submit, so RS-P-7 barely reached the crate: \ + {coverage:?}" + ); + assert!( + coverage.outstanding_checks_with_work > PLANS / 8, + "P-4 was rarely checked with a non-empty ring, where it cannot catch a disagreement: \ + {coverage:?}" + ); + assert!( + coverage.empty_pop_withins > PLANS / 8, + "pop_within almost always found something, so its bound rarely had anything to do: \ + {coverage:?}" + ); +} + +// ------------------------------------------------------------------- P-3 --- + +/// A resolution in which nothing has completed yet. +/// +/// Not a violation of `RS-C-1`, which says every operation completes +/// *eventually* -- no finite observation can distinguish "eventually" from +/// "never", so this is the prefix of a satisfying resolution rather than a +/// counter-example to one. That prefix is the only window in which a bound has +/// anything to do, which is why P-3's non-vacuous case needs it. +/// +/// Deliberately not a [`Resolver`] with a switch: a knob that made a resolver +/// stop satisfying a constraint is exactly what `D-61` refuses, because a test +/// could then assert against a platform that cannot exist. This is a separate, +/// degenerate responder that claims to be nothing else. +struct Stalled; + +impl Responses for Stalled { + unsafe fn submit( + &mut self, + _ring: *mut std::ffi::c_void, + _wait_operations: u32, + _milliseconds: u32, + submitted: *mut u32, + ) -> HRESULT { + // SAFETY: a valid out-pointer, as the Win32 call requires. + unsafe { submitted.write(0) }; + 0 + } + + unsafe fn pop( + &mut self, + _ring: *mut std::ffi::c_void, + _cqe: *mut windows_sys::Win32::Storage::FileSystem::IORING_CQE, + ) -> HRESULT { + // `S_FALSE`: the queue is empty. + 1 + } + + unsafe fn build_flush( + &mut self, + _ring: *mut std::ffi::c_void, + _file: windows_sys::Win32::Storage::FileSystem::IORING_HANDLE_REF, + _mode: windows_sys::Win32::Storage::FileSystem::FILE_FLUSH_MODE, + _user_data: usize, + _flags: i32, + ) -> HRESULT { + 0 + } +} + +#[test] +fn pop_within_honours_its_bound_when_nothing_completes() { + // P-3, in the only window where it is not nearly vacuous. Under an + // ordinary resolution a completion is usually waiting almost at once, so + // "it did not exceed the bound" says little; here nothing arrives at all + // and the deadline is the only thing that can end the call. + let guard = windows_ioring_sys::sys::install(Box::new(Stalled)); + let mut ring = IoRing::new(64, 128).expect("a ring"); + let path = std::env::temp_dir().join(format!( + "windows-ioring-sys-m26-4-stalled-{}.tmp", + std::process::id() + )); + let file = SharedFile::new( + std::fs::OpenOptions::new() + .create(true) + .truncate(true) + .write(true) + .open(&path) + .expect("a scratch file") + .into(), + ); + + let token = { + let mut batch = Batch::new(&mut ring); + let token = batch + .flush(&file, FlushCoverage::Unordered, FlushMode::Default) + .expect("a flush builds"); + batch.submit().expect("the submit is answered"); + token + }; + assert_eq!( + ring.outstanding(), + 1, + "the bound only has work to do while something is outstanding -- with nothing \ + outstanding pop_within returns at once by documented design, which would make this \ + test measure the early return instead" + ); + + for millis in [0_u64, 5, 30, 120] { + let bound = Duration::from_millis(millis); + let started = Instant::now(); + let outcome = ring.pop_within(bound).expect("pop_within does not fail"); + let elapsed = started.elapsed(); + assert!( + outcome.is_none(), + "nothing completed, so nothing may be popped" + ); + assert!( + elapsed >= bound.saturating_sub(Duration::from_millis(20)), + "pop_within({millis}ms) returned after {elapsed:?}, well short of its bound -- a \ + bound honoured by returning early is not honoured" + ); + assert!( + elapsed < bound + Duration::from_millis(400), + "pop_within({millis}ms) took {elapsed:?}, past its bound" + ); + } + + // The operation never completed, so the token is abandoned on purpose and + // the ring is left to its own teardown; calling `run_down` here would park + // forever, which is this responder working as described rather than a + // defect to route around. + std::mem::forget(token); + std::mem::forget(ring); + drop(guard); + drop(file); + let _ = std::fs::remove_file(&path); +} diff --git a/crates/windows-ioring-sys/tests/registration.rs b/crates/windows-ioring-sys/tests/registration.rs index d5df79761..54ce636aa 100644 --- a/crates/windows-ioring-sys/tests/registration.rs +++ b/crates/windows-ioring-sys/tests/registration.rs @@ -12,6 +12,18 @@ use windows_ioring_sys::{ Batch, IoBuf, IoBufMut, IoRing, PushOptions, RegisteredSpan, SharedFile, WriteCaching, }; +/// How long a completion this test caused is allowed to take to arrive. +/// +/// M26.7 replaced a `try_pop` here that asserted the completion was *already* +/// queued when `submit_and_wait` returned. That is not something this crate +/// promises -- `pop_within`'s own documentation says a submit-side wait's +/// return "promises nothing about poppability", and `RESPONSE-SPACE.md` states +/// it as `RS-P-5`. It was true on the handle this test happens to use and is +/// false on others, which is the definition of a frozen observation. +/// +/// Stated as this crate's own contract instead: the completion arrives within +/// a bound we choose. Generous, because the bound is not what is under test. +const POP_BOUND: std::time::Duration = std::time::Duration::from_secs(10); /// Registration is the path where this crate hands the kernel a pointer and /// the kernel dereferences it *later* -- which is precisely how /// [D-32](../DESIGN-NOTES.md#d-32) shipped a use-after-free that surfaced as a @@ -66,9 +78,9 @@ fn a_read_addressing_a_registered_file_and_a_registered_buffer_round_trips() { unsafe { batch.register_files(&[handle]) }.expect("queue file registration"); batch.submit_and_wait(1, 5_000).expect("submit and wait"); let completion = ring - .try_pop() + .pop_within(POP_BOUND) .expect("pop completion") - .expect("a completion is ready"); + .expect("a completion arrives within the bound"); let registered_files = files_pending .claim_if(&completion) .expect("id matches") @@ -82,9 +94,9 @@ fn a_read_addressing_a_registered_file_and_a_registered_buffer_round_trips() { .expect("queue buffer registration"); batch.submit_and_wait(1, 5_000).expect("submit and wait"); let completion = ring - .try_pop() + .pop_within(POP_BOUND) .expect("pop completion") - .expect("a completion is ready"); + .expect("a completion arrives within the bound"); let mut registered_buffers = buffers_pending .claim_if(&completion) .expect("id matches") @@ -110,10 +122,14 @@ fn a_read_addressing_a_registered_file_and_a_registered_buffer_round_trips() { batch.submit_and_wait(1, 5_000).expect("submit and wait"); let completion = ring - .try_pop() + .pop_within(POP_BOUND) .expect("pop completion") - .expect("a completion is ready"); + .expect("a completion arrives within the bound"); let transferred = completion.result().expect("registered read succeeded"); + // CONFIRMS: RS-P-8 -- a full count here is a property of the handle + // this test chose (an ordinary file on a local volume, where a + // successful completion carries the whole length and a full volume is + // an error instead), not of the space, which permits a short one. assert_eq!(transferred, 256); let _ = token .claim_if(&completion) @@ -140,9 +156,9 @@ fn a_read_addressing_a_registered_file_and_a_registered_buffer_round_trips() { .expect("queue another registered read"); batch.submit_and_wait(1, 5_000).expect("submit and wait"); let completion = ring - .try_pop() + .pop_within(POP_BOUND) .expect("pop completion") - .expect("a completion is ready"); + .expect("a completion arrives within the bound"); completion.result().expect("registered read succeeded"); let _ = token .claim_if(&completion) @@ -174,9 +190,9 @@ fn a_second_file_or_buffer_registration_on_the_same_ring_is_refused() { unsafe { batch.register_files(&[handle]) }.expect("queue first file registration"); batch.submit_and_wait(1, 5_000).expect("submit and wait"); let completion = ring - .try_pop() + .pop_within(POP_BOUND) .expect("pop completion") - .expect("a completion is ready"); + .expect("a completion arrives within the bound"); let _registered_files = files_pending .claim_if(&completion) .expect("id matches") @@ -196,9 +212,9 @@ fn a_second_file_or_buffer_registration_on_the_same_ring_is_refused() { .expect("queue first buffer registration"); batch.submit_and_wait(1, 5_000).expect("submit and wait"); let completion = ring - .try_pop() + .pop_within(POP_BOUND) .expect("pop completion") - .expect("a completion is ready"); + .expect("a completion arrives within the bound"); let registered_buffers = buffers_pending .claim_if(&completion) .expect("id matches") @@ -252,9 +268,9 @@ fn a_buffer_registration_survives_heap_churn_between_the_push_and_the_submit() { batch.submit_and_wait(1, 5_000).expect("submit and wait"); let completion = ring - .try_pop() + .pop_within(POP_BOUND) .expect("pop completion") - .expect("a completion is ready"); + .expect("a completion arrives within the bound"); let mut registered_buffers = pending .claim_if(&completion) .expect("id matches") @@ -273,9 +289,9 @@ fn a_buffer_registration_survives_heap_churn_between_the_push_and_the_submit() { .expect("queue read against the registered buffer"); batch.submit_and_wait(1, 5_000).expect("submit and wait"); let completion = ring - .try_pop() + .pop_within(POP_BOUND) .expect("pop completion") - .expect("a completion is ready"); + .expect("a completion arrives within the bound"); let bytes = completion.result().expect("read succeeded"); assert_eq!(bytes, 256); let _ = token.claim_if(&completion); @@ -361,9 +377,9 @@ fn a_zero_length_registration_does_not_spend_the_ring_s_one_registration() { .expect("a real registration must still be accepted after a zero-length one"); batch.submit_and_wait(1, 5_000).expect("submit and wait"); let completion = ring - .try_pop() + .pop_within(POP_BOUND) .expect("pop completion") - .expect("a completion is ready"); + .expect("a completion arrives within the bound"); let registered = pending .claim_if(&completion) .expect("id matches") @@ -431,9 +447,9 @@ fn dropping_a_registration_with_an_operation_in_flight_leaks_rather_than_frees() .expect("queue buffer registration"); batch.submit_and_wait(1, 5_000).expect("submit and wait"); let completion = ring - .try_pop() + .pop_within(POP_BOUND) .expect("pop completion") - .expect("a completion is ready"); + .expect("a completion arrives within the bound"); let registered_buffers = buffers_pending .claim_if(&completion) .expect("id matches") @@ -492,9 +508,9 @@ fn a_registered_file_from_a_different_ring_is_rejected() { unsafe { batch.register_files(&[handle]) }.expect("queue file registration on ring a"); batch.submit_and_wait(1, 5_000).expect("submit and wait"); let completion = ring_a - .try_pop() + .pop_within(POP_BOUND) .expect("pop completion") - .expect("a completion is ready"); + .expect("a completion arrives within the bound"); let registered_files = files_pending .claim_if(&completion) .expect("id matches") @@ -539,9 +555,9 @@ fn a_registered_file_is_readable_through_the_safe_api_without_unsafe() { unsafe { batch.register_files(&[handle]) }.expect("queue file registration"); batch.submit_and_wait(1, 5_000).expect("submit and wait"); let completion = ring - .try_pop() + .pop_within(POP_BOUND) .expect("pop completion") - .expect("a completion is ready"); + .expect("a completion arrives within the bound"); let registered_file = files_pending .claim_if(&completion) .expect("id matches") @@ -555,9 +571,9 @@ fn a_registered_file_is_readable_through_the_safe_api_without_unsafe() { .expect("queue a safe read against the registered file"); batch.submit_and_wait(1, 5_000).expect("submit and wait"); let completion = ring - .try_pop() + .pop_within(POP_BOUND) .expect("pop completion") - .expect("a completion is ready"); + .expect("a completion arrives within the bound"); assert_eq!(completion.result().expect("read succeeded"), 128); let (buffer, returned_file) = token .claim_if(&completion) @@ -592,9 +608,9 @@ fn a_registered_file_and_a_registered_buffer_compose_through_the_safe_api() { unsafe { batch.register_files(&[handle]) }.expect("queue file registration"); batch.submit_and_wait(1, 5_000).expect("submit and wait"); let completion = ring - .try_pop() + .pop_within(POP_BOUND) .expect("pop completion") - .expect("a completion is ready"); + .expect("a completion arrives within the bound"); let registered_file = files_pending .claim_if(&completion) .expect("id matches") @@ -608,9 +624,9 @@ fn a_registered_file_and_a_registered_buffer_compose_through_the_safe_api() { .expect("queue buffer registration"); batch.submit_and_wait(1, 5_000).expect("submit and wait"); let completion = ring - .try_pop() + .pop_within(POP_BOUND) .expect("pop completion") - .expect("a completion is ready"); + .expect("a completion arrives within the bound"); let mut registered_buffers = buffers_pending .claim_if(&completion) .expect("id matches") @@ -633,9 +649,9 @@ fn a_registered_file_and_a_registered_buffer_compose_through_the_safe_api() { .expect("queue a safe fully-registered read"); batch.submit_and_wait(1, 5_000).expect("submit and wait"); let completion = ring - .try_pop() + .pop_within(POP_BOUND) .expect("pop completion") - .expect("a completion is ready"); + .expect("a completion arrives within the bound"); assert_eq!(completion.result().expect("read succeeded"), 64); let (registered_use, returned_file) = token .claim_if(&completion) @@ -674,9 +690,9 @@ fn the_safe_api_rejects_a_registered_file_from_a_different_ring() { unsafe { batch.register_files(&[handle]) }.expect("queue file registration on ring a"); batch.submit_and_wait(1, 5_000).expect("submit and wait"); let completion = ring_a - .try_pop() + .pop_within(POP_BOUND) .expect("pop completion") - .expect("a completion is ready"); + .expect("a completion arrives within the bound"); let registered_file = files_pending .claim_if(&completion) .expect("id matches") @@ -718,9 +734,9 @@ fn a_registered_buffers_from_a_different_ring_is_rejected() { .expect("queue buffer registration on ring a"); batch.submit_and_wait(1, 5_000).expect("submit and wait"); let completion = ring_a - .try_pop() + .pop_within(POP_BOUND) .expect("pop completion") - .expect("a completion is ready"); + .expect("a completion arrives within the bound"); let registered_buffers = buffers_pending .claim_if(&completion) .expect("id matches") @@ -834,9 +850,9 @@ fn a_registered_buffer_can_be_filled_and_written_back_out() { .expect("queue buffer registration"); batch.submit_and_wait(1, 5_000).expect("submit and wait"); let completion = ring - .try_pop() + .pop_within(POP_BOUND) .expect("pop completion") - .expect("a completion is ready"); + .expect("a completion arrives within the bound"); let mut buffers = pending .claim_if(&completion) .expect("id matches") @@ -866,9 +882,9 @@ fn a_registered_buffer_can_be_filled_and_written_back_out() { .expect("queue registered write"); batch.submit_and_wait(1, 5_000).expect("submit and wait"); let completion = ring - .try_pop() + .pop_within(POP_BOUND) .expect("pop completion") - .expect("a completion is ready"); + .expect("a completion arrives within the bound"); assert_eq!( completion.result().expect("registered write succeeded"), record.len() @@ -907,9 +923,9 @@ fn get_mut_refuses_a_buffer_with_an_operation_outstanding_but_allows_its_neighbo .expect("queue buffer registration"); batch.submit_and_wait(1, 5_000).expect("submit and wait"); let completion = ring - .try_pop() + .pop_within(POP_BOUND) .expect("pop completion") - .expect("a completion is ready"); + .expect("a completion arrives within the bound"); let mut buffers = pending .claim_if(&completion) .expect("id matches") @@ -959,11 +975,15 @@ fn get_mut_refuses_a_buffer_with_an_operation_outstanding_but_allows_its_neighbo // Drain, claim, and the buffer becomes fillable again -- the count only // returns to zero against a real popped completion. - let mut popped = None; - while popped.is_none() { - popped = ring.try_pop().expect("pop completion"); - } - let completion = popped.expect("a completion is ready"); + // + // Bounded rather than spun (M26.7). This was an unbounded `try_pop` loop, + // which is the shape `IoRing::pop_within` was introduced to replace: it + // never states how long the operation is allowed to take, so a kernel that + // stopped completing turns a failing test into a hung one. + let completion = ring + .pop_within(POP_BOUND) + .expect("pop completion") + .expect("a completion arrives within the bound"); completion.result().expect("registered write succeeded"); let _ = token .claim_if(&completion) @@ -1005,9 +1025,9 @@ fn get_mut_yields_only_the_registered_bytes_and_cannot_move_the_allocation() { .expect("queue buffer registration"); batch.submit_and_wait(1, 5_000).expect("submit and wait"); let completion = ring - .try_pop() + .pop_within(POP_BOUND) .expect("pop completion") - .expect("a completion is ready"); + .expect("a completion arrives within the bound"); let mut buffers = pending .claim_if(&completion) .expect("id matches") @@ -1104,9 +1124,9 @@ fn get_refuses_a_buffer_a_read_is_landing_into_but_allows_one_a_write_is_reading .expect("queue buffer registration"); batch.submit_and_wait(1, 5_000).expect("submit and wait"); let completion = ring - .try_pop() + .pop_within(POP_BOUND) .expect("pop completion") - .expect("a completion is ready"); + .expect("a completion arrives within the bound"); let mut buffers = pending .claim_if(&completion) .expect("id matches") @@ -1198,3 +1218,27 @@ fn get_refuses_a_buffer_a_read_is_landing_into_but_allows_one_a_write_is_reading let _ = std::fs::remove_file(&path); } + +// ------------------------------------------------------------------------ +// Relocated from `src/batch/tests.rs` at 9bc0350e (M24.3). + +#[test] +fn the_debug_rendering_names_the_registration_and_its_identity() { + // `>::fmt -> Ok(Default::default())` + // survived: that mutation writes nothing to the formatter, so the + // rendering comes back empty regardless of what the registration holds. + let mut ring = IoRing::new(8, 8).expect("create ring"); + let mut batch = Batch::new(&mut ring); + let pending = batch + .register_buffers(vec![vec![0_u8; 64]]) + .expect("queue buffer registration"); + let rendering = format!("{pending:?}"); + assert!( + rendering.contains("PendingBufferRegistration"), + "got {rendering}" + ); + assert!( + rendering.contains(&pending.user_data().to_string()), + "the operation's identity must appear: {rendering}" + ); +} diff --git a/crates/windows-ioring-sys/tests/resolver_over_a_real_ring.rs b/crates/windows-ioring-sys/tests/resolver_over_a_real_ring.rs new file mode 100644 index 000000000..efa10fa41 --- /dev/null +++ b/crates/windows-ioring-sys/tests/resolver_over_a_real_ring.rs @@ -0,0 +1,490 @@ +// Copyright (c) Mike Grier +//! The response-space resolver, answering for a real `IoRing` (M26.3). +//! +//! # Why this is an integration test +//! +//! `M26.2`'s seam deliberately leaves the ring's *lifecycle* calls real +//! ([D-60](../DESIGN-NOTES.md#d-60)), so a resolver runs against a ring the +//! kernel actually created. That is an operating-system boundary, which this +//! repository's Quality rule places in `tests/` rather than in the lib suite +//! -- and `check-ring-tests.ps1` exists to keep that population from drifting +//! back the other way. +//! +//! # What this file is for, and what it is not +//! +//! The resolver's own unit tests drive [`Responses`] directly, because whether +//! it occupies the space it claims to is a property of the resolver alone. +//! What they cannot show is that it is **reachable** -- that installing one +//! actually diverts a real `IoRing`'s calls, that an operation the kernel +//! never saw still completes through the crate's ordinary accounting, and that +//! the ring afterwards runs down rather than blocking on work that does not +//! exist. +//! +//! That is this file. It is a reachability check, not a property suite; +//! `M26.4` is where the properties that must hold under *every* resolution are +//! written, and `M26.5` is where the resolution is calibrated by re-injecting +//! defects this crate actually shipped. + +#![cfg(all(windows, feature = "kernel-seam"))] + +use std::os::windows::io::AsRawHandle; + +use windows_ioring_sys::sys::{Resolver, ResolverConfig}; +use windows_ioring_sys::{Batch, FlushCoverage, FlushMode, IoRing, SharedFile}; + +/// A file to aim flushes at. Its contents never matter: under a resolver the +/// operation never reaches the kernel, and the point of the handle is that the +/// crate's own `Build*` path is exercised exactly as it would be otherwise. +fn scratch(tag: &str) -> (SharedFile, std::path::PathBuf) { + let path = std::env::temp_dir().join(format!( + "windows-ioring-sys-m26-3-{}-{tag}.tmp", + std::process::id() + )); + let file = std::fs::OpenOptions::new() + .create(true) + .truncate(true) + .write(true) + .open(&path) + .expect("a scratch file"); + assert!(!file.as_raw_handle().is_null()); + (SharedFile::new(file.into()), path) +} + +#[test] +fn an_installed_resolver_answers_a_real_rings_operations() { + // The narrowest point in the space: this asserts reachability, so every + // freedom that could make the outcome depend on the seed is off. A test + // about reachability that could fail for an ordering reason would be two + // tests wearing one name. + let resolver = Resolver::with_config(0x5EED, ResolverConfig::narrowest()); + let replay = resolver.replay_hint(); + + let (path, outcome) = resolver.scoped(|watch| { + // Created *inside* the scope, so the ring's rundown is answered by the + // resolver that owns its operations. Outside it, rundown would ask a + // real ring to wait for completions the kernel has no record of. + let mut ring = IoRing::new(64, 128).expect("a ring"); + let (file, path) = scratch("reachable"); + + let token = { + let mut batch = Batch::new(&mut ring); + let token = batch + .flush(&file, FlushCoverage::Unordered, FlushMode::Default) + .expect("a flush builds"); + batch.submit().expect("the submit is answered"); + token + }; + + assert_eq!( + ring.outstanding(), + 1, + "the crate's accounting counts the operation whether the kernel saw it or not" + ); + + let completion = ring + .pop_within(std::time::Duration::from_secs(5)) + .expect("the pop is answered") + .expect("a completion arrives within the bound"); + assert!( + token.claim_if(&completion).is_ok(), + "RS-C-2: the completion must identify the operation that produced it" + ); + assert_eq!( + ring.outstanding(), + 0, + "and the crate's accounting must clear on a resolved completion" + ); + + ring.run_down().expect("a ring under a resolver runs down"); + (path, watch.stats()) + }); + let _ = std::fs::remove_file(path); + + // The assertion that makes this a reachability test rather than a ring + // test: had the seam not diverted the call, the kernel would have answered + // and these counters would all be zero while every assertion above still + // passed. + assert_eq!( + outcome.built, 1, + "{replay}: the resolver did not see the Build* call, so the seam did not divert it" + ); + assert_eq!(outcome.submitted, 1, "{replay}: nor the submit"); + assert_eq!( + outcome.posted, 1, + "{replay}: nor did it post the completion" + ); + assert_eq!(outcome.popped, 1, "{replay}: nor hand it over"); +} + +#[test] +fn a_ring_runs_down_under_the_widest_resolution() { + // Rundown is where a resolver that fails to satisfy RS-C-1 stops being a + // failing test and becomes a hung one: `IoRing::run_down` loops while + // anything is outstanding. Driving it at the widest point -- operations + // pend, completions reorder, individual operations fail, waits expire and + // wake empty -- is the case where a resolver that made no progress would + // park here. + // + // Swept over seeds rather than run once, because "it terminated" at one + // seed says nothing about a space whose whole purpose is that the choices + // differ. + // + // RS-P-7 IS DELIBERATELY NARROWED OFF, AND THAT IS A FINDING RATHER THAN + // A CONVENIENCE. At the genuinely widest point this test fails: a submit + // declined under RS-P-7 propagates out of `run_down` as an error, leaving + // `outstanding() > 0`, after which `Drop` asserts and calls `CloseIoRing` + // anyway -- the exact hazard M21.6 fixed for `ERROR_TIMEOUT`, reachable + // again through a different HRESULT. Measured: with this one permission + // off, all 32 seeds pass; with it on, seed 0x1A fails at `0x80070008`. + // + // That is not fixed here, because the fix is a decision rather than a + // correction: `run_down`'s own documentation argues that blocking is the + // safe failure mode and closing early is not, which says it should keep + // looping -- but a permanently failing submit then hangs, and "no hang" is + // one of the properties M26.4 is about to write. The two pull opposite + // ways and an engineer chooses. Queued as `M26.8`. + for seed in 0..32_u64 { + let resolver = Resolver::with_config( + seed, + ResolverConfig { + may_fail_submits: false, + ..ResolverConfig::default() + }, + ); + let replay = resolver.replay_hint(); + let path = resolver.scoped(|watch| { + let mut ring = IoRing::new(64, 128).expect("a ring"); + let (file, path) = scratch("rundown"); + + let mut tokens = Vec::new(); + let mut batch = Batch::new(&mut ring); + for _ in 0..6 { + tokens.push( + batch + .flush(&file, FlushCoverage::Unordered, FlushMode::Default) + .expect("a flush builds"), + ); + } + batch.submit().expect("the submit is answered"); + + ring.run_down() + .unwrap_or_else(|error| panic!("{replay}: rundown failed: {error}")); + assert_eq!( + ring.outstanding(), + 0, + "{replay}: rundown returned with work still outstanding" + ); + + let stats = watch.stats(); + assert_eq!( + stats.built, 6, + "{replay}: the resolver must have seen every build" + ); + assert_eq!( + stats.posted, 6, + "{replay}: RS-C-1 -- every submitted operation completes exactly once" + ); + path + }); + let _ = std::fs::remove_file(path); + } +} + +#[test] +fn a_declined_submit_leaves_the_ring_resumable_and_the_policy_to_the_caller() { + // What M26.8 settled, replacing the test that pinned the old behaviour. + // + // `SubmitIoRing` documents that an error other than IORING_E_WAIT_TIMEOUT + // leaves **all entries in the submission queue**. So a declined submit has + // not lost the operations, and rundown reporting the error is correct -- + // what was missing was the guarantee that makes reporting it useful: the + // ring is resumable, and calling again is what runs the queued entries. + // + // Deciding *when* to call again is deliberately not this crate's business. + // This test therefore plays the caller: it sees the error, chooses to try + // again, and the work completes. + // + // Seed 0x1A is the one M26.3's sweep found. Pinned rather than swept + // because the point is reproducing one observation exactly. + let resolver = Resolver::new(0x1A); + let replay = resolver.replay_hint(); + let (path, refusals, finished) = resolver.scoped(|_| { + let mut ring = IoRing::new(64, 128).expect("a ring"); + let (file, path) = scratch("declined"); + + let mut batch = Batch::new(&mut ring); + for _ in 0..6 { + let _token = batch + .flush(&file, FlushCoverage::Unordered, FlushMode::Default) + .expect("a flush builds"); + } + let _ = batch.submit(); + + // The caller's policy, which is all this crate asks of it: keep + // going while it chooses to. A real consumer would back off here; the + // point is that it is *their* loop and not ours. + let mut refusals = 0_usize; + let mut finished = false; + for _ in 0..64 { + match ring.run_down_within(std::time::Duration::from_millis(50)) { + Ok(true) => { + finished = true; + break; + } + Ok(false) => {} + Err(_) => refusals += 1, + } + } + (path, refusals, finished) + }); + let _ = std::fs::remove_file(path); + + assert!( + refusals > 0, + "{replay}: this seed declines a submit, which is the condition under test" + ); + assert!( + finished, + "{replay}: after a declined submit the entries remain queued, so a caller that tries \ + again must be able to finish -- that is the guarantee SubmitIoRing's documentation \ + gives and what makes reporting the error useful rather than terminal" + ); +} + +#[test] +fn an_expired_wait_is_a_successful_submit() { + // M26.8's correction, and the reason it is not a judgement call: + // `SubmitIoRing` documents IORING_E_WAIT_TIMEOUT as "All operations were + // submitted without error and the subsequent wait timed out". Reporting + // that as an Err is M21.6's defect at the one site that sweep missed -- + // and the damage is not merely a wrong sign, because an Err from a submit + // means the entries are still queued, so a caller who frees their buffers + // on seeing one hands the kernel freed memory next time. + // + // RS-P-4 makes an expired wait reachable on demand, so this is a test + // rather than an argument about a rare timing. Swept rather than pinned to + // one seed: the clause is a permission the resolver takes sometimes, and + // picking a seed that happens to take it would make the test a hostage to + // the mixer. Every seed that expires must report success. + let mut expiring_seeds = 0_usize; + for seed in 0..64_u64 { + let resolver = Resolver::with_config( + seed, + ResolverConfig { + may_expire_waits: true, + may_pend: true, + ..ResolverConfig::narrowest() + }, + ); + let replay = resolver.replay_hint(); + let path = resolver.scoped(|watch| { + let mut ring = IoRing::new(64, 128).expect("a ring"); + let (file, path) = scratch("expired-submit"); + + let outcome = { + let mut batch = Batch::new(&mut ring); + for _ in 0..4 { + let _token = batch + .flush(&file, FlushCoverage::Unordered, FlushMode::Default) + .expect("a flush builds"); + } + // Ask to wait, which is what lets RS-P-4 apply. + batch.submit_and_wait(4, 50) + }; + + if watch.stats().expired_waits > 0 { + expiring_seeds += 1; + assert!( + outcome.is_ok(), + "{replay}: an expired wait means every entry was submitted, so \ + submit_and_wait must report success -- see SubmitIoRing's documented \ + return values. Got {outcome:?}" + ); + } + + // Whatever the wait did, the operations were submitted -- so they + // run down normally rather than needing recovery. + ring.run_down().expect("rundown"); + path + }); + let _ = std::fs::remove_file(path); + } + + assert!( + expiring_seeds > 0, + "no seed expired a wait, so this test checked nothing about IORING_E_WAIT_TIMEOUT" + ); +} + +#[test] +fn run_down_within_honours_its_bound_and_reports_rather_than_deciding() { + // The shape the audit found wrong: `run_down` waited in segments with no + // period to sit inside, which made "how long to keep trying" this crate's + // policy. The bounded form hands that back. + // + // A resolution that has not completed the work yet is the case where the + // bound has anything to do, so RS-P-1 supplies one. + let resolver = Resolver::with_config( + 11, + ResolverConfig { + may_pend: true, + ..ResolverConfig::narrowest() + }, + ); + let replay = resolver.replay_hint(); + let path = resolver.scoped(|_| { + let mut ring = IoRing::new(64, 128).expect("a ring"); + let (file, path) = scratch("bounded-rundown"); + { + let mut batch = Batch::new(&mut ring); + for _ in 0..4 { + let _token = batch + .flush(&file, FlushCoverage::Unordered, FlushMode::Default) + .expect("a flush builds"); + } + batch.submit().expect("the submit is answered"); + } + + // A zero bound is the honest spelling of "do not wait": it reports + // rather than blocking, which is the whole point of the shape. + let started = std::time::Instant::now(); + let finished = ring + .run_down_within(std::time::Duration::ZERO) + .expect("a zero-bound rundown does not fail"); + assert!( + started.elapsed() < std::time::Duration::from_millis(500), + "{replay}: a zero bound must not block" + ); + assert!( + !finished || ring.outstanding() == 0, + "{replay}: reporting finished must mean nothing is outstanding" + ); + + // And the caller's own loop finishes it, because that is their policy. + for _ in 0..256 { + if ring + .run_down_within(std::time::Duration::from_millis(20)) + .expect("rundown") + { + break; + } + } + assert_eq!( + ring.outstanding(), + 0, + "{replay}: a caller who keeps calling reaches quiescence" + ); + path + }); + let _ = std::fs::remove_file(path); +} + +#[test] +fn a_pending_completion_defeats_try_pop_and_not_pop_within() { + // The defect class M26.7 audited, demonstrated rather than described. + // + // Thirty-one assertions across five kernel tests read `try_pop()` straight + // after `submit_and_wait` and expected a completion to be *there*. That is + // not something this crate promises -- `pop_within`'s own documentation + // says a submit-side wait's return "promises nothing about poppability", + // and RS-P-5 states it as a permission the platform holds. The assertions + // passed anyway, because on a buffered handle the operation completes + // inside the submit; D-40 measured that at 80 of 80 attempts. Change the + // handle and the same assertion gives the opposite answer, which is what + // makes it a frozen observation rather than a contract. + // + // A resolver makes the pending case reachable on demand, so the difference + // between the two spellings is a test rather than an argument. This is the + // guard for the restatement: if `try_pop` ever starts satisfying this, the + // premise of the audit was wrong and this test says so. + let resolver = Resolver::with_config( + 0x9, + ResolverConfig { + may_pend: true, + ..ResolverConfig::narrowest() + }, + ); + let replay = resolver.replay_hint(); + + let path = resolver.scoped(|_| { + let mut ring = IoRing::new(64, 128).expect("a ring"); + let (file, path) = scratch("pending"); + + let token = { + let mut batch = Batch::new(&mut ring); + let token = batch + .flush(&file, FlushCoverage::Unordered, FlushMode::Default) + .expect("a flush builds"); + batch.submit().expect("the submit is answered"); + token + }; + + // The frozen-observation spelling. Under a resolution that pends, the + // completion is not there yet -- so a test written this way would have + // failed here rather than at anything it meant to check. + assert!( + ring.try_pop().expect("try_pop").is_none(), + "{replay}: this resolution pends, so nothing is poppable the instant the submit \ + returns -- if that changed, this test's premise is gone" + ); + assert_eq!( + ring.outstanding(), + 1, + "{replay}: the operation is outstanding, not lost" + ); + + // The restated spelling: this crate's own contract, which holds under + // every resolution rather than on one kind of handle. + let completion = ring + .pop_within(std::time::Duration::from_secs(5)) + .expect("pop_within") + .expect("a completion arrives within the bound"); + assert!( + token.claim_if(&completion).is_ok(), + "{replay}: the completion identifies its operation" + ); + + ring.run_down().expect("rundown"); + path + }); + let _ = std::fs::remove_file(path); +} + +#[test] +fn a_thread_with_nothing_installed_still_talks_to_the_kernel() { + // The seam's transparency, asserted from the far side. This is the same + // property `M26.2` verified by running its suite both ways, restated here + // because a resolver is the thing most likely to break it: a responder + // that leaked past its guard would divert every ring in the process, and + // the failure would look like a flaky kernel rather than like a harness + // defect. + let (file, path) = scratch("uninstalled"); + let mut ring = IoRing::new(64, 128).expect("a ring"); + + { + let resolver = Resolver::with_config(1, ResolverConfig::narrowest()); + let watch = resolver.scoped(|watch| watch.clone()); + // The guard has dropped. Nothing installed on this thread now. + assert_eq!(watch.stats().built, 0, "the scope did nothing of its own"); + + let mut batch = Batch::new(&mut ring); + let token = batch + .flush(&file, FlushCoverage::Unordered, FlushMode::Default) + .expect("a flush builds"); + batch.submit().expect("the kernel accepts the submit"); + let completion = ring + .pop_within(std::time::Duration::from_secs(5)) + .expect("the kernel answers") + .expect("a real completion arrives"); + assert!(token.claim_if(&completion).is_ok()); + assert_eq!( + watch.stats().built, + 0, + "an uninstalled resolver must not have seen the call -- the seam leaked" + ); + } + + ring.run_down().expect("rundown"); + drop(file); + let _ = std::fs::remove_file(path); +} diff --git a/crates/windows-ioring-sys/tests/response_space_census.rs b/crates/windows-ioring-sys/tests/response_space_census.rs new file mode 100644 index 000000000..71cd6cb27 --- /dev/null +++ b/crates/windows-ioring-sys/tests/response_space_census.rs @@ -0,0 +1,374 @@ +// Copyright (c) Mike Grier +//! Every clause in the response space is checked by something (M26.6). +//! +//! # What this is for +//! +//! `M26` split one job in two. The resolver sweeps the *permissions* -- the +//! `RS-P-n` clauses, which say what a platform may do -- and the kernel tests +//! confirm that a real Windows stays inside the *constraints*, the `RS-C-n` +//! clauses, which say what this crate requires of it. That division is stated +//! in [RESPONSE-SPACE.md](../RESPONSE-SPACE.md), and `RS-C-4` in particular is +//! justified there **on the strength of these tests existing**: the resolver is +//! forbidden to break the drain half of `DRAIN_PRECEDING_OPS`, so a Windows +//! that broke it would be caught by nothing the resolver does. +//! +//! A division of labour recorded only in prose is enforced by whoever +//! remembers it, which over a long change is nobody. This file is the rung +//! below prose: it reads the space, finds every clause, and fails when one is +//! cited by nothing on the side that owes it a check. +//! +//! # Why a marker, and not a search for the clause ID +//! +//! The first version of this census searched each file for the clause ID +//! anywhere in its text, and it was **measured green while broken**. Removing +//! `RS-C-4`'s check from the only test that performs it did not turn it red, +//! because a second file mentioned that clause only to say the check was +//! somebody else's -- and a disclaimer reads identically to a claim under a +//! substring search. That is the same trap this repository already recorded +//! once, where a bare substring matched a probe whose only mention of a tag +//! was a comment. +//! +//! So a claim is now a **structured marker** -- `CONFIRMS:` for a constraint +//! checked against a real kernel, `EXERCISES:` for a permission the resolver +//! takes -- and prose mentioning a clause means nothing. The markers are +//! verified to go red in both directions before being trusted. +//! +//! # Why a census over source is sound here, when it usually is not +//! +//! This repository has been burned by proxies: a test that walked `src/bin` +//! and grepped for a substring was replaced because emitting a row is a +//! property of a program's *output*, which no read of its source can +//! establish. The distinction is that **the claim here is itself about the +//! source**. "Does a kernel test claim this clause" is a fact about what is +//! written in `tests/`, so reading `tests/` is the direct measurement rather +//! than a stand-in for one. +//! +//! What that buys is narrow, and saying so is the point: this proves a clause +//! is *claimed* by a file, never that the file's assertions are adequate. That +//! second question is what calibration is for -- `M26.5` re-injected two real +//! defects to show the instruments go red -- and no census can answer it. + +#![cfg(windows)] + +use std::collections::{BTreeMap, BTreeSet}; +use std::path::PathBuf; + +/// Marker for "this file confirms a real kernel stays inside this clause". +const CONFIRMS: &str = "CONFIRMS: "; +/// Marker for "this file exercises this freedom of the resolver". +const EXERCISES: &str = "EXERCISES: "; + +/// The crate root, so this reads the same files whatever the working +/// directory is -- including the scratch copy `cargo-mutants` and the sabotage +/// harness build from. +fn crate_root() -> PathBuf { + PathBuf::from(env!("CARGO_MANIFEST_DIR")) +} + +/// Clause IDs declared in the space, read from their headings. +/// +/// Derived from the document rather than listed here, so a clause added to the +/// space is immediately owed a check rather than silently exempt. That is the +/// same rule the resolver follows against the same document. +fn clauses_in_the_space() -> BTreeSet { + let path = crate_root().join("RESPONSE-SPACE.md"); + let text = std::fs::read_to_string(&path) + .unwrap_or_else(|error| panic!("reading {}: {error}", path.display())); + text.lines() + .filter_map(|line| { + let heading = line.strip_prefix("### ")?; + let id = heading.split_whitespace().next()?; + (id.starts_with("RS-P-") || id.starts_with("RS-C-")).then(|| id.to_owned()) + }) + .collect() +} + +/// Every clause claimed by `marker` in `text`. +/// +/// A marker must be the whole of what follows it on its line, so a sentence +/// that happens to contain the word cannot become a claim by accident. +fn claims(text: &str, marker: &str) -> BTreeSet { + text.lines() + .filter_map(|line| { + let at = line.find(marker)?; + let id = line[at + marker.len()..].trim(); + (!id.is_empty() && id.chars().all(|ch| ch.is_ascii_alphanumeric() || ch == '-')) + .then(|| id.to_owned()) + }) + .collect() +} + +/// Every `.rs` file directly under `tests/`, with its text. +fn test_files() -> BTreeMap { + let dir = crate_root().join("tests"); + std::fs::read_dir(&dir) + .unwrap_or_else(|error| panic!("reading {}: {error}", dir.display())) + .filter_map(Result::ok) + .map(|entry| entry.path()) + .filter(|path| path.extension().is_some_and(|ext| ext == "rs")) + .map(|path| { + let name = path + .file_name() + .and_then(|name| name.to_str()) + .expect("a UTF-8 file name") + .to_owned(); + let text = std::fs::read_to_string(&path) + .unwrap_or_else(|error| panic!("reading {}: {error}", path.display())); + (name, text) + }) + .collect() +} + +/// The resolver's own unit tests, which live in `src/` and carry the +/// `EXERCISES:` markers. +fn resolver_unit_tests() -> String { + let path = crate_root() + .join("src") + .join("sys") + .join("resolver") + .join("tests.rs"); + std::fs::read_to_string(&path) + .unwrap_or_else(|error| panic!("reading {}: {error}", path.display())) +} + +/// This file, whose own text names the markers and must never be counted as +/// claiming anything. +const SELF: &str = "response_space_census.rs"; + +#[test] +fn every_constraint_is_confirmed_against_a_real_kernel() { + // RS-C-n says what this crate requires of the platform. The resolver is + // forbidden to violate these, so the resolver can never be the thing that + // checks them -- only a test against a real ring can. + let files = test_files(); + let constraints: Vec = clauses_in_the_space() + .into_iter() + .filter(|id| id.starts_with("RS-C-")) + .collect(); + + // Vacuity guard, and not a formality: a parse that silently found nothing + // would make every assertion below pass while checking no clause at all. + assert!( + constraints.len() >= 4, + "only {} constraints parsed out of RESPONSE-SPACE.md, which means the heading format \ + changed and this census is reading nothing", + constraints.len() + ); + + let confirmed: BTreeMap> = files + .iter() + .filter(|(name, _)| name.as_str() != SELF) + .fold(BTreeMap::new(), |mut acc, (name, text)| { + for clause in claims(text, CONFIRMS) { + acc.entry(clause).or_default().push(name.clone()); + } + acc + }); + + let unchecked: Vec<&String> = constraints + .iter() + .filter(|clause| !confirmed.contains_key(*clause)) + .collect(); + + assert!( + unchecked.is_empty(), + "these constraints carry no CONFIRMS marker in any kernel test, so nothing confirms \ + Windows stays inside them: {unchecked:?}\n\ + RESPONSE-SPACE.md justifies constraining the resolver away from these on the grounds \ + that the kernel tests cover them. A constraint checked on neither side is untested in \ + both halves at once, which is the hole M26.6 exists to close.\n\ + Claimed today: {confirmed:?}" + ); +} + +#[test] +fn every_permission_is_exercised_by_a_resolver_test() { + // The other direction, and the reason to have both: a permission the + // resolver never takes is indistinguishable from one it never implemented, + // and a suite written against it would be hardened for nothing. + let permissions: Vec = clauses_in_the_space() + .into_iter() + .filter(|id| id.starts_with("RS-P-")) + .collect(); + assert!( + permissions.len() >= 7, + "only {} permissions parsed out of RESPONSE-SPACE.md; the heading format changed", + permissions.len() + ); + + let mut exercised = claims(&resolver_unit_tests(), EXERCISES); + for (name, text) in test_files() { + if name.as_str() != SELF { + exercised.extend(claims(&text, EXERCISES)); + } + } + + let unexercised: Vec<&String> = permissions + .iter() + .filter(|clause| !exercised.contains(*clause)) + .collect(); + + assert!( + unexercised.is_empty(), + "these permissions carry no EXERCISES marker: {unexercised:?}\n\ + A clause the resolver never exercises is permitted on paper and implemented nowhere, \ + which is the failure RESPONSE-SPACE.md's clause IDs exist to make visible.\n\ + Claimed today: {exercised:?}" + ); +} + +#[test] +fn a_marker_is_a_claim_and_a_mention_is_not() { + // The distinction this census was rebuilt around, asserted rather than + // trusted -- because the version that conflated the two was green while + // failing to detect a removed check. + let disclaimer = "//! The drain clause RS-C-4 is somebody else's job entirely."; + assert!( + claims(disclaimer, CONFIRMS).is_empty(), + "prose naming a clause must not count as claiming it" + ); + + let claim = "//! CONFIRMS: RS-C-4"; + assert_eq!( + claims(claim, CONFIRMS), + BTreeSet::from(["RS-C-4".to_owned()]), + "a marker line must be read as a claim" + ); + + // A marker with trailing prose is not a claim either: allowing it would + // let "CONFIRMS: RS-C-4 (eventually, once someone writes it)" pass. + let hedged = "//! CONFIRMS: RS-C-4 eventually"; + assert!( + claims(hedged, CONFIRMS).is_empty(), + "a marker must name a clause and nothing else" + ); +} + +#[test] +fn no_kernel_test_asserts_a_completion_is_already_poppable() { + // The defect class M26.7 audited, guarded so it cannot return. + // + // `try_pop()` immediately after a submit, with the `Option` unwrapped, + // asserts that the kernel has *already* queued the completion. RS-P-5 + // permits it not to have, and `pop_within`'s documentation says the same + // in this crate's own words -- so the assertion is about one kind of + // handle rather than about anything promised. D-40 measured why it passed + // regardless: a buffered read completes inside the submit in 80 of 80 + // attempts, while an unbuffered one genuinely pends. + // + // Thirty-one of these were restated as `pop_within` in M26.7. Nothing + // prevented them being written, and nothing would prevent the next one, + // because on the handles these tests use the assertion is simply true. + // That is what this census is for: the shape is refused at the source, + // since no run can be relied on to object to it. + // + // Resolver-driven tests are exempt, and one of them uses the shape on + // purpose -- `a_pending_completion_defeats_try_pop_and_not_pop_within` + // demonstrates the failure this rule exists to prevent, which it can only + // do by writing it. + let shape = regex_lite_matches; + let mut offenders = Vec::new(); + for (name, text) in test_files() { + if name.as_str() == SELF || text.contains("Resolver") { + continue; + } + let hits = shape(&text); + if hits > 0 { + offenders.push(format!("{name}: {hits}")); + } + } + + assert!( + offenders.is_empty(), + "these kernel tests assert a completion is poppable the instant a submit returns: \ + {offenders:?}\n\ + That is RS-P-5's freedom being treated as a guarantee. State it as this crate's own \ + contract instead -- `pop_within(bound)` -- which holds on every handle rather than on \ + the one the test happens to open." + ); +} + +/// Count `try_pop()` occurrences whose `Option` is unwrapped, which is the +/// "already poppable" assertion. +/// +/// Hand-rolled rather than pulled in as a dependency: the shape is two method +/// calls in sequence, and a scanner for it is shorter than the argument for +/// adding a regex crate to a test. +fn regex_lite_matches(text: &str) -> usize { + let mut count = 0; + let mut rest = text; + while let Some(at) = rest.find("try_pop()") { + rest = &rest[at + "try_pop()".len()..]; + // Look at the next two chained calls, skipping whitespace and dots. + let tail: String = rest.chars().take(200).collect(); + let mut calls = tail + .split('.') + .skip(1) + .map(|piece| piece.trim_start()) + .filter(|piece| !piece.is_empty()); + let first = calls.next().unwrap_or_default(); + let second = calls.next().unwrap_or_default(); + let unwraps = |call: &str| call.starts_with("expect(") || call.starts_with("unwrap()"); + if unwraps(first) && unwraps(second) { + count += 1; + } + } + count +} + +#[test] +fn the_already_poppable_scanner_recognises_the_shape_and_nothing_else() { + // The scanner decides what the census above means, so it is checked in + // both directions. A scanner that matched nothing would leave that test + // permanently, silently green. + assert_eq!( + regex_lite_matches("let c = ring.try_pop().expect(\"pop\").expect(\"ready\");"), + 1, + "the two-unwrap shape is the assertion being refused" + ); + assert_eq!( + regex_lite_matches("let c = ring.try_pop().unwrap().unwrap();"), + 1, + "unwrap spells the same assertion as expect" + ); + // One unwrap is the honest form: it takes the `Result` and leaves the + // `Option` for the caller to handle, which is what "empty at this instant" + // means. + assert_eq!( + regex_lite_matches("while let Some(c) = ring.try_pop().expect(\"pop\") { }"), + 0, + "handling the Option rather than unwrapping it is not the refused shape" + ); + assert_eq!( + regex_lite_matches("let maybe = ring.try_pop()?;"), + 0, + "propagating the Result is not the refused shape either" + ); +} +#[test] +fn every_marker_names_a_clause_the_space_declares() { + // A clause renamed in the document and not in the markers would leave the + // two censuses above checking a clause that no longer exists, and passing. + let declared = clauses_in_the_space(); + let mut sources = test_files(); + sources.insert("resolver/tests.rs".to_owned(), resolver_unit_tests()); + + let mut dangling: BTreeSet = BTreeSet::new(); + for (name, text) in &sources { + if name.as_str() == SELF { + continue; + } + for marker in [CONFIRMS, EXERCISES] { + for id in claims(text, marker) { + if !declared.contains(&id) { + dangling.insert(format!("{name}: {marker}{id}")); + } + } + } + } + + assert!( + dangling.is_empty(), + "these markers name clauses RESPONSE-SPACE.md does not declare: {dangling:?}" + ); +} diff --git a/crates/windows-ioring-sys/tests/ring_lifecycle.rs b/crates/windows-ioring-sys/tests/ring_lifecycle.rs new file mode 100644 index 000000000..9c002598d --- /dev/null +++ b/crates/windows-ioring-sys/tests/ring_lifecycle.rs @@ -0,0 +1,42 @@ +// Copyright (c) 2026 Mike Grier +//! A ring's own lifecycle, through public API only (M24.3). +//! +//! Relocated from `src/ring/tests.rs` at 9bc0350e. These open a real kernel ring, +//! which the repository's Quality rule classifies as an external boundary, so +//! `src/` was never where they belonged -- see +//! [D-49](../DESIGN-NOTES.md#d-49). They reach nothing crate-private, which is +//! what made them relocatable where the other 39 ring-opening lib tests are not. +//! +//! Pure relocation: the bodies are unchanged. + +use windows_ioring_sys::{IoRing, RingVersion, capabilities}; + +#[test] +fn a_ring_negotiates_a_version_no_higher_than_the_hosts_maximum() { + let ring = IoRing::new(64, 128).expect("create ring"); + let caps = capabilities().expect("capabilities"); + assert!(ring.version() <= caps.max_version); + assert!(ring.version() <= RingVersion::HIGHEST_KNOWN); +} + +#[test] +fn a_negotiated_ring_reports_its_version_back_through_get_ring_info() { + let ring = IoRing::new(64, 128).expect("create ring"); + let info = ring.info().expect("GetIoRingInfo"); + assert_eq!(info.version, ring.version()); +} + +#[test] +fn run_down_is_a_no_op_when_nothing_is_outstanding() { + let mut ring = IoRing::new(64, 128).expect("create ring"); + ring.run_down().expect("run_down with nothing outstanding"); + assert_eq!(ring.outstanding(), 0); +} + +#[test] +fn dropping_a_ring_with_nothing_outstanding_does_not_hang() { + // The ordinary path: no tokens were ever minted, so Drop's run_down must + // return immediately rather than waiting on SubmitIoRing at all. + let ring = IoRing::new(64, 128).expect("create ring"); + drop(ring); +} diff --git a/crates/windows-ioring-sys/tests/submission_lifecycle.rs b/crates/windows-ioring-sys/tests/submission_lifecycle.rs index 638eee326..0a496a514 100644 --- a/crates/windows-ioring-sys/tests/submission_lifecycle.rs +++ b/crates/windows-ioring-sys/tests/submission_lifecycle.rs @@ -10,11 +10,23 @@ use std::path::PathBuf; use windows_ioring_sys::contract::RingContract; use windows_ioring_sys::{ - Batch, FlushCoverage, FlushMode, IoRing, IoRingErrorExt, PushOptions, RingCondition, - SharedFile, Token, + Batch, FlushCoverage, FlushMode, IoBuf, IoBufMut, IoRing, IoRingErrorExt, PushOptions, + RingCondition, SharedFile, Token, WriteCaching, }; -use windows_sys::Win32::Foundation::ERROR_NOT_FOUND; +use windows_sys::Win32::Foundation::{ERROR_NOT_FOUND, HANDLE}; +/// How long a completion this test caused is allowed to take to arrive. +/// +/// M26.7 replaced a `try_pop` here that asserted the completion was *already* +/// queued when `submit_and_wait` returned. That is not something this crate +/// promises -- `pop_within`'s own documentation says a submit-side wait's +/// return "promises nothing about poppability", and `RESPONSE-SPACE.md` states +/// it as `RS-P-5`. It was true on the handle this test happens to use and is +/// false on others, which is the definition of a frozen observation. +/// +/// Stated as this crate's own contract instead: the completion arrives within +/// a bound we choose. Generous, because the bound is not what is under test. +const POP_BOUND: std::time::Duration = std::time::Duration::from_secs(10); const CHUNKS: usize = 8; const CHUNK_LEN: usize = 512; @@ -93,6 +105,10 @@ fn many_reads_round_trip_every_user_data_and_buffer() { .claim_if(&completion) .expect("a token claims its own completion"); contract.observe_claim(user_data); + // CONFIRMS: RS-P-8 -- a full count here is a property of the handle + // this test chose (an ordinary file on a local volume, where a + // successful completion carries the whole length and a full volume is + // an error instead), not of the space, which permits a short one. assert_eq!(transferred, CHUNK_LEN); assert_eq!( buffer, @@ -186,9 +202,9 @@ fn pushing_past_submission_queue_capacity_reports_backpressure_and_the_ring_stay contract.observe_tokenless_push(user_data); batch.submit_and_wait(1, 5_000).expect("submit and wait"); let completion = ring - .try_pop() + .pop_within(POP_BOUND) .expect("pop completion") - .expect("a completion is ready"); + .expect("a completion arrives within the bound"); assert_eq!(completion.user_data(), user_data); contract.observe_completion(completion.user_data()); completion.result().expect("flush succeeded"); @@ -225,9 +241,9 @@ fn a_dropped_batch_still_submits_its_queued_operations() { .expect("submit and wait"); let completion = ring - .try_pop() + .pop_within(POP_BOUND) .expect("pop completion") - .expect("a completion is ready"); + .expect("a completion arrives within the bound"); assert_eq!(completion.user_data(), user_data); assert_eq!(completion.result().expect("read succeeded"), content.len()); let buffer = token @@ -257,9 +273,9 @@ fn cancelling_a_target_that_is_not_outstanding_reports_error_not_found_through_c }; let completion = ring - .try_pop() + .pop_within(POP_BOUND) .expect("pop completion") - .expect("a completion is ready"); + .expect("a completion arrives within the bound"); assert_eq!(completion.user_data(), cancel_user_data); let error = completion .result() @@ -302,9 +318,9 @@ fn dropping_the_callers_own_sharedfile_clone_does_not_close_a_still_outstanding_ .submit_and_wait(1, 5_000) .expect("submit and wait"); let completion = ring - .try_pop() + .pop_within(POP_BOUND) .expect("pop completion") - .expect("a completion is ready"); + .expect("a completion arrives within the bound"); assert_eq!( completion .result() @@ -316,3 +332,77 @@ fn dropping_the_callers_own_sharedfile_clone_does_not_close_a_still_outstanding_ .expect("token claims its own completion"); assert_eq!(buffer, content); } + +// ------------------------------------------------------------------------ +// Relocated from `src/batch/tests.rs` at 9bc0350e (M24.3). `HugeBuffer` and +// `NULL_FILE` came with them: they were used by these two tests and +// nothing else. + +/// A buffer that claims a length no real allocation could ever have, to +/// exercise `checked_len`'s rejection without needing a real file: the +/// rejection must happen before the buffer's pointer is ever read. +struct HugeBuffer; + +// SAFETY: `stable_ptr`/`stable_mut_ptr` are never dereferenced in the tests +// that use this type -- `checked_len` rejects the operation first. +unsafe impl IoBuf for HugeBuffer { + fn stable_ptr(&self) -> *const u8 { + std::ptr::NonNull::dangling().as_ptr() + } + + fn bytes_len(&self) -> usize { + usize::MAX + } +} + +// SAFETY: see the `IoBuf` impl above. +unsafe impl IoBufMut for HugeBuffer { + fn stable_mut_ptr(&mut self) -> *mut u8 { + std::ptr::NonNull::dangling().as_ptr() + } +} + +const NULL_FILE: HANDLE = std::ptr::null_mut(); + +#[test] +fn read_rejects_a_buffer_longer_than_u32_max_without_touching_the_ring() { + let mut ring = IoRing::new(8, 8).expect("create ring"); + let outstanding_before = ring.outstanding(); + let mut batch = Batch::new(&mut ring); + // SAFETY: NULL_FILE is never dereferenced -- the oversized buffer is + // rejected before the handle would be used. + let error = unsafe { batch.read_raw(NULL_FILE, HugeBuffer, 0, PushOptions::new()) } + .expect_err("an oversized buffer must be rejected"); + assert_eq!(error.kind(), std::io::ErrorKind::InvalidInput); + drop(batch); + assert_eq!( + ring.outstanding(), + outstanding_before, + "a rejected push must not reserve an identity" + ); +} + +#[test] +fn write_rejects_a_buffer_longer_than_u32_max_without_touching_the_ring() { + let mut ring = IoRing::new(8, 8).expect("create ring"); + let outstanding_before = ring.outstanding(); + let mut batch = Batch::new(&mut ring); + // SAFETY: as above. + let error = unsafe { + batch.write_raw( + NULL_FILE, + HugeBuffer, + 0, + PushOptions::new(), + WriteCaching::Cached, + ) + } + .expect_err("an oversized buffer must be rejected"); + assert_eq!(error.kind(), std::io::ErrorKind::InvalidInput); + drop(batch); + assert_eq!( + ring.outstanding(), + outstanding_before, + "a rejected push must not reserve an identity" + ); +} diff --git a/crates/windows-threadpool-sys/Cargo.toml b/crates/windows-threadpool-sys/Cargo.toml index 0d516665a..ba7cb070f 100644 --- a/crates/windows-threadpool-sys/Cargo.toml +++ b/crates/windows-threadpool-sys/Cargo.toml @@ -33,6 +33,11 @@ targets = ["x86_64-pc-windows-msvc"] # the core completion machinery, not any operation-family adapter (fs, socket, device). windows-overlapped-io-sys = { version = "0.1.3", path = "../windows-overlapped-io-sys", default-features = false } +[features] +# A trace for defects that only appear under concurrency. Off by default and +# absent from the compiled output when off, because the instrument for that +# class of problem must not change the schedule it is measuring. +trace = [] [dependencies.windows-sys] version = "0.61.2" default-features = false diff --git a/crates/windows-threadpool-sys/src/lib.rs b/crates/windows-threadpool-sys/src/lib.rs index 8a5b00f63..c97a05aef 100644 --- a/crates/windows-threadpool-sys/src/lib.rs +++ b/crates/windows-threadpool-sys/src/lib.rs @@ -143,6 +143,17 @@ pub mod io; pub mod pool; #[cfg(windows)] pub mod timer; + +/// A trace for defects that only appear under concurrency, compiled out +/// unless the `trace` feature is on and narrowed by environment variable when +/// it is. See the module documentation for why an `eprintln!` is the wrong +/// instrument for that class of problem. +/// +/// Gated on Windows like every other Win32-backed module here: the traced +/// build calls `GetCurrentThreadId`, so leaving it ungated would let +/// `--features trace` break this crate's empty-on-other-targets behaviour. +#[cfg(windows)] +pub mod trace; #[cfg(windows)] pub mod wait; #[cfg(windows)] diff --git a/crates/windows-threadpool-sys/src/trace.rs b/crates/windows-threadpool-sys/src/trace.rs new file mode 100644 index 000000000..5a1bb2cb0 --- /dev/null +++ b/crates/windows-threadpool-sys/src/trace.rs @@ -0,0 +1,245 @@ +// Copyright (c) Mike Grier +//! A trace that is cheap enough to use on a timing-sensitive defect. +//! +//! # Why this is not `eprintln!` +//! +//! It exists for one class of problem: a fault that only appears under +//! concurrency, where the obvious instrument destroys the thing it is +//! measuring. Formatting and writing a line takes microseconds and a lock on +//! stderr; the window being investigated may be shorter than that, so an +//! `eprintln!` in the wrong place does not observe a race, it prevents one. +//! +//! Three properties follow, and each is a deliberate cost: +//! +//! 1. **It compiles to nothing unless the `trace` feature is on.** Every entry +//! point below is `#[inline(always)]` and has an empty body without the +//! feature, so a published build carries no branch, no atomic, and no +//! storage. This is what makes it safe to leave the call sites in place. +//! 2. **Recording does not format and does not allocate.** An entry is a +//! timestamp, a thread id, two `&'static str` labels and two `u64` slots. +//! Text is produced only when a dump is asked for, which happens after the +//! interesting moment has passed. +//! 3. **It is off at runtime even when compiled in**, and when on it is +//! *narrowed* rather than global -- see [`crate::trace::enabled`]. A trace that records +//! everything is a trace that changes the schedule of everything. +//! +//! # Narrowing to one scenario +//! +//! Set `WINDOWS_THREADPOOL_TRACE` to a comma-separated list of target +//! substrings. Only records whose target contains one of them are kept: +//! +//! ```text +//! $env:WINDOWS_THREADPOOL_TRACE = 'wait' # just the wait object +//! $env:WINDOWS_THREADPOOL_TRACE = 'wait,delivery' # two subsystems +//! $env:WINDOWS_THREADPOOL_TRACE = '*' # everything +//! ``` +//! +//! Unset, empty, or matching nothing means no record is kept and the cost is +//! one relaxed atomic load per call site. + +#[cfg(feature = "trace")] +mod imp { + use std::sync::Mutex; + use std::sync::atomic::{AtomicU8, Ordering}; + use std::time::Instant; + + /// How many records are retained. Fixed and pre-allocated: growing a + /// buffer mid-trace would allocate on the path being measured, which is + /// the one thing this module exists to avoid. + /// + /// Oldest records are dropped first. The failures this was built for show + /// up in the first moments of a run, so keeping the most recent entries is + /// the wrong bias -- but keeping a bounded window is what stops a long run + /// consuming the machine, and a dump reports how many were lost. + const CAPACITY: usize = 8192; + + /// One observation. Deliberately `Copy` and free of owned data, so + /// recording is a push and never an allocation. + #[derive(Clone, Copy)] + pub struct Record { + pub at: std::time::Duration, + pub thread: u32, + pub target: &'static str, + pub event: &'static str, + pub a: u64, + pub b: u64, + } + + struct State { + records: Vec, + dropped: usize, + } + + static ARMED: AtomicU8 = AtomicU8::new(ARMED_UNKNOWN); + const ARMED_UNKNOWN: u8 = 0; + const ARMED_OFF: u8 = 1; + const ARMED_ON: u8 = 2; + + fn state() -> &'static Mutex { + static STATE: std::sync::OnceLock> = std::sync::OnceLock::new(); + STATE.get_or_init(|| { + Mutex::new(State { + records: Vec::with_capacity(CAPACITY), + dropped: 0, + }) + }) + } + + fn started() -> Instant { + static STARTED: std::sync::OnceLock = std::sync::OnceLock::new(); + *STARTED.get_or_init(Instant::now) + } + + fn filters() -> &'static Vec { + static FILTERS: std::sync::OnceLock> = std::sync::OnceLock::new(); + FILTERS.get_or_init(|| { + std::env::var("WINDOWS_THREADPOOL_TRACE") + .unwrap_or_default() + .split(',') + .map(|piece| piece.trim().to_owned()) + .filter(|piece| !piece.is_empty()) + .collect() + }) + } + + /// Whether anything is being traced at all. + /// + /// Cached in a relaxed atomic after the first call, so the steady-state + /// cost at a disabled call site is one load and a branch. The environment + /// is read once: re-reading it per call would put a lock and a lookup on + /// the measured path. + #[inline(always)] + pub fn enabled() -> bool { + match ARMED.load(Ordering::Relaxed) { + ARMED_ON => true, + ARMED_OFF => false, + _ => { + let on = !filters().is_empty(); + ARMED.store(if on { ARMED_ON } else { ARMED_OFF }, Ordering::Relaxed); + on + } + } + } + + /// Whether this specific target is being traced. + #[inline(always)] + pub fn wants(target: &str) -> bool { + enabled() + && filters() + .iter() + .any(|filter| filter == "*" || target.contains(filter.as_str())) + } + + /// Record one observation. Cheap by construction: a clock read, a thread + /// id, and a push under a lock that is never held across anything slow. + #[inline] + pub fn record(target: &'static str, event: &'static str, a: u64, b: u64) { + if !wants(target) { + return; + } + let at = started().elapsed(); + // SAFETY: no preconditions; returns the calling thread's id. + let thread = unsafe { windows_sys::Win32::System::Threading::GetCurrentThreadId() }; + let Ok(mut state) = state().lock() else { + return; + }; + if state.records.len() == CAPACITY { + state.records.remove(0); + state.dropped += 1; + } + state.records.push(Record { + at, + thread, + target, + event, + a, + b, + }); + } + + /// Every record so far, oldest first, formatted for reading. + pub fn dump() -> String { + let Ok(state) = state().lock() else { + return "".to_owned(); + }; + let mut out = String::with_capacity(state.records.len() * 64); + if state.dropped > 0 { + out.push_str(&format!( + " ... {} earlier record(s) dropped; raise CAPACITY to keep them\n", + state.dropped + )); + } + for record in &state.records { + out.push_str(&format!( + " {:>12.6}s t{:<6} {:<22} {:<28} {:>6} {:>6}\n", + record.at.as_secs_f64(), + record.thread, + record.target, + record.event, + record.a, + record.b + )); + } + out + } + + /// Discard everything recorded so far. + pub fn clear() { + if let Ok(mut state) = state().lock() { + state.records.clear(); + state.dropped = 0; + } + } +} + +#[cfg(not(feature = "trace"))] +mod imp { + /// Always false in this build: the feature is off. + #[inline(always)] + pub fn enabled() -> bool { + false + } + /// Always false in this build: the feature is off. + #[inline(always)] + pub fn wants(_target: &str) -> bool { + false + } + /// Does nothing in this build, and is expected to compile away entirely. + #[inline(always)] + pub fn record(_target: &'static str, _event: &'static str, _a: u64, _b: u64) {} + /// Reports that the build carries no trace, which is a different finding + /// from a build that traced and saw nothing. + pub fn dump() -> String { + "".to_owned() + } + /// Does nothing in this build. + pub fn clear() {} +} + +pub use imp::{clear, dump, enabled, record, wants}; + +/// Record one observation, evaluating its arguments only when the target is +/// being traced. +/// +/// Prefer this to calling [`record`] directly: the macro keeps argument +/// evaluation behind the filter check, so a call site can compute a value that +/// would itself be too expensive for the measured path. +#[macro_export] +macro_rules! trace_record { + ($target:expr, $event:expr) => { + $crate::trace::record($target, $event, 0, 0) + }; + ($target:expr, $event:expr, $a:expr) => { + if $crate::trace::wants($target) { + $crate::trace::record($target, $event, $a as u64, 0); + } + }; + ($target:expr, $event:expr, $a:expr, $b:expr) => { + if $crate::trace::wants($target) { + $crate::trace::record($target, $event, $a as u64, $b as u64); + } + }; +} + +#[cfg(test)] +mod tests; diff --git a/crates/windows-threadpool-sys/src/trace/tests.rs b/crates/windows-threadpool-sys/src/trace/tests.rs new file mode 100644 index 000000000..25950b6eb --- /dev/null +++ b/crates/windows-threadpool-sys/src/trace/tests.rs @@ -0,0 +1,56 @@ +// Copyright (c) Mike Grier +//! Tests for the trace facility. +//! +//! These run in both builds. Without the `trace` feature every entry point is +//! a no-op, and asserting *that* is the point: a call site left in shipping +//! code must cost nothing and must not misreport. + +use super::{dump, enabled, wants}; + +#[test] +fn a_build_without_the_feature_reports_nothing_and_says_so() { + if cfg!(feature = "trace") { + return; + } + assert!(!enabled(), "tracing cannot be on without the feature"); + assert!( + !wants("anything"), + "no target is traced without the feature" + ); + assert!( + dump().contains("without the `trace` feature"), + "a dump must say why it is empty rather than look like a clean run -- an empty trace and \ + a trace that was never compiled in are very different findings" + ); +} + +#[test] +fn recording_is_inert_without_the_feature() { + if cfg!(feature = "trace") { + return; + } + // The call must compile and do nothing. If this ever starts recording, + // the shipping build has acquired a lock and an allocation on a path that + // is meant to carry neither. + crate::trace_record!("test", "inert", 1, 2); + assert!(!dump().contains("inert")); +} + +#[test] +fn the_filter_narrows_to_the_targets_named() { + if !cfg!(feature = "trace") { + return; + } + // `wants` is driven by the process environment, which is read once and + // cached, so this asserts the relationship that holds whatever the + // variable is set to rather than mutating it: a target that is wanted + // implies tracing is on at all. + for target in ["wait", "delivery", "something-else"] { + if wants(target) { + assert!( + enabled(), + "a wanted target implies the trace is enabled: {target}" + ); + } + } +} diff --git a/crates/windows-threadpool-sys/src/wait.rs b/crates/windows-threadpool-sys/src/wait.rs index afc69c56e..8abb3d216 100644 --- a/crates/windows-threadpool-sys/src/wait.rs +++ b/crates/windows-threadpool-sys/src/wait.rs @@ -345,6 +345,7 @@ impl WaitContext { let mut suppressed = self.suppression(); *suppressed = suppressed.saturating_add(1); let wait = self.wait.load(Ordering::Acquire); + crate::trace_record!("wait", "suppress-and-disarm", wait, *suppressed); if wait != 0 { // SAFETY: `wait` is this object's live PTP_WAIT, published before any // callback could run and valid until Drop closes it. @@ -454,6 +455,21 @@ impl WaitActivation<'_> { /// /// [`TimerFiring::rearm_after`]: crate::timer::TimerFiring::rearm_after /// + /// # Re-arm before the next signal + /// + /// The ordering rule described on [`ThreadpoolWait::arm`] applies to + /// every arming, not just the first -- `SetThreadpoolWait` says the event + /// must be re-registered "before signaling it each time". A producer that + /// signals in the window after an activation consumed the arming but + /// before this call re-establishes it may not get a callback for that + /// signal. + /// + /// Re-arming *before* draining closes that window, at the cost of + /// callbacks that find nothing to do; draining first and re-arming after + /// leaves it open. A caller who cannot order the two can drain, re-arm, + /// then drain again, so that anything which landed in the window is + /// picked up by the second pass rather than waited for. + /// /// # Teardown /// /// Re-arming after the object has begun tearing down does nothing, so a @@ -524,6 +540,7 @@ unsafe fn arm_raw(wait: PTP_WAIT, handle: HANDLE, timeout: Option) { // SAFETY: forwarded; a null timeout means "wait indefinitely". None => unsafe { SetThreadpoolWait(wait, handle, ptr::null()) }, } + crate::trace_record!("wait", "armed", wait, handle as usize); } /// Trampoline from the raw `PTP_WAIT_CALLBACK` ABI into the boxed closure. @@ -537,6 +554,7 @@ unsafe extern "system" fn wait_trampoline( _wait: PTP_WAIT, wait_result: u32, ) { + crate::trace_record!("wait", "trampoline-entered", _wait, wait_result); // SAFETY: context is a valid *mut WaitContext for the full callback duration. let ctx = unsafe { &*(context as *const WaitContext) }; let activation = WaitActivation { @@ -688,6 +706,7 @@ impl ThreadpoolWait { // so no callback can observe the unpublished value. // SAFETY: context is live and exclusively ours until the first arming. unsafe { (*context).wait.store(wait, Ordering::Release) }; + crate::trace_record!("wait", "created", wait, target.raw() as usize); Ok(Self { wait, @@ -708,6 +727,25 @@ impl ThreadpoolWait { /// arming rather than adding to it, and an activation consumes the arming -- /// rearm from inside the callback with [`WaitActivation::rearm`] to keep /// watching. + /// + /// # Arm before you signal + /// + /// `SetThreadpoolWait` documents that "you must re-register the event with + /// the wait object before signaling it each time to trigger the wait + /// callback". Signal a handle that is not currently armed -- including in + /// the window between constructing a [`ThreadpoolWait`] and this call -- + /// and the callback is not guaranteed to run for that signal. + /// + /// An auto-reset event makes a dropped signal permanent rather than merely + /// late, because the signal is consumed and there is nothing left for a + /// subsequent arming to observe. If the handle is only ever signalled once, + /// as a wakeup for state that is already present, that lost signal is the + /// last one the waiter will ever get. + /// + /// Every example in this module arms first for that reason; so does every + /// caller in this workspace. A caller who cannot control the ordering + /// should watch a manual-reset event kept in agreement with the state it + /// reports, which is level-triggered and so has no signal to lose. pub fn arm(&self, timeout: Option) { // SAFETY: `wait` is valid for the lifetime of self, and the handle is // owned by self so it is still open. @@ -847,10 +885,12 @@ impl Drop for ThreadpoolWait { let ctx = unsafe { &*self.context }; // Raised and never released: unlike `stop_and_drain`, there is no // afterwards for this object. + crate::trace_record!("wait", "drop-begin", self.wait); ctx.suppress_and_disarm(); // The lock is released before draining: a callback blocked on it would // otherwise never finish, and this wait would never return. self.cancel_pending(); + crate::trace_record!("wait", "drop-drained", self.wait); // SAFETY: no callback can be queued or executing, so the object can be // closed and the context freed exactly once. `target` is dropped after @@ -861,6 +901,7 @@ impl Drop for ThreadpoolWait { CloseThreadpoolWait(self.wait); drop(Box::from_raw(self.context)); } + crate::trace_record!("wait", "drop-closed", self.wait); } } diff --git a/design-sessions/DESIGN-SESSION-2026-09-22-principles-triggers-coverage.md b/design-sessions/DESIGN-SESSION-2026-09-22-principles-triggers-coverage.md new file mode 100644 index 000000000..3f314acaf --- /dev/null +++ b/design-sessions/DESIGN-SESSION-2026-09-22-principles-triggers-coverage.md @@ -0,0 +1,124 @@ +# Design session 2026-09-22: principles, triggers, and measuring the gap between them + +**Status: parked. Not approved, not scheduled, no tooling built.** Recorded so the idea and the +concern about it both survive to a later conversation. The absence of a checklist item is deliberate, +not an oversight. + +**The engineer's open concern:** it is not clear this is operationalizable to the degree claimed. +Raised after the proposal below was made, and explicitly *a concern rather than an objection* -- it +does not block the idea and is not a position to be refuted. It marks the part that is unproven. +Anyone picking this up should treat feasibility as the open question and start there, rather than +arguing the idea is either dead or settled. + +## The problem this came from + +Over one long session, eight process errors were made against +[copilot-instructions.md](../.github/copilot-instructions.md). Sorting them by kind produced a +pattern sharp enough to be worth recording independently of any fix. + +**Five of the eight were claims about this repository's own state** -- each answerable by a command +in seconds, each asserted instead from recollection: + +| # | The claim | The fact | +|---|---|---| +| 1 | the epoch-commit trigger was "armed" for a reader raising `EPOCH_SIZE` | unreachable at any constants; measured over five values | +| 2 | a blast-radius sweep covered "13 files" | 14; the figure was read off `rg`'s grouped summary rather than counted | +| 3 | an item's two named sites were the population | a census found six | +| 4 | "`M22` is a testing-heavy milestone" | all three `M22` items touch only `examples/` | +| 5 | "`M23.1` touches the crate's contract surface" | it names the *sample's* `contract.rs` | + +The other three were claims about Windows behaviour -- an error-kind mapping, `SubmitIoRing`'s +timeout result, and inline completion on a synchronous handle. Those are the class the kernel tests +and an independent review exist to catch, and they were caught. + +Claims 4 and 5 are the sharpest, because they were written *into a milestone's sequencing rationale* +and claim 4 reversed that milestone's conclusion about when to schedule the work. + +## The diagnosis + +**Rules with a mechanical trigger fired reliably in that session. Rules requiring recognition that +they applied did not -- and the most carefully argued rules in the file are the ones breached.** + +Not once forgotten: `--no-pager`, LF line endings, writing scratch output under `.scratch/`, +`cargo fmt` before a commit, the scratch-file commit message. Each fires on a discrete, visible act: +*running a git command*, *creating a file*, *committing*. + +Breached repeatedly: CONTRACT INTEGRITY and FAIL FAST -- the two most heavily reasoned sections. +Their triggers are properties of text being generated, which is the worst possible moment to +introspect, because generation is the thing in flight. + +A secondary effect: attention to any one principle decays as a session fills. The sweep census in +`M21.1` was run unprompted early on; twenty-odd tool calls later a file count was read off a summary +line without a thought. Same rule, same session. + +## Why writing the trigger is not simply the answer + +The obvious fix is to write the trigger beside the principle. The reservation about that is precise +and +was raised before any of this was proposed: + +> The problem with *me* writing the triggers rather than the principles is that the triggers are +> invariably narrower than the principle so then when the trigger is too narrow, it becomes like a +> legal argument about definitional terms rather than intent. + +That is the rules-lawyering failure mode, and it is real. A trigger narrow enough to be mechanical +is narrow enough to be argued around. So a trigger cannot *replace* a principle; at best it samples +it. + +## The proposal, such as it is + +The idea is to stop arguing about whether a trigger is too narrow and measure it instead. + +**Why it is tractable at all:** the space a principle describes cannot be enumerated -- that is what +makes it a principle. But the incidents that actually occurred can be. Each incident is a sampled +point in that space, and a trigger either would or would not have fired on it. + +**The metric.** Coverage of *past* incidents is worthless as a health measure: it can be driven to +100% by adding one trigger per incident, which is the legalism trap with a green dashboard. The +measure that means something is the **escape rate on incidents recorded after a trigger was +written**. A trigger set that only catches what it was built for is narrow, by measurement rather +than by argument; one that catches novel incidents generalises. + +**The twin failure is already covered by a rule this repository holds.** +[README-sabotage.md](../tools/README-sabotage.md) requires a manifest to carry an +`expect: "survives"` control, on the grounds that without one it "can only tell you the tests are +sensitive, never that they are sensitive to the right things". The same applies: the corpus needs +text that *looks* like a violation and is not. Then a trigger that is too narrow shows up as escapes, +and one that is too broad shows up as hits on the controls. + +**Shape**, modelled on the sabotage harness because that pattern is already trusted here: + +- principles: the existing `##` sections, referenced by heading and never copied; +- a trigger list: id, the tell, which principles it serves, the date it was added; +- an incident list: id, date, the actual offending text, the principle breached, how it was caught; +- a script that replays incidents against triggers, runs the controls, and reports the escape rate + split by whether each incident predates or postdates the trigger. + +## What was already conceded about it + +Stated when the proposal was made, not extracted afterwards: + +- **A script checks committed text; the triggers that matter fire during generation.** It can + measure trigger quality. It cannot make a trigger fire. This is the same limit the sabotage + harness has -- it tells you the tests are sensitive, it does not write them. +- **A subset does graduate to real enforcement.** "`because` followed by a reference to another file + or item, inside a planning document" is greppable on committed text. That subset is a genuine gain + regardless of whether the escape-rate signal ever proves useful. +- **The corpus would start at eight incidents, self-reported, from one session.** It says nothing for + several sessions. +- **It is more machinery in a repository that already carries a great deal.** + +## Kill criterion, agreed in advance + +If after an agreed number of sessions every new incident still escapes every trigger, the approach +has failed and is deleted rather than extended. Written down first precisely because the instinct to +systematise is what produced a 146 KiB instruction file whose best-argued sections went unheeded. + +## What happened instead, and what is actually in effect + +One change landed, and it is not this: CONTRACT INTEGRITY rule 6 gained a clause covering +characterisations of sibling checklist items, with claims 4 and 5 above cited as its worked +examples, and naming **because** followed by a reference to another file or item as the tell. + +That is a single trigger written beside a single principle. Whether that generalises, or merely +relocates the problem, is the question this note exists to reopen later. diff --git a/design-sessions/DESIGN-SESSION-2026-09-23-adoption-thesis.md b/design-sessions/DESIGN-SESSION-2026-09-23-adoption-thesis.md new file mode 100644 index 000000000..9c9127aa0 --- /dev/null +++ b/design-sessions/DESIGN-SESSION-2026-09-23-adoption-thesis.md @@ -0,0 +1,217 @@ +# Design session -- the adoption thesis behind the ring, queue and locality work (2026-09-23) + +**Summary.** The engineer stated the strategic intent behind the whole ring / queue / +topology / durability line of work: that making queues and rings reachable, and having +locality benefits arrive adaptively out of that structure, could start a virtuous cycle +ending with consumer hardware exposing what only server hardware exposes today. The +session produced the Tier 1 section +[The adoption thesis](../DESIGN-NOTES.md#the-adoption-thesis), the Tier 2 entry +[Why no option is foreclosed while the hardware gap lasts](../DESIGN-RATIONALE.md#why-no-option-is-foreclosed), +and milestone `M27` in +[CHECKLIST.md](../crates/windows-ioring-sys/CHECKLIST.md). It also supplies the reason +behind the **OPTION INTEGRITY** rule recorded in +[copilot-instructions.md](../.github/copilot-instructions.md) one commit earlier. + +## Why it was stated now + +It followed directly from a correction. The prior commit had concluded that a commit +strategy's justification was "dead on structural grounds" on the strength of a harness +that could not have shown the strategy working. The engineer's response was to reject the +move in general rather than only in that instance: + +> Unless we can determine that the option has no possible value, we should expose it and +> provide the tools for clients to be able to choose it when applicable and to gather the +> appropriate data to make the wisest choice possible. + +That produced OPTION INTEGRITY as a rule. This session supplies what the rule was missing: +the reason such foreclosures are *specifically* costly in this repository, which is not a +general preference for breadth but a consequence of what this work is for and what +hardware it is being built on. + +## The engineer's framing, recorded + +### The hardware trend + +Non-uniform memory has been relegated to very expensive machines. As Moore's law has +plateaued, the need to introduce wider multiprocessors in contexts closer to consumers has +risen: you can only go so far with a uniform memory architecture. Even the consumer AMD +memory architectures are non-uniform with a uniform facade in front of them, and arguably +Intel also. + +### Why NUMA is not a consumer-level concept + +One of the main reasons is that it is difficult to program for, and the benefits are +difficult to measure in comparison to how you would have to program your system otherwise. +Today you have to make a significant architectural decision at the beginning of your +system architecture: do I want to take advantage of NUMA or not? Who would make such a +choice unless targeting larger datacenter-class hardware? + +The framing worth preserving here is that the **decision point**, not the difficulty, is +the principal barrier. It is a commitment demanded at the moment the least is known, and +its payoff is invisible to anyone not already buying datacenter hardware. + +### The parallel problem with queues + +The programming techniques of `IoRing` and queues in general are of general purpose +utility but are not available as readily as one might wish. There are a lot of building +blocks, but except in certain small domains they are not easily grasped for how to +structure your system from the beginning. + +This is the same shape as +[The value is existence, not cleverness](../DESIGN-NOTES.md#the-value-is-existence-not-cleverness), +which was already recorded: the correct construction is not within reach, so people reach +for what is. + +### The thesis + +There is a potential virtuous-cycle synergy to launch. Enabling applications to more +easily tap into the benefits of rings, queues and the Windows `IoRing`, and then to have +adaptive NUMA benefits without having to write specialized code to receive those benefits, +would: + +- make "normal" application writers attracted to the techniques; +- yield benefits to I/O-bound application writers who would benefit from the gains of + `IoRing` and the epoch-based durability idiom; +- provide a useful substrate for high-end application designers who actually do want to + target NUMA systems; +- and then, if there is adoption, give reason for systems designers to expose more system + NUMA features at the consumer hardware level, so that lower-capability hardware can + receive the same kinds of benefits. + +The operative phrase, and the one with consequences for the code, is **"without having to +write specialized code to receive those benefits"**. That is a claim about adaptivity, and +it is not what the tree does today. + +### The honest expectation + +Today the engineer expects low benefits to all but server-class hardware, to be fair. And +even though this work is being developed on a machine that is logically server-class +hardware, it is nonetheless just a slice, which does not exhibit the non-uniform memory +characteristics. + +Thus the reason to avoid all early foreclosures of techniques which may yield benefits to +application authors: we are not the application authors, so we do not actually know what +they may want to do; and we do not have the hardware to make significant analyses of +performance tradeoffs on real NUMA hardware. + +## What this changes, and what it deliberately does not + +**Unchanged.** The hardware-gap analysis in +[DESIGN-SESSION-2026-08-30-numa-sharded-io-execution-domains.md](DESIGN-SESSION-2026-08-30-numa-sharded-io-execution-domains.md) +already worked out what is blocked on real NUMA hardware and what is not, and concluded +that the blocked column "threatens the justification and the tuning", not the structure. +Nothing in this session revises that; the thesis explains why the justification is worth +keeping open rather than settling locally, which is a different question. + +**Unchanged.** [D-8](../crates/windows-ioring-sys/DESIGN-NOTES.md#d-8) leaves partitioning +policy with the consumer. The thesis does not overturn it, and `M27` is explicitly written +so that "the library decides for you" is one candidate answer rather than the presumed +one. A thesis about making a benefit reachable is not automatically a thesis about making +it automatic, and conflating the two would be the crate taking a workload decision it has +repeatedly refused to take. + +**Changed.** The gap between "benefits arrive adaptively out of the structure" and a tree +where partitioning is an explicit `Policy` a consumer selects is now a queued question +rather than an unstated aspiration. Per the repository's rule that design notes are not a +work queue, an intent that implies work on existing code has to become a checklist item or +say explicitly that it schedules none. This one implies work, so it is `M27`. + +## The distinction that keeps the thesis honest + +A thesis is not evidence, and this one is unusually exposed to being mistaken for +evidence, because the repository's other claims are unusually well measured. Two guards +apply, both already established here: + +- The 2026-08-30 session's practice of **marking documented-but-unwitnessed claims + distinctly from measured ones** applies to the thesis in full. The Tier 1 section says + "thesis, not a measurement" in its first sentence for that reason. +- The repository rule to **present what was observed and not write the conclusion** means + the thesis must not leak into places that report measurements. `cache_domains.rs` prints + `L1: 8`, `L2: 8`, `L3: 1` and marks which the heuristic chose; it does not tell the + reader what that implies for their design. That separation is the thesis working + correctly, not a limitation of the sample. + +The measured observations the thesis leans on are genuinely measured, and are cited rather +than restated: the ARM laptop reporting no L3 and zero nodes +([D-48](../crates/windows-ioring-sys/DESIGN-NOTES.md#d-48)), and this host's L3 spanning +all 16 processors over a real 8-way L2 partition. What is *not* measured is every forward +step of the cycle, and the Tier 1 section says so. + +## Later the same day: the mechanism has a name, and it already existed + +The thesis above was recorded without naming the component that delivers it, and the gap was queued +as an adaptivity question inside `windows-ioring-sys`. The engineer's correction: + +> This is why it's the "topology planner". The idea is to have the developer give a sufficiently +> abstract definition of the application's input, output, and processing code paths and then the +> topology planner would be able to infer both the general connectivity / directed flow of data +> needed to realize the graph and then when given a physical machine model would respond with one or +> more suggested realizations of the graph in terms of specific execution threads pinned on which +> processor groups, numbers of queues of which types, etc. + +[topology-planner](../crates/topology-planner/COMPONENT.md) had been planned since 2026-09-03 and +its input was the one part left open: `EP-D-4` recorded the goal as an input whose shape was +"deferred for litigation". That deferral is what this statement discharges, and it is recorded as +[EP-D-6](../crates/topology-planner/DESIGN-NOTES.md#ep-d-6). + +**Three things arrived at once, and only the first was the deferred question.** + +1. **The input is a dataflow description** -- the application's input, output, and processing code + paths. Not a topology preference and not a set of tuning hints. +2. **Planning is two stages.** Connectivity and directed flow are inferred with no machine in hand; + the machine enters only at the second stage. This was not asked for separately and follows from + the first: a derivation that needs no machine should not be entangled with one. +3. **The answer is plural** -- one or more *suggested* realizations, which the developer chooses + between. + +**The third is the one worth guarding.** A single returned plan is easier to consume, test and +document, and those pressures will argue for collapsing the set at some later convenient moment. The +reason not to is structural: ranking candidates requires knowing what the developer values, and this +component was handed a description of an application rather than a statement of preference. A +planner that returns one arrangement has either acquired a preference it was not given or hidden a +choice it was not entitled to make. + +### What this corrected about the morning's work + +**`M27` was in the wrong crate.** It had been written that morning, asking whether +`windows-ioring-sys` should derive a partition for a consumer who expresses no preference. Answering +that there would have produced a second policy surface beside the planner's -- which is precisely the +`outermost_partitioning_cache` defect the planner exists to avoid, a policy answer landing in a crate +whose job is something else. `M27` was re-planned the same day into what the ring crate genuinely +owes: being **realizable from** a plan it did not choose. The checklist rules require saying that +plainly rather than silently rewriting the milestone, which is why both the milestone and the two +PLANS rows describing it now carry the correction. + +**The thesis section was incomplete rather than wrong.** It named the barrier -- an architectural +commitment demanded when the least is known -- and did not name what removes it. The mechanism +section added afterwards is the answer, and the three properties it lists are each doing work: the +developer never makes a topology decision, stage 1 is stable across every machine the application +will run on, and the plural answer is OPTION INTEGRITY at component scale. + +**And the deferral held up -- but the reason is not the one first recorded here.** `EP-D-4`'s goal +input had been undefined for twenty days across four documents, and nothing was built on a guess in +the meantime, because the deferral was *named* at every site that mentioned it rather than being an +absence. Correcting it was a sweep of four sites, all found by grepping the phrase the deferral was +recorded under; an unnamed omission would have left nothing to grep for. + +That much is true and is the smaller half. **The first version of this paragraph stopped there, and +in doing so implied the answer had existed since 2026-09-03 and was waiting to be stated.** It had +not. The engineer's correction, recorded verbatim because the distinction is the point: + +> I wasn't "holding out" on my perfectly formed design from before, it was the work that we've done +> that helped me reach this clarity. + +The sequence that produced it was known in outline from the start -- build some building blocks, +build some measurement tools, then experiment with how those tools could be used to infer things -- +and the shape of the planner's input is an **output** of having done that, not an input withheld +from it. The way forward is now much clearer than it was and is still not crystal clear, which is +the expected state and not a gap to be closed by questioning. + +**The lesson generalises past this deferral**, and is now recorded as RESOLUTION GRADIENT in +[copilot-instructions.md](../.github/copilot-instructions.md): a plan is sharp at the front and +deliberately coarse behind, and an assistant that puts a very specific question to a general sense +manufactures a low-confidence answer which then gets recorded as a decision and mis-placed. Several +questions in this session were of that shape. The remedy is to calibrate a question's specificity to +the resolution actually available, to offer "too early to say" as a real answer, and to record a +hedged direction as a **working position** rather than a decision -- a device this repository already +had, in the 2026-08-30 session's "Working position on domain counts (not a decision)". \ No newline at end of file diff --git a/release-please-config.json b/release-please-config.json index fc19ecb93..f7cbe028b 100644 --- a/release-please-config.json +++ b/release-please-config.json @@ -21,7 +21,11 @@ "component": "windows-platform-probes", "versioning": "probe-calver" }, - "crates/windows-file-enumeration-sys": { + "crates/win-numa-sys": { + "package-name": "win-numa-sys", + "component": "win-numa-sys" + }, + "crates/windows-file-enumeration-sys": { "package-name": "windows-file-enumeration-sys", "component": "windows-file-enumeration-sys" }, diff --git a/tools/check-borrow-surface.ps1 b/tools/check-borrow-surface.ps1 index 5830f44dc..25fb16b72 100644 --- a/tools/check-borrow-surface.ps1 +++ b/tools/check-borrow-surface.ps1 @@ -5,8 +5,8 @@ # # Population C -- what safe code is *permitted* to do -- is the defect class no # runtime technique reaches, because nothing has to execute for the hole to -# exist. windows-ioring-sys has shipped three of them, all the same shape: a -# public method whose return type allowed an operation nobody intended. +# exist. windows-ioring-sys has shipped four of them, all the same shape: a +# public signature that allowed an operation nobody intended. # # D-35 `get_mut` returned `&mut Vec`, which permits `reserve`, `resize` # and reassignment, where only byte writes were intended. @@ -15,23 +15,43 @@ # D-43 `EventDelivery::ring` returned `&Mutex`, and any `&mut IoRing` # permits whole-value assignment -- so safe code could replace the ring # and silently stop delivery. +# D-45 `get` returned a slice living as long as the borrow, while the check +# that guarded it held only at the instant of the call. # -# All three arrived through ordinary, well-reviewed changes. What was missing +# All four arrived through ordinary, well-reviewed changes. What was missing # was not diligence, it was a specific question being asked at a specific # moment. M18.1 asked it once, over the whole surface; this script is what makes # it recur, because a rule that lives only in a document is a rule that depends -# on somebody remembering to apply it -- which is exactly how the three above +# on somebody remembering to apply it -- which is exactly how the four above # got in. # -# The mechanism is a committed inventory. Every public function in -# `crates/windows-ioring-sys/src` whose return type carries a borrow -- a -# reference, or a lifetime parameter such as `Batch<'_>` -- is listed in +# The mechanism is a committed inventory. Every public signature in +# `crates/windows-ioring-sys/src` that carries a borrow is listed in # BORROW-SURFACE.txt. This script regenerates that list from the source and -# fails if it differs. Adding or widening such a method therefore cannot land +# fails if it differs. Adding or widening such a signature therefore cannot land # quietly: CI stops, and the author has to answer the question in # DESIGN-INSTRUCTIONS.md and record the answer before the inventory can be # updated. # +# THREE SHAPES ARE INSPECTED, and the second and third were added by M21+.1 +# after a review found the check silent on a change that used both: +# +# 1. An inherent `pub fn` whose RETURN type carries a borrow. The original +# rule, and what D-35/D-36/D-43/D-45 all were. +# 2. Any method of a `pub trait`, on either side of the signature. Trait +# items are declared `fn`, not `pub fn`, so the rule above never matched +# one -- a public trait method returning `&[u8]` was invisible. +# 3. A PARAMETER carrying an explicit lifetime, such as `&mut RingWait<'_>`. +# This is the direction the original rule could not see, and it is the +# wider exposure of the two: a return value goes to a known caller, while +# a borrow-carrying parameter of a public trait method is handed to +# arbitrary safe code the crate has never seen. +# +# A plain `&T` parameter is deliberately NOT reported. Lending a reference to a +# callee is the caller's business and not this defect class; what matters is a +# borrow-carrying wrapper whose lifetime the crate chose. The explicit-lifetime +# test is what separates them. +# # This deliberately checks *shape*, not correctness. It cannot tell a safe # accessor from a dangerous one -- only that the surface changed and a human # owes an answer. That is the whole job: the question, asked reliably. @@ -57,10 +77,11 @@ if (-not (Test-Path $sourceRoot)) { exit 2 } -# Collect one entry per public function whose return type carries a borrow. +# Collect one entry per public signature that carries a borrow. # -# Signatures wrap across lines, so accumulate from `pub fn` until the line that -# closes the signature -- the one ending in `{` (a body) or `;` (a trait item). +# Signatures wrap across lines, so accumulate from the `fn` line until the line +# that closes the signature -- the one ending in `{` (a body) or `;` (a trait +# item without one). function Get-BorrowSurface { param([string]$Root) @@ -72,24 +93,62 @@ function Get-BorrowSurface { $lines = [System.IO.File]::ReadAllLines($file.FullName) $index = 0 + # Depth tracking for `pub trait` blocks. Trait items inherit the + # trait's visibility and are declared `fn`, not `pub fn`, so the only + # way to recognise one is to know we are inside such a block. + $traitDepth = -1 + $depth = 0 + while ($index -lt $lines.Length) { $line = $lines[$index] - if ($line -notmatch '^\s*pub(\s+(unsafe|const|async))*\s+fn\s') { + # A `pub trait` opens a block whose items are public. `unsafe` and + # `auto` may sit between; a supertrait list may follow. + if ($traitDepth -lt 0 -and $line -match '^\s*pub(\s+(unsafe|auto))*\s+trait\s') { + $traitDepth = $depth + } + + $isTraitItem = ($traitDepth -ge 0) -and ($line -match '^\s*(unsafe\s+)?fn\s') + $isInherent = $line -match '^\s*pub(\s+(unsafe|const|async))*\s+fn\s' + + if (-not $isTraitItem -and -not $isInherent) { + # Track braces only on lines that are not signature starts; a + # signature's own braces are consumed by the accumulator below. + $depth += ([regex]::Matches($line, '\{')).Count + $depth -= ([regex]::Matches($line, '\}')).Count + if ($traitDepth -ge 0 -and $depth -le $traitDepth) { + $traitDepth = -1 + } $index++ continue } # Accumulate the whole signature. + # + # A signature ends at the first `{` -- wherever it appears, including + # a one-line body such as `pub fn f() -> &[u8] { &[] }` -- or at a + # `;` for a trait item with no body. Testing the accumulated text + # rather than "the line ends with `{`" is what handles the one-line + # form; the earlier test ran off the end of the file on it. $signature = '' $cursor = $index while ($cursor -lt $lines.Length) { $signature += ' ' + $lines[$cursor].Trim() - if ($lines[$cursor] -match '\{\s*$' -or $lines[$cursor] -match ';\s*$') { + if ($signature -match '\{' -or $signature -match ';\s*$') { break } $cursor++ } + if ($cursor -ge $lines.Length) { + $cursor = $lines.Length - 1 + } + foreach ($consumed in $index..$cursor) { + $depth += ([regex]::Matches($lines[$consumed], '\{')).Count + $depth -= ([regex]::Matches($lines[$consumed], '\}')).Count + } + if ($traitDepth -ge 0 -and $depth -le $traitDepth) { + $traitDepth = -1 + } $index = $cursor + 1 $signature = ($signature -replace '\s+', ' ').Trim() @@ -100,28 +159,87 @@ function Get-BorrowSurface { $name = $Matches[1] # The return type is what follows the last `->` before the body. + # `-1` distinguishes "no arrow" from "arrow at position 0". $arrow = $signature.LastIndexOf('->') - if ($arrow -lt 0) { - continue + $returns = '' + if ($arrow -ge 0) { + $returns = $signature.Substring($arrow + 2) + # Strip a body, including the one-line form `{ &[] }` whose + # braces sit on the same line as the return type. + $brace = $returns.IndexOf('{') + if ($brace -ge 0) { + $returns = $returns.Substring(0, $brace) + } + $returns = ($returns -replace '\s*;\s*$', '') + $returns = ($returns -replace '\s*where\b.*$', '').Trim() } - $returns = $signature.Substring($arrow + 2) - $returns = ($returns -replace '\s*\{\s*$', '') -replace '\s*;\s*$', '' - $returns = ($returns -replace '\s*where\b.*$', '').Trim() # A borrow is a reference, or a lifetime parameter carried by a # wrapper such as `Batch<'_>` or `RingScope<'_>` -- which is exactly # where D-43's fix lives, so a `&`-only scan would miss it. - if ($returns -notmatch '&' -and $returns -notmatch "'") { + $returnBorrows = $returns -and ($returns -match '&' -or $returns -match "'") + + $parameters = Get-ParameterList -Signature $signature + # Only an EXPLICIT lifetime counts in parameter position. A plain + # `&T` is the caller lending to us, which is not this defect class; + # a wrapper whose lifetime the crate chose, such as + # `&mut RingWait<'_>`, is. + $parameterBorrows = $parameters -and ($parameters -match "'") + + if (-not $returnBorrows -and -not $parameterBorrows) { continue } - $entries += "{0} :: {1} -> {2}" -f $relative, $name, $returns + # Returns and parameters stay distinguishable in the inventory, and + # a return-only entry keeps the format it had before M21+.1 so + # widening the check did not churn the rows it already covered. + if ($parameterBorrows) { + $entries += "{0} :: {1}({2}) -> {3}" -f $relative, $name, $parameters, ($returns ? $returns : '()') + } + else { + $entries += "{0} :: {1} -> {2}" -f $relative, $name, $returns + } } } return , ($entries | Sort-Object) } +# The parameter list of `$Signature`, with the receiver removed. +# +# `&self` and `&mut self` are on every method and say nothing about this defect +# class, so reporting them would bury the entries that matter. +function Get-ParameterList { + param([string]$Signature) + + $open = $Signature.IndexOf('(') + if ($open -lt 0) { + return '' + } + + # Balance parentheses: a parameter type may contain its own, as in + # `impl FnOnce(*mut c_void, usize) -> HRESULT`. + $depth = 0 + $close = -1 + for ($i = $open; $i -lt $Signature.Length; $i++) { + if ($Signature[$i] -eq '(') { $depth++ } + elseif ($Signature[$i] -eq ')') { + $depth-- + if ($depth -eq 0) { $close = $i; break } + } + } + if ($close -lt 0) { + return '' + } + + $inner = $Signature.Substring($open + 1, $close - $open - 1).Trim() + # Drop the receiver, however it is spelled. + $inner = $inner -replace "^&\s*('[a-z_][a-z0-9_]*\s*)?(mut\s+)?self\s*,?\s*", '' + $inner = $inner -replace '^mut\s+self\s*,?\s*', '' + $inner = $inner -replace '^self\s*,?\s*', '' + return $inner.Trim() +} + $current = Get-BorrowSurface -Root $sourceRoot if ($Update) { @@ -172,8 +290,8 @@ foreach ($entry in $removed) { Write-Host '' Write-Host 'This is not an error in itself. It is the moment the question has to be' -ForegroundColor Cyan -Write-Host 'asked, because three shipped defects (D-35, D-36, D-43) all entered as' -ForegroundColor Cyan -Write-Host 'ordinary reviewed changes to this surface:' -ForegroundColor Cyan +Write-Host 'asked, because four shipped defects (D-35, D-36, D-43, D-45) all entered' -ForegroundColor Cyan +Write-Host 'as ordinary reviewed changes to this surface:' -ForegroundColor Cyan Write-Host '' Write-Host ' What can safe code do with this, and does the registration or the' -ForegroundColor White Write-Host ' kernel still hold anything it could invalidate?' -ForegroundColor White diff --git a/tools/check-ring-tests.ps1 b/tools/check-ring-tests.ps1 new file mode 100644 index 000000000..6d88856a5 --- /dev/null +++ b/tools/check-ring-tests.ps1 @@ -0,0 +1,191 @@ +# Copyright (c) Mike Grier +# +# tools/check-ring-tests.ps1 -- keeps windows-ioring-sys's population of +# ring-opening *lib* tests from growing unnoticed. +# +# D-49 recorded the defect: 63 of 131 lib tests opened a real kernel ring, so +# `cargo test --lib` did not mean what its name implies. The repository's own +# Quality rule already classifies an operating-system API as an external +# boundary, which makes those integration tests living in the unit-test +# location. +# +# M24.2, M24.7 and M24.3 took it to 41. What matters now is that it does not +# climb back, and the reason it climbed in the first place is that nothing was +# watching -- not that anyone was careless. +# +# WHY THIS IS AN INVENTORY AND NOT A ZERO-CHECK. The obvious rule, and the one +# M24.5 originally assumed, is "no lib test constructs an IoRing". That rule is +# false and cannot be made true by effort: +# +# * `event_delivery` needs a real ring and the thread pool. +# * `ring`'s injected-failure cluster transforms a REAL completion on +# purpose -- fabricating one is the unsoundness the seam exists to avoid, +# and one of those tests says so in its own assertion message. +# * `batch` needs a `Batch`, which needs the handle for its `Build*` calls. +# * Several reach `#[cfg(test)] pub(crate)` helpers that exist only inside +# the crate. +# +# A zero-check would fail on day one and could only be satisfied by deleting +# real coverage. So the check records *which* tests open a ring, and fails when +# that set changes -- the same mechanism, and for the same reason, as +# check-borrow-surface.ps1. +# +# WHY PER-TEST RATHER THAN PER-FILE. Two thirds of the remainder lives in +# `ring/tests.rs`. A file-level allow-list would permit that file to grow +# without limit, which is where a new ring-opening test would most naturally +# land. A bare count was rejected too: add-one-remove-one nets to zero and +# passes, and a number in a file is derived data nobody can check by reading. +# +# This deliberately checks *population*, not correctness. It cannot tell a test +# that needs a ring from one that merely uses it -- only that the set changed +# and a human owes an answer: +# +# Does this test need the kernel, or only a ring-shaped thing? If the +# latter, narrow what it reaches for (M24.7) or move it to tests/ (M24.3). +# +# ./tools/check-ring-tests.ps1 # verify (CI) +# ./tools/check-ring-tests.ps1 -Update # regenerate after answering + +[CmdletBinding()] +param( + [switch]$Update +) + +Set-StrictMode -Version Latest +$ErrorActionPreference = 'Stop' + +$repoRoot = Split-Path -Parent $PSScriptRoot +$sourceRoot = Join-Path $repoRoot 'crates\windows-ioring-sys\src' +$inventoryPath = Join-Path $repoRoot 'crates\windows-ioring-sys\RING-OPENING-LIB-TESTS.txt' + +if (-not (Test-Path $sourceRoot)) { + Write-Host "CONFIG ERROR: source root not found: $sourceRoot" -ForegroundColor Red + exit 2 +} + +# One entry per `#[test]` in `src/**/tests.rs` whose body reaches a ring, either +# directly or through a helper in the same file that does. +function Get-RingOpeningTests { + param([string]$Root) + + $entries = New-Object System.Collections.Generic.List[string] + + foreach ($file in (Get-ChildItem -Path $Root -Recurse -Filter 'tests.rs' | Sort-Object FullName)) { + $text = [System.IO.File]::ReadAllText($file.FullName) + $module = $file.Directory.Name + + # Helpers in this file that construct a ring themselves. A test calling + # one of these opens a ring just as surely as one saying so inline, + # which is the shape a per-file grep for `IoRing::new` would miss. + $helpers = New-Object System.Collections.Generic.List[string] + foreach ($match in [regex]::Matches($text, '(?m)^\s*fn\s+(\w+)')) { + $start = $match.Index + $match.Length + $stop = $text.IndexOf("`n}", $start) + if ($stop -lt 0) { $stop = $text.Length } + $body = $text.Substring($start, [Math]::Min(4000, $stop - $start)) + if ($body -match 'IoRing::new') { $helpers.Add($match.Groups[1].Value) | Out-Null } + } + + $blocks = $text -split '#\[test\]' + for ($i = 1; $i -lt $blocks.Count; $i++) { + $block = $blocks[$i] + $named = [regex]::Match($block, 'fn\s+(\w+)') + if (-not $named.Success) { continue } + $name = $named.Groups[1].Value + + # The test's own body ends at the first closing brace in column 0. + # NOT at the next `#[`: a plain helper defined after the last test + # in a file would otherwise be swallowed into that test's body, and + # a helper containing `IoRing::new` would report the innocent test + # above it as ring-opening. That false positive was produced by this + # script's own bidirectional check before it was fixed. `cargo fmt` + # is enforced here, so a column-0 `}` is reliably a function end. + $end = $block.IndexOf("`n}") + $body = if ($end -gt 0) { $block.Substring(0, $end) } else { $block } + + $opensRing = $body -match 'IoRing::new' + if (-not $opensRing) { + foreach ($helper in $helpers) { + if ($helper -eq $name) { continue } + if ($body -match "\b$([regex]::Escape($helper))\s*\(") { $opensRing = $true; break } + } + } + + if ($opensRing) { $entries.Add("$module::$name") | Out-Null } + } + } + + return @($entries | Sort-Object) +} + +$current = Get-RingOpeningTests -Root $sourceRoot + +if ($Update) { + $header = @( + '# windows-ioring-sys: lib tests that open a real kernel ring.', + '#', + '# GENERATED by tools/check-ring-tests.ps1 -Update. Do not hand-edit.', + '#', + '# These are integration tests living in the unit-test location (D-49).', + '# The list exists so the population cannot grow unnoticed, which is how', + '# it reached 63 before anyone counted. Adding an entry obliges an', + '# answer: does this test need the kernel, or only a ring-shaped thing?' + ) + $body = $header + $current + [System.IO.File]::WriteAllText($inventoryPath, ($body -join "`n") + "`n") + Write-Host "Updated $inventoryPath ($($current.Count) entries)." -ForegroundColor Green + exit 0 +} + +if (-not (Test-Path $inventoryPath)) { + Write-Host "CONFIG ERROR: inventory not found: $inventoryPath" -ForegroundColor Red + Write-Host "Create it with: ./tools/check-ring-tests.ps1 -Update" -ForegroundColor Yellow + exit 2 +} + +$recorded = @([System.IO.File]::ReadAllLines($inventoryPath) | + Where-Object { $_ -and -not $_.StartsWith('#') }) + +# `@(...)` on both: under StrictMode a pipeline yielding nothing is `$null` and +# one yielding a single string is a bare string, neither of which has `.Count`. +$added = @($current | Where-Object { $recorded -notcontains $_ }) +$removed = @($recorded | Where-Object { $current -notcontains $_ }) + +if ($added.Count -eq 0 -and $removed.Count -eq 0) { + Write-Host "Ring-opening lib tests unchanged ($($current.Count) entries)." -ForegroundColor Green + exit 0 +} + +Write-Host '' +Write-Host 'windows-ioring-sys: the set of lib tests that open a kernel ring changed.' -ForegroundColor Red +Write-Host '' + +foreach ($entry in $added) { + Write-Host " ADDED $entry" -ForegroundColor Yellow +} +foreach ($entry in $removed) { + Write-Host " REMOVED $entry" -ForegroundColor Yellow +} + +Write-Host '' +if ($added.Count -gt 0) { + Write-Host 'An ADDED entry is the moment the question has to be asked, because a' -ForegroundColor Cyan + Write-Host 'lib test that opens a ring is an integration test in the unit-test' -ForegroundColor Cyan + Write-Host 'location (D-49), and the population reached 63 that way:' -ForegroundColor Cyan + Write-Host '' + Write-Host ' Does this test need the KERNEL, or only a ring-shaped thing?' -ForegroundColor White + Write-Host '' + Write-Host ' If only the latter, narrow what it reaches for -- Token::new takes' -ForegroundColor White + Write-Host ' the ledger rather than the ring for exactly this reason (M24.7) --' -ForegroundColor White + Write-Host ' or move it to tests/ if it uses only public API (M24.3).' -ForegroundColor White + Write-Host '' +} +if ($removed.Count -gt 0) { + Write-Host 'A REMOVED entry is progress and needs no justification, only the' -ForegroundColor Cyan + Write-Host 'regeneration below so the inventory keeps describing the source.' -ForegroundColor Cyan + Write-Host '' +} +Write-Host 'Then run:' -ForegroundColor Cyan +Write-Host ' ./tools/check-ring-tests.ps1 -Update' -ForegroundColor White +Write-Host '' +exit 1