Repository navigation
pool: the page-stepping half of #94 recovered nothing on a live walk (recovered_bytes = 0) #104
Description
Activity
Still relevant —
walk_region's stall path is untouched since. But the alternative
reading is disproved without needing another run, and by something already in the tree.The accounting is right.
stalled_hereis set in the stall branch and never
cleared, so it latches for the rest of the region, and every later successful
read_extentin that region adds torecovered_bytes. The case the issue worried
about — "a stall in a region walked no further would never contribute" — is real but is
not a mis-count: a region walked no further has nothing to contribute.
test_a_stalled_query_costs_a_page_not_the_regionalready pins
recovered_bytes: 2 * PAGE_SIZEfor a stall with two committed pages behind it, so a
stall that is followed by a read does show up.So the benign reading is the only one left: on 26100 every one of those 1,619 stalls
sits at the end of its region's readable content, and the page-stepping recovers
nothing there. The extra bool this issue proposed would report0 of 1619and tell us
whatrecovered_bytes: 0already does.I would not remove it, though. It costs at most eight queries per dead region, and
what it replaced was losing every committed page behind one bad one — the samples in
#94 fired part-way through regions of 0x10fd0 and 0x3afc0 bytes.What is genuinely missing is why the query stalls at all, and the walk knew it and
threw it away. Two different answers land in that branch and need opposite fixes:- the engine reporting a valid region behind the cursor (
reported_base + size <= cursor) — it is answering about memory already walked, and the page step is pointless
because the region is over; - a zero-length region reported ahead of the cursor — the engine found something and
could not size it, and stepping past it is the right move.
#107 makes the diagnostic carry
reported_baseandreported_size.PoolDiagnostics
folds numbers into#, so it stays one shape with one verbatim sample, and one live run
settles which of the two it is — and therefore whether stepping should become stopping.
That measurement is worth having before deleting anything.Separately: that diagnostic has been emitting twenty-two spaces mid-sentence since its
source line was wrapped without a continuation. Fixed in the same PR.- the engine reporting a valid region behind the cursor (
Confirmed on live 26100, three independent runs of the tier:
run stalled_pagesskipped_bytesrecovered_bytesthis issue 1,619 6,627,520 0 chain fix 2,174 8,872,160 0 + bound fix 1,116 4,541,344 0 Three different stall populations,
recovered_byteszero every time. Together with the
accounting being provably right —stalled_herelatches for the rest of the region and
every later read adds, whichtest_a_stalled_query_costs_a_page_not_the_regionpins at
two pages — the benign reading is settled: on this target every stall sits at the end of
its region's readable content and stepping over it buys nothing.The stall count roughly halved in the last run because the VS bound fix (#103) makes every
VS region one page shorter, so there is one less page per subsegment to stall on. That is
worth noting for anyone comparing runs:stalled_pagesmoves with unrelated changes, and
onlyrecovered_bytesanswers this issue.Still not removing the page-stepping — eight queries per dead region against losing every
committed page behind one bad one — but the diagnostic now carries the engine's own
reported_base/reported_size, so the next run can say which of the two stall shapes it
is and therefore whether stepping should become stopping.#107 has that, and the fix for the 22-space mangling in the message.
- added a commit that references this issue
on Aug 14, 2026 Ran it. The diagnostic paid for itself immediately — every sample carried the same
answer, and it was neither of the two shapes I listed:made no progress at 0xffffce89caf79000: the engine answered 0x0+0x0 ...caf7a000 0x0+0x0 ...caf7b000 ... through ...caf80000 0x0+0x00x0+0x0is not a region reported behind the cursor, and not a zero-length region
reported ahead of it. It is the engine saying it found nothing valid in the span it
was asked about — which is a statement about every byte through to the end of that
span. And those eight samples are eight consecutive pages: one dead region stepping
itself into the give-up bound, eight KD round trips to be told eight times what the
first answer already said.So the page step was asking an answered question, which is why
recovered_byteswas
0 on four separate runs. It is the same fact as the sibling answer #94's quiet half
already handles (a base named past the end), in the engine's other encoding, and it
now gets the same response and the same silence.The step stays for the other shape — a query that reports a region and cannot size it
has said nothing about what is behind it, and that is the case #94's coverage half was
written for.before after stalled_pages/skipped_bytes1,121 / 4,558,800 0 / 0 "valid-region query made no progress" 1,121 0 "giving up after consecutive pages" 59 0 diagnostics / categories 2,249 / 12 1,014 / 9 chunks walked 446,906 446,749 walk wall-clock 19.9s 18.3s ~1,100 round trips gone, chunk count flat. #108 has it.
- added a commit that references this issue
on Aug 14, 2026
#94 was fixed two ways: the quiet half, which stopped reporting a region running out of committed pages as a failure to advance, and the coverage half, which steps over a stalled page and keeps probing rather than abandoning everything behind it.
WalkStallswas added in the same change specifically so the second half could be judged by what it recovers rather than by whether a diagnostic category shrank.First live 26100 walk since. The first half worked. The second recovered nothing.
The measurement
valid-region query made no progressfell from 3,285 (the walk that raised pool: a valid-region query that cannot advance abandons the rest of the region, not just the page #94) to ~1,600, and a newregion # giving up after consecutive pages that would not advancefired 24 times — so the quiet half and the consecutive bound are both doing what they were built to do.skipped_bytes / stalled_pages= 4093.6, so nearly every skip was a whole page and a handful were partial — the unaligned-region case the page-boundary fix was for, firing rarely but firing.recovered_bytesis 0. Across 1,619 stalls, not one was followed by a successful read in the same region.What it probably means, and what would tell
The likely explanation is benign: after a stall, the next query reports the next valid region beyond the end of the span, the walk breaks, and there was never anything behind the stall to recover. In other words the remaining stalls all sit at the end of their region's readable content, and stepping over them buys nothing because nothing is there.
If that is so, the coverage half of #94 is dead weight on this target — it costs at most eight extra queries per dead region and returns nothing — and the honest thing is to say so rather than leave a fix in place whose stated benefit has never been observed.
The alternative is that the accounting is wrong:
stalled_hereis set in the stall branch andrecovered_bytesis added after a successfulread_extentin the samewalk_regioncall, so a stall in a region walked no further would never contribute. A region whose stall is followed by a successful read should contribute, and zero across 1,619 of them is worth one look before the benign reading is accepted.Deciding it
Record, per stall, whether the region went on to read anything — one bool, reported as a count.
n of 1619 stalls had committed memory behind themdistinguishes the two readings in a single run, whererecovered_bytes: 0cannot: it is the same number whether nothing was there or nothing was counted.Reproduce
with
WINDBG_MCP_SMOKE_KERNELsourced from the configured profile. The figures above come frompool_censusdriven over the same link, which surfaceswalk.gaps(glslang/windbg-mcp#121).