Skip to content

pool: the page-stepping half of #94 recovered nothing on a live walk (recovered_bytes = 0) #104

Description

@glslang

#94 was fixed two ways: the quiet half, which stopped reporting a region running out of committed pages as a failure to advance, and the coverage half, which steps over a stalled page and keeps probing rather than abandoning everything behind it. WalkStalls was added in the same change specifically so the second half could be judged by what it recovers rather than by whether a diagnostic category shrank.

First live 26100 walk since. The first half worked. The second recovered nothing.

The measurement

"gaps": {
  "stalled_pages": 1619,
  "skipped_bytes": 6627520,
  "recovered_bytes": 0,
  "refused_chunks": 106516
}
  • valid-region query made no progress fell from 3,285 (the walk that raised pool: a valid-region query that cannot advance abandons the rest of the region, not just the page #94) to ~1,600, and a new region # giving up after consecutive pages that would not advance fired 24 times — so the quiet half and the consecutive bound are both doing what they were built to do.
  • skipped_bytes / stalled_pages = 4093.6, so nearly every skip was a whole page and a handful were partial — the unaligned-region case the page-boundary fix was for, firing rarely but firing.
  • recovered_bytes is 0. Across 1,619 stalls, not one was followed by a successful read in the same region.

What it probably means, and what would tell

The likely explanation is benign: after a stall, the next query reports the next valid region beyond the end of the span, the walk breaks, and there was never anything behind the stall to recover. In other words the remaining stalls all sit at the end of their region's readable content, and stepping over them buys nothing because nothing is there.

If that is so, the coverage half of #94 is dead weight on this target — it costs at most eight extra queries per dead region and returns nothing — and the honest thing is to say so rather than leave a fix in place whose stated benefit has never been observed.

The alternative is that the accounting is wrong: stalled_here is set in the stall branch and recovered_bytes is added after a successful read_extent in the same walk_region call, so a stall in a region walked no further would never contribute. A region whose stall is followed by a successful read should contribute, and zero across 1,619 of them is worth one look before the benign reading is accepted.

Deciding it

Record, per stall, whether the region went on to read anything — one bool, reported as a count. n of 1619 stalls had committed memory behind them distinguishes the two readings in a single run, where recovered_bytes: 0 cannot: it is the same number whether nothing was there or nothing was counted.

Reproduce

cargo test --test mcp_smoke -- --ignored --nocapture --test-threads=1 live_kernel

with WINDBG_MCP_SMOKE_KERNEL sourced from the configured profile. The figures above come from pool_census driven over the same link, which surfaces walk.gaps (glslang/windbg-mcp#121).

Activity

  1. glslang commented on Aug 14, 2026

    @glslang
    OwnerAuthor

    Still relevant — walk_region's stall path is untouched since. But the alternative
    reading is disproved without needing another run, and by something already in the tree.

    The accounting is right. stalled_here is set in the stall branch and never
    cleared, so it latches for the rest of the region, and every later successful
    read_extent in that region adds to recovered_bytes. The case the issue worried
    about — "a stall in a region walked no further would never contribute" — is real but is
    not a mis-count: a region walked no further has nothing to contribute.
    test_a_stalled_query_costs_a_page_not_the_region already pins
    recovered_bytes: 2 * PAGE_SIZE for a stall with two committed pages behind it, so a
    stall that is followed by a read does show up.

    So the benign reading is the only one left: on 26100 every one of those 1,619 stalls
    sits at the end of its region's readable content, and the page-stepping recovers
    nothing there. The extra bool this issue proposed would report 0 of 1619 and tell us
    what recovered_bytes: 0 already does.

    I would not remove it, though. It costs at most eight queries per dead region, and
    what it replaced was losing every committed page behind one bad one — the samples in
    #94 fired part-way through regions of 0x10fd0 and 0x3afc0 bytes.

    What is genuinely missing is why the query stalls at all, and the walk knew it and
    threw it away. Two different answers land in that branch and need opposite fixes:

    • the engine reporting a valid region behind the cursor (reported_base + size <= cursor) — it is answering about memory already walked, and the page step is pointless
      because the region is over;
    • a zero-length region reported ahead of the cursor — the engine found something and
      could not size it, and stepping past it is the right move.

    #107 makes the diagnostic carry reported_base and reported_size. PoolDiagnostics
    folds numbers into #, so it stays one shape with one verbatim sample, and one live run
    settles which of the two it is — and therefore whether stepping should become stopping.
    That measurement is worth having before deleting anything.

    Separately: that diagnostic has been emitting twenty-two spaces mid-sentence since its
    source line was wrapped without a continuation. Fixed in the same PR.

  2. glslang commented on Aug 14, 2026

    @glslang
    OwnerAuthor

    Confirmed on live 26100, three independent runs of the tier:

    run stalled_pages skipped_bytes recovered_bytes
    this issue 1,619 6,627,520 0
    chain fix 2,174 8,872,160 0
    + bound fix 1,116 4,541,344 0

    Three different stall populations, recovered_bytes zero every time. Together with the
    accounting being provably right — stalled_here latches for the rest of the region and
    every later read adds, which test_a_stalled_query_costs_a_page_not_the_region pins at
    two pages — the benign reading is settled: on this target every stall sits at the end of
    its region's readable content and stepping over it buys nothing.

    The stall count roughly halved in the last run because the VS bound fix (#103) makes every
    VS region one page shorter, so there is one less page per subsegment to stall on. That is
    worth noting for anyone comparing runs: stalled_pages moves with unrelated changes, and
    only recovered_bytes answers this issue.

    Still not removing the page-stepping — eight queries per dead region against losing every
    committed page behind one bad one — but the diagnostic now carries the engine's own
    reported_base/reported_size, so the next run can say which of the two stall shapes it
    is and therefore whether stepping should become stopping.

    #107 has that, and the fix for the 22-space mangling in the message.

  3. glslang commented on Aug 14, 2026

    @glslang
    OwnerAuthor

    Ran it. The diagnostic paid for itself immediately — every sample carried the same
    answer, and it was neither of the two shapes I listed:

    made no progress at 0xffffce89caf79000: the engine answered 0x0+0x0
                        ...caf7a000                              0x0+0x0
                        ...caf7b000  ... through ...caf80000      0x0+0x0
    

    0x0+0x0 is not a region reported behind the cursor, and not a zero-length region
    reported ahead of it. It is the engine saying it found nothing valid in the span it
    was asked about
    — which is a statement about every byte through to the end of that
    span. And those eight samples are eight consecutive pages: one dead region stepping
    itself into the give-up bound, eight KD round trips to be told eight times what the
    first answer already said.

    So the page step was asking an answered question, which is why recovered_bytes was
    0 on four separate runs. It is the same fact as the sibling answer #94's quiet half
    already handles (a base named past the end), in the engine's other encoding, and it
    now gets the same response and the same silence.

    The step stays for the other shape — a query that reports a region and cannot size it
    has said nothing about what is behind it, and that is the case #94's coverage half was
    written for.

    before after
    stalled_pages / skipped_bytes 1,121 / 4,558,800 0 / 0
    "valid-region query made no progress" 1,121 0
    "giving up after consecutive pages" 59 0
    diagnostics / categories 2,249 / 12 1,014 / 9
    chunks walked 446,906 446,749
    walk wall-clock 19.9s 18.3s

    ~1,100 round trips gone, chunk count flat. #108 has it.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions