Skip to content

67% of CI compute cannot move to our own runners: 17 container: jobs blocked on podman-in-podman on pulseengine-ci-01 #419

Description

@avrabe

[fathom (gale) — measured, and the remaining work is admin-side rather than repo-side]

#418 moves the cheap required gates onto pulseengine-ci-01. It does not touch the part that matters for queue depth, because that part cannot move yet.

The measurement

Last merge to main (aa48860), 76 jobs: 630 job-minutes, 59 minutes wall-clock. Where it goes:

job-min share
Zephyr C Coverage 39.3
Zephyr C Coverage (gale-enabled) 23.4
34 × qemu_cortex_m3 twister jobs 360 (median 10.4 each)
Zephyr subtotal ~423 67%
everything else ~207 33%

And the queue behaviour that makes it hurt: while #416 and #417 were open, 78 checks were queued at once, and the thread sanitizer sat QUEUED for over two hours with no runner assigned. The two self-hosted sanitizer jobs in the same workflow were picked up in three minutes (pulseengine-ci-01-6 and -10, concurrently). Instances -5, -6, -7, -9, -10, -11, -12 have all taken jobs recently, so the pool has real width.

So the machine is there and idle enough; the work that would use it is exactly the work that cannot run there.

Why it cannot run there

17 job definitions use container: — every Zephyr job, LLVM-LTO, all four Renode engine benches, the wasm dist builds, renode-test, size-comparison. The self-hosted runners are themselves containers, so those need container-in-container, which is the parked work on pulseengine-ci-01:

  1. podman graphroot onto the containers-ralf volume (the root filesystem cannot hold Zephyr's SDK + workspace — this is the same disk pressure that produced the /mnt design in zephyr-tests.yml on hosted runners, where /mnt is GitHub's ~70 GB ephemeral scratch and does not exist on ours).
  2. subuid / newuidmap for rootless podman.

Both are host configuration on pulseengine-ci-01, not changes to this repo.

The smaller, separate ask

Three jobs are movable the moment three packages exist in the rust-cpu image, because they only fail on apt under a container with no passwordless sudo:

  • qemu-system-arm — unblocks mpu-enforcement (a REQUIRED context: "a denied write really faults (REQ-OS-MPU-001 kill-criterion)")
  • binutils-arm-none-eabi — unblocks the cross-arch seam gate (REQUIRED)
  • curl — unblocks Rust Coverage (REQUIRED)

With those three, 19 of the 22 required contexts run on our own hardware and a merge stops depending on hosted capacity at all. The remaining three are the container ones above.

What this issue is asking for

Nothing in the repo. Two host changes on pulseengine-ci-01 (podman graphroot + subuid/newuidmap) and optionally three packages in the runner image. #418 is the repo-side half and is independent of both — it lands value now and does not depend on this.

Kill-criterion for the container half, so it is checkable rather than declared done: a container: job (start with renode-test, the smallest) completes on a pulseengine-ci-01-* runner, and zephyr-tests.yml's workspace fits without the /mnt assumptions it currently carries.

Activity

  1. avrabe commented on Sep 23, 2026

    @avrabe
    ContributorAuthor

    [fathom (gale) — add gh to the image ask, and a finding about WHY it matters]

    #418's first run on pulseengine-ci-01 turned up a fourth missing tool, and the way it surfaced is worth recording.

    gh is not in the rust-cpu image. zephyr-fork-pin failed loudly with exit 127, gh: command not found — it uses gh api for the Zephyr-fork ancestry query and its negative control. Ordinary; it moves back to hosted for now.

    required-contexts did not fail. It PASSED, having checked half of what it claims. That job also uses gh api, to read main's branch protection, and it degrades rather than fails when it cannot:

    if gh api ".../branches/main/protection" > /tmp/protection.json; then
      readable=true    # committed list checked against what main REQUIRES
    else
      readable=false   # committed list checked against the workflows only
    fi
    

    With no gh binary the else branch runs, the list-vs-protection drift check never happens, and a required context reports success. In the job whose entire purpose is "every required context can actually report".

    The degradation is deliberate and correct for its intended case — a token that lacks permission to read protection. It does not distinguish that from an image that lacks the binary. The first is a permission boundary worth degrading across; the second is a broken environment, and degrading across it silently narrows a gate.

    #418 now carries a command -v gh guard that fails before the branch is reached, so the missing-binary case can never take the permission path again — on any runner, hosted or not. That guard is the part that outlives this issue.

    So the image ask is four packages, not three:

    package unblocks required context?
    qemu-system-arm mpu-enforcement yes
    binutils-arm-none-eabi cross-arch seam gate yes
    curl Rust Coverage yes
    gh zephyr-fork-pin, required-contexts one of two

    And one observation from the same run, relevant to the podman half of this issue: the post-job hook reported disk: 70% used (threshold 70%), runner=runner5 and ran a cleanup that freed nothing measurable (70% → 70%, 1 paths removed, just the cargo registry cache). That is before any Zephyr workload exists on the box. Whatever the graphroot ends up on will need real headroom, not the current root filesystem.

  2. avrabe commented on Sep 30, 2026

    @avrabe
    ContributorAuthor

    [fathom (gale) — the runner queue is now blocking a release tag, measured]

    Escalating with a concrete instance rather than a projection, because this issue has been sitting on "67% of compute cannot move" and the cost has become specific.

    v0.10.0 is verified, its notes are merged, and it cannot be tagged, because main's CI at the tag commit (0c4f089) has five checks that have now been queued for 63 minutes (17:50:43Z → 18:54:06Z) with no runner assigned:

    queued  a Rocq proof that is merely STATED cannot pass as proven
    queued  committed objects are not older than their sources
    queued  fused data segments are disjoint (gale#266)
    queued  VER-OS-WCET-001 kill-criterion can still fail
    queued  no CI retry loop can swallow its own failure
    

    All five are [self-hosted, linux, x64, rust-cpu]. In the same run, four jobs did get runners — pulseengine-ci-01-11 (×2), -6, -7 — so the pool is alive and serving roughly three instances concurrently while five wait.

    This is not an argument that #418 was wrong. Before that change these same gates queued behind GitHub-hosted capacity, where one job was measured waiting over two hours. The move made the wait shorter and ours to fix. It did not make it bounded, and I corrected that claim on #418 when I realised timeout-minutes bounds run time rather than queue time.

    It is an argument that the capacity half is now the binding constraint on releasing, not just on iteration speed. Two levers, both on the host and both in this issue:

    1. More instances online. Cheapest, no repo change. rivet#1018's probe hit the same pool at 74 minutes queued, so this is not specific to gale.
    2. The podman work, which would let the 17 container: jobs — Zephyr, LLVM-LTO, the Renode engine benches — leave the hosted pool entirely and stop the two queues competing for the same wall clock.

    One thing worth knowing for the first lever: rivet#1018 records that the runner list needs the administration scope, which cannot be granted to GITHUB_TOKEN, so no workflow in any repo can assess pool capacity by itself. Whatever monitoring exists has to be run with an org-scoped credential — which is also why neither that probe nor this comment can tell you how many instances are actually online, only how many took work.

    No change requested in gale. Reporting the cost with a number attached.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions