Repository navigation
67% of CI compute cannot move to our own runners: 17 container: jobs blocked on podman-in-podman on pulseengine-ci-01 #419
Description
Activity
[fathom (gale) — add
ghto the image ask, and a finding about WHY it matters]#418's first run on
pulseengine-ci-01turned up a fourth missing tool, and the way it surfaced is worth recording.ghis not in the rust-cpu image.zephyr-fork-pinfailed loudly withexit 127, gh: command not found— it usesgh apifor the Zephyr-fork ancestry query and its negative control. Ordinary; it moves back to hosted for now.required-contextsdid not fail. It PASSED, having checked half of what it claims. That job also usesgh api, to readmain's branch protection, and it degrades rather than fails when it cannot:if gh api ".../branches/main/protection" > /tmp/protection.json; then readable=true # committed list checked against what main REQUIRES else readable=false # committed list checked against the workflows only fiWith no
ghbinary the else branch runs, the list-vs-protection drift check never happens, and a required context reports success. In the job whose entire purpose is "every required context can actually report".The degradation is deliberate and correct for its intended case — a token that lacks permission to read protection. It does not distinguish that from an image that lacks the binary. The first is a permission boundary worth degrading across; the second is a broken environment, and degrading across it silently narrows a gate.
#418 now carries a
command -v ghguard that fails before the branch is reached, so the missing-binary case can never take the permission path again — on any runner, hosted or not. That guard is the part that outlives this issue.So the image ask is four packages, not three:
package unblocks required context? qemu-system-armmpu-enforcementyes binutils-arm-none-eabicross-arch seam gate yes curlRust Coverageyes ghzephyr-fork-pin,required-contextsone of two And one observation from the same run, relevant to the podman half of this issue: the post-job hook reported
disk: 70% used (threshold 70%), runner=runner5and ran a cleanup that freed nothing measurable (70% → 70%, 1 paths removed, just the cargo registry cache). That is before any Zephyr workload exists on the box. Whatever the graphroot ends up on will need real headroom, not the current root filesystem.[fathom (gale) — the runner queue is now blocking a release tag, measured]
Escalating with a concrete instance rather than a projection, because this issue has been sitting on "67% of compute cannot move" and the cost has become specific.
v0.10.0 is verified, its notes are merged, and it cannot be tagged, because
main's CI at the tag commit (0c4f089) has five checks that have now been queued for 63 minutes (17:50:43Z → 18:54:06Z) with no runner assigned:queued a Rocq proof that is merely STATED cannot pass as proven queued committed objects are not older than their sources queued fused data segments are disjoint (gale#266) queued VER-OS-WCET-001 kill-criterion can still fail queued no CI retry loop can swallow its own failureAll five are
[self-hosted, linux, x64, rust-cpu]. In the same run, four jobs did get runners —pulseengine-ci-01-11(×2),-6,-7— so the pool is alive and serving roughly three instances concurrently while five wait.This is not an argument that #418 was wrong. Before that change these same gates queued behind GitHub-hosted capacity, where one job was measured waiting over two hours. The move made the wait shorter and ours to fix. It did not make it bounded, and I corrected that claim on #418 when I realised
timeout-minutesbounds run time rather than queue time.It is an argument that the capacity half is now the binding constraint on releasing, not just on iteration speed. Two levers, both on the host and both in this issue:
- More instances online. Cheapest, no repo change. rivet#1018's probe hit the same pool at 74 minutes queued, so this is not specific to gale.
- The podman work, which would let the 17
container:jobs — Zephyr, LLVM-LTO, the Renode engine benches — leave the hosted pool entirely and stop the two queues competing for the same wall clock.
One thing worth knowing for the first lever: rivet#1018 records that the runner list needs the
administrationscope, which cannot be granted toGITHUB_TOKEN, so no workflow in any repo can assess pool capacity by itself. Whatever monitoring exists has to be run with an org-scoped credential — which is also why neither that probe nor this comment can tell you how many instances are actually online, only how many took work.No change requested in gale. Reporting the cost with a number attached.
- added 4 commits that reference this issue
on Sep 30, 2026
[fathom (gale) — measured, and the remaining work is admin-side rather than repo-side]
#418 moves the cheap required gates onto
pulseengine-ci-01. It does not touch the part that matters for queue depth, because that part cannot move yet.The measurement
Last merge to
main(aa48860), 76 jobs: 630 job-minutes, 59 minutes wall-clock. Where it goes:Zephyr C CoverageZephyr C Coverage (gale-enabled)qemu_cortex_m3twister jobsAnd the queue behaviour that makes it hurt: while #416 and #417 were open, 78 checks were queued at once, and the
threadsanitizer sat QUEUED for over two hours with no runner assigned. The two self-hosted sanitizer jobs in the same workflow were picked up in three minutes (pulseengine-ci-01-6and-10, concurrently). Instances-5, -6, -7, -9, -10, -11, -12have all taken jobs recently, so the pool has real width.So the machine is there and idle enough; the work that would use it is exactly the work that cannot run there.
Why it cannot run there
17 job definitions use
container:— every Zephyr job, LLVM-LTO, all four Renode engine benches, the wasm dist builds,renode-test,size-comparison. The self-hosted runners are themselves containers, so those need container-in-container, which is the parked work onpulseengine-ci-01:containers-ralfvolume (the root filesystem cannot hold Zephyr's SDK + workspace — this is the same disk pressure that produced the/mntdesign inzephyr-tests.ymlon hosted runners, where/mntis GitHub's ~70 GB ephemeral scratch and does not exist on ours).Both are host configuration on
pulseengine-ci-01, not changes to this repo.The smaller, separate ask
Three jobs are movable the moment three packages exist in the
rust-cpuimage, because they only fail onaptunder a container with no passwordless sudo:qemu-system-arm— unblocksmpu-enforcement(a REQUIRED context: "a denied write really faults (REQ-OS-MPU-001 kill-criterion)")binutils-arm-none-eabi— unblocks the cross-arch seam gate (REQUIRED)curl— unblocksRust Coverage(REQUIRED)With those three, 19 of the 22 required contexts run on our own hardware and a merge stops depending on hosted capacity at all. The remaining three are the container ones above.
What this issue is asking for
Nothing in the repo. Two host changes on
pulseengine-ci-01(podman graphroot + subuid/newuidmap) and optionally three packages in the runner image. #418 is the repo-side half and is independent of both — it lands value now and does not depend on this.Kill-criterion for the container half, so it is checkable rather than declared done: a
container:job (start withrenode-test, the smallest) completes on apulseengine-ci-01-*runner, andzephyr-tests.yml's workspace fits without the/mntassumptions it currently carries.