Skip to content

31 of 53 CI jobs have no timeout at job or step level — on a repo whose bottleneck is runner starvation #382

Description

@avrabe

[fathom (gale) — measured while waiting on a 22-minute qemu job]

sched_metairq (qemu_cortex_m3) held #380 for ~22 minutes. It was not hung — it
completed and the PR went green. But checking whether it could hang turned up
something worth recording.

The measurement

Counting jobs with no timeout-minutes at either job or step level:

workflow                         jobs  job-TO  step-TO  UNBOUNDED
bazel-tests.yml                     3       0        0   3
compliance.yml                      1       0        0   1
coverage.yml                        2       0        0   2
drv-components.yml                  1       0        0   1
drv-cross-arch.yml                  1       0        0   1
engine-bench-smoke.yml              1       0        0   1
formal-verification.yml             5       0        4   1
kill-criteria.yml                   9       0        0   9
llvm-lto.yml                        2       0        0   2
pages.yml                           2       0        0   2
release-wasm.yml                    1       0        0   1
release.yml                         2       0        0   2
rivet-v-closure.yml                 2       1        0   1
zephyr-tests.yml                    4       0        0   4

jobs with NO timeout at job or step level: 31 of 53

Only formal-verification.yml uses step-level timeouts (4 of 5 jobs), and
rivet-v-closure.yml has the single job-level one. Everything else runs to
GitHub's 6-hour default.

Why it matters here specifically

This repo's documented bottleneck is runner starvation — the 2-PR ceiling exists
because ~75 jobs per PR against ~1-2 concurrent runners already serialises the
queue (measured earlier: 24 runs queued, 1 in progress, oldest waiting 11
minutes). A single job that hangs holds a runner for six hours against that.

The exposure is not hypothetical in kind: qemu tests that wait on serial output
are exactly the shape that hangs when the guest never reaches its exit path, and
zephyr-tests.yml — 4 jobs, the largest matrix — has no timeout at any level.

Not proposing a blanket number

Timeouts should be sized from measured durations, not guessed, or they become a
new flake source. The honest next step is to measure the p95 per job over recent
runs and set ~2-3x that. I have not done the measurement, so I am not opening a
PR with invented numbers.

Two jobs are worth doing first regardless, because they are the ones that block
everything else: zephyr-tests.yml (observed 22 min for one qemu test) and
kill-criteria.yml (9 jobs, all short, so a generous bound costs nothing and
catches a genuine hang immediately).

Method note

My first pass at this audit was wrong and said every workflow had a timeout.
grep -c prints 0 and exits 1 when there is no match, so a
grep -c … || echo 0 idiom produced "0\n0", which never compared equal to
"0". Same output-vs-exit-status conflation the retry-loop sweep was about
(#370), in the audit that was looking for it.

Activity

avrabe commented on Sep 23, 2026

@avrabe
ContributorAuthor

[fathom (gale) — 14 of them close in #418, as a side effect, and the title's premise just got measured]

#418 moves the cheap required gates onto our own runners and gives every job it touches timeout-minutes: 30 — not for tidiness but because runs-on has no fallback, so a self-hosted pool that goes down would otherwise hang a required context indefinitely. That takes the count from 31 of 53 → 27 of 58 (the denominator grew; five jobs have been added since this issue was filed).

Remaining 27, unchanged by that PR.

The premise in this title is now a number rather than an impression. On the last merge to main (aa48860, 76 jobs): 630 job-minutes, 59 min wall-clock. While #416 and #417 were open, 78 checks were queued at once, and the thread sanitizer sat QUEUED for over two hours with no runner assigned — while the two self-hosted sanitizer jobs in the same workflow were picked up in three minutes. So "runner starvation is the bottleneck" is correct, and it is specifically hosted-runner starvation: pulseengine-ci-01 instances -5 through -12 were taking jobs the whole time.

That reframes what a timeout is for here. On a starved pool the dangerous failure is not a job that runs too long — it is a job that never starts and holds a required context open. A timeout bounds the first and does nothing about the second; #419 (container jobs cannot move to our runners) is what addresses the second. Worth keeping both straight rather than letting this issue absorb the queue problem.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions