Repository navigation
[Bug][CI]: a dispatched run shares its concurrency group with pushes, so the macOS control lane is cancelled by the next merge #5037
Description
Activity
- addedbugSomething isn't workingSomething isn't workingplatformOS/service/tray/ACL (Windows-heavy, not Windows-only)OS/service/tray/ACL (Windows-heavy, not Windows-only)
on Sep 18, 2026 Confirmed from the workflow contract and the run timestamps. This is separate from the macOS timeout: the dispatched control job was cancelled before it executed.
The fix should preserve supersession for ordinary
push/pull_requestruns while giving eachworkflow_dispatchan immutable group. Keying dispatches bygithub.run_idis safer than adding onlygithub.event_name: event-name separation prevents pushes from cancelling a dispatch, but two operator dispatches on the same ref would still cancel one another even though each asks for evidence about a specific selected SHA.Please keep this as a narrow workflow/security-boundary PR and add a workflow-source regression proving:
- two pushes on the same ref still share a group and the newer one cancels the older;
- a pull-request head keeps its existing supersession behavior;
- a manual dispatch uses a run-unique group;
- a push to
devcannot cancel that dispatch; and - the dispatch does not gain permissions or secret exposure.
The temporary immutable probe ref is a valid workaround, not the final contract.
리뷰 · 우선순위 75 / 80
이 이슈는 CI 버그입니다.
ci.yml이concurrency.group: cross-platform-ci-${{ github.ref }}에cancel-in-progress: true를 걸고 있어서,workflow_dispatch로dev에 돌린 긴 레인(특히macos control)이 같은 ref의push(merge)와 그룹을 공유합니다. 그래서 다음 merge가 dispatch를 취소합니다. 지금dev(facd2b6ca) checkout의.github/workflows/ci.yml62-63행 근처가 그대로입니다.증거도 분명합니다. run 35318264610은
lane=macos-controldispatch였는데, jobstarted_at과completed_at이 같은 초이고 conclusion이 cancelled입니다. 타임아웃이 아니라 동시성 취소입니다. #4905/#5028은 timeout-minutes(75분) 예산 문제였고, 예산을 올려도 이 취소는 안 사라집니다. 50분짜리 레인이 바쁜dev에서 끝까지 갈 확률은 거의 없습니다. cancelled는 pass도 fail도 아니라서, 릴리스 증거용 dispatch가 조용히 증발합니다. 본문이 제안한 frozen-ref 워크어라운드(ci/control-probe-2590)는 임시로 맞고, 최종 계약은 아닙니다.고치는 방향도 이슈·Ingwannu 확인 댓글과 같습니다. push/PR의 supersession은 유지하고, dispatch만 불변 그룹을 가져야 합니다.
github.run_id로 키를 나누는 편이event_name만 나누는 것보다 낫습니다. 같은 ref에 dispatch를 두 번 해도 서로 취소하지 않아야, “특정 SHA 증거”라는 질문이 두 개일 때 의미가 살아납니다. 워크플로 파일이라AGENTS.md보안 리뷰 경계에 걸리고, 권한/시크릿을 늘리면 안 됩니다.이미 열린 수정 PR이 있습니다. #5048
fix(ci): give a dispatched run its own concurrency group이 Closes #5037을 달고, run_id 그룹 + push/PR 취소 유지 + 표현식 평가 테스트를 넣었습니다. grok-bot이 #5048에 76/80으로 리뷰해 두었습니다. 이 이슈 자체는 추가 설계 논쟁 없이 #5048 랜딩으로 닫히면 됩니다.라인 62-63 (.github/workflows/ci.yml, 현재
dev) -cross-platform-ci-${{ github.ref }}+ cancel-in-progress. dispatch와 push 충돌의 직접 원인.
이슈 본문 run 35318264610 - started=completed=cancelled. 예산(#5028)과 다른 축의 증거.
경로 PR #5048 - 이미 Closes #5037. event_name/run_id 분기 +tests/ci-workflows/ci-concurrency-groups.test.ts.
경로 #4905 / #5028 - timeout 예산. 이 이슈와 분리해서 생각해야 한다.메인테이너의 판단이 필요한 지점
- fix(ci): give a dispatched run its own concurrency group #5048 보안 리뷰 서명만 남았는지, 추가 워크플로 변경을 원하는지.
- merge 후 frozen-ref 브랜치 정리 여부.
- dispatch 병렬로 러너 분이 겹쳐도 괜찮은지(의도적 비용).
너의 추천
이 이슈는 #5048 merge로 닫기. 별도 패치나 rebase를 새로 받지 말고, #5048이 green + 보안 한 줄이면 랜딩. merge 뒤dev에macos-controldispatch 한 번으로 “merge가 와도 cancelled가 아니다”를 확인하면 증거가 끝난다.이 댓글은 grok-bot이 작성했습니다
- added a commit that references this issue
on Sep 18, 2026 Fixed on
devby #5048, merged as1a135f112e.ci.ymlnow keys aworkflow_dispatchongithub.run_id, which is unique per run, so every dispatch is a group of one: a push to the same ref cannot cancel it, and neither can a second dispatch of that ref.pushandpull_requestkeep the ref-keyed group andcancel-in-progress, because superseding an older head is exactly right for them.tests/ci-workflows/ci-concurrency-groups.test.tspins it by evaluating the two expressions rather than matching their text — a spelling assertion would survive any rewrite that kept the words. It implements the narrow grammar the workflow uses and refuses anything outside it, with a case asserting that refusal, so the model cannot quietly stop describing the thing it models. A fourth trigger on this workflow also fails it, since such a trigger would inherit the push answer with no decision recorded.The frozen-ref workaround this issue described (
ci/control-probe-2590at56a99d3848) is no longer needed. It is still in use by run 35318628931, which as of this writing has held itsmacos controljob for over eighty minutes — the first time that lane has been allowed to run instead of being cancelled at startup. The ref will be removed once that run reports.
Client or integration
CI only.
Area
CI and release
Summary
The unsharded
macos controllane cannot complete on a branch that is receiving merges, independently of its job budget.ci.ymlsetsconcurrency: cross-platform-ci-${{ github.ref }}withcancel-in-progress: true, and aworkflow_dispatchagainstdevlands in the same group as thepushruns ondev. So the next merge cancels the dispatch.This is separate from #4905, which was about the budget, and fixing that one did not fix this one. Immediately after the budget went to 75 minutes I dispatched
lane=macos-controlondevat89b7d880d9: run 35318264610, job started 07:14:37Z and was cancelled at 07:14:37Z — the same second. That is a concurrency cancellation at startup, not a timeout, and the run had been queued since 07:11:41Z.The consequence is that this lane is effectively uncompletable during any active development window. It is the longest job in the workflow at roughly fifty minutes, so on a busy day the probability that no merge lands inside its window is close to zero. That, together with the budget, explains why it has produced cancellations for months and why those cancellations were read as runner capacity.
It also means a maintainer who dispatches the control lane to get release evidence will usually get nothing back, and may not notice, because a cancelled job reports neither pass nor fail.
Reproduction
ci.ymlwithlane=macos-controlagainstdev.devwhile it runs.started_atequal tocompleted_atwhen the merge arrives before it starts work.Current workaround, which is what release evidence should use in the meantime: push the candidate SHA to a ref nothing else writes to and dispatch against that.
ci/control-probe-2590at56a99d3848is a live example.Version
2.59.0 (dev).
Operating system
Not OS-specific, though it is the macOS control lane that suffers because it is the longest.
Provider and model
Not provider-specific.
Logs or error output
Screenshots and supporting files
Not applicable.
Redacted configuration
Not applicable.
On the fix, as analysis rather than a decision: the concurrency group is right for
pushandpull_request, where superseding an older head is exactly what you want. It is wrong forworkflow_dispatch, where the operator has asked for evidence about one specific SHA. Keying the group ongithub.event_nameas well asgithub.ref, or ongithub.run_idfor dispatches, would separate them. That is a GitHub Actions workflow change and falls under the explicit security-review boundary inAGENTS.md, so it belongs in its own pull request.