Repository navigation
[Bug][Windows]: reboot leaves native-main gate stuck; GPT-only 503 until ocx restart #2108
Description
Activity
coderabbitai commented
on Aug 19, 2026 coderabbitaiboton Aug 19, 2026 – with coderabbitaiContributorMore actions🔗 Related PRs
#780 - fix(service): make Windows scheduler stop actually stop the proxy [merged]
#805 - fix(service): bake WinSW admin env and retry stop-path probes [merged]
#1154 - fix(probe): walk Windows service definition chains (Task Scheduler + WinSW) for ownership [merged]
#1626 - fix(windows): remove native service on fresh scheduler install [open]
#1647 - fix(tray): let the service wrapper exit 0 on an already-live proxy [open]
📝 Issue Planner
Check the box below or use the
@coderabbitai plancommand to generate an implementation plan and prompts that you can use with your favorite coding assistant.- Create Plan
🧪 Issue enrichment is currently in open beta.
To disable automatic issue enrichment, add the following to your
.coderabbit.yaml:issue_enrichment: auto_enrich: enabled: false
💬 Have feedback or questions? Drop into our discord!
- addedbugSomething isn't workingSomething isn't workingcliCLI, config inject, packaging flagsCLI, config inject, packaging flagsplatformOS/service/tray/ACL (Windows-heavy, not Windows-only)OS/service/tray/ACL (Windows-heavy, not Windows-only)serviceService lifecycle (WinSW/launchd/scheduler)Service lifecycle (WinSW/launchd/scheduler)
on Aug 19, 2026 리뷰 · 우선순위 78 / 80
Windows 재부팅 뒤에 프록시는 살아 있는데 ChatGPT 네이티브 모델만 503이 나는 버그다.
/healthz는 200, Anthropic 같은 비네이티브 경로도 200인데gpt-5.6-sol만OpenCodex local native-main profile maintenance is active; retry this request로 막힌다. 혼자 풀리지 않고ocx restart면 바로 살아난다. 지금dev가 잡고 있는 native-main / Windows 서비스 / doctor-reclaim 불이랑 정확히 겹쳐서 78이다. 다른 프로바이더는 되고 수동 restart로 복구는 돼서 80은 안 넘긴다.503은 프록시가 죽은 게 아니다.
src/codex/native-profile-startup.ts의 프로세스 전역 게이트다. 스냅샷이ready가 아니면blocked이고, reason은foreign-ownership/ownership-unknown/recovery-pending/manual-recovery/owner-conflict/owner-unavailable/stage-cleanup-required중 하나다.startNativeMainStartupLifecycle()은 시작하자마자recovery-pending으로 펜스를 친 뒤retainNativeMainOwner()+observeOwner()로 수렴한다.acquiring이면 계속recovery-pending,held면convergeOwnedStartup(), 그 외는owner-unavailable또는owner-conflict로 굳는다.요청이 실제로 거절되는 지점은 게이트 파일이 아니라
src/codex/auth-context.ts다.isNativeMainTrafficBlocked()는 reason을 보지 않는다. 서비스홈이foreign-ownership/ownership-unknown이거나 스냅샷이blocked면 true다.resolveCodexAuthContext()가 네이티브 메인 트래픽을 ineligible로 보고, 고른 계정이__main__이거나 풀 대체 계정이 없으면CodexMainProfileDrainingError를 던진다. Responses/live/search/images/compact가 이걸 받아codexMainProfileDrainingResponse()로 HTTP 503,server_busy,Retry-After: 1을 만든다. 비네이티브 프로바이더는 이 경로를 안 타니까 healthz/Anthropic이 200인 게 정상이다. 문제는 reason이 요청 로그에 안 나온다는 거다.recovery-pending인지owner-unavailable인지ownership-unknown인지manual-recovery인지 한 줄로 찍어야 재현이 닫힌다.리포터 로그의
Previous session (PID ) did not shut down cleanly. Codex state restored from journal.는 native-main 암호화 저널이 아니다.src/codex/journal.ts의~/.codex/opencodex-journal.json이고, Codexconfig.toml/ 프로필을 복구하는 주입 저널이다.(PID )가 비어 있으면journal.pid가 비어 있다는 뜻이다. 게이트가 보는 저널은 따로 있다.src/codex/native-profile-store.ts의probeNativeProfileRecoveryState()가recoveryBlockPath/ nativejournalPath를 보고none | journal | manual | unreadable을 준다.journal만 자동 복구하고 나머지는manual-recovery로 정착한다. 둘을 한 원인으로 쓰면 안 된다. 다만 재부팅 +AllowHardTerminate=True면 둘 다 남을 수는 있다.스케줄러 첫 기동이 native-main을 아예 안 잡는 건 아니다. 래퍼(
src/service.ts)가OCX_SERVICE=1로ocx start --port를 띄우고,startServer()는 ownership이owned면startNativeMainStartupLifecycle()으로 재획득을 시도한다.inspectNativeCodexOwnership()이unknown/foreign이면 그때만 복구 없이blockNativeMainStartupForUnownedServiceHome()으로 영구 503이다. 프로브 타임아웃은SERVICE_PROBE_TIMEOUT_MS = 2000이다. 오너 락도 게이트를 고정할 수 있다.retainNativeMainOwner()가CODEX_HOME/.opencodex-native-main.owner.sqlite에BEGIN IMMEDIATE를 걸고,SQLITE_BUSY면 250ms 재시도, ACLETIMEDOUT한 번 더 재시도한 뒤unavailable으로 끝난다. 하드킬이면 exit hook이 DB를 못 닫는다.ocx restart가 고치는 이유는 서비스를 지웠다 다시 깔아서가 아니다. 라이브면POST /api/system/restart→acceptSystemRestart()이고, 슈퍼바이즈드 자식은process.exit(1)해서 래퍼:loop가 다시 띄운다. 그래서 두 번째service wrapper start에는 unclean-journal 줄이 없고, 새 프로세스 + 새 모듈 스냅샷 + 두 번째inspectStartupOwnership이 생긴다. 스케줄러가 재획득을 안 해서가 아니라, 첫 기동이 fail-closed로 정착한 뒤 혼자 안 풀리는 쪽에 가깝다. 그래서 고칠 것도 명확하다. 막힌 reason을 로그에 남기고, stale journal/ownership을 치운 뒤 복구를 다시 돌리고, 첫 기동이unknown/unavailable으로 굳으면 한 번 더 재획득하게 만들면 된다. #2107이 Win10+WSL 쪽 쌍둥이다. #780/#805/#1154는 stop/ownership-probe 위생 히스토리고, 이 게이트의 직접 원인은 아니다.해결방안
src/codex/native-profile-startup.ts에서 blocked로 정착한 순간 reason을 요청 로그와 service.log에 한 줄로 남겨라.recovery-pending/owner-unavailable/ownership-unknown/manual-recovery/owner-conflict를 구분하지 못하면 재부팅 재현이 닫히지 않는다. 스케줄러 첫 기동이unknown/unavailable으로 굳으면startNativeMainStartupLifecycle()을 한 번 더 돌려 stalerecoveryBlockPath와CODEX_HOME/.opencodex-native-main.owner.sqlite를 치운 뒤retainNativeMainOwner()+observeOwner()를 재시도한다.src/service.ts래퍼의 hard-terminate 경로(AllowHardTerminate=True)는 exit hook이 owner DB를 못 닫으므로, 기동 시SQLITE_BUSY/빈 PID 저널을 자동 복구 대상으로 보고manual-recovery로 영구 펜스하지 않게 한다.src/codex/auth-context.ts의 503은 reason을 숨기지 말고, healthz와 달리 native-main 게이트가 열린 뒤에만 ready로 보이게 맞춘다.ocx restart가 고치는 두 번째inspectStartupOwnership을 첫 기동에도 한 번 더 주는 것이 패치의 핵심이다.이 댓글은 grok-bot이 작성했습니다
Phase 1 of this is up as #2121 — it does not fix the sticking fence, it makes the next occurrence diagnosable. Explaining why that ordering, since it is the difference between a fix and a guess.
Your instinct about the source was right, and the ownership verdict is the structural cause:
src/server/index.ts:702-710 owned -> startNativeMainStartupLifecycle() anything else -> blockNativeMainStartupForUnownedServiceHome(...) // process lifetimestartServer()takes that verdict once and never revisits it. That is exactly why waiting never helped you andocx restartalways did — restart is the only thing that re-runs the decision.But there are two paths that produce that verdict, and your report cannot distinguish them:
- ACL fail-closed. A second
ETIMEDOUTin the icacls hardening is terminal — it publishes{ status: "unavailable", reason: "lock-unavailable" }atnative-main-owner.ts:205-212, andobserveOwner()settles toowner-unavailableand stops. Your timing fits: the first 503 is ~74s after wrapper start, past the ~60s owner budget. - Probe fail-closed.
SERVICE_PROBE_TIMEOUT_MSis 2000ms. A scheduler-only install still runssc.exe queryfor WinSW; if that times out with WinSW assets absent, the chain returnsunknownrather thanabsent(service-manager-probe.ts:732-736).
Those want different fixes. So #2121 makes the settled reason say which one, on stdout — the same stream your excerpt was already quoting, since it interleaves
[17:25:23] opencodex service wrapper startwith the 503 lines:native-main admission is fenced (reason: ownership-unknown); native model requests return 503 until it clearsOne line per distinct reason, not per request — your log shows three 503s in eleven seconds and a real client retries harder than that.
Two things you asked for that deliberately did not happen, because both would have broken something:
- The 503 message is unchanged.
claude-messages.ts:818identifies this exact response by matching that string, to keep the fence a 503 instead of remapping it to an Anthropic 529. Appending the reason there would have started telling Claude Code to back off from an upstream that was never involved. - No header, and nothing new in
/api/logs. That log populates its error text by readingerror.messageback out of the response body; headers are already gone by then, and the 503 error code is pinned toserver_is_overloadedby design. A header would have looked like a fix and shown you nothing.
One detail worth knowing when you next hit this: a 503 with no reason in the log is meaningful. The turn-drain claim race throws the same error while the startup gate reads
ready, and that path stays deliberately silent — so a reasonless 503 tells you it is not the startup fence at all. That distinction did not exist before.What would help. Next time it happens, the wrapper stdout around startup — specifically that one
native-main admission is fencedline. That single word decides phase 2, which is making a boot-timeunknownretryable whileOCX_SERVICE=1instead of a process-lifetime fence, keeping genuineforeignfail-closed. Two narrower fixes fall out of it regardless: a timed-outsc.exe querywith both WinSW xml and exe absent should not mark the machineunknown, and a second ACLETIMEDOUTshould back off and retry so a warm icacls reopens the gate withoutocx restart.Being straight about coverage: this was verified on macOS.
owner-unavailable— the branch your report most likely hit — is a Windows icacls path, and it turns out nothing in the suite asserts it today. Which is rather the point of shipping the log line first.- ACL fail-closed. A second
- added a commit that references this issue
on Aug 19, 2026 Phase 1 is on
devas of #2121 (fbc6f26a2).The 503 now names the gate reason that actually settled, on stdout:
native-main admission is fenced (reason: ownership-unknown); native model requests return 503 until it clearsThis does not fix the sticking fence and the issue stays open for that. It makes the next occurrence diagnosable, which is the prerequisite: the two candidate triggers settle to different reasons, and until one of them is named in a real report, phase 2 would be aiming at a coin flip.
So the ask is unchanged from my earlier comment — next time it happens, the wrapper stdout around startup, specifically that one
native-main admission is fencedline. That word decides whether phase 2 targets the ACL retry or the probe classification.Fixed on
devas of #2130 (890d339de). This closes the mechanism half; #2121 shipped the diagnostic earlier.Two things changed, and between them the reboot case should no longer stick.
The fence re-asks now.
startServertook the ownership verdict once, at boot, and held it for the process lifetime — which is exactly why waiting never helped you andocx restartalways did. That is right for a foreign owner, which is a fact worth refusing on. It is wrong forunknown, which only means the probe could not answer. An unknown fence now re-probes when a native request arrives and lifts itself when the host becomes answerable. It never retries for foreign, and it is capped so a permanently unaskable machine stops asking rather than probing forever.The WinSW query stops fencing over a service that cannot exist. Your install is scheduler-only, so WinSW has neither its XML nor its exe on disk — but a timed-out
sc.exe querystill returnedunknown, and that outranked the disk. With both assets absent there is nothing for a registration to belong to. Either one present keeps the oldunknown.Being straight about verification: these are Windows paths and the suites that exercise them ran on macOS and Linux. The platform CI legs are green, but the specific trigger on your machine has not been reproduced on real hardware — which is why the diagnostic from #2121 still matters.
Could you send a follow-up probe? If it recurs after a reboot, the wrapper stdout around startup, specifically this line:
native-main admission is fenced (reason: <...>); native model requests return 503 until it clearsThat word decides everything.
ownership-unknownmeans the probe path, and this fix should now be retrying it — so seeing it persist would mean the retry is not reaching your trigger.owner-unavailablemeans the ACL path, which is a different fix and is still open work: a second icaclsETIMEDOUTsettles terminal rather than backing off, and nothing in the suite asserts that branch today.And if the 503 carries no reason at all, that is informative too — it means the fence is not the startup gate.
Closing manually since PRs here target
dev. Reopen if a reboot still wedges it.- added a commit that references this issue
on Aug 19, 2026 4 remaining items
- added 12 commits that reference this issue
on Sep 17, 2026
Client or integration
Direct HTTP/API client
Area
Service lifecycle
Summary
On Windows 11 with the Task Scheduler service backend, a reboot/login can leave only the native OpenAI/Codex-login models blocked behind the process-wide
native-mainstartup gate.The proxy itself is healthy (
/healthzreturns 200), the model remains in the catalog, and non-native providers such as Anthropic continue returning 200. Native GPT requests fail locally with HTTP 503:The state does not recover on its own. Running
ocx restart(or pressing the dashboard/launcher restart control) immediately restores the same GPT model without reauthentication or configuration changes.This has recurred after multiple Windows reboots. I expected startup recovery to reacquire native-main ownership and reopen admission automatically.
Reproduction
GET /healthzreturns 200.stream: true,store: false, list-forminput) returns local HTTP 503 with the native-main maintenance message.ocx restart.The Windows scheduled task currently reports
AllowHardTerminate=TrueandMultipleInstances=IgnoreNew.Version
2.26.0 (installed and running runtime both 2.26.0)
Codex runtime reported by OpenCodex: 0.148.0
Operating system
Windows 11 Home, version 10.0.26200, build 26200
Provider and model
OpenAI (Codex login / native main) / gpt-5.6-sol
Logs or error output
The exact internal blocked reason is not exposed in the request log. The installed source suggests the relevant path may be:
native-profile-startup.ts: owner acquisition/recovery updates the process-wide snapshot toblocked.isNativeMainTrafficBlocked()keeps native traffic fenced while that snapshot remains blocked.auth-context.tsconverts this state to the generic native-main maintenance 503.The preceding unclean-shutdown/journal-recovery line is consistently associated with the post-reboot failure, but I cannot confirm whether the specific retained reason is
recovery-pending,owner-unavailable, or another startup block because it is not logged.It would help if OpenCodex:
Screenshots and supporting files
Screenshots are available if needed. They show repeated client-visible “Selected model is at capacity” errors, while OpenCodex request logs reveal the actual native-main maintenance 503.
Redacted configuration
{ "serviceBackend": "Windows Task Scheduler", "trigger": "logon", "allowHardTerminate": true, "multipleInstances": "IgnoreNew", "proxyPort": 10100, "nativeProvider": "openai (Codex login)", "nativeModel": "gpt-5.6-sol", "otherProvidersRemainHealthy": true }Checks