Skip to content

[Bug][Windows]: CPU-saturated host starves the NORMAL-priority proxy, so /healthz misses the 750 ms liveness ceiling and every CLI verb reports "Proxy not reachable" #6208

Description

@MeroZemory

Client or integration

Other

Area

Platform (Windows / macOS / Linux)

Summary

On a CPU-saturated Windows host, the proxy running as the Task Scheduler service is reported as down by every CLI surface while it is actually listening and healthy:

  • ocx service start → Service started, but no proxy answered on port 10100 after 48s
  • ocx service status → proxy not running
  • opencodex account list (and every other verb that goes through findLiveProxy) → Proxy not reachable

At the same time curl http://127.0.0.1:10100/healthz returns 200, just slowly: 1–9 s per request, p90 about 3 s. The default CLI liveness ceiling is 750 ms (SHARED_PROBE_FLOOR_MS in src/server/proxy-liveness.ts), so every probe counts as a miss.

The cause is CPU starvation of the proxy process, not a blocking call inside it:

  • / and /healthz stalled together, so the whole event loop was waiting, not one handler.
  • During most stalls the proxy had no child process, its own CPU use was low (4–25 % of one core), and its disk I/O was about zero.
  • Sampling the Bun main thread (Process.Threads from .NET) showed it in ThreadState=Ready most of the time: runnable, but not scheduled.
  • The host was saturated by unrelated NORMAL-priority work. The mix changed over the session: first a headless Android emulator (about 7 cores) plus Defender MsMpEng; after the emulator was stopped, Defender about 3 cores, then ffmpeg about 7.5 cores and python about 6 cores. The E-cores (logical 16–19) were at 100 %, and total CPU was 81–100 %.
  • The proxy ran at NORMAL priority. That is the task XML <Priority>4</Priority> from fix(service): restore normal Windows scheduler priority #3682, which fixed the older BELOW_NORMAL/background default. NORMAL still puts the proxy in the same scheduling class as the saturating work, so it waits in line with it.

Raising the live proxy to ABOVE_NORMAL ((Get-Process -Id <pid>).PriorityClass = 'AboveNormal') fixed it immediately:

Condition /healthz latency CLI verbs
NORMAL, saturated host 1–9 s, p90 ~3 s every verb: "Proxy not reachable"
ABOVE_NORMAL, same load 0.01–0.65 s work
ABOVE_NORMAL, heavier 100 % saturation p50 0.40 s, p90 3.0 s intermittent; with OCX_PROBE_TIMEOUT_MS=5000, account list succeeded 8/8

Expected: a mostly idle, latency-sensitive local proxy stays responsive to its own health probe when other processes saturate the CPU.

Related but different:

Proposed fix: at startup on win32, the proxy process raises its own priority to ABOVE_NORMAL (os.setPriority(0, PRIORITY_ABOVE_NORMAL)), best-effort and never fatal. I tested it in Bun on Windows: after the call, getPriority() returns -7. HIGH is intentionally not proposed, because it can starve the interactive desktop and ABOVE_NORMAL was enough. PR to follow.

Reproduction

  1. Windows 11, opencodex installed via npm and running as the Task Scheduler service (ocx service install). The proxy is at NORMAL priority.
  2. Saturate the CPU with NORMAL-priority work so total CPU stays at 80–100 %. I observed it with a real mix (Defender scan, ffmpeg encode, a Python job, a headless Android emulator). I have not tried a synthetic load, but one busy loop per logical CPU should have the same effect.
  3. Run ocx service status, opencodex account list, or ocx service start: they report the proxy as down or unreachable.
  4. At the same time run curl -w "%{time_total}\n" http://127.0.0.1:10100/healthz in a loop: 200 responses taking 1–9 s.
  5. (Get-Process -Id <proxy pid>).PriorityClass = 'AboveNormal', then repeat steps 3–4: /healthz answers in well under a second and the CLI verbs work.

The problem depends on load. On an idle host everything is fine.

Version

2.69.0 (also verified on dev at 09f8e5e: no setPriority anywhere in src/, SHARED_PROBE_FLOOR_MS = 750)

Operating system

Windows 11 Pro 10.0.26200, Intel i7-12700K (8P + 4E, 20 logical)

Provider and model

No response

Logs or error output

> ocx service start
Service started, but no proxy answered on port 10100 after 48s

> ocx service status
proxy not running

> opencodex account list
Proxy not reachable

Meanwhile, /healthz returned HTTP 200 on every sample, taking 1–9 s each (p90 about 3 s).

Screenshots and supporting files

No response

Redacted configuration

No response

Checks

  • I searched existing issues and documentation.
  • I removed secrets, tokens, account details, request credentials, and personal data.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't workingplatformOS/service/tray/ACL (Windows-heavy, not Windows-only)priority: P2Medium: provider/client-specific bug with a workaround, bounded enhancement tied to a tracked issue,serviceService lifecycle (WinSW/launchd/scheduler)

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions