Skip to content

[Bug]: Quota-limited previous-model compact blocks Sol → routed DeepSeek handoff after manual /compact succeeds #2723

Description

@juzijia

Client or integration

Codex App (Desktop) through OpenCodex loopback proxy integration.

Area

Proxy and routing

Summary

A thread that previously used gpt-5.6-sol can be switched in-place to the routed model deepseek/deepseek-v4-flash, and a manual /compact succeeds with the current routed DeepSeek model. However, the next normal user turn triggers Codex's pre-sampling previous-model compaction. The first compaction attempt intentionally uses the previous model (gpt-5.6-sol). If the Sol/ChatGPT quota is exhausted, OpenCodex records that compact attempt as an upstream failure and the turn stops before the current routed DeepSeek model can continue.

Observed failing compact record:

provider=openai
model=gpt-5.6-sol
status=502
code=upstream_server_error
termination=incomplete
error=The usage limit has been reached

No subsequent compact retry to deepseek/deepseek-v4-flash was observed for that blocked turn.

This report is not claiming that the first previous-model compact attempt is incorrect. Upstream Codex intentionally performs previous-model pre-sampling compaction during some model transitions. The compatibility problem is that, through OpenCodex's same-thread routed-model setup, a quota failure on the previous OpenAI model prevents the handoff to the already-working current routed model.

Upstream Codex has also broadened previous-model compact fallback to include usage-limit, unexpected-status, server and exhausted-retry failures:

Expected OpenCodex behavior is one of the following, depending on what context is available at the proxy layer:

  1. Preserve/translate the previous-model quota failure so Codex can trigger its selected/current-model compact fallback; or
  2. provide an equivalent OpenCodex compatibility fallback from the unavailable previous route to the current routed model; or
  3. if neither is safe, surface/document this limitation explicitly instead of leaving the routed thread blocked by the previous provider's quota.

Reproduction

  1. Run OpenCodex 2.31.0 with Codex App loopback integration and a routed DeepSeek model available as deepseek/deepseek-v4-flash.
  2. Start or resume a thread that has been using gpt-5.6-sol.
  3. Exhaust the ChatGPT/OpenAI usage allowance for gpt-5.6-sol.
  4. In the same thread, switch the current model to deepseek/deepseek-v4-flash.
  5. Run /compact manually.
  6. Observe that manual compaction succeeds with the current routed model.
  7. Send a normal user message to continue the thread.
  8. Codex performs pre-sampling model-transition compaction against the previous model gpt-5.6-sol.
  9. OpenCodex records the compact request as provider=openai, model=gpt-5.6-sol, status=502, error=The usage limit has been reached.
  10. The turn stops; no DeepSeek compact fallback is observed.
  11. After the Sol quota resets, the same model-switch workflow can continue normally again.

This reproduces a narrow handoff failure: DeepSeek compaction itself is usable, but the previous-model pre-sampling compact failure blocks continuation before the routed current model gets control.

Version

OpenCodex 2.31.0

Operating system

Windows x64 (NT 10.0; exact edition/build not captured)

Provider and model

Previous model/provider:

OpenAI / ChatGPT
model=gpt-5.6-sol

Current routed model:

DeepSeek via OpenCodex
model=deepseek/deepseek-v4-flash

Logs or error output

Error running remote compact task: stream disconnected before completion:
stream closed before response.completed

Matching OpenCodex usage record at the compact failure timestamp:

provider=openai
model=gpt-5.6-sol
status=502
code=upstream_server_error
termination=incomplete
error=The usage limit has been reached

Control observation:

manual /compact on current routed DeepSeek -> succeeds
next normal turn -> previous-model Sol pre-compact -> quota failure -> turn blocked
Sol quota reset -> model-switch workflow works again

Screenshots and supporting files

Relevant upstream Codex behavior/fixes:

Relevant OpenCodex compact routing implementation:

Redacted configuration

{
  "note": "Credentials, account identifiers, unrelated providers, and secrets removed",
  "routedModel": "deepseek/deepseek-v4-flash",
  "integration": "Codex App loopback proxy"
}

Checks

  • I searched existing issues and documentation.
  • I removed secrets, tokens, account details, request credentials, and personal data.

Activity

github-actions commented on Aug 27, 2026

@github-actions
Contributor

Issue reopened

The report now contains the information required by the automated check. Thanks for updating it.

changed the title [-][Bug] Remote auto-compaction ignores routed thread model and falls back to gpt-5.6-sol, causing quota-bound 502s[/-] [+][Bug]: Quota-limited previous-model compact blocks Sol → routed DeepSeek handoff after manual /compact succeeds[/+] on Aug 27, 2026
added
bugSomething isn't working
proxyHTTP proxy, routing, reverse-proxy / management auth
on Aug 27, 2026

lidge-jun commented on Aug 27, 2026

@lidge-jun
Owner

리뷰 · 우선순위 64 / 80

이 버그는 Codex 앱이 같은 스레드에서 이전 모델(gpt-5.6-sol) → 라우트된 DeepSeek로 바꾼 뒤, 수동 /compact는 DeepSeek로 성공하는데, 다음 일반 턴의 "이전 모델 사전 압축"이 Sol 쿼터 한도에 걸려 턴이 멈추는 이야기입니다. 보고된 사용 기록은 provider=openai, model=gpt-5.6-sol, status=502, code=upstream_server_error, 메시지 The usage limit has been reached입니다. 이슈도 말하듯, 이전 모델로 먼저 compact를 시도하는 것 자체는 업스트림 Codex 동작입니다. 문제는 OpenCodex를 통과한 뒤 그 실패가 Codex의 "선택/현재 모델 compact 폴백"을 깨우는 모양으로 보존되지 않거나, 프록시가 대신 현재 라우트 모델로 넘겨 주지 않아 DeepSeek 경로가 한 번도 안 타는 점입니다.

src/server/responses/compact.ts - 네이티브 compact는 요청된 모델 라우트로 그대로 보냅니다. 이전 모델 Sol이면 Sol 계정/쿼터로 갑니다. 풀 계정 대체는 주로 429/402에서만 같은 요청 안 재시도를 합니다. "usage limit" 문구가 다른 상태 코드로 오면 그 경로를 못 탈 수 있습니다.
src/lib/errors.ts - classifyError는 quota exhausted / insufficient_quota 같은 문구는 쿼터로 분류하지만, 보고된 문장 The usage limit has been reached를 전용 분기로 두지 않습니다. 상태 코드가 5xx면 upstream_server_error로 떨어지기 쉽습니다. 사용 기록의 502 / upstream_server_error와 맞습니다.
업스트림 Codex PR #30319 / #32881 - usage-limit·서버·재시도 소진을 selected-model compact 폴백 조건으로 넓혔습니다. 프록시가 쿼터 실패를 일반 502 서버 오류처럼 보이면, Codex 쪽 폴백이 안 켜질 수 있습니다.
수동 /compact 성공 - 현재 라우트(DeepSeek) compact 자체는 살아 있다는 뜻입니다. 막히는 지점은 "이전 모델 사전 압축 실패 이후의 핸드오프"입니다.

메인테이너의 판단이 필요한 지점

  • (1) compact 응답에서 usage-limit을 Codex가 알아보는 코드/상태로 보존·번역할지, (2) OpenCodex가 previous→current 라우트 compact 폴백을 직접 할지, (3) 안전하게 못 하면 문서/제한으로만 알릴지. 이슈가 고른 세 선택지가 그대로 설계 분기입니다.
  • isRateLimitOrQuotaFailureMessage는 "usage limit" 문구를 건강/차단용으로 이미 봅니다. 클라이언트에 돌려주는 classify/status와 이 헬퍼가 어긋나 있는지 먼저 확인하는 편이 싸습니다.
  • preview 배포는 계획에 없습니다. 수정은 dev에만 올리면 됩니다.

너의 추천
닫지 말고 needs-design 성격으로 유지하세요. 첫 조사는 compact 경로에서 The usage limit has been reached가 어떤 HTTP 상태·error.code로 클라이언트에 나가는지 로컬로 한 번 찍는 것입니다. 그 결과가 Codex 폴백 조건과 다르면, 쿼터 분류를 compact/passthrough에 맞추는 작은 PR이 (1)안이고, 그보다 크면 (2)안 설계 이슈로 쪼개면 됩니다.

이 댓글은 grok-bot이 작성했습니다

lidge-jun commented on Aug 28, 2026

@lidge-jun
Owner

Fixed on dev by commit 676a3c0 (PR #2858).

The cause was narrower than "compact is broken": compaction routed the client-supplied previous model — the quota-limited Sol — directly, and upstream Codex's own current-model fallback is not reachable through the custom-provider path. So once Sol was quota-exhausted, automatic compaction kept aiming at it and the handoff never got a chance, even though a manual /compact worked.

Worth noting for anyone tracking this: commit a52d00a6e from earlier in this campaign ("skip failing quota candidates") does not cover this. That one fixes quota-based account candidate selection; it does not touch compact model or provider routing. Two different code paths.

The fix remembers the last successful compact model per bounded session lane, and on a body-confirmed quota failure retries that same-thread handoff target. Separate threads stay isolated, and explicit 402/429 attribution is preserved so a genuine payment or rate error still surfaces as itself rather than being masked by a retry.

Verified in tests/responses-compaction-routing.test.ts driven RED (expected 200, got 502 — your exact symptom) then GREEN at 45 pass, with 204 further passes across the routing, quota, and error focused files. tsc and privacy:scan clean.

This PR targeted dev, so GitHub did not auto-close the issue — closing manually.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    account-poolOAuth, credentials, Codex pool, quota, failover, plansbugSomething isn't workingproxyHTTP proxy, routing, reverse-proxy / management auth

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions