Skip to content

problem with gemini #4856

Description

@heygoodluck

Client or integration

Claude Code

Area

CLI

Summary

제미나이 한도가 남아있는데 429 Provider error 429: Antigravity rate limit exceeded: Resource has been exhausted (e.g. check quota). 이 오류가 지속적으로 발생합니다

Reproduction

  1. ocx start
  2. send message to proxy

Version

2.57.0

Operating system

ubuntu 24

Provider and model

antigravity / google-antigravity/gemini-3.8-flash

Logs or error output

모델
gemini-3.8-flash
프로바이더
Google Antigravity
추론 강도
high
오류
rate_limit_exceeded
업스트림 원인
Provider error 429: Antigravity rate limit exceeded: Resource has been exhausted (e.g. check quota).

Screenshots and supporting files

No response

Redacted configuration

Checks

  • I searched existing issues and documentation.
  • I removed secrets, tokens, account details, request credentials, and personal data.

Activity

lidge-jun commented on Sep 17, 2026

@lidge-jun
Owner

리뷰 · 우선순위 45 / 80

이 이슈는 Claude Code에서 google-antigravity / gemini-3.8-flash로 보낼 때, 제미나이 한도가 남아 있는데도 429 / rate_limit_exceeded / Antigravity rate limit exceeded: Resource has been exhausted (e.g. check quota).가 반복된다는 짧은 보고입니다. 지금 dev HEAD는 19bdcaaf6(패키지 2.58.0)이고, 이 문구는 업스트림 Google Cloud Code Assist가 돌려준 본문을 src/adapters/google-errors.ts의 safeAntigravityHttpErrorMessage가 짧게 다시 쓴 결과와 맞습니다. 429이거나 status가 RESOURCE_EXHAUSTED이면, 본문에 “hard quota”로 보는 특정 needle이 없을 때 Antigravity rate limit exceeded로 분류합니다. "Resource has been exhausted (e.g. check quota)." 같은 일반 문구는 GOOGLE_QUOTA_EXHAUSTED_NEEDLES에 없어서, 코드가 “일시 rate limit” 버킷으로 넣는 것이 현재 설계입니다.

여기서 헷갈리기 쉬운 점이 두 가지입니다. 첫째, Antigravity OAuth 쿼타는 AI Studio/Gemini API 키 쿼타와 별개입니다. src/providers/quota/antigravity.ts는 retrieveUserQuotaSummary / fetchAvailableModels로 Gem과 Cla 창을 따로 만듭니다. 화면 어딘가에서 “제미나이 한도”가 남아 보여도, 그 숫자가 Antigravity 계정의 Gem 5h/weekly 창이 아니면 429와 동시에 성립할 수 있습니다. 둘째, 같은 Antigravity 안에서도 Claude 계열(Cla) 창이 비었어도 Gemini 모델은 Gem 창을 씁니다. 리포터가 본 UI가 어느 창인지, ocx 계정 풀에 계정이 여러 개인지, 429 뒤 재시도/계정 로테이션이 돌았는지가 본문에 없습니다. 예전 #1413·#1062·#2568 쪽이 Antigravity 멀티계정·429 failover를 다루었고, 지금 풀/페일오버 인프라는 있지만 “한도가 남았는데도 429”를 코드 버그로 단정할 증거는 이 이슈만으로는 부족합니다.

재현 절차도 ocx start → 메시지 한 줄뿐이라, 버전 2.57.0 / ubuntu 24 / 모델명 말고는 로그·쿼타 스냅샷·계정 수가 없습니다. 그래서 지금은 “어댑터가 429를 잘못 만든다”기보다 “업스트림이 RESOURCE_EXHAUSTED를 냈고, 클라이언트가 그걸 rate_limit로 보여 준다” 쪽이 기본 가설입니다. 코드 쪽을 의심할 여지는 분류 메시지가 “quota” 단어를 포함한 일반 exhausted 문구를 hard-quota가 아니라 rate-limit으로 보여 줘서, 사용자가 AI Studio 잔여 한도와 더 쉽게 충돌해 보이는 정도입니다. 그건 UX/분류 개선 후보이지, 이번 증상만으로 바로 고칠 한 줄은 아닙니다. types.ts/config.ts 분리와도 무관합니다.

경로 src/adapters/google-errors.ts (classifyGoogle / isGoogleQuotaExhaustedText) - bare Resource has been exhausted는 hard-quota needle이 아니라서 429+RESOURCE_EXHAUSTED → Antigravity rate limit exceeded로 표시된다. 리포트 문구와 일치한다.
경로 src/adapters/google-http.ts - 429 body가 hard-quota면 재시도하지 않고, 아니면 재시도 버킷이다. 본 이슈 메시지는 hard-quota needle이 없어 재시도 가능한 쪽으로 간다.
경로 src/providers/quota/antigravity.ts - Gem/Cla(및 Weekly) 창이 분리되어 있다. “제미나이 한도 남음”이 이 창을 가리키는지 확인이 필요하다.
경로 이슈 본문 - redacted config·계정 수·ocx 쿼타 표시 스크린·429 직전 Gem percent가 없어 재현/분류가 막힌다.
심볼 관련 이슈 - #1413(Antigravity pooling/429 failover, closed), #1080(Gem 쿼타 표시)과 같은 계열. 새 어댑터 버그로 열기 전에 계정/창 확인이 먼저다.

메인테이너의 판단이 필요한 지점

  • 이 이슈를 needs-info로 두고 리포터에게 Antigravity Gem 창 수치·계정 수·전체 upstream 본문을 요청할지
  • bare Resource has been exhausted를 hard-quota로 승격할지(잘못 승격하면 일시 429 재시도가 죽을 수 있음 — 현재 테스트도 transient 문구를 rate-limit으로 고정)
  • 단일 계정에서 Gem이 실제로 비었는데 UI만 다른 소스를 보여 주는 표시 버그인지, 업스트림 일시 한도인지

너의 추천
지금은 코드 수정 PR보다 정보 요청이 맞다. 리포터에게 (1) google-antigravity 로그인 계정 개수, (2) ocx/GUI에 보이는 Gem·Cla 퍼센트, (3) “제미나이 한도”를 본 화면이 Antigravity인지 AI Studio인지, (4) 가능하면 429 raw body를 부탁하자. 그 답이 Gem 창 고갈이면 문서/안내로 닫고, Gem이 여유인데도 반복이면 그때 google-http 재시도·계정 로테이션 경로를 좁혀 보면 된다. 우선순위는 사용자 체감은 있지만 재현 재료가 부족해서 중간 아래다.

이 댓글은 grok-bot이 작성했습니다

heygoodluck commented on Sep 17, 2026

@heygoodluck
Author
  1. 4개
  2. 대부분 0%
  3. Antigravity oauth

{
  "requestId": "ocx-",
  "timestamp": 1789614513354,
  "model": "gemini-3.8-flash",
  "provider": "google-antigravity",
  "surface": "claude",
  "admissionKind": "loopback",
  "inboundProtocol": "messages",
  "accountLogLabel": "oe85b7e",
  "conversationId": "",
  "requestedModel": "google-antigravity/gemini-3.8-flash",
  "requestedAlias": "google-antigravity/gemini-3.8-flash",
  "requestedEffort": "high",
  "tierOutcome": {
    "wireKind": null,
    "wireValue": null,
    "fastOutcome": "unknown",
    "confirmation": "unknown"
  },
  "status": 429,
  "durationMs": 1783,
  "errorCode": "rate_limit_exceeded",
  "closeReason": "non_stream",
  "upstreamError": "Provider error 429: Antigravity rate limit exceeded: Resource has been exhausted (e.g. check quota).",
  "usageStatus": "unreported",
  "attempts": [
    {
      "ordinal": 1,
      "provider": "google-antigravity",
      "model": "gemini-3.8-flash",
      "adapter": "google",
      "status": 429,
      "durationMs": 1776,
      "sendCount": 1,
      "recoveryKinds": [],
      "usageStatus": "unreported",
      "accountLogLabel": "",
      "requestedEffort": "high",
      "tierOutcome": {
        "wireKind": null,
        "wireValue": null,
        "fastOutcome": "unknown",
        "confirmation": "unknown"
      },
      "errorCode": "rate_limit_exceeded",
      "displayMetrics": {
        "tokPerSecond": {
          "kind": "unavailable",
          "reason": "usage_missing"
        },
        "decodeTokPerSecond": {
          "kind": "unavailable",
          "reason": "usage_missing"
        },
        "cost": {
          "kind": "unavailable",
          "reason": "usage_missing"
        }
      }
    }
  ],
  "spend": {
    "sends": 1,
    "settled": 1,
    "unresolved": 0
  },
  "routeDecision": {
    "version": 1,
    "decisionId": "",
    "createdAt": 1789614513360,
    "requestedModel": "google-antigravity/gemini-3.8-flash",
    "routeKind": "explicit-provider",
    "requirements": [],
    "candidates": [
      {
        "provider": "google-antigravity",
        "model": "gemini-3.8-flash",
        "eligible": true,
        "exclusions": []
      }
    ],
    "selected": {
      "candidateIndex": 0,
      "provider": "google-antigravity",
      "model": "gemini-3.8-flash",
      "reason": "explicit-provider-namespace"
    }
  },
  "displayMetrics": {
    "tokPerSecond": {
      "kind": "unavailable",
      "reason": "usage_missing"
    },
    "decodeTokPerSecond": {
      "kind": "unavailable",
      "reason": "usage_missing"
    },
    "cost": {
      "kind": "unavailable",
      "reason": "combo_attempt_unavailable"
    }
  }
}

heygoodluck commented on Sep 17, 2026

@heygoodluck
Author

Antigravity cli를 사용하면 문제 계정에서도 정상적인 사용이 가능합니다

heygoodluck commented on Sep 17, 2026

@heygoodluck
Author

Same active google-antigravity account and same gemini-3.8-flash model:

  1. OpenCodex /v1/chat/completions -> HTTP 200
  2. OpenCodex /v1/messages -> HTTP 200
  3. /v1/messages with stream=true, max_tokens=32000,
    metadata.user_id, Claude Code anthropic-beta headers,
    and output_config.effort=medium -> HTTP 200
  4. Codex surface using the same Antigravity model -> works
  5. Actual Claude Code surface -> reproducible HTTP 429 RESOURCE_EXHAUSTED

Ingwannu commented on Sep 17, 2026

@Ingwannu
Owner

The follow-up materially narrows this: the same provider/model succeeds through synthetic Chat, synthetic Messages (including streaming/Claude headers), the Codex surface, and Antigravity CLI, while the real Claude Code turn alone receives RESOURCE_EXHAUSTED. That makes this worth keeping open as a surface-specific integration bug rather than an ordinary exhausted-quota report.

The current log still cannot distinguish the two strongest candidates: (1) Claude Code carries a stable thread identity that is affined to a different/exhausted pool account even after the active account changes, or (2) its real system/tool/history envelope crosses an upstream request-specific resource limit that the minimal Messages control does not. The single attempt and empty attempt-level account label are also important; a transient 429 did not visibly rotate.

Please add one matched, redacted control pair from the same minute: real Claude Code and a successful minimal request, including accountLogLabel at request and attempt level, whether a client thread/conversation affinity key existed, approximate input/history byte or token count, tool count/schema bytes, image/file presence, requested max tokens/effort, and whether a fresh Claude Code thread behaves differently. Do not attach bodies, prompts, tokens, account IDs, cookies, or raw credentials.

I will check the Claude-surface account-affinity/429 recovery boundary once those matched fields are available. A fix should target the demonstrated difference; promoting the generic RESOURCE_EXHAUSTED phrase to hard quota now would be unsafe because the same account/model is demonstrably usable and could disable valid transient recovery.

lidge-jun commented on Sep 17, 2026

@lidge-jun
Owner

Separating this report by cause

This thread contains several distinct failures. They are listed separately below because a single
patch that makes all of them disappear would be hiding at least one of them, and because only one
of them currently has no owner.

Assessed against dev at 2f025814f3.

Cause Owner State on dev
A Cloud Code Assist rejects Claude Code's private billing metadata carried in systemInstruction #4913 not fixed
B A real Gemini or Claude window is spent, and headroomOf ranks on the worst window across all of them, so a spent Claude window disqualifies the account for Gemini #4676 per-account quota visibility landed in ef7b3c9cf4; family-aware ranking did not
C Provider save rejected before persistence when Clash/Mihomo resolves the canonical host into 198.18.0.0/15 #4723 not fixed
D Four eligible accounts, one physical send, and an empty attempt-level account label none generic 429 failover itself landed in 816f3a159d and 8bfac71466
E Some other envelope-specific upstream limit (tool schemas, history bytes, media, output ceiling) none unproven

Cause C is included because it shares the Antigravity surface, not because it is reported here.
Note that model-allowlist persistence goes through model-routes.ts, which runs no destination DNS
validation, so #4723 cannot explain a selected-model-only persistence failure.

Why the 429s here cannot be told apart from the wire

classifyGoogle decides between "quota exhausted" and "rate limit exceeded" by looking for
hard-quota phrases in the body. The message reported in this thread —
Resource has been exhausted (e.g. check quota) — contains none of the needles in
isGoogleQuotaExhaustedText, so it is classified as a retryable rate limit, exactly as a genuinely
transient limit would be.

That matters because this repository has already established that Cloud Code Assist answers
429 RESOURCE_EXHAUSTED for a policy rejection: ANTIGRAVITY_CLAUDE_SDK_PARAGRAPH_REJECTORS in
src/adapters/google.ts documents 3.7 and 3.8 Flash returning 429 with the Claude-Agent paragraph
present and 200 with it stripped, same account, seconds apart. Its comment names the hazard
directly: "a policy rejection wearing a quota error's clothing sends users hunting a quota problem
that does not exist".

So status, body, and headers cannot separate cause A from cause B. Only a matched envelope toggle
can: same account, same minute, one field changed, 429 becomes 200 and back. That is the standard
#4913 is being held to, and it is why #4913 should not be described as resolving this issue.

Cause D is the part with no owner, and it needs a trace before it needs code

The log in this thread shows a request-level account label, an empty attempt-level label, one
attempt, no recovery kinds, and an empty conversation id. Two very different failures produce that:

  1. rotation never happened, which is a defect in the recovery path; or
  2. rotation happened and the attribution was not recorded, which is a logging defect.

The fixes point in opposite directions, and adding another rotator on the second reading would stack
a second mechanism on top of one that already works. The empty conversation id also weakens the
sticky-thread-affinity explanation rather than supporting it.

What would settle it, from a single reproduction with OCX_DEBUG provider diagnostics on: the
eligible-account count at selection, any reauth flags, the send budget, the account actually chosen
for each send, and whether every attempted account received the same rejection. If four eligible
accounts produce one physical send, that is a recovery-path defect; if the accounts differ per
attempt and only the label is blank, it is attribution.

Suggested disposition

Keep this issue open for cause D, since A, B and C are tracked by the pull requests above and
would otherwise close this thread while D went unrecorded.

No local verification was run for this analysis; it is source reasoning against the commit named
above, and the per-cause state should be re-checked when those pull requests land.

lidge-jun commented on Sep 17, 2026

@lidge-jun
Owner

Update on the split posted above. Three of the five causes have landed on dev and will ship in the next release:

  • A — Cloud Code Assist rejecting Claude Code's private billing metadata: fix(google): strip x-anthropic-billing-header for Cloud Code Assist models #4913, squashed as a5aac78f892d2e87e36a28784e29ac9fd9731c38. The x-anthropic-billing-header paragraph is now stripped for Cloud Code Assist models, so the policy rejection that wore a quota error's clothing no longer happens.
  • B — a spent Claude window disqualifying the account for Gemini: feat(oauth): rank Antigravity failover by Gemini vs Claude quota family #4676, squashed as 513bb5656dad1332816cc54d37e9e7699a474948. Failover now ranks Antigravity custom windows by the requested family (Gemini vs Claude, with GPT-OSS in the Claude family) instead of on the worst window across all of them. An unknown model keeps all-window ranking, and absent matching evidence keeps the existing unranked behavior.
  • C — provider save rejected when Clash/Mihomo resolves the canonical host into 198.18.0.0/15: fix(providers): allow canonical Antigravity through fake-IP save validation #4723, squashed as 7e864362b18291a6145bd9c71690d84c94436cc6. This one was included for surface adjacency rather than because it explains your report; model-allowlist persistence goes through model-routes.ts, which runs no destination DNS validation.

D and E are still open, and this issue stays open for them. D is the one that matters for your report: four eligible accounts, one physical send, and an empty attempt-level account label. Generic 429 failover itself is on dev (816f3a159d, 8bfac71466), so the question is why this path took a single send instead of rotating. Nothing in the 429 status, body, or headers separates it from a genuine limit — classifyGoogle sees Resource has been exhausted (e.g. check quota), which contains none of the hard-quota needles, so it is classified as retryable exactly as a transient limit would be.

What would settle it: retest on the next release, and if the 429 persists, attach a trace from a single failing request showing the attempt count and which account each attempt used. With A and B removed from the picture, a remaining failure is much more likely to be D than a real quota, and the attempt-level account labels are what distinguish them.

heygoodluck commented on Sep 18, 2026

@heygoodluck
Author

Update after retesting:

  • I reduced the configured google-antigravity accounts to a single account, and the behavior is still the same.
  • I also updated OpenCodex to the latest available version and retested, but the result is unchanged.
  • The actual Claude Code surface still reproduces the same 429 RESOURCE_EXHAUSTED failure.

So this does not appear to depend on multi-account rotation/failover, and the issue still persists after the update.

heygoodluck commented on Sep 18, 2026

@heygoodluck
Author

Resolved for me with the latest dev build containing #4913.

The issue appears to have been caused by Claude Code injecting an internal line like:

x-anthropic-billing-header: cc_version=...

into the beginning of the system prompt. When that was forwarded to Google Antigravity / Cloud Code Assist, Gemini 3.8 Flash returned the misleading 429 RESOURCE_EXHAUSTED error even though quota was still available.

After using the fix from #4913, google-antigravity/gemini-3.8-flash works correctly from Claude Code again.

I also confirmed that 2.58.0 stable still reproduced the issue, which makes sense because #4913 was merged after the 2.58.0 release. Thanks!

lidge-jun commented on Sep 18, 2026

@lidge-jun
Owner

Cause D, assessed against dev at 11bc4f708c

Static source reading only; no local verification was run. Causes A, B and C are not revisited here.

Neither reading of cause D survives contact with the current source. The attribution is not
missing, and the single send is not a rotation defect — but the reason the two could not be told
apart is real, and it is the only part worth keeping.

The attribution half is not a defect

beginRequestAttempt creates an unlabeled provisional attempt (src/server/request-log.ts:1547),
which is why an empty attempt label looked like a plausible ordering bug. On the Antigravity path it
is not: the resolved OAuth account stamps the request label first
(src/server/responses/request-transport.ts:532), and the attempt is created and sealed with that
same label immediately after (:651). stampOAuthAccountLabel excludes only openai and
anthropic (src/providers/label.ts:59), so google-antigravity receives an o<hex6> label.

A literal accountLogLabel: "" is also not producible by the writer. Persistence omits the property
unless it validates (src/usage/log.ts:640), so the empty string in the pasted log is most likely an
absent property rendered as "" by whatever extracted it. Establishing that needs the raw
usage.jsonl row.

Rotation was present in the version you ran

816f3a159d and 8bfac71466 are both dated 2026-08-26 and are ancestors of v2.57.0
(44de45dfdc33, 2026-09-17). Reactive rotation activates on two eligible accounts
(hasFailoverAccountQuorum, src/oauth/generic-account-failover.ts:128) and allows three
alternates per request (GENERIC_OAUTH_MAX_FAILOVERS_PER_REQUEST, :40). With four accounts the
dispatch loop at src/server/responses/adapter-dispatch.ts:769 had every input it needed.

Why one physical send cannot be reproduced from current source

This is the part that changes the conclusion. Resource has been exhausted (e.g. check quota)
matches none of GOOGLE_QUOTA_EXHAUSTED_NEEDLES (src/adapters/google-errors.ts:21), and 429 is in
retryableGoogleStatus (:111), so isQuotaExhaustedBody is false and fetchGoogleWithRetry does
not terminalize on the first 429. It performs up to three same-account sends, each recording
rate-limit-429 (src/adapters/google-http.ts:90-108). Your log shows sendCount: 1 and
recoveryKinds: [].

Exactly one path in current source produces that shape: a send-budget denial.
createAdapterPhysicalSend throws SendBudgetExhaustedError when reserveDispatch refuses
(src/adapters/physical-send.ts:24), and fetchGoogleWithRetry catches it and returns the already
received 429 while pendingResponse is set (src/adapters/google-http.ts:112) — one send, no
recovery kind. The rotation loop then breaks on a refused hop without recording anything
(adapter-dispatch.ts:777-790).

But neither budget can do that here. The count policy allows four total sends with a three-send base
allowance (CODEX_TEXT_GUARDED_BUDGET_POLICY, src/lib/request-execution-budget.ts:43), and nothing
on this path spends it before the first inference send — OAuth refresh, loadCodeAssist project
discovery, cached quota ranking and routed count_tokens all run outside the ledger. The durable
token ceiling that could refuse is unreachable in a released build: sharedSpendLedger() is
constructed with no policy argument (src/lib/spend-reservation-ledger.ts:944),
DEFAULT_SPEND_RESERVATION_POLICY leaves every maxTokens undefined (:119), and
src/types/config.ts declares no spend key, so limitFor returns undefined for every production
scope and the refusal branch cannot be reached. (Same finding as the #4546 assessment.)

So the log implies a build that differs from the current source rather than a defect in it. Combined
with your single-account reproduction and your confirmation that #4913 resolved it, the 429 was the
deterministic policy rejection, and rotating four accounts would only have repeated it.

The one thing cause D found that is still true

When the budget refuses the rotation hop, the loop breaks and nothing records that it was
refused
. "Rotation never happened" and "rotation was withheld" are therefore indistinguishable in
the log — which is precisely why this cause could not be decided. The same is true of the
budget-denied adapter retry: it returns the original 429 with no recovery kind, so the attempt reads
as if no retry was ever attempted.

If you want this closed for good, the remaining work is observability — record the withheld failover
and the budget-denied retry — not another rotator stacked on one that works. I can take that as a
separate change if you want it.

Suggested disposition

Close cause D as not reproducible against current source and resolved in practice by #4913, and open
the withheld-failover attribution as its own item if it is worth tracking. What nobody here can
produce is the reporter-side evidence that would distinguish the remaining possibilities: the exact
installed package SHA, the raw 429 body, and the raw persisted log row.

Ingwannu commented on Sep 18, 2026

@Ingwannu
Owner

Closing as resolved by #4913, based on the reporter's confirmation against the fixed dev build. The misleading 429 came from Claude Code's private x-anthropic-billing-header paragraph being forwarded in systemInstruction, not from exhausted quota or account-pool rotation.

Stable 2.58.0 predates that fix, so it is expected to keep reproducing until the next release from dev. The separate withheld-retry observability gap identified in the final analysis should be tracked independently rather than keeping this resolved provider failure open.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't workingcliCLI, config inject, packaging flags

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions