Skip to content

[Bug]: Shadow Call Intercept forces effort low on every gpt-5.6-luna request when luna is the target (or main model) — max turns silently downgraded #2706

Description

@keohanoi

Client or integration

Codex CLI

Area

Proxy and routing

Summary

gpt-5.6-luna selected with [max] reasoning returns near-instant, shallow responses through the proxy, while the same model/effort directly against OpenAI thinks for a long time. Root cause: Shadow Call Intercept was enabled with gpt-5.6-luna itself as the target model, and the intercept's slug-only matching treats every luna request — main turns included — as a helper call, forcing reasoning.effort to low on the wire.

Mechanism (v2.33.0):

  1. shouldInterceptShadowCall (src/lib/shadow-call.ts) matches the model slug alone, unconditionally (deliberate since In Codex 0.147.0, background helper requests use gpt-5.6-luna. Even with OpenCodex's Shadow Call Intercept enabled, some Luna requests still reached OpenAI. #1684 — helper calls cannot be distinguished by headers). gpt-5.6-luna is both the helper slug and a legitimate main model.
  2. The rewrite in src/server/responses/core.ts (~2306–2330) replaces the model and forces effort: parsed.options.reasoning = "low" + _rawBody.reasoning = { effort: "low" }. When the configured target is luna (self-target), the model rewrite is a no-op but the effort clamp still applies to every request. Self-interception is pure loss.
  3. logCtx.requestedEffort is captured at ~2296, before the rewrite — so usage.jsonl / the Logs page show requestedEffort: "max" while the wire effort was low. Unlike applyEffortCap (which annotates from->to), the shadow rewrite leaves the log stale, which makes this very hard to diagnose from logs.

Evidence from usage.jsonl (incident window Aug 26 20:29–22:21 local, 258 consecutive requests): every entry has resolvedModel: "gpt-5.6-luna", provider: "openai", requestedEffort: "max", and shadowCallRewrittenFrom: "gpt-5.6-luna". reasoningOutputTokens on those max turns: 13–31 (median ~24). Earlier non-intercepted max turns on the same setup reach 5,000+ reasoning tokens and 30+ minute turns. The four hard turns the user actually noticed were all inside the window.

Four concrete defects:

  1. No self-target guard (runtime): shadowCallIntercept.model matching a shadow source model is accepted and silently degrades every request for that model to low. The management API (/api/shadow-call-settings) validates shape but not self-intersection.
  2. Slug-only matching hijacks luna main turns (design gap): with any target configured, a user whose main model is luna gets all main turns rerouted at low effort. [Bug]: Codex App now sends gpt-5.6-luna helper requests on every message and turn completion (previously only title generation) #2157 established Codex App emits luna helper calls on every turn, so operators are encouraged to enable the intercept — but nothing warns that this is incompatible with luna as a main model.
  3. GUI offers the source model as target: shadowCallModelOptions (gui/src/pages/dashboard-shared.ts:383) lists every active model, including openai/gpt-5.6-luna itself. That is exactly how this config was reached from the dashboard.
  4. Misleading logs: requestedEffort is recorded pre-rewrite (see above).

Suggested fixes: reject (or skip at runtime) a target that matches a shadow source model; exclude shadow source models from the target dropdown; annotate the effort rewrite in the request log (max->low) like applyEffortCap does; document/warn that enabling the intercept while luna is the main model downgrades main turns.

Reproduction

  1. ocx start with the built-in openai forward provider logged in (2.33.0).
  2. Dashboard → Models → enable Shadow Call Intercept; in the target dropdown pick openai/gpt-5.6-luna (it is offered).
  3. In Codex CLI, select gpt-5.6-luna [max] as the main model and send a genuinely hard task (one that normally thinks for minutes).
  4. Observe: responses return in seconds with shallow reasoning. Direct OpenAI with the same model/effort on the same task thinks for a long time.
  5. Inspect ~/.opencodex/usage.jsonl: every turn carries shadowCallRewrittenFrom: "gpt-5.6-luna" and requestedEffort: "max", with reasoningOutputTokens in the tens.

Variant of step 2 with any other target (e.g. zai/glm-5.3-flash) still reproduces the user-visible problem for luna-main users: every main turn is rerouted to the target model at low effort (per the unconditional matching introduced for #1684).

Version

2.33.0

Operating system

macOS 26.5.1

Provider and model

openai / gpt-5.6-luna

Logs or error output

# usage.jsonl entries during the incident window (aggregate; identifying fields removed)
# all 258 requests in the window, including the user's [max] main turns:
{
  "provider": "openai",
  "resolvedModel": "gpt-5.6-luna",
  "requestedModel": "gpt-5.6-luna",
  "shadowCallRewrittenFrom": "gpt-5.6-luna",
  "requestedEffort": "max",          # captured BEFORE the rewrite — misleading
  "status": 200,
  "durationMs": 9396,
  "usage": { "inputTokens": 194311, "outputTokens": 101,
             "reasoningOutputTokens": 13 }   # wire effort was "low"
}

# earlier non-intercepted [max] turns, same provider/model/setup:
#   reasoningOutputTokens up to 5083, durationMs up to 2203997 (~37 min)

Screenshots and supporting files

n/a

Redacted configuration

{
  "defaultProvider": "openai",
  "shadowCallIntercept": {
    "enabled": true,
    "model": "openai/gpt-5.6-luna"
  }
}

(State at incident time; intercept has since been disabled, after which max reasoning behaves normally.)

Checks

  • I searched existing issues and documentation.
  • I removed secrets, tokens, account details, request credentials, and personal data.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't workingcatalogModel catalog, slugs, visibility, routed entriesproxyHTTP proxy, routing, reverse-proxy / management auth

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions