You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
[Bug]: Shadow Call Intercept forces effort low on every gpt-5.6-luna request when luna is the target (or main model) — max turns silently downgraded #2706
gpt-5.6-luna selected with [max] reasoning returns near-instant, shallow responses through the proxy, while the same model/effort directly against OpenAI thinks for a long time. Root cause: Shadow Call Intercept was enabled with gpt-5.6-luna itself as the target model, and the intercept's slug-only matching treats every luna request — main turns included — as a helper call, forcing reasoning.effort to low on the wire.
The rewrite in src/server/responses/core.ts (~2306–2330) replaces the model and forces effort: parsed.options.reasoning = "low" + _rawBody.reasoning = { effort: "low" }. When the configured target is luna (self-target), the model rewrite is a no-op but the effort clamp still applies to every request. Self-interception is pure loss.
logCtx.requestedEffort is captured at ~2296, before the rewrite — so usage.jsonl / the Logs page show requestedEffort: "max" while the wire effort was low. Unlike applyEffortCap (which annotates from->to), the shadow rewrite leaves the log stale, which makes this very hard to diagnose from logs.
Evidence from usage.jsonl (incident window Aug 26 20:29–22:21 local, 258 consecutive requests): every entry has resolvedModel: "gpt-5.6-luna", provider: "openai", requestedEffort: "max", and shadowCallRewrittenFrom: "gpt-5.6-luna". reasoningOutputTokens on those max turns: 13–31 (median ~24). Earlier non-intercepted max turns on the same setup reach 5,000+ reasoning tokens and 30+ minute turns. The four hard turns the user actually noticed were all inside the window.
Four concrete defects:
No self-target guard (runtime): shadowCallIntercept.model matching a shadow source model is accepted and silently degrades every request for that model to low. The management API (/api/shadow-call-settings) validates shape but not self-intersection.
GUI offers the source model as target: shadowCallModelOptions (gui/src/pages/dashboard-shared.ts:383) lists every active model, including openai/gpt-5.6-luna itself. That is exactly how this config was reached from the dashboard.
Misleading logs: requestedEffort is recorded pre-rewrite (see above).
Suggested fixes: reject (or skip at runtime) a target that matches a shadow source model; exclude shadow source models from the target dropdown; annotate the effort rewrite in the request log (max->low) like applyEffortCap does; document/warn that enabling the intercept while luna is the main model downgrades main turns.
Reproduction
ocx start with the built-in openai forward provider logged in (2.33.0).
Dashboard → Models → enable Shadow Call Intercept; in the target dropdown pick openai/gpt-5.6-luna (it is offered).
In Codex CLI, select gpt-5.6-luna [max] as the main model and send a genuinely hard task (one that normally thinks for minutes).
Observe: responses return in seconds with shallow reasoning. Direct OpenAI with the same model/effort on the same task thinks for a long time.
Inspect ~/.opencodex/usage.jsonl: every turn carries shadowCallRewrittenFrom: "gpt-5.6-luna" and requestedEffort: "max", with reasoningOutputTokens in the tens.
Variant of step 2 with any other target (e.g. zai/glm-5.3-flash) still reproduces the user-visible problem for luna-main users: every main turn is rerouted to the target model at low effort (per the unconditional matching introduced for #1684).
Version
2.33.0
Operating system
macOS 26.5.1
Provider and model
openai / gpt-5.6-luna
Logs or error output
# usage.jsonl entries during the incident window (aggregate; identifying fields removed)# all 258 requests in the window, including the user's [max] main turns:
{
"provider": "openai",
"resolvedModel": "gpt-5.6-luna",
"requestedModel": "gpt-5.6-luna",
"shadowCallRewrittenFrom": "gpt-5.6-luna",
"requestedEffort": "max", # captured BEFORE the rewrite — misleading"status": 200,
"durationMs": 9396,
"usage": { "inputTokens": 194311, "outputTokens": 101,
"reasoningOutputTokens": 13 } # wire effort was "low"
}
# earlier non-intercepted [max] turns, same provider/model/setup:# reasoningOutputTokens up to 5083, durationMs up to 2203997 (~37 min)
Client or integration
Codex CLI
Area
Proxy and routing
Summary
gpt-5.6-lunaselected with[max]reasoning returns near-instant, shallow responses through the proxy, while the same model/effort directly against OpenAI thinks for a long time. Root cause: Shadow Call Intercept was enabled withgpt-5.6-lunaitself as the target model, and the intercept's slug-only matching treats every luna request — main turns included — as a helper call, forcingreasoning.efforttolowon the wire.Mechanism (v2.33.0):
shouldInterceptShadowCall(src/lib/shadow-call.ts) matches the model slug alone, unconditionally (deliberate since In Codex 0.147.0, background helper requests use gpt-5.6-luna. Even with OpenCodex's Shadow Call Intercept enabled, some Luna requests still reached OpenAI. #1684 — helper calls cannot be distinguished by headers).gpt-5.6-lunais both the helper slug and a legitimate main model.src/server/responses/core.ts(~2306–2330) replaces the model and forces effort:parsed.options.reasoning = "low"+_rawBody.reasoning = { effort: "low" }. When the configured target is luna (self-target), the model rewrite is a no-op but the effort clamp still applies to every request. Self-interception is pure loss.logCtx.requestedEffortis captured at ~2296, before the rewrite — sousage.jsonl/ the Logs page showrequestedEffort: "max"while the wire effort waslow. UnlikeapplyEffortCap(which annotatesfrom->to), the shadow rewrite leaves the log stale, which makes this very hard to diagnose from logs.Evidence from
usage.jsonl(incident window Aug 26 20:29–22:21 local, 258 consecutive requests): every entry hasresolvedModel: "gpt-5.6-luna",provider: "openai",requestedEffort: "max", andshadowCallRewrittenFrom: "gpt-5.6-luna".reasoningOutputTokenson thosemaxturns: 13–31 (median ~24). Earlier non-interceptedmaxturns on the same setup reach 5,000+ reasoning tokens and 30+ minute turns. The four hard turns the user actually noticed were all inside the window.Four concrete defects:
shadowCallIntercept.modelmatching a shadow source model is accepted and silently degrades every request for that model tolow. The management API (/api/shadow-call-settings) validates shape but not self-intersection.loweffort. [Bug]: Codex App now sends gpt-5.6-luna helper requests on every message and turn completion (previously only title generation) #2157 established Codex App emits luna helper calls on every turn, so operators are encouraged to enable the intercept — but nothing warns that this is incompatible with luna as a main model.shadowCallModelOptions(gui/src/pages/dashboard-shared.ts:383) lists every active model, includingopenai/gpt-5.6-lunaitself. That is exactly how this config was reached from the dashboard.requestedEffortis recorded pre-rewrite (see above).Suggested fixes: reject (or skip at runtime) a target that matches a shadow source model; exclude shadow source models from the target dropdown; annotate the effort rewrite in the request log (
max->low) likeapplyEffortCapdoes; document/warn that enabling the intercept while luna is the main model downgrades main turns.Reproduction
ocx startwith the built-inopenaiforward provider logged in (2.33.0).openai/gpt-5.6-luna(it is offered).gpt-5.6-luna [max]as the main model and send a genuinely hard task (one that normally thinks for minutes).~/.opencodex/usage.jsonl: every turn carriesshadowCallRewrittenFrom: "gpt-5.6-luna"andrequestedEffort: "max", withreasoningOutputTokensin the tens.Variant of step 2 with any other target (e.g.
zai/glm-5.3-flash) still reproduces the user-visible problem for luna-main users: every main turn is rerouted to the target model atloweffort (per the unconditional matching introduced for #1684).Version
2.33.0
Operating system
macOS 26.5.1
Provider and model
openai / gpt-5.6-luna
Logs or error output
Screenshots and supporting files
n/a
Redacted configuration
{ "defaultProvider": "openai", "shadowCallIntercept": { "enabled": true, "model": "openai/gpt-5.6-luna" } }(State at incident time; intercept has since been disabled, after which
maxreasoning behaves normally.)Checks