Client or integration
OMO / senpi (openai-completions provider pointed at the ocx loopback /v1/chat/completions)
Area
Anthropic adapter (adaptive thinking request shape)
Summary
For Claude Opus 5.x adaptive-thinking models the Anthropic adapter sends thinking: { type: "adaptive" } without display. On Opus 4.7+ the API default is display: "omitted", so the upstream stream contains empty, signature-only thinking blocks and no thinking text. At effort: max on a large context (~220k tokens) Opus routinely thinks for 2-5 minutes before the first text or tool call. During that time a Chat Completions client gets only the role chunk and : opencodex heartbeat SSE comments (#5806), with no semantic event.
Clients whose stream-start or idle watchdog counts semantic events, not transport bytes, therefore see a 5-minute silence and cancel. In OMO this surfaces as Provider stream start timed out after 300000ms. #5806's ADR says this outcome is intentional ("a client's separate semantic-progress watchdog may still expire"). The fix needs real reasoning deltas.
Requesting display: "summarized" would stream summarized thinking and relay it as reasoning deltas on the Chat surface. Long-thinking turns would then look alive to every client. It is also what the Anthropic SDK-based harnesses request by default.
Reproduction
- Messages surface, with an explicit
display in the request:
curl -sN http://127.0.0.1:10100/v1/messages -H 'content-type: application/json' -H 'anthropic-version: 2023-06-01' -H 'x-api-key: x' \
-d '{"model":"anthropic/claude-opus-5-5","max_tokens":4000,"stream":true,"thinking":{"type":"adaptive","display":"summarized"},"output_config":{"effort":"high"},"messages":[{"role":"user","content":"Think carefully: how many primes are below 200?"}]}'
The thinking block arrives as {"type":"thinking","thinking":"","signature":""} followed only by signature_delta. There is no thinking_delta text, even though the caller explicitly asked for summarized.
- Chat surface at
reasoning_effort: max: 53 s of : opencodex heartbeat every 2 s, then the answer. completion_tokens was 3397, reasoning_tokens was 0, and no reasoning delta was ever emitted.
Expected: summarized thinking text is streamed, either always or at least when the caller requests display: "summarized". The Chat bridge would then emit reasoning deltas during the thinking phase.
Version
opencodex 2.65.0 (includes #5806)
Operating system
Linux x64
Provider and model
anthropic (OAuth) / claude-opus-5-5, effort max
Logs or error output
usage.jsonl on the failing session. There were six 502s on one conversation, each upstreamError: "The connection was closed.". Four ended at ~300 s, when the client gave up, with failureStage: protocol-prelude and semanticBytes: 0. Successful turns in the same conversation took 120-295 s.
Checks
Client or integration
OMO / senpi (
openai-completionsprovider pointed at the ocx loopback/v1/chat/completions)Area
Anthropic adapter (adaptive thinking request shape)
Summary
For Claude Opus 5.x adaptive-thinking models the Anthropic adapter sends
thinking: { type: "adaptive" }withoutdisplay. On Opus 4.7+ the API default isdisplay: "omitted", so the upstream stream contains empty, signature-only thinking blocks and no thinking text. Ateffort: maxon a large context (~220k tokens) Opus routinely thinks for 2-5 minutes before the first text or tool call. During that time a Chat Completions client gets only the role chunk and: opencodex heartbeatSSE comments (#5806), with no semantic event.Clients whose stream-start or idle watchdog counts semantic events, not transport bytes, therefore see a 5-minute silence and cancel. In OMO this surfaces as
Provider stream start timed out after 300000ms. #5806's ADR says this outcome is intentional ("a client's separate semantic-progress watchdog may still expire"). The fix needs real reasoning deltas.Requesting
display: "summarized"would stream summarized thinking and relay it asreasoningdeltas on the Chat surface. Long-thinking turns would then look alive to every client. It is also what the Anthropic SDK-based harnesses request by default.Reproduction
displayin the request:{"type":"thinking","thinking":"","signature":""}followed only bysignature_delta. There is nothinking_deltatext, even though the caller explicitly asked forsummarized.reasoning_effort: max: 53 s of: opencodex heartbeatevery 2 s, then the answer.completion_tokenswas 3397,reasoning_tokenswas 0, and noreasoningdelta was ever emitted.Expected: summarized thinking text is streamed, either always or at least when the caller requests
display: "summarized". The Chat bridge would then emit reasoning deltas during the thinking phase.Version
opencodex 2.65.0 (includes #5806)
Operating system
Linux x64
Provider and model
anthropic (OAuth) / claude-opus-5-5, effort max
Logs or error output
usage.jsonl on the failing session. There were six 502s on one conversation, each
upstreamError: "The connection was closed.". Four ended at ~300 s, when the client gave up, withfailureStage: protocol-preludeandsemanticBytes: 0. Successful turns in the same conversation took 120-295 s.Checks