feat(providers): GPT-6 family and Claude Fable 5.1; fix prompt_cache_key length and Anthropic keep-alives - #455
Merged
Merged
Conversation
Add claude-fable-5-1 (1M context, 128K output, $10/$50, cache reads $0.25/M, adaptive-only) with a fable_51 constructor, and list it as an adaptive-thinking model for the Anthropic and Vertex providers so budget thinking fails fast instead of returning a 400. Add gpt-6-astra, gpt-6-sol and gpt-6-luna (1.05M context, 922K input, 128K output) to the capability and feature registries, with constructors on OpenAIProvider and OpenAIResponsesProvider. GPT-6 keeps the GPT-5.6 reasoning mode, context, summary and prompt_cache_options surface, so the GPT-5.6 gate becomes is_gpt56_or_later_model: GPT-6 requests route to Responses on the official base URL and accept exact cache controls. Astra rejects `none` effort and cannot call tools on Chat Completions; Sol and Luna call tools there only with effort `none`. Only gpt-6-astra joins the `detail: original` image list, because the vision guide does not list Sol or Luna.
OpenAI rejects a prompt_cache_key longer than 64 characters with string_above_max_length. The Chat Completions, Responses and Codex providers sent request.session_id as the key unchanged, and callers use thread ids such as eval-bipa-premium-data-query-001-<uuid> (69 chars), so every request with such an id failed. Route all three sites through one helper, bounded_prompt_cache_key. It passes ids of up to 64 bytes through unchanged and maps longer ones to their SHA-256 hex digest (exactly 64 chars, via the foundation crate's sha256_hex). The digest is the same on every turn, so cache routing still groups one session's requests.
Opus 5.5 and Sonnet 5 at max effort think with display=omitted and can stream no content delta for over two minutes. Anthropic keeps the connection alive with `ping` events, but parse_sse_event dropped them, so the agent loop's 2-minute LLM_STREAM_INACTIVITY_TIMEOUT saw silence and killed healthy streams: 51 kills in one 77-case Opus eval run, each retry dying the same way until connectivity retries ran out. Add StreamDelta::KeepAlive and emit it for `ping`. Vertex-Claude shares the parser, so it gets the same fix. A keep-alive restarts the inactivity window because stream.next() yields, and is otherwise ignored: the agent loop and the agent-server root-turn loop skip it before delta_count, chunk timing, the accumulator, usage beacons, and events. Record/replay forwards it but never writes it to a cassette, and the fallback wrapper does not treat it as committing output, so a provider that only pinged can still fail over. The timeout constant is unchanged: a stream that stops pinging still times out.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Three changes found while evaluating the newest models in Satoshi against real production tasks.
1. GPT-6 family and Claude Fable 5.1 (
ebc09f4)gpt-6-astra,gpt-6-sol,gpt-6-lunato the capability and feature registries (1.05M context, 922K input, 128K output), with constructors onOpenAIProviderandOpenAIResponsesProvider.reasoning_effort: none; Astra doesn't accept them there at all.is_gpt56_modelbecomesis_gpt56_or_later_model. Only Astra getsdetail: originalimages, per OpenAI's vision guide.claude-fable-5-1(1M context, 128K output, $10/$50, adaptive-only) to the Anthropic and Vertex adaptive-thinking lists.2. Bound
prompt_cache_keyto 64 characters (d403445)OpenAI rejects longer keys (
string_above_max_length), and callers pass thread ids as the session id, so every GPT-6 request failed. Keys of 64 bytes or fewer pass through unchanged; longer ones become their SHA-256 hex digest, which is stable per session so caching still works. Applies to Chat Completions, Responses and Codex Responses.3. Treat Anthropic
pingkeep-alives as stream liveness (e2ed16a)While Claude 5 models think at high effort with
thinking.display: "omitted", the stream carries onlypingevents for minutes.parse_sse_eventdropped them, so the agent loop's 2-minute inactivity watchdog, andagent-server's 120s stall budget inroot_turn, killed healthy streams and retried them until the turn failed. In one eval run this happened 51 times in 77 Opus 5.5 cases; with the fix, zero.Adds
StreamDelta::KeepAlive(the enum is#[non_exhaustive]), emitted forping. It resets the inactivity window and nothing else: nodelta_count, chunk timing, accumulation, events, journaling or recording, and it doesn't pin the fallback wrapper to a provider. The timeout constants are unchanged, so a genuinely stalled connection still times out.Testing
cargo fmt --all --checkcargo clippy --workspace --all-targets --all-features -- -D warningscargo test --all-featuresforagent-sdk-providers,agent-sdkandagent-serverRUSTDOCFLAGS=-D warnings cargo doc -p agent-sdk-providers --all-features --no-depsRUSTFLAGS=-D warningsKnown follow-up
Structured output on
claude-fable-5-1will fail: the Anthropic provider forces arespondtool, and Fable 5.1 rejects forcedtool_choicewith a 400.