Skip to content

feat(providers): GPT-6 family and Claude Fable 5.1; fix prompt_cache_key length and Anthropic keep-alives - #455

Merged
luizParreira merged 3 commits into
mainfrom
feat/newer-model-families
Sep 28, 2026
Merged

luizParreira merged 3 commits into
mainfrom
feat/newer-model-families

Conversation

@luizParreira

Copy link
Copy Markdown
Member

Summary

Three changes found while evaluating the newest models in Satoshi against real production tasks.

1. GPT-6 family and Claude Fable 5.1 (ebc09f4)

  • Adds gpt-6-astra, gpt-6-sol, gpt-6-luna to the capability and feature registries (1.05M context, 922K input, 128K output), with constructors on OpenAIProvider and OpenAIResponsesProvider.
  • GPT-6 shares the GPT-5.6 routing: agentic requests on the official base URL go to the Responses API. GPT-6 Sol and Luna accept function tools on Chat Completions only with reasoning_effort: none; Astra doesn't accept them there at all.
  • is_gpt56_model becomes is_gpt56_or_later_model. Only Astra gets detail: original images, per OpenAI's vision guide.
  • Adds claude-fable-5-1 (1M context, 128K output, $10/$50, adaptive-only) to the Anthropic and Vertex adaptive-thinking lists.

2. Bound prompt_cache_key to 64 characters (d403445)

OpenAI rejects longer keys (string_above_max_length), and callers pass thread ids as the session id, so every GPT-6 request failed. Keys of 64 bytes or fewer pass through unchanged; longer ones become their SHA-256 hex digest, which is stable per session so caching still works. Applies to Chat Completions, Responses and Codex Responses.

3. Treat Anthropic ping keep-alives as stream liveness (e2ed16a)

While Claude 5 models think at high effort with thinking.display: "omitted", the stream carries only ping events for minutes. parse_sse_event dropped them, so the agent loop's 2-minute inactivity watchdog, and agent-server's 120s stall budget in root_turn, killed healthy streams and retried them until the turn failed. In one eval run this happened 51 times in 77 Opus 5.5 cases; with the fix, zero.

Adds StreamDelta::KeepAlive (the enum is #[non_exhaustive]), emitted for ping. It resets the inactivity window and nothing else: no delta_count, chunk timing, accumulation, events, journaling or recording, and it doesn't pin the fallback wrapper to a provider. The timeout constants are unchanged, so a genuinely stalled connection still times out.

Testing

  • cargo fmt --all --check
  • cargo clippy --workspace --all-targets --all-features -- -D warnings
  • cargo test --all-features for agent-sdk-providers, agent-sdk and agent-server
  • RUSTDOCFLAGS=-D warnings cargo doc -p agent-sdk-providers --all-features --no-deps
  • The CI feature matrix with RUSTFLAGS=-D warnings
  • Live: Satoshi's eval harness on GPT-6 Sol/Luna (Responses API with tools) and Opus 5.5 / Sonnet 5 at max effort; zero inactivity kills after the fix.

Known follow-up

Structured output on claude-fable-5-1 will fail: the Anthropic provider forces a respond tool, and Fable 5.1 rejects forced tool_choice with a 400.

Add claude-fable-5-1 (1M context, 128K output, $10/$50, cache reads
$0.25/M, adaptive-only) with a fable_51 constructor, and list it as an
adaptive-thinking model for the Anthropic and Vertex providers so budget
thinking fails fast instead of returning a 400.

Add gpt-6-astra, gpt-6-sol and gpt-6-luna (1.05M context, 922K input,
128K output) to the capability and feature registries, with constructors
on OpenAIProvider and OpenAIResponsesProvider. GPT-6 keeps the GPT-5.6
reasoning mode, context, summary and prompt_cache_options surface, so the
GPT-5.6 gate becomes is_gpt56_or_later_model: GPT-6 requests route to
Responses on the official base URL and accept exact cache controls.
Astra rejects `none` effort and cannot call tools on Chat Completions;
Sol and Luna call tools there only with effort `none`.

Only gpt-6-astra joins the `detail: original` image list, because the
vision guide does not list Sol or Luna.
OpenAI rejects a prompt_cache_key longer than 64 characters with
string_above_max_length. The Chat Completions, Responses and Codex
providers sent request.session_id as the key unchanged, and callers use
thread ids such as eval-bipa-premium-data-query-001-<uuid> (69 chars),
so every request with such an id failed.

Route all three sites through one helper, bounded_prompt_cache_key. It
passes ids of up to 64 bytes through unchanged and maps longer ones to
their SHA-256 hex digest (exactly 64 chars, via the foundation crate's
sha256_hex). The digest is the same on every turn, so cache routing
still groups one session's requests.
Opus 5.5 and Sonnet 5 at max effort think with display=omitted and can
stream no content delta for over two minutes. Anthropic keeps the
connection alive with `ping` events, but parse_sse_event dropped them,
so the agent loop's 2-minute LLM_STREAM_INACTIVITY_TIMEOUT saw silence
and killed healthy streams: 51 kills in one 77-case Opus eval run, each
retry dying the same way until connectivity retries ran out.

Add StreamDelta::KeepAlive and emit it for `ping`. Vertex-Claude shares
the parser, so it gets the same fix. A keep-alive restarts the
inactivity window because stream.next() yields, and is otherwise
ignored: the agent loop and the agent-server root-turn loop skip it
before delta_count, chunk timing, the accumulator, usage beacons, and
events. Record/replay forwards it but never writes it to a cassette,
and the fallback wrapper does not treat it as committing output, so a
provider that only pinged can still fail over. The timeout constant is
unchanged: a stream that stops pinging still times out.
@luizParreira
luizParreira merged commit 889da23 into main Sep 28, 2026
40 checks passed
@luizParreira
luizParreira deleted the feat/newer-model-families branch September 28, 2026 14:50
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant