Skip to content

refactor(ai): one model-profile table per provider for Anthropic and OpenAI - #1013

Open
sroussey wants to merge 2 commits into
mainfrom
claude/determined-franklin-7ikqd3
Open

sroussey wants to merge 2 commits into
mainfrom
claude/determined-franklin-7ikqd3

Conversation

@sroussey

@sroussey sroussey commented Oct 5, 2026

Copy link
Copy Markdown
Collaborator

Summary

Finishes the part of #1006 left open by #1011: per-model capability decisions now live in one profile table per provider, read by request building, capability inference and (OpenAI) model search, instead of being re-derived from the id in each place.

  • Anthropic: Anthropic_ModelProfiles.ts holds {supportsOutputFormat, acceptsForcedToolChoice, vision} per canonical id, with resolveAnthropicProfile and one defaultAnthropicProfile (unreadable id → legacy profile; generation 5+ → newest-known profile; older → legacy). New Anthropic_ModelId.ts holds the id parser and normalizeAnthropicModelId, which maps anthropic/ (OpenRouter), Bedrock ([us.]anthropic.…-v1:0) and Vertex (@date/@latest) spellings and trailing dates to the canonical id. The resolve… functions still return {value, source} and still honour the provider_config overrides from Remove surface gate; fix Anthropic/OpenAI capability, AgentTask retry/budget, FetchUrl error leak #1011.
  • OpenAI: OpenAI_ModelProfiles.ts holds {kind, vision, acceptsTemperatureWithReasoning} per id with normalizeOpenAiModelId and one defaultOpenAiProfile. The temperature rule, inferOpenAiCapabilities and model-search image capabilities read it; the OPENAI_IMAGE_MODELS list is gone.
  • Contract test: profile ids equal the pricing ids plus an explicit list of unpriced rows (ANTHROPIC_UNPRICED_PROFILE_IDS: opus-4-1, opus-4-5, mythos-5, mythos-5-1; OPENAI_UNPRICED_PROFILE_IDS: gpt-5.6) whose flags differ from the default; every row resolves every flag explicitly; per-id agreement between flags, inference and profile; normaliser tables. No prices were invented.

Behaviour changes for ids no existing test covers

  • Anthropic: unlisted 5.x minors (e.g. opus-5-3) now take the newest profile (no forced tool choice); previously the minor-below-5 rule answered true. This is the safer direction.
  • OpenAI: an unlisted gpt-5.6-* no longer keeps a pinned temperature (the ^gpt-5\.6 regex is gone), so a new 5.6 model needs a profile row. source for unlisted OpenAI ids is "default" where it was "table".
  • OpenAI model search gives any gpt-image-* / dall-e* id its image capabilities, not only the four hard-coded ids.
  • Effort policy and parsedModelName now go through the normaliser, so OpenRouter and Vertex spellings resolve there too.

Not done

Sampling params (anthropicAcceptsSamplingParams) still parse the raw id and ignore gateway spellings; Anthropic model search still stamps capabilities: []; pricing lookup is untouched; other providers are not covered.

Closes #1006

Verification

ai-provider + ai-provider-api sections: 135 files / 1698 tests pass; oxlint and oxfmt clean on the changed files; bun run build:types 47/47.

🤖 Generated with Claude Code

https://claude.ai/code/session_01Wm2G7fm2T9HzFWoj66EH7B


Generated by Claude Code

…OpenAI

Per-model capability flags, task capabilities and id normalisation now live
in a profile table each provider's request building, capability inference and
model search read. Unknown ids take one documented default profile.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Wm2G7fm2T9HzFWoj66EH7B

@mergestorm-vortex mergestorm-vortex Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review progress ██████████ 14/14 files

Comment — found 2 issue(s) at e6dcea8.

Actionable comments posted: 2

🤖 Prompt for AI agents
Verify each finding against current code. Fix only still-valid concrete bugs, skip the
rest with a brief reason, keep changes minimal, and validate. Skip Decision required,
policy forks, and "consider X" alternatives. Do not add new features, refactors, or
architecture beyond the fix; prefer the smallest diff.

Findings to address:
1. In `@packages/test/src/test/ai-provider/ProviderCapabilityTables.test.ts` (line 172, Important):
   `resolveAnthropicOutputFormatSupport(claude(id)).value` and `resolveAnthropicForcedToolChoice(claude(id)).value` are compared against `resolveAnthropicProfile(id).profile.{supportsOutputFormat,acceptsForcedToolChoice}`. Both resolvers are defined as `resolveAnthropicProfile(modelNameOf(model))` with no override in `claude(id)` (only `model_name` is passed), so each comparison is literally `resolveAnthropicProfile(id).profile.X === resolveAnthropicProfile(id).profile.X` — it is an identity and can never fail, regardless of whether request building still reads the table. The same holds for `resolveOpenAiTemperatureWithReasoning(gpt(id)).value` vs `profile!.acceptsTemperatureWithReasoning` at lines 189-191 (once `source === "table"` guarantees `profile` is defined). The meaningful checks in these cases are the `source === "table"` / capability-array assertions; the flag-agreement lines give false confidence that request building and inference share the table when they only re-verify the table against itself.

2. In `@providers/anthropic/src/ai/common/Anthropic_ModelId.ts` (line 103, Important):
   `normalizeAnthropicModelId` now centralizes gateway-spelling handling, but `anthropicSupportsAdaptiveThinking` (`Anthropic_Thinking.ts:63-69`) still does `parseAnthropicModelId(id)` on the raw `provider_config.model_name`. For a gateway spelling such as `us.anthropic.claude-opus-5-5-v1:0` (or the OpenRouter/Vertex forms this normalizer handles), `parseAnthropicModelId` returns `undefined` because the string does not start with `claude-`, so it answers `false`. Downstream, `buildAnthropicThinkingParams` (line 119) then takes the legacy branch and sends `thinking: { type: "enabled", budget_tokens }` instead of `{ type: "adaptive" }` + `output_config.effort` for a Claude 5.x model whenever an `effort` is set, contradicting the PR's claim that OpenRouter/Vertex spellings resolve through the normalizer. Route this call through `normalizeAnthropicModelId` (as `parsedModelName`/`anthropicEffortPolicy` now do) so the two thinking paths agree.

Try MergeStorm

(id) => {
const { profile, source } = resolveAnthropicProfile(id);
expect(source).toBe("table");
expect(resolveAnthropicOutputFormatSupport(claude(id)).value).toBe(

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue · Important

Self-referential "agrees" assertions cannot fail

resolveAnthropicOutputFormatSupport(claude(id)).value and resolveAnthropicForcedToolChoice(claude(id)).value are compared against resolveAnthropicProfile(id).profile.{supportsOutputFormat,acceptsForcedToolChoice}. Both resolvers are defined as resolveAnthropicProfile(modelNameOf(model)) with no override in claude(id) (only model_name is passed), so each comparison is literally resolveAnthropicProfile(id).profile.X === resolveAnthropicProfile(id).profile.X — it is an identity and can never fail, regardless of whether request building still reads the table. The same holds for resolveOpenAiTemperatureWithReasoning(gpt(id)).value vs profile!.acceptsTemperatureWithReasoning at lines 189-191 (once source === "table" guarantees profile is defined). The meaningful checks in these cases are the source === "table" / capability-array assertions; the flag-agreement lines give false confidence that request building and inference share the table when they only re-verify the table against itself.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid concrete bugs, skip the
rest with a brief reason, keep changes minimal, and validate. Skip Decision required,
policy forks, and "consider X" alternatives. Do not add new features, refactors, or
architecture beyond the fix; prefer the smallest diff.

In `@packages/test/src/test/ai-provider/ProviderCapabilityTables.test.ts` (line 172, Important):
`resolveAnthropicOutputFormatSupport(claude(id)).value` and `resolveAnthropicForcedToolChoice(claude(id)).value` are compared against `resolveAnthropicProfile(id).profile.{supportsOutputFormat,acceptsForcedToolChoice}`. Both resolvers are defined as `resolveAnthropicProfile(modelNameOf(model))` with no override in `claude(id)` (only `model_name` is passed), so each comparison is literally `resolveAnthropicProfile(id).profile.X === resolveAnthropicProfile(id).profile.X` — it is an identity and can never fail, regardless of whether request building still reads the table. The same holds for `resolveOpenAiTemperatureWithReasoning(gpt(id)).value` vs `profile!.acceptsTemperatureWithReasoning` at lines 189-191 (once `source === "table"` guarantees `profile` is defined). The meaningful checks in these cases are the `source === "table"` / capability-array assertions; the flag-agreement lines give false confidence that request building and inference share the table when they only re-verify the table against itself.

.replace(/@(?:\d{8}|latest)$/, "")
.replace(/-v\d+(?::\d+)?$/, "")
.replace(/-\d{8}$/, "");
if (!/^claude-(?:instant-)?\d+\.\d+$/.test(out)) out = out.replace(/(\d)\.(\d)/g, "$1-$2");

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue · Important

adaptive-thinking decision still parses the raw id, bypassing this normalizer

normalizeAnthropicModelId now centralizes gateway-spelling handling, but anthropicSupportsAdaptiveThinking (Anthropic_Thinking.ts:63-69) still does parseAnthropicModelId(id) on the raw provider_config.model_name. For a gateway spelling such as us.anthropic.claude-opus-5-5-v1:0 (or the OpenRouter/Vertex forms this normalizer handles), parseAnthropicModelId returns undefined because the string does not start with claude-, so it answers false. Downstream, buildAnthropicThinkingParams (line 119) then takes the legacy branch and sends thinking: { type: "enabled", budget_tokens } instead of { type: "adaptive" } + output_config.effort for a Claude 5.x model whenever an effort is set, contradicting the PR's claim that OpenRouter/Vertex spellings resolve through the normalizer. Route this call through normalizeAnthropicModelId (as parsedModelName/anthropicEffortPolicy now do) so the two thinking paths agree.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid concrete bugs, skip the
rest with a brief reason, keep changes minimal, and validate. Skip Decision required,
policy forks, and "consider X" alternatives. Do not add new features, refactors, or
architecture beyond the fix; prefer the smallest diff.

In `@providers/anthropic/src/ai/common/Anthropic_ModelId.ts` (line 103, Important):
`normalizeAnthropicModelId` now centralizes gateway-spelling handling, but `anthropicSupportsAdaptiveThinking` (`Anthropic_Thinking.ts:63-69`) still does `parseAnthropicModelId(id)` on the raw `provider_config.model_name`. For a gateway spelling such as `us.anthropic.claude-opus-5-5-v1:0` (or the OpenRouter/Vertex forms this normalizer handles), `parseAnthropicModelId` returns `undefined` because the string does not start with `claude-`, so it answers `false`. Downstream, `buildAnthropicThinkingParams` (line 119) then takes the legacy branch and sends `thinking: { type: "enabled", budget_tokens }` instead of `{ type: "adaptive" }` + `output_config.effort` for a Claude 5.x model whenever an `effort` is set, contradicting the PR's claim that OpenRouter/Vertex spellings resolve through the normalizer. Route this call through `normalizeAnthropicModelId` (as `parsedModelName`/`anthropicEffortPolicy` now do) so the two thinking paths agree.

…ser; make the table-agreement test non-tautological

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Wm2G7fm2T9HzFWoj66EH7B

@mergestorm-vortex mergestorm-vortex Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review progress ██████████ 15/15 files

Approve — reviewed files look good at a3c1f5a.

✅ All clear — nothing to fix.


Try MergeStorm

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

2 participants