Repository navigation
docs: document usage.completion_tokens_details.reasoning_tokens - #56
Conversation
The reasoning token count now also appears in usage.completion_tokens_details.reasoning_tokens, the location OpenAI uses and OpenAI-compatible SDKs read. Update the reasoning-models guide and the quickstart example, and mark the top-level usage.reasoning_tokens as a deprecated alias.
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 08f724bd1e
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
|
|
||
| - `reasoning_content` — the model's internal thinking (streamed as `delta.reasoning_content` chunks when `stream: true`) | ||
| - `usage.reasoning_tokens` — how many completion tokens were spent thinking | ||
| - `usage.completion_tokens_details.reasoning_tokens` — how many of the completion tokens were spent thinking. It is present in non-streaming responses and in the final usage chunk of a stream when you set `stream_options: {"include_usage": true}`. The top-level `usage.reasoning_tokens` carries the same number and is deprecated. |
There was a problem hiding this comment.
Update Best Practices to use the canonical token field
This declares the top-level usage.reasoning_tokens alias deprecated, but the Best Practices checklist at line 215 still tells readers to monitor that alias. Readers following the checklist will therefore build new monitoring against the field this change is trying to retire; update that item to recommend usage.completion_tokens_details.reasoning_tokens as well.
Useful? React with 👍 / 👎.
Why
OpenAI-compatible clients (the Vercel AI SDK `@ai-sdk/openai-compatible`, OpenCode, OpenRouter tooling) read the reasoning token count from `usage.completion_tokens_details.reasoning_tokens`. Our responses only carried SGLang's top-level `usage.reasoning_tokens`, and streams carried no count at all, so those clients showed zero reasoning tokens for GLM 5.3 Flash. The gateway and proxy are being changed to emit the standard field on every path while keeping the top-level field as a deprecated alias.
What
Merge after
Both code changes are merged but not yet in production. Merge this PR once both are live:
Follow-up outside this PR: the quickstart still uses the retired `zai-org/GLM-5.1-FP8` id in its examples.