Skip to content

fix: price Claude Kimi context aliases without crossing providers - #3259

Merged
steipete merged 3 commits into
mainfrom
codex/claude-kimi-context-pricing
Aug 29, 2026
Merged

steipete merged 3 commits into
mainfrom
codex/claude-kimi-context-pricing

Conversation

@steipete

@steipete steipete commented Aug 28, 2026 •

Copy link
Copy Markdown
Owner

Summary

Repair an omission in the existing Claude local-cost path: the documented Kimi Code k3[1m] alias does not reach the canonical kimi-for-coding/k3 price row. Without recognizing its vendor, a bare alias can also match an unrelated vendor's same-named row.

Append the canonical candidate only inside Claude's existing Kimi Code routes, after all exact candidates. Keep transcript model names, Codex lookup, other explicit routes, context variants, and paid Moonshot routing unchanged. Known catalog zero remains distinct from missing pricing.

After grouped unknown-model refreshes, recheck the shared catalog so a canonical price downloaded for a preceding group is used immediately. Preserve the existing retry/TTL guards and one rescan limit; actual provider routing still decides the price.

Preserve compatible native Codex SQLite rows, cursors, and retained reports when the generated parser hash changes. Remove the retired Pi scheduler-only compatibility exception: Pi's Claude costs use the changed pricing path and must be recalculated. Tests retain independent parser/formula/custom-price/catalog invalidation checks. Documentation and the 0.56.1 Unreleased changelog are updated.

Maintainer decision: accept the one-time Pi/OMP report recalculation. These are derived pricing caches, and retaining the predecessor key would preserve incorrect or missing estimates after the pricing change. Raw transcripts are not modified; unaffected native Codex cache state has a separate, tested preservation path.

During full validation, repair a test-runner failure in process inventory: use native macOS PID enumeration and Linux /proc instead of spawning ps on every ownership poll. Preserve required-PID inclusion and existing birth/session cleanup checks, reject incomplete inventories, and cover the observed failure plus native error/growth boundaries. This changes test infrastructure only, not app behavior or test assertions.

Related to #2374; thanks @joeVenner for the alias lead. This does not close or replace that PR's broader Kimi/Moonshot Pi-provider and CLI feature work, nor touch its contributor branch or the adjacent pricing-series branches.

Proof

  • Validated before-fix reproduction with the pricing file byte-identical to main: real temporary Claude JSONL retains the model and 160 tokens but leaves cost unknown. Exact-route controls work; bare alias/vendor collision and missing-canonical lookup regressions fail as expected.
  • New regressions cover scanner/snapshot output, exact-route precedence, conflicting catalogs, known zero versus absent/null prices, one injected catalog refresh, fresh warm/cold cache repricing, and existing cache-token/long-context arithmetic. Native Codex predecessor adoption and affected Pi Claude repricing have separate cache tests.
  • Isolated actual CLI before/after using the same synthetic transcript and persisted cache: shipped 0.56.0 reports 160 tokens with one unpriced row; the freshly built CLI reports the same k3[1m] row and 160 tokens with an expected synthetic estimate of 0.000385 USD and listPriceEstimate provenance. No force-refresh was needed. These are deliberately synthetic rates, not a claim about current Kimi prices.
  • CLI command: CodexBarCLI cost --provider claude --provider-native-only --format json --pretty. Both runs used an empty inherited environment, explicit temporary config/transcript/cache roots, disabled Keychain access, Codex credential-file isolation, and a sandbox denying network access and real-home data reads/writes. No accounts, credentials, browser cookies, or live provider APIs were used. The packaged app was not relaunched.
  • make check: passed, zero violations across 2,038 files; generated parser hash verified.
  • Codex autoreview at the configured P0 threshold: no actionable findings; independent source review also checked routing, cache boundaries, and fixture isolation.
  • Broader focused swift test --no-parallel --filter 'CostUsageClaudeKimiAliasTests|CostUsagePricingTests|CostUsageScannerClaudeMemoTests|CostUsageFetcherUnknownModelPricingTests|CostUsageStoreTests|PiSessionCostCompatibilityTests|PiSessionCostScannerTests|ProviderArchitectureGatekeeperTests': all 218 tests in eight suites passed (166.6 seconds). The initial parallel run exceeded five SQLite lock-test deadlines and exposed stale architecture line anchors plus a test-only refresh-interval mismatch; the corrected serial run passed every case without changing product lock behavior or test deadlines.
  • The qualified-route first-refresh finding is addressed in 44729e8736c830e4a5ae84e5611f4a7e3530c556: all three alias spellings now pass the first-call refresh regression with exactly one catalog request. Follow-up make check, blocking review, and repeated isolated CLI proof passed. The final focused rebuild passed all 112 tests in five suites (50.6 seconds), including the corrected architecture anchors. The latest source review of this head confirms no actionable findings remain.
  • App-fix CI passed every job on 44729e8736c830e4a5ae84e5611f4a7e3530c556. Fresh full CI including the harness repair passed every job on 1740f8aa71d32cfd3e87ed57248a93e7cadf4cf2: both macOS shards, Linux x64/arm64, musl, lint, and the aggregate gate.
  • The first local full run stopped on group watchdog limits and a synthetic Python CLI initialization timeout. The unchanged fallback suite then passed all 13 tests in 22 seconds, including strict hung-request assertions. A second run used the existing CODEXBAR_TEST_SUITE_TIMEOUT=600 outer watchdog and passed all first 47 groups without failures or retries, then stopped at group 48 when the runner's ps subprocess exceeded its separate two-second timeout.
  • Harness repair: both new no-ps regressions failed on the old implementation and passed after the repair. All 73 cleanup tests passed (one platform-specific skip), including real-process lifecycle and unrelated-process protection; follow-up make check passed with zero violations, and Codex autoreview plus independent source review found no actionable defect.
  • Local Swift coverage is complete: all 957 selections across 80 groups. The app and Swift test sources remain byte-identical to the 47-group checkpoint; a guarded continuation verified all 564 completed selection filters against fresh discovery, then passed the remaining 393 selections in 33 groups on 1740f8a (800.6 seconds, zero failures/retries/timeouts). This is verified checkpoint-plus-continuation coverage, not a claim that the interrupted make test invocation completed. No production timeout or test assertion was relaxed; the fresh exact-head CI matrix also passed in full.

The official guide establishes the alias contract; which client versions naturally persist that spelling in transcripts was not measured. No new provider support, paid-region inference, authentication change, or release publication is included.

@clawsweeper

clawsweeper Bot commented Aug 28, 2026

Copy link
Copy Markdown

🦞👀
ClawSweeper picked this up.

Pull request received. I will update this pull request when review starts.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 70a204242d

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

Comment thread Sources/CodexBarCore/Vendored/CostUsage/CostUsagePricing.swift
@clawsweeper clawsweeper Bot added P2 Normal priority bug or improvement with limited blast radius. rating: 🦐 gold shrimp Decent PR readiness signal, but merge confidence is limited. status: ⏳ waiting on author ClawSweeper has contributor-facing work open and is waiting for author action. labels Aug 29, 2026
@clawsweeper

clawsweeper Bot commented Aug 29, 2026 •

Copy link
Copy Markdown

Codex review: needs maintainer review before merge. Reviewed August 28, 2026, 10:13 PM ET / August 29, 2026, 02:13 UTC.

ClawSweeper review

What this changes

This PR prices Claude transcript k3[1m] aliases through Kimi Code’s scoped catalog entry, refreshes affected local pricing caches, and removes ps polling from the Swift test runner.

Merge readiness

⚠️ Ready for maintainer review - 3 items remain

Keep open for maintainer review: this owner-authored PR has no remaining actionable correctness finding, but it intentionally invalidates affected derived Pi/OMP pricing caches and the exact-head macOS test shards are still running.

Priority: P2
Reviewed head: 1740f8aa71d32cfd3e87ed57248a93e7cadf4cf2

Review scores

Measure Result What it means
Overall readiness 🐚 platinum hermit (4/6) The patch is well-scoped and heavily regression-tested, with remaining merge confidence depending on the in-progress macOS test shards and the documented cache-repricing effect.
Proof confidence 🌊 off-meta tidepool Not applicable: This owner-authored PR is exempt from the external-contributor proof gate; its supplied body nevertheless documents an isolated synthetic CLI before/after for the changed Claude pricing path.
Patch quality 🐚 platinum hermit (4/6) No actionable review findings were identified.

Verification

Check Result Evidence
Real behavior Not applicable Not applicable: This owner-authored PR is exempt from the external-contributor proof gate; its supplied body nevertheless documents an isolated synthetic CLI before/after for the changed Claude pricing path.
Evidence reviewed 7 items Introduced alias routing: The introduced Claude-only route appends kimi-for-coding/k3 only after an exact k3[1m] candidate is already scoped to Kimi Code; explicit routes and other providers remain governed by the existing lookup path.
Provider isolation remains enforced: Lookup returns the first scoped match for explicit or recognizable vendor routes, while unrecognized bare identifiers require exactly one matching provider; this supports the PR’s no-cross-provider claim.
Cache upgrade behavior is explicit: Pi reports reject a prior pricing key and rescan before saving the current key, while native Codex SQLite caches may adopt the listed compatible predecessor hash; focused tests cover both paths.
Findings None None.
Security None None.

How this fits together

Claude local-usage scans produce model-token rows, then provider-scoped pricing maps those rows to catalog prices and cached reports. The fetcher can refresh a missing catalog entry, while Pi/OMP session reports retain derived pricing aggregates separately from native Codex caches.

flowchart LR
  A[Claude and Pi transcripts] --> B[Local usage scanners]
  B --> C[Provider-scoped pricing routes]
  C --> D[Models.dev catalog]
  D --> E[Cost reports]
  E --> F[Cached Pi and OMP reports]
  B --> G[Unknown-price refresh]
  G --> D
Loading

Before merge

  • Resolve merge risk (P1) - Upgrading invalidates Pi/OMP report caches carrying the predecessor pricing key, causing a one-time rescan and potentially changed local cost estimates.
  • Resolve merge risk (P1) - The test harness now relies on platform-specific PID enumeration; exact-head macOS Swift-test validation is still in progress.
  • Complete next step (P2) - This owner-authored PR needs normal maintainer merge review and completion of exact-head macOS validation; no mechanical repair is identified.
Agent review details

Security

None.

Review metrics

Metric Value Why it matters
Change distribution production +33/-27, test/CI +487/-32, docs/changelog +3/-1 Most added code is targeted regression and test-runner coverage; the production alias and cache changes remain small.
Affected paths 14 files affected The PR combines a focused pricing repair with a separately scoped test-harness reliability change.

Merge-risk options

Maintainer options:

  1. Confirm upgrade and macOS behavior (recommended)
    Wait for the exact-head macOS test shards, then merge with the documented one-time Pi/OMP cache recalculation policy.
  2. Pause the harness portion if macOS diverges
    If macOS process-inventory checks fail, split or repair the test-runner change before landing the pricing fix.

Technical review

Best possible solution:

Retain the provider-scoped alias mapping and explicit one-time derived-cache repricing, then merge once the exact-head macOS test shards confirm the process-inventory change.

Do we have a high-confidence way to reproduce the issue?

Yes: the new focused tests and source path establish the missing k3[1m]-to-Kimi catalog route, and the supplied PR evidence describes an isolated CLI before/after scenario. This read-only review did not execute that scenario.

Is this the best way to solve the issue?

Yes: appending the canonical Kimi candidate only within Claude’s existing Kimi-scoped targets preserves exact-route precedence and avoids a cross-provider fallback.

AGENTS.md: found and applied where relevant.

Codex review notes: model internal, reasoning high; reviewed against 8a20919a3e0b.

Labels

Label justifications:

  • P2: This repairs local cost-estimate accuracy without evidence of data loss, authentication failure, or core availability loss.
  • merge-risk: 🚨 compatibility: Existing Pi/OMP derived pricing reports are deliberately invalidated and recalculated during upgrade.
  • merge-risk: 🚨 automation: The PR changes the cross-platform process-inventory implementation used by the sharded Swift test runner.
  • rating: 🐚 platinum hermit: Overall readiness is 🐚 platinum hermit; proof is 🌊 off-meta tidepool and patch quality is 🐚 platinum hermit.
  • status: 👀 ready for maintainer look: ClawSweeper has no concrete contributor-facing blocker left for this PR. Not applicable: This owner-authored PR is exempt from the external-contributor proof gate; its supplied body nevertheless documents an isolated synthetic CLI before/after for the changed Claude pricing path.

Evidence

What I checked:

Likely related people:

  • steipete: Authored all three commits on this PR and has recent merged-history ownership across the pricing and test-runner surfaces. (role: recent area contributor; confidence: high; commits: 70a204242da2, 44729e8736c8, 1740f8aa71d3; files: Sources/CodexBarCore/Vendored/CostUsage/CostUsagePricing.swift, Sources/CodexBarCore/CostUsageFetcher.swift, Scripts/ci_swift_test_by_suite.py)
  • Yuxin Qiao: Recent history credits Claude first-party vendor routing and its pricing/cache fingerprinting, the main behavior this PR extends. (role: feature-history contributor; confidence: high; commits: e0b22be9a4fa, fbe70c3e9c65; files: Sources/CodexBarCore/Vendored/CostUsage/CostUsagePricing.swift, Sources/CodexBarCore/PiSessionCostScanner.swift)

Rank-up moves

Optional improvements that raise the rating; they are not merge blockers.

  • Let the exact-head macOS Swift-test shards complete before merge.

Rating scale

Score Internal tier Crab rank Meaning
6/6 S 🦀 challenger crab Exceptional readiness
5/6 A 🦞 diamond lobster Very strong readiness
4/6 B 🐚 platinum hermit Good normal PR; ordinary maintainer review
3/6 C 🦐 gold shrimp Useful, but confidence is limited
2/6 D 🦪 silver shellfish Proof or implementation needs work
1/6 F 🧂 unranked krab Not merge-ready
N/A NA 🌊 off-meta tidepool Rating does not apply

Overall follows the weaker of proof and patch quality.
Shiny media proof means a screenshot, video, or linked artifact directly shows the changed behavior. Runtime, network, CSP, and security claims still need visible diagnostics.

Workflow

  • ClawSweeper keeps one durable marker-backed review comment per issue or PR.
  • Re-runs edit this comment so the latest verdict, findings, and automation markers stay together instead of adding duplicate bot comments.
  • A fresh review can be triggered by eligible @clawsweeper re-review comments, exact-item GitHub events, scheduled/background review runs, or manual workflow dispatch.
  • PR/issue authors and users with repository write access can comment @clawsweeper re-review or @clawsweeper re-run on an open PR or issue to request a fresh review only.
  • Maintainers can also comment @clawsweeper review to request a fresh review only.
  • Fresh-review commands do not start repair, autofix, rebase, CI repair, or automerge.
  • Maintainer-only repair and merge flows require explicit commands such as @clawsweeper autofix, @clawsweeper automerge, @clawsweeper fix ci, or @clawsweeper address review.
  • Maintainers can comment @clawsweeper explain to ask for more context, or @clawsweeper stop to stop active automation.

History

Review history (6 earlier review cycles)
  • reviewed 2026-08-29T00:02:08.755Z sha 70a2042 :: needs changes before merge. :: [P2] Check all targets after a shared catalog refresh
  • reviewed 2026-08-29T00:08:06.656Z sha 44729e8 :: needs maintainer review before merge. :: none
  • reviewed 2026-08-29T00:14:27.317Z sha 44729e8 :: needs maintainer review before merge. :: none
  • reviewed 2026-08-29T00:51:47.983Z sha 44729e8 :: needs maintainer review before merge. :: none
  • reviewed 2026-08-29T01:02:51.473Z sha 44729e8 :: needs maintainer review before merge. :: none
  • reviewed 2026-08-29T02:01:54.723Z sha 1740f8a :: needs maintainer review before merge. :: none

@clawsweeper clawsweeper Bot added rating: 🐚 platinum hermit Good normal PR readiness with ordinary maintainer review expected. status: 👀 ready for maintainer look ClawSweeper has no concrete contributor-facing blocker left for this PR. rating: 🦞 diamond lobster Very strong PR readiness with only minor maintainer review expected. merge-risk: 🚨 compatibility 🚨 Merging this PR could break existing users, config, migrations, defaults, or upgrades. and removed status: ⏳ waiting on author ClawSweeper has contributor-facing work open and is waiting for author action. rating: 🦐 gold shrimp Decent PR readiness signal, but merge confidence is limited. rating: 🐚 platinum hermit Good normal PR readiness with ordinary maintainer review expected. labels Aug 29, 2026
@clawsweeper clawsweeper Bot added merge-risk: 🚨 automation 🚨 Merging this PR could break CI, automerge, proof capture, label sync, or automation. rating: 🐚 platinum hermit Good normal PR readiness with ordinary maintainer review expected. and removed rating: 🦞 diamond lobster Very strong PR readiness with only minor maintainer review expected. labels Aug 29, 2026
@steipete
steipete merged commit 69df341 into main Aug 29, 2026
9 checks passed
@steipete

Copy link
Copy Markdown
Owner Author

Landed as 69df3415adf9 after all checks passed on 1740f8aa71d32cfd3e87ed57248a93e7cadf4cf2.

The fix preserves exact provider routes and recorded model names while resolving the documented Kimi alias. It also uses a canonical price immediately when an earlier provider group downloaded it. Maintainer decision: accept the tested one-time Pi/OMP derived-cache recalculation; raw transcripts remain untouched, and unaffected native Codex cache state is preserved.

Validation is complete: full exact-head CI passed both macOS shards, Linux x64/arm64, musl, lint, and the aggregate gate; make check, focused regression suites, and independent reviews passed. Local Swift coverage includes all 957 selections across a verified 47-group checkpoint and 33-group continuation. The interrupted attempts and recovery are documented in the PR body, including the test-runner ps timeout that was repaired with native process inventory. All 73 cleanup tests passed with one platform-specific skip, and the 393-selection continuation passed with no failures, retries, or timeouts.

The isolated actual-CLI before/after preserves the same synthetic transcript's 160 tokens and raw model name while changing its unknown price to the expected synthetic 0.000385 USD estimate. This is not live-account or current-price validation; no credentials, browser cookies, provider APIs, or app relaunch were involved.

The 0.56.1 Unreleased changelog and pricing documentation are updated. Local main is fast-forwarded and clean, and its tree is identical to the tested head. #2374 and its broader feature work remain open and separate. No release was published.

olddonkey added a commit to olddonkey/CodexBar that referenced this pull request Aug 30, 2026
steipete#3259 moved the Codex parser hash to `d9a91f31d0addc15` and removed the Pi
cache's reviewed-predecessor adoption entirely, so this branch no longer carries
a Pi gate: `PiSessionCostScanner` and `PiSessionCostCompatibilityTests` are
byte-identical to main again, which is the conservative option the maintainer
offered on the previous head.

- Regenerate `CodexParserHash.value` to `ac4862abcdfe21a8`.
- Record main's `d9a91f31d0addc15` in `CostUsageStore.compatiblePredecessorParserHashes`
  and in its exact-set assertion.
olddonkey added a commit to olddonkey/CodexBar that referenced this pull request Aug 31, 2026
steipete#3259 removed the Pi cache's reviewed-predecessor adoption entirely, so this
branch no longer carries a Pi gate: `PiSessionCostScanner` and
`PiSessionCostCompatibilityTests` are byte-identical to main again, which is the
conservative option the maintainer offered on an earlier head.

- Regenerate `CodexParserHash.value` to `7757495e4cc975df`.
- Record main's `b77d4ec72e14ea63` in `CostUsageStore.compatiblePredecessorParserHashes`
  and in its exact-set assertion.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

merge-risk: 🚨 automation 🚨 Merging this PR could break CI, automerge, proof capture, label sync, or automation. merge-risk: 🚨 compatibility 🚨 Merging this PR could break existing users, config, migrations, defaults, or upgrades. P2 Normal priority bug or improvement with limited blast radius. rating: 🐚 platinum hermit Good normal PR readiness with ordinary maintainer review expected. status: 👀 ready for maintainer look ClawSweeper has no concrete contributor-facing blocker left for this PR.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant