Skip to content

Grok local scan: signals.json undercounts by 54.5x, and costUsdTicks is a usable real-spend source (divisor verified as 1e10) #3345

Description

@initH271

Grok local scan: signals.json undercounts by 54.5x, and costUsdTicks is a usable real-spend source (divisor verified as 1e10)

Context

Not a PR proposal. #3135 and #3236 are already rewriting the Grok local scanner and are in review; I'm not opening a competing branch. I measured an independent corpus and got two results worth putting on the record — one corroborating #3135, one that I think argues against a design choice all three open PRs currently share.

Corpus: 51 session directories under ~/.grok/sessions, 934 completed turns, 2026-07-22 → 2026-08-25, models grok-4.6-build and grok-4.5-build.

1. main still reads signals.json — 54.5x undercount

docs/grok.md:115 and GrokLocalSessionScanner.swift on current main aggregate totalTokensBeforeCompaction, contextTokensUsed, and modelsUsed. Those are context-window occupancy snapshots, not per-turn consumption.

Source Tokens
signals.json (totalTokensBeforeCompaction + contextTokensUsed) 44.9M
turn_completed in updates.jsonl (934 turns) 2.45B (2.44B in / 7.9M out)

54.5x. #3135 reports 653K vs 54.1M (83x) on its own corpus — different ratio, same failure mode, so this is not specific to one usage pattern.

Mechanism, stated explicitly because it explains why the ratio varies: totalTokensBeforeCompaction counts each compaction event's context snapshot once, while billed input is the full context re-sent every turn. My corpus is 97.6% cachedReadTokens (2.38B of 2.44B) — cache reads are real billed input that signals.json has no field to represent. The ratio therefore scales with turns-per-compaction, which is why my 54.5x and #3135's 83x differ while both being correct.

2. costUsdTicks / 1e10 == USD — verified, and it is the only real-spend source

All three open PRs avoid costUsdTicks (#2407: "missing ticks are not estimated"; #3135: models.dev list price labeled an estimate; #3236: token-only, no dollars), and ccusage documents ignoring it because the scale is undocumented. I want to submit evidence that the scale is determinable, because the consequence is material.

Method. Price each turn from the published xAI rates and compare against costUsdTicks / 1e10. Two structural facts must be modeled or the check fails:

  1. Long-context tiering. At 200k tokens the entire request bills at 2x ($2/$0.50/$6 → $4/$1.00/$12 for 4.6).
  2. A promotional rate from 2026-08-15, at 17% of list. (xAI shipped Grok 4.6 on 2026-08-12 with a launch promotion for Grok Build users.)

Taking the best of {tier1, tier2} × {list, 17%} per turn:

Model turns median error within 1%
grok-4.6-build 645 0.0000% 82.2% (530/645)
grok-4.5-build 286 0.0000% 79.4% (227/286)

The promotional rate is not a fitted parameter — it falls out of the data as a clean date boundary:

Date turns priced at list turns priced at 17%
07-22 → 08-14 (8 active days) 445 0
08-15 → 08-25 (6 active days) 9 477

Zero discount-priced turns before 08-15; 9 list-priced turns after. A fitted constant does not produce that.

A caution on my own result: the 200k tier does not key off inputTokens as recorded. 212 turns with inputTokens >= 200k price at tier 1. So the tier predicate is something else (uncached portion, or per-modelCalls context). This does not affect the conclusion — it means the tier must be read from costUsdTicks rather than recomputed, which is precisely the argument for using the field.

Why this matters for #3135 and #3236

List-price reconstruction cannot see a promotional rate. On my corpus:

Period Real spend (costUsdTicks) List-price reconstruction Overstatement
through 08-14 (list-rate period) $697.56 $822.92 1.18x
08-15 onward (promo period) $264.82 $1,762.74 6.66x
total $962.38 $2,585.66 2.69x

A user inside the promo window would see about 6.7x their actual spend. #3135's "list-price estimate, explicitly labeled" framing is honest about being an estimate, but a 6.7x gap is not a rounding artifact — it is the difference between a bill and a sticker price during any promotional period, and xAI runs these at model launches. The residual 1.18x in the list-rate period comes from the tier predicate noted above: reconstructing the 200k tier from inputTokens over-applies the 2x rate.

costUsdTicks already encodes tier and promo. Suggested handling: use it when present, fall back to list price only when absent, and label which one produced each figure. Absence is rare — 2 of 934 turns (0.21%) had costUsdTicks == 0, both early in the corpus.

I'd also gently flag that this cuts against #3236's "do not invent dollar/API-rate estimates for subscription credits" rationale: reading costUsdTicks is not inventing an estimate, it is reading the number the CLI recorded. Whether to display dollars for subscription users stays a product call — but the data is there and it is accurate.

3. Two bounds notes for the scanners under review

  • costUsdTicks coverage is high: 2/934 turns (0.21%) missing. Add Grok local token cost tracking and main-menu usage chart #2407's "missing ticks are not estimated" stays correct but will rarely trigger on a mid-2026 corpus.
  • 7 of 51 session directories have no turn_completed rows at all (aborted/short sessions). Keying off turn_completed presence rather than directory existence avoids carrying empty entries.

Environment

  • CodexBar main as of 2026-09-01 (docs/grok.md:115, GrokLocalSessionScanner.swift)
  • Grok Build: 51 session dirs, 934 turns, 2026-07-22 → 2026-08-25
  • macOS 15.6 (Darwin 25.6.0)

Ask

Nothing beyond recording this, and specifically that the costUsdTicks scale question be treated as answered rather than open. Happy to re-run any query against my corpus if it would settle a review question — field distributions, per-version behavior, specific edge-case records. Not planning a competing PR.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    P2Normal priority bug or improvement with limited blast radius.clawsweeper:needs-live-reproClawSweeper needs live local, crabbox, or manual validation to confirm this issue.clawsweeper:needs-product-decisionClawSweeper marked this issue as needing a product or behavior decision.clawsweeper:no-new-fix-prClawSweeper does not recommend queueing a new automated fix PR for this issue.impact:otherThis issue has meaningful maintainer-visible impact outside the owned taxonomy.issue-rating: 🐚 platinum hermitGood issue quality with a plausible reproduction path needing some confirmation.

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions