You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Grok local scan: signals.json undercounts by 54.5x, and costUsdTicks is a usable real-spend source (divisor verified as 1e10)
Context
Not a PR proposal. #3135 and #3236 are already rewriting the Grok local scanner and are in review; I'm not opening a competing branch. I measured an independent corpus and got two results worth putting on the record — one corroborating #3135, one that I think argues against a design choice all three open PRs currently share.
Corpus: 51 session directories under ~/.grok/sessions, 934 completed turns, 2026-07-22 → 2026-08-25, models grok-4.6-build and grok-4.5-build.
1. main still reads signals.json — 54.5x undercount
docs/grok.md:115 and GrokLocalSessionScanner.swift on current main aggregate totalTokensBeforeCompaction, contextTokensUsed, and modelsUsed. Those are context-window occupancy snapshots, not per-turn consumption.
54.5x.#3135 reports 653K vs 54.1M (83x) on its own corpus — different ratio, same failure mode, so this is not specific to one usage pattern.
Mechanism, stated explicitly because it explains why the ratio varies: totalTokensBeforeCompaction counts each compaction event's context snapshot once, while billed input is the full context re-sent every turn. My corpus is 97.6% cachedReadTokens (2.38B of 2.44B) — cache reads are real billed input that signals.json has no field to represent. The ratio therefore scales with turns-per-compaction, which is why my 54.5x and #3135's 83x differ while both being correct.
2. costUsdTicks / 1e10 == USD — verified, and it is the only real-spend source
All three open PRs avoid costUsdTicks (#2407: "missing ticks are not estimated"; #3135: models.dev list price labeled an estimate; #3236: token-only, no dollars), and ccusage documents ignoring it because the scale is undocumented. I want to submit evidence that the scale is determinable, because the consequence is material.
Method. Price each turn from the published xAI rates and compare against costUsdTicks / 1e10. Two structural facts must be modeled or the check fails:
Long-context tiering. At 200k tokens the entire request bills at 2x ($2/$0.50/$6 → $4/$1.00/$12 for 4.6).
Taking the best of {tier1, tier2} × {list, 17%} per turn:
Model
turns
median error
within 1%
grok-4.6-build
645
0.0000%
82.2% (530/645)
grok-4.5-build
286
0.0000%
79.4% (227/286)
The promotional rate is not a fitted parameter — it falls out of the data as a clean date boundary:
Date
turns priced at list
turns priced at 17%
07-22 → 08-14 (8 active days)
445
0
08-15 → 08-25 (6 active days)
9
477
Zero discount-priced turns before 08-15; 9 list-priced turns after. A fitted constant does not produce that.
A caution on my own result: the 200k tier does not key off inputTokens as recorded. 212 turns with inputTokens >= 200k price at tier 1. So the tier predicate is something else (uncached portion, or per-modelCalls context). This does not affect the conclusion — it means the tier must be read from costUsdTicks rather than recomputed, which is precisely the argument for using the field.
List-price reconstruction cannot see a promotional rate. On my corpus:
Period
Real spend (costUsdTicks)
List-price reconstruction
Overstatement
through 08-14 (list-rate period)
$697.56
$822.92
1.18x
08-15 onward (promo period)
$264.82
$1,762.74
6.66x
total
$962.38
$2,585.66
2.69x
A user inside the promo window would see about 6.7x their actual spend. #3135's "list-price estimate, explicitly labeled" framing is honest about being an estimate, but a 6.7x gap is not a rounding artifact — it is the difference between a bill and a sticker price during any promotional period, and xAI runs these at model launches. The residual 1.18x in the list-rate period comes from the tier predicate noted above: reconstructing the 200k tier from inputTokens over-applies the 2x rate.
costUsdTicks already encodes tier and promo. Suggested handling: use it when present, fall back to list price only when absent, and label which one produced each figure. Absence is rare — 2 of 934 turns (0.21%) had costUsdTicks == 0, both early in the corpus.
I'd also gently flag that this cuts against #3236's "do not invent dollar/API-rate estimates for subscription credits" rationale: reading costUsdTicks is not inventing an estimate, it is reading the number the CLI recorded. Whether to display dollars for subscription users stays a product call — but the data is there and it is accurate.
7 of 51 session directories have no turn_completed rows at all (aborted/short sessions). Keying off turn_completed presence rather than directory existence avoids carrying empty entries.
Environment
CodexBar main as of 2026-09-01 (docs/grok.md:115, GrokLocalSessionScanner.swift)
Nothing beyond recording this, and specifically that the costUsdTicks scale question be treated as answered rather than open. Happy to re-run any query against my corpus if it would settle a review question — field distributions, per-version behavior, specific edge-case records. Not planning a competing PR.
Grok local scan:
signals.jsonundercounts by 54.5x, andcostUsdTicksis a usable real-spend source (divisor verified as 1e10)Context
Not a PR proposal. #3135 and #3236 are already rewriting the Grok local scanner and are in review; I'm not opening a competing branch. I measured an independent corpus and got two results worth putting on the record — one corroborating #3135, one that I think argues against a design choice all three open PRs currently share.
Corpus: 51 session directories under
~/.grok/sessions, 934 completed turns, 2026-07-22 → 2026-08-25, modelsgrok-4.6-buildandgrok-4.5-build.1.
mainstill readssignals.json— 54.5x undercountdocs/grok.md:115andGrokLocalSessionScanner.swifton currentmainaggregatetotalTokensBeforeCompaction,contextTokensUsed, andmodelsUsed. Those are context-window occupancy snapshots, not per-turn consumption.signals.json(totalTokensBeforeCompaction+contextTokensUsed)turn_completedinupdates.jsonl(934 turns)54.5x. #3135 reports 653K vs 54.1M (83x) on its own corpus — different ratio, same failure mode, so this is not specific to one usage pattern.
Mechanism, stated explicitly because it explains why the ratio varies:
totalTokensBeforeCompactioncounts each compaction event's context snapshot once, while billed input is the full context re-sent every turn. My corpus is 97.6%cachedReadTokens(2.38B of 2.44B) — cache reads are real billed input thatsignals.jsonhas no field to represent. The ratio therefore scales with turns-per-compaction, which is why my 54.5x and #3135's 83x differ while both being correct.2.
costUsdTicks / 1e10 == USD— verified, and it is the only real-spend sourceAll three open PRs avoid
costUsdTicks(#2407: "missing ticks are not estimated"; #3135: models.dev list price labeled an estimate; #3236: token-only, no dollars), and ccusage documents ignoring it because the scale is undocumented. I want to submit evidence that the scale is determinable, because the consequence is material.Method. Price each turn from the published xAI rates and compare against
costUsdTicks / 1e10. Two structural facts must be modeled or the check fails:Taking the best of {tier1, tier2} × {list, 17%} per turn:
The promotional rate is not a fitted parameter — it falls out of the data as a clean date boundary:
Zero discount-priced turns before 08-15; 9 list-priced turns after. A fitted constant does not produce that.
A caution on my own result: the 200k tier does not key off
inputTokensas recorded. 212 turns withinputTokens >= 200kprice at tier 1. So the tier predicate is something else (uncached portion, or per-modelCallscontext). This does not affect the conclusion — it means the tier must be read fromcostUsdTicksrather than recomputed, which is precisely the argument for using the field.Why this matters for #3135 and #3236
List-price reconstruction cannot see a promotional rate. On my corpus:
costUsdTicks)A user inside the promo window would see about 6.7x their actual spend. #3135's "list-price estimate, explicitly labeled" framing is honest about being an estimate, but a 6.7x gap is not a rounding artifact — it is the difference between a bill and a sticker price during any promotional period, and xAI runs these at model launches. The residual 1.18x in the list-rate period comes from the tier predicate noted above: reconstructing the 200k tier from
inputTokensover-applies the 2x rate.costUsdTicksalready encodes tier and promo. Suggested handling: use it when present, fall back to list price only when absent, and label which one produced each figure. Absence is rare — 2 of 934 turns (0.21%) hadcostUsdTicks == 0, both early in the corpus.I'd also gently flag that this cuts against #3236's "do not invent dollar/API-rate estimates for subscription credits" rationale: reading
costUsdTicksis not inventing an estimate, it is reading the number the CLI recorded. Whether to display dollars for subscription users stays a product call — but the data is there and it is accurate.3. Two bounds notes for the scanners under review
costUsdTickscoverage is high: 2/934 turns (0.21%) missing. Add Grok local token cost tracking and main-menu usage chart #2407's "missing ticks are not estimated" stays correct but will rarely trigger on a mid-2026 corpus.turn_completedrows at all (aborted/short sessions). Keying offturn_completedpresence rather than directory existence avoids carrying empty entries.Environment
mainas of 2026-09-01 (docs/grok.md:115,GrokLocalSessionScanner.swift)Ask
Nothing beyond recording this, and specifically that the
costUsdTicksscale question be treated as answered rather than open. Happy to re-run any query against my corpus if it would settle a review question — field distributions, per-version behavior, specific edge-case records. Not planning a competing PR.