Skip to content

feat(events): context and cost per agent step, ACP-aware run detail - #72

Open
FreshlyBrewedCode wants to merge 3 commits into
acp/authoringfrom
acp/usage-ui
Open

FreshlyBrewedCode wants to merge 3 commits into
acp/authoringfrom
acp/usage-ui

Conversation

@FreshlyBrewedCode

@FreshlyBrewedCode FreshlyBrewedCode commented Oct 2, 2026 •

Copy link
Copy Markdown
Owner

What

This PR implements ADR 0013 §5 and the SPA side of it: context and cost per agent step, and an ACP-aware run detail.

Events and runtime

  • AgentStepFinished gains two optional fields:
    • context: { used, size }: from the step's last usage signal;
    • cost: { amount, currency }: the latest cost any usage signal carried.
  • Cancelled and failed steps keep both. Old logs without the fields still decode.
  • The usage doc comment now says "as the agent reports it": the last message on opencode, the whole turn on Claude. agentStepContextTokens is documented as the fallback for old logs.
  • acp.tool-call chunk (adapter, src/runtime/acp-adapter.ts). translateAcpStream keeps only the first title of a tool call, and on Claude that title is generic (Edit, Read File, Terminal). The descriptive title (Edit math.ts) and the real input arrive in later tool_call_updates, and the translator drops them. Opencode's titles (git status, math.ts) are dropped the same way. So the adapter now follows each tool-call update that has a title with a CUSTOM acp.tool-call chunk, { toolCallId, title, input? }. It carries no signal.

SPA

  • Step rows show 68 chunks · 22.3k/200k ctx · $0.039.
  • Step detail has these rows:
    • context (22.3k / 200k · 11%, with a small meter);
    • cost;
    • tokens (the four usage components, as reported).
  • Run meta has a cost row: the step costs summed per currency, running steps included.
  • Live context: deriveSteps reads acp.usage chunks, so a running step's context and cost update before AgentStepFinished arrives.
  • Fallbacks:
    • logs written by the ACP adapter before this PR use their acp.usage chunks;
    • pre-ACP logs use agentStepContextTokens, with no window, marked "(final turn)".
  • Transcript tool calls are labelled with the agent's title. The kind and state go in the hint, e.g. Edit math.ts with edit · complete. The args show the agent's real input, which also fixes the run-together JSON that re-sent ACP inputs used to produce. When a log has no acp.tool-call chunks, the label falls back to the title inside the arguments. Old opencode logs are unchanged.
  • The chunk names are spelled out in src/web, so the SPA bundle doesn't import the runtime. Tests pin them to ACP_CHUNK.

Testing

  • bun run check is green: 459 tests.
  • nix develop --command bun run test:e2e: 17 passed. Two of them are new, in e2e/acp.e2e.ts, over a seeded ACP-shaped run:
    • context and cost per step, including a cancelled step, the run's summed cost, and the transcript tool title;
    • an old opencode log that falls back to the derived context, with no cost.
  • New unit tests:
    • events.test.ts: round-trip, an old log without the fields still decodes, rejections.
    • agent-step.usage.test.ts: the last context wins; cost is the latest reported, and an update without one keeps it; 0 USD is recorded; a step with no usage has neither field.
    • run.test.ts: a cancelled step keeps context and cost; a completed step records them, and omits them when none came.
    • acp-adapter.test.ts: the fake agent's new tools verb sends Claude-shaped updates; the test checks the acp.tool-call chunks, their order, and that they carry no signal.
    • run-events.test.ts: live context from chunks, cost carry-over, recorded fields win, chunk fallback for pre-feat(events): context and cost per agent step, ACP-aware run detail #65 ACP logs, the agentStepContextTokens fallback, cancelled steps, per-step isolation, runCost.
    • transcript.test.ts: the latest title and input, title-only updates keep the input, the args-title fallback, no extra rows, old corpus logs have no title.
    • format.test.ts: formatCost and formatContext.

Live validation

The setup is a scratch project with a local bare origin, a slug that is not on GitHub, and no write-back. The two-step workflow (implement: read, three separate edits, cat; then review) ran on factory serve --port 3065 with a test db. The browser checks used playwright-cli.

run step rows (final) run cost live
claude / haiku 68 chunks · 22.3k/200k ctx · $0.039, 23 chunks · 20.7k/200k ctx · $0.014 $0.053 the row ticked 20.5k → 20.7k → 20.8k → 21.2k → 21.6k → 21.9k → 22.2k while implement ran
opencode / opencode/big-pickle 173 chunks · 9.1k/200k ctx · $0, 64 chunks · 8.2k/200k ctx · $0 $0 none: opencode sends a single usage update, at the end
claude, cancelled at ~7 s AgentStepFinished{outcome:"cancelled", context:{used:20784,…}}, no cost — —

Transcript labels:

  • Claude: Read math.ts, Edit math.ts ×3, cat …/math.ts (execute).
  • opencode: glob (search), math.ts (read/edit), cat math.ts (execute). The glob args show {"pattern":"**/math.ts"}.

Deviations and notes

Closes #65
Part of #62

🤖 Generated with Claude Code

FreshlyBrewedCode and others added 3 commits October 2, 2026 19:59
AgentStepFinished gains optional `context: {used, size}` and
`cost: {amount, currency}` (ADR 0013 §5), taken from the step's `usage`
signals: the last one's context, and the latest cost any of them carried,
since the agents report the session's cost so far and Claude sends it only
at the end of a turn. Cancelled and failed steps keep what was reported.
The `usage` doc comment now says what it means per agent.

The ACP adapter adds an `acp.tool-call` CUSTOM chunk after each tool call
update that carries a title: translateAcpStream keeps only the first,
generic title (`Edit`, `Read File`), and drops the descriptive one
(`Edit math.ts`) and the real input that arrive in later updates.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- Steps show context against the window and cost; the run's cost is
  summed per currency. Old logs fall back to agentStepContextTokens.
- A running step's context and cost follow its `acp.usage` chunks.
- Transcript tool calls show the agent's title (`Edit math.ts`) and input
  from `acp.tool-call` chunks, else the title in the arguments.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

feat(events): context and cost per agent step, ACP-aware run detail

1 participant