Skip to content

[FEAT] Session-audit CLI (ccusage-style) packaged as a Claude Code plugin + skill #2

Description

@axisrow

Problem

claude-devtools is GUI-first: you browse sessions in the Electron app (or the standalone HTTP UI). But a large part of the target audience lives in the terminal — and in Claude Code itself. The ecosystem has ccusage for macro cost aggregates (daily/monthly/blocks), yet nothing answers the micro question: for this one session, where inside the agent loop did tokens actually leak, and what was wasted?

Concretely, there is no terminal/agent-consumable way to get:

  • a per-turn / per-round token ledger from session JSONL (input vs cache_read vs cache_write vs output, context growth, reread share — in one recent 2-hour session, 94% of 25M billed tokens were context re-reads);
  • waste findings: duplicate tool calls, failed calls, oversized tool outputs, context-growth spikes ("agent scanned the whole disk"), dead prompt caching, thinking-heavy turns;
  • slow-subagent flags (a subagent running > 5 min is worth questioning);
  • a sessions inventory across projects: duration, models used, token totals (i.e. "which sessions ran 2h+ and on which model" — currently not visible anywhere).

Proposal

Two layers:

1. A CLI surface reusing machinery the repo already has — parseJsonlFile / deduplicateByRequestId, pathDecoder, SubagentResolver are all Electron-free, so this adds zero coupling to the app:

pnpm analyze:session <file.jsonl> | --project <dir> --last   # ledger + findings + subagents + est. cost
pnpm analyze:sessions [--min-minutes N] [--sort duration|tokens|date] [--json]

Merged: #1 — src/cli/, ~1.2k lines incl. vitest coverage, built on the existing parser stack.

2. Claude Code plugin + skill packaging, so the analysis is consumable by Claude Code itself — the agent can run the audit on the current session and reason about its own waste:

  • .claude-plugin/marketplace.json at the repo root + .claude-plugin/plugin.json for a session-audit plugin;
  • skills/audit-session/SKILL.md — when to run the audit, how to read the ledger/findings, what to recommend (deny rules for culprit commands, splitting long turns, slimming the parent context);
  • slash-command wrappers (/audit-session, /sessions-inventory) with allowed-tools pre-approving the CLI invocation;
  • install becomes: /plugin marketplace add axisrow/claude-devtools → /plugin install session-audit@…; CI-checkable via claude plugin validate.

Reference: https://code.claude.com/docs/en/plugins.md, https://code.claude.com/docs/en/plugin-marketplaces.md

Why it fits the project

This is priority #2 from CONTRIBUTING ("context engineering insight — how tokens flow through a session") extended to the place where the insight is actionable: the terminal, where the sessions are produced. It also covers headless environments (SSH/docker) where the Electron app can't run.

Prior art / non-goals

  • ccusage answers a different question (macro cost aggregates across days/months/billing windows, many agent sources). Complementary, not a replacement; no published Claude Code plugin wrapping a per-session audit CLI is known to me — this would fill that gap.
  • Not proposing provider-specific billing logic in v1 — findings and ledger are provider-agnostic (exact usage fields from JSONL); cost estimation is optional and pricing-table based.

Roadmap (fork takeover — upstream dormant since 2026-05)

Activity

  1. axisrow commented on Sep 20, 2026

    @axisrow
    OwnerAuthor

    Layer 2 (plugin packaging) split out into #3; this issue stays as the umbrella.

  2. added
    epicТребует декомпозиции
    on Sep 20, 2026
  3. axisrow commented on Sep 20, 2026

    @axisrow
    OwnerAuthor

    Decomposition

    Layer-1 core (ledger, 6 finding types, slow-subagent table, cost estimation, sessions inventory) already shipped via #1 (merged). The umbrella now tracks two native sub-issues:

    # Подзадача LOC Размер Код, мин Review, мин Итого, мин Приоритет
    #4 CLI: unify command surface in ccusage style (flags/ergonomics, not scope) ~250–450 M–L 40–90 15–30 55–120 priority/high
    #3 Layer 2: package the CLI as a Claude Code plugin (marketplace + plugin.json + skills) ~200–300 S–M 25–50 15–25 40–75 —

    Суммарная оценка: ~95–195 мин агентной работы (по «итого», с учётом review-цикла).

    Рекомендованный порядок: #4 → #3. Сначала стабилизировать флаг-грамматику CLI — SKILL-файлы из #3 документируют эти команды, так surface надо зафиксировать до написания скиллов.

    Граница скоупа (для всех субишью): структура команд берётся у ccusage, но функции не смешиваются — ccusage остаётся макро-аналитикой юсейджа по всем агентам, claude-devtools — глубоким аудитом сессий именно Claude Code. Никаких daily/monthly/blocks-агрегатов, мульти-агентных источников и statusline здесь.

    Для исполнителей: реализация через скилл ponytail — минимальный рабочий дифф, без новых зависимостей и спекулятивных абстракций.

  4. axisrow commented on Sep 25, 2026

    @axisrow
    OwnerAuthor

    Все субишью выполнены и смержены: 3-8 через PR 9-13, 21. Эпик завершён; живые follow-ups — 19, 20, 22, 28.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or requestepicТребует декомпозиции

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions