Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
35 changes: 29 additions & 6 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -113,7 +113,9 @@ the arrows keep working inside `claude`.
| `↵` | open / attach |
| `n` | new session |
| `a` | add project |
| `e` | rename project |
| `x` | close session |
| `c` | connect this session to one on another project |
| `t` | theme picker |
| `?` | help |
| `q` | quit |
Expand Down Expand Up @@ -226,12 +228,33 @@ you would want to watch it. The reviewer is handed the diff, given no tools,
and run in a mode that answers but cannot act. `analysis` collects the answer
along with what it cost.

The sidebar shows `⊙ n` for claims held, `✉ n` for messages waiting, and
`⚗ n · $x.xx` for reviews running and what they have cost. That last one is
the only thing in Deck that spends money with nobody watching it: the review
has no pane, and the session that asked for it has moved on. The figure stays
after the last review finishes, so a total is not lost the moment it stops
moving.
The sidebar shows `⊙ n` for claims held, `✉ n` for messages waiting, and `⚗` for
reviews. That last one is the only thing in Deck that spends money with nobody
watching it: the review has no pane, and the session that asked for it has
moved on. While a review runs the badge shows how many are in flight and the
tokens they have produced; once they land it shows the total spent, and that
figure stays so it is not lost the moment it stops moving. It is tokens during
and dollars after because the CLI reports a cost only when a turn ends — a
dollar figure while the review ran would sit at zero and read as free.

## Connecting sessions across projects

Everything above is scoped to one project, which is where sessions share a
repository. Press `c` to connect the selected session to one on **another**
project — an API changing in one repository while its consumer changes in
another is the case that scope cannot express.

A connection is a pair, and it widens rather than narrows. The two sessions
appear in each other's `sessions`, can read each other's `work`, can `analyse`
it, can `message` each other, and see each other's notes. Connect A to B and A
to C and A sees both, while B and C stay invisible to each other. Pressing `c`
on a session you are already connected to disconnects it. Connections are
saved, so they survive restarting Deck.

Claims are the deliberate exception: they stay inside one project. A claim is a
path relative to a repository root, so two sessions in different repositories
touching `internal/api/client.go` are not in each other's way, and saying they
were would make the whole mechanism worth ignoring.

## Status

Expand Down
57 changes: 56 additions & 1 deletion docs/architecture.md
Original file line number Diff line number Diff line change
Expand Up @@ -271,7 +271,8 @@ Sessions are isolated by construction — separate worktrees, separate claude
transcripts — so two agents will happily refactor the same interface in
parallel and find out at merge time. `internal/coord` is the one channel
through that isolation: an in-process MCP server exposing `sessions`, `claim`,
`release`, `note`, and `notes`.
`release`, `note`, `notes`, `message`, `inbox`, `work`, `analyse` and
`analysis`.

It is modelled on cathode's `approvals.go` (hand-rolled JSON-RPC over
Streamable HTTP, honouring the client's `Accept` header for SSE framing) with
Expand Down Expand Up @@ -373,6 +374,44 @@ with anything else lying around in the tree.
later, for the reason in (1): by the time anyone asks, the branch it came from
has moved.

### Connections widen a project, never narrow one (`coord.sees`)

Everything above scopes to a project, because that is where sessions share a
repository. A **connection** is the one exception, and it goes outward: it
joins two sessions on *different* projects so each can see the other's work.
The case it exists for is an API changing in one repository while its consumer
changes in another, which project scope excludes by construction.

It is a pair — `store.Connection{A, B}` — rather than a named set. Sessions on
one project already see each other, so the motivating case is exactly two
sessions, and a set would need a name, a member editor and a rule for the last
member leaving. Connecting A to B and A to C lets A see both without making B
and C visible to each other. The pair lives on `State`, not on `Session`, so
there is no two-way link to keep consistent.

Three things follow, and the first is the one a later reader is most likely to
"fix":

1. **Claiming does not go through `sees`.** A claim is a repo-relative path, so
two sessions in different repositories both claiming
`internal/api/client.go` would be told they collide over a file they do not
share. One false conflict is enough to teach an agent to ignore the
mechanism. A connection widens what a session can read; it must never widen
the soft lock.
2. **The shared log widens on read only.** A note still goes to the writer's
own project log. A reader gets that merged with what connected sessions
wrote in theirs, filtered to those sessions — a connection joins two
sessions, so handing over the far project's whole log would publish the
notes of every session there.
3. **The coordinator is given the whole set, never a delta.** `SetConnections`
replaces. The store owns the document and the coordinator holds a copy of
it; a copy maintained by deltas drifts the first time an update is missed,
and nothing observes the drift.

A result that crosses the boundary says so. `work` and the `sessions` rows
carry a `project` only when the far session is on another one, so a row without
it is on yours and its paths resolve against the tree you are looking at.

### A spawned review (`coord.Analyse`, `agent.RunClaude`)

`analyse` starts a **separate agent** to review a sibling's work. Four
Expand Down Expand Up @@ -406,6 +445,22 @@ on whether its context was read from cache or written to it — so the running
total is what makes a pattern visible. It is dropped with the session, like
claims and the inbox, because a review belongs to the session that paid for it.

**A run in flight reports tokens, not dollars.** The CLI is read as
line-delimited events (`--output-format stream-json` with
`--include-partial-messages`), and `RunClaude` takes a callback that fires on
every event carrying usage. Two facts about that stream decide what the UI can
show. Only the **output** count moves: the first event of a turn already
carries the final input, cache-read and cache-write figures — one measured run
knew 15,888 read and 7,954 written before generating a word. And **cost appears
only in the result event**, so a running badge showing dollars would sit at
zero for the whole review and read as free. The sidebar therefore shows tokens
while a review runs and the total once it lands.

Reading stdout to the end before waiting on the process is what makes the
paragraph above about failed turns actually true. It used to be `cmd.Output`,
which turns any non-zero exit into a Go error and takes the result envelope
down with it — the exact case the design says must keep its accounting.

Jobs are bounded like the inbox and the log. Dropping the oldest is safe:
`Spend` is a running total kept separately, so a discarded record costs the
reader an old answer and never the bill.
Expand Down
132 changes: 92 additions & 40 deletions docs/backlog.md
Original file line number Diff line number Diff line change
Expand Up @@ -143,46 +143,98 @@ currently awkward without a mouse. It lived under "Not built yet" until the
deferral reason expired, and was listed in both places for a while — a deferred
item and a planned one are different claims.

## 10. Connections between sessions in a project

Sessions on one project can already read each other's work (`work`) and have it
reviewed by a spawned agent (`analyse`). Both are scoped to the project, which
is the same scope `Siblings`, `notes` and `message` use.

A **connection** would be a smaller grouping inside that: a set of sessions
that share context and can review each other, persisted as
`Connections []Connection` on `store.State` rather than as a field on
`Session`, so there is no two-way link to keep consistent. `Load` already
back-fills missing fields, so an older state file stays readable.

It was deliberately deferred rather than built first. A connection gates two
things — the shared log and the analysis — and until the analysis existed a
`Connection` type would have been stored, rendered and read by nothing. Now
that both exist, the question is answerable from use rather than from
prediction: **is project scope actually too coarse?** With three or four
sessions on one repository it is not, and a grouping inside it would add a
concept without removing a problem.

The case that project scope genuinely cannot express is a connection *across*
projects — an API changing in one repository while its consumer is updated in
another. `Siblings` excludes that by construction. If connections are built,
that is the motivating case, and it inverts the framing: a connection is not a
narrowing of the project, it is an escape from it.

One question decides the shape and should be answered before any code: does a
connection **narrow the shared notes log**, or only gate the analysis?
Narrowing changes behaviour every sibling relies on today; gating is additive.

## 11. Live token counts while a review runs

The sidebar shows elapsed time while a spawned analysis is in flight and the
exact cost once it lands, because the figures arrive only in the final result
envelope. Showing them as they accumulate needs `--output-format stream-json`
with `--include-partial-messages`, which is a parser rather than a field read.

Worth doing only if a review ever runs long enough that watching the number
move tells you something a spinner does not. The measured runs so far finish in
a few seconds.
## ~~10. Connections between sessions~~ — done 2026-08-28

Shipped, and the entry that stood here answered its own question. It asked
whether a connection should **narrow** a project, and concluded that the case
project scope cannot express is the opposite one: an API changing in one
repository while its consumer changes in another. So a connection is an escape
from the project, never a subdivision of it, and the shared log question it
left open resolves the same way — nothing a sibling relies on today changed.

Four decisions are worth keeping.

**A pair, not a named set.** `store.Connection{A, B}`. Sessions on one project
already see each other, so the motivating case is exactly two sessions. A set
would need a name, a member editor and a rule for the last member leaving, and
none of that is asked for by the case that justified the feature. Connecting A
to B and A to C lets A see both without making B and C visible to each other.

**Claims deliberately do not widen.** Every other scoping test became `sees`;
`Claim` kept `ProjectID`. A claim is a repo-relative path, so two sessions in
different repositories both claiming `internal/api/client.go` would be reported
as colliding over a file they do not share, and an agent that meets one false
conflict stops trusting the mechanism. This is the exception most likely to be
tidied away by a later reader, which is why it has a test named for it.

**The read widens, the write does not.** A note still goes to the writer's own
project log. A reader gets that log merged with what connected sessions wrote
in theirs, filtered to those sessions — a connection joins two sessions, so
handing over the far project's whole log would publish the notes of every
session there.

**The coordinator is told the whole set, never a delta.** `SetConnections`
replaces. The store owns the document and the coordinator holds a copy; a copy
updated by deltas is free to drift the first time an update is missed, and the
drift is invisible.

Still open, and deliberately not built: connecting sessions whose agents are
not running. The registry is live, so a connection to a stopped session is
recorded and does nothing until it starts. That matches `work`, which cannot
read an exited session either, and it is the same underlying question — what a
session means after its agent stops.

## ~~11. Live token counts while a review runs~~ — done 2026-08-28

The run now uses `--output-format stream-json --verbose
--include-partial-messages` and `agent.RunClaude` takes a callback that fires
on every event carrying usage.

What a captured stream showed, and what the design follows from: **only the
output count moves.** The first event of a turn already carries the final
input, cache-read and cache-write figures — in one measured run, 15,888 read
and 7,954 written were known before a single word was generated. And **cost is
not in any event but the last**, so a run in flight can report what it is using
and not what it will cost. The sidebar therefore shows tokens while a review
runs and dollars once it lands, rather than a dollar figure that would sit at
zero for the whole run and read as free.

One format, not two. The non-live path could have kept `--output-format json`,
but a second parser is a second place for claude's field names to drift, and
the totals are the thing least affordable to get quietly wrong.

It also closed a gap the old code's own doc comment denied. `RunClaude`
promised to return an error "only when there is no accounting at all", while
`cmd.Output` turned any non-zero exit into an error and discarded the result
envelope with it. Reading stdout to the end before waiting means a result that
arrived is returned whatever the process does afterwards.

## 13. A shell as a session

Not every session wants an agent. Reaching a server, running a migration, or
watching a log is work that belongs beside the agents rather than in a separate
terminal.

Almost all of it already works: `agent.Start` runs whatever `store.Session.Agent`
names, `coordArgs` gives no coordination flags to a program it does not know,
and `willResume` refuses `--continue` to anything but claude. `deck -agent
/bin/zsh` is a working shell session today. What is missing is the menu entry,
and three decisions around it.

1. **Which command.** `$SHELL`, falling back to `/bin/sh`. It holds the login
shell on both macOS and Linux and Deck inherits it from the terminal it was
started in. The authoritative record is per-platform and needs a subprocess
to read — `getent passwd` on Linux, Open Directory on macOS, where
`/etc/passwd` holds only system accounts — to reproduce a value already in
hand.
2. **`-agent-args` must not reach it.** `agentArgsFor` prepends them
unconditionally. That is harmless while `-agent` and `-agent-args` are set
together, and stops being harmless once a shell is on the menu.
3. **The choice must not stick.** Submitting the form writes the agent to
settings as the next session's default, which is wrong for a one-off.

Numbered 13 rather than reusing 12: that number already names the public-tree
notice, and a closed entry should not change meaning.

## ~~12. A public-repo notice a cloner will meet~~ — done 2026-08-28

Expand Down
131 changes: 131 additions & 0 deletions internal/agent/claudestream.go
Original file line number Diff line number Diff line change
@@ -0,0 +1,131 @@
package agent

// Reading what `claude -p --output-format stream-json` emits: one JSON object
// per line, and the totals in the result event at the end.

import (
"bufio"
"encoding/json"
"fmt"
"io"
"strings"
"time"
)

// maxEventLine bounds one line of the stream.
//
// The result event carries the model's whole answer, so the longest line is as
// long as a review. Four megabytes is far above any answer measured and far
// below a size that would matter; a line past it is reported rather than
// truncated, because half a JSON object parses as nothing and would otherwise
// read as "the run produced no result".
const maxEventLine = 4 << 20

// claudeEvent is one line of the stream.
//
// Usage arrives in four different places depending on the event, which is why
// this carries the shape rather than a flat set of fields: the result event
// holds it at the top level, an assistant event under message, and a
// stream_event under event or under event.message.
type claudeEvent struct {
Type string `json:"type"`

// The result event's own fields. Absent everywhere else.
Subtype string `json:"subtype"`
IsError bool `json:"is_error"`
Result string `json:"result"`
DurationMS int `json:"duration_ms"`
CostUSD float64 `json:"total_cost_usd"`
Usage *Tokens `json:"usage"`

Message *usageHolder `json:"message"`
Event *struct {
Type string `json:"type"`
Usage *Tokens `json:"usage"`
Message *usageHolder `json:"message"`
} `json:"event"`
}

type usageHolder struct {
Usage *Tokens `json:"usage"`
}

// usage is the accounting this event carries, or nil.
func (e claudeEvent) usage() *Tokens {
switch {
case e.Usage != nil:
return e.Usage
case e.Message != nil && e.Message.Usage != nil:
return e.Message.Usage
case e.Event == nil:
return nil
case e.Event.Usage != nil:
return e.Event.Usage
case e.Event.Message != nil && e.Event.Message.Usage != nil:
return e.Event.Message.Usage
}
return nil
}

// readClaudeStream consumes the stream and returns what the run reported,
// calling onUsage with each accounting it passes.
//
// Only Output moves once the run has started: the first event of a turn
// already carries the final input, cache-read and cache-write counts, which is
// why a live display fills in almost at once and then creeps. Cost is not in
// any of them — it appears only in the result — so a run in flight can report
// how much it is using and not what it will cost.
//
// A line that does not parse is skipped rather than fatal. The stream is
// several event kinds wide and gains more over time, and refusing a run
// because one line was unfamiliar would throw away the accounting on the next.
func readClaudeStream(r io.Reader, onUsage func(Tokens)) (ClaudeRun, error) {
sc := bufio.NewScanner(r)
sc.Buffer(make([]byte, 0, 64<<10), maxEventLine)

var run ClaudeRun
var got bool
for sc.Scan() {
line := sc.Bytes()
if len(line) == 0 {
continue
}
var e claudeEvent
if json.Unmarshal(line, &e) != nil {
continue
}
if u := e.usage(); u != nil && onUsage != nil {
onUsage(*u)
}
if e.Type != "result" {
continue
}
run = ClaudeRun{
Text: e.Result,
CostUSD: e.CostUSD,
Took: time.Duration(e.DurationMS) * time.Millisecond,
}
if e.Usage != nil {
run.Tokens = *e.Usage
}
if e.IsError {
run.Text = ""
run.Failure = strings.TrimSpace(e.Subtype + ": " + e.Result)
}
got = true
}
// The read error is subordinate to the result, not the other way round. got
// is set only by a fully parsed result event, so once one has arrived the
// accounting is in hand and a later oversized or unreadable line cannot
// take it back. Returning the error here regardless would lose the cost,
// the tokens and the answer of a run that had already reported all three —
// the exact case ClaudeRun's shape exists to prevent.
if err := sc.Err(); err != nil && !got {
return ClaudeRun{}, fmt.Errorf("read claude output: %w", err)
}
if !got {
// No result event means no accounting: whatever it spent, we cannot say.
return ClaudeRun{}, fmt.Errorf("claude produced no result event")
}
return run, nil
}
Loading
Loading