Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
11 changes: 11 additions & 0 deletions .memory/subagent-live-context-window.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,11 @@
# Live subagent context window

Status: desktop implementation in PR #271; CI confirmation pending.

After a child provider response, `SubagentEventProjector.usage()` publishes the latest reported token total and model context window through the `chat:subagent-context` notification. Unknown windows, empty usage and finished runs do not publish. Observer failures cannot fail the run.

This is an ephemeral desktop side channel. It adds no durable snapshot fields, revisions or run-store writes. The renderer retains at most 32 chats × 64 runs, and displays readings only for active runs. Roster percentages and the detail token count use existing context formatting and warn at 80% of the full model window.

Coverage lives in the projector, context-usage-store and subagents-panel suites. The new store test is registered in the package scripts and CI registry. Recovery review found this memory note missing from the original commit and added it to match the plan and PR description; no implementation changed in that follow-up.

Follow-ups recorded in `docs/plans/subagent-live-context-window-plan.md`: Remote/iOS/Android need a separate live event and contract revision; message-list chips do not yet show the reading; percentages use the full context window rather than the usable input budget; notifications emitted before the first panel mounts are not replayed.
6 changes: 6 additions & 0 deletions .papercuts/troubleshooting.md
Original file line number Diff line number Diff line change
Expand Up @@ -1407,6 +1407,11 @@ because their native file-mutator test binary had not been built. Run
## 2026-09-26 PR #121 merge of #251 (Remote contract revision 14)
- A PR that adds to the Remote contract has to renumber when main bumps `contractRevision`. The conflicts show up in 7 files: both fixtures, the TS/iOS/Android fixture assertions and the iOS fixture CodingKeys. After resolving, `cmp` the Android copy against the shared fixture. Plan docs that name the revision also go stale.

## 2026-09-27 Live subagent context window
- A fresh workflow worktree has no `node_modules`, so `tsc` reports hundreds of misleading pi-ai type errors. Run `npm ci` before the first type-check.
- The worktree guard refuses a heredoc and a `python3` run in the same Bash command. Write the script to the scratchpad and run it as a separate command.
- Adding a field to `SubagentRunSnapshot` is costly: the exact-key parsers reject it on older builds, and revision/replay checks gate history reads. Live, per-response data belongs on a separate notification channel.

## 2026-09-27 Mobile remembered model selection
- The worktree-isolation guard refuses any command whose arguments contain a runtime value (`$HOME`, `$f`, `$(ls ...)`), including `ANDROID_HOME=$HOME/...` before `./gradlew` and `xcresulttool --path "$(ls -t ...)"`. Spell out absolute paths and list first, then pass the literal name.
- Stable `xcrun simctl` is missing in agent shells ("unable to find utility simctl"); prefix both `simctl` and `xcodebuild test` with `DEVELOPER_DIR=/Applications/Xcode-beta.app/Contents/Developer`. `xcodebuild test -quiet` prints little, so read pass/fail from `xcresulttool get test-results summary`.
Expand Down Expand Up @@ -1445,3 +1450,4 @@ because their native file-mutator test binary had not been built. Run
- Hosted `iOS simulator tests` failed twice on PR #284 (Android-only diff, head `b01cd474`) with different XCTest cases each time: run [36494766134 attempt 1](https://github.com/sambitcreate/aiden-agent/actions/runs/36494766134/job/109171707959) failed `AidenNativeIntegrationTests.testPhysicalActivityKitLifecycleUsesPrivateBoundedStateAndImmediateCleanup()` and `AidenChatTests.testProgressObservationReleasesCompletedHandleAndCanRestart()`; the single rerun ([job 109190587242](https://github.com/sambitcreate/aiden-agent/actions/runs/36494766134/job/109190587242)) failed `AidenChatTests.testCompletedUploadsRemainOwnedThroughRemovalAndExplicitCleanup()`. Main run [36490573933](https://github.com/sambitcreate/aiden-agent/actions/runs/36490573933/job/109158009385) (`cb169c60`) also failed `AidenChatTests.testRemovalCancelsAdmittedConsumerBeforeHeldEventsPublish()`. Timing-sensitive iOS chat tests under parallel simulator clones on the hosted macos-26 runner; needs a real fix, not retries or longer timeouts.
- PR #284 head `5826ff70`: `Electron E2E (3/3)` failed `tests/e2e/chat-message-queue.spec.ts:346 › rejected and unknown Steer receipts keep the draft without replaying it` (`toBeVisible` element not found; Playwright marked it flaky, `--fail-on-flaky-tests` failed the job) in [run 36502234240](https://github.com/sambitcreate/aiden-agent/actions/runs/36502234240/job/109195573859); the PR diff was Android-only and the single rerun passed. PR #266 head `0b2975a1`: `Electron E2E (2/3)` was killed by a hosted-runner shutdown signal ([job 109195654649](https://github.com/sambitcreate/aiden-agent/actions/runs/36502260004/job/109195654649)), infrastructure rather than a test failure.
- PR #261 (head `5fa946a6`, git-executable diff): `iOS simulator tests` failed `AidenNativeIntegrationTests.testPhysicalActivityKitLifecycleUsesPrivateBoundedStateAndImmediateCleanup()` and `testCacheWriteCannotPublishPreSendSnapshotAfterSendStarts` in [run 36512227985](https://github.com/sambitcreate/aiden-agent/actions/runs/36512227985/job/109226750116); the single rerun passed. Same flaky iOS family as above.
- PR #278 (head `2c1a75f0`, desktop-only web-search diff): `iOS simulator tests` failed `AidenChatTests.testProgressObservationReleasesCompletedHandleAndCanRestart()` in [run 36522573181](https://github.com/sambitcreate/aiden-agent/actions/runs/36522573181/job/109258515469), and the single rerun ([job 109267060281](https://github.com/sambitcreate/aiden-agent/actions/runs/36522573181/job/109267060281)) failed `AidenChatTests.testCacheWriteCannotPublishPreSendSnapshotAfterSendStarts()`. Left unmerged during the merge train; this iOS flake family now blocks PRs that do not touch iOS.
1 change: 1 addition & 0 deletions docs/plans/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -20,6 +20,7 @@ This directory is the source of truth for Aiden's implementation plans. The engi
| [Transcript polish: sticky headers, preparing stage, turn footers](transcript-polish-sticky-headers-plan.md) | Implemented for review | Opened Thinking disclosures and activity/compaction trails keep a sticky header under the toolbar and flow at full height; pending tool calls read `Preparing <tool>` on desktop, iOS and Android; settled responses show a duration/model/token footer from new content-free `turnStats`. Mobile footer and live streaming footer are follow-ups. |
| [Chronological Chat Motion](chronological-chat-motion-plan.md) | In review | Readable Thinking stretches, tool rows, and prose project in sequence across desktop and regular native chats. [PR #224](https://github.com/sambitcreate/aiden-agent/pull/224) and visual acceptance remain. |
| [Composer Context Meter](composer-context-meter-plan.md) | In review | Composer gauge and popover driven by `projectNextContextUsage` (the runtime compaction projection) with live `chat:context-pressure` pushes; PR #187 kept only this part after PR #224 shipped ordered thinking traces. Visual acceptance pending. |
| [Live subagent context window](subagent-live-context-window-plan.md) | In review | Each running child shows its latest provider-reported context use (percent in the roster row, `used / window` in detail, warning tone at 80%) through a non-durable `chat:subagent-context` side channel; snapshots, revisions and Remote are unchanged. Desktop only; Remote/iOS/Android live context is a follow-up. |
| [Managed Worktree Lifecycle](managed-worktree-lifecycle-plan.md) | Partial | P0 lifecycle merged in #185; admission hardening and [old-stack disposition](managed-worktree-stack-disposition.md) under review: hook-free managed creation, free-space admission, `.worktreeinclude` provisioning, durable snapshot records + synthetic `refs/aiden/snapshots/*` commits, byte-exact provisioned-ignored blob restore, journal v4 snapshot-aware quarantine deletion with force semantics, and first-class restore. GC/owner-kind/setup-script/UI phases remain open. |
| [Durable chat pull requests](chat-pull-requests-plan.md) | Implemented | PR #184 hardens durable multi-link identities, exact create intent, selected push repository, atomic recovery and persistent unlink dismissal; local full suite and two independent final reviews pass; hosted acceptance tracked on PR #184. |
| [Trusted AGENTS refresh](trusted-agents-refresh-plan.md) | Implemented | Bounded global/workspace instruction refresh and per-dispatch scope fence; local loader/runtime/Bot/scheduled/onboarding checks and both reviews pass. PR CI pending. |
Expand Down
28 changes: 28 additions & 0 deletions docs/plans/subagent-live-context-window-plan.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,28 @@
# Live subagent context window

Status: In review (desktop). Source: pi-subagents #2448, reusing the #187 context meter's formatting and 80% warning threshold.

## Goal

While a child agent runs, show how full its context window is, so a user can see which child is close to its limit before it fails or compacts.

## Design

- **Figure.** After each child provider response, `SubagentEventProjector.usage()` builds `{tokens, window}`. `tokens` is the response's reported usage total, the same figure Pi uses as a response's context size. `window` is the child model's `contextWindow`. No reading is made when the model has no window or the provider reported nothing.
- **Transport.** The reading goes on a separate `chat:subagent-context` notification, emitted by `llm-client` through `sendGeneration`. It is **not** a snapshot field. The reasons:
- Snapshot parsers use exact key sets, so older builds would reject a new field.
- History and detail reads gate on revision monotonicity and exact replay, and a reading per response would bump revisions.
- Durable writes are capped per run.
- Remote progress would churn on every child response.
- **Renderer.** `renderer/lib/subagent-context-usage-store.ts` is a bounded, in-memory LRU of up to 32 chats × 64 runs. The panel shows a reading only while the run's view state is active. Terminal runs drop their reading.
- **UI.**
- Each active roster row shows a gauge icon with the percentage, and says it in the row's accessible name.
- The detail pane adds `Context window: 84K / 200K tokens (42%)` under the model line.
- Both use the warning tone at 80% or more.

## Follow-ups

- Remote, iOS and Android: the Remote roster reads persisted snapshots, so live readings need a Remote side channel and a contract revision.
- Show readings in the message-list subagent chips.
- The percentage is of the full window, not the usable input budget (window minus reserved output).
- Readings sent before any subagents panel has mounted in the window are not replayed.
11 changes: 11 additions & 0 deletions main/services/llm-client.ts
Original file line number Diff line number Diff line change
Expand Up @@ -994,6 +994,17 @@ async function prepareGeneration(
chatId: params.chatId,
workspaceId: workspace.id,
modelId: model.id,
contextWindow: model.contextWindow,
// Presentation-only: bypasses persistence and the Remote progress
// revision, which only track durable snapshots.
onContextUsage: (runId, usage) => {
sendGeneration(streamId, "chat:subagent-context", {
streamId,
chatId: params.chatId,
runId,
usage,
});
},
prepareSnapshot: (snapshot) => subagentPersistence.prepare(snapshot),
onControlSnapshot: async (snapshot) => {
subagentPersistence.projectControlSnapshot(snapshot);
Expand Down
91 changes: 91 additions & 0 deletions main/services/subagents/subagent-event-projector.test.ts
Original file line number Diff line number Diff line change
Expand Up @@ -993,6 +993,97 @@ test("turn, usage, and tool telemetry has a hard durable-write bound", async ()
assert.equal(emitted[emitted.length - 1]?.state, "completed");
});

test("live context usage follows the latest response without durable writes or revisions", async () => {
const emitted: SubagentRunSnapshotV1[] = [];
const readings: Array<{ runId: string; tokens: number; window: number }> = [];
const projector = new SubagentEventProjector({
generationId: "generation-context",
chatId: "chat-context",
workspaceId: "workspace-context",
modelId: "model-context",
contextWindow: 200_000,
onContextUsage: (runId, usage) => {
readings.push({ runId, ...usage });
},
onSnapshot: (snapshot) => {
emitted.push(snapshot);
},
});
const runId = "run-context";
projector.begin(
{ runId, groupId: "generation-context:group", childId: "child-context" },
{ role: "scout", label: "Context", task: "Measure context." },
);
projector.running(runId);
await projector.flush();
const durableBefore = emitted.length;
const revisionBefore = projector.snapshot()[0]!.revision;

projector.usage(runId, usageMessage(12_000));
projector.usage(runId, usageMessage(84_000));
// A child compaction shrinks context; the reading is the latest, not a sum.
projector.usage(runId, usageMessage(30_000));
projector.usage(runId, usageMessage(0));
await projector.flush();

assert.deepEqual(readings, [
{ runId, tokens: 12_000, window: 200_000 },
{ runId, tokens: 84_000, window: 200_000 },
{ runId, tokens: 30_000, window: 200_000 },
]);
// Cumulative spend still accrues on the snapshot, which never carries context.
const live = projector.snapshot()[0]!;
assert.equal(live.tokens, 126_000);
assert.equal("contextUsage" in live, false);
assert.equal(emitted.length, durableBefore);
assert.equal(live.revision, revisionBefore + 4);

projector.finish(runId, {
role: "scout",
label: "Context",
status: "completed",
summary: "Done.",
});
projector.usage(runId, usageMessage(99_000));
assert.equal(readings.length, 3, "a finished run publishes no further readings");
});

test("live context usage is skipped without a known window and never breaks the run", () => {
let calls = 0;
const unknownWindow = new SubagentEventProjector({
generationId: "generation-no-window",
chatId: "chat-no-window",
workspaceId: "workspace-no-window",
modelId: "model-no-window",
onContextUsage: () => {
calls += 1;
},
});
unknownWindow.begin(
{ runId: "run-no-window", groupId: "generation-no-window:group", childId: "child" },
{ role: "scout", label: "No window", task: "Measure context." },
);
unknownWindow.usage("run-no-window", usageMessage(500));
assert.equal(calls, 0);

const throwing = new SubagentEventProjector({
generationId: "generation-throwing",
chatId: "chat-throwing",
workspaceId: "workspace-throwing",
modelId: "model-throwing",
contextWindow: 1_000,
onContextUsage: () => {
throw new Error("renderer gone");
},
});
throwing.begin(
{ runId: "run-throwing", groupId: "generation-throwing:group", childId: "child" },
{ role: "scout", label: "Throwing", task: "Measure context." },
);
throwing.usage("run-throwing", usageMessage(500));
assert.equal(throwing.snapshot()[0]!.tokens, 500);
});

test("projector withholds text across stream boundaries until terminal sanitization", () => {
const cases = [
{
Expand Down
23 changes: 23 additions & 0 deletions main/services/subagents/subagent-event-projector.ts
Original file line number Diff line number Diff line change
Expand Up @@ -19,6 +19,10 @@ import {
type SubagentProjectionNoticeKind,
} from "../../../renderer/shared/subagent-runs.js";
import { reportedTokens } from "../usage-accounting.js";
import {
subagentContextUsageFromReport,
type SubagentContextUsageV1,
} from "../../../renderer/shared/subagent-context-usage.js";
import type { SubagentTaskRequest, SubagentTaskResult } from "./contracts.js";
import { sanitizeSubagentSnapshotTextWithFacts } from "../../../renderer/shared/subagent-safe-text.js";

Expand All @@ -35,6 +39,13 @@ export interface SubagentRunProjectorInput {
chatId: string;
workspaceId: string;
modelId: string;
/** The child model's context window, for live context-usage readings. */
contextWindow?: number;
/**
* Live, non-durable context reading after each child response. It never
* enters the run snapshot, so it cannot bump a revision or reach storage.
*/
onContextUsage?: (runId: string, usage: SubagentContextUsageV1) => void;
/** Synchronous authority/admission seam that runs before a new run is published. */
prepareSnapshot?: (snapshot: SubagentRunSnapshotV1) => void;
onSnapshot?: (snapshot: SubagentRunSnapshotV1) => void | Promise<void>;
Expand Down Expand Up @@ -327,6 +338,7 @@ export class SubagentEventProjector {
usage(runId: string, message: AssistantMessage): void {
const current = this.require(runId);
const reported = reportedTokens(message.usage)?.total ?? 0;
this.publishContextUsage(current, reported);
const tokens = Number.isSafeInteger(reported)
? reported
: Math.min(Number.MAX_SAFE_INTEGER, Math.max(0, Math.floor(reported)));
Expand Down Expand Up @@ -456,6 +468,17 @@ export class SubagentEventProjector {
if (this.persistenceError) throw this.persistenceError;
}

private publishContextUsage(current: SubagentRunSnapshotV1, reported: number): void {
if (current.finishedAt !== undefined || !this.input.onContextUsage) return;
const usage = subagentContextUsageFromReport(reported, this.input.contextWindow);
if (!usage) return;
try {
this.input.onContextUsage(current.runId, usage);
} catch {
// A presentation-only observer must never affect the child run.
}
}

private require(runId: string): SubagentRunSnapshotV1 {
const record = this.records.get(runId);
if (!record) throw new Error("Unknown subagent run.");
Expand Down
Loading
Loading