Skip to content

feat(ai): AgentTask turns that end in a checked answer, under budgets - #1001

Merged
sroussey merged 4 commits into
mainfrom
claude/agent-loop-finish-parallel-budgets
Oct 1, 2026
Merged

sroussey merged 4 commits into
mainfrom
claude/agent-loop-finish-parallel-budgets

Conversation

@sroussey

@sroussey sroussey commented Oct 1, 2026

Copy link
Copy Markdown
Collaborator

What

AgentTask gains what a batch host reading documents needs from the turn loop:

  • outputSchema adds a submit_answer tool taking that schema. A passing call ends the turn (stopReason: "submitted", the answer on object); a failing call is answered as a correction, so the model resubmits instead of the host failing the turn. A text-only reply is reminded twice to submit.
  • checkSubmission: a host function judges an answer the schema passed (a section left empty that the document has, two figures that disagree) and sends it back with a reason, at most twice. Reasons are on submissionRejections.
  • toolConcurrency: a round's calls run at once, with results in the order asked. A round with a call that needs a person's approval still runs one at a time.
  • steps[]: a per-round record of timings, attempts, each tool's outcome and size, usage and cost. costUsd totals them. Steps also ride on the snapshot event.
  • Budgets (maxInputTokens, maxCostUsd, maxDurationMs) end the turn as "budget", never between a tool_use and its result. maxCostUsd is refused for a model with no price card.
  • maxRoundRetries (default 2) retries a RetryableJobError as the provider's retry-after asks, within 1 to 60 s. roundTimeoutMs abandons a round that a provider accepted and never answered (seen live on OpenAI) and treats it as a retryable failure.

An unset object port reaches the task as {}; an empty outputSchema counts as none.

Why

embarc-data is replacing an external coding agent with this loop for reading SEC filings. Across ~2,600 recorded sessions and a 63-filing labelled set, these were the gaps:

  • no typed finish;
  • tool calls run one at a time;
  • no per-round accounting;
  • no budgets;
  • no retry or timeout for a hung or rate-limited call.

Tests

  • New packages/test/src/test/ai/AgentTaskTurnControls.test.ts, 21 tests.
  • Existing AgentTask and a2a suites pass unchanged, 122 in all.
  • Clean: build:types (47/47), lint, format-check.

🤖 Generated with Claude Code

What a batch host reading documents needs from the turn loop, measured
against thousands of recorded agent sessions:

- outputSchema adds a submit_answer tool taking that schema. A passing call
  ends the turn ("submitted", the answer on `object`); a failing one is
  answered as a correction, so the model resubmits instead of the host
  failing the turn. A text reply is reminded twice to submit.
- checkSubmission lets the host judge an answer the schema passed — a
  section left empty that the document has, two figures that disagree — and
  send it back with a reason, at most twice.
- toolConcurrency runs a round's calls at once, results in the order asked;
  a round with a call put to a person still runs one at a time.
- steps[] records each round: timings, attempts, each tool's outcome and
  size, usage and cost; costUsd totals them.
- Budgets (maxInputTokens, maxCostUsd, maxDurationMs) end the turn "budget",
  never between a tool_use and its result.
- maxRoundRetries retries a retryable round failure as the provider's
  retry-after asks; roundTimeoutMs abandons a round a provider accepted and
  never answered, as a retryable failure.

An unset object port reaches the task as {}, and an empty schema counts as
none.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

@mergestorm-vortex mergestorm-vortex Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review progress ████████░░ 4/5 files

Comment — found 2 issue(s) at 51ab60f.

Actionable comments posted: 2

🤖 Prompt for AI agents
Verify each finding against current code. Fix only still-valid concrete bugs, skip the
rest with a brief reason, keep changes minimal, and validate. Skip Decision required,
policy forks, and "consider X" alternatives. Do not add new features, refactors, or
architecture beyond the fix; prefer the smallest diff.

Findings to address:
1. In `@packages/ai/src/task/AgentTask.ts` (line 861, Important):
   `if (timedOut && !context.signal.aborted) { throw new RetryableJobError(...) }`
   
   The timeout only becomes a retryable failure if `runRound()` actually rejects after `abandon()` fires. `abandon` is `roundAbort.abort()` for a host `AGENT_ROUND_RUNNER` and `turn.abort()` for the owned `ToolCallingTask` — both are cooperative, and the doc comment on `streamRound` itself says "if the host runner ignores the signal, the round may hang despite the timer". When it does ignore it, `runRound()` never settles, so the `catch` never runs, `run` never settles, and `await run` at line 876 blocks forever: the turn hangs instead of being retried as `roundTimeoutMs` promises (and `maxDurationMs` cannot rescue it, since that is only checked at the top of a round). The test at AgentTaskTurnControls.test.ts:490 passes only because the mock run-fn rejects on `signal.abort`.
   
   Racing the round against the timer (reject with the `RetryableJobError` when the timer fires, rather than relying on the abandoned call to reject) makes the timeout hold for a runner that ignores the signal.

2. In `@packages/test/src/test/ai/AgentTaskTurnControls.test.ts` (line 312, Important):
   `new AgentTask().run({ model: MODEL, prompt: "?", tools: [...], toolConcurrency: 2 }, { registry })`
   
   This is the only test in the file that omits `approval: "never"`, so `approval` defaults to `"beyond-inference"` (AgentTask.ts:490). `toolCallNeedsApproval` (AgentToolExecution.ts:137-148) returns `true` for the `slow` tool only via its explicit `requiresApproval: true` — but the `fast` tool is a host function (`execute` is a function), so `mayRelax` is true and `mode === "never"` would have returned `false` for it. With the default mode, `fast` falls through to `getTaskConstructors(registry).get(backingTaskType(tool))`, which is undefined for a function tool, so it also returns `false`. The `width = 1` branch is therefore entered solely because of `slow`'s `requiresApproval: true`, and the assertion `overlap.peak === 1` would hold even if the `calls.some(...)` guard at AgentTask.ts:683-688 were deleted — the test cannot fail for the behavior it names.
   
   Add `approval: "never"` to the input (as every other test in the file does) so the round's width is decided by the approval guard rather than by the default mode, and the test actually pins the "a round with a call that needs a person runs one at a time" contract.

Try MergeStorm

try {
captured.output = await runRound();
} catch (err) {
if (timedOut && !context.signal.aborted) {

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue · Important

roundTimeoutMs does not retry when the round runner ignores the abort

if (timedOut && !context.signal.aborted) { throw new RetryableJobError(...) }

The timeout only becomes a retryable failure if runRound() actually rejects after abandon() fires. abandon is roundAbort.abort() for a host AGENT_ROUND_RUNNER and turn.abort() for the owned ToolCallingTask — both are cooperative, and the doc comment on streamRound itself says "if the host runner ignores the signal, the round may hang despite the timer". When it does ignore it, runRound() never settles, so the catch never runs, run never settles, and await run at line 876 blocks forever: the turn hangs instead of being retried as roundTimeoutMs promises (and maxDurationMs cannot rescue it, since that is only checked at the top of a round). The test at AgentTaskTurnControls.test.ts:490 passes only because the mock run-fn rejects on signal.abort.

Racing the round against the timer (reject with the RetryableJobError when the timer fires, rather than relying on the abandoned call to reject) makes the timeout hold for a runner that ignores the signal.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid concrete bugs, skip the
rest with a brief reason, keep changes minimal, and validate. Skip Decision required,
policy forks, and "consider X" alternatives. Do not add new features, refactors, or
architecture beyond the fix; prefer the smallest diff.

In `@packages/ai/src/task/AgentTask.ts` (line 861, Important):
`if (timedOut && !context.signal.aborted) { throw new RetryableJobError(...) }`

The timeout only becomes a retryable failure if `runRound()` actually rejects after `abandon()` fires. `abandon` is `roundAbort.abort()` for a host `AGENT_ROUND_RUNNER` and `turn.abort()` for the owned `ToolCallingTask` — both are cooperative, and the doc comment on `streamRound` itself says "if the host runner ignores the signal, the round may hang despite the timer". When it does ignore it, `runRound()` never settles, so the `catch` never runs, `run` never settles, and `await run` at line 876 blocks forever: the turn hangs instead of being retried as `roundTimeoutMs` promises (and `maxDurationMs` cannot rescue it, since that is only checked at the top of a round). The test at AgentTaskTurnControls.test.ts:490 passes only because the mock run-fn rejects on `signal.abort`.

Racing the round against the timer (reject with the `RetryableJobError` when the timer fires, rather than relying on the abandoned call to reject) makes the timeout hold for a runner that ignores the signal.

};
registry.registerInstance(HUMAN_CONNECTOR, approve);
const overlap = { active: 0, peak: 0 };
const output = await new AgentTask().run(

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue · Important

Approval test never exercises the approval path

new AgentTask().run({ model: MODEL, prompt: "?", tools: [...], toolConcurrency: 2 }, { registry })

This is the only test in the file that omits approval: "never", so approval defaults to "beyond-inference" (AgentTask.ts:490). toolCallNeedsApproval (AgentToolExecution.ts:137-148) returns true for the slow tool only via its explicit requiresApproval: true — but the fast tool is a host function (execute is a function), so mayRelax is true and mode === "never" would have returned false for it. With the default mode, fast falls through to getTaskConstructors(registry).get(backingTaskType(tool)), which is undefined for a function tool, so it also returns false. The width = 1 branch is therefore entered solely because of slow's requiresApproval: true, and the assertion overlap.peak === 1 would hold even if the calls.some(...) guard at AgentTask.ts:683-688 were deleted — the test cannot fail for the behavior it names.

Add approval: "never" to the input (as every other test in the file does) so the round's width is decided by the approval guard rather than by the default mode, and the test actually pins the "a round with a call that needs a person runs one at a time" contract.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid concrete bugs, skip the
rest with a brief reason, keep changes minimal, and validate. Skip Decision required,
policy forks, and "consider X" alternatives. Do not add new features, refactors, or
architecture beyond the fix; prefer the smallest diff.

In `@packages/test/src/test/ai/AgentTaskTurnControls.test.ts` (line 312, Important):
`new AgentTask().run({ model: MODEL, prompt: "?", tools: [...], toolConcurrency: 2 }, { registry })`

This is the only test in the file that omits `approval: "never"`, so `approval` defaults to `"beyond-inference"` (AgentTask.ts:490). `toolCallNeedsApproval` (AgentToolExecution.ts:137-148) returns `true` for the `slow` tool only via its explicit `requiresApproval: true` — but the `fast` tool is a host function (`execute` is a function), so `mayRelax` is true and `mode === "never"` would have returned `false` for it. With the default mode, `fast` falls through to `getTaskConstructors(registry).get(backingTaskType(tool))`, which is undefined for a function tool, so it also returns `false`. The `width = 1` branch is therefore entered solely because of `slow`'s `requiresApproval: true`, and the assertion `overlap.peak === 1` would hold even if the `calls.some(...)` guard at AgentTask.ts:683-688 were deleted — the test cannot fail for the behavior it names.

Add `approval: "never"` to the input (as every other test in the file does) so the round's width is decided by the approval guard rather than by the default mode, and the test actually pins the "a round with a call that needs a person runs one at a time" contract.

…t out of turbo agent guidance

DeepSeek's v4-pro card still carried the retired 16:30-00:30 discount; every
model is now billed at half rate outside 01:00-04:00 and 06:00-10:00 UTC, as
the published table states. Checked against billing: 2,933 deepseek-flash
calls priced at $6.07 on this card, against $6.08 off the account balance.

turbo 2.11 rewrites AGENTS.md with a guidance block when it detects an AI
agent, which replaces the symlink to .claude/CLAUDE.md with a regular file.
agentGuidance: false turns that off.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

@mergestorm-vortex mergestorm-vortex Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review progress ██████████ 9/9 files

Comment — found 1 issue(s) at a8abd4e.

Actionable comment posted: 1

🤖 Prompt for AI agents
Verify each finding against current code. Fix only still-valid concrete bugs, skip the
rest with a brief reason, keep changes minimal, and validate. Skip Decision required,
policy forks, and "consider X" alternatives. Do not add new features, refactors, or
architecture beyond the fix; prefer the smallest diff.

Findings to address:
1. In `@providers/deepseek/src/ai/common/DeepSeek_Pricing.ts` (line 24, Important):
   `const DEEPSEEK_OFF_PEAK: ModelTimingTier[] = offPeak({ input: 0.66, output: 1.98, cached: 0.022 });`
   
   Before this change `DEEPSEEK_OFF_PEAK` (used by `DEEPSEEK_PRO`) was the single window `16:30–00:30`; routing it through `offPeak()` makes the Pro card carry the Flash windows instead — `10:00–01:00` and `04:00–06:00` — i.e. ~15 discounted hours/day instead of 8.
   
   The new unit test covers `deepseek-v4-pro`, so this repricing is presumably intended, but `examples/eval/src/test/models.test.ts:23-34` (`"prices the DeepSeek V4 Pro 0813 GA snapshot at the published cache-miss rates"`) still asserts the old value via `resolveModelConfig("deepseek-v4-pro-0813", "extract").pricing` (which reads `getDeepSeekModelPricing`, `examples/eval/src/models.ts:106`) and `toEqual`s `timingTiers: [{ start: "16:30", end: "00:30", ... }]`. It will now fail under the `eval` vitest project (`examples/eval/package.json` test script). Update that assertion to the two windows — or, if V4 Pro really keeps its own `16:30–00:30` window distinct from Flash, give it a separate tier list rather than sharing `offPeak()`.

Try MergeStorm

];
}

const DEEPSEEK_OFF_PEAK: ModelTimingTier[] = offPeak({ input: 0.66, output: 1.98, cached: 0.022 });

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue · Important

offPeak() changes the V4 Pro window, breaking an existing eval test

const DEEPSEEK_OFF_PEAK: ModelTimingTier[] = offPeak({ input: 0.66, output: 1.98, cached: 0.022 });

Before this change DEEPSEEK_OFF_PEAK (used by DEEPSEEK_PRO) was the single window 16:30–00:30; routing it through offPeak() makes the Pro card carry the Flash windows instead — 10:00–01:00 and 04:00–06:00 — i.e. ~15 discounted hours/day instead of 8.

The new unit test covers deepseek-v4-pro, so this repricing is presumably intended, but examples/eval/src/test/models.test.ts:23-34 ("prices the DeepSeek V4 Pro 0813 GA snapshot at the published cache-miss rates") still asserts the old value via resolveModelConfig("deepseek-v4-pro-0813", "extract").pricing (which reads getDeepSeekModelPricing, examples/eval/src/models.ts:106) and toEquals timingTiers: [{ start: "16:30", end: "00:30", ... }]. It will now fail under the eval vitest project (examples/eval/package.json test script). Update that assertion to the two windows — or, if V4 Pro really keeps its own 16:30–00:30 window distinct from Flash, give it a separate tier list rather than sharing offPeak().

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid concrete bugs, skip the
rest with a brief reason, keep changes minimal, and validate. Skip Decision required,
policy forks, and "consider X" alternatives. Do not add new features, refactors, or
architecture beyond the fix; prefer the smallest diff.

In `@providers/deepseek/src/ai/common/DeepSeek_Pricing.ts` (line 24, Important):
`const DEEPSEEK_OFF_PEAK: ModelTimingTier[] = offPeak({ input: 0.66, output: 1.98, cached: 0.022 });`

Before this change `DEEPSEEK_OFF_PEAK` (used by `DEEPSEEK_PRO`) was the single window `16:30–00:30`; routing it through `offPeak()` makes the Pro card carry the Flash windows instead — `10:00–01:00` and `04:00–06:00` — i.e. ~15 discounted hours/day instead of 8.

The new unit test covers `deepseek-v4-pro`, so this repricing is presumably intended, but `examples/eval/src/test/models.test.ts:23-34` (`"prices the DeepSeek V4 Pro 0813 GA snapshot at the published cache-miss rates"`) still asserts the old value via `resolveModelConfig("deepseek-v4-pro-0813", "extract").pricing` (which reads `getDeepSeekModelPricing`, `examples/eval/src/models.ts:106`) and `toEqual`s `timingTiers: [{ start: "16:30", end: "00:30", ... }]`. It will now fail under the `eval` vitest project (`examples/eval/package.json` test script). Update that assertion to the two windows — or, if V4 Pro really keeps its own `16:30–00:30` window distinct from Flash, give it a separate tier list rather than sharing `offPeak()`.

sroussey and others added 2 commits October 1, 2026 12:52
…e answer back

OpenAI reports a server failure mid-stream as a server_error event with no
HTTP status ("An error occurred while processing the request."), which fell
through to the unknown-error default and failed the turn outright; it is
classified retryable like a 5xx now.

A rejected submission now asks for the complete answer again. Told only what
to fix, a model resubmitted that part and dropped sections it had already
filled — thirteen underwriters to none.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The card now bills every DeepSeek model at half rate outside the weekday peak
hours (01:00-04:00, 06:00-10:00 UTC); this test still pinned the retired
16:30-00:30 window.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

@mergestorm-vortex mergestorm-vortex Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review progress ██████████ 12/12 files

Approve — reviewed files look good at a07972d.

✅ All clear — nothing to fix.


Try MergeStorm

@sroussey
sroussey merged commit 6fd708b into main Oct 1, 2026
32 of 33 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant