Repository navigation
feat(ai): AgentTask turns that end in a checked answer, under budgets - #1001
Conversation
What a batch host reading documents needs from the turn loop, measured
against thousands of recorded agent sessions:
- outputSchema adds a submit_answer tool taking that schema. A passing call
ends the turn ("submitted", the answer on `object`); a failing one is
answered as a correction, so the model resubmits instead of the host
failing the turn. A text reply is reminded twice to submit.
- checkSubmission lets the host judge an answer the schema passed — a
section left empty that the document has, two figures that disagree — and
send it back with a reason, at most twice.
- toolConcurrency runs a round's calls at once, results in the order asked;
a round with a call put to a person still runs one at a time.
- steps[] records each round: timings, attempts, each tool's outcome and
size, usage and cost; costUsd totals them.
- Budgets (maxInputTokens, maxCostUsd, maxDurationMs) end the turn "budget",
never between a tool_use and its result.
- maxRoundRetries retries a retryable round failure as the provider's
retry-after asks; roundTimeoutMs abandons a round a provider accepted and
never answered, as a retryable failure.
An unset object port reaches the task as {}, and an empty schema counts as
none.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
There was a problem hiding this comment.
Review progress ████████░░ 4/5 files
Comment — found 2 issue(s) at 51ab60f.
Actionable comments posted: 2
🤖 Prompt for AI agents
Verify each finding against current code. Fix only still-valid concrete bugs, skip the
rest with a brief reason, keep changes minimal, and validate. Skip Decision required,
policy forks, and "consider X" alternatives. Do not add new features, refactors, or
architecture beyond the fix; prefer the smallest diff.
Findings to address:
1. In `@packages/ai/src/task/AgentTask.ts` (line 861, Important):
`if (timedOut && !context.signal.aborted) { throw new RetryableJobError(...) }`
The timeout only becomes a retryable failure if `runRound()` actually rejects after `abandon()` fires. `abandon` is `roundAbort.abort()` for a host `AGENT_ROUND_RUNNER` and `turn.abort()` for the owned `ToolCallingTask` — both are cooperative, and the doc comment on `streamRound` itself says "if the host runner ignores the signal, the round may hang despite the timer". When it does ignore it, `runRound()` never settles, so the `catch` never runs, `run` never settles, and `await run` at line 876 blocks forever: the turn hangs instead of being retried as `roundTimeoutMs` promises (and `maxDurationMs` cannot rescue it, since that is only checked at the top of a round). The test at AgentTaskTurnControls.test.ts:490 passes only because the mock run-fn rejects on `signal.abort`.
Racing the round against the timer (reject with the `RetryableJobError` when the timer fires, rather than relying on the abandoned call to reject) makes the timeout hold for a runner that ignores the signal.
2. In `@packages/test/src/test/ai/AgentTaskTurnControls.test.ts` (line 312, Important):
`new AgentTask().run({ model: MODEL, prompt: "?", tools: [...], toolConcurrency: 2 }, { registry })`
This is the only test in the file that omits `approval: "never"`, so `approval` defaults to `"beyond-inference"` (AgentTask.ts:490). `toolCallNeedsApproval` (AgentToolExecution.ts:137-148) returns `true` for the `slow` tool only via its explicit `requiresApproval: true` — but the `fast` tool is a host function (`execute` is a function), so `mayRelax` is true and `mode === "never"` would have returned `false` for it. With the default mode, `fast` falls through to `getTaskConstructors(registry).get(backingTaskType(tool))`, which is undefined for a function tool, so it also returns `false`. The `width = 1` branch is therefore entered solely because of `slow`'s `requiresApproval: true`, and the assertion `overlap.peak === 1` would hold even if the `calls.some(...)` guard at AgentTask.ts:683-688 were deleted — the test cannot fail for the behavior it names.
Add `approval: "never"` to the input (as every other test in the file does) so the round's width is decided by the approval guard rather than by the default mode, and the test actually pins the "a round with a call that needs a person runs one at a time" contract.
| try { | ||
| captured.output = await runRound(); | ||
| } catch (err) { | ||
| if (timedOut && !context.signal.aborted) { |
There was a problem hiding this comment.
roundTimeoutMs does not retry when the round runner ignores the abort
if (timedOut && !context.signal.aborted) { throw new RetryableJobError(...) }
The timeout only becomes a retryable failure if runRound() actually rejects after abandon() fires. abandon is roundAbort.abort() for a host AGENT_ROUND_RUNNER and turn.abort() for the owned ToolCallingTask — both are cooperative, and the doc comment on streamRound itself says "if the host runner ignores the signal, the round may hang despite the timer". When it does ignore it, runRound() never settles, so the catch never runs, run never settles, and await run at line 876 blocks forever: the turn hangs instead of being retried as roundTimeoutMs promises (and maxDurationMs cannot rescue it, since that is only checked at the top of a round). The test at AgentTaskTurnControls.test.ts:490 passes only because the mock run-fn rejects on signal.abort.
Racing the round against the timer (reject with the RetryableJobError when the timer fires, rather than relying on the abandoned call to reject) makes the timeout hold for a runner that ignores the signal.
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid concrete bugs, skip the
rest with a brief reason, keep changes minimal, and validate. Skip Decision required,
policy forks, and "consider X" alternatives. Do not add new features, refactors, or
architecture beyond the fix; prefer the smallest diff.
In `@packages/ai/src/task/AgentTask.ts` (line 861, Important):
`if (timedOut && !context.signal.aborted) { throw new RetryableJobError(...) }`
The timeout only becomes a retryable failure if `runRound()` actually rejects after `abandon()` fires. `abandon` is `roundAbort.abort()` for a host `AGENT_ROUND_RUNNER` and `turn.abort()` for the owned `ToolCallingTask` — both are cooperative, and the doc comment on `streamRound` itself says "if the host runner ignores the signal, the round may hang despite the timer". When it does ignore it, `runRound()` never settles, so the `catch` never runs, `run` never settles, and `await run` at line 876 blocks forever: the turn hangs instead of being retried as `roundTimeoutMs` promises (and `maxDurationMs` cannot rescue it, since that is only checked at the top of a round). The test at AgentTaskTurnControls.test.ts:490 passes only because the mock run-fn rejects on `signal.abort`.
Racing the round against the timer (reject with the `RetryableJobError` when the timer fires, rather than relying on the abandoned call to reject) makes the timeout hold for a runner that ignores the signal.
| }; | ||
| registry.registerInstance(HUMAN_CONNECTOR, approve); | ||
| const overlap = { active: 0, peak: 0 }; | ||
| const output = await new AgentTask().run( |
There was a problem hiding this comment.
Approval test never exercises the approval path
new AgentTask().run({ model: MODEL, prompt: "?", tools: [...], toolConcurrency: 2 }, { registry })
This is the only test in the file that omits approval: "never", so approval defaults to "beyond-inference" (AgentTask.ts:490). toolCallNeedsApproval (AgentToolExecution.ts:137-148) returns true for the slow tool only via its explicit requiresApproval: true — but the fast tool is a host function (execute is a function), so mayRelax is true and mode === "never" would have returned false for it. With the default mode, fast falls through to getTaskConstructors(registry).get(backingTaskType(tool)), which is undefined for a function tool, so it also returns false. The width = 1 branch is therefore entered solely because of slow's requiresApproval: true, and the assertion overlap.peak === 1 would hold even if the calls.some(...) guard at AgentTask.ts:683-688 were deleted — the test cannot fail for the behavior it names.
Add approval: "never" to the input (as every other test in the file does) so the round's width is decided by the approval guard rather than by the default mode, and the test actually pins the "a round with a call that needs a person runs one at a time" contract.
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid concrete bugs, skip the
rest with a brief reason, keep changes minimal, and validate. Skip Decision required,
policy forks, and "consider X" alternatives. Do not add new features, refactors, or
architecture beyond the fix; prefer the smallest diff.
In `@packages/test/src/test/ai/AgentTaskTurnControls.test.ts` (line 312, Important):
`new AgentTask().run({ model: MODEL, prompt: "?", tools: [...], toolConcurrency: 2 }, { registry })`
This is the only test in the file that omits `approval: "never"`, so `approval` defaults to `"beyond-inference"` (AgentTask.ts:490). `toolCallNeedsApproval` (AgentToolExecution.ts:137-148) returns `true` for the `slow` tool only via its explicit `requiresApproval: true` — but the `fast` tool is a host function (`execute` is a function), so `mayRelax` is true and `mode === "never"` would have returned `false` for it. With the default mode, `fast` falls through to `getTaskConstructors(registry).get(backingTaskType(tool))`, which is undefined for a function tool, so it also returns `false`. The `width = 1` branch is therefore entered solely because of `slow`'s `requiresApproval: true`, and the assertion `overlap.peak === 1` would hold even if the `calls.some(...)` guard at AgentTask.ts:683-688 were deleted — the test cannot fail for the behavior it names.
Add `approval: "never"` to the input (as every other test in the file does) so the round's width is decided by the approval guard rather than by the default mode, and the test actually pins the "a round with a call that needs a person runs one at a time" contract.
…t out of turbo agent guidance DeepSeek's v4-pro card still carried the retired 16:30-00:30 discount; every model is now billed at half rate outside 01:00-04:00 and 06:00-10:00 UTC, as the published table states. Checked against billing: 2,933 deepseek-flash calls priced at $6.07 on this card, against $6.08 off the account balance. turbo 2.11 rewrites AGENTS.md with a guidance block when it detects an AI agent, which replaces the symlink to .claude/CLAUDE.md with a regular file. agentGuidance: false turns that off. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
There was a problem hiding this comment.
Review progress ██████████ 9/9 files
Comment — found 1 issue(s) at a8abd4e.
Actionable comment posted: 1
🤖 Prompt for AI agents
Verify each finding against current code. Fix only still-valid concrete bugs, skip the
rest with a brief reason, keep changes minimal, and validate. Skip Decision required,
policy forks, and "consider X" alternatives. Do not add new features, refactors, or
architecture beyond the fix; prefer the smallest diff.
Findings to address:
1. In `@providers/deepseek/src/ai/common/DeepSeek_Pricing.ts` (line 24, Important):
`const DEEPSEEK_OFF_PEAK: ModelTimingTier[] = offPeak({ input: 0.66, output: 1.98, cached: 0.022 });`
Before this change `DEEPSEEK_OFF_PEAK` (used by `DEEPSEEK_PRO`) was the single window `16:30–00:30`; routing it through `offPeak()` makes the Pro card carry the Flash windows instead — `10:00–01:00` and `04:00–06:00` — i.e. ~15 discounted hours/day instead of 8.
The new unit test covers `deepseek-v4-pro`, so this repricing is presumably intended, but `examples/eval/src/test/models.test.ts:23-34` (`"prices the DeepSeek V4 Pro 0813 GA snapshot at the published cache-miss rates"`) still asserts the old value via `resolveModelConfig("deepseek-v4-pro-0813", "extract").pricing` (which reads `getDeepSeekModelPricing`, `examples/eval/src/models.ts:106`) and `toEqual`s `timingTiers: [{ start: "16:30", end: "00:30", ... }]`. It will now fail under the `eval` vitest project (`examples/eval/package.json` test script). Update that assertion to the two windows — or, if V4 Pro really keeps its own `16:30–00:30` window distinct from Flash, give it a separate tier list rather than sharing `offPeak()`.
| ]; | ||
| } | ||
|
|
||
| const DEEPSEEK_OFF_PEAK: ModelTimingTier[] = offPeak({ input: 0.66, output: 1.98, cached: 0.022 }); |
There was a problem hiding this comment.
offPeak() changes the V4 Pro window, breaking an existing eval test
const DEEPSEEK_OFF_PEAK: ModelTimingTier[] = offPeak({ input: 0.66, output: 1.98, cached: 0.022 });
Before this change DEEPSEEK_OFF_PEAK (used by DEEPSEEK_PRO) was the single window 16:30–00:30; routing it through offPeak() makes the Pro card carry the Flash windows instead — 10:00–01:00 and 04:00–06:00 — i.e. ~15 discounted hours/day instead of 8.
The new unit test covers deepseek-v4-pro, so this repricing is presumably intended, but examples/eval/src/test/models.test.ts:23-34 ("prices the DeepSeek V4 Pro 0813 GA snapshot at the published cache-miss rates") still asserts the old value via resolveModelConfig("deepseek-v4-pro-0813", "extract").pricing (which reads getDeepSeekModelPricing, examples/eval/src/models.ts:106) and toEquals timingTiers: [{ start: "16:30", end: "00:30", ... }]. It will now fail under the eval vitest project (examples/eval/package.json test script). Update that assertion to the two windows — or, if V4 Pro really keeps its own 16:30–00:30 window distinct from Flash, give it a separate tier list rather than sharing offPeak().
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid concrete bugs, skip the
rest with a brief reason, keep changes minimal, and validate. Skip Decision required,
policy forks, and "consider X" alternatives. Do not add new features, refactors, or
architecture beyond the fix; prefer the smallest diff.
In `@providers/deepseek/src/ai/common/DeepSeek_Pricing.ts` (line 24, Important):
`const DEEPSEEK_OFF_PEAK: ModelTimingTier[] = offPeak({ input: 0.66, output: 1.98, cached: 0.022 });`
Before this change `DEEPSEEK_OFF_PEAK` (used by `DEEPSEEK_PRO`) was the single window `16:30–00:30`; routing it through `offPeak()` makes the Pro card carry the Flash windows instead — `10:00–01:00` and `04:00–06:00` — i.e. ~15 discounted hours/day instead of 8.
The new unit test covers `deepseek-v4-pro`, so this repricing is presumably intended, but `examples/eval/src/test/models.test.ts:23-34` (`"prices the DeepSeek V4 Pro 0813 GA snapshot at the published cache-miss rates"`) still asserts the old value via `resolveModelConfig("deepseek-v4-pro-0813", "extract").pricing` (which reads `getDeepSeekModelPricing`, `examples/eval/src/models.ts:106`) and `toEqual`s `timingTiers: [{ start: "16:30", end: "00:30", ... }]`. It will now fail under the `eval` vitest project (`examples/eval/package.json` test script). Update that assertion to the two windows — or, if V4 Pro really keeps its own `16:30–00:30` window distinct from Flash, give it a separate tier list rather than sharing `offPeak()`.
…e answer back
OpenAI reports a server failure mid-stream as a server_error event with no
HTTP status ("An error occurred while processing the request."), which fell
through to the unknown-error default and failed the turn outright; it is
classified retryable like a 5xx now.
A rejected submission now asks for the complete answer again. Told only what
to fix, a model resubmitted that part and dropped sections it had already
filled — thirteen underwriters to none.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The card now bills every DeepSeek model at half rate outside the weekday peak hours (01:00-04:00, 06:00-10:00 UTC); this test still pinned the retired 16:30-00:30 window. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
What
AgentTaskgains what a batch host reading documents needs from the turn loop:outputSchemaadds asubmit_answertool taking that schema. A passing call ends the turn (stopReason: "submitted", the answer onobject); a failing call is answered as a correction, so the model resubmits instead of the host failing the turn. A text-only reply is reminded twice to submit.checkSubmission: a host function judges an answer the schema passed (a section left empty that the document has, two figures that disagree) and sends it back with a reason, at most twice. Reasons are onsubmissionRejections.toolConcurrency: a round's calls run at once, with results in the order asked. A round with a call that needs a person's approval still runs one at a time.steps[]: a per-round record of timings, attempts, each tool's outcome and size, usage and cost.costUsdtotals them. Steps also ride on thesnapshotevent.maxInputTokens,maxCostUsd,maxDurationMs) end the turn as"budget", never between atool_useand its result.maxCostUsdis refused for a model with no price card.maxRoundRetries(default 2) retries aRetryableJobErroras the provider's retry-after asks, within 1 to 60 s.roundTimeoutMsabandons a round that a provider accepted and never answered (seen live on OpenAI) and treats it as a retryable failure.An unset object port reaches the task as
{}; an emptyoutputSchemacounts as none.Why
embarc-data is replacing an external coding agent with this loop for reading SEC filings. Across ~2,600 recorded sessions and a 63-filing labelled set, these were the gaps:
Tests
packages/test/src/test/ai/AgentTaskTurnControls.test.ts, 21 tests.AgentTaskand a2a suites pass unchanged, 122 in all.build:types(47/47), lint, format-check.🤖 Generated with Claude Code