Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
22 changes: 22 additions & 0 deletions .claude/CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -228,6 +228,28 @@ approval), so a host draws one lifecycle rather than one per way a call can end.
`approval: "never"` turns it off for a headless run, and with no connector registered such a
call is refused rather than run.

**A turn that must end in a typed answer** passes `outputSchema`: the loop adds a
`submit_answer` tool (`AGENT_SUBMIT_TOOL_NAME`) taking that schema, ends with
`stopReason: "submitted"` and the answer on `object` when a call passes it, and answers a
failing one as a correction to make, so the model resubmits instead of the host failing the
turn. A model that replies in text is reminded twice, then the turn ends `"answered"` with no
`object`. An unset object port reaches the task as `{}`, and an empty schema counts as none.
`checkSubmission` is where a host enforces what a schema cannot say — a section left empty
that the document plainly has, two figures that must agree: given an answer that passed the
schema and the turn so far, a returned reason goes back to the model as the submit call's
error. It is turned away at most twice, then accepted, so a check can never hold a turn; the
reasons are on `submissionRejections`.
The rest are controls a batch host needs: `toolConcurrency` (opt-in; a round holding a call
put to a person still runs one at a time, and results return in the order asked),
`maxRoundRetries` (default 2, for a `RetryableJobError`, waiting as the provider's
retry-after says within 1–60 s), `roundTimeoutMs` (a provider can accept a request and never answer;
past it the round is abandoned as a retryable failure, so the retries cover it), and budgets — `maxInputTokens`, `maxCostUsd` (refused
without a price card, since a budget it cannot measure never stops anything),
`maxDurationMs` — each ending the turn `"budget"`, never between a `tool_use` and its result.
Every round leaves an `AgentStep` on `steps` (timings, attempts, each tool's outcome and
size, usage, cost) and rides on the `snapshot` beside `messages`; `costUsd` totals them when
every round could be priced.

**The loop and the model call can live in different places.** A host whose tools are closures
(they draw on a screen, ask a person, read state only that process holds) but whose model is
reachable only through a backend binds `AGENT_ROUND_RUNNER` on the run's registry: every round is
Expand Down
4 changes: 3 additions & 1 deletion examples/eval/src/test/models.test.ts
Original file line number Diff line number Diff line change
Expand Up @@ -27,8 +27,10 @@ describe("resolveModelConfig", () => {
input: 1.32,
output: 3.96,
cached: 0.044,
// Half rate outside DeepSeek's weekday peak hours, 01:00-04:00 and 06:00-10:00 UTC.
timingTiers: [
{ start: "16:30", end: "00:30", pricing: { input: 0.66, output: 1.98, cached: 0.022 } },
{ start: "10:00", end: "01:00", pricing: { input: 0.66, output: 1.98, cached: 0.022 } },
{ start: "04:00", end: "06:00", pricing: { input: 0.66, output: 1.98, cached: 0.022 } },
],
});
});
Expand Down
26 changes: 26 additions & 0 deletions packages/ai/src/job/AiJob.ts
Original file line number Diff line number Diff line change
Expand Up @@ -202,6 +202,15 @@ export function classifyProviderError(err: unknown, taskType: string, provider:
);
}

// A provider's own server failure reported in the body or mid-stream rather
// than as an HTTP status — OpenAI's streaming "server_error" event, "An error
// occurred while processing the request." — is as transient as a 5xx.
if (isServerErrorReport(err, message)) {
return new RetryableJobError(
withJobErrorDiagnostics(`Server error from ${provider} for ${taskType}: ${message}`, err)
);
}

if (
message.includes("ECONNREFUSED") ||
message.includes("ECONNRESET") ||
Expand All @@ -227,6 +236,23 @@ export function classifyProviderError(err: unknown, taskType: string, provider:
);
}

/**
* Whether an error is a provider reporting its own server failure without an
* HTTP status: a `server_error` code or type on the error or the body it
* carries, or the message providers send with one.
*/
function isServerErrorReport(err: unknown, message: string): boolean {
const fields = (value: unknown): unknown[] => {
if (value === null || typeof value !== "object") return [];
const record = value as { code?: unknown; type?: unknown };
return [record.code, record.type];
};
const body =
err !== null && typeof err === "object" ? (err as { error?: unknown }).error : undefined;
if ([...fields(err), ...fields(body)].includes("server_error")) return true;
return /an error occurred while processing (the|your) request/i.test(message);
}

export class AiJob<
Input extends AiJobInput<TaskInput> = AiJobInput<TaskInput>,
Output extends TaskOutput = TaskOutput,
Expand Down
2 changes: 1 addition & 1 deletion packages/ai/src/model/ModelPricing.ts
Original file line number Diff line number Diff line change
Expand Up @@ -53,7 +53,7 @@ export interface ModelUsageTier {

/**
* A rate card that replaces the base one inside a daily clock window, which is
* how time-of-day discounts are published (DeepSeek's runs 16:30-00:30 UTC).
* how time-of-day discounts are published (DeepSeek's off-peak hours are an example).
*
* `start` and `end` are `HH:MM` in **UTC** — providers publish these windows in
* UTC and a local-time reading would silently misprice by the host's offset.
Expand Down
Loading
Loading