Skip to content

Let the node that owns a decision model answer for it - #285

Merged
webdevtodayjason merged 2 commits into
mainfrom
fable/decide-forward-to-owner
Sep 26, 2026
Merged

webdevtodayjason merged 2 commits into
mainfrom
fable/decide-forward-to-owner

Conversation

@webdevtodayjason

Copy link
Copy Markdown
Contributor

A decision request is now answered by the node that owns the model, so calibration and warm state apply wherever the request enters.

Why

A decision model's temperatures.json and its grammar warm state live on the node that serves it. /v1/decide and /v1/systemone on a node that does not serve the model, the master above all, called the owner's engine directly. They answered raw with calibration: {applied: false, temperatures: null}, while the same request sent to the owner was tempered. Seen on 2026-09-26 with frontier-infra/jebadiah-9b-v2: tempered on Spark-4, raw through Spark-1. The master is the endpoint clients use.

What changed

  • decide.py::forward_to_owner, called by both routes once the model is resolved and candidates exist.
    • When no candidate is local, it posts the whole request to the serving node's AINode, on the same path, with headers=fleet_headers(app) and X-AINode-Forwarded-By: <node_id>. The body carries the resolved model, so a request that named no model is answered for the one this node resolved.
    • The owner's web port is the one it announces (owner_web_ports: fabric IP to member, so Atlas's 3100 works). A node with two replicas stacked on it is asked once.
  • Failover. Owners are tried in routing order. A transport error, a timeout, a 5xx (an owner whose engine is down answers 503) or a 401/403/404 (a key it does not share, or a release without the route) moves on to the next replica. Any other status is the owner's own answer, a 422 naming a field included, and is returned as it came.
  • Loops and the owner's rate limit.
    • A request carrying X-AINode-Forwarded-By is answered where it lands and is never forwarded again.
    • On the owner, ratelimit/middleware.py::is_forwarded_by_fleet skips the limiter for it. Every forwarded request arrives under the one fleet key, so counting them would turn the whole fleet into one client capped at 32 in flight, and the node the caller reached has already counted the caller. Only the fleet key stamps api_key_id == "fleet", so a client cannot claim the exemption by sending the header.
  • Fallback. When no owner answers, the receiving node calls the engines itself, exactly as before, and the answer is raw rather than missing.
  • The local path is unchanged. A model served on this node never forwards.
  • Docs. README's calibration paragraph and the AGENTS.md /v1/decide bullet describe the rule.

Relation to PR 284

PR 284's third commit (the routing node fetches the owner's table) is complementary. With this PR, the peer fetch only matters on the fallback path, when no owner's AINode answers. Both edit the same README paragraph and the same AGENTS.md bullet. Whichever merges second should take this PR's text for both and keep 284's sentence about the fetch as the fallback detail.

Rollout

The owner must run this release (or any release with the route; /v1/systemone exists since 0.5.31 and /v1/decide before that) and must share the cluster_secret, which the whole fleet does. An owner on an older release still answers the forwarded request, because the routes exist; it just applies temperatures only if it runs 0.5.32 or later. The whole fleet is on 0.5.32.

Tests

New tests/test_decide_forward.py (8 tests):

  • the master answers /v1/decide and /v1/systemone with the owner's calibration, and raw still opts out (two real AINode apps, owner and entry);
  • the owner's 422 comes back as it is;
  • a dead owner and then a 503 owner fail over to the healthy replica, which receives the fleet key, the forwarded-by header and the resolved model;
  • an owner without the route (404) falls back to the engines, raw;
  • a forwarded request is answered where it lands and never forwarded again;
  • a locally served model takes the local path;
  • owner ports come from each node's announcement;
  • only the fleet key skips the owner's rate limit.

pytest tests/: 3235 passed, 2 skipped, 1 xfailed, including the fleet-key call-site guard. ruff check ainode tests is clean.

Changelog text for the release PR

Fixed

  • A decision is answered by the node that owns the model, wherever the request enters. /v1/decide and /v1/systemone for a model the receiving node does not serve (the master, usually) are forwarded whole to the serving node's AINode over the fleet key, so the answer carries that node's temperatures and warm grammar instead of calibration: {applied: false, temperatures: null}. Replicas are tried in turn when one is down; a forwarded request is never forwarded again and does not count against the owner's rate limit; with no owner answering, the node calls the engines itself as before. A locally served model is unchanged.

/v1/decide and /v1/systemone for a model this node does not serve now post the
whole request to the serving node's AINode under the fleet key, instead of
calling its engine directly, so the owner applies its own temperatures.json and
warm state wherever the request entered. Replicas are tried in routing order; a
forwarded request is never forwarded again and does not count against the
owner's rate limiter. With no owner answering, the node calls the engines itself
as before. A model served locally takes the local path unchanged.
@webdevtodayjason
webdevtodayjason force-pushed the fable/decide-forward-to-owner branch from 50e7a04 to 92cfaec Compare September 26, 2026 14:22
@webdevtodayjason
webdevtodayjason merged commit 13ea48c into main Sep 26, 2026
1 check passed
@webdevtodayjason
webdevtodayjason deleted the fable/decide-forward-to-owner branch September 26, 2026 14:25
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant