Let the node that owns a decision model answer for it - #285
Merged
Merged
Conversation
/v1/decide and /v1/systemone for a model this node does not serve now post the whole request to the serving node's AINode under the fleet key, instead of calling its engine directly, so the owner applies its own temperatures.json and warm state wherever the request entered. Replicas are tried in routing order; a forwarded request is never forwarded again and does not count against the owner's rate limiter. With no owner answering, the node calls the engines itself as before. A model served locally takes the local path unchanged.
webdevtodayjason
force-pushed
the
fable/decide-forward-to-owner
branch
from
September 26, 2026 14:22
50e7a04 to
92cfaec
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
A decision request is now answered by the node that owns the model, so calibration and warm state apply wherever the request enters.
Why
A decision model's
temperatures.jsonand its grammar warm state live on the node that serves it./v1/decideand/v1/systemoneon a node that does not serve the model, the master above all, called the owner's engine directly. They answered raw withcalibration: {applied: false, temperatures: null}, while the same request sent to the owner was tempered. Seen on 2026-09-26 with frontier-infra/jebadiah-9b-v2: tempered on Spark-4, raw through Spark-1. The master is the endpoint clients use.What changed
decide.py::forward_to_owner, called by both routes once the model is resolved and candidates exist.headers=fleet_headers(app)andX-AINode-Forwarded-By: <node_id>. The body carries the resolvedmodel, so a request that named no model is answered for the one this node resolved.owner_web_ports: fabric IP to member, so Atlas's 3100 works). A node with two replicas stacked on it is asked once.X-AINode-Forwarded-Byis answered where it lands and is never forwarded again.ratelimit/middleware.py::is_forwarded_by_fleetskips the limiter for it. Every forwarded request arrives under the one fleet key, so counting them would turn the whole fleet into one client capped at 32 in flight, and the node the caller reached has already counted the caller. Only the fleet key stampsapi_key_id == "fleet", so a client cannot claim the exemption by sending the header./v1/decidebullet describe the rule.Relation to PR 284
PR 284's third commit (the routing node fetches the owner's table) is complementary. With this PR, the peer fetch only matters on the fallback path, when no owner's AINode answers. Both edit the same README paragraph and the same AGENTS.md bullet. Whichever merges second should take this PR's text for both and keep 284's sentence about the fetch as the fallback detail.
Rollout
The owner must run this release (or any release with the route;
/v1/systemoneexists since 0.5.31 and/v1/decidebefore that) and must share thecluster_secret, which the whole fleet does. An owner on an older release still answers the forwarded request, because the routes exist; it just applies temperatures only if it runs 0.5.32 or later. The whole fleet is on 0.5.32.Tests
New
tests/test_decide_forward.py(8 tests):/v1/decideand/v1/systemonewith the owner's calibration, andrawstill opts out (two real AINode apps, owner and entry);pytest tests/: 3235 passed, 2 skipped, 1 xfailed, including the fleet-key call-site guard.ruff check ainode testsis clean.Changelog text for the release PR
Fixed
/v1/decideand/v1/systemonefor a model the receiving node does not serve (the master, usually) are forwarded whole to the serving node's AINode over the fleet key, so the answer carries that node's temperatures and warm grammar instead ofcalibration: {applied: false, temperatures: null}. Replicas are tried in turn when one is down; a forwarded request is never forwarded again and does not count against the owner's rate limit; with no owner answering, the node calls the engines itself as before. A locally served model is unchanged.