Keep the hostname lookup off the event loop, warm adopted stacked decision models, and temper decisions wherever they enter - #284
Merged
Conversation
… decision models /api/server/status resolved this host's name with gethostbyname_ex on the event loop. With the hostname missing from /etc/hosts and a nameserver that does not answer, each call took 20 s and froze every route on the node; on 2026-09-26 the Spark-1 master's heartbeats went stale and it stopped routing to its peers. The lookup now runs in a thread, one at a time, cached for five minutes, and a request waits at most one second for it. A stacked decision model adopted at restart never binds, so the #277 warm-up never ran for it and /api/status reported warm: null. Adoption now warms every adopted stacked instance that is serving.
A model's temperatures.json lives on the node that serves it, so /v1/decide and /v1/systemone sent to a node without a copy (the master) answered raw with temperatures: null while the serving node answered tempered. A routing node with no local table now asks the node it routes to, GET /api/decide/calibration over the fleet key, and caches the answer for five minutes (a peer that does not answer for 30 s). A failure there never fails a decision.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Three fixes found during the 0.5.32 roll. The first one took fleet routing down, and the third meant the calibration fix in 0.5.32 did not reach clients that call the master.
What changed
1. The hostname lookup behind
/api/server/statusno longer runs on the event loop (api/server_routes.py::_reachable_urls,_host_addrs)_reachable_urlscalledsocket.gethostbyname_ex(socket.gethostname())inline in the handler. The Sparks don't have their hostname in /etc/hosts, and since 2026-09-24 the resolver on each of them includes the NetBird (Brev) nameserver, which never answers. So the lookup took 20 s inside the container on Spark-1 and Spark-4. The fleet dashboard and a browser tab were both polling this route. After the 0.5.32 restart, the master's loop was blocked almost all the time (30 of 30/api/healthsamples took 15 to 40 s). Its peers went stale and every peer model returned 404model_not_foundthrough the master.The lookup now runs in the default executor, with only one in flight at a time. Its answer is cached for
HOST_ADDRS_TTL_SECONDS(300). A request waits at mostHOST_ADDRS_WAIT_SECONDS(1.0). If the lookup hasn't finished by then, the request gets the last known list, or no addresses the first time. A slow lookup still fills the cache when it completes.reachable_atkeeps its shape.The fleet was restored before this PR by adding
127.0.1.1 <hostname>to /etc/hosts on the four Sparks. This PR is what stops it happening on any node with a slow resolver.2. Adopted stacked decision models get their grammar warm-up (
models/api_routes.py::_warm_adopted_stacked)The #277 warm-up is started from the bind wait, and from
_await_primary_bindfor an adopted primary. A stacked engine kept across a restart goes through neither. After the roll, both jebadiah adapters on Spark-4 reportedwarm: null. Startup now warms every adopted stacked instance that isserving. It skips the primary's port, which is already warmed, and any instance stillstarting, which will bind later.3. A decision is tempered wherever the request enters (
api/decide.py::fetch_peer_temperatures,handle_decide_calibration)A model's
temperatures.jsonis on the node that serves it. When/v1/decideor/v1/systemoneentered on a node without a copy, the answer came back raw withcalibration: {applied: false, temperatures: null}, while the same call sent directly to the serving node was tempered. The master is exactly such a node, and it is the endpoint clients use. Proven on 2026-09-26 with jebadiah-9b-v2: tempered on Spark-4, raw through Spark-1.Now, when
run_questionsfinds no table locally, it asks the node it is routing to. It maps each remote candidate's fabric IP to that node's announced web port and callsGET /api/decide/calibration?model=<id>with the fleet key. The answer is cached forPEER_TEMPERATURES_TTL_S(300 s), and it is cached whether or not the model has a table. A peer that does not answer (down, or on a release before this route) is remembered forPEER_TEMPERATURES_RETRY_S(30 s), so it cannot add a timeout to every request. None of this can fail a decision: the worst case is the old raw answer."calibration": "raw"still opts out, and the block still lists the temperatures that were skipped. Both the peer's answer and the file are validated by the same filter,valid_temperatures.Other blocking socket/DNS calls, checked: none are left on a request path.
tls/certs.py(getfqdn,gethostbyname_ex,gethostbyaddr) runs from the CLI only; the reverse lookup is already on a daemon thread with a timeout.api/server.py::_fetch_latest_ghcr_tag,registry.fetch_openrouter_popularand the/v1/modelsprobe inapi_routesalready run in executors.replication.http_jsonis called from the CLI.engine/backends/eugr.py::_local_ip_for_interfacefalls back togethostbynameinside the opt-in eugr launch path. It was left as is.socket.gethostname()alone does no DNS.AGENTS.md: a new rule, "No resolver call runs on the event loop", next to the NVML one. The #277 warm-up bullet now names both adoption paths, and the #276 temperatures bullet says where a routing node gets the table. README: the calibration paragraph describes the peer lookup instead of "a routed model answers raw".
Tests
tests/test_reachable_urls.py(3 tests):tests/test_decide_warmup.py:tests/test_decide_calibration.py:/v1/decideand/v1/systemone, honoursraw, asks once with the fleet key, and caches the answer;model.pytest tests/: 3232 passed, 2 skipped, 1 xfailed.ruff check ainode testsis clean.Changelog text for the release PR
Fixed
/api/server/statuslooked up this host's name on the event loop, so a node whose hostname is not in /etc/hosts, with a nameserver that does not answer, stopped serving every route for 20 s per poll. On 2026-09-26 that made the master's peers go stale, and it stopped routing to them. The lookup now runs in a thread, is cached for five minutes, and a request waits at most one second for it./api/statusreportedwarm: nullfor it and the first real question paid the compile./v1/decideand/v1/systemoneraw withtemperatures: null, while the node serving the model applied itstemperatures.json. The routing node now fetches the table from the serving node over the fleet key (GET /api/decide/calibration?model=<id>, new) and caches it for five minutes.