Skip to content

Keep the hostname lookup off the event loop, warm adopted stacked decision models, and temper decisions wherever they enter - #284

Merged
webdevtodayjason merged 2 commits into
mainfrom
fable/nonblocking-host-lookup
Sep 26, 2026
Merged

webdevtodayjason merged 2 commits into
mainfrom
fable/nonblocking-host-lookup

Conversation

@webdevtodayjason

@webdevtodayjason webdevtodayjason commented Sep 26, 2026 •

Copy link
Copy Markdown
Contributor

Three fixes found during the 0.5.32 roll. The first one took fleet routing down, and the third meant the calibration fix in 0.5.32 did not reach clients that call the master.

What changed

1. The hostname lookup behind /api/server/status no longer runs on the event loop (api/server_routes.py::_reachable_urls, _host_addrs)
_reachable_urls called socket.gethostbyname_ex(socket.gethostname()) inline in the handler. The Sparks don't have their hostname in /etc/hosts, and since 2026-09-24 the resolver on each of them includes the NetBird (Brev) nameserver, which never answers. So the lookup took 20 s inside the container on Spark-1 and Spark-4. The fleet dashboard and a browser tab were both polling this route. After the 0.5.32 restart, the master's loop was blocked almost all the time (30 of 30 /api/health samples took 15 to 40 s). Its peers went stale and every peer model returned 404 model_not_found through the master.

The lookup now runs in the default executor, with only one in flight at a time. Its answer is cached for HOST_ADDRS_TTL_SECONDS (300). A request waits at most HOST_ADDRS_WAIT_SECONDS (1.0). If the lookup hasn't finished by then, the request gets the last known list, or no addresses the first time. A slow lookup still fills the cache when it completes. reachable_at keeps its shape.

The fleet was restored before this PR by adding 127.0.1.1 <hostname> to /etc/hosts on the four Sparks. This PR is what stops it happening on any node with a slow resolver.

2. Adopted stacked decision models get their grammar warm-up (models/api_routes.py::_warm_adopted_stacked)
The #277 warm-up is started from the bind wait, and from _await_primary_bind for an adopted primary. A stacked engine kept across a restart goes through neither. After the roll, both jebadiah adapters on Spark-4 reported warm: null. Startup now warms every adopted stacked instance that is serving. It skips the primary's port, which is already warmed, and any instance still starting, which will bind later.

3. A decision is tempered wherever the request enters (api/decide.py::fetch_peer_temperatures, handle_decide_calibration)
A model's temperatures.json is on the node that serves it. When /v1/decide or /v1/systemone entered on a node without a copy, the answer came back raw with calibration: {applied: false, temperatures: null}, while the same call sent directly to the serving node was tempered. The master is exactly such a node, and it is the endpoint clients use. Proven on 2026-09-26 with jebadiah-9b-v2: tempered on Spark-4, raw through Spark-1.

Now, when run_questions finds no table locally, it asks the node it is routing to. It maps each remote candidate's fabric IP to that node's announced web port and calls GET /api/decide/calibration?model=<id> with the fleet key. The answer is cached for PEER_TEMPERATURES_TTL_S (300 s), and it is cached whether or not the model has a table. A peer that does not answer (down, or on a release before this route) is remembered for PEER_TEMPERATURES_RETRY_S (30 s), so it cannot add a timeout to every request. None of this can fail a decision: the worst case is the old raw answer. "calibration": "raw" still opts out, and the block still lists the temperatures that were skipped. Both the peer's answer and the file are validated by the same filter, valid_temperatures.

Other blocking socket/DNS calls, checked: none are left on a request path.

  • tls/certs.py (getfqdn, gethostbyname_ex, gethostbyaddr) runs from the CLI only; the reverse lookup is already on a daemon thread with a timeout.
  • api/server.py::_fetch_latest_ghcr_tag, registry.fetch_openrouter_popular and the /v1/models probe in api_routes already run in executors.
  • replication.http_json is called from the CLI.
  • engine/backends/eugr.py::_local_ip_for_interface falls back to gethostbyname inside the opt-in eugr launch path. It was left as is.
  • socket.gethostname() alone does no DNS.

AGENTS.md: a new rule, "No resolver call runs on the event loop", next to the NVML one. The #277 warm-up bullet now names both adoption paths, and the #276 temperatures bullet says where a routing node gets the table. README: the calibration paragraph describes the peer lookup instead of "a routed model answers raw".

Tests

  • New tests/test_reachable_urls.py (3 tests):
    • a lookup that hangs does not stop the loop, and a second request does not start a second lookup;
    • the late answer is cached and expires;
    • a failing lookup still answers.
  • 2 new tests in tests/test_decide_warmup.py:
    • an adopted stacked decision model is warmed;
    • the primary's port and an instance still starting are not.
  • 3 new tests in tests/test_decide_calibration.py:
    • a routing node with no copy applies the serving node's temperatures on /v1/decide and /v1/systemone, honours raw, asks once with the fleet key, and caches the answer;
    • a peer without the route leaves the answer raw and is not asked again;
    • the route serves the node's own table, null for an unknown model, and 400 without model.
  • pytest tests/: 3232 passed, 2 skipped, 1 xfailed. ruff check ainode tests is clean.

Changelog text for the release PR

Fixed

  • A slow resolver can no longer freeze a node. /api/server/status looked up this host's name on the event loop, so a node whose hostname is not in /etc/hosts, with a nameserver that does not answer, stopped serving every route for 20 s per poll. On 2026-09-26 that made the master's peers go stale, and it stopped routing to them. The lookup now runs in a thread, is cached for five minutes, and a request waits at most one second for it.
  • A decision model kept across a restart on a stacked port warms its answer grammar again. Adoption skipped the Warm a decision engine on load: the first grammar-constrained request costs about 80 s #277 warm-up, so /api/status reported warm: null for it and the first real question paid the compile.
  • A decision is tempered the same wherever the request enters. A node with no copy of the model, usually the master, answered /v1/decide and /v1/systemone raw with temperatures: null, while the node serving the model applied its temperatures.json. The routing node now fetches the table from the serving node over the fleet key (GET /api/decide/calibration?model=<id>, new) and caches it for five minutes.

… decision models

/api/server/status resolved this host's name with gethostbyname_ex on the event
loop. With the hostname missing from /etc/hosts and a nameserver that does not
answer, each call took 20 s and froze every route on the node; on 2026-09-26 the
Spark-1 master's heartbeats went stale and it stopped routing to its peers. The
lookup now runs in a thread, one at a time, cached for five minutes, and a
request waits at most one second for it.

A stacked decision model adopted at restart never binds, so the #277 warm-up
never ran for it and /api/status reported warm: null. Adoption now warms every
adopted stacked instance that is serving.
A model's temperatures.json lives on the node that serves it, so /v1/decide and
/v1/systemone sent to a node without a copy (the master) answered raw with
temperatures: null while the serving node answered tempered. A routing node
with no local table now asks the node it routes to, GET
/api/decide/calibration over the fleet key, and caches the answer for five
minutes (a peer that does not answer for 30 s). A failure there never fails a
decision.
@webdevtodayjason webdevtodayjason changed the title Keep the hostname lookup off the event loop, and warm adopted stacked decision models Keep the hostname lookup off the event loop, warm adopted stacked decision models, and temper decisions wherever they enter Sep 26, 2026
@webdevtodayjason
webdevtodayjason merged commit 3bb0edf into main Sep 26, 2026
1 check passed
@webdevtodayjason
webdevtodayjason deleted the fable/nonblocking-host-lookup branch September 26, 2026 14:20
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant