Route to a member's LAN address when this node is off its fabric - #298
Merged
Merged
Conversation
A master on the R750 has no route to the Sparks' fabric IPs, so every request it proxied to a Spark timed out (#295). Each member's announcement already arrives from its LAN address (peer_ip). When this node's own fabric IP and the member's are on different /24s, routing now uses the LAN address first and keeps the fabric IP as a failover target. Nodes that share the fabric route exactly as before. Covers the proxy, /v1/models, decide owner lookup, chat model cards, the bench's placement lookup, pinned targets, cluster load/unload dispatch and the unload fan-out. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Vwc6zpV1YgK718opwKjuEU
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes #295.
With Atlas (the R750) as master tonight, every Spark model timed out through it: the proxy targets each member's announced fabric IP (10.100.0.x) and the R750 is not on that switch. Pollux's model worked because Pollux's fabric address is its LAN address.
Each member's announcement already arrives from its LAN address, stored as
peer_ip. Newmember_hosts(node, local_fabric)inainode/api/server.pyreturns the addresses to try, best first: when this node's own fabric IP and the member's are on different /24s, the LAN address goes first and the fabric IP stays as a failover target. A node with no fabric IP of its own known keeps the old order. Nodes that share the fabric (every Spark, Castor, Pollux today) route exactly as before._routing_candidatesnow lists every node's best address before any second guess, so a wrong guess costs a retry, not the request. Applied at every HTTP path that usedfabric_ip: the proxy,_routing_table(/v1/models, decide),fleet_instances, decide's owner and peer web-port lookups (they now match either address), the bench's_owning_node, pinned targets,/api/cluster/load|unloaddispatch and the unload fan-out. Engine traffic (NCCL, distributed launch) still uses the fabric and is untouched.Checked on Atlas before writing it: from the R750,
192.168.0.10:8000,:8001,192.168.0.11:8000and192.168.0.199:8000|8002all answer/v1/models200.Tests: six new cases in
tests/test_federation.py(off-fabric prefers LAN, on-fabric keeps fabric, best guesses before fallbacks, unknown own fabric keeps old order, LAN-only member now routable, decide owner lookup by LAN address).Changelog: Routing works from a node that is not on the cluster fabric (a master on the R750), using each member's LAN address (#295).
Need help on this PR? Tag
@codesmith-botwith what you need. Autofix is disabled.