Make a master that serves no model safe (Atlas control plane, phase 1) - #282
Merged
Merged
Conversation
Four fixes so AINode can run as a routing-only master on Atlas: - A request with no model field no longer falls back to config.model on a node that serves nothing; it answers 400 missing_model_field. A node that serves a primary keeps its default. - Unloading or ejecting the primary never promotes a survivor. The slot and its per-load overrides are cleared and stacked instances stay stacked (Whisper was promoted into Spark-4's config.model with another model's image and flags). - ainode doctor passes on a node that loads no model: "no model loaded" is INFO, and docker and engine backend findings are INFO there. - An empty account list is never replicated: the master's push, the worker's pull, the CLI push and /api/auth/users/sync all refuse it and log.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Phase 1 of the Atlas control-plane move: the code fixes that make a master which serves no model safe to run. Atlas (the R750) will join as a worker first and then take over from Spark-1 as a routing-only master, so each of these is a path that misbehaves today on a node with nothing loaded.
What changed
1. No
modelin the body no longer means "whatever config.model says" (api/server.py::own_model,proxy_to_vllm)A JSON body with no
model(or an empty one) used to route toconfig.model. On a node that serves nothing, that name is the dataclass default or whatever config.json was copied from, so the caller got a 404 or 502 about a model they never asked for. Now the fallback only applies when the node serves a primary (app["engine"]set andconfig.modelnamed). Otherwise the answer is a 400 withcode: missing_model_field,param: model, and nothing is forwarded. A node that serves a primary behaves exactly as before. The multipart audio paths already did this.2. Losing the primary never promotes a survivor (
models/api_routes.py::release_primary, used by/api/models/unloadand the server view eject)Unloading the primary used to repoint
app["engine"]andconfig.modelatsurvivors[0]. On Spark-4 that made Whisper (stacked on :8001) the node'sconfig.modelwhile config.json still held the old primary's image, vLLM flags and memory share. The node then advertised Whisper on :8000, where nothing served it, and the next boot would have launched Whisper with another model's engine parameters next to the manifest's own copy. Now the slot is cleared, the per-load overrides reset to NodeConfig defaults, and a stacked instance stays stacked on its own port. Eject now clears the overrides as well.3.
ainode doctorpasses on a node that loads no model (cli/doctor.py::engine_expected,as_no_model_info)"No model loaded" (
config.model,port.engine) is now INFO instead of OK. When nothing is pinned, no stacked instance is ininstances.jsonand there is nodistributed.json, the engine-launch checks (docker.daemon,config.engine_backendand its image) report INFO with the reason and keep their fix line. So a bare-metal routing-only node whose GPU belongs to something else (Atlas's A40 runs llama.cpp) no longer fails on docker or on an eugr backend without vllm.data.status_with_a_modelrecords what the finding would be on a node that does load a model. A node that pins a model still FAILs exactly as before.4. An empty account list is never replicated (
auth/replication.py,auth/session_routes.py::handle_sync_users)import_usersreplaces the whole list, so a master whoseusers.jsonis empty would have logged every dashboard user out of the fleet on its first push. There are now four guards, each of which refuses the empty list and logs the fix (copyusers.jsonfirst). On the master's push and the worker's pull the warning is logged once, not every 60 s tick:ainode auth user removeof the last account) sends nothing and prints the reason;/api/auth/users/syncanswers 409empty_account_list, which covers a sender still on an older release.Removing the last account is now a per-node act.
AGENTS.md records the four invariants on their existing bullets (proxy, stacked/primary, accounts, doctor).
Tests
tests/test_proxy_no_model.py(10),tests/test_primary_release.py(5),tests/test_doctor_no_model.py(8), 4 intests/test_auth_replication.py, 1 intests/test_session_routes.py. That is 28 new tests. 24 of them fail againstmain. The other 4 guard behaviour that must not change: a serving node keeps its default, and a named model still routes.test_no_model_pinned_is_infoand the idle port shape now expect INFO.pytest tests/gives 3227 passed, 2 skipped, 1 xfailed.ruff check ainode testsis clean.Not in this PR
gpu_memory_utilizationand does not become the primary even when it lands on the free api_port, so/api/statusshows no primary until the stack is empty. That behaviour is unchanged here. Making a load that takes the primary port become the primary would also change stacked admission control, so it is left as its own change.fabric.interfacewhen a copied config names acluster_interfacewith no address. Atlas's config should setcluster_interfaceto empty (or its own NIC) rather than copying Spark-1's.Changelog
modelon a node that serves no model of its own now gets a 400missing_model_fieldnaming the field, instead of being routed to a model the node does not serve. Nodes that serve a primary keep their default.config.modelwith the old primary's engine parameters. The primary slot and its overrides are cleared, and stacked instances stay on their own ports./api/auth/users/sync(409) all refuse it and log how to fix the store.