Skip to content

Make a master that serves no model safe (Atlas control plane, phase 1) - #282

Merged
webdevtodayjason merged 1 commit into
mainfrom
fable/atlas-control-plane-master
Sep 26, 2026
Merged

webdevtodayjason merged 1 commit into
mainfrom
fable/atlas-control-plane-master

Conversation

@webdevtodayjason

Copy link
Copy Markdown
Contributor

Phase 1 of the Atlas control-plane move: the code fixes that make a master which serves no model safe to run. Atlas (the R750) will join as a worker first and then take over from Spark-1 as a routing-only master, so each of these is a path that misbehaves today on a node with nothing loaded.

What changed

1. No model in the body no longer means "whatever config.model says" (api/server.py::own_model, proxy_to_vllm)
A JSON body with no model (or an empty one) used to route to config.model. On a node that serves nothing, that name is the dataclass default or whatever config.json was copied from, so the caller got a 404 or 502 about a model they never asked for. Now the fallback only applies when the node serves a primary (app["engine"] set and config.model named). Otherwise the answer is a 400 with code: missing_model_field, param: model, and nothing is forwarded. A node that serves a primary behaves exactly as before. The multipart audio paths already did this.

2. Losing the primary never promotes a survivor (models/api_routes.py::release_primary, used by /api/models/unload and the server view eject)
Unloading the primary used to repoint app["engine"] and config.model at survivors[0]. On Spark-4 that made Whisper (stacked on :8001) the node's config.model while config.json still held the old primary's image, vLLM flags and memory share. The node then advertised Whisper on :8000, where nothing served it, and the next boot would have launched Whisper with another model's engine parameters next to the manifest's own copy. Now the slot is cleared, the per-load overrides reset to NodeConfig defaults, and a stacked instance stays stacked on its own port. Eject now clears the overrides as well.

3. ainode doctor passes on a node that loads no model (cli/doctor.py::engine_expected, as_no_model_info)
"No model loaded" (config.model, port.engine) is now INFO instead of OK. When nothing is pinned, no stacked instance is in instances.json and there is no distributed.json, the engine-launch checks (docker.daemon, config.engine_backend and its image) report INFO with the reason and keep their fix line. So a bare-metal routing-only node whose GPU belongs to something else (Atlas's A40 runs llama.cpp) no longer fails on docker or on an eugr backend without vllm. data.status_with_a_model records what the finding would be on a node that does load a model. A node that pins a model still FAILs exactly as before.

4. An empty account list is never replicated (auth/replication.py, auth/session_routes.py::handle_sync_users)
import_users replaces the whole list, so a master whose users.json is empty would have logged every dashboard user out of the fleet on its first push. There are now four guards, each of which refuses the empty list and logs the fix (copy users.json first). On the master's push and the worker's pull the warning is logged once, not every 60 s tick:

  • the master's push sends nothing;
  • a worker's pull does not import an empty list and records no sync state;
  • the CLI push (ainode auth user remove of the last account) sends nothing and prints the reason;
  • /api/auth/users/sync answers 409 empty_account_list, which covers a sender still on an older release.

Removing the last account is now a per-node act.

AGENTS.md records the four invariants on their existing bullets (proxy, stacked/primary, accounts, doctor).

Tests

  • New: tests/test_proxy_no_model.py (10), tests/test_primary_release.py (5), tests/test_doctor_no_model.py (8), 4 in tests/test_auth_replication.py, 1 in tests/test_session_routes.py. That is 28 new tests. 24 of them fail against main. The other 4 guard behaviour that must not change: a serving node keeps its default, and a named model still routes.
  • Updated: test_no_model_pinned_is_info and the idle port shape now expect INFO.
  • Full suite: pytest tests/ gives 3227 passed, 2 skipped, 1 xfailed. ruff check ainode tests is clean.

Not in this PR

  • After the primary is released and stacked instances remain, the next load still takes the stacked path. It needs an explicit gpu_memory_utilization and does not become the primary even when it lands on the free api_port, so /api/status shows no primary until the stack is empty. That behaviour is unchanged here. Making a load that takes the primary port become the primary would also change stacked admission control, so it is left as its own change.
  • The doctor still FAILs fabric.interface when a copied config names a cluster_interface with no address. Atlas's config should set cluster_interface to empty (or its own NIC) rather than copying Spark-1's.

Changelog

  • Proxy: a request with no model on a node that serves no model of its own now gets a 400 missing_model_field naming the field, instead of being routed to a model the node does not serve. Nodes that serve a primary keep their default.
  • Instances: unloading or ejecting a node's primary no longer promotes another instance (for example Whisper) into config.model with the old primary's engine parameters. The primary slot and its overrides are cleared, and stacked instances stay on their own ports.
  • Doctor: a node that loads no model under AINode now passes. "No model loaded" is INFO, and docker and engine-backend findings are INFO on such a node.
  • Accounts: an empty account list is never replicated. The master's push, a worker's pull, the CLI push and /api/auth/users/sync (409) all refuse it and log how to fix the store.

Four fixes so AINode can run as a routing-only master on Atlas:

- A request with no model field no longer falls back to config.model on a
  node that serves nothing; it answers 400 missing_model_field. A node that
  serves a primary keeps its default.
- Unloading or ejecting the primary never promotes a survivor. The slot and
  its per-load overrides are cleared and stacked instances stay stacked
  (Whisper was promoted into Spark-4's config.model with another model's
  image and flags).
- ainode doctor passes on a node that loads no model: "no model loaded" is
  INFO, and docker and engine backend findings are INFO there.
- An empty account list is never replicated: the master's push, the
  worker's pull, the CLI push and /api/auth/users/sync all refuse it and log.
@webdevtodayjason
webdevtodayjason merged commit 83567ea into main Sep 26, 2026
1 check passed
@webdevtodayjason
webdevtodayjason deleted the fable/atlas-control-plane-master branch September 26, 2026 09:26
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant