Skip to content

fix: startup scheduler reconciliation replaces blind active_jobs DEL (#455, v3) - #492

Merged
stevenweaver merged 1 commit into
masterfrom
fix/v3-startup-reconcile-455
Sep 21, 2026
Merged

stevenweaver merged 1 commit into
masterfrom
fix/v3-startup-reconcile-455

Conversation

@stevenweaver

Copy link
Copy Markdown
Member

Closes #455.

v3-only port to master of the v2 lib/reconcile.js design that shipped on release/2.6.x in PR #462 (merge a937ec3), adapted to v3's architecture: redis@5 promise API via the shared lib/redis-client.js (TTL helpers from #454), async/await, and the v3 boot sequence in server.js.

What changes

server.js used to run a blind boot-time client.del("active_jobs"), orphaning every job still live on the scheduler and leaving their TTL-less hashes behind forever. Boot now takes a single scheduler snapshot (squeue/qstat) and reconciles:

  • live entries are kept exactly once — deduped via an atomic MULTI DEL+RPUSH rebuild, purging the historical duplicate-push backlog, with a defensive PERSIST against the Ensure proper cleanup of transient data in Redis (bound memory growth) #453 resurrect-same-id TTL trap;
  • already-terminal entries are reaped with the fix: bound Redis transient-data growth with terminal-only TTLs (#453) #454 retention TTLs (completed 24h, error/aborted/cancelled 1h);
  • dead-scheduler-id zombies are marked aborted (orphaned at server restart) with the terminal TTL;
  • a SCAN sweep (TYPE-guarded, no WRONGTYPE on result blobs/lists) applies the same treatment to in-flight-claiming hashes outside active_jobs (issue gap 1 — the pre-fix restarts deleted the list but not the hashes);
  • fail-open on any scheduler error: active_jobs is left untouched, never mass-aborted. v3-only hardening on top of the v2 module: a 15s exec timeout so a hung lister cannot wedge the boot gate, and submit_type: local resolves an empty snapshot without exec'ing a lister.

Both spawn surfaces — socket handler registration and the MCP server — are gated on reconciliation completing, so no new submission can race the snapshot or the sweep. reconcileActiveJobs never rejects; the gate depends on that contract (asserted by tests).

Verification (silverback, isolated Redis :6390 — never prod)

Unit tier — 34 new tests, all green; wired into test:ci (shimmed suites + static tripwire) and test:slurm-unit (real sbatch suite):

  • test/reconcile.js — real sbatch --wrap 'sleep 180' survivor: seeded [live, zombie, zombie, terminal] → entries=4 unique=3 kept=1 reaped=2, live kept exactly once with no TTL, zombie aborted + 1h TTL, terminal reaped + 24h TTL.
  • test/reconcile-edges.js — fail-open (nonzero exit), 15s hung-lister timeout, resolve-on-every-path contract, atomic dedup rebuild, no-hash/unparseable-torque_id edges, qsub qstat parsing branch, local branch.
  • test/reconcile-sweep.js — SCAN sweep semantics incl. TYPE guard (string/list keys untouched, no HGETALL ever issued on them), scheduler-side hashes, batch iteration (redis@5 scanIterator yields arrays).
  • test/regression/boot-reconcile-gate.js — source tripwire: no blind del("active_jobs"), reconcile precedes both spawn surfaces.
  • Full npm run test:ci green (315 + 151 + 7 passing) with the CI fixture config; npm run test:slurm-unit green (22 passing); npm run lint 0 errors; typecheck error count identical to master (no new errors); the 4 new test files are prettier-clean.

End-to-end restart survival (dev checkout, submit_type: slurm, real GARD on CD2.nex via the WebSocket path):

  1. Live survivor: GARD submitted, SLURM job RUNNING, active_jobs=[D], TTL(D)=-1 → kill -9 the server mid-run → restart → boot log reconcile : active_jobs entries=4 unique=3 kept=1 reaped=2 zombie hashes swept=1; D kept exactly once, hash untouched, TTL still -1, SLURM job kept RUNNING (+10s/+30s checks) and completed normally (97s wall) with a valid GARD.json. Re-engaging the same id post-restart finalized the lifecycle: status=completed, results in Redis, TTL 86396s; delivered via the job:status socket route and MCP HTTP get_job_results (full OAuth dance) — the exact lifecycle the blind DEL destroyed.
  2. Zombie + dedup: seeded duplicate running entry with dead scheduler id → removed from list, aborted, error orphaned at server restart, TTL 3600.
  3. Terminal reap: seeded completed entry → dropped from list, TTL 86400.
  4. Sweep + TYPE guard: seeded queued hash with dead scheduler id outside the list, plus a plain string key and a list key → zombie hashes swept=1, orphan aborted + terminal TTL, string/list untouched, no WRONGTYPE anywhere in the logs.
  5. Fail-open + timeout: boot with squeue exiting 2 → could not snapshot scheduler queue, leaving active_jobs untouched, list byte-identical (duplicate intact, nothing zombified), WS + MCP surfaces serving; boot with squeue hung 60s → boot completed after ~17s with the same fail-open warn.
  6. Gate order: in every boot log the reconcile summary precedes MCP listening; handler registration is code-gated on the same promise.
  7. Real-world zombie reap: a later restart with three genuinely dead jobs (SLURM COMPLETED/FAILED, watchers gone) reaped all three: entries=3 unique=3 kept=0 reaped=3, each aborted + TTL.
  8. MCP stdio spawn_analysis (FEL) immediately post-boot spawned cleanly through the gated surface; the job COMPLETED on SLURM with a valid FEL.json.

v2/release/2.6.x already shipped this in #462 and is untouched here.

…lind DEL (#455)

v3 port of the v2 lib/reconcile.js design that shipped on release/2.6.x in
PR #462 (merge a937ec3), adapted to v3's architecture: redis@5 promise API
via the shared lib/redis-client.js helpers (#454), async/await, and the v3
boot sequence in server.js.

server.js previously ran client.del("active_jobs") at boot, orphaning
every job still live on the scheduler and leaving their TTL-less hashes
behind forever (issue #455 gap 2). Now boot takes a single scheduler
snapshot (squeue/qstat) and reconciles:

- entries whose scheduler job is live are KEPT exactly once (deduped,
  atomic multi rebuild), with a defensive PERSIST against the #453
  resurrect-same-id TTL trap;
- already-terminal entries are dropped with retention TTLs (completed 24h,
  other terminals 1h);
- dead-scheduler-id zombies are marked aborted (orphaned-at-restart error)
  with the terminal TTL;
- a SCAN sweep (TYPE-guarded against WRONGTYPE) applies the same treatment
  to in-flight-claiming hashes outside active_jobs (issue #455 gap 1);
- FAIL OPEN on any scheduler error — active_jobs is left untouched, and a
  15s exec timeout keeps a hung lister from wedging boot (v3-only
  hardening; submit_type local resolves an empty snapshot without exec).

Both spawn surfaces — socket handler registration AND the MCP server — are
gated on reconciliation completing, so no new submission can race the
snapshot or the sweep. reconcileActiveJobs never rejects; the boot gate
depends on that contract.

Tests: test/reconcile.js (real sbatch survivor), test/reconcile-edges.js
(fail-open, timeout, dedup, qsub branch, local branch), test/
reconcile-sweep.js (SCAN sweep + TYPE guard), test/regression/
boot-reconcile-gate.js (static tripwire: no blind DEL, gate ordering).

Closes #455
@stevenweaver
stevenweaver merged commit 3f594d0 into master Sep 21, 2026
8 checks passed
@stevenweaver
stevenweaver deleted the fix/v3-startup-reconcile-455 branch September 21, 2026 19:16
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Reconcile queued/running job state against the scheduler on startup (bound the last Redis-leak class)

1 participant