Skip to content

feat: qualify native model swapping on AMD - #3

Closed
dbourdea wants to merge 569 commits into
mainfrom
codex/freetoken-swap-lan215
Closed

dbourdea wants to merge 569 commits into
mainfrom
codex/freetoken-swap-lan215

Conversation

@dbourdea

Copy link
Copy Markdown
Owner

Summary

  • harden native model routing lifecycle, cancellation, rollback, readiness, and adjacent port allocation
  • add Linux AMD process-memory telemetry and Qwen 3.5/3.6 GGUF MTP-aware geometry handling
  • expand the private qualification harness and regression coverage for live two-model AMD swapping

Verification

  • ruff check passes for all changed Python files
  • 183 focused tests pass
  • full repository run: 1947 passed, 81 skipped, 21 failed; all 21 failures reproduce unchanged at the parent commit on the same AMD host
  • live LAN-215 two-model qualification passed with restoration, direct/warm/cold/A-B-A routing, concurrency, cancellation, failed-switch rollback, restart re-adoption, TTL, metrics, logs, auth, and cleanup

Scope and safety

  • no llama-swap or llama.cpp source changes
  • raw host artifacts remain private
  • service deployment is handled separately from this source review

David added 30 commits September 4, 2026 12:12
FreeToken contributor and others added 28 commits September 14, 2026 19:28
@dbourdea dbourdea closed this Sep 24, 2026
@dbourdea
dbourdea deleted the codex/freetoken-swap-lan215 branch September 24, 2026 05:10
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant