Skip to content

feat: qualify native model swapping on AMD - #4

Merged
dbourdea merged 574 commits into
mainfrom
codex/freetoken-swap-amd
Sep 30, 2026
Merged

dbourdea merged 574 commits into
mainfrom
codex/freetoken-swap-amd

Conversation

@dbourdea

@dbourdea dbourdea commented Sep 24, 2026 •

Copy link
Copy Markdown
Owner

Summary

  • harden native model routing lifecycle, cancellation, rollback, readiness, and adjacent port allocation
  • add Linux AMD process-memory telemetry and Qwen 3.5/3.6 GGUF MTP-aware geometry handling
  • expand the private two-model qualification harness and regression coverage
  • add sanitized GMKtek EVO-X2 user-service templates and the final qualification audit

Verification

  • changed-file ruff check passes
  • 183 focused daemon/router/metrics/qualification/Qwen tests pass
  • full repository run: 1947 passed, 81 skipped, 21 failed; all 21 failures reproduce unchanged at the parent commit on the same AMD environment
  • private two-model qualification passed with clean restoration and cleanup across direct/warm/cold/A-B-A routing, concurrency, cancellation, failed-switch rollback, restart re-adoption, reload conflict, persistence capacity, TTL, metrics, logs, auth, and positive owned-process AMD memory
  • the production user service passed authenticated health, exact resident identity, deterministic completion, LAN reachability, and a real unattended reboot/startup test
  • service/catalog/power-profile templates pass TOML parsing, Bash syntax, systemd verification, privacy review, and git diff --check

Scope and safety

  • no llama-swap or llama.cpp source changes
  • raw host artifacts, credentials, prompts, responses, and identifiers remain private
  • unsupported backend modalities and deferred MCP/Tailcat/peer surfaces remain explicit rather than being advertised as parity
  • this pull request remains draft and is not merged

David added 30 commits September 4, 2026 12:29
@dbourdea

Copy link
Copy Markdown
Owner Author

Final review completed on head 36de0a3. Release blockers found by the adversarial review were remediated: network authentication now fails closed, shutdown deadlines are bounded, image preprocessing rejects unsafe aspect ratios, publication artifacts are privacy-safe, benchmark semantics are corrected and schema-versioned, and qualification assertions remain active under Python optimization. Verification: 395 focused daemon/AMD safety tests passed; GitHub daemon-linux passed on the exact head; compile, shell syntax, diff whitespace, secret, and private-host scans passed. Remaining items are non-blocking CI-coverage enhancements recorded in project deferred work.

@dbourdea
dbourdea marked this pull request as ready for review September 30, 2026 03:03
Copilot AI balanced review requested due to automatic review settings September 30, 2026 03:03
@dbourdea
dbourdea merged commit da1af4f into main Sep 30, 2026
1 check passed
@dbourdea
dbourdea deleted the codex/freetoken-swap-amd branch September 30, 2026 03:03
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 30, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ⚠️ Failed 2026-09-30T03:03:56.544256Z 36de0a3 Draft marked ready
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

The MiniMax-M3 backend has a stale call that crashes at runtime, the Q4_K regression fixture is invalid, and deployment safeguards do not reliably restore or validate hardware state.

Review effort: Balanced
Findings: 2 High severity · 5 Medium severity · 3 Low severity

Open (10)
What changed in this PR

Expands AMD/ROCm support and native model swapping, alongside new model architectures, GGUF paths, multimodal handling, telemetry, deployment tooling, and qualification evidence.

Changes:

  • Adds AMD/HIP compatibility, kernels, model loaders, and cache behavior.
  • Extends daemon routing, readiness, telemetry, image input, and diagnostics.
  • Adds deployment/benchmark tooling and broad regression coverage.
File Description
tests/​utils/​test_rocm_runtime.py Tests ROCm capability gates.
tests/​server/​test_openai_image_input.py Tests image marker ordering.
tests/​scheduler/​test_hybrid_cache_manager.py Tests hybrid chunk alignment.
tests/​scheduler/​test_abort_inflight_prefill.py Updates gate fixture.
tests/​reproduce/​test_run_local_api_benchmark.py Tests benchmark safeguards.
tests/​reproduce/​test_collect_host_manifest.py Tests manifest privacy.
tests/​models/​test_qwen36_gdn_grouped_output.py Tests Qwen GDN layouts.
tests/​models/​test_qwen35_gguf_expert_banks.py Tests mixed expert banks.
tests/​models/​test_muse_glimmer.py Updates backend resolution test.
tests/​models/​test_minimax_m3.py Updates M3 backend test.
tests/​models/​test_glm_dsa.py Updates DSA input structure.
tests/​models/​test_gguf_tokenizer_specials.py Tests GGUF special tokens.
tests/​models/​test_gguf_q4_k.py Tests Q4_K decoding.
tests/​models/​test_gemma4_mmproj_mapping.py Tests projector mapping.
tests/​models/​qwen4_exp/​ple_hf_ref.py Adds PLE reference runner.
tests/​models/​qwen4_exp/​conftest.py Adds Qwen runtime fixture.
tests/​models/​qwen4_exp/​__init__.py Creates test package.
tests/​kvcache/​test_pool_sizing_surface.py Updates gate fixture.
tests/​kvcache/​test_kv_cache_rebuild.py Updates gate fixture.
tests/​kernels/​test_swiglu_clamp.py Tests clamped SwiGLU.
tests/​kernels/​test_pinned_tensor.py Tests AOT variants and HIP behavior.
tests/​kernels/​test_gguf_hip_flags.py Tests HIP GGUF flags.
tests/​kernels/​test_gguf_hip_build_flags.py Tests HIP target selection.
tests/​engine/​test_cache_budget.py Tests pinned-memory budget.
tests/​e2e/​test_aime.py Adds MoE benchmark controls.
tests/​daemon/​test_daemon_import_safety.py Extends daemon safety tests.
tests/​benchmarks/​test_public_document_privacy.py Checks public-document privacy.
SECURITY.md Adds vulnerability reporting guidance.
scripts/​publish-wheels.sh Publishes platform manifests.
scripts/​gmk-evo-x2/​verify_rocm_kernel_cache.py Verifies no-JIT ROCm cache.
scripts/​gmk-evo-x2/​stop_qwen_recovery_server.sh Stops isolated recovery server.
scripts/​gmk-evo-x2/​run_qwen_dpm_policy_benchmark.sh Wraps DPM benchmark policy.
scripts/​gmk-evo-x2/​freetoken-swap-service/​README.md Documents service installation.
scripts/​gmk-evo-x2/​freetoken-swap-service/​freetoken-swap.service Adds user-service template.
scripts/​gmk-evo-x2/​freetoken-swap-service/​freetoken-power-profile.sh Applies power profile.
scripts/​gmk-evo-x2/​deepseek_route_transfer_projection.py Calculates transfer bounds.
scripts/​gmk-evo-x2/​benchmark_qwen_router.py Benchmarks Qwen router.
scripts/​gmk-evo-x2-rocprof-wheel-sdk.sh Selects wheel ROCm profiler SDK.
README.md Documents AMD and swapping support.
python/​freetoken/​version.py Bumps package version.
python/​freetoken/​utils/​hf.py Narrows shard downloads.
python/​freetoken/​utils/​arch.py Adds ROCm runtime detection.
python/​freetoken/​utils/​__init__.py Exports runtime detection.
python/​freetoken/​tokenizer/​server.py Forwards cache stats and images.
python/​freetoken/​server/​openai_api.py Handles images and token insertion.
python/​freetoken/​server/​generation.py Adds image-aware generation specs.
python/​freetoken/​server/​control_api.py Adds readiness endpoint.
python/​freetoken/​server/​args.py Adds PLE and metrics options.
python/​freetoken/​server/​api_server.py Adds cache-statistics API.
python/​freetoken/​server/​api_models.py Adds special-token control.
python/​freetoken/​scheduler/​prefill.py Aligns hybrid prefill chunks.
python/​freetoken/​scheduler/​cache.py Exposes chunk alignment.
python/​freetoken/​moe/​fused_q4_k_q6_k.py Adds mixed Q4/Q6 execution.
python/​freetoken/​moe/​fused_q4_k_q5_k.py Adds mixed Q4/Q5 execution.
python/​freetoken/​moe/​expert_banks.py Adds mixed GGUF banks.
python/​freetoken/​moe/​cpu_executor.py Registers clamped SwiGLU.
python/​freetoken/​moe/​benchbw.py Adds GLM workload profile.
python/​freetoken/​models/​register.py Registers new architectures.
python/​freetoken/​models/​qwen4_exp/​moe.py Adds Qwen4 MoE implementation.
python/​freetoken/​models/​qwen4_exp/​__init__.py Exports Qwen4 model support.
python/​freetoken/​models/​qwen3_5_moe/​weight.py Pads FP8 scale banks.
python/​freetoken/​models/​qwen3_5_moe/​moe.py Honors router normalization.
python/​freetoken/​models/​qwen3_5_moe/​model.py Installs GGUF replacements.
python/​freetoken/​models/​qwen3_5_moe/​attention.py Adds mixed GGUF projections.
python/​freetoken/​models/​qwen3_5_moe/​__init__.py Exports GGUF support.
python/​freetoken/​models/​glm5_next/​mlp.py Adds GLM clamped MLP.
python/​freetoken/​models/​glm5_next/​__init__.py Exports GLM5 model.
python/​freetoken/​models/​glm_moe_dsa/​attention.py Structures DSA inputs.
python/​freetoken/​models/​gguf/​tokenizer.py Registers embedded special tokens.
python/​freetoken/​models/​gguf/​config.py Maps Qwen GGUF architectures.
python/​freetoken/​models/​gemma4/​model.py Adds vision diagnostics.
python/​freetoken/​models/​gemma4/​attention.py Enables image bidirectionality.
python/​freetoken/​models/​blocks.py Adds host-forward context hook.
python/​freetoken/​message/​utils.py Serializes shaped CPU tensors.
python/​freetoken/​message/​tokenizer.py Adds image and stats messages.
python/​freetoken/​message/​frontend.py Adds cache-statistics reply.
python/​freetoken/​message/​backend.py Adds image and stats payloads.
python/​freetoken/​message/​__init__.py Exports new message types.
python/​freetoken/​layers/​gguf.py Enables Q4_K dispatch.
python/​freetoken/​layers/​activation.py Exposes clamped SwiGLU.
python/​freetoken/​layers/​__init__.py Exports new activation.
python/​freetoken/​kvcache/​base.py Accounts for compressed indices.
python/​freetoken/​kernel/​utils.py Filters CUDA-only HIP flags.
python/​freetoken/​kernel/​triton/​qsa/​__init__.py Exports QSA kernels.
python/​freetoken/​kernel/​triton/​norm.py Guards CUDA-only PDL arguments.
python/​freetoken/​kernel/​triton/​e4m3_compat.py Adds HIP FP8 emulation.
python/​freetoken/​kernel/​fla/​utils.py Adds KDA/TMA configuration.
python/​freetoken/​kernel/​fla/​__init__.py Exports KDA kernels.
python/​freetoken/​kernel/​csrc/​pinned_tensor.cpp Uses HIP compatibility header.
python/​freetoken/​kernel/​csrc/​jit/​store.cu Accepts ROCm tensors.
python/​freetoken/​kernel/​csrc/​jit/​index.cu Accepts ROCm tensors.
python/​freetoken/​kernel/​csrc/​hip_compat.h Adds CUDA/HIP API bridge.
python/​freetoken/​kernel/​csrc/​gguf/​ggml-common.h Adds HIP dot-product handling.
python/​freetoken/​kernel/​csrc/​gguf/​dispatch.h Adds HIP shuffle wrappers.
python/​freetoken/​kernel/​backend.py Disables CUDA binaries on ROCm.
python/​freetoken/​kernel/​aot.py Filters unsupported AOT shapes.
python/​freetoken/​engine/​config.py Adds PLE backend setting.
python/​freetoken/​daemon/​proxy.py Adds uncached readiness probes.
python/​freetoken/​daemon/​osproc.py Distinguishes unavailable PSS.
python/​freetoken/​attention/​base.py Adds QSA and multimodal metadata.
python/​freetoken/​attention/​__init__.py Registers QSA backend.
pyproject.toml Expands ROCm dependencies.
paper-draft/​RELEASE_PAYLOAD.md Lists release artifacts.
paper-draft/​RELEASE_NOTES_v0.1.0-rc1.md Adds release notes.
paper-draft/​references.bib Adds bibliography.
paper-draft/​README.md Documents paper package.
paper-draft/​PUBLICATION_CHECKLIST.md Adds publication checklist.
examples/​freetoken-swap.yaml Adds swapping example.
docs/​upstream-qwen-paper-protocol.md Records upstream protocol.
docs/​gmktec-evo-x2-rocm-transfer-prototype.md Documents transfer results.
docs/​gmktec-evo-x2-real-qwen-nvfp4-route-matrix.md Documents route parity.
docs/​gmktec-evo-x2-persistent-overlap-prototype.md Documents overlap rejection.
docs/​gmktec-evo-x2-overlap-prototype.md Documents stream experiment.
docs/​gmktec-evo-x2-nvfp4-shape-prototype.md Documents NVFP4 shape test.
docs/​gmktec-evo-x2-nvfp4-marlin-warps8-rejection.md Records warp rejection.
docs/​gmktec-evo-x2-nvfp4-marlin-tile8-rejection.md Records tile rejection.
docs/​gmktec-evo-x2-nvfp4-marlin-stages2-rejection.md Records staging rejection.
docs/​gmktec-evo-x2-nvfp4-marlin-parity-test.md Records Marlin parity.
docs/​gmktec-evo-x2-mapped-host-gather.md Documents mapped-host gather.
docs/​gmktec-evo-x2-hip-gather-prototype.md Documents HIP gather.
docs/​gmktec-evo-x2-gemma4-q4-text-control-20260830.md Records Gemma control.
docs/​gmktec-evo-x2-fused-moe-prototype.md Documents fused MoE prototype.
docs/​gmktec-evo-x2-expert-block-prototype.md Documents block transfers.
docs/​gmktec-evo-x2-deepseek-route-transfer-projection-20260905.json Stores transfer projection.
docs/​gmktec-evo-x2-deepseek-expert-slice-metadata-20260905.json Stores expert geometry.
docs/​gmktec-evo-x2-deepseek-capacity-gate-result-20260905.json Stores capacity decision.
docs/​gmktec-evo-x2-batched-expert-transfer-prototype.md Documents batched transfers.
docs/​gmktec-evo-x2-amd-validation-program.md Defines validation program.
docs/​gmktec-evo-x2-amd-paper-protocol-ledger.md Tracks protocol gaps.
CONTRIBUTING.md Updates community link.
CLAUDE.md Points agents to instructions.
CITATION.cff Adds citation metadata.
benchmarks/​swap/​direct_model_canary.py Adds direct model canary.
benchmarks/​README.md Documents fixed cache sizing.
benchmarks/​gmk_evo_x2/​README.md Documents API harness.
benchmarks/​gmk_evo_x2/​quality_suite.json Adds quality cases.
benchmarks/​gmk_evo_x2/​multiturn_state_suite.json Adds state-retention cases.
benchmarks/​bench_offload_cache_copy.py Adds Qwen NVFP4 profile.
benchmarks/​bench_decode_moe.py Adds KV-pool override.
.zenodo.json Adds archival metadata.
.gitignore Ignores generated HIP/PDF artifacts.
.github/​workflows/​nightly-wheels.yml Uses platform manifest stamps.
.github/​workflows/​issue-labels.yml Adds issue automation.
.github/​issue-labeler.yml Defines issue labels.
.github/​ISSUE_TEMPLATE/​model_checkpoint.yml Adds checkpoint template.
.github/​ISSUE_TEMPLATE/​feature_request.yml Adds feature template.
.github/​ISSUE_TEMPLATE/​config.yml Configures issue routing.
.gitattributes Marks PDFs binary.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment on lines +63 to +64
if not torch.cuda.is_available():
raise RuntimeError("this diagnostic requires GMKtek EVO-X2's native ROCm device")
@@ -234,7 +234,7 @@ def test_auto_backend_resolution():
cfg = parse_config(_hf_config())
required = _required_attn_types(cfg)
assert required == frozenset({AttnType.BSA})
assert _resolve_auto_attention_backend(required, False) == "m3_sparse"
assert _resolve_auto_attention_backend(required) == "m3_sparse"
Comment on lines +842 to +850
@app.get("/v1/cache/stats")
async def cache_stats():
"""Read accumulated MoE cache hit and miss counters without modifying the cache."""

state = get_global_state()
if state.maintenance_state != "serving":
return JSONResponse({"error": "server is not serving"}, status_code=503)
try:
return await state.cache_stats()
Comment on lines +9 to +15
gpu=/sys/bus/pci/devices/0000:64:00.0
# What: continue only when the qualified GPU exists; why: startup should not write to an absent or different device.
if [[ -d "$gpu" ]]; then
# What: restore dynamic GPU clock scaling; why: automatic scaling cooperates with the thermal watchdog and avoids unsafe forced clocks.
echo auto | sudo -n tee "$gpu/power_dpm_force_performance_level" >/dev/null
# What: close the GPU-presence guard; why: the conditional must bound only the device-specific write.
fi
Comment on lines +40 to +44
# Restore the safe default policy and append post-run telemetry. Each command
# is best-effort so a benchmark failure cannot conceal the restoration attempt.
restore_policy() {
sudo rocm-smi --setperflevel auto || true
rocm-smi --showperflevel | tee -a "${POLICY_LOG}" || true
root = Path(__file__).resolve().parents[2]
patterns = [
re.compile(r"\bLAN-\d+\b", re.I),
re.compile(r"\b192\.168\.\d+\.\d+\b"),
raw[0, 8] = 4
raw[0, 5] = 3
raw[0, 9] = 4
raw[0, 16:32] = 0xF1 # low nibble 1, high nibble 15 for the first 32-value group.
Comment thread README.md
Comment on lines +67 to +68
for scope and platform-specific boundaries, and [Reproducibility and independent
extension](docs/reproducibility.md) for the portable public evidence workflow.
Comment on lines +61 to +64
# What: register GET /ready on the application router; why: clients reach ready's handler only through this method-and-path binding.
@app.get("/ready")
# What: define ready around the current object state; why: the registered API client call ready for ready and rely on this exact input and result contract.
async def ready():
Comment on lines +1 to +5
# What: begin unit metadata; why: systemd needs ordering and dependency declarations before starting inference.
[Unit]
# What: describe the service; why: operators must distinguish native FreeToken routing from the retired llama runner.
Description=FreeToken native AMD model router
# What: start after online networking; why: LAN binding must be ordered without creating a cycle through the watchdog default-target ordering.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants