feat: qualify native model swapping on AMD - #4
Conversation
|
Final review completed on head 36de0a3. Release blockers found by the adversarial review were remediated: network authentication now fails closed, shutdown deadlines are bounded, image preprocessing rejects unsafe aspect ratios, publication artifacts are privacy-safe, benchmark semantics are corrected and schema-versioned, and qualification assertions remain active under Python optimization. Verification: 395 focused daemon/AMD safety tests passed; GitHub daemon-linux passed on the exact head; compile, shell syntax, diff whitespace, secret, and private-host scans passed. Remaining items are non-blocking CI-coverage enhancements recorded in project deferred work. |
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
The MiniMax-M3 backend has a stale call that crashes at runtime, the Q4_K regression fixture is invalid, and deployment safeguards do not reliably restore or validate hardware state.
Review effort: Balanced
Findings: 2
Open (10)
Require AMD HIP and validate the target GPU architecture · New Update M3 sparse caller for new backend resolver signature · New Report statistics availability only when collection is enabled · New Fail pre-start when the qualified GPU device is missing · New Restore the previously configured policy after benchmarking · New Recognize all RFC 1918 private IPv4 address ranges · New Initialize all packed bytes shared by decoder groups · New Rewrite the continuation as a complete direct-link sentence · New Remove redundant generated comments from the handler · New Remove mechanical comments from the configuration file · New
What changed in this PR
Expands AMD/ROCm support and native model swapping, alongside new model architectures, GGUF paths, multimodal handling, telemetry, deployment tooling, and qualification evidence.
Changes:
- Adds AMD/HIP compatibility, kernels, model loaders, and cache behavior.
- Extends daemon routing, readiness, telemetry, image input, and diagnostics.
- Adds deployment/benchmark tooling and broad regression coverage.
| File | Description |
|---|---|
tests/utils/test_rocm_runtime.py |
Tests ROCm capability gates. |
tests/server/test_openai_image_input.py |
Tests image marker ordering. |
tests/scheduler/test_hybrid_cache_manager.py |
Tests hybrid chunk alignment. |
tests/scheduler/test_abort_inflight_prefill.py |
Updates gate fixture. |
tests/reproduce/test_run_local_api_benchmark.py |
Tests benchmark safeguards. |
tests/reproduce/test_collect_host_manifest.py |
Tests manifest privacy. |
tests/models/test_qwen36_gdn_grouped_output.py |
Tests Qwen GDN layouts. |
tests/models/test_qwen35_gguf_expert_banks.py |
Tests mixed expert banks. |
tests/models/test_muse_glimmer.py |
Updates backend resolution test. |
tests/models/test_minimax_m3.py |
Updates M3 backend test. |
tests/models/test_glm_dsa.py |
Updates DSA input structure. |
tests/models/test_gguf_tokenizer_specials.py |
Tests GGUF special tokens. |
tests/models/test_gguf_q4_k.py |
Tests Q4_K decoding. |
tests/models/test_gemma4_mmproj_mapping.py |
Tests projector mapping. |
tests/models/qwen4_exp/ple_hf_ref.py |
Adds PLE reference runner. |
tests/models/qwen4_exp/conftest.py |
Adds Qwen runtime fixture. |
tests/models/qwen4_exp/__init__.py |
Creates test package. |
tests/kvcache/test_pool_sizing_surface.py |
Updates gate fixture. |
tests/kvcache/test_kv_cache_rebuild.py |
Updates gate fixture. |
tests/kernels/test_swiglu_clamp.py |
Tests clamped SwiGLU. |
tests/kernels/test_pinned_tensor.py |
Tests AOT variants and HIP behavior. |
tests/kernels/test_gguf_hip_flags.py |
Tests HIP GGUF flags. |
tests/kernels/test_gguf_hip_build_flags.py |
Tests HIP target selection. |
tests/engine/test_cache_budget.py |
Tests pinned-memory budget. |
tests/e2e/test_aime.py |
Adds MoE benchmark controls. |
tests/daemon/test_daemon_import_safety.py |
Extends daemon safety tests. |
tests/benchmarks/test_public_document_privacy.py |
Checks public-document privacy. |
SECURITY.md |
Adds vulnerability reporting guidance. |
scripts/publish-wheels.sh |
Publishes platform manifests. |
scripts/gmk-evo-x2/verify_rocm_kernel_cache.py |
Verifies no-JIT ROCm cache. |
scripts/gmk-evo-x2/stop_qwen_recovery_server.sh |
Stops isolated recovery server. |
scripts/gmk-evo-x2/run_qwen_dpm_policy_benchmark.sh |
Wraps DPM benchmark policy. |
scripts/gmk-evo-x2/freetoken-swap-service/README.md |
Documents service installation. |
scripts/gmk-evo-x2/freetoken-swap-service/freetoken-swap.service |
Adds user-service template. |
scripts/gmk-evo-x2/freetoken-swap-service/freetoken-power-profile.sh |
Applies power profile. |
scripts/gmk-evo-x2/deepseek_route_transfer_projection.py |
Calculates transfer bounds. |
scripts/gmk-evo-x2/benchmark_qwen_router.py |
Benchmarks Qwen router. |
scripts/gmk-evo-x2-rocprof-wheel-sdk.sh |
Selects wheel ROCm profiler SDK. |
README.md |
Documents AMD and swapping support. |
python/freetoken/version.py |
Bumps package version. |
python/freetoken/utils/hf.py |
Narrows shard downloads. |
python/freetoken/utils/arch.py |
Adds ROCm runtime detection. |
python/freetoken/utils/__init__.py |
Exports runtime detection. |
python/freetoken/tokenizer/server.py |
Forwards cache stats and images. |
python/freetoken/server/openai_api.py |
Handles images and token insertion. |
python/freetoken/server/generation.py |
Adds image-aware generation specs. |
python/freetoken/server/control_api.py |
Adds readiness endpoint. |
python/freetoken/server/args.py |
Adds PLE and metrics options. |
python/freetoken/server/api_server.py |
Adds cache-statistics API. |
python/freetoken/server/api_models.py |
Adds special-token control. |
python/freetoken/scheduler/prefill.py |
Aligns hybrid prefill chunks. |
python/freetoken/scheduler/cache.py |
Exposes chunk alignment. |
python/freetoken/moe/fused_q4_k_q6_k.py |
Adds mixed Q4/Q6 execution. |
python/freetoken/moe/fused_q4_k_q5_k.py |
Adds mixed Q4/Q5 execution. |
python/freetoken/moe/expert_banks.py |
Adds mixed GGUF banks. |
python/freetoken/moe/cpu_executor.py |
Registers clamped SwiGLU. |
python/freetoken/moe/benchbw.py |
Adds GLM workload profile. |
python/freetoken/models/register.py |
Registers new architectures. |
python/freetoken/models/qwen4_exp/moe.py |
Adds Qwen4 MoE implementation. |
python/freetoken/models/qwen4_exp/__init__.py |
Exports Qwen4 model support. |
python/freetoken/models/qwen3_5_moe/weight.py |
Pads FP8 scale banks. |
python/freetoken/models/qwen3_5_moe/moe.py |
Honors router normalization. |
python/freetoken/models/qwen3_5_moe/model.py |
Installs GGUF replacements. |
python/freetoken/models/qwen3_5_moe/attention.py |
Adds mixed GGUF projections. |
python/freetoken/models/qwen3_5_moe/__init__.py |
Exports GGUF support. |
python/freetoken/models/glm5_next/mlp.py |
Adds GLM clamped MLP. |
python/freetoken/models/glm5_next/__init__.py |
Exports GLM5 model. |
python/freetoken/models/glm_moe_dsa/attention.py |
Structures DSA inputs. |
python/freetoken/models/gguf/tokenizer.py |
Registers embedded special tokens. |
python/freetoken/models/gguf/config.py |
Maps Qwen GGUF architectures. |
python/freetoken/models/gemma4/model.py |
Adds vision diagnostics. |
python/freetoken/models/gemma4/attention.py |
Enables image bidirectionality. |
python/freetoken/models/blocks.py |
Adds host-forward context hook. |
python/freetoken/message/utils.py |
Serializes shaped CPU tensors. |
python/freetoken/message/tokenizer.py |
Adds image and stats messages. |
python/freetoken/message/frontend.py |
Adds cache-statistics reply. |
python/freetoken/message/backend.py |
Adds image and stats payloads. |
python/freetoken/message/__init__.py |
Exports new message types. |
python/freetoken/layers/gguf.py |
Enables Q4_K dispatch. |
python/freetoken/layers/activation.py |
Exposes clamped SwiGLU. |
python/freetoken/layers/__init__.py |
Exports new activation. |
python/freetoken/kvcache/base.py |
Accounts for compressed indices. |
python/freetoken/kernel/utils.py |
Filters CUDA-only HIP flags. |
python/freetoken/kernel/triton/qsa/__init__.py |
Exports QSA kernels. |
python/freetoken/kernel/triton/norm.py |
Guards CUDA-only PDL arguments. |
python/freetoken/kernel/triton/e4m3_compat.py |
Adds HIP FP8 emulation. |
python/freetoken/kernel/fla/utils.py |
Adds KDA/TMA configuration. |
python/freetoken/kernel/fla/__init__.py |
Exports KDA kernels. |
python/freetoken/kernel/csrc/pinned_tensor.cpp |
Uses HIP compatibility header. |
python/freetoken/kernel/csrc/jit/store.cu |
Accepts ROCm tensors. |
python/freetoken/kernel/csrc/jit/index.cu |
Accepts ROCm tensors. |
python/freetoken/kernel/csrc/hip_compat.h |
Adds CUDA/HIP API bridge. |
python/freetoken/kernel/csrc/gguf/ggml-common.h |
Adds HIP dot-product handling. |
python/freetoken/kernel/csrc/gguf/dispatch.h |
Adds HIP shuffle wrappers. |
python/freetoken/kernel/backend.py |
Disables CUDA binaries on ROCm. |
python/freetoken/kernel/aot.py |
Filters unsupported AOT shapes. |
python/freetoken/engine/config.py |
Adds PLE backend setting. |
python/freetoken/daemon/proxy.py |
Adds uncached readiness probes. |
python/freetoken/daemon/osproc.py |
Distinguishes unavailable PSS. |
python/freetoken/attention/base.py |
Adds QSA and multimodal metadata. |
python/freetoken/attention/__init__.py |
Registers QSA backend. |
pyproject.toml |
Expands ROCm dependencies. |
paper-draft/RELEASE_PAYLOAD.md |
Lists release artifacts. |
paper-draft/RELEASE_NOTES_v0.1.0-rc1.md |
Adds release notes. |
paper-draft/references.bib |
Adds bibliography. |
paper-draft/README.md |
Documents paper package. |
paper-draft/PUBLICATION_CHECKLIST.md |
Adds publication checklist. |
examples/freetoken-swap.yaml |
Adds swapping example. |
docs/upstream-qwen-paper-protocol.md |
Records upstream protocol. |
docs/gmktec-evo-x2-rocm-transfer-prototype.md |
Documents transfer results. |
docs/gmktec-evo-x2-real-qwen-nvfp4-route-matrix.md |
Documents route parity. |
docs/gmktec-evo-x2-persistent-overlap-prototype.md |
Documents overlap rejection. |
docs/gmktec-evo-x2-overlap-prototype.md |
Documents stream experiment. |
docs/gmktec-evo-x2-nvfp4-shape-prototype.md |
Documents NVFP4 shape test. |
docs/gmktec-evo-x2-nvfp4-marlin-warps8-rejection.md |
Records warp rejection. |
docs/gmktec-evo-x2-nvfp4-marlin-tile8-rejection.md |
Records tile rejection. |
docs/gmktec-evo-x2-nvfp4-marlin-stages2-rejection.md |
Records staging rejection. |
docs/gmktec-evo-x2-nvfp4-marlin-parity-test.md |
Records Marlin parity. |
docs/gmktec-evo-x2-mapped-host-gather.md |
Documents mapped-host gather. |
docs/gmktec-evo-x2-hip-gather-prototype.md |
Documents HIP gather. |
docs/gmktec-evo-x2-gemma4-q4-text-control-20260830.md |
Records Gemma control. |
docs/gmktec-evo-x2-fused-moe-prototype.md |
Documents fused MoE prototype. |
docs/gmktec-evo-x2-expert-block-prototype.md |
Documents block transfers. |
docs/gmktec-evo-x2-deepseek-route-transfer-projection-20260905.json |
Stores transfer projection. |
docs/gmktec-evo-x2-deepseek-expert-slice-metadata-20260905.json |
Stores expert geometry. |
docs/gmktec-evo-x2-deepseek-capacity-gate-result-20260905.json |
Stores capacity decision. |
docs/gmktec-evo-x2-batched-expert-transfer-prototype.md |
Documents batched transfers. |
docs/gmktec-evo-x2-amd-validation-program.md |
Defines validation program. |
docs/gmktec-evo-x2-amd-paper-protocol-ledger.md |
Tracks protocol gaps. |
CONTRIBUTING.md |
Updates community link. |
CLAUDE.md |
Points agents to instructions. |
CITATION.cff |
Adds citation metadata. |
benchmarks/swap/direct_model_canary.py |
Adds direct model canary. |
benchmarks/README.md |
Documents fixed cache sizing. |
benchmarks/gmk_evo_x2/README.md |
Documents API harness. |
benchmarks/gmk_evo_x2/quality_suite.json |
Adds quality cases. |
benchmarks/gmk_evo_x2/multiturn_state_suite.json |
Adds state-retention cases. |
benchmarks/bench_offload_cache_copy.py |
Adds Qwen NVFP4 profile. |
benchmarks/bench_decode_moe.py |
Adds KV-pool override. |
.zenodo.json |
Adds archival metadata. |
.gitignore |
Ignores generated HIP/PDF artifacts. |
.github/workflows/nightly-wheels.yml |
Uses platform manifest stamps. |
.github/workflows/issue-labels.yml |
Adds issue automation. |
.github/issue-labeler.yml |
Defines issue labels. |
.github/ISSUE_TEMPLATE/model_checkpoint.yml |
Adds checkpoint template. |
.github/ISSUE_TEMPLATE/feature_request.yml |
Adds feature template. |
.github/ISSUE_TEMPLATE/config.yml |
Configures issue routing. |
.gitattributes |
Marks PDFs binary. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
| if not torch.cuda.is_available(): | ||
| raise RuntimeError("this diagnostic requires GMKtek EVO-X2's native ROCm device") |
| @@ -234,7 +234,7 @@ def test_auto_backend_resolution(): | |||
| cfg = parse_config(_hf_config()) | |||
| required = _required_attn_types(cfg) | |||
| assert required == frozenset({AttnType.BSA}) | |||
| assert _resolve_auto_attention_backend(required, False) == "m3_sparse" | |||
| assert _resolve_auto_attention_backend(required) == "m3_sparse" | |||
| @app.get("/v1/cache/stats") | ||
| async def cache_stats(): | ||
| """Read accumulated MoE cache hit and miss counters without modifying the cache.""" | ||
|
|
||
| state = get_global_state() | ||
| if state.maintenance_state != "serving": | ||
| return JSONResponse({"error": "server is not serving"}, status_code=503) | ||
| try: | ||
| return await state.cache_stats() |
| gpu=/sys/bus/pci/devices/0000:64:00.0 | ||
| # What: continue only when the qualified GPU exists; why: startup should not write to an absent or different device. | ||
| if [[ -d "$gpu" ]]; then | ||
| # What: restore dynamic GPU clock scaling; why: automatic scaling cooperates with the thermal watchdog and avoids unsafe forced clocks. | ||
| echo auto | sudo -n tee "$gpu/power_dpm_force_performance_level" >/dev/null | ||
| # What: close the GPU-presence guard; why: the conditional must bound only the device-specific write. | ||
| fi |
| # Restore the safe default policy and append post-run telemetry. Each command | ||
| # is best-effort so a benchmark failure cannot conceal the restoration attempt. | ||
| restore_policy() { | ||
| sudo rocm-smi --setperflevel auto || true | ||
| rocm-smi --showperflevel | tee -a "${POLICY_LOG}" || true |
| root = Path(__file__).resolve().parents[2] | ||
| patterns = [ | ||
| re.compile(r"\bLAN-\d+\b", re.I), | ||
| re.compile(r"\b192\.168\.\d+\.\d+\b"), |
| raw[0, 8] = 4 | ||
| raw[0, 5] = 3 | ||
| raw[0, 9] = 4 | ||
| raw[0, 16:32] = 0xF1 # low nibble 1, high nibble 15 for the first 32-value group. |
| for scope and platform-specific boundaries, and [Reproducibility and independent | ||
| extension](docs/reproducibility.md) for the portable public evidence workflow. |
| # What: register GET /ready on the application router; why: clients reach ready's handler only through this method-and-path binding. | ||
| @app.get("/ready") | ||
| # What: define ready around the current object state; why: the registered API client call ready for ready and rely on this exact input and result contract. | ||
| async def ready(): |
| # What: begin unit metadata; why: systemd needs ordering and dependency declarations before starting inference. | ||
| [Unit] | ||
| # What: describe the service; why: operators must distinguish native FreeToken routing from the retired llama runner. | ||
| Description=FreeToken native AMD model router | ||
| # What: start after online networking; why: LAN binding must be ordered without creating a cycle through the watchdog default-target ordering. |



Summary
Verification
ruff checkpassesgit diff --checkScope and safety