Repository navigation
fix(dspark): preserve target quantization for folded drafts - #36
Conversation
|
@jasl This is ready for review. The pre-run gate requires a |
|
Reviewed and reproduced on GB10 (SM121) 2-node TP=2. This is correct and I'm merging it. Thank you — you also found a red test of ours that our own gates were not running. The bug reproduces exactlyControlled A/B, same node pair, same runner, same spec shape, patch as the only variable:
Word for word the failure you described. One correction to the PR descriptionThe trigger is not "the draft is folded" — it is " With the implicit shape — The failure needs the explicit shape, This matters for anyone reading the commit later: your discriminator is a superset of the failing condition. That is fine — it is safe, it is derived from the same place that decides foldedness ( The test is not vacuousChecked directly rather than assumed — with your test but Exactly the right shape: only the folded parametrization fails, You found a test we had been shipping red
It went unnoticed because that file was not in our unit list at all. It is now ( One residual, not a blockerThe patch leaves So the patch trades a hard crash for a possible subtle inconsistency. That is the right trade, and empirically it does not bite: the patched explicit-model serve passes the #19 gate and generates coherently. I am not asking you to extend the fix to Scope for other users of this branch
|
|
Merged as Closing manually rather than by GitHub's auto-close: the branch had moved on (a 41-commit upstream merge plus a FlashInfer 0.6.16 bump landed while this was open), so it went in through Full review is in the comment above — the short version is that the fix is right, the test is non-vacuous, and the root cause is one step to the left of where the description puts it (the trigger is an explicitly-named speculative Thank you especially for the test rewrite. |
Summary
DSparkDraftModelcheckpointsRoot cause
load_dspark_modelnow replaces the copiedVllmConfig.quant_configwith the draft model quantization config. That is correct for standalone drafts, but DeepSeek V4 uses a folded draft whose weights live in the target checkpoint and whose model-specificDeepseekV4FP8Configselects FP4/MXFP4 experts.The folded draft ModelConfig resolves a generic FP8 config. Applying it to the whole draft registers
w13_weight_scale_inv/w2_weight_scale_inv, while the folded target checkpoint supplies FP4 expert scales and the DSpark loader maps them tow13_weight_scale/w2_weight_scale. Cold startup then fails with a missingw13_weight_scaleparameter.DSparkDraftModelis the explicit architecture marker assigned to this folded DeepSeek path, so use it to retain the copied target quantization config. Other DSpark architectures keep the new standalone-draft behavior.Tests
python -m pytest -q tests/v1/spec_decode/test_dspark.py -k load_dspark_model_shares_direct_draft_embedding_and_head(2 passed)