Conversation
…ct#55629) Signed-off-by: levius <2114377220@qq.com> Signed-off-by: Isotr0py <Isotr0py@outlook.com> Co-authored-by: Codex <codex@openai.com> Co-authored-by: Isotr0py <Isotr0py@outlook.com>
…ect#52156) Signed-off-by: Thomas Parnell <tpa@zurich.ibm.com> Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com> Co-authored-by: Claude <noreply@anthropic.com> Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
…n of EAGLE resume (vllm-project#53945) Signed-off-by: wzhao18 <wzhao18.sz@gmail.com> Signed-off-by: Adam Shaver <ashaver@nvidia.com> Signed-off-by: akshaver <168006157+akshaver@users.noreply.github.com> Co-authored-by: wzhao18 <wzhao18.sz@gmail.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-authored-by: roikoren755 <26850796+roikoren755@users.noreply.github.com>
Signed-off-by: mohit-sarvam <mohit@sarvam.ai> Co-authored-by: Codex <noreply@openai.com>
…oject#53379) Signed-off-by: Sherif Waly <sherif.waly@mistral.ai> Co-authored-by: Nicolò Lucchesi <nicolo.lucchesi@mistral.ai>
…-project#55774) Signed-off-by: yewentao256 <zhyanwentao@126.com>
Signed-off-by: linitra24 <renshuang.zhou@daocloud.io>
…t#55890) Signed-off-by: Canlin <canlinguosdu@gmail.com> Co-authored-by: OpenAI Codex <noreply@openai.com>
Signed-off-by: QiuChunshuo <qiuchunshuo@huawei.com>
Signed-off-by: yewentao256 <zhyanwentao@126.com>
…irection (vllm-project#55643) Signed-off-by: zjy0516 <riverclouds.zhu@qq.com> Co-authored-by: khluu <khluu000@gmail.com> Co-authored-by: jiahao <jxia77@terpmail.umd.edu> Co-authored-by: Kimi Code <noreply@moonshot.cn>
…es (vllm-project#55908) Signed-off-by: Raya Elena Solano <raya.solano@mbinf.de> Co-authored-by: mgoin <mgoin64@gmail.com>
Signed-off-by: Shiyang Chen <shiychen@nvidia.com> Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk>
vllm-project#54112) Signed-off-by: Shiksha Patel <shiksha.patel@amd.com>
…roject#54809) Signed-off-by: Roderick-Wu <roderickwu2003@gmail.com> Signed-off-by: Roderick Wu <roderickwu2003@gmail.com> Signed-off-by: mgoin <mgoin64@gmail.com> Co-authored-by: mgoin <mgoin64@gmail.com>
…project#53780) Signed-off-by: Matthew Bonanni <mbonanni@redhat.com> Co-authored-by: OpenAI Codex <codex@openai.com> Co-authored-by: OpenAI Codex <noreply@openai.com>
Signed-off-by: khluu <khluu000@gmail.com> Co-authored-by: Kimi Code <noreply@moonshot.cn> Co-authored-by: Codex <noreply@openai.com>
…the KV block LCM (vllm-project#53007) Signed-off-by: Bill Nell <bnell@redhat.com> Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com> Co-authored-by: Lucas Wilkinson <lwilkins@redhat.com> Co-authored-by: OpenAI Codex <codex@openai.com>
Signed-off-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
…ns (vllm-project#55780) Signed-off-by: Andreas Karatzas <akaratza@amd.com> Co-authored-by: OpenAI Codex <codex@openai.com>
Signed-off-by: Djordje Ramic <djoramic@amd.com> Co-authored-by: Kevin H. Luu <khluu000@gmail.com>
…llm-project#55223) Signed-off-by: sfeng33 <4florafeng@gmail.com> Signed-off-by: Flora Feng <4florafeng@gmail.com> Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com>
…llm-project#55223) Signed-off-by: sfeng33 <4florafeng@gmail.com> Signed-off-by: Flora Feng <4florafeng@gmail.com> Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com>
… was encountered` (vllm-project#55924) Signed-off-by: yewentao256 <zhyanwentao@126.com>
…llm-project#49104) Signed-off-by: cjackal <44624812+cjackal@users.noreply.github.com> Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com> Co-authored-by: Flora Feng <4florafeng@gmail.com>
Signed-off-by: khluu <khluu000@gmail.com>
Signed-off-by: lcskrishna <lollachaitanya@gmail.com>
Signed-off-by: Shiyang Chen <shiychen@nvidia.com> Signed-off-by: Isotr0py <Isotr0py@outlook.com> Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com> Co-authored-by: Isotr0py <Isotr0py@outlook.com>
…ject#56401) Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com>
vllm-project#55522) Signed-off-by: JartX <sagformas@epdcenter.es>
Signed-off-by: zhenwei-intel <zhenwei.liu@intel.com>
…54821) Signed-off-by: Clinton Thomas <1033162+KernelClint@users.noreply.github.com> Co-authored-by: Lucas Bourtoule <35483370+dhalf@users.noreply.github.com>
…pers (vllm-project#56594) Signed-off-by: LopezCastroRoberto <rocastro@redhat.com>
Signed-off-by: Artem Perevedentsev <aperevedents@nvidia.com>
…56332) Signed-off-by: Joe Cotant <joe@inferact.ai> Co-authored-by: Claude <noreply@anthropic.com>
…5799) Signed-off-by: yangzeyu <532183776@qq.com>
…ect#56061) Signed-off-by: NickLucche <nicolo.lucchesi@mistral.ai> Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com> Co-authored-by: Lucas Wilkinson <lwilkins@redhat.com> Co-authored-by: OpenAI Codex <codex@openai.com>
Signed-off-by: Ayushman Singh <40520701+ayush1399@users.noreply.github.com> Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Signed-off-by: Kevin Luu <51931015+khluu@users.noreply.github.com> Co-authored-by: OpenAI Codex <codex@openai.com>
Signed-off-by: Kevin Luu <51931015+khluu@users.noreply.github.com> Co-authored-by: OpenAI Codex <codex@openai.com>
…ing support (vllm-project#56122) Signed-off-by: Raphael Rialland <raphael.rialland@mistral.ai> Signed-off-by: Simon Veitner <sveitner@redhat.com> Signed-off-by: Tomas Ruiz <tomas.ruiz.te@gmail.com> Co-authored-by: OpenAI Codex <codex@openai.com> Co-authored-by: Simon Veitner <sveitner@redhat.com> Co-authored-by: Tomas Ruiz <tomas.ruiz.te@gmail.com>
Clear a failed cudaHostRegister status before later warmup kernels run, while preserving degraded unpinned operation when registration fails. Co-authored-by: OpenAI Codex <noreply@openai.com> Signed-off-by: alexbi29 <alexbi29@users.noreply.github.com>
Support adaptive DSpark verification with capture-stable metadata and preserve pre-norm confidence inputs. Respect disabled image limits when sizing packed-prefill metadata for language-model-only serving. Co-authored-by: OpenAI Codex <noreply@openai.com> Signed-off-by: alexbi29 <alexbi29@users.noreply.github.com>
Reconcile upstream V4.1, JIT warmup, watermarking, HiSparse, and KV-offload refactors while preserving the validated SM120 DS4, DSpark, compact-offload, and KV-accounting behavior. Co-authored-by: OpenAI Codex <noreply@openai.com> Signed-off-by: alexbi29 <alexbi29@users.noreply.github.com>
Preserve the compact-layout positional invariant while using upstream's selected host-group projection, so corrupt transported signatures still fail closed. Co-authored-by: OpenAI Codex <noreply@openai.com> Signed-off-by: alexbi29 <alexbi29@users.noreply.github.com>
Carry upstream's DeepSeek V4.1 model-type support into the fork's rank-consistent dummy-forward warmup path. Co-authored-by: OpenAI Codex <noreply@openai.com> Signed-off-by: alexbi29 <alexbi29@users.noreply.github.com>
Forward upstream replay and alignment metadata through the local MLA and HiSparse managers, preserve complete per-group hybrid tail blocks, and update affected synthetic fixtures. Co-authored-by: OpenAI Codex <noreply@openai.com> Signed-off-by: alexbi29 <alexbi29@users.noreply.github.com>
Author
|
2026-09-12 upstream refresh
|
Keep the Jasl DSpark layer compatible with upstream's explicit DeepseekV4MoE routing boundary and cover the constructor contract. Co-authored-by: OpenAI Codex <noreply@openai.com> Signed-off-by: alexbi29 <alexbi29@users.noreply.github.com>
Replace the removed upstream KV shard helper with the runner's configured DCP size so SM120 long-context decode kernels stay prewarmed. Co-authored-by: OpenAI Codex <noreply@openai.com> Signed-off-by: alexbi29 <alexbi29@users.noreply.github.com>
Author
|
2026-09-12 live validation and promotion update The refreshed branch is now at
Validation:
Only the llama-swap stanza's source, binary, and cache paths changed; all serving flags remain matched. The previous tree/config/cache are retained for rollback. Post-push CI is running. |
Adapt vllm-project#50737 to the current DS4 integration while preserving top-k and adaptive-verification behavior.
Adapt vllm-project#52187 to retain current allocation-derived page geometry and overlaid-buffer widening.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Merge
vllm-project/vllm:mainat537af2c3a4ba7462ddc9bc94ec7a4ea496da6d2eintojasl/vllm:codex/ds4-sm120-min-enableat9ad62027bc84ca0ccbcc40853179312de770220c.This is a draft integration PR, not a production promotion. The merge commit is
25900c381414d7d8bad3ca8e0f5d1782134a8b85, with both original histories retained. The merge required resolving 46 conflicted files, plus semantic overlaps in automatically merged files.Integration choices and fixes
KVCacheTensor.layers/layer_striderepresentation, including aliased cache groups. Use the resolved layout to distinguish packed from layer-major allocations. Add a per-layer-stride accounting regression.Duplicate-work check
Checked open PRs on
jasl/vllmimmediately before publishing. #44 is a focused SM12x MQA dispatch fix and #46 concerns CPU mmap registration; neither is an upstream-main integration PR. This PR does not cherry-pick either open PR. It is intended for this fork branch, not as another replacement forvllm-project/vllm#41834.Validation
Performed in a new worktree and independent virtual environment, not the production installation, on an RTX PRO 6000 Blackwell / SM120 host.
uv pip check --python .venv/bin/python: all installed packages compatible.git diff --checkpassed..venv/bin/python tools/generate_versions_json.py --check: Docker dependency metadata in sync.test_swap_blocks_batch.py,test_compact_transfer.py,test_compact_worker_spec.py,test_gpu_worker.py): 92 passed, 1 skipped.Representative commands
309-pass focused unit command
.venv/bin/python -m pytest -q --maxfail=15 \ tests/config/test_deepseek_v4_cudagraph_config.py \ tests/models/test_deepseek_v4_rope.py \ tests/models/test_deepseek_v4_mega_moe.py \ tests/models/test_deepseek_v4_fi_moe_ep.py \ tests/models/test_deepseek_v4_nvfp4_draft_routing.py \ tests/models/test_dspark_v2_speculator_hooks.py \ tests/models/test_dspark_shared_expert_pad.py \ tests/v1/spec_decode/test_dspark_config.py \ tests/v1/spec_decode/test_dspark_aux_layer_ids.py \ tests/transformers_utils/test_dspark_mla_config.py \ tests/transformers_utils/test_speculators_dspark_config.py \ tests/v1/core/test_ghost_block_guard.py \ tests/v1/core/test_single_type_kv_cache_manager.py \ tests/v1/kv_offload/cpu/test_compact_manager.py \ tests/v1/kv_offload/cpu/test_compact_accounting.py \ tests/v1/kv_offload/cpu/test_fixed_page_allocator.py \ tests/v1/kv_offload/cpu/test_canonical_layout.py \ tests/v1/kv_offload/cpu/test_manager.py \ tests/v1/kv_offload/cpu/policies \ tests/entrypoints/openai/test_deepseek_v4_thinking_kwargs.py \ tests/tokenizers_/test_deepseek_v4.py \ tests/model_executor/test_deepseek_v4_sparse_mla_metadata.py \ tests/model_executor/test_deepseek_v4_kernel_warmup.py \ tests/model_executor/test_deepseek_v4_flashmla_decode_dispatch.py \ tests/model_executor/test_deepseek_v4_moe_metadata.py \ --deselect=tests/models/test_deepseek_v4_mega_moe.py \ -k 'not gpu_roundtrip and not cross_topology_roundtrip and not writer_rotation_submits and not mhc_warmup_drives_dummy_runs'Remaining gate
AI assistance
OpenAI Codex assisted with conflict resolution, semantic review, regression tests, the isolated build, and this draft. This PR is being opened at the fork owner's explicit request; it is not represented as having completed human line-by-line review.
September 12 follow-up: selected perf backports promoted
This section supersedes the older Remaining gates deployment statements above.
The head is now abda028. Three Python-only commits were added after the validated upstream-sync head 3361f28:
Combined validation: 41/41 DSpark/SM120 tests; 177/178 broad KV/cache tests, with the only failure caused by gated google/gemma-3-1b-it fixture access returning HTTP 403; Ruff/format clean for all 17 changed Python files; cold boot; exact 187; structured tool call; health 200; unchanged 9.63 GiB / 1,778,916-token GPU KV allocation; native 32 GiB CPU offload; matched hard8 x15 exactly 115/120 with the same five length caps and no incorrect completed answers. Reverse-A/B warm decode was +1.04% mean. Repeated short prefill was effectively flat. A 141,317-token extreme-disambiguation recall prompt completed correctly in 26.9 seconds without truncation.
Upstream vllm-project#54674 was tested but intentionally excluded: it reduced KV capacity by 3.8%, slowed C=1 decode about 3%, and did not show a repeatable C=8 benefit.
The build is live on epyc from /home/ubuntu/llm-src-ds4-upstream-20260912. Rollback branch backup/ds4-sm120-upstream-20260912-pre-perf retains 3361f28.