forked from vllm-project/vllm
-
Notifications
You must be signed in to change notification settings - Fork 52
Pull requests: ROCm/vllm
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
ci: bump build-rocm-wheels to torch 2.13, fix stale-line version resolution
#1237
opened Aug 21, 2026 by
marcusr-amd
Loading…
rocm: implement tie_weights for the dynamic INT8 lm_head
#1213
opened Aug 18, 2026 by
roberteg16
•
Draft
3 tasks done
[perf]: fuse forward_static() RoPE fallback ops via torch.compile
#1212
opened Aug 18, 2026 by
serged-amd
Loading…
[ROCm][Attention] Wide MHA decode attention kernel for gfx1151
#1087
opened Aug 12, 2026 by
roberteg16
•
Draft
5 of 6 tasks
[Model] Complete port of Transformers v5 heterogeneous config fix
#1086
opened Aug 11, 2026 by
amd-callumm
•
Draft
[ROCm] Stop ROCPROFILER_REGISTER_LIBRARY leaking into engine subprocesses
#1077
opened Aug 3, 2026 by
roberteg16
•
Draft
5 tasks done
[bench] vit_attention: model-matching V stride + controllable cache-eviction (--flush-mib)
#1070
opened Jul 30, 2026 by
parthash0804
Loading…
[ROCm][Kernel] W4A16 MoE prefill: M-aware moe_align block size (1.6x TTFT)
#1069
opened Jul 29, 2026 by
stoivonen-amd
•
Draft
3 of 5 tasks
[ROCm][Kernel] W4A16 skinny GEMM: hand-asm kernel for gfx1151
#1045
opened Jul 1, 2026 by
mgehre-amd
•
Draft
[ROCm] Split skinny_gemms_int8.cu into per-N translation units
#1028
opened Jun 26, 2026 by
marcusr-amd
•
Draft
3 of 5 tasks
tune Qwen3-VL-4B prefill unified-attention on gfx1150
#1024
opened Jun 26, 2026 by
qingxuamd
Loading…
[ROCm][MoE] W4A16 MoE routing-distribution benchmark suite for the gfx11 prefill GEMM
#1020
opened Jun 25, 2026 by
roberteg16
•
Draft
feat: Add NPU+GPU async pipelining for vision-language models
#936
opened May 14, 2026 by
liangliangchang
•
Draft
4 of 5 tasks
Auto-build flash-attn wheels on push, upload to S3
#910
opened Apr 30, 2026 by
mgehre-amd
•
Draft
1 task
[ROCm][DSv4] Share AITER decode dequant + fp8-cast buffers across layers (rebased, stacked on #902)
#903
opened Apr 27, 2026 by
ChuanLi1101
•
Draft
2 of 4 tasks
[ROCm][DSv4] Make AITER sparse decode cudagraph-clean (rebased, stacked on #901)
#902
opened Apr 27, 2026 by
ChuanLi1101
•
Draft
2 of 5 tasks
[ROCm][DSv4] AITER-accelerated MLA decode for DeepSeek V4 on MI355X (rebased on tj/dsv4prrebase)
#901
opened Apr 27, 2026 by
ChuanLi1101
•
Draft
1 of 4 tasks
[Do Not Merge] For review purpose: Rocm/aiter mla dsv4 decode cudagraph
#900
opened Apr 26, 2026 by
tjtanaavllm
•
Draft
5 tasks
Previous Next
ProTip!
Mix and match filters to narrow down what you’re looking for.