-
-
Notifications
You must be signed in to change notification settings - Fork 20k
Open Issues sidebar navigation
All issues
Issue creation is restricted in this repository
- #42770 · WoosukKwon opened
on May 16on May 15, 2026 21 - #44280 · BugenZhao opened
on Jun 2on Jun 2, 2026 20 - #48168 · simon-mo opened
3w agoon Jul 9, 2026 3
To pick up a draggable item, press the space bar.
While dragging, use the arrow keys to move the item.
Press space again to drop the item in its new position, or press escape to cancel.
Issues
is:issue state:open
is:issue state:open
Search results
- Status: Open.#50295 In vllm-project/vllm;
- Status: Open.#50292 In vllm-project/vllm;1 comment
[Feature]: Expose startup status and health endpoint before model engine is ready
feature requestNew feature or requestNew feature or requestStatus: Open.#50282 In vllm-project/vllm;- Status: Open.#50281 In vllm-project/vllm;
[Bug]: Rust HF benchmark hard-codes the Dataset Viewer endpoint and can stall when it is unreachable
bugSomething isn't workingSomething isn't workingStatus: Open.#50271 In vllm-project/vllm;2 comments[Bug]: Host memory is not reducing after the model is loaded into Intel XPU
bugSomething isn't workingSomething isn't workingintel-gpuRelated to Intel GPURelated to Intel GPUStatus: Open.#50269 In vllm-project/vllm;[Performance][ROCm]: Fused bf16→fp32 router GEMM falls back to a standalone copy kernel on ROCm
performancePerformance-related issuesPerformance-related issuesrocmRelated to AMD ROCmRelated to AMD ROCmStatus: Open.#50267 In vllm-project/vllm;1 comment[Performance]: On RDNA, hybrid-Mamba models fall back to Triton paged attention and decode collapses at long context
performancePerformance-related issuesPerformance-related issuesrocmRelated to AMD ROCmRelated to AMD ROCmStatus: Open.#50264 In vllm-project/vllm;2 comments[Doc]: Add feature documentation for undocumented vLLM-specific SamplingParams fields
documentationImprovements or additions to documentationImprovements or additions to documentationStatus: Open.#50259 In vllm-project/vllm;- Status: Open.#50258 In vllm-project/vllm;
[Bug]: Responses API validation errors bypass VLLMValidationError — 4 sites in harmony_utils.py and responses/utils.py
bugSomething isn't workingSomething isn't workingStatus: Open.#50256 In vllm-project/vllm;[Kernel][Model] Gemma4: optimize FA4 mm_prefix range lookup and CuTe JIT stability
feature requestNew feature or requestNew feature or requestStatus: Open.#50255 In vllm-project/vllm;3 comments[Fix]: 7 user-facing errors in chat_utils.py bypass VLLMValidationError after PR #49665
documentationImprovements or additions to documentationImprovements or additions to documentationStatus: Open.#50253 In vllm-project/vllm;[Bug]: CUDA error: no kernel image is available for execution on the device when running Kimi-K3 on Ampere (A100) with vllm/vllm-openai:kimi-k3
bugSomething isn't workingSomething isn't workingStatus: Open.#50249 In vllm-project/vllm;[Feature]: Add DFly speculative decoding and D-Cut support
feature requestNew feature or requestNew feature or requestStatus: Open.#50245 In vllm-project/vllm;[Bug]: UVA is not available on WSL2 with RTX 5050 (sm_120) — crashes on basic model load, no CPU offload
bugSomething isn't workingSomething isn't workingStatus: Open.#50239 In vllm-project/vllm;[Bug]: ROCm mixed KV-cache layouts abort speculative decoding with AssertionError
rocmRelated to AMD ROCmRelated to AMD ROCmStatus: Open.#50237 In vllm-project/vllm;3 comments[Bug]: Kimi-K3 prefix cache miss when prompt length is exactly a 1536-token boundary
bugSomething isn't workingSomething isn't workingStatus: Open.#50235 In vllm-project/vllm;1 comment- Status: Open.#50217 In vllm-project/vllm;2 linked PRs
[Installation]: VLLM_PRECOMPILED_WHEEL_COMMIT=nightly is rejected
installationInstallation problemsInstallation problemsStatus: Open.#50215 In vllm-project/vllm;- Status: Open.#50206 In vllm-project/vllm;
[Bug]: vllm-k3-toolcall repeat and repeat error
bugSomething isn't workingSomething isn't workingStatus: Open.#50203 In vllm-project/vllm;- Status: Open.#50189 In vllm-project/vllm;
- Status: Open.#50188 In vllm-project/vllm;4 comments
[CI Failure]: test_deep_ep_v2_moe FP8 output tolerance failure
ci-failureIssue about an unexpected test failure in CIIssue about an unexpected test failure in CIStatus: Open.#50184 In vllm-project/vllm;