Skip to content
Change the repository type filter

All

    Repositories list

    • compressed-tensors

      Public
      A safetensors extension to efficiently store sparse quantized tensors on disk
      Python
      Apache License 2.0
      108308742Updated Jul 29, 2026Jul 29, 2026
    • vllm-omni

      Public
      A framework for efficient model inference with omni-modality models
      Python
      Apache License 2.0
      1.4k5.7k630681Updated Jul 29, 2026Jul 29, 2026
    • tpu-inference

      Public
      TPU inference for vLLM, with unified JAX and PyTorch support.
      Python
      Apache License 2.0
      27139573320Updated Jul 29, 2026Jul 29, 2026
    • vllm

      Public
      A high-throughput and memory-efficient inference and serving engine for LLMs
      Python
      Apache License 2.0
      20k88k2k4.1kUpdated Jul 29, 2026Jul 29, 2026
    • perf-eval

      Public
      Performance benchmark & accuracy evaluation for vLLM
      Python
      1511023Updated Jul 29, 2026Jul 29, 2026
    • vllm-ascend

      Public
      Community maintained hardware plugin for vLLM on Ascend
      C++
      Apache License 2.0
      1.9k2.5k1.5k1kUpdated Jul 29, 2026Jul 29, 2026
    • speculators

      Public
      A unified library for building, evaluating, and storing speculative decoding algorithms for LLM inference in vLLM
      Python
      Apache License 2.0
      1726694066Updated Jul 29, 2026Jul 29, 2026
    • vllm-project.github.io

      Public
      HTML
      12358413Updated Jul 29, 2026Jul 29, 2026
    • afd-plugin

      Public
      vLLM plugin for attention-ffn disaggregation support
      Python
      Apache License 2.0
      1655166Updated Jul 29, 2026Jul 29, 2026
    • guidellm

      Public
      Evaluate and Enhance Your LLM Deployments for Real-World Inference Needs
      Python
      Apache License 2.0
      2011.4k5020Updated Jul 29, 2026Jul 29, 2026
    • llm-compressor

      Public
      Transformers-compatible library for applying various compression algorithms to LLMs for optimized deployment with vLLM
      Python
      Apache License 2.0
      5923.6k4578Updated Jul 29, 2026Jul 29, 2026
    • vllm-gaudi

      Public
      Community maintained hardware plugin for vLLM on Intel Gaudi
      Python
      Apache License 2.0
      14649462Updated Jul 29, 2026Jul 29, 2026
    • agentic-api

      Public
      Stateful API logic for agentic applications using vLLM
      Rust
      Apache License 2.0
      2357259Updated Jul 29, 2026Jul 29, 2026
    • recipes

      Public
      Common recipes to run vLLM
      JavaScript
      Apache License 2.0
      34993935109Updated Jul 29, 2026Jul 29, 2026
    • semantic-router

      Public
      Intelligent Mixture-of-Models Router for Efficient Heterogeneous LLMs Inference
      Go
      Apache License 2.0
      7855.1k19183Updated Jul 29, 2026Jul 29, 2026
    • vllm-bnb-plugin

      Public
      vLLM Quantization plugin for bitsandbytes
      Python
      Apache License 2.0
      1201Updated Jul 29, 2026Jul 29, 2026
    • DeepGEMM

      Public
      DeepGEMM: clean and efficient FP8 GEMM kernels with fine-grained scaling
      Cuda
      MIT License
      1.1k000Updated Jul 29, 2026Jul 29, 2026
    • vllm-xpu-kernels

      Public
      The vLLM XPU kernels for Intel GPU
      C++
      Apache License 2.0
      92561249Updated Jul 29, 2026Jul 29, 2026
    • vllm-dashboard

      Public
      TypeScript
      91012Updated Jul 29, 2026Jul 29, 2026
    • ci-infra

      Public
      This repo hosts code for vLLM CI & Performance Benchmark infrastructure.
      Python
      Apache License 2.0
      7445055Updated Jul 29, 2026Jul 29, 2026
    • vllm-daily

      Public
      vLLM Daily Summarization of Merged PRs
      55100Updated Jul 29, 2026Jul 29, 2026
    • aibrix

      Public
      Cost-efficient and pluggable Infrastructure components for GenAI inference
      Go
      Apache License 2.0
      6325k31238Updated Jul 29, 2026Jul 29, 2026
    • production-stack

      Public
      vLLM’s reference system for K8S-native cluster-wide deployment with community-driven performance optimization
      Python
      Apache License 2.0
      4542.5k9882Updated Jul 28, 2026Jul 28, 2026
    • vllm-metal

      Public
      Community maintained hardware plugin for vLLM on Apple Silicon
      Python
      Apache License 2.0
      1991.5k86Updated Jul 28, 2026Jul 28, 2026
    • vllm-bench

      Public
      High-performance Rust benchmark client for vLLM serving endpoints.
      Rust
      Apache License 2.0
      114913Updated Jul 28, 2026Jul 28, 2026
    • vime

      Public
      An LLM post-training framework with vLLM for RL Scaling
      Python
      Apache License 2.0
      693933213Updated Jul 28, 2026Jul 28, 2026
    • FlashKDA

      Public
      Cuda
      MIT License
      0800Updated Jul 28, 2026Jul 28, 2026
    • MSA

      Public
      Python
      MIT License
      46001Updated Jul 24, 2026Jul 24, 2026
    • router

      Public
      A high-performance and light-weight router for vLLM large scale deployment
      Rust
      Apache License 2.0
      1133331927Updated Jul 24, 2026Jul 24, 2026
    • flash-attention

      Public
      Fast and memory-efficient exact attention
      Python
      BSD 3-Clause "New" or "Revised" License
      2.9k133043Updated Jul 22, 2026Jul 22, 2026
    ProTip! When viewing an organization's repositories, you can use the props. filter to filter by custom property.