vllm

A high-throughput and memory-efficient inference and serving engine for LLMs

90.5k
Stars
+27.7k
Gained
44.2%
Growth
Python
Language

🎭 Best For

⚖️ Compare With

🏷️ Topics & Ecosystem

amd blackwell cuda deepseek deepseek-v3 gpt gpt-oss inference kimi llama llm llm-serving model-serving moe openai pytorch qwen qwen3 tpu transformer

📊 Activity

Latest commit: 2026-08-30. Over the past 291 days, this repository gained 27.7k stars (+44.2% growth). Activity data is based on daily RepoPi snapshots of the GitHub repository.