Best Open Source LLM Runtimes
Ranked by GitHub stars and 252 days of tracked growth. The top open source LLM runtimes below are updated daily β pick two to see a full side-by-side comparison.
1
vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
2
ray
Ray is an AI compute engine. Ray consists of a core distributed runtime and a set of AI Libraries for accelerating ML workloads.
3
DeepSpeed
DeepSpeed is a deep learning optimization library that makes distributed training and inference easy, efficient, and effective.
4
ColossalAI
Making large AI models cheaper, faster and more accessible
5
ml-engineering
Machine Learning Engineering Open Book
6
OpenLLM
Run any open-source LLMs, such as DeepSeek and Llama, as OpenAI compatible API endpoint in the cloud.
7
amazon-sagemaker-examples
Example π Jupyter notebooks that demonstrate how to build, train, and deploy machine learning models using π§ Amazon SageMaker.
8
openvino
OpenVINOβ’ is an open source toolkit for optimizing and deploying AI inference
9
skypilot
Run, manage, and scale AI workloads on any AI infrastructure. Use one system to access & manage all AI compute (Kubernetes, Slurm, 20+ clouds, on-prem).
10
BentoML
The easiest way to serve AI apps and models - Build Model Inference APIs, Job queues, LLM apps, Multi-model pipelines, and more!
11
optillm
Optimizing inference proxy for LLMs
βοΈ Compare LLM runtimes head-to-head
Real adoption data, side by side.