Inference & Local Deploy
High-throughput serving and local runtimes — Ollama, vLLM, Triton, and more.
183 projects
See methodology for ranking rules; order uses public GitHub metrics within this scenario.
| Rank | Project | Stars | Forks | Updated | License |
|---|---|---|---|---|---|
| 1 |
ollama
Get up and running with Kimi-K2.6, GLM-5.2, MiniMax, DeepSeek, gpt-oss, Qwen, Gemma and other models. |
179.7K | 17.6K | 08/29/26 | MIT |
| 2 |
llama.cpp
LLM inference in C/C++ |
126.3K | 22.4K | 08/29/26 | MIT |
| 3 |
vllm
A high-throughput and memory-efficient inference and serving engine for LLMs |
90.4K | 21.4K | 08/29/26 | Apache-2.0 |
| 4 |
gpt4all
GPT4All: Run Local LLMs on Any Device. Open-source and available for commercial use. |
77.4K | 8.3K | 05/27/25 | MIT |
| 5 |
LocalAI
LocalAI is the open-source AI engine. Run any model - LLMs, vision, voice, image, video - on any hardware. No GPU required. |
48.8K | 4.4K | 08/29/26 | MIT |
| 6 |
textgen
Open-source desktop app for local LLMs. Text, vision, tool-calling, OpenAI/Anthropic-compatible API. 100% private. |
47.6K | 6.0K | 08/17/26 | AGPL-3.0 |
| 7 |
exo
Run frontier AI locally. |
47.1K | 3.5K | 08/25/26 | Apache-2.0 |
| 8 |
sglang
SGLang is a high-performance serving framework for large language models and multimodal models. |
32.8K | 8.3K | 08/29/26 | Apache-2.0 |
| 9 |
BitNet
Official inference framework for 1-bit LLMs |
40.2K | 3.7K | 07/27/26 | MIT |
| 10 |
llmfit
Hundreds of models & providers. One command to find what runs on your hardware. |
34.5K | 2.2K | 08/28/26 | MIT |
| 11 |
airllm
AirLLM 70B inference with single 4GB GPU |
33.1K | 3.5K | 08/29/26 | Apache-2.0 |
| 12 |
modular
The Modular Platform (includes MAX & Mojo) |
29.3K | 3.1K | 08/29/26 | Other |
| 13 |
onnxruntime
ONNX Runtime: cross-platform, high performance ML inferencing and training accelerator |
21.7K | 4.2K | 08/29/26 | MIT |
| 14 |
ncnn
ncnn is a high-performance neural network inference framework optimized for the mobile platform |
23.8K | 4.5K | 08/28/26 | Other |
| 15 |
llamafile
Distribute and run LLMs with a single file. |
25.7K | 1.6K | 08/26/26 | Other |
| 16 |
mlc-llm
Universal LLM Deployment Engine with ML Compilation |
23.1K | 2.1K | 08/17/26 | Apache-2.0 |
| 17 |
omlx
LLM inference server with continuous batching & SSD caching for Apple Silicon — managed from the macOS menu bar |
21.0K | 1.8K | 08/29/26 | Apache-2.0 |
| 18 |
ktransformers
A Flexible Framework for Experiencing Heterogeneous LLM Inference/Fine-tune Optimizations |
19.3K | 1.5K | 08/28/26 | Apache-2.0 |
| 19 |
web-llm
High-performance In-browser LLM Inference Engine |
18.6K | 1.4K | 08/04/26 | Apache-2.0 |
| 20 |
TensorRT-LLM
Optimize and deploy large language models with state-of-the-art inference performance on NVIDIA GPUs. |
14.5K | 2.7K | 08/29/26 | Other |