Inference & Local Deploy
High-throughput serving and local runtimes — Ollama, vLLM, Triton, and more.
183 projects
See methodology for ranking rules; order uses public GitHub metrics within this scenario.
| Rank | Project | Stars | Forks | Updated | License |
|---|---|---|---|---|---|
| 141 |
ollama-docker
Welcome to the Ollama Docker Compose Setup! This project simplifies the deployment of Ollama using Docker Compose, making it easy to run Ollama with all its dependencies in a containerized environment |
1.4K | 254 | 05/26/26 | Other |
| 142 |
SageAttention
[ICLR2025, ICML2025, NeurIPS2025 Spotlight] Quantized Attention achieves speedup of 2-5x compared to FlashAttention, without losing end-to-end metrics across language, image, and video models. |
3.7K | 493 | 01/17/26 | Apache-2.0 |
| 143 |
TensorRT-YOLO
🚀 Easier & Faster YOLO Deployment Toolkit for NVIDIA 🛠️ |
1.9K | 194 | 03/22/26 | GPL-3.0 |
| 144 |
turboquant
TurboQuant: Near-optimal KV cache quantization for LLM inference (3-bit keys, 2-bit values) with Triton kernels + vLLM integration |
1.7K | 195 | 03/27/26 | GPL-3.0 |
| 145 |
rotorquant
KV cache compression via block-diagonal rotation. Beats TurboQuant: better PPL (6.91 vs 7.07), 28% faster decode, 5.3x faster prefill, 44x fewer params. Drop-in llama.cpp integration. |
1.0K | 89 | 04/23/26 | Other |
| 146 |
deepseek-ocr.rs
Rust multi‑backend OCR/VLM engine (DeepSeek‑OCR-1/2, PaddleOCR‑VL, DotsOCR) with DSQ quantization and an OpenAI‑compatible server & CLI – run locally without Python. |
2.2K | 169 | 02/21/26 | Apache-2.0 |
| 147 |
gaianet-node
Install, run and deploy your own decentralized AI agent service |
5.0K | 326 | 10/13/25 | GPL-3.0 |
| 148 |
picolm
Run a 1-billion parameter LLM on a $10 board with 256MB RAM |
1.9K | 240 | 02/22/26 | MIT |
| 149 |
LLMFarm
llama and other large language models on iOS and MacOS offline using GGML library. |
2.1K | 180 | 01/30/26 | MIT |
| 150 |
LlamaEdge
The easiest & fastest way to run customized and fine-tuned LLMs locally or on the edge |
1.7K | 149 | 02/08/26 | Apache-2.0 |
| 151 |
comfyui-deploy
An open source `vercel` like deployment platform for Comfy UI |
1.5K | 219 | 11/13/25 | AGPL-3.0 |
| 152 |
hummingbird
Hummingbird compiles trained ML models into tensor computation for faster inference. |
3.5K | 292 | 07/17/25 | MIT |
| 153 |
Jlama
Jlama is a modern LLM inference engine for Java |
1.3K | 164 | 10/12/25 | Apache-2.0 |
| 154 |
dalai
The simplest way to run LLaMA on your local machine |
12.9K | 1.3K | 06/18/24 | Other |
| 155 |
petals
🌸 Run LLMs at home, BitTorrent-style. Fine-tuning and inference up to 10x faster than offloading |
10.5K | 643 | 09/07/24 | MIT |
| 156 |
DeepSpeed-MII
MII makes low-latency and high-throughput inference possible, powered by DeepSpeed. |
2.1K | 192 | 06/30/25 | Apache-2.0 |
| 157 |
Tengine
Tengine is a lite, high performance, modular inference engine for embedded device |
4.5K | 982 | 03/06/25 | Apache-2.0 |
| 158 |
TurboTransformers
a fast and user-friendly runtime for transformer inference (Bert, Albert, GPT2, Decoders, etc) on CPU and GPU. |
1.5K | 208 | 07/18/25 | Other |
| 159 |
uTensor
TinyML AI inference library |
1.9K | 250 | 05/10/25 | Apache-2.0 |
| 160 |
rwkv.cpp
INT4/INT5/INT8 and FP16 inference on CPU for RWKV language model |
1.6K | 129 | 03/23/25 | MIT |