Inference & Local Deploy
High-throughput serving and local runtimes — Ollama, vLLM, Triton, and more.
183 projects
See methodology for ranking rules; order uses public GitHub metrics within this scenario.
| Rank | Project | Stars | Forks | Updated | License |
|---|---|---|---|---|---|
| 101 |
Bonsai-demo
Bonsai Demo |
2.3K | 233 | 08/28/26 | Apache-2.0 |
| 102 |
spark-vllm-docker
Docker configuration for running VLLM on dual DGX Sparks |
2.2K | 373 | 08/27/26 | MIT |
| 103 |
club-3090
Community recipes for serving LLMs on RTX 3090/4090/5090 CUDA gpus. Multi-engine (vLLM, llama.cpp, ik_llama) and model-agnostic. Currently shipping Qwen3.6-27B Qwen3.6 35B Gemma 4 26B Gemma 4 31B configs for 1× and 2× cards. |
2.1K | 128 | 08/27/26 | Apache-2.0 |
| 104 |
vllm-metal
Community maintained hardware plugin for vLLM on Apple Silicon |
1.7K | 233 | 08/29/26 | Apache-2.0 |
| 105 |
lorax
Multi-LoRA inference server that scales to 1000s of fine-tuned LLMs |
3.8K | 326 | 05/28/26 | Apache-2.0 |
| 106 |
FastFlowLM
Run LLMs on AMD Ryzen™ AI NPUs in minutes; purpose-built and deeply optimized for the AMD NPUs. |
1.8K | 145 | 08/28/26 | MIT |
| 107 |
FastFlowLM
Run LLMs on AMD Ryzen™ AI NPUs in minutes; purpose-built and deeply optimized for the AMD NPUs. |
1.8K | 145 | 08/28/26 | MIT |
| 108 |
GPTQModel
LLM model quantization (compression) toolkit with HW acceleration support for Nvidia, AMD, Intel GPU and Intel/AMD/Apple CPU via HF, vLLM, and SGLang. |
1.2K | 204 | 08/28/26 | Other |
| 109 |
react-native-executorch
Declarative way to run AI models in React Native on device, powered by ExecuTorch. |
1.7K | 95 | 08/28/26 | Other |
| 110 |
truss
The simplest way to serve AI/ML models in production |
1.2K | 122 | 08/28/26 | MIT |
| 111 |
vllm-mlx
High-performance OpenAI and Anthropic compatible LLM inference server for Apple Silicon. Native MLX, continuous batching, multimodal models, MCP tool calling, and Claude Code support. |
1.6K | 219 | 08/26/26 | Apache-2.0 |
| 112 |
auto-round
A SOTA quantization toolkit for high-accuracy low-bit LLM inference, seamlessly optimized for CPU/XPU/CUDA, with multi-datatype support and full compatibility with vLLM, SGLang, and Transformers|简洁且高效的量化工具包 |
1.6K | 168 | 08/29/26 | Apache-2.0 |
| 113 |
nndeploy
An Easy-to-Use and High-Performance AI Deployment Framework |
1.9K | 233 | 08/15/26 | Apache-2.0 |
| 114 |
node-llama-cpp
Run AI models locally on your machine with node.js bindings for llama.cpp. Enforce a JSON schema on the model output on the generation level |
2.2K | 215 | 08/11/26 | MIT |
| 115 |
distributed-llama
Distributed LLM inference. Connect home devices into a powerful cluster to accelerate LLM inference. More devices means faster inference. |
3.0K | 248 | 07/05/26 | MIT |
| 116 |
uzu
A high-performance inference engine for AI models |
1.7K | 73 | 08/29/26 | MIT |
| 117 |
sonar
Large-scale LLM inference engine |
1.8K | 208 | 08/13/26 | AGPL-3.0 |
| 118 |
local-studio
Control panel for VLLM, Sglang, llama.cpp, exllamav3 |
1.7K | 155 | 08/28/26 | Apache-2.0 |
| 119 |
ComfyUI-Docker
🐳Dockerfile for 🎨ComfyUI. | 容器镜像与启动脚本 |
1.6K | 232 | 08/26/26 | Other |
| 120 |
mllm
Fast Multimodal LLM on Mobile Devices |
1.6K | 213 | 08/19/26 | MIT |