Inference & Local Deploy
High-throughput serving and local runtimes — Ollama, vLLM, Triton, and more.
183 projects
See methodology for ranking rules; order uses public GitHub metrics within this scenario.
| Rank | Project | Stars | Forks | Updated | License |
|---|---|---|---|---|---|
| 41 |
Mooncake
Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI. |
6.4K | 1.1K | 08/29/26 | Apache-2.0 |
| 42 |
kserve
Standardized Distributed Generative and Predictive AI Inference Platform for Scalable, Multi-Framework Deployment on Kubernetes |
5.8K | 1.6K | 08/29/26 | Apache-2.0 |
| 43 |
mlx-lm
Run LLMs with MLX |
6.8K | 1.0K | 08/29/26 | MIT |
| 44 |
GenieX
Run frontier LLMs and VLMs locally on Qualcomm devices across NPU, GPU, and CPU with a few lines of code |
8.3K | 1.0K | 08/28/26 | BSD-3-Clause |
| 45 |
mistral.rs
Fast, flexible LLM inference |
7.6K | 689 | 08/29/26 | MIT |
| 46 |
serving
A flexible, high-performance serving system for machine learning models |
6.4K | 2.2K | 08/28/26 | Apache-2.0 |
| 47 |
stable-diffusion.cpp
Diffusion model(SD,Flux,Wan,Qwen Image,Z-Image,...) inference in pure C/C++ |
6.9K | 761 | 08/27/26 | MIT |
| 48 |
vllm-ascend
Community maintained hardware plugin for vLLM on Ascend |
2.7K | 2.1K | 08/29/26 | Apache-2.0 |
| 49 |
coremltools
Core ML tools contain supporting tools for Core ML model conversion, editing, and validation. |
5.4K | 838 | 08/27/26 | BSD-3-Clause |
| 50 |
kimi-k3-in-c
A 2.78-trillion-parameter Kimi K3 running inference on a single CPU in 8.24 GB of RAM. Portable C99: no BLAS, no framework, no GPU. |
6.7K | 1.1K | 08/26/26 | Apache-2.0 |
| 51 |
mlx-vlm
MLX-VLM is a package for inference and fine-tuning of Vision Language Models (VLMs) on your Mac using MLX. |
5.4K | 751 | 08/29/26 | MIT |
| 52 |
lemonade
Lemonade helps users discover and run local AI apps by serving optimized LLMs right from their own GPUs and NPUs. Join our discord: https://discord.gg/5xXzkMu8Zk |
5.5K | 479 | 08/30/26 | Apache-2.0 |
| 53 |
turbo-fieldfare
Gemma 4 26B-A4B inference in ~2 GB of RAM on any M-series MacBook |
6.5K | 408 | 08/29/26 | Apache-2.0 |
| 54 |
iree
A retargetable MLIR-based machine learning compiler and runtime toolkit. |
3.9K | 995 | 08/29/26 | Apache-2.0 |
| 55 |
whichllm
Find the local LLM that actually runs and performs best on your hardware. Ranked by real, recency-aware benchmarks, not parameter count. One command, run it instantly. |
6.5K | 357 | 08/14/26 | MIT |
| 56 |
cactus
Quantization, kernels, runtime and inference engine for mobiles, wearables, smart home and robots. |
6.0K | 497 | 08/26/26 | Other |
| 57 |
gpustack
A GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU instances. |
5.6K | 631 | 08/28/26 | Apache-2.0 |
| 58 |
apfel
The free AI already on your Mac. CLI tool, OpenAI-compatible server, and interactive chat — all on-device via Apple Intelligence. No API keys, no cloud, no downloads. |
6.3K | 243 | 08/05/26 | MIT |
| 59 |
shimmy
⚡ Pure-Rust WebGPU inference engine — OpenAI-API compatible, GGUF native, runs on any GPU. No Python. No llama.cpp. Single binary. |
5.8K | 561 | 08/29/26 | Apache-2.0 |
| 60 |
llm-d
Achieve state of the art inference performance with modern accelerators on Kubernetes |
4.3K | 727 | 08/29/26 | Apache-2.0 |