Inference & Local Deploy
High-throughput serving and local runtimes — Ollama, vLLM, Triton, and more.
183 projects
See methodology for ranking rules; order uses public GitHub metrics within this scenario.
| Rank | Project | Stars | Forks | Updated | License |
|---|---|---|---|---|---|
| 61 |
Paddle-Lite
PaddlePaddle High Performance Deep Learning Inference Engine for Mobile and Edge (飞桨高性能深度学习端侧推理引擎) |
7.3K | 1.6K | 04/27/26 | Apache-2.0 |
| 62 |
FastDeploy
High-performance Inference and Deployment Toolkit for LLMs and VLMs based on PaddlePaddle |
3.7K | 760 | 08/26/26 | Apache-2.0 |
| 63 |
fastllm
fastllm是后端无依赖的高性能大模型推理库。同时支持张量并行推理稠密模型和混合模式推理MOE模型,任意10G以上显卡即可推理满血DeepSeek。双路9004/9005服务器+单显卡部署DeepSeek满血满精度原版模型,单并发20tps;INT4量化模型单并发30tps,多并发可达60+。 |
4.9K | 485 | 08/29/26 | Apache-2.0 |
| 64 |
ODS
Turn your PC, Mac, or Linux box into an AI server. LLM inference, chat UI, voice, agents, workflows, RAG, and image generation. |
4.9K | 741 | 08/30/26 | Apache-2.0 |
| 65 |
RWKV-Runner
A RWKV management and startup tool, full automation, only 8MB. And provides an interface compatible with the OpenAI API. RWKV is a large language model that is fully open source and available for commercial use. |
6.5K | 600 | 07/07/26 | MIT |
| 66 |
llm-compressor
Transformers-compatible library for applying various compression algorithms to LLMs for optimized deployment with vLLM |
3.7K | 644 | 08/28/26 | Apache-2.0 |
| 67 |
lms
LM Studio CLI |
5.2K | 447 | 08/18/26 | MIT |
| 68 |
jetson-containers
Machine Learning Containers for NVIDIA Jetson and JetPack-L4T |
4.8K | 845 | 08/10/26 | Other |
| 69 |
CTranslate2
Fast inference engine for Transformer models |
4.7K | 522 | 08/16/26 | MIT |
| 70 |
lollms-webui
Lord of Large Language and Multi modal Systems Web User Interface |
4.8K | 587 | 08/13/26 | Apache-2.0 |
| 71 |
text-embeddings-inference
A blazing fast inference solution for text embeddings models |
5.0K | 426 | 07/24/26 | Apache-2.0 |
| 72 |
LightLLM
LightLLM is a Python-based LLM (Large Language Model) inference and serving framework, notable for its lightweight design, easy scalability, and high-speed performance. |
4.3K | 357 | 08/29/26 | Apache-2.0 |
| 73 |
optimum
🚀 Accelerate inference and training of 🤗 Transformers, Diffusers, TIMM and Sentence Transformers with easy to use hardware optimization tools |
3.5K | 680 | 08/24/26 | Apache-2.0 |
| 74 |
ComputeLibrary
The Compute Library is a set of computer vision and machine learning functions optimised for both Arm CPUs and GPUs using SIMD technologies. |
3.2K | 819 | 08/27/26 | Other |
| 75 |
Model-Optimizer
A unified library of state-of-the-art model optimization techniques (quantization, pruning, NAS, distillation, speculative decoding) to compress and accelerate deep learning models for deployment in frameworks like TensorRT-LLM, vLLM, and SGLang. |
3.6K | 565 | 08/30/26 | Apache-2.0 |
| 76 |
zml
Any model. Any hardware. Zero compromise. Built with @ziglang / @openxla / MLIR / @bazelbuild |
4.0K | 179 | 08/29/26 | Apache-2.0 |
| 77 |
ramalama
RamaLama is an open-source developer tool that simplifies the local serving of AI models from any source and facilitates their use for inference in production, all through the familiar language of containers. |
3.0K | 360 | 08/29/26 | MIT |
| 78 |
TensorRT
PyTorch/TorchScript/FX compiler for NVIDIA GPUs using TensorRT |
3.0K | 409 | 08/29/26 | BSD-3-Clause |
| 79 |
LitServe
A minimal Python framework for building custom AI inference servers with full control over logic, batching, and scaling. |
3.9K | 301 | 08/17/26 | Apache-2.0 |
| 80 |
Rapid-MLX
The fastest local AI engine for Apple Silicon. 4.2x faster than Ollama, 0.08s cached TTFT, 100% tool calling. 17 tool parsers, prompt cache, reasoning separation, cloud routing. Drop-in OpenAI replacement. Works with Claude Code, Cursor, Aider. |
3.6K | 402 | 08/30/26 | Other |