Inference & Local Deploy
High-throughput serving and local runtimes — Ollama, vLLM, Triton, and more.
183 projects
See methodology for ranking rules; order uses public GitHub metrics within this scenario.
| Rank | Project | Stars | Forks | Updated | License |
|---|---|---|---|---|---|
| 21 |
MNN
MNN: A blazing-fast, lightweight inference engine battle-tested by Alibaba, powering high-performance on-device LLMs and Edge AI. |
16.0K | 2.4K | 08/28/26 | Apache-2.0 |
| 22 |
openvino
OpenVINO™ is an open source toolkit for optimizing and deploying AI inference |
10.8K | 3.3K | 08/28/26 | Apache-2.0 |
| 23 |
TensorRT
NVIDIA® TensorRT™ is an SDK for high-performance deep learning inference on NVIDIA GPUs. This repository contains the open source components of TensorRT. |
13.3K | 2.4K | 08/25/26 | Apache-2.0 |
| 24 |
LMCache
LMCache: Supercharge Your LLM with the Fastest KV Cache Layer |
11.6K | 1.8K | 08/29/26 | Apache-2.0 |
| 25 |
WasmEdge
WasmEdge is a lightweight, high-performance, and extensible WebAssembly runtime for cloud native, edge, and decentralized applications. It powers serverless apps, embedded functions, microservices, smart contracts, and IoT devices. |
10.8K | 1.2K | 08/28/26 | Apache-2.0 |
| 26 |
OpenLLM
Run any open-source LLMs, such as DeepSeek and Llama, as OpenAI compatible API endpoint in the cloud. |
12.5K | 836 | 08/24/26 | Apache-2.0 |
| 27 |
nano-vllm
Nano vLLM |
15.2K | 2.5K | 04/26/26 | MIT |
| 28 |
server
The Triton Inference Server provides an optimized cloud and edge inferencing solution. |
10.9K | 1.8K | 08/28/26 | BSD-3-Clause |
| 29 |
llama-cpp-python
Python bindings for llama.cpp |
10.6K | 1.5K | 08/17/26 | MIT |
| 30 |
inference
Swap GPT for any LLM by changing a single line of code. Xinference lets you run open-source, speech, and multimodal models on cloud, on-prem, or your laptop — all through one unified, production-ready inference API. |
9.5K | 865 | 08/29/26 | Apache-2.0 |
| 31 |
runanywhere-sdks
Production ready toolkit to run AI locally |
10.3K | 370 | 08/29/26 | Other |
| 32 |
BentoML
The easiest way to serve AI apps and models - Build Model Inference APIs, Job queues, LLM apps, Multi-model pipelines, and more! |
8.8K | 1.0K | 08/28/26 | Apache-2.0 |
| 33 |
serve
☁️ Build multimodal AI applications with cloud-native stack |
21.9K | 2.2K | 03/24/25 | Apache-2.0 |
| 34 |
ollama-python
Ollama Python library |
10.5K | 1.2K | 08/12/26 | MIT |
| 35 |
PowerInfer
High-speed Large Language Model Serving for Local Deployment |
9.8K | 597 | 05/11/26 | MIT |
| 36 |
cog
Containers for machine learning |
9.5K | 697 | 08/26/26 | Apache-2.0 |
| 37 |
ik_llama.cpp
llama.cpp fork with additional SOTA quants and improved performance |
3.1K | 444 | 08/28/26 | MIT |
| 38 |
executorch
On-device AI across mobile, embedded and edge for PyTorch |
5.0K | 1.1K | 08/29/26 | Other |
| 39 |
vllm-omni
A framework for efficient model inference with omni-modality models |
6.5K | 1.6K | 08/29/26 | Apache-2.0 |
| 40 |
lmdeploy
LMDeploy is a toolkit for compressing, deploying, and serving LLMs. |
8.0K | 733 | 08/28/26 | Apache-2.0 |