Inference & Local Deploy
High-throughput serving and local runtimes — Ollama, vLLM, Triton, and more.
183 projects
See methodology for ranking rules; order uses public GitHub metrics within this scenario.
| Rank | Project | Stars | Forks | Updated | License |
|---|---|---|---|---|---|
| 81 |
tract
Tiny, no-nonsense, self-contained, Tensorflow and ONNX inference |
3.0K | 281 | 08/29/26 | Other |
| 82 |
production-stack
vLLM’s reference system for K8S-native cluster-wide deployment with community-driven performance optimization |
2.5K | 476 | 08/29/26 | Apache-2.0 |
| 83 |
mesh-llm
Distributed AI/LLM for the people. Share compute privately or publicly to power your agents and chat. |
3.3K | 402 | 08/30/26 | Apache-2.0 |
| 84 |
spiceai
Add a real-time analytics node to your operational database. Spice is a portable, accelerated SQL query, search, and LLM-inference engine in Rust for data-grounded AI apps and agents. |
3.1K | 224 | 08/30/26 | Apache-2.0 |
| 85 |
optillm
Optimizing inference proxy for LLMs |
4.3K | 385 | 07/18/26 | Apache-2.0 |
| 86 |
harbor
Stop configuring your AI stack. Start using it. One command brings a complete pre-wired LLM stack with hundreds of services to explore. |
3.2K | 224 | 08/29/26 | Apache-2.0 |
| 87 |
chitu
High-performance inference framework for large language models, focusing on efficiency, flexibility, and availability. |
3.0K | 256 | 08/28/26 | Apache-2.0 |
| 88 |
inference
Turn any computer or edge device into a command center for your computer vision projects. |
2.4K | 311 | 08/28/26 | Other |
| 89 |
claude-code-local
Run Claude Code fully on-device with local AI on Apple Silicon via an MLX-native Anthropic-API server, offering private, offline, airgap-ready inference. |
3.2K | 620 | 08/22/26 | MIT |
| 90 |
algernon
Small self-contained pure-Go web server with Lua, Teal, Markdown, HTTP/2, QUIC, Redis, TypeScript, npm-less React 19, SQLite, and PostgreSQL support ++ |
3.0K | 151 | 08/27/26 | BSD-3-Clause |
| 91 |
opyrator
🪄 Turns your machine learning code into microservices with web API, interactive GUI, and more. |
3.1K | 167 | 08/25/26 | MIT |
| 92 |
Olive
Olive: Simplify ML Model Finetuning, Conversion, Quantization, and Optimization for CPUs, GPUs and NPUs. |
2.4K | 312 | 08/28/26 | MIT |
| 93 |
ort
Fast ML inference & training for ONNX models in Rust |
2.5K | 264 | 08/27/26 | Apache-2.0 |
| 94 |
llm-checker
Advanced CLI tool that scans your hardware and tells you exactly which LLM or sLLM models you can run locally, with full Ollama integration. |
2.9K | 198 | 08/28/26 | Other |
| 95 |
sie
Open-source inference server and production cluster for all the models your agent needs. |
2.9K | 287 | 08/27/26 | Apache-2.0 |
| 96 |
onnx-tensorrt
ONNX-TensorRT: TensorRT backend for ONNX |
3.2K | 548 | 08/03/26 | Apache-2.0 |
| 97 |
hls4ml
Machine learning on FPGAs using HLS |
2.1K | 583 | 08/28/26 | Apache-2.0 |
| 98 |
deepdetect
Deep Learning Server and CLI for Torch and TensorRT |
2.6K | 546 | 08/28/26 | Other |
| 99 |
Atomic-Chat
Local AI app and inference engine for agents. Run open-weight LLMs locally — private, 100% offline on your computer. Join our Discord: https://discord.com/invite/8wGSsvmg4V |
1.4K | 160 | 08/28/26 | Other |
| 100 |
tokenspeed
TokenSpeed is a speed-of-light LLM inference engine. |
2.0K | 260 | 08/29/26 | MIT |