Inference & Local Deploy
High-throughput serving and local runtimes — Ollama, vLLM, Triton, and more.
183 projects
See methodology for ranking rules; order uses public GitHub metrics within this scenario.
| Rank | Project | Stars | Forks | Updated | License |
|---|---|---|---|---|---|
| 121 |
pruna
Pruna is a model optimization framework built for developers, enabling you to deliver faster, more efficient models with minimal overhead. |
1.3K | 103 | 08/27/26 | Apache-2.0 |
| 122 |
local-ai-packaged
Run all your local AI together in one package - Ollama, Supabase, n8n, Open WebUI, and more! |
3.8K | 1.4K | 05/25/26 | Apache-2.0 |
| 123 |
kubeai
AI Inference Operator for Kubernetes. The easiest way to serve ML models in production. Supports VLMs, LLMs, embeddings, and speech-to-text. |
1.3K | 134 | 08/24/26 | Apache-2.0 |
| 124 |
mleap
MLeap: Deploy ML Pipelines to Production |
1.5K | 316 | 07/21/26 | Apache-2.0 |
| 125 |
seldon-core
An MLOps framework to package, deploy, monitor and manage thousands of production machine learning models |
4.8K | 868 | 03/23/26 | Other |
| 126 |
kvcached
Virtualized Elastic KV Cache for Dynamic GPU Sharing and Beyond |
1.1K | 133 | 08/23/26 | Apache-2.0 |
| 127 |
DeepSeek-v4-Flash-DSpark-2x-DGX-Spark
DeepSeek-v4-Flash 0731 recipe for 2x DGX Sparks |
1.1K | 150 | 08/29/26 | MIT |
| 128 |
wllama
WebAssembly binding for llama.cpp - Enabling on-browser LLM inference |
1.2K | 119 | 08/27/26 | MIT |
| 129 |
kvpress
LLM KV cache compression made easy |
1.2K | 172 | 08/18/26 | Apache-2.0 |
| 130 |
llama.rn
React Native binding of llama.cpp |
1.0K | 115 | 08/28/26 | MIT |
| 131 |
locally-uncensored
Plug-and-play local AI studio: uncensored chat, image & video generation, coding agent. Runs abliterated LLMs + ComfyUI 100% offline. One installer, no Docker, no cloud. |
1.2K | 199 | 08/24/26 | AGPL-3.0 |
| 132 |
Comfy-WaveSpeed
https://wavespeed.ai/ [WIP] The all in one inference optimization solution for ComfyUI, universal, flexible, and fast. |
1.2K | 68 | 08/21/26 | MIT |
| 133 |
gollama
Go manage your Ollama models |
1.8K | 109 | 07/20/26 | MIT |
| 134 |
paddler
Open-source LLM/VLM load balancer and serving platform for self-hosting LLMs (and VLMs) at scale 🏓🦙 Alternative to projects like llm-d, Docker Model Runner, etc but with less moving parts and simple deployments built around ggml ecosystem. Runs on CPU and GPU. |
1.7K | 97 | 07/19/26 | Apache-2.0 |
| 135 |
BrowserAI
Run local LLMs like llama, deepseek-distill, kokoro and more inside your browser |
1.4K | 137 | 07/21/26 | MIT |
| 136 |
OnnxStream
Lightweight inference library for ONNX files, written in C++. It can run Stable Diffusion XL 1.0 on a RPI Zero 2 (or in 298MB of RAM) but also Mistral 7B on desktops and servers. ARM, x86, WASM, RISC-V supported. Accelerated by XNNPACK. Python, C# and JS(WASM) bindings available. |
2.1K | 99 | 06/18/26 | Other |
| 137 |
parallax
Parallax is a distributed model serving framework that lets you build your own AI cluster anywhere |
1.4K | 146 | 07/01/26 | Apache-2.0 |
| 138 |
ollama-js
Ollama JavaScript library |
4.4K | 471 | 02/18/26 | MIT |
| 139 |
HuggingFaceModelDownloader
Simple go utility to download HuggingFace Models and Datasets |
1.2K | 124 | 06/13/26 | Apache-2.0 |
| 140 |
infinity
Infinity is a high-throughput, low-latency serving engine for text-embeddings, reranking models, clip, clap and colpali |
2.9K | 198 | 03/24/26 | MIT |