Inference & Local Deploy

High-throughput serving and local runtimes — Ollama, vLLM, Triton, and more.

183 projects

See methodology for ranking rules; order uses public GitHub metrics within this scenario.

121–140 of 183

Rank Project Stars Forks
121 pruna

Pruna is a model optimization framework built for developers, enabling you to deliver faster, more efficient models with minimal overhead.

1.3K 103
122 local-ai-packaged

Run all your local AI together in one package - Ollama, Supabase, n8n, Open WebUI, and more!

3.8K 1.4K
123 kubeai

AI Inference Operator for Kubernetes. The easiest way to serve ML models in production. Supports VLMs, LLMs, embeddings, and speech-to-text.

1.3K 134
124 mleap

MLeap: Deploy ML Pipelines to Production

1.5K 316
125 seldon-core

An MLOps framework to package, deploy, monitor and manage thousands of production machine learning models

4.8K 868
126 kvcached

Virtualized Elastic KV Cache for Dynamic GPU Sharing and Beyond

1.1K 133
127 DeepSeek-v4-Flash-DSpark-2x-DGX-Spark

DeepSeek-v4-Flash 0731 recipe for 2x DGX Sparks

1.1K 150
128 wllama

WebAssembly binding for llama.cpp - Enabling on-browser LLM inference

1.2K 119
129 kvpress

LLM KV cache compression made easy

1.2K 172
130 llama.rn

React Native binding of llama.cpp

1.0K 115
131 locally-uncensored

Plug-and-play local AI studio: uncensored chat, image & video generation, coding agent. Runs abliterated LLMs + ComfyUI 100% offline. One installer, no Docker, no cloud.

1.2K 199
132 Comfy-WaveSpeed

https://wavespeed.ai/ [WIP] The all in one inference optimization solution for ComfyUI, universal, flexible, and fast.

1.2K 68
133 gollama

Go manage your Ollama models

1.8K 109
134 paddler

Open-source LLM/VLM load balancer and serving platform for self-hosting LLMs (and VLMs) at scale 🏓🦙 Alternative to projects like llm-d, Docker Model Runner, etc but with less moving parts and simple deployments built around ggml ecosystem. Runs on CPU and GPU.

1.7K 97
135 BrowserAI

Run local LLMs like llama, deepseek-distill, kokoro and more inside your browser

1.4K 137
136 OnnxStream

Lightweight inference library for ONNX files, written in C++. It can run Stable Diffusion XL 1.0 on a RPI Zero 2 (or in 298MB of RAM) but also Mistral 7B on desktops and servers. ARM, x86, WASM, RISC-V supported. Accelerated by XNNPACK. Python, C# and JS(WASM) bindings available.

2.1K 99
137 parallax

Parallax is a distributed model serving framework that lets you build your own AI cluster anywhere

1.4K 146
138 ollama-js

Ollama JavaScript library

4.4K 471
139 HuggingFaceModelDownloader

Simple go utility to download HuggingFace Models and Datasets

1.2K 124
140 infinity

Infinity is a high-throughput, low-latency serving engine for text-embeddings, reranking models, clip, clap and colpali

2.9K 198