Inference & Local Deploy

High-throughput serving and local runtimes — Ollama, vLLM, Triton, and more.

183 projects

See methodology for ranking rules; order uses public GitHub metrics within this scenario.

141–160 of 183

Rank Project Stars Forks
141 ollama-docker

Welcome to the Ollama Docker Compose Setup! This project simplifies the deployment of Ollama using Docker Compose, making it easy to run Ollama with all its dependencies in a containerized environment

1.4K 254
142 SageAttention

[ICLR2025, ICML2025, NeurIPS2025 Spotlight] Quantized Attention achieves speedup of 2-5x compared to FlashAttention, without losing end-to-end metrics across language, image, and video models.

3.7K 493
143 TensorRT-YOLO

🚀 Easier & Faster YOLO Deployment Toolkit for NVIDIA 🛠️

1.9K 194
144 turboquant

TurboQuant: Near-optimal KV cache quantization for LLM inference (3-bit keys, 2-bit values) with Triton kernels + vLLM integration

1.7K 195
145 rotorquant

KV cache compression via block-diagonal rotation. Beats TurboQuant: better PPL (6.91 vs 7.07), 28% faster decode, 5.3x faster prefill, 44x fewer params. Drop-in llama.cpp integration.

1.0K 89
146 deepseek-ocr.rs

Rust multi‑backend OCR/VLM engine (DeepSeek‑OCR-1/2, PaddleOCR‑VL, DotsOCR) with DSQ quantization and an OpenAI‑compatible server & CLI – run locally without Python.

2.2K 169
147 gaianet-node

Install, run and deploy your own decentralized AI agent service

5.0K 326
148 picolm

Run a 1-billion parameter LLM on a $10 board with 256MB RAM

1.9K 240
149 LLMFarm

llama and other large language models on iOS and MacOS offline using GGML library.

2.1K 180
150 LlamaEdge

The easiest & fastest way to run customized and fine-tuned LLMs locally or on the edge

1.7K 149
151 comfyui-deploy

An open source `vercel` like deployment platform for Comfy UI

1.5K 219
152 hummingbird

Hummingbird compiles trained ML models into tensor computation for faster inference.

3.5K 292
153 Jlama

Jlama is a modern LLM inference engine for Java

1.3K 164
154 dalai

The simplest way to run LLaMA on your local machine

12.9K 1.3K
155 petals

🌸 Run LLMs at home, BitTorrent-style. Fine-tuning and inference up to 10x faster than offloading

10.5K 643
156 DeepSpeed-MII

MII makes low-latency and high-throughput inference possible, powered by DeepSpeed.

2.1K 192
157 Tengine

Tengine is a lite, high performance, modular inference engine for embedded device

4.5K 982
158 TurboTransformers

a fast and user-friendly runtime for transformer inference (Bert, Albert, GPT2, Decoders, etc) on CPU and GPU.

1.5K 208
159 uTensor

TinyML AI inference library

1.9K 250
160 rwkv.cpp

INT4/INT5/INT8 and FP16 inference on CPU for RWKV language model

1.6K 129