Inference & Local Deploy

High-throughput serving and local runtimes — Ollama, vLLM, Triton, and more.

183 projects

See methodology for ranking rules; order uses public GitHub metrics within this scenario.

101–120 of 183

Rank Project Stars Forks
101 Bonsai-demo

Bonsai Demo

2.3K 233
102 spark-vllm-docker

Docker configuration for running VLLM on dual DGX Sparks

2.2K 373
103 club-3090

Community recipes for serving LLMs on RTX 3090/4090/5090 CUDA gpus. Multi-engine (vLLM, llama.cpp, ik_llama) and model-agnostic. Currently shipping Qwen3.6-27B Qwen3.6 35B Gemma 4 26B Gemma 4 31B configs for 1× and 2× cards.

2.1K 128
104 vllm-metal

Community maintained hardware plugin for vLLM on Apple Silicon

1.7K 233
105 lorax

Multi-LoRA inference server that scales to 1000s of fine-tuned LLMs

3.8K 326
106 FastFlowLM

Run LLMs on AMD Ryzen™ AI NPUs in minutes; purpose-built and deeply optimized for the AMD NPUs.

1.8K 145
107 FastFlowLM

Run LLMs on AMD Ryzen™ AI NPUs in minutes; purpose-built and deeply optimized for the AMD NPUs.

1.8K 145
108 GPTQModel

LLM model quantization (compression) toolkit with HW acceleration support for Nvidia, AMD, Intel GPU and Intel/AMD/Apple CPU via HF, vLLM, and SGLang.

1.2K 204
109 react-native-executorch

Declarative way to run AI models in React Native on device, powered by ExecuTorch.

1.7K 95
110 truss

The simplest way to serve AI/ML models in production

1.2K 122
111 vllm-mlx

High-performance OpenAI and Anthropic compatible LLM inference server for Apple Silicon. Native MLX, continuous batching, multimodal models, MCP tool calling, and Claude Code support.

1.6K 219
112 auto-round

A SOTA quantization toolkit for high-accuracy low-bit LLM inference, seamlessly optimized for CPU/XPU/CUDA, with multi-datatype support and full compatibility with vLLM, SGLang, and Transformers|简洁且高效的量化工具包

1.6K 168
113 nndeploy

An Easy-to-Use and High-Performance AI Deployment Framework

1.9K 233
114 node-llama-cpp

Run AI models locally on your machine with node.js bindings for llama.cpp. Enforce a JSON schema on the model output on the generation level

2.2K 215
115 distributed-llama

Distributed LLM inference. Connect home devices into a powerful cluster to accelerate LLM inference. More devices means faster inference.

3.0K 248
116 uzu

A high-performance inference engine for AI models

1.7K 73
117 sonar

Large-scale LLM inference engine

1.8K 208
118 local-studio

Control panel for VLLM, Sglang, llama.cpp, exllamav3

1.7K 155
119 ComfyUI-Docker

🐳Dockerfile for 🎨ComfyUI. | 容器镜像与启动脚本

1.6K 232
120 mllm

Fast Multimodal LLM on Mobile Devices

1.6K 213