Inference & Local Deploy

High-throughput serving and local runtimes — Ollama, vLLM, Triton, and more.

183 projects

See methodology for ranking rules; order uses public GitHub metrics within this scenario.

1–20 of 183

Rank Project Stars Forks
1 ollama

Get up and running with Kimi-K2.6, GLM-5.2, MiniMax, DeepSeek, gpt-oss, Qwen, Gemma and other models.

179.7K 17.6K
2 llama.cpp

LLM inference in C/C++

126.3K 22.4K
3 vllm

A high-throughput and memory-efficient inference and serving engine for LLMs

90.4K 21.4K
4 gpt4all

GPT4All: Run Local LLMs on Any Device. Open-source and available for commercial use.

77.4K 8.3K
5 LocalAI

LocalAI is the open-source AI engine. Run any model - LLMs, vision, voice, image, video - on any hardware. No GPU required.

48.8K 4.4K
6 textgen

Open-source desktop app for local LLMs. Text, vision, tool-calling, OpenAI/Anthropic-compatible API. 100% private.

47.6K 6.0K
7 exo

Run frontier AI locally.

47.1K 3.5K
8 sglang

SGLang is a high-performance serving framework for large language models and multimodal models.

32.8K 8.3K
9 BitNet

Official inference framework for 1-bit LLMs

40.2K 3.7K
10 llmfit

Hundreds of models & providers. One command to find what runs on your hardware.

34.5K 2.2K
11 airllm

AirLLM 70B inference with single 4GB GPU

33.1K 3.5K
12 modular

The Modular Platform (includes MAX & Mojo)

29.3K 3.1K
13 onnxruntime

ONNX Runtime: cross-platform, high performance ML inferencing and training accelerator

21.7K 4.2K
14 ncnn

ncnn is a high-performance neural network inference framework optimized for the mobile platform

23.8K 4.5K
15 llamafile

Distribute and run LLMs with a single file.

25.7K 1.6K
16 mlc-llm

Universal LLM Deployment Engine with ML Compilation

23.1K 2.1K
17 omlx

LLM inference server with continuous batching & SSD caching for Apple Silicon — managed from the macOS menu bar

21.0K 1.8K
18 ktransformers

A Flexible Framework for Experiencing Heterogeneous LLM Inference/Fine-tune Optimizations

19.3K 1.5K
19 web-llm

High-performance In-browser LLM Inference Engine

18.6K 1.4K
20 TensorRT-LLM

Optimize and deploy large language models with state-of-the-art inference performance on NVIDIA GPUs.

14.5K 2.7K
< Previous
/ 10
Next >