Inference & Local Deploy

High-throughput serving and local runtimes — Ollama, vLLM, Triton, and more.

183 projects

See methodology for ranking rules; order uses public GitHub metrics within this scenario.

81–100 of 183

Rank Project Stars Forks
81 tract

Tiny, no-nonsense, self-contained, Tensorflow and ONNX inference

3.0K 281
82 production-stack

vLLM’s reference system for K8S-native cluster-wide deployment with community-driven performance optimization

2.5K 476
83 mesh-llm

Distributed AI/LLM for the people. Share compute privately or publicly to power your agents and chat.

3.3K 402
84 spiceai

Add a real-time analytics node to your operational database. Spice is a portable, accelerated SQL query, search, and LLM-inference engine in Rust for data-grounded AI apps and agents.

3.1K 224
85 optillm

Optimizing inference proxy for LLMs

4.3K 385
86 harbor

Stop configuring your AI stack. Start using it. One command brings a complete pre-wired LLM stack with hundreds of services to explore.

3.2K 224
87 chitu

High-performance inference framework for large language models, focusing on efficiency, flexibility, and availability.

3.0K 256
88 inference

Turn any computer or edge device into a command center for your computer vision projects.

2.4K 311
89 claude-code-local

Run Claude Code fully on-device with local AI on Apple Silicon via an MLX-native Anthropic-API server, offering private, offline, airgap-ready inference.

3.2K 620
90 algernon

Small self-contained pure-Go web server with Lua, Teal, Markdown, HTTP/2, QUIC, Redis, TypeScript, npm-less React 19, SQLite, and PostgreSQL support ++

3.0K 151
91 opyrator

🪄 Turns your machine learning code into microservices with web API, interactive GUI, and more.

3.1K 167
92 Olive

Olive: Simplify ML Model Finetuning, Conversion, Quantization, and Optimization for CPUs, GPUs and NPUs.

2.4K 312
93 ort

Fast ML inference & training for ONNX models in Rust

2.5K 264
94 llm-checker

Advanced CLI tool that scans your hardware and tells you exactly which LLM or sLLM models you can run locally, with full Ollama integration.

2.9K 198
95 sie

Open-source inference server and production cluster for all the models your agent needs.

2.9K 287
96 onnx-tensorrt

ONNX-TensorRT: TensorRT backend for ONNX

3.2K 548
97 hls4ml

Machine learning on FPGAs using HLS

2.1K 583
98 deepdetect

Deep Learning Server and CLI for Torch and TensorRT

2.6K 546
99 Atomic-Chat

Local AI app and inference engine for agents. Run open-weight LLMs locally — private, 100% offline on your computer. Join our Discord: https://discord.com/invite/8wGSsvmg4V

1.4K 160
100 tokenspeed

TokenSpeed is a speed-of-light LLM inference engine.

2.0K 260