Inference & Local Deploy

High-throughput serving and local runtimes — Ollama, vLLM, Triton, and more.

183 projects

See methodology for ranking rules; order uses public GitHub metrics within this scenario.

61–80 of 183

Rank Project Stars Forks
61 Paddle-Lite

PaddlePaddle High Performance Deep Learning Inference Engine for Mobile and Edge (飞桨高性能深度学习端侧推理引擎)

7.3K 1.6K
62 FastDeploy

High-performance Inference and Deployment Toolkit for LLMs and VLMs based on PaddlePaddle

3.7K 760
63 fastllm

fastllm是后端无依赖的高性能大模型推理库。同时支持张量并行推理稠密模型和混合模式推理MOE模型,任意10G以上显卡即可推理满血DeepSeek。双路9004/9005服务器+单显卡部署DeepSeek满血满精度原版模型,单并发20tps;INT4量化模型单并发30tps,多并发可达60+。

4.9K 485
64 ODS

Turn your PC, Mac, or Linux box into an AI server. LLM inference, chat UI, voice, agents, workflows, RAG, and image generation.

4.9K 741
65 RWKV-Runner

A RWKV management and startup tool, full automation, only 8MB. And provides an interface compatible with the OpenAI API. RWKV is a large language model that is fully open source and available for commercial use.

6.5K 600
66 llm-compressor

Transformers-compatible library for applying various compression algorithms to LLMs for optimized deployment with vLLM

3.7K 644
67 lms

LM Studio CLI

5.2K 447
68 jetson-containers

Machine Learning Containers for NVIDIA Jetson and JetPack-L4T

4.8K 845
69 CTranslate2

Fast inference engine for Transformer models

4.7K 522
70 lollms-webui

Lord of Large Language and Multi modal Systems Web User Interface

4.8K 587
71 text-embeddings-inference

A blazing fast inference solution for text embeddings models

5.0K 426
72 LightLLM

LightLLM is a Python-based LLM (Large Language Model) inference and serving framework, notable for its lightweight design, easy scalability, and high-speed performance.

4.3K 357
73 optimum

🚀 Accelerate inference and training of 🤗 Transformers, Diffusers, TIMM and Sentence Transformers with easy to use hardware optimization tools

3.5K 680
74 ComputeLibrary

The Compute Library is a set of computer vision and machine learning functions optimised for both Arm CPUs and GPUs using SIMD technologies.

3.2K 819
75 Model-Optimizer

A unified library of state-of-the-art model optimization techniques (quantization, pruning, NAS, distillation, speculative decoding) to compress and accelerate deep learning models for deployment in frameworks like TensorRT-LLM, vLLM, and SGLang.

3.6K 565
76 zml

Any model. Any hardware. Zero compromise. Built with @ziglang / @openxla / MLIR / @bazelbuild

4.0K 179
77 ramalama

RamaLama is an open-source developer tool that simplifies the local serving of AI models from any source and facilitates their use for inference in production, all through the familiar language of containers.

3.0K 360
78 TensorRT

PyTorch/TorchScript/FX compiler for NVIDIA GPUs using TensorRT

3.0K 409
79 LitServe

A minimal Python framework for building custom AI inference servers with full control over logic, batching, and scaling.

3.9K 301
80 Rapid-MLX

The fastest local AI engine for Apple Silicon. 4.2x faster than Ollama, 0.08s cached TTFT, 100% tool calling. 17 tool parsers, prompt cache, reasoning separation, cloud routing. Drop-in OpenAI replacement. Works with Claude Code, Cursor, Aider.

3.6K 402