Inference & Local Deploy
High-throughput serving and local runtimes — Ollama, vLLM, Triton, and more.
183 projects
See methodology for ranking rules; order uses public GitHub metrics within this scenario.
| Rank | Project | Stars | Forks | Updated | License |
|---|---|---|---|---|---|
| 161 |
cortex
Production infrastructure for machine learning at scale |
8.0K | 594 | 06/12/24 | Apache-2.0 |
| 162 |
streaming-llm
[ICLR 2024] Efficient Streaming Language Models with Attention Sinks |
7.3K | 399 | 07/11/24 | MIT |
| 163 |
mmdeploy
OpenMMLab Model Deployment Framework |
3.1K | 715 | 09/30/24 | Apache-2.0 |
| 164 |
mace
MACE is a deep learning inference framework optimized for mobile heterogeneous computing platforms. |
5.0K | 819 | 06/17/24 | Apache-2.0 |
| 165 |
gpu_poor
Calculate token/s & GPU memory requirement for any LLM. Supports llama.cpp/ggml/bnb/QLoRA quantization |
1.4K | 88 | 12/03/24 | Other |
| 166 |
api-for-open-llm
Openai style api for open large language models, using LLMs just as chatgpt! Support for LLaMA, LLaMA-2, BLOOM, Falcon, Baichuan, Qwen, Xverse, SqlCoder, CodeLLaMA, ChatGLM, ChatGLM2, ChatGLM3 etc. |
2.5K | 271 | 09/26/24 | Apache-2.0 |
| 167 |
transformer-deploy
Efficient, scalable and enterprise-grade CPU/GPU inference server for 🤗 Hugging Face transformer models 🚀 |
1.7K | 153 | 10/23/24 | Apache-2.0 |
| 168 |
chatglm.cpp
C++ implementation of ChatGLM-6B & ChatGLM2-6B & ChatGLM3 & GLM4(V) |
3.0K | 326 | 07/31/24 | MIT |
| 169 |
llama.go
llama.go is like llama.cpp in pure Golang! |
1.4K | 73 | 09/20/24 | Other |
| 170 |
Medusa
Medusa: Simple Framework for Accelerating LLM Generation with Multiple Decoding Heads |
2.8K | 205 | 06/25/24 | Apache-2.0 |
| 171 |
FreedomGPT
This codebase is for a React and Electron-based app that executes the FreedomGPT LLM locally (offline and private) on Mac and Windows using a chat-based interface |
2.7K | 363 | 06/13/24 | GPL-3.0 |
| 172 |
mixtral-offloading
Run Mixtral-8x7B models in Colab or consumer desktops |
2.3K | 225 | 04/08/24 | MIT |
| 173 |
punica
Serving multiple LoRA finetuned LLM as one |
1.2K | 64 | 05/08/24 | Apache-2.0 |
| 174 |
ppq
PPL Quantization Tool (PPQ) is a powerful offline neural network quantization tool. |
1.8K | 285 | 03/28/24 | Apache-2.0 |
| 175 |
WebGPT
Run GPT model on the browser with WebGPU. An implementation of GPT inference in less than ~1500 lines of vanilla Javascript. |
3.8K | 223 | 01/12/24 | Other |
| 176 |
llama2-webui
Run any Llama 2 locally with gradio UI on GPU or CPU from anywhere (Linux/Windows/Mac). Use `llama2-wrapper` as your local llama2 backend for Generative Agents/Apps. |
1.9K | 199 | 03/22/24 | MIT |
| 177 |
ctransformers
Python bindings for the Transformer models implemented in C/C++ using GGML library. |
1.9K | 143 | 01/28/24 | MIT |
| 178 |
budgetml
Deploy a ML inference service on a budget in less than 10 lines of code. |
1.3K | 65 | 02/12/24 | Apache-2.0 |
| 179 |
text-generation-webui-colab
A colab gradio web UI for running Large Language Models |
2.1K | 356 | 12/22/23 | Unlicense |
| 180 |
Bender
Easily craft fast Neural Networks on iOS! Use TensorFlow models. Metal under the hood. |
1.8K | 89 | 11/07/23 | MIT |