Inference & Local Deploy

High-throughput serving and local runtimes — Ollama, vLLM, Triton, and more.

183 projects

See methodology for ranking rules; order uses public GitHub metrics within this scenario.

161–180 of 183

Rank Project Stars Forks
161 cortex

Production infrastructure for machine learning at scale

8.0K 594
162 streaming-llm

[ICLR 2024] Efficient Streaming Language Models with Attention Sinks

7.3K 399
163 mmdeploy

OpenMMLab Model Deployment Framework

3.1K 715
164 mace

MACE is a deep learning inference framework optimized for mobile heterogeneous computing platforms.

5.0K 819
165 gpu_poor

Calculate token/s & GPU memory requirement for any LLM. Supports llama.cpp/ggml/bnb/QLoRA quantization

1.4K 88
166 api-for-open-llm

Openai style api for open large language models, using LLMs just as chatgpt! Support for LLaMA, LLaMA-2, BLOOM, Falcon, Baichuan, Qwen, Xverse, SqlCoder, CodeLLaMA, ChatGLM, ChatGLM2, ChatGLM3 etc.

2.5K 271
167 transformer-deploy

Efficient, scalable and enterprise-grade CPU/GPU inference server for 🤗 Hugging Face transformer models 🚀

1.7K 153
168 chatglm.cpp

C++ implementation of ChatGLM-6B & ChatGLM2-6B & ChatGLM3 & GLM4(V)

3.0K 326
169 llama.go

llama.go is like llama.cpp in pure Golang!

1.4K 73
170 Medusa

Medusa: Simple Framework for Accelerating LLM Generation with Multiple Decoding Heads

2.8K 205
171 FreedomGPT

This codebase is for a React and Electron-based app that executes the FreedomGPT LLM locally (offline and private) on Mac and Windows using a chat-based interface

2.7K 363
172 mixtral-offloading

Run Mixtral-8x7B models in Colab or consumer desktops

2.3K 225
173 punica

Serving multiple LoRA finetuned LLM as one

1.2K 64
174 ppq

PPL Quantization Tool (PPQ) is a powerful offline neural network quantization tool.

1.8K 285
175 WebGPT

Run GPT model on the browser with WebGPU. An implementation of GPT inference in less than ~1500 lines of vanilla Javascript.

3.8K 223
176 llama2-webui

Run any Llama 2 locally with gradio UI on GPU or CPU from anywhere (Linux/Windows/Mac). Use `llama2-wrapper` as your local llama2 backend for Generative Agents/Apps.

1.9K 199
177 ctransformers

Python bindings for the Transformer models implemented in C/C++ using GGML library.

1.9K 143
178 budgetml

Deploy a ML inference service on a budget in less than 10 lines of code.

1.3K 65
179 text-generation-webui-colab

A colab gradio web UI for running Large Language Models

2.1K 356
180 Bender

Easily craft fast Neural Networks on iOS! Use TensorFlow models. Metal under the hood.

1.8K 89