#36 · Primary category: AI Gateway & API Infrastructure
llama-swap
Reliable model swapping for any local OpenAI/Anthropic compatible server - llama.cpp, vllm, etc
Project last updated:08/29/26
GitHub Stars
5.5K
Forks
444
Contributors
70
License
MIT
Why we included this project
Keeping several local models on one machine usually means running separate servers and configs. llama-swap sits in front of any OpenAI- or Anthropic-compatible backend and loads the model you ask for, then unloads it to free VRAM when you're done. One binary and one config file replace the mix of llama.cpp, vLLM, tabbyAPI, or stable-diffusion.cpp instances you'd otherwise manage. If you rotate between models during the day, your clients keep hitting the same endpoint while the upstream server changes underneath. A small web UI and a handful of admin endpoints show what's loaded and let you force-unload models manually.
Articles for this project
No articles for this project yet.
To suggest a topic or contribute an article, contact us.
Related projects in this category
OmniRoute
Open-source AI gateway: one endpoint, 290+ providers (90+ free), 500+ models, auto-fallback, token compression, MCP/A2A, works with Claude Code, Codex, Cursor, Cline, Copilot.
CLIProxyAPI
Wrap Antigravity, ChatGPT Codex, Claude Code, Grok Build as an OpenAI/Gemini/Claude/Codex compatible API service, allowing you to enjoy the free Gemini 3.1 Pro, GPT 5.6 Series, Grok 4.5, Claude model through API
kong
🦍 The API and AI Gateway
novu
The open-source communication infrastructure for agents and products
GPT_API_free
Free API for large models, supports GPT, DeepSeek, etc., 10k free points daily, paid plans at 10-20% of official price.