#107 · Primary category: Inference & Local Deploy
FastFlowLM
Run LLMs on AMD Ryzen™ AI NPUs in minutes; purpose-built and deeply optimized for the AMD NPUs.
Project last updated:08/28/26
GitHub Stars
1.8K
Forks
145
Contributors
30
License
MIT
Why we included this project
Developers with Ryzen AI laptops that have an XDNA2 NPU will find FastFlowLM a straightforward way to run LLMs without a discrete GPU. It installs as a small signed package, pulls models like Ollama, and exposes an OpenAI-compatible server plus a CLI for streaming tokens. The runtime covers text, vision, audio, embeddings, and MoE models, with context windows up to 256k tokens, so it suits local chat, RAG pipelines, and on-device transcription. The real value is that it puts otherwise idle NPU hardware to work as a low-power inference engine, which matters for private, offline use on machines people already own. Teams comparing local deployment options on AMD hardware will appreciate the drop-in API and benchmark telemetry for side-by-side tests against GPU-based stacks.
Articles for this project
No articles for this project yet.
To suggest a topic or contribute an article, contact us.
Related projects in this category
ollama
Get up and running with Kimi-K2.6, GLM-5.2, MiniMax, DeepSeek, gpt-oss, Qwen, Gemma and other models.
llama.cpp
LLM inference in C/C++
vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
gpt4all
GPT4All: Run Local LLMs on Any Device. Open-source and available for commercial use.
LocalAI
LocalAI is the open-source AI engine. Run any model - LLMs, vision, voice, image, video - on any hardware. No GPU required.