#106 · Primary category: Inference & Local Deploy
FastFlowLM
Run LLMs on AMD Ryzen™ AI NPUs in minutes; purpose-built and deeply optimized for the AMD NPUs.
Project last updated:08/28/26
GitHub Stars
1.8K
Forks
145
Contributors
30
License
MIT
Why we included this project
Anyone running a Ryzen AI laptop or mini-PC with an XDNA2 NPU has probably noticed that most local-LLM tooling ignores the NPU and leans on a GPU instead. FastFlowLM is built for that hardware. A small Windows installer gets a model serving on the NPU within 20 seconds, covering vision, audio, embedding, and MoE model families with context windows up to 256k tokens. Because it targets the NPU rather than a GPU, power draw is over ten times lower, which matters for battery-powered laptops and cramped edge boxes. The project publishes benchmark pages, a model list, and a Linux getting-started guide, so you can check real token throughput before committing hardware to a workload. For teams that want LLM inference to stay on-device, this is a practical way to get there on AMD hardware.
Articles for this project
No articles for this project yet.
To suggest a topic or contribute an article, contact us.
Related projects in this category
ollama
Get up and running with Kimi-K2.6, GLM-5.2, MiniMax, DeepSeek, gpt-oss, Qwen, Gemma and other models.
llama.cpp
LLM inference in C/C++
vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
gpt4all
GPT4All: Run Local LLMs on Any Device. Open-source and available for commercial use.
LocalAI
LocalAI is the open-source AI engine. Run any model - LLMs, vision, voice, image, video - on any hardware. No GPU required.