#106 · Primary category: Inference & Local Deploy

FastFlowLM

amd deepseek llama llm npu

Run LLMs on AMD Ryzen™ AI NPUs in minutes; purpose-built and deeply optimized for the AMD NPUs.

Project last updated:08/28/26

GitHub Stars

1.8K

Forks

145

Contributors

30

License

MIT

Why we included this project

Anyone running a Ryzen AI laptop or mini-PC with an XDNA2 NPU has probably noticed that most local-LLM tooling ignores the NPU and leans on a GPU instead. FastFlowLM is built for that hardware. A small Windows installer gets a model serving on the NPU within 20 seconds, covering vision, audio, embedding, and MoE model families with context windows up to 256k tokens. Because it targets the NPU rather than a GPU, power draw is over ten times lower, which matters for battery-powered laptops and cramped edge boxes. The project publishes benchmark pages, a model list, and a Linux getting-started guide, so you can check real token throughput before committing hardware to a workload. For teams that want LLM inference to stay on-device, this is a practical way to get there on AMD hardware.

Articles for this project

No articles for this project yet.

To suggest a topic or contribute an article, contact us.

Related projects in this category