#107 · Primary category: Inference & Local Deploy

FastFlowLM

amd deepseek llama llm npu

Run LLMs on AMD Ryzen™ AI NPUs in minutes; purpose-built and deeply optimized for the AMD NPUs.

Project last updated:08/28/26

GitHub Stars

1.8K

Forks

145

Contributors

30

License

MIT

Why we included this project

Developers with Ryzen AI laptops that have an XDNA2 NPU will find FastFlowLM a straightforward way to run LLMs without a discrete GPU. It installs as a small signed package, pulls models like Ollama, and exposes an OpenAI-compatible server plus a CLI for streaming tokens. The runtime covers text, vision, audio, embeddings, and MoE models, with context windows up to 256k tokens, so it suits local chat, RAG pipelines, and on-device transcription. The real value is that it puts otherwise idle NPU hardware to work as a low-power inference engine, which matters for private, offline use on machines people already own. Teams comparing local deployment options on AMD hardware will appreciate the drop-in API and benchmark telemetry for side-by-side tests against GPU-based stacks.

Articles for this project

No articles for this project yet.

To suggest a topic or contribute an article, contact us.

Related projects in this category