# FastFlowLM: an NPU-first LLM runtime for AMD Ryzen AI, installed in one command

> FastFlowLM (FLM) runs LLMs, vision, audio and embedding models on AMD Ryzen AI XDNA2 NPUs through a single-command CLI and a local server. The install is short; the hardware and driver requirements are not negotiable.

**ROCm/FastFlowLM** — Run LLMs on AMD Ryzen™ AI NPUs in minutes; purpose-built and deeply optimized for the AMD NPUs.

- Repository: https://github.com/ROCm/FastFlowLM
- Website: http://fastflowlm.com/
- Stars: 1,913 · Forks: 156
- Language: C++
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/rocm-fastflowlm

## What FastFlowLM solves, and for whom

Most local inference stacks assume a discrete GPU or a CPU with enough memory bandwidth to make token generation tolerable. FastFlowLM takes the opposite position: it targets the NPU that already sits in AMD's Ryzen AI processors. The README states the project runs models fully on the AMD Ryzen AI NPU with no GPU or CPU load, and describes the runtime as lightweight at 17 MB with an install that completes within 20 seconds.

The audience is narrow and specific. The README says FLM supports all Ryzen AI Series chips with XDNA2 NPUs, naming Strix, Strix Halo, Kraken and Gorgon Point. If your laptop or mini PC has one of those parts, you are the target user. If it does not, nothing here applies to you, and no amount of configuration will change that, because the accelerated kernels are compiled for that hardware.

The pitch rests on power, not peak throughput. The README claims the NPU path is faster and over 10x more power-efficient than the GPU alternative, and advertises context windows up to 256k tokens on models such as Qwen3-4B-Thinking-2507. Those are vendor figures hosted on the project's own benchmark page, so treat them as claims to reproduce on your own workload rather than as neutral measurements.

## The NPU-first architecture: a small runtime over prebuilt kernels

FLM is not a training framework and not a general inference engine. The repository layout makes the split visible: src/ holds the C++ runtime, third_party/ vendors dependencies, and the model weights and NPU-accelerated binary kernels are downloaded separately from HuggingFace at first use. The README is explicit that internet access to HuggingFace is required to fetch the optimized model kernels, which is why the first run of a model is slower than later ones.

That download step is also the main failure mode. The README warns that downloads from HuggingFace sometimes arrive corrupted and gives the remedy: rerun the pull with the force flag, for example flm pull llama3.2:1b --force. There is no documented integrity check you can run yourself, so the project's own answer to a bad kernel is to fetch it again.

On disk, models land in C:\Users\<USER>\.flm\models\ on Windows and under ~/.config/flm/ on Linux. The Windows installer lets you pick a different base folder at install time, and on Linux the FLM_MODEL_PATH environment variable overrides the default. Because kernels are tied to model tags, switching models means another download rather than a conversion step on your side.

The runtime exposes two surfaces. CLI mode is an interactive chat session; server mode starts a local HTTP server on port 52625 by default and the README says it speaks REST and the OpenAI API. A model tag passed to the serve command only sets the initial model, and the server will switch automatically if a client requests a different one. For anyone wiring FLM into an existing application, that OpenAI-compatible endpoint is the part that matters, since it means an existing client library can point at localhost instead of a hosted API.

## Installing FastFlowLM on Windows and running your first model

The Windows path is a packaged MSI from the releases page. Before you install, check the NPU driver: the README requires version 32.0.203.311 or above and states that earlier versions are no longer supported. You can read the version in Task Manager under Performance, or in Device Manager. AMD's own install documentation for NPU drivers is linked from the README and requires an AMD account; the README also points at Windows Update and the AMD support site as the recommended routes, and flags a third-party forum thread as unverified.

Once the driver is current, install flm-setup.msi and open PowerShell with Win + X then I. The single command to run a model interactively is:

```powershell
flm run llama3.2:1b
```

The first invocation downloads the optimized kernels from HuggingFace, so expect a pause before the prompt appears. Inside the session, /verbose toggles performance reporting and /bye exits. To see what else is available locally, run:

```powershell
flm list
```

For an application rather than a chat session, start the server instead:

```powershell
flm serve llama3.2:1b
```

The server listens on port 52625. Open Task Manager, go to the Performance tab and click NPU to confirm the accelerator is actually carrying the work; that is the quickest sanity check that you are on the NPU path and not falling back to something else.

On Linux the README does not give a single copy-paste recipe. It points to a Linux getting started guide in docs/ and to the Lemonade Server documentation and video walkthrough. The repository also contains a Dockerfile that builds a FastFlowLM build environment with dependencies such as cmake, ninja-build, rustc, cargo and libxrt-dev preinstalled, with /workspace as the working directory. That image is a build environment, not a turnkey runtime image, so it is useful if you intend to compile FLM and not if you just want to chat with a model.

## Where FastFlowLM is the wrong tool

The hardware requirement is absolute. No XDNA2 NPU means no acceleration, and the README does not describe a CPU or GPU fallback path. A desktop with a discrete Radeon or GeForce card is better served by an inference engine built around that GPU, even if the CPU is an AMD part.

The second constraint is the model list. FLM ships prebuilt kernels for a curated set of models rather than accepting arbitrary GGUF files. The README links to a model list on the project site, and the recent releases show the list moving quickly, with v1.0.4 adding Gemma4-12B-IT and v1.0.3 revising Qwen3.5 and Qwen3.6-MoE weights for higher accuracy. If your application depends on a fine-tune or a niche architecture, check the list before you plan around FLM, because you cannot bring your own weights in the way you can with a general runtime.

Third, the driver floor is a real operational cost. The README states plainly that earlier driver versions are no longer supported, which means a fleet of machines needs driver updates before FLM works at all. On managed Windows hardware where driver versions are pinned by an IT policy, that is a blocking dependency rather than a footnote. The README's own caution about unverified third-party driver downloads is worth heeding in that context.

Finally, the project is young in organisational terms. The README notes that FLM joined the ROCm organisation at v1.0.0 in August 2026, following a 2025 university project, and that it was integrated into AMD's Lemonade Server in October 2025. The repository is not archived and the last push was on 2026-09-09, with v1.0.5 released the same day, so the code is moving. Rapid releases also mean the model list and kernel set can shift under you between versions.

## FastFlowLM compared with llama.cpp and Ollama

The honest comparison is about where the arithmetic happens. llama.cpp is a portable C/C++ inference engine with broad backend coverage and a large GGUF ecosystem; you can run it on CPU, CUDA, ROCm or Metal, and you choose the quantisation. Ollama wraps a similar engine in a model-management daemon with a friendly CLI and an OpenAI-compatible API, and it runs on whatever hardware backend it was built for. Both let you run almost any model you can find weights for.

FLM inverts those priorities. It supports one accelerator family, the AMD Ryzen AI NPU, and in exchange it does not ask you to pick a quantisation or tune anything. The README frames this as no model rewrites and no tuning, and lists no low-level tuning required as a highlight. The trade is flexibility for a smaller download, lower power draw and a shorter path from install to a running model on the specific hardware it targets.

If you already run Ollama or llama.cpp on a Ryzen AI laptop, FLM is not a replacement for the whole stack. It is a second backend for the case where you want the NPU doing the work and the GPU and CPU left alone. The README also documents an integration path through Lemonade Server, which is the route to take if you want FLM's NPU kernels behind a server that already manages multiple backends, particularly on Linux where the README directs Linux users to the Lemonade documentation for getting started.

## Licence, packaging and what upgrading costs you

The split licence is the detail to read carefully. The README states that all orchestration code and CLI tools are open source under the MIT License, pointing at LICENSE_RUNTIME.txt, while the NPU-accelerated binary kernels are distributed separately and are described as completely free for any use, including commercial use. The repository also carries a TERMS.md file alongside the licence. The README asks that you credit the project with a specific line, Powered by FastFlowLM with a link to the repository, in your README, project page or product. That is a request as written, not a condition stated in the MIT text, but it is cheap to honour and it is the project's stated expectation. This is a description of what the documents say, not legal advice; if the binary kernel terms matter to your legal team, they should read TERMS.md and the kernel distribution terms directly.

Upgrade cost is dominated by the driver, not the package. Because the README ties the current release to AMD NPU driver 32.0.203.311 or above, a FLM upgrade can force a driver upgrade on every machine you support. Disable the startup version check with FLM_DISABLE_UPDATE_CHECK=1 if you need deterministic behaviour on machines that should not reach out at launch, and pre-seed the model directory if HuggingFace is unreachable from your network, which the README addresses by pointing at an issue thread describing manual model placement. On Linux, the Dockerfile gives you a reproducible build environment, but the runtime still needs the host NPU driver, which a container cannot supply for you.

## Conclusion

Adopt FastFlowLM if you have a Ryzen AI machine with an XDNA2 NPU (Strix, Strix Halo, Kraken, Gorgon Point) on Windows, or on Linux through the guide and Lemonade integration, and you want local inference without a discrete GPU. Do not adopt it if your machine has no XDNA2 NPU, if you cannot move to AMD NPU driver 32.0.203.311 or above, or if you need a model outside the maintained list. Verify three things first: your NPU driver version in Task Manager, that your chip is in the supported set, and that the model tag you need appears in flm list after install.

## FAQ

### What is FastFlowLM?

FastFlowLM is an NPU-first runtime that runs large language models, plus vision, audio, embedding and MoE models, on AMD Ryzen AI NPUs. It ships a single-command CLI and a local server on port 52625 that speaks REST and the OpenAI API. The README describes it as the only out-of-box, NPU-first runtime built exclusively for Ryzen AI.

### How do I use FastFlowLM?

Install the flm-setup.msi package, open PowerShell, and run flm run llama3.2:1b for a terminal chat session; the first run downloads optimized kernels from HuggingFace. Use flm serve llama3.2:1b to start the local server instead. Inside a session, /verbose toggles performance reporting and /bye exits.

### How does FastFlowLM compare with llama.cpp?

llama.cpp is a portable inference engine that runs across CPU, CUDA, ROCm and Metal and accepts arbitrary GGUF weights, so you choose the quantisation. FastFlowLM supports only AMD Ryzen AI chips with XDNA2 NPUs and uses prebuilt optimized kernels for a curated model list, which the README presents as requiring no model rewrites and no tuning.

### How does FastFlowLM compare with Ollama?

Ollama is a model-management daemon with a friendly CLI and an OpenAI-compatible API that runs on whatever backend it was built for. FastFlowLM targets one accelerator family, the Ryzen AI NPU, and trades that breadth for a 17 MB runtime and an install the README says completes within 20 seconds.

### What is the difference between Lemonade and FastFlowLM?

They are complementary rather than competing. FastFlowLM provides the NPU-optimized runtime and kernels, and the README states FLM was integrated into AMD's Lemonade Server in October 2025. On Linux the README directs users to the Lemonade Server documentation for getting started with FLM.

## Sources

- [License: MIT](https://github.com/ROCm/FastFlowLM/blob/main/LICENSE)
- [Project website](http://fastflowlm.com/)
- [README](https://github.com/ROCm/FastFlowLM/blob/main/README.md)
- [Releases](https://github.com/ROCm/FastFlowLM/releases)
- [ROCm/FastFlowLM on GitHub](https://github.com/ROCm/FastFlowLM)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/rocm-fastflowlm
