hipfire's headline decode figure is a single-turn number, and Redline fails closed five ways
RDNA-native LLM inference engine in Rust.
At a glance
- What is it?
- hipfire is a Rust inference engine for AMD RDNA cards with an Ollama-style command surface and an in-house dispatch substrate called Redline. Its own table shows the newest GPU in the benchmark as the slowest, and its fastest single-turn figure drops by a third once a session runs eight turns.
- Who is it for?
- hipfire suits someone with an AMD RDNA card who wants local quantised models with draft-model speculative decoding and does not mind reading a per-architecture performance table before choosing a tag. It does not suit an Intel or Nvidia machine, and the flagship tag needs a 24 GB card or Strix Halo, so smaller cards start several model sizes down.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository received new commits within the last day.
- What is it written in?
- Mainly Rust, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 3, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Five kinds of failure all fall back to plain HIP dispatch
Redline is the part of hipfire with no external equivalent, described as an in-tree dispatch and retained-replay substrate. It records the actual kernel graph, derives resource dependencies from it, retains invariant command state, and lowers validated paths through public ROCr queue interfaces. The design brief is unusually blunt about what happens when it cannot help. Unsupported graphs, failed shadow validation, ABI mismatches, queue faults, and model changes all fail closed to ordinary HIP dispatch. That is five distinct failure classes with one shared outcome, and the word closed is doing real work: the fast path is abandoned rather than partially applied. Two qualifications follow from the same paragraph. Redline spans RDNA1 through RDNA4, but optimised routes remain architecture-specific and workload-specific, so architecture coverage is not route coverage. Reproduction is treated as part of the contract rather than an afterthought: a sealed route, model, and sampling fixture can be replayed with a single module invocation, and a separate guide covers making that validated fixture the serve default and connecting an OpenAI-compatible client to it. A performance claim you can re-run is worth more than a faster number you cannot.
The headline decode figure is a single-turn measurement
The MQ4R performance table carries three columns that tell different stories about the same run. On a Radeon RX 7900 XTX, ordinary autoregressive decode with Q8 KV measures 253.3 tokens per second at a 128-token generation target, 191.0 averaged over eight user-facing turns, and 160.3 on the final turn. So the headline number is roughly a third higher than what a long session actually sustains. The same shape holds on the other two cards: 115.1 becomes 92.2 and then 82.5 on Strix Halo, and 203.9 becomes 169.5 and then 146.7 on the R9700. The readme is candid about why the context column is not uniform, noting that the three recorded sessions did not all terminate at exactly 22K. Health is reported as eight clean runs out of eight on every card.
The newest architecture in the table is the slowest
Read down the architecture column and the ranking inverts. The Radeon RX 7900 XTX on gfx1100 produces 253.31 tokens per second at the median, Strix Halo on gfx1151 produces 115.10, and the Radeon AI PRO R9700 on gfx1201 produces 203.93, so the most recent part in the table sits between the other two rather than above them. The spread is narrow inside each card, with the 7900 XTX ranging from 253.04 to 253.48 and the R9700 from 203.42 to 204.04. The readme also reports a campaign that raised ordinary MQ4R autoregressive decode on gfx1201 from roughly 110 tokens per second to 203.9, which means the R9700 figure reflects work done on that architecture rather than a hardware ceiling. Comparing cards on these rows compares software maturity as much as silicon. The prefill numbers tell a different story again: on one R9700 the flagship model reaches roughly 5,120 tokens per second at an 8192-token prompt with native fp8 key-value caching as the default, dropping to 4,464, 3,804, and 2,940 at 32K, 64K, and 128K, while gfx11 prefill lands near 2,985 on the 7900 XTX and 1,141 on Strix Halo. Prefill scales far better than decode across architectures, which is why context-heavy work and chat-heavy work are not the same purchase.
The readme says 80 models and points you at a command for the real list
The registry section states that it currently contains 80 curated model entries and then immediately tells you to run the listing command to see the authoritative live list:
hipfire list -rThat ordering is the honest one, because a curated quantisation registry changes whenever a new weight set is refitted, and a number printed in prose would be stale within a release. The families are organised by model generation rather than by quantisation scheme, which is where the naming gets interesting. Qwen 3.5 offers dense sizes with MQ3 and MQ6 suffixes, a legacy HF6 tag, and separate draft tags. Qwen 3.6 35B-A3B offers nine variants including an MQ4P default and the MQ4R performance tag. Curated weights are published on a Hugging Face organisation rather than inside the repository, with per-model repositories recorded in the registry itself. The same family column also carries the release notes' other v0.4.0 items: a mixture-of-experts model at its full 262K context on a single R9700 with tensor parallelism at one and host-mapped experts, the same model on Strix Halo, and prebuilt kernel packs shipped in the installer rather than compiled on first run.
The tag grammar changes between model generations
A user who learned the naming in one generation has to relearn it in the next. The Qwen 3.5 and 3.6 tags use a compact suffix scheme: mq3, mq6, mfp4, mq4p, mq4r, and a legacy hf6. The Qwen 3.8 dense ladder adds a second axis. Alongside mq3, mq3-pro, mq4 and mq4-pro there are xt and symmetric xt variants, with mq4-xts marked as the flagship and the bare tag documented as the MQ4V2 default, plus corresponding MQ5 and MQ6 forms in the same three shapes. Draft models follow the same families with their own suffixes and an MQ4 recommendation. The practical consequence is that an upgrade script that constructs tags by string concatenation has to know which grammar the target generation uses, because the same suffix means different precisions across families. That choice is a quality decision the readme makes explicit for the one model it profiles. MQ4R is described as the throughput-oriented variant of the 35B-A3B model, combining uniform four-bit attention and gate-side weights with graded routed experts and a fused gate path; the weights are 18.7 GB and need roughly 22 GB of available memory. When quality matters more than decode speed, the guidance is to fall back to the default quantisation or move up to the four-bit floating-point, five-bit, or six-bit variants of the same model.
Apache-2.0 at the workspace, MIT headers on some files
The licence situation is layered. The workspace manifest sets the outbound licence for the work as a whole to Apache-2.0, and a comment above that line explains the exception: individual files whose authors have not elected Apache-2.0 keep an MIT SPDX header, with LICENSE and NOTICE named as the places to check. The root carries both licence texts plus NOTICE, CREDITS.md, and a PRIOR-ART.md, and the workspace includes a deny.toml, which is the configuration cargo-deny uses to fail a build on licence or advisory problems rather than report them after the fact. The hosting site still shows no recognised licence for the repository, because the dual arrangement at the file level defeats automatic detection. That is a reason to read LICENSE and NOTICE directly if you intend to vendor the crates rather than run the binary.
Loopback by default on port 11435, with the rest hardened in v0.4.0
The daemon exposes an OpenAI-compatible API, which is what lets any existing client talk to a local model without change. The bind address is loopback only by default, and the one setting that changes it is serve.host, so a machine on a shared network does not publish an inference endpoint by accident. The v0.4.0 notes group the same hardening under one heading: the loopback default, daemon respawn, drain on SIGTERM, and typed errors. Those four are the difference between a crash that leaves the port busy, a process that outlives its usefulness, and an error a client can parse. The three-command quick path is pull, serve in the background, and chat:
hipfire pull qwen3.8:27b-mq4-xts
hipfire serve qwen3.8:27b-mq4-xts -d
hipfire chat qwen3.8:27b-mq4-xtsA one-shot command uses the same registry and serving stack for a single prompt, so nothing has to be running first.
Fifty workspace members, sixteen of them one per model family
The Cargo workspace lists 50 members, and the shape of that list explains the architecture. Sixteen are named hipfire-arch-something, one per model family: Qwen 3.5, Qwen 4, Qwen 3.5 vision, Llama, DeepSeek 4, Gemma 4, Muse Glimmer, Qwen 2, LFM2 mixture of experts and its vision variant, MiniMax, Cohere 2 MoE, Maple, a document OCR family, and a diffusion family for image generation. Underneath sit the bridges, one each for HIP, VAAPI, HSA, and RDNA compute, plus the Redline trio of redline, redline-rocr, and redline-dispatch alongside a plain hipfire-dispatch and its test crate. The workspace version is 0.4.0 on edition 2021. Around the crates sit a Containerfile rather than a Dockerfile, a flake.nix with a lock file, a kernels directory, a third_party directory with a submodule declaration, and separate trees for benchmarks, experiments, and an autoresearch directory.
Editorial conclusion
hipfire suits someone with an AMD RDNA card who wants local quantised models with draft-model speculative decoding and does not mind reading a per-architecture performance table before choosing a tag. It does not suit an Intel or Nvidia machine, and the flagship tag needs a 24 GB card or Strix Halo, so smaller cards start several model sizes down. Before pulling anything, pick the quantisation from your own quality requirement rather than from the fastest row, check whether Redline has a validated route for your architecture at all, and treat the single-turn decode figures as an upper bound rather than a throughput you will see in a long session.
Frequently asked questions
What is hipfire?
It is an AMD-native LLM inference engine written in Rust, described as RDNA-first and CDNA-supported, built on HIP with an in-tree dispatch layer called Redline and no Python in the hot path. The command surface is Ollama-style, with pull, serve, chat, run, img, and list subcommands.
What does hipfire run on?
AMD RDNA graphics cards from RDNA1 through RDNA4, with CDNA support claimed. The benchmark table covers a Radeon RX 7900 XTX on gfx1100, a Radeon 8060S Strix Halo on gfx1151, and a Radeon AI PRO R9700 on gfx1201. The figures are self-measured on ROCm 10.0.
What hardware does the hipfire flagship model need?
The flagship qwen3.8:27b-mq4-xts tag pulls about 15 GB and wants a 24 GB card or Strix Halo. The readme suggests starting with qwen3.5:4b or qwen3.5:9b on smaller cards, using the same commands.
How does hipfire expose its models to other tools?
The daemon serves an OpenAI-compatible API on 127.0.0.1:11435, bound to loopback by default. Setting serve.host makes it listen on other interfaces. There is also a Golden Redline guide covering connecting an OpenAI-compatible client.
How many models does hipfire offer?
The readme says the curated registry contains 80 entries and tells you to run hipfire list -r for the authoritative live list, since the registry changes as new weight sets are refitted. Families cover Qwen 3.5, 3.6, 3.8, and the Flash-Next mixture of experts, plus a Flux image model.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/warpfront-hipfire)