# Strix Halo Guide separates its measured numbers from everyone else's

> The Strix Halo Guide is an independent documentation repository for running large language models on AMD Ryzen AI MAX+ 395 hardware, and its real subject is evidence hygiene. It records one machine's measurements, keeps community results in a different file, labels speculative decoding when it was used, and lists newer runtimes without promoting them.

**hogeheer499-commits/strix-halo-guide** — Evidence-backed AMD Strix Halo local-AI setup and benchmarks: Qwen3.8, Ollama, llama.cpp, Vulkan/ROCm, large GGUFs, and cross-OEM results.

- Repository: https://github.com/hogeheer499-commits/strix-halo-guide
- Website: https://strixhaloguide.com/
- Stars: 350 · Forks: 24
- Language: Python
- License: MIT
- Published: 2026-09-15 · Updated: 2026-09-15 · Language: en
- Canonical page: https://hysenlabs.com/projects/hogeheer499-commits-strix-halo-guide

## One Beelink GTR9 Pro, and everyone else in another file

The guide is explicit about its base of measurement. The local numbers come primarily from one Beelink GTR9 Pro, and community results live in COMMUNITY_RESULTS.md rather than in the headline table. Cross-OEM coverage is reported as 13 systems or independent sources drawn from 10 credited community benchmark contributors.

That separation is the thing to understand about this project. Most local-AI leaderboards blend a maintainer's own machine with numbers pasted from a forum post. Here a headline row is either measured locally or it is not in the headline table at all.

The hardware underneath is specific: AMD Strix Halo and Ryzen AI MAX+ 395 systems with the Radeon 8060S, which reports as gfx1151, and 96GB or 128GB of unified memory. The documented setup path is a configured BIOS and Ubuntu 24.04 with kernel work, because the memory configuration is a firmware decision before it is a Linux one.

What the repository ships is narrow on purpose. It carries docs, scripts, data and charts, and states that it contains no .exe, no binary .zip, no browser extensions and no model weights.

The commands are not printed in the README summary either. A six-step Quick Start, a short setup answer in STRIX_HALO_LOCAL_LLM_SETUP.md, a setup.sh script and BEST_KNOWN_PROFILES.md are the entry points, and the README names this GitHub repository rather than the project website as the source of truth for setup commands, benchmark claims and raw evidence.

## 20.42 t/s is not the same claim as a community 22 to 65 t/s

The Qwen3.8 page is the clearest example of the method, and it starts by refusing to merge two numbers that look comparable. The locally measured figure is 20.42 t/s, reached through the Ollama API with Ollama-default MTP drafting. Community routes report 22 to 65 t/s.

Those are different claims, and the page exists to say why. MTP drafting is speculative decoding: the draft model proposes tokens that the target model verifies in a batch, which raises throughput and changes the shape of the measurement depending on acceptance rate. An API route also adds server overhead that a direct backend does not. Quoting the community range next to the local number without those qualifiers would be the standard mistake, and the guide names it.

So the Qwen3.8 page splits into what is verified locally and what still needs reproduction. The same discipline runs across the project: direct, API/server, MTP/speculative, concurrency, capacity, power, NPU, RPC, first-party and community claims are each kept apart, and headline rows link to CSVs, raw logs, charts or an explicit n/a.

## Newer builds are indexed, not promoted

Freshness is handled by an explicit snapshot rather than by editing old pages. The README's current evidence state is dated September 26, 2026: Qwen3.8 27B measured through the official Ollama route and labelled as using Ollama-default MTP drafting, with an active-evidence review filed under that date. An August 30 Vulkan sentinel on b10687 and a Flash-Next scout are indexed with their own stack and caveats attached.

Then there is the September 25 availability check, which lists newer candidates including Ollama 0.34.4 and llama.cpp v0.5.0 and does not promote them. That distinction is the whole point. A list of what exists is useful; a list of what is recommended is a claim, and the claim is what has to be measured.

Underneath all of it is a machine-readable record at data/public_state.json, so the freshness of the guide is a file you can diff rather than a promise in prose.

ROCm and HIP work is described as experiments, and vLLM appears as notes. Neither is presented as a finished route.

## The repository root is a filing cabinet of dated evidence

Around sixty markdown files sit at the top level, and the naming convention does most of the explaining. ACTIVE_EVIDENCE_REVIEW_2026-09-19.md and ACTIVE_EVIDENCE_REVIEW_2026-09-26.md are two days apart. BUYER_SNAPSHOT_2026-09-13.md and BUYER_SNAPSHOT_2026-09-19.md are a week apart. RUNTIME_QUALIFICATION_2026-09-19.md and MAX_PERFORMANCE_RESULTS_2026-05-07.md are four months apart.

Dated snapshots mean a reader can see what the position was on a date instead of inferring it from a page that has since moved. The cost is a root directory that is closer to an archive than a manual.

Three files carry the load. REPRODUCIBILITY.md holds the exact commands, metadata, raw evidence and caveats, with data/headline_claims.csv as the index and data/raw/ underneath it. EVIDENCE_CORRECTIONS.md records the times a published number changed. COMMUNITY_RESULTS.md holds the third-party rows. CONTRIBUTING.md points at a benchmark-report issue template, which is the mechanism by which someone else's machine becomes a row.

The reader-routing table in the README is the interface to all of it, sending a buyer, an owner, an evaluator, a reproducer and a reviewer to different first pages.

## Vendor money is allowed here, and it is named

Most independent benchmark projects leave that question unasked. This one has PARTNERSHIP.md, SPONSORSHIP.md, SPONSOR_ROADMAP.md and VENDOR_DISCLOSURE.md at the root, plus BEELINK_OUTREACH.md and OUTREACH_TEMPLATES.md, and the README states that users need independent buyer guidance that can coexist with future disclosed vendor or affiliate relationships.

Coexistence is the honest word. Vendors make these machines, and a guide that could not survive a sponsorship would have a short life. What matters is that the terms are written down in advance rather than negotiated after the money arrives.

SERVICE_INTAKE.md and SERVICES.md extend the same logic into paid work, with TRACTION.md recording what has actually happened. That is an unusual amount of process around what is essentially a documentation repository, and it is a deliberate choice by a maintainer who is also selling something.

A buyer who cares about independence should read the disclosure policy rather than the sponsorship page.

## Upstream pull requests are a credential, not evidence

The maintainer's argument for their own competence is a count of merged work upstream: 15 engineering pull requests across 10 projects, reconciled on 2026-09-13 and itemised in UPSTREAM_CONTRIBUTIONS.md. The named targets are a llama.cpp change, the AMD-sponsored Lemonade local-AI server, a Strix Halo detection fix in llmfit, OpenAI's official .NET SDK, and Kubernetes SIG inference-perf, plus three merged listing and documentation pull requests.

That is a real signal about whether the person measuring understands the stack. It is not evidence about throughput, and the guide says so in the same breath: upstream acceptance strengthens confidence in the engineering process, and it does not replace the raw evidence required for each benchmark claim.

The distinction matters more here than in most projects, because the subject is hardware where a plausible-sounding number is easy to produce. Somebody who can land a patch in llama.cpp is harder to fool with a table than somebody who cannot.

CITATION.cff is there too, which is what you want if you intend to cite any of this.

## What this guide is not

Three limits are stated in the guide itself and are worth repeating before anyone plans a machine around it.

It is not official. The README says AMD now publicly frames Ryzen AI Halo-class systems as a local-AI and developer-platform direction, and positions this repository as the independent practical layer underneath that, explicitly not official AMD or OEM endorsement, with the context filed separately.

It is not one leaderboard. The stated output is workload-to-model and backend recommendations rather than a single undifferentiated table, and capacity evidence is presented for 70B, 120B and selected 284B GGUF files rather than a promise about every size.

And its evidence does not transfer automatically. The guide's own framing of its purpose is a working route plus the evidence needed to understand where it transfers, or fails to transfer, to another system. One machine, one firmware configuration and one Linux path produce numbers that other machines can contradict.

The website at strixhaloguide.com is the readable front door; the repository is named as the source of truth for setup commands, benchmark claims and raw evidence.

## Conclusion

This guide suits someone who already owns or is about to buy Ryzen AI MAX hardware and wants to know which runtime to pick for a specific workload, along with the receipts. It does not suit someone looking for a single headline number or an official AMD position, because the maintainer says plainly that it is neither. Verify two things before you plan a build around it: that the row you care about is a local measurement rather than a community figure, and that your machine matches the hardware the guide measured, one Beelink GTR9 Pro with a Radeon 8060S reporting as gfx1151. If a figure is marked n/a, that is the guide telling you it has no measurement for it. The last push on the repository was on 2026-09-26 and the licence is MIT.

## FAQ

### What hardware does the Strix Halo Guide cover?

AMD Strix Halo and Ryzen AI MAX+ 395 systems with the Radeon 8060S, which reports as gfx1151, and 96GB or 128GB of unified memory. The documented path is a configured BIOS followed by Ubuntu 24.04 with kernel setup.

### Which runtimes does the Strix Halo Guide set up?

Ollama and llama.cpp over Vulkan with RADV, plus ROCm and HIP experiments and notes on vLLM. Newer candidates such as Ollama 0.34.4 and llama.cpp v0.5.0 are listed in an availability check without being promoted to a recommended route.

### Is the Strix Halo Guide official AMD documentation?

No. The README states it is the independent practical layer and not official AMD or OEM endorsement, and that context is recorded separately in RYZEN_AI_HALO_CONTEXT.md.

### Do the Strix Halo Guide benchmarks transfer to my machine?

Measurements were made primarily on one Beelink GTR9 Pro. Cross-OEM evidence covers 13 systems or independent sources from 10 credited community benchmark contributors, kept in a separate file from the local headline claims.

### Does the Strix Halo Guide accept vendor sponsorship?

The repository carries PARTNERSHIP.md, SPONSORSHIP.md and VENDOR_DISCLOSURE.md at the root, and the README states the goal is buyer guidance that can coexist with future disclosed vendor or affiliate relationships. Upstream pull requests are presented as confidence in the engineering process, not as a substitute for benchmark evidence.

## Sources

- [hogeheer499-commits/strix-halo-guide on GitHub](https://github.com/hogeheer499-commits/strix-halo-guide)
- [License: MIT](https://github.com/hogeheer499-commits/strix-halo-guide/blob/main/LICENSE)
- [Project website](https://strixhaloguide.com/)
- [README](https://github.com/hogeheer499-commits/strix-halo-guide/blob/main/README.md)
- [Releases](https://github.com/hogeheer499-commits/strix-halo-guide/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/hogeheer499-commits-strix-halo-guide
