# GenieX: running GGUF and AI Hub models on Snapdragon NPU, GPU or CPU

> Qualcomm's developer-preview runtime loads GGUF files from Hugging Face or pre-compiled AI Hub bundles and dispatches them to the Hexagon NPU, Adreno GPU or CPU through one C SDK with CLI, Python, Kotlin/Java and an OpenAI-compatible server on top.

**qualcomm/GenieX** — Run frontier LLMs and VLMs locally on Qualcomm devices across NPU, GPU, and CPU with a few lines of code

- Repository: https://github.com/qualcomm/GenieX
- Website: https://geniex.aihub.qualcomm.com/en/get-started/what-is-geniex
- Stars: 8,406 · Forks: 1,065
- Language: Rust
- License: BSD-3-Clause
- Published: 2026-09-09 · Updated: 2026-09-09 · Language: en
- Canonical page: https://hysenlabs.com/projects/qualcomm-geniex

## The gap GenieX fills: Snapdragon inference without writing a backend per accelerator

Running a quantised language model on a phone or an ARM laptop usually means picking an accelerator and living with it. CPU inference through a generic runtime is portable and slow. NPU inference through a vendor SDK is fast and tied to that vendor's toolchain, model format and build steps. Moving between the two means a second integration. GenieX's stated goal is to remove that choice: the README describes it as an on-device Gen AI inference runtime for Qualcomm devices that runs models on the Hexagon NPU, Adreno GPU or CPU, and the architecture diagram shows one GenieX SDK dispatching to either the llama.cpp runtime (GGML over CPU, GPU and Hexagon HTP kernels) or the Qualcomm AI Engine Direct runtime on the NPU.

The audience is narrow and specific. This is for developers shipping on Snapdragon X and X Elite Windows ARM64 machines, Snapdragon 8 Elite class Android phones, and Linux ARM64 IoT boards such as Dragonwing QCS9075. If your deployment target is an x86 server with an NVIDIA card, nothing here applies. The README is explicit that GenieX runs only on Qualcomm Snapdragon.

The project describes itself as the community version of Qualcomm GENIE, and the repository carries a developer-preview badge. Treat the badge as load-bearing: it tells you the API surface is still moving, which the release cadence supports (v0.5.0 on 2026-08-22, then v0.6.0 and v0.6.1 both on 2026-09-03).

## Two runtimes, one SDK: how GGUF and AI Hub bundles reach the hardware

The mechanism visible in the README is a two-path dispatcher. The first path takes a GGUF file, from Hugging Face or from Docker Hub, and runs it through llama.cpp, where GGML kernels cover CPU, Adreno GPU and Hexagon HTP. The second path takes a pre-compiled bundle from Qualcomm AI Hub and runs it through Qualcomm AI Engine Direct, which the README associates with the NPU.

The choice is not cosmetic. A GGUF file is portable and easy to obtain; you can point at almost any quantised model on Hugging Face. An AI Hub bundle is compiled ahead of time for the target, which is why the README lists the NPU only on that path. So the practical rule is: GGUF when you want breadth of models and are content with GPU or CPU, AI Hub bundles when you want the NPU and can accept the models Qualcomm has published.

Above that dispatcher sits one C SDK, with bindings and front ends layered on top. The README names a CLI, Python, Kotlin/Java, Docker and an OpenAI-compatible server. That layering is the actual product claim: the accelerator decision lives below the API you call, so switching from a GGUF build to an AI Hub build is a change of model identifier rather than a rewrite.

## Installing GenieX and running a first model

There are three install routes in the README, one per front end. On Linux ARM64 the CLI installs with a single shell line and no sudo, and the README notes you should open a new terminal afterwards on Windows.

```bash
curl -fsSL https://qaihub-public-assets.s3.us-west-2.amazonaws.com/qai-hub-geniex/install.sh | sh
```

The Python binding is a pip package, which is the shortest path if you already work in Python.

```bash
pip install geniex
```

Once installed, the CLI runs a model in one line. Passing a Hugging Face repository identifier selects the GGUF path; passing an ai-hub-models identifier selects the pre-compiled bundle path. The README also shows a Docker Hub identifier form, docker.io/ai/gemma3.

```bash
geniex infer google/gemma-4-E4B-it-qat-q4_0-gguf
```

For programmatic use, the Python API deliberately mirrors Hugging Face transformers: from_pretrained() to load, generate() to decode. The README's example loads a GGUF repository with an explicit precision and streams tokens.

```python
from geniex import AutoModelForCausalLM

model = AutoModelForCausalLM.from_pretrained("unsloth/Qwen3.5-2B-GGUF", precision="Q4_0")
messages = [{"role": "user", "content": "What is 2+2?"}]
prompt = model.tokenizer.apply_chat_template(messages, add_generation_prompt=True)

for chunk in model.generate(prompt, max_new_tokens=256, stream=True):
    print(chunk, end="", flush=True)

model.close()
```

If you want an HTTP endpoint instead, the server ships with the CLI. Pull a model, start the server, and point any OpenAI client at the local address the README gives.

```bash
geniex pull ai-hub-models/Qwen3-4B-Instruct-2507
geniex serve   # serves http://127.0.0.1:18181/v1
```

The README states the server is OpenAI-compatible and that no client code changes are needed beyond the base URL. It does not document authentication on that endpoint, so treat it as a local development server rather than something to expose.

## Where GenieX will not help you

The strongest limitation is the one the README states outright: Qualcomm Snapdragon only. That rules out the most common place people run local models today, which is an x86 desktop or server. It also rules out Apple silicon and non-Qualcomm ARM parts.

Hardware support is not uniform across Snapdragon either. The platform table lists three groups with example devices, not an exhaustive compatibility list, and the README does not document a fallback when your specific part is absent. The search questions around "why isn't Geniex working on my device" suggest this is a real point of friction, though the README gives no troubleshooting section to answer it.

The NPU path carries a second constraint. Because it depends on pre-compiled AI Hub bundles, the set of models you can run on the NPU is the set Qualcomm has published, not the set on Hugging Face. If your model of choice has no bundle, you are on the llama.cpp path and therefore on GPU or CPU.

Finally, the developer-preview status matters. Three releases in the month before the last push on 2026-09-09 is a fast cadence, and fast cadences on a preview SDK mean interfaces can shift between versions. Nothing in the README documents a deprecation policy or a stability guarantee for the Python or Kotlin APIs.

## How GenieX compares with llama.cpp and ONNX Runtime on device

The honest comparison is with llama.cpp itself, because GenieX uses it underneath for the GGUF path. If your target is CPU or Adreno GPU and you are comfortable building llama.cpp for ARM64, you get the same kernels without an extra SDK layer. What you would not get is the AI Engine Direct path to the Hexagon NPU, the unified from_pretrained() style Python API, the Kotlin/Java binding, or the bundled OpenAI-compatible server. GenieX's value is the packaging and the NPU route, not the GGUF execution engine.

The other reference point is ONNX Runtime with a vendor execution provider. That route is model-format-first: you convert to ONNX and target an execution provider. GenieX is model-identifier-first: you name a Hugging Face repo or an AI Hub bundle and it picks the runtime. ONNX Runtime gives you a wider set of supported accelerators across vendors; GenieX gives you a shorter path to Qualcomm silicon specifically. Neither is a superset of the other, and the deciding question is simply whether your hardware is Snapdragon.

## Licence, release cadence and what an upgrade costs you

GenieX is BSD-3-Clause, a permissive licence that allows commercial use and modification provided the copyright notice and disclaimer are retained. The repository also carries a NOTICE file, and BSD-3-Clause includes a clause restricting use of contributor names for endorsement, so check NOTICE alongside LICENSE before redistributing. That is a pointer to the files, not legal advice; your own counsel decides what your product needs.

The upgrade cost is tied to the preview cadence. The last push was on 2026-09-09, and v0.6.0 and v0.6.1 landed on the same day, 2026-09-03, which suggests patches arrive quickly after a feature release. The Android binding is versioned separately from the runtime (the README shows com.qualcomm.qti:geniex-android:0.3.1 while the repository is at v0.6.1), so a runtime upgrade does not automatically mean a binding upgrade. Pin both versions explicitly and re-run your model on each bump rather than tracking latest. The README does not document a rollback procedure or a compatibility matrix between runtime and binding versions, so budget for testing on every upgrade.

## Conclusion

Adopt GenieX if your target hardware is Snapdragon and you want to move a model between NPU, GPU and CPU without rewriting the calling code: the same from_pretrained() call covers both a Hugging Face GGUF and an AI Hub bundle. Do not adopt it for x86 servers, Intel or AMD GPUs, or anything where the developer-preview label is a blocker for production. Verify three things before committing: that your exact Snapdragon part appears in the platform table, that the model you need exists as an AI Hub bundle if you want the NPU path, and that the Python or Android binding version you install matches the runtime release you tested against.

## FAQ

### What is GenieX used for?

It is an on-device inference runtime for running large language models and vision-language models locally on Qualcomm Snapdragon devices. It loads GGUF files from Hugging Face or pre-compiled bundles from Qualcomm AI Hub and runs them on the Hexagon NPU, Adreno GPU or CPU.

### Can I use GenieX without a SIM card?

The README does not mention SIM cards, cellular connectivity or any network requirement for inference. Models run locally on the device, and the only network access shown is fetching the installer, the pip package and the model files themselves.

### Why isn't GenieX working on my device?

The README states GenieX runs only on Qualcomm Snapdragon, and lists three platform groups: Windows ARM64, Android with Snapdragon 8 Elite class chips, and Linux ARM64 IoT boards. If your device is not Snapdragon, or not in one of those groups, the README offers no fallback. It also does not include a troubleshooting section.

### Is there any app like GenieX?

The README does not compare GenieX with other applications. The closest documented alternatives are the runtimes it uses or parallels: llama.cpp, which handles its GGUF path, and ONNX Runtime with a vendor execution provider, which is model-format-first rather than model-identifier-first.

### What is GenieX service?

The README documents a local server rather than a hosted service: geniex serve exposes an OpenAI-compatible endpoint at http://127.0.0.1:18181/v1 after you pull a model with geniex pull. It does not describe any remote or account-based GenieX service.

### How to use GenieX on Android?

For Android the README documents a Kotlin/Java SDK added to your app module's build.gradle.kts as com.qualcomm.qti:geniex-android:0.3.1. The CLI, Python and local server routes are listed for Windows ARM64 and Linux ARM64 instead.

## Sources

- [License: BSD-3-Clause](https://github.com/qualcomm/GenieX/blob/main/LICENSE)
- [Project website](https://geniex.aihub.qualcomm.com/en/get-started/what-is-geniex)
- [qualcomm/GenieX on GitHub](https://github.com/qualcomm/GenieX)
- [README](https://github.com/qualcomm/GenieX/blob/main/README.md)
- [Releases](https://github.com/qualcomm/GenieX/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/qualcomm-geniex
