CLI tool
qualcomm/ai-hub-models avatar
qualcomm/ai-hub-models

Qualcomm AI Hub Models: a Python toolkit for exporting ONNX, TFLite and QNN models to Snapdragon hardware

Qualcomm® AI Hub Models is our collection of state-of-the-art machine learning models optimized for performance (latency, memory etc.) and ready to deploy on Qualcomm® devices.

1,222 stars215 forksPythonBSD-3-Clause

At a glance

What is it?
The repository packages a catalog of pretrained models with export scripts that compile, quantize and profile them on cloud-hosted Qualcomm devices. It is aimed at engineers who need measured on-device latency before shipping, not at people looking for a general-purpose inference runtime.
Who is it for?
Adopt it if your target is a Snapdragon chipset and you need compiled assets plus measured on-device latency before committing to a model choice. Do not adopt it if you are deploying to non-Qualcomm silicon, if you cannot use AI Hub Workbench, or if you need a runtime library rather than a model preparation pipeline.
Can I use it commercially?
Yes. BSD-3-Clause is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 1 day ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What problem the Qualcomm AI Hub Models catalog solves

Moving a PyTorch checkpoint onto a phone NPU is mostly paperwork. You need a runtime-specific compiled artifact, a quantization choice, and evidence that the result is fast enough on the exact chipset you ship. The ai-hub-models repository exists to remove that paperwork for a fixed catalog of models. Each entry ships the original PyTorch definition, a Python App class that wraps preprocessing and postprocessing, and export tooling that produces deployable assets for Qualcomm AI Engine Direct, LiteRT (TensorFlow Lite) and ONNX. The intended reader is an engineer who has already picked a task (detection, classification, segmentation) and wants a measured answer about latency and memory on a named device rather than a paper claim. The repository is not a runtime. It does not ship an inference engine for your app; it ships the models, the conversion path and the measurement loop. The homepage points to per-model on-device performance data, and the README states that the collection is optimized for deployment on Qualcomm devices. If your deployment target is an Apple Neural Engine or a desktop GPU, nothing here applies.

How export, quantization and profiling actually flow through AI Hub Workbench

The pipeline is split between local code and a hosted service. Locally, the package holds model definitions and demo scripts. Remotely, Qualcomm AI Hub Workbench performs the heavy steps. According to the README, an export script will compile the model for the chosen device and target runtime, quantize it if applicable, profile the compiled model on a real device in the cloud, run inference with sample input and compare the on-device output against the PyTorch output, then download the compiled model to disk. That comparison step is the part worth noting: the tooling does not simply hand you a converted file, it reports whether the converted model still produces the same answer as the source model. The runtime matrix is narrow by design. Qualcomm AI Engine Direct covers Android, Linux and Windows; LiteRT covers Android and Linux; ONNX covers all three. Precision is tied to the compute unit: CPU supports FP32, INT16 and INT8; GPU supports FP32 and FP16; the NPU supports FP16, INT16 and INT8, with the caveat in the README that some older chipsets do not support fp16 inference on their NPU. That caveat matters more than it looks. A model that profiles well on a current Snapdragon 8 series part may need a different precision on an older one, and the repository's answer is to re-export rather than to guess.

Installing qai_hub_models and running a first export

The package is on PyPI as qai_hub_models. The README lists Python 3.10 as recommended, with 3.11, 3.12 and 3.13 also supported. One installation note is easy to miss: on Snapdragon X Elite and X2 Elite machines, only 64-bit x86 Python is supported on Windows, and installation fails under Windows ARM64 Python. Start with the base install.

bash
pip install qai_hub_models

Individual models can pull extra dependencies, and the README points to the per-model README for those instructions. YOLOv7 is the example used throughout, and it installs as an extra.

bash
pip install "qai_hub_models[yolov7]"

Compilation, quantization and profiling all require access to AI Hub Workbench. You create a Qualcomm ID, log in to Workbench, generate an API token from your account page, and configure it locally. Without this step the export command has nothing to talk to.

bash
qai-hub configure --api_token API_TOKEN

With the token in place, the export command takes a model id, a target runtime, a precision and a device string. The README's example targets TFLite at float precision on a Samsung Galaxy S25 family device.

bash
qai-hub-models export yolov7 --target-runtime tflite --precision float --device "Samsung Galaxy S25 (Family)"

Expect the command to compile, optionally quantize, profile on a cloud-hosted device, run a sample inference and compare the output against PyTorch, then write the compiled model to disk. For a local sanity check that skips Workbench, the demo command can run in PyTorch mode.

bash
qai-hub-models demo yolov7 --eval-mode fp

The demo preprocesses input, runs inference and postprocesses the output into something human-readable. The README notes that many demos can also run on-device with --eval-mode on-device, which routes inference through a real cloud-hosted device. If you only want to look before you install, the CLI can browse and fetch without the full package.

bash
pip install qai_hub_models_cli
qai-hub-models models
qai-hub-models info mobilenet_v2
qai-hub-models fetch mobilenet_v2 --runtime tflite --precision float

The fetch command downloads a deployable asset directly, which is the fastest way to confirm that a model exists for your runtime before you invest in the export path.

Where the workflow breaks down

The dependency on AI Hub Workbench is the biggest constraint. Compilation, quantization, on-device profiling and the output comparison all run through the hosted service, so an air-gapped build environment cannot use the main path. The README does not document an offline equivalent for those steps. Local PyTorch inference through the demo command works without Workbench, but that tells you nothing about NPU behaviour, which is usually the reason to be here. The second limitation is coverage. The catalog is finite. If your model is not in the model directory, you are not adapting an existing entry so much as writing a new one, and the export tooling assumes a model definition it understands. The third is the App classes. The README is explicit that the Python apps are written to be an easy-to-follow example rather than to minimize prediction time. Anyone benchmarking the App wrapper and concluding something about model speed is measuring the wrong thing; the compiled asset profiled on the device is the number that counts. Finally, precision support is not uniform across chipsets. The fp16 NPU caveat on older chips means a working export on one device is not proof of a working export on another, and the repository gives no compatibility shortcut beyond re-running the export against the target device string.

How this differs from ONNX Runtime or LiteRT on their own

ONNX Runtime and LiteRT are runtimes. You bring a model, convert it yourself, and manage quantization and device-specific tuning on your own. Qualcomm AI Hub Models sits one layer above that: it supplies the model, the conversion recipe and a hosted measurement loop, and it produces artifacts for those same runtimes. The practical difference shows up in what you get back. With a plain runtime conversion you get a file and whatever accuracy you can verify locally. With this repository's export path you get a file plus a profile on real hardware plus a comparison between on-device output and PyTorch output, which is the check that catches a quantization mistake before it reaches a phone. The trade is control. You are limited to the models in the catalog, the runtimes in the support table and the devices Qualcomm hosts. A team with an in-house model and a local Snapdragon board may find direct ONNX Runtime with the QNN execution provider faster to iterate on, because nothing in the loop requires a network round trip or an account. The repository's value is highest when the model you want is already listed and you need defensible latency numbers without buying a device lab.

Release cadence, licence and upgrade cost

The repository is not archived, and the last push was on 2026-09-10. Releases are frequent and versioned in the v0.x series, with v0.61.0 published on 2026-08-25, v0.60.0 on 2026-08-13 and v0.59.0 on 2026-07-28. That pace is the upgrade cost. Pinning qai_hub_models is the only way to keep a build reproducible, because a floating version will move under you within weeks. Because model definitions, export flags and the supported device list can all change between releases, treat a version bump as a change to your build inputs and re-run the export and the output comparison rather than assuming the previous artifact still matches. The licence is BSD-3-Clause, which is permissive and generally straightforward for commercial use, but the licence covers the repository code. It does not grant rights to Qualcomm trademarks or to the hosted Workbench service, and model weights may carry their own upstream terms that this repository does not restate. Check the terms attached to each model you ship rather than assuming the repository licence settles it. This is a description of what the files say, not legal advice.

Editorial conclusion

Adopt it if your target is a Snapdragon chipset and you need compiled assets plus measured on-device latency before committing to a model choice. Do not adopt it if you are deploying to non-Qualcomm silicon, if you cannot use AI Hub Workbench, or if you need a runtime library rather than a model preparation pipeline. Before anything else, verify that your chipset appears in the repository's supported list, that your model has an entry in the model directory, and that your API token is configured with qai-hub configure --api_token API_TOKEN, because every compile and profile step depends on that access.

Frequently asked questions

What is Qualcomm AI Hub Models?

It is a collection of pretrained machine learning models optimized for deployment on Qualcomm devices, distributed as the qai_hub_models Python package. Each entry includes a model definition, a demo app and export tooling that compiles, quantizes and profiles the model through AI Hub Workbench.

How do I install Qualcomm AI Hub Models?

Install the base package with pip install qai_hub_models, then add per-model extras such as qai_hub_models[yolov7] when the model README calls for them. Python 3.10 is recommended, and Windows ARM64 Python is not supported on Snapdragon X Elite and X2 Elite machines.

Does Qualcomm AI Hub Models require an account?

Compilation, quantization and on-device profiling require access to Qualcomm AI Hub Workbench, which needs a Qualcomm ID and an API token configured with qai-hub configure --api_token API_TOKEN. Local PyTorch demos can run without that access, but they do not measure device performance.

Which runtimes and precisions does Qualcomm AI Hub Models support?

The support table lists Qualcomm AI Engine Direct, LiteRT (TensorFlow Lite) and ONNX as target runtimes. CPU supports FP32, INT16 and INT8, GPU supports FP32 and FP16, and the NPU supports FP16, INT16 and INT8, with the README noting that some older chipsets do not support fp16 inference on their NPU.

Can I browse and download Qualcomm AI Hub Models without exporting?

Yes. The qai_hub_models_cli package provides the same CLI, and commands such as qai-hub-models models, qai-hub-models info mobilenet_v2 and qai-hub-models fetch mobilenet_v2 --runtime tflite --precision float let you browse the catalog and download a deployable asset directly.

Official sources

  1. License: BSD-3-Clause
  2. Project website
  3. qualcomm/ai-hub-models on GitHub
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/qualcomm-ai-hub-models.svg)](https://hysenlabs.com/projects/qualcomm-ai-hub-models)