# ONNX Runtime GenAI: Running LLMs On Device Without a Python Server

> Microsoft's C++ generative AI loop for ONNX models ships Python, C#, C/C++ and Java APIs, with CPU, CUDA, DirectML, OpenVINO and QNN backends. Here is what the repository actually documents, and where it stops.

**microsoft/onnxruntime-genai** — Generative AI extensions for onnxruntime

- Repository: https://github.com/microsoft/onnxruntime-genai
- Stars: 1,128 · Forks: 362
- Language: C++
- License: MIT
- Published: 2026-08-04 · Updated: 2026-08-18 · Language: en
- Canonical page: https://hysenlabs.com/projects/microsoft-onnxruntime-genai

## What onnxruntime-genai actually adds on top of ONNX Runtime

ONNX Runtime executes graphs. It does not decide what token comes next. ONNX Runtime GenAI supplies that missing layer: the generative loop that wraps a model, from prompt formatting through token sampling to KV cache reuse. The README describes it as implementing "the generative AI loop for ONNX models, including pre and post processing, inference with ONNX Runtime, logits processing, search and sampling, KV cache management, and grammar specification for tool calling."

That list matters because each item is work you would otherwise write yourself. KV cache management in particular is the difference between a chat session that stays responsive and one that reprocesses the whole conversation on every turn. Constrained decoding and grammar specification are what make tool calling reliable enough to parse.

The audience is narrow and specific. If you are shipping a desktop application, a mobile app, or an edge device and you want a language model running locally with no Python process in the loop, this is aimed at you. The README states the project powers Foundry Local, Windows ML, and the Visual Studio Code AI Toolkit, which tells you the intended shape of a deployment: a library embedded in a larger product, not a standalone server.

## The API surface: Python, C#, C/C++, Java, and where Java breaks

The support matrix lists Python, C#, C/C++ and Java as supported APIs, with Objective-C under development. That spread is the project's main structural advantage over Python-only inference stacks. A Windows application written in C# can call the same underlying runtime as a Python notebook, and the NuGet package Microsoft.ML.OnnxRuntimeGenAI.Managed appears in the status badges at the top of the README.

Java carries an asterisk. The matrix footnote reads "Requires build from source", so there is no prebuilt Java artifact to pull. If your deployment target is a JVM service, budget for a C++ toolchain and a CMake build before you can evaluate anything.

Operating system coverage is Linux, Windows, Mac and Android, with iOS listed under development. Architectures are x86, x64 and arm64. Hardware acceleration is the longest row: CPU, CUDA, DirectML, NvTensorRtRtx (TRT-RTX), OpenVINO, QNN and WebGPU are marked supported, while AMD GPU sits on the roadmap. That means an AMD discrete GPU is not a supported acceleration target yet, even though DirectML covers many Windows GPUs.

## Installing onnxruntime-genai and running Phi-3 in Python

The README points to the ONNX Runtime website for installation instructions and for building from source. It also gives a complete worked example with Phi-3, which is the fastest way to confirm the package works on your machine before you commit to it.

The first step downloads a pre-converted INT4 model from Hugging Face. Note the include path: only the cpu_and_mobile/cpu-int4-rtn-block-32-acc-level-4 directory is fetched, not the whole repository.

```shell
huggingface-cli download microsoft/Phi-3-mini-4k-instruct-onnx --include cpu_and_mobile/cpu-int4-rtn-block-32-acc-level-4/* --local-dir .
```

The second step installs numpy and a pre-release build of the package. The --pre flag is in the README's own example, so the published stable wheel and the documented quickstart are not the same artifact.

```shell
pip install numpy
pip install --pre onnxruntime-genai
```

The third step loads the model, creates a tokenizer and a stream, and sets search options. Setting max_length matters: the README comment says that otherwise it will be set to the entire context length.

```python
import onnxruntime_genai as og

model = og.Model('cpu_and_mobile/cpu-int4-rtn-block-32-acc-level-4')
tokenizer = og.Tokenizer(model)
stream = tokenizer.create_stream()

search_options = {}
search_options['max_length'] = 2048
search_options['batch_size'] = 1

chat_template = '<|user|>\n{input} <|end|>\n<|assistant|>'
```

What you should see is a model object, a tokenizer bound to it, and a stream ready to consume generated tokens. The README's snippet stops mid-expression at the prompt construction, so the generation call itself is not shown there; the examples/python directory in the repository is where the complete loop lives.

## Matching examples to the package version you installed

This is the part of the README most people skim, and it is the part that costs an afternoon. The project states plainly that examples on main may not align with the latest stable release, because features are added continuously. If you clone main and run its examples against a stable wheel, an API mismatch is expected, not a bug report.

The documented remedy is to read your installed version and check out the matching tag. On Linux or Mac the README uses pip list with grep; on Windows it uses findstr.

```bash
pip list | grep onnxruntime-genai
```

Then clone and check out the branch for that version, using the README's own example of v0.11.5 as the form to follow:

```bash
git clone https://github.com/microsoft/onnxruntime-genai.git && cd onnxruntime-genai
git checkout v0.11.5
cd examples
```

If you want the nightly build instead, the README gives a separate index URL and a build step. The nightly package comes from the ORT-Nightly feed, and building the Python wheel from a source checkout is done with python build.py.

```bash
python build.py
pip install --index-url https://aiinfra.pkgs.visualstudio.com/PublicPackages/_packaging/ORT-Nightly/pypi/simple/ onnxruntime-genai
```

One consequence of this versioning scheme: the example code you find in a blog post or a search result may target a different release than the one you installed. Reading your own version number first is the only reliable anchor.

## Where onnxruntime-genai is the wrong choice

The support matrix is honest about gaps, and they are large enough to disqualify the project for some workloads. Stable diffusion is listed as under development. Multi-modal models are on the roadmap, not in the supported column. If your product needs image generation or a vision-language model as a first-class feature, this is not the library for it yet, even though Phi and Qwen are listed with vision support in the model row.

AMD GPU acceleration is also on the roadmap. DirectML is supported and covers a range of Windows hardware, but a team standardised on ROCm or on AMD datacenter cards will not find a supported path here.

Speculative decoding is on the roadmap as well. If your workload is latency-bound and you were counting on speculative decoding to close the gap, the feature is not available.

There is also a telemetry note in the README: the project "may collect usage data and send it to Microsoft", with a link to docs/Privacy.md. For a library that runs on device, that is worth reading before you ship it inside a regulated or air-gapped product. The README does not state what is collected or how to disable it, so the privacy document is the place to look.

Finally, iOS is under development. A Mac build is supported, an iPhone build is not.

## How it differs from llama.cpp and from calling a hosted API

The nearest alternative in spirit is llama.cpp, and the difference is architectural rather than cosmetic. llama.cpp defines its own model format, GGUF, and its own inference kernels; you convert weights into that format and the project owns the whole stack from quantization to sampling. ONNX Runtime GenAI does the opposite. It consumes ONNX graphs produced by an external conversion pipeline and delegates execution to ONNX Runtime, which already has execution providers for CUDA, DirectML, OpenVINO, QNN and WebGPU. You inherit that provider ecosystem, and you inherit its constraints: your model has to be convertible to ONNX, and the architecture has to appear in the support matrix.

The other alternative is not a library at all. A hosted inference API removes the model download, the hardware requirement and the version-matching problem described above. It also removes on-device execution, which is the entire point of this project. If your requirement is that prompts never leave the machine, or that the application works offline, a hosted API is not a substitute regardless of how it performs.

Between the two local options, the deciding question is usually the model. If your model already exists as an ONNX graph, or you need a specific execution provider for your hardware, ONNX Runtime GenAI fits. If you need a model that only exists in GGUF, the conversion path runs the other way.

## Maintenance, licensing and the cost of tracking a moving target

The repository is not archived, and the last push was on 2026-08-07, the same day as the v0.15.2 release. Releases v0.15.0, v0.15.1 and v0.15.2 all landed within roughly a week at the end of July and the start of August 2026, which indicates active release work rather than a dormant tree.

That cadence has a cost. The README's own warning about examples drifting from stable releases is a direct consequence of it. A team that pins a version and upgrades deliberately will spend less time on this than a team that tracks main. The version-checkout workflow above is the mechanism the project provides for staying pinned.

The licence is MIT. That is permissive and imposes few conditions on redistribution, but it is not the whole picture for a shipped product. The repository contains a ThirdPartyNotices.txt at the top level, and the runtime links against ONNX Runtime plus whichever execution providers you enable. Those carry their own terms. The README also notes that most contributions require a Contributor License Agreement, which affects you only if you plan to send patches upstream.

The trademark section is worth a glance if you intend to use the name in a product: Microsoft's brand guidelines govern use of its marks, and modified versions must not imply Microsoft sponsorship. This is not legal advice; read LICENSE, ThirdPartyNotices.txt and the privacy statement yourself.

## Conclusion

Adopt onnxruntime-genai when you need local inference inside an existing application and one of the listed model families and execution providers matches your target hardware. Do not adopt it if you need multi-modal models, Stable Diffusion or AMD GPU acceleration today, since the support matrix lists those as under development or on the roadmap. Before writing code, verify three things: that your model architecture appears in the support matrix, that your execution provider is listed as supported rather than planned, and that the examples branch you check out matches the package version you installed, because the README warns that main-branch examples may not align with the latest stable release.

## FAQ

### What is the difference between ONNX and ONNX Runtime?

ONNX is the model format, and ONNX Runtime is the engine that executes ONNX graphs. ONNX Runtime GenAI sits above the runtime and implements the generative loop: pre and post processing, logits processing, search and sampling, KV cache management and grammar specification for tool calling.

### What does ONNX Runtime do?

It executes ONNX model graphs. In this project it is the inference layer underneath the generative AI loop, and it is what provides the execution providers for CPU, CUDA, DirectML, NvTensorRtRtx, OpenVINO, QNN and WebGPU.

### How do I install onnxruntime-genai for Python?

The README's Phi-3 example installs numpy and then a pre-release build with pip install --pre onnxruntime-genai. The README also links to the ONNX Runtime website for full installation instructions and for building from source.

### Which model architectures does onnxruntime-genai support?

The support matrix lists AMD OLMo, ChatGLM, DeepSeek, ERNIE 4.5, Fara, Gemma, gpt-oss, Granite, Granite MoE Hybrid, HunYuan Dense V1, InternLM2, Llama, Mistral, Nemotron, Phi (language and vision), Qwen (language and vision), SmolLM3 and Whisper. Stable diffusion is listed as under development and multi-modal models as on the roadmap.

### Why do the examples in the onnxruntime-genai main branch not match my installed package?

The README states that examples on main may not align with the latest stable release because features are added continuously. It tells you to read your installed version with pip list, then check out the branch with the matching version tag before using the examples.

## Sources

- [Official README](https://github.com/microsoft/onnxruntime-genai#readme)
- [Project repository](https://github.com/microsoft/onnxruntime-genai)
- [Release notes](https://github.com/microsoft/onnxruntime-genai/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/microsoft-onnxruntime-genai
