ADI-Stable-Diffusion: Running Stable Diffusion from C++ with ONNXRuntime
Accelerate your Stable Diffusion inference with the library's universal C/C++ framework design, powered by ONNXRuntime & across platforms.
At a glance
- What is it?
- ADI is a C++17 library and CLI that runs Stable Diffusion, SD3.5, FLUX and SVD through ONNXRuntime with no Python at inference time. It is aimed at engineers deploying diffusion models inside native applications, and its own README reports a 164 second wall time for SD3.5-turbo on CPU.
- Who is it for?
- Adopt ADI if you need diffusion inference inside a native C++ application or a desktop and mobile artifact where shipping a Python runtime is not acceptable, and if your models are already converted to ONNX. Do not adopt it if you want a prompt box in a browser, a model marketplace, or a Python pipeline you can edit in a notebook; ADI is a library and a CLI, not a web UI.
- Can I use it commercially?
- Yes, with conditions. GPL-3.0 is a copyleft licence: if you distribute software that includes it, you must release that software's source code under the same licence. Running it internally without distributing it does not trigger that obligation.
- Is it still maintained?
- Yes. The repository last received commits 29 days ago.
- What is it written in?
- Mainly C++, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What ADI solves, and who it is actually for
Stable Diffusion tooling is overwhelmingly Python. That is fine when the output is an image on a workstation, and awkward when the output is a feature inside a shipped application. A Python runtime, a diffusers install and a CUDA stack are not things most desktop or mobile products want to carry. ADI (Agile Diffusers Inference) targets that gap: a C++17 library with a CLI tool, built on ONNXRuntime, that loads .onnx model files and generates images or short video clips without Python in the inference path. The README states this directly: pure C++17, zero Python at inference time, one CLI, nine platform targets.
The intended user is an engineer who already has ONNX weights, or is willing to produce them, and needs to call diffusion from native code. The repository ships an include/adi.h header, a static library per platform, and the ONNXRuntime library alongside it. That layout tells you the project expects to be linked into something, not just run from a terminal. The CLI exists so you can validate the pipeline before writing any integration code.
This is not a consumer image generator. There is no web interface in the repository layout, no model browser, no queue. If you want a browser tab where you type a prompt, ADI is the wrong layer of the stack.
How the ONNXRuntime pipeline is put together
The architecture is a fairly direct mapping of the diffusion pipeline onto ONNX graphs. ONNXRuntime executes the models; ADI supplies the orchestration around them: text encoding, the denoising loop, scheduler sigma computation, latent packing and unpacking, and decoding back to pixels. The repository separates this into engine/, sd/, source/, include/ and clitools/ directories, with ARCHITECTURE.md at the top level describing the split.
The v2.0.0 release notes describe two model families that need different handling. MMDiT-era models (SD3.5-turbo, FLUX.1-schnell) use triple text encoders including T5-XXL via SentencePiece, and FLUX uses packed latents with rotary ids. SVD img2vid runs a spatio-temporal UNet with a CLIP vision encoder and a dedicated euler_svd scheduler. Those are not cosmetic differences; each family implies a different graph topology and a different set of inputs the runtime must assemble.
Scheduling is handled as a discrete component. The release notes claim 14 discrete schedulers plus Karras sigmas and a rectified-flow family, described as numpy-verified against diffusers. A precision policy flag, --precision auto|fp32|fp16, probes available RAM and derives fp16 model copies on demand. That is a memory-for-time trade decided at runtime rather than at build time, which is a reasonable choice for a tool that has to run on machines you do not control.
Installing ADI and generating a first image
There are three documented installation routes. The package manager route is the fastest where it applies. The README notes that since v2.0.0 the Homebrew tap serves Apple Silicon (arm64) only, because upstream ONNXRuntime discontinued the osx-x86_64 prebuilt; Intel Macs are told to build from source.
## macOS (Homebrew, Apple Silicon):
brew tap windsander/adi-stable-diffusion
brew install adiOn Windows, the documented route uses git-Bash plus Chocolatey, downloading the .nupkg from the deploy branch and installing it locally.
curl -L -o adi.2.0.0.nupkg "https://raw.githubusercontent.com/Windsander/ADI-Stable-Diffusion/deploy/adi.2.0.0.nupkg"
choco install adi.2.0.0.nupkg -yThe second method is a manual download from the Releases page. The package tree contains bin/adi, lib/ with the platform ADI library and the ONNXRuntime library, include/adi.h, and the changelog, README and licence files. The README says you can install bin and lib into your system, or simply enter the unzipped bin directory and run adi from there.
The third method builds both the library and the CLI locally using the provided auto_build.sh script. If you omit the BUILD_TYPE parameter the script defaults to Debug, and if you do not enable an ORT provider explicitly the script picks a default for the platform.
bash ./auto_build.sh --platform macos --build-type debug
bash ./auto_build.sh --platform linux --build-type debugAfter installation the README's workflow is short: pick a scheduler, point at the ONNX models, and generate. The README does not print a full example command line with model paths, so the exact flag names for pointing at weights are not something I can give you here. Check the CLI's own help output after installing.
Where ADI is the wrong tool, and what it costs
The performance table in the README is the honest part of the project. On an Apple M4 Max with 128 GB of RAM, ONNXRuntime 1.28.0, default provider, single cold run, the README reports approximately 9.7 seconds for sd-turbo at 512x512 with 4 steps, and approximately 164 seconds for SD3.5-turbo at 1024x1024 with 4 steps. Those are CPU numbers, and the second one is nearly three minutes for a single 1024px image. Anyone expecting interactive latency from the CPU path should read that row twice. The README states no GPU benchmark, so what a GPU provider buys you is not documented.
The model coverage is also narrower than the marketing line suggests. The README says "From SD v1.5 to SD3.5 / FLUX / SVD", but the v2.0.0 notes list specific families: SD3.5-turbo, FLUX.1-schnell, SVD img2vid, plus SDXL-turbo, sd-turbo and SD v2.1 in the showcase. A model outside those families may or may not load. The README does not document a general model compatibility matrix.
Rollback is undocumented. The README does not describe how to revert to a previous ADI version, and the package manager channels are described as refreshed automatically by the deploy chain on every release, which means the channel may move forward without an obvious way back. If you pin a version, pin it by artifact, not by tap.
Finally, the licence is GPL-3.0. That is a strong copyleft licence, and it matters for a library you link into a distributed application. The README does not discuss linking exceptions or commercial terms.
ADI against the Python diffusers stack
The obvious alternative is Python with diffusers, which is what ADI's own schedulers are verified against. The difference is not speed in the abstract; it is where the runtime lives. Diffusers gives you a Python object graph you can modify, a large ecosystem of pipelines and community scripts, and access to PyTorch's GPU kernels. ADI gives you a compiled binary, a fixed set of supported model families, and ONNXRuntime as the only execution engine.
That trade runs in both directions. If you need to prototype a new sampler or patch a pipeline mid-experiment, Python wins outright. If you need to ship a 200 MB desktop binary that generates an image without asking the user to install Python, ADI is the shape of thing you want. The README's own framing, "engineering deployment", is accurate.
A second comparison point is the model format. Diffusers consumes PyTorch checkpoints; ADI consumes ONNX. Conversion is a step you own. ADI does not hide that, and the repository includes an ort_sd_py_imp.py file at the top level, which suggests the Python side is used for conversion and reference work rather than inference.
Maintenance, platform coverage and licence terms
The repository is not archived. The last push was on 2026-09-02, and the most recent release is v2.0.0 from 2026-08-27, following v1.2.0 on 2026-08-01 and v1.0.1 back in 2024. The gap between v1.0.1 and v1.2.0 is roughly two years, which is worth knowing if you adopt early: this project has had long quiet periods before.
Upgrade cost is concentrated in the ONNX model files rather than the library. v2.0.0 moved to ONNXRuntime 1.28.0 and the release notes claim bit-identical output against the previous engine baseline across 25 local regression cases. If that claim holds for your models, an engine upgrade is low risk. The larger cost is when a new model family arrives, since it can bring new encoders, new latent layouts and a new scheduler, as SD3.5 and FLUX did.
Platform coverage is documented as Android x4, Linux x2, macOS arm64 and Windows x2 in the automated release chain. The Homebrew note about dropping Intel Macs is a concrete example of upstream pressure changing what ADI can ship, and it will not be the last one.
On licensing: GPL-3.0 is a copyleft licence. Linking ADI into a proprietary application and distributing that application raises obligations the README does not discuss. That is a question for your own legal review, not something this article can settle.
Editorial conclusion
Adopt ADI if you need diffusion inference inside a native C++ application or a desktop and mobile artifact where shipping a Python runtime is not acceptable, and if your models are already converted to ONNX. Do not adopt it if you want a prompt box in a browser, a model marketplace, or a Python pipeline you can edit in a notebook; ADI is a library and a CLI, not a web UI. Before committing, verify three things: that a prebuilt artifact exists for your exact platform and architecture, since Homebrew serves Apple Silicon only from v2.0.0, that your target model family is covered by the v2.0.0 release notes, and that the GPL-3.0 terms fit how you plan to distribute the binary.
Frequently asked questions
What is Stable Diffusion and how does it work?
Stable Diffusion is a diffusion model that generates images by iteratively denoising a latent representation. ADI runs that denoising loop in C++ through ONNXRuntime, with the scheduler and text encoding handled by the library and the UNet graphs executed as ONNX models.
What are the different versions of Stable Diffusion that ADI supports?
The README's showcase covers SD v1.5, SD v2.1, SDXL-turbo, sd-turbo, SD3.5-turbo, FLUX.1-schnell and SVD img2vid. The v2.0.0 release notes group these into MMDiT-era models with triple text encoders and the SVD spatio-temporal UNet path.
Is Stable Diffusion AI free to use?
ADI itself is released under GPL-3.0, which is free to use and modify but carries copyleft obligations when you distribute linked software. The licence terms of the individual model weights you load are separate and are not covered by the ADI repository.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/windsander-adi-stable-diffusion)