OnnxStream: running Stable Diffusion and LLMs in a few hundred megabytes of RAM
Lightweight inference library for ONNX files, written in C++. It can run Stable Diffusion XL 1.0 on a RPI Zero 2 (or in 298MB of RAM) but also Mistral 7B on desktops and servers. ARM, x86, WASM, RISC-V supported. Accelerated by XNNPACK. Python, C# and JS(WASM) bindings available.
At a glance
- What is it?
- OnnxStream is a C++ ONNX inference library built around one goal: cutting memory use, not latency. It runs SDXL 1.0 in under 300MB of RAM and ships Python, C# and WASM bindings.
- Who is it for?
- Adopt OnnxStream when memory, not latency, is the binding constraint: a Raspberry Pi Zero 2 with 512MB of RAM, a browser tab running the YOLOv8 or Whisper WASM demos, or a machine where the model simply will not fit under OnnxRuntime. Do not adopt it for interactive image generation on ordinary hardware, where a 10-step SDXL image takes about 11 hours on the author's RPI Zero 2, or for any pipeline that expects a stable C API.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository last received commits 105 days ago.
- What is it written in?
- Mainly C++, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The constraint OnnxStream was written against
Most ONNX runtimes optimize for latency or throughput, and they spend RAM to do it. OnnxStream inverts that priority. The README states the challenge directly: run Stable Diffusion 1.5, a model with almost 1 billion parameters, on a Raspberry Pi Zero 2 with 512MB of RAM, without adding swap and without offloading intermediate results to disk. The recommended minimum for SD 1.5 is typically 8GB.
That framing tells you who the library is for. It is not for a server with a GPU that wants more images per second. It is for the person holding a 512MB board, a browser tab, or a container with a hard memory cap, who would rather wait than fail. The author describes OnnxStream as a small and hackable inference library, and the repository layout matches that description: a src directory, an onnx2txt tool, and example directories for Whisper and YOLOv8 compiled to WebAssembly.
The project is not archived, and the last push was on 2026-06-18. Releases are sparse: v0.1 in July 2023 and v0.2 in January 2024. Newer work, including the Python and C# bindings announced on September 28, 2025 and the Whisper WASM demo from April 15, 2025, appears in the repository rather than in tagged releases. If you depend on versioned artifacts, that gap matters.
WeightsProvider: the one abstraction that explains the design
OnnxStream decouples the inference engine from whatever supplies the model weights. That supplier is a class derived from WeightsProvider, and it can implement any loading, caching and prefetching strategy. A custom provider could fetch parameters from an HTTP server and never write a byte to disk, which is where the "Stream" in the name comes from. Three default providers ship with the library: DiskNoCache, DiskPrefetch and Ram.
The consequence is that peak memory is a property of your provider choice, not of the model file size. DiskNoCache reads weights as the graph needs them and keeps little resident. DiskPrefetch trades some of that back for speed. Ram loads everything and behaves like a conventional runtime. This is the mechanism behind the README's headline claim that OnnxStream can consume 55x less memory than OnnxRuntime with a 50% to 200% increase in latency on CPU with a good SSD, measured on SD 1.5's UNET.
Quantization is the second lever, and it is not applied uniformly. For SD 1.5 the VAE decoder is statically quantized to UINT8, which brings it to 260MB; the README explains that the decoder is the one SD 1.5 component that would not fit on the Pi Zero 2 in single or half precision because of residual connections and large tensors. For SDXL 1.0 the approach differs: UINT8 dynamic quantization is limited to a specific subset of large intermediate tensors, and the VAE decoder, four times the size of SD 1.5's and 4.4GB in FP32, needs more than one technique. The README's section on attention slicing and quantization is the place to read before assuming a model will fit.
Building and running the Stable Diffusion example
The repository ships a Dockerfile that builds the sd binary from src and produces a runtime image based on Ubuntu 22.04. The comment at the top of the file gives the build command:
sudo docker build -t sd .The same comment gives an example run. Weights and the output image are written to the current directory, which is mounted at /current, and the flags shown are --models-path, --output, --steps and --turbo:
sudo docker run -t -i --init -v .:/current sd --models-path /current --output /current/result.png --steps 1 --turboWith --steps 1 and --turbo you are generating a single-step image, which is the SDXL Turbo path rather than the full 10-step SDXL 1.0 pipeline. Expect the first run to spend its time fetching weights into the mounted directory before any image appears. For native builds the README has a section titled How to Build the Stable Diffusion example on Linux/Mac/Windows/Termux/FreeBSD; the Dockerfile is the shortcut, not the only route.
The README also mentions a MAX_SPEED compile option. The SD 1.5 image generated on the author's RPI Zero 2 is described as taking about 1.5 hours with MAX_SPEED, against roughly 3 hours without it. That is the clearest signal that build flags change runtime by a factor, not a few percent.
What the memory savings actually cost you
The 55x figure is not free, and the README is unusually honest about the price: 50% to 200% more latency, on CPU, with a good SSD, measured on one component of SD 1.5. A good SSD is doing real work in that sentence. DiskNoCache and DiskPrefetch read weights during inference, so storage throughput becomes part of your critical path. On a Pi Zero 2 with a slow SD card, the bottleneck is not the CPU.
The SDXL 1.0 numbers make the trade explicit. Generating a 10-step image takes about 11 hours on the author's RPI Zero 2. The same 10-step image takes 26 minutes on a 12-core PC with 32GB of RAM using Hugging Face Diffusers. OnnxStream's advantage is that the 12-core PC is not required at all; its disadvantage is that you have moved from a coffee break to an overnight job.
There is a second limitation that the README does not address: API stability. The releases are v0.1 and v0.2, and the Python and C# bindings arrived years after the C++ core. If you are writing bindings into a product, the surface you are binding to is a research-oriented C++ project, and the release history does not suggest a compatibility guarantee. The README also does not document rollback or version pinning for converted models, so plan on pinning a commit rather than a tag.
Where OnnxStream is the wrong tool
If your goal is interactive generation on a workstation or a GPU server, OnnxStream is the wrong choice. OnnxRuntime will be faster, and the memory it consumes is memory you have. The README's own comparison is framed around CPU inference with a good SSD, which is a different deployment shape from a CUDA box.
If your model is not one of the supported ones and has no conversion path, the library will not help either. The README documents conversion for a custom Stable Diffusion 1.5 model, contributed by a third party, and the examples cover SD 1.5, SDXL 1.0 Base, SDXL Turbo 1.0, TinyLlama 1.1B, Mistral 7B, YOLOv8 and Whisper. That is a specific list. General ONNX graph coverage is not claimed anywhere in the README, and the onnx2txt tool exists because inspecting and converting graphs is part of the workflow rather than an afterthought.
Finally, if you need a supported, versioned SDK with a deprecation policy, this is not it. The project is active in the sense that the last push was on 2026-06-18, but the tagged release cadence is two releases in three years, and the bindings are announced in the news list rather than in a release.
OnnxRuntime and Diffusers: the difference in approach
The natural alternative is OnnxRuntime. It is the general-purpose ONNX execution engine, and it targets latency and throughput. OnnxStream's README positions itself against exactly that: major frameworks minimize latency or maximize throughput at the cost of RAM, and OnnxStream was written to minimize memory instead. The README's comparison is the concrete difference: up to 55x less memory than OnnxRuntime on SD 1.5's UNET, with 50% to 200% more latency on CPU with a good SSD. If you have the RAM, OnnxRuntime gives you the speed back.
The second comparison is Hugging Face Diffusers, which the README uses as the reference implementation for SDXL 1.0. The ONNX files in this project were exported from Diffusers version 0.19.3, so the two are related rather than unrelated: Diffusers is the source of the graph, OnnxStream is a way to execute it under a memory ceiling. Diffusers on a 12-core PC with 32GB of RAM produced a 10-step SDXL image in 26 minutes, against about 11 hours on the RPI Zero 2. That is roughly the exchange rate for going from 12GB of recommended VRAM to under 300MB of RAM.
A third option worth knowing about is Fastsdcpu, which appears among the searches people run against this project. The README does not compare itself to it, so treat any claim about that pairing as outside what the repository documents.
Licence and upgrade cost
The repository metadata reports the licence as NOASSERTION, which means the automated classifier could not map the LICENSE file to a known SPDX identifier. The file exists at the top level, so the terms are written down, but you have to read them yourself. Nothing in the README describes commercial use, redistribution of converted weights, or what happens to models you export with the onnx2txt tooling. Those are the questions to take to the LICENSE file before shipping anything, and to a lawyer if the answer is not obvious.
Upgrade cost is dominated by the weight conversion pipeline rather than by the C++ code. If you upgrade the library, you may need to re-export models, and the README pins its SDXL export to Diffusers 0.19.3, which tells you the conversion is sensitive to the exporter version. Budget for regenerating and re-validating weights, not just for pulling a new commit. On the runtime side, the Dockerfile pins UBUNTU_VERSION to 22.04, so container rebuilds are reproducible until you choose to move that argument.
Editorial conclusion
Adopt OnnxStream when memory, not latency, is the binding constraint: a Raspberry Pi Zero 2 with 512MB of RAM, a browser tab running the YOLOv8 or Whisper WASM demos, or a machine where the model simply will not fit under OnnxRuntime. Do not adopt it for interactive image generation on ordinary hardware, where a 10-step SDXL image takes about 11 hours on the author's RPI Zero 2, or for any pipeline that expects a stable C API. Before committing, verify that the licence file matches your intended use, confirm your target model has a conversion path in the repository, and read the Performance section to see what the 55x memory reduction costs you in latency on your own hardware.
Frequently asked questions
What is ONNX used for?
ONNX is the model format OnnxStream reads. The README describes the library as a lightweight inference library for ONNX files, and the SDXL 1.0 graphs it runs were exported from Hugging Face Diffusers version 0.19.3.
Is ONNX free to use?
The repository carries a LICENSE file, but the metadata reports the licence as NOASSERTION, meaning it could not be matched to a standard identifier. Read the LICENSE file directly before relying on any assumption about cost or reuse.
Is ONNX for CPU?
OnnxStream's performance claims are stated for CPU inference with a good SSD, and it is accelerated by XNNPACK. The LLM chat application in v0.2 also lists optional cuBLAS support for GPU.
What is the ONNX file format used for?
In this project the ONNX file holds the model graph and its parameters. OnnxStream separates the two: a WeightsProvider such as DiskNoCache, DiskPrefetch or Ram decides how those parameters are loaded, cached or prefetched during inference.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/vitoplantamura-onnxstream)