Microsoft Olive: an ONNX Runtime optimization toolkit for turning models into deployable artifacts
Olive: Simplify ML Model Finetuning, Conversion, Quantization, and Optimization for CPUs, GPUs and NPUs.
At a glance
- What is it?
- Olive composes quantization, conversion and graph optimization into a single pass over a model, targeting CPUs, GPUs and NPUs. It is aimed at engineers who already know which runtime they will deploy on and want the export handled for them.
- Who is it for?
- Adopt Olive if you have already committed to ONNX Runtime and want quantization, graph capture and optimization run as one pipeline instead of three hand-written scripts. Do not adopt it if your deployment target is PyTorch, TensorRT or a vendor runtime Olive has no pass for; the output is an ONNX model and the toolkit has no path back.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 1 day ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem Olive solves: model export as a pipeline, not a script
Getting a trained model onto a target device is normally a chain of separate jobs. Export the graph. Quantize the weights. Check that the quantized graph still produces acceptable output. Optimize the graph for the target execution provider. Each step has its own tool, its own config format and its own failure mode, and the glue between them is usually a shell script that nobody maintains.
Olive's premise is that these steps are one job with one input and one output. The README describes it as composing "the best suitable optimization techniques to output the most efficient ONNX model(s) for inferencing on the cloud or edge, while taking a set of constraints such as accuracy and latency into consideration." The name is an abbreviation of Onnx LIVE.
The audience is narrow and worth stating plainly. This is for engineers who have already decided that ONNX Runtime is the inference engine, whether in a Python service, a C# application or an edge deployment. If that decision is not made, Olive is the wrong first step, because every technique it applies produces ONNX.
There is a second audience the README speaks to directly: people who prefer a command line over notebooks. The quickstart exists precisely because the project's earlier entry points were Jupyter-based, and the news list includes a November 2024 post titled "Democratizing AI Model optimization with the new Olive CLI." The CLI is the intended path for scripted, repeatable runs.
The constraints the README names are accuracy and latency. Those are the two axes a deployment team actually argues about, and they pull in opposite directions: aggressive quantization reduces latency and can move accuracy. Olive's answer is to treat the precision choice as a search input rather than a fixed setting, which is why optuna appears in requirements.txt. That design choice is also what makes the toolkit harder to reason about than a single export script, because the configuration that produced a given artifact is not always obvious from the artifact itself.
What the automatic optimizer actually does to a model
The quickstart walks a model through four stages, and the README lists them in order: acquire the model from the Hugging Face repo, quantize it to int4 using GPTQ, capture the ONNX graph and store the weights in an ONNX data file, then optimize the ONNX graph.
That sequence is the architecture in miniature. Acquisition and quantization happen on the source framework's terms, which is why torch and transformers sit in requirements.txt alongside onnx. Graph capture converts the result into an ONNX graph plus a separate weights file. Optimization then rewrites that graph, and the passes available at that stage depend on the execution provider you are targeting. The split between a graph file and a weights data file matters operationally: two files have to be shipped together, and a deployment that copies only the .onnx file will fail at load time.
The four stages are not independent. Quantization changes the numerical behaviour of the model, and the accuracy constraint the README mentions is what ties the choice of precision back to whether the output model is usable. Olive can search over configurations rather than applying one fixed recipe, which is where optuna in requirements.txt comes from.
One detail matters for planning: the README notes that the quickstart model, Qwen/Qwen2.5-0.5B-Instruct, "has many model files in the Hugging Face repo for different precisions that are not required by Olive." Olive pulls what it needs rather than the whole repository, but it still downloads from the Hub, so a machine without network access to Hugging Face cannot run the quickstart as written. Air-gapped build environments need a pre-populated cache or a local model path instead of a Hub identifier.
The repository layout hints at where the pipeline lives. The olive/ package holds the passes and the configuration model, and pydantic>=2.0 in requirements.txt is the mechanism: an Olive run is described by a typed config object, which is what allows the CLI to validate a run before executing any of it. That is a real advantage over a shell script, where a typo in a flag surfaces only after the model has finished downloading.
Installing olive-ai and running a first optimization
The README recommends a virtual environment or a conda environment before installing. The base package is olive-ai, and the quickstart installs two more packages alongside it so that the downloaded model can be converted and then run.
pip install olive-ai
pip install transformers onnxruntime-genaiThe README adds that Olive has optional dependencies enabling additional features, and points at olive/olive_config.json for the list of extras and their dependencies. That file is read by setup.py through a get_extra_deps function, so the extras are declared in the repository rather than in the README text. If a pass you want is missing, that config file is the place to look.
The first real use is the automatic optimizer. This command downloads the model, quantizes it, captures the ONNX graph and optimizes it:
olive optimize \
--model_name_or_path Qwen/Qwen2.5-0.5B-Instruct \
--precision int4 \
--output_path models/qwenThe README gives a PowerShell variant because line continuation differs between shells, so the same command with backticks instead of backslashes:
olive optimize `
--model_name_or_path Qwen/Qwen2.5-0.5B-Instruct `
--output_path models/qwen `
--precision int4After the run, the output path holds the optimized model. The README's next step is inference on the ONNX Runtime, using a sample chat application it links as model-chat.py. On Windows, the README documents a huggingface_hub warning about symlink support and lists four ways to handle it, including setting HF_HUB_DISABLE_SYMLINKS_WARNING, which suppresses the warning but does not recover the disk space. The README is explicit that this is a huggingface_hub limitation rather than an Olive one, and it frames the choice in terms of what your environment permits, for example company policy.
Where the supported-architecture list stops
The README states that Olive can automatically optimize popular model architectures such as Llama, Phi, Qwen and Gemma out of the box, and links to the ONNX exporter's overview page for the detailed list. That link is the real boundary. Olive's automatic path is bounded by what the exporter underneath it can convert, so "Olive supports your model" and "the ONNX exporter supports your model" are the same statement. A model that the exporter handles only through a contributed patch is, for practical purposes, outside the automatic path.
For anything outside that list, the README describes the escape hatch: provide details on the model's inputs and outputs through an io_config. That is a manual step and it is where the toolkit stops being automatic. You need to know the input and output names, shapes and dtypes of the graph you are producing, which in practice means reading the model's forward signature and its tokenizer configuration. Nothing in the README suggests this is a small amount of work, and getting it wrong produces a graph that captures successfully but returns the wrong tensors, which is a failure that shows up at inference time rather than at conversion time.
The second limitation is structural rather than architectural. Olive produces ONNX. If your serving stack is PyTorch with torch.compile, or TensorRT with its own engine format, or a vendor SDK that expects its own intermediate representation, the output of an Olive run is not directly consumable. There is no documented path back to the source framework's format.
A third constraint is the alpha classifier in setup.py. It is the project's own statement about maturity, and it sits awkwardly next to a release history that already includes NPU support and multiple quantization algorithms. Treat it as a warning that configuration keys and pass names can move between minor versions.
Olive compared with writing the ONNX export yourself
The obvious alternative is the ONNX exporter that Olive already depends on, used directly. The difference is what each one owns. The exporter owns the conversion from a framework model to an ONNX graph. It does not own quantization, it does not own graph-level optimization passes, and it does not own the decision about which precision to use.
Using the exporter directly means writing the quantization call yourself, choosing the calibration data, and running the optimization passes as a separate step. That is more code, but every decision is visible in your repository instead of being resolved inside a configuration search. For a single model and a single target, that visibility is often worth more than the automation, because a reviewer can read the script and see exactly which precision was chosen and why.
Olive's value shows up when the number of combinations grows: several precisions, several execution providers, and an accuracy constraint that has to hold across them. At that point the search over configurations is doing work you would otherwise do by hand. The cost is that the pipeline's behaviour is described across the README, the documentation site and the recipes repository, and the README notes that Olive examples were relocated to microsoft/olive-recipes. Anyone following an older tutorial will find the example paths have moved, and a moved example is indistinguishable from a broken one until you check the recipes repository.
A second alternative worth naming is the ONNX Runtime itself. If your model is already available in ONNX form from the model repository, running it through Olive adds a conversion step you may not need. Quantization is the part that is hard to replace by hand, so the honest test is whether you need quantization or graph optimization at all. If the model already ships in the precision you want, Olive's main selling point does not apply.
Licence, release cadence and upgrade cost
Olive is MIT licensed, and the classifier in setup.py confirms "License :: OSI Approved :: MIT License". The same classifier block declares "Development Status :: 3 - Alpha". That is the project's own description of its maturity, and it is worth weighing against the release history: v0.11.0 in January 2026, v0.12.0 in April 2026 and v0.13.0 in June 2026, with the last push to main on 2026-09-27. The version numbers stay in the 0.x range, which is consistent with the alpha classifier. The cadence is roughly quarterly for tagged releases, with commits in between.
MIT is permissive. It places no conditions on the models you produce with the toolkit, and it does not reach the licences of the models you download. Those come from the model repositories, and the README's quickstart pulls from the Hugging Face Hub, so the licence you have to check is the one attached to the model, not the one attached to Olive. A model with a non-commercial licence stays non-commercial after quantization. This is not legal advice.
Upgrade cost is driven by the extras. Because optional dependencies are declared in olive/olive_config.json and read at install time by setup.py, the set of packages a given Olive version pulls in can change between releases. Pinning the Olive version and the extras together is the only way to keep an environment reproducible across an upgrade. The repository also carries requirements.txt with torch, transformers, onnx, onnxscript, optuna and opentelemetry-sdk among others, so the base dependency set is not small. A pinned lockfile for the whole environment, not just olive-ai, is the practical answer.
The open questions around upgrades are the ones the README does not answer. It does not document a deprecation policy for config keys, and it does not describe how to migrate a saved optimization config between versions. If you keep Olive configs in version control, plan to re-run them after each upgrade rather than assuming they still mean the same thing.
What the repository layout tells you about the project's scope
The top level contains olive/ for the package, docs/, notebooks/, test/, scripts/, mcp/ and skills/. The presence of mcp/ and skills/ alongside the optimizer is a signal about direction: the project is being packaged for use by agents and assistants, not only by humans typing CLI commands. That is separate from the optimization pipeline and does not change how a conversion run behaves.
AGENTS.md at the repository root is another marker of the same shift. Neither directory is documented in the README excerpt, so anyone evaluating Olive for an agent workflow should read those directories rather than assume the README covers them. The README's news list, meanwhile, points outward to notebooks and to the VS Code AI Toolkit extension, which suggests the CLI is one of several front ends rather than the only one.
The test configuration in pyproject.toml defines a single marker, openvino, for "tests requiring openvino specific packages". That tells you OpenVINO is treated as an optional backend with its own test slice, which matches the extras model. It also tells you the test suite is not uniform across targets: some paths are exercised only when the OpenVINO packages are present, so a bug in an OpenVINO pass may not be caught by a default test run.
The lint configuration is unusually permissive. The pylint disable list in pyproject.toml turns off too-many-arguments, too-many-locals, too-many-branches and a long list of similar checks, with a comment noting that import order is handled by the formatter instead. Ruff is configured at a line length of 120 with a target of py39. None of this affects users of the toolkit, but it does tell you the codebase is written for breadth of supported passes over uniformity of style, and that reading a pass to understand its behaviour will sometimes mean holding a long function in your head.
Editorial conclusion
Adopt Olive if you have already committed to ONNX Runtime and want quantization, graph capture and optimization run as one pipeline instead of three hand-written scripts. Do not adopt it if your deployment target is PyTorch, TensorRT or a vendor runtime Olive has no pass for; the output is an ONNX model and the toolkit has no path back. Before committing, verify three things: that your model architecture appears in the supported list Olive inherits from the ONNX exporter, that the precision you want is one the automatic optimizer actually applies, and that the extras named in olive/olive_config.json install cleanly on your platform, since the base pip install olive-ai does not pull in every optimization backend.
Frequently asked questions
What is Microsoft Olive?
It is an AI model optimization toolkit for the ONNX Runtime, described in the README as composing suitable optimization techniques to output efficient ONNX models for cloud or edge inference. The name is short for Onnx LIVE.
How do I install Microsoft Olive?
The README recommends a virtual environment or conda environment, then pip install olive-ai, followed by pip install transformers onnxruntime-genai for the quickstart. Additional features come from optional dependencies listed in olive/olive_config.json.
What does the olive optimize command do?
According to the README, the automatic optimizer acquires the model from the Hugging Face repo, quantizes it to the requested precision using GPTQ, captures the ONNX graph with weights in an ONNX data file, and optimizes that graph.
Which model architectures can Olive optimize automatically?
The README names Llama, Phi, Qwen and Gemma as architectures that work out of the box, and links to the ONNX exporter's overview page for the detailed list. Other architectures require you to supply input and output details through an io_config.
What licence does Microsoft Olive use?
The repository is MIT licensed, and setup.py carries the classifier "License :: OSI Approved :: MIT License". The same classifier block declares the project's development status as Alpha.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/microsoft-olive)