huggingface/optimum: A Hardware Backend Hub for Transformers Models
🚀 Accelerate inference and training of 🤗 Transformers, Diffusers, TIMM and Sentence Transformers with easy to use hardware optimization tools
At a glance
- What is it?
- Optimum is a thin Python layer that sits between Hugging Face training and inference libraries and vendor-specific accelerator stacks. It is useful when your target hardware has an official backend; it is not a general-purpose speedup for arbitrary PyTorch code.
- Who is it for?
- Adopt Optimum if your deployment target already has a supported backend, such as OpenVINO on Intel, Gaudi on Habana, or Inferentia on AWS, and you want to keep the Transformers API while exporting or quantizing. Do not adopt it expecting a generic speedup on plain CUDA: the base package is a dispatcher, and the heavy lifting lives in separate packages such as optimum-onnx and optimum-intel.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 5 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem Optimum solves is backend fragmentation, not model quality
A team that trains a Transformer and then has to serve it on Intel CPUs, Habana Gaudi cards, AWS Inferentia chips and NVIDIA GPUs ends up maintaining four separate export and quantization pipelines. Each vendor ships its own converter, its own calibration format and its own runtime wrapper. Optimum's answer is to keep the model definition in the Transformers, Diffusers, TIMM or Sentence-Transformers ecosystem and move only the hardware-specific work behind a consistent set of classes and command-line entry points.
The audience is therefore narrow and specific: engineers who own the deployment step and already know which accelerator they are buying. If you are still choosing hardware, Optimum gives you little, because the interesting code lives in per-vendor packages rather than in the base install. The README describes the project as an extension of Transformers, Diffusers, TIMM and Sentence-Transformers that provides optimization tools for training and running models on targeted hardware.
How the package is layered: a dispatcher plus vendor packages
The repository layout shows a single `optimum/` package alongside `docs/`, `notebooks/` and `tests/`. The base install is deliberately small. According to setup.py, the required dependencies are `transformers>=4.29`, `torch>=1.11`, `packaging`, `numpy` and `huggingface_hub>=1.31.0`. Everything accelerator-specific is an extra that resolves to a separate distribution: `optimum-onnx`, `optimum-intel`, `optimum-habana`, `optimum-neuronx`, `optimum-amd`, `optimum-furiosa`, `optimum-graphcore` and `optimum-quanto`.
That layering explains most of the confusing behaviour people hit. Installing `optimum` alone gives you the shared abstractions and the version checks, not the ONNX exporter or the OpenVINO quantizer. The extras are pinned to minimum versions in setup.py, for example `optimum-habana>=1.17.0` and `optimum-intel>=1.23.0`, which means the base library and the vendor package can drift apart if you upgrade one without the other. The `--upgrade --upgrade-strategy eager` flag the README repeats in every install line exists precisely to pull those vendor packages forward together.
Installing Optimum and exporting a model for ONNX Runtime
Start with the base package. The README gives this as the plain install:
python -m pip install optimumFor ONNX work the README points at the ONNX extra, and repeats the eager upgrade strategy so the vendor package is not left behind:
pip install --upgrade --upgrade-strategy eager optimum[onnx]A warning in the README states that the ONNX integration was moved to the `optimum-onnx` repository, so the installation instructions for that path live there rather than in this repository. The same README notes that ONNX Runtime itself is installed separately, following the instructions on the ONNX Runtime site.
If you need a build that is not on PyPI, the README documents installing from source, and appending the extra to the same URL:
python -m pip install git+https://github.com/huggingface/optimum.git
python -m pip install optimum[onnxruntime]@git+https://github.com/huggingface/optimum.gitAfter installation, the entry point you use depends on the backend. The README describes exporting Transformers, Diffusers, Sentence Transformers and TIMM models to ONNX and then running them through `ORTModelForXXX` classes with ONNX Runtime in the backend. The export step is documented as available both programmatically and from a command line. What you should see after a successful export is a directory of ONNX graph files plus the tokenizer and config artifacts, which the ORT model classes then load.
The ONNX split is the biggest practical trap
The most consequential limitation is organisational. The README carries an explicit warning that ONNX integration moved to `optimum-onnx`, and setup.py confirms it: the `onnx`, `onnxruntime` and `onnxruntime-gpu` extras all resolve to `optimum-onnx` variants rather than to code in this repository. Anyone reading an older tutorial that imports ONNX helpers from `optimum` directly will find the code path gone or redirected.
A second limitation is version coupling. Because the extras carry minimum version constraints, a working ONNX or OpenVINO pipeline can break when you upgrade `optimum` alone, or when you upgrade the vendor package without the base library. The README's insistence on `--upgrade --upgrade-strategy eager` is a mitigation, not a guarantee, and it will also pull unrelated packages forward in the same environment.
Third, Optimum is the wrong tool if your model architecture is not already supported by the upstream library. Optimum routes and converts; it does not implement new architectures. If your model is not in Transformers, Diffusers, TIMM or Sentence-Transformers, there is nothing here to route.
Optimum compared with exporting to ONNX by hand
The direct alternative is writing your own `torch.onnx.export` call plus a calibration script for quantization, then loading the result with a bare ONNX Runtime session. That approach gives you full control over dynamic axes, opset version and input names, and it has no dependency on Hugging Face release cadence.
The difference in approach is where the complexity sits. Hand-rolled export puts graph-shape decisions and preprocessing alignment on you, and you will rediscover the tokenizer and config plumbing that the ORT model classes already handle. Optimum moves that plumbing into maintained classes at the cost of accepting its version constraints and its package split. For a single model on a single backend, hand-rolling is often less machinery. For a fleet of models across several accelerators, the shared abstraction is the point.
Maintenance, licence and upgrade cost
The repository is not archived, and the last push was on 2026-09-18, days before this writing, so the project is being worked on. The release history shows v2.2.0 titled "Transformers v5 Support & Deprecation Cleanup" and v2.1.0 titled "Transformers v5 compatibility and full transition to GPTQModel". Two consecutive release notes about compatibility and deprecation cleanup tell you the upgrade cost is real: staying current means tracking Transformers major versions and, in the GPTQ case, a change of quantization backend.
The licence is Apache-2.0, the same permissive licence used across the Hugging Face stack, which keeps it compatible with most commercial deployment. That is a statement about the licence text, not legal advice; if you redistribute a modified vendor package, read the actual terms.
Ongoing cost is dominated by the extras matrix. Each accelerator extra is a separate package with its own release cadence, and the base library only pins minimum versions. Budget for periodic dependency reconciliation rather than a one-time install.
Editorial conclusion
Adopt Optimum if your deployment target already has a supported backend, such as OpenVINO on Intel, Gaudi on Habana, or Inferentia on AWS, and you want to keep the Transformers API while exporting or quantizing. Do not adopt it expecting a generic speedup on plain CUDA: the base package is a dispatcher, and the heavy lifting lives in separate packages such as optimum-onnx and optimum-intel. Before committing, check the extras table in setup.py against your accelerator and confirm which package actually owns the code path you need, because the ONNX integration has moved out of this repository into optimum-onnx.
Frequently asked questions
What is huggingface/optimum used for?
It provides optimization tools that let you train and run Transformers, Diffusers, TIMM and Sentence-Transformers models on specific hardware backends, including ONNX Runtime, OpenVINO, Gaudi, Inferentia and TensorRT-LLM. The README describes it as an extension of those libraries rather than a replacement.
How do I install huggingface/optimum?
The README gives `python -m pip install optimum` for the base package. Accelerator features come from extras such as `optimum[onnx]`, `optimum[openvino]` or `optimum[habana]`, installed with `--upgrade --upgrade-strategy eager` so the vendor packages are upgraded together.
Where did the ONNX support in huggingface/optimum go?
The README states that ONNX integration was moved to the `optimum-onnx` repository, and setup.py shows the `onnx`, `onnxruntime` and `onnxruntime-gpu` extras resolving to `optimum-onnx` packages. Follow the installation instructions in that repository for ONNX work.
Does installing huggingface/optimum alone give me GPU acceleration?
No. The base dependencies listed in setup.py are `transformers`, `torch`, `packaging`, `numpy` and `huggingface_hub`. Hardware-specific behaviour comes from extras such as `optimum[onnxruntime-gpu]`, `optimum[openvino]` or `optimum[habana]`, which install separate packages.
Which Python and Transformers versions does huggingface/optimum require?
setup.py lists `transformers>=4.29` and `torch>=1.11` as required dependencies. The PyPI badge in the README tracks the supported Python versions, and the v2.2.0 and v2.1.0 release notes are titled around Transformers v5 support, so the upper end of that range moves with Transformers releases.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/huggingface-optimum)