Library / SDK
pykeio/ort avatar
pykeio/ort

ort: ONNX Runtime Bindings for Rust Inference and Training

Fast ML inference & training for ONNX models in Rust

2,533 stars275 forksRustApache-2.0

At a glance

What is it?
ort wraps Microsoft's ONNX Runtime in safe Rust, so you can run or fine-tune ONNX models on CPU, CUDA and other accelerators. It is still on 2.0.0 release candidates, and the docs are thin on rollback and packaging details.
Who is it for?
Adopt ort when your models are already exported to ONNX and your application is Rust, especially if you need to pick an execution provider per deployment target. Do not adopt it if you need a stable, non-RC API surface, if your team wants a Python-first toolchain, or if you cannot ship the native ONNX Runtime libraries alongside your binary.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 6 days ago.
What is it written in?
Mainly Rust, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What ort solves, and who is meant to use it

Rust has no shortage of tensor libraries, but it has few ways to run a model that was trained somewhere else. That is the gap ort fills. It is a Rust interface for inference and training on models in the Open Neural Network Exchange format, and it is primarily a wrapper around Microsoft's ONNX Runtime, with support for other pure-Rust runtimes listed under backends on the project site. The README frames the audience plainly: when you need to deploy a PyTorch, TensorFlow, Keras, scikit-learn or PaddlePaddle model either on-device or in the datacenter, ort is the layer that loads the exported ONNX graph and executes it.

The practical consequence is that you keep your training stack wherever it already is, export to ONNX, and write the serving or embedding side in Rust. The README lists projects doing exactly that: Text Embeddings Inference uses ort for text embedding inference, Magika uses it for file type detection, FastEmbed-rs uses it to generate embeddings and rerank locally, and Ultralytics ships a pure Rust YOLO inference library and CLI built on ort. Those are different domains, but the same shape: a model trained elsewhere, executed inside a Rust process, often on a user's machine rather than a server.

How ort talks to ONNX Runtime

ort is split into two crates. The workspace Cargo.toml declares members = [ 'ort-sys' ] and default-members = [ '.' ], so ort-sys holds the low-level bindings and ort is the safe layer above it. The package description in Cargo.toml calls it a safe Rust wrapper for ONNX Runtime 1.28, and the README badge pins the same version. That pairing matters: ort tracks a specific ONNX Runtime release, so upgrading ort can mean moving the native runtime underneath it.

Execution providers are the mechanism that makes the crate portable across hardware. They are Cargo features rather than runtime strings. The docs.rs metadata alone lists cuda, tensorrt, openvino, onednn, directml, nnapi, coreml, xnnpack, rocm, acl, armnn, tvm, migraphx, rknpu, vitis, cann, qnn, webgpu, azure, nvrtx and vsinpu. You enable the ones you need at build time and the session picks a provider from there. The default feature set is std, ndarray, tracing, download-binaries, tls-native, copy-dylibs and api-27, which tells you the intended baseline is a CPU build with binaries fetched for you and ndarray as the tensor type.

The repository layout reinforces the two-layer design. ort-sys sits at the top level next to src/, and backends/ holds separate runtime integrations (candle, tract, web) that are excluded from the main workspace, meaning they are built on their own terms rather than as part of the core crate.

Installing ort and running a first session

The crate is published as ort, so the install step is a normal Cargo dependency. The version in Cargo.toml is 2.0.0-rc.13, and the release list shows v2.0.0-rc.13 dated 2026-07-28, so pin that release candidate explicitly rather than expecting a stable 2.0.0.

toml
ort = "2.0.0-rc.13"

That line goes in the [dependencies] table of your own Cargo.toml. With the default features you get download-binaries, which fetches the ONNX Runtime native libraries for your platform, and copy-dylibs, which copies them next to your build output. If you would rather point at a runtime you installed yourself, the docs.rs feature list includes load-dynamic for that case. The README points to the guide at ort.pyke.io and to the examples directory for runnable code; the examples include gpt2, yolov8, sentence-transformers, model-info and training, and each has its own directory under examples/. Those examples are excluded from the workspace, so they are built individually rather than with a plain workspace build.

For a first real use, start with the model-info example rather than a full inference pipeline. It answers the question you will hit immediately: which inputs and outputs does this ONNX file actually expose, and what shapes do they have? Getting that from the model itself beats guessing from the training code. Once you know the input names and shapes, the sentence-transformers or yolov8 example is the shortest path to a working session, because both show the load, prepare input, run, read output sequence end to end.

The 2.0 release candidates and the migration tax

The most important limitation is not technical, it is scheduling. The newest published version is 2.0.0-rc.13, released 2026-07-28. The two before it are v2.0.0-rc.12 from 2026-03-05 and v2.0.0-rc.11 from 2026-01-07. That is a release candidate line that has been running for a while, and the README links a dedicated page for migrating from v1.x to v2.0. If you are starting a project today you are starting on an RC, and you should expect the API to shift between candidates. The repository's last push was on 2026-09-02, so work is ongoing, but ongoing work on an RC line is not the same as a frozen API.

There is a second constraint that is easy to miss until you try to ship. The default features include download-binaries and copy-dylibs. That is convenient during development and awkward during packaging: your build now depends on fetching native libraries, and your output needs those shared libraries beside it. The load-dynamic feature exists for teams that want to manage the runtime themselves, but choosing it means you own version matching between ort-sys and the ONNX Runtime you supply. The README does not document rollback or how to pin a specific ONNX Runtime build, so verify that against your own distribution requirements rather than assuming it is handled.

Where ort is the wrong tool

If your serving path is already Python, ort is probably the wrong layer. ONNX Runtime has first-class Python bindings, and adding a Rust crate between your model and your service buys you nothing unless the surrounding application is Rust. The same applies if your team's strength is Python tooling and the deployment target is a container you control; you would be paying a build-system cost for a language boundary you do not need.

A subtler case is dynamic model loading. ort's execution providers are compile-time Cargo features. That is a deliberate design choice and it produces smaller, more predictable binaries, but it means a single build cannot switch from CPU to CUDA based on a runtime config value. If your product needs one artifact that adapts to whatever accelerator is present, you are looking at either multiple builds or a different runtime. The backends directory shows the project does care about alternative runtimes, but those integrations are excluded from the workspace and are not the default path.

Finally, if your model is not in ONNX and converting it is expensive or lossy, that conversion cost sits outside ort entirely. The crate executes ONNX graphs; it does not train PyTorch checkpoints in place or convert them for you.

ort against a pure-Rust runtime such as tract

The clearest alternative in the same ecosystem is a pure-Rust inference runtime, and the repository itself names tract: backends/tract is one of the excluded workspace members. The difference in approach is fundamental rather than cosmetic. ort delegates to ONNX Runtime, a mature C++ library with a long list of hardware execution providers, so you inherit its operator coverage and its accelerator support, at the cost of shipping and version-matching native libraries. A pure-Rust runtime removes the native dependency entirely, which simplifies cross-compilation and packaging, but you are relying on that project's own operator implementations and its own set of supported backends.

That trade-off decides itself based on your deployment target. If you need CUDA, TensorRT, CoreML or one of the other providers in ort's feature list, the native dependency is the price of admission and a pure-Rust runtime is unlikely to match the coverage. If you are shipping to a constrained or unusual target where pulling in a C++ runtime is the hard part, the pure-Rust route is the one that gets you there. ort also lists candle and web under backends/, so the project treats ONNX Runtime as the primary path rather than the only one.

Licence, maintenance and upgrade cost

The repository carries both LICENSE-APACHE and LICENSE-MIT at the top level, and Cargo.toml declares license = "MIT OR Apache-2.0". The GitHub metadata lists Apache-2.0, so the dual grant is the accurate picture. That is the same permissive pairing used across much of the Rust ecosystem, which keeps it compatible with most application licences, but it says nothing about the licence of ONNX Runtime itself or of the models you load. Those are separate questions, and the README does not address them.

On maintenance, the last push was on 2026-09-02, which is recent, and the repository is not archived. The Cargo.toml even carries a [badges.maintenance] entry with status = "actively-developed". The upgrade cost is the part to budget for. Because ort-sys tracks ONNX Runtime 1.28 and execution providers are compile-time features, a version bump can change both the Rust API and the native runtime in one step. The migration guide at ort.pyke.io/migrating/v2 exists precisely because v1 to v2 is not a drop-in change. For a project on a long support window, pinning to a specific 2.0.0-rc release and treating upgrades as scheduled work is the realistic approach.

Editorial conclusion

Adopt ort when your models are already exported to ONNX and your application is Rust, especially if you need to pick an execution provider per deployment target. Do not adopt it if you need a stable, non-RC API surface, if your team wants a Python-first toolchain, or if you cannot ship the native ONNX Runtime libraries alongside your binary. Before committing, check on docs.rs whether the feature flags you need are present in the 2.0.0-rc.13 API, and read the v1.x to v2.0 migration guide at ort.pyke.io/migrating/v2, because the crate version in Cargo.toml is 2.0.0-rc.13 and the API is still moving between release candidates.

Frequently asked questions

What is ONNX used for?

ONNX is the Open Neural Network Exchange format, and ort uses it as the model format it loads and executes. The README describes exporting models from PyTorch, TensorFlow, Keras, scikit-learn or PaddlePaddle and deploying them on-device or in the datacenter.

Does ort support GPU acceleration?

Yes, through execution providers enabled as Cargo features. The docs.rs feature list includes cuda, tensorrt, rocm, directml, coreml, openvino, webgpu and others, and the README says ort supports almost any hardware accelerator.

Is ort stable enough for production?

The newest release is 2.0.0-rc.13, so the 2.0 line is still in release candidates as of the release list. A migration guide for v1.x to v2.0 exists, which indicates the API changed between major versions.

Can ort train models, or only run inference?

The README describes ort as an interface for inference and training, and Cargo.toml exposes a training feature that enables ort-sys/training. There is also a training example directory in the repository.

Official sources

  1. License: Apache-2.0
  2. Project website
  3. pykeio/ort on GitHub
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/pykeio-ort.svg)](https://hysenlabs.com/projects/pykeio-ort)