Model or dataset
sonos/tract avatar
sonos/tract

sonos/tract: a tiny Rust inference runtime for ONNX and NNEF models

Tiny, no-nonsense, self-contained, Tensorflow and ONNX inference

3,077 stars291 forksRustNOASSERTION

At a glance

What is it?
Tract is Sonos' self-contained neural-network inference engine in Rust. It loads ONNX, NNEF and legacy TensorFlow Lite models, optimises them, and runs them on CPUs, GPUs, WebAssembly, or ships them as a small NNEF-based runtime.
Who is it for?
Adopt tract if you need a small Rust runtime that can execute an optimised NNEF model on a CPU, a GPU or in WebAssembly, and you are willing to convert models at build time. Do not adopt it if you require TensorFlow 2 import, ONNX export, or a stable internal crate API: the README points to the tract crate as the only stable surface.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository last received commits 1 day ago.
What is it written in?
Mainly Rust, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem tract solves: inference without a training framework

Running a trained model on a device usually means dragging a training framework along. Tract takes the opposite position. It is a self-contained inference engine written in Rust, and the README describes it as loading ONNX and NNEF models, optimising them, and running them anywhere from embedded ARM CPUs to NVIDIA and Apple GPUs, in the browser through WebAssembly, or on a Linux, macOS or Windows workstation. The target user is an engineer shipping a model into a product, not a researcher training one. Sonos uses it in production for wake-word and streaming speech-recognition workloads, and the repository also carries examples for LLM text generation, text-to-image, object detection and face recognition. The design goal stated in the README is the translate-once / ship-tiny-runtime story enabled by the NNEF-based intermediate format called tract-OPL.

How tract works: load, optimise, prepare, run

The pipeline has four visible stages. A loader reads a model format and produces a model; prepare() then optimises and compiles that model for a chosen runtime; the resulting runnable object executes it. The README's quick-start example shows the shape of the Rust API, with the runtime selected by name. All runtimes share the same TypedModel intermediate representation and the same loaders, so according to the README a model optimised on one platform can be moved to another. The runtime table lists four entries: the default CPU runtime in tract-linalg with hand-rolled SIMD micro-kernels covering x86, ARMv6/7/8 and ARM SVE; metal in tract-metal for Apple GPUs; cuda in tract-cuda for NVIDIA GPUs; and WebAssembly through standard wasm32 targets using the default runtime. There is also a separate pulsification mechanism. tract-pulse translates a network that operates on full sequences into one that processes a fixed-size pulse along its streaming axis at each step, so the same model serves batch evaluation and low-latency real-time inference. The translate-time logic lives in tract-pulse, while the runtime ships only the small tract-pulse-opl crate.

Installing tract and running a first ONNX model

The README points at the tract crate on crates.io as the authoritative public API and warns that the internal crates such as tract-core, tract-nnef and tract-onnx are not stable surface and should not be depended on directly. The Rust example below is the one from examples/onnx-mobilenet-v2. It loads a MobileNet v2 ONNX file, asks for the runtime named "default", prepares the model, and runs it with a single input. The call to tract::impl_ndarray_interop!() brings in the ndarray interop used by the .tract() conversion on the input.

rust
use tract::prelude::*;
tract::impl_ndarray_interop!();

let model = tract::onnx()?
    .load("mobilenetv2-7.onnx")?
    .into_model()?;

let runtime = tract::runtime_for_name("default")?;
let runnable = runtime.prepare(model)?;

let result = runnable.run([input.tract()?])?;

For Python, the README gives a single install command and says the API mirrors the Rust pipeline: load a model, set input facts, optimise, then run. Documentation for that binding lives at sonos.github.io/tract, and the source is in api/py/.

sh
pip install tract

The deployment workflow the README recommends is different from the development one. Convert once at build time with the tract CLI, then ship only the small runtime crates. The command below dumps an ONNX model to a compressed NNEF archive; the README shows the same command with the output named model.nnef.tgz.

sh
tract model.onnx dump --nnef model.nnef.tgz

At runtime, the README says to ship tract-core plus tract-nnef, adding tract-onnx-opl if the model uses ONNX-only operators and tract-pulse-opl if it is pulsified. That keeps protobuf and the training-framework loaders out of the binary.

Format coverage and the tract-OPL versioning rule

The format table is asymmetric and worth reading carefully. ONNX can be loaded but not saved. NNEF with tract-OPL extensions can be both loaded and saved. TensorFlow Lite is listed as legacy and supports both directions. TensorFlow 1 frozen graphs can be loaded but not saved. TensorFlow 2 is not directly supported: the README says to convert to ONNX first. For PyTorch, the project points at torch-to-nnef, an open-source converter maintained alongside tract, which lets you skip the detour through ONNX entirely. On stability, the README separates the two halves of the format. NNEF parts are tied to the NNEF specification and described as very stable; tract-OPL extensions are described as a bit more in flux. The stated rule is that a model serialised with tract 0.x.y should work with tract 0.x.z where z is greater than or equal to y. Models embed a tract_nnef_ser_version property identifying the generating version, but tract itself does not enforce a version check, so the application has to do it. The CHANGELOG is the running list of serialisation-format changes.

Where tract is the wrong tool

The clearest boundary is training and model authoring. Tract loads and runs models; it does not train them, and the README describes no training path. TensorFlow 2 is a second boundary: the README states it is not directly supported and that you must convert to ONNX first, which means a TF2 shop inherits a conversion step it may not want. A third is the internal crate API. Anyone tempted to build directly against tract-core, tract-onnx or tract-nnef is told these are not stable surface, so a project that needs a frozen low-level interface has no guarantee here. There is also a deployment constraint hiding in the recommended workflow: converting to NNEF at build time is what keeps the runtime small, but it also means the model your runtime sees is the result of a conversion, and tract-OPL extensions are the part of the format the README calls more in flux. If your model leans on ONNX-only operators, you keep tract-onnx-opl in the runtime, which eats into the footprint argument. Finally, the pulsification path is aimed at streaming workloads; a purely batch model gains nothing from it and pays the extra conceptual layer.

Tract compared with ONNX Runtime

ONNX Runtime is the obvious alternative, and the difference is architectural rather than a matter of speed. ONNX Runtime is built around ONNX as its native and only real interchange format, with execution providers plugging into that graph. Tract treats ONNX as an import format and NNEF, extended by tract-OPL, as the format it can both read and write. That is what makes the translate-once workflow possible: you convert at build time and ship a runtime that contains no protobuf and no training-framework loaders. The trade is that you now own a conversion step and a serialisation format whose extension half the README describes as more in flux. Tract also carries pulsification as a first-class concept for streaming models, which is a specific answer to wake-word and streaming ASR workloads rather than a general graph optimisation feature. If your deployment is a server with room for a large runtime and your models stay in ONNX, the conversion detour buys you less.

Maintenance, licence and upgrade cost

The repository is not archived, and the last push was on 2026-09-23, one day before this review's date, with releases v0.23.8 on 2026-09-21 and v0.23.7 on 2026-09-08. That is a recent cadence, and the project is developed in the open under the Sonos organisation. Upgrade cost depends on which surface you touch. The Rust crate is the stable API; the internal crates are explicitly not. Serialised models are the other surface, and the compatibility rule is one-directional: a model written by 0.x.y should load in 0.x.z for z at or above y, but nothing in the engine enforces it, so an application that stores models needs its own check against tract_nnef_ser_version. The CHANGELOG is where serialisation changes are listed. On licensing, the README's licence section is truncated at the TensorFlow proto files, and the repository root carries LICENSE, LICENSE-APACHE and LICENSE-MIT alongside a crates.io badge reading MIT/Apache 2. The GitHub metadata reports the licence as NOASSERTION, so the practical reading is a dual MIT or Apache-2.0 grant for the project code with separate terms for the vendored TensorFlow protobuf files. Confirm the terms of those vendored files before redistributing, since they are not covered by the badge.

Editorial conclusion

Adopt tract if you need a small Rust runtime that can execute an optimised NNEF model on a CPU, a GPU or in WebAssembly, and you are willing to convert models at build time. Do not adopt it if you require TensorFlow 2 import, ONNX export, or a stable internal crate API: the README points to the tract crate as the only stable surface. Before committing, run tract model.onnx dump --nnef model.nnef.tgz on your own model and check that the operators you rely on survive the conversion, since tract-OPL extensions are described as more in flux than the NNEF parts.

Frequently asked questions

Does sonos/tract support TensorFlow 2 models?

No. The README states that TensorFlow 2 is not directly supported and that you should convert to ONNX first. TensorFlow 1 frozen graphs are still loadable, with the operator set needed for the classical CV and wake-word models that originally drove the design.

Can sonos/tract export a model back to ONNX?

No. The format table lists ONNX as load-only. NNEF with tract-OPL extensions and TensorFlow Lite are the two formats that can be both loaded and saved.

Which crates should I depend on in a sonos/tract deployment?

The README names the tract crate as the authoritative public API and says the internal crates such as tract-core, tract-nnef and tract-onnx are not stable surface. For the small-runtime workflow it recommends shipping tract-core plus tract-nnef, adding tract-onnx-opl for ONNX-only operators and tract-pulse-opl for pulsified models.

How do I install sonos/tract from Python?

The README gives pip install tract for the PyPI package, which is built on the same Rust core. The Python API mirrors the Rust pipeline: load a model, set input facts, optimise, then run.

Official sources

  1. Issues
  2. README
  3. Releases
  4. sonos/tract on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/sonos-tract.svg)](https://hysenlabs.com/projects/sonos-tract)