Open-source project
raskr/rust-autograd avatar
raskr/rust-autograd

rust-autograd is a 2.0.0 release candidate with a blas feature you wire yourself

Tensors and differentiable operations (like TensorFlow) in Rust

503 stars40 forksRustMIT

At a glance

What is it?
Tensors and reverse-mode automatic differentiation for Rust, built on ndarray rather than on a tensor runtime of its own, with neural network support described as low-level and inspired by TensorFlow and Theano. The manifest reads 2.0.0-rc3, the repository publishes no releases, and the default branch was last pushed on 2 June 2026.
Who is it for?
Adopt rust-autograd if you want automatic differentiation and a small tensor API inside an existing Rust program rather than a machine learning framework, and you are willing to write the training loop yourself. Because it is built on ndarray, anything you already do with arrays composes with it, and the blas feature is the difference between usable and slow on matrix-heavy work.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 124 days ago.
What is it written in?
Mainly Rust, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 3, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The manifest says 2.0.0-rc3 and the repository has published no releases

Start with the packaging, because it sets expectations for everything else. The crate is named autograd, the version in Cargo.toml is 2.0.0-rc3, and the edition is 2021, so the Rust toolchain expectation is a recent one. GitHub lists no releases for the repository at all, which means the release candidate in the manifest is what you would be tracking rather than a tagged series you can pin to with confidence. The licence is MIT and the manifest points at a LICENSE file rather than declaring a licence expression inline. Dependencies are unremarkable and mostly small: rand, rand_distr and rand_xorshift for randomness, ndarray at 0.16.1 with the serde and approx features enabled, rayon for parallelism, libc, matrixmultiply for the matrix kernels, num-traits and num, rustc-hash, smallvec, uuid with the v4 feature, and the serde stack. A crate called special is there for the awkward maths. The keyword list is numerics, machine-learning, ndarray, multidimensional and neural-network, which is an accurate summary of the scope. The default branch is master, and it was last pushed on 2 June 2026.

Tensors wrap ndarray, so Tensor::map() is the integration point

The design decision that shapes everything else is that tensors are backed by ndarray rather than by a bespoke array type, and the README shows the consequence directly. Tensor::map() applies ndarray's own methods to the underlying data, so a fold along an axis is written against ndarray's API and not against something autograd invented:

rust
use autograd as ag;
use ag::tensor_ops::*;
use ag::ndarray;

// `Tensor::map()`
ag::run(|ctx| {
    let x = ones(&[2, 3], ctx);
    // apply ndarray's methods
    let y = x.map(|x| x.fold_axis(ndarray::Axis(0), 0.0, |acc, x| acc + x));
    let z = x.map(|x| ag::ndarray_ext::zeros(x.shape()));
});

// Hooks
ag::run(|ctx| {
    let x: ag::Tensor<f32> = ones(&[2, 3], ctx).show_shape();
    let y: ag::Tensor<f32> = ones(&[2, 3], ctx).raw_hook(|x| println!("{}", x));
});

Two things sit next to map() in that snippet. show_shape() returns a tensor carrying the shape as a value, which is the kind of primitive that lets a program treat metadata as part of the graph, and raw_hook() takes a closure over the raw data, which is how you print or inspect intermediates without leaving the automatic differentiation machinery. The second half of the snippet is the debugging story: hooks exist because a tape-based gradient you cannot inspect is a tape-based gradient you cannot debug.

The blas feature is an empty flag, and the implementation is your problem

Linear algebra is where a tensor library either keeps its promise or does not, and here the decision is pushed onto the user. The manifest declares a feature named blas that is empty, meaning enabling it changes no dependency by itself. What it does is make three sibling features available, each of which pulls in the bindings for a specific implementation: accelerate for macOS only, intel-mkl for Intel and AMD CPUs only and including Vector Mathematics operations, and openblas. The README states the reason plainly: if you use basic linear algebra operations, especially matrix multiplications, the blas feature matters for speed. So the dependency line in your own manifest names the crate and the two features you want, for example autograd with features blas and your chosen implementation, and the angle-bracket placeholders in the README are there because the value differs per platform. The optional dependencies behind those features are blas-src, intel-mkl-src and cblas-sys, all declared with default features off. Two consequences follow. The crate compiles without any of them, so a first build will not fail on a missing BLAS, and the performance you get by default is the unaccelerated one.

Reverse-mode differentiation through placeholders and a second grad pass

The worked example in the README is small on purpose. It computes partial derivatives of z = 2x^2 + 3y + 1 inside ag::run(), which takes a context, and it creates x and y as placeholders rather than as constants. The first derivative, dz/dy, comes from grad(&[z], &[y]) followed by eval on the context, and the printed result is Ok(3.). The derivative with respect to x is the same call, and because x is a placeholder you have to feed it a value first, which the example does with an ndarray scalar of 2. fed into the evaluator and pushed through the graph, giving Ok(8.). The third line is the interesting one for anyone planning to use this for anything beyond a first-order model: grad is applied to the already-computed gradient gx, producing ddz/dx, which is higher-order differentiation obtained by differentiating the tape again rather than by a separate implementation. The result type is a Result, and the Ok wrapper is visible in every printed value, which tells you error handling is part of the return path rather than something bolted on.

Neural network support is a VariableEnvironment, an Adam and a loop you write

The training example is where the TensorFlow and Theano lineage shows, and the description is careful about the level: various low-level features inspired by those libraries, with the claim that computation graphs require only a bare minimum of heap allocations, so the overhead stays small even for complex networks. The code matches that description. You create an ag::VariableEnvironment, take an rng from ag::ndarray_ext as an ArrayRng over f32, and register named variables in it, setting w from rng.glorot_uniform with the shape [28 * 28, 10] and b from zeros with the shape [1, 10], which is the standard fan-in fan-out initialisation for a 784 to 10 layer. The optimiser is constructed by hand with Adam::default given a name, the current variable ids from the environment's default namespace, and a mutable borrow of the environment, after which training is an ordinary epoch loop calling env.run(). One number is attached to that loop as a comment, 0.11 seconds per epoch on a 2.7GHz Intel Core i5, which is a figure from the machine that wrote the example rather than a benchmark, and it is the only performance statement in the README.

Four declared examples, and MNIST arrives through a shell script

The manifest declares four example targets explicitly, and each one is a named entry point you can run with cargo: mlp_mnist for the multi-layer perceptron digit classification, lstm_lm for a language model, cnn_mnist for a convolutional network on the same dataset, and sine for the small synthetic curve. Two more files sit in the examples directory without being declared as targets, and they are the more revealing pair. mnist_data.rs is the shared data handling, and download_mnist.sh is a shell script that fetches the dataset. That is a design choice with consequences: there is no dataset pipeline in the crate, no async downloader, and no cache layer, so a first run needs the script to have been run, and anyone embedding this in a service has to solve data fetching themselves. The repository layout is correspondingly spare, with src/, tests/, examples/, a Cargo.toml, a CHANGELOG.md, a licence file and a .github directory holding the workflow that builds the crate. There is no separate documentation tree in the repository; the published API reference lives on docs.rs instead.

Editorial conclusion

Adopt rust-autograd if you want automatic differentiation and a small tensor API inside an existing Rust program rather than a machine learning framework, and you are willing to write the training loop yourself. Because it is built on ndarray, anything you already do with arrays composes with it, and the blas feature is the difference between usable and slow on matrix-heavy work. Do not adopt it expecting a framework: there is no GPU backend in the dependency list, no model zoo, no serving story, and the neural network support is described as low-level and inspired by TensorFlow and Theano rather than as a replacement for them. Four things to check before you commit. Whether the crates.io version you would depend on matches the 2.0.0-rc3 in the manifest, since the repository publishes no releases and a release candidate is what the default branch carries. Which BLAS you will point the feature at, because the choice is constrained by platform. Whether the ndarray version it pins still lines up with the rest of your dependency tree. And where your data comes from, since one of the examples fetches MNIST with a shell script rather than a library call.

Frequently asked questions

What is an AutoGrad?

In this repository it is the autograd crate, which provides tensors and differentiable operations in Rust backed by ndarray. It implements reverse-mode automatic differentiation, so you compute partial derivatives by calling grad over the outputs and the inputs you care about, and it can be applied again to a gradient for higher-order derivatives.

How do I make rust-autograd fast for matrix multiplication?

Enable the blas feature together with an implementation feature for your platform: accelerate on macOS, intel-mkl on Intel or AMD CPUs, or openblas. The blas feature on its own is an empty flag, so the implementation choice is what pulls in the bindings.

Which Rust version does the autograd crate target?

The manifest declares edition 2021 and version 2.0.0-rc3. The repository publishes no GitHub releases, and its default branch is master, which was last pushed on 2 June 2026.

Does rust-autograd include neural network training?

It provides low-level features inspired by TensorFlow and Theano, including an optimizers module with Adam, a VariableEnvironment for registering named variables, and random initialisation helpers such as glorot_uniform. The examples cover a multi-layer perceptron on MNIST, a convolutional network on MNIST and an LSTM language model.

How do the autograd examples get their data?

The examples directory contains a download_mnist.sh script alongside mnist_data.rs, so the MNIST dataset is fetched by running the script rather than by a library call inside the crate. The manifest declares mlp_mnist, lstm_lm, cnn_mnist and sine as example targets.

Official sources

  1. Issues
  2. License: MIT
  3. raskr/rust-autograd on GitHub
  4. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/raskr-rust-autograd.svg)](https://hysenlabs.com/projects/raskr-rust-autograd)