Open-source project
HIPS/autograd avatar
HIPS/autograd

HIPS/autograd: reverse-mode and forward-mode differentiation for plain NumPy code

Efficiently computes derivatives of NumPy code.

7,538 stars946 forksPythonMIT

At a glance

What is it?
Autograd differentiates native Python and NumPy functions without a graph you have to build by hand, and it can nest derivatives to arbitrary depth. It is a small, MIT-licensed library for gradient-based optimization and scientific work, not a replacement for a full deep-learning framework.
Who is it for?
Adopt Autograd when your model is already written as NumPy plus Python control flow and you need gradients, Hessians, or higher-order derivatives without rewriting into a framework's tensor type. Skip it if you need GPU execution, distributed training, or a large prebuilt layer zoo; the README lists no accelerator support and the dependency set is numpy<3 with optional SciPy.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The gap Autograd fills: gradients for code you already wrote in NumPy

Most gradient tooling asks you to adopt its tensor type, its device model, and its execution model before you get a single derivative. Autograd takes the opposite position. The README describes it as automatically differentiating native Python and NumPy code, and the intended application it names is gradient-based optimization. That framing matters: the library is for people who have a numerical function already written in NumPy and want its derivative, not for people starting a training pipeline from scratch.

The audience follows from that. Researchers working on Gaussian processes, variational inference, Hamiltonian Monte Carlo, or simulation code that happens to be differentiable are the natural users. The examples directory supports this read: gaussian_process.py, bayesian_neural_net.py, black_box_svi.py, hmm_em.py, and a fluid simulation all sit alongside more conventional neural network examples. Several of those are numerical methods where the derivative is the interesting quantity, not a means to train a network.

The scope claim is narrow and testable. The README says Autograd handles a large subset of Python's features including loops, ifs, recursion, and closures, and that it can take derivatives of derivatives of derivatives. Those are the properties that separate it from a tape-based system that only records tensor operations.

How the tracing works: numpy wrapping, grad, and composable modes

The mechanism visible in the README is a thinly-wrapped NumPy. You import autograd.numpy as np instead of numpy, and that module forwards to NumPy while recording the operations a traced function performs. The README calls this a "Thinly-wrapped numpy". Because the wrapper is thin, ordinary NumPy semantics carry over, which is why Python control flow works: a loop or an if statement executes normally during the trace, and the recorded operations are whatever that execution path actually touched.

From there, grad is the entry point. The README calls it "The only autograd function you may ever need". You pass it a function and get back a function that returns the gradient. The README states Autograd supports reverse-mode differentiation, which it equates with backpropagation, and that reverse mode efficiently takes gradients of scalar-valued functions with respect to array-valued arguments. It also supports forward-mode differentiation, and the two modes can be composed arbitrarily.

That composability is the part worth pausing on. Because modes compose, higher-order derivatives fall out of nesting rather than requiring separate machinery. The README demonstrates this by applying elementwise_grad repeatedly to get first through fourth derivatives of the same function, and separately notes that you can differentiate as many times as you like. The cost is that each nesting level adds trace overhead, and the README gives no performance figures for nested differentiation.

Installing Autograd and taking a first derivative

Installation is a single pip command. The README gives it as pip install autograd. The package metadata in pyproject.toml sets requires-python to >=3.10 and declares a single runtime dependency, numpy<3. There is no compiled extension to build and no CUDA toolkit to install.

The optional SciPy extra is separate. The README says some features require SciPy and offers pip install "autograd[scipy]", which maps to the scipy optional-dependency group in pyproject.toml.

bash
pip install autograd
pip install "autograd[scipy]"

With that installed, the smallest useful program is the README's own tanh example. You import the wrapped NumPy as np, import grad, define a function, and ask grad for its derivative. Note that the function body uses np.exp from the wrapped module, not from plain numpy; that is what makes the trace possible.

python
import autograd.numpy as np
from autograd import grad

def tanh(x):
    return (1.0 - np.exp((-2 * x))) / (1.0 + np.exp(-(2 * x)))

grad_tanh = grad(tanh)
print(grad_tanh(1.0))

The README shows this printing np.float64(0.419974341614026) at x = 1.0, and compares it against a finite-difference estimate of 0.41997434264973155 over the interval 0.9999 to 1.0001. Running that comparison on your own machine is the fastest way to confirm the install is behaving. If the two numbers disagree beyond floating-point noise, the wrapped NumPy import is the first thing to check.

For vectorized functions, the README points to elementwise_grad, imported as egrad, which it describes as being for functions that vectorize over inputs. The same section of the README uses it to plot first through fourth derivatives of tanh across a linspace from -7 to 7. The example file for this is examples/tanh.py.

Where Autograd stops: no accelerator story and a narrow dependency base

The dependency list is the clearest statement of scope. pyproject.toml declares numpy<3 and nothing else at runtime. There is no JAX, no CuPy, no torch, no device abstraction. Everything runs where NumPy runs, which in practice means CPU. If your training loop needs a GPU, Autograd is the wrong tool, and the README does not claim otherwise.

There is a second constraint that is easy to underestimate. Because differentiation works by tracing Python execution, any operation that leaves the wrapped NumPy world is invisible to the trace. A call into a library that accepts plain NumPy arrays and returns plain NumPy arrays will not produce a gradient path through itself unless that library is written against autograd.numpy. The README's own example list is instructive here: it links out to Sampyl and to an external Neural Turing Machine implementation rather than shipping those as part of the package.

The project's own classifier is also worth reading literally. pyproject.toml lists "Development Status :: 4 - Beta". The maintainers have not labeled this production-stable, and the README points readers to the tutorial and examples for depth rather than to a formal API stability guarantee. Treat the interface as stable in practice but not contractually frozen.

How Autograd differs from PyTorch and JAX

PyTorch and JAX both offer automatic differentiation, so the comparison is about what each one asks of your code. PyTorch builds a dynamic graph from tensor operations and expects your model to be expressed in torch tensors; its autograd machinery is tied to that tensor type and to its device and serialization ecosystem. JAX composes function transformations, including grad, over its own array type and targets accelerators through XLA.

Autograd's difference is that the array type is NumPy. You keep writing np.linspace, np.exp, and ordinary Python loops, and you get a gradient function back. That is a much smaller conceptual step if your code already exists and you do not want to port it. The trade is that you inherit NumPy's execution characteristics, with no XLA compilation and no GPU kernels.

The second difference is higher-order derivatives. Autograd's README makes repeated differentiation a first-class demonstration, showing fourth derivatives of a scalar function through nested elementwise_grad calls. PyTorch supports higher-order gradients through create_graph, but that is a mode you opt into; in Autograd, nesting grad calls is the documented way to work. If your method needs third or fourth derivatives as a matter of course, that ergonomics difference is real.

Maintenance, releases, and what the MIT licence means here

The repository is not archived, and the last push was on 2026-09-08. Releases are infrequent but real: v1.8.0 on 2025-05-03, then v1.9.0 on 2026-06-22 and v1.9.1 on 2026-06-29. The README names current maintainers, and the repository carries a CONTRIBUTING.md, a noxfile.py, a pre-commit configuration, and GitHub Actions workflows for checks, tests, and publishing. That is a maintained project with a small, deliberate release cadence rather than a fast-moving one.

Upgrade cost is low by construction. The runtime dependency is numpy<3, so an upgrade is bounded by NumPy compatibility and the Python floor of 3.10. There are no compiled artifacts to rebuild and no migration tooling to run. The practical risk of upgrading is behavioural rather than mechanical: a new version could change how a particular operation is traced, and the README does not document a rollback procedure or a deprecation policy, so pinning a version in your own requirements is the only rollback mechanism available to you.

The licence is MIT, declared both in pyproject.toml via license = "MIT" and in license.txt at the repository root. MIT is permissive: it allows commercial and closed-source use with attribution and without a copyleft obligation on your own code. This is a description of the licence text, not legal advice; if your organization has specific compliance requirements, have counsel read license.txt rather than relying on a summary.

Editorial conclusion

Adopt Autograd when your model is already written as NumPy plus Python control flow and you need gradients, Hessians, or higher-order derivatives without rewriting into a framework's tensor type. Skip it if you need GPU execution, distributed training, or a large prebuilt layer zoo; the README lists no accelerator support and the dependency set is numpy<3 with optional SciPy. Before committing, verify the Python floor of 3.10, check that your NumPy version satisfies numpy<3, and run the tanh gradient example against finite differences to confirm the derivative matches on your install.

Frequently asked questions

What is autograd in machine learning?

It is the general technique of computing derivatives of code automatically, and in this project's case the README describes a library that differentiates native Python and NumPy code. Autograd supports reverse-mode differentiation, which the README equates with backpropagation, and names gradient-based optimization as its main intended application.

How to install autograd?

The README gives pip install autograd. Some features need SciPy, which the README offers as an optional dependency with pip install "autograd[scipy]". The package requires Python 3.10 or later and depends on numpy<3.

How to use autograd?

Import autograd.numpy as np instead of plain numpy, define your function against that module, then pass the function to grad from autograd to get a gradient function. The README's tanh example shows grad_tanh = grad(tanh) followed by grad_tanh(1.0).

What is an autograd engine?

In this project, the differentiating machinery sits behind the wrapped NumPy module and the grad function. The README describes autograd.numpy as a thinly-wrapped NumPy and grad as the only autograd function you may ever need, with reverse-mode and forward-mode differentiation that can be composed arbitrarily.

What is an autograd?

The README describes Autograd as a library that automatically differentiates native Python and NumPy code, handling loops, ifs, recursion and closures, and supporting reverse-mode and forward-mode differentiation. Its main intended application is gradient-based optimization.

What is autograd in PyTorch?

This project is separate from PyTorch. Autograd here operates on NumPy arrays through the wrapped autograd.numpy module rather than on torch tensors, and its only runtime dependency is numpy<3.

Official sources

  1. HIPS/autograd on GitHub
  2. Issues
  3. License: MIT
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/hips-autograd.svg)](https://hysenlabs.com/projects/hips-autograd)