micrograd: a scalar autograd engine you can read in one sitting
A tiny scalar-valued autograd engine and a neural net library on top of it with PyTorch-like API
At a glance
- What is it?
- micrograd is a roughly 100-line reverse-mode autodiff engine plus a 50-line neural network library with a PyTorch-like API. It is a teaching artifact, and its limits follow directly from that.
- Who is it for?
- Adopt micrograd if your goal is to understand or teach backpropagation, or to check a gradient implementation against something small enough to read end to end. Do not adopt it for real training workloads: every operation is scalar, so a 16-node hidden layer becomes thousands of Value objects, and the README itself frames the project as useful for educational purposes.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 58 days ago.
- What is it written in?
- Mainly Jupyter Notebook, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem micrograd solves is comprehension, not throughput
Most autodiff systems are optimized for speed, which means their core loops are written in C++ or CUDA and their graphs are built from tensors. Reading one top to bottom is not a realistic afternoon project. micrograd inverts that priority. The README states that the engine and the neural network library are about 100 and 50 lines of code respectively, and that the DAG operates only over scalar values. That single design decision is the whole point: because every node is one number, the backward pass is a small chain of local derivative rules you can trace by hand. The audience is anyone who has used PyTorch and wants to know what backward() is actually doing, plus instructors who need a codebase small enough for students to modify. The README is explicit that it is potentially useful for educational purposes, and that framing should be taken literally.
How the DAG, Value and backward pass fit together
Everything starts from the Value class in micrograd.engine. A Value wraps a Python float in .data, and each operation on it returns a new Value that remembers the inputs it came from, so a dynamically built directed acyclic graph forms as you write ordinary arithmetic. The README's example chains addition, multiplication, exponentiation, division, in-place +=, and relu, which shows that the operators are overloaded rather than exposed as explicit graph-building calls. Calling backward() on the final node walks that graph in reverse and accumulates into each node's .grad. The README gives the concrete numbers: a forward pass over its example expression prints 24.7041, and after g.backward() the gradients print 138.8338 for a and 645.5773 for b. Those printed values are the useful part for a learner, because you can hand-check them against the chain rule. On top of the engine sits micrograd.nn, which provides Neuron, Layer and MLP objects with a PyTorch-like calling convention. The README notes that the engine stores per-op backward closures, and points to microgpt as a variant that instead stores local gradients at forward time. That is a real architectural difference worth knowing about, because it changes where the memory goes.
Installing micrograd and running a first forward and backward pass
The README gives a single install command, and the package targets Python 3.6 or newer according to setup.py. Install it with pip, then build a small expression and inspect both the forward value and the gradients. The numbers below are the ones the README documents, so if you see them you know the engine is behaving as described. Note that the import path is micrograd.engine for Value and micrograd for nn. The install command is exactly the one the README lists:
pip install microgradWith that done, here is the README's example expression, which chains addition, multiplication, exponentiation, division, in-place +=, and relu before calling backward:
from micrograd.engine import Value
a = Value(-4.0)
b = Value(2.0)
c = a + b
d = a * b + b**3
c += c + 1
c += 1 + c + (-a)
d += d * 2 + (b + a).relu()
d += 3 * d + (b - a).relu()
e = c - d
f = e**2
g = f / 2.0
g += 10.0 / f
print(f'{g.data:.4f}') # prints 24.7041
g.backward()
print(f'{a.grad:.4f}') # prints 138.8338
print(f'{b.grad:.4f}') # prints 645.5773For a full training loop, the README points to demo.ipynb, which trains a two-layer MLP binary classifier on the moon dataset. The recipe there is a max-margin SVM-style loss plus SGD, and the README states that two hidden layers of 16 nodes each produce the decision boundary shown in moon_mlp.png. To verify gradients independently, install PyTorch and run the test suite, since the README says the tests use PyTorch as a reference for verifying the correctness of the calculated gradients:
python -m pytestVisualizing the graph with trace_graph.ipynb and draw_dot
The second notebook, trace_graph.ipynb, produces Graphviz renderings of the computation graph. The README's example builds a two-input Neuron, feeds it a list of two Values, and calls draw_dot on the output:
from micrograd import nn
n = nn.Neuron(2)
x = [Value(1.0), Value(-2.0)]
y = n(x)
dot = draw_dot(y)Each rendered node shows the data on the left and the gradient on the right, which makes the backward pass inspectable rather than abstract. The repository ships gout.svg as an example of the output. This is the most practical feature for debugging a hand-written derivative rule: if a gradient looks wrong, the picture shows which node introduced it. The caveat is that Graphviz is an external dependency the README does not walk through installing, and draw_dot is demonstrated inside the notebook rather than documented as a public API in the README text.
Where micrograd stops being the right tool
Scalar-only execution is a hard ceiling. A network with two hidden layers of 16 nodes is already thousands of individual Value objects, each carrying its own Python object overhead and its own backward closure. The README acknowledges the cost directly: each neuron is chopped into all of its individual tiny adds and multiplies. For the moon dataset that is fine. For anything approaching real data, the Python interpreter dominates and training becomes impractical. There is also no GPU path, no tensor abstraction, and no broadcasting, because there is nothing to broadcast over. If you need to train something that matters, this is the wrong tool. The README itself redirects readers to microgpt for a more advanced example, describing it as a more efficient and better version of the autograd engine here. That is a useful signal about where the author thinks the interesting work continues.
micrograd compared with PyTorch and Tinygrad
The difference with PyTorch is not scale, it is what the unit of computation is. PyTorch builds graphs of tensors and pushes the heavy arithmetic into optimized kernels, so a single node can represent a whole matrix multiply. micrograd builds graphs of scalars, so the graph structure mirrors the arithmetic literally. That is why micrograd is readable and PyTorch is not, and equally why PyTorch is fast and micrograd is not. The comparison with Tinygrad runs along a different axis. Tinygrad keeps a compact, readable codebase but targets real hardware backends and tensor operations, so it sits between the two: still small enough to study, but aimed at execution rather than exposition. If your goal is to understand the chain rule applied to a graph, micrograd's scalar granularity is a feature. If your goal is to understand how a compiler lowers operations onto a GPU, micrograd has nothing to say.
Maintenance, packaging and the MIT licence
The repository is not archived, and the last push was on 2026-08-03. The version in setup.py is 0.1.0, with python_requires set to >=3.6, and the classifier declares OS Independent. There are no retrieved releases, so pip installs whatever the current source produces rather than a tagged artifact. Practically, that means upgrade risk is low in the sense that the API surface is tiny and stable, and also low in the sense that there is little churn to track. The MIT licence permits use, modification and redistribution with the licence text retained; that is a permissive arrangement, but it is not legal advice and anyone embedding the code in a product should read the LICENSE file themselves. The main cost of adoption is not maintenance, it is the time to read the engine and the two notebooks.
Editorial conclusion
Adopt micrograd if your goal is to understand or teach backpropagation, or to check a gradient implementation against something small enough to read end to end. Do not adopt it for real training workloads: every operation is scalar, so a 16-node hidden layer becomes thousands of Value objects, and the README itself frames the project as useful for educational purposes. Before relying on it, verify first that the gradients you compute match a reference, which is exactly what the test suite does by comparing against PyTorch after installing it.
Frequently asked questions
What is micrograd?
It is a tiny scalar-valued autograd engine implementing backpropagation over a dynamically built DAG, plus a small neural network library with a PyTorch-like API. The README states the two parts are about 100 and 50 lines of code.
What does micrograd do?
It computes forward values and gradients for expressions built from Value objects, and provides Neuron, Layer and MLP classes on top of that. The README frames it as potentially useful for educational purposes.
How does micrograd compare with PyTorch?
micrograd's DAG operates only over scalar values, so each neuron is broken into individual adds and multiplies, whereas PyTorch works on tensors. The README notes that the test suite uses PyTorch as a reference for verifying calculated gradients.
How does micrograd compare with autograd?
The README describes micrograd as implementing backpropagation, which is reverse-mode autodiff, over a dynamically built DAG. It does not draw a comparison with any specific library called autograd, so the precise differences are not documented there.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/karpathy-micrograd)