Library / SDK
tinygrad/tinygrad avatar
tinygrad/tinygrad

tinygrad: A Hackable End-to-End Deep Learning Stack

You like pytorch? You like micrograd? You love tinygrad. Neural networks As it turns out, 90% of what you need for neural networks are a decent autograd/tensor library.

33,658 stars4,348 forksPythonMIT

At a glance

What is it?
tinygrad is an MIT-licensed Python deep learning stack that combines a PyTorch-style tensor API with a visible IR-based compiler, kernel fusion through lazy evaluation, and JIT graph execution. It targets developers who want to read, understand, and modify the full stack from front-end to hardware.
Who is it for?
tinygrad suits developers who need to work with the full compiler pipeline: researchers writing new accelerator backends, engineers studying how deep learning frameworks compile to hardware, or teams that want a minimal dependency footprint for an embedded or edge deployment.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 4 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What tinygrad Is and Who It Is For

tinygrad positions itself between PyTorch and karpathy/micrograd. PyTorch is large and production-grade; micrograd is a minimal autograd toy. tinygrad occupies the space where the full stack, including tensor operations, a compiler IR, scheduling, kernel fusion, and codegen, fits in a codebase small enough to read in a day.

The README describes tinygrad as an end-to-end deep learning stack with four layers: a tensor library with autograd, an IR and compiler that fuse and lower kernels, JIT plus graph execution, and standard nn/optim/datasets for real training. It is maintained by tiny corp.

The audience is developers who want to understand what happens between a Python matrix multiply and a GPU kernel, engineers building new accelerator backends, and researchers who need a training framework where the entire compilation path is modifiable without patching a multi-million-line codebase. The examples/ directory includes complete model implementations for GPT-2, LLaMA, LLaMA 3, Mamba, Mixtral, and a minimal diffusion model, all written against tinygrad's own API. An examples/openpilot/ directory is also present, reflecting real-world use of tinygrad in the openpilot autonomous driving project.

The shipped package includes a tinygrad.viz module with a browser-based viewer (index.html plus assets and JavaScript) and a tinygrad.llm module with a chat.html interface, as listed in pyproject.toml. These are part of the installed package, not separate tools.

Lazy Evaluation and Kernel Fusion

tinygrad uses lazy evaluation to fuse operations. Operations are not executed when called; instead, they build a computation graph that is realized later. This allows the compiler to merge multiple operations into a single kernel rather than launching one kernel per op.

The README demonstrates this with a matmul using DEBUG=3:

sh
DEBUG=3 python3 -c "from tinygrad import Tensor;
N = 1024; a, b = Tensor.empty(N, N), Tensor.empty(N, N);
(a.reshape(N, 1, N) * b.T.reshape(1, N, N)).sum(axis=2).realize()"

This shows the reshape, broadcast multiply, and reduction fused into one kernel. Setting DEBUG=4 prints the generated kernel code for the target device. The fusion is not a post-hoc optimization applied to a traced graph; it falls out of how the lazy IR represents operations.

TinyJit captures and replays kernels at the function level, similar in concept to JAX's jit, so a training loop can be compiled once and replayed without re-tracing. The tinygrad/codegen/ directory contains the lowering passes, organized as tinygrad/codegen/decomp/ for decomposition, tinygrad/codegen/opt/ for optimization, and tinygrad/codegen/late/ for late-stage passes, all readable as ordinary Python. The tinygrad/uop/ module holds the universal operation representation, and tinygrad/schedule/ manages the execution ordering of the realized graph. The tinygrad/renderer/ module contains per-target code generators, including renderers for AMD ISA and general ISA targets.

Installing tinygrad and Running a First Model

tinygrad requires Python 3.11 or later and has no runtime dependencies listed in pyproject.toml. The recommended install is from source:

sh
git clone https://github.com/tinygrad/tinygrad.git
cd tinygrad
python3 -m pip install -e .

For a direct install from master without cloning:

sh
python3 -m pip install git+https://github.com/tinygrad/tinygrad.git

A minimal neural network in tinygrad looks like this example from the README:

python
from tinygrad import Tensor, nn, Context

class LinearNet:
  def __init__(self):
    self.l1 = Tensor.kaiming_uniform(784, 128)
    self.l2 = Tensor.kaiming_uniform(128, 10)
  def __call__(self, x:Tensor) -> Tensor:
    return x.flatten(1).dot(self.l1).relu().dot(self.l2)

model = LinearNet()
optim = nn.optim.Adam([model.l1, model.l2], lr=0.001)

The full MNIST example in examples/beautiful_mnist.py reaches 98% accuracy in about five seconds, according to the README. The tensor API is close enough to PyTorch that training loops transfer with minor changes.

The autograd behavior mirrors PyTorch's interface. The README gives a direct comparison showing that x.grad.tolist() produces the same result with both frameworks for a given matmul backward pass.

Accelerator Support and the Backend Contract

tinygrad supports a range of accelerators: OpenCL, CPU, METAL, CUDA, AMD, NV, QCOM, and WEBGPU, each implemented in the tinygrad/runtime/ directory. To check which accelerator is selected by default on a given machine:

python
from tinygrad import Device; print(Device.DEFAULT)

Adding a new accelerator requires implementing about 25 low-level ops, according to the README. This is the primary selling point for hardware teams: the abstraction surface is small enough that a new backend can be written without understanding the entire codebase. The existing backends serve as reference implementations.

The repository layout makes the boundary clear. tinygrad/runtime/ holds one file per accelerator: ops_cl.py for OpenCL, ops_metal.py for Metal, ops_cuda.py for CUDA, ops_amd.py and ops_nv.py for AMD and NVIDIA GPUs respectively, and ops_webgpu.py for WebGPU. A hardware team adding support for a new chip implements the same ~25-op interface against that chip's native compute API.

The pyproject.toml lists torch==2.9.1 as a testing dependency, not a runtime dependency, so a production install does not pull in PyTorch. For AMD and NVIDIA GPUs, tinygrad generates auto-generated hardware bindings stored in tinygrad/runtime/autogen/. These include bindings for AMD RDNA3, RDNA4, and CDNA architectures, plus NVIDIA register definitions, which allows the framework to talk to the hardware at a lower level than typical compute APIs.

The NV and AMD backends bypass vendor-supplied compute runtimes in favor of direct hardware access via these generated bindings. This is a deliberate design choice: removing the vendor runtime layer reduces code that tinygrad developers cannot modify.

How tinygrad Compares to PyTorch

The README makes the comparison explicit. tinygrad and PyTorch share an eager Tensor API, autograd, and an optimizer library. A familiar PyTorch training loop transfers with minimal changes.

The difference is what is visible. In PyTorch, the compiler and IR are internal implementation details; the public API does not expose them. In tinygrad, the entire compiler and IR are part of the codebase intended to be read and modified. This makes tinygrad useful for understanding how a framework compiles to hardware.

The README also compares tinygrad to JAX, noting shared ideas like IR-based autodiff and function-level JIT, and to TVM, noting that tinygrad ships the full front-end framework alongside the compiler, which TVM does not. The specific gap relative to JAX is that tinygrad does not yet have full vmap or pmap support.

For production workloads that depend on PyTorch's ecosystem of pre-built models, ONNX export, and third-party integrations, tinygrad is not a substitute. Its value is in the compiler layer, not the model zoo.

Contribution Rules, Bounties, and Real Limitations

tinygrad's contribution policy is unusually strict and worth reading before planning to rely on a missing feature. The README lists several automatic rejection criteria: code that looks AI-generated, changes that resemble code golf, documentation or whitespace changes from non-established contributors, performance claims without benchmarks, and diffs that are too large.

The project offers cash bounties for certain improvements, listed in an external spreadsheet linked from the README. This is the primary mechanism for incentivizing work on specific gaps.

The practical limitation is API stability. tinygrad does not guarantee a stable public API across versions. Code written against 0.14.0 may not run against 0.15.0 without changes. Teams that need long-term reproducibility should pin the version and budget for migration work between releases. The project is not at 1.0 yet, and the README says so directly.

A second limitation: the no-AI-written-code rule applies to contributions, not to using tinygrad in your own projects. But it signals that the maintainers prioritize human-readable, carefully written code over fast feature accumulation, which affects the pace of new capability additions.

Editorial conclusion

tinygrad suits developers who need to work with the full compiler pipeline: researchers writing new accelerator backends, engineers studying how deep learning frameworks compile to hardware, or teams that want a minimal dependency footprint for an embedded or edge deployment. It is the wrong choice for production workloads where ecosystem breadth matters: no stable release API, a strict no-AI-written-code contribution policy, and a codebase that the maintainers deliberately keep small mean that missing features stay missing until someone writes clean, benchmarked code to add them. Verify Python 3.11 compatibility and your hardware accelerator against the ops_*.py files in tinygrad/runtime/ before committing.

Frequently asked questions

how to install tinygrad

Clone the repository and install with pip install -e . for a development install, or use pip install git+https://github.com/tinygrad/tinygrad.git for a direct install from master. Python 3.11 or later is required.

How does tinygrad compare to PyTorch?

Both share an eager Tensor API, autograd, and an optimizer library. The difference is that tinygrad exposes its entire IR and compiler as readable, modifiable code, while PyTorch's compilation internals are not intended to be modified. tinygrad is for developers who want to understand or change the full stack; PyTorch is better for production workloads with ecosystem dependencies.

What does tinygrad do?

tinygrad is an end-to-end deep learning stack with a tensor library and autograd, an IR-based compiler that fuses and lowers kernels, JIT graph execution, and standard nn, optim, and dataset utilities. It supports multiple hardware accelerators including CUDA, AMD, Metal, and WebGPU, each implemented as about 25 low-level ops.

Official sources

  1. Official README
  2. Project repository
  3. Release notes
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/tinygrad-tinygrad.svg)](https://hysenlabs.com/projects/tinygrad-tinygrad)