Model or dataset
tenstorrent/tt-metal avatar
tenstorrent/tt-metal

tt-metal: TT-NN Operators and TT-Metalium Kernels for Tenstorrent Hardware

TT-NN operator library, and TT-Metalium low level kernel programming model.

1,687 stars705 forksC++Apache-2.0

At a glance

What is it?
tt-metal is Tenstorrent's C++ and Python stack for its AI accelerators, split between the TT-NN operator library and the TT-Metalium kernel programming model. It is hardware-bound work, and the repository's own performance tables show how much depends on which board you own.
Who is it for?
Adopt tt-metal if you already own Wormhole or Blackhole hardware and want to run or port models against TT-NN, or if you intend to write kernels against TT-Metalium. Do not adopt it as a portable ML framework; there is no CPU or GPU path here, and the README's own numbers show Wormhole and Blackhole performing very differently on the same model.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 5 days ago.
What is it written in?
Mainly C++, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What tt-metal solves, and who it is actually for

tt-metal is the software layer between a Tenstorrent accelerator and the code you want to run on it. The repository ships two things that share one build: TT-NN, described in the README as a Python and C++ neural network operator library, and TT-Metalium, which the README calls the low-level programming model for kernel development on Tenstorrent hardware. The audience is narrow and specific. If you write PyTorch and expect a device-agnostic backend, this is not your entry point. If you have a Wormhole or Blackhole board and need either a set of operators to build a model from, or direct control over what runs on the compute cores, both halves are here in one repository.

The split matters because the two halves have different contracts. TT-NN is an operator surface: you compose ops and the library handles placement and execution. TT-Metalium is the layer underneath, where the README points to a programming guide and a set of simple kernel examples. Choosing between them is the first decision a new user makes, and the repository does not hide that. The README links the TT-NN API reference and the TT-Metalium programming guide as separate documents, and the tech reports are filed under the two names rather than merged.

How TT-NN and TT-Metalium divide the work

The repository layout reflects the split directly. The top level contains ttnn/, tt_metal/ and tt_stl/ as separate trees, with models/, tests/ and tech_reports/ alongside them. The Python package is named ttnn and is built from the same source tree, which is why the build produces both a C++ library and an importable Python module rather than two independent projects.

TT-NN's declared dependencies in pyproject.toml are ordinary Python data and graph tooling: numpy pinned below 2, loguru, networkx, graphviz, pyyaml, click, pandas, seaborn and ml_dtypes, on Python 3.10 or newer. networkx and graphviz are the interesting entries. They point at graph construction and visualisation as first-class concerns of the operator library rather than afterthoughts, which fits a model where ops are assembled into a graph before execution. The package also declares a console script, tt-run, mapped to ttnn.distributed.ttrun:main, so multi-device execution has a command-line entry point rather than being Python-only.

On the Metalium side, the README routes readers to the programming guide and to simple kernel examples in the published documentation. The tech reports listed there cover the matrix engine, data formats and reconfiguring data formats, all dated in 2024. Those are the pages that explain what the hardware does with your data, and their age is worth noticing: the repository's last push was on 2026-08-29, so the reports predate the current code by a wide margin.

Installing tt-metal and running a first model

The README does not carry install commands. It links to INSTALLING.md at the repository root, and the top-level tree also contains install_dependencies.sh, create_venv.sh and build_metal.sh, which is where the actual setup steps live. Read INSTALLING.md before running anything; the scripts assume a specific order and a Linux host, and setup.py contains Linux-specific logic for choosing between lib and lib64.

The Python package itself is named ttnn, and the project metadata declares Python 3.10 or newer. The repository defines the package under the project name ttnn in pyproject.toml:

toml
[project]
name = "ttnn"
requires-python = ">=3.10"

For a source build, the repository ships the scripts referenced above. They sit at the repository root:

bash
./create_venv.sh
./install_dependencies.sh

The README points model work at models/demos/ and the full list at the Model Matrix in models/README.md. That matrix is the first thing to check, because it tells you which models have demos and which hardware they were tuned for. The README's featured table lists Llama 3.3 70B, Qwen 2.5 7B and 72B, Mixtral 8x7B and Whisper, each with a specific board and parallelisation factor, so a first real run means picking one of those pairings rather than assuming any model runs on any board. The declared console script is tt-run, mapped to ttnn.distributed.ttrun:main, so distributed execution is reached through that entry point rather than through a Python-only path:

toml
[project.scripts]
tt-run = "ttnn.distributed.ttrun:main"

What you should see after a successful build is an importable ttnn module and, for the demos, output from the model scripts under models/. If the import fails, the problem is almost always the native library build rather than the Python layer.

The hardware lock-in is the point, not a side effect

tt-metal has no CPU fallback and no GPU backend. Every operator and every kernel in this repository targets Tenstorrent silicon. That is not a limitation the maintainers are working around; it is the reason the project exists, and it means the usual escape hatch of running on a laptop to debug does not apply. If you do not have a board, you cannot meaningfully evaluate this code.

The README's own performance tables make the second constraint plain. Whisper distil-large-v3 at batch size 1 reports 163 ms TTFT on n150 (Wormhole) and 63 ms on p150 (Blackhole), with throughput of 105.0 and 263.4 tokens per second per user respectively. Same model, same repository, a factor of roughly two and a half between boards. The README also states that Blackhole software optimization is under active development, which is a candid way of saying the newer hardware's numbers are still moving. Anyone budgeting on the basis of a published figure should check which board that figure came from.

The performance tables carry a further caveat in the README's own notes: metrics were collected using the tt-metal model demos, and results may vary when using other runtimes such as the vLLM inference server. Each row also pins a specific TT-Metalium release tag and a specific vLLM commit, which tells you the numbers are tied to a version pair rather than to the project in general.

Where TT-NN stops and vLLM takes over

If your goal is serving an LLM behind an HTTP endpoint, tt-metal is the wrong layer to start at. The repository's own material points elsewhere: the README directs readers to the Tenstorrent vLLM plugin repository for vLLM installation and environment creation, and the performance tables list a vLLM Tenstorrent repo release alongside each TT-Metalium release. The two are versioned together.

The difference in approach is straightforward. TT-NN gives you operators and a graph you assemble yourself, which is what you want when a model is not already supported or when you need to control placement and execution. vLLM gives you a serving stack with its own scheduling and batching, built on top of the same device code. The README's note that results may differ between the model demos and the vLLM server is the honest version of this: the same hardware, two runtimes, two sets of numbers. Pick based on whether you are building a model or operating one.

A second boundary is worth naming. TT-Metalium is for people writing kernels, not for people consuming them. The README describes it as enabling kernel development and points to a programming guide and examples rather than to an operator catalogue. If you never intend to write a kernel, TT-NN is the half of this repository you care about.

Maintenance, versioning and licence

The repository is not archived, and its last push was on 2026-08-29. Releases follow a clear pattern visible in the tags: v0.78.0-dev20260829 and v0.77.0-rc3 both landed on 2026-08-29, with v0.78.0-dev20260828 the day before. Development builds carry a date suffix and release candidates carry an rc number, which means you can pin to a dated dev tag or to an rc, but you should expect to move. The README's performance tables pin old tags such as v0.65.0-rc7 and v0.62.0-rc35 next to current model code, so a published number and the current main branch are not the same thing.

Versioning is driven by setuptools-scm, with a tag regex that accepts v1.2.3 and v1.2.3-rc1 forms and a fallback of 0.0.0.dev0. The tag discovery command explicitly excludes tags matching *-dev*, so the dated dev tags do not feed version calculation. That is a deliberate separation between nightly-style builds and versioned releases, and it is the mechanism behind the two tag families you see in the release list.

The licence is Apache-2.0, declared in pyproject.toml and present as a LICENSE file at the root. The repository also carries a NOTICE file and a file named LICENSE_understanding.txt, which is unusual enough to be worth reading if you plan to redistribute. Apache-2.0 includes an explicit patent grant and requires attribution and notice retention. Whether that fits your distribution model is a question for your own legal review, not something the repository answers.

Editorial conclusion

Adopt tt-metal if you already own Wormhole or Blackhole hardware and want to run or port models against TT-NN, or if you intend to write kernels against TT-Metalium. Do not adopt it as a portable ML framework; there is no CPU or GPU path here, and the README's own numbers show Wormhole and Blackhole performing very differently on the same model. Before committing, read INSTALLING.md, open the Model Matrix under models/README.md to confirm your target model is listed, and check the tech report dates on the pages you plan to rely on, since several were last updated in 2024.

Frequently asked questions

What is tt-metal?

tt-metal is Tenstorrent's repository containing TT-NN, a Python and C++ neural network operator library, and TT-Metalium, a low-level programming model for writing kernels on Tenstorrent hardware. The Python package is named ttnn and requires Python 3.10 or newer.

Which Tenstorrent hardware does tt-metal support?

The README's featured model tables reference Wormhole boards (n150, n300, QuietBox, Galaxy) and Blackhole boards (p150). The README notes that Blackhole software optimization is under active development.

How do I install tt-metal?

The README links to INSTALLING.md at the repository root rather than listing steps inline. The top-level tree also contains install_dependencies.sh, create_venv.sh and build_metal.sh, and the Python package is named ttnn.

Which models does tt-metal ship demos for?

The README's featured list covers Llama 3.3 70B, Qwen 2.5 7B and 72B, Mixtral 8x7B and Whisper distil-large-v3, each tied to specific hardware and parallelisation factors. The README points to the Model Matrix in models/README.md for the full list.

What licence does tt-metal use?

Apache-2.0, declared in pyproject.toml and present as a LICENSE file at the repository root. The repository also contains a NOTICE file and a file named LICENSE_understanding.txt.

Official sources

  1. Official documentation
  2. Official README
  3. Project repository
  4. Release notes
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/tenstorrent-tt-metal.svg)](https://hysenlabs.com/projects/tenstorrent-tt-metal)