# PyCuTe: a pure-Python reference implementation of the CuTe layout algebra

> PyCuTe reimplements CuTe's layout algebra in plain Python so you can learn the algebra, prototype transformations and generate test vectors without a GPU. The core package has no third-party dependencies, but the layout algebra is all it gives you.

**NVlabs/CuTe** — Reference implementation and examples of the CuTe Layout representation and algebra.

- Repository: https://github.com/NVlabs/CuTe
- Stars: 359 · Forks: 40
- Language: Python
- License: Apache-2.0
- Published: 2026-09-17 · Updated: 2026-09-17 · Language: en
- Canonical page: https://hysenlabs.com/projects/nvlabs-cute

## What PyCuTe solves, and who it is written for

CuTe is the layout and tensor algebra underneath CUTLASS 3.x and the CuTe DSL. In C++ it is a header-only template library tied to CUDA, so reading a definition means reading template code, and trying an expression means compiling it. PyCuTe removes both steps. The README describes it as "plain Python you can `import` from any script" and names three jobs for it: learn the algebra, prototype new transformations, and generate test vectors for the C++ and DSL implementations.

The audience follows from that. If you are writing a kernel in CUTLASS or the CuTe DSL and want to know what `logical_divide` does to a particular shape before you write it, PyCuTe lets you evaluate the expression and print the result. If you maintain the C++ implementation and want a second implementation to compare against, the README positions PyCuTe as exactly that: a reference. It is not a runtime for shipping kernels, and nothing in the README suggests it is meant to be one. The package itself is labelled "Development Status :: 4 - Beta" in pyproject.toml, which is consistent with a reference implementation rather than a production library.

## Layout as a function from coordinates to offsets

The README compresses the whole model into one sentence: a Layout is a function from coordinates to offsets, defined by a Shape and a Stride "of the same hierarchical profile". A stride turns a shape into row-major, column-major or arbitrarily nested addressing, and the README's own table gives three cases: `Layout((4, 8), (8, 1))` for 4x8 row-major, `Layout((4, 8), (1, 4))` for column-major, and `Layout(((2, 4), 8), ((1, 16), 2))` for nested modes.

The algebra is a set of pure functions that take layouts and return layouts. The README lists `coalesce`, `composition`, `complement`, `logical_divide`, `logical_product`, `right_inverse`, `left_inverse`, `nullspace`, `recast`, `layout_add` and `greatest_common_domain`, and says they operate over integer and coordinate strides, with limited support for F2 (XOR-swizzle) strides. Above the layout sits a thin Tensor and Accessor layer: evaluating a tensor at a coordinate evaluates the layout to an offset and dereferences the accessor there. That is the whole data flow. There is no scheduler, no memory allocator and no device code in the picture.

The authority for definitions is not the code. The README states that the CuTe Whitepaper (Cris Cecka, arXiv:2603.02298) "is the authoritative source for every definition and post-condition; PyCuTe defers to it throughout". If you find a discrepancy, that sentence tells you which side the project considers correct.

## Installing PyCuTe in a virtual environment

PyCuTe is installed in place from a source checkout. The README notes that many systems mark the system Python as externally managed under PEP 668, so it recommends a virtual environment first. Python 3.10 or newer is the only hard requirement.

```bash
python3 -m venv .venv
source .venv/bin/activate

pip install -e .              # core layout algebra only (no third-party deps)
pip install -e ".[viz]"       # + visualization helpers (svgwrite, tabulate)
pip install -e ".[test]"      # + everything needed to run the test suite
```

The bare `pip install -e .` gives you the algebra and nothing else; pyproject.toml declares an empty `dependencies` list for the core. The extras are separate: `viz` adds `svgwrite` and `tabulate`, `symbolic` adds `sympy`, and `test` adds `pytest>=7` along with the others. The README also notes that `import pycute` works straight from a checkout as long as the repository is on `PYTHONPATH`, so you can skip the install entirely.

A first real use is to build a layout, evaluate it, and simplify it. The README's quick start shows a 3x4 row-major layout indexed both by coordinate and by a flat offset, then three algebra calls whose printed results are given in the README:

```python
>>> from pycute import *
>>> A = Layout((3, 4), (4, 1))   # 3x4 row-major matrix
>>> A(2, 3)
11
>>> A(11)
11
>>> size(A), rank(A)
(12, 2)
>>> coalesce(Layout((2, (1, 6)), (1, (6, 2))))
Layout(12, 1)
>>> composition(Layout(12), Layout((4, 3)))
Layout((4, 3), (1, 4))
>>> logical_divide(Layout(24), Layout(4, 2))
Layout((4, (2, 3)), (2, (1, 8)))
```

If you want data rather than offsets, `make_tensor(Layout((4, 4), (4, 1)))` builds a 4x4 row-major tensor whose elements you read and write with `T[1, 2]`. To see a layout rather than compute with it, `print_tensor(Layout((4, 8), (1, 4)))` prints an ASCII table of offsets and `draw_svg` writes an SVG file, which the README says is saved as `layout.svg`.

## Where PyCuTe stops being the right tool

The clearest boundary is speed and scope. Nothing in the README presents PyCuTe as a layout engine you would run inside a training loop or a kernel launch path. It is a reference implementation: the README's framing is learning, prototyping and generating test vectors for the C++ and DSL implementations. If you need the layout evaluated on device as part of a kernel, the C++ CuTe or the CuTe DSL is the layer that does that work, and PyCuTe is the thing you consult while writing it.

The second boundary is stride support. The README says the algebra works over integer and coordinate (`ArithTuple`/basis) strides "plus limited support for `F2` (XOR-swizzle) strides". Limited is the operative word, and it is stated rather than implied. If your layout depends on swizzled addressing, check `docs/06_swizzle.md` and the API reference before assuming the operation you need is implemented.

The third is verification. The README points at a test suite and at `docs/08_api_reference.md`, which it describes as an "Index of every function, class, and unit test". That index is the honest way to check coverage, because a reference implementation is only as useful as the operations it actually carries. There are no releases listed for the repository, so there is no versioned changelog to consult; the last push was on 2026-09-11, and the package version in pyproject.toml is 0.1.0.

## PyCuTe against the C++ CuTe and the CuTe DSL

The obvious alternative is CuTe itself. The difference is not features, it is where the code runs and what it costs to try an idea. C++ CuTe is header-only and coupled to CUDA: you get the real implementation with device execution, but expressing a layout means writing template code and compiling it. The CuTe DSL is the Python-facing path to the same machinery, and it still assumes the GPU stack. PyCuTe is pure Python with no GPU requirement, which is why the README calls it the place to learn the algebra and generate test vectors. The trade is that you get the algebra and a thin Tensor/Accessor data model, not the kernel-facing surface.

Within Python, the alternative to building on PyCuTe is to reimplement the few layout operations you need by hand, usually as index arithmetic over shape and stride tuples. That works until composition, complement or logical_divide enters the picture, at which point you are re-deriving definitions that the Whitepaper already fixes and that PyCuTe already implements. The README's list of operations is the honest measure of what you would otherwise be writing yourself.

One more comparison is worth naming because the search data mentions it: Cutile-rs, a Rust implementation of CuTe. It shares the goal of a standalone implementation of the algebra outside the C++ headers, but it is a different language and a different ecosystem. If your surrounding code is Python, PyCuTe is the one you can import directly.

## Licence, maintenance and what an upgrade costs

PyCuTe is Apache-2.0, stated in LICENSE.txt, in the README badge and in the `license` field of pyproject.toml. That is a permissive licence with an explicit patent grant, and it is the same licence family as the CUTLASS project the README points to. If you are embedding PyCuTe in a larger system, the usual Apache-2.0 obligations apply: keep the licence and notices. This is a description of the terms, not legal advice, and the terms themselves are in LICENSE.txt.

Upgrade cost is low by construction. The core has no third-party runtime dependencies, so `pip install -e .` cannot pull in a conflicting version of anything, and the extras are opt-in. The package version is 0.1.0 and the repository carries no releases, so there is no published changelog to diff across; updates arrive as commits. The last push was on 2026-09-11, and the repository is not archived.

Because the Whitepaper is named as the authority for definitions and post-conditions, the semantics of the algebra should not drift with the code. What can change is the API surface and the set of implemented operations. That makes the API reference in `docs/08_api_reference.md` the file to re-read after pulling, rather than a version number.

## Conclusion

Adopt PyCuTe if you need to understand, prototype or test CuTe layouts on a machine without a GPU: the core algebra installs with pip install -e . and pulls in no third-party packages. Do not adopt it as a fast layout engine or as a substitute for the CuTe DSL on real hardware; it is a reference implementation and its own README frames it as the place to learn, prototype and generate test vectors. Before committing, check docs/08_api_reference.md for the functions you actually need, and confirm whether your strides are integer or F2, because the README describes F2 support as limited.

## FAQ

### Do I need a GPU or CUDA installed to use PyCuTe?

No. The README states that PyCuTe is a pure-Python reference implementation and that no GPU is required. Python 3.10 or newer is the only hard requirement, and the core layout algebra has no third-party dependencies.

### How do I install PyCuTe?

It is installed in place from a source checkout. The README recommends a virtual environment because many systems mark the system Python as externally managed under PEP 668, then `pip install -e .` for the core algebra, with `.[viz]` or `.[test]` for the optional extras. It can also be imported straight from a checkout if the repository is on PYTHONPATH.

### What is the difference between PyCuTe and the C++ CuTe in CUTLASS?

The README describes C++ CuTe as a header-only template library tightly coupled to CUDA, while PyCuTe is plain Python you can import from any script. PyCuTe is positioned for learning the algebra, prototyping transformations and generating test vectors for the C++ and DSL implementations, and it defers to the CuTe Whitepaper for every definition and post-condition.

## Sources

- [Issues](https://github.com/NVlabs/CuTe/issues)
- [License: Apache-2.0](https://github.com/NVlabs/CuTe/blob/main/LICENSE)
- [NVlabs/CuTe on GitHub](https://github.com/NVlabs/CuTe)
- [README](https://github.com/NVlabs/CuTe/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/nvlabs-cute
