# BBuf/tvm_mlir_learn: a TVM, MLIR and Relay learning archive

> The repository is a set of compiler study notes and runnable examples rather than a library. It is useful if you want to read scheduling, codegen and Relay experiments, and it is the wrong choice if you need something to install and depend on.

**BBuf/tvm_mlir_learn** — compiler learning resources collect.

- Repository: https://github.com/BBuf/tvm_mlir_learn
- Stars: 2,778 · Forks: 371
- Language: Python
- License: not declared
- Published: 2026-09-28 · Updated: 2026-09-28 · Language: en
- Canonical page: https://hysenlabs.com/projects/bbuf-tvm-mlir-learn

## What problem the tvm_mlir_learn repository solves

Deep learning compiler work is hard to learn from papers alone. The gap between a scheduling paper and a working schedule, or between an MLIR dialect description and a pass that runs, is where most beginners stall. This repository is an attempt to fill that gap with concrete files: a scheduler directory of TVM scheduling examples, a relay directory with custom pass experiments and deployment demos, a codegen directory built on tensor expressions and Relay IR, a torchscript directory, an optimize_gemm directory of GEMM experiments, and paper_reading notes on PET, Ansor and MLIR-related work. The README calls the whole thing a public learning archive for AI compiler systems. That framing matters. It is a record of one person learning, not a product with a support contract. If you are the intended reader, a student or an engineer moving from framework work into compiler work, the value is in seeing the pieces side by side in one tree rather than in any single abstraction the project exposes.

## How the examples are organised and how data flows through them

There is no runtime and no framework here, so the architecture is the directory layout. Each top-level directory isolates one stage of the compiler pipeline. The relay directory holds Relay IR examples and custom pass experiments, which is the level where you write transformations over a graph rather than over generated code. The codegen directory takes tensor expressions and Relay IR and produces generated code, so it sits downstream of Relay in the same pipeline. The scheduler directory covers scheduling, the step that decides how a computation is mapped onto hardware after the computation itself is expressed. The optimize_gemm directory is the narrowest slice: one kernel family, GEMM, optimized with compiler ideas, which makes it the most readable entry point if you want to see the whole loop from expression to tuned kernel in a small space. The dataflow_controlflow directory is conceptual rather than executable pipeline code, comparing the two ideas in small examples. At the repository root sit three standalone scripts, pytorch_resnet18_export_onnx.py, tvm_onnx_resnet18_inference.py and tvm_pytorch_resnet18_inference.py, which together describe the shortest path from a PyTorch model to inference through TVM, with ONNX as the intermediate format in one of the two routes. That is the clearest end-to-end story the repository tells.

## Installing nothing: how to run a first TVM example from this archive

The repository is not distributed as a package. There is no setup.py, no pyproject.toml and no release in the repository, so there is nothing to pip install from it. The README points at compile_tvm_in_docker.md for building TVM in Docker, which is the closest thing to installation instructions the project provides, and it is about building TVM, not about installing this repository. The practical route is to clone the repository and run the scripts against an environment where TVM is already available. The README does not pin a TVM version, so the version an example expects is something you have to infer from the code.

```bash
git clone https://github.com/BBuf/tvm_mlir_learn
cd tvm_mlir_learn
```

After cloning, the root-level inference script is the shortest thing to try, because it does not depend on any of the subdirectories:

```bash
python tvm_pytorch_resnet18_inference.py
```

The README does not document the expected output of that script, so what you see depends on your TVM build and on whether the model weights are available locally. The ONNX route is the alternative, and it splits into an export step and an inference step:

```bash
python pytorch_resnet18_export_onnx.py
python tvm_onnx_resnet18_inference.py
```

If you want to build TVM itself first, compile_tvm_in_docker.md is the file to read. Treat the subdirectories as reading material and run their scripts individually; nothing in the repository suggests a single entry point that exercises all of them.

## Where this archive stops being useful

The README states plainly that the repository is a legacy learning archive and that new public-facing documentation will use English entry points elsewhere. The last push was on 2026-05-20. That combination sets expectations: this is a snapshot, not a maintained tool. There is no test suite described and no CI configuration visible in the top-level entries, so an example that no longer runs against a current TVM or MLIR build will not be caught for you. The licence is the sharper problem. The repository gives no licence identifier, and the repository root listing contains no LICENSE file, so if you plan to copy code from here into something you ship, you have no stated terms to rely on and should ask the author directly. The paper_reading notes are prose, not code, and they will age like any reading notes. Finally, the repository is pedagogical by design: it does not attempt to be a library with a stable API, so any attempt to vendor it as a dependency will meet code written to demonstrate an idea rather than to be called from elsewhere.

## TVM versus MLIR, and where this repository sits in that comparison

People searching around this project often arrive at the TVM versus MLIR question, and the repository is a reasonable place to see the two in the same tree. TVM is a full stack aimed at tensor programs, with Relay as its graph-level IR, a scheduler for mapping computations to hardware, and code generation downstream. MLIR is a compiler infrastructure for defining dialects and transformation passes over them, and it is used inside many projects rather than being an end-to-end deep learning compiler on its own. The difference shows up in the directory names here. The relay, scheduler and codegen directories are TVM work, expressed in TVM's own abstractions. The mlir-ods directory and the MLIR-related paper notes are about the infrastructure side, where you define operations and their ODS descriptions rather than tune a kernel. If your goal is to get a model running fast on a specific device, the TVM examples are the relevant ones. If your goal is to build or extend a compiler with custom IR, the MLIR material is the starting point, and you should expect to leave this repository quickly for the upstream MLIR documentation, because the notes here are reading notes, not a course.

## Maintenance, licence and upgrade cost

The last push was on 2026-05-20, and the README describes the repository as a legacy learning archive, so there is no upgrade path to plan around. Nothing in the repository indicates version pinning for TVM or MLIR, and neither dependency is declared in a manifest, which means any breakage from an upstream API change lands on you when you try to run an example. The cost of using the repository is therefore not installation but triage: reading a script, working out which compiler version it assumes, and deciding whether to fix it or read it as pseudocode. On licensing, the repository provides no licence identifier and no LICENSE file appears in the top-level listing, so the terms under which the code can be reused are unstated. That is a question for the author, not something to assume, and it matters most if you intend to copy an example into a commercial codebase. For private study the question is less pressing, but it does not disappear.

## Conclusion

Use this repository if you are studying deep learning compilers and want worked examples of TVM scheduling, Relay passes, codegen and GEMM optimization in one place, and you are comfortable reading code without a documented API surface. Do not adopt it as a dependency, a build system or a source of supported tooling: there is no package, no release, no licence file in the repository root and no test suite described. Before you spend time on it, check the contents of the directory that matches your interest, confirm the TVM or MLIR version the example was written against, and read compile_tvm_in_docker.md if you intend to build TVM yourself.

## FAQ

### What is the purpose of MLIR?

MLIR is compiler infrastructure for defining dialects and transformation passes over them. In this repository it appears through the mlir-ods directory and the MLIR-related paper notes, alongside TVM work that uses Relay, scheduling and codegen instead.

### What is LLVM in simple words?

The repository does not explain LLVM in introductory terms. LLVM appears only in the README's list of topics the repository collects examples around, so the repository is not the place to learn what LLVM is.

### Why is LLVM slow?

The repository does not address LLVM performance. LLVM is listed as one of the topics the repository touches, but no file in it discusses why it might be slow.

### Is LLVM written in C or C++?

The repository does not cover LLVM's implementation language. It is Python-based learning material about TVM, MLIR, Relay and related compiler topics, and it does not describe LLVM's source.

## Sources

- [BBuf/tvm_mlir_learn on GitHub](https://github.com/BBuf/tvm_mlir_learn)
- [Issues](https://github.com/BBuf/tvm_mlir_learn/issues)
- [README](https://github.com/BBuf/tvm_mlir_learn/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/bbuf-tvm-mlir-learn
