# TPU-MLIR: compiling PyTorch, ONNX and HuggingFace LLMs into Sophgo bmodel files

> TPU-MLIR is an MLIR-based compiler that turns pre-trained networks into bmodel files for Sophgo TPUs. It ships as a Docker image and a PyPI wheel, and its LLM path is the part most teams will actually care about.

**sophgo/tpu-mlir** — Machine learning compiler based on MLIR for Sophgo TPU.

- Repository: https://github.com/sophgo/tpu-mlir
- Stars: 993 · Forks: 226
- Language: C++
- License: NOASSERTION
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/sophgo-tpu-mlir

## What TPU-MLIR actually converts, and for whom

TPU-MLIR is a compiler, not a runtime and not a training framework. Its job is narrow and stated plainly: it takes a pre-trained neural network and produces a bmodel file that runs on a Sophgo TPU. The repository describes it as an MLIR-based machine-learning compiler for TPUs, and the input side covers PyTorch, ONNX, TFLite and Caffe, with anything else expected to arrive through ONNX first.

The audience is correspondingly narrow. If you are writing CUDA kernels or shipping to a cloud accelerator, this project has nothing for you. It is for engineers who have chosen Sophgo hardware and now need to get an existing model onto it without hand-writing a backend. The chip strings that appear in the LLM conversion examples are bm1684x, bm1688 and cv186ah, which tells you the target surface is a specific family of edge and inference parts rather than a general accelerator abstraction.

The second audience is LLM deployers. The README calls the project LLM-ready and points at llm_convert.py for one-shot conversion of HuggingFace models, naming Qwen and MiniCPM-V as examples. That is a different workflow from the classic vision path, and the documentation treats it as such: a separate quick start, separate flags, and a separate set of trade-offs around KV cache handling.

One more thing about scope. The project ships production tooling alongside the compiler itself, including model_runner, model_tool, accuracy validation, a simulator and a visualizer. Those are not incidental. They are what makes the output verifiable without a physical board in the loop, and they are the reason the compiler is usable as a deployment step rather than as an academic exercise.

## The lowering pipeline: Top and Tpu dialects

The architecture is the standard MLIR shape. Front ends import a model into a high-level dialect, and a sequence of lowering passes rewrites it toward the hardware. The README names two dialects, Top and Tpu, and describes the pipeline as clean dialects plus pattern rewrites plus layer-group memory planning. That last item is the interesting one: memory planning at the layer-group level is where a compiler stops being a translator and starts making scheduling decisions that affect whether a model fits at all.

Quantization is not a bolt-on stage here. The toolchain covers F32, BF16, F16 and INT8 in both symmetric and asymmetric form, and it accepts AWQ, GPTQ and AutoRound checkpoints as passthrough, alongside calibration and QAT. In practice this means the compiler is designed around the assumption that you will hand it a quantized model rather than a float one. The Qwen walkthrough makes that explicit by recommending a pre-quantized AWQ, GPTQ or AutoRound build as the starting point.

The LLM path diverges from the vision path in a way worth understanding before you pick flags. With --use_history_kv disabled, the compiler emits two instruction groups, block_ for prefill and block_cache_ for decode, and the README recommends this for single-turn conversations with short context, giving 4K as the rough boundary. With --use_history_kv enabled, a third group appears, block_kv_, which handles prefill with history. That third group is what makes multi-turn and long-context work viable, and the README recommends it when you are unsure because it is more flexible with good overall performance.

The chunking argument interacts with all of this. --chunk_length only has meaning together with --use_history_kv, and it sets the segment length for chunked inference. The README's own example: with 1K chunks, a 7K input runs prefill as block_ plus seven block_kv_ passes. Decode is segmented by KV-cache length as well, which is why the README notes that performance differs at 1K, 2K, 4K and 8K lengths. If you are benchmarking, you are benchmarking a configuration, not the compiler.

## Installing TPU-MLIR and running a first LLM conversion

The supported path is a prebuilt Docker image rather than a bare host install. Pull it first:

```bash
docker pull sophgo/tpuc_dev:latest
```

If that pull fails, the README gives a tarball fallback and a docker load command for it. Either way you end up with an image that already satisfies the Python requirement, which the README states as Python 3.10 or newer on Ubuntu 22.04. Then create a container with the working directory mounted:

```bash
docker run --privileged --name tpu-mlir -v $PWD:/workspace -it sophgo/tpuc_dev:latest
```

Inside the container you have two options. The quick one is the prebuilt wheel:

```bash
pip install tpu_mlir
```

The README labels the other option, building from source, as recommended. That sequence is clone, install requirements, source the environment script, build:

```bash
cd /workspace/tpu-mlir
pip install -r requirements.txt
source ./envsetup.sh
./build.sh
```

Be aware of what requirements.txt drags in. It pins torch, torchvision and torchaudio to +cpu builds from the PyTorch CPU index specifically to avoid pulling multi-GB CUDA packages, and it keeps a find-links entry for PaddlePaddle's archive because the pinned paddlepaddle version is no longer on PyPI. Budget disk space accordingly, and expect the source build to be the slow part of a first setup.

For a first real run, the README's LLM walkthrough starts by cloning a pre-quantized model:

```bash
git lfs install
git clone https://huggingface.co/Intel/Qwen3.5-2B-int4-AutoRound
```

Then compile it. The two-scenario choice is controlled by --use_history_kv; this is the multi-turn variant, which the README recommends when you are unsure:

```bash
llm_convert.py \
  -m /workspace/Qwen3.5/Qwen3.5-2B-int4-AutoRound \
  -s 8192 \
  --use_history_kv \
  --chunk_length 1024 \
  -c bm1684x \
  -o qwen3.5_4b_history
```

The required arguments are the model path (-m), the maximum sequence length (-s), the chip (-c) and the output directory (-o). Everything else has a default or is optional. If you omit --max_input_length it defaults to the value of -s. The README also lists a --dynamic flag for dynamic-shape compilation, which it recommends for all new conversions and notes that Qwen3.5 forces. The output is a bmodel file in the directory you named, and model_runner is what the README points at for running it.

## Where TPU-MLIR is the wrong tool

The clearest limitation is that the output is not portable. A bmodel runs on Sophgo TPUs. If your deployment plan changes to a different accelerator, the compiled artifact is worthless and you are back at the source model. Nothing in the repository suggests an export path to another backend, and the Tpu dialect exists precisely because the target is fixed.

The second limitation is the Docker-first posture. The README's install section is built around pulling sophgo/tpuc_dev:latest and running it with --privileged. That is a reasonable choice for a compiler that needs device access, but it means you are not dropping this into a locked-down CI runner without thinking about the privilege flag and the image size.

Third, the LLM path is genuinely configuration-sensitive. The README itself frames the choice between the two compilation scenarios as a recommendation rather than a rule, and the reason is that the trade-off is real: the no-history variant produces fewer instruction groups and is aimed at short single-turn context, while the history variant adds block_kv_ and is aimed at multi-turn and longer contexts. Pick wrong and you either lose capability you needed or carry instruction groups you never execute. The chunk length compounds this, since it changes the prefill and decode segmentation and therefore the performance profile at each context length.

Finally, the front-end list is finite. PyTorch, ONNX, TFLite and Caffe are named; the README says other frameworks go through ONNX. That conversion step is yours to own, and it is a common place for operator mismatches to appear before TPU-MLIR is even involved. A model that fails to export cleanly to ONNX never reaches the compiler, and the compiler cannot help you there.

## How this differs from TVM and the vendor SDK route

The obvious comparison is Apache TVM, which also compiles models to multiple accelerator backends. The difference in approach is the IR. TVM builds around its own TensorIR and Relay representations and targets a broad set of backends through a BYOC-style pattern. TPU-MLIR builds on MLIR and defines Top and Tpu dialects that exist to serve one vendor's silicon. That is a narrower bet with a shallower abstraction: you get a pipeline tuned for Sophgo memory planning and quantization rather than a general backend framework you have to teach about your chip.

The other alternative is the vendor's own deployment SDK, which typically expects you to work with the compiled artifact and the runtime rather than the compiler. TPU-MLIR sits upstream of that. Choosing between them is not really a choice about compiler quality; it is about whether you need to compile arbitrary models at all, or whether a fixed set of prebuilt bmodels would do.

If what you actually want is an MLIR project to learn from or extend, TPU-MLIR is a working example of dialects, pattern rewrites and memory planning, but it is not a neutral one. The Tpu dialect encodes Sophgo hardware assumptions, so the reusable lessons are structural rather than directly transferable. Someone searching for a general MLIR tutorial will find a vendor compiler instead.

## Maintenance, releases and licence status

The repository is not archived, and the last push was on 2026-08-31. The most recent release, v1.30.2, is dated 2026-08-31, following v1.28.1 on 2026-04-29 and v1.27 on 2026-02-06. The release cadence visible in that list is roughly every two to three months, with a patch release landing the same day as the latest push.

Upgrade cost is dominated by the pinned dependency set rather than by the compiler API. requirements.txt pins exact versions across a wide surface, including numpy, onnx, protobuf, transformers and the torch trio, and it carries two special index directives to work around packages that are no longer on PyPI as pinned. Any environment that installs from this file inherits those pins, so moving TPU-MLIR versions may mean moving a large dependency graph with it. The Docker image sidesteps most of that, which is presumably why it is the documented default.

The licence is listed as NOASSERTION in the repository metadata, which means the licence could not be automatically classified. The repository does contain a LICENSE file at the top level, so the terms are available to read, but the repository metadata does not state which licence it is. Read that file before you ship anything. The README also carries an arXiv reference, 2210.15016, for readers who want the design rationale in paper form.

## Conclusion

Adopt TPU-MLIR if you are already committed to Sophgo silicon and need to move HuggingFace or ONNX models onto it; the Docker image, the prebuilt wheel and the llm_convert.py path cover the common cases. Do not adopt it as a general-purpose MLIR playground or as a way to target non-Sophgo accelerators, because the lowering pipeline ends in Top and Tpu dialects that only those chips consume. Before committing, verify that your exact chip string (bm1684x, bm1688 or cv186ah) is listed for your model family, and check whether your quantized checkpoint format (AWQ, GPTQ or AutoRound) is one the toolchain accepts as passthrough.

## FAQ

### How do I install TPU-MLIR?

The documented path is the prebuilt Docker image: pull sophgo/tpuc_dev:latest, then run a container with --privileged and your working directory mounted at /workspace. Inside the container you can either run pip install tpu_mlir for the prebuilt wheel or build from source with pip install -r requirements.txt, source ./envsetup.sh and ./build.sh.

### What Python version does TPU-MLIR require?

The README states Python 3.10 or newer on Ubuntu 22.04, and notes that the Docker image already satisfies this requirement.

### Which model formats can TPU-MLIR convert?

The front ends named in the README are PyTorch, ONNX, TFLite and Caffe, with other frameworks expected to go through ONNX first. The output is a bmodel file for Sophgo TPUs.

### What is the --use_history_kv flag in TPU-MLIR?

It controls the LLM compilation scenario. Without it, the compiler emits two instruction groups, block_ for prefill and block_cache_ for decode, which the README recommends for single-turn conversations with short context. With it, a third group block_kv_ is added for prefill with history, which the README recommends for multi-turn conversations, long contexts, or when you are unsure.

### Which chips does TPU-MLIR target?

The chip argument in the LLM conversion examples accepts bm1684x, bm1688 and cv186ah.

## Sources

- [Issues](https://github.com/sophgo/tpu-mlir/issues)
- [README](https://github.com/sophgo/tpu-mlir/blob/master/README.md)
- [Releases](https://github.com/sophgo/tpu-mlir/releases)
- [sophgo/tpu-mlir on GitHub](https://github.com/sophgo/tpu-mlir)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/sophgo-tpu-mlir
