LLM-TPU turns HuggingFace weights into bmodels for SOPHGO chips
Run generative AI models in sophgo BM1684X/BM1688
At a glance
- What is it?
- sophgo/LLM-TPU exports HuggingFace weights into the bmodel format that SOPHGO BM1684X and BM1688 parts consume, wires three demos to a single run.sh flag, and documents far more models than one command can reach. Most of its usable detail sits in the per-model directories rather than the top-level entry point.
- Who is it for?
- LLM-TPU fits an engineer who already holds a SOPHGO BM1684X or BM1688 board and wants a generative model running without building a compiler stack first, because the pre-compiled bmodels and the three wired run.sh targets keep that first session short. It does not fit CV186X owners, anyone who needs a C++ demo outside Qwen3.5, Qwen3-VL, and Qwen2.5-VL, or teams that pin deployments to tagged releases.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository last received commits 8 days ago.
- What is it written in?
- Mainly C++, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Converting HuggingFace weights is the step that gates everything else
LLM-TPU does not load a HuggingFace checkpoint at inference time. Its one-click compilation entry point is `llm_convert.py`, which exports HuggingFace weights directly to bmodel, and that export needs a TPU-MLIR environment. TPU-MLIR itself can arrive as a Docker image or as a source build, so the compiler is the first real dependency a new user meets. The alternative is stated plainly in the blockquote under the feature list: use the pre-compiled bmodels provided in each demo, which takes the compiler off the path entirely but also takes away the step where you would build the model yourself. Compilation is therefore an option you opt into rather than a default stage. The chips named for this work are the BM1684X and BM1688 parts in the introduction, and the full model list lives under `models/`.
Only three of the listed models are reachable through run.sh
The top-level entry point is `run.sh`, and it exposes one flag worth knowing, `--model`. The demo table gives it exactly three values. `qwen3` starts Qwen3-4B, `qwen3.5` starts Qwen3.5-2B, and `internvl3` starts InternVL3-2B.
git clone https://github.com/sophgo/LLM-TPU.git
cd LLM-TPU
./run.sh --model qwen3.5Everything else under `models/`, from QwQ-32B to Gemma4 to Qwen3-TTS, is reached by entering its own directory instead of by naming it on the command line. That gap matters when you read the tables and count a dozen families. The tables describe what has been ported to these chips, while `run.sh` describes what has been wired to a one-command path. The other two demo commands are `./run.sh --model qwen3` for Qwen3-4B and `./run.sh --model internvl3` for InternVL3-2B. Three rows out of a much longer table also means the wired models are the ones with the shortest documented path from clone to running model.
CV186X appears in the tagline and then disappears from the tables
The subtitle names three chips, BM1684X, BM1688, and CV186X, and the introduction repeats all three. The tables below it do not. Every row in the multimodal table lists BM1684X or BM1688, with the second one abbreviated to `1688` in the chip column, and every family row in the LLM table carries a single generic compile checkmark instead of a chip list. The repository description is narrower again, naming only BM1684X and BM1688. The third chip in the tagline therefore has no per-model evidence behind it in the visible documentation. If you own a CV186X board, nothing here tells you which model directories target it, and the honest reading is that the support matrix is a BM1684X and BM1688 table with a broader subtitle attached.
One-click Compile and Deployed mark two different states
The multimodal table carries two columns that are easy to conflate. One-click Compile is a yes or no, and a yes means `llm_convert.py` can export that model. The Notes column carries states outside that binary. Qwen-VL, InternVL2, MiniCPM-V-2.6, and Llama3.2-Vision share one row marked Deployed with no compile check, which reads as ports that exist on the chips without a scripted export. Between the poles sit models with no compile path and a working demo. Step3_VL covers image understanding, Falcon-Perception covers referring segmentation with box and mask output, and LocateAnything-3B covers visual grounding with box or point output, each Python only and each with an empty compile cell. Qwen2-VL and Gemma3 clear the compile check while their Notes cells say nothing, so the table leaves their runtime unspecified.
Each models/ directory has a dated news entry and its own naming convention
The Latest News table is an index into `models/`, one dated row and one detail link per entry. The two most recent are 2026.09.17 for Qwen3-Embedding-0.6B and 2026.09.16 for Qwen3-TTS. Neither is a generative demo in the usual sense. The embedding model is described by its pipeline, last-token pooling, L2 normalize, MRL truncation, and retrieval on BM1684X. Qwen3-TTS covers text-to-speech with 3-second voice cloning across 10 languages, broken into an ECAPA speaker encoder, a 28-layer Talker language model, a CodePredictor, and a Mimi codec. Directory naming is not uniform: most entries use underscores, Qwen3_5, InternVL3, MiniCPMV4_6, while Falcon-Perception keeps its hyphen, so anything that walks the tree cannot assume one separator. The table also dates MiniCPM-V-4.6 to 2026.06.30 with image and video support while the multimodal table still groups MiniCPM-V-2.6 into its Deployed row.
Python demos are everywhere, C++ demos are limited to three rows
The feature list advertises dual-language reference implementations, and the repository language is C++. The multimodal table shows where that claim narrows. Only three rows say Python + C++: Qwen3.5, Qwen3-VL, and Qwen2.5-VL. Everything else is Python only or marked with a dash, including Gemma4, InternVL3, Qwen3-TTS, Mage-VL, and the three grounding and segmentation demos. The news entries repeat the split, with Qwen3.5 dated 2026.04.15 carrying Python and C++ demos with image and video support, and Qwen3-VL dated 2025.10.15 doing the same. The C++ path is therefore a reference implementation for a short list rather than the default runtime. The layout backs that reading: `harness/`, `tools/`, `template/`, and `support/` sit at the top level beside `models/`, and `.clang-format` and `.clang-tidy` are committed for the C++ sources.
Quantization and parallelism are named without per-model values
The efficiency bullet names AWQ and GPTQ quantized models, dynamic compilation, KV cache, and multi-chip parallelism. No table cell breaks any of those down. The multimodal table has no quantization column, no cache setting, and no chip count, so the Notes cells carry the only model-specific detail: video supported for InternVL3, image, video, and audio for Gemma4, image and video for Qwen3.5. The clearest evidence for multi-chip work sits in the news table, where the 2025.03.07 row records QwQ-32B and DeepSeek-R1-Distill-Qwen-32B multi-chip demos adapted and points at `./models/Qwen2_5/`. Two 32B models share a directory with the smaller 1.5B, 7B, and 14B distilled variants dated 2025.02.05. `config.json` at the repository root is the one obvious place to look for those knobs, and the visible documentation never describes what it holds.
There is no release channel and no resolved license field
Two repository facts sit outside the model story and both matter when you pin a deployment. The project has no GitHub releases, so there is no versioned artifact to track, and the only dated handle on the code is the last recorded push, 2026-09-23. The README also links a LICENSE file at the repository root while the repository metadata reports the license as NOASSERTION, the value used when a license cannot be determined automatically. For a stack that vendors converted weights into chip-specific binaries, that combination means you follow a moving branch instead of a release, and you resolve licensing from the file yourself. Both are worth checking against your own compliance requirements before you build on top of it.
Editorial conclusion
LLM-TPU fits an engineer who already holds a SOPHGO BM1684X or BM1688 board and wants a generative model running without building a compiler stack first, because the pre-compiled bmodels and the three wired run.sh targets keep that first session short. It does not fit CV186X owners, anyone who needs a C++ demo outside Qwen3.5, Qwen3-VL, and Qwen2.5-VL, or teams that pin deployments to tagged releases. Before building around it, read the LICENSE file directly, confirm your model directory ships a demo for your chip, and inspect what config.json actually holds.
Frequently asked questions
Which chips does LLM-TPU target?
The subtitle and introduction name BM1684X, BM1688, and CV186X, but every row of both model tables lists only BM1684X or BM1688, and CV186X does not appear there.
Do I need TPU-MLIR to run a model with LLM-TPU?
Only for compilation, because `llm_convert.py` needs a TPU-MLIR environment that can be a Docker image or a source build. Pre-compiled bmodels ship with each demo and require no compilation step.
Which LLM-TPU models ship a C++ demo alongside the Python one?
Three rows in the multimodal table say Python + C++: Qwen3.5, Qwen3-VL, and Qwen2.5-VL. The rest are Python only.
What values does ./run.sh --model accept in LLM-TPU?
Three: qwen3 for Qwen3-4B, qwen3.5 for Qwen3.5-2B, and internvl3 for InternVL3-2B. Other models under models/ are launched from their own directories.
How does LLM-TPU handle quantization, caching, and multi-chip work?
The feature list names AWQ and GPTQ quantized models, dynamic compilation, KV cache, and multi-chip parallelism, and the 2025.03.07 news row records adapted multi-chip demos for QwQ-32B and DeepSeek-R1-Distill-Qwen-32B.
Is Qwen3-Embedding-0.6B a text-to-text generative demo in LLM-TPU?
No. Its entry dated 2026.09.17 describes a non-generative text embedding pipeline on BM1684X using last-token pooling, L2 normalize, MRL truncation, and retrieval.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/sophgo-llm-tpu)