# MLLM: a multimodal LLM inference engine for phones, NPUs and Jetson boards

> MLLM (UbiquitousLearning/mllm) is a C++ inference engine for running multimodal models on mobile and edge hardware, with a Python layer called pymllm and an Android demo built on an in-app Go server. This review covers how the pieces fit together, how pymllm is packaged, and where the project is still thin.

**UbiquitousLearning/mllm** — Fast Multimodal LLM on Mobile Devices

- Repository: https://github.com/UbiquitousLearning/mllm
- Website: https://ubiquitouslearning.github.io/mllm/
- Stars: 1,617 · Forks: 218
- Language: C++
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/ubiquitouslearning-mllm

## The problem MLLM targets: multimodal inference on hardware you cannot rent

Most multimodal inference tutorials assume a data center GPU. MLLM assumes the opposite: a phone, a Jetson board, or an edge box with a Qualcomm Hexagon or Ascend NPU. The README describes it as a "Fast and lightweight multimodal LLM inference engine for mobile and edge devices", and the supported-model table makes the constraint concrete. Models listed for mllm v2 include Qwen3-0.6B, Qwen3-1.7B, Qwen3-4B, Qwen3.5-0.8B and the newly added MiniCPM5-2B, with quantized variants such as w4a8, W4A16 and W8A8 rather than full-precision weights.

The audience is therefore narrow but specific: an Android or embedded engineer who has a model picked out, a target chip, and no tolerance for a Python-only runtime. The repository is C++ with a CMake build, and the Python package pymllm sits on top of it. If your deployment target is an x86 server with an A100, this project is not aimed at you, and the README never claims otherwise.

## How the pieces fit: convertor, runtime, and an in-app Go server on Android

The data flow described in the README has three stages. First, mllm-convertor ingests PyTorch and SafeTensors checkpoints from community frameworks, quantizes them, and writes them into mllm format. Second, the mllm Runtime loads and executes those converted files. Third, on Android the project does not use the usual JNI integration. Instead it ships an in-app server built with Golang (mllm_server.aar), which the README calls a Client-Server architecture entirely on-device. The stated reason is decoupling the UI from inference computation, and the November 2025 demo update ties that architecture to stable Qwen3 and DeepSeek-OCR streaming on Android.

Above the runtime sit optimization passes the README groups as speculative decoding, pruning and quantization; below it sit compiler and runtime layers named CANN, CUDA and MLIR. The project positions itself as the node between those two halves. That is a fair description of the repository layout: algorithms/, mllm-kernel/, mllm-ext-opset/ and third_party/ are separate directories, and the examples/ tree contains one folder per model family, including qwen3_qnn_aot, qwen_ascend, qwen_npu and qwen3_service.

## Installing pymllm and converting a checkpoint for a QNN NPU target

The Python package is published as pymllm, version 2.0.2 in pyproject.toml, and requires Python 3.10 or newer. The build backend is scikit-build-core, so installing from source triggers a CMake build with MLLM_ENABLE_PY_MLLM=on and CMAKE_BUILD_TYPE=Release. Note the environment variable the file documents for skipping that build: setting SKBUILD_WHEEL_CMAKE=false skips the CMake step, which is only useful if you already have the native library.

```bash
pip install pymllm
```

That installs the console scripts declared in pyproject.toml: pymllm, mllm-convertor, mllm-service and pymllm-server. The base dependency list includes torch, torchao, modelscope, fastapi, uvicorn, typer and apache-tvm-ffi pinned at 0.1.8.post2. CUDA extras (tilelang, flashinfer-python, pyzmq) are optional and separate, declared under the cuda extra.

For a first real task, convert a checkpoint and then run inference. The README states that mllm-convertor reads PyTorch and SafeTensors checkpoints and writes mllm format, so the conversion step comes before any runtime call. The convertor entry point is defined in pyproject.toml as pymllm.mobile.utils.mllm_convertor:main, and the same file defines the pymllm CLI as pymllm.__main__:main.

If you are targeting a Qualcomm NPU, the February 2026 release added QNN AOT support for full graph execution, with a quick start page at docs/qnn_backend/aot_execute.html and a worked example directory at examples/qwen3_qnn_aot/. The repository also ships requirements-qnn-aot.txt as a separate dependency file, which tells you the QNN path pulls in packages the base install does not. The README does not document rollback or downgrade steps for a converted checkpoint, so keep the original SafeTensors files.

## What the benchmarks do and do not say

The README reports prefill and decode numbers for pymllm on Jetson Orin. For input_len=2048 and output_len=128, it states that Qwen3-VL-2B W8A8 reaches up to 3.12x prefill speedup on AGX Orin 32GB and about 12243 tok/s prefill throughput, while decode throughput stays broadly close to llama.cpp with small wins or losses depending on model, device and quantization. A separate multimodal table measures the full vision-encoding plus image/text token prefill path via bench_one_batch --image, and reports mean TPS: 4875.75 FP16, 4700.28 W4A16 and 6443.59 W8A8 for Qwen3-VL-2B on AGX Orin 32GB.

Read those numbers carefully. They are the project's own measurements on its own hardware, not an independent result, and the README itself concedes that decode is roughly at parity with llama.cpp. The honest summary is that MLLM's advantage on Jetson is concentrated in prefill, and that a decode-bound workload will not see the same gain. The W4A16 and W8A8 rows also show that quantization is not uniformly faster: Qwen3-VL-2B on AGX Orin is slower at W4A16 (4700.28) than at FP16 (4875.75).

## Where MLLM is the wrong choice

The v1 to v2 transition is the largest practical risk. The README states that support for MLLM V1 was ending, that V1 would integrate GPT-OSS before retirement, and that V2 introduces a different model authoring approach with eager execution, compilation for NPU integration, and parallel execution of multiple models. Any code written against V1 semantics will need rework, and the release history shows the gap: 1.0.0 in January 2024, then 2.0.0 in February 2026.

The second limitation is backend coverage. The v2 support table has empty cells. Qwen3-0.6B lists CPU and Ascend NPU but no Hexagon NPU entry; Qwen3-4B lists CPU only; Qwen3-1.7B is the one row with a Hexagon NPU W4A16-SM8650 build. So "unified hardware support" in the feature list means the framework has backends, not that every model runs on every backend. Check the row for your model before planning a port.

Third, the CUDA path for Jetson Orin and Jetson Thor is described in the March 2026 note as experimental and still under active development. Treat it as such. Finally, the README documents no rollback procedure for a failed conversion and no compatibility guarantee for converted checkpoints across pymllm versions, which matters if you plan to ship an app that updates the runtime independently of the model files.

## llama.cpp, ONNX Runtime and ExecuTorch: what actually differs

llama.cpp is the closest comparison and the one the README benchmarks against. Both are C++ inference engines, but llama.cpp centers on GGUF quantized weights and CPU or GPU execution, while MLLM routes models through a dedicated convertor into its own format and puts more weight on vendor NPU paths: QNN AOT for Hexagon, ATB graph execution for Ascend, and a CUDA path for Jetson. If your target is a plain Arm CPU and a Llama-family text model, llama.cpp covers that ground with a wider model catalogue. MLLM's differentiator is the multimodal prefill path and the NPU compilers, not text-only CPU inference.

ONNX Runtime is a general graph runtime with execution providers; MLLM is a model-specific engine with per-family example directories. Choosing ONNX Runtime means you own the graph export and the operator coverage problem. Choosing MLLM means you accept its supported-model table as the boundary of what you can ship. ExecuTorch, from the PyTorch side, keeps you inside the PyTorch export story; MLLM instead ingests PyTorch and SafeTensors checkpoints through mllm-convertor and does its own quantization. The practical question is whether your model is in MLLM's table. If it is not, the convertor will not help you.

## Licence, maintenance and the cost of upgrading

MLLM is MIT licensed, which permits commercial use and modification with attribution and without a copyleft obligation on your own code. The licence text is in the repository root as LICENSE. That covers the framework; it does not cover the model weights, which come from separate sources such as Hugging Face and ModelScope and carry their own terms. If you ship a converted Qwen or MiniCPM checkpoint inside an app, check that model's licence separately. This is not legal advice.

The last push to the repository was on 2026-09-08, and the most recent release is 2.0.0 from 2026-02-16. The README's news entries run from July 2025 through September 2026, so the project is being extended, but the release cadence and the commit cadence do not match: much of the recent work appears in pymllm and in example directories rather than in tagged releases. For an adopter this means pinning a commit or a pymllm version rather than tracking main. The upgrade cost is dominated by re-conversion, since the runtime reads mllm-format files produced by mllm-convertor, and the README does not promise that files converted by one pymllm version load in the next.

## Conclusion

Adopt MLLM if you are deploying a supported model (Qwen3, Qwen3-VL, Qwen3.5, MiniCPM5-2B, DeepSeek-OCR) to Arm CPUs, Hexagon NPUs, Ascend NPUs or Jetson Orin, and you are willing to build from source with CMake. Do not adopt it if you need a stable API contract across releases: v1 support ended, v2 shipped in November 2025, and the CUDA path for Jetson is still marked experimental. Before committing, verify that a converted checkpoint exists for your exact model and quantization on ModelScope, and check that the backend you need appears in the v2 support table, because the table has empty cells for several model and NPU combinations.

## FAQ

### What is MLLM?

MLLM is a C++ inference engine for running multimodal LLMs on mobile and edge devices, described in the README as fast and lightweight. It ships a Python package called pymllm, a convertor for PyTorch and SafeTensors checkpoints, and an Android demo built on an in-app Go server.

### Which models does MLLM support?

The v2 support table lists Qwen3-0.6B, Qwen3-1.7B, Qwen3-4B and Qwen3.5-0.8B, and the news entries add MiniCPM5-2B, Qwen3-VL and DeepSeek-OCR. Backend coverage differs per model: Qwen3-1.7B has a Hexagon NPU W4A16-SM8650 build, while Qwen3-4B lists CPU only.

### How do I install MLLM?

The Python package is pymllm, version 2.0.2, requiring Python 3.10 or newer. Installing it runs a CMake build through scikit-build-core unless you set SKBUILD_WHEEL_CMAKE=false, and the install exposes the pymllm, mllm-convertor, mllm-service and pymllm-server commands.

### Does MLLM run on Android?

Yes. The README describes an Android demo that uses a Client-Server architecture entirely on-device, with an in-app server built in Golang shipped as mllm_server.aar. The November 2025 demo update enabled Qwen3 and DeepSeek-OCR streaming on Android through that architecture.

### Is MLLM faster than llama.cpp?

The README reports up to 3.12x prefill speedup for Qwen3-VL-2B W8A8 on AGX Orin 32GB at input_len=2048 and output_len=128, but states that decode throughput is generally close to llama.cpp with small wins or losses depending on model, device and quantization.

## Sources

- [License: MIT](https://github.com/UbiquitousLearning/mllm/blob/main/LICENSE)
- [Project website](https://ubiquitouslearning.github.io/mllm/)
- [README](https://github.com/UbiquitousLearning/mllm/blob/main/README.md)
- [Releases](https://github.com/UbiquitousLearning/mllm/releases)
- [UbiquitousLearning/mllm on GitHub](https://github.com/UbiquitousLearning/mllm)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/ubiquitouslearning-mllm
