# nndeploy: a visual workflow and multi-backend deployment framework for edge AI

> nndeploy wraps 13 inference runtimes behind a drag-and-drop workflow graph and a JSON export, so one pipeline can target desktop, mobile, edge devices and servers. The trade-off is a C++/CMake build with optional submodules and a documentation set that is mostly Chinese.

**nndeploy/nndeploy** — 一款简单易用和高性能的AI部署框架 | An Easy-to-Use and High-Performance AI Deployment Framework

- Repository: https://github.com/nndeploy/nndeploy
- Website: https://nndeploy-zh.readthedocs.io/zh-cn/latest/
- Stars: 1,882 · Forks: 230
- Language: C++
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/nndeploy-nndeploy

## The problem nndeploy targets: one algorithm, many runtimes

A model that runs in PyTorch on a workstation rarely runs unchanged on a Jetson, an Ascend310B board, an iPhone and a T4 server. Each target has its own runtime, its own tensor layout, and its own preprocessing quirks. Teams usually absorb that cost by writing a separate integration per device, which means the detection pipeline on Android drifts away from the one on Linux until nobody can diff them.

nndeploy approaches this as a deployment framework rather than a converter. The README describes it as solving AI algorithm deployment on desktop (Windows, macOS), mobile (Android, iOS), edge devices (NVIDIA Jetson, Ascend310B, RK) and single-machine servers (RTX series, T4, Ascend310P). The intended user is an engineer who has a working model and needs it running on several of those targets without rewriting the surrounding pipeline each time.

The README also draws a boundary worth reading carefully. For models above 10B parameters, including large language models and AIGC generative models, it positions itself as a visual workflow tool rather than a latency-optimised inference engine. That is a different promise from the one made to a YOLOv8 detection pipeline, and the two should not be evaluated the same way.

## How the workflow graph, nodes and JSON export fit together

The core abstraction is a directed graph of nodes. Each node is one step: preprocess, infer, postprocess, codec, and so on. The README states that over 100 visual nodes are available and that the graph is edited by dragging nodes and adjusting parameters while the result is visible. Because the graph is the unit of work, the same topology can be re-pointed at a different inference backend instead of being rebuilt.

Execution is not necessarily sequential. The README lists serial, pipeline-parallel and task-parallel execution modes, plus memory strategies described as zero-copy, memory pooling and memory reuse. Those matter most in video pipelines, where a codec node, a detector and a tracker can overlap instead of running in lockstep.

Nodes are extensible in both languages. The README states that Python and C++ custom nodes are supported, so a preprocessing step can be written in Python while a latency-sensitive step is written in C++ or CUDA, and both attach to the same graph. The graph itself exports to JSON, which is then loaded through the C++ or Python API on Linux, Windows, macOS or Android. That JSON is the deployment artifact: the thing you version, review and ship, rather than a hand-written main() per platform.

The backend list is the other half of the design. ONNXRuntime, TensorRT, OpenVINO, MNN, TNN, ncnn, CoreML, AscendCL, RKNN, SNPE, TVM and PyTorch are all marked supported, along with an internal inference submodule. The README notes that engines can be selected at compile time to reduce dependencies, and that a custom inference framework can be integrated in a standalone mode.

## Installing nndeploy and running a first workflow

The repository is a CMake project, and the top level contains build_linux.py, build_mac_arm64.py, build_win.py, clone_submodule.py and clean.py alongside CMakeLists.txt. The README does not spell out a single copy-paste install sequence, so treat the scripts and the ReadTheDocs site as the source of truth rather than the README alone.

Start by cloning with submodules, since third_party and .gitmodules are present at the top level and the inference backends are pulled in that way:

```bash
git clone --recursive https://github.com/nndeploy/nndeploy.git
cd nndeploy
```

If you already cloned without --recursive, the repository ships a helper for that case:

```bash
python clone_submodule.py
```

Platform builds go through the Python wrappers. On Linux the entry point is build_linux.py:

```bash
python build_linux.py
```

The Python side has its own dependency list in requirements.txt, installed with the command the file documents at the top:

```bash
pip install -r requirements.txt
```

That file separates required packages (cython, packaging, Pillow, numpy, opencv-python>=4.10.0, plus the server stack: modelscope, multiprocess, requests>=2.31.0, fastapi>=0.104.0, uvicorn>=0.24.0, websockets>=11.0, python-multipart>=0.0.6, pydantic>=2.0.0, chardet>=5.2.0) from optional ones (torch>=2.0.0, torchvision>=0.15.0, onnx>=1.16.0, onnxruntime>=1.18.0, accelerate, diffusers, transformers>=4.51.0, triton>=3.0.0, flash-attn, xxhash>=3.0.0). Installing everything is not the intended path; the optional section exists so you can skip the generative and GPU-heavy parts if you only need detection.

For a first real run, the demo directory is the map. demo/ contains per-task folders such as detect, classification, segment, ocr, llm, matting, depth, keypoint and interpret, plus demo/infer and demo/inference. The README's own model table points at YOLOv5 through YOLOv11, Paddle OCR, Segment Anything, RBMGv1.4, QWen-2.5 and QWen-3, and the Stable Diffusion family via diffusers. A reasonable first target is a detection demo, because it exercises preprocessing, an inference backend and postprocessing without pulling in the diffusers or LLM stacks.

What you should see after a successful build is a runnable demo binary or Python entry point that loads a graph, reports the selected inference backend, and produces annotated output. The README does not document a rollback procedure for a workflow JSON that fails to load, so keep the previous exported graph under version control before you edit the visual one.

## Where nndeploy is the wrong tool

The first limitation is build cost. A C++ framework with submodules for 13 runtimes is not a pip install away from a working pipeline. Selecting backends at compile time reduces the dependency surface, but the initial configuration still assumes you can build C++ on your target, and the README does not claim that every backend works on every listed platform. If your deployment target is a managed inference service, the graph layer is overhead you will not use.

The second is documentation language. The homepage is nndeploy-zh.readthedocs.io and the primary README is Simplified Chinese; README_EN.md exists, but the deeper documents referenced from the Chinese README, such as the deployed model list and the internal inference submodule page, are linked under docs/zh_cn. Expect to read Chinese for anything beyond the front page. That is a real cost for an English-speaking team, and it is not a temporary state you can assume away.

The third is the large-model positioning. The README recommends nndeploy as a visual workflow tool for models above 10B parameters. A visual graph is a good way to experiment with a diffusion pipeline; it is not a claim about serving throughput or memory efficiency at that scale. If your requirement is maximum tokens per second on a single model, a dedicated serving stack is the better comparison.

The fourth is that the README does not document rollback, versioning of workflow JSON, or a migration path between releases. Release notes are the only signal for breaking changes, and there is no statement about API stability across the v3.0.x line.

## How nndeploy differs from ONNX Runtime or TensorRT alone

The honest alternative is not another framework of the same shape; it is using one runtime directly. ONNX Runtime gives you a session, tensors and execution providers. TensorRT gives you an optimised engine for NVIDIA hardware. Both are mature, both have English documentation, and both let you write exactly the pipeline you need in C++ or Python.

The difference is where the abstraction sits. With ONNX Runtime alone, the pipeline is your code: preprocessing, the session call, postprocessing, and a separate branch for each platform you support. With nndeploy, the pipeline is a graph of nodes that exports to JSON, and the runtime becomes a selectable backend behind the inference node. You trade direct control of the session for a topology that can be re-pointed at MNN on Android and TensorRT on a T4 without rewriting the surrounding code.

That trade favours teams with several targets and repeated pipelines. It works against teams with one target, or teams that need a runtime feature nndeploy's node does not expose, because you are then writing a custom node anyway and paying for the graph on top. The README's support for custom inference frameworks in standalone mode is the escape hatch for that case, and it is worth checking before you assume a missing feature blocks you.

## Maintenance, licence and the cost of staying current

The repository is not archived, and the last push was on 2026-08-15. Recent releases are v3.0.10 and v3.0.9, both dated 2026-04-04, with v3.0.8 on 2025-12-04. So the release cadence in the v3.0.x line has been roughly quarterly, while commits continue between releases. There is no long-term support branch mentioned, and no stated deprecation policy for node APIs or for the workflow JSON schema.

Upgrade cost is dominated by the backends, not by nndeploy itself. Because engines are pulled in as submodules and can be selected at compile time, bumping TensorRT or AscendCL is a submodule and build-configuration change, and each backend carries its own compatibility matrix with drivers and SDKs. A team enabling three backends is maintaining three of those matrices.

nndeploy is Apache-2.0, which is permissive and includes an explicit patent grant. That licence covers nndeploy's own code. It does not relicense the inference engines it links: TensorRT, OpenVINO, MNN, TNN, ncnn, CoreML, SNPE and the rest each ship under their own terms, and some have redistribution conditions that matter if you ship a binary. Check the licence of every backend you enable before distributing, and treat this as an engineering checklist item rather than a legal opinion.

## Conclusion

nndeploy fits teams that must ship the same vision or generative pipeline onto several runtimes and want the graph, not the glue code, to be the artifact they maintain. It is a poor fit if you only ever target one backend, if you need a fully English documentation set, or if your team will not touch CMake and third_party submodules. Before committing, verify three things: that your target backend appears in the supported list, that the workflow JSON your graph exports still loads through the C++ or Python API, and that the licence terms of every backend you enable are acceptable for your distribution, since nndeploy itself is Apache-2.0 but the inference engines it links are not covered by that licence.

## FAQ

### What is deployment in software, and how does nndeploy relate to it?

In software, deployment is the step that takes a built artifact and puts it where it runs. nndeploy addresses the AI-specific part of that: moving a trained model onto desktop, mobile, edge or server hardware and running it through a pipeline, using a workflow graph exported as JSON and a choice of 13 inference backends.

### What does it mean to deploy an AI model with nndeploy?

It means building a graph of nodes (preprocess, inference, postprocess and others), exporting that graph to JSON, and loading it through the C++ or Python API on Linux, Windows, macOS or Android. The inference backend behind the graph is selected at compile time from the supported list.

### Which inference frameworks does nndeploy support?

The README lists ONNXRuntime, TensorRT, OpenVINO, MNN, TNN, ncnn, CoreML, AscendCL, RKNN, SNPE, TVM and PyTorch, plus an internal inference submodule. Backends can be selected at compile time to reduce dependencies, and a custom framework can be integrated in standalone mode.

### Is nndeploy suitable for large language models?

The README states that for models above 10B parameters, including large language models and AIGC generative models, nndeploy is suitable as a visual workflow tool. It does not present itself as a serving engine optimised for throughput at that scale.

### What licence does nndeploy use?

The repository is Apache-2.0. That covers nndeploy's own code; the inference engines it links, such as TensorRT, OpenVINO, MNN and the others, keep their own licences and need separate review before you redistribute a binary.

## Sources

- [License: Apache-2.0](https://github.com/nndeploy/nndeploy/blob/main/LICENSE)
- [nndeploy/nndeploy on GitHub](https://github.com/nndeploy/nndeploy)
- [Project website](https://nndeploy-zh.readthedocs.io/zh-cn/latest/)
- [README](https://github.com/nndeploy/nndeploy/blob/main/README.md)
- [Releases](https://github.com/nndeploy/nndeploy/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/nndeploy-nndeploy
