Tencent/ncnn: deploying neural networks on phones and edge devices
ncnn is a high-performance neural network inference framework optimized for the mobile platform
At a glance
- What is it?
- ncnn is a C++ inference framework with no third-party runtime dependencies, a Vulkan GPU backend and the pnnx converter for PyTorch and ONNX models. Here is how the pipeline works, how to install it, and where it stops being the right choice.
- Who is it for?
- Adopt ncnn when you ship a model inside a mobile, embedded or desktop binary and cannot afford a Python runtime or a large dependency tree; the pnnx route from PyTorch is the documented beginner path and the prebuilt release archives cover Android, iOS, macOS, Linux, Windows, WebAssembly, watchOS, tvOS and visionOS.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository last received commits 13 days ago.
- What is it written in?
- Mainly C++, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 17, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The deployment gap ncnn was built to close
Training frameworks assume a host with a Python interpreter, a package manager and gigabytes of spare memory. A phone application, a Raspberry Pi, a browser tab or a set-top box has none of that. ncnn targets exactly that gap: it is a C++ inference framework whose README states it has no third-party runtime dependencies, so the model and the runtime travel together in one binary. The audience is the engineer who has a trained model and needs it to run inside an app, not behind an HTTP endpoint. The README names Tencent applications that already ship it, including QQ, Qzone, WeChat and Pitu, which tells you the framework was shaped by consumer mobile constraints rather than data-centre ones. The supported targets listed in the release notes span Android, HarmonyOS, iOS, macOS, Linux, Windows, WebAssembly, watchOS, tvOS and visionOS. That breadth is the point: one inference API, many deployment surfaces.
How the pnnx conversion pipeline and the extractor API fit together
The documented beginner path is PyTorch to pnnx to ncnn. pnnx traces a model and emits two files, a .param text file describing the graph and a .bin file holding the weights. At runtime ncnn::Net loads both, and inference goes through an extractor: you create one, feed named input blobs with ex.input, and pull named output blobs with ex.extract. The blob names, in0 and out0 in the README example, come from the converter, not from you, so a mismatch between the exported graph and your calling code is the first thing to check when output shapes look wrong.
Backends are selected at the ncnn::Net level, with CPU and Vulkan GPU paths available according to the README. The CPU path is built around SIMD kernels for ARM NEON and RISC-V, which is why the topics list reads like a hardware catalogue. The practical consequence is that the same .param and .bin pair runs on a phone CPU, a desktop Vulkan device or a WebAssembly build, and performance tuning happens through option flags rather than a second export. That is a genuinely different architecture from runtimes that compile a graph to a device-specific plan ahead of time.
Installing ncnn and running a first model
The README recommends the PyTorch route for beginners. pnnx installs as a Python package inside an existing PyTorch environment:
pip3 install pnnxYou then export a model by calling pnnx.export on an eval-mode module with a sample input tensor. The README example defines a small conv, relu, mean and linear stack, builds it with Model().eval(), creates a torch.rand tensor of shape (1, 3, 224, 224) and exports:
import torch
import torch.nn as nn
import pnnx
class Model(nn.Module):
def __init__(self):
super().__init__()
self.conv = nn.Conv2d(3, 8, 1)
self.relu = nn.ReLU()
self.fc = nn.Linear(8, 4)
def forward(self, x):
x = self.conv(x)
x = self.relu(x)
x = x.mean((2, 3))
return self.fc(x)
model = Model().eval()
x = torch.rand(1, 3, 224, 224)
pnnx.export(model, "model.pt", (x,))According to the README this generates model.ncnn.param and model.ncnn.bin in the working directory. Those two files are the artefact you ship. Loading them from C++ uses the header net.h:
#include "net.h"
ncnn::Net net;
net.load_param("model.ncnn.param");
net.load_model("model.ncnn.bin");
ncnn::Mat in(224, 224, 3);
auto ex = net.create_extractor();
ex.input("in0", in);
ncnn::Mat out;
ex.extract("out0", out);If you would rather stay in Python, the repository ships a Python binding. The README shows the same flow with a NumPy array wrapped in ncnn.Mat, and the extract call returns a status code plus the output blob:
import numpy as np
import ncnn
net = ncnn.Net()
net.load_param("model.ncnn.param")
net.load_model("model.ncnn.bin")
x = np.zeros((3, 224, 224), np.float32)
mat = ncnn.Mat(x)
ex = net.create_extractor()
ex.input("in0", mat)
ret, out = ex.extract("out0")
print(np.array(out).shape)The README points to tools/pnnx, the wiki page on using ncnn with PyTorch or ONNX, the python directory and the examples directory for complete workflows. For a first real use, the examples folder is the better starting point than the snippet above: it contains complete programs for mobilenetv2ssdlite, nanodet, scrfd, ppocrv5, rvm and piper, each of which pairs a model with the pre- and post-processing the model expects. Running one of those end to end tells you more about input normalisation and blob naming than any minimal example. Prebuilt libraries for the listed platforms are attached to the releases; the README links a how-to-build wiki page for Linux, Windows, macOS, Raspberry Pi 3 and 4, POWER, Android, NVIDIA Jetson, iOS, WebAssembly, AllWinner D1 and Loongson 2K1000.
Where ncnn is the wrong tool
The framework assumes you can rebuild and reship your application when the model changes. There is no model server, no hot reload, no versioned endpoint. If your deployment model is a fleet of backend services where the model is updated independently of the client, ncnn adds a compilation step and a binary size cost for nothing.
The README does not document rollback. If a newly converted model produces worse output on a subset of devices, nothing in the documented workflow lets you pin a previous graph version at runtime; you fall back to whatever your own release process provides. That is a real operational gap for teams used to server-side model registries.
Operator coverage is the second boundary. pnnx has to translate every operation in your graph, and the repository's examples skew toward vision and speech models: detection, face recognition, OCR, pose, image restoration, text to speech. There is no example of a large language model in the file listing, and the README does not claim transformer-scale text generation support, so anyone searching for large language model deployment should treat that as unverified rather than assumed. If your model relies on a dynamic control flow pattern that the tracer cannot capture, the export is where you find out, not the runtime.
ncnn versus ONNX Runtime: two different bets
The comparison people actually search for is ncnn against ONNX. ONNX Runtime is the natural alternative, and the difference is architectural rather than a matter of speed claims. ONNX Runtime takes an ONNX graph and executes it through pluggable execution providers, which lets a single runtime delegate to CUDA, TensorRT, DirectML or a CPU backend depending on what is installed on the host. The model format is the interchange standard, and the runtime is the thing you install.
ncnn inverts that. The model is converted into ncnn's own .param and .bin pair by pnnx, and the runtime is compiled into your application with no third-party runtime dependencies. You give up the ability to swap execution providers at deployment time, and you gain a build that has no external runtime to ship or version. ONNX Runtime is the better fit when the host is a server you control and the model is one of several artefacts you deploy. ncnn is the better fit when the host is a device you do not control and the binary is the product. Note that the two are not mutually exclusive at the source: the README links a wiki page on using ncnn with PyTorch or ONNX, so ONNX can be an input format on the way into ncnn.
Release cadence, licence and the cost of upgrading
The repository is not archived, and the last push was on 2026-09-10. Releases are dated rather than semantically versioned: 20250916, 20260113 and 20260526, each with prebuilt libraries for the platform list. That naming tells you the cadence is roughly quarterly and that an upgrade means moving from one date-stamped archive to the next, not from 1.x to 2.x. There is no long-term support branch described in the README.
The upgrade cost sits in the converter, not the library. A new ncnn release may change how pnnx emits a graph, so a model exported against an older pnnx may need re-exporting to match a newer runtime. Because the .param file is a text description of the graph, a diff between the old and new export is readable, which helps, but the README does not describe a compatibility guarantee between pnnx output and older runtime versions. Treat the converter and the runtime as a matched pair and upgrade them together.
The licence situation needs a direct look. The repository metadata reports NOASSERTION, while the README carries a badge reading BSD 3-Clause and the repository root contains LICENSE.txt. Those two signals disagree at the metadata level, and the only way to resolve it is to read LICENSE.txt in the tree you are vendoring. This is not legal advice; it is a note that the automated licence field is not the authority here.
Editorial conclusion
Adopt ncnn when you ship a model inside a mobile, embedded or desktop binary and cannot afford a Python runtime or a large dependency tree; the pnnx route from PyTorch is the documented beginner path and the prebuilt release archives cover Android, iOS, macOS, Linux, Windows, WebAssembly, watchOS, tvOS and visionOS. Do not adopt it if your workload is server-side batched inference on NVIDIA GPUs, where CUDA-oriented runtimes fit better, or if you need a documented rollback story for model upgrades, because the README does not describe one. Before committing, verify that pnnx converts your specific operator set without silent graph rewrites, and check the licence file in the repository rather than the badge, since the metadata reports NOASSERTION while the README badge says BSD 3-Clause.
Frequently asked questions
What is ncnn used for?
It is a C++ neural network inference framework for deploying models on mobile, embedded and desktop targets, including phones, PCs, browsers and edge devices. The README names Tencent applications such as QQ, Qzone, WeChat and Pitu as users.
What does ncnn stand for?
The README and repository description do not spell out the acronym or give an expansion for it, so the name is best treated as a project name rather than an initialism with a documented meaning.
How do I convert a PyTorch model to ncnn?
Install pnnx with pip3 install pnnx in a PyTorch environment, then call pnnx.export on an eval-mode module with a sample input tensor. The README states this generates model.ncnn.param and model.ncnn.bin, which ncnn::Net loads with load_param and load_model.
Which platforms do the ncnn prebuilt releases cover?
The release notes for 20260526 list Android, HarmonyOS, iOS, macOS, Linux, Windows, WebAssembly, watchOS, tvOS and visionOS. The README also links a how-to-build wiki page covering Raspberry Pi, POWER, NVIDIA Jetson, AllWinner D1 and Loongson 2K1000.
Does ncnn have a Python API?
Yes. The repository contains a python directory, and the README shows a Python example that builds an ncnn.Net, loads the param and bin files, wraps a NumPy array in ncnn.Mat and calls extract through an extractor.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/tencent-ncnn)