nncase: compiling neural networks for Kendryte K210, K510 and K230 accelerators
Open deep learning compiler stack for Kendryte AI accelerators ✨
At a glance
- What is it?
- nncase is an Apache-2.0 neural network compiler that turns TFLite, Caffe and ONNX models into code for Kendryte AI accelerators. It targets a narrow set of chips, and that narrowness is both its value and its main constraint.
- Who is it for?
- Adopt nncase if your target is a Kendryte K210, K510 or K230 and your model is already in TFLite, Caffe or ONNX form. Do not adopt it for x86, GPU or general-purpose embedded inference; it compiles for Kendryte accelerators, and that is the whole scope.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 75 days ago.
- What is it written in?
- Mainly C#, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What nncase solves, and who it is actually for
Training frameworks emit models for CPUs and GPUs. Kendryte accelerators are neither. nncase sits between the two: it takes a trained model and produces something the KPU inside a K210, K510 or K230 can execute, handling operator mapping, quantization and memory layout along the way. The README describes it in one line as "a neural network compiler for AI accelerators."
The audience is narrow and specific. You are building firmware or an application for a Kendryte board, you have a model in TFLite, Caffe or ONNX, and you need it to run on that silicon rather than on a host. If you are deploying to a Raspberry Pi, a Jetson or a phone, nncase has nothing to offer you. The repository topics list k210, k230 and k510, and the README splits its usage documentation along the same lines: K230 material under docs/USAGE_v2_EN.md, and K210/K510 material on the release/1.0 branch. That split is the first thing to understand, because the two branches are not interchangeable.
How the compiler pipeline is put together
The repository layout shows a conventional compiler split. src/ holds the core, targets/ holds the per-chip backends, modules/ holds reusable pieces, and python/ holds the bindings that pip installs. The build is CMake-driven (CMakeLists.txt, cmake/, conanfile.py) with setuptools wrapping it for wheel production, which is why setup.py contains a CMakeExtension class that shells out to CMake rather than compiling sources directly.
The feature list in the README names the properties that matter at runtime: multiple inputs and outputs with multi-branch structures, static memory allocation with no heap memory acquired, and operator fusion. Static allocation is the design decision with the widest consequences. A model's peak memory is fixed at compile time, so the runtime never asks the allocator for more. On a microcontroller with a few megabytes of RAM that predictability is the point. It also means the compiler must resolve every buffer up front, and a model whose activations do not fit will fail at compile time rather than degrade at run time.
The input surface is documented per format. The README links separate operator tables for TFLite, Caffe and ONNX, and those tables, not the feature list, are what determine whether a given model compiles.
Installing nncase and compiling a first model
The README gives the install path for K230 directly. On Linux, one pip command pulls both the compiler and the KPU plugin:
pip install nncase nncase-kpuWindows needs two steps because the KPU wheel is distributed separately. The README instructs you to install nncase first, then download nncase_kpu-2.x.x-py2.py3-none-win_amd64.whl from the releases page and install that file by path. The Python requirement in pyproject.toml is >=3.9, and the classifiers list 3.9 through 3.13, so an older interpreter will not work.
Once installed, the documented entry point for a first run is the K230 simulation notebook at examples/user_guide/k230_simulate-EN.ipynb. It is the only worked example the README points to for the current branch, and the same page offers a Colab link so you can run it without a board attached. That is the realistic way to find out whether your operators are supported before you flash anything.
Version alignment is a separate install step that the README treats as its own topic. It links a page on the version relationship between nncase and K230_SDK, plus a guide for updating the nncase runtime library inside an SDK. Compiler and runtime are versioned separately from the SDK, so mismatches are a documented concern rather than an edge case. The current release line is v2.11.0, published on 2026-02-28; v2.10.0 came out on 2025-08-08 and v2.9.0 on 2024-07-31, which is a cadence of roughly one release per year on the 2.x branch.
Where nncase stops being the right tool
The operator tables are the hard boundary. A model that uses an operator absent from docs/onnx_ops.md, docs/tflite_ops.md or docs/caffe_ops.md will not compile, and no amount of configuration changes that. Custom layers and unusual attention variants are the usual casualties. The README's own feature list is careful to say "Support f" before the text is cut off, so the full list of supported constructs is not visible from the README alone; the operator tables are the authoritative source.
The K210 and K510 story is the second boundary. The README points those chips at the release/1.0 branch for usage, FAQ and examples, while K230 lives on the default branch. Anyone arriving with a K210 and following the top-level instructions will land on documentation written for a different generation of the tool. That is a documentation-organisation problem rather than a technical one, but it costs time.
Finally, the project is not a runtime for general embedded Linux. It compiles for Kendryte accelerators. If your deployment target changes to a different SoC, the compiled artifact is worthless and you are back to a different toolchain. Treat nncase as a commitment to one vendor's silicon.
How nncase differs from generic edge inference runtimes
The natural comparison is with runtimes such as TensorFlow Lite for Microcontrollers or ONNX Runtime. Those take an already-prepared model and execute it, with the operator set and memory behaviour defined by the runtime. nncase is a compiler: it reads the model, lowers it to the target's instruction set, and emits an artifact plus a runtime library that the K230_SDK links against. The README's own benchmark table makes the division visible, listing nncase_fps alongside tflite_onnx_result accuracy for the same models, so you can see what quantization to u8/u8 costs in top-1 and mAP terms before you commit.
That table is worth reading as a design statement rather than a scoreboard. For mobilenetv2 at [1,224,224,3], the listed top-1 is 71.1% against a reference 71.3%. For yolov8s_det at [1,3,640,640], mAP50-90 is listed as 0.404 against 0.446. The gaps are small in classification and larger in detection, which is the expected shape when activations are quantized to eight bits. The trade-off is deliberate: you accept some accuracy loss to get a model that fits and runs on the accelerator at all.
Maintenance, licensing and the cost of upgrading
The repository is not archived, and the last push was on 2026-07-17, so there is recent activity. Releases are less frequent: v2.11.0 landed on 2026-02-28, roughly six months before that last push. Between releases, the default branch moves, which means anyone building from source is tracking something ahead of the packaged wheels.
Upgrade cost concentrates in two places. The first is the compiler-to-runtime pairing: because the README devotes a page to the nncase and K230_SDK version relationship, moving nncase forward without moving the SDK is a known source of trouble. The second is operator coverage, which changes between releases and can alter what compiles.
Licensing is Apache-2.0, declared both in the LICENSE file and in the pyproject.toml classifier. Apache-2.0 permits commercial use and modification and includes an explicit patent grant, which matters for a compiler that generates code for hardware. It also requires that you preserve notices and state significant changes. That is a summary of the licence text, not legal advice; read LICENSE and your own counsel's view before shipping a product.
Editorial conclusion
Adopt nncase if your target is a Kendryte K210, K510 or K230 and your model is already in TFLite, Caffe or ONNX form. Do not adopt it for x86, GPU or general-purpose embedded inference; it compiles for Kendryte accelerators, and that is the whole scope. Before committing, check your operators against docs/onnx_ops.md, docs/tflite_ops.md or docs/caffe_ops.md, and confirm that the nncase version you install matches the one your K230_SDK ships with, since the README links to a version-relationship page for exactly that reason.
Frequently asked questions
How do I install nncase on Linux?
The README gives a single command for Linux: pip install nncase nncase-kpu. Python 3.9 or newer is required according to pyproject.toml, which lists classifiers for 3.9 through 3.13.
How do I install nncase on Windows?
Windows takes two steps. Install nncase with pip first, then download nncase_kpu-2.x.x-py2.py3-none-win_amd64.whl from the releases page and install that wheel file by path, because the KPU plugin is distributed separately for Windows.
Which model formats does nncase accept?
The README links operator tables for TFLite, Caffe and ONNX, so those are the three input formats the documentation covers. Whether a specific model compiles depends on whether its operators appear in the relevant table.
Which Kendryte chips does nncase support?
The README and repository topics name the K210, K510 and K230. K230 usage is documented on the default branch, while K210 and K510 usage, FAQ and examples are pointed at the release/1.0 branch.
Does nncase allocate memory at runtime?
The README's feature list states that nncase uses static memory allocation and acquires no heap memory. Peak memory is therefore determined when the model is compiled rather than when it runs.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/kendryte-nncase)