Paddle Lite: a mobile and edge inference engine for PaddlePaddle models
PaddlePaddle High Performance Deep Learning Inference Engine for Mobile and Edge (飞桨高性能深度学习端侧推理引擎)
At a glance
- What is it?
- Paddle Lite turns PaddlePaddle inference models into optimized binaries for Android, iOS, embedded Linux and NPUs. It is a strong fit if your models are already PaddlePaddle, and the opt tool is a step you cannot skip.
- Who is it for?
- Adopt Paddle Lite if your models are already PaddlePaddle inference models and your targets are Android, iOS, embedded Linux or one of the NPUs listed in the README. Do not adopt it if your models live in PyTorch or TensorFlow and you have no reason to convert.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 155 days ago.
- What is it written in?
- Mainly C++, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 17, 2026, and from our analysis. They are not legal advice.
DEEP OPEN-SOURCE ANALYSIS
What Paddle Lite solves, and who it is for
A trained model is not a deployable artifact. Paddle Lite is the runtime that takes a PaddlePaddle inference model and executes it on a phone, a single-board computer or an edge box. The README describes it as a high performance, lightweight inference framework positioned for mobile, embedded and edge hardware, and says it is used in Baidu's internal business as well as by external users in production.
The audience is narrow and specific. You are an engineer shipping an on-device model to Android or iOS, or to an ARM Linux board, and your model already comes out of PaddlePaddle. If that is your situation, Paddle Lite removes the work of writing your own kernels and of hand-tuning for each chip. If your models come from PyTorch or TensorFlow, the README points you at X2Paddle for conversion, which is a real extra step rather than a detail.
The supported target list is unusually long: Android, iOS, embedded Linux, Windows, macOS and Linux hosts, with backends for ARM CPU, x86, OpenCL, Metal, Huawei Kirin NPU, Huawei Ascend NPU, Kunlunxin XPU and XTCL, Qualcomm QNN, Cambricon MLU, Verisilicon TIM-VX, Android NNAPI, MediaTek APU, Imagination NNA, Intel OpenVINO and an EEasyTech NPU. That breadth is the pitch. It is also where the maintenance questions start, because a long backend list is only useful if the backends you need still build.
The four-step deployment path: model, opt, library, API
The workflow in the README has four stages, and the third one is where most people get stuck.
First, prepare the model. Paddle Lite consumes models saved by PaddlePaddle's save_inference_model API. Models from other frameworks go through X2Paddle first.
Second, optimize the model with the opt tool. This is not optional in practice. The README states that opt performs quantization, subgraph fusion and kernel selection, and that it can also list the operators in a model and report whether a given hardware platform supports them. That second function is the useful one during evaluation: you can check operator coverage before you commit to a target.
Third, get the runtime. Prebuilt libraries are published for Android, iOS, x86 and macOS, and the README recommends downloading those first rather than compiling. If you must compile, it recommends the Docker environment to avoid assembling the toolchain by hand, with per-platform source build guides as the alternative.
Fourth, call the API. C++, Java and Python bindings exist, each with a full example, plus platform-specific demo apps for Android, iOS, Linux, ARM, x86, OpenCL, Metal and the NPU backends. The data flow is fixed: a PaddlePaddle inference model goes through opt, the optimized model is loaded by the runtime on the device, and the chosen backend executes the kernels.
Installing Paddle Lite and running a first inference
The README does not give a single install command for the runtime. It recommends the prebuilt library downloads first, so the realistic path is: download the prebuilt library for your target, then call it from your application. The Python binding is the fastest way to confirm the library works before you touch Android or iOS.
The README's Python demo guide is the reference for the exact import and API names; the top-level README does not quote that code, so check the demo page rather than guessing at symbols. What you are looking for is a predictor object created from the optimized model and the parameter file, an input tensor you fill, and a run call whose output you read back.
For Android, the README lists prebuilt demo APKs for image classification, object detection, mask detection, face keypoints and human segmentation. Those are the quickest way to see the engine running on a real device without writing any code, and the README links them directly from the Android apps demo page.
If you need to build from source instead, the README points at the Docker environment page for the unified build environment. That page, not this article, is where the actual build commands live, and they differ per target CPU architecture and OS.
One thing to plan for: opt runs on your development machine, not on the device. The optimized model is the artifact you ship.
Where Paddle Lite is the wrong choice
The strongest limitation is the model format. Paddle Lite does not load PyTorch or TensorFlow models directly. The README routes you through X2Paddle, and every conversion step is a place where an operator can fail to map. If your team's models are PyTorch and nothing else, adding a conversion layer to reach Paddle Lite buys you nothing over a runtime that reads your format natively.
The second limitation is operator coverage on accelerated backends. The README's CI table is the honest signal here: 32-bit and 64-bit CPU builds pass on x86 Linux, ARM Linux, Android and iOS, but most NPU backends pass on exactly one or two platforms. Huawei Kirin NPU, Qualcomm QNN, Android NNAPI and MediaTek APU are Android only. Ascend NPU, Kunlunxin XPU and XTCL are x86 and ARM Linux. Cambricon MLU is x86 Linux only. If your product ships on two of those targets, you are maintaining two backends, and the CI table tells you which ones have coverage.
The third issue is release cadence. The most recent release listed is v2.14-rc from 2024-07-19, and the one before that, v2.13-rc, is from 2023-03-27. The repository is not archived and the last push was on 2026-04-27, so work continues on the develop branch, but tagged releases are infrequent and two of the three most recent carry the rc suffix. If your process requires a stable tagged release, budget for that gap.
Finally, the documentation is split between the repository and the paddlepaddle.org.cn/lite site. The README is a routing table of links. Anything about API signatures, build flags or backend configuration lives on the site, which means you cannot evaluate the project from the repository alone.
Paddle Lite compared with ONNX Runtime and TFLite
The natural alternative for on-device inference is ONNX Runtime or TensorFlow Lite, and the difference is mostly about where your model comes from.
ONNX Runtime takes ONNX as its input format. Almost every training framework exports to ONNX, so it is format-neutral in a way Paddle Lite is not. If your models come from PyTorch, ONNX Runtime is the shorter path: export, then run. With Paddle Lite you would export to PaddlePaddle format first, and the README's own answer to a PyTorch model is X2Paddle. The trade is that Paddle Lite is tuned around PaddlePaddle's operator set and its own opt pipeline, so a model that is already PaddlePaddle has fewer moving parts there.
TensorFlow Lite is closer to Paddle Lite in design intent: a converted flatbuffer model, a small runtime, delegates for GPU and NPU acceleration. The difference is the ecosystem lock. TFLite expects TensorFlow or a converter path into its format. Paddle Lite expects PaddlePaddle. Both are good at what they are built around and awkward outside it.
The hardware coverage is where Paddle Lite stands out. Its README lists backends for Kirin, Ascend, Kunlunxin, QNN, Cambricon, TIM-VX, MediaTek APU, Imagination NNA and OpenVINO. Many of those are chips that matter in the Chinese market and are less commonly covered by the other two runtimes. If your target board is one of those, that list is the reason to look here.
Licence, maintenance and upgrade cost
Paddle Lite is Apache-2.0. That is a permissive licence, so linking the runtime into a closed-source mobile application is the normal use, and the repository ships the LICENSE file at the top level. This is a factual note about the licence identifier, not legal advice; if you modify the runtime itself or redistribute it, read the licence text and your own legal review.
The maintenance picture is mixed. The repository is not archived and the last push was on 2026-04-27, so the develop branch is being touched. Tagged releases are another matter: v2.14-rc dates from 2024-07-19, v2.13-rc from 2023-03-27, and v2.12 from 2022-11-18. Anyone pinning to a release should look at that spacing before planning an upgrade cycle.
Upgrade cost is dominated by the opt step, not by the library. When you move to a newer Paddle Lite, you generally re-run opt on your models, because the optimized model is produced by a specific version of the tool. That means an upgrade is a model pipeline change, and it needs the same operator coverage check you did the first time. The prebuilt library download is the cheap part. The C++ API also means an ABI surface to track if you ship a shared library to other teams.
There is no migration or rollback procedure documented in the README. If you need one, it is not there.
Editorial conclusion
Adopt Paddle Lite if your models are already PaddlePaddle inference models and your targets are Android, iOS, embedded Linux or one of the NPUs listed in the README. Do not adopt it if your models live in PyTorch or TensorFlow and you have no reason to convert. Before committing, verify that a prebuilt library exists for your exact target, confirm that opt supports every operator in your model, and check whether your NPU backend is still listed in the CI table.
Frequently asked questions
What is Paddle Lite used for?
It is an inference engine for running PaddlePaddle models on mobile, embedded and edge devices. The README lists Android, iOS, embedded Linux, Windows, macOS and Linux hosts as targets, with C++, Java and Python APIs.
How do I install Paddle Lite?
The README recommends downloading a prebuilt library for Android, iOS, x86 or macOS first, or taking one from the GitHub releases. If you compile from source, it recommends the Docker unified build environment instead of assembling the toolchain manually.
Is Paddle Lite the same thing as PaddlePaddle?
No. PaddlePaddle is the deep learning framework that produces the inference model via save_inference_model; Paddle Lite is the separate runtime that executes that model on device after it passes through the opt tool.
Does Paddle Lite need a GPU?
No. CPU builds pass CI on x86 Linux, ARM Linux, Android and iOS in both 32-bit and 64-bit. GPU and NPU paths are separate backends such as OpenCL, Metal, QNN or NNAPI, each with its own platform coverage in the CI table.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/paddlepaddle-paddle-lite)
Community notes