ExecuTorch: On-Device AI Inference for PyTorch Models
On-device AI across mobile, embedded and edge for PyTorch
At a glance
- What is it?
- ExecuTorch is PyTorch's open-source runtime for deploying neural network models on mobile phones, embedded systems, and microcontrollers. It extends the torch.export pipeline with backend-specific compilation and a C++ runtime that can be trimmed to only the operators a deployment actually uses.
- Who is it for?
- Engineers with existing PyTorch models who need production edge deployment for iOS, Android, or microcontrollers will find ExecuTorch the most direct path, because the export flow stays inside the PyTorch toolchain without a separate format conversion. Those running GPU-only cloud inference have no reason to adopt it.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository last received commits 1 day ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 28, 2026, and from our analysis. They are not legal advice.
DEEP OPEN-SOURCE ANALYSIS
Deploying PyTorch Models Without a Cloud Connection
ExecuTorch addresses the gap between training a model in PyTorch and running it on a phone, a wearable, or a microcontroller without a network call. The README describes it as powering on-device AI across Instagram, WhatsApp, Facebook, and Messenger, as well as Meta Quest and Ray-Ban Meta devices. The core problem it solves is not just size compression but full lifecycle management: a consistent export format, per-target hardware delegation, a versioned binary contract, and a C++ runtime small enough for embedded deployment.
The primary audience is ML engineers who already have a PyTorch model and need to ship it to a constrained device. It is not an automated model optimization service. The engineer controls each step of the export and backend selection, which means understanding operator coverage on the target becomes the engineer's responsibility.
The torch.export-to-.pte Pipeline
The pipeline starts with a standard PyTorch model. The README describes the flow as: capture the model with torch.export, optimize it for the target hardware, and then run it through C++, Python, Swift or Objective-C, Kotlin or Java, or JavaScript APIs.
torch.export produces an ExportedProgram, a graph representation that preserves PyTorch program metadata and source mappings. The next step is partitioning: graph regions that a backend supports are delegated to that backend, while unsupported regions remain as portable CPU kernels. The output is a .pte file, one per target, containing the compiled operator kernels and the delegation metadata. The README makes the one-source-many-targets model explicit: 'reuse the PyTorch model and export flow, while producing a backend-specific .pte for each target that needs hardware specialization.'
Selective build then reduces the runtime footprint. The build system links only the operators, kernels, and delegates the deployment actually uses. The documentation calls this the 'right-size' runtime property.
Installing ExecuTorch and Exporting a First Model
ExecuTorch publishes stable Python wheels for Linux x86-64 and AArch64, macOS arm64, and Windows x86-64. Install the latest stable release into a Python 3.10 to 3.14 environment:
pip install executorchTo use the newest features from the main branch before a stable release, install a nightly build. The README notes that the torch package must be listed explicitly because nightly ExecuTorch wheels do not declare it as a dependency:
pip install --upgrade --pre executorch torch --extra-index-url https://download.pytorch.org/whl/nightly/cpuSome backend export tools require optional extras. For Ethos-U ahead-of-time export, the README gives:
pip install 'executorch[ethos_u]'For Android, iOS, and embedded targets, the Quick Start documentation covers additional platform-specific setup steps. Embedded toolchains, simulators, and target runtimes are installed separately and are not bundled with the wheel.
The Makefile in the repository shows the per-model, per-backend build convention. For example:
make voxtral-cuda
make llama-cpu
make whisper-metalEach of these targets builds the ExecuTorch core libraries with the selected backend and the model-specific runner. A new model does not have to appear in the Makefile: the README directs users to the standard export guide or the closest model example in the examples/ directory.
Backend Coverage: Qualcomm, ARM, MediaTek, and the Portable Fallback
The examples/ directory lists backends by vendor: apple/, arm/, cadence/, espressif/, mediatek/, nxp/, qualcomm/, samsung/, and riscv/. For embedded targets specifically, the README lists Cortex-M CPU kernels, Ethos-U and NXP NPUs, Cadence DSPs, and workflows for Zephyr, Arduino, and Raspberry Pi Pico.
When a graph region cannot be delegated to the selected backend because an operator is not supported, ExecuTorch falls back to portable CPU kernels. That fallback exists by design, but it carries a performance cost that must be measured on the target device. The README is explicit that model examples are 'starting points, not a compatibility list,' and that deployment requires validating numerical accuracy, memory use, and performance on the specific target.
On Linux with an NVIDIA GPU, the installation note in the README warns that using the nightly CPU index can silently replace a CUDA-enabled PyTorch installation. The CUDA wheels are published for Linux only.
Where ExecuTorch Falls Short
ExecuTorch depends on torch.export successfully capturing the model. Models with dynamic control flow that export cannot trace, such as those using Python conditionals that depend on tensor values at trace time, require rewriting before export. The README does not document a workaround for this case.
Operator coverage varies by backend. A model that runs without issues on the portable CPU kernel may fail or produce different numerical results on a specific NPU. The README instructs users to validate numerical accuracy on the target before treating a deployment as production-ready, but it gives no automated testing tool for this in the README itself. The devtools/ directory and ETDump and ETRecord are available for post-deployment debugging, not pre-deployment validation.
The prebuilt wheel covers a limited set of host platforms. Building from source is required for other host configurations. For iOS and Android native integration, the documentation describes prebuilt C++ libraries and CMake packages as an alternative to building from source, but these are separate from the pip installation.
ExecuTorch vs ONNX Runtime and TFLite
The most common alternative for on-device inference is ONNX Runtime. ONNX Runtime is a cross-platform inference engine that consumes the ONNX format, which requires an explicit torch.onnx.export step before ExecuTorch's torch.export-based workflow. ONNX Runtime supports a broader set of ML frameworks as inputs, which is an advantage for teams with models trained in frameworks other than PyTorch. ExecuTorch's advantage is that it stays inside the PyTorch toolchain and preserves program metadata and source mappings that torch.onnx.export does not carry forward.
TFLite is the inference runtime for the TensorFlow ecosystem. It uses the .tflite format, which requires conversion from TensorFlow or PyTorch via third-party converters. Like ONNX Runtime, TFLite is framework-agnostic, but it has no direct support for torch.export semantics. ExecuTorch is the better choice when the model will remain in active PyTorch development, because the export flow does not require maintaining a separate conversion step as the model changes.
TorchScript, the older PyTorch serialization format, is effectively superseded by ExecuTorch for mobile deployments. TorchScript targeted mobile use cases through torch.jit.trace and torch.jit.script, but the README positions torch.export as the current standard. The key difference is that ExecuTorch's .pte has a formal versioned compatibility guarantee: a .pte created with stable APIs is guaranteed to load and execute for at least one following non-patch runtime release.
Deployment Contract, License, and Upgrade Cost
ExecuTorch releases follow semantic versioning. The pyproject.toml lists it as Production/Stable on PyPI. The .pte compatibility contract, described in runtime/COMPATIBILITY.md, guarantees that a binary produced with stable APIs will load on at least the next non-patch runtime version. This limits the downside of updating the runtime in isolation, but teams using nightly or main features do not have that guarantee.
The license is BSD-3-Clause, as declared in pyproject.toml. This permits use in commercial products without requiring source disclosure. The last push to the repository was on 2026-09-17, and version v1.5.1 was released on 2026-09-23.
Python 3.10 through 3.14 are the supported host environments. The pyproject.toml lists cmake 3.26 to 4.0 as a build dependency. Teams building from source for embedded targets must install and maintain the appropriate vendor toolchain separately, which adds ongoing setup cost beyond the Python package itself.
Editorial conclusion
Engineers with existing PyTorch models who need production edge deployment for iOS, Android, or microcontrollers will find ExecuTorch the most direct path, because the export flow stays inside the PyTorch toolchain without a separate format conversion. Those running GPU-only cloud inference have no reason to adopt it. Before committing to a target backend, run torch.export on your model and confirm operator coverage against that backend's supported set. Unsupported operators fall back to the portable CPU kernel, but that fallback must be measured for latency and memory on the actual hardware, not assumed.
Frequently asked questions
What is ExecuTorch?
ExecuTorch is PyTorch's open-source stack for running AI models locally on phones, wearables, laptops, browsers, embedded systems, and microcontrollers. It takes a PyTorch model through torch.export, compiles it into a backend-specific .pte file, and executes it through a C++ runtime with optional language bindings for Swift, Kotlin, and JavaScript.
How do I install ExecuTorch?
Run pip install executorch in a Python 3.10 to 3.14 environment to install the latest stable release. For nightly builds, use pip install --upgrade --pre executorch torch --extra-index-url https://download.pytorch.org/whl/nightly/cpu, adding torch explicitly because nightly ExecuTorch wheels do not declare it as a dependency.
How does ExecuTorch compare to ONNX Runtime?
ONNX Runtime accepts the ONNX format and supports multiple ML frameworks as inputs. ExecuTorch works directly from torch.export and preserves PyTorch program metadata and source mappings that a torch.onnx.export step discards. ExecuTorch is the better fit when the model will remain in active PyTorch development and the team wants to avoid a separate format conversion step.
How does ExecuTorch compare to TFLite?
TFLite is the inference runtime for the TensorFlow ecosystem and uses the .tflite format, which requires conversion from TensorFlow or PyTorch. ExecuTorch stays inside the PyTorch toolchain with torch.export, carries a versioned .pte compatibility contract, and offers direct integration with PyTorch's debugging and profiling tools through ETDump and ETRecord.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/pytorch-executorch)
Community notes