LiteRT Torch: getting PyTorch models onto a device as .tflite
Support PyTorch model conversion with LiteRT.
At a glance
- What is it?
- A Google AI Edge library that converts PyTorch modules into .tflite flatbuffers for on-device inference, with a beta converter aimed at vision models and an alpha generative API aimed at transformers.
- Who is it for?
- LiteRT Torch earns its place when a model has to run on Android, iOS or a device with an NPU and your team already thinks in PyTorch, because it removes the detour through a second graph format and lands on a runtime Google already documents. It is the wrong tool if you need macOS or Windows conversion, if the model never reaches a device, or if you cannot accept a converter the project itself calls Beta while its generative path is Alpha.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 18 days ago.
- What is it written in?
- Mainly Jupyter Notebook, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 20, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Converting a PyTorch model into a .tflite flatbuffer
The whole conversion path in the README fits on one screen. It loads a pretrained ResNet 18 through torchvision, captures a sample input under torch.no_grad, calls litert_torch.convert on the module in evaluation mode, and writes the result to a file:
import torch
import torchvision
import litert_torch
# Use resnet18 with pre-trained weights.
resnet18 = torchvision.models.resnet18(torchvision.models.ResNet18_Weights.IMAGENET1K_V1)
with torch.no_grad():
sample_inputs = (torch.randn(1, 3, 224, 224),)
# Convert and serialize PyTorch model to a .tflite flatbuffer. Note that we
# are setting the model to evaluation mode prior to conversion.
edge_model = litert_torch.convert(resnet18.eval(), sample_inputs)
edge_model.export("resnet18.tflite")Two details in that snippet are load-bearing. The first is that sample_inputs is a tuple, which is how torch.export expects its example arguments, and the second is the call to .eval() before conversion, so dropout and batch norm folding behave the way they will at inference. The result is a single .tflite flatbuffer rather than a checkpoint directory, and the project describes the point of that format plainly: it is what LiteRT runs, for applications on Android, iOS and IoT that keep inference entirely on the device.
The repository's primary language is listed as Jupyter Notebook, which matches how this project expects to be approached. Rather than only a README, the tree carries a docs directory whose first entry is a notebook walkthrough at docs/pytorch_converter/getting_started.ipynb that the README says can be tried out in Google Colab. Everything deeper than the conversion call lives behind that notebook and behind docs/pytorch_converter/README.md.
Why the PyTorch path is built on torch.export
LiteRT Torch does not invent its own graph capture. The README says it builds on top of torch.export() and that it aims for good coverage of Core ATen operators. That choice has visible consequences, and it is the single most useful thing to understand before you start converting your own models.
torch.export produces a program representation rather than a trace tied to one set of concrete values, so shapes and control flow are recorded in a form a compiler can reason about. A converter built on it can therefore do operator-level lowering rather than freezing a graph around whatever numbers happened to be in your sample tensor. It also means the boundary where conversion fails is an operator boundary. If your model reaches for something outside the covered ATen set, you get a conversion failure naming that operator rather than a silently wrong result, which is the better of the two failure modes.
The project treats that coverage as something that moves. The repository has a nightly model coverage workflow alongside unit tests, a generative API build and a nightly release job, and a ci directory with a script that rewrites the version badges in the README to track the latest nightly builds. In other words, the set of models known to convert is measured in CI on every run, and the README's claim of broad CPU coverage is a snapshot of a process that runs nightly rather than a promise about your architecture.
A beta converter sitting next to an alpha generative API
One PyPi package ships two very different things at two different maturity levels, and the README is explicit about which is which. The PyTorch converter is a Beta release. The Generative API is an Alpha release. Both live in the same package, which means a `pip install` gives you both, and a reader has to work out from the docs which half they are actually using.
The Generative API is described as a Torch native library for authoring mobile-optimized PyTorch transformer models, with conversion to LiteRT-LM models as the destination. It currently supports CPU and GPU, with NPU support planned. That word, planned, is doing real work in that sentence: anyone whose requirement is specifically an NPU path is reading a roadmap, not a shipped capability. The stated longer-term direction is collaboration with the PyTorch community so frequently used transformer abstractions can be supported directly, without reauthoring, which is an honest admission that today you author models with this library rather than drop an existing Hugging Face model in unchanged.
The packaging story has its own tool. When you have a converted .tflite file and a tokenizer, the tip in the README is to bundle them into a deployment-ready .litertlm container using the litert-lm-builder CLI, which arrives with the nightly package. The stable requirements list already names litert-lm and litert-lm-builder as dependencies, and the v0.9.3 release notes added a LitertLmBundle package for object-oriented container management, so this part of the pipeline is being treated as a first-class surface rather than a script.
What the v0.9.x releases actually changed
Three releases are on record and they show where the effort is going. v0.9.1 shipped on 2026-05-19 with an empty description, which tells you nothing except that the tag exists. v0.9.3 on 2026-08-04 is a long list: Gemma 4 unified model support, a MoE layer in the Generative API with FP32 and INT8 lowerings, export support for speech and ASR models including Whisper, Parakeet and Qwen3-ASR, ConvTranspose1d lowering with output_padding, support for torch.uint8 tensors, a BPE tokenizer patch with a Metaspace pre-tokenizer, and an experimental LiteRT-LM NPU compiler tool.
v0.9.4 on 2026-08-24 is narrower and reads like a response to real export failures. It adds Gemma 4 architecture export, ASR models in the export_hf pipeline, dynamic context length support through --enable_gpu_dynamic_prefill and --enable_gpu_dynamic_cache, automatic expansion of stop token prefixes with inference of extra stop tokens from the chat template, and LiteRT-LM ExecutorMetadata for sliding window attention and non-SDPA attention shapes.
Two things stand out. Split-cache export appears in both releases, first as a feature and then reworked so a single valid_mask input tensor is used, which is the signature of a problem being iterated on rather than designed once. And the density of experimental flags, including --experimental_use_mixed_precision, --externalize_embedder and --experimental_transpile_chat_template_for_minijinja, means the path to a working LLM export is still partly a negotiation with the tool. The repository's last push was on 2026-09-18, so this line is moving.
Linux, Python 3.11 and pinned nightly dependencies
The installation section is specific about its environment, and the requirements are worth reading literally because the README and the repository files do not say quite the same thing. Python 3.10 or newer, with 3.11 strongly recommended. Linux as the operating system. PyTorch 2.4.0 or newer, and a TensorFlow nightly build, which the README badges as tf-nightly latest.
python3.11 -m venv --prompt litert-torch venv
source venv/bin/activateThe stable install pulls three things together, and the reasons are spelled out in the sentence above it:
pip install litert-torch torchvision ai-edge-literttorchvision is there to run the quickstart example, and ai-edge-litert is there for CLI benchmarking tools. The nightly variant swaps in litert-torch-nightly and adds litert-cli-nightly. Inside the repository, requirements.txt is far more specific than the badge: torch is pinned to 2.12.0, torchvision to 0.27.0, torchaudio to 2.11.0, with ai-edge-litert-nightly carrying the model-utils extra, ai-edge-quantizer-nightly and litert-converter at a dev version, then jax on CPU, transformers, sentencepiece, safetensors, kagglehub, tabulate, fire and rich. That gap between a README minimum of 2.4.0 and a pinned 2.12.0 is the kind of thing that decides whether your install works on the first try.
Two build systems coexist here. setup.py is the packaging entry point and reads its version out of litert_torch/version.py with a regular expression, appending .dev plus the date from the NIGHTLY_RELEASE_DATE environment variable and renaming the distribution with a -nightly suffix when that variable is set. Alongside it sit MODULE.bazel, WORKSPACE, a .bazelrc, a pinned .bazelversion and a bazel directory, plus run_tests.sh, format.sh and test/. Contributors are pointed at CONTRIBUTING.md, and the repository even ships a SKILL.md.
Where the README stops and the docs take over
Read this README as a signpost rather than a manual. It gives you one conversion example, one installation recipe and a build status table with four badges, and then it hands you off: technical details for the converter are in docs/pytorch_converter/README.md, the walkthrough is the notebook in the same directory, and generative documentation sits under the litert_torch/generative path inside the package.
The runtime side is documented elsewhere too, and it matters for planning. Once a model is exported, the README points at the LiteRT compiled model API for native C++ or Java deployments so you can hit CPU, GPU and NPU paths, and the converted generative models are run through the separate LiteRT-LM repository. Questions go to a GitHub issue chooser rather than a discussion forum, which tells you the project's support model is issue triage.
The honest summary is that LiteRT Torch covers the authoring and conversion half of on-device inference well and leaves the serving half to a runtime you adopt separately. If your work is a vision model with ordinary ATen operators, the README plus the notebook is nearly enough to get through conversion today. If your work is a transformer, budget for reading the generative documentation and for expecting the flags to move, because that path is Alpha and its release notes read like a running list of the cases that still need handling.
Editorial conclusion
LiteRT Torch earns its place when a model has to run on Android, iOS or a device with an NPU and your team already thinks in PyTorch, because it removes the detour through a second graph format and lands on a runtime Google already documents. It is the wrong tool if you need macOS or Windows conversion, if the model never reaches a device, or if you cannot accept a converter the project itself calls Beta while its generative path is Alpha. Three things are worth checking before committing: that every operator in your graph falls inside the covered Core ATen set, that the torch version pinned in requirements.txt matches your training environment, and that the accelerator you care about is real on your hardware rather than described as coming later. Start from docs/pytorch_converter/getting_started.ipynb, then read litert_torch/generative if a transformer is what you are shipping.
Frequently asked questions
What is LiteRT, in the context of a PyTorch converter?
LiteRT is the runtime that executes what this project produces. LiteRT Torch converts a PyTorch module into a .tflite flatbuffer, and that file is then run with LiteRT for applications on Android, iOS and IoT that keep inference on the device. The README links the runtime documentation on the Google AI Edge developer site rather than restating it here.
Is LiteRT open source?
The repository this library lives in, litert-torch, is released under Apache-2.0. For LiteRT itself, the README links out to the Google AI Edge documentation for the runtime instead of stating licensing terms, so the runtime's own licence has to be checked at that destination.
Which Python and PyTorch versions does LiteRT Torch need?
The README asks for Python 3.10 or newer with 3.11 strongly recommended, Linux, PyTorch 2.4.0 or newer and a TensorFlow nightly build. The repository's own requirements.txt pins more tightly, at torch 2.12.0 and torchvision 0.27.0, so match your training environment to that pin rather than to the README minimum.
How mature is the PyTorch converter compared with the generative API?
They ship in one PyPi package at two different levels. The PyTorch converter is a Beta release with broad CPU coverage and initial GPU and NPU support, while the Generative API is an Alpha release that currently covers CPU and GPU for transformer and LLM models, with NPU support described as planned.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/google-ai-edge-litert-torch)