CLI tool
google-ai-edge/LiteRT avatar
google-ai-edge/LiteRT

LiteRT: Google's On-Device Runtime, From .tflite to NPU Delegation

Project brief: LiteRT, successor to TensorFlow Lite. is Google's On-device framework for high-performance ML & GenAI deployment on edge platforms, via efficient conversion, runtime, and optimization.

3,459 stars464 forksC++Apache-2.0

At a glance

What is it?
LiteRT is the successor to TensorFlow Lite and the runtime layer behind Google's on-device AI stack. This review covers what the V2 Compiled Model API actually changes, how to install the CLI, and where the project is still thin.
Who is it for?
Adopt LiteRT if you are shipping inference to Android, iOS, Linux, macOS, Windows or the browser and you want one runtime with CPU, GPU and NPU paths rather than a per-vendor integration. Do not adopt it if you need a documented, stable Python inference API on desktop: the README points Python users at LiteRT-LM and the CLI, and the runtime's own bindings are C++ and Kotlin.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly C++, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What LiteRT is replacing, and who feels that most

TensorFlow Lite shipped one model format (.tflite), one interpreter, and a delegate system where the developer picked the accelerator by hand. LiteRT keeps the format and the interpreter lineage but adds a second surface: the Compiled Model API in V2. The README describes it as offering "automated accelerator selection (no explicit delegates needed), true asynchronous execution, easy NPU distribution, and highly efficient I/O buffer handling." That is the whole pitch in one line. The delegate plumbing that used to sit in your app moves into the runtime.

The audience is narrow and specific. You are shipping inference onto a device you do not control, on silicon you did not choose, and you want the same code to run on a phone CPU, a phone GPU and whatever NPU the chipset vendor supplied. If your models run in a data centre on a GPU you provisioned, LiteRT is the wrong layer and the README does not pretend otherwise. If you are already on TensorFlow Lite and your app works, the migration guide is the entry point rather than a rewrite.

How the runtime, converter and model formats fit together

The README's diagram is the clearest statement of the data flow. A PyTorch model or a Hugging Face safetensors checkpoint goes into LiteRT Torch, which emits either a .tflite file or a .litertlm file. The AI-Edge Quantizer takes that artifact and produces an optimized version of the same format. From there the paths split by format: .litertlm goes to LiteRT-LM, which the README lists with Python, C++, Kotlin, Swift and JS bindings, while the runtime itself (C++, Kotlin, JS) consumes the compiled model. Execution lands on CPU through XNNPack, GPU through ML Drift, or a supported TPU/NPU.

Two things are worth noticing. First, the generative path and the classic path are separate products with separate bindings, not one API with a flag. .litertlm is not a .tflite with extra tensors. Second, the quantizer sits between the converter and the runtime, which means the artifact you benchmark is not the artifact the converter produced. That is a normal pipeline, but it means a quantization regression and a runtime regression look identical from the outside unless you keep the pre-quantization file.

Installing litert-cli and running the first command

The README's installation section covers the CLI, not the C++ runtime. The steps assume Python 3.13 and uv. The first block creates a clean virtual environment with the interpreter seeded, then activates it. The README notes that setting the UV_INDEX_URL environment variable sometimes helps dependency resolution.

bash
uv venv --clear --python=3.13 --seed
source .venv/bin/activate

With the environment active, install the nightly package and run the help command. The package name is litert-cli-nightly and the entry point is litert.

bash
uv pip install litert-cli-nightly
litert --help

What you should see is the CLI's command list. The README links the CLI to agentic coding workflows and points at a separate repository, LiteRT-CLI, for the full command surface. If you expected a Python inference API from this install, you will not find one: the package is the command-line toolkit, and the README does not present a pip-installable Python runtime binding. For the C++ runtime, the repository root carries Bazel files (.bazelrc, .bazelversion, WORKSPACE) and a cmake_example directory, which tells you the supported build paths are Bazel and CMake rather than a package manager.

Platform coverage, and the asterisks that matter

The support table is broad: Android, iOS, Linux, macOS, Windows, Web and IoT, with CPU on every row. GPU is OpenCL and OpenGL on Android, Metal on iOS and macOS, WebGPU on Linux, macOS, Windows, Web and IoT. NPU support is where the asterisks cluster. Android lists Broadcom, Google Tensor, Intel, MediaTek and Qualcomm as shipping, with S.LSI marked coming soon. Linux and Windows list Broadcom and Intel. iOS and macOS list ANE as coming soon. Web lists WebNN as coming soon. IoT lists Raspberry Pi as coming soon.

Read that table as a procurement document, not a feature list. The Qualcomm entry links to a vendor README under litert/vendors/qualcomm, which is the pattern to expect: NPU support is per-vendor code with per-vendor documentation, wrapped by a common API. That is a real improvement over hand-wiring delegates, but it does not make the vendors interchangeable. A model that runs on Google Tensor may still fail on MediaTek, and the README does not claim otherwise. If your shipping matrix includes an S.LSI device or an iPhone where you need the ANE, the table says wait.

Where LiteRT is the wrong tool

The clearest limitation is the one the search data keeps circling: LiteRT and LiteRT-LM are not the same thing, and neither is LiteRT.js. LiteRT-LM handles the .litertlm generative path with its own bindings. LiteRT.js is the browser path over WebGPU and WASM. If you are building a desktop Python service that loads a model and returns predictions, none of the three is aimed at you, and the README's install section will leave you holding a CLI you do not need.

The second limitation is build weight. The repository root carries Bazel configuration, a WORKSPACE file, a configure script, Docker build files, and patches for third-party dependencies including flatbuffers, protobuf, sentencepiece and perfetto. That is the shape of a project compiled into other people's products, not one you drop into a script. Building the runtime from source is a real commitment, and the release cadence of six to eight weeks for stable releases, with nightly wheels in between, means you have to decide deliberately whether you track stable or nightly.

The third is that the README does not document rollback or version compatibility between a compiled model and a runtime version. The migration guide is referenced for TensorFlow Lite to V2, but nothing in the README states whether a .tflite produced for an older runtime keeps working after an upgrade. Verify that against the release notes for the version you pin.

LiteRT against ONNX Runtime and llama.cpp

ONNX Runtime and LiteRT solve overlapping problems from opposite directions. ONNX Runtime starts from a framework-neutral graph format and targets server, desktop and edge through execution providers. LiteRT starts from Google's own lineage, keeps .tflite as the classic artifact, and treats the accelerator as something the runtime selects rather than something the developer configures. If your models already exist as ONNX and your deployment targets are x86 servers with occasional edge boxes, converting to .tflite to gain automated NPU selection is work you probably do not need.

llama.cpp is the sharper comparison for the generative case. It is built around GGUF weights and a C/C++ inference path that runs on CPU first, with GPU offload as an option. LiteRT-LM consumes .litertlm, a format produced by LiteRT Torch and then optimized by the quantizer, and the README frames the target as quantized LLMs and diffusion models on-device with NPU access. So the difference is not speed, which the README does not quantify, but where the weights come from and what hardware you are allowed to assume. If your deployment is a phone with a vendor NPU, LiteRT-LM is the path with the vendor integrations already written. If your deployment is a laptop and you want to run whatever GGUF file you downloaded, llama.cpp is the format-native choice.

Licence, cadence and the cost of staying current

LiteRT is Apache-2.0. That is a permissive licence, and the practical implication is that you can ship it inside a closed product without publishing your application code, provided you keep the licence and notice files intact. It says nothing about the models you run through it: the Hugging Face LiteRT Community collections are separate artifacts under their own terms, and the README links to them as external resources rather than bundling them. Check the licence on each model independently. This is not legal advice.

The maintenance picture is straightforward. The repository is not archived, and the last push was on 2026-08-13, which is the same date as the v2.2.0 release. Releases land on a six to eight week stable cadence with nightly builds in between, and the README's build status table shows nightly wheels for Linux, macOS and Windows plus continuous builds for macOS arm64, Linux x86_64 and Windows x86_64. Upgrading means re-validating your compiled models against the new runtime, and because the README does not promise artifact compatibility across versions, budget for that re-validation rather than assuming a drop-in swap.

Editorial conclusion

Adopt LiteRT if you are shipping inference to Android, iOS, Linux, macOS, Windows or the browser and you want one runtime with CPU, GPU and NPU paths rather than a per-vendor integration. Do not adopt it if you need a documented, stable Python inference API on desktop: the README points Python users at LiteRT-LM and the CLI, and the runtime's own bindings are C++ and Kotlin. Before committing, verify on your target device that the accelerator you care about is not one of the entries marked coming soon (ANE on iOS and macOS, S.LSI on Android, WebNN in the browser, Raspberry Pi on IoT), and check that a converter path exists for your source framework, since LiteRT Torch is a separate repository.

Frequently asked questions

How does LiteRT work?

A model is converted to .tflite or .litertlm by LiteRT Torch, optionally optimized by the AI-Edge Quantizer, and then executed by the LiteRT runtime or LiteRT-LM on CPU through XNNPack, GPU through ML Drift, or a supported NPU. The V2 Compiled Model API handles accelerator selection inside the runtime instead of requiring explicit delegates.

Is TensorFlow Lite deprecated?

The README describes LiteRT as continuing the legacy of TensorFlow Lite and provides a migration guide for upgrading from TensorFlow Lite or LiteRT V1.x to LiteRT V2.x. The repository does not use the word deprecated, so treat this as an upgrade path rather than a removal notice.

Which models are supported by LiteRT-LM?

LiteRT-LM consumes .litertlm files, which the README associates with quantized LLMs and diffusion models produced through LiteRT Torch and the AI-Edge Quantizer. The README also lists recently added model families on Hugging Face, including Gemma, ASR and image classification collections.

What is the difference between LiteRT and LiteRT-LM?

LiteRT is the runtime, with C++, Kotlin and JS bindings, that executes compiled models on CPU, GPU or NPU. LiteRT-LM is the separate path for the .litertlm generative format and lists Python, C++, Kotlin, Swift and JS bindings.

How do I install LiteRT?

The README's install section uses uv to create a Python 3.13 virtual environment, installs litert-cli-nightly with uv pip, and runs litert --help. That installs the CLI; the repository provides Bazel and CMake build paths for the C++ runtime itself.

What is LiteRT.js?

LiteRT.js is the browser path, running client-side ML through WebGPU and WASM. The platform table lists WebGPU as the GPU API on the Web row and marks WebNN as coming soon.

Official sources

  1. Official documentation
  2. Official README
  3. Project repository
  4. Release notes
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/google-ai-edge-litert.svg)](https://hysenlabs.com/projects/google-ai-edge-litert)