LiteRT-LM: Google's C++ Runtime for On-Device LLM Inference
LiteRT-LM is Google's production-ready, high-performance, open-source inference framework for deploying Large Language Models on edge devices.
At a glance
- What is it?
- LiteRT-LM is an Apache-2.0 orchestration layer that runs Gemma, Llama, Phi-4 and Qwen models through LiteRT on Android, iOS, Web, desktop and IoT targets. It installs as a CLI in one uv command, but the C++ core and its Bazel build are where the real adoption cost sits.
- Who is it for?
- Adopt LiteRT-LM if you are shipping a mobile, browser or wearable product that needs Gemma-class models running locally with GPU or NPU acceleration, and you can accept a Bazel and CMake build with pinned patches. Do not adopt it if you only need batch inference on a server GPU, or if you want a runtime that treats every model format as a first-class citizen.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 4 days ago.
- What is it written in?
- Mainly C++, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem LiteRT-LM solves, and who it is actually for
Running a language model on a phone is not the same problem as running one on a server. The weights have to fit in memory that other applications share, the accelerator is a GPU or NPU with a vendor-specific driver, and the process can be killed at any moment by the operating system. LiteRT-LM is Google's answer to that specific set of constraints. The README describes it as an "orchestration layer to run LLMs with LiteRT", which is a precise word choice: LiteRT handles the tensor execution and hardware delegates, and LiteRT-LM sits above it handling model loading, tokenization, sampling, multi-modality and tool calling.
The audience is narrow and identifiable. It is the engineer building a chat or agent feature into an Android or iOS application, a browser page, or a device like a Raspberry Pi. The README states that LiteRT-LM powers on-device GenAI in Chrome, Chromebook Plus and Pixel Watch, so the framework has been exercised in shipping Google products rather than only in a research repository. That matters more than any feature list, because most edge inference stacks break on the last mile of platform integration rather than on the math.
The project is written primarily in C++ and carries an Apache-2.0 licence. There is no hosted service in the loop: the model file is downloaded or bundled, and inference happens in the local process. If your workload is a nightly batch job over a million documents, this is the wrong tool and the rest of this article will not change that.
How the runtime is put together: LiteRT, delegates and the Rust parser layer
The repository layout tells you more about the architecture than the README does. Alongside the C++ sources there are directories for c/, python/, js/, kotlin/, rust/, android_ndk_env.bzl, and a set of BUILD files that vendor third-party components directly into the build graph: BUILD.antlr4, BUILD.espeak_ng, BUILD.llguidance, BUILD.miniaudio, BUILD.minizip, BUILD.minja, BUILD.nanobind_json, BUILD.sentencepiece, BUILD.stb, BUILD.tokenizers_cpp. The presence of espeak_ng points at text-to-speech support, miniaudio at audio capture, and llguidance at constrained decoding.
The Rust side is not incidental. Cargo.toml declares a static library named litert_lm_deps that pulls in llguidance at a pinned version 1.3.0 with the lark, rayon and ahash features, minijinja pinned at 2.14.0 for chat template rendering, and tokenizers pinned at 0.21.0 with the onig feature. It also declares five local parser crates: antlr_fc_tool_call_parser, antlr_python_tool_call_parser, json_parser, python_parser and fc_parser. That is the tool-calling machinery. When the model emits a function call, these parsers turn the generated text into a structured call, and the comment at the top of Cargo.toml shows the lockfile is regenerated through CARGO_BAZEL_REPIN=1 bazel sync --only=crate_index.
Acceleration is delegated. The v0.16.0 release notes add an experimental YNNPACK delegate, described as enabled for linux arm64 builds in the CLI and Python API. YNNPACK is a CPU-side backend from the XNNPACK project, so it is a different path from the GPU and NPU delegates the README advertises under hardware acceleration. The word experimental in the release notes should be taken literally: it is not presented as the default path on any platform.
Installing LiteRT-LM and running a first prompt from the terminal
The README offers a no-code path through uv, a Python tool installer. This is the fastest way to see whether the runtime behaves on your machine before you touch the C++ build.
uv tool install litert-lm
litert-lm run \
--from-huggingface-repo=google/gemma-3n-E2B-it-litert-lm \
gemma-3n-E2B-it-int4 \
--prompt="What is the capital of France?"The first command installs the CLI as a uv tool. The second downloads the quantized Gemma 3n model from Hugging Face and runs a single prompt. The --from-huggingface-repo flag is what pulls the weights, and the positional argument names the model file inside that repository. Expect the download to dominate the first run.
The README's newer example shows the same CLI with additional flags, including speculative decoding and an explicit backend selection. Note that the repository name and the model filename in this snippet do not match each other in the README as published, so treat the flags as the reliable part and confirm the exact model identifier against the Hugging Face repository before you script it.
litert-lm run \
--from-huggingface-repo=litert-community/gemma-4-E2B-it-litert-lm \
gemma-4-E4B-it.litertlm \
--backend=gpu \
--enable-speculative-decoding=true \
--prompt="What is the capital of France?"If you are integrating rather than experimenting, v0.16.0 added the first versioned C API shared library prebuilts for all supported platforms. The README's stated purpose is to let you integrate LiteRT-LM natively and build language bindings without compiling shared libraries yourself. The archive is published on the release page as litert_lm_c_api-0.1.0.zip. The README does not document the ABI stability policy for that C API, and the 0.1.0 version number suggests you should not assume one.
Where LiteRT-LM gets expensive: the build, the patches and the platform matrix
The top-level entries include PATCH.abseil, PATCH.llguidance, PATCH.llguidance_grammar, PATCH.llguidance_numeric, PATCH.llguidance_parser, PATCH.llguidance_perf, PATCH.llguidance_regexvec, PATCH.minja, PATCH.nanobind_json, PATCH.rules_rust, PATCH.sentencepiece, PATCH.skia, PATCH.skia_user_config, PATCH.tensorflow and PATCH.toktrie. That is fifteen patched dependencies, and it is the single clearest signal about what building from source involves. Each patch has to keep applying as the upstream project moves. Upgrading LiteRT-LM is therefore not just bumping a version: it is re-resolving a dependency graph that includes TensorFlow, Skia, abseil and a fork of llguidance.
There is a second cost hidden in the platform matrix. The README claims cross-platform support across Android, iOS, Web, Desktop and IoT, with hardware acceleration through GPU and NPU accelerators. Those are different code paths with different failure modes, and the release notes only call out YNNPACK for linux arm64. Nothing in the README states which delegates are available on which platform, and the README does not document a rollback or fallback procedure when a delegate fails to initialise.
The third constraint is the model format. Model files carry the .litertlm extension in the CLI examples, which means the weights have to be converted into a LiteRT-LM-specific artifact rather than consumed as GGUF. The README lists Gemma, Llama, Phi-4 and Qwen under broad model support, but it does not document which of those have published .litertlm conversions or who maintains them. Verify that before you design around a specific model.
LiteRT-LM against llama.cpp and Ollama: different jobs, different trade-offs
The comparison people reach for most often is llama.cpp, and the difference is structural rather than a matter of speed. llama.cpp is a portable inference engine built around the GGUF format, with a large ecosystem of community quantizations and a CPU-first execution model that also supports GPU offload. LiteRT-LM is a layer over LiteRT that assumes a mobile or embedded target, ships vendor delegates for GPU and NPU, and expects models in its own .litertlm container. If your distribution channel is a phone app store or a browser, the delegate story is the reason to pick LiteRT-LM. If your distribution channel is a tarball on a Linux server, GGUF and llama.cpp give you a far wider choice of weights today.
Ollama is a different kind of difference again. It is a model-serving product with a daemon, a registry and an HTTP API, aimed at developers who want a local endpoint rather than an embedded library. LiteRT-LM does have a serve command in its CLI, judging by the search interest in it, but the README's quick start is built around one-shot prompts and the project's stated purpose is on-device deployment. Using it as a desktop model server means discarding the parts of the design that make it interesting.
The server-side comparison is even cleaner. vLLM targets throughput on data-center GPUs with paged attention and continuous batching. LiteRT-LM is not competing there and its README makes no such claim. Multi-modality, tool calling and speculative decoding through MTP drafters are the features it advertises, and all three are about interactive latency on a device that is also running a user interface.
Maintenance, releases and what the Apache-2.0 licence means for shipping
The release cadence is fast. v0.15.0 landed on 2026-08-04, v0.16.0 on 2026-08-11, and v0.16.1 on 2026-08-18. Three releases in fifteen days, with the last push to main on 2026-08-18. The repository is not archived. A cadence that tight on a 0.x version line means the API surface is still moving, and the README itself links v0.16.0 back to v0.15.0 as a "quick follow up", which is a reasonable description of what a patch-level release should be.
For planning purposes, treat the version number as the contract. A 0.x project can change the C API or the CLI flags between minor releases, and the C API prebuilt is at 0.1.0. If you are embedding the shared library in a shipped application, pin the exact prebuilt archive and the exact model file, because neither the README nor the release notes promise compatibility across minor versions.
The licence is Apache-2.0, which is permissive and includes an explicit patent grant. That is the same licence family Google uses across many of its open source projects, and it is generally compatible with proprietary application code. The repository also vendors third-party components under their own licences, including Skia, sentencepiece, abseil, miniaudio, minizip and espeak_ng, plus llguidance and the Rust crates listed in Cargo.toml. Those carry their own terms, and the patch files exist precisely because the vendored copies are modified. Anyone shipping a binary should review the licence of each vendored component rather than assuming the Apache-2.0 file at the repository root covers the whole build. That is a statement about the repository layout, not legal advice.
Editorial conclusion
Adopt LiteRT-LM if you are shipping a mobile, browser or wearable product that needs Gemma-class models running locally with GPU or NPU acceleration, and you can accept a Bazel and CMake build with pinned patches. Do not adopt it if you only need batch inference on a server GPU, or if you want a runtime that treats every model format as a first-class citizen. Before committing, verify that your target platform has a supported delegate, check whether the C API prebuilt covers your architecture, and confirm which model family you intend to ship, because the README's supported-model list is broader than the models the CLI examples actually demonstrate.
Frequently asked questions
What is LiteRT-LM?
It is Google's orchestration layer for running large language models with LiteRT, aimed at edge devices rather than servers. The README describes it as production-ready and cross-platform, and notes it powers on-device GenAI in Chrome, Chromebook Plus and Pixel Watch.
Which models are supported by LiteRT-LM?
The README lists Gemma, Llama, Phi-4 and Qwen under broad model support. The CLI examples in the README use Gemma 3n and Gemma 4 model identifiers from Hugging Face repositories.
Is LiteRT open source?
Yes. The repository carries an Apache-2.0 licence, and the primary language is C++. The build also vendors third-party components under their own licences.
How do I install LiteRT-LM?
The README's quick start installs the CLI with uv tool install litert-lm, then runs a prompt through litert-lm run with a Hugging Face repository flag. For native integration, v0.16.0 added versioned C API shared library prebuilts for all supported platforms.
What is Google AI Edge Gallery?
It is a separate Google AI Edge application that the README links to for running models immediately on a device. It is distributed through Google Play and the App Store, and the README mentions on-device function calling in it powered by LiteRT-LM tool use APIs.
What is LiteRT-LM for on-device AI?
It handles model loading, tokenization, sampling, multi-modality and tool calling above LiteRT, which performs the tensor execution on GPU or NPU delegates. The README lists cross-platform support, hardware acceleration, vision and audio inputs, and function calling as its key features.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/google-ai-edge-litert-lm)