# Google AI Edge Gallery: running open-source LLMs on Android, iOS and macOS

> Google AI Edge Gallery is an Apache-2.0 Kotlin app that runs open-source language models on your phone or Mac, with no server in the loop. The model downloads are the real cost, and the app is still labelled an experimental Beta.

**google-ai-edge/gallery** — Project brief: A gallery of on-device ML and generative-AI demos that lets users run curated local models directly.

- Repository: https://github.com/google-ai-edge/gallery
- Stars: 24,805 · Forks: 2,699
- Language: Kotlin
- License: Apache-2.0
- Published: 2026-08-08 · Updated: 2026-08-18 · Language: en
- Canonical page: https://hysenlabs.com/projects/google-ai-edge-gallery

## What Google AI Edge Gallery is actually for

The repository describes itself as a gallery that showcases on-device ML and GenAI use cases and lets people try models locally. That framing matters more than it first appears. This is not a framework you import into another app, and it is not an inference server. It is a finished application whose job is to put a model on your hardware and give you surfaces to poke at it: chat, image input, audio transcription, a prompt workbench, a benchmark runner.

The intended audience is anyone who wants to know how a given open-source model performs before building around it. The README lists the model management tile as a place to download models from a list, load custom models, and run benchmark tests to understand how each model performs on your specific hardware. That last clause is the whole point. Published model cards describe datacenter behaviour. A phone is a different machine, with different memory ceilings and thermal limits, and the only way to learn what a model does there is to run it there.

The secondary audience is privacy-driven. The README states that all model inferences happen directly on the device hardware and that no internet is required, which means prompts and images never leave the handset. If your constraint is that data cannot cross a network boundary, that property is the reason to look at this project rather than a hosted API.

## How the app runs models: LiteRT, LiteRT-LM and the allowlist

The technology stack is named in the README: Google AI Edge for the core on-device ML APIs and tools, LiteRT as the lightweight runtime for model execution, and Hugging Face integration for model discovery and download. The linked LiteRT-LM repository is where the language-model side of that runtime lives.

The repository layout tells you a little more about how model choice is governed. There is a top-level model_allowlist.json alongside a model_allowlists/ directory, plus skills/ and mcp/ directories. The allowlist is the mechanism that decides which models the app will accept, which is a reasonable design for a Beta: it keeps users from pointing the app at a format the runtime cannot execute and then filing a bug. The trade-off is that a model you want may simply not be on the list, and the README does not describe the process for getting one added.

The skills/ and mcp/ directories correspond to the Agent Skills feature. The README says skills can be loaded modularly from a URL, and that community contributions are browsed on GitHub Discussions. A separate Function_Calling_Guide.md sits at the top level. Between the allowlist, the skills directory and the function-calling guide, the repository is structured so that model capability and tool capability are extended through files rather than through a rebuild.

Data flow is straightforward and entirely local. You download a model through the app, the runtime loads it, and inference happens on the device. The README frames the absence of a network round trip as the privacy guarantee, and it is also the performance story: no request latency, but every token is bounded by your phone's compute.

## Installing Google AI Edge Gallery and running a first prompt

The README gives two installation paths. The primary one is the app stores: Google Play under the package id com.google.ai.edge.gallery, or the App Store listing. The README states the OS requirement as Android 12 and up and iOS 17 and up, so check that before anything else.

For devices without Google Play access, the README points at the latest release page and says to install the APK from there. That is the path for a corporate handset or a device outside Play coverage. The README does not print the APK filename, so check the release page for the asset name before installing. The app also ships a macOS build, published as a DMG and linked from the README as GoogleAIEdgeGallery-0.1.0.dmg. Note the version mismatch: the macOS link is 0.1.0 while the Android and iOS releases are at 1.0.18. Treat the desktop build as earlier in its life.

Once installed, the first real use is a model download followed by a chat. Open the Model Management tile, pick a model from the list, and let it download. Then open AI Chat and send a prompt. According to the README, the Thinking Mode toggle shows the model's step-by-step reasoning, and it currently works with supported models starting with the Gemma 4 family, so if that toggle appears inert, the model is the likely reason. For a more controlled first test, Prompt Lab gives single-turn control over parameters such as temperature and top-k.

For anything beyond the store install, the README directs readers to the Project Wiki for detailed installation instructions, including for corporate devices, and a full user guide. The wiki is where the gaps in the README are meant to be filled.

## Thinking Mode, Agent Skills and the features that carry caveats

Several features come with conditions attached in the README, and those conditions are worth reading before you plan around them.

Thinking Mode is explicitly scoped: it currently works with supported models, starting with the Gemma 4 family. If you load a custom model outside that set, expect the toggle to do nothing useful. The README does not list which other models are supported.

Agent Skills is the most interesting part of the project and the least specified. The README describes augmenting model capabilities with tools such as Wikipedia for fact-grounding, interactive maps, and visual summary cards, and says you can load modular skills from a URL. What the README does not document is the skill format itself, how a loaded skill is sandboxed, or what happens when a skill's remote endpoint is unreachable. The Function_Calling_Guide.md in the repository is presumably where that is answered, but the README does not say so.

Mobile Actions and Tiny Garden are both described as built on a finetune of FunctionGemma 270m. A 270m model is small by design, and the README presents these as offline device controls and a mini-game rather than as general-purpose assistants. That is a fair presentation, but it also means these features demonstrate a pattern rather than deliver a capability you would rely on.

Audio Scribe and Ask Image are the two features with the clearest practical value: transcription and translation of voice recordings, and object identification or description from the camera or photo library. Both inherit the same on-device constraint, so both are limited by the model you have downloaded.

## Where Google AI Edge Gallery is the wrong tool

The most important limitation is stated by the project itself. The README describes this as an experimental Beta release. That is not a formality. It means the model list, the skill format and the app surfaces can change between releases, and the release cadence visible in the repository (1.0.16 in June, 1.0.17 in early August, 1.0.18 on 2026-08-10) is fast enough that behaviour you depend on may move.

Beyond the Beta label, the architecture sets hard boundaries. Because inference is on-device, throughput is whatever your handset can sustain, and the README makes no performance claims you could plan capacity around. Model download size is the other practical wall: the README does not state storage requirements for any model, so the only way to know whether a model fits is to check the model itself. On a device with limited free storage, that is a real constraint.

The app is also not embeddable. Nothing in the README describes an API, an SDK, or a way to call the runtime from your own application. If you want on-device inference inside a product you are shipping, this project is a reference for what the underlying Google AI Edge and LiteRT stack can do, not the integration path. The README points at the Google AI Edge documentation and the LiteRT-LM repository for that.

Finally, the allowlist cuts both ways. It protects you from loading something the runtime cannot run, and it also means the app is not a general model loader. Custom models are supported, but within whatever the allowlist and runtime permit.

## Alternatives and the actual difference in approach

The relevant comparison is not another gallery app. It is llama.cpp, which occupies the same problem space (run open-source language models on hardware you own) with a different design centre.

llama.cpp is a runtime and a set of command-line tools. You build it, you point it at a GGUF model file, and you run inference from a shell or link it into your own program. There is no curated model list, no allowlist, no app store distribution and no graphical surfaces for chat, image input or transcription unless someone else builds them on top. Its strength is exactly the flexibility Google AI Edge Gallery trades away: any model in a supported format, any integration, any platform you can compile for.

Google AI Edge Gallery inverts that. It ships a finished, store-distributed application with a fixed set of tiles, a vetted model catalogue and a runtime (LiteRT) chosen by the vendor. You give up control over which model you run and how you call it, and in exchange you get something a non-engineer can install from Google Play or the App Store and use in five minutes. The README's own framing supports this reading: it is a gallery for trying and evaluating models, not a runtime for building on.

If your goal is to evaluate a model on a phone you already carry, the app is the shorter path. If your goal is to embed on-device inference in software you are writing, llama.cpp or the LiteRT-LM repository is the more direct route.

## Maintenance, licensing and what an upgrade costs you

The repository is not archived, and the last push was on 2026-08-10, which lines up with the 1.0.18 release. Three releases landed between late June and mid August 2026, so the project is moving, but the README's own Beta label is the more useful signal about stability than the commit dates.

Upgrade cost is asymmetric between platforms. On Android and iOS, the app store handles the update and your downloaded models persist in the app's storage, so the practical cost is re-checking whether a feature you depend on changed scope. The README's Thinking Mode note, which ties the feature to supported models starting with Gemma 4, is the kind of statement that can shift between releases. On macOS the README links a 0.1.0 DMG while mobile is at 1.0.18, so desktop users should expect to re-download the disk image rather than receive an in-place update.

Licensing is Apache-2.0, stated in the README and present as a LICENSE file at the top level. That covers the application source. It does not automatically cover the models you download through the app: those come from the Hugging Face LiteRT community and carry their own terms, which the README does not enumerate. If you plan to redistribute anything built from this, check the licence of each model separately. Nothing here is legal advice; read the LICENSE file and each model's terms.

## Conclusion

Adopt Google AI Edge Gallery if you want to see how a specific open-source model behaves on hardware you already own, or if offline inference is a hard requirement: the README states that all model inferences happen directly on the device and that no internet is required. Do not adopt it as a production inference layer or as a substitute for your own evaluation harness. Before committing, verify three things: the OS floor (Android 12 and up, iOS 17 and up), the size of the model you intend to download, and whether your target device appears in the Project Wiki's installation instructions for corporate devices. The macOS build is a 0.1.0 DMG while the mobile app is at 1.0.18, so treat desktop as the less-travelled path.

## FAQ

### What are the OS requirements for Google AI Edge Gallery?

The README states Android 12 and up, and iOS 17 and up. A macOS build is also linked from the README as a DMG.

### How do I install Google AI Edge Gallery without Google Play?

The README says users without Google Play access should install the APK from the latest release page on GitHub. Detailed installation instructions, including for corporate devices, are in the Project Wiki.

### Does Google AI Edge Gallery send my prompts to a server?

The README states that all model inferences happen directly on the device hardware and that no internet is required, so prompts, images and other data stay on the device.

### Which models does Thinking Mode work with in Google AI Edge Gallery?

The README says Thinking Mode currently works with supported models, starting with the Gemma 4 family. It does not list the full set of supported models.

### What licence does Google AI Edge Gallery use?

The README states the project is licensed under the Apache License, Version 2.0, with a LICENSE file at the top level. Models downloaded through the app come from the Hugging Face LiteRT community and carry their own terms.

## Sources

- [Official README](https://github.com/google-ai-edge/gallery#readme)
- [Project repository](https://github.com/google-ai-edge/gallery)
- [Release notes](https://github.com/google-ai-edge/gallery/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/google-ai-edge-gallery
