# Off Grid AI (OGAM): an offline AI suite for Android, iOS and macOS

> Off Grid AI bundles GGUF chat, vision, Whisper transcription, Stable Diffusion and tool calling into one React Native app that runs on phone or Mac hardware. The README is candid about NPU limits, and that candour is the most useful thing in it.

**off-grid-ai/OGAM** — The Swiss Army Knife of Offline AI. Chat, see, speak, and generate images on your phone or Mac — GGUF LLMs, vision, Whisper speech-to-text, Stable Diffusion, tool calling, and local-network servers. Runs on your CPU, GPU, or NPU. No account, no API key, zero data leaves your device.

- Repository: https://github.com/off-grid-ai/OGAM
- Website: https://getoffgridai.co/pro/
- Stars: 3,164 · Forks: 302
- Language: TypeScript
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/off-grid-ai-ogam

## What Off Grid AI solves, and for whom

Most local LLM apps on phones are chat wrappers. Off Grid AI, published as the OGAM repository by off-grid-ai, positions itself as a suite instead: text generation, image generation, vision, voice transcription, tool calling and document analysis in one application, running on the device. The README's own framing is that "most 'local LLM' apps give you a text chatbot and call it a day." Whatever you think of the marketing, the feature list backs the claim. The same app runs GGUF models, Stable Diffusion, SmolVLM or Qwen3-VL for vision, and Whisper for speech-to-text.

The audience is narrow and specific. It is people who cannot or will not send prompts to a hosted API: field workers, privacy-conscious users, anyone on a plane or in a basement, and engineers who want to test on-device inference without building the whole stack themselves. The README states plainly that there is no account and no API key, and that zero data leaves the device. That is the product's entire reason to exist, and everything else is secondary to it.

There is a commercial layer. Off Grid AI Pro costs $69 lifetime or $49 annual and adds on-device Kokoro text-to-speech, custom personas with persistent memory, draft-then-approve actions against Calendar, email and MCP servers, and Personal Mesh sync between phone and Mac. The free tier is not crippled: it transcribes speech, generates images and runs models. Pro adds the voice that talks back and the action layer. One license covers phone and Mac.

## How the inference stack is put together

The repository is a React Native application written in TypeScript, with native Android and iOS directories alongside it, plus a shared package that the mobile build pulls in. The package.json scripts make the boundary visible: `prepare:shared` runs `../shared`'s build and then verifies a consumer contract before the Android or iOS build starts. That is a deliberate guard against the mobile app drifting from the shared code, and it means a clean checkout needs the sibling `shared` directory present before `npm run android` will work.

The inference side is llama.cpp for text, whisper.cpp for speech, and Stable Diffusion for images, all wrapped in native modules. Model files are GGUF, and the app accepts your own `.gguf` files as well as its own catalogue. Backend selection is automatic: the README says the app detects what the device has and defaults to the fastest backend that works, with an override in Settings. Adreno GPUs go through OpenCL and Apple Silicon through Metal.

Memory is treated as a first-class problem rather than an afterthought. A model manager shows what is resident and what each model costs in RAM, with a per-model eject. A Model Loading policy setting chooses between Lean (one model at a time), Balanced (co-resident models that fit, swap the rest) and Aggressive (commit a larger RAM share so bigger models load). On a phone, that is the difference between a working session and an out-of-memory kill, and it is more control than most mobile inference apps expose.

Projects add a retrieval layer. Uploaded PDFs and text files are chunked, embedded on-device with a bundled MiniLM model, and retrieved by cosine similarity, stored locally in SQLite. The `search_knowledge_base` tool becomes available in project conversations. Remote servers are also supported: any OpenAI-compatible endpoint on the local network (Ollama, LM Studio, LocalAI) can be discovered and streamed from via SSE, with API keys stored in the system keychain.

## Installing Off Grid AI and running a first model

There are two routes. The README links Google Play (`ai.offgridmobile`) and the App Store (`id6759299882`) for the packaged app, and that is the route most people should take, because the native toolchains for llama.cpp, whisper.cpp and Stable Diffusion are not trivial to assemble. A macOS desktop build is linked separately from the off-grid-ai/desktop repository, and the README says one Pro license covers both.

Building from source is for contributors. The repository is private in the sense that package.json sets `"private": true`, so this is an application, not a published library. Clone it with the sibling `shared` repository, install dependencies, and note that the prestart, preandroid and preios hooks all run `prepare:shared` first, which builds the shared package and checks the contract.

```bash
git clone https://github.com/off-grid-ai/OGAM.git
cd OGAM
npm install
```

The Android debug run targets a separate application id, `ai.offgridmobile.dev`, so a development build can sit beside a store install:

```bash
npm run android
```

On iOS the script is simply `npm run ios`, with `npm run ios:device` for a physical device. Expect the first build to take a while: Gradle builds the native inference libraries, and the README's model catalogue is not bundled with the app.

For a first real use, the onboarding flow walks you through picking a model. Choose a small GGUF quant first. The README's throughput figures are 15-30 tok/s on CPU on flagship devices and 20-40 tok/s on an Adreno GPU via OpenCL on a Snapdragon 8 Gen 2 or newer. If you have a GPU-capable device, check the badge in the model list before downloading, because the badge tells you whether that model can use the GPU or NPU at all. Then open a chat, type a prompt, and watch the streaming response render as markdown. If you want to test retrieval instead, create a project, upload a PDF, and ask a question that only the PDF answers; the app should route through `search_knowledge_base`.

## The Hexagon NPU is the honest part of the README

The README does something rare: it documents its own weak spot without hedging. The Hexagon NPU on Snapdragon is "marked experimental because it is." Three specific constraints follow. It only accelerates `Q4_0` and `Q8_0` quants. A K-quant silently falls back to CPU, which means you can believe you are on the NPU and be measuring CPU throughput instead. And some model architectures come out garbled on it.

That third point is the one to take seriously. Garbled output is not a performance regression; it is a correctness failure, and it can look like the model being bad at the task rather than the backend being wrong. If you hit nonsense from a model that should handle the prompt, switching the backend in Settings to CPU or GPU is the first thing to try. The README's mitigation, badging GPU- and NPU-capable models in the list, only helps if you read the badge before downloading. The app does not appear to warn you after the fact.

There are other boundaries worth naming. Image generation on the NPU is quoted at 5-10s per image, and vision at roughly 7s on flagship devices, both of which are usable but not interactive. The throughput numbers throughout the README are for flagship hardware, and a mid-range phone will be slower. The README does not document a rollback path for a failed model download or an upgrade that misbehaves, so treat model files as something you manage yourself. And the Pro tier's MCP integration, which the README describes as drafting a reply or filing a ticket and waiting for your tap, is only as useful as the servers you connect; the free tier has no action layer at all.

## How it differs from Ollama and LM Studio

The obvious comparison is Ollama or LM Studio, and the difference is not model support. Both of those run GGUF models well, and Off Grid AI can talk to them: the README describes connecting to any OpenAI-compatible server on the local network, discovering models and streaming via SSE. So on a laptop, Ollama is the more conventional choice, with a CLI, a large model library and a server API that other tools already speak.

The difference is where the compute lives. Ollama and LM Studio assume a desktop or server with a few gigabytes of headroom and a stable power supply. Off Grid AI assumes a phone: finite RAM, thermal throttling, a battery, and a GPU or NPU that may or may not be usable for a given quant. The model manager, the three loading policies and the per-model eject exist because of that assumption. Neither Ollama nor LM Studio has an equivalent problem to solve, and neither ships vision, Whisper transcription and Stable Diffusion in the same binary.

If your goal is a local model server for other applications to call, Ollama is the better fit and Off Grid AI's remote-server mode makes it a client rather than a competitor. If your goal is one app on a phone that does chat, camera questions, dictation and image generation without a network, the comparison does not really apply, because the desktop tools were never trying to run there.

## Maintenance, licence and upgrade cost

The repository is not archived, and the last push was on 2026-09-10, a week before this writing. Releases are frequent and versioned at patch level: v0.0.107 landed on 2026-08-21, preceded by two beta builds the same day. The 0.0.x versioning is worth reading literally. This is pre-1.0 software, and the package.json version matches the release tag, so an upgrade can change native behaviour as well as JavaScript.

The practical upgrade cost is model files, not code. Because the app reads GGUF files from disk, a release that changes the inference backend can change which quants run well, and the NPU's `Q4_0`/`Q8_0` restriction means a quant you chose for quality may not be the one that accelerates. Budget for re-testing your chosen model after a version bump, especially if you rely on the NPU or GPU path.

The licence is MIT, which is permissive: you can use, modify and redistribute the code, including commercially, provided the copyright notice and permission notice are preserved. That covers the application source in this repository. It does not automatically cover the models you download, which carry their own licences (Llama, Qwen, Gemma and Stable Diffusion derivatives each differ), and it does not cover the Pro tier, which is a paid product governed by its own terms rather than the MIT grant. The README points at a separate desktop repository and a Pro directory inside the tree; if you plan to redistribute a build, check which parts are actually under the MIT file before you ship. This is a description of the licence, not legal advice.

## Conclusion

Adopt Off Grid AI if you want a single on-device app that covers chat, vision, speech and image generation without an account or API key, and you accept that the Hexagon NPU path is experimental and quant-sensitive. Do not adopt it if your workflow depends on a model architecture the NPU garbles, or if you need a documented rollback path the README does not provide. Verify first that your device has a GPU backend the app detects (Adreno via OpenCL, Apple Silicon via Metal), and check the model list badges before downloading a 4GB quant that will silently fall back to CPU.

## FAQ

### Is there a free offline AI app for Android?

Off Grid AI has a free tier on Google Play under the application id ai.offgridmobile, and the README states it runs without an account or API key and that no data leaves the device. The free app includes text generation, image generation, vision and Whisper transcription; Pro adds text-to-speech, personas and an action layer.

### Is there any AI app that works without internet?

Off Grid AI is built for that case. Models are GGUF files run through llama.cpp on the device, Whisper handles speech-to-text locally, and Stable Diffusion generates images on-device, so no request goes to a hosted API. The README also notes it can connect to a local-network OpenAI-compatible server, which still requires no internet.

### Does Off Grid AI use AI?

Off Grid AI is the AI: it runs large language models, a vision model, Whisper speech-to-text and Stable Diffusion image generation locally on your phone or Mac. The README describes the app as a complete offline AI suite rather than a client for a hosted service.

### What are the top 5 AI applications?

The README does not rank AI applications or make any claim about which are the top five, so the material cannot answer this. What it does describe is Off Grid AI's own feature set: text generation, image generation, vision AI, voice transcription, tool calling and document analysis.

## Sources

- [License: MIT](https://github.com/off-grid-ai/OGAM/blob/main/LICENSE)
- [off-grid-ai/OGAM on GitHub](https://github.com/off-grid-ai/OGAM)
- [Project website](https://getoffgridai.co/pro/)
- [README](https://github.com/off-grid-ai/OGAM/blob/main/README.md)
- [Releases](https://github.com/off-grid-ai/OGAM/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/off-grid-ai-ogam
