Model or dataset
orailnoor/cross-platform-llm-client avatar
orailnoor/cross-platform-llm-client

PrivateLM: a Flutter LLM client that runs GGUF models on Android and iOS

A unified cross-platform AI client supporting seamless transitions between standard cloud APIs and on-device, offline execution of custom and uncensored language models.

1,110 stars224 forksC++MIT

At a glance

What is it?
PrivateLM is a cross-platform AI chat client built with Flutter that pairs on-device GGUF inference on Android and iOS with fallback to OpenAI, Anthropic, Gemini and Kimi. This review covers how the inference pipeline is wired, how to build it, and where it breaks down.
Who is it for?
Adopt PrivateLM if you want a phone-first chat client where the default path is a local GGUF model and cloud providers are an explicit opt-in, and if you are prepared to build it yourself: the repository ships source, not store listings, and the iOS build is a sideloaded ZIP. Skip it if you need a desktop Linux or macOS client, since the platform table lists only Android, iOS and web, and the web target is cloud-only.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 71 days ago.
What is it written in?
Mainly C++, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What PrivateLM actually solves for Android and iOS users

Most chat clients assume a network call. PrivateLM starts from the opposite assumption: on Android and iOS, the model runs on the device. The README describes local inference as the primary mode, with cloud providers as a fallback for what a phone cannot do well. That ordering matters, because it changes what the app has to get right: model download, memory budgeting, GPU offload, and a UI that stays responsive while tokens stream from a local engine.

The audience is narrower than the phrase cross-platform suggests. The supported-platform table lists Android (minSdk 28), iOS, and web. Android gets full llama.cpp inference with CPU offload via NEON. iOS uses Metal GPU acceleration, distributed as a standalone ZIP for sideloading through AltStore, Sideloadly or Xcode. The README states that iPad is the recommended iOS target because of the RAM local models need, and calls iPhone support experimental. Web is cloud-only, with local inference described as coming soon.

So the real user is someone with a recent Android phone or an iPad who wants conversations to stay on the device unless they choose otherwise. Persistent sessions, tasks and settings are stored locally via Hive, and the README says nothing leaves the device unless cloud mode is selected explicitly.

How the inference pipeline moves from UI to tokens

The README lays out a four-layer architecture. Views (ChatView, TaskView, ModelView, SettingsView) sit above GetX controllers (ChatController, TaskController, ModelController, SettingsController, HomeController). Below those are services: InferenceService for local GGUF, CloudService for the four providers, DownloadService for model files, HiveService for persistence, DeviceInfoService for RAM and GPU tier, and ExecutionService for background tasks.

Local generation is delegated to llama_flutter_android, a custom Flutter plugin in the local_plugins/ directory that wraps llama.cpp for ARM64. At runtime the sequence the README gives is: detect GPU capabilities via Vulkan to decide offload layers, pick a thread count from the device tier (ultra, high, mid, low), load the GGUF model with progress streaming, then generate tokens through generateChat() with native chat-template support for ChatML, Llama-3, Gemma and Phi. If a native template fails, the app falls back to manual prompt construction.

Two timeouts are worth noting. Idle detection fires at 5 seconds and a hard timeout at 180 seconds, which the README frames as keeping the UX responsive on underpowered hardware. That is a design admission: on a slow device, generation can be cut off rather than allowed to run to completion.

CloudService normalizes four different API shapes into one interface: OpenAI's /v1/chat/completions, Anthropic's Messages API with a separate system parameter, Gemini's generateContent with inline base64 images, and Kimi's OpenAI-compatible endpoint from Moonshot AI. API keys live in Hive and, per the README, are sent only to the provider's endpoint.

Cross-platform abstraction is handled by conditional compilation: inference_android.dart for Android and iOS, inference_stub.dart for web. InferenceService exposes supportsLocalInference so the UI can hide local-model controls where they do not apply.

Building PrivateLM from source: prerequisites and a first APK

There is no app store listing in the README. The project points to a live web demo and to the Releases page for the iOS ZIP, and the rest is a Flutter build. Prerequisites are Flutter SDK >=3.3.0, Android SDK (API 26+), JDK 17, and the NDK bundled with the Android SDK. The pubspec requires Dart >=3.3.0.

Start by fetching dependencies and producing a debug APK:

bash
flutter pub get
flutter build apk --debug

The first command resolves the Flutter packages, including the local plugin under local_plugins/. The second produces a debug APK you can install directly. Expect the build to take a while on first run, since the native llama.cpp plugin has to compile.

Release builds need a signing key. The README says to copy the example properties file and fill in the keystore values:

bash
cp android/key.properties.example android/key.properties
flutter build apk --release --split-per-abi

The --split-per-abi flag produces separate APKs per architecture. The README adds two constraints that are easy to miss: never rotate the signing key between GitHub releases, and keep increasing the build number in pubspec.yaml. Android only accepts an APK upgrade when it is signed with the same key as the installed APK, so a rotated key means users must uninstall first.

For iOS, the build path is different:

bash
flutter pub get
cd ios
pod install
flutter build ios

pod install resolves the CocoaPods dependencies for the iOS target. If you would rather not build, the README says the iPad release is a standalone PrivateLM-iOS.zip on the Releases page, extracted and installed as an .ipa via AltStore, Sideloadly or Xcode. For web, flutter build web --release is the command given.

Where PrivateLM is the wrong tool

The platform table is the first limit. There is no Linux or macOS target listed, so anyone searching for a desktop client will not find one here. The web build exists but is cloud-only: inference_stub.dart compiles in place of the real engine, and the README describes local inference on web as coming soon. If your requirement is offline chat in a browser, this build does not meet it.

iPhone support is described as experimental, with iPad recommended instead. That is a RAM constraint, not a polish issue. Local GGUF inference on a phone-sized memory budget limits both model size and context length, and the README's smart auto-configuration, which detects device RAM on first launch and recommends context size and token limits, is a workaround for that ceiling rather than a fix.

The 180-second hard timeout is a second boundary. For long generations on mid or low tier hardware, the app may cut off before the model finishes. The README presents this as a responsiveness trade-off, which it is, but it also means PrivateLM is a poor fit for workloads that need a complete long answer in one pass.

Finally, the local engine is ARM64 through llama_flutter_android. Nothing in the README claims x86 support, and the per-ABI release build implies you should check which ABI you are shipping before distributing.

PrivateLM against llama.cpp and Ollama

The closest comparison is llama.cpp itself. PrivateLM does not replace it; the README says the Android and iOS paths are a Flutter plugin wrapping llama.cpp. The difference is the layer above. llama.cpp gives you a binary and a command line. PrivateLM gives you Flutter views, GetX controllers, Hive-backed session persistence, a model download service, and a cloud fallback that switches providers without changing the chat UI. If you want a library to embed in your own app, llama.cpp is the substrate and PrivateLM is a reference application built on it.

Ollama takes a different approach again: it runs as a local server exposing an HTTP API, typically on a desktop or a machine with more memory, and clients talk to it over the network. PrivateLM runs the model inside the phone app process, with no server hop, and its cloud branch speaks directly to OpenAI, Anthropic, Gemini and Kimi. The trade-off is capacity. A server-hosted runner can hold a much larger model than a phone, but it is not offline in the same sense, and it is not the same device. PrivateLM's premise is that the device you already carry is the host.

For the cloud-only path, PrivateLM is one client among many, and its distinguishing feature there is that it normalizes four provider APIs behind one interface while keeping the local option in the same app.

Maintenance, licence and what a fork inherits

The repository is not archived. The last push was on 2026-07-21, and the most recent release, PrivateLM App v1.0.5, was tagged the same day. The previous release, listed as 1.0.3 with the tag PrivateLM App v1.0.4, dates to 2026-04-27. That is a gap of roughly three months between the two releases shown, so the cadence is not weekly, and the version numbers in the release list do not line up cleanly with the tag names.

The licence is MIT, which is permissive and places few obligations on a fork beyond retaining the notice. This is not legal advice, and anyone redistributing should read LICENSE in the repository root rather than rely on the README's one-line summary.

Upgrade cost is dominated by the signing key rule. The README is explicit that the signing key must not rotate between releases and that the build number in pubspec.yaml must keep increasing. Get either wrong and existing installs will not upgrade in place. A fork that wants to ship its own builds therefore needs its own keystore from the start; you cannot adopt the upstream key. Beyond that, the dependency surface is a normal Flutter stack: GetX, Hive, Dio, flutter_background_service, flutter_local_notifications, Firebase Core and Firebase Messaging. Firebase is listed for push notifications and background handling, which means a fork that removes Firebase has work to do in the ExecutionService path.

Editorial conclusion

Adopt PrivateLM if you want a phone-first chat client where the default path is a local GGUF model and cloud providers are an explicit opt-in, and if you are prepared to build it yourself: the repository ships source, not store listings, and the iOS build is a sideloaded ZIP. Skip it if you need a desktop Linux or macOS client, since the platform table lists only Android, iOS and web, and the web target is cloud-only. Before committing, verify two things in the source tree: that local_plugins/llama_flutter_android actually wraps llama.cpp for your target ABI, and that android/key.properties is set up as the README describes, because an APK upgrade is rejected unless it is signed with the same key as the installed build.

Frequently asked questions

Is there a private LLM app available in PrivateLM?

Yes. PrivateLM runs GGUF models on-device on Android and iOS, and the README states that chats, tasks and settings are stored locally via Hive and that nothing leaves the device unless cloud mode is chosen explicitly.

What does cross-platform mean for PrivateLM?

The supported-platform table lists Android, iOS and web. Local inference is available on Android and iOS through llama_flutter_android, while the web build uses inference_stub.dart and is cloud-only, with local inference described as coming soon.

How do I build PrivateLM for Android?

The README lists Flutter SDK >=3.3.0, Android SDK (API 26+), JDK 17 and the bundled NDK, then flutter pub get followed by flutter build apk --debug. Release builds need android/key.properties copied from the example file and a stable signing key.

Does PrivateLM run local models on the web or on iPhone?

The web build is cloud-only, with local inference marked as coming soon. On iOS the README says iPad is the recommended target because of RAM requirements, and iPhone support is experimental.

Which cloud providers can PrivateLM talk to?

CloudService normalizes OpenAI, Anthropic, Google Gemini and Kimi from Moonshot AI behind one interface. API keys are stored in Hive and, per the README, are only transmitted to the provider's endpoint.

Official sources

  1. Issues
  2. License: MIT
  3. orailnoor/cross-platform-llm-client on GitHub
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/orailnoor-cross-platform-llm-client.svg)](https://hysenlabs.com/projects/orailnoor-cross-platform-llm-client)