# SmolChat-Android: GGUF Model Inference on Android Without a Server

> SmolChat-Android is an Android app and Kotlin/JNI wrapper for running GGUF small language models on-device through llama.cpp. It is built for people who want a chat UI and a reusable inference module, not a hosted API.

**shubham0204/SmolChat-Android** — Running any GGUF SLMs/LLMs locally, on-device in Android

- Repository: https://github.com/shubham0204/SmolChat-Android
- Stars: 894 · Forks: 146
- Language: Kotlin
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/shubham0204-smolchat-android

## What SmolChat-Android Solves, and for Whom

Running a language model on an Android phone usually means one of two compromises: send the prompt to a remote API and lose the privacy and offline properties, or embed a C/C++ inference stack and write the JNI plumbing yourself. SmolChat-Android takes the second path and packages the plumbing. The README describes the goal as providing "a usable user interface to interact with local SLMs (small language models) locally, on-device", plus the ability to add and remove GGUF models and edit their system prompts and inference parameters such as temperature and min-p.

The audience is narrow and identifiable. First, Android developers who want a working reference for llama.cpp through the NDK, because the repository separates the inference wrapper (the smollm module) from the application logic (the app module). Second, users who want a chat app that does not require an account or a network round trip. Third, people evaluating small models on real phone hardware, since the app loads GGUF files directly and exposes the sampling parameters.

It is not a general-purpose assistant platform. There is no server, no model marketplace, and no tool-calling layer described in the README. The repository also contains an hf-model-hub-api entry and an smolvectordb entry at the top level, but the README does not document what either one does, so treat them as unexplored until you read the source.

## How the llama.cpp and JNI Layer Fit Together

The architecture has three layers, and the README is explicit about each. At the bottom, llama.cpp loads and executes GGUF models. The README notes that llama.cpp is pure C/C++, which makes it straightforward to compile for Android targets with the NDK. In the middle, the smollm module holds llm_inference.cpp, which talks to llama.cpp's C-style API, and smollm.cpp, which is the JNI binding. On the Kotlin side, the SmolLM class exposes the methods that reach the native code.

The data flow in the app module is the part worth reading closely. When a new chat opens, the app instantiates SmolLM and hands it the model file path stored by the LLMModel entity. It then retrieves messages with the user and system roles from the database and pushes them into the native side with LLMInference::addChatMessage. That means chat history lives in a local database and is replayed into the inference context at chat open time, rather than being held only in memory.

Tasks take a different route. The README states that for tasks, messages are not persisted, and the app signals this by passing _storeChats=false to LLMInference::loadModel. This is a deliberate split: interactive chat is durable, one-shot task generation is not. If you are building on the smollm module directly, that boolean is the switch you will care about most, because it decides whether the native layer keeps conversation state across calls.

## Installing SmolChat-Android and Loading a First Model

There are three installation routes. The README lists Google Play, a direct APK from GitHub Releases, and Obtainium for update notifications. The Play listing is io.shubham0204.smollmandroid. For the APK route, the README says to download the latest APK from GitHub Releases and transfer it to the device, and to search for how to allow installing APKs from unknown sources if the device blocks it.

Obtainium is the option for people who want release notifications without checking the repository. The README gives the exact URL to paste:

```bash
https://github.com/shubham0204/SmolChat-Android
```

After adding that URL in Obtainium's Add App screen, SmolChat appears in the Apps screen and newer releases can be downloaded from there.

Building from source is the path for anyone modifying the inference layer. The README gives a shallow clone plus submodule init, because llama.cpp arrives as a submodule:

```bash
git clone --depth=1 https://github.com/shubham0204/SmolChat-Android
cd SmolChat-Android
git submodule update --init --recursive
```

Android Studio then starts building the project automatically. If it does not, the README says to select Build > Rebuild Project. After a successful build, connect an Android device; its name should appear in the top menu bar before you run the app. Note that the README's install section does not walk through downloading a GGUF file or adding it in the UI, so the first model you load is something you supply yourself. The app's stated purpose is to let you add and remove GGUF models, and the model path is stored in the LLMModel entity, but the README stops short of a file-picker walkthrough.

## Where SmolChat-Android Falls Short

The clearest limitation is stated by the project itself. Under Future, the README lists checking whether llama.cpp can be compiled to use Vulkan for inference on Android devices, with the parenthetical "and use the mobile GPU". Until that happens, the documented execution path is llama.cpp on the CPU through the NDK. On a phone, that means token generation speed is bounded by the CPU and by thermal throttling, and it degrades as the model grows. The app is aimed at SLMs, and the naming is honest about it.

The second limitation is the model supply chain. SmolChat does not ship models and the README does not describe a download flow. You bring a GGUF file, which means you also own the responsibility for picking a quantization that fits your device's memory and for verifying that the file is the model you think it is.

The third is the documentation boundary. The README explains the module split and the _storeChats flag, but it does not document error handling for failed model loads, memory limits, or what happens when a chat's history exceeds the context window. The Future list also mentions a background service for Bluetooth, HTTP, or WiFi query forwarding from a desktop, and automatic chat naming, which tells you neither exists yet. If your requirement is a desktop-to-phone inference bridge, this project is not the wrong tool because of quality; it is the wrong tool because that feature is on a planning list, not in the code path the README describes.

## How It Compares to Shipping Your Own llama.cpp Wrapper

The real alternative for an Android developer is not another chat app. It is writing the JNI bridge yourself against llama.cpp, or using a runtime such as ONNX Runtime Mobile with a converted model. The difference in approach is concrete. With a hand-rolled llama.cpp integration, you control the CMake setup, the NDK version, the threading strategy, and the exact subset of the C API you expose. You also own every one of those decisions, including the ones SmolChat has already made for you.

SmolChat's bet is that the smollm module is small enough to read and adopt: llm_inference.cpp against the C API, smollm.cpp as the JNI binding, and SmolLM.kt as the Kotlin surface. The app module then demonstrates a database-backed chat history and a separate non-persisted path for tasks. If that split matches your design, adopting the module saves you the JNI work. If it does not, the module is a readable starting point rather than a framework you have to fight.

The ONNX Runtime route differs more fundamentally: it targets ONNX graphs rather than GGUF, so the model ecosystem you draw from is different, and the GGUF-specific tooling around llama.cpp does not apply. The README does not compare the two, and it should not; the choice depends on which model files you already have.

## Licence, Maintenance, and the Cost of Upgrading

SmolChat-Android is Apache-2.0. That is a permissive licence, and for an application that links native code, the practical question is whether your distribution obligations are acceptable; the repository includes a LICENSE file at the top level, and you should read it rather than rely on the identifier alone. The README does not discuss licence compatibility between the app and the models you load, which is a separate question you have to answer per model.

The project is not archived, and the last push was on 2026-06-21. Releases are frequent enough to matter for upgrade planning: v16 on 2026-06-21, v15 on 2026-04-17, and v14 on 2026-03-01. That cadence means if you fork the smollm module, you should expect to rebase against upstream changes rather than treat it as a frozen dependency.

The upgrade cost has a specific shape here. llama.cpp is a submodule, so bumping the app's inference engine means updating the submodule pointer and rebuilding the native side, not just changing a Gradle version. Any local patches to llm_inference.cpp or smollm.cpp will need to be replayed. The README's build instructions assume the submodule is present, so a source build without git submodule update --init --recursive will not produce a working native library.

## Conclusion

Adopt SmolChat-Android if you need on-device GGUF chat on Android and want a small Kotlin surface (SmolLM) over llama.cpp, or if you are building an app that must keep prompts and responses local. Skip it if you need GPU-accelerated inference today: the README lists Vulkan support only as a future item, and the current implementation is CPU-side llama.cpp through the NDK. Before committing, verify that your target GGUF models load on your device, that the smollm submodule builds against your NDK version, and that the app module's database-backed chat history matches your privacy requirements.

## FAQ

### What is SmolChat?

SmolChat-Android is an Android app that runs GGUF small language models locally, on-device, using llama.cpp. It provides a chat interface and lets users add or remove models and adjust system prompts and inference parameters such as temperature and min-p.

### What is the best local LLM app for Android?

SmolChat-Android is one candidate: it loads GGUF models on-device through llama.cpp, keeps chat history in a local database, and exposes inference parameters such as temperature and min-p. The README does not compare it against other apps, so the choice depends on whether its GGUF and llama.cpp approach fits your models.

### Which is the best AI for Android?

The README does not rank models or assistants. It states that SmolChat-Android executes GGUF SLMs on-device with llama.cpp and lets users add or remove models, so the model choice is left to the user rather than recommended by the project.

## Sources

- [Issues](https://github.com/shubham0204/SmolChat-Android/issues)
- [License: Apache-2.0](https://github.com/shubham0204/SmolChat-Android/blob/main/LICENSE)
- [README](https://github.com/shubham0204/SmolChat-Android/blob/main/README.md)
- [Releases](https://github.com/shubham0204/SmolChat-Android/releases)
- [shubham0204/SmolChat-Android on GitHub](https://github.com/shubham0204/SmolChat-Android)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/shubham0204-smolchat-android
