# MagicWX: an Android 17 prototype for offline LLM chat and CPU Stable Diffusion

> MagicWX is a Kotlin and Jetpack Compose Android app that runs registered language models offline through ONNX Runtime, MediaPipe LiteRT and RWKV, and generates images locally with an SD1.5 CPU pipeline. Its own README is unusually strict about what counts as supported, and that distinction is the main thing to understand before adopting it.

**Pangu-Immortal/MagicWX** — Android 17 local LLM prototype with Jetpack Compose and ONNX Runtime for offline AI inference experiments.

- Repository: https://github.com/Pangu-Immortal/MagicWX
- Stars: 856 · Forks: 362
- Language: Kotlin
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/pangu-immortal-magicwx

## What MagicWX actually solves, and for whom

Most Android AI demos assume a network call. MagicWX assumes the opposite: the phone has the weights, and inference happens on the device. The app is built with Kotlin, Jetpack Compose and Material3, targets API 37 (Android 17), and pairs offline text chat with on-device Stable Diffusion image generation. The README describes the scope as offline model selection, verified package checks, runtime gating, chat, and a full image workspace covering txt2img, img2img, inpaint, mask painting, cropping and persistent history.

The audience is narrow and specific. This is for Android engineers who want to see how a text-generation runtime is wired up on a phone, how a model download is gated behind verification, or how a diffusion pipeline is kept out of the UI process. It is not aimed at people who want an app to install and chat with. The README states plainly that the repository intentionally distinguishes registered model candidates from verified GA support, and that a model should not be described as supported until it has a complete package manifest, required assets, a successful load, a fixed-input dry-run and device verification evidence. That sentence is the project's thesis. Everything else follows from it.

## Runtime adapters, download verification and the image backend process

The architecture separates three concerns: model metadata, download and validation, and runtime dispatch. The package layout puts tokenizers, model wrappers and the downloader under model/, with an adapter/ subdirectory holding ModelRuntimeAdapter.kt for dispatch and load gates and MediaPipeLlmAdapter.kt for the MediaPipe/LiteRT .task text runtime. A foreground service, ModelDownloadService.kt, handles downloads and reports progress through ModelDownloadEvents.kt, with state surfaced by MainViewModel.kt as MVVM state and user actions.

Four runtime adapters are named in the README: BUILTIN_TEXT, ONNX_TEXT_GENERATION, LITERT_LM (MediaPipe), and an isolated native image backend. That last one is the interesting design decision. The diffusion pipeline lives in a separate native library, libmagicwx_image_backend.so, running as a foreground-service-managed process on localhost:18081 with SSE streaming. Keeping a CPU-bound image pipeline out of the UI process is a reasonable call on a phone, and the localhost boundary means the Compose layer talks to it over a socket rather than through JNI. The cost is a second process to manage and an HTTP surface to keep alive.

The model list is deliberately short. Verified entries include the built-in experience model, RWKV-7 World 0.4B, Qwen3 0.6B, Qwen2.5 0.5B, SmolLM2 360M, and TinyLlama 1.1B Chat LiteRT. A longer candidate pool exists in code with capability, runtime adapter, visibility and asset metadata, but candidates stay hidden until the adapter and device tests pass. Several candidates are explicitly hidden for concrete failures: Qwen2 0.5B for garbled output, SmolLM2 135M for abnormal output, SmolLM2 135M MHA for an ORT shape mismatch, and Gemma 3 270M for repeated-fragment output. Naming the failure rather than quietly dropping the model is more useful than a longer support table would be.

## Building MagicWX and getting a first offline chat response

The repository is a standard Gradle Android project. The top level contains build.gradle.kts, settings.gradle.kts, gradle.properties, the gradle wrapper scripts gradlew and gradlew.bat, and an app/ module. The README does not spell out a build command, so the wrapper scripts it ships are the entry point. From the repository root, the wrapper is invoked like this, and on Windows the same directory holds gradlew.bat:

```bash
./gradlew
```

The README does not document the exact task names, so check the app module's build.gradle.kts before assuming a target. The README states the Android target is API 37 / Android 17, so an emulator or device below that is not what the project was built against.

Once the app is running, the first useful action is the built-in model. The README describes it as a deterministic local experience engine, not a bundled large model weight, and it is the only model that needs no download. Sending a message there confirms the Compose UI, the tokenizer path and the generation interface are wired correctly before you spend time on a multi-hundred-megabyte download.

The external models are fetched through a user-started foreground service with a persistent progress notification and model-card progress state. The README records that RWKV-7 World 0.4B and several transformer models completed a full download on a Samsung test device through an HF mirror fallback, and that TinyLlama 1.1B Chat LiteRT pulled a 1.1GB .task file the same way. If your network cannot reach the primary source, that mirror fallback is the path the project documents, not a workaround you have to invent.

The image side is separate and heavier. The README reports the SD1.5 CPU pipeline verified end to end on a Samsung SM-A566E, with txt2img, img2img and inpaint all producing images at roughly 10 to 14 seconds per 512x512 step using MNN with -O3 and a fp16 UNet. Treat that as a single-device figure reported by the project, not a general expectation.

## Where MagicWX stops being the right tool

The most obvious limitation is the modality gap. The README lists ASR, VAD, TTS, image understanding and VLM as separate roadmap items. If your use case needs speech input or a model that reads a screenshot, this app does not do it today, and the roadmap entry is a statement of intent rather than a schedule.

The second limitation is the verification bar itself. It is a strength for trust and a cost for coverage. Models that would work fine for a narrow task are hidden because they failed a general check. SmolLM2 135M is hidden for abnormal output, and MobileLLM 125M for producing webpage fragments. A researcher who wants to study exactly why a 135M model misbehaves on a phone cannot reach it through the visible model list.

The third is hardware. The image pipeline is GA on CPU. The README states that NPU and QNN pipelines for SD1.5 NPU, SDXL, Anima and an upscaler are registered and will open after QNN runtime integration. Registered is not the same as working. If your plan depends on NPU acceleration, the integration is not there yet, and the CPU figure of 10 to 14 seconds per step is what the project reports today.

Finally, this is a prototype with no homepage, and the README gives no public API stability promise. The package namespace, com.qihao.open.rwkv, still reflects the RWKV origin rather than the current scope, which is a small signal that the surface is still moving.

## How MagicWX differs from llama.cpp and MLC LLM on Android

The closest alternatives are llama.cpp with a GGUF model on Android and MLC LLM. The difference is in where the runtime work sits.

llama.cpp ships its own inference engine and its own quantization format. You build or embed the native library, load a GGUF file, and the model format and the runtime come from the same project. MagicWX does not do that. It treats the runtime as an interchangeable adapter behind ModelRuntimeAdapter, and the README lists GGUF paths as candidates that need a native llama.cpp adapter before they can be used. In other words, MagicWX is closer to a host application that can dispatch to several runtimes than to an inference engine. That is why the same app can list an ONNX q4f16 model, a MediaPipe .task model and a bundled RWKV vocabulary model side by side.

MLC LLM takes a third approach, compiling models ahead of time for a target backend. MagicWX instead registers candidates with metadata and opens them only after a device-level check. The practical consequence: with llama.cpp you pick a model and a quantization and you are responsible for whether it runs. With MagicWX, the project has already decided for you, and the price of that decision is a shorter list. If you want breadth, the engine route gives you more. If you want a worked example of gating and dispatch on Android, MagicWX is the more instructive codebase.

## Licence, releases and what an upgrade costs

MagicWX is MIT licensed, with the LICENSE file at the repository root. MIT is permissive: it allows use, modification and redistribution provided the copyright notice and permission notice are retained. That matters here because the app bundles or downloads third-party model weights and runtimes. The MIT licence covers the project's own code, not the models it loads. The README itself flags this for at least one entry, noting that Llama 3.2 1B ONNX is a direct candidate requiring license and runtime validation. If you redistribute an APK with weights inside it, the model licences are your problem, not the MIT grant. This is a general observation about how model and code licences differ, not legal advice.

The release history is short and recent. v1.1.1 arrived on 2026-07-31, v1.1.3 on 2026-08-02 with the Android 17 adapter-gated prototype label, and v1.2.0 on 2026-08-07 described as the LocalDream NPU on-device image generation release. The README's current status section names v1.1.5 as the latest release, which does not match the release list, so treat the README status block as possibly ahead of or behind the published tags and check both. The last push to the repository was on 2026-09-06.

Upgrade cost is mostly model-side. Because visibility is driven by adapter metadata and device validation, a version bump can change which models appear without changing the UI. If you fork the app, expect to re-run the validation path for any model you add, since the README's bar requires a complete package manifest, required assets, a successful load, a fixed-input dry-run and device verification evidence. The test_automation.sh script at the repository root is where that kind of check would live.

## Conclusion

MagicWX is worth adopting if you want a working reference for Android-side runtime dispatch, model package verification and a foreground-service download flow, and you accept that the visible model list is short because the project hides anything that failed device validation. It is the wrong choice if you need multimodal input, a stable API, or a shipped product rather than a prototype. Before building on it, check the app module for the adapter load gates, confirm which ONNX and MediaPipe artifacts your target device can load, and read the README's validation bar, which requires a complete package manifest, required assets, a successful load, a fixed-input dry-run and device verification evidence before a model is described as supported.

## FAQ

### What is MagicWX?

MagicWX is an Android local AI app built with Kotlin, Jetpack Compose and Material3 that combines fully offline text chat through RWKV, ONNX and MediaPipe runtimes with on-device Stable Diffusion image generation using an SD1.5 CPU pipeline on a native MNN backend.

### Which models does MagicWX actually support?

The app exposes only models that passed device validation: the built-in experience model, RWKV-7 World 0.4B, Qwen3 0.6B, Qwen2.5 0.5B, SmolLM2 360M, and TinyLlama 1.1B Chat LiteRT. The README states that a model should not be described as supported until it has a complete package manifest, required assets, a successful load, a fixed-input dry-run and device verification evidence.

### Does MagicWX need an internet connection to run?

Inference is offline once the weights are on the device, but the external models are not bundled. They are downloaded through a user-started foreground service with a persistent progress notification, and the README records that downloads completed on a Samsung test device via an HF mirror fallback.

### How fast is image generation in MagicWX?

The README reports the SD1.5 CPU pipeline verified end to end on a Samsung SM-A566E at roughly 10 to 14 seconds per 512x512 step, using MNN with -O3 and an fp16 UNet. NPU and QNN pipelines are registered but will only open after QNN runtime integration.

### What licence does MagicWX use?

The repository is MIT licensed, with the LICENSE file at the root. That covers the project's own code; the model weights and runtimes it loads carry their own terms, and the README notes that at least one candidate requires license and runtime validation.

## Sources

- [Issues](https://github.com/Pangu-Immortal/MagicWX/issues)
- [License: MIT](https://github.com/Pangu-Immortal/MagicWX/blob/Ai/LICENSE)
- [Pangu-Immortal/MagicWX on GitHub](https://github.com/Pangu-Immortal/MagicWX)
- [README](https://github.com/Pangu-Immortal/MagicWX/blob/Ai/README.md)
- [Releases](https://github.com/Pangu-Immortal/MagicWX/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/pangu-immortal-magicwx
