Model or dataset
Skythinker616/gpt-assistant-android avatar
Skythinker616/gpt-assistant-android

GPT Assistant for Android: a volume-key GPT client with agent mode

【新增智能体模式】安卓端全场景GPT助手,可用音量键唤起并进行语音交流,支持联网、拍照、模板、附件解析、智能体模式等 | GPT assistant for Android, activated via volume keys for voice interaction, supporting features such as networking, taking photos, templates, parsing PDF and Office documents, and agent mode.

906 stars122 forksJavaGPL-3.0

At a glance

What is it?
An open source Android app that puts GPT behind the volume keys, handles photos and Office documents, and can drive the phone through Android accessibility. Here is what the repository documents, and where it stops.
Who is it for?
Adopt it if you want a system-wide GPT entry point on Android that reads PDF, DOCX, PPTX and XLSX files and does not lock you into one vendor, provided you are willing to hold your own OpenAI API key and grant accessibility permissions. Skip it if you need a stable automation tool: agent mode is labelled experimental, and the README warns that some apps do not expose full control information.
Can I use it commercially?
Yes, with conditions. GPL-3.0 is a copyleft licence: if you distribute software that includes it, you must release that software's source code under the same licence. Running it internally without distributing it does not trigger that obligation.
Is it still maintained?
Yes. The repository last received commits 149 days ago.
What is it written in?
Mainly Java, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem: getting GPT out of its own app

Most Android chat clients are destinations. You open the app, type, wait, and leave. GPT Assistant inverts that. The README describes it as a "全场景 GPT 助手", a whole-scene assistant, and the entry points are the point: a long press of the volume down key, a quick-settings tile labelled GPT, and a text-selection menu item that appears inside other apps. The intended user is someone who wants to ask a question while reading a WeChat message or looking at a photo, without switching context.

The second problem it targets is input. Typing on a phone is slow, and pasting a 40-page PDF into a chat box is impossible. The app accepts files through five different routes according to the README: an in-app attachment button (camera, gallery, file picker), the system share sheet, the "open with" dialog, drag and drop from apps that support it, and text selection. Supported formats are images, TXT, PDF, DOCX, PPTX and XLSX.

That combination, a global hotkey plus document ingestion, is a narrower and more opinionated product than a generic ChatGPT wrapper. It also explains why the project needs accessibility permissions at all, which is the part that will decide whether you can use it.

How the pieces fit: accessibility service, WebView, ASR modules

The repository is a Gradle Android project with a single app module and a separate asr_core module, which holds the speech recognition backends. Four ASR providers are selectable in settings: Huawei HMS (the default), Baidu, Google and Whisper. The README states the Huawei key is bundled so the default works without configuration, and that a self-compiled build must supply its own key because the shipped one is bound to a package name and signature.

Networking is not a separate HTTP client bolted on. The README says the app uses Android WebView to load pages, and exposes that through the OpenAI Function interface. GPT decides to fetch a URL, the app loads it in WebView, extracts text, and returns it. For a list of adapted sites (Baidu, Bing, Google, Google Scholar, Zhihu, Weibo, JD, GitHub, Bilibili, CNKI) the extraction also returns result links. Everything else yields plain text only. This design inherits the phone's browser session, so pages that need a login can sometimes be read, and it also inherits the browser's failure modes.

Agent mode uses the accessibility service differently. Instead of screenshotting the screen, the README says the AI reads control information exposed by the accessibility tree and then performs clicks, scrolls and text input. That is why it works with any model that supports tool calling rather than only vision models, and it is also why it fails on apps that do not expose their controls.

Installing the APK and getting a first answer

There is no Play Store listing and no build step required for normal use. The README says to download the APK from the latest release and install it. The current release line is v3.0.0, published on 2026-04-19.

The app ships without credentials. You supply an OpenAI API key in settings, along with the base URL. The README gives the official endpoint, `https://api.openai.com/`, and notes that third-party relay services can be substituted when the official one is unreachable. It also points at GPT_API_free as one relay that offers `gpt-3.5-turbo` and `gpt-4o-mini`.

Once the key is in place, the volume-key flow is the fastest path to a result. The README describes it in four steps: long press volume down to show the interface, hold the volume key to record, release and short press again to send, and the reply is spoken automatically.

If the phone vibrates but nothing appears, the README points at one cause: the "background pop-up interface" permission is missing. That permission is not present on every ROM, and the README says to ignore it when it is absent.

Voice input is the weakest documented link

The README's own comparison of the four ASR backends is unusually candid, and it should shape your expectations more than the feature list does. Huawei's real-time recognition is described as the most accurate for single sentences and for mixed Chinese and English. Baidu handles long sentences well with sensible segmentation but cannot do mixed Chinese and English because the selected model is Mandarin. Google covers many languages, performs only adequately on Chinese, and adds no punctuation. Whisper covers many languages and does reasonably on Chinese, but the README notes that simplified and traditional characters can drift unpredictably, and that it does not produce output while you are still speaking.

The cost picture is equally uneven. Huawei and Google are free, with the caveat that Google is generally unreachable from mainland China. Baidu gives new users 150,000 calls or 180 days, then charges per call. Whisper bills through the same OpenAI account as chat. If you switch to Baidu, the README adds a configuration detail that is easy to miss: the application's speech package name must be set to "Android" and the app package name `com.skythinker.gptassistant` must be entered, and the long-speech toggle decides whether you are using the short or the real-time recognition service.

None of this is a bug. It is the cost of not shipping a model. But it means the perceived quality of the assistant is decided in a settings screen most users will never open.

Where it breaks, and where the documentation goes quiet

Three failure modes are documented, and all three are structural rather than fixable by configuration.

Web access depends on WebView loading a live page. The README states that a 15-second timeout, a login wall or a CAPTCHA can all produce "cannot get content", and that GPT may simply pick a wrong URL on its own. It also warns that sending page text to the model burns a lot of tokens, and that `gpt-4` should be used carefully for that reason. Models without Function support, such as `gpt-4-vision-preview`, cannot use networking at all.

Agent mode is labelled experimental in the README, and the caveats are explicit: not every app or control is recognised reliably, and a model with strong agent capability is recommended. The README also advises against using it on payment, verification-code and password screens. That is a sensible warning, but it is advice, not a technical guard: the documentation does not describe a mechanism that blocks those screens.

Interoperability is deliberately out of scope. Asked about Claude, Gemini and Ollama endpoints, the README says the project has no plan to adapt to each one and points users at OneAPI or NewAPI to convert them into OpenAI format. If you run a local model server, that extra hop is on you.

Two things the README does not cover: there is no documented rollback or downgrade path between releases, and no statement about what happens to the accessibility service when the app is force-stopped.

Alternatives and the real difference

The obvious comparison is the official ChatGPT Android app. It is maintained by OpenAI, needs no API key, and has none of the permission friction. The difference is reach: it is an app you open, not a layer over the system. It does not wake on a volume key, it does not accept a DOCX through the share sheet, and it does not read the accessibility tree of other apps. If your use is "open an app and chat", the official client wins on every axis that matters, including stability.

A closer comparison is TTS Server, which the README itself recommends as a replacement for the system text-to-speech engine. That project addresses one slice, speech output, and does it with more configurability than GPT Assistant's single "open system speech settings" route. GPT Assistant treats TTS as a system dependency and hands you off to Android's settings. If voice quality is your main concern, the two projects are complementary rather than competing.

For anyone who wants to route Claude, Gemini or a local model through this interface, OneAPI and NewAPI are the intended answer, and they are a different kind of tool: a gateway, not a client. The honest framing is that GPT Assistant is a front end. It bets on OpenAI's API shape and treats every other provider as an adapter problem for someone else.

Licence, maintenance and what an upgrade costs you

The project is GPL-3.0. For an end user installing an APK this changes nothing. For anyone embedding the code in a closed product it is decisive: GPL-3.0 requires derivative works distributed to others to be licensed under the same terms, and the bundled Huawei credentials cannot simply be reused because the README states they are bound to a package name and signature. A fork that keeps the package name is not a clean path either. This is a description of the licence, not legal advice; read the LICENSE file before shipping anything.

The last push to the default branch was on 2026-05-05, and the repository is not archived. Releases are infrequent rather than continuous: v2.0.0 in April 2025, v2.0.1 a fortnight later, then v3.0.0 in April 2026. That cadence suggests a project that ships when a feature set is ready, not one that patches weekly. Budget accordingly: if you depend on agent mode, you are depending on a feature the README itself calls experimental, and the gap between releases means a fix may not arrive quickly.

Upgrade cost is low in the ordinary case. Settings live in the app, the API key is yours, and there is no server component to migrate. The one upgrade hazard is the Huawei ASR key: if the bundled credential stops working, the README's own answer is to configure Baidu, Google or Whisper instead, or to build from source with your own AppGallery credentials.

Editorial conclusion

Adopt it if you want a system-wide GPT entry point on Android that reads PDF, DOCX, PPTX and XLSX files and does not lock you into one vendor, provided you are willing to hold your own OpenAI API key and grant accessibility permissions. Skip it if you need a stable automation tool: agent mode is labelled experimental, and the README warns that some apps do not expose full control information. Before installing, confirm that your phone allows background pop-up windows, that your chosen model supports Function calling if you want networking or agent mode, and that you have a reachable API endpoint if you are outside the regions OpenAI serves directly.

Frequently asked questions

How do I install GPT Assistant on Android?

Download the APK from the latest release and install it, as the README instructs. There is no Play Store listing and no build step needed for normal use. After installing, open settings and enter your OpenAI API URL and API key.

Does GPT Assistant for Android need an API key?

Yes. The app uses the OpenAI API, and the README says the user fills in their own API_KEY in settings, either against the official service or a third-party relay. The speech recognition default is separate and works without configuration.

Can GPT Assistant control my phone?

Agent mode lets the AI read controls exposed by the accessibility service and perform clicks, scrolling and text input, and the README describes sending messages and ordering food as examples. It is labelled experimental and the README warns that some apps do not expose full control information and that payment, verification-code and password screens should be avoided.

How do I use GPT Assistant as an assistant while using other apps?

Enable the volume-key accessibility service and allow the app to run in the background, then long press volume down to bring up the interface. The README also lists a quick-settings tile and a text-selection menu entry as triggers.

Official sources

  1. Issues
  2. License: GPL-3.0
  3. README
  4. Releases
  5. Skythinker616/gpt-assistant-android on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/skythinker616-gpt-assistant-android.svg)](https://hysenlabs.com/projects/skythinker616-gpt-assistant-android)