# TTS Server for Android: a network speech bridge that keeps working when the built-in engine fails

> A Kotlin and Compose application with 4,500 stars that exposes a Microsoft Edge read-aloud style endpoint on your phone, ships a configurable HTTP request for any service, can import other local TTS engines, and reads Chinese narration from punctuation alone.

**jing332/tts-server-android** — 这是一个Android系统TTS应用，内置微软演示接口，可自定义HTTP请求，可导入其他本地TTS引擎，以及根据中文双引号的简单旁白/对话识别朗读 ，还有自动重试，备用配置，文本替换等更多功能。

- Repository: https://github.com/jing332/tts-server-android
- Stars: 4,508 · Forks: 430
- Language: Kotlin
- License: not declared
- Published: 2026-10-07 · Updated: 2026-10-07 · Language: en
- Canonical page: https://hysenlabs.com/projects/jing332-tts-server-android

## A phone that answers text to speech requests over HTTP

The project description is written in Chinese and is worth reading literally, because it describes a bridge rather than a synthesiser. It is an Android system TTS application with a built-in Microsoft demo interface, support for custom HTTP requests, the ability to import other local TTS engines, and a simple narration and dialogue detection based on Chinese double quotation marks. Automatic retries, fallback configurations and text replacement are listed as further features.

The Microsoft demo interface is the part that makes it interoperable. Rather than inventing a protocol, the application presents an endpoint shaped like the Edge read-aloud service, which means a client written against that service can point at your phone instead. That is the whole architectural bet: reuse a widely implemented request and response shape so the number of clients you have to modify is close to zero.

The scale suggests it works. The repository has 4,508 stars and 430 forks, which is substantial for a utility of this kind, with 92 open issues. It is written in Kotlin, the default branch is `compose`, and the last push was 2026-09-24.

The description also states the limits of the cloud path plainly: whether a particular voice works depends on that service being configured and reachable, and a bundled interface described as a demo interface should not be treated as a permanent free service. The local engine import path exists precisely because relying on a remote demo endpoint is not a plan you want to build on permanently.

## Configuration, retries and fallbacks are the actual feature

Once the endpoint shape is settled, the project's depth shows up in the configuration story, which is where most of the listed features live.

Custom HTTP requests are the general escape hatch. Instead of the built-in interface, you describe the request to a service you control, which is how you point the application at a self-hosted synthesiser, a different vendor, or a local server on your network. Automatic retry handles the case where a request fails transiently, which matters when the phone is on mobile data and the endpoint is not. Fallback configuration covers the harder case: when the primary configuration cannot serve a request, another one takes over, so a voice failure does not end the read-aloud session.

Text replacement is the quiet feature that matters most in practice. Novels, documentation and screenplays all contain abbreviations, symbols and transliterated words that a synthesiser mispronounces, and a global replacement list applied before synthesis is the cheapest way to fix them without editing the source text.

The narration detection feature is specific to Chinese. The application uses double quotation marks in the source text to distinguish narration from dialogue, and reads them with different voice treatment. That is a simple heuristic, but it is the kind of thing that is genuinely annoying to do by hand across a long book, and pairing it with a scriptable rule system is what the 0.6 release notes describe as replacing an earlier built-in version.

One further capability is worth calling out because it is easy to miss: callers can specify which pronunciation configuration to use through the API, rather than the application always applying its default. For a reading app that serves several books with different narrator voices, that removes a round of manual switching.

## Reading rules turned into a scripting layer

The repository name mentions a server, but the feature that most distinguishes it from a thin endpoint wrapper is the reading rules system, which is scriptable.

Two dependencies explain how that works. `gedoor/rhino-android` provides RhinoScriptEngine access through the JSR 223 interfaces on the Android JRE, which means rules can be written as JavaScript and evaluated on device. `Rosemoe/sora-editor` provides the code editor, so editing a rule on a phone is a supported activity rather than something you do in a desktop text editor and push over.

The combination turns voice settings into a programmable surface. A rule can branch on the content of a line, apply different voices for narration and dialogue, handle the punctuation conventions of a particular language, or preprocess text before it reaches the synthesiser. That is a meaningfully different design from a screen of dropdowns, and it explains the presence of a code editor and a scripting engine in a TTS application.

The 0.6 release notes describe the introduction of this capability and warn in the same breath to export a configuration backup before experimenting, because a bad rule can crash the application and the warning also asks you to clear app data if it does. The 0.9 notes show the maturing of the feature set: rules can be exported individually, the caller can choose the pronunciation configuration via the API, and a fix landed for local TTS engines not previewing in the editor screen.

Importantly, the rules layer and the engine layer are separate. A rule describes what should happen to a piece of text, and the built-in interface, a custom HTTP service, or an imported local engine decides who actually speaks it.

## A Kotlin and Compose codebase split by responsibility

The repository tree is a Gradle multi-module Android project and the module names describe the architecture more clearly than any document would.

`lib-tts` is the speech side: interfaces, voices and the engines that implement them, which is the module that has to change when you add a new service. `lib-server` is the HTTP server that exposes the endpoint, the module that defines the request and response contract every client depends on. `lib-script` holds the reading rules and the scripting engine integration. `lib-database` covers persistence, and `lib-common` holds what the rest of the project shares. `lib-compose` provides the shared UI components, and `app/` is the application module tying them together.

The split is sensible because it puts the compatibility-sensitive part, the HTTP contract, in its own module behind the speech interface. Adding an engine should not require touching the server, and the presence of a dedicated scripting module confirms that rules are meant to be a platform rather than a feature of one screen.

The build uses a version catalog, with dependency versions collected in `libs.versions.toml`, and there is a separate `build-logic/` module, which is the convention for encapsulating convention plugins. Gradle wrapper scripts are present, so a clone builds without installing Gradle yourself. `crowdin.yml` is the localisation configuration, matching the Crowdin badge in the README, so translation runs through Crowdin rather than through hand-edited resource files.

Two GitHub Actions workflows are declared as badges. One is the release workflow and one is the test workflow, so the CI story is visible from the README even though the details live in the workflow files.

## Releases stopped in 2023 while development continued

The single most important practical fact about this project is the gap between its release feed and its development activity.

The most recent tagged release is version 0.9, published 2023-07-26. Before that came 0.8 in June 2023 and 0.6 in mid-June 2023. Nothing has been tagged in the more than three years since, while the last push to the repository was 2026-09-24. Those two facts are not in conflict, because the README explains the distribution model: stable builds come from the releases page, and development builds come from the Actions artifacts instead of from tags.

The README also lists a QQ group number and a set of Actions mirrors on a third-party file host, with a password given in the text. Mirrors exist for a reason, which is that GitHub Actions artifacts require a signed-in GitHub account to download. That is an unusual and slightly awkward arrangement, and it means the newest code is genuinely harder to obtain than the older tagged release.

Open issues sit at 92, which is a high absolute number against 4,508 stars and points to a project with a steady user base and an active support load. The CHANGELOG file is present in the tree, so per-change detail is available even when the tag history is not.

On licensing, the README displays an MIT badge at the top while the repository tree contains no licence file. The metadata reports no licence. An MIT badge in a README is a statement of intent by the author and is widely treated as sufficient, but if your use involves redistribution or embedding, it is worth asking or waiting for a licence file rather than assuming one.

## Conclusion

The problem this application solves is narrow and real: many text to speech tools, readers and automation scripts expect an HTTP endpoint, and Android does not provide one. TTS Server supplies that endpoint from a phone you already own, and because it exposes a Microsoft-style read-aloud interface out of the box, existing consumers often work without modification. What you give up for that convenience is visible in the numbers. The latest tagged release is 0.9 from 2023-07-26 while the last push was 2026-09-24, so the release feed is not where current work shows up, and there are 92 open issues against 4,508 stars. The repository also presents an MIT badge in the README without a licence file in the tree, which is a small inconsistency worth confirming before you redistribute anything. Treat it as an actively developed personal tool rather than a versioned library, and check the dev builds if you need current fixes.

## FAQ

### What problem does TTS Server for Android solve?

It exposes text to speech over HTTP from an Android device. The application presents a built-in Microsoft read-aloud style interface by default, and supports custom HTTP requests so you can point it at a different service or at something you host yourself, with automatic retries and fallback configurations when a request cannot be served.

### Can I use a local TTS engine instead of the built-in interface?

Yes. Importing other local TTS engines is one of the listed capabilities of the project. The speech and engine layers are separated from the server layer in the module structure, so a locally provided voice and a remote HTTP service are alternatives that serve the same interface.

### How does the app decide what is narration and what is dialogue?

For Chinese text, it uses the double quotation marks in the source to distinguish narration from spoken dialogue, and reads the two with different treatment. This heuristic became scriptable reading rules, which replaced an earlier built-in version and can be edited on the device through an embedded code editor using a Rhino scripting engine.

### Where do I get the newest version, given that the latest tag is from 2023?

The last tagged release is version 0.9 from 2023-07-26, but the repository continues to receive pushes, with the most recent on 2026-09-24. The README says stable builds come from the releases page while development builds come from GitHub Actions artifacts, which require a signed-in GitHub account, and it also lists third-party mirrors with a password in the README text.

### Is TTS Server for Android open source?

The source is published and the README displays an MIT licence badge, but the repository tree contains no licence file and the repository metadata reports no licence. The build scripts, module layout, Crowdin localisation setup and GitHub Actions workflows are all in the repository, so the code is available to read and build. Confirm the licence position directly with the author before redistributing.

## Sources

- [Issues](https://github.com/jing332/tts-server-android/issues)
- [jing332/tts-server-android on GitHub](https://github.com/jing332/tts-server-android)
- [README](https://github.com/jing332/tts-server-android/blob/compose/README.md)
- [Releases](https://github.com/jing332/tts-server-android/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/jing332-tts-server-android
