Model or dataset
jedzqer/manga-translator-android avatar
jedzqer/manga-translator-android

Manga Translator picks its output language from its own interface, and OCRs twice on purpose

安卓手机端的即时自动漫画翻译软件,由LLM驱动。Instant automatic manga translation app for mobile devices, powered by LLM.

472 stars30 forksKotlinMIT

At a glance

What is it?
Manga Translator is a Kotlin app for Android that detects speech bubbles locally, recognises text with local or API OCR, sends the result to an OpenAI compatible endpoint for translation, and draws the translation back over the original page as a draggable bubble. Two design choices are worth reading before you install it: the target language is decided by the app's interface language rather than by a setting, and the build ships five model files into assets that a source build has to supply by hand.
Who is it for?
Manga Translator fits a reader who has a folder of images in reading order, an OpenAI compatible endpoint, and a habit of reading on a phone rather than a laptop, and who wants to fix a mistranslation by dragging the bubble. It does not fit someone who needs batch translation of a thousand pages without watching it, since long runs are advised to be split or given a longer timeout.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 2 days ago.
What is it written in?
Mainly Kotlin, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Detect, OCR, translate, overlay

The pipeline has four stages and the interesting boundary is which of them need a network.

Bubble detection is local. The app finds the speech bubbles and the loose text on the page with a bundled model, which is what lets it draw a box in the right place later. Recognition is configurable: you can use local OCR, or point the app at any OpenAI compatible OCR endpoint with an address, a key and a model name. Translation is the only stage that requires a language model, and it goes through the same kind of OpenAI compatible interface.

The output is an overlay rather than a new image. The translated text is drawn on top of the original page, and the bubble positions are draggable, so a mistranslation or a bad line break is corrected by moving a box rather than by regenerating a page and losing the art. The reading view also keeps the zoom in step between the webtoon mode and the normal mode, so a page does not jump when the reader switches.

The same four stages are available outside the library. The screen translation mode puts a floating window over any app or the home screen, detects and translates whatever manga text is on the current screen, and it shares one font configuration with the normal bubble overlay, which is why a custom bubble font applies to both places.

The stated purpose of the cancellable key and the plus and minus controls in the reading view is fine adjustment, which is the same theme as the drag: the translation is a draft laid over a fixed page.

The interface language decides the output language

There is no target language picker, and that is a deliberate simplification with a consequence worth understanding before you install it.

The target is derived from the app's own interface language. A Simplified Chinese interface produces Simplified Chinese, a Traditional Chinese interface produces Traditional Chinese, an English interface produces English, a Russian interface produces Russian, and a Brazilian Portuguese interface produces Brazilian Portuguese. Five mappings, and no setting to override them.

The source language is separate and is set per folder in the library, so a library can hold a Japanese series, a Korean series and a Chinese one and translate each correctly. The list of source languages is Japanese, English, Korean, Simplified Chinese, Traditional Chinese, mixed Chinese and English, French, Spanish, Portuguese, German, Italian and Russian. Reading the two lists together gives the actual coverage: the first set of source languages is translated into Chinese, and Chinese in turn is translated into English or Russian.

There is a related prompt detail. When the interface is switched to Traditional Chinese, the app prefers Traditional Chinese prompts, so the instruction sent to the model changes with the interface rather than only the output language. A prompt is part of the translation quality, and a tool that lets you pick an output language but not a prompt language is making a choice for you.

Per folder is also where the glossary lives, which is the other reason the library is organised this way.

You give it a base URL and it appends the path

The failure mode in the FAQ is the one every user hits, so the answer is specific about which string goes in which field.

A failed translation or an empty result means the API address is wrong. The address has to be the provider's OpenAI compatible base address, with the two examples given being `https://api.deepseek.com/v1` and `https://open.bigmodel.cn/api/paas/v4`. The app appends `/chat/completions` itself, so pasting the full completions path is the mistake. The model name has to match what the provider calls that model, and the network has to be able to reach it.

The question of which provider to use gets a one line answer: go and find out how to get access to one. That is not a brush off so much as a statement of the design, since the app speaks the OpenAI compatible shape and nothing more, so any provider with that endpoint is interchangeable.

The other FAQ entry is about translation coming out in the wrong order, and the answer is not a setting: rename the images so the filenames follow the reading order, for example 1.jpg then 2.jpg. The app uses the filename order and has no other way to know which page is which, which is the one piece of manual work the whole flow requires.

For long runs there is a piece of advice rather than a feature: when a work has many pages, upload in batches for the whole text quick translation mode, or raise the API timeout in settings.

A 1472 pixel dual label detector, and PaddleOCR for the text

The models section is the part of this repository that a source build has to get right by hand, and the file names are the documentation.

Detection uses `models/detection/mixed-dual-s-e5_float16.tflite`, described as a dual label segmentation model for bubbles and loose text, a YOLO26s-seg at 1472 by 1472 in TFLite float16. This one is described as shipping with the app assets, so a normal install already has it.

Recognition is PaddleOCR, and the split by language is the interesting part. `models/ocr/PP-OCRv6_small_rec.onnx` handles Japanese, English, Chinese and mixed Chinese and English. Korean has its own pair, `models/ocr/korean_PP-OCRv5_mobile_rec.onnx` plus a matching dictionary file, `korean_PP-OCRv5_mobile_rec_dict.txt`, which is a reminder that a recogniser is a model and a character table rather than a single file.

Text line detection inside an OCR region is `models/detection/PP-OCRv6_det_mobile_infer.onnx`, and it appears twice in the list, once described as Paddle text line detection inside the OCR region and once as English line detection. The two descriptions are of one path, so a builder reading the list top to bottom will look for two files and find one. The download links for the recognition models, the English detector and the Korean recogniser are given as Hugging Face paths under the PaddlePaddle organisation.

Prompts, fonts and the OCR configuration also live in assets subdirectories, with the note that the names have to match what the code expects.

A release touches three files, and the data lives beside the images

Two small sections of the README describe the parts of the project that are not features, and both are more useful than the feature list.

The first is release version synchronisation. A version bump has to change three files together: the Kotlin source holding the version, the app build file, and a JSON file named `update.json` at the repository root. Two of those are obvious; the third is the update feed, which is the same file that makes the app's startup update check work. A release that updates only the build file would ship an APK whose internal version disagrees with what the updater advertises.

The second is where the data goes. The library lives under the app's external files directory, at `/Android/data/<package>/files/manga_library/`. Each image gets a same named JSON file with the translation result, and the OCR result is cached separately in a file with an `ocr` extension in the middle of the name. The glossary is one `glossary.json` per folder, and the reading progress and the quick translation switch live in SharedPreferences.

That two file split is the design point worth noticing. Keeping the OCR result separately from the translation means a re-translation with a different model does not have to re-run recognition, and recognition is the slow, local, expensive half. On a volume of two hundred pages that is the difference between a minute and an hour.

The build itself needs JDK 17.0.17 or newer, Kotlin 2.0.0 or newer, Gradle 8.11.1 or newer, and Android SDK platform 36 with build tools 36.0.0, and produces either flavour through the Gradle wrapper:

bash
./gradlew :app:assembleDebug
./gradlew :app:assembleRelease

CBZ, ZIP and PDF, with a reader that guesses the format

The library side of the app is about getting files in and keeping track of where you were, and the reader side is about the two very different shapes a comic page comes in.

The library supports creating folders, importing images in batches, importing whole manga folders, and importing or exporting CBZ, ZIP and PDF archives. So a volume you already have as an archive is not a special case, and the export side means a translated volume can leave the app in a form your other reader understands.

The reader guesses the shape. The app decides whether a work is closer to a webtoon, the long vertical strip format, and switches its reading mode accordingly, and in the long image and webtoon modes it merges bubbles across page boundaries so a sentence split over two panels is not lost at the join. Zoom stays in step between the two modes, which is the part that makes the switch bearable.

Two smaller conveniences close the loop. The translated text sits in draggable bubbles and the progress saves itself as you read, so closing the app mid volume does not cost you your place. And when a folder or batch translation finishes, the app sends a high priority notification with sound that takes you back to the library, which is what makes it reasonable to start a batch and go do something else.

Editorial conclusion

Manga Translator fits a reader who has a folder of images in reading order, an OpenAI compatible endpoint, and a habit of reading on a phone rather than a laptop, and who wants to fix a mistranslation by dragging the bubble. It does not fit someone who needs batch translation of a thousand pages without watching it, since long runs are advised to be split or given a longer timeout. Before you start, name your files in reading order, decide whether local or API OCR suits you, and read the glossary behaviour, because per folder naming of characters is how you stop the same name being translated three different ways across a volume.

Frequently asked questions

Which manga translator is best?

That is a preference question, but the Manga Translator app on Android is specific about what it covers: local bubble detection, local or API OCR, LLM translation through any OpenAI compatible endpoint, draggable translation overlays, screen and floating window translation, CBZ, ZIP and PDF import and export, a per folder glossary, and a source build that needs JDK 17, Kotlin 2 and Gradle 8.11.1.

Is there a way to auto translate manga?

Yes, that is the whole app. It detects bubbles in the image, recognises the text with local or API OCR, sends it to a language model, and draws the translation back over the original page as a bubble you can drag. You create a folder, import images named in reading order, choose local or API OCR in settings, and tap translate folder.

Why is the translation empty or failing in Manga Translator?

The API address has to be the provider's OpenAI compatible base address, for example https://api.deepseek.com/v1 or https://open.bigmodel.cn/api/paas/v4, because the app appends /chat/completions itself. The model name must match what the provider calls it, and the endpoint has to be reachable. Out of order translations are a different problem and are fixed by renaming the images in reading order.

Which languages can Manga Translator translate into?

The target language comes from the app's interface language rather than from a picker: a Simplified Chinese interface produces Simplified Chinese, Traditional Chinese produces Traditional Chinese, English produces English, Russian produces Russian, and Brazilian Portuguese produces Brazilian Portuguese. The source language is set per folder and covers Japanese, English, Korean, both Chinese scripts, mixed Chinese and English, French, Spanish, Portuguese, German, Italian and Russian.

What does building Manga Translator from source need?

JDK 17.0.17 or newer, Kotlin 2.0.0 or newer, Gradle 8.11.1 or newer, and Android SDK platform 36 with build tools 36.0.0, producing a build with the Gradle assemble tasks. The model files have to be placed in the assets subdirectories, including the YOLO26s-seg detector at 1472 by 1472, the PaddleOCR recognisers for the general and Korean cases, and the Korean character table, and a version bump has to be applied to three files at once.

Official sources

  1. Issues
  2. jedzqer/manga-translator-android on GitHub
  3. License: MIT
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/jedzqer-manga-translator-android.svg)](https://hysenlabs.com/projects/jedzqer-manga-translator-android)