zyddnys/manga-image-translator: self-hosted OCR, inpainting and typesetting for manga pages
Translate manga/image 一键翻译各类图片内文字 https://cotrans.touhou.ai/ (no longer working)
At a glance
- What is it?
- A GPL-3.0 Python pipeline that detects text in an image, erases it, translates it and redraws it in the original lettering. It is aimed at untranslated group-chat and imageboard pages, and its own README calls it an early-stage project with many shortcomings.
- Who is it for?
- Adopt it if you can run PyTorch locally or in Docker and want a scriptable, self-hosted pipeline with a documented config file, glossary and replacement dictionary, and if you accept GPL-3.0 terms for whatever you distribute. Do not adopt it if you need a maintained service, a supported browser extension, or a promise that the hosted cotrans.touhou.ai endpoint will answer.
- Can I use it commercially?
- Yes, with conditions. GPL-3.0 is a copyleft licence: if you distribute software that includes it, you must release that software's source code under the same licence. Running it internally without distributing it does not trigger that obligation.
- Is it still maintained?
- Yes. The repository last received commits 72 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 17, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What manga-image-translator actually does, and who it is for
The README states the goal plainly: translate images that are unlikely to be professionally translated, such as comics and images posted in group chats and image boards, so that a reader who does not know Japanese can follow the content. The primary target is Japanese, with Simplified and Traditional Chinese, English and about 20 other minor languages supported. Alongside translation the project handles image repair, meaning removal of the original text, and typesetting, meaning putting the translated text back onto the page.
The intended user is not a localization studio. It is someone with a folder of pages, a machine with a GPU or a willingness to wait on CPU, and no expectation of publication-quality lettering. The README is explicit that the project is still in an early stage of development and has many shortcomings, and it asks for help improving it. The release history backs that up: the newest release listed is beta-0.3 from 2022-04-23, while the repository itself has continued to receive commits, with the last push on 2026-07-20. Releases and code have moved at different speeds.
The homepage is https://cotrans.touhou.ai/, and the repository description notes the online version is no longer working. Treat the hosted service as gone and the local code as the product.
The pipeline: detection, OCR, translation, inpainting, rendering
Each stage is a separately configured component, which is the main architectural fact about this project. The configuration file exposes sections for render options, upscale options, translator options, detector options, inpainter options, colorizer options and OCR options. The detector finds text regions and produces a mask. The OCR component reads the characters inside those regions. The translator converts the recognized text into the target language. The inpainter fills the masked region so the original lettering disappears. The renderer draws the translated text back, and the colorizer is an optional pass for pages where the mask leaves flat or damaged areas.
Because the stages are swappable, quality is not a single number. The README's showcase marks one output with --detector ctd and another with --translator none, which shows both that detector choice changes the mask and that you can run the pipeline purely as a text eraser and typesetter when you want to supply your own translation. The README also warns that the showcase examples may not be frequently updated and may not represent the effect of the current main branch.
Translation backends are pluggable too. requirements.txt pulls in deepl, openai, google-genai, groq, ctranslate2, transformers, sentencepiece and manga-ocr, so the same pipeline can call a hosted API or run a local model. A glossary and a replacement dictionary are documented, which is the practical lever for keeping character names consistent across pages.
Installing manga-image-translator locally with pip and venv
The README recommends the Pip/venv route for a local setup and has a separate notes section for Windows users. The package metadata pins the interpreter range, so check your Python before anything else. pyproject.toml declares requires-python = ">=3.10, <3.12", which rules out 3.12 and newer.
pip install -r requirements.txtOn Windows the README points to run.bat and MangaStudioMainRun.bat at the repository root rather than a shell script. Once dependencies are installed, run.sh and run.bat are the entry points for the local modes, and the module form works as well because the Dockerfile sets PYTHONPATH="/app" and uses ENTRYPOINT ["python", "-m", "manga_translator"].
Expect the first run to be slow. The Dockerfile runs docker_prepare.py to fetch models into /app/models before the container is usable, and a local install does the same kind of download on demand. Nothing in the README suggests the models ship inside the repository.
A first real run: batch mode, web mode and the API
The README documents four usage modes: local batch mode, web mode with an old UI and a new UI, API mode with its own documentation, and a config-help mode for inspecting options. Batch mode is the one to try first, because it fails loudly on a single directory instead of behind a browser tab.
python -m manga_translatorThe flags follow the README's option tables: --translator selects the translation backend, --detector selects the text detector, and --use-gpu and --verbose are listed under basic options. The README's recommended-options section carries a tips subsection on improving translation quality, and the config-help mode exists specifically so you can list the valid keys instead of guessing them.
For a browser workflow, Docker Compose brings up the front end on port 3000:
docker compose upThe compose file defines a single front service built from the front directory with NODE_ENV: production and the mapping 3000:3000. Running the web server with an Nvidia GPU is a documented variant, and the README also covers using the container as a CLI and building it locally. If you would rather call it programmatically, API mode has its own documentation section, and the requirements list fastapi, uvicorn, python-multipart and websockets, which is consistent with an HTTP and WebSocket server.
Where manga-image-translator falls short
The README's own warning is the first limitation: early stage, many shortcomings. The second is the dead hosted instance. Anyone arriving from an old bookmark to cotrans.touhou.ai has no working service to fall back on, and the repository description says so directly.
There is no browser extension in this repository. The related searches people type include "manga image translator extension" and "manga image translator chrome extension", but the top-level entries are a Python package, a front-end directory, notebooks, and Docker files. The front end is a web UI served on port 3000, not a content-script extension.
Licensing is a real constraint rather than a footnote. pyproject.toml declares license = "GPL-3.0-only" and the repository carries a GPL-3.0 LICENSE file. If you plan to redistribute a modified version, or to ship it inside a product, the copyleft terms apply to what you distribute. That is a factual property of the licence, not legal advice, and it is the kind of thing to raise with counsel before you build on it.
The dependency list is heavy and version-pinned in places: numpy==1.26.4, openai==1.63.0, httpx==0.27.2 with a comment about a blocking change in 0.28.0, pydantic==2.5.0, py3langid==0.2.2. There are separate requirements files for ROCm and XPU, plus a rust wheel index in requirements.txt, which tells you the hardware matrix is wide but not uniform. A failed install is more likely to come from a conflicting pin than from the code. And with the newest listed release dated 2022-04-23, there is no tagged artifact that matches current main; you are tracking a branch.
Alternatives and how their approach differs
The most direct alternative named by this project itself is its Rust version. The README has a Rust Version section, requirements.txt adds an extra index at frederik-uni.github.io/manga-image-translator-rust and the package rusty-manga-image-translator, and the repository topics and search traffic both point at it. The difference is implementation strategy rather than features: the Rust port aims to replace the Python inference path, while this repository is a Python and PyTorch codebase whose Dockerfile starts from pytorch/pytorch:2.5.1-cuda11.8-cudnn9-runtime.
For the OCR stage specifically, requirements.txt includes manga-ocr, a dedicated model for Japanese manga text. A user who only needs recognition, not inpainting or typesetting, can call that model directly and skip the rest of the pipeline. The trade-off is losing the mask, the repair step and the renderer, which is exactly what this project adds on top.
Against a general-purpose vision or translation API, the difference is where the work happens. This project runs detection, OCR and inpainting locally and only sends recognized strings to a hosted translator if you configure one. That keeps page images on your own machine, at the cost of installing PyTorch and downloading models.
Maintenance, upgrade cost and licence obligations
The last push to the default branch was on 2026-07-20, and the repository is not archived, so the code is moving. The release list is a different story: beta-0.3 dates to 2022-04-23, beta-0.2.1 to 2021-07-27 and alpha-v3.0.0 to 2021-05-21. There is no recent tag to pin to, which means an upgrade is a pull of main plus a dependency re-resolve, and a regression can appear without a version number changing.
The pinned dependencies raise the cost of that re-resolve. numpy==1.26.4, openai==1.63.0, httpx==0.27.2 and pydantic==2.5.0 are all fixed versions, so a transitive requirement that moves past them will conflict. The README's changelog link points at CHANGELOG_CN.md, and there is also a CHANGELOG.md, so read the changelog before pulling. The README itself carries a Last Updated line of 2025/05/10, which is older than the last commit, so documentation and code can drift apart.
On licence: GPL-3.0-only means derivatives you distribute must be under the same terms. Running it privately to translate pages you keep to yourself is a different situation from shipping a modified binary or bundling it into a service. The repository gives you the licence text and the SPDX identifier; it does not give you advice, and neither does this article.
Editorial conclusion
Adopt it if you can run PyTorch locally or in Docker and want a scriptable, self-hosted pipeline with a documented config file, glossary and replacement dictionary, and if you accept GPL-3.0 terms for whatever you distribute. Do not adopt it if you need a maintained service, a supported browser extension, or a promise that the hosted cotrans.touhou.ai endpoint will answer. Verify first that your Python version satisfies requires-python = ">=3.10, <3.12", that you have the GPU memory the detector and inpainter you pick actually need, and that your chosen translator backend has working credentials before you batch a directory.
Frequently asked questions
How do I install manga-image-translator?
The README recommends a Pip/venv local setup, and pyproject.toml requires Python ">=3.10, <3.12". Create a virtual environment, then install with pip install -r requirements.txt, or use the Dockerfile and docker-compose.yml if you prefer a container.
How do I use manga-image-translator?
The README documents local batch mode, web mode with an old and a new UI, API mode with its own documentation, and a config-help mode. Batch mode takes an input directory and an output directory, and flags such as --translator and --detector select the stage implementations.
How can I translate Japanese characters from an image?
That is the core pipeline: the detector finds text regions, the OCR component reads the characters, the configured translator converts them, the inpainter removes the original lettering, and the renderer draws the result back. The project mainly supports Japanese, with Simplified and Traditional Chinese, English and about 20 other minor languages.
Is translating manga legal?
The repository does not address the legality of translating or redistributing manga, so this cannot be answered from what it documents. What it does state is that the code is licensed GPL-3.0-only, which governs the software, not the images you run through it.
What is the best manga translator?
That depends on what you need, and the repository does not rank itself against other tools. It positions itself for images unlikely to be professionally translated, and its own README says the project is in an early stage of development with many shortcomings.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/zyddnys-manga-image-translator)