# BallonsTranslator: a PyQt desktop tool that runs OCR, inpainting and LLM translation over comic pages

> BallonsTranslator is a GPL-3.0 Python application from dmMaze that chains text detection, OCR, inpainting and machine translation into one click, then lets you fix the result on a canvas. It is aimed at people who translate manga and webtoons and are willing to tune four separate models.

**dmMaze/BallonsTranslator** — 深度学习辅助漫画翻译工具, 支持一键机翻和简单的图像/文本编辑 | Yet another computer-aided comic/manga translation tool powered by deeplearning

- Repository: https://github.com/dmMaze/BallonsTranslator
- Stars: 5,130 · Forks: 353
- Language: Python
- License: GPL-3.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/dmmaze-ballonstranslator

## What BallonsTranslator actually automates

The README describes the pipeline as four modules: text detection, recognition, erasing (inpainting), and machine translation. The final page quality depends on the combined behaviour of those four, which the project states plainly: the effect depends on how well detection, recognition, erasing and translation each perform. That framing matters, because it means the tool is not a single model you can swap out. It is an orchestration layer over several, and a weak link in any one of them shows up in the rendered page.

The target user is someone translating Japanese or American comics page by page and willing to correct output. The README says the tool supports Japanese and American comics, and that English-to-Chinese and Japanese-to-English typesetting are already optimized, with Chinese sentence segmentation handled by pkuseg. Vertical Japanese-to-Chinese layout is listed as still needing improvement. That last point is the honest boundary of the project: if your source is vertical Japanese and your target is Chinese, the layout half of the problem is not solved yet.

## How the four modules and the LLM layer fit together

Detection finds text regions. Recognition turns them into strings. Inpainting removes the original lettering from the image. Translation produces the target text, and the renderer puts it back using an estimate of the original layout: color, outline, angle, orientation and alignment. The README notes that typesetting defaults such as size and color are decided by the program, and can be overridden globally in the settings panel under the typesetting menu.

Text detection is the narrowest part. The README states it currently supports Japanese and English detection only, with training code at dmMaze/comic-text-detector. Everything downstream inherits that limit.

The LLM layer is more developed than the rest. `LLMTranslator` can be given translation history as context, so earlier completed pages inform names, terminology and tone. A token budget controls how much history is included, defaulting to `4096`, and the README suggests roughly 70 percent of a model's context limit as a reasonable ceiling, about `90000` for a 128K model. Glossaries can be supplied as UTF-8 `.json`, `.txt` or `.tsv` files, with a match-only mode that sends entries whose source term appears on the relevant page, or a full-table mode that sends everything. The README warns that full-table mode can noticeably increase token usage.

Two details are worth flagging. First, history and glossary injection only apply to `LLMTranslator`; other translators ignore those settings. Second, the README states that conflicting entries, malformed files, unsupported extensions or a missing file stop translation before any LLM request is sent. That is a fail-fast design, and it is the right one for a paid API.

## Installing BallonsTranslator on Windows, macOS and Linux

On Windows the README offers a one-line PowerShell installer that sets up a local environment in the current directory. It requires PowerShell and does not support Windows 7.

```powershell
irm https://raw.githubusercontent.com/dmMaze/BallonsTranslator/dev/scripts/install.ps1 | iex
```

From `cmd.exe` the README gives an equivalent invocation.

```cmd
powershell -NoProfile -ExecutionPolicy Bypass -Command "irm https://raw.githubusercontent.com/dmMaze/BallonsTranslator/dev/scripts/install.ps1 | iex"
```

The alternative is a preconfigured archive: download `Ballonstranslator_win_minium.zip` from GitHub Releases, extract it, and run `launch_win.bat`. If you hit `msvcp140.dll`, `c10.dll` or `[WinError 1114]` errors, the README points to the Microsoft Visual C++ Redistributable x64.

On macOS and Linux the install script is downloaded and executed, then the program starts automatically.

```bash
curl -fLO https://raw.githubusercontent.com/dmMaze/BallonsTranslator/dev/scripts/install.sh && chmod +x install.sh && ./install.sh
```

Afterwards, `cd BallonsTranslator && ./launch.sh` starts it again. The README notes that `wget -O ...` works if `curl` is unavailable.

For a first run, the README recommends launching from a terminal, setting source and target language, opening a folder containing images, and pressing Run. When you select a module that needs extra libraries, the program prompts you to install the missing optional dependencies, or you can enable automatic installation in settings. The packaging file `pyproject.toml` pins `requires-python = ">=3.8,<3.13"`, so a Python 3.13 environment will not install the package.

## Headless mode and what config/config.json controls

There is a command-line mode without the GUI. The README gives this form, where directories are passed as a comma-separated list.

```python
python -m ballontranslator --headless --exec_dirs "[DIR_1],[DIR_2]..."
```

All settings, including the detection model and source and target languages, are read from `config/config.json`. If the rendered font size comes out wrong, the README says to set the logical DPI with `--ldpi`, usually 96 or 72. That flag exists because the renderer needs to know the pixel density the layout estimate assumes; getting it wrong scales every text block.

This is the mode to look at if you want to script translation across many folders. It is also where the tool is least forgiving: there is no canvas to correct a bad detection, so a page where the detector misses a balloon stays wrong.

## The inpainting brush is the weak link, and the README says so

The canvas includes a repair brush with a rectangle tool. You drag a left-button rectangle to erase text inside it, and a right-button rectangle to clear the repair result inside it. Checking the auto option repairs immediately on release; otherwise you press the repair button or the space bar, and `Ctrl+D` deletes the rectangle.

The README is unusually direct about reliability here. It states that the result depends on how accurately the algorithm estimates the text region, that the box should generally be slightly larger than the text block you want removed, and that both methods are somewhat unpredictable. It claims they handle most simple text on simple backgrounds and some combinations of complex backgrounds with simple text or simple backgrounds with complex text, and only a minority of complex text on complex backgrounds. The suggested workaround is to drag the box several times.

That is a real constraint, not a footnote. Screentone, gradients and busy line art are exactly where inpainting fails, and those are common in printed manga. Budget manual cleanup time for any page that is not flat white behind the lettering.

## Editing, shortcuts and the undo boundary

The text editor is a rich-text canvas with font style presets, effects and deformation, plus find and replace across the whole document, the source text or the translation, and Word document import and export. Batch format adjustment and automatic typesetting are supported, as is running OCR and translation on selected text boxes only.

The shortcut list is dense: `A`/`D` or PageUp/PageDown page through the document and auto-save an unsaved page, `T` switches to text editing, `P` switches to the canvas with an opacity slider for the original image, `Ctrl+F` searches the current page and `Ctrl+G` searches globally, and `0` through `9` adjust lettering or original-image opacity. While editing a text block, `Alt+Arrow` or `Alt+WASD` moves between blocks.

One limitation is documented and easy to forget: `Ctrl+Z` and `Ctrl+Y` undo and redo most operations, but the undo and redo stack is cleared after you turn the page. If you make a mistake, notice it two pages later, and page back, the history is gone. Work on one page at a time and save deliberately.

## Alternatives and the licence question

The closest point of comparison is manga-image-translator, which BallonsTranslator depends on heavily. The README states this dependency openly and asks users to consider supporting its author through Ko-fi, Patreon or Afdian. The difference in approach is the interface: manga-image-translator is a pipeline you invoke, while BallonsTranslator wraps a comparable set of modules in a PyQt6 desktop application with a canvas, manual inpainting, and per-block text editing. If you want to script translation and never touch a mouse, the upstream project is the more direct route. If you want to fix what the models got wrong, the canvas is the reason to use this one.

The licence is GPL-3.0-or-later, declared in both `pyproject.toml` and the repository. For anyone embedding this in another product, that is a copyleft licence and the practical obligations are worth reading in full rather than summarizing here. The README also carries a notice about machine translation output: if you plan to share machine-translated results publicly without a complete pass by an experienced translator, mark them as machine translated in a visible place. That is a request from the project, not a legal term, but it reflects how the maintainers expect the output to be used.

## Conclusion

Adopt BallonsTranslator if you translate comics yourself and want OCR, inpainting and an LLM translator in one canvas you can correct by hand, and if you are comfortable installing PyQt and per-module backends. Do not adopt it if you need unattended batch translation of licensed material, or if Windows 7 is your platform. Before committing, verify three things: that your target language pair is covered by the text detector, that your GPU or CPU can run the detection and inpainting models at acceptable speed, and that your LLM provider's context window fits the token budget you set in the run dialog.

## FAQ

### What are translation softwares?

BallonsTranslator is one example: a PyQt desktop application for computer-aided comic and manga translation, written in Python and licensed GPL-3.0-or-later. It chains text detection, OCR, inpainting and machine translation, then lets you edit the result on a canvas.

### What are translation softwares?

The README frames BallonsTranslator as a computer-aided translation tool rather than a fully automatic one: it runs text detection, recognition, erasing and machine translation, and the final result depends on how well all four modules perform. The canvas exists so a human can correct what they get wrong.

### What are translation softwares?

BallonsTranslator belongs to the category of tools that wrap an automated pipeline in an editor. Its README lists one-click machine translation, mask editing and a repair brush, and rich-text editing with font style presets, effects and Word document import and export.

## Sources

- [dmMaze/BallonsTranslator on GitHub](https://github.com/dmMaze/BallonsTranslator)
- [Issues](https://github.com/dmMaze/BallonsTranslator/issues)
- [License: GPL-3.0](https://github.com/dmMaze/BallonsTranslator/blob/dev/LICENSE)
- [README](https://github.com/dmMaze/BallonsTranslator/blob/dev/README.md)
- [Releases](https://github.com/dmMaze/BallonsTranslator/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/dmmaze-ballonstranslator
