# Manga OCR recognises a whole text bubble at once, and invents text when there is none

> A Japanese OCR model built as an end-to-end Vision Encoder Decoder, aimed at manga pages with furigana, vertical text and overlaid lettering. It handles multi-line bubbles in a single pass, and it will return a plausible sentence for an image with no writing on it.

**kha-white/manga-ocr** — Optical character recognition for Japanese text, with the main focus being Japanese manga

- Repository: https://github.com/kha-white/manga-ocr
- Stars: 2,798 · Forks: 142
- Language: Python
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/kha-white-manga-ocr

## It returns a sentence for an image that carries no text

The most important thing to know about this model is not in the feature list. The README states that the model always attempts to recognize some text on the image, even when there is none, and explains the mechanism: it uses a transformer decoder, and therefore has some understanding of the Japanese language, so it may dream up sentences that look entirely realistic. That is a hallucinating OCR rather than a detector, and it has no confidence threshold exposed for filtering. The documentation calls it acceptable for most use cases and says it might be improved in the next version. For a lookup workflow driven by a screenshot you took yourself, a stray reading is noise you notice. For anything scanning pages in bulk, it is a false positive rate you have to design around.

## One pass over a whole bubble, with a length ceiling attached

The headline capability is genuine and specific: unlike many OCR models, Manga OCR recognizes multi-line text in a single forward pass, so text bubbles in a page can be processed at once instead of being split into individual lines. That is what makes it work on manga, where a bubble is the natural unit. The ceiling comes in the usage tips, which say plainly that the longer the text, the more likely some errors are to occur, and that if recognition failed for part of a longer text you should run it on a smaller portion of the image. So the single-pass design and the advice to split are in tension, and the resolution is per call site rather than per page.

```python
from manga_ocr import MangaOcr

mocr = MangaOcr()
text = mocr('/path/to/img')
```

## Clipboard mode overwrites every image you copy

The recommended reading pipeline runs through the system clipboard, and that has a consequence worth stating before you set it up. In clipboard scanning mode, any image you copy is processed by the OCR and replaced by the recognised text. Copy and paste of images stops working while the mode is on. The documented fix is to use folder scanning instead and define a separate task in a screenshot tool such as ShareX or Flameshot that saves to a folder without touching the clipboard, then point the OCR at that folder.

```commandline
manga_ocr "/path/to/sharex/screenshot/folder"
```

The full chain the README describes runs capture region, write image to clipboard, Manga OCR, write text to clipboard, then a dictionary such as Yomitan reads it. By default the recognised text goes to the clipboard.

## Linux clipboard output depends on which session type you are in

Clipboard mode is not portable across Linux desktops, and the dependency differs by display server. Wayland sessions need wl-copy, X11 sessions need xclip. The README does not ask you to guess: it tells you to run `echo $XDG_SESSION_TYPE` in the terminal and work out which one your system needs from the answer. Output then goes through pyperclip, which is a declared dependency, so the failure mode on a bare window manager is a missing external binary rather than a Python error. First use also costs a download: the model is roughly 400 MB and takes a few minutes to fetch, after which the OCR is ready once the message OCR ready appears in the logs. The tool exposes its own reference through `manga_ocr --help`, and the README notes that `python -m manga_ocr` is the fallback if the console script does not work.

## The ARM workaround names a package the project does not depend on

The troubleshooting section has two entries and neither is obvious from the dependency list. One reports problems installing mecab-python3 on ARM architecture and links a workaround. That package is not in pyproject.toml: the declared dependencies are fire, fugashi, jaconv, loguru, numpy, Pillow, pyperclip, torch, transformers, and unidic_lite. The other entry covers Windows, where Python installed from the Microsoft Store can fail with `ImportError: DLL load failed while importing fugashi: The specified module could not be found.` The fix given is to install from the official python.org download instead. Note which library is named in that error: fugashi, a Japanese tokenizer, not the OCR model, so a failure that looks like a model problem is really a broken Python installation.

## The metadata sets no Python ceiling while the README warns about new ones

pyproject.toml declares requires-python as >=3.9 with no upper bound, and the README opens by telling you the newest Python release might not be supported because of the PyTorch dependency, which often breaks with new Python releases and needs time to catch up. The two statements do not contradict each other but they point in opposite directions, and a resolver will happily pick the newest interpreter your machine has. The dependency pins are lopsided in the same spirit: transformers is held at 4.45.0 or newer, which is the floor for the Vision Encoder Decoder model class this project builds on, while torch is left at 1.0 or newer. A torch requirement that old predates most of the Python versions the metadata permits.

## The wheel ships one package, and the model download has no named source

The build configuration is worth reading before you plan around the repository layout. The wheel packages a single directory, manga_ocr, and takes its version dynamically from the attribute manga_ocr._version.__version__, with setuptools_scm in the build requirements. The repository itself also holds manga_ocr_dev, which the README points to for the training code and the synthetic data generation pipeline, plus a tests directory and an assets directory, and none of those are packaged. So the model weights are not in the repository and the README names no host for the roughly 400 MB download: it says only that the first run downloads it. Training data is credited to the Manga109-s and CC-100 datasets, which is consistent with a synthetic data pipeline feeding an end-to-end model.

## The examples table lists ten readings and no images

The examples section is introduced as cherry-picked examples showing the capability of the model, and it is a two column table pairing an image with the Manga OCR result. Ten rows carry Japanese output, including a long shouted line with punctuation and a line mixing full width Latin letters. The image column of every row is empty, so the page presents the claims without the evidence a reader would need to judge them, and no accuracy figure appears anywhere to fill the gap. Two limits from the usage tips are the honest counterweight. The model was trained for manga and should handle other printed text such as novels or video games decently, but it will not handle handwriting.

## Conclusion

Manga OCR suits a reading workflow where a screenshot becomes a lookup key, and its handling of furigana, vertical text and lettering over artwork is the reason to reach for it rather than a general OCR engine. Two limits decide whether it fits your use. It always attempts to read something, so an image with no text on it comes back with a sentence the model composed, which rules it out anywhere a false positive costs you. And accuracy falls as the recognised block grows, so long passages need chunking rather than one call. Before installing, confirm your Python version is one PyTorch already supports, since the project metadata sets no ceiling and the README warns that new releases routinely break the dependency. On Windows, install from python.org rather than the Microsoft Store, or the fugashi import fails outright.

## FAQ

### what is manga ocr

Optical character recognition for Japanese text, with the main focus on Japanese manga, built as a custom end-to-end model on the Transformers Vision Encoder Decoder framework. It is aimed at vertical and horizontal text, furigana, text overlaid on images, varied fonts, and low quality scans, and it recognises multi-line text in a single forward pass.

### how to install manga ocr

Python 3.9 or newer is required, with the caveat that the newest release may not yet work because of the PyTorch dependency. Three routes are given: pip install manga-ocr, uv add manga-ocr, or running it without installing into the project using uvx manga_ocr. If you want GPU support, install PyTorch first as its own step.

### How do I run manga-ocr over a folder of images instead of the clipboard?

Pass the directory as the single argument, for example manga_ocr "/path/to/sharex/screenshot/folder". Folder scanning is also the mode the README recommends if you want to keep copying and pasting images normally, since clipboard scanning mode replaces every copied image with its recognised text.

### Does manga-ocr return text for an image with no writing on it?

Yes, and this is a known behaviour rather than a bug report. The model always attempts to recognize some text, and because it uses a transformer decoder with some understanding of Japanese, it may produce sentences that look realistic. The README says this should not matter for most use cases and might be improved in the next version.

### manga ocr vs paddleocr

The repository does not compare itself to PaddleOCR anywhere. What it names instead are projects built on it: Poricom, a GUI reader, mokuro, which generates an HTML overlay, and Xelieu's guide to setting up a reading and mining workflow. For manga specifically, the comparison it offers is against general purpose printed Japanese OCR, where the differences claimed are vertical text, furigana, overlaid text and low quality images.

## Sources

- [Issues](https://github.com/kha-white/manga-ocr/issues)
- [kha-white/manga-ocr on GitHub](https://github.com/kha-white/manga-ocr)
- [License: Apache-2.0](https://github.com/kha-white/manga-ocr/blob/master/LICENSE)
- [README](https://github.com/kha-white/manga-ocr/blob/master/README.md)
- [Releases](https://github.com/kha-white/manga-ocr/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/kha-white-manga-ocr
