Open-source project
koharu-rs/koharu avatar
koharu-rs/koharu

Koharu: a local-first manga translator built in Rust

AI-powered manga translator, written in Rust.

5,693 stars398 forksRustApache-2.0

At a glance

What is it?
Koharu chains detection, OCR, inpainting and LLM translation into one desktop pipeline, running the vision models on your own GPU. It is a serious tool for people who translate manga pages, and a poor fit for anyone without a supported graphics driver.
Who is it for?
Adopt Koharu if you already translate manga pages, have a Turing-class NVIDIA card, an Apple silicon Mac, or a Vulkan-capable Windows or Linux machine, and want the source images to stay on your own disk. Do not adopt it if you need a headless batch service, if your only hardware is a CPU-only laptop, or if your target language is not covered by a model you are willing to run.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly Rust, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What Koharu actually automates, and for whom

Translating a manga page by hand is four jobs stacked on top of each other. You find the text regions, you read the Japanese, you translate it, and then you erase the original lettering and set new type in the same bubble. Koharu's README describes a local-first workflow that automates that chain with ML, combining object detection, OCR, inpainting and LLMs. The audience is not a general reader who wants a page translated once. It is someone who produces translated pages and cares about the intermediate artifacts: region masks, extracted source text, corrected translations, and a layered export.

The privacy claim is the differentiator. The README states that Koharu runs its vision models and LLMs locally on your machine. That matters if you are working on licensed material, unpublished doujin, or anything you would rather not upload to a hosted OCR or translation API. The trade-off is hardware: local inference means your GPU does the work, and the README is explicit that CPU inference is available but substantially slower.

The pipeline: detection, OCR, inpainting, translation, typesetting

The repository is a Cargo workspace whose members map closely onto the pipeline stages. There are separate crates for koharu-ml, koharu-pipeline, koharu-translator, koharu-renderer, koharu-rasterizer, koharu-psd, koharu-diffusion, koharu-llama and koharu-torch, with a koharu-app crate as the default workspace member. That layout tells you the stages are not one monolith: the diffusion stack and the llama stack are separate bindings, so inpainting and translation can be swapped independently.

The README describes the stages as selective. Detection and segmentation find text regions, speech bubbles and cleanup regions. A multimodal OCR model reads dialogue, captions and general page text from those regions. Inpainting reconstructs the artwork behind the source text before anything new is drawn. Translation runs either through a local GGUF model or a hosted provider. Then the renderer handles multilingual text shaping and layout with automatic fitting, font fallback, vertical CJK and right-to-left text. The output can be a flattened image or a layered PSD.

Models are chosen per stage, not as a bundle. Detection defaults to Koharu Layout RF-DETR Seg 2XL. OCR lists PaddleOCR VL 1.6, Manga OCR, Baberu OCR and Hayai OCR. Inpainting lists FLUX.2 Klein, RORem mixed, LaMa and AOT GAN. That is a real design decision with a real cost: a fast small model and a heavy generative model produce visibly different pages, and the README does not rank them.

Installing Koharu and running one page

The README points to https://koharu.rs/en/installation for getting started, and the releases page carries prebuilt assets. The repository itself is a Tauri desktop application, so the build path goes through the Rust toolchain and the Tauri CLI rather than a plain cargo install.

The package.json defines the developer entry points. The postinstall step installs the Tauri CLI from a pinned git revision, and dev launches the desktop shell.

bash
bun install
bun run dev

The dev script runs cargo tauri dev, so the first run compiles the Rust workspace before the window appears. For a distributable build, the same file defines a no-bundle build:

bash
bun run build

Before any of that, check the driver situation, because the README ties acceleration to specific generations. CUDA 13.0 needs an NVIDIA Turing-class or newer GPU and an R580 or newer driver. ROCm 10.0 depends on the exact AMD GPU, operating system and driver combination, and AMD publishes a compatibility matrix. Metal is Apple silicon only. Vulkan covers Windows and Linux as an alternative. The editor canvas itself uses WebGPU and requires a current graphics driver even when inference falls back to the CPU.

Once the window opens, the first real use is a single page rather than a chapter. Create a project, add the image, run detection, confirm the regions look right, run OCR and read the extracted text, then run translation and inpainting. The README documents a proofreading step for correcting OCR and translation output, and a WebGPU canvas for manual cleanup, text placement and page composition. Do the first page manually end to end before touching project-scope processing, because a bad detection model will silently ruin every page after it.

Where Koharu breaks down

The hardware requirements are the first wall. A machine without a supported GPU can run CPU inference, but the README calls it substantially slower, and the editor canvas still wants a current graphics driver for WebGPU. If your only machine is an older integrated-graphics laptop, this is the wrong tool regardless of how good the models are.

The second wall is model selection. Koharu does not ship one blessed configuration. Detection, OCR, inpainting and translation are all chosen separately, and the README lists the candidates without saying which combination works for which kind of page. Dense four-koma with handwritten sound effects and a full-color splash page are different problems, and nothing in the documentation tells you which inpainting model preserves screentones. You will find out by running them.

The third is the LLM side. The README distinguishes general-purpose local models from uncensored local models, and lists Qwen, Gemma, Ministral and LFM families at sizes from 0.8B up to 35B. A 35B model is not going to run comfortably alongside a diffusion inpainting model on the same consumer card. Local GGUF inference and hosted providers are both supported, but choosing a hosted provider gives up the privacy property that the project leads with.

Finally, the README does not document rollback. If inpainting overwrites a region badly, the documented recovery path is the layered PSD export and manual canvas work, not an undo of the model output.

Koharu compared with scripted OCR pipelines

The obvious alternative is assembling the same stages yourself from separate tools: an OCR engine, a segmentation model, a translation API and an image editor, glued together with a script. That approach wins on flexibility and loses on integration. You control every model choice and every intermediate file, and you can run it headless on a server. You also own the text shaping, the font fallback, the vertical CJK layout and the PSD layering, which is exactly the work Koharu's renderer and rasterizer crates exist to absorb.

The difference in approach is architectural. A scripted pipeline is a batch process: images in, images out, no interface. Koharu is a desktop editor with a project model, a WebGPU canvas and a proofreading step, which means a human stays in the loop at each stage. That is better for a chapter you care about and worse for a thousand pages you do not want to look at. The README also documents an agent-based workflow for project inspection, editing and pipeline control, which narrows the gap, but the primary interface is still the application window.

The other real difference is that the scripted route lets you send text to a hosted translation service that may be far better than any model you can fit locally. Koharu supports hosted providers too, so this is a configuration choice rather than a hard limit, but the project's stated reason for existing is the local path.

Licence, maintenance and the cost of keeping up

The workspace Cargo.toml declares license = "MIT OR Apache-2.0" and the repository root carries both LICENSE-MIT and LICENSE-APACHE. The GitHub metadata lists Apache-2.0. The dual grant is the permissive Rust convention: you pick either licence and comply with that one. Note that publish = false is set in the workspace package block, so the crates are not intended for crates.io distribution, and the package.json license field repeats the same MIT OR Apache-2.0 choice for the JavaScript side. None of this is legal advice; if you are redistributing bundled model weights, check the licence of each model separately, because the README links to Hugging Face repositories rather than vendoring them.

Maintenance is active by any reasonable reading. The last push to main was on 2026-09-10, and releases 0.81.8, 0.81.9 and 0.81.10 all landed on 2026-09-08. The workspace version in Cargo.toml reads 0.83.0, ahead of the published release tags, which is normal for a repository that tags after merging.

The upgrade cost is not the application binary, it is the model set. The README lists FLUX.2 Klein, PaddleOCR VL 1.6, Qwen 3.5, 3.6 and 3.8, and Gemma 4 variants. Model families turn over faster than the application does, and each new inpainting or OCR model changes what your pages look like. Budget for re-running a sample chapter after any model swap, and keep the layered PSD exports so you can fall back to manual typesetting when a new model does worse than the old one.

Editorial conclusion

Adopt Koharu if you already translate manga pages, have a Turing-class NVIDIA card, an Apple silicon Mac, or a Vulkan-capable Windows or Linux machine, and want the source images to stay on your own disk. Do not adopt it if you need a headless batch service, if your only hardware is a CPU-only laptop, or if your target language is not covered by a model you are willing to run. Before committing a chapter, verify three things on one page: that the OCR model you pick reads your source text correctly, that the inpainting model reconstructs screentones without smearing them, and that the LLM you select actually emits the target language rather than repeating the Japanese.

Frequently asked questions

What is Koharu?

Koharu is an ML-powered manga translator written in Rust, distributed as a desktop application. It combines object detection, OCR, inpainting and LLMs so that detection, translation and cleanup happen in one project-based workflow.

How do I use the Koharu manga translator?

Install it, open a project with your page images, then run the pipeline stages in order: detection and segmentation, OCR, translation, and inpainting. The README documents a proofreading step for correcting OCR and translation output, and a WebGPU canvas for manual cleanup and text placement.

How do I use Koharu to translate a page?

The README describes the pipeline as selective, so you can run detection, OCR, translation and inpainting at page scope or project scope. Start at page scope on one page, check the detected regions and the extracted text, then move to project scope once the models behave.

How do I use Koharu?

The README points to https://koharu.rs/en/installation for getting started, and the releases page carries prebuilt assets. From a source checkout, bun install runs a postinstall step that installs the Tauri CLI, and bun run dev launches the desktop shell.

Official sources

  1. koharu-rs/koharu on GitHub
  2. License: Apache-2.0
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/koharu-rs-koharu.svg)](https://hysenlabs.com/projects/koharu-rs-koharu)