# RetainPDF: PDF translation aimed at the layout problems, with a monorepo behind the UI

> A Chinese-first project for translating scanned and image-based PDFs while keeping formulas, columns and code intact, shipped as desktop installers and a Docker service.

**wxyhgk/retain-pdf** — 在保留版面、公式与结构的前提下进行 PDF 翻译，适用于科研与技术文档

- Repository: https://github.com/wxyhgk/retain-pdf
- Stars: 2,476 · Forks: 290
- Language: Python
- License: MIT
- Published: 2026-10-07 · Updated: 2026-10-07 · Language: en
- Canonical page: https://hysenlabs.com/projects/wxyhgk-retain-pdf

## The problem is the text box, not the language

The project's stated claim is that it is the only open source tool for translating image-based and scanned PDFs while preserving layout, and that its output matches or exceeds comparable commercial products. Bold marketing aside, the interesting part is the stated mechanism.

The translation system is described as working on reassembled semantic units rather than individual text boxes: first restore the complete meaning, then translate. That single design choice explains most of the claimed behavior. Translating box by box is what produces the familiar failure mode where a sentence continues onto the next column, or a footnote on the previous page, and both halves get translated as if they were separate statements.

The font layout algorithm is also described as custom, with support for complex formulas and multi-column paper layouts. And the claim the README makes most strongly is about inline formulas: that it preserves the formula body, its relation to the surrounding text, and its inline layout after translation.

Whether any of this is true for your documents is an empirical question. The README provides sample before-and-after comparisons across SCI papers, scanned formula-dense documents and textbooks, which is the right kind of evidence to look at, but it is chosen by the project.

## A comparison table that deserves a skeptical read

The README includes a feature matrix comparing RetainPDF against PDFMathTranslate, PolyglotPDF and Doc2X across nine columns: scanned PDF support, complex inline formulas, code not mistranslated, table control, custom translation strategy, layout preservation, PDF compression optimization and API automation.

In that table, RetainPDF is the only row with a checkmark in every column. The other three are marked partial or absent across most rows, and Doc2X is listed as having no open API.

Two things are worth knowing about a self-authored comparison like this. First, the criteria were chosen by the party being compared, and the categories are defined in terms of capabilities the project set out to build. Second, a feature marked absent in someone else's tool often means absent from that tool's public documentation, which is different from absent from the tool.

That said, the categories themselves are a useful checklist even if you discount the checkmarks entirely. Code not being mistranslated, table control being switchable, and custom translation strategy being configurable per rule are all real problems in this category of tool, and they are the three you should test first with any of these projects.

## Two frontends, a Rust backend and seven workspace crates

GitHub reports this repository's language as Python, which is the first thing worth untangling, because the tree and the manifests describe something more layered.

The root `Cargo.toml` declares a Rust workspace with seven members: `database/retain-db`, `backend/api`, `backend/jobs`, and four packages under `backend/packages` named `retain-core`, `retain-data`, `retain-jobs` and `retain-proc`. That split has the shape of a service: a database layer, an HTTP API, a job runner for the long translation work, and a core library separating shared logic from the pipeline steps.

Alongside that sits a Node monorepo. The root `package.json` is marked private, named `retain-pdf-monorepo`, version 4.2.5, and declares five workspaces: `frontend/web`, `frontend/desktop`, `frontend/packages/*`, `contracts` and `backend/packages/retainpdf2doc`. The naming convention there tells you the intended split, with web and desktop frontends sharing a packages directory and a contracts workspace defining the API boundary between them.

Python appears in the toolchain rather than the product. Several scripts shell out to it: `test:api` runs `python3 .github/scripts/run_with_backend_source.py` before invoking `uv run` and `cargo test`, and the translation test suites run through `python3 backend/pipeline/devtools/run_translation_tests.py`. So Python is the harness that drives Rust, not the implementation. If you are deciding whether to contribute, the language you need is Rust.

## Docker deployment and the one URL that matters

Self-hosting is the documented path, and the clone-and-compose sequence is short:

```bash
git clone https://github.com/wxyhgk/retain-pdf.git
cd retain-pdf/ops/deployment/docker/delivery
docker compose up -d
```

Once it is up, the README says to visit `http://127.0.0.1:40001`. That port is bound to loopback, which is the right default for a service that holds document content and translation state but does mean you need a reverse proxy or a tunnel for anything beyond the machine it runs on.

Updates are a two-step pull and up:

```bash
docker compose pull
docker compose up -d
```

Further configuration is documented in a README inside the same deployment directory. The directory structure is worth noting: deployment lives under `ops/deployment/docker/delivery`, which suggests the Compose file is one option among several rather than the only supported shape.

There are desktop builds too, distributed through GitHub Releases as a Windows `Setup.exe`, a macOS `.dmg` and a Linux `.deb`. The macOS build hits Gatekeeper's quarantine attribute on first launch, and the README's fix is:

```bash
sudo xattr -r -d com.apple.quarantine /Applications/RetainPDF.app
```

## Product surface: library, side-by-side reader and document Q&A

The README documents four screens, which tells you what the product is actually for rather than what the translation engine can technically do.

The library view organizes papers, books and scanned documents on a bookshelf, with translation status, collections and favorites. The side-by-side reader shows the original PDF and the translated PDF on the same page, which exists so a reader can verify formulas, figures, citations and layout position rather than take the output on trust. The Markdown view puts the PDF and a Markdown rendering together, preserving formulas and images.

The fourth is document Q&A: ask questions about the currently open document, get answers carrying page numbers and quotes from the source. When a question implies an edit to the PDF, the interface switches explicitly to a PDF agent.

The existence of the side-by-side reader is a design statement worth taking seriously. A translation tool that cannot be checked against the original is asking for too much trust, and building verification into the product rather than leaving it to the reader is the right call. It also sets a higher bar for the project, because the comparison view is where any layout regression becomes visible.

The repository description is in Chinese and scoped to research and technical documents, and the tree carries three READMEs: `README.md`, `README.en.md` and `README.vi.md`, so English and Vietnamese versions exist alongside the Chinese original.

## Signals worth checking before you rely on it

A few things about the repository state are worth knowing.

The releases carry no release notes. v4.2.5, v4.2.4 and v4.2.3 all have empty bodies, published on 2026-09-19, 2026-09-18 and 2026-09-14 respectively. Three patch releases in six days with no description means you cannot tell what changed or whether an upgrade is safe without diffing the source yourself.

The MIT license is consistent between the README's closing section and the `LICENSE` file in the tree, which is unusual for a project of this type and worth noting in your favor. The topic list includes `typst`, which is a reasonable signal about the layout approach, and `ocr` and `document-ai` alongside `layout-preserving` and `scientific-papers`.

The test layout in `package.json` is more revealing than the README. Beyond the desktop and web test targets there are translation-specific suites, including separate tooling and benchmark suites invoked through `run_translation_tests.py`, and a Windows remote script for running desktop tests on another machine. A translation quality regression is hard to catch with unit tests, so having a benchmark suite wired into CI is a better sign than it sounds.

The last push was 2026-09-21 and the repository is not archived. The community contact listed in the README is a QQ group rather than an issue tracker or a chat platform with an English-speaking audience.

## Conclusion

RetainPDF is interesting for one specific reason: it treats the text box as the wrong unit of work. Reassembling a semantic unit before translating, rather than translating box by box, is what makes cross-column and cross-page text come out coherent, and the inline formula handling follows from the same idea. Whether that holds up depends on documents you care about, which is not something the README can settle, and it ships with a self-assessment table that you should read with some skepticism given who wrote it. It is the wrong tool for plain digital PDFs where text extraction already works, and it is worth remembering that the translation backend is a moving part the repository does not pin. Start with a Docker deployment on a document you already know the contents of, so you can check the output against ground truth before trusting it with anything you cannot verify by eye.

## FAQ

### Can RetainPDF handle scanned PDFs, not just digital ones?

That is the project's main claim. It targets image-based and scanned PDFs specifically, and the README's comparison table marks scanned PDF support as present for RetainPDF and absent for PDFMathTranslate, PolyglotPDF and Doc2X. The repository topics include `ocr` and `document-ai`, which is consistent with that scope.

### How do I self-host it?

Clone the repository, change into `ops/deployment/docker/delivery` and run `docker compose up -d`, then open `http://127.0.0.1:40001`. The port binds to loopback, so remote access needs a proxy or tunnel. There are also desktop builds distributed as a Windows installer, a macOS dmg and a Linux deb package.

### What language is RetainPDF written in?

GitHub reports Python, but the manifests show a Rust workspace holding the database layer, API, job runner and core packages, with Python used to drive tests and tooling. The user interface is a Node monorepo covering web and desktop frontends plus a contracts workspace. Read the manifests rather than the language badge.

## Sources

- [Issues](https://github.com/wxyhgk/retain-pdf/issues)
- [License: MIT](https://github.com/wxyhgk/retain-pdf/blob/main/LICENSE)
- [README](https://github.com/wxyhgk/retain-pdf/blob/main/README.md)
- [Releases](https://github.com/wxyhgk/retain-pdf/releases)
- [wxyhgk/retain-pdf on GitHub](https://github.com/wxyhgk/retain-pdf)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/wxyhgk-retain-pdf
