Model or dataset
xunbu/docutranslate avatar
xunbu/docutranslate

DocuTranslate: LLM-based file translation for PDF, Word, Excel, EPUB and SRT

文档(小说、论文、字幕)翻译工具(支持 pdf/word/excel/json/epub/srt...)Document (Novel, Thesis, Subtitle) Translation Tool (Supports pdf/word/excel/json/epub/srt...)

1,315 stars185 forksPythonMPL-2.0

At a glance

What is it?
DocuTranslate turns documents into markdown or structured text, sends chunks to an LLM, and writes the translation back into the original file format. It is a local, Python-first tool with a web UI and an MCP server, and its PDF path drops the original layout.
Who is it for?
Adopt DocuTranslate if you translate many documents in formats such as docx, xlsx, json or srt and you already have an API key for an OpenAI-compatible endpoint, because the format-preserving paths and the jsonpath-ng value selection are the parts that save real work.
Can I use it commercially?
Yes, with conditions. MPL-2.0 is a weak copyleft licence: you can use it inside commercial and closed-source software, but if you distribute changes to its own files, you must publish those changes under the same licence.
Is it still maintained?
Yes. The repository last received commits 26 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What DocuTranslate solves, and for whom

Most translation tools handle one shape of input. A subtitle tool reads srt and ass. A PDF tool reads PDF. DocuTranslate targets the person who has a folder of mixed files: a thesis in PDF, a manuscript in docx, a glossary in xlsx, a localisation file in json, and a season of subtitles in srt. The project describes itself as a lightweight local file translation tool based on Large Language Models, and the format list in the README covers pdf, docx, xlsx, md, txt, json, epub, srt and ass.

The intended user is technical enough to hold an API key and run a Python package or a container, but does not want to write a translation pipeline. The README also notes LAN and multi-user support, so a small team can share one instance rather than each person configuring their own. If you only ever translate plain text, this is more machinery than you need.

The pipeline: format readers, chunking, and an LLM client

The dependency list in pyproject.toml is the clearest description of the architecture. Each format has a dedicated reader or writer: python-docx for docx, openpyxl and xlsx2html for xlsx, python-pptx for pptx, pypdf for PDF, mammoth for Word conversion, srt and pysubs2 for subtitles, beautifulsoup4 and html2text for HTML, plus markdown and pymdown-extensions. Translation itself goes through httpx against an OpenAI-compatible endpoint, and json-repair handles malformed model output, which tells you the project expects the model to occasionally return broken JSON.

Text is split before it reaches the model. .env.example lists DOCUTRANSLATE_CHUNK_SIZE with a default of 4000, DOCUTRANSLATE_CONCURRENT with a default of 30, and separate timeout settings for streaming and non-streaming calls. DOCUTRANSLATE_LLM_STREAMING defaults to true, and DOCUTRANSLATE_STREAM_IDLE_TIMEOUT (300 seconds) restarts on every received chunk, so a long stream has no total wall-clock limit while a silent connection is cut. Non-streaming calls use DOCUTRANSLATE_NON_STREAM_TIMEOUT, which falls back to the older DOCUTRANSLATE_TIMEOUT value.

For PDFs the flow is different. The README states that a PDF is first converted to markdown, using mineru either online or locally deployed, and that this loses the original layout. Tables, formulas and code in academic papers are the reason mineru is in the path at all.

Installing DocuTranslate with pip and running the web UI

The README gives four install routes: pip, uv, git clone, and Docker. The pip path is the shortest. Python 3.11 or newer is required, as declared in pyproject.toml.

bash
pip install docutranslate
pip install docutranslate[mcp]
docutranslate -i

The first command installs the core package. The second adds the optional mcp extra, which pulls in the mcp package and python-dotenv. The third starts the interactive service. By default it binds locally, and the README says to open http://127.0.0.1:8010 in a browser, with Swagger UI for the REST API at http://127.0.0.1:8010/docs.

To let other machines on your network reach it, the README documents a host flag and a port flag:

bash
docutranslate -i --host 0.0.0.0
docutranslate -i -p 8081
docutranslate -i --cors

Before the first translation you need a model endpoint. Copy .env.example to .env and set the three required keys. The example file lists these names:

bash
DOCUTRANSLATE_API_KEY=sk-xxxxxxxx
DOCUTRANSLATE_BASE_URL=https://api.openai.com/v1
DOCUTRANSLATE_MODEL_ID=gpt-4o
DOCUTRANSLATE_TO_LANG=中文
DOCUTRANSLATE_CONCURRENT=30

The provider list in .env.example includes default, google, minimax, ollama, aliyuncs, volces, siliconflow, bigmodel, deepseek, openrouter, mimo and litellm. Ollama in that list is the route to a fully local model. Once the service is up, upload a file in the web UI and the translation is queued and executed concurrently. The README also mentions integration packages under 40MB on GitHub Releases for Windows and Mac, which skip the Python setup entirely.

Where DocuTranslate is the wrong tool

The PDF limitation is stated plainly in the README: translating a PDF first converts it to markdown, which loses the original layout, and users with strict layout requirements are told to take note. That rules it out for scanned contracts, designed brochures, or anything where a figure and its caption must stay in place. A tool that edits the PDF in place is the right shape for that job.

The format list has two visible gaps. The README says docx and xlsx are supported but doc and xls are not, so legacy binary Office files need converting first. There is also no documented rollback: if a translation is wrong, the README does not describe an undo or a diff step, so keeping the source file is the only safety net.

Cost and privacy are the other boundaries. Every chunk goes to the configured endpoint unless you point DOCUTRANSLATE_BASE_URL at a local server such as Ollama. The README does not document a cost estimator, and the concurrency default of 30 in .env.example means a large document can fan out into many simultaneous requests.

DocuTranslate compared with PDFMathTranslate and BabelDOC

The related searches for this project include PDFMathTranslate and BabelDOC, and the difference in approach is worth stating. PDFMathTranslate and BabelDOC are PDF translation tools. Their working assumption is a PDF in and a PDF out, with the layout preserved, which is exactly the constraint DocuTranslate gives up when it routes a PDF through markdown.

The trade is deliberate. By converting to markdown first, DocuTranslate gets uniform text for any format and can use one chunking and prompting path for a thesis, a subtitle file and a spreadsheet. Tools built around PDF layout cannot reuse that path for docx or srt. If your work is overwhelmingly PDF and layout matters, the PDF-first tools are closer to the problem. If your work is mixed formats and you want one queue, one glossary and one API key, DocuTranslate's approach is the one that scales across them.

MCP server mode and environment variables

DocuTranslate can run as a Model Context Protocol server, which lets an MCP client call it as a tool. The README documents three transports. The stdio mode is the default for docutranslate --mcp, SSE mode adds --transport sse with --mcp-host and --mcp-port, and streamable HTTP uses --transport streamable-http. When you start the web UI with --with-mcp, the MCP SSE endpoint shares the same port at http://127.0.0.1:8010/mcp/sse; a standalone MCP server defaults to http://127.0.0.1:8000/mcp/sse.

The uvx configuration in the README runs the server without a prior install:

json
{
  "mcpServers": {
    "docutranslate": {
      "command": "uvx",
      "args": ["--from", "docutranslate[mcp]", "docutranslate", "--mcp"],
      "env": {
        "DOCUTRANSLATE_API_KEY": "sk-xxxxxx",
        "DOCUTRANSLATE_BASE_URL": "https://api.openai.com/v1",
        "DOCUTRANSLATE_MODEL_ID": "gpt-4o"
      }
    }
  }
}

The environment variables in the README table are DOCUTRANSLATE_API_KEY, DOCUTRANSLATE_BASE_URL and DOCUTRANSLATE_MODEL_ID as required, with DOCUTRANSLATE_TO_LANG, DOCUTRANSLATE_CONCURRENT, DOCUTRANSLATE_CONVERT_ENGINE and DOCUTRANSLATE_MINERU_TOKEN optional. .env.example adds two behaviour switches worth knowing: DOCUTRANSLATE_WEB_SKIP_VALIDATION, which makes the web front end skip its empty-value check and fall back to the .env defaults, and DOCUTRANSLATE_ENV_FORCE_OVERRIDE, which makes .env values win over values sent by the front end. The default of the latter is false, meaning request parameters take precedence and .env only fills gaps.

Docker, licence and the cost of keeping up

The published image is run with a single command, mapping port 8010:

bash
docker run -d -p 8010:8010 xunbu/docutranslate:latest

The Dockerfile builds in two stages on python:3.11-slim, installs uv in the builder, runs uv sync --frozen --extra mcp, and copies the virtual environment into the runtime stage. The runtime stage installs pandoc, ca-certificates and curl, sets DOCUTRANSLATE_PORT to 8010, exposes that port, and defines a health check that curls /service/meta every 30 seconds. The entrypoint is docutranslate -i --with-mcp, so a container gives you the web UI and the MCP endpoint together. Note that the image is built with the mcp extra but not the docling extra, which pyproject.toml keeps separate along with opencv-python and hf-xet; a container that needs the docling conversion path would need a different build.

The licence is MPL-2.0, a file-level copyleft licence. Modifying DocuTranslate's own files and distributing them carries obligations; calling it as a service or a dependency does not change your application's licence. That is a summary of the licence family, not legal advice, and the LICENSE file in the repository is the authority.

Upgrade cost looks low. The last push was on 2026-09-04, the same day as the v1.7.9 release, and the two releases before it landed on 2026-07-10 and 2026-06-25, so the cadence is roughly monthly. The pinned dependency set in pyproject.toml is broad and moves with the Python ecosystem, and uv.lock plus the frozen sync in the Dockerfile mean the container will not drift underneath you. There is no documented migration guide between releases, so read the changelog file in the repository before upgrading a shared instance.

Editorial conclusion

Adopt DocuTranslate if you translate many documents in formats such as docx, xlsx, json or srt and you already have an API key for an OpenAI-compatible endpoint, because the format-preserving paths and the jsonpath-ng value selection are the parts that save real work. Do not adopt it if your PDFs must keep their original layout, since the README states that PDF translation first converts to markdown and loses the layout, or if you cannot send document text to a third-party model. Before committing, verify two things: that your provider is listed under DOCUTRANSLATE_PROVIDER in .env.example, and that your PDFs survive the mineru conversion stage acceptably.

Frequently asked questions

What is the best free document translator?

DocuTranslate itself is free software under MPL-2.0, but it is not free to run: it calls a Large Language Model, and the README requires an API key, a base URL and a model ID. You can avoid per-token cost by pointing DOCUTRANSLATE_BASE_URL at a local server such as Ollama, which the provider list in .env.example includes.

How much do document translation services cost?

The README and .env.example do not give any pricing figures or a cost estimator. DocuTranslate bills through whichever model endpoint you configure, so the cost depends on your provider and on how many chunks the document produces at the configured DOCUTRANSLATE_CHUNK_SIZE.

Can Chatgpt translate a PDF document?

DocuTranslate does not use ChatGPT directly, but it speaks to any OpenAI-compatible endpoint, and the example configuration in the README uses DOCUTRANSLATE_BASE_URL=https://api.openai.com/v1 with DOCUTRANSLATE_MODEL_ID=gpt-4o. For PDFs the README states the file is first converted to markdown, using mineru online or locally deployed, and that the original layout is lost.

Is Google Translate document free?

That question is about Google Translate, not about DocuTranslate, and the repository material says nothing about it. DocuTranslate lists google among the provider identifiers in .env.example, which means you can route through a Google endpoint, but the project does not document that provider's pricing.

Official sources

  1. Issues
  2. License: MPL-2.0
  3. README
  4. Releases
  5. xunbu/docutranslate on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/xunbu-docutranslate.svg)](https://hysenlabs.com/projects/xunbu-docutranslate)