MTranServer: an offline translation server that runs without a GPU
Offline translation model server with low resource consumption, fast speed, and private deployment capability. 低资源占用速度快可私有部署的离线翻译模型服务器
At a glance
- What is it?
- MTranServer is an Apache-2.0 translation server built around Bergamot-family models, packaged as an npm CLI, a Docker image and a desktop app. It trades translation quality for latency, memory and full local control, and that trade is the whole point.
- Who is it for?
- Adopt MTranServer if you want translation that never leaves your machine, runs on CPU, and answers in tens of milliseconds, and if you can accept quality below a large online model. Do not adopt it if translated output is customer-facing prose where fluency matters more than latency, or if you cannot host a process that downloads model files on first use.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Activity is slowing. The repository last received commits 6 months ago.
- What is it written in?
- Mainly C++, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What MTranServer solves, and for whom
The project describes itself as an offline translation model server with low resource consumption, fast speed and private deployment, requiring no graphics card. The README states an average response time of 50 milliseconds per request and says the server supports translation between the world's major languages. Those two claims define the audience: people who need a lot of short translations, quickly, on hardware they already own, and who are unwilling or unable to send that text to a third-party API.
The README is explicit about the trade-off rather than hiding it. It says the model server focuses on offline translation, response speed, cross-platform deployment and local execution to reach unlimited free translation, and that because of model size and the degree of optimisation the translation quality is not as good as large-model translation. Readers who need high quality are told to use an online large-model API instead. That is an unusually honest framing, and it should be taken literally: this is a latency and privacy tool, not a quality tool.
The people it fits are developers translating code comments, documentation, chat messages or web pages in bulk, and anyone running a personal machine or a small server where a GPU is not available. The people it does not fit are those producing published prose in a second language, where the difference between a Bergamot-class model and a large language model is visible in every paragraph.
The mechanism: local models, HTTP compatibility endpoints
MTranServer is a server process that loads translation models locally and exposes them over HTTP. The package keywords list bergamot, wasm, machine-translation and translation-server, which places the inference stack in the Bergamot family of compact neural translation models rather than in a transformer served from a GPU. The repository is written primarily in C++ with a TypeScript layer: package.json declares "type": "module", a bin entry named mtranserver pointing at ./dist/main.js, and a build script that produces a Node build (build:node) and an Electron build (electron:build). The Dockerfile reflects the same split: a builder stage on oven/bun:1 runs bun run build:node, and the runtime stage is node:22-alpine running node main.js.
Models are not bundled. The README states that the first time a language pair is translated the server downloads the corresponding model, unless offline mode is enabled, and that this can take a while depending on network speed and model size. After the download, requests are answered at millisecond latency. The README advises testing one translation before real use so the server pre-downloads and pre-loads the model. Model files live in a directory the user controls: the CLI default is ~/.config/mtran/models, and the compose file mounts ./models into /app/models.
The interesting design choice is the endpoint surface. Rather than inventing one API and asking clients to adapt, MTranServer implements the request shapes of existing translation clients: /imme for Immersive Translate, /kiss for Kiss Translator, /deepl for DeepL API v2, /deeplx for DeepLX, /hcfy for a selection-translation client, plus /google/language/translate/v2 and /google/translate_a/single. Each accepts a token parameter or Bearer authentication. This is the practical reason the project is usable at all: you point an existing browser extension at http://localhost:8989/imme and it works, with no plugin development. The cost is that the server must track several third-party contracts it does not control.
Installing MTranServer and translating once
The README gives three installation paths. For a quick look, the CLI can be started directly with npx, and the README notes that npx can be replaced with another package manager such as bunx or pnpx:
npx mtranserver@latestFor a persistent install, the package is installed globally and then started by name:
npm i -g mtranserver@latestAfter that, running mtranserver starts the server. The default listen address is 0.0.0.0 and the default port is 8989, and the Web UI is enabled by default (-ui, default true). According to the README, the first translation of a given language pair triggers a model download, so the first request for that pair will be slow and later ones fast. The README recommends making one test translation before real use for exactly this reason.
For a server deployment, the repository ships a compose.yml that pins the image, maps port 8989, sets MT_HOST, MT_PORT, MT_ENABLE_UI and MT_OFFLINE, mounts ./models into /app/models, and defines a healthcheck against http://localhost:8989/health. The README's own compose example uses MT_OFFLINE=false so models can be fetched on demand:
services:
mtranserver:
image: xxnuo/mtranserver:latest
container_name: mtranserver
restart: unless-stopped
ports:
- "8989:8989"
environment:
- MT_HOST=0.0.0.0
- MT_PORT=8989
- MT_OFFLINE=false
volumes:
- ./models:/app/modelsStart it with docker compose up -d, then point a client at one of the compatibility endpoints. The README's table gives, for example, http://localhost:8989/imme for Immersive Translate without a password, or http://localhost:8989/imme?token=your_token when MT_API_TOKEN is set. The README also says the CLI supports --download pairs and --languages, and that both require network access. Note that the truncated CLI help ends mid-sentence on that point, so check the current output of --languages before scripting against it.
Where MTranServer is the wrong choice
Quality is the first limit and the project states it plainly. The README says translation quality is certainly not as good as large-model translation and directs users who need high quality to an online large-model API. If your output is marketing copy, legal text or anything a human will read as finished prose, this is the wrong tool, and no amount of tuning changes that, because the constraint is model size and the optimisation choices that buy the 50 ms figure.
The second limit is the model download. The README says the first translation of a language pair downloads the corresponding model unless offline mode is enabled, and that the wait depends on network speed and model size. That means a fresh deployment is not actually offline until the models are present. If you run in an air-gapped network, you must pre-populate the model directory, and the README does not describe a supported way to fetch models on a connected machine and move them; the CLI offers --download pairs, but the help text notes it requires network access. Verify that path yourself before designing around it.
The third limit is the compatibility surface. The server emulates DeepL, DeepLX, Google Translate, Immersive Translate, Kiss Translator and a selection-translation client. Each of those is an external contract that can change. A client update that alters request or response shape can break the corresponding endpoint here, and the fix depends on this project tracking someone else's API. Treat each endpoint as a separate integration with its own risk, not as one stable API.
MTranServer compared with LibreTranslate and hosted APIs
The closest comparison is LibreTranslate, and the difference is the inference stack rather than the deployment story. Both are self-hosted translation servers that expose an HTTP API and keep text on your machine. LibreTranslate is a Python application built around Argos Translate and OpenNMT-family models, which makes it straightforward to extend and to run on a CPU, but its per-request latency is typically well above the 50 ms average the MTranServer README claims. MTranServer compiles its inference path in C++, ships a single Node binary and an Alpine image, and its whole pitch is that latency plus a small memory footprint. If you want a Python service you can patch, LibreTranslate is the more natural base. If you want a small binary answering short requests fast, MTranServer is the more direct fit.
Against hosted APIs the difference is categorical. DeepL and Google Translate are services: you send text to a third party, you pay per character or per month, and quality is high. MTranServer sends nothing anywhere and costs nothing per request once the models are on disk, and the README frames the goal as unlimited free translation. The trade is quality and, for anyone behind a proxy or in a regulated environment, the absence of a data-processing agreement to negotiate. The /deepl and /google endpoints here are wire-compatible shims, not the real services, and clients that assume DeepL-specific glossaries, document translation or formality controls will find those features absent.
Maintenance, versions and the Apache-2.0 licence
The repository is not archived. Its last push was on 2026-03-08, and the most recent release, v4.0.33, was published the same day. The two releases before it, v4.0.32 and v4.0.31, both landed on 2026-01-18. That is a gap of roughly seven weeks between the January pair and the March release, with nothing after March. The README's v4 note says that v4 improved memory usage, further increased speed and improved stability, and that users on an older version should upgrade immediately. The CHANGELOG.md and HISTORY.md files at the repository root are where release detail lives, and the README explicitly warns that the program is updated frequently and that updating to the latest version is the first thing to try when something breaks.
Upgrade cost is low in the common cases. The npm package is installed by name at @latest, the Docker image is xxnuo/mtranserver:latest, and the desktop builds are downloaded from the Releases page, so an upgrade is a re-pull or a re-install. The one thing to preserve across upgrades is the model directory, which is why the compose file mounts it as a volume: models are large and re-downloading them is the slow part. If you pin a version, pin the model directory too.
On licensing, the project is Apache-2.0, which permits commercial use and modification and includes an explicit patent grant. The models the server downloads are separate artefacts with their own provenance, and the README does not state their licence. If you redistribute a container image with models baked in, or ship the desktop build to customers, confirm the model licences yourself; the repository's LICENSE file covers the code, not necessarily the weights. This is not legal advice.
Editorial conclusion
Adopt MTranServer if you want translation that never leaves your machine, runs on CPU, and answers in tens of milliseconds, and if you can accept quality below a large online model. Do not adopt it if translated output is customer-facing prose where fluency matters more than latency, or if you cannot host a process that downloads model files on first use. Before committing, run npx mtranserver@latest, translate one pair such as en_zh once to force the model download, then measure that pair against your own text; check the /imme, /kiss and /deepl endpoints against the client you actually plan to use, and decide whether MT_API_TOKEN belongs in your compose.yml.
Frequently asked questions
How do I install MTranServer?
The README gives three routes: run npx mtranserver@latest for a one-off start, install globally with npm i -g mtranserver@latest and then run mtranserver, or deploy the xxnuo/mtranserver:latest image with the provided compose.yml. Desktop installers for Windows, Mac and Linux are also available from the Releases page.
Does MTranServer need a GPU?
No. The README describes it as an offline translation model server that requires no graphics card, and the Docker image runs on node:22-alpine with models loaded from a local directory.
Why is the first translation so slow?
The README states that the first time a language pair is translated the server downloads the corresponding model unless offline mode is enabled, and that the wait depends on network speed and model size. After that download, the README says requests get millisecond-level responses, so it recommends testing one translation before real use.
How do I use MTranServer with Immersive Translate?
Point the extension's custom API URL at http://localhost:8989/imme, or at http://localhost:8989/imme?token=your_token when MT_API_TOKEN is set. The README also suggests raising the extension's maximum requests per second to make use of the server's throughput.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/xxnuo-mtranserver)