Model or dataset
rockbenben/subtitle-translator avatar
rockbenben/subtitle-translator

Subtitle Translator: Batch SRT, ASS, VTT and LRC Translation Without Touching Timecodes

Translate a whole season of subtitles in one pass — .srt/.ass/.vtt/.lrc, 120+ languages, 27 LLM providers, timing untouched | 整季字幕一次译完,时轴不动,支持 120+ 语言

1,144 stars155 forksTypeScriptMIT

At a glance

What is it?
A browser-based and CLI tool that pulls dialogue out of subtitle files locally, sends only text to 8 machine translation APIs or 27 LLM providers, and writes the timing back unchanged. Here is how the pipeline works, how to run it, and where it breaks.
Who is it for?
Adopt it if you routinely translate whole seasons of subtitles and want the timeline to stay byte-identical while you pick your own engine. Skip it if you need burned-in captioning, speech recognition from video, or a guarantee of literary quality from a small local model, because the README itself warns that models under 70B parameters may misalign context-mode output.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 4 days ago.
What is it written in?
Mainly TypeScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The two failures Subtitle Translator was built around

Pasting an .srt file into a general-purpose chat translator usually produces two distinct problems. The model rewrites or drops timecodes, and the workflow is one file at a time. Subtitle Translator addresses both by splitting the file locally before any network call happens. The README describes the split directly: timecodes, cue numbers, ASS headers and VTT cue IDs are extracted on the client, and only dialogue text is sent to the engine, so "the timeline physically cannot be touched."

That framing matters because it defines the tool's scope. It is not a speech-to-text system and it does not read video. You bring subtitle files that already exist, in .srt, .ass, .vtt or .lrc form, and it returns translated files with the same structure. The audience is therefore narrow and specific: people who already have captions and need them in another language, often a whole season at once, often into more than one target language. Fansub workflows, course material, archival transcription projects, and anyone sitting on a folder of untranslated .srt files fit that description. Someone who has a video and no subtitles is not the user.

How the parse, translate and rebuild pipeline is arranged

The architecture is a client-side Next.js application with a headless CLI sharing the same engine. Parsing and reassembly happen in the browser, which is why the privacy claim is structural rather than a policy promise: the README states that subtitle content and API keys never touch a server, and that LLM requests go directly from the browser to whatever endpoint you configure.

Translation is chunked and parallel. The README describes chunked compression plus parallel processing producing roughly one second per episode, with GTX noted as slightly slower. For LLM modes, two settings govern the batch shape: Concurrent Lines, the maximum number of lines translated in parallel, defaults to 20, and Context Lines, the number of surrounding lines included per batch as context, defaults to 50. The README warns that raising concurrency too high triggers rate limits, and that raising context increases coherence at the cost of tokens. That is the central tuning trade-off of the whole tool, and it is exposed rather than hidden.

Caching runs through IndexedDB, described as having no browser-storage size limit, so a page refresh does not lose translated files. The dependency list is consistent with this design: idb and idb-keyval for storage, p-limit and p-retry for concurrency and retries, jschardet for character-set detection, spark-md5 for hashing. Nothing in the dependency set suggests a server-side translation proxy, which matches the client-only claim.

Installing Subtitle Translator with Docker or running it locally

The README points to a hosted instance at tools.newzone.top, so the fastest path is no installation at all. For self-hosting, the repository ships a Dockerfile that builds in standalone mode with local API support enabled. The Dockerfile comments give the build and run commands directly.

bash
docker build -t subtitle-translator .
docker run -d -p 3000:3000 --name subtitle-translator subtitle-translator

After the container starts, the app listens on port 3000 on the host. The image runs as a non-root user and sets PORT=3000 and HOSTNAME="0.0.0.0" internally, so the port mapping above is the one to keep.

For a source checkout, package.json pins yarn 1.22.22 and requires Node 20.9.0 or newer. The scripts block exposes dev, build, start and cli.

bash
yarn install
yarn dev
yarn cli

The first command installs dependencies, the second starts the Next.js development server, and the third launches the command-line entry point at scripts/cli.ts through tsx. The README describes the CLI as running "the same engine, parsers, and cache" from a terminal. A first real use is to drop a single .srt into the web interface, leave the source language on Auto-detect, pick a target, and watch the exported file appear with its original filename. Batch mode is the same gesture applied to a folder: the README says each file translates and downloads independently and the run ends with an aggregated summary such as "Exported (3/5)".

Where the format handling gets opinionated

Format support is not uniform, and the differences are worth reading before you plan a workflow. .srt, .ass, .vtt and .lrc are auto-detected. WebVTT NOTE, STYLE and REGION blocks are skipped rather than treated as dialogue, which is the correct behavior and also the kind that is easy to get wrong. One-click conversion is offered during translation: SRT to VTT, and SRT or VTT to ASS.

Bilingual output inserts the translation above or below the original with alignment preserved. For SRT and VTT sources specifically, you can export ASS with separate styles for original and translation, described as Default 70pt white plus Secondary 55pt cyan, editable in any subtitle editor. That is a concrete, inspectable default rather than a vague "bilingual support" claim, and it tells you the intended use: burned-in or player-rendered dual-language playback where the two tracks must be visually distinguishable.

Multi-language output is the other notable choice. A single pass can target several languages, and each is exported as its own file with the language code appended, for example movie.zh.srt and movie.fr.srt. If you need five languages from one source season, that is one run rather than five. The cost is that all five consume API quota in parallel, which interacts badly with the concurrency limits mentioned above.

The provider list is wide, and that width is the point

Subtitle Translator separates traditional machine translation APIs from LLM providers. The traditional side has eight: DeepL, Google Translate, Azure Translate, DeepLX, Qwen-MT, TranslateGemma, GTX and Edge. GTX and Edge require no configuration and serve as each other's fallback, which is why a first run works before you have entered a single key. The README's own table lists free tiers, including 500K characters per month for DeepL and Google, 2M characters per month for Azure during the first 12 months, and self-hosting for DeepLX and TranslateGemma.

The LLM side is larger: DeepSeek, OpenAI, Claude, Gemini, Qwen, Moonshot, Doubao, Xiaomi MiMo, Zhipu GLM, MiniMax, Baidu ERNIE, Tencent Hunyuan, Mistral, xAI, Perplexity, Cohere and YandexGPT, plus gateways including OpenRouter, OpenCode Zen, Groq, SiliconFlow, Atlas Cloud, GitHub Models, Nvidia NIM, Azure OpenAI and LiteLLM, and any OpenAI-compatible custom endpoint such as Ollama, LM Studio, vLLM, Together AI or Fireworks AI.

The practical consequence is that provider choice becomes a real decision instead of a constraint. LLM modes add configurable system and user prompts, a temperature control on a 0 to 1 scale, and a per-provider thinking-mode toggle for reasoning-capable models. Traditional APIs give you none of that. If you want a house style for character names, you need an LLM mode and you need to write the prompt yourself.

CORS, relays, and the limits the README states plainly

The most consequential limitation is network topology, not translation quality. Because requests originate in the browser, any provider that blocks browser-origin requests through CORS cannot be called directly. The README acknowledges this and offers a relay: a built-in relay works out of the box, and API Settings exposes a Relay address field that points every relayed provider at your own deployment of the relay Worker. That is a real escape hatch, but it means a self-hosted deployment behind a restrictive network may need a second piece of infrastructure that is not part of the Docker image.

The second limitation is stated as a warning. In context mode, models under 70B parameters may produce misaligned output, and the README recommends mainstream online large models such as Claude, GPT, DeepSeek and Gemini. This is a direct admission that the feature which most improves dialogue coherence is also the feature most sensitive to model capability. If your plan is to point this at a small local model through Ollama to avoid API costs, context mode is the wrong setting to enable.

The third is throughput tuning. Concurrent Lines defaults to 20 and the README says too high a value triggers rate limits. There is no documented auto-throttling that adapts to a provider's limits, so the burden of finding a safe concurrency per provider sits with the user. The README also does not document rollback or a way to revert a completed batch, which matters if a bad prompt style is applied across a whole season before anyone checks the output.

How it differs from subtitle editors with translation built in

The obvious alternative is a desktop subtitle editor with a translation plugin, such as Subtitle Edit or the translation features bundled into similar tools. The difference is architectural. An editor loads a file into a timeline UI, holds it in memory, and translation is one operation among many; batch behavior depends on the editor's own scripting or command-line mode, and the translation engine is whatever the plugin author wired up.

Subtitle Translator inverts that. There is no timeline editor. The file is a unit of work, the engine is a pluggable choice among 8 APIs and 27 LLM providers plus gateways, and the batch is the primary interface. The CLI exists to script the same engine rather than to drive a GUI. If your job is to review and hand-fix timing drift, an editor is the better tool. If your job is to convert four hundred files from Japanese into English and French with the timing untouched, the editor's single-file model is the obstacle.

A second alternative is a general-purpose LLM chat interface with file upload. That gets you good translation and no batch management, no caching, no per-file export naming, and no structural guarantee about timecodes. The README's opening argument is aimed exactly at that workflow.

Editorial conclusion

Adopt it if you routinely translate whole seasons of subtitles and want the timeline to stay byte-identical while you pick your own engine. Skip it if you need burned-in captioning, speech recognition from video, or a guarantee of literary quality from a small local model, because the README itself warns that models under 70B parameters may misalign context-mode output. Before committing, verify three things: that your chosen provider accepts browser-origin requests or has a relay configured under API Settings, that your target format survives the bilingual export you plan to use, and that the Node version on your machine is at least 20.9.0.

Frequently asked questions

What is Subtitle Translator and what files does it accept?

It is a free, browser-based batch subtitle translation tool that auto-detects .srt, .ass, .vtt and .lrc files and translates them into 120+ languages. It extracts timecodes, cue numbers, ASS headers and VTT cue IDs locally so only dialogue text is sent to the translation engine.

Is there a free AI subtitle translator available?

Yes. The GTX and Edge APIs need no configuration at all and are described as the zero-setup defaults, with each serving as the other's fallback. Several configured providers also list free tiers, including DeepL and Google at 500K characters per month and Azure at 2M characters per month for the first 12 months.

How do I translate a subtitle file with Subtitle Translator?

Open the hosted instance or a self-hosted copy, drop one or more subtitle files in, leave the source language on Auto-detect, choose a target language, and start the run. Each file translates and downloads independently with its original filename, and the run ends with an aggregated summary such as "Exported (3/5)".

What is the best subtitle translator?

The README does not rank tools against each other, so it offers no comparative verdict. What it does describe is a client-side pipeline that strips timecodes locally before sending dialogue to 8 traditional APIs or 27 LLM providers, which is the basis on which you would compare it with any other option.

What is the best app for translating subtitles?

The repository presents two interfaces rather than one app: a browser-based batch interface and a CLI reachable through yarn cli that the README says runs the same engine, parsers and cache. Which one suits you depends on whether you prefer dropping files into a page or scripting the run.

Official sources

  1. License: MIT
  2. Project website
  3. README
  4. Releases
  5. rockbenben/subtitle-translator on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/rockbenben-subtitle-translator.svg)](https://hysenlabs.com/projects/rockbenben-subtitle-translator)