SmartSub: local subtitles, translation and dubbing in one desktop app
视频转字幕、字幕翻译、AI 配音与声音克隆、字幕烧录——免费开源的一站式桌面工具。基于 Whisper / FunASR 等本地模型离线语音转文字,批量处理 + 全平台 GPU 加速,跨 Windows / macOS / Linux。Free, open-source desktop app to generate, translate, dub & burn video subtitles — local Whisper speech-to-text, AI dubbing & voice cloning, offline, GPU-accelerated.
At a glance
- What is it?
- SmartSub is a free, MIT-licensed desktop app wrapping the whole subtitle pipeline, download, transcription, translation, proofreading, TTS dubbing and burning, around local models with files that never leave the machine. It also exposes 111 MCP tools and a CLI for automation.
- Who is it for?
- Use SmartSub if you process video in bulk and want transcription, translation and dubbing without sending files to a service or paying for API keys. Skip it when a one-off transcript is the whole job, where a plain whisper.cpp invocation is less machinery, or when your work already centers on a professional video editor.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 6 days ago.
- What is it written in?
- Mainly TypeScript, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 28, 2026, and from our analysis. They are not legal advice.
Editorial analysis
A five-stage pipeline inside one Electron window
SmartSub, whose Chinese name is Miaomu, packs speech-to-text, subtitle translation, proofreading, TTS dubbing and final burning into a single desktop application, with online video download bolted on the front so a pasted Bilibili or YouTube link can enter the pipeline directly. Transcription runs on local models such as whisper.cpp and sherpa-onnx, files stay on the machine, batches are processed with adjustable concurrency, and hardware acceleration covers NVIDIA, AMD, Intel and Apple Silicon on Windows, macOS and Linux. The licence is MIT and the app is built on Nextron, the Next.js plus Electron combination visible in the repository. The documented free path needs no API key anywhere: local models transcribe, built-in free sources translate, local engines handle dubbing including cloning, and ffmpeg burns, with no usage limits on the local steps. Optional cloud capacity, 20 translation services, 9 dictation providers and 6 dubbing services, sits beside that rather than replacing it.
Eight transcription engines, switched per task
Transcription is not one implementation but a roster: whisper.cpp, faster-whisper, FunASR, Qwen3-ASR, FireRedASR, NVIDIA Parakeet and a local Whisper CLI, plus GPU-free cloud dictation across 9 providers, with the engine selectable per task. Guidance in the documentation is language-specific, FunASR or FireRedASR for Chinese, Parakeet models for English, other European languages and Japanese. Local engines run fully offline with no upload. An optional AI refinement pass fixes what raw transcripts get wrong: semantic re-segmentation that keeps timing accurate to the word, so conjunctions do not dangle at line ends and numbers are not split by pauses, followed by homophone correction, filler removal and punctuation normalization. The refinement provider defaults to the AI translation configuration, so a local Ollama model makes it free, and failures fall back to rule-based segmentation. Running whisper.cpp by hand gets you a transcript; this is the same engine plus everything after it.
Twenty translation services and a bilingual mode
Translation offers 20 services. The built-in free option uses Bing and Google free endpoints with automatic fallback and rate limiting, and the paid and model-based list runs from Baidu, Aliyun, Tencent, iFlytek, Volcano Engine, Doubao and Niutrans through DeepLX, Azure and Google, to LLM services including Ollama running locally, DeepSeek, Gemini, Qwen, SiliconFlow, Azure OpenAI and DeerAPI. Any OpenAI-style endpoint can be registered, so self-hosted models slot in. Output is either a pure translation or bilingual subtitles carrying original and translated lines together, which is the common need for language-learning and export content. Each service exposes custom request parameters in the interface, with import and export of configurations, so tuning a provider never becomes a code change. That breadth is the design point: translation is a commodity with many price points, and the app refuses to pick one for you.
Local voices, zero-shot cloning, and the over-limit list
The dubbing workbench takes one subtitle file plus an optional video and synthesizes speech line by line, aligned back to the timeline. Local engines are offline and free: Kokoro with 103 multilingual voices and VITS with 174 Chinese voices. Voice cloning runs locally through ZipVoice in zero-shot mode, one reference recording builds the voice, while cloud cloning is available through Volcano Engine Voice Replication 2.0 and ElevenLabs Instant Cloning, with Edge TTS free tier, OpenAI-compatible endpoints, Azure Speech, Volcano Doubao and Xiaomi MiMo as further cloud TTS. Alignment is engineered rather than assumed: speech rate is pre-controlled, durations are re-measured, silence gaps are borrowed, and lines that still exceed their slot land on a manual handling list with three choices, rewrite the text, regenerate the single line, or accept a speed change. Output covers pure wav or mp3 audio, track replacement, mixed video or MKV dual audio tracks, with ducking or muting for the original audio.
111 MCP tools, a CLI, and a headless daemon
Version 3.9.0's headline is automation depth. SmartSub exposes 111 standard MCP tools with matching CLI commands covering download, transcription, translation, proofreading, dubbing, muxing, format conversion, audio extraction and model configuration. The bundled Electron/Node runtime means clients and external agents need no Node.js install. Integration is one-click by client: Cursor through the official MCP protocol, Codex via a copyable TOML configuration, Claude Code via one-click registration, plus a generic JSON block. A headless daemon starts on first call, shares task state, configuration and model resources with the desktop client, and is reclaimed when idle, so scripted and interactive use coexist. The smartsub CLI supports pipeline.run orchestration, JSON and stdin parameter piping, task status polling and deduplication, which is what makes the pipeline scriptable inside CI rather than merely inside the app.
A copilot that screenshots its own app to debug it
The built-in AI assistant opens from a hotkey, Command+J on macOS and Ctrl+J elsewhere, as a resizable drawer with locally persisted sessions. Its context awareness is specific: it sees the current video playback position, the selected subtitle line, neighboring lines, task list state and recent error logs. Commands like converting the current video to bilingual subtitles and composing it, listing installed models, or diagnosing a failed task map onto the same 111 tools the MCP surface exposes. Subtitle edits render into the editor immediately, enter the undo and redo stack, and only touch the physical file when you say save, a versioning discipline many editors lack. A camera icon captures the workspace so a multimodal model can read error dialogs, toasts and control states, and file chat accepts audio, video, subtitles, reference texts and images, passing local absolute paths instead of uploading media. Providers include DeepSeek, Qwen, Gemini, SiliconFlow, DeerAPI and any OpenAI-compatible endpoint.
Burning, downloading, and CUDA simulated in dev
Composition offers two subtitle modes: hard burn, permanent in the picture and playable anywhere, and soft mux, a stream-copy that keeps quality lossless and the track switchable, with font, size, color, stroke, shadow and nine-position placement under live preview. Downloading pairs two engines automatically, yt-dlp for YouTube and 1800-plus sites, lux for Chinese platforms such as Bilibili, Douyin and Xiaohongshu, fetches platform-provided subtitles including auto-generated ones and pairs them automatically, and imports site cookies, stored locally only, for member content. Releases are frequent and labeled in Chinese: v3.9.0 on 2026-09-24 brought the AI assistant and MCP toolbox, v3.8.0 on 2026-09-10 added multi-format subtitle export, and v3.7.0 on 2026-08-06 added multi-voice dubbing and speaker recognition, with the last push on 2026-09-24. A detail from package.json shows the engineering behind GPU support: dev scripts simulate CUDA with DEV_SIMULATE_CUDA, DEV_SIMULATE_PLATFORM and version overrides, so the Windows GPU paths are testable without the hardware.
Editorial conclusion
Use SmartSub if you process video in bulk and want transcription, translation and dubbing without sending files to a service or paying for API keys. Skip it when a one-off transcript is the whole job, where a plain whisper.cpp invocation is less machinery, or when your work already centers on a professional video editor. Verify first that your GPU is covered by the acceleration paths, NVIDIA CUDA, AMD or Intel Vulkan, Apple Core ML or Metal, and note that the app documents an automatic fallback to CPU when an accelerator fails to load.
Frequently asked questions
What is SmartSub?
SmartSub is a free, open-source desktop app that turns video into subtitles, translates them, adds AI dubbing with voice cloning, and burns subtitles into the final file. Transcription runs on local models such as whisper.cpp and FunASR, offline, with GPU acceleration on Windows, macOS and Linux.
Does SmartSub require an API key?
No. The whole pipeline can run for free: local model transcription, built-in free translation sources, local TTS dubbing including voice cloning, and local ffmpeg burning, with no API key and no usage limits on the local steps. Cloud services for translation, dictation and dubbing are optional enhancements.
Can SmartSub download online videos?
Yes. Paste Bilibili or YouTube links and the app downloads them using two engines matched automatically: yt-dlp for YouTube and 1800-plus sites, and lux for Chinese platforms. Platform-provided subtitles can be fetched and paired automatically, and site cookies can be imported and are stored locally only.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/buxuku-smartsub)