# SoniTranslate's install depends on two gated model licences, and two of its dependencies come from pull request refs

> An Apache licensed Gradio application for translating and dubbing video with synchronized audio, with about eighty-seven languages available for transcription and a separate list of roughly twenty-three that are translation targets only. Local installation needs an NVIDIA driver, two accepted model licences, a token with a specific scope, and a pinned CUDA generation, and the primary documentation is a video tutorial.

**R3gm/SoniTranslate** — Synchronized Translation for Videos. Video dubbing

- Repository: https://github.com/R3gm/SoniTranslate
- Stars: 1,417 · Forks: 349
- Language: Python
- License: Apache-2.0
- Published: 2026-09-16 · Updated: 2026-09-16 · Language: en
- Canonical page: https://hysenlabs.com/projects/r3gm-sonitranslate

## The readme opens with a paid placement and spells the name three ways

The first thing in the readme is not about the project.

Before any content there is a block recommending a commercial meeting transcription API, with a tracking-tagged link, described as something to consider if you are looking for a meeting transcription service that records video conferencing calls and in-person meetings. That is a sponsored placement, and the tracking parameters in the link are visible in the source.

Then the name, which appears three ways. The repository, the headings and the project links all say one thing, with the second syllable written the way it sounds. The prose says it a second way, with the vowel swapped, and a heading repeats that second spelling. So the introduction describes a differently spelled product from the one you arrived at.

What the project is, once you get past that, is a web application for translating videos into other languages with synchronized audio. It is built on a Python web UI framework rather than being a library or a command line tool, and the description is marketing copy about being user friendly.

Three ways to run it are offered: a hosted notebook, the repository, and an online demo hosted as a space on a model platform. The primary documentation, though, is a video tutorial by a contributor outside the project, which the readme recommends for understanding how it works.

## Two releases, seven minutes apart, in one day in 2024

The release history is two entries long, and both landed on the same afternoon.

Version 0.4 and version 0.5 were published seven minutes apart on one day in May 2024. There is nothing before them and nothing after them.

Meanwhile the last commit to the main branch is from August 2026, which puts the branch more than two years ahead of the newest tag. That is the normal state for a project whose releases are occasional, but it means the version you get depends on which of three things you install: the newest tag, the main branch, or one of the two notebooks.

There is no version constraint anywhere in the documentation to help you choose. For a project with a CUDA-pinned dependency tree and dependencies installed from moving references, that distinction matters more than it would for a library with a small surface.

The repository description is one line, and it names the two things the project does: synchronized translation for videos, and video dubbing. Everything else in the readme elaborates on those.

## Local install needs a driver, two accepted licences, and a token with a named scope

The install section opens with four prerequisites, and only one of them is a package manager.

The first is an NVIDIA driver for a specific older CUDA generation, with a link to the toolkit archive rather than to current instructions. That pins the whole installation to one generation of GPU runtime, and the dependency file agrees: the PyTorch packages are requested with a build suffix for that same generation, alongside an extra package index URL serving those wheels.

The second is a licence acceptance. Speaker diarization and segmentation models are gated on a model host, so you need an account there and must accept the licence for both of them before anything will download. This is not an optional extra; it is on the critical path.

The third is a token, and the readme specifies its scope precisely. When creating it you are told to tick a specific permission: read access to the contents of all public gated repositories you can reach. Without that scope the automatic model downloads fail, and the error will be about permissions rather than about the app.

The fourth is a Python distribution, with the full and minimal versions both offered.

And the section heading is honest about coverage: installation is tested on Linux. Nothing in the readme says what happens on Windows or macOS, which for a project pinned to an NVIDIA CUDA generation is the expected answer but is not stated.

## Two dependencies install from pull request merge references

The dependency file is where this project is least conventional, and the details matter for anyone who intends to build it.

One package is installed directly from a pull request merge reference rather than from a release. Another, which is commented out, does the same. A pull request merge reference is not a version: it points at whatever the head of that pull request is, and if the branch is rebased or the pull request is closed, a fresh install resolves to something different or fails outright.

A third dependency comes from a branch with a one-word name suggesting an in-progress state, rather than from a tag. A fourth is pinned to a specific commit, which is the correct approach and makes the other three look careless by comparison.

Then there is the pinned-versus-floating split across the rest. The web framework, the audio library, the translation client, the speaker-similarity index and the tensor runtime are all pinned to exact versions with no upper bound, while several model libraries are unpinned entirely. For a project that pulls large model weights, exact pins on the small packages and no pins on the large ones is an odd risk profile.

The net effect is that a working local build is not reproducible from the file alone, which is why building it once and keeping the environment is the practical advice.

## Three text-to-speech clients and two speech analysis libraries, because this is dubbing

The dependency list is dense, but it clusters into one theme: this project treats the voice as the deliverable.

Three separate text-to-speech stacks are present. One is a cloud edge service client, one is Google's, and a third requirement file exists specifically for a neural synthesis model. So voice output can come from a cloud service, a Google service, or a local neural model, and the presence of a dedicated extra file suggests the local option is the heaviest and is installed separately.

Voice conversion, rather than only synthesis, is what two root scripts are for: one application module named for a voice conversion technique, and a pipeline script. A model directory at the repository root holds the assets such conversion needs. So the pipeline is not read this, write that; it is generate or convert a voice, then align and mix it.

The supporting cast tells you how the alignment is done. A phonetics analysis library and a speech analysis library are both present, which is what you install when you need pitch and formant measurements. A pitch-shifting library is pinned exactly. A speaker-similarity index on the CPU suggests voice matching against a reference, which is how you tell whether the converted voice resembles the original speaker.

Rounding it out: a video downloader, a subtitle format library, document parsers for text and word files, an archive reader, and a downloader for hosted files. This is a full media pipeline.

## Eighty-seven languages can be transcribed and about twenty-three cannot, and nobody says why

The language table is the longest part of the readme and the most interesting, because it is split in two and the split is unexplained.

The first table lists language codes available for transcription, running from English through to Yoruba, with several dozen entries covering European, South Asian, East Asian, Middle Eastern and African languages. Two details stand out. Simplified and traditional Chinese are separate codes rather than variants of one, and one language is listed with two acceptable codes, which is a sign the list was assembled from a source that allowed synonyms.

The second table is headed as non-transcription and lists roughly twenty-three more languages: Aymara, Bambara, Cebuano, Divehi, Dogri, Ewe, Guarani, Iloko, Kinyarwanda, Krio, Kurdish, Kirghiz, Maithili, Quechua, Samoan, Tigrinya, Akan, Uighor and others.

So roughly twenty-three of these languages can be translated into but not transcribed from. That is a meaningful capability boundary for anyone building a localisation pipeline, and the readme offers no explanation for it. It is almost certainly about the availability of multilingual speech models, but the readme does not say so.

The practical consequence is that the language list is not a statement about what the app can translate. It is two lists with a boundary in the middle, and you have to notice the boundary to know which languages work end to end.

## The only worked example is two audio files, and the demo runs on someone else's space

The evidence offered for the project working is minimal, and that is worth calibrating expectations against.

The example section contains two links and two labels: original audio, and translated audio. Both are attachments hosted by the repository's own asset store. There is no description of what the source was, no settings used, no output transcript, and no runtime. So the example demonstrates that the output sounds like speech, and nothing else.

The online demo is a hosted space on a model platform. That is genuinely useful for evaluating whether the voice quality is acceptable before you install an NVIDIA driver, and it is the lowest-friction entry point in the whole project.

For running without a local GPU, there are two notebooks at the repository root, one of them a variant with something embedded. Both are linked from the readme, so the hosted-notebook route is a first-class option rather than a workaround.

Beyond that the repository has an assets directory for media, a documentation directory, a library directory and the main package directory. Four separate requirement files split the install: a base set, an extra set, a specific one for the neural synthesis model, and the combined one you would actually use. That split is the clearest signal that the dependency surface grew faster than anyone wanted to maintain in one file.

## Conclusion

SoniTranslate is aimed at people who want a translated audio track rather than subtitles, and the pipeline behind it, transcription, translation, voice conversion and mixing, is more machinery than a translation script. Two things to check before you commit to a local install. The dependency list is pinned to an old CUDA generation and installs two packages from pull request merge references, which are mutable in a way a commit hash is not, so build it and keep the artefact. And note that the documented path depends on accepting licences for two third-party models, so this is not a zero-setup install even though a hosted notebook exists for trying it.

## FAQ

### How do I install SoniTranslate?

Locally, the install is stated as tested on Linux and needs four things first: NVIDIA drivers for a specific CUDA generation, acceptance of the licences for two gated speaker-diarization and segmentation models on a model host, a token from that host with read access to gated repositories ticked, and either a full or minimal Python distribution. Two hosted notebooks and an online demo let you try it without any of that.

### What is SoniTranslate?

An Apache-2.0 licensed Python application, built as a web interface on the Gradio framework, for translating videos into other languages with synchronized audio. Its description calls it synchronized translation for videos and video dubbing, and it can be run from the repository, from a hosted notebook, or as a hosted demo space.

### Is SoniTranslate free to use?

The code is under a permissive licence and there is a hosted demo plus notebooks you can run without a local GPU. There is no licence fee. The readme does carry a sponsored placement at the very top recommending a commercial meeting transcription API, with tracking parameters in the link, which is worth knowing about when you read the documentation.

### Which languages does SoniTranslate support?

Two lists rather than one. Around eighty-seven language codes are available for transcription, including Simplified and traditional Chinese as separate entries and one language listed with two accepted codes. A second list of about twenty-three languages, from Aymara to Uighor, is marked as non-transcription, meaning they can be translated into but not transcribed from. The readme does not explain the boundary.

### What are the voice-related scripts in SoniTranslate for?

The pipeline treats the voice as the output rather than the text. The repository root holds an application module for voice conversion, a dubbing pipeline script and a main voice module, with a model directory for the assets conversion needs. Dependencies include three separate text-to-speech clients, two speech analysis libraries, a pitch-shifting library and a speaker-similarity index, which together suggest generate or convert, then align and mix.

### Can I reproduce a SoniTranslate installation from its dependency file?

Not reliably. Two dependencies are installed from pull request merge references, one of them commented out, a third from a branch with an in-progress name, and only one from a pinned commit. Several model libraries are unpinned entirely while the small packages are pinned exactly, and the deep learning stack is pinned to one CUDA generation. Build it once and keep the resulting environment rather than rebuilding from the file.

## Sources

- [Issues](https://github.com/R3gm/SoniTranslate/issues)
- [License: Apache-2.0](https://github.com/R3gm/SoniTranslate/blob/main/LICENSE)
- [R3gm/SoniTranslate on GitHub](https://github.com/R3gm/SoniTranslate)
- [README](https://github.com/R3gm/SoniTranslate/blob/main/README.md)
- [Releases](https://github.com/R3gm/SoniTranslate/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/r3gm-sonitranslate
