CLI tool
HAKORADev/VODER avatar
HAKORADev/VODER

VODER: eight audio modes, offline voice cloning and a DLC system that adds image and video

Voice Operation and Design Engine with Reproduction capabilities

519 stars92 forksPythonAGPL-3.0

At a glance

What is it?
VODER is a Python toolkit that puts speech-to-text, text-to-speech, voice conversion, music generation, enhancement, sound effects, vocal separation and diarization behind one local CLI, with translation into 76 languages and dubbing that keeps the speaker's voice. The Project Eva DLC extends it into image, video, 3D world and chat generation, and is marked in development.
Who is it for?
VODER suits people who do a lot of voice work and want it on their own hardware with no subscription and no upload: podcasters assembling multi-speaker dialogue, translators dubbing videos, and anyone who needs cloning, separation and music generation reachable from one command.
Can I use it commercially?
Yes, with strict conditions. AGPL-3.0 is a network copyleft licence: if people use a modified version over a network, for example as a hosted service, you must offer them its source code under the same licence.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Eight modes behind one interface

Most audio tools do one job. VODER bundles eight and puts them behind a single command surface: speech-to-text, text-to-speech, voice conversion, music generation, speech enhancement, sound effects, vocal separation and speaker diarization. The pitch in the README is local, free and offline, with no subscription and no requirement for a GPU.

That framing matters for who this is for. Sending audio to a hosted API is fine until the material is confidential, the volume is high, or the work has to happen on a machine with no reliable network. VODER takes the other side of that trade: you own the compute and the model downloads, and nothing leaves the machine.

The cost of breadth is surface area. A tool that wraps eight model families has eight upgrade paths, and the requirements file shows it: pinned versions of torch 2.8.0, torchaudio 2.8.0 and transformers 4.57.3 alongside openai-whisper, pyannote for diarization, gradio, yt-dlp and easyocr for OCR. Expect the install to be the hard part.

Voice design, cloning and the dialogue system

Two ways to get a voice are supported and can be mixed. You can describe one in plain English, the README example being a deep and authoritative male voice, or supply a reference clip and have the speaker cloned from it. Characters in the same dialogue can use either approach, so a script can combine invented voices with cloned ones.

The dialogue layer is more capable than a text-to-speech call. Scripts support directives for timing, level and duration, written as /time, /level and /duration, and sound effects can be embedded in a line with the special sfx: character, generated from a text description. Background music can be produced to match the spoken duration exactly, mixed at a configurable volume with fade transitions, and an optional reference audio, video or URL can steer its style, with the reference run through vocal separation first to get clean instrumental.

Clones are reusable rather than ephemeral. The train task saves them as .tts or .ttse files, so a voice you like becomes an asset you can load into later work instead of something you re-describe each time.

Translation, dubbing and transcribe-edit-resynthesize

Translation is decoupled from transcription, which is the design decision worth noting. Any-to-any translation runs through TranslateGemma 12B using a translate (source-target) syntax, covering 76 languages and independent of whichever ASR engine produced the transcript. Speech-to-text itself accepts audio, video, images and direct URLs, and can batch process multiple files with speaker diarization to identify who spoke when.

Dubbing is the combination that most tools make you assemble yourself: speech is translated while preserving the original speaker's voice identity, with per-segment timing alignment and background music preserved across the whole video.

There is also a correction path. Transcribe-edit-resynthesize is built into the interactive TTS mode, and the tts svc command re-reads transcribed speech in a different voice, with an optional sts: prefix that switches to high-fidelity conversion through Seed-VC v2. That is the right shape for fixing a line: transcribe, edit the text, regenerate just that segment in the same voice.

Installing VODER and running a first task

The repository ships setup.py and requirements.txt, and setup.py is where the optional weight lives. It provisions separate environments for the Eva models, one per capability, so you only download what you intend to use:

python
EVA_ENVS = {
    "flux2":   "Flux 2 Dev / Klein 9B (image gen / edit / nbg / mini)",
    "h3":      "MiniMax H3 (video gen)",
    "vace":    "Wan 2.1 VACE 14B (video edit)",
    "animate": "Wan 2.2 Animate 14B + S2V 14B (video animify + lipsync)",
}

The full list also covers hyworld for world generation, trellis for image to 3D, sam3 for segmentation and siglip2 as a vision encoder. Each is a separate install decision rather than one monolithic download.

Eva modes are reached through the main entry point:

bash
python voder.py eva <mode>

with the modes given as tti, ttv, ttt and ttw. Utility tasks live outside the engine as side-quests, covering URL download, format conversion, cutting, merging, mixing, silence stripping and effects such as speed, pitch, reverb and loudness normalisation. To see them all grouped by category:

bash
python voder.py quest

Chains are the piece that turns the modes into workflows: each chain is named, its output is captured to a temporary location, and later chains can refer to earlier chain names as input paths, so building a song, isolating its vocals, training a voice from them and dubbing a video can be one command. There is also a Colab notebook linked from the README for anyone without suitable hardware.

What local execution costs you

Running everything locally is the feature and the tax. There is no queue and no per-minute billing, but every capability arrives as a model you download and store, and the Eva side of the project multiplies that: Flux 2 Dev for images, MiniMax H3 for video, Wan models for editing and animation, HY-World 2.0 for 3D scenes and TRELLIS.2 for image to 3D, each with its own environment.

Hardware decides how usable it is. The README says it works with or without a GPU, but it does not claim CPU execution is fast, and the voice conversion path switches to a 44.1kHz model for music, which is a signal that quality settings change the compute profile. Video generation and 3D world generation on CPU are a different proposition from speech enhancement.

Maturity is uneven across the surface. The release dated 2026-08-16 is marked stable with all features working, and added MiniMax Music 3 as an extreme text-to-music mode capable of songs up to five minutes. The release dated 2026-08-20, which introduced Project Eva and the DLC system, is marked in development. Treat the audio core and the image and video expansion as two different risk levels.

Hosted voice APIs sit on the other side

The obvious alternative is a commercial voice API such as ElevenLabs, which offers synthesis and voice cloning as a metered service. The difference is not quality, it is where the work and the data sit.

A hosted API gives you a stable endpoint, no model management, consistent latency and an uptime expectation, paid for by usage and by handing your audio and any reference voice to someone else. Cloning a specific person's voice on those platforms is gated by their consent and verification rules, which is a real constraint if your use case involves a voice you do not own.

VODER gives you the opposite: everything stays on the machine, there is no per-character cost, and cloning is limited only by what the models can do and by your own judgement. You take on model downloads, dependency maintenance and hardware, and you get no service guarantee. If your volumes are low and the content is not sensitive, the API is less work. If the content cannot leave, or the volume makes per-minute pricing absurd, local is the only answer.

Licence, releases and upkeep

VODER is AGPL-3.0. That is the strongest copyleft of the common open source licences, and the clause that matters here is the network one: obligations that GPL attaches to distribution also apply when users interact with the software over a network. Wrapping VODER in a web service and letting others use it means offering them the source of your version, modifications included. Internal use with no external users triggers nothing. This is not legal advice, but it makes AGPL projects a deliberate choice for anything company-facing.

Release activity is recent and frequent, with tags dated by day rather than by semantic version: 2026-08-16 for the MiniMax Music 3 addition, published on 2026-08-21, and 2026-08-20 for Project Eva plus the Klarify DLC, published on 2026-08-21. The last push was on 2026-08-23. Dated tags are good for ordering and awkward for dependency management, so expect to track commits rather than pin a version number.

Upkeep concentrates in dependencies. The requirements file pins torch 2.8.0, torchaudio 2.8.0, torchvision 0.23.0, transformers 4.57.3 and huggingface-hub 0.34.0, and pulls in pyqt5 for the GUI, pyannote for diarization, modelscope and funasr, plus ollama for the VADAR chat mode. Any one of those can block a Python or CUDA upgrade, which is the standing cost of a tool that integrates this many models.

Editorial conclusion

VODER suits people who do a lot of voice work and want it on their own hardware with no subscription and no upload: podcasters assembling multi-speaker dialogue, translators dubbing videos, and anyone who needs cloning, separation and music generation reachable from one command. It is a poor fit if you want a narrow tool that does one thing well, since the dependency set is large and the value is in the breadth, or if you need the Eva image and video features today, because that release is marked in development. Before committing, check what the installer wants to download: setup.py provisions separate environments per Eva model, and the README notes the whole thing runs with or without a GPU, so decide which modes you need before you install all of them.

Frequently asked questions

Does VODER need a GPU?

The README says it runs entirely on your machine and works with or without a GPU. It does not state relative speeds, and heavier modes such as video generation have their own model requirements.

How many languages can VODER translate?

Any-to-any translation covers 76 languages through TranslateGemma 12B, invoked with a translate (source-target) syntax that is decoupled from the speech recognition engine.

Can VODER work with video files?

Yes. Voice conversion accepts an MP4 and returns video with the converted voice, dubbing translates and retimes whole videos while preserving the speaker's voice identity and background music, and speech-to-text accepts video files and URLs.

What is Project Eva in VODER?

Project Eva is the first DLC expansion, adding text-to-image, text-to-video, text-to-world 3D generation and a local chat mode called VADAR. It lives in the DLC folder and is reached with voder.py eva using the tti, ttv, ttt or ttw modes.

Official sources

  1. HAKORADev/VODER on GitHub
  2. Issues
  3. License: AGPL-3.0
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/hakoradev-voder.svg)](https://hysenlabs.com/projects/hakoradev-voder)