# VODER runs eight voice and vision modes locally, with ten model environments under src/envs

> An AGPL-3.0 Python toolbox that bundles speech to text, text to speech, voice conversion, music generation, separation, enhancement, sound effects and diarization behind one script, plus a Project Eva expansion for image, video, 3D and chat generation. Everything runs on your own machine, and the weight sits in a dependency file that pins nine packages exactly.

**HAKORADev/VODER** — Voice Operation and Design Engine with Reproduction capabilities

- Repository: https://github.com/HAKORADev/VODER
- Stars: 519 · Forks: 92
- Language: Python
- License: AGPL-3.0
- Published: 2026-09-18 · Updated: 2026-09-18 · Language: en
- Canonical page: https://hysenlabs.com/projects/hakoradev-voder

## Eight modes, three task layers, one script

The design is a flat list of processing modes under a single entry point, `voder.py`. The eight modes are speech-to-text, text-to-speech, voice conversion, music generation, speech enhancement, sound effects, vocal separation and speaker diarization. Language dubbing is reached through `tts dub`, and any-to-any translation runs through TranslateGemma 12B with a `translate (source-target)` syntax that is decoupled from the speech recognition engine. Three task-layer features sit above the modes rather than beside them. `train` saves a reusable voice clone to a `.tts` or `.ttse` file. `quest` covers side tasks that have nothing to do with the engine, such as URL download, format conversion, cutting, merging, mixing, silence stripping and effects like reverb and loudnorm, and `python voder.py quest` lists them grouped by category. Chains are the third, and they are user-defined pipelines that wire any number of tasks together. The Project Eva expansion adds image, video, 3D world and chat generation through `python voder.py eva <mode>` using the same infrastructure and the same command patterns.

## setup.py ships ten model environments under src/envs

The installer is where the model inventory is visible, because it keeps a dictionary of environment names against the models they hold. `ENVS_DIR` points at `src/envs`, and `EVA_ENVS` maps ten keys to their payloads: `flux2` for Flux 2 Dev with Klein 9B, `qwen-image2.1-uc` for Qwen-Image-2.1 UC, `h3` for MiniMax H3, `vace` for Wan 2.1 VACE 14B, `animate` for Wan 2.2 Animate 14B with S2V 14B, `hyworld` for Tencent HY-World 2.0, `trellis` for Microsoft TRELLIS.2, `lyra2` for NVIDIA Lyra 2.0, `sam3` for Meta SAM 3.1 segmentation and `siglip2` for a SigLIP 2 giant vision encoder. Each of those is a multi-gigabyte download waiting to happen, and the fact that segmentation and a vision encoder are included even though no mode is described as segmentation says the environments are shared infrastructure rather than per-feature assets. The rest of the file is a thin platform layer: `is_linux`, `is_windows` and `is_macos` wrappers, a `command_exists` check built on `shutil.which`, and a `have_sudo` helper that only returns true on Linux.

## Chains capture output to temp and later chains read it by name

Chains are the feature that turns the tool from a menu into a pipeline. Each chain is given a name, its output is captured to a temporary location, and a later chain can reference an earlier chain by name as its input path. The documented example is the one worth copying: build a song, isolate its vocals from it, train a voice from those vocals, then dub a video with the result, all in a single command. That ordering matters because each step consumes the artefact of the previous one, which is something a menu-driven tool cannot express. The trade-off is that chains depend on the temp directory surviving between steps, so a chain interrupted part way leaves intermediate files behind. Dialogue generation works the same way as a script: multiple characters with distinct voices, per-line control over timing, level and duration through script directives written `/time`, `/level` and `/duration`, sound effects embedded inline with the `sfx:` marker and generated from a text description, and an automatic background track mixed to match the spoken length.

## The input pipeline checks the URL before it downloads anything

Pasting a link is the fastest path in, and the supported sources are named: YouTube, TikTok, Bilibili, Snapchat, Instagram, Facebook and X or Twitter. What makes this more than a convenience is the verification step, since VODER checks that the link actually resolves to a video before it downloads anything. That check has a history. The release tagged 2026-07-30 is a bugfix for `is_supported_url()` treating local file paths as URLs, which is exactly the kind of bug that appears when a URL classifier also has to handle paths on different operating systems. Two other input paths sit alongside it: an image goes in and VODER extracts text with OCR, and multi-speaker audio is automatically split into voice clips so one of them can be turned into a clone without leaving the flow. Downloads go through `yt-dlp` and `gallery-dl`, which is why those two packages are hard requirements rather than optional extras.

## Voice conversion has a mimic mode and a 44.1kHz music model

Conversion comes in a speech flavour and a music flavour, and they are not the same path. The speech path transforms one voice into another while preserving the original words, emotion and timing, and it accepts and returns video, so an MP4 in produces a video with the converted voice. A `mimic` mode goes further and transfers accent and speaking style along with timbre, which is the setting to try when the source and target speaker share a language. For music, the tool switches to a high-fidelity 44.1kHz model instead. The re-synthesis route is separate again: `tts svc` transcribes speech and re-reads it in a different voice, and an optional `sts:` prefix routes to Seed-VC v2 for what is described as high-fidelity conversion. Speaker separation sits on the other side of the same pipeline, pulling individual speakers out of a multi-speaker recording into separate files, each paired with a speaker-labelled transcript that can then be reused for cloning or conversion.

## Nine exact pins and two user interface frameworks

The requirements file is where the real cost of this project shows. It pins `torch` at 2.8.0, `torchaudio` at 2.8.0 and `torchvision` at 0.23.0, `transformers` at 4.57.3, `huggingface-hub` at 0.34.0, `hydra-core` at 1.3.2, `lightning` at 2.4, `x-transformers` at 2.3.1 and `torchlibrosa` at 0.0.9, which means an existing environment with different torch versions will conflict rather than resolve. Two user interface frameworks are listed, `pyqt5` for the desktop side and `gradio` for the web side, so the tool carries both. Beyond that, the file is organised by which mode needs what: `openai-whisper` for transcription, `diffusers` for the generative models, `librosa` and `descript-audiotools` for audio analysis, the `pyannote` family plus `asteroid-filterbanks` for diarization, `funasr` and `modelscope` for multilingual speech, `easyocr` for image text, `langdetect` for picking the source language automatically, and `ollama` for the local chat model behind the Eva text-to-text mode. OpenTelemetry protocol exporters are present as well, so trace export is wired in even though the documentation excerpt does not describe it.

## Tags are dates, and the newest one is 2026-08-20

Version tracking works by dating the tag, which is unusual enough to be worth knowing before you look for a changelog. The three most recent tags are `2026-08-20`, covering the Project Eva expansion covering image, video, world and chat generation plus a separate Klarify expansion; `2026-08-16`, described as an extreme TTM covering a MiniMax Music 3 integration; and `2026-07-30`, the `is_supported_url()` bugfix described above. All three were published on 2026-08-21, so the dates in the tag names are the date the work happened rather than the date it was cut. The last push to `main` is 2026-10-01, and the repository is not archived. The tree itself is minimal for something this large: `.gitignore`, `LICENSE`, `README.md`, `requirements.txt`, `setup.py`, `src/` and `docs/`. There is no test directory, no CI configuration and no issue templates at the top level, which tells you where to look for verification before you rely on it. The licence is AGPL-3.0, which is the term to read before this sits behind anything a paying customer touches.

## Conclusion

VODER is for someone who wants speech and music models running locally and is willing to accept a large dependency surface in exchange for one interface instead of eight repositories. The chains feature and the reusable voice clones are the two things that make it more than a launcher, since they let you script a song, isolate its vocals, train a voice and dub a video as a single command. It is the wrong shape for a laptop with a small disk or for anyone who needs predictable install times, because nine exact pins in the requirements file tie you to specific torch and transformers versions, and the Eva expansion adds ten more model environments on top. Before you commit, read the requirements file against your existing Python environment, confirm the disk and GPU headroom for the models you actually intend to use rather than all of them, and check the AGPL-3.0 terms against your own deployment plans before putting it behind a service.

## FAQ

### What can VODER do without a GPU or a subscription?

It runs entirely on your own machine, needs no subscription, and works with or without a GPU. Eight processing modes are bundled under one interface: speech-to-text, text-to-speech, voice conversion, music generation, speech enhancement, sound effects, vocal separation and speaker diarization, with dubbing, translation and re-synthesis on top.

### What are chains and quests in VODER?

Chains are user-defined pipelines that wire any number of tasks together. Each chain is named, its output is captured to a temp location, and a later chain can reference an earlier chain name as its input path. Quests are lightweight utilities outside the main engine, such as URL download, format conversion, cutting, merging and effects, listed with python voder.py quest.

### Which models does the VODER Eva expansion download?

setup.py keeps a dictionary of ten environments under src/envs, including Flux 2 Dev, Qwen-Image-2.1 UC, MiniMax H3, Wan 2.1 VACE 14B, Wan 2.2 Animate 14B, Tencent HY-World 2.0, Microsoft TRELLIS.2, NVIDIA Lyra 2.0, Meta SAM 3.1 and a SigLIP 2 giant vision encoder. They are reached with python voder.py eva followed by the mode.

### How do I check which version of VODER I have?

Releases are tagged by the date the work happened rather than by semver: the newest tags are 2026-08-20, 2026-08-16 and 2026-07-30, all published on 2026-08-21. The repository is not archived and the last push to main is dated 2026-10-01, so commits continue past the newest tag.

## Sources

- [HAKORADev/VODER on GitHub](https://github.com/HAKORADev/VODER)
- [Issues](https://github.com/HAKORADev/VODER/issues)
- [License: AGPL-3.0](https://github.com/HAKORADev/VODER/blob/main/LICENSE)
- [README](https://github.com/HAKORADev/VODER/blob/main/README.md)
- [Releases](https://github.com/HAKORADev/VODER/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/hakoradev-voder
