# LuxTTS: A ZipVoice-Derived Voice Cloning Model That Fits in 1 GB of VRAM

> LuxTTS is an Apache-2.0 Python text-to-speech model built on ZipVoice, distilled to four sampling steps with a custom 48 kHz vocoder. It is aimed at people who want local voice cloning without a large GPU, and the README is explicit about where the quality and speed trade-offs sit.

**ysharma3501/LuxTTS** — A high-quality rapid TTS voice cloning model that reaches speeds of 150x realtime.

- Repository: https://github.com/ysharma3501/LuxTTS
- Stars: 5,413 · Forks: 685
- Language: Python
- License: Apache-2.0
- Published: 2026-09-22 · Updated: 2026-09-22 · Language: en
- Canonical page: https://hysenlabs.com/projects/ysharma3501-luxtts

## What LuxTTS solves, and who it is actually for

Most voice cloning models that produce convincing output assume you have a datacenter GPU. LuxTTS takes the opposite position. The README describes it as a "lightweight zipvoice based text-to-speech model" that fits within 1 GB of VRAM, which puts it in range of a modest local card, and it also runs on CPU and on Apple MPS. The stated target is 150x realtime on a single GPU and faster than realtime on CPUs.

The second selling point is output resolution. The README notes that most TTS models are limited to 24 kHz and positions the custom vocoder at 48 kHz as the differentiator. That matters if you are feeding generated speech into a video or podcast pipeline where resampling artifacts are audible.

The audience is therefore narrow and specific: Python developers who want to clone a voice from a short reference clip on their own hardware, without sending audio to a third party. If you need a managed endpoint, this is not that. The repository has no homepage and no releases, so the distribution channel is the git repository plus the Hugging Face model and Space.

## ZipVoice distillation, the four-step sampler and the 48 kHz vocoder

The README's own FAQ answers the architecture question directly: LuxTTS uses the same architecture as ZipVoice but is distilled to 4 steps with an improved sampling technique, and swaps the default 24 kHz vocoder for a custom 48 kHz one. That is the whole design story, and it explains the parameter surface.

The data flow at inference is two-stage. First, encode_prompt takes a reference audio file and turns it into an encoded prompt, with an rms parameter that the README recommends setting around 0.01 and describes as controlling loudness. The encoding step is where librosa is first loaded, and the README warns it takes about 10 seconds to initialize the first time. Second, generate_speech takes the text and that encoded prompt and produces a waveform, with num_steps controlling the sampler. The README recommends 3 to 4 steps as the efficiency sweet spot, which is consistent with the distillation claim.

Two more sampler parameters are exposed. t_shift is described as trading quality against word error rate: higher can sound better but raises WER, lower reduces pronunciation errors at the cost of quality. speed scales the tempo of the generated audio. A return_smooth flag exists specifically for metallic-sounding output. The repository layout is small: a single zipvoice/ package directory alongside requirements.txt, pyproject.toml and LICENSE. Note that pyproject.toml still carries the name "Zipvoice" and version 0.0.11, and its description field talks about LinaCodec rather than LuxTTS. That is inherited packaging metadata, not a LuxTTS release, and it is worth knowing before you read anything into the version number.

## Installing LuxTTS and cloning a voice from a reference clip

The README gives a three-command install from source. Clone the repository, enter it, and install the requirements file.

```bash
git clone https://github.com/ysharma3501/LuxTTS.git
cd LuxTTS
pip install -r requirements.txt
```

That requirements file is not trivial. It pins transformers<=4.57.6 and setuptools<81, pulls piper_phonemize from a custom find-links index at k2-fsa.github.io, and installs LinaCodec straight from a git URL. Expect the install to take a while and to need network access to more than PyPI. The pyproject.toml also declares requires-python >=3.10, so a 3.9 environment will not work.

With the install done, load the model. The README shows three device variants, and the CPU one exposes a threads argument.

```python
from zipvoice.luxvoice import LuxTTS

# load model on GPU
lux_tts = LuxTTS('YatharthS/LuxTTS', device='cuda')

# load model on CPU
# lux_tts = LuxTTS('YatharthS/LuxTTS', device='cpu', threads=2)

# load model on MPS for macs
# lux_tts = LuxTTS('YatharthS/LuxTTS', device='mps')
```

The model identifier is a Hugging Face repo, so weights download on first use. The README's tips say to use a reference clip of at least three seconds.

A first real generation looks like this. The prompt_audio path accepts wav or mp3, and the output is written at 48000 Hz, which matches the 48 kHz vocoder claim.

```python
import soundfile as sf

text = "Hey, what's up? I'm feeling really great if you ask me honestly!"
prompt_audio = 'audio_file.wav'

encoded_prompt = lux_tts.encode_prompt(prompt_audio, rms=0.01)
final_wav = lux_tts.generate_speech(text, encoded_prompt, num_steps=4)

final_wav = final_wav.numpy().squeeze()
sf.write('output.wav', final_wav, 48000)
```

If the result sounds metallic, the README's tip is return_smooth=True. If inference is slower than you want, the README suggests lowering ref_duration, with the caveat that setting it to 1000 avoids artifacts when a short reference causes them.

## Where the quality and speed knobs stop being free

The sampling parameters are not independent, and the README says so in a way that is easy to skim past. t_shift is a direct trade between how good the audio sounds and how likely the model is to mispronounce words. There is no setting that improves both. If you are generating narration for a product demo, you will probably accept a higher t_shift. If you are generating text where a mispronounced name is a defect, you will lower it and accept flatter audio.

num_steps has a similar shape. The README recommends 3 to 4 for efficiency, which implies that going higher buys quality at a linear cost in generation time. The 150x realtime figure is stated for the model as shipped, and the README's own FAQ says it currently uses float32 and that float16 "should be significantly faster (almost 2x)". Float16 inference is on the roadmap as unreleased work, so the speed you get today is the float32 speed.

There is also a hard input constraint that the README states as a tip rather than a validation rule: use a reference clip of at least three seconds. Nothing in the README indicates the library rejects shorter clips. It appears to be on you to check.

The most significant limitation is structural. The README does not document an API stability policy, a versioning scheme for the Python interface, or a rollback path. The roadmap lists "Release LuxTTS v1.5" as unchecked, and pyproject.toml still identifies the package as Zipvoice 0.0.11. If you build a service on top of this, pin the git commit you installed from, because there is no release artifact to pin to.

## LuxTTS versus ZipVoice and the hosted TTS options

The clearest comparison is with ZipVoice itself, and the README supplies it. Same architecture, different post-training: LuxTTS is distilled to 4 steps with a modified sampling technique and a 48 kHz vocoder instead of the default 24 kHz one. If you already run ZipVoice and your complaint is inference cost or output sample rate, LuxTTS is the direct answer. If your complaint is something else, the shared architecture means you will inherit it.

Against hosted services, the difference is not quality but control and cost structure. A hosted TTS API gives you an endpoint, rate limits and someone else's uptime. LuxTTS gives you weights that run locally under 1 GB of VRAM, which means no per-character billing and no audio leaving your machine. The cost you take on instead is the install: a pinned transformers version, a custom find-links index for piper_phonemize, and a git dependency on LinaCodec. That is a real maintenance burden, and it is the price of the local model.

The community ecosystem is worth noting because it changes what you have to build. The README lists a Gradio app, a ComfyUI node set, an ONNX port, and a hosted demo on fal.ai. If you want a UI, the Gradio project exists. If you want to run without PyTorch, the ONNX port is the relevant path. LuxTTS itself ships neither a UI nor an ONNX export.

## Licence, dependencies and what upgrades will cost you

The model and code are Apache-2.0, and the README points to the LICENSE file. That is a permissive licence, which generally means commercial use is permitted, but the repository does not state a position on voice cloning consent, and the licence text governs the software, not the voices you clone. Whether you have the right to clone a particular speaker is a separate question that the project does not answer. That is not legal advice, and if you plan to ship generated audio commercially, the consent question is the one to resolve before the licensing question.

Upgrade cost is dominated by the dependency pins. requirements.txt pins transformers<=4.57.6 and setuptools<81, and installs LinaCodec from a git URL rather than a released package. A git dependency has no version number, so an upgrade of LuxTTS can silently pull different LinaCodec code. The piper_phonemize wheel comes from a third-party find-links index, which is another external point of failure.

There are no releases in the repository, so "upgrading" means pulling master and reinstalling requirements. The roadmap's unchecked items (LuxTTS v1.5, float16 inference code) suggest the interface may change. If you deploy this, record the commit hash and the resolved dependency versions alongside your code.

## Conclusion

Adopt LuxTTS if you want a local, Apache-2.0 voice cloning model that runs in under 1 GB of VRAM and you are willing to tune rms, t_shift and num_steps yourself. Do not adopt it if you need a hosted API with an SLA, a stable public interface, or a model whose sampling defaults are documented beyond the README. Before committing, verify that your reference clip is at least three seconds long, that your torch install matches your CUDA or MPS backend, and that the git-sourced LinaCodec dependency resolves in your environment.

## FAQ

### How do I install LuxTTS?

Clone the repository, change into it, and run pip install -r requirements.txt. The requirements file pulls piper_phonemize from a custom find-links index and installs LinaCodec from a git URL, so the install needs access to more than PyPI.

### How much VRAM does LuxTTS need?

The README states it fits within 1 GB of VRAM, which is the basis for the claim that it can run on any local GPU. It also documents CPU and Apple MPS device options.

### How is LuxTTS different from ZipVoice?

According to the README's FAQ, LuxTTS uses the same architecture as ZipVoice but is distilled to 4 steps with an improved sampling technique, and replaces the default 24 kHz vocoder with a custom 48 kHz one.

### What is the most realistic TTS voice?

The README does not compare LuxTTS against other TTS systems on realism, and its own claim is that its voice cloning is on par with models 10x larger. That is the project's stated position, not an independent measurement.

### Which AI is best for voice cloning?

The repository does not benchmark LuxTTS against other cloning systems, so it offers no basis for ranking them. What it does state is that LuxTTS clones from a reference clip and recommends using at least a 3 second audio file.

## Sources

- [Issues](https://github.com/ysharma3501/LuxTTS/issues)
- [License: Apache-2.0](https://github.com/ysharma3501/LuxTTS/blob/master/LICENSE)
- [README](https://github.com/ysharma3501/LuxTTS/blob/master/README.md)
- [ysharma3501/LuxTTS on GitHub](https://github.com/ysharma3501/LuxTTS)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/ysharma3501-luxtts
