# franken_tts: A pure-Rust CPU runtime that turns Qwen3-TTS's hidden microdecoder into its best asset

> franken_tts reimplements Qwen3-TTS-12Hz-0.6B-Base entirely in safe Rust, targeting the model's 15-step residual-code microdecoder as the main optimization lever. It runs faster than real time on a Mac Mini M4 Pro with no GPU, no Python, and no ML framework, but its most aggressive optimizations are still unimplemented.

**Dicklesworthstone/franken_tts** — Pure-Rust, CPU-hyper-optimized runtime for Qwen3-TTS zero-shot voice cloning — turns the model's hidden 15-step residual-code microdecoder from its largest CPU liability into its largest optimization advantage

- Repository: https://github.com/Dicklesworthstone/franken_tts
- Stars: 44 · Forks: 9
- Language: Rust
- License: NOASSERTION
- Published: 2026-08-23 · Updated: 2026-08-23 · Language: en
- Canonical page: https://hysenlabs.com/projects/dicklesworthstone-franken-tts

## The problem: a 0.6B model that hides a 15-step CPU killer

Running Qwen3-TTS for voice cloning normally means Python, PyTorch, and a GPU, even for a 0.6B-parameter model. The public story of a 12.5 Hz model hides the real cost: every 80 ms frame contains a 15-step autoregressive residual-code microdecoder. The 28-layer talker runs once, then a 5-layer code predictor runs fifteen sequential times, each with per-depth embeddings and 2,048-way heads, followed by a causal codec. The README quantifies the traffic: first-order Q8 weight movement is about 1.65 GB per frame, and the microdecoder body accounts for roughly 1.18 GB of that, reread fifteen times. That hidden loop is the largest CPU liability in the model, and franken_tts treats it as the primary optimization target. The project is for engineers who want zero-shot voice cloning on commodity hardware, without the usual ML stack, and who are willing to accept a single-model, single-revision commitment.

## How it works: seam-by-seam verification and a speculative draft that verifies itself

franken_tts reimplements the entire pipeline in safe Rust, and the README claims it is verified seam-by-seam against the upstream PyTorch model. The verification is specific: argmax-exact talker parity through all 28 layers, whole-utterance codec codes exact versus the oracle, and an audio envelope check against the pinned PyTorch reference. That is not a fuzzy 'sounds close' claim; it is a bit-level contract. The optimization strategy centers on the microdecoder. Per-depth quantization is the default int8 route today. Two other techniques are described but not yet wired in: a cache-resident hot pack that keeps frequently used weights in cache, and FrankenMTP, which the README describes as 'speculative block drafting that the 5-layer predictor itself verifies exactly in a single causal pass.' The drafter and block-verification primitives exist in-tree, but they are not connected to the generation path. A bit-exact f32 reference engine remains available via one environment variable, which is useful for debugging or for users who distrust quantization.

## Getting it running: three commands and no configuration

Installation is a one-liner. On macOS or Linux, run the curl script or use Homebrew with brew install dicklesworthstone/tap/franken-tts. On Windows, a PowerShell command with -EasyMode is provided. After installation, the workflow is three commands: ftts pull fetches the quantized model, about 2.0 GB, with SHA-256 verification, into ~/.cache/franken_tts/model. Then ftts say "text" out.m4a speaks with the default voice matt. Eighteen built-in voices ship in the binary, selectable with --voice: masculine names like james, leo, robert, and denzel, feminine names like judy, aria, ember, and jodie. Voice cloning is ftts enroll voice_memo.m4a --default to replace the default, or ftts say --voice my_voice.spk to use a specific file. Output format follows the extension: .wav is encoded in pure Rust, while .m4a, .mp3, and .flac require a system encoder such as afconvert, ffmpeg, lame, or flac. If none is found, you get an error naming the missing tools, not a silent format change. The model stops at EOS with a text-proportional frame cap as a backstop; FTTS_MAX_FRAMES sets an exact cap when needed.

## Beyond speech: video and voice cards with a shared format

The CLI is not limited to audio. ftts make-video turns a sentence into a 1080p video with a Ken Burns drift, a live waveform, and the voice name on a pill. Every frame is drawn in memory-safe Rust, including the embedded illustration and rasterized text via fmd-font. Only the final H.264+AAC encode uses a system encoder, under the same contract as .m4a: .mp4 needs ffmpeg, while .y4m renders natively with a .wav beside it and needs no encoder. Voice cards are a more distinctive feature. A card is a PNG that carries the voice itself: the 1,024-float speaker embedding is written at two bits per cell across a 144x144 grid with QR-style finder patterns and interleaved Reed-Solomon error correction. That survives screenshots and messaging-app recompression. A lossless PNG chunk is used first when bytes arrive intact. The format is shared with an iOS app, so a card exported here imports on a phone from Photos, and vice versa. The two encoders are bit-identical by pinned test.

## Deterministic streaming and agent-first output

Two properties matter for production use. First, deterministic streaming: the README states that codec streaming output is bit-identical to offline decoding under every packet schedule. That means you can stream audio in chunks without worrying about packet boundaries changing the result. Second, agent-first design: ftts robot schema emits NDJSON events with stable exit codes. That is a concrete interface for automation, not a vague promise. The CLI is built for scripting, with a single binary and no runtime dependencies. The project positions itself as a sibling of franken_ocr and franken_whisper, with one fixed model revision and model-specific kernels. It explicitly makes no pretense of being a general speech framework. That focus is a strength for reproducibility, but it also means you cannot swap in a different TTS model or even a newer revision of Qwen3-TTS without waiting for the project to update.

## Limitations and the wrong-tool cases

The most honest limitation is in the README itself: the biggest optimization is still ahead. The cache-resident hot pack is designed but unimplemented, and FrankenMTP's drafter and block-verification primitives exist in-tree but are not wired into the generation path. That means the advertised speedup from those techniques is not available yet. Per-depth quantization is the only default optimization. Another limitation is the output format dependency: .m4a, .mp3, .flac, and .mp4 require external encoders. If you are on a minimal system without ffmpeg or lame, you are stuck with .wav and .y4m. The license is NOASSERTION, which is a red flag for legal review. You must inspect the repository's LICENCE file yourself to determine what you can do with the code. The project pins one model revision, so if Qwen releases a better TTS, you cannot simply point franken_tts at it. Finally, voice cloning raises ethical and legal questions about consent; the README says 'any recording you have the right to use,' but enforcement is on you.

## Alternatives and maintenance cost

The obvious alternative is the upstream Qwen3-TTS with PyTorch and CUDA. That approach gives you the full model, flexibility to use other models, and a larger ecosystem, but it requires a GPU for practical speed and a Python environment. Another alternative is a general TTS framework like Coqui TTS or Piper, but those do not offer the same zero-shot voice cloning quality as Qwen3-TTS, and they are not pure Rust. The difference in approach: franken_tts optimizes the microdecoder specifically, whereas a general framework treats the model as a black box and relies on framework-level optimizations. Maintenance cost is a real concern. The project has had recent releases, with v0.1.9 in August 2026 and a fix for model downloads surviving GitHub throttling in v0.1.8. That suggests active maintenance. But the license is NOASSERTION, so you cannot assume you can redistribute or modify it without review. The pinned model revision means you are dependent on the project's release cadence for model updates. The install script from a raw GitHub URL is convenient but also a supply-chain risk; you should inspect it before running.

## Conclusion

Adopt franken_tts if you need zero-shot voice cloning on CPU-only hardware, want a single binary with no Python or GPU, and value deterministic, verifiable output. Do not adopt it if you need a general TTS framework, support for other models, or the full speed of the advertised optimizations, since the cache-resident hot pack and FrankenMTP are not yet wired into generation. Before relying on it, verify that the pinned model revision matches your use case, check the license status (NOASSERTION means you must review the source for terms), and confirm that the system encoder requirements for .m4a, .mp3, .flac, and .mp4 outputs are acceptable. The project's own documentation admits the biggest optimization is still ahead, so treat current performance as a baseline, not the ceiling.

## Sources

- [Official README](https://github.com/Dicklesworthstone/franken_tts#readme)
- [Project repository](https://github.com/Dicklesworthstone/franken_tts)
- [Release notes](https://github.com/Dicklesworthstone/franken_tts/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/dicklesworthstone-franken-tts
