AuK: Tencent Hunyuan's 1.5B Speech Generation and Editing Model
AuK: An Open-Source Foundational Model for Speech Generation and Editing
At a glance
- What is it?
- AuK is a 1.5B foundation model from Tencent Hunyuan that handles zero-shot TTS, speech editing, enhancement and separation behind one natural-language instruction interface. This article covers how it installs, what the instruction model actually does, and where it is the wrong tool.
- Who is it for?
- Adopt AuK if you have a CUDA or CPU PyTorch 2.7 environment, need several speech tasks behind one instruction interface, and can accept a 1.5B checkpoint plus a separately downloaded SenseVoice fallback for ASR. Skip it if you need a hosted API with an SLA, or if your pipeline is already built around a single-task model you do not want to re-prompt.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository last received commits 5 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What AuK Does That a Single-Task TTS Model Cannot
Most open speech models do one thing. A zero-shot TTS checkpoint clones a voice. A separate model denoises. Another separates speakers. AuK's stated goal is to collapse that stack into one 1.5B model that reads a natural-language instruction and decides which operation to perform. The README lists the tasks it covers: zero-shot TTS, instruct TTS, speech content editing, lyric editing, pitch, speed and volume editing, emotion, timbre, de-accent, nonverbal editing, whisper conversion, speech enhancement, speech separation and music separation. That breadth is the product. The audience is engineers building voice pipelines who would rather maintain one checkpoint and one prompt format than five models with five preprocessing conventions. The trade-off is obvious and the README does not hide it: a 1.5B generalist is unlikely to match a specialist model on any single benchmark, and the repository's performance figure is an image, not a table you can read in text form.
One Instruction Interface Over Many Speech Tasks
The mechanism the README describes is a unified natural-language instruction interface. Every supported task is expressed as an instruction, and the model produces the edited or generated audio. The repository points to docs/COOKBOOK.md for instruction templates plus CLI and Python examples for each task, which is where the actual prompt strings live. Two variants ship: AuK, described as the base model for high-quality generation, and AuK-Flash, described as a distilled model for fast 4-step inference. That distinction matters operationally. If you are generating long audio in a loop, the 4-step variant changes your latency budget in a way the base model does not. The architecture diagram is in assets/arch.png and the README does not restate it in prose, so anyone evaluating the design should open that file rather than expect a textual description. On the data side, the README states the model was trained on millions of hours of diverse audio, a claim that cannot be checked from the repository alone.
Installing AuK with uv or Conda and Running a First Command
The README documents two installation paths, uv and Conda, under Quick Start. The package is named auk and pyproject.toml requires Python 3.10 or newer. The PyTorch family is pinned to the 2.7 line: torch>=2.7,<2.8, torchaudio>=2.7,<2.8 and torchvision>=0.22,<0.23. The comment in pyproject.toml notes that users may preinstall a platform-specific CPU or CUDA build before installing AuK, which is the normal way to avoid pulling a default wheel that does not match your GPU. Transformers is pinned to >=4.52.0,<5, with a note that Qwen2.5-Omni support landed in 4.52 and the release was validated on 4.57.
The editable install is the documented form, and the optional extras are separated by interface. The gradio extra pulls the Gradio UI, the Prompt Enhancer, cloud ASR and a local SenseVoice fallback:
pip install -e ".[gradio]"A separate comfyui extra exists for the ComfyUI nodes, and the repository has a top-level comfyui/ directory. After installation, weights come from Hugging Face or ModelScope: tencent/AuK and tencent/AuK-Flash on Hugging Face, with matching ModelScope entries under Tencent-Hunyuan. The README's Quick Start lists Download the weights, Command-line inference, Interactive Gradio demo, ComfyUI, Prompt Enhancer and Python API as the entry points, and docs/COOKBOOK.md carries the per-task instruction templates you need before any CLI call will do something useful.
The .env.example shows the optional LLM and ASR configuration. Copy it and fill in one OpenAI-compatible provider:
cp .env.example .envLLM_API_KEY=your-llm-api-key
LLM_BASE_URL=https://tokenhub.tencentmaas.com/v1
LLM_MODEL_NAME=hy3The file also shows DeepSeek V4 Flash as an alternative provider, with LLM_BASE_URL=https://api.deepseek.com and LLM_MODEL_NAME=deepseek-v4-flash. For ASR, TENCENTCLOUD_SECRET_ID, TENCENTCLOUD_SECRET_KEY and ASR_ENGINE_MODEL_TYPE=16k_zh_en are optional: the comment states that without these credentials, or if the cloud request fails, AuK lazily downloads and uses SenseVoiceSmall on CPU. That lazy download is a real first-run cost on a machine without network access to the model host.
Where AuK Fails or Is the Wrong Choice
The dependency pins are the first constraint. torch>=2.7,<2.8 is narrow. If your environment is locked to an older PyTorch because of another model in the same process, you cannot install AuK alongside it without resolving that conflict. The same applies to transformers>=4.52.0,<5. The pyproject.toml comment about numpy is worth reading: numpy is no longer imported directly and comes in transitively, with a commented-out pin for 1.x on Python 3.10. If your stack needs numpy 1.x locked, you are re-enabling a pin the maintainers left disabled.
The second failure mode is task ambiguity. A single instruction interface means the model must infer intent from text. The README does not describe a schema, an enum or a validation layer for instructions, so malformed or ambiguous prompts are a prompt-engineering problem rather than a caught error. For production systems that need deterministic routing, one model per task with a typed API is easier to test.
The third is the ASR fallback. Cloud ASR is optional and SenseVoiceSmall runs on CPU when credentials are missing or the request fails. On a CPU-only host, that fallback competes with the model itself for the same cores. The README does not state expected transcription latency, so treat it as an unknown until you measure it.
Finally, the licence metadata is inconsistent. pyproject.toml declares license = { text = "MIT License" } and carries the MIT classifier, but the repository's licence field reads NOASSERTION and there is a LICENSE file at the top level. The README's own License section is the place to check, and the two sources should be reconciled before you depend on the terms.
AuK Compared with a Specialist Speech Stack
The realistic alternative is not another foundation model. It is assembling specialists: a zero-shot TTS model for cloning, a separate separation model such as a dedicated source-separation network, and a denoiser for enhancement. That stack wins on predictability. Each component has one input contract, one output contract and benchmarks you can compare across versions. It loses on integration: you own the glue, the resampling, the chunking and the failure handling between stages. AuK's bet is that one instruction interface plus one checkpoint is worth more than per-task accuracy. Whether that bet pays off depends on your workload. If you run one task at volume, a specialist is the safer call. If you run many tasks at low volume, or you want to prototype editing operations without standing up five services, the unified interface removes real work. The README does not publish a task-by-task comparison against named alternatives, so the performance figure in assets/performance.png is the only evidence offered and it is not machine-readable in the repository text.
Maintenance, Upgrades and Licence Questions
The repository is not archived and the last push was on 2026-09-10, which is recent. There are no retrieved releases, so versioning currently runs off the main branch and the version string in pyproject.toml, which is 0.1.0. That means upgrades are git pulls rather than tagged installs, and you should pin a commit for anything you deploy. The SGLang-Omni cookbook linked in the News section is a second serving path if you want to run AuK through that stack rather than the bundled CLI or Gradio app.
Upgrade cost concentrates in three places: the torch 2.7 window, the transformers 4.x window and the model weights. Any bump to those pins is a potential breaking change for code that imports the model directly through the Python API. The optional extras are independent, so a change to the gradio extra does not touch a comfyui-only install.
On licensing, the repository metadata says NOASSERTION while pyproject.toml says MIT and includes the MIT classifier. The LICENSE file at the top level is the document that governs, and the README has a License section. Read both before shipping, and note that model weights are distributed separately on Hugging Face and ModelScope, where terms can differ from the code licence. This is not legal advice; it is a list of the files you need to read.
Editorial conclusion
Adopt AuK if you have a CUDA or CPU PyTorch 2.7 environment, need several speech tasks behind one instruction interface, and can accept a 1.5B checkpoint plus a separately downloaded SenseVoice fallback for ASR. Skip it if you need a hosted API with an SLA, or if your pipeline is already built around a single-task model you do not want to re-prompt. Before committing, verify that the Python version and torch range in pyproject.toml match your environment, that the weights download completes from Hugging Face or ModelScope, and that the LICENSE file contents match the MIT text declared in pyproject.toml, since the repository metadata is marked NOASSERTION.
Frequently asked questions
How do I install AuK?
The README documents two paths, uv and Conda, both under Quick Start. The package is auk, requires Python 3.10 or newer, and the editable install with the Gradio extra is pip install -e ".[gradio]"; a comfyui extra exists for the ComfyUI nodes.
How do I download the AuK model weights?
Weights are hosted on Hugging Face as tencent/AuK and tencent/AuK-Flash, with matching ModelScope entries under Tencent-Hunyuan. The README lists Download the weights as a separate step before command-line inference.
What is the difference between AuK and AuK-Flash?
The README describes AuK as the base model for high-quality generation and AuK-Flash as a distilled model for fast 4-step inference. They are published as separate weight repositories on Hugging Face and ModelScope.
Does AuK need an API key to run?
The .env.example file configures an optional OpenAI-compatible LLM provider for the Prompt Enhancer and optional Tencent Cloud ASR credentials. Without the ASR credentials, or if the cloud request fails, the file states that AuK lazily downloads and uses SenseVoiceSmall on CPU instead.
What Python and PyTorch versions does AuK require?
pyproject.toml requires Python 3.10 or newer and pins torch>=2.7,<2.8, torchaudio>=2.7,<2.8 and torchvision>=0.22,<0.23, with transformers>=4.52.0,<5. A comment notes users may preinstall a platform-specific CPU or CUDA build first.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/tencent-hunyuan-auk)