# Step-Audio-EditX: a 3B speech editing model you run yourself

> Step-Audio-EditX edits emotion, speaking style and paralinguistic detail in existing speech, and also does zero-shot TTS. It is a Python project with a Dockerfile, a Gradio app and published weights, but the README leaves deployment details thin.

**stepfun-ai/Step-Audio-EditX** — A powerful 3B-parameter, LLM-based Reinforcement Learning audio edit model excels at editing emotion, speaking style, and paralinguistics, and features robust zero-shot text-to-speech

- Repository: https://github.com/stepfun-ai/Step-Audio-EditX
- Stars: 980 · Forks: 80
- Language: Python
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/stepfun-ai-step-audio-editx

## What Step-Audio-EditX actually changes about a recording

Most speech tooling starts from text and produces audio. Step-Audio-EditX starts from audio you already have and rewrites selected properties of it. The README describes the model as a 3B-parameter, LLM-based reinforcement-learning audio model for expressive and iterative editing, covering emotion, speaking style and paralinguistics, with zero-shot TTS as a second capability.

The tag vocabulary is the interface. The README lists emotion tags including happy, angry, sad, fear, surprised, confusion, empathy, embarrass, excited, depressed, admiration, coldness, disgusted and humour, and speaking style tags such as serious, arrogant, act_coy, older, child, whisper, generous and exaggerated. A January 29, 2026 release note adds paralinguistic tags: exhale, snort, inhale, chuckle, clears throat and giggle. The original paralinguistic set covers breathing, laughter, sigh and several language-specific interjections.

That is a narrow but real job. If you are producing a dubbed track and need the same line delivered cold instead of warm, or you want a breath placed before a sentence, this is the class of tool that does it. If you need to remove filler words, the open-source plan still lists that as unchecked.

## How the editing loop is wired

The repository is organized as a Python package with a tokenizer component and a model component, which matches the two published checkpoint families: Step-Audio-Tokenizer and Step-Audio-EditX. Editing therefore runs through a discrete audio token space rather than a waveform-to-waveform model. The tokenizer converts audio into tokens, the language model rewrites those tokens under the tag instruction, and a vocoder (the stepvocoder directory) turns them back into audio.

That structure explains the iterative claim in the README. Because the model operates on tokens, an edit can be applied to the output of a previous edit, so you can stack an emotion change and then a speaking style change. It also explains why the project ships SFT, DPO and GRPO training code: the behaviour is learned through preference optimization, not hand-coded signal processing.

Two supporting modules sit around the core. whisper_wrapper.py suggests transcription is used in the pipeline, and funasr_detach is bundled, which points to speech recognition and voice activity handling on the input side. The config directory holds the runtime configuration, and quantization holds the Int4 path. The README does not document the internal data flow step by step, so treat the module names as the map rather than a specification.

## Installing Step-Audio-EditX and running a first edit

The project targets Python 3.12 to 3.13 (pyproject.toml declares requires-python >=3.12,<3.14). Dependencies include torch>=2.9.1, transformers>=4.57.3, gradio>=6.3.0, funasr>=1.3.0 and vllm. The Dockerfile is the most complete install recipe in the repository: it starts from nvidia/cuda:12.4.1-cudnn-runtime-ubuntu22.04, installs Python 3.12 from the deadsnakes PPA, installs uv, and runs uv sync against pyproject.toml and uv.lock.

If you build the image yourself, this is the command the Dockerfile ends with, which starts the Gradio interface and expects weights mounted at /model:

```bash
uv run python app.py --model-path /model --model-source local
```

Before that, uv sync must resolve the dependency set. Note that pyproject.toml pins vllm to a specific wheel URL under tool.uv.sources rather than a version range, so a resolution failure on that URL stops the install.

```toml
[tool.uv.sources]
vllm = { url = "https://wheels.vllm.ai/c826c72a9633454679871fcb81fbc31fe03fb150/vllm-0.14.0rc2.dev125%2Bgc826c72a9-cp38-abi3-manylinux_2_31_x86_64.whl" }
```

Weights are published on Hugging Face and ModelScope under stepfun-ai/Step-Audio-EditX and stepfun-ai/Step-Audio-Tokenizer. The README also points to a Hugging Face Space playground and a StepFun Audio Studio page, which lets you evaluate tag behaviour before you commit a GPU. For a first run, the repository ships prompt audio under examples/, including en_happy_prompt.wav, whisper_prompt.wav, fear_zh_female_prompt.wav and paralingustic_prompt.wav. The README does not give a CLI example for a single edit, so the shortest verified path is the Gradio app with one of those files.

## Language coverage and the pronunciation workaround

Zero-shot cloning is documented for Mandarin, English, Sichuanese and Cantonese, with Japanese and Korean added in a November 28, 2025 release. Other languages are listed as planned, with Arabic, French, Russian and Spanish still unchecked. If your content is in a language outside that set, the model is the wrong tool today, and no amount of prompt engineering changes the training data.

Polyphonic pronunciation control is the most concrete feature in the README, and it is a workaround rather than a learned behaviour. You replace ambiguous characters with pinyin. The README gives this example: [我也想过过过儿过过的生活] becomes [我也想guo4guo4guo1儿guo4guo4的生活]. That only helps for languages where you can write the pronunciation, and it puts the burden on you to find the characters that need marking. For Cantonese, Sichuanese, Japanese or Korean output, the README says to prefix the text with a tag such as [Cantonese] or [Japanese].

The tag list is also a hard boundary. Ten paralinguistic features were supported at launch; the January 2026 release added six more. Anything outside the documented tags has no defined behaviour, and the README does not describe what happens when you pass an unsupported tag.

## Where the deployment story is thin

The README is a release log with a feature table, not an operations guide. There is no documented API contract, no versioned endpoint, and no statement about serving concurrency or latency. The homepage field is empty, so the only entry points are the repository, the Hugging Face Space and the StepFun Studio page.

The dependency set is heavy and opinionated. It includes deepspeed, bitsandbytes, llmcompressor, wandb, modelscope and a CUDA toolkit package, which is a training stack as much as an inference stack. onnxruntime-gpu and a vLLM wheel pinned to a development build mean the install is sensitive to your CUDA and Python versions. The Dockerfile's base image is CUDA 12.4.1, while pyproject.toml pulls nvidia-cuda-nvrtc-cu12>=12.8.93, and the README does not reconcile those. Expect to spend time on the environment before you spend time on audio.

Training code coverage is also incomplete. SFT, DPO and GRPO are checked in the open-source plan; PPO is not. If your workflow assumes PPO, it is not there yet.

## Step-Audio-EditX compared with a conventional editing chain

The realistic alternative is not another model of the same kind. It is a conventional chain: a TTS engine for synthesis plus a DAW or a DSP library for pitch, time and level changes. That chain gives you deterministic, inspectable transformations. Changing tempo or adding a pause is arithmetic, and the result is reproducible byte for byte.

Step-Audio-EditX trades that determinism for semantic control. Asking for a cold or empathetic delivery is not something a pitch shift can express, and stacking an emotion edit on top of a style edit is the specific capability the token-based design enables. The cost is that the output is generated, so the same input can produce different audio, and the model can alter parts of the recording you did not intend to touch. The README does not describe an ablation or a way to constrain the edit to a region.

Within the StepFun family, the related searches point at Step-Audio 2, Step-Audio-R1 and Step-Audio-TTS, but the README does not describe those projects, so the only comparison that can be made from this repository is against the conventional chain above.

## Licence and maintenance cost

The repository is Apache-2.0, which is permissive and generally allows commercial use, modification and redistribution provided you keep the licence and notice files. That is the licence on the code. The README links model weights hosted on Hugging Face and ModelScope, and the terms attached to those weights are a separate question from the code licence; the README does not state them. Check the model card before shipping anything commercial. This is a description of what the licence file says, not legal advice.

The last push to the repository was on 2026-04-09. The release log is dense through January 2026 and the open-source plan still carries unchecked items: PPO training, filler word removal, and additional languages. Upgrading is not a routine operation. Because vLLM is pinned to a wheel URL and the dependency list spans CUDA, DeepSpeed and quantization libraries, a bump to torch or transformers can invalidate the environment. Budget for rebuilding the container rather than patching it.

## Conclusion

Adopt Step-Audio-EditX if you need programmatic control over emotion, speaking style and paralinguistic tags in speech you already have, and you can supply a CUDA machine with Python 3.12. Do not adopt it if you need CPU-only inference, a stable versioned API, or languages outside Mandarin, English, Sichuanese, Cantonese, Japanese and Korean. Before committing, verify which checkpoint your GPU can hold, whether the pinned vLLM wheel resolves on your platform, and how the tag syntax behaves on your own audio.

## FAQ

### Can AI edit my audio with Step-Audio-EditX?

Yes, that is the model's primary purpose. The README describes it as a 3B-parameter LLM-based reinforcement-learning audio model for expressive and iterative editing of emotion, speaking style and paralinguistics, with zero-shot TTS as an additional capability.

### Can I edit audio for free with Step-Audio-EditX?

The code is Apache-2.0 and the weights are published on Hugging Face and ModelScope, so you can run it yourself without paying for a hosted service. The README does not state the terms attached to the model weights, and running it requires a CUDA GPU.

### Which app is best to edit audio, and where does Step-Audio-EditX fit?

The README does not compare Step-Audio-EditX with other editors, so no ranking can be drawn from it. What it does state is the scope: tag-driven editing of emotion, speaking style and paralinguistic features, plus zero-shot TTS, exposed through a Gradio app, a Hugging Face Space and the StepFun Audio Studio page.

### What is the best software for editing audiobooks, and is Step-Audio-EditX one?

The README does not position the project for audiobook work and does not compare it with other software. It documents iterative editing of emotion, speaking style and paralinguistic tags in speech, which is a narrower job than a full audiobook editing workflow.

## Sources

- [Issues](https://github.com/stepfun-ai/Step-Audio-EditX/issues)
- [License: Apache-2.0](https://github.com/stepfun-ai/Step-Audio-EditX/blob/main/LICENSE)
- [README](https://github.com/stepfun-ai/Step-Audio-EditX/blob/main/README.md)
- [stepfun-ai/Step-Audio-EditX on GitHub](https://github.com/stepfun-ai/Step-Audio-EditX)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/stepfun-ai-step-audio-editx
