# GPA: One Autoregressive Model for ASR, TTS, and Voice Conversion

> AutoArk's GPA trains a single autoregressive transformer to handle speech recognition, text-to-speech, and voice conversion together. Version 1.5 delivers near-SOTA performance on ASR and TTS, but Voice Conversion and full deployment tooling remain on the roadmap.

**AutoArk/GPA** — [AutoArk] GPA (General Purpose Audio) can do ASR, TTS and voice conversion with one tiny model!

- Repository: https://github.com/AutoArk/GPA
- Website: https://autoark.github.io/GPA/
- Stars: 3,107 · Forks: 206
- Language: Python
- License: Apache-2.0
- Published: 2026-09-16 · Updated: 2026-09-16 · Language: en
- Canonical page: https://hysenlabs.com/projects/autoark-gpa

## What GPA Is and Who It Is For

GPA stands for General Purpose Audio. The project's stated goal is to do for audio what a unified language model does for text: handle multiple tasks through one set of weights trained on one training objective. In GPA-v1.5, the two completed tasks are automatic speech recognition (ASR) and text-to-speech synthesis (TTS). Voice Conversion is listed on the roadmap but has no native path in the current version.

The primary audience is AI researchers and engineers who want to evaluate whether a single unified model can match task-specific baselines on ASR and TTS. It is also relevant to teams building voice products who want to reduce the number of model checkpoints they maintain in production. GPA is not aimed at casual users: the setup requires Python 3.10 or later, PyTorch 2.1 or later, and a working understanding of Hugging Face model management.

The repository sits under AutoArk on GitHub and publishes checkpoints to Hugging Face under the AutoArk-AI organization. The paper is available on arXiv (2601.10770). The last push to the repository was on 2026-05-25.

## The Autoregressive Architecture Behind the Unified Model

GPA-v1.5 is described in the README as a larger, cleaner, and more capable unified audio-language model compared to the earlier GPA-0.3B-preview. The core claim is that speech understanding (ASR) and speech generation (TTS) share enough structure that a single autoregressive transformer, trained jointly, can perform both without a separate encoder or decoder for each task.

The README describes it as an auto-regressive transformer that unifies audio generation and understanding in a single model. The training pipeline uses Hugging Face Trainer, which means it follows the standard supervised fine-tuning workflow familiar from LLM training, with mixed-precision and distributed training support through deepspeed and accelerate. The repository's pyproject.toml specifies deepspeed >= 0.12.0 and accelerate as core training dependencies.

For inference, GPA-v1.5 supports several execution backends: native PyTorch and Hugging Face pipelines, vLLM, llama-cpp, sglang, and mlx-lm. RKNN is listed on the roadmap. This means users with different inference infrastructure can try the model without switching frameworks.

## Installing GPA and Running the First Inference

The GPA repository does not have GitHub releases. Checkpoints are distributed through Hugging Face (AutoArk-AI/GPA-v1.5) and ModelScope (AutoArk/GPA-v1.5). The README points to separate documentation files for inference and training: GPA_1.5/docs/infer.md for inference and GPA_1.5/docs/train.md for training. Those documents are not included in the prompt material, so the steps below come from what the root README and pyproject.toml state directly.

For the training environment, install from the pyproject.toml at the repo root:

```bash
pip install -e .
```

The core runtime dependencies declared in requirements.txt are:

```bash
torch>=2.1.0
torchaudio
transformers>=4.57.0
acceleate
librosa
soundfile
soxr
```

For ONNX-based inference, the README describes a separate runtime bundle: GPA-v1.5 ONNX runtime, available as a CLI tool, a FastAPI service, and a browser UI. The ONNX runtime assets are published to Hugging Face under AutoArk-AI/GPA-v1.5-onnx-runtime. The ONNX runtime documentation lives in GPA_1.5/onnx_runtime/README.md in the repository. This path is appropriate for deployment scenarios where PyTorch is not available or where binary size matters.

## GPA-TTS: The Standalone Edge TTS Runtime

Because TTS was the most used feature in the online demo, the AutoArk team extracted the TTS component from GPA-v1.5 into a self-contained runtime called GPA-TTS. This is a separate subdirectory (GPA_TTS/) in the same repository with its own README and its own Hugging Face checkpoint.

GPA-TTS supports zero-shot voice cloning from a short reference audio clip. Its quantization scheme is Qwen INT4 for the language model component and INT8, FP16, or FP32 for the SparkDetokenizer decoder. The decoder precision is selectable at runtime through the CLI, the API, or the web UI. INT8 targets edge and CPU inference; FP16 offers a balance of quality and speed; FP32 is the highest quality option for users with compute headroom. The README describes GPA-TTS as being among the smallest open-source TTS runtimes that support voice cloning.

GPA-TTS is built to run on CPU on Mac, Linux, and edge hardware. It does not require a full PyTorch training environment. The checkpoint is available under the main GPA Hugging Face repository at AutoArk-AI/GPA/tree/main/GPA_TTS. For teams who want lightweight TTS without committing to the full unified model, GPA-TTS is the more practical entry point.

The GPA-TTS announcement was on 2026-03-31, with FP16 and FP32 decoder options added on 2026-04-07.

## What Is Missing: Roadmap Gaps and Practical Constraints

GPA has several well-defined gaps that engineers should confirm before choosing it for production use.

First, Voice Conversion has no working path in GPA-v1.5. The roadmap marks it as planned but incomplete. Users who need voice conversion today must use a separate tool. The repository's roadmap table shows it as an empty checkbox under GPA-v1.5 Next Steps.

Second, the basic service deployment recipes for vLLM and FastAPI are listed on the roadmap as incomplete (empty checkbox). This means production serving requires engineers to write their own serving layer on top of the inference code. The ONNX runtime provides a FastAPI service path, but that path is separate from the native PyTorch inference pipeline.

Third, there are no GitHub releases. All versioning is handled through checkpoint names and direct Hugging Face downloads. Teams that depend on a versioned release artifact with a changelog will need to track commits manually.

Fourth, the training environment pulls heavy dependencies. deepspeed >= 0.12.0, transformers >= 4.57.0, vllm >= 0.6.0, and accelerate are listed in pyproject.toml. These are not trivial to install on every platform, particularly on non-CUDA hardware.

Whisper from OpenAI provides a point of comparison for the ASR component. Whisper is a purpose-built ASR-only model distributed in several sizes with well-documented deployment tooling. The difference with GPA is that GPA covers ASR and TTS in one set of weights, whereas Whisper requires a separate TTS system alongside it. The trade-off is that Whisper's single-task focus means a narrower dependency graph and a larger ecosystem of integrations.

## License and Maintenance

The repository is licensed under Apache-2.0, which permits commercial use, modification, and distribution, as long as the original license and notice are included. The Apache-2.0 license is permissive enough for most production deployments without legal complexity.

The last push was on 2026-05-25, which is within six months of 2026-09-28, and the repository shows active development: the announcements section in the README records a series of releases through April 2026. The project has a roadmap with concrete open items, which suggests ongoing development rather than a finished archive. There is no indication of a formal release cadence, and all model versions are distributed through Hugging Face rather than GitHub releases.

## Conclusion

GPA-v1.5 is worth evaluating for teams that want a single model weight covering both speech recognition and text-to-speech, and who are comfortable working from checkpoints rather than packaged releases. Verify that your deployment target is covered before committing: vLLM and FastAPI service recipes are listed as incomplete in the roadmap, and Voice Conversion has no native v1.5 path yet. The GPA-TTS spinoff is the better starting point for production edge TTS deployments today, since its quantization options and standalone design require fewer dependencies than the full training environment.

## FAQ

### How do I run GPA-v1.5 inference without setting up the training environment?

The GPA-v1.5 ONNX runtime provides a CLI, a FastAPI service, and a browser UI that do not require the full deepspeed and training stack. The runtime assets are available on Hugging Face under AutoArk-AI/GPA-v1.5-onnx-runtime, and the documentation is at GPA_1.5/onnx_runtime/README.md in the repository.

### Does GPA-v1.5 support Voice Conversion?

Not in the current release. The roadmap lists Voice Conversion as a planned native path for GPA-v1.5, but the entry is marked incomplete. The original GPA-0.3B-preview included VC support; the v1.5 branch has not yet implemented it.

### What hardware does GPA require for inference?

The README does not specify minimum GPU VRAM for inference. The training environment requires CUDA-compatible hardware given the deepspeed and vllm dependencies. GPA-TTS, the standalone TTS component, is explicitly designed for local CPU inference on Mac, Linux, and edge devices.

## Sources

- [AutoArk/GPA on GitHub](https://github.com/AutoArk/GPA)
- [Issues](https://github.com/AutoArk/GPA/issues)
- [License: Apache-2.0](https://github.com/AutoArk/GPA/blob/main/LICENSE)
- [Project website](https://autoark.github.io/GPA/)
- [README](https://github.com/AutoArk/GPA/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/autoark-gpa
