# open-audio-opd compares its 0.6B ASR student against published Qwen3 numbers, ships a vendored framework inside the wheel, and has no TTS yet

> open-audio-opd is a training stack for online policy distillation in speech recognition: a small autoregressive student rolls out transcripts, a stronger teacher scores the same audio and transcript, and the student is updated with a token-level divergence over the union of both models' top-k support. What the repository adds on top of the method is specific and worth reading in the manifest rather than the prose, including a wheel that declares no Python package at all.

**AutoArk/open-audio-opd** — Industrial audio online policy distillation (OPD) training stack for ASR and TTS, distilling compact audio models from stronger   teacher models.

- Repository: https://github.com/AutoArk/open-audio-opd
- Stars: 1,128 · Forks: 72
- Language: Python
- License: NOASSERTION
- Published: 2026-09-18 · Updated: 2026-09-18 · Language: en
- Canonical page: https://hysenlabs.com/projects/autoark-open-audio-opd

## The wheel declares no Python package and ships scripts, markdown and a vendored tree

The packaging configuration is the most surprising thing in this repository, and it is not a typo. The project builds with hatchling, and the wheel target sets its packages list to empty. What it ships instead is an artifacts list containing four directories and three files: scripts, the vendored framework directory, configs, and then the English readme, the Chinese readme, and the licence file. So the installable artefact is a wheel containing directories and markdown, with no importable module inside it and no console script declared. Anyone who installs this with a package manager gets a set of files on disk and nothing to import, which is consistent with a training stack meant to be run from a checkout rather than consumed as a library, but it also means the wheel is not a meaningful distribution channel for a framework that needs scripts, configuration templates, and a vendored tree to work at all. The licence adds a second oddity. The manifest names Apache-2.0 in its project table, while the project metadata carries no licence value at all, and the readme links a licence file. Three signals, two of which agree and one that says nothing.

## A trimmed copy of an upstream reinforcement learning library is committed in

The training stack declares its lineage openly, citing two upstream projects, and then vendors one of them. A trimmed vendored copy of the reinforcement learning library sits in the repository root as its own directory, and the stated reason is specific: so the training script can use the sharding wrapper, gradient clipping, and checkpoint management without depending on another local checkout. The practical consequence is that this repository now owns a fork of somebody else's framework and ships it inside its own wheel. Vendoring buys reproducibility, since the training run cannot break because an upstream release changed, and it costs drift, because nobody is merging fixes back in and the copy is described as trimmed, meaning parts were removed and the gap between this tree and upstream will widen in both directions. The manifest lists the vendored directory as a shipped artifact, so it is not merely present in the source tree. Two other dependencies in the same list are worth watching for a stack of this kind: an object store library version-bounded to a narrow range with one version explicitly excluded, and a build backend chosen for speed rather than for the metadata conventions the ecosystem has settled on.

## numpy is held below 2.0 while pyarrow and the serving engine are capped at older floors

The dependency list has one pin that stands out and two floors that do not. numpy is held below the 2.0 line, which in a stack that asks for a recent transformers, a recent torch, and a pyarrow from the 19 series means one dependency is pinned to a major version behind everything around it. Whatever forced that, it is a constraint you inherit, and it is the kind that decides whether an unrelated library can be upgraded in your environment. The pyarrow floor is recent by contrast, and torch and transformers both have high floors too. Then there are three optional groups, which is good practice and also where the real constraints live. A development group adds mypy, pytest, and ruff. An evaluation group adds a single accuracy-metric package, which tells you that character and word error rates are measured by an extra you have to ask for rather than by the training stack itself. A serving group pins an inference engine to a range ending at 0.11.0, which is a ceiling rather than a floor and will stop you from upgrading it. The interesting part is that the shipped inference script is described as a transformers script, so the core use case does not need that engine at all.

## The competing numbers are copied from a technical report, and a larger model still wins

The results section is carefully worded and the wording matters. The student model is 0.6B parameters trained on 100k hours of audio. The comparison is against figures that public technical-report material reports for a larger model family, and the data-scale comparison is framed as the encoder pretraining stage alone using about 40M hours of pseudo-labelled audio, which is where the roughly one-in-400 figure comes from. So the baseline rows are published numbers taken on someone else's terms, not the same models re-measured in the same harness by these authors. Within that framing the results hold: average English word error rate improves from 6.93% to 6.55% against the same-size baseline, and average Chinese character error rate from 4.36% to 4.30%. Two caveats are in the text itself. The larger 1.7B model is better on both averages, and the text says so. And the per-column detail shows the student is behind on one Chinese test set, so the average improves while a specific case regresses. Reading the column table rather than the takeaway list is the difference between an accurate picture and a flattering one.

## The project is named for ASR and TTS, and every TTS item is still planned

The repository description covers both automatic speech recognition and text to speech, and the abstract narrows the current release to recognition. The roadmap settles it with a status column. Everything on the recognition side is done: the sharded online policy distillation trainer, a teacher scoring backend in the style of a named open ASR model, resumable checkpointing, a multi-node launcher driven by a hostfile, and inference plus error rate evaluation scripts. All three text to speech items are marked planned: the online rollout and teacher scoring recipe, objectives for speech generation quality and alignment, and acoustic-token supervision support. The announcement section describes the intended reuse rather than the delivered work, saying the planned recipe will reuse the online student rollout and the teacher scoring adapted for generation quality, alignment, and acoustic token supervision. That is a reasonable description of what carrying this method over to synthesis would involve, since the two halves of the loop are generic, but it is a plan. Anyone expecting a text to speech recipe in this repository today will not find one, and the reasoning about why the method transfers is an argument rather than a result.

## No audio, no dataset, and no machine paths ship with it

The repository is code and nothing else, stated plainly: no audio files, no line-delimited JSON datasets, and no private machine paths are included, and all model, data, and output paths are explicit command-line arguments. That is the right call for a repository that others will adapt, and it has a cost worth naming, which is that reproducing the released checkpoint requires obtaining the data yourself and reconstructing the pipeline around these scripts. There is no data preparation code in the visible layout and no documented corpus, so the 100k hours behind the checkpoint are described by scale rather than by composition. The layout listing is itself a small artefact, and it is worth looking at directly:
```text
scripts/train/train_ark_asr_opd_fsdp2_resume.py      # main FSDP2 ASR OPD trainer
scripts/run/run_ark_asr_opd_fsdp2_resume_hostfile.sh # multi-node launcher
scripts/infer/ark_asr_transformers.py                # ASR inference
scripts/eval/eval_jwer_ark_asr_transformers.py       # J/WER evaluation
scripts/eval/run_arkasr_eval.sh                      # multi-GPU evaluation launcher
configs/hostfile.example                             # hostfile format example
https://arxiv.org/abs/2605.28139                     # arXiv paper
assets/opd_overview.png                              # OPD over
```
It is presented as a file tree, and among the script paths, the hostfile example, and the overview image, one line is an arXiv identifier rather than a path, and the comment on the last line stops partway through the word it was describing. Read as a tree, it mixes a link in with the files, which is the kind of copy and paste artefact that makes a layout listing slightly less reliable than it looks.

## Strict typing, enforced line length, and a release cadence that stopped in June

The quality configuration is stricter than most research stacks. The type checker runs in strict mode pinned to the minimum supported Python version, and the linter selects the pycodestyle error set, the static checks, import sorting, modern syntax upgrades, and bugbear, with a line length of 100 characters. Selecting the error set is what makes that line length binding, which is not true of every project that declares one. Development dependencies for the checker, the test runner, and the linter are all declared rather than assumed. The training side is a single entry-point script for the sharded recognition trainer, with a companion shell script for multi-node launching driven by a hostfile format that has its own example file in the configs directory, and resumable checkpointing is a completed item rather than a wish. Against that, the release record is thin. There are no GitHub releases at all, so there is no version to pin, and the last push to the default branch is dated 2026-06-05, roughly two weeks after the three announcements in late May covering the repository, the weights, the paper, and the hosted demo.

## Conclusion

open-audio-opd is worth reading if you are training a compact speech recognition model with a teacher you already have, because the release is honest about what is finished and what is not, the training entry point is one script, and the multi-node launcher plus resumable checkpoints mean a long run does not have to start over. Treat the benchmark table with more care than it invites. The competing rows are figures copied from a public technical report rather than measured here, so the comparison is against published numbers under someone else's evaluation setup, and a 1.7B model is better on both the English and Chinese averages in that same table. Before you build on it, check four things. Which base framework you are actually running, since a trimmed copy of the upstream reinforcement learning library is committed into the repository and shipped in the wheel, and it will drift. Which numpy you resolve, because the manifest holds it below 2.0 while its pyarrow and tensordict floors are much newer. What your evaluation harness is, since the accuracy metric package is an optional extra and the scripts are transformers-based rather than built on the optional serving engine. And whether you need text to speech at all, because all three text to speech roadmap items are still marked planned.

## FAQ

### What is open-audio-opd and what does it do?

It is a training stack for online policy distillation in speech recognition. A 0.6B autoregressive student model rolls out transcripts on its own audio, a stronger teacher scores the same audio and transcript, and the student is updated with a token-level divergence over the union of the two models' top-k support.

### How does Ark-ASR compare with Qwen3-ASR?

On average English word error rate it improves from 6.93% to 6.55% against the same-size baseline, and on average Chinese character error rate from 4.36% to 4.30%. Those baseline figures come from public technical-report material rather than a re-measured run, and a larger 1.7B model remains better on both averages in the same table.

### How much audio data does Ark-ASR training use?

The OPD experiments use 100k hours of ASR audio, which the documentation compares against the roughly 40M hours of pseudo-labelled audio reported for the encoder pretraining stage of the larger baseline family, putting this at around one in 400 of that disclosed scale.

### Does open-audio-opd support text to speech?

Not yet. All three text to speech roadmap items are marked planned, including the online rollout and teacher scoring recipe, speech generation quality and alignment objectives, and acoustic-token supervision support. Only the recognition side is complete.

### How do I install open-audio-opd?

The manifest declares Python 3.10 or newer with hatchling as the build backend, and a wheel that contains no Python package, only the scripts, configuration, vendored framework, readmes and licence as artifacts. That means the practical path is running the training script from a checkout rather than importing an installed library.

### What is the licence of open-audio-opd?

The project manifest names Apache-2.0 and a licence file is linked from the readme, while the project metadata carries no licence value. The manifest and the file agree, and the missing metadata value is the outlier rather than a second licence.

## Sources

- [AutoArk/open-audio-opd on GitHub](https://github.com/AutoArk/open-audio-opd)
- [Issues](https://github.com/AutoArk/open-audio-opd/issues)
- [README](https://github.com/AutoArk/open-audio-opd/blob/master/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/autoark-open-audio-opd
