# MAW's stable download and its latest tag are different builds, and its peak kernel has no Python fallback

> MAW is an API-driven subtitle workflow that transcribes local media, produces SRT plus a .mosp project, and edits in a local server-based editor. The packaging metadata is more interesting than the README: a Rust extension as the only peak path, two Transformers majors split across two requirement files, and CUDA wheels routed to every platform that is not macOS.

**Moyf/moys-asr-workflow** — Moy 的 ASR 字幕生成工作流及编辑器，简称 MAW

- Repository: https://github.com/Moyf/moys-asr-workflow
- Stars: 489 · Forks: 44
- Language: Python
- License: AGPL-3.0
- Published: 2026-09-17 · Updated: 2026-09-17 · Language: en
- Canonical page: https://hysenlabs.com/projects/moyf-moys-asr-workflow

## The latest tag is a beta and the stable entry point is not

The release cadence is fast and the versioning is not hidden. Three tags land within about a week: v1.6.1 on 2026-09-23, v1.7.0 on 2026-09-24 and v1.8.0-beta.1 on 2026-10-02, the last push to the repository being the day before. The packaging metadata agrees with the beta: the project version in pyproject is 1.8.0-beta.1. That creates a split a new user meets immediately, because the badges and the first quick-start step both send you to the latest release, which is the beta, while the README states that the stable download entry point stays where it was. So the default download is a test build and the stable path is a link the reader has to notice separately. Given an AGPL licence and a tool that touches paid API endpoints, that distinction is worth more attention than a single sentence in the middle of a Chinese README gives it.

## Seven provider scripts live at the repository root

The tree is unusual in a way that tells you how the tool grew. Instead of one entry point per provider inside the package, there are seven top-level scripts named for the service they call, covering bcut, Doubao, a local model, an OpenAI-compatible endpoint, Qwen, Soniox and Tencent, alongside `edit.py` and `maw_gui.py` as the editing and launcher entry points. The actual package lives in a separate `maw/` directory, so the importable code and the runnable scripts are split across two locations at the root. There are also two server directories, one for the editor and one named for alignment, a `desktop/` tree, a `web/` tree and a `website/` tree, plus a bundled `blank-editor.html` and a PyInstaller spec file for building the desktop app. None of that is documented in the README, which is four sections and a link list. The CLI documentation is the place the parameter surface is described, along with server management and exit codes.

## The peak kernel is a Rust extension with no Python fallback

Waveform peak generation is delegated to a separate package called quapeaks, described in the dependency comments as a Rust kernel licensed MIT or Apache-2.0, kept in its own repository and renamed from an earlier name. Three details in that comment block matter more than the dependency line itself. It is the only generation path, with no Python reference implementation to fall back on. It is imported lazily at the point of use rather than at import time, so a missing kernel logs a message, skips cache generation and does not block the transcription entry point, with the cache rebuildable from an environment that has it. And the version is a calendar-style floor rather than a normal release pin, with a note that the header fingerprint takes the low 32 bits of the file's size and modification time so large files work. The design is deliberate, but the failure mode is quiet: an install without the kernel produces a running application with degraded output.

## Two Transformers majors, so a second requirements file

The dependency comment block explains a conflict rather than a preference. The MOSS runtime needs Transformers 5.x, while the local ASR dependency group pins an older Transformers release, and the two cannot coexist in the same dependency group of one lock file. The resolution is to declare the MOSS side in a separate requirements input file and freeze it with a separate compile command, running alongside the export pipeline for the local and OCR groups. That is a pragmatic answer, and it has a consequence: the repository now has two dependency resolutions to keep in step, and an environment built from one file is not the environment the other file describes. The main project itself is much lighter, needing Python 3.11 or newer and a short list including font tools, a Chinese segmenter, a traditional-to-simplified converter, a requests client, a recycle-bin helper for deletions, and pywebview for the desktop shell, which is declared twice with different extras depending on whether the platform is Linux.

## CUDA wheels for every platform that is not macOS

Torch and torchaudio are routed through an explicit package index built for CUDA 13.0, and the marker on that routing excludes only macOS. Every other platform, meaning Linux and Windows, resolves torch from the CUDA index rather than from the default source. That is the right call if the local ASR path assumes GPU hardware and you want a consistent wheel, and it is a heavy one otherwise, since a CPU-only machine on either platform pulls a CUDA build through the local ASR dependency group. The faster-whisper entry in that group is annotated as using the CTranslate2 backend and not depending on Transformers, which is the escape hatch for anyone who wants local transcription without the framework conflict described above. The project also declares build and development groups, pinning the packaging tool exactly and holding the linter to a lower bound.

## Exported subtitles keep none of the timing data

The important-notes section is where the real boundaries are, and it is unusually specific. Choosing a cloud transcription service means media is uploaded directly to that provider; the application has no server of its own and does not take custody of keys. Cost, retention and availability follow the provider's current policy. Then the export rules: the .mosp project is the source of truth for subtitles, SRT suits ordinary delivery, and ASS uses a default output scheme and associated styles from the editor's style manager, with script coordinates set from the source video resolution recorded in the project. The Launcher uses a separate default SRT style when it burns subtitles in. The localhost editor and the Launcher share a user-level style library while the portable file:// editor uses a browser-local copy. And neither subtitle format preserves word-level timecodes, waveforms or the rest of the project data, so a round trip through SRT or ASS loses them permanently.

## Hotwords require a table created in the provider console first

The environment file is the most detailed document in the repository and it reveals how much of the feature surface is really a two-sided workflow. The Qwen section carries an API key, a region choice between a Beijing default and Singapore, a workspace identifier that is optional in one region and required in the other, an optional default language, and switches for word-level timestamps and for inverse text normalisation, the latter fixed on for one model family. The Qwen-Audio section then requires a prebuilt vocabulary table that must already exist in the provider console under a specific model name, plus an immediate hotword weight that accepts either a one-to-five scale or a single high value, and a context file whose contents are capped at four hundred characters per request. The Fun-ASR hotword section repeats the constraint in the other direction: its table has to be created for that target model and cannot be reused from the Qwen one. That is the shape of a tool whose defaults come from a paid API.

## A committed debug log and two copies of the same FAQ

Repository hygiene is mixed, and the top level shows it. A `debug.log` file is committed next to the README, which is the kind of artefact that should be ignored rather than published. The FAQ exists twice, as a plain text file with a Chinese name at the root and as a markdown document inside the docs directory, so a reader has two sources to keep in step. The rest is disciplined: a third-party notices file, a changelog, a contributing guide, a design document, an agents file, a coding-assistant instruction file, a Python version file, a lock file for the main resolution, a lock file for the dev tooling, and two Playwright configuration files, one of them dedicated to scroll behaviour. The JavaScript side is explicitly labelled dev-only browser regression tooling, private, version zero, and described as having no production dependency, which is the honest way to ship test infrastructure in a repository that also ships a desktop application.

## Conclusion

MAW is a serious piece of work for anyone producing subtitles from Chinese or multilingual media, and its honesty about data is the first thing to note: media goes straight to whichever provider you pick, there is no server in the middle, and keys are never held for you. Three things to check before adopting it. Export is lossy, since SRT and ASS keep neither word-level timing nor waveforms, so treat the .mosp project as the only round-trippable artefact. The peak generation path is a Rust extension with no Python fallback, imported lazily so the app still starts without it, which means cache generation quietly degrades rather than failing loudly. And the source tree is on a beta version while the stable download entry is unchanged, so decide deliberately whether you are testing 1.8.0-beta.1 or building from the 1.7.0 line. The licence is AGPL-3.0-only, which matters if you plan to modify it.

## FAQ

### What is Moy's ASR Workflow?

An API-first subtitle generation and editing workflow, usually shortened to MAW. Media stays local, transcription goes through an external speech service, and the output is an SRT file plus a .mosp project that the MAWE editor opens for correction before export. It ships a Windows and macOS desktop app, a public CLI and a local server editor, under AGPL-3.0-only.

### Which speech recognition services does MAW support?

Qwen, Fun-ASR, Soniox, Tencent Cloud recording file recognition and Doubao from Volcano Engine, plus any OpenAI-format compatible endpoint. Media is uploaded directly to whichever provider you configure, and MAW runs no server of its own and does not hold your API keys.

### What do I lose by exporting SRT or ASS instead of the .mosp project?

Neither subtitle format keeps word-level timecodes, waveforms or other project data, so the .mosp project is the only round-trippable artefact. ASS styling comes from the editor's style manager and uses the source resolution recorded in the project, while the Launcher applies its own default SRT style when burning subtitles in.

### Are the local ASR options in MAW production ready?

Not according to the documentation. Local Qwen3-ASR, FunASR and Faster-Whisper, along with the key-free Jianying ASR entry, are described as experimental and suitable only for trying the workflow out, as opposed to the cloud providers listed as the main path.

## Sources

- [Issues](https://github.com/Moyf/moys-asr-workflow/issues)
- [License: AGPL-3.0](https://github.com/Moyf/moys-asr-workflow/blob/main/LICENSE)
- [Moyf/moys-asr-workflow on GitHub](https://github.com/Moyf/moys-asr-workflow)
- [README](https://github.com/Moyf/moys-asr-workflow/blob/main/README.md)
- [Releases](https://github.com/Moyf/moys-asr-workflow/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/moyf-moys-asr-workflow
