Open-source project
yeyupiaoling/MASR avatar
yeyupiaoling/MASR

MASR: A PyTorch Framework for Streaming Chinese Speech Recognition

Pytorch实现的流式与非流式的自动语音识别框架,同时兼容在线和离线识别,目前支持Conformer、Squeezeformer、DeepSpeech2模型,支持多种数据增强方法。

728 stars115 forksPythonApache-2.0

At a glance

What is it?
MASR is an open-source automatic speech recognition toolkit built on PyTorch, offering both streaming and non-streaming inference with Conformer, Squeezeformer and DeepSpeech2 architectures. Pretrained models cover Mandarin, Cantonese, English, Chinese-English mixed speech, and Uyghur, though downloading them requires a paid subscription to the author's knowledge platform.
Who is it for?
Engineers who need to train or fine-tune a streaming ASR model for Mandarin or Cantonese on their own data will find MASR's configuration-driven approach practical. Those who only need inference and do not want to train from scratch should verify first that they are willing to subscribe to the author's paid knowledge platform, because pretrained model files are not distributed through the repository.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 86 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Streaming and Non-Streaming ASR in One Toolkit

MASR, which the README expands as Magical Automatic Speech Recognition, is a PyTorch-based toolkit aimed at engineers who want to train, evaluate, or deploy automatic speech recognition models without assembling components from separate research repositories. Its primary focus is Mandarin, but V3 added support for Cantonese, English, Chinese-English mixed speech, and Uyghur through a new tokenization system.

The defining design choice is that streaming and non-streaming inference share the same model weights. A single parameter, the `streaming` flag in the model configuration file, switches between real-time frame-by-frame decoding and full-utterance processing. A deployment that starts with offline batch transcription can switch to streaming by editing one line, with no retraining. The README states the framework targets server deployment and Nvidia Jetson edge devices; mobile device support for Android and similar platforms is listed as a future goal but is not present in V3.

Four Architectures and Four Decoders

MASR supports four model architectures: Conformer, Squeezeformer, EfficientConformer, and DeepSpeech2. Conformer is the architecture used in all reported evaluations across WenetSpeech, AIShell, and Librispeech. DeepSpeech2 offers a simpler recurrent design for teams familiar with that lineage.

Decoding is separate from the model and selectable at inference time. The four options cover different accuracy-latency trade-offs. ctc_greedy_search picks the single most likely token at each step, the fastest option. ctc_prefix_beam_search extends that with a beam search over prefix sequences; the README's evaluation tables use a beam size of 10 for this decoder. attention_rescoring runs ctc_prefix_beam_search first, then rescores the beam candidates with an attention decoder; across all published evaluations in the README, this decoder achieves the lowest character error rate when used with Conformer. ctc_beam_search requires an external CTC beam search library and appears less universally applicable than the others.

The preprocessing layer uses either fbank or mfcc features. The underlying library is kaldi_native_fbank, which replaced an earlier library in V3. According to the README, that switch improved data preprocessing speed and added multi-platform compatibility.

Installing MASR and Running Your First Inference

The README lists the supported environment as Anaconda 3, Python 3.11, and PyTorch 2.5.1 on Windows 11 or Ubuntu 22.04. Clone the repository and install its Python dependencies from requirements.txt:

bash
git clone https://github.com/yeyupiaoling/MASR
cd MASR
pip install -r requirements.txt

The requirements.txt included in the repository specifies kaldi_native_fbank, sentencepiece, onnxruntime, scikit-learn, soundfile, soundcard, resampy, and yeaudio among the core packages.

After installation, several top-level scripts serve different inference scenarios. infer_path.py handles single-file recognition. infer_server.py starts a web server for network requests. infer_gui.py opens a graphical interface. infer_sd_asr.py adds speaker-separation-aware inference. Model evaluation uses eval.py, which the README cites as the tool for measuring character error rate and word error rate on test sets. For teams that want to run inference without the full PyTorch dependency, export_model.py exports trained weights to ONNX format.

Eight Augmentation Methods and sentencepiece Tokenization

MASR includes eight built-in data augmentation methods for training: noise augmentation, reverb augmentation, speed augmentation, volume augmentation, resampling augmentation, shift augmentation, SpecAugmentor, and SpecSubAugmentor. The README notes that the WenetSpeech pretrained models were trained with noise and reverb augmentation among others; the complete augmentation settings for any training run live in configs/augmentation.yml.

The switch to sentencepiece tokenization is the other major training-side change in V3. The README states that sentencepiece substantially lowered the difficulty of handling multiple languages and made Chinese-English mixed training possible. Token vocabularies built with sentencepiece are not compatible with V2's character-level tokens, which is why V3 models and V2 models cannot be exchanged. The tools/ directory in the repository holds the utilities for building new sentencepiece vocabularies from custom datasets, which is the practical path for anyone adding a new language.

Pretrained Models and the Subscription Barrier

All pretrained models listed in the README, covering WenetSpeech (10,000 hours of Mandarin), AIShell (179 hours of Mandarin), Librispeech (960 hours of English), Cantonese, Chinese-English mixed data, and Uyghur, are distributed only to subscribers of the author's paid knowledge platform. The download column in every model table in the README reads 加入知识星球获取, meaning all weights are gated behind a paid membership.

This matters for evaluation. The repository has no free model weights. An engineer who wants to measure recognition quality on their own audio before committing to training must either subscribe or train a model themselves, which requires the matching training data. The README mentions an online demo at tools.yeyupiaoling.cn that allows browser-based inference without weights, but it gives no control over decoder selection, models, or audio format.

Where MASR Is the Wrong Choice

MASR has no integration path for frameworks other than PyTorch. Teams whose pipelines run on PaddlePaddle, TensorFlow, or JAX need a different toolkit.

Mobile deployment is not supported in the current version. The README explicitly lists mobile device support as a future goal. For Nvidia Jetson, ONNX export provides a route, but the README does not detail the full Jetson setup.

The speaker-separation inference path (infer_sd_asr.py) exists in the repository, but the README does not document the full configuration pipeline for this mode. Teams that need a documented diarization workflow should verify that the documentation covers their case before relying on this mode.

V3 is not backward-compatible with V2. Any existing trained weights, custom vocabularies, or data pipelines built on V2 require rebuilding under V3's sentencepiece approach. The V2 branch remains on GitHub but receives no new development.

MASR Against WeNet: Two Approaches to Chinese ASR

WeNet is another open-source PyTorch ASR toolkit that also targets Mandarin and also implements the Conformer architecture with CTC-based decoding. The central difference is deployment packaging. WeNet ships a C++ runtime and a gRPC server layer intended for production environments where Python dependencies are undesirable or where low-latency serving matters. MASR stays in Python throughout and uses infer_server.py as its server entry point.

Teams that need a framework they can embed in a C++ service or deploy at low latency with minimal runtime overhead should examine WeNet. Teams that want to iterate quickly in Python, run custom data augmentation, or train multilingual models using sentencepiece fit MASR's design better. The two frameworks are not mutually exclusive at the architecture level: both implement Conformer and CTC decoding, so a team could prototype in MASR and consider WeNet later for production without retraining from a fundamentally different architecture.

Maintenance, V3 Compatibility Break, and Apache-2.0 License

The last push to the repository was on 2026-07-06. The repository is not archived. The README labels V3 as the final version of the V3 line and states it was formally released in March 2025. The project has no GitHub releases; versioning is tracked through commit history and branch naming.

The Apache-2.0 license permits commercial use, modification, and redistribution with attribution. Teams building a product on top of MASR should also review the licenses for kaldi_native_fbank, sentencepiece, and onnxruntime, since those dependencies are listed in requirements.txt and carry their own terms. The ONNX export path can help teams separate PyTorch-dependent training from inference under a different runtime environment, which may simplify the licensing review for the inference component.

Editorial conclusion

Engineers who need to train or fine-tune a streaming ASR model for Mandarin or Cantonese on their own data will find MASR's configuration-driven approach practical. Those who only need inference and do not want to train from scratch should verify first that they are willing to subscribe to the author's paid knowledge platform, because pretrained model files are not distributed through the repository. Anyone locked into a framework other than PyTorch should look elsewhere; MASR runs on PyTorch 2.5.1 and offers no support for other backends. The last push was on 2026-07-06, under the Apache-2.0 license.

Frequently asked questions

What model architectures does MASR support?

MASR supports Conformer, Squeezeformer, EfficientConformer, and DeepSpeech2. Each architecture works in both streaming and non-streaming modes, controlled by the streaming flag in the model configuration file.

How do I download MASR pretrained models?

All pretrained models listed in the README require joining the author's paid knowledge platform (知识星球). The repository does not host any free model weights, and there is no academic download link.

Is MASR V3 backward compatible with V2?

No. The README states that V3 is incompatible with V2. The two versions use different tokenization approaches: V3 uses sentencepiece while V2 used a character-level method. The V2 branch remains on GitHub for existing users, but new development is on V3 only.

Official sources

  1. Issues
  2. License: Apache-2.0
  3. README
  4. yeyupiaoling/MASR on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/yeyupiaoling-masr.svg)](https://hysenlabs.com/projects/yeyupiaoling-masr)