# Microsoft Muzic: a research monorepo for music understanding and generation

> Muzic is a Microsoft Research Asia project that collects more than a dozen music AI models in one repository. It is a paper companion, not a product, and the README points you to a per-model folder for each one.

**microsoft/muzic** — Muzic: Music Understanding and Generation with Artificial Intelligence

- Repository: https://github.com/microsoft/muzic
- Stars: 4,962 · Forks: 501
- Language: Python
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/microsoft-muzic

## What Muzic actually is, and who it is written for

Muzic is described in its README as a research project on AI music that applies deep learning to music understanding and generation. It was started by researchers from Microsoft Research Asia, with outside collaborators, and it is published under the MIT licence. The repository is a collection rather than a single tool: the README lists work across symbolic music understanding (MusicBERT), automatic lyrics transcription (PDAugment), contrastive language-music pre-training (CLaMP), song writing (SongMASS, DeepRapper, TeleMelody, ReLyMe, ROC), music form and structure generation (MeloForm, Museformer), multi-track generation (PopMAG, GETMusic), text-to-music generation (MuseCoco), singing voice synthesis (HiFiSinger), and an AI agent for music processing (MusicAgent).

The audience is narrow and specific. Each entry is tied to a paper with a venue and year, so the primary reader is someone reproducing or extending published work. If you want a library that takes a text prompt and returns an audio file, the README does not promise that. It says the code of several research works is released and that each folder has its own README with detailed instructions. That sentence is the real contract: the top level tells you what exists, and the folder tells you how to run it. Everything else is on you.

The repository is not archived, and the last push was on 2026-08-05. That is recent enough that the project is still receiving commits, but the only tagged release listed is DeepRapper-v1.0 from 2021-08-24, so do not read commit activity as a stability guarantee for any individual model.

## How the code is organised across the model folders

The top level of the repository is a directory per research work: clamp, deeprapper, emogen, getmusic, meloform, musecoco, museformer, musicagent, musicbert, pdaugment, relyme, roc, songmass, telemelody, alongside img, requirements.txt and the usual Microsoft open source files (LICENSE, SECURITY.md, SUPPORT.md, CODE_OF_CONDUCT.md). There is no shared package, no single entry point and no top-level CLI. The coupling between models is essentially the Python environment.

That environment is pinned in requirements.txt, and the pins are the most consequential detail in the repository. It fixes torch==1.7.1, fairseq==0.10.0, transformers==3.5.1, scikit-learn==0.21.3, scipy==1.3.1, tensorboard==2.4.0, miditoolkit==0.1.14, pypianoroll, pretty_midi, pypinyin==0.39.1, jieba==0.42.1, pkuseg==0.0.25, textrank4zh==0.3, keras, nltk, librosa, pyworld, soundfile and others. Several of these are old enough that a modern Python will not build them without help. The README states the tested combination explicitly: Linux, Ubuntu 16.04.6 LTS, CUDA 10, Python 3.6.12. Treat that line as the supported configuration, because nothing in the repository suggests a newer one has been validated.

The data flow differs per model, but a common shape is visible from the dependency list. Symbolic work loads MIDI through miditoolkit, pretty_midi or pypianoroll, tokenises it, and feeds a fairseq or transformer model. Chinese lyric work pulls in pypinyin, jieba, pkuseg and textrank4zh for segmentation and rhyme handling. Audio-side work uses librosa, pyworld and soundfile. So the practical integration point between models is files on disk, not an API.

## Installing Muzic and running a first model

The README gives one install command for the whole repository, run from the top level. It installs the shared requirements; it does not install any model, and it does not check your CUDA version.

```bash
pip install -r requirements.txt
```

Expect this to be the slow part. The pins include torch==1.7.1 and fairseq==0.10.0, and on a recent Python the resolver will either fail or try to build packages from source. The README's stated target is Python 3.6.12 on Ubuntu 16.04.6 LTS with CUDA 10, so a matching container is the path of least resistance.

After that, pick one model folder and read its README. The top-level README does not contain run commands for any model; it says only that you can find the README in the corresponding folder for detailed instructions. For example, the MusicBERT work on symbolic music understanding lives in the musicbert directory.

```bash
cd musicbert
cat README.md
```

The README for that folder is where the actual invocation lives, along with any data preparation and checkpoint download steps. Do not assume the command you find there will work unchanged against the top-level requirements: some folders carry their own instructions, and the top-level file is a union of the dependencies of many models, not a curated set for one.

The samples the project points to are hosted outside the repository, at the ai-muzic.github.io page linked from the README. That page is the fastest way to hear what the models are supposed to produce before you spend time on the environment.

## Where Muzic breaks down in practice

The dependency pins are the first real limitation. torch==1.7.1 and fairseq==0.10.0 are old, and the README's tested stack is Ubuntu 16.04.6 LTS with CUDA 10 and Python 3.6.12. If you are on a current GPU driver or a Python that no longer ships the modules those versions expect, installation becomes a compatibility exercise rather than a setup step. Nothing in the README claims otherwise.

The second limitation is that the repository is a set of research releases, not a maintained product. The only release tag listed is DeepRapper-v1.0 from 2021-08-24, and the models span papers from 2020 to 2023. Some folders may be complete and runnable; others may assume data you do not have. The README does not document a support policy, a deprecation process, or rollback for any of it, and SUPPORT.md is the standard Microsoft open source support file rather than a project-specific channel.

The third is scope mismatch. The README's concept map and its list of current work describe the research programme, and the sentence about released code names a subset of folders. If the model you care about appears in the paper list but you cannot find a matching directory, the README does not tell you when or whether it will appear. Check the directory listing before you plan around a model.

Finally, this is the wrong tool if you want a hosted service or an end-to-end application. There is no server, no API, no web UI and no packaged model in the repository. You get source code, a requirements file, and per-folder instructions.

## Muzic compared with a general audio generation toolkit

The closest thing to a general-purpose alternative in this space is a toolkit built around audio generation from text or audio prompts, where one install gives you a model you can call directly. The difference in approach is structural rather than a matter of quality. A general toolkit optimises for a single supported pipeline: one environment, one inference entry point, one set of checkpoints, and a promise that the thing runs.

Muzic optimises for breadth of research coverage. It carries symbolic music understanding, lyric transcription, lyric-to-melody, structure modelling, multi-track generation, singing voice synthesis and an agent, each with its own paper and its own folder. That breadth is the reason to choose it, and it is also why there is no single entry point. If your task is symbolic music and you need a model that was trained on MIDI tokens rather than waveforms, a waveform-first toolkit is the wrong shape, and Muzic's musicbert or museformer directories are the relevant ones.

Conversely, if your task is generating audio from a text description and you do not care which architecture does it, the per-model research structure is overhead. You would spend your time reconciling torch==1.7.1 and fairseq==0.10.0 instead of generating anything. The honest framing is that Muzic is a reference implementation collection, and general toolkits are products. Pick based on whether you need to read and modify the model, or just call it.

## Licence, maintenance and the cost of upgrading

The repository is MIT licensed, which is permissive and places few conditions on reuse beyond attribution and the licence text. Note that this covers the code in the repository. The README links to papers and to a samples page, and it does not state a licence for model checkpoints, training data or generated audio. If you plan to ship something built on a checkpoint, that is the question to resolve before you build, and the repository does not answer it for you. This is a description of what the licence file says, not legal advice.

On maintenance, the last push was on 2026-08-05, so the repository is still receiving changes. That does not translate into a stable interface. There is no versioned API, no changelog beyond the What is New list in the README, and the single release tag is from 2021. Upgrading means re-reading the folder README and re-resolving requirements.txt, and the pins mean you cannot simply move to a current PyTorch and expect the fairseq-based models to load.

The practical upgrade cost is therefore a container rebuild rather than a dependency bump. If you vendor Muzic into a product, budget for holding the old environment alongside it, because the alternative is porting model code off fairseq==0.10.0 and torch==1.7.1 yourself.

## Conclusion

Adopt Muzic if you are reproducing a specific paper from the list, or if you need a reference implementation of symbolic music understanding or generation and can supply your own data pipeline. Do not adopt it if you want a supported library with a stable API, a single install that gives you a working music generator, or anything that runs without a GPU and Linux. Before committing, check three things: whether the folder for the model you want has a README with runnable instructions, whether the pinned versions in requirements.txt resolve on your Python and CUDA combination, and whether the model you need is actually released as code, since the README describes the scope of the project more broadly than the list of released folders.

## FAQ

### What does Muzic mean?

The README says Muzic is pronounced [ˈmjuːzeik] and describes it as a research project on AI music. The name is the project's own, not a term with a separate definition in the repository.

### What is Microsoft Muzic?

It is a research project from Microsoft Research Asia, released under the MIT licence, that covers music understanding and generation with deep learning. The repository collects models including MusicBERT, CLaMP, MuseCoco, GETMusic, Museformer and MusicAgent, each in its own folder.

### Does Muzic come with a music player or downloadable songs?

No. The repository contains research code and per-model folders, and the README points to ai-muzic.github.io for music samples generated by the systems. There is no player, no download service and no packaged audio in the repository.

### Which operating system and Python version does Muzic require?

The README states the operating system is Linux and that the project is tested on Ubuntu 16.04.6 LTS with CUDA 10 and Python 3.6.12. The requirements are listed in requirements.txt and installed with pip install -r requirements.txt.

## Sources

- [Issues](https://github.com/microsoft/muzic/issues)
- [License: MIT](https://github.com/microsoft/muzic/blob/main/LICENSE)
- [microsoft/muzic on GitHub](https://github.com/microsoft/muzic)
- [README](https://github.com/microsoft/muzic/blob/main/README.md)
- [Releases](https://github.com/microsoft/muzic/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/microsoft-muzic
