YuE2: an editable score before the song, and what that costs you
YuE2: frontier music generation with symbolic planning, zero-shot covers, and agentic music editing.
At a glance
- What is it?
- YuE2 splits music generation into a symbolic plan and an audio render, so melody and chords can be inspected and rewritten before anything is sung. The trade-off is a 24 GB VRAM floor and a separate environment for transcription.
- Who is it for?
- YuE2 fits teams and individuals with an NVIDIA GPU of at least 24 GB VRAM who want to inspect and edit a composition rather than only prompt for audio, and who are willing to run a second environment for SheetSage2 when covering existing recordings. It does not fit CPU-only machines, GPUs below the stated memory floor, or anyone who wants a hosted service where the model files never land on local disk.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 1 day ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem YuE2 solves: a song you cannot edit after you hear it
Text-to-audio music models hand you a waveform. If the chorus is wrong, you change the prompt and roll the dice again. YuE2's README frames the project around the opposite: "White-box music generation through symbolic planning." Melody and chords become explicit controls that a person or an agent can inspect and edit before rendering.
The intended user is not someone who wants a background loop for a video. It is someone who already has lyrics and a style in mind, and wants to see the composition that results, change a chord, then render again. The README's three headline workflows are creation, zero-shot covers, and agentic editing, and all three differ only in where the score comes from: YuE2 itself, a transcribed recording, or an edited composition.
How the AR-NAR backbone turns a style prompt into a score and then audio
The architecture is described as one AR-NAR Mixture-of-Transformers backbone. It predicts the score and semantic tokens autoregressively, then generates acoustic latents with flow matching. A VAE decodes those latents into stereo audio, at 48 kHz and without quantization according to the README.
The staged Python API exposes the same pipeline as four named steps: `plan()` then `generate_semantic()` then `synthesize()` then `decode()`. That naming matters more than it looks. Because planning is a separate stage, an exact plan can be reused, and the generation guide in docs/generation.md is where the project says decoder selection and plan reuse are documented.
The `cot` setting decides how much of the score is fixed up front. `cot="full"` generates an editable melody-and-chord plan and is the default for new songs. `cot="melody"` uses a melody plan with free accompaniment and is what the README recommends for covers. `cot="off"` skips planning and generates directly from lyrics and style. You can also pass `abc=...` to supply your own score in `full` or `melody` mode. Those four options are the real control surface of the project.
Installing YuE2 and generating a first song
The stated requirements are Linux, Python 3.12, and an NVIDIA GPU with BF16 support and 24 GB VRAM. Model files download from Hugging Face on first use, so the first run needs network access and enough disk for the checkpoint.
Clone the repository, create the virtual environment, install the package, and run the bundled example. The example writes to a directory you name on the command line.
git clone https://github.com/multimodal-art-projection/YuE.git
cd YuE
python3.12 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install .
python examples/generate.py --output outputs/first-songThe README says to open `outputs/first-song/audio.flac` when it finishes. The same directory also retains the score, semantic tokens, acoustic latents, generation settings, and model identities, which is the part worth checking first: if the score file is not there, the run did not exercise the planning path you came for.
The Python interface is short enough to paste into a script. It reads a request from examples/song.json, runs the pipeline on CUDA, saves artifacts, and prints whether the output was truncated.
import json
from pathlib import Path
from yue2 import YuE2Pipeline
request = json.loads(Path("examples/song.json").read_text(encoding="utf-8"))
with YuE2Pipeline.from_pretrained("m-a-p/YuE2-3B", device="cuda") as pipe:
song = pipe(**request)
song.save_artifacts("outputs/my-song")
print(song.truncated)The package installs as `yue2-infer` and also registers a console script named `yue2`, so a shell entry point exists if you prefer it to Python calls. An optional `fast` extra pulls in vllm and triton; the README does not present it as required for the quick start.
Covering an existing recording, and why it needs a second environment
A cover starts from a transcription, not from lyrics. The README points at SheetSage2 on Hugging Face to transcribe a source recording, then asks you to review the resulting melody ABC before generating. For covers it gives a specific instruction: use `cot="melody"` and a score without chord symbols, so the accompaniment can adapt to the new style.
The call takes a style string, lyrics read from a file, the ABC score read from a file, the `cot` mode, and a seed.
from pathlib import Path
from yue2 import YuE2Pipeline
with YuE2Pipeline.from_pretrained("m-a-p/YuE2-3B", device="cuda") as pipe:
cover = pipe(
style="English, jazz-funk, warm lead vocal, Rhodes, bass and drums",
lyrics=Path("cover-lyrics.txt").read_text(encoding="utf-8"),
abc=Path("cover-score/score.abc").read_text(encoding="utf-8"),
cot="melody",
seed=42,
)
cover.save_artifacts("outputs/cover")The constraint the README states plainly is that SheetSage2 runs in a separate environment and loads its MERT2 encoder automatically. That is an operational cost, not a footnote: the cover workflow is two installs, two dependency sets, and a manual review step between them. If you only want score-conditioned generation without transcribing anything, the repository ships examples/melody.abc so you can try it immediately.
The 24 GB floor and the other places YuE2 is the wrong tool
The hardware requirement is the first real limitation, and it is stated as a requirement rather than a suggestion: NVIDIA GPU, BF16 support, 24 GB VRAM. There is no CPU path in the quick start and no quantized variant mentioned. If your machine does not clear that bar, the project does not have an answer for you.
The second limitation is environmental. The quick start targets Linux, and the README does not document Windows or macOS installation. Nothing in the repository layout suggests a platform-agnostic build.
The third is the cover path itself. Transcription is not part of the main package. You install SheetSage2 separately, and the README asks you to review the melody ABC by hand before generating. That review step is a feature if you care about the result and a burden if you expected one command from recording to finished cover.
Finally, editing is not free-form. The edit workflow exports a plan, you revise the musical details, and you render the edited score. That means your revisions have to be expressible in the score format. If your mental model of the fix is "make the vocal warmer," the symbolic layer may not be where that change lives.
YuE2 against YuE v1 and against direct text-to-audio models
The most direct alternative is YuE itself. The README opens by pointing at the YuE-v1 branch, where the original code, documentation, and license are preserved. So the honest comparison is version against version, not project against unrelated tool. YuE2 is described as unifying symbolic and audio generation; v1 is the earlier line, kept intact on its own branch. If you have a working v1 setup and no need for symbolic planning, the branch still exists and its license is preserved there.
Against a direct text-to-audio model, the difference is where the intermediate representation sits. A direct model goes from prompt to waveform with no artifact in between that a person or an agent can read. YuE2 stops at a melody-and-chord plan first, which is what makes the agentic editing demo possible: the README describes a demo that follows one song through 9 steps and 14 versions, with the conversation, score, prompt, and lyrics inspectable at each version. You cannot build that on top of a model that never emits a score.
The cost of that advantage is the staged pipeline and its memory floor. A direct model may run on smaller hardware precisely because it has fewer stages.
Licence, model weights, and what upgrading looks like
The repository is Apache-2.0, and pyproject.toml declares `license = "Apache-2.0"` with license files spanning LICENSE, MODEL_LICENSE, THIRD_PARTY_NOTICES.md, and everything under licenses/. The presence of a separate MODEL_LICENSE file is the detail that matters. The code licence and the terms attached to the m-a-p/YuE2-3B weights are listed as different files, so do not assume the Apache-2.0 grant on the source covers the checkpoint. Read MODEL_LICENSE and THIRD_PARTY_NOTICES.md before you ship anything generated, and get your own legal read rather than treating this paragraph as advice.
Upgrade cost is mostly pinned dependencies. The project pins exact versions rather than ranges: torch 2.10.0, transformers 4.57.6, huggingface-hub 0.36.2, safetensors 0.7.0, tiktoken 0.12.0, numpy 2.2.6, soundfile 0.13.1, accelerate 1.13.0. That is predictable but rigid. If another tool in your environment needs a different transformers or numpy, you are resolving a conflict, not adjusting a floor. The optional `fast` extra pins vllm 0.19.0 and triton 3.6.0, and the test extra pins pytest 9.0.3.
On release cadence, the repository is not archived and the last push was on 2026-09-10. The most recent release listed is yue2-v0.1.6 from 2026-09-09, and the README links a wheel archive for that tag. The version in pyproject.toml is 0.1.6, so the package and the release line up. Everything is still 0.1.x, which is worth weighing if you need stability guarantees.
Editorial conclusion
YuE2 fits teams and individuals with an NVIDIA GPU of at least 24 GB VRAM who want to inspect and edit a composition rather than only prompt for audio, and who are willing to run a second environment for SheetSage2 when covering existing recordings. It does not fit CPU-only machines, GPUs below the stated memory floor, or anyone who wants a hosted service where the model files never land on local disk. Before committing, verify three things: that your card supports BF16 and has the memory the README requires, that your use of the m-a-p/YuE2-3B weights complies with MODEL_LICENSE rather than assuming Apache-2.0 covers them, and that the example run in examples/generate.py produces a score you can actually open and edit, because that artifact is the whole reason to pick this project over a direct text-to-audio model.
Frequently asked questions
What is the best local AI for music generation?
That depends on your hardware, and YuE2 sets a high floor: the README requires Linux, Python 3.12, and an NVIDIA GPU with BF16 support and 24 GB VRAM. If you clear that, YuE2 offers something most local generators do not, an editable melody-and-chord plan before the audio render. If you do not clear it, the README documents no CPU or reduced-memory path.
Which open source model is considered the best for generating music?
The README does not rank open source models against each other. It reports that YuE2 best-of-8 reaches 6.9632 SongBench Avg on WildSongBench and that this is the highest observed mean among the settings the project evaluated, with scores and protocol in docs/benchmarks.md. Treat that as the project's own evaluation, not a general ranking.
What AI music generator is available on GitHub?
YuE2 is on GitHub under multimodal-art-projection/YuE, installable with pip from a clone of the repository, and it also publishes weights on Hugging Face as m-a-p/YuE2-3B. The README points to SheetSage2 and MERT2 on Hugging Face for the cover workflow.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/multimodal-art-projection-yue)