Model or dataset
HeartMuLa/heartlib avatar
HeartMuLa/heartlib

HeartMuLa heartlib: running the open source music generation stack locally

HeartMuLa Official Repo: The Most Powerful Open-Source Music Generation Model of 2026

3,722 stars468 forksPythonApache-2.0

At a glance

What is it?
HeartMuLa's heartlib package bundles the generation model, the HeartCodec audio tokenizer and the transcription model behind one pip install. Here is what the repository documents, what it leaves out, and who should wait.
Who is it for?
Adopt heartlib if you need a locally runnable music generation stack with a permissive licence and you accept an RTF near 1.0, a 3B checkpoint rather than the internal 7B, and a README that documents the happy path only. Do not adopt it if you need streaming inference, reference-audio conditioning, or a 7B model today: the TODOs list all three as unreleased.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 7 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 2, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What heartlib actually is, and who it is for

heartlib is the Python package that ships HeartMuLa's inference code. The repository describes HeartMuLa as a family of music foundation models, and the package is the entry point to four of them: HeartMuLa itself, a music language model that generates music conditioned on lyrics and tags; HeartCodec, described as a 12.5 hz music codec with high reconstruction fidelity; HeartTranscriptor, a whisper-based model tuned for lyrics transcription; and HeartCLAP, an audio-text alignment model for cross-modal retrieval.

The intended user is someone who wants to run music generation on their own hardware rather than through a hosted demo. The README points to online demo spaces on Hugging Face and ModelScope for people who only want to try it, and to a separate MuLaCover repository for cover-song and remix work. heartlib is the local path.

There is a gap worth naming. The README's highlight claims the internal HeartMuLa-7B version reaches comparable performance with Suno on musicality, fidelity and controllability, but that 7B model is not in this repository. The released checkpoint is HeartMuLa-oss-3B-happy-new-year, which the release notes call the best open-sourced model for lyrics controllability and music quality as of 13 Feb. 2026. Everything you can actually download and run here is the 3B line.

How the pipeline is put together: language model plus codec

The architecture visible in the repository separates generation from audio rendering. HeartMuLa produces music from lyrics and tags, and HeartCodec handles the audio side: the release notes state that HeartCodec-oss-encoder paired with HeartCodec-oss-20260123 converts mono or stereo audio into discrete tokens and reconstructs 48 kHz stereo audio. The encoder and decoder checkpoints can be downloaded and loaded separately.

That split is why the download step pulls three artifacts rather than one. HeartMuLaGen is the generation component, HeartMuLa-oss-3B is the language model checkpoint, and HeartCodec-oss is the codec. The example scripts in the examples directory map onto the three capabilities: run_music_generation.py, run_music_reconstruction.py and run_lyrics_transcription.py.

The package metadata reinforces how tightly the stack is pinned. pyproject.toml fixes numpy at 2.0.2, transformers at 4.57.0, tokenizers at 0.22.1, accelerate at 1.12.0 and bitsandbytes at 0.49.0, while torch, torchaudio and torchvision are given ranges (torch>=2.4,<2.11). Exact pins on the tokenizer and transformers pair are a reasonable choice for reproducibility and a poor one for coexistence: if another project in the same environment needs a different transformers version, you are resolving that conflict by hand.

Installing heartlib and running your first generation

The README recommends python=3.10 for local deployment. Clone the repository, install the package in editable mode, then pull the checkpoints. The install is a single command from the repository root.

bash
git clone https://github.com/HeartMuLa/heartlib.git
cd heartlib
pip install -e .

Checkpoint retrieval is documented for both Hugging Face and ModelScope. The README gives the Hugging Face route first, using the hf download command with --local-dir to place each artifact under ./ckpt. Note the directory names: the model id HeartMuLa-oss-3B-happy-new-year is downloaded into a folder called HeartMuLa-oss-3B, not into a folder matching the model id.

bash
hf download --local-dir './ckpt' 'HeartMuLa/HeartMuLaGen'
hf download --local-dir './ckpt/HeartMuLa-oss-3B' 'HeartMuLa/HeartMuLa-oss-3B-happy-new-year'
hf download --local-dir './ckpt/HeartCodec-oss' HeartMuLa/HeartCodec-oss-20260123

If you are on ModelScope, the equivalent commands use modelscope download with --model and --local_dir, and the README shows the same three artifacts. After the downloads finish, the ./ckpt folder should contain HeartCodec-oss and HeartMuLa-oss-3B subfolders alongside the HeartMuLaGen files. The README's directory listing is cut off mid-tree, so treat the three download commands as the specification for the layout rather than the truncated output.

For a first real run, the repository ships examples/run_music_generation.py, which is the generation entry point for lyrics and tags. The README does not print the script's flags or its default output path, so read the file before running it; the checkpoint paths it expects are the ones you just created under ./ckpt.

The RTF near 1.0 problem and other documented limits

The TODOs section is the most honest part of the README, and it should shape your expectations more than the highlight does. Inference acceleration and streaming inference are listed as pending, with the current inference speed stated as roughly RTF 1.0. Real-time factor near 1.0 means generating a minute of audio takes about a minute of compute on the hardware the team measured on. That is workable for offline batch generation and awkward for anything interactive.

Reference audio conditioning and fine-grained controllable music generation are also pending. So is the HeartMuLa-oss-7B release. If your use case depends on conditioning generation on an existing recording, this repository is not the tool yet; MuLaCover is the separate project aimed at that, and it is a different repository with its own weights and generation guide.

The README does not document rollback, checkpoint compatibility between HeartMuLa versions, or what happens when you mix an encoder checkpoint with a mismatched decoder. It also gives no memory or VRAM figures for the 3B checkpoint, which is the first thing you will want to know before downloading several gigabytes of weights. Those are real gaps, not oversights in this article: the repository is silent on them.

heartlib compared with hosted generation services

The obvious alternative is not another library but a hosted service, including HeartMuLa's own demo spaces on Hugging Face and ModelScope. The difference in approach is where the model runs and what you can change. A hosted space gives you a prompt box and returns audio; you cannot swap the codec, adjust the checkpoint directory layout, or run the pipeline inside your own batch job. heartlib gives you the components and the weights, and in exchange you own the environment, the GPU, and the version pinning.

Within the local category, the comparison the README invites is against the internal 7B model, which is not available to download. The released 3B checkpoint is what you get. The release notes frame the 3B happy-new-year version as the best open-sourced model for lyrics controllability and music quality at its release date, which is a claim about the open-source field rather than a measurement you can reproduce from the README alone. The paper and the HeartMuLa-Benchmark dataset on ModelScope are the places to check that claim rather than take it.

One more local route exists: a ComfyUI custom node created by a community contributor, credited in the 20 Jan. 2026 news entry. If your workflow already lives in ComfyUI, that node may be a shorter path than driving the Python examples directly, though it is a third-party repository and not part of heartlib.

Licence terms and what upgrades will cost you

The repository is Apache-2.0, and pyproject.toml declares the same licence for the package. The 20 Jan. 2026 news entry states that the licence of the repository and all related model weights was updated to Apache 2.0. That matters if you plan to ship generated audio or embed the models in a product, because it removes the source-available restrictions that often accompany music generation weights. It does not settle questions about the training data or about rights in the output itself; the repository says nothing about either, and that silence is the thing to resolve with your own counsel rather than from this page.

Upgrade cost is dominated by the pinned dependency set. transformers and tokenizers are pinned to exact versions, so a future HeartMuLa checkpoint that requires a newer transformers will force a coordinated bump of tokenizers, accelerate and possibly bitsandbytes. The torch range (>=2.4,<2.11) gives some room, but numpy at exactly 2.0.2 is a hard pin that will collide with environments still on the 1.x line. Budget for a dedicated virtual environment per project rather than a shared one.

The last push to the default branch was on 2026-09-17, six days before this writing, and the repository is not archived. The release cadence visible in the news entries has been steady through 2026, with the most recent entry on 16 Sep. 2026 announcing MuLaCover. There are no retrieved releases, so version tracking happens through the news list and the checkpoint names rather than through tagged releases.

Editorial conclusion

Adopt heartlib if you need a locally runnable music generation stack with a permissive licence and you accept an RTF near 1.0, a 3B checkpoint rather than the internal 7B, and a README that documents the happy path only. Do not adopt it if you need streaming inference, reference-audio conditioning, or a 7B model today: the TODOs list all three as unreleased. Before committing, verify that the three checkpoint downloads land in ./ckpt/HeartCodec-oss, ./ckpt/HeartMuLa-oss-3B and ./ckpt/HeartMuLaGen, because the README's own directory listing is truncated and the example scripts resolve paths from there.

Frequently asked questions

What is heartlib and how is it different from the HeartMuLa demo?

heartlib is the Python package in the HeartMuLa repository that contains the inference code for the HeartMuLa music language model, HeartCodec, HeartTranscriptor and HeartCLAP. The demo spaces on Hugging Face and ModelScope let you generate music without installing anything, while heartlib runs the models on your own machine with the checkpoints you download.

How do I install HeartMuLa heartlib and download the checkpoints?

The README recommends python=3.10, then git clone of the heartlib repository, cd into it, and pip install -e . . Checkpoints come from three downloads: HeartMuLaGen, HeartMuLa-oss-3B-happy-new-year into ./ckpt/HeartMuLa-oss-3B, and HeartCodec-oss-20260123 into ./ckpt/HeartCodec-oss, using either hf download or modelscope download.

Does heartlib support streaming generation or reference audio conditioning?

No. The TODOs section lists release scripts for inference acceleration and streaming inference as pending, along with reference audio conditioning and fine-grained controllable music generation. The README states the current inference speed is around RTF 1.0, and cover-song or remix work points to the separate MuLaCover repository instead.

Official sources

  1. HeartMuLa/heartlib on GitHub
  2. Issues
  3. License: Apache-2.0
  4. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/heartmula-heartlib.svg)](https://hysenlabs.com/projects/heartmula-heartlib)