ACE-Step 1.5: a local music generation model that fits under 4GB of VRAM
The most powerful local music generation model that outperforms almost all commercial alternatives, supporting Mac, AMD, Intel, and CUDA devices.
At a glance
- What is it?
- ACE-Step 1.5 pairs a language-model planner with a diffusion transformer to generate songs locally on CUDA, ROCm, Intel XPU, Apple Silicon or CPU. It is MIT licensed, ships a Gradio UI and a REST API, and its own README claims quality between Suno v4.5 and Suno v5.
- Who is it for?
- Adopt ACE-Step 1.5 if you have a CUDA, ROCm, Intel XPU or Apple Silicon machine and you want generated audio to stay on your own disk, with the Gradio UI at port 7860 as the fastest way in. Do not adopt it if you need the 4B XL decoder on a machine below 12GB of VRAM, or if you expect a documented uninstall path: the README covers installation and launch scripts but not removal.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 27 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What ACE-Step 1.5 is for, and who ends up using it
ACE-Step 1.5 is a local music generation model. The README describes it as a music foundation model that runs on consumer hardware, with a stated floor of less than 4GB of VRAM for the base model. The audience it names is music artists, producers and content creators, but the repository layout suggests a second audience: people who want an HTTP endpoint rather than a UI. There is a cli.py at the top level, a run_api_server.sh, a start_api_server.sh and a Dockerfile, so the project can be driven from a script or a service as well as from a browser.
The practical reason to run it locally is control. Generation happens on your own GPU, the checkpoints live in ./checkpoints, and the outputs land in ./gradio_outputs when the Docker Compose file is used. Nothing in the described flow requires uploading audio to a hosted service. The README also points to acemusic.ai as a free hosted alternative, which is a fair signal that the maintainers know the local install is the harder path.
The feature list is broad rather than deep in any single direction: text-to-music, reference audio input, cover generation, repainting and local editing, track separation, multi-track layering, vocal-to-BGM conversion, metadata control for BPM and key, LRC timestamp generation, quality scoring and LoRA training. If you only need short instrumental loops, most of that surface is irrelevant to you.
The LM planner and DiT decoder split
The architecture described in the README is a two-stage pipeline. A language model acts as a planner: it takes a short user query and expands it into a song blueprint, synthesizing metadata, lyrics and captions. That blueprint then conditions a Diffusion Transformer (DiT), which produces the audio. The README says the LM can scale the output from short loops up to 10-minute compositions, and that the alignment between planner and decoder is trained with intrinsic reinforcement learning rather than an external reward model or human preference data.
That last claim is the interesting part and also the least verifiable from the repository alone. The README links a technical report on arXiv, which is where the training method would actually be documented. Nothing in the file listing or the dependency set lets you confirm how the intrinsic reward is computed.
The model zoo splits along two axes. The base line is acestep-v15-turbo, which is the default in docker-compose.yml via ACESTEP_CONFIG_PATH. The XL series, released 2026-04-02, uses a 4B-parameter DiT decoder and comes in xl-base, xl-sft and xl-turbo variants. The README states the XL models require at least 12GB of VRAM with offload, and recommends 20GB. The LM side is separate: ACESTEP_LM_MODEL_PATH defaults to acestep-5Hz-lm-4B, and the README says all LM models are compatible with the XL series. So the VRAM question is really two questions, one for the decoder and one for the planner.
Installing ACE-Step 1.5 locally with uv
The README's Quick Start requires Python 3.11 to 3.12, which matches requires-python = ">=3.11,<3.13" in pyproject.toml. A CUDA GPU is recommended, and the project also supports MPS, ROCm, Intel XPU and CPU. The first step is installing uv, the package manager the README uses for the local path:
curl -LsSf https://astral.sh/uv/install.sh | sh # macOS / Linux
# powershell -ExecutionPolicy ByPass -c "irm https://astral.sh/uv/install.ps1 | iex" # WindowsThere are also install_uv.sh and install_uv.bat at the repository root if you prefer a checked-in script over a curl pipe. Note the ROCm caveat the README gives: ROCm on Windows requires Python 3.12, because AMD ships Python 3.12 wheels only. Separate requirement files exist for the non-CUDA paths: requirements-rocm.txt, requirements-rocm-linux.txt and requirements-xpu.txt, plus a README-XPU.md and a setup_xpu.bat.
Once the environment resolves, the launch scripts are the intended entry point. The repository ships start_gradio_ui.sh, start_gradio_ui.bat, start_gradio_ui_macos.sh and start_gradio_ui_rocm.bat, alongside manual variants such as start_gradio_ui_manual.sh. For the API instead of the UI, start_api_server.sh and start_api_server_rocm.sh exist. The Docker Compose path is the least ambiguous:
docker compose up # Start Gradio UI (default)
docker compose up -d # Run in background
docker compose logs -f # View logs
ACESTEP_MODE=api docker compose upThe Compose file maps GRADIO_PORT (default 7860) to 7860 and API_PORT (default 8001) to 8001, and mounts ./checkpoints, ./gradio_outputs and a named hf_cache volume. Two environment variables decide what actually loads: ACESTEP_CONFIG_PATH, default acestep-v15-turbo, and ACESTEP_LM_MODEL_PATH, default acestep-5Hz-lm-4B. If you want the XL decoder, that is the variable you change. There is also ACESTEP_LLM_BACKEND, defaulting to pt, and ACESTEP_INIT_SERVICE, defaulting to true. A .env.example sits at the repository root and the Compose file marks the .env file as optional via required: false.
A first generation and where the output goes
The repository ships two example directories, examples/simple_mode/ and examples/text2music/, which is the clearest starting point for a first run. Simple Mode is described in the feature table as generating full songs from simple descriptions, and Query Rewriting is the LM expanding your tags and lyrics automatically. So the first test does not need a long prompt: a short description of style and mood should be enough for the planner to fill in lyrics, captions and structure.
For a scripted run, run_generate_test.py and generate_examples.py at the repository root are the files to read before writing your own loop. There is also quick_test.sh and quick_test.bat for a smoke test, and profile_inference.py if you want to see where time goes. The README's performance claim is under 2 seconds per full song on an A100 and under 10 seconds on an RTX 3090, with a range of 0.5s to 10s on A100 depending on think mode and diffusion steps. Those numbers come from the project's own README, not from independent measurement.
Two controls are worth knowing early because they change what you get. Metadata Control lets you set duration, BPM, key or scale and time signature instead of letting the planner guess. Batch Generation produces up to 8 songs at once, which is useful when you are exploring a prompt rather than finishing one track. If you are running the API server, the endpoint is on port 8001; if you are in the Gradio UI, it is 7860. Files written under Docker land in ./gradio_outputs on the host.
The dependency pins are where installs break
The most concrete warning in the repository is a comment in requirements.txt about torchao. It says the pin must match pyproject.toml, because an unpinned torchao resolves to 0.18.x, which imports torch.nn.functional.ScalingType, a symbol that does not exist in the torch 2.7.1+cu128 build pinned for Windows. The result is a model load failure with "cannot import name 'ScalingType'". The project pins torchao>=0.16.0,<0.17.0 in pyproject.toml for exactly this reason.
That comment tells you something useful about the project's support matrix. The torch pins are platform-specific and not uniform: Windows gets torch 2.7.1+cu128, Linux x86_64 gets 2.10.0+cu128, Linux aarch64 (the comment names NVIDIA DGX Spark) gets 2.10.0+cu130, and macOS arm64 gets torch>=2.9.1 with no CUDA suffix. transformers is bounded on both ends at >=4.51.0,<4.58.0, and gradio is pinned exactly at 6.2.0. If you already have a working PyTorch environment for another project, do not assume ACE-Step 1.5 will slot into it.
A second constraint is torchcodec>=0.9.1, which is excluded on aarch64 via platform_machine != 'aarch64'. That means the aarch64 path, including the DGX Spark case, is a genuinely different dependency set rather than a variation on the same one.
Where ACE-Step 1.5 is the wrong choice
The XL decoder is the clearest boundary. The README states it requires at least 12GB of VRAM with offload and recommends 20GB. If your GPU sits below that, the XL variants are out and you are on the turbo base model. The README does not describe what quality difference to expect between turbo and XL beyond the general claim that XL targets higher audio quality, so the decision is partly a hardware decision and partly an unknown.
LoRA training has its own floor. The feature table states one-click annotation and training in Gradio, with 8 songs taking about 1 hour on a 3090 with 12GB of VRAM. That is a specific, useful number, and it also means training is not a laptop activity on integrated graphics. The README does not document how to evaluate a trained LoRA, how to merge it, or how to remove it once it is registered.
The README also does not document uninstall or rollback. There is a check_update.sh and a check_update.bat for checking updates, but no removal script, and no migration notes between model versions. If you need a clean, reversible install, plan for that yourself.
Finally, the quality claim. The README says output quality sits between Suno v4.5 and Suno v5. That is the project's own framing, and the benchmark section is where the supporting numbers would live. Treat it as a claim to test with your own prompts rather than a settled result.
ACE-Step 1.5 versus Suno and the hosted route
The obvious alternative is not another open source model but Suno, which the README itself uses as the comparison point. The difference is not only quality. Suno is a hosted service: you send a prompt and get audio back, with no GPU, no checkpoint download and no dependency resolution. ACE-Step 1.5 runs the model on your machine, which means the 4GB VRAM floor, the platform-specific torch pins and the checkpoint downloads are yours to manage.
What you get in exchange is that the pipeline is inspectable and scriptable. You can set BPM, key and time signature directly, run batch generation of up to 8 songs, drive the REST API on port 8001 from your own code, and train a LoRA from a few songs to capture a style. Suno Studio's "Add Layer" feature is explicitly named in the ACE-Step feature table as the analogue for multi-track generation, which is a useful signal about where the project sees itself.
If you want the model without the install, the README points to acemusic.ai and to a Hugging Face Space demo. Both are the same model behind someone else's hardware. The trade is the usual one: less control, no local files, and a dependency on a service you do not run. The MIT licence applies to the code you install, not to whatever the hosted endpoints do with your prompts.
For a middle path, the Docker Compose file is the closest thing to a reproducible deployment. It pins the image to ghcr.io/ace-step/ace-step-1.5:latest, requests all NVIDIA GPUs, and keeps checkpoints and outputs on mounted volumes, so a rebuild does not discard your models.
Editorial conclusion
Adopt ACE-Step 1.5 if you have a CUDA, ROCm, Intel XPU or Apple Silicon machine and you want generated audio to stay on your own disk, with the Gradio UI at port 7860 as the fastest way in. Do not adopt it if you need the 4B XL decoder on a machine below 12GB of VRAM, or if you expect a documented uninstall path: the README covers installation and launch scripts but not removal. Before committing, check the pinned torch and torchao versions in requirements.txt against your own driver, and confirm which checkpoint acestep-v15-turbo resolves to for your hardware.
Frequently asked questions
What is ACE-Step 1.5?
It is an MIT-licensed open-source music generation model that runs locally on CUDA, ROCm, Intel XPU, Apple Silicon or CPU. A language model expands a short query into a song blueprint, and a Diffusion Transformer decodes that blueprint into audio.
How do I install ACE-Step 1.5 locally?
The README's Quick Start requires Python 3.11 to 3.12 and installs uv first, then uses the shipped launch scripts such as start_gradio_ui.sh. Docker Compose is the alternative: docker compose up starts the Gradio UI and docker compose up -d runs it in the background.
What are the key differences between ACE-Step 1.5 and XL?
The XL series uses a 4B-parameter DiT decoder and comes in xl-base, xl-sft and xl-turbo variants, released 2026-04-02. The README states XL requires at least 12GB of VRAM with offload and recommends 20GB, while the base model is documented as running under 4GB of VRAM.
How do I install ACE-Step 1.5 on a Mac?
The Quick Start lists MPS support alongside CUDA, ROCm, Intel XPU and CPU, and pyproject.toml pins torch>=2.9.1 for darwin arm64. The repository ships start_gradio_ui_macos.sh and start_gradio_ui_macos_manual.sh for that path.
How do I train a LoRA with ACE-Step 1.5?
The feature table describes one-click annotation and training in Gradio, and states that 8 songs take about 1 hour on a 3090 with 12GB of VRAM. The README does not document how to evaluate or remove a trained LoRA afterwards.
How do I uninstall ACE-Step 1.5?
The README does not document an uninstall or rollback procedure. The repository ships check_update.sh and check_update.bat for checking updates, but no removal script, so cleanup of the environment, checkpoints and outputs is left to the user.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/ace-step-ace-step-1-5)