AudioMuse-AI: sonic playlists for Navidrome, Jellyfin, LMS, Lyrion, Emby and Plex
AudioMuse-AI uses sonic analysis to rediscover forgotten songs, uncover hidden connections in your music library, and generate intelligent playlists for Navidrome, Jellyfin, LMS, Lyrion, Emby and Plex: no metadata or external services required.
At a glance
- What is it?
- AudioMuse-AI is a self-hosted Python application that analyses the audio itself, not the tags, and turns the result into playlists for six media servers. The hard part is not the concept, it is the first analysis run.
- Who is it for?
- Adopt it if you already run Navidrome, Jellyfin, LMS, Lyrion, Emby or Plex on your own hardware and you are willing to spend disk and CPU on a one-off analysis pass. Skip it if you want metadata-driven tagging, or if you cannot give the worker container a persistent volume and a long uninterrupted run.
- Can I use it commercially?
- Yes, with strict conditions. AGPL-3.0 is a network copyleft licence: if people use a modified version over a network, for example as a hosted service, you must offer them its source code under the same licence.
- Is it still maintained?
- Yes. The repository last received commits 4 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What AudioMuse-AI solves that tags cannot
Most self-hosted music setups organise themselves around metadata. Genre tags, artist fields, years, ratings. That works until the tags are wrong, missing, or describe the release rather than the recording. A track tagged "Electronic" may be ambient drift or peak-time techno, and no tag tells you which.
AudioMuse-AI takes the other route. According to the README, it uses sonic analysis to "rediscover forgotten songs" and generate playlists "without relying on metadata or external APIs". The features follow from that: Clustering groups sonically similar songs into genre-defying playlists, Music Map renders the collection as a 2D map, Playlist from Similar Songs finds tracks sharing a sonic signature, and Song Paths builds a listening journey between two chosen tracks.
The audience is specific. You are running one of Navidrome, Jellyfin, LMS, Lyrion, Emby or Plex on your own hardware, you have a library large enough that you no longer remember what is in it, and you are comfortable with Docker Compose or Kubernetes. If you stream from a commercial service, there is nothing here for you.
Flask, a worker, and an analysis pass that runs once per track
The README names the shape of the deployment: AudioMuse-AI "run[s] Flask and Worker containers to actually run all the feature". The Flask side serves the UI and the API. The worker does the audio analysis. The top-level repository listing matches that split: app_analysis.py, app_clustering.py, app_map.py, app_path.py, app_sync.py and app_music_servers.py sit alongside flask_app.py and a gunicorn.conf.py, with config.py and database.py holding shared state.
Since v3.2.0 the queue lives on PostgreSQL, and the release notes state that Redis is no longer needed. That removes one moving part from a homelab stack and is the single most useful change for anyone who tried an earlier version.
The multi-server behaviour from v3.0.0 is the part worth understanding before you size the machine. You can connect several media servers, any mix of the six, to one deployment. Duplicate detection recognises the same song across servers, so each track is analysed only once and every server shares the result. If you run both Navidrome and Jellyfin over the same files, you pay the analysis cost once.
The model stack is visible in the Dockerfile. A first build stage downloads the ML models into /app/model, and the GPU variant pins onnxruntime-gpu 1.28.0 with CUDA 13 plus cupy-cuda13x and cuml-cu13. The plain build uses ubuntu:24.04 as its base. There is also a Dockerfile-noavx2, which tells you the default image expects a CPU with AVX2.
Installing AudioMuse-AI with Docker Compose
The README points at deployment/ for Compose examples and at the Helm chart repository for Kubernetes. The Dockerfile documents the two build paths directly, and these are the commands it gives:
docker build -t audiomuse-ai .
docker build --build-arg BASE_IMAGE=nvidia/cuda:13.3.1-cudnn-runtime-ubuntu24.04 -t audiomuse-ai-gpu .The first produces a CPU image, the second a GPU image. The build is not quick: an early stage installs wget, ca-certificates and curl with retries, then downloads the models into /app/model. Models are cached in their own stage so later rebuilds skip them.
For ARM64 with a GB10 GPU, the Dockerfile notes there is no official linux/aarch64 onnxruntime-gpu wheel on PyPI or pypi.nvidia.com. That build takes an extra argument instead:
docker build --build-arg BASE_IMAGE=nvidia/cuda:13.3.1-cudnn-runtime-ubuntu24.04 \
--build-arg ONNXRUNTIME_WHEEL_URL=<prebuilt-wheel-url> -t audiomuse-ai-gpu-arm64 .The Dockerfile states that when ONNXRUNTIME_WHEEL_URL is set, BASE_IMAGE must be a CUDA 13.x image. The README describes the resulting -nvidia-arm image as experimental.
Once the containers are up, the workflow the README describes is one step: "start with an initial analysis". Everything else (clustering, Instant Playlists, Music Map, Song Paths, Text Search, Lyrics Search) becomes available after that pass finishes. For a first real use, point the deployment at a single server, let the analysis run, then open the Music Map to see whether the grouping matches how you actually hear your library. If it does not, the analysis parameters are the thing to change, not the server connection.
One warning comes straight from the README: the plugin system "requires a persistent volume mounted on both the Flask and worker containers, otherwise installed plugins are lost whenever the containers restart". Mount it before you install anything.
Where AudioMuse-AI gets expensive or simply does not fit
The analysis pass is the whole product, and it is also the cost. Every track in your library has to be decoded and embedded, and the README treats the initial analysis as a milestone rather than a background detail. On a large library this is a long job, and it competes for CPU with anything else on the same machine.
The GPU path is not a shortcut for everyone. The Dockerfile pins CUDA 13 and cuDNN 9 for GPU builds, and the ARM64 route depends on a prebuilt wheel URL that the repository hosts as a release asset rather than publishing to an index. The README calls the DGX Spark image experimental. If your hardware is an older NVIDIA card, a CPU-only build may be the honest choice.
The AVX2 assumption is easy to miss. A default build on a CPU without AVX2 is the case Dockerfile-noavx2 exists for, and the README does not present that as the normal path.
Lyrics Search has a hard boundary the README states plainly: it works only with the 72 listed languages. A library heavy in a language outside that list gets nothing from that feature, though Clustering, Music Map and the sonic features are unaffected because they never read lyrics.
Finally, this is not a tagger. It will not fix your metadata, and it does not try. If your actual problem is that your albums are mislabelled, AudioMuse-AI is the wrong tool and a metadata editor is the right one.
How it differs from metadata-driven playlist tools
The obvious comparison is with the smart-playlist features built into the servers themselves. Navidrome, Jellyfin and Plex all ship some form of rule-based playlist: filter by genre, year, rating, play count. The difference is the input. Those rules read the database fields you already have. AudioMuse-AI reads the audio, which is why its Clustering feature can produce "genre-defying playlists based on the music's actual sound" rather than grouping by a string that a ripper wrote years ago.
The trade-off is real in both directions. Rule-based playlists update instantly when your library changes and cost nothing to compute. AudioMuse-AI needs the analysis pass first, and the quality of its output depends on that pass being complete. In exchange it can answer questions your tags cannot: what sounds like this track, what sits between these two tracks, what have I been playing that resembles this.
The README also points at a hosted option: Elestio offers AudioMuse-AI as a managed cloud service. That is the pragmatic answer if you want the features without running the worker yourself, and it is worth knowing the project acknowledges the case rather than pretending self-hosting is the only route.
Licence, maintenance and what an upgrade costs
AudioMuse-AI is AGPL-3.0. For a homelab deployment that changes nothing practical: you run it, you do not distribute it. It matters if you embed it in a product or expose a modified version over a network, because the AGPL's network clause reaches further than the GPL. That is a description of the licence, not legal advice; read the LICENSE file and talk to a lawyer if you are building on it commercially.
The repository is not archived, and the last push was on 2026-09-09. Releases arrive often: v3.5.0 on 2026-08-27, v3.5.1 on 2026-08-30, v3.5.2 on 2026-09-04, the last described as a maintenance release. That cadence is good for fixes and awkward for operators, because a Compose file pinned to a tag will drift behind quickly.
The upgrade cost is concentrated in two places. First, the database: v3.2.0 moved the queue onto PostgreSQL and dropped the Redis requirement, so any deployment older than that needs its Compose file rewritten rather than just re-pulled. Second, the model stage: the Dockerfile downloads models into /app/model during build, so an image rebuild can mean a fresh download unless the layer is cached.
There is a backup path in the repository, app_backup.py, so the state worth preserving is not only the PostgreSQL volume. The README does not document rollback procedure, and the docs folder is where deployment questions are meant to be answered.
Editorial conclusion
Adopt it if you already run Navidrome, Jellyfin, LMS, Lyrion, Emby or Plex on your own hardware and you are willing to spend disk and CPU on a one-off analysis pass. Skip it if you want metadata-driven tagging, or if you cannot give the worker container a persistent volume and a long uninterrupted run. Before committing, check docs/PARAMETERS.md for the analysis settings that match your CPU budget, confirm which image tag suits your hardware (the plain one or Dockerfile-noavx2) and verify that your server version is one of the six the README badge lists, since the integration depends on that server's API.
Frequently asked questions
What is AudioMuse-AI?
It is an open source, self-hosted tool that uses sonic analysis to rediscover forgotten songs and generate playlists without relying on metadata or external APIs. It runs Flask and worker containers and integrates with Navidrome, Jellyfin, LMS, Lyrion, Emby and Plex.
How do I use AudioMuse-AI?
According to the README you start with an initial analysis, and that unlocks the rest: Clustering, Instant Playlists, Music Map, Playlist from Similar Songs, Song Paths, Sonic Fingerprint, Song Alchemy, Text Search and Lyrics Search. Point the deployment at a media server first, then let the analysis run.
Is there an alternative to AudioMuse-AI?
The README notes that Elestio offers AudioMuse-AI as a managed cloud service if you prefer not to self-host. For playlist generation without the analysis pass, the rule-based smart playlists built into Navidrome, Jellyfin and Plex work from metadata instead of audio.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/neptunehub-audiomuse-ai)