Self-hosted service
rzru/nightingale avatar
rzru/nightingale

Nightingale (rzru/nightingale): karaoke from your own music library with WhisperX and Demucs

Machine learning powered Karaoke app (with scores!)

1,481 stars108 forksTypeScriptGPL-3.0

At a glance

What is it?
A Tauri desktop app that separates vocals, transcribes word-level lyrics and scores your pitch. Here is how it works, how to run it, and where it falls short.
Who is it for?
Adopt Nightingale if you already keep a music library on a disk or a media server and want karaoke without hunting for instrumental tracks. Skip it if you need a hosted service, a phone-first experience, or a headless pipeline you can script.
Can I use it commercially?
Yes, with conditions. GPL-3.0 is a copyleft licence: if you distribute software that includes it, you must release that software's source code under the same licence. Running it internally without distributing it does not trigger that obligation.
Is it still maintained?
Yes. The repository last received commits 1 day ago.
What is it written in?
Mainly TypeScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What Nightingale actually replaces

Karaoke has always been a content problem before it is a software problem. You either buy tracks that already have the vocals stripped, or you hunt for instrumental versions that may not exist for the song you want. Nightingale takes the position that any song in your library is a candidate, and it does the stripping itself. The README describes the pipeline plainly: scan a folder, Plex, Jellyfin, Navidrome or a self-hosted web library, separate lead vocals from the instrumental with the UVR Karaoke model or Demucs, transcribe lyrics with WhisperX at word level, then play it back with synchronized highlighting and pitch scoring. The audience is narrow and specific. It is for someone who already has a music collection and a microphone, and who is willing to spend disk space and compute time converting that collection into karaoke. It is not for someone who wants to open a website and sing. The README also notes the app ships as a single binary and claims no manual installation of Python, ffmpeg or ML models is required, with everything downloaded and bootstrapped on first launch. That claim is the whole product thesis. If the bootstrap works, the barrier to entry is a folder path and a wait.

The analysis pipeline: UVR, Demucs, WhisperX and the alignment backends

The architecture splits into a Rust workspace and a React client. The root Cargo.toml lists four members: app-core, client/src-tauri, client/src-server and xtask. So the core logic lives in app-core, the Tauri shell wraps it for desktop, and a separate server binary exists for the self-hosted web mode the README mentions. The workspace lints are unusually strict for a hobby-adjacent app: unsafe_code is denied, unwrap_used is denied, todo and unimplemented are denied, and print_stdout and print_stderr are denied. That last pair matters in a GUI app, because it means the codebase is expected to route output through logging rather than console prints. The analysis itself is not one model. Stem separation defaults to the UVR Karaoke model, with Demucs as the alternative, and the README makes a specific claim about the default: the karaoke model preserves backing vocals in the instrumental, which sounds more natural than a hard vocal removal. Transcription defaults to Whisper via WhisperX for word-level timestamps. Two experimental alternatives are documented. Parakeet v3 covers roughly 25 European languages and runs on NeMo with CUDA or ONNX Runtime elsewhere. For alignment, you can keep WhisperX's aligner, switch to GPU forced alignment through torchaudio forced_align on CUDA and Apple Silicon, or use the Qwen aligner (Qwen3-ForcedAligner-0.6B), which the README says timestamps 11 languages including CJK in a single pass on CUDA, MPS or CPU. Both experimental aligners fall back to WhisperX automatically. That fallback design is the sensible part: a failed experiment degrades to the default rather than breaking the song. CJK handling is separate and more involved than the English path, using per-character forced alignment plus romanized readings in Hepburn, pinyin, Jyutping or Revised Romanization.

Installing Nightingale and analyzing a first song

The README points to two distribution routes. Desktop builds come from GitHub Releases, and a Docker image is published as razzaru/nightingale for CPU or CUDA/GPU. The self-hosted web mode is documented under docs/self-hosted, and the Docker route under docs/docker. The README does not give a package-manager command, so the honest starting point is the releases page or the image. For the Docker route, the image name is the only concrete identifier the README provides:

bash
docker pull razzaru/nightingale

What you do with it afterward is in docs/docker, which the README links rather than inlines. For the desktop route, install from the release for your platform and launch it. The README states the first launch downloads ffmpeg, uv, Python, PyTorch and the ML packages, and pre-downloads video backgrounds. During setup you choose the main data folder, and Settings lets you split cache, models, videos and vendor tools into separate folders later. That split is worth using if your system drive is small, because model weights and cached playback variants are the parts that grow. Once the library is pointed at a folder or a server, the README describes an Analyze All action plus optional auto-analysis to queue the library, and a per-song Actions button that opens the lyrics editor with an LRCLIB browser. One detail changes the cost of a song significantly: if you paste timed LRC or Enhanced LRC, Nightingale uses it as-is and can skip stem separation entirely, which means you sing over the original mix and avoid the analysis pass. Plain lyrics, by contrast, go through alignment. If your goal is a fast first session, pasting timed LRC for two or three songs is the shortest path to hearing something.

Where the design costs you: disk, waiting and microphone calibration

The single-binary claim has a hidden bill. Downloading PyTorch and ML packages on first launch is convenient, and it is also a large download over a connection you may not control, with no documented offline or air-gapped path in the README. The same applies to the analysis itself: separating stems and transcribing lyrics takes time per song, and the README does not publish a per-song estimate, so the only way to know your throughput is to run Analyze All on a subset and watch. Storage compounds this. The README lists separate folders for cache, models, videos and vendor tools precisely because these accumulate, and cached playback variants for key and tempo shifts add more. Scoring is the second place where the design assumes something of you. Pitch scoring uses real-time microphone input with pitch detection, and the README includes a beep-based latency test because scoring lines up with your room only if the delay is measured. A USB microphone, a Bluetooth headset and a laptop's built-in mic will produce different offsets, and the README does not claim to detect them automatically. If you skip the latency test, the scoreboard will be wrong in a way that is hard to diagnose, because the app is working as designed and your input path is the variable. Finally, the README marks two features experimental: UltraStar Deluxe song folders, and the Parakeet v3 ASR engine. Experimental in a karaoke app means a song you planned to sing may need to fall back to the default path, which it does automatically for the aligners but which still costs you the analysis time.

How Nightingale differs from UltraStar Deluxe and Ultimate Vocal Remover

The two obvious alternatives take opposite approaches. UltraStar Deluxe is the long-standing karaoke game format, and Nightingale supports its song folders as an experimental import. The difference is where the data comes from. In UltraStar, pitch and lyric data are authored by hand and shipped with the song folder, so the game is only as good as the community's catalog. Nightingale generates that data from your audio instead, which means any track in your library is playable but also that the timing quality depends on the alignment backend rather than on a human who checked it. Ultimate Vocal Remover is the other reference point, and the README names its karaoke model as Nightingale's default separator. UVR is a tool for producing a stem file you then use elsewhere. Nightingale is a player that happens to do separation as a step. If you want a clean instrumental to export and edit, UVR gives you the file and the model choice directly. If you want to sing tonight from the songs you already own, the separation is a means to an end, and Nightingale's value is in the playback, scoring and profiles layered on top. Neither comparison favors Nightingale on audio quality alone, because both use overlapping models. The distinction is the integration.

Licence, updates and the cost of keeping it current

The repository is GPL-3.0-or-later, and the README badge and LICENSE file agree. For personal use that changes nothing. For anyone embedding Nightingale in another product, the copyleft terms apply to the combined work, and the README's own dependency list is worth reading in that light: it links to UVR, Demucs, WhisperX and LRCLIB, each with its own terms. I am not giving legal advice; if you plan to redistribute, read the licences of the models and the upstream projects rather than only this repository's. Maintenance is visible in the release history. v1.0.0 landed on 2026-07-25, v1.1.0 on 2026-08-14, and v1.2.0 on 2026-09-02, with the last push to the repository on 2026-09-05. That is a steady cadence over roughly six weeks, and the CHANGELOG.md at the repository root is where the details live. Upgrades differ by platform, and the README is explicit about the asymmetry: on macOS and Windows the app checks for new releases at launch, badges the sidebar avatar, and downloads and installs signed updates with one click, while on Linux the Update entry opens GitHub Releases for a manual download. If you run the self-hosted or Docker route, in-app updates do not apply at all, and you are pulling a new image or rebuilding. Budget for that: the bootstrap artifacts are the expensive part, and the README does not document whether they are reused across versions or re-downloaded.

Editorial conclusion

Adopt Nightingale if you already keep a music library on a disk or a media server and want karaoke without hunting for instrumental tracks. Skip it if you need a hosted service, a phone-first experience, or a headless pipeline you can script. Before your first party, verify three things on your own machine: that the bootstrap step completes and the models land in the folder you chose, that the beep-based latency test in Settings produces a delay that makes the scoring line up with your microphone, and that your GPU is actually being used if you picked the CUDA Docker image, because the README does not document what happens when it is not.

Frequently asked questions

How do I install Nightingale?

Desktop builds come from GitHub Releases, and a Docker image is published as razzaru/nightingale for CPU or CUDA/GPU. The README links to docs/self-hosted and docs/docker for the server routes rather than inlining the steps.

What sources can Nightingale import songs from?

A plain folder, Plex Media Server, Jellyfin, Navidrome, or a self-hosted web library. Folder libraries also read .m3u, .m3u8 and .pls playlists, and UltraStar Deluxe song folders are supported experimentally.

Do I need to install Python, ffmpeg or the ML models myself?

The README states no manual installation is required: ffmpeg, uv, Python, PyTorch and the ML packages are downloaded automatically during setup. You choose the main data folder, and Settings lets you move cache, models, videos and vendor tools to separate locations.

How does Nightingale score my singing?

It uses real-time microphone input with pitch detection, star ratings and per-song scoreboards, tracked per profile. The README includes a beep-based latency test in Settings so the scoring lines up with your room.

Official sources

  1. License: GPL-3.0
  2. Project website
  3. README
  4. Releases
  5. rzru/nightingale on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/rzru-nightingale.svg)](https://hysenlabs.com/projects/rzru-nightingale)