Open-source project
danielravina/stemkit avatar
danielravina/stemkit

StemKit: a local YouTube stem splitter in an Electron app

Split any YouTube song into stems

918 stars114 forksTypeScriptMIT

At a glance

What is it?
StemKit is a desktop app from danielravina that downloads a YouTube song, splits it into vocals, drums, bass, guitar and piano, and plays the stems back in sync. The trade-off is a roughly 2 GB first-launch setup and a pipeline that depends on yt-dlp staying ahead of YouTube.
Who is it for?
Adopt StemKit if you want a local, account-free way to pull vocals or an instrumental out of a YouTube link and you can absorb a roughly 2 GB first-launch download plus the occasional yt-dlp update. Skip it if you need a headless or scriptable pipeline, batch processing, or a service that keeps working unattended, because the separation engine and the downloader are both moving parts the app manages inside its own UI.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly TypeScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What StemKit solves, and who it is actually for

Say you want the vocal line out of a song, or a backing track to sing over. The usual route is a web service: upload the file, wait, download a zip, and accept that someone else's server has your audio. StemKit takes the opposite route. It is an Electron desktop app that searches YouTube or accepts a pasted link, downloads the audio, runs source separation on your own machine, and then plays the result back as a small multitrack player with a fader per stem. The README lists vocals, drums, bass, guitar, piano and more, with one-click presets for All, Karaoke, Acapella and Drums + Bass.

The audience is narrower than the feature list suggests. This is a personal-use tool, and the README says so plainly: downloading audio from YouTube violates their terms of service for public products, with the note to keep it personal. If you are building something that other people use, the download step is the part that disqualifies StemKit, not the separation. For a musician practising at home, a hobbyist making karaoke tracks, or someone who wants to hear how a mix is put together, the fit is direct. There are no accounts and no API keys, and the app sends one anonymous ping per install containing a random install id, version, OS and architecture, which the README documents along with instructions for removing the telemetry source if you would rather send nothing.

The pipeline: yt-dlp, ffmpeg, and two separation models

The README's diagram is worth reading closely because it shows that StemKit is not one model doing everything. A YouTube URL goes through yt-dlp, which the README notes runs with a JS runtime, and then through a bundled ffmpeg. From there the audio branches: a mel-band roformer handles vocals, while demucs htdemucs handles drums, bass and other, with shift averaging. The two outputs are written as stems/*.wav.

The split matters. Vocals get the heavier, dedicated model, and the remaining instruments come from demucs. That is why the optional Studio-quality vocals upgrade (Mel-Band Roformer, +913 MB) is described as improving vocals specifically rather than the whole mix. The other optional downloads change accuracy or cost rather than routing: fine-tuned demucs (htdemucs_ft) is +~320 MB and up to 4x slower, and refinement passes use 2 shifts instead of 1 and are up to 3x slower.

The playback side is a separate design from the separation side. The renderer is React, the main process owns the pipeline and the library, and they talk over IPC. The video plays in a muted iframe while the audio comes from Web Audio stem playback, with the README stating the master clock is the audio itself. That choice is what makes per-stem mute, solo and volume work without the video drifting, and it is also why seeking can be described as instant and artifact-free: the audio timeline is the reference, not the video element.

Separation runs on Apple Silicon via MPS, on NVIDIA GPUs via CUDA, or on CPU. The README does not publish per-machine timings, so how long a split takes depends on your hardware and on which optional models you have enabled. The CPU path is documented as slower for the studio-quality vocal model.

Installing StemKit and running a first split

There is no package manager install. The README points to GitHub Releases, with named artifacts per platform: StemKit-x.y.z-mac-arm64.dmg for Apple Silicon macOS, StemKit-Setup-x.y.z.exe or a portable .zip for Windows, and StemKit-x.y.z-linux-x86_64.AppImage or StemKit-x.y.z-linux-amd64.deb for Linux x64. Download the one matching your machine and run it.

First launch is the slow part. The app creates a private Python environment and downloads the separation engine, about 2 GB, one time. If no Python 3.9+ is detected, StemKit downloads its own runtime (python-build-standalone) during that setup. ffmpeg is bundled, so there is nothing else to install.

On macOS the build is signed with a Developer ID but not notarized, so Gatekeeper may refuse the first open. The README gives the one-time fix:

bash
xattr -cr /Applications/StemKit.app

Or use System Settings, then Privacy & Security, then Open Anyway. On Windows, SmartScreen may warn on first run, and the README's instruction is More info, then Run anyway.

Once the app is open, the workflow is: search YouTube or paste a link, pick the instruments you want, and start the split. The README describes splitting as running in parallel in the background with live progress. When it finishes, the video sits on one side and each stem gets its own fader and colour-coded waveform; clicking a waveform seeks. Presets cover the common cases without touching individual faders. Export writes any stem, or all of them, as WAV.

If you want to build from source rather than use an installer, the README gives two commands and notes that Node.js 20+ is required only for that path:

bash
npm install
npm run dev

The README adds that the scripts auto-relaunch with a suitable Node version via nvm or nvm-windows if you are on the wrong one. Building installers yourself needs the ffmpeg fetch script for your OS first (scripts/fetch-ffmpeg.sh on mac and Linux, scripts/fetch-ffmpeg.ps1 on Windows), then one of npm run dist, dist:win, dist:linux or dist:all. The Linux .deb build needs dpkg and fakeroot on the host, and running the AppImage needs FUSE.

Where StemKit breaks: yt-dlp drift and the 2 GB entry fee

The most likely failure is not the separation. It is the download. The README states that yt-dlp breaks occasionally when YouTube changes things, and that the error dialog offers a one-click update which refreshes yt-dlp and the challenge solver together. That is a reasonable mitigation, but it means the app has a dependency that can stop working between your sessions and needs a manual click to recover. If you are splitting a song the night before a rehearsal, that is a real risk.

The second constraint is size and time. The base engine is about 2 GB on first launch, and the optional models add 913 MB for studio-quality vocals or around 320 MB for fine-tuned demucs. Those optional models also cost time: up to 4x slower for htdemucs_ft and up to 3x slower for 2-shift refinement. The README does not say what hardware those multipliers assume, so treat them as relative, not as a schedule.

Platform coverage has gaps worth naming. macOS builds are described as Apple Silicon only, so Intel Macs are not covered by the listed installer. Linux is x64 only. The .deb does not self-update in place: the README says AppImage installs update themselves in-app while .deb installs update by re-downloading.

Finally, StemKit is the wrong tool if you need automation. There is no documented CLI or HTTP API for splitting, and the pipeline lives inside the Electron main process behind an IPC bridge. If your goal is to process a folder of files on a server, or to call separation from a script, this app does not offer that surface, and the README does not document one.

How it compares to Demucs and to hosted splitters

The closest reference point is Demucs itself, the model family StemKit uses for drums, bass and other. Running Demucs directly means a Python environment, a command-line invocation, and you supply the audio. You get scripting, batch runs and headless operation, and you get to choose your own input source. What you do not get is the rest of StemKit: the YouTube search, the bundled ffmpeg, the model download management, the synced multitrack player and the WAV export UI. The difference is not the separation quality, since StemKit is calling the same class of models. It is everything around the model.

Against hosted splitting services, the difference is where your audio goes and what you pay. StemKit runs locally and requires no account or API key, and the README states your songs, searches and audio never leave the machine. Hosted services handle the compute for you and usually work from a browser on any machine, including one that cannot run a 2 GB local model stack. If you are on a low-spec laptop, or you want to split from a phone, the local approach is the wrong trade. If you would rather not upload unreleased material, it is the right one.

One more distinction: StemKit is also a player. The README describes it as working like a mini DAW, with the video on one side and every stem on its own fader, in sync. Demucs gives you files. StemKit gives you files plus a place to listen to them against the original.

Maintenance, updates and what the MIT licence covers

The repository is not archived, and the last push was on 2026-09-17. Releases have been frequent and small: v0.1.19 on 2026-09-06, v0.1.20 on 2026-09-12, v0.1.21 on 2026-09-14. The version in package.json is 0.1.21, and the project is still on a 0.x line, so expect the shape of the app to move.

Upgrade cost depends on how you installed it. The Linux AppImage self-updates in-app, and the README says .deb installs update by re-downloading. The dependency on electron-updater in package.json is consistent with in-app updating being part of the design. Beyond the app itself, two things can need refreshing independently: the separation models, which are one-time downloads behind the gear icon in Settings, and yt-dlp, which the error dialog can update in one click when YouTube changes break downloads.

The licence is MIT, which is permissive and places few obligations on you if you fork or redistribute. That said, the licence covers the code, not the audio you feed it. The README's own note that downloading from YouTube violates their terms of service for public products is the constraint that matters if you plan to ship anything built on this, and it is a question for your own legal advice rather than something the MIT grant resolves.

Editorial conclusion

Adopt StemKit if you want a local, account-free way to pull vocals or an instrumental out of a YouTube link and you can absorb a roughly 2 GB first-launch download plus the occasional yt-dlp update. Skip it if you need a headless or scriptable pipeline, batch processing, or a service that keeps working unattended, because the separation engine and the downloader are both moving parts the app manages inside its own UI. Verify first that your machine meets the stated requirements (macOS 12+ on Apple Silicon, Windows 10/11 x64, or Linux x64), that you have the disk space for the base engine plus any optional model, and that your use of downloaded audio fits the terms the README itself flags.

Frequently asked questions

How do I install and use StemKit?

Download the installer for your platform from GitHub Releases (a .dmg for Apple Silicon macOS, an .exe or .zip for Windows, an AppImage or .deb for Linux x64) and run it. First launch creates a private Python environment and downloads the separation engine, about 2 GB; after that, search YouTube or paste a link, pick instruments, and split.

What is StemKit used for?

It splits a YouTube song into isolated stems such as vocals, drums, bass, guitar and piano, then plays them back in a multitrack player with per-stem mute, solo and volume. The README describes one-click presets for karaoke, acapella and drums plus bass, and WAV export of any stem.

Are STEM kits worth the money?

This does not apply to StemKit. It is a free, MIT-licensed desktop application from danielravina, not a purchasable kit, and the README describes no pricing, subscription or paid tier.

Where can I find free STEM kits?

StemKit itself is free software under the MIT licence, and its installers are published on GitHub Releases. The README does not describe any other distribution channel or paid version.

Official sources

  1. danielravina/stemkit on GitHub
  2. Issues
  3. License: MIT
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/danielravina-stemkit.svg)](https://hysenlabs.com/projects/danielravina-stemkit)