vocal-remover: the newest release is a 2024 beta and requirements.txt leaves torch commented out
Vocal Remover using Deep Neural Networks
At a glance
- What is it?
- vocal-remover is a Python tool that separates a song into an instrumental and a vocal track with a deep model, run from the command line on CPU or GPU. Its install story is split across two steps, its newest tag is a two year old beta, and two of its runtime dependencies are tooling.
- Who is it for?
- Judgement: vocal-remover is a sound choice if you want a local, scriptable separation step and are willing to install PyTorch yourself, because the inference and training entry points are plain scripts with a documented command line. Two things to check first.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 29 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 9, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The newest release is a 2024 beta while the branch has moved since
The install section sends readers to the releases page for the latest version, and the newest tag there is v6.0.0b4 from 2024-07-23. Before it sit v6.0.0b3 from 2024-07-14 and v5.1.1 from 2024-07-12, which is the most recent release without a beta suffix.
So the latest version offered by that link is a pre-release, and the last non-beta artifact predates it by eleven days. The repository itself has moved on: the default branch is develop, and the last push is dated 2026-09-10, more than two years after the newest tag.
Nothing in the file connects the download to the branch, so a reader who takes the instruction literally gets a 2024 beta binary, while a reader who clones develop gets code that has not been packaged at all. Which of the two you want is the first decision this repository forces on you.
requirements.txt leaves torch commented out and adds ruff at runtime
The install step is two commands:
cd vocal-remover
pip install -r requirements.txtThe dependency file that command reads starts with a comment pointing at the PyTorch install page, and the two framework lines beneath it are commented out:
# install from https://pytorch.org/get-started/locally/
# torch==2.14.0
# torchvision==0.29.0
librosa~=0.10.0So the framework the model runs on is not installed by this file, which matches the separate Install PyTorch section above it. The catch is that the pairing is manual: you choose the PyTorch build yourself, and the exact pins written in the comment are only a suggestion. Note also that torchvision is named there while no documented command uses it.
Everything else in the file is a compatible release constraint, librosa, matplotlib, opencv_python, resampy, numpy and tqdm, so a fresh install takes the newest patch in each line rather than a known good set. Two of those are not runtime libraries at all: ruff is a linter, and it sits in the same list as the audio stack.
pyproject.toml configures a linter and nothing else
The project metadata file holds one setting:
[tool.ruff]
line-length = 99There is no build system table, no project name, no version and no dependency list, so this repository is not packaged or installable in the ordinary way. The documented install is the requirements file and the two scripts it drives, and the metadata file exists purely to tell the linter how long a line may be.
That also explains why ruff appears in requirements.txt. It is a development tool, reachable through npm-style scripts in a Python project that otherwise has no development extras, and the two facts are the same fact.
For anyone expecting a pip installable library with a version, this is not one. The version lives in release tags, and inference.py and train.py are run as scripts from a checkout.
postprocess is the only option the file marks experimental
Two options change the quality of the separation, and they are presented with different weight.
python inference.py --input path/to/an/audio/file --tta --gpu 0The tta option performs Test-Time Augmentation to improve separation quality. The other one is introduced with a warning callout:
python inference.py --input path/to/an/audio/file --postprocess --gpu 0The postprocess option masks the instrumental track based on the vocal volume to improve separation quality, and the warning says it is an experimental feature and to disable it if any problems appear. That is a different contract from tta, which is described as a plain improvement with no caveat attached.
The mechanism also differs. Test-time augmentation runs the model more than once and combines results, while postprocess operates on the finished instrumental track using the vocal volume as a mask. One changes how the model runs, the other changes the output after the fact, which is where an experimental label belongs.
The dataset layout never names a vocals directory
Training takes a dataset directory and two rates:
python train.py --dataset path/to/dataset --mixup_rate 0.5 --reduction_rate 0.5 --gpu 0The expected layout has two subdirectories:
path/to/dataset/
+- instruments/
| +- 01_foo_inst.wav
| +- 02_bar_inst.mp3
| +- ...
+- mixtures/
+- 01_foo_mix.wav
+- 02_bar_mix.mp3
+- ...Three things are fixed by that layout. Pairs are matched by the numeric prefix, since 01_foo appears in both directories. Both directories accept wav and mp3, so a dataset does not have to be converted first. And the vocal stem is never named anywhere: there is no vocals directory, even though inference writes a file called *_Vocals.wav. The target is implied by the mixture minus the instrumental rather than supplied as a third input.
The two rates default to 0.5 in the documented invocation, so the example is also the only guidance on what they should be set to.
augment.py, pseudo.py and appendix/ have no counterpart in the documentation
The top level holds four scripts and four directories: augment.py, pseudo.py, inference.py and train.py, plus lib/, models/ and appendix/.
Two of those scripts have visible counterparts. augment.py is what the tta option and the mixup rate belong to, and train.py has its own section with its own command. inference.py is the third documented entry point.
pseudo.py has no section, no command and no mention anywhere in the file. appendix/ is a directory with no description either. A reader working from the documentation alone cannot tell what either one does, and pseudo-labelled scripts in an audio separation project usually mean label generation, but the file does not say so and it would be wrong to assume.
The lib/ and models/ directories are at least predictable from context, one for shared code and one for network definitions, which is what a train and infer pair needs.
The reference list is a stack of published architectures, not a method of its own
Six references are given, and reading them in order describes the architecture rather than a new contribution. Jansson et al. on singing voice separation with deep U-Net convolutional networks comes first, then Takahashi et al. on multi-scale multi-band DenseNets, then their MMDENSELSTM combining convolutional and recurrent networks, then Choi et al. on phase-aware speech enhancement with a deep complex U-Net, then Jansson et al. on learned complex masks for multi-instrument separation.
The sixth is different in kind. Liutkus et al. on the 2016 Signal Separation Evaluation Campaign is the task definition that the rest of the list is measured against, and it is the only entry with no link attached.
So the honest description of this project is an assembly of published separation architectures with a training script, rather than a new model. That is a strength when you want a known approach, and it is the thing to weigh when a paper is what you needed.
Inference takes one input and names two outputs
The base command takes a single input path:
python inference.py --input path/to/an/audio/fileIt separates that file into two tracks, saved as *_Instruments.wav and *_Vocals.wav. There is no output option in the documented command line, and the file does not say which directory the two files are written to, nor which input formats the loader accepts beyond the mp3 and wav that appear in the training layout.
That is the whole contract for the tool as documented: a path in, two suffixed wav files out, on CPU by default or with a GPU index when one is present. Anything beyond that, batch processing, format conversion or choosing an output location, is not in the file.
Editorial conclusion
Judgement: vocal-remover is a sound choice if you want a local, scriptable separation step and are willing to install PyTorch yourself, because the inference and training entry points are plain scripts with a documented command line. Two things to check first. The release you are offered by the file is a beta tag from 2024-07-23, while the branch has moved since, so packaged binaries and current source are not the same thing. And the dependency file that the install step points at does not install the framework the model needs, since torch is commented out. If you were looking for a windowed app or a web service, this repository is neither.
Frequently asked questions
What is tsurumeso/vocal-remover?
A deep-learning based tool for extracting the instrumental track from songs, written in Python and MIT licensed. It is run from the command line as inference.py, on CPU by default or with a GPU index, and it is trained with train.py against your own dataset.
How do I install the tsurumeso vocal-remover tool?
Download the latest version from the releases page, install PyTorch separately using the instructions at pytorch.org, then run cd vocal-remover followed by pip install -r requirements.txt. The torch and torchvision lines in requirements.txt are commented out, so that file alone will not install the framework the model needs.
Does vocal-remover have a window or an online version?
This repository does not provide one. It is command line scripts, inference.py and train.py, with no graphical interface, installer or web service mentioned anywhere in its documentation. Questions about an app, an APK or an online vocal remover refer to a different product.
What do the tta and postprocess options do in vocal-remover?
The tta option performs Test-Time Augmentation to improve separation quality. The postprocess option masks the instrumental track based on the vocal volume, and it is marked experimental, with a warning to disable it if you run into problems with it.
How do I train my own model with vocal-remover?
Run train.py with a dataset path, a mixup rate, a reduction rate and a GPU index. The dataset needs instruments/ and mixtures/ subdirectories whose files share a numeric prefix such as 01_foo_inst.wav and 01_foo_mix.wav, and both directories accept wav and mp3. The documented invocation sets both rates to 0.5.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/tsurumeso-vocal-remover)