BS-RoFormer: lucidrains' Band-Split RoFormer for Music Source Separation
Implementation of Band Split Roformer, SOTA Attention network for music source separation out of ByteDance AI Labs
At a glance
- What is it?
- A PyTorch implementation of ByteDance's band-split, rotary-position transformer for separating music into stems. It is a model library, not a finished stem-splitting app, and the README is explicit that you train it yourself.
- Who is it for?
- Adopt BS-RoFormer if you are a PyTorch engineer who wants the architecture in your own training loop, or if you plan to run the replicated weights published by ZFTurbo and Kimberley Jensen. Do not adopt it if you want a desktop app that turns a song into stems in one click, because the README documents no inference checkpoint, no CLI and no pretrained download; that path runs through Ultimate Vocal Remover or the Music-Source-Separation-Training repository instead.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 108 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What BS-RoFormer actually is, and who should reach for it
BS-RoFormer is a PyTorch implementation of the Band Split RoFormer described in the paper Music Source Separation with Band-Split RoPE Transformer from ByteDance AI Labs. The README states the technique uses axial attention across frequency, hence multi-band, and time, and that rotary positional encoding gave a large improvement over learned absolute positions in the authors' experiments. The repository is the model definition plus training utilities, not a product. There is no command line entry point, no bundled checkpoint and no inference script in the layout: the top level holds bs_roformer/, tests/, pyproject.toml, LICENSE and README.md.
The person this fits is an engineer who already has a training loop. You get a module that returns a loss when you hand it a waveform and a target, and you decide the data pipeline, the optimiser, the checkpointing and the inference wrapper. The README's own usage example ends with the comment "after much training" before calling the model on input, which is the honest summary of what this package does for you. If you want separated stems today, this is the wrong entry point, and the README points elsewhere: it credits ZFTurbo with replicating the paper and open sourcing training code and weights, and Kimberley Jensen with a MelBand Roformer trained on vocals.
Band splitting, axial attention and the three model variants
The architecture splits the input waveform into frequency bands and runs attention separately along the frequency axis and the time axis rather than attending over the full spectrogram at once. That is what keeps the cost tractable for long audio, and it is the reason the band configuration is a hyperparameter worth getting right. The README credits @chenht2010 and ZFTurbo with working out the default band splitting hyperparameter, and the Todo list shows that n_fft and the band split plus mask estimation modules were reviewed after the first release.
Three classes are exported. BSRoformer is the base model from the 2023 paper. MelBandRoformer is the follow-up from the Mel-Band RoFormer paper, imported from the same package. FlowBSRoformer is a flow-matching variant that, in the README's words, replaces masking with predicting the flow between pure noise and the target audio, and it is the only one of the three whose inference call is model.sample(x) rather than model(x).
Stereo training and multiple stems are supported, and the README credits contributors with fixing stereo and multi-stem issues in the Mel-Band variant. Two flags appear in every example: use_pope, described as a successor to rotary embeddings, and the separate time_transformer_depth and freq_transformer_depth settings that control how deep each axial branch is. The defaults in the README examples pair a large dim with depth 12 for the base model and a very small dim with depth 1 for the Mel-Band example, so do not read those snippets as recommended training configurations.
Installing BS-RoFormer and running a first training step
The README gives a single install command. It pulls the package from PyPI under the name BS-RoFormer, and the dependency list in pyproject.toml requires torch>=2.0 along with einops, einx, librosa, rotary-embedding-torch, PoPE-pytorch, hyper-connections and torch-einops-utils.
pip install BS-RoFormerThe first real use is the README's own example. It builds a BSRoformer, feeds a batch of two random waveforms of 352800 samples, computes the loss against a target and backpropagates. After that loop converges, calling the model on input alone returns the separated output. Note that the example passes use_pope=False explicitly, so the successor embedding is opt-in.
import torch
from bs_roformer import BSRoformer
model = BSRoformer(
dim = 512,
depth = 12,
time_transformer_depth = 1,
freq_transformer_depth = 1,
use_pope = False
)
x = torch.randn(2, 352800)
target = torch.randn(2, 352800)
loss = model(x, target = target)
loss.backward()Swapping the import to MelBandRoformer changes nothing else about the call shape, and FlowBSRoformer takes the same arguments minus use_pope, with model.sample(x) as the generation call. If you want to check the package against its own test suite, pyproject.toml configures pytest with pythonpath set to the repository root and lists pytest under the test extra.
Where BS-RoFormer stops and you have to start
The largest limitation is stated by omission. The README documents no pretrained weights, no inference CLI and no audio loading path, so a fresh install gives you an untrained network and a loss function. Everything between a folder of songs and a folder of stems is yours to write. The repository is a library, and the README's Todo list, all three items checked, is about internal modules rather than user-facing tooling.
The second constraint is hardware and data. This is a transformer over audio at 44.1 kHz with a default dim of 512 and depth 12 in the base example, and training a music source separator to a useful quality needs far more than the two random tensors in the snippet. The README does not publish training times, dataset requirements or memory figures, so anyone planning a run should treat those as unknown until measured on their own hardware.
The third is version drift. pyproject.toml declares version 1.2.4 and requires-python >= 3.6, while the most recent release listed is 1.1.0 from 2026-02-01. The classifier still reads Development Status :: 4 - Beta. The last push to the repository was on 2026-06-14. None of that is disqualifying for a research library, but it means you should pin a version rather than track the branch if you are embedding this in a pipeline.
BS-RoFormer versus Mel-Band RoFormer and versus Demucs
The comparison the README invites is internal. MelBandRoformer implements the follow-up paper, also from the same lab, and the repository ships both. The difference is in how the frequency axis is divided: the base model uses the band-split scheme from the first paper, while the Mel-Band variant partitions according to a mel scale. The README does not benchmark the two against each other, so the honest position is that the choice depends on which published weights you intend to run. In practice the ecosystem answers this: ZFTurbo's repository hosts weights for the vocal models, and Kimberley Jensen's Mel-Band Roformer vocal model is a separate open-source release, so the variant you pick is usually the variant someone has already trained.
Against Demucs, the difference is architectural. Demucs is a waveform-domain convolutional model; BS-RoFormer is a spectrogram-domain transformer with axial attention and rotary position encoding. The README's claim is about the paper's result, that the authors beat the previous first place by a large margin, not about this implementation beating any specific tool. If you need a model that runs without a training step, Demucs and the Ultimate Vocal Remover ecosystem are the practical route, and the search interest in bs roformer uvr and bs roformer ultimate vocal remover reflects that people reach the architecture through that app rather than through the pip package.
Licence and the cost of keeping up
The repository is MIT licensed, declared both in the LICENSE file at the top level and in pyproject.toml under license = { text = "MIT" }. The practical implication is that you can use, modify and redistribute the code, including in commercial products, provided you keep the copyright notice and the licence text. This is a description of what the licence says, not legal advice; if you are shipping a product, read the LICENSE file and the licences of the dependencies, which include torch and librosa under their own terms.
The upgrade cost is low but not zero. Releases are frequent and versioned, with 1.0.5, 1.0.6 and 1.1.0 landing within about a month of each other earlier in 2026, and pyproject.toml already sits at 1.2.4. Because the package is a model definition, breaking changes tend to appear as new constructor arguments or changed defaults rather than as API removals, and the README's examples show that optional behaviour such as use_pope is exposed as a flag. Pin the version in your requirements, and re-run the pytest suite after any bump. The dependency on torch>=2.0 is the real upgrade constraint: moving this package forward may force a torch upgrade, and that is the change most likely to break a working training environment.
Editorial conclusion
Adopt BS-RoFormer if you are a PyTorch engineer who wants the architecture in your own training loop, or if you plan to run the replicated weights published by ZFTurbo and Kimberley Jensen. Do not adopt it if you want a desktop app that turns a song into stems in one click, because the README documents no inference checkpoint, no CLI and no pretrained download; that path runs through Ultimate Vocal Remover or the Music-Source-Separation-Training repository instead. Before you commit, verify three things: that torch>=2.0 is available in your environment, that your audio hyperparameters match the ones the community settled on for band splitting, and that the loss returned by model(x, target=target) behaves as expected on your own data.
Frequently asked questions
How do I install BS-RoFormer?
The README gives one command: pip install BS-RoFormer. The package requires Python 3.6 or later and torch>=2.0, along with the other dependencies listed in pyproject.toml such as einops, einx and librosa.
How do I use BS-RoFormer?
Import BSRoformer from bs_roformer, construct it with your chosen dim, depth and transformer depths, then pass an input waveform and a target to get a loss you can backpropagate. After training, calling the model on the input alone returns the separated output.
What is BS-RoFormer?
It is a PyTorch implementation of the Band Split RoFormer from ByteDance AI Labs, a music source separation network that uses axial attention across frequency and time with rotary positional encoding. The package also includes the Mel-Band Roformer and a flow-matching variant.
How does BS-RoFormer compare with Mel-Band RoFormer?
Both are shipped in the same package: MelBandRoformer implements the follow-up Mel-Band RoFormer paper, while BSRoformer implements the original band-split scheme. The README does not benchmark them against each other, so the practical difference is which published weights you intend to run.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/lucidrains-bs-roformer)