# MOABB: a benchmark harness for EEG motor imagery and P300 pipelines

> MOABB wraps freely available EEG datasets and scikit-learn pipelines into a reproducible benchmark. It is aimed at BCI researchers who need comparable numbers across datasets, not at people shipping a real-time decoder.

**NeuroTechX/moabb** — Mother of All BCI Benchmarks

- Repository: https://github.com/NeuroTechX/moabb
- Website: https://moabb.neurotechx.com/docs/index.html
- Stars: 1,059 · Forks: 266
- Language: Python
- License: BSD-3-Clause
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/neurotechx-moabb

## What MOABB fixes about BCI reproducibility

The README states the problem plainly: many BCI datasets are freely available, but researchers rarely publish the code that produced their numbers, and preprocessing parameters, toolbox choices and implementation tricks that are almost never reported can shift performance significantly. The result is that a newcomer has no single place to learn which algorithm works best and on which dataset. MOABB's answer is to fix the surrounding machinery (dataset download, subject and session selection, epoch extraction, cross-validation split, scoring) so that the only thing that varies between two runs is the pipeline you plug in. That is a narrower claim than it sounds. MOABB does not certify that a published result is correct; it makes the loading and splitting identical, which removes one large class of irreproducibility. The audience is research groups and students in EEG-based BCI, plus anyone reviewing a motor imagery or P300 claim who wants to re-run it on public data. The README also frames success as a future in which papers report a MOABB score the way they report accuracy on a named dataset.

## Paradigms, evaluations and pipelines: the three moving parts

The architecture visible in the README and the repository layout separates three concerns. A paradigm defines the experimental setting and the filtering window, for example LeftRightImagery(fmin=8, fmax=35), and is responsible for turning raw recordings into epochs with labels. An evaluation defines how those epochs are split and scored; the README's quickstart uses CrossSessionEvaluation, and the examples directory also contains a learning_curve example set, which implies evaluation strategies are pluggable objects rather than one hardcoded loop. A pipeline is a scikit-learn estimator, so anything that follows the fit/predict contract can compete: the quickstart builds one with make_pipeline(LogVariance(), LDA()) and passes a dictionary of named pipelines to evaluation.process(). The dictionary key becomes the pipeline name in the results, which is why the quickstart calls it "LogVar+LDA". Results come back as a pandas DataFrame, so the comparison step is ordinary DataFrame work rather than a bespoke result format. Datasets are objects too: BNCI2014_001() is instantiated and then narrowed with dataset.subject_list = dataset.subject_list[:2], which is the pattern for a quick trial run before committing to a full download.

## Installing MOABB and running a first two-subject benchmark

Installation is a single pip command, as given in the README quickstart. MOABB requires Python 3.11 or newer according to pyproject.toml, and it pulls in numpy>=2.0, scipy>=1.9.3, mne>=1.10.0, pandas>=1.5.2 and h5py>=3.10.0 as dependencies, so an environment with MNE-Python already present will not need much more.

```bash
pip install moabb
```

The README's quickstart is the shortest path to a real result. It sets the log level, defines one pipeline, restricts the dataset to two subjects, and runs a cross-session evaluation. Expect the first run to spend most of its time downloading and caching the recordings for BNCI2014_001; the print(results.head()) at the end shows the per-subject score rows.

```python
import moabb
from moabb.datasets import BNCI2014_001
from moabb.evaluations import CrossSessionEvaluation
from moabb.paradigms import LeftRightImagery
from moabb.pipelines.features import LogVariance

from sklearn.discriminant_analysis import LinearDiscriminantAnalysis as LDA
from sklearn.pipeline import make_pipeline

moabb.set_log_level("info")

pipelines = {"LogVar+LDA": make_pipeline(LogVariance(), LDA())}

dataset = BNCI2014_001()
dataset.subject_list = dataset.subject_list[:2]

paradigm = LeftRightImagery(fmin=8, fmax=35)
evaluation = CrossSessionEvaluation(paradigm=paradigm, datasets=[dataset])
results = evaluation.process(pipelines)

print(results.head())
```

To add a second contender, add another entry to the pipelines dictionary with a distinct key and re-run process(). The evaluation object handles the split, so the two pipelines see identical folds. The repository also ships runnable material under examples/, including how_to_benchmark and paradigm_examples, which is the place to look when the quickstart's paradigm does not match your labels.

## Where MOABB stops being the right tool

MOABB is an offline benchmarking library. Nothing in the README describes online decoding, streaming acquisition, latency budgets or a real-time loop, so if you are building an actual BCI application that must classify a window while the subject is still wearing the cap, MOABB is not that layer. Its unit of work is a downloaded recording, not a live stream. Two further constraints follow from the design. First, the cost is dominated by data: the datasets are fetched and cached locally, and a full multi-dataset benchmark is a storage and bandwidth commitment before it is a compute commitment. Second, the project describes itself in pyproject.toml as Development Status 4 - Beta, and the README carries an explicit disclaimer that it is an open science project that may evolve with community need. That is an honest label: APIs around paradigms and evaluations have changed across releases, so pinning a version for a paper you intend to reproduce later is sensible. Finally, MOABB does not solve the problem it names at the top of the README. It fixes loading and splitting. If your pipeline's advantage comes from a preprocessing step that is not expressed inside the estimator, the benchmark will not capture it.

## MOABB versus Braindecode and raw MNE-Python

The related searches for this project include Braindecode, and the comparison is worth making concrete. Braindecode is a deep learning library for EEG: it provides PyTorch model definitions and training utilities, and its centre of gravity is the network and the training loop. MOABB's centre of gravity is the evaluation protocol: which subjects, which sessions, which epochs, which split, which score. The two are not mutually exclusive, since a Braindecode model wrapped as a scikit-learn compatible estimator is the kind of object MOABB's pipelines dictionary accepts, but they answer different questions. If your question is "does this architecture learn motor imagery better than that one", Braindecode gives you the architecture. If your question is "does this architecture beat a log-variance plus LDA baseline across several public datasets under one fixed protocol", MOABB gives you the protocol and the baseline. The other alternative is doing it by hand with MNE-Python: load each dataset, write your own epoch extraction, write your own cross-session split, and keep it consistent across datasets. That is entirely feasible and gives you full control, but the consistency is on you, and it is exactly the consistency that the README identifies as missing from the literature.

## Maintenance, licensing and the cost of upgrading

The repository is not archived, and the last push was on 2026-09-08. Releases are frequent: v1.6.1 on 2026-08-27, v1.7.0 and v1.7.1 both on 2026-08-31. For a research library, that cadence cuts both ways. You get fixes and new datasets, and you also get a moving target if you are trying to reproduce a number from a paper written a year earlier. Pin the version in your environment file and record it alongside your results. The licence is BSD-3-Clause, declared in pyproject.toml and in the LICENSE file, and the classifiers list it as OSI Approved. BSD-3-Clause is permissive: it allows use and redistribution with the copyright notice and disclaimer retained, and it does not carry the copyleft obligations of a GPL-style licence. That is a statement about the licence text, not legal advice; if you are embedding MOABB in a commercial product, read the LICENSE file and the licences of the datasets it downloads, which are separate from the software licence and are not covered by the README. Upgrading also has a data cost: new releases can change which datasets are available and how paradigms window the signal, so re-running a benchmark after an upgrade is not guaranteed to reproduce the previous numbers.

## Conclusion

Adopt MOABB if you are comparing decoding pipelines across several public EEG datasets and need the dataset loading, session splitting and scoring to be identical between runs. Do not adopt it as a real-time BCI framework: it processes offline recordings, and the repository labels the project Development Status 4 - Beta. Before committing, verify three things: that your target datasets are on the dataset summary page, that the paradigm you need (LeftRightImagery or another) matches your labels, and that the download size of those datasets fits your storage, since MOABB fetches the raw recordings rather than shipping them. The first run is the expensive one; after the cache is warm, comparing a second pipeline against the first is a few lines of code.

## FAQ

### What Python version does MOABB need?

pyproject.toml sets requires-python to >=3.11 and lists classifiers for Python 3.11 through 3.14. The README's quickstart installs it with pip install moabb.

### Which datasets does MOABB include?

The README links a dataset summary page in the documentation, and the quickstart uses BNCI2014_001. The related searches also mention Cho2017, OpenBMI and PhysionetMI, but the authoritative list is the documentation's dataset summary page.

### Can I use my own scikit-learn pipeline with MOABB?

Yes. The quickstart passes a dictionary of named pipelines to evaluation.process(), and each value is an ordinary scikit-learn estimator built with make_pipeline. The dictionary key becomes the pipeline name in the results DataFrame.

### Is MOABB suitable for a real-time BCI application?

Nothing in the README describes streaming, online decoding or latency handling. MOABB works on downloaded recordings and evaluates pipelines offline, so a real-time decoder needs a different layer.

## Sources

- [License: BSD-3-Clause](https://github.com/NeuroTechX/moabb/blob/develop/LICENSE)
- [NeuroTechX/moabb on GitHub](https://github.com/NeuroTechX/moabb)
- [Project website](https://moabb.neurotechx.com/docs/index.html)
- [README](https://github.com/NeuroTechX/moabb/blob/develop/README.md)
- [Releases](https://github.com/NeuroTechX/moabb/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/neurotechx-moabb
