# NeuroKit2: a Python toolbox for ECG, EDA, PPG and HRV signal processing

> NeuroKit2 wraps filtering, peak detection and feature extraction for physiological signals behind two functions, bio_process and bio_analyze. It is aimed at researchers and clinicians who want to analyze ECG, RSP and EDA data without writing their own pipeline.

**neuropsychology/NeuroKit** — NeuroKit2: The Python Toolbox for Neurophysiological Signal Processing

- Repository: https://github.com/neuropsychology/NeuroKit
- Website: https://neuropsychology.github.io/NeuroKit
- Stars: 2,385 · Forks: 550
- Language: Python
- License: MIT
- Published: 2026-09-28 · Updated: 2026-09-28 · Language: en
- Canonical page: https://hysenlabs.com/projects/neuropsychology-neurokit

## The problem NeuroKit2 targets: physiological recordings are messy before they are analyzable

A raw ECG or EDA recording is not a dataset. Before any statistic can be computed, the signal has to be filtered, peaks have to be located, artifacts have to be flagged, and the intervals between peaks have to be turned into features such as heart rate variability. Each of those steps has its own literature and its own parameter choices. Researchers who are not signal processing specialists often end up writing the same bandpass filter, the same R-peak detector and the same windowing code that a dozen other labs have already written.

NeuroKit2 exists to collapse that work into a small number of calls. The README describes the target audience directly: researchers and clinicians without extensive knowledge of programming or biomedical signal processing who want to analyze physiological data with only two lines of code. The package covers ECG, RSP, EDA, PPG, EMG, EOG and EEG, and the topics list on the repository adds entropy and heart-rate variability to that coverage. The scope is deliberately broad rather than deep in one modality.

## How bio_process and bio_analyze move a recording from raw channels to features

The architecture visible in the README is a two-stage pipeline. The first stage, nk.bio_process, takes one or more signal arrays plus a sampling rate and returns a tuple: a processed dataframe and an info dictionary. Filtering, peak detection and the per-modality cleaning steps happen inside that call. The second stage, nk.bio_analyze, consumes the processed dataframe and the sampling rate and returns a results object with the derived features.

The README example passes ECG, RSP and EDA together and a sampling rate of 100. That is the design pattern: the multi-modal call is the default path, not a special case. The info dictionary returned by bio_process is where the per-signal metadata ends up, which matters because the analysis stage needs to know what was found before it can compute intervals.

Because the two functions are thin wrappers over a larger library, the customization route is to call the underlying per-modality functions instead. The README links a tutorial titled Customize your Processing Pipeline, and the repository ships separate example notebooks for EDA peaks, respiratory rate variability, individual heartbeat extraction and ECG wave delineation. Those examples are the practical documentation for anyone who needs to change a filter cutoff or a peak detection method rather than accept the defaults.

## Installing NeuroKit2 and running a first multi-modal analysis

The README gives two install paths. PyPI is the primary one, and conda-forge is the alternative. The package requires Python 3.10 or newer according to pyproject.toml, so an older interpreter will fail at install time rather than at import time.

```bash
pip install neurokit2
```

If you prefer conda, the README gives the conda-forge channel form:

```bash
conda install -c conda-forge neurokit2
```

The README points users who are unsure which route to take at a dedicated installation guide on the documentation site. Once installed, the first real use is the example from the README itself. It downloads a bundled dataset, processes three channels, and computes features:

```python
import neurokit2 as nk

# Download example data
data = nk.data("bio_eventrelated_100hz")

# Preprocess the data (filter, find peaks, etc.)
processed_data, info = nk.bio_process(ecg=data["ECG"], rsp=data["RSP"], eda=data["EDA"], sampling_rate=100)

# Compute relevant features
results = nk.bio_analyze(processed_data, sampling_rate=100)
```

What you should see is a processed dataframe with the cleaned and annotated channels, an info dictionary, and a results object holding the extracted features. Note the channel names in the example: ECG, RSP and EDA, with a sampling rate of 100. If your own dataframe uses different column names, the call will need to be adapted, and the sampling rate you pass must match the rate the data was actually recorded at, because the peak detection and interval calculations depend on it.

## Where NeuroKit2 stops being the right tool

The package is built around offline batch processing of arrays and dataframes. Nothing in the README or the repository layout describes a streaming interface, a socket, a device driver or a real-time callback. If your application needs to detect an R-peak within milliseconds of it arriving from a sensor, this is not the library for that job, and no amount of customization of bio_process will turn it into one.

There is a second boundary that the project itself acknowledges. The pyproject.toml classifier reads Development Status :: 2 - Pre-Alpha, even though the package is on its thirteenth 0.2.x release and the last push to master was on 2026-09-27. Pre-Alpha is a conservative label, but it is the label the maintainers chose, and it signals that function signatures and default parameters can still move between minor versions. Anyone pinning NeuroKit2 in a production pipeline should pin the version explicitly rather than track the latest release.

A third constraint is the dependency set. The core install pulls numpy>=2.0.0, pandas<3.0.0, scipy, scikit-learn, matplotlib, PyWavelets and setuptools. The optional full extra adds a much longer list including mne, opencv-python, bioread, pyxdf, PyEMD, plotly and cvxopt. Several of those are heavy or have their own build requirements, so the full extra is not a casual addition to a lightweight environment.

## NeuroKit2 against MNE-Python: overlapping EEG ground, different centers of gravity

MNE-Python is the obvious comparison point, and the difference is one of emphasis rather than a strict either-or. MNE is built around EEG and MEG: montages, channel layouts, source localization, evoked responses, time-frequency decomposition. Its data model is an object hierarchy of Raw, Epochs and Evoked containers.

NeuroKit2 approaches the same recordings from the peripheral physiology side. The README example is ECG, RSP and EDA, not EEG epochs. Its output is a dataframe plus an info dictionary, which fits a workflow where the next step is a statistical model in pandas or R rather than a topographic plot. For a study that is mostly EEG with an occasional ECG channel for heart rate, MNE plus a small amount of custom peak detection is a reasonable path. For a study built around autonomic measures, NeuroKit2 is the shorter route. The two are not mutually exclusive: mne appears in NeuroKit2's own optional full dependency list, which suggests the maintainers expect users to combine them rather than pick one.

## Maintenance, licence and the cost of upgrading

The repository is not archived, and the last push to master was on 2026-09-27. Releases in the 0.2.x line have arrived at a steady pace: v0.2.11 on 2025-05-13, v0.2.12 on 2025-07-08 and v0.2.13 on 2026-03-02. The project also carries a NEWS.rst file at the repository root, which is where release notes live, and a CITATION.cff for academic citation.

The practical upgrade cost comes from the version floor and ceiling in pyproject.toml. NeuroKit2 requires numpy>=2.0.0 and pandas<3.0.0. If your environment is still on numpy 1.x, installing NeuroKit2 will force an upgrade that may break other packages in the same environment. If pandas 3.0 is released and you move to it, NeuroKit2's upper bound will block the upgrade until the project widens it. The optional full extra compounds this, since packages such as mne and opencv-python have their own compatibility ranges.

The licence is MIT, stated in both the LICENSE file and the pyproject.toml metadata. MIT is permissive: it allows commercial and closed-source use, modification and redistribution provided the copyright notice and licence text are retained. That is the general shape of the licence, not legal advice, and anyone embedding NeuroKit2 in a distributed product should read the LICENSE file at the repository root and, where it matters, get proper counsel.

## Conclusion

Adopt NeuroKit2 if you have physiological recordings in arrays or dataframes and want a documented pipeline for filtering, peak detection and feature extraction without assembling it from scipy and custom code. Do not adopt it for real-time streaming or for clinical diagnosis, since the README describes a research-oriented package and the classifiers still mark the project as Pre-Alpha. Before committing, verify that your sampling rate and channel names match what bio_process expects, and check the version you install against the NEWS.rst changelog, because the dependency floor is numpy>=2.0.0 and pandas<3.0.0.

## FAQ

### What is NeuroKit2 used for?

NeuroKit2 is a Python toolbox for neurophysiological signal processing. It covers modalities including ECG, EDA, PPG, EMG, EOG, EEG and RSP, and its README example shows preprocessing and feature extraction for ECG, RSP and EDA in two function calls.

### How do I install NeuroKit2?

The README gives two routes: pip install neurokit2 from PyPI, or conda install -c conda-forge neurokit2 from conda-forge. The package requires Python 3.10 or newer according to pyproject.toml.

### Which signals does NeuroKit2 process?

The repository topics list biosignals, cardiac, ECG, EDA, EEG, EMG, EOG, PPG, HRV, entropy and skin conductance. The README's quick example processes ECG, RSP and EDA together at a sampling rate of 100.

### What is the difference between bio_process and bio_analyze in NeuroKit2?

bio_process handles the preprocessing stage: filtering, peak detection and the per-modality cleaning steps, returning a processed dataframe and an info dictionary. bio_analyze takes that processed dataframe and the sampling rate and returns the relevant features.

### Which Python versions and dependencies does NeuroKit2 need?

pyproject.toml sets requires-python to >=3.10 and lists numpy>=2.0.0, pandas<3.0.0, scipy, scikit-learn>=1.0.0, matplotlib>=3.5.0, PyWavelets>=1.4.0, setuptools<82.0.0 and requests as core dependencies. An optional full extra adds packages including mne, opencv-python, bioread, pyxdf, PyEMD, plotly and cvxopt.

### What licence does NeuroKit2 use?

The licence is MIT, stated in the LICENSE file and in the pyproject.toml metadata. MIT permits commercial and closed-source use, modification and redistribution as long as the copyright notice and licence text are retained.

## Sources

- [License: MIT](https://github.com/neuropsychology/NeuroKit/blob/master/LICENSE)
- [neuropsychology/NeuroKit on GitHub](https://github.com/neuropsychology/NeuroKit)
- [Project website](https://neuropsychology.github.io/NeuroKit)
- [README](https://github.com/neuropsychology/NeuroKit/blob/master/README.md)
- [Releases](https://github.com/neuropsychology/NeuroKit/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/neuropsychology-neurokit
