# River: online machine learning in Python for streaming data

> River is a Python library for online machine learning, built around learn_one and predict_one. It fits event-based pipelines that must adapt to concept drift, and it is the wrong tool when a periodic batch refit already works.

**online-ml/river** — 🌊 Online machine learning in Python

- Repository: https://github.com/online-ml/river
- Website: https://riverml.xyz
- Stars: 6,116 · Forks: 830
- Language: Python
- License: BSD-3-Clause
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/online-ml-river

## The problem River solves, and the reader it assumes

Batch learning assumes you can collect a dataset, fit a model, and repeat later. That cycle breaks when data arrives as an event stream and the relationship between features and target shifts while the model is serving. River targets that case. The README is blunt about the audience: "You should ask yourself if you need online machine learning. The answer is likely no. Most of the time batch learning does the job just fine."

The library is for engineers who want a model that learns from new data without revisiting past data, who need robustness to concept drift, and who want to develop in a way closer to production, which the README describes as usually event-based. River is the merger of creme and scikit-multiflow, so its API inherits conventions from both. The unit of work is a single observation, not a matrix.

## How the learn_one and predict_one loop works

River's core interface is two methods. predict_one takes a dictionary of features and returns a prediction; learn_one takes the same dictionary plus the true label and updates the model in place. There is no fit call that consumes a whole array, and no separate transform step that must be applied to a held-out set. State lives inside the estimator objects and changes with every call.

Composition follows the same pattern. compose.Pipeline chains a transformer and an estimator, and calling learn_one on the pipeline routes the sample through the transformer first. Metrics work the same way: metrics.Accuracy has an update method that takes the true label and the predicted label, and it accumulates a running score. Because every component is updated per sample, you can interleave prediction, metric update and model update in one loop, which is exactly the order the README's quickstart uses.

The repository shows a mixed implementation. pyproject.toml declares a maturin build backend and a module named river._river_rust, and Cargo.toml defines a cdylib crate with pyo3 0.29.2. The Rust side carries benchmarks for statistics, ADWIN, rolling metrics, sorted windows, covariance and expected mutual information. So parts of the numeric core are compiled, while the Python layer keeps the user-facing API. The README states that River focuses on clarity and user experience more so than performance, and that it is very fast at processing one sample at a time.

## Installing River and running a first streaming loop

River requires Python 3.11 and above. The README gives a plain pip install, and notes that wheels exist for Linux, macOS and Windows, so building from source is usually unnecessary.

```bash
pip install river
```

The core learn_one and predict_one interface has no pandas dependency. The mini-batch interface, which covers learn_many, predict_many, predict_proba_many and transform_many, is built on pandas and is opt-in through an extra.

```bash
pip install "river[pandas]"
```

A first real use is the website phishing dataset bundled with the library. The README prints the first observation, which is a dictionary of numeric features such as age_of_domain, https and long_url, paired with a boolean label. You build a pipeline, attach a metric, and loop.

```python
from river import compose, datasets, linear_model, metrics, preprocessing

model = compose.Pipeline(
    preprocessing.StandardScaler(),
    linear_model.LogisticRegression(),
)
metric = metrics.Accuracy()

for x, y in datasets.Phishing():
    y_pred = model.predict_one(x)
    metric.update(y, y_pred)
    model.learn_one(x, y)

print(metric)
```

The README reports that this exact setup reaches an Accuracy of 89.28 percent on that dataset. Note the ordering: predict, score, then learn. Reversing the last two steps trains on the label before predicting it, and the metric will look better than the model is. For a development install, the README also gives a git-based path, but states that it requires Cython and Rust on the machine.

```bash
pip install git+https://github.com/online-ml/river --upgrade
```

## Where River is the wrong tool

The README answers the adoption question before you ask it: if you do not specifically need online learning, you probably do not need River. Periodic batch retraining with a conventional library covers a large share of production cases, and it gives you cross-validation, hyperparameter search tooling and a stable artefact you can version and roll back. River's model state is mutable and moves with every sample, so reproducing a past prediction means replaying the stream.

Performance is a stated trade-off rather than an accident. The README says the project focuses on clarity and user experience more so than performance. If your bottleneck is throughput on large batches rather than latency on individual events, the per-sample API is the wrong shape. The pandas mini-batch interface exists, but it is an opt-in extra, which tells you where the design centre sits.

There is also an ecosystem gap. River ships online implementations of linear models, trees and forests, approximate nearest neighbours, anomaly detection, drift detection, recommenders, forecasting, bandits, factorization machines, imbalanced learning, clustering, ensembles and active learning. What it does not offer is the breadth of pretrained models and third-party integrations that batch-oriented libraries accumulate. You are assembling the pipeline yourself.

## River compared with scikit-learn plus periodic retraining

The honest alternative is scikit-learn with a scheduled refit. The difference is not accuracy on a static split; it is what happens between refits. A scikit-learn estimator is fitted once and frozen. When the distribution moves, you either wait for the next retrain or accept degraded predictions. River's estimators keep updating, so the model tracks the stream continuously, and river.drift provides detectors so you can observe when the distribution has moved rather than inferring it from a dashboard.

The cost is operational. A frozen model is a file you can hash, ship and roll back. A continuously updated model is a process, and any bug in the update path is silent: the model keeps predicting, just worse. River's progressive model validation utilities exist for this reason, but they do not remove the need to monitor the metric stream itself.

The two are not mutually exclusive. A common shape is a batch model as a baseline and an online model alongside it, with the online model promoted only once its running metric holds up. River's compatibility extras include scikit-learn and SQLAlchemy, which suggests interoperation is an intended path rather than an afterthought.

## Maintenance, releases and the BSD-3-Clause licence

The repository is not archived and the last push was on 2026-09-03, which is recent. Releases are frequent: 0.26.1 on 2026-08-21, 0.26.0 the same day, and 0.25.0 on 2026-05-31. The version still sits below 1.0, so the API can move between minor releases, and the changelog is the place to check before upgrading. There is no long-term support branch documented in the README.

Upgrade cost is shaped by the dependency floor. pyproject.toml pins scipy to >=1.14.1,<2, numpy to >=2.2.5,<3, and narwhals to >=2.0.0, with pandas >=2.2,<3 only in the optional extra. Those upper bounds mean a major numpy or scipy release will require a River release before you can take it, which is a constraint worth knowing if you upgrade dependencies aggressively.

Building from source is heavier than the pip path. The build backend is maturin, the crate uses pyo3, and Cargo.toml sets unsafe_code to forbid and warnings to deny, so a warning in the Rust layer fails the build. There is also an abi3 feature used only for a fallback wheel in CI, which the Cargo.toml comments say is enabled only for that wheel because the default per-version wheels are built without it for full-speed native code.

The licence is BSD-3-Clause. That is a permissive licence, and the practical implication is that redistribution in source or binary form requires retaining the copyright notice and disclaimer. This is a description of the licence text, not legal advice; check the LICENSE file and your own obligations.

## Conclusion

Adopt River if your model must update from each event without revisiting past data, and if you accept that the project favours clarity and user experience over raw performance. Do not adopt it if periodic batch retraining already meets your latency and accuracy needs, or if you need a wide catalogue of pretrained models. Before committing, verify that Python 3.11 or above is available in your environment, that the pandas mini-batch interface is worth the extra dependency for your workload, and that the drift detectors in river.drift match the kind of drift you actually expect.

## FAQ

### What Python version does River require?

River is intended to work with Python 3.11 and above, and pyproject.toml sets requires-python to >=3.11. Wheels are available for Linux, macOS and Windows.

### How do I install River?

The README gives pip install river for the core library, which has no pandas dependency for learn_one and predict_one. The mini-batch interface needs the extra: pip install "river[pandas]".

### Do I need online machine learning at all?

River's own README says you should ask yourself that question and that the answer is likely no, because batch learning usually does the job. It lists three cases where an online approach fits: learning from new data without revisiting past data, robustness to concept drift, and development closer to event-based production behaviour.

## Sources

- [License: BSD-3-Clause](https://github.com/online-ml/river/blob/main/LICENSE)
- [online-ml/river on GitHub](https://github.com/online-ml/river)
- [Project website](https://riverml.xyz)
- [README](https://github.com/online-ml/river/blob/main/README.md)
- [Releases](https://github.com/online-ml/river/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/online-ml-river
