# Retentioneering Tools: clickstream analytics in Python with an MCP server for AI agents

> Retentioneering turns timestamped event logs into interactive user-flow graphs, funnels and step matrices inside a notebook. Version 5.0 replaced the pandas engine with DuckDB and added an MCP server so agents can run the same analyses.

**retentioneering/retentioneering-tools** — Python toolkit, MCP server, and agent skills for reproducible, auditable clickstream and event log analytics. Helps AI agents, data scientists and analysts build, validate, and cross-check product analytics, quantitative UX, customer journeys, graph-based user flows, behavioral segmentation, A/B tests, process mining models, Markov chain simulation

- Repository: https://github.com/retentioneering/retentioneering-tools
- Website: https://retentioneering.com/docs/
- Stars: 919 · Forks: 136
- Language: Python
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/retentioneering-retentioneering-tools

## What Retentioneering solves, and for whom

Funnels answer whether users passed through predefined stages. They do not answer how users moved between those stages. The gap between those two questions is where most product analytics work actually happens: which alternative routes reach checkout, where loops form, which paths dead-end, and how any of that differs between two cohorts. Retentioneering is built for that second question. The README frames the project as an alternative to one-off scripts written for a single question, and the pitch is that tested analytical primitives replace generated code.

The intended users are data scientists and analysts who work in Python, plus AI agents that operate through the bundled MCP server and agent skills. That second audience shapes the design more than the first. A toolkit that an agent can call needs deterministic outputs and reusable components, which is why the README talks about reducing token usage and cross-checking results through independent computations rather than about dashboards. If you want a point-and-click interface for a marketing team, this is not that product.

## The DuckDB Eventstream and the processor chain

Version 5.0 is described in the README as a ground-up rewrite. The pandas engine and CDN-loaded widgets of 3.x were replaced by a DuckDB-backed Eventstream object and a new generation of anywidget-based widgets. The legacy engine still lives on the 3.x branch, which tells you the maintainers expect some users to stay behind.

Everything starts from a DataFrame with three columns: a path identifier, an event name and a timestamp. If your column names differ, the documentation points to a schema parameter on the Eventstream constructor. From there, data processors chain. Each call returns a new Eventstream and the original is never modified, so a filtering step cannot silently corrupt the object you loaded. The README shows filter_events, collapse_events with loops=True, and split_sessions with a 30m timeout composed in a single expression.

The output side has two forms. Widgets render interactive views in Jupyter, Colab, Cursor, VS Code, Codex or Claude Code. Every widget also has a headless twin that returns plain data: transition_graph_data returns a DataFrame, funnel_data returns a dict. That pairing is the most useful design decision in the project, because it means a chart and a number can come from the same computation instead of two implementations that drift.

## Installing Retentioneering and running a first graph

Python 3.10 or newer is required. The README gives a single pip command for a normal environment and the same command with a leading exclamation mark inside a notebook cell.

```bash
pip install retentioneering
```

Once installed, the smallest useful program needs a CSV with user_id, event and timestamp columns. Loading it produces an Eventstream, and transition_graph renders the interactive behavior graph in the notebook output.

```python
import pandas as pd
import retentioneering as rete

df = pd.read_csv("events.csv")   # columns: user_id, event, timestamp
stream = rete.Eventstream(df)

stream.transition_graph()        # interactive behavior graph, right in the notebook
```

If you have no data yet, the package ships a synthetic e-commerce dataset. The README uses it to build a funnel and to compare two platform segments in diff mode, which is the fastest way to see what the widgets actually look like before pointing them at production data.

```python
import retentioneering as rete

ecom = rete.datasets.load_ecom()

ecom.funnel(steps=["catalog", "add_to_cart", "purchase"])

ecom.transition_graph(diff=["platform", "mobile", "desktop"])
```

The README notes that anonymous product telemetry may be enabled by default and can be disabled at any time. Raw and analysed event data stay in your environment, which is stated explicitly. If your organisation treats outbound telemetry as a blocker, that switch is the first thing to locate, and the README does not name the exact setting.

## Where Retentioneering is the wrong tool

The honest limitation is the one the README admits indirectly: 5.0 is a rewrite, and the 3.x engine is still maintained on a separate branch. Migrations of individual features from 3.x into 5.x are handled through issues and pull requests, which means feature parity is a work in progress rather than a guarantee. If your existing pipelines depend on a 3.x behaviour that has not been ported, upgrading is a project, not a version bump.

The second limitation is the interface. Retentioneering renders widgets in Python-compatible environments. There is no hosted SaaS platform, and the README presents that as a feature: your data never leaves your machine. The cost is that anyone who does not run Python cannot open the analysis. An exported HTML report is shareable, but interactive exploration is not.

Third, the toolkit assumes event data is already clean enough to have a path identifier, an event name and a timestamp. Sessionisation is available through split_sessions with a timeout, but identity resolution, bot filtering and schema drift upstream of the Eventstream are your problem. filter_events can drop events by name, which is a blunt instrument compared with a proper ingestion layer.

## How it differs from session-replay and BI dashboard tools

The closest category is session replay and product analytics platforms that ship a JavaScript snippet and a hosted dashboard. Those tools collect the data for you and give non-engineers a UI. Retentioneering does neither. It reads event logs you already have, from CSV, TSV, Parquet or a custom export out of BigQuery or ClickHouse, and it returns graphs and DataFrames inside your own environment. The difference in approach is collection versus analysis: a replay tool owns the pipeline, Retentioneering owns the method.

Against general BI dashboards the split is similar. A BI tool aggregates into charts defined ahead of time. Retentioneering computes path structures on demand, including step matrices along a path pattern such as add_to_cart followed by anything followed by purchase. The README also positions the project against writing analysis code from scratch for each question, arguing that reusing domain-specific components reduces implementation effort and the risk of subtle analytical errors. That claim is plausible for repeated analyses and weaker for a one-off question, where a short script may be faster than learning the processor API.

## Maintenance, releases and the licence split

The last push to the default branch was on 2026-09-05, and the most recent release listed is v5.2.0 on 2026-08-19, preceded by v5.1.0 on 2026-07-22 and v5.0.1 on 2026-07-15. The cadence over that window is roughly monthly. The repository is not archived.

Upgrade cost has a concrete shape here. The Makefile documents that master is protected and takes pull requests only, and that merging to master does not trigger a release: only pushing a v* tag does. For a user rather than a contributor, that means releases are deliberate. It also means the CHANGELOG is the place to read before upgrading, since the 3.3.0 to 5.0 delta is large by the maintainers' own description. Running from source additionally requires uv sync plus an npm install in js, and building widgets regenerates a metric schema first so the TypeScript definitions cannot go stale.

The licence is Apache-2.0, and the repository also carries a COMMERCIAL.md file at the top level. Apache-2.0 permits commercial use of the code under its own terms. What COMMERCIAL.md adds is not described in the README, so read it before assuming the Apache grant is the whole story. Nothing here is legal advice.

## Conclusion

Adopt Retentioneering if your questions are about paths, loops and segments rather than aggregate metrics, and if you want the analysis to live in a notebook next to the code. Skip it if you need a hosted dashboard that non-technical stakeholders open without Python, or if you are still on pandas 3.x workflows that you do not want to migrate. Before committing, verify two things: that your event log has a path identifier, an event name and a timestamp column, and that the DuckDB-backed Eventstream loads your file size without you needing the 3.x branch. Check the licence and COMMERCIAL.md together, because Apache-2.0 covers the code while the commercial file exists for a reason the README does not spell out.

## FAQ

### What does retention analysis mean in Retentioneering?

The README describes the toolkit as turning raw sequences of user and system events into answers about where users get stuck and which journeys lead to conversion or churn. Retention is treated as one question among several, alongside funnels, segments and A/B comparisons, rather than as the only output.

### How do I install Retentioneering tools in Python?

Python 3.10 or newer is required, and the README gives pip install retentioneering as the install command, with the same command prefixed by an exclamation mark inside a notebook cell.

### Can Retentioneering run without sending my event data to a server?

Yes. The README states that the analysis runs in your own environment and that raw and analysed event data never leave your machine. Anonymous product telemetry may be enabled by default and can be disabled at any time.

### What data format does an Eventstream need?

The README says all you need is a DataFrame with three columns: a path identifier, an event name and a timestamp. Different column names are handled by passing a schema, which the documentation covers on the Eventstream page.

### Does Retentioneering work with AI agents?

The project ships an MCP server and a collection of agent skills, and the README describes asking an AI agent to run and cross-check the analysis through them. The repository also contains .agents/, .claude/, AGENTS.md and CLAUDE.md entries at the top level.

## Sources

- [License: Apache-2.0](https://github.com/retentioneering/retentioneering-tools/blob/master/LICENSE)
- [Project website](https://retentioneering.com/docs/)
- [README](https://github.com/retentioneering/retentioneering-tools/blob/master/README.md)
- [Releases](https://github.com/retentioneering/retentioneering-tools/releases)
- [retentioneering/retentioneering-tools on GitHub](https://github.com/retentioneering/retentioneering-tools)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/retentioneering-retentioneering-tools
