# Pathway Live Data Framework: a Python ETL framework that runs your pipeline on a Rust engine

> Pathway is a Python ETL framework for stream processing, real-time analytics, LLM pipelines and RAG. It targets engineers who want one pipeline to work for batch and streaming, and it pays for that with a Rust engine, a BUSL-1.1 core and a licensing split between the free and enterprise editions.

**pathwaycom/pathway** — Pathway is a Python ETL framework for stream processing, real-time analytics, LLM pipelines and RAG, driven by a scalable Rust engine using Differential Dataflow.

- Repository: https://github.com/pathwaycom/pathway
- Website: https://pathway.com
- Stars: 62,215 · Forks: 1,682
- Language: Python
- License: not declared
- Published: 2026-08-08 · Updated: 2026-08-18 · Language: en
- Canonical page: https://hysenlabs.com/projects/pathwaycom-pathway

## What Pathway Live Data Framework actually replaces

Most Python data teams end up maintaining two code paths for the same logic: a batch job that runs on a schedule and a streaming job that keeps a result current. Pathway's pitch is that you write one pipeline and the same code covers local development, CI/CD tests, batch jobs, stream replays and live streams. That is the claim in the README, and it is the reason the project exists.

The audience is narrower than "Python developers". It is engineers who already know they need incremental results: a dashboard that must reflect a Kafka topic within seconds, a RAG index that must pick up new documents without a nightly rebuild, an alerting pipeline that reacts to log lines as they arrive. If your data lands once a day and a cron job finishes in ten minutes, Pathway adds a Rust build dependency and a new execution model for no benefit. The README frames the framework around "stream processing, real-time analytics, LLM pipelines, and RAG", and that framing is honest about where it fits.

## Differential dataflow under a Python API

The split is the whole design. You write Python; the Rust engine executes it. The README states the engine is "based on Differential Dataflow and performs incremental computation", and that your Python code is run by the Rust engine, which enables multithreading, multiprocessing and distributed computations. Cargo.toml confirms the shape of that: the crate is named pathway_engine, declared as both cdylib and lib, and it depends on differential-dataflow through a vendored path at ./external/differential-dataflow rather than a crates.io release.

Incremental computation is the part worth understanding before you adopt anything. Instead of recomputing an aggregate from scratch when a row arrives, the engine propagates the change through the operators that depend on it. That is what makes late and out-of-order data tractable: the README says Pathway "manages late and out-of-order points by updating its results whenever new (or late, in this case) data points come into the system". A batch-oriented framework cannot do this without a full rerun, and a hand-written streaming job usually gets it wrong at the first late event.

The trade-off is memory. The README states plainly that "all the pipeline is kept in memory". There is no spill-to-disk story in the README, so the working set of your state has to fit in the memory you give the process. Persistence is offered as a feature to save computation state and restart after an update or a crash, but that is checkpointing, not a disk-backed operator state. Size your state before you size your cluster.

The dependency list in pyproject.toml also tells you what you are buying. Alongside pandas, numpy and pyarrow, the base install pulls in fastapi and uvicorn, panel and jupyter_bokeh, sqlalchemy and sqlmodel, deltalake, google-cloud-bigquery, google-cloud-pubsub and the OpenTelemetry SDK. That is a wide default footprint for an ETL library, and it is a direct consequence of shipping connectors and LLM tooling in the same package.

## Installing Pathway and running a first pipeline

Pathway requires Python 3.10 or above, per the README and the requires-python field in pyproject.toml. The install is a single pip command, and the README uses the -U flag to pull the current release:

```bash
pip install -U pathway
```

The package is built with maturin, so a source install compiles the Rust engine; the Cargo.toml pins rust-version to 1.97, and rust-toolchain.toml at the repository root fixes the toolchain. A wheel install avoids that entirely, which is what most users should do first.

Optional extras are declared as extras in pyproject.toml. The pyfilesystem extra pulls in fs and a pinned setuptools, the milvus extra pulls in pymilvus and milvus-lite, and there is a sql extra. Install only the ones your pipeline touches:

```bash
pip install -U pathway
```

Once installed, the entry point is a Python script that defines tables and transformations and then runs the pipeline. The README points to runnable templates in notebook and Docker form at pathway.com/developers/templates, covering Kafka ETL, event-driven alerting, real-time analytics, and RAG variants including a private RAG with Ollama and Mistral AI. Those templates are the fastest way to see a working pipeline, because the README does not inline a minimal code example in the section shown here. Start from a template rather than from a blank file: you will get the connector configuration and the run call in the correct shape, and you can strip it down from there.

## Consistency, persistence and the line the free edition does not cross

This is the constraint that decides most adoption questions. The README states that the free version of Pathway gives "at least once" consistency while the enterprise version provides "exactly once". If your pipeline writes to a sink where a duplicate is a real problem, such as a ledger or a billing table, at-least-once semantics push the deduplication work back onto you, or push you toward the enterprise edition. The README does not describe how exactly-once is implemented, so there is nothing here to evaluate beyond the statement itself.

The second limitation is the licence. Cargo.toml declares license = "BUSL-1.1" for the engine crate, and pyproject.toml classifies the package as "License :: Other/Proprietary License". The README links to LICENSE.txt rather than stating terms inline, and the repository carries a library_licenses/ directory, which suggests the dependency licensing is tracked deliberately. A Business Source License is source-available, not open source in the OSI sense, and it typically carries use restrictions and a change date. The README does not give the change date or the restricted uses. Anyone evaluating this for a commercial product needs to read LICENSE.txt directly rather than rely on the "open source" framing that surrounds most Python data tooling.

The third limitation is operational. Because the pipeline lives in memory and the engine handles the time dimension for you, the failure modes are different from a Spark job. A crash without a recent checkpoint means re-reading from the source, and whether that is cheap depends entirely on the connector. The README lists connectors for Kafka, GDrive, PostgreSQL and SharePoint, plus an Airbyte connector for more than 300 sources, but it does not document rollback behaviour or replay guarantees per connector. That gap is worth testing against your own source before you commit.

## How Pathway differs from Flink and Spark Structured Streaming

The honest comparison is Apache Flink. Both are stateful stream processors with a notion of event time and a way to handle late data. The difference is the surface. Flink asks you to write Java, Scala or Python against its DataStream or Table API, and its Python support runs through PyFlink, which is a binding over the JVM runtime. Pathway inverts that: the API is Python first, and the Rust engine is an implementation detail you do not write against. If your team is Python-only and your transformations lean on scikit-learn, pandas or an LLM client, that inversion is the entire argument.

Spark Structured Streaming is the other reference point, and the difference is the execution model. Spark processes micro-batches; Pathway performs incremental computation over a differential dataflow, and the README contrasts its unified engine for batch and streaming against the usual two-code-path setup. Spark's advantage is the ecosystem: a mature SQL layer, a huge connector catalogue, and operational knowledge that exists in most data teams. Pathway's connector list is real but shorter, and the Airbyte connector is the escape hatch for the long tail.

The third alternative is not a framework at all: a plain Python consumer loop with a queue and a database. It is simpler, it has no Rust toolchain, and for a few thousand events per second with simple aggregation it is often enough. Pathway becomes the right answer when you need joins, windowing and sorting over a stream that keeps arriving, and you want the late-data handling to be someone else's problem.

## Maintenance, releases and what upgrading costs

The last push to the default branch was on 2026-08-01, and the most recent release, v0.32.1, carries the same timestamp. Before that, v0.31.1 landed on 2026-06-12 and v0.31.0 on 2026-05-25. The cadence visible in those three releases is roughly monthly, with minor version bumps rather than patch-only churn.

The upgrade cost is dominated by the dependency surface, not by Pathway's own API. pyproject.toml pins pyarrow to >=10.0.0, < 26.0.0, beartype to >=0.14.0, < 0.23.0, and deltalake to >=1.6.0, < 2.0.0, while pydantic is pinned with a compatible-release operator at ~=2.9. Those upper bounds are there to keep the engine and the Python layer in step, and they mean a Pathway upgrade can force a pyarrow or deltalake upgrade in the same environment. If you co-install Pathway with a warehouse client that pins pyarrow differently, you will feel it.

The Rust side adds a second upgrade axis. Cargo.toml pins rust-version to 1.97 and rust-toolchain.toml fixes the toolchain, so a source build is reproducible but not flexible. The Cargo.toml also carries a comment explaining that chrono-tz 0.8 is pulled in under an alias, chrono-tz-0_8, purely because the ClickHouse client expects that major version while the rest of the crate uses 0.10. That is the kind of detail that tells you the engine integrates with a lot of external systems, and that each of those integrations is a place where a version bump can break something.

## Conclusion

Adopt Pathway if you want a single Python pipeline that handles batch and streaming, and you accept the Rust engine, the in-memory model and the BUSL-1.1 core. Skip it if you need exactly-once consistency without an enterprise agreement, or if your workload does not need incremental recomputation. Before committing, check LICENSE.txt for the terms that apply to your deployment, and confirm whether the persistence and consistency guarantees you need sit in the free edition or the enterprise one.

## FAQ

### What is Pathway Live Data Framework?

It is a Python ETL framework for stream processing, real-time analytics, LLM pipelines and RAG, published by pathwaycom. Your Python code is executed by a Rust engine based on Differential Dataflow that performs incremental computation.

### How do I install Pathway Live Data Framework?

The README gives pip install -U pathway, and the package requires Python 3.10 or above. Optional extras such as sql, milvus and pyfilesystem are declared in pyproject.toml and are installed alongside the base package.

### Does Pathway Live Data Framework support exactly-once processing?

The README states that the free version gives at-least-once consistency while the enterprise version provides exactly-once consistency. The README does not describe how the exactly-once implementation works.

### What licence does Pathway Live Data Framework use?

Cargo.toml declares BUSL-1.1 for the engine crate, and pyproject.toml classifies the package as an Other/Proprietary License. The README links to LICENSE.txt for the terms rather than stating them inline.

### Which data sources can Pathway Live Data Framework connect to?

The README lists connectors for Kafka, GDrive, PostgreSQL and SharePoint, and an Airbyte connector that reaches more than 300 data sources. It also says you can build a custom connector in Python if the one you need is missing.

## Sources

- [Official documentation](https://pathway.com)
- [Official README](https://github.com/pathwaycom/pathway#readme)
- [Project repository](https://github.com/pathwaycom/pathway)
- [Release notes](https://github.com/pathwaycom/pathway/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/pathwaycom-pathway
