# Kedro: a Python framework for pipelines that survive handover

> Kedro wraps data science code in a project template, a Data Catalog and a node graph so pipelines stay reproducible. It is a strong fit for teams, and overkill for a single notebook.

**kedro-org/kedro** — Kedro is a toolbox for production-ready data science. It uses software engineering best practices to help you create data engineering and data science pipelines that are reproducible, maintainable, and modular.

- Repository: https://github.com/kedro-org/kedro
- Website: https://kedro.org
- Stars: 11,003 · Forks: 1,078
- Language: Python
- License: NOASSERTION
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/kedro-org-kedro

## The problem Kedro solves: notebooks that cannot be handed over

The README is unusually blunt about its own origin story. Kedro was built, in the maintainers' words, from "best-practice (and mistakes) trying to deliver real-world ML applications that have vast amounts of raw unvetted data", and it exists to address "the main shortcomings of Jupyter notebooks, one-off scripts, and glue-code". That is the whole pitch. The failure mode it targets is not bad modelling. It is a pipeline that only its author can run, where the order of cell execution is the documentation and where rerunning an analysis produces a different number.

The intended user is a team, not an individual. The README frames the goals as maintainable data engineering and data science code and enhanced team collaboration. If you work alone on a throwaway analysis, the structure Kedro imposes is cost with no return. If two or more people touch the same pipeline, or if the pipeline has to run on a schedule after the person who wrote it has moved to another project, the trade flips.

## Nodes, the Data Catalog and how a Kedro pipeline actually runs

A Kedro pipeline is a graph of nodes. A node is a pure Python function plus declarations of which data it consumes and which it produces. The README describes this as "automatic resolution of dependencies between pure Python functions". You do not write the execution order; the framework derives it from the names of the inputs and outputs.

Those names are not variables in memory. They are entries in the Data Catalog, a set of lightweight connectors that save and load data across file formats and file systems, including local and network file systems, cloud object stores and HDFS. The catalog also provides data and model versioning for file-based systems. So the data flow is: a node asks the catalog for a dataset by name, the catalog resolves that name to a concrete loader, the function returns a value, and the catalog writes it back under the output name. Swapping a local CSV for a cloud object store is a catalog edit, not a code change.

Two consequences follow. First, purity is load-bearing. A node function that reads a file directly, or writes to a path it was not given, bypasses the graph and breaks the dependency resolution the framework is selling. Second, the catalog is a configuration surface, and configuration surfaces drift. The project template keeps these definitions in YAML under the conf directory, which is also where environment-specific overrides live.

## Installing Kedro and running a first pipeline

The README gives two installation routes. From PyPI with uv:

```bash
uv pip install kedro
```

Or with conda:

```bash
conda install -c conda-forge kedro
```

The package requires Python 3.10 or later, per the classifiers and requires-python field in pyproject.toml. The README also documents installing from source when you want an unreleased version:

```bash
uv pip install git+https://github.com/kedro-org/kedro@main
```

That last command pulls the main branch, so treat it as a moving target rather than a pinned dependency. The README points to the Get Started guide for full instructions, including how to set up Python virtual environments, which is where you should look for the project scaffolding command rather than guessing at it.

Once a project exists, the README directs new users to the spaceflights tutorial for hands-on experience. That tutorial is the honest starting point: it walks through building a project rather than describing one. For visual inspection of the resulting graph, the README points to Kedro-Viz, a separate project in the same organisation, and to documentation on working with Kedro and Jupyter notebooks if you want to keep an interactive loop.

## Where Kedro gets in the way

The project template is based on Cookiecutter Data Science, and it is opinionated. Directories, configuration layout and the separation between code and conf are decided for you. Teams that already have a working repository structure will spend real time reconciling the two, and the framework does not offer a lighter mode that keeps the node graph without the template.

Purity is the second constraint. Nodes are supposed to be functions of their inputs. In practice, data science code reaches for global state, random seeds, model objects loaded from disk and environment variables. Each of those has to be routed through the catalog or configuration to stay inside the model, and every shortcut you take is a place where the graph stops describing what actually runs.

There is also a deployment story to read carefully. The README lists flexible deployment with strategies including single or distributed-machine deployment and additional support for Argo, Prefect, Kubeflow, AWS Batch and Databricks. Those are integrations, and the README does not document what happens when a deployment target changes its API or when a plugin lags a Kedro release. If your platform is not on that list, you are writing the deployment layer yourself.

Finally, the repository's licence metadata is inconsistent. The README badge and the pyproject.toml licence field both say Apache 2.0, but the repository-level licence identifier is NOASSERTION. That is a signal to read LICENSE.md directly rather than trusting a badge.

## Kedro versus dbt: two different layers

The comparison people search for is Kedro against dbt, and the honest answer is that they solve different problems. dbt is a SQL transformation tool. Its unit of work is a SELECT statement against tables in a warehouse, and its dependency graph is resolved from ref() calls inside SQL. Kedro's unit of work is an arbitrary Python function, and its graph is resolved from named dataset inputs and outputs that may be files, dataframes or model objects.

That means dbt is the better tool when your transformation logic is expressible in SQL and lives in a warehouse. You get the warehouse's optimiser, its scheduling and its lineage for free. Kedro is the better tool when the work is Python: feature engineering that needs pandas or numpy, model training, or a step that calls an external service. The two can coexist, with dbt handling warehouse-side transforms and Kedro orchestrating the Python stages around them. Choosing Kedro for pure SQL work adds a Python runtime and a catalog layer to do what dbt already does.

## Maintenance, releases and the upgrade bill

The repository is not archived, and the last push was on 2026-09-10. The release cadence visible in the release list is roughly every six to ten weeks: 1.4.0 on 2026-05-22, 1.5.0 on 2026-06-29, and 1.6.0 on 2026-09-09. That is frequent enough that pinning matters. A pipeline that tracks the framework loosely will absorb breaking changes on someone else's schedule.

The dependency list in pyproject.toml is broad for a framework of this kind: attrs, click, cookiecutter, dynaconf, fsspec, gitpython, more_itertools, omegaconf, parse, pluggy, PyYAML, rich, tomli and tomli-w, plus typing_extensions. Each is a version constraint you inherit, and each is a potential conflict with an existing environment. The plugin system is built on pluggy, which is why the ecosystem extends through separate packages such as kedro-datasets rather than through the core. That is good for the core's size and bad for version alignment: a plugin that has not been updated for 1.6.0 is your problem, not the framework's.

The Makefile shows the project's own expectations of contributors: pre-commit hooks, mypy in strict mode over the kedro package, pytest run with four workers, and behave for end-to-end tests. Matching that bar in your own fork is a real cost if you intend to patch the framework rather than use it.

## Kedro's telemetry and what ships by default

One detail worth noticing before you install: kedro-telemetry is a direct dependency in pyproject.toml, not an optional extra. Installing Kedro installs the telemetry package. The README does not discuss what is collected or how to disable it, so if your organisation has rules about outbound telemetry from developer machines, check the telemetry package's own documentation before rolling Kedro out across a team. This is the kind of thing that is cheap to resolve at evaluation time and expensive to discover during a security review.

## Conclusion

Adopt Kedro if more than one person has to run, review or hand over your pipeline, and if you accept its project layout as the cost of that reproducibility. Skip it for exploratory work that will never be rerun, and for teams unwilling to keep node functions pure. Before committing, install it with uv pip install kedro, scaffold a project, and check that the Data Catalog connectors you need for your storage exist in kedro-datasets. Then read the pyproject.toml licence field: the repository ships Apache 2.0 text while the package metadata declares the Apache Software License, so confirm the terms with your own legal review rather than assuming.

## FAQ

### What is Kedro used for?

It is a Python toolbox for building production-ready data engineering and data science pipelines, with a project template, a Data Catalog for loading and saving data, and automatic dependency resolution between pure Python functions. The README states it was created to address the shortcomings of notebooks, one-off scripts and glue code.

### How do I install Kedro?

The README gives two routes: uv pip install kedro from PyPI, or conda install -c conda-forge kedro. It also documents installing from source with uv pip install git+https://github.com/kedro-org/kedro@main for unreleased versions.

### What is a Kedro pipeline?

A pipeline is a graph of nodes, where each node is a pure Python function with declared inputs and outputs. Kedro resolves the execution order automatically from those declarations, and the names refer to Data Catalog entries rather than in-memory variables.

### How do I use Kedro-Viz?

Kedro-Viz is a separate project in the kedro-org organisation that generates pipeline visualisations. The README links to a dedicated documentation section on visualising Kedro projects with it, and shows an example visualisation, but does not give install or launch commands.

### How do I use Kedro?

The README points to the documentation, which explains installation and then introduces key Kedro concepts, followed by the spaceflights tutorial for hands-on experience. There is also a section on working with Kedro and Jupyter notebooks, plus advanced user guides and API reference documentation.

### What is the Kedro framework?

Kedro is an open-source Python framework hosted by the LF AI & Data Foundation. Its stated purpose is to help you create data engineering and data science pipelines that are reproducible, maintainable and modular.

## Sources

- [Issues](https://github.com/kedro-org/kedro/issues)
- [kedro-org/kedro on GitHub](https://github.com/kedro-org/kedro)
- [Project website](https://kedro.org)
- [README](https://github.com/kedro-org/kedro/blob/main/README.md)
- [Releases](https://github.com/kedro-org/kedro/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/kedro-org-kedro
