# Featuretools: Automated Feature Engineering from Relational Data

> Featuretools generates features from multi-table datasets using the Deep Feature Synthesis algorithm. Point it at related tables with timestamps and relationships, and it creates hundreds of computed features automatically. Designed for machine learning workflows where feature creation is the bottleneck.

**alteryx/featuretools** — An open source python library for automated feature engineering

- Repository: https://github.com/alteryx/featuretools
- Website: https://www.featuretools.com
- Stars: 7,684 · Forks: 915
- Language: Python
- License: BSD-3-Clause
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/alteryx-featuretools

## Automated feature engineering with Deep Feature Synthesis

Featuretools is a Python library that generates features automatically from multi-table datasets. Given a collection of related tables (like customers, transactions, and orders joined by foreign keys), Featuretools creates a single flat feature table where each row is an entity (customer, product, account) and each column is a computed statistic. It uses the Deep Feature Synthesis algorithm to combine aggregations and transformations across table relationships. The tool targets data scientists who spend weeks manually writing SQL queries and Python code to transform raw data into a feature matrix for machine learning. Instead of creating features by hand, you define the structure of your data and let Featuretools generate hundreds of candidates. The library runs on Python 3.9 through 3.12. Core dependencies include pandas (version 2.0 or later), numpy, scipy, cloudpickle, and woodwork for data structure management.

## Deep Feature Synthesis: generating features from relationships

Deep Feature Synthesis (DFS) is the core algorithm. It discovers relationships between tables through foreign keys, then applies built-in primitives (functions) to create features. Primitives include aggregation functions like COUNT, SUM, MIN, MAX, MODE, MEAN, SKEW, and STD (standard deviation). It also includes time-based functions like YEAR and DAY that extract temporal components from dates. For example, with a customers table and a transactions table (linked by customer ID), DFS automatically creates features like COUNT(transactions), SUM(transactions.amount), MAX(transactions.amount), and MIN(transactions.amount) for each customer. These become columns in the output feature matrix. If transactions also links to a products table, DFS creates nested aggregations like SUM(transactions.products.price). You can extend DFS with custom primitives if the built-in set does not cover your domain. The README includes an example with a Predict Next Purchase demo using 3 million Instacart grocery orders, which shows how to build a feature matrix and machine learning pipeline that can be reused for multiple prediction problems.

## Installing from PyPI and generating features with demo data

Install Featuretools from PyPI with pip or from Conda-forge:

```bash
python -m pip install featuretools
```

or

```bash
conda install -c conda-forge featuretools
```

Load a demo dataset and generate features:

```python
import featuretools as ft
es = ft.demo.load_mock_customer(return_entityset=True)
feature_matrix, feature_defs = ft.dfs(entityset=es, target_dataframe_name="customers")
```

The dfs function returns a feature matrix (rows are customers, columns are computed features) and a list of feature definitions (what each column means). The output is a pandas DataFrame ready for machine learning. The example output shows columns like COUNT(transactions), COUNT(sessions), SUM(transactions.amount), MODE(sessions.device), MIN(transactions.amount), MAX(transactions.amount), YEAR(join_date), SKEW(transactions.amount), DAY(join_date), and many more. For custom data, you create an EntitySet by loading DataFrames and declaring relationships, then call dfs on it.

## Building entitysets from relational data

An entityset is Featuretools' representation of your data structure. You create one by loading DataFrames and declaring relationships via foreign keys. Featuretools then infers which table is the parent and which is the child. The tool maps table columns to Woodwork semantic types, enabling proper handling of categorical and numerical data. Time-based relationships (where a child row is associated with a parent at a specific time) enable time-aware feature creation, crucial for domains like transaction history or event logs. These relationships ensure that Featuretools only combines data that is temporally coherent. The entityset abstraction requires that your data fits the relational model: no nested arrays, no unstructured blobs. If your data is JSON or nested, you must flatten it first. The repository structure includes a featuretools/ directory with the core library and docs/ with documentation. A pyproject.toml file specifies dependencies and optional features.

## Add-on support for specialized cases

Featuretools has optional add-ons for extended functionality. Premium Primitives add domain-specific aggregations beyond the built-in set. NLP Primitives handle text features (word counts, TF-IDF). Dask Support lets you run DFS in parallel with njobs > 1 for large datasets, distributing computation across cores or clusters. Install add-ons as separate packages:

```bash
python -m pip install "featuretools[premium]"
python -m pip install "featuretools[nlp]"
python -m pip install "featuretools[dask]"
```

Or install all at once with featuretools[complete]. Premium Primitives come from a separate repository and may require additional setup. NLP Primitives support natural language processing operations on text columns. The Dask add-on is critical for scaling DFS to large datasets, allowing parallelization across multiple machines.

## Limitations and what Featuretools cannot do

Featuretools requires data to fit a relational model. Unstructured data (images, free text, audio) cannot be handled directly; you must extract or preprocess features first. The tool generates candidate features but does not select the best ones. You still need feature selection or regularization to avoid overfitting on high-cardinality features. Time-series data with many lags requires custom primitives or preprocessing. Featuretools does not scale to datasets with billions of rows without careful partitioning. For very large data, you may hit memory or performance limits. The primitives it generates are statistical summaries; if your domain requires domain-specific logic (e.g., financial ratios, biomedical indices), you must code those as custom primitives. The README notes that the library allows you to define your own custom primitives if the built-in set is insufficient.

## Comparison with tsfresh

tsfresh is another automated feature library, focused on time-series data. It extracts features from univariate time series by computing hundreds of statistics on rolling windows. Featuretools works with relational tables and creates aggregations across linked entities. If your data is a single time series (stock prices over time), tsfresh is simpler. If your data is a collection of time series (one per customer, one per sensor) linked to metadata (customer demographics, sensor location), Featuretools handles the relational structure while tsfresh does not. The libraries solve different problems: tsfresh focuses on temporal patterns, Featuretools focuses on entity relationships. Some workflows use both: Featuretools to generate relational features, then tsfresh on time-series sub-data.

## Maintained by Alteryx under BSD 3-Clause license

The last push was on 2026-09-11, indicating active maintenance. Recent releases are from 2024 (v1.31.0 in May 2024, v1.30.0 in February 2024), showing steady updates. Featuretools is maintained by Alteryx and is available under the BSD 3-Clause license, permitting commercial use with attribution. Support channels include Stack Overflow (tag: featuretools), GitHub issues, Slack (join.slack.com/t/alteryx-oss), and email (open_source_support@alteryx.com). The project is marked as Production/Stable in its development status classifiers. The repository includes a Makefile with targets for linting, testing, and packaging. The library also cites a peer-reviewed paper by Kanter and Veeramachaneni on Deep Feature Synthesis published at IEEE DSAA 2015.

## Conclusion

Adopt Featuretools if your data is already in a relational structure with timestamps and you need to generate baseline features for machine learning models without manual engineering. It is not for you if your data is unstructured (images, text, audio) or if you already have high-quality hand-crafted features. Before starting, verify that your table relationships can be expressed in the entityset abstraction and that the primitive set covers your domain.

## FAQ

### How do you define relationships in an entityset?

Create an EntitySet, load DataFrames into it, and declare foreign key relationships. Featuretools infers parent-child relationships: the child table has a column that references a primary key in the parent table. For time-based relationships, specify the time column on the child row that matches the parent.

### Can Featuretools handle text or image features?

No. Featuretools works with numerical and categorical data in relational tables. Unstructured data (text, images, audio) must be converted to features first through other tools or manual preprocessing.

### How do you select which features to use after DFS?

Featuretools generates candidates but does not select them. Use feature selection methods afterward: correlation analysis, mutual information, model-based selection, or regularization. High-cardinality features often need filtering to avoid overfitting.

### What is the difference between Featuretools and tsfresh?

Featuretools creates features from relational data by aggregating across linked tables. tsfresh extracts features from univariate time series. Use Featuretools for data with multiple related tables, tsfresh for single or multiple independent time series.

### Does Featuretools scale to large datasets?

Featuretools can handle medium-sized datasets in memory. For very large data (billions of rows), use the Dask add-on to parallelize DFS across machines. Memory and runtime grow with the number of tables and rows; benchmark on your data first.

### What Python versions does Featuretools support?

Featuretools requires Python 3.9 or later and supports 3.9, 3.10, 3.11, and 3.12. It also runs on Windows, Linux, and macOS.

## Sources

- [alteryx/featuretools on GitHub](https://github.com/alteryx/featuretools)
- [License: BSD-3-Clause](https://github.com/alteryx/featuretools/blob/main/LICENSE)
- [Project website](https://www.featuretools.com)
- [README](https://github.com/alteryx/featuretools/blob/main/README.md)
- [Releases](https://github.com/alteryx/featuretools/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/alteryx-featuretools
