Model or dataset
aimclub/FEDOT avatar
aimclub/FEDOT

FEDOT: evolutionary AutoML for composite pipelines

Automated modeling and machine learning framework FEDOT

711 stars98 forksPythonBSD-3-Clause

At a glance

What is it?
FEDOT builds machine learning pipelines as graphs and evolves them with a genetic algorithm. It suits engineers who need structural search over preprocessing and models, not just hyperparameter tuning.
Who is it for?
Use FEDOT when the pipeline structure itself is the unknown: mixed preprocessing plus several model families, time series with gaps, or hybrid models you want to wrap as custom operations. Skip it if you only need to tune hyperparameters of one gradient boosting model, since the evolutionary search adds cost without changing the structure.
Can I use it commercially?
Yes. BSD-3-Clause is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What FEDOT solves and who it is aimed at

Most AutoML tools search over hyperparameters inside a fixed model. FEDOT searches over the structure of the pipeline itself. The README describes it as a framework for "automatic generative design of machine learning pipelines", where the pipeline is represented as a graph whose nodes are data preprocessing blocks and model blocks, and whose edges define how data flows between them. The framework supports classification (binary and multiclass), regression, clustering, and time series prediction.

The audience is narrower than "anyone doing machine learning". FEDOT is for engineers and researchers who already know which model families are plausible but do not know how to wire them together: which imputation step before which encoder, whether a dimensionality reduction block belongs before a classifier, whether a time series model should be combined with a regression model on exogenous features. The project is maintained by the Natural Systems Simulation Lab at ITMO University, and the topics list includes genetic-programming and structural-learning alongside automl, which tells you where the emphasis sits.

If your problem is a single tabular dataset and a single gradient boosting model, FEDOT is more machinery than you need. If your problem is a pipeline you keep rebuilding by hand for each new dataset, the graph representation is the part that pays off.

How the evolutionary search over pipeline graphs works

The mechanism is a genetic algorithm operating on graphs. A population of candidate pipelines is generated, each one evaluated on the training data, and the better candidates are used to produce the next generation through mutation and crossover of the graph structure. The README calls this the core of FEDOT and states that the optimization algorithms are data- and task-independent.

That independence is the design decision worth noticing. The same search procedure is reused across tabular, text, image, and time series problems; what changes is the set of available operations and the strategies applied. The documentation points to special strategies for time-series forecasting, NLP, and tabular data, which adjust how the search behaves rather than replacing it.

Two consequences follow. First, the search space is defined by the operations you make available, so a small operation set produces fast, shallow searches and a large one produces slower, more varied graphs. Second, because the search is structural, it is not a replacement for hyperparameter tuning: the README lists hyperparameter tuning as a separate capability with its own methods, custom evaluation metrics, and search spaces. In practice you are running two nested searches, and the timeout you pass to the API bounds the whole thing.

The framework integrates Scikit-learn, CatBoost, XGBoost, and others as pipeline blocks, and the documentation describes an extension point for custom operations, so a model that FEDOT does not ship with can still appear in the graph.

Installing FEDOT and running a first classification pipeline

The README gives pip as the simplest installation path. The base package is installed with a single command:

bash
pip install fedot

Optional dependencies for image processing, text processing, and deep neural networks are installed with the extra marker:

bash
pip install fedot[extra]

Docker images also exist; the README points to the docker directory in the repository for the available tags rather than listing them inline. Note the Python requirement from pyproject.toml: Python 3.10 through 3.14. The dependency list is long and includes several version pins, so installing into a fresh virtual environment is the safer route.

Once installed, the high-level API is three calls. Import the class, construct it with a problem type and a timeout, then fit and predict:

python
from fedot.api.main import Fedot

model = Fedot(problem='classification', timeout=5, preset='best_quality', n_jobs=-1)
model.fit(features=x_train, target=y_train)
prediction = model.predict(features=x_test)
metrics = model.get_metrics(target=y_test)

This is the example the README gives, with x_train, y_train, and x_test as numpy arrays. Input can also be a Pandas DataFrame or a file path. The timeout parameter is in minutes and bounds the optimization, so a five-minute run returns whatever the best pipeline found so far is, not necessarily a converged one. After fit returns, the object holds the composite pipeline; predict applies it to new features and get_metrics scores the predictions against a target you supply.

The repository ships an examples directory with simple and advanced subdirectories, and the README points to a separate notebooks repository for tutorials covering AutoML basics, time series forecasting, gap-filling, and hybrid modelling with custom models.

Where FEDOT is the wrong tool

The evolutionary search is the cost. Every candidate pipeline in every generation has to be fitted and evaluated, and the number of candidates is bounded by the timeout rather than by a convergence criterion you control directly. On a large dataset, a short timeout can end the run before the population has moved far from its initial random graphs. The README does not document what happens to the returned pipeline in that case beyond the fact that fit returns the resulting composite pipeline, so you should treat the timeout as a hard budget and check the resulting graph rather than assuming it is the best achievable.

A second limitation is the dependency surface. The requirements file pins hyperopt to exactly 0.2.7, caps scipy below 1.13.0 for Python versions before 3.13, caps scikit-learn below 1.7.0 on the same range, and pins pyvis to exactly 0.2.1. Projects that already depend on newer versions of those libraries will have to resolve conflicts. The numpy constraint also differs by Python version, with a 2.x line only enabled from 3.13 onward.

A third case: if your pipeline is fixed by regulation, by a serving constraint, or by a vendor contract, structural search has nothing to offer. You would be paying for a search whose answer you already know. Similarly, if you need a model you can explain block by block to a reviewer, an evolved graph is harder to justify than a pipeline you wrote yourself, even though FEDOT can export the graph for inspection.

FEDOT compared with fixed-pipeline AutoML

The natural comparison is with AutoML systems that keep the pipeline shape fixed and optimize hyperparameters inside it. Tools built on that model search a much smaller space, which makes them faster per unit of data and easier to reason about: the output is always the same sequence of steps, with different parameter values. FEDOT's answer is the opposite trade. It spends its budget on deciding which steps exist at all, and it accepts that some of those evaluations will be wasted on structures that a human would have ruled out immediately.

The practical difference shows up when preprocessing is the hard part. If your dataset needs imputation, categorical encoding, and a scaling step, and you are not sure of the order or whether all three belong, a fixed-pipeline tool leaves that decision to you. FEDOT makes it part of the search. The README's framing of the key feature is exactly this: "complex management of interactions between various blocks of pipelines", represented as a graph.

FEDOT also overlaps with general-purpose hyperparameter optimizers, since it wraps Optuna, hyperopt, SALib, and scikit-optimize as tuning backends. The distinction is that those libraries optimize a function you define, while FEDOT optimizes a structure it generates. For a single model with a handful of parameters, calling the optimizer directly is less indirection.

One more comparison point is reproducibility. FEDOT can export a resulting pipeline as JSON on its own, or together with the input data as a ZIP archive. That matters when you need to hand a result to someone else or re-run an experiment later, and it is a capability the README lists explicitly rather than leaving implicit.

Maintenance, release cadence, and licence

The repository is not archived, and the last push was on 2026-09-09. The most recent release listed is v0.7.5 from 2025-03-10, preceded by v0.7.4 in August 2024 and v0.7.3.2 in May 2024. The gap between the last tagged release and the last commit is roughly six months, so the version you install from PyPI may lag the master branch. If you need a fix that landed after v0.7.5, installing from the repository rather than from the package index is the route, and you should expect to track master yourself.

Upgrade cost is driven by the pinned dependencies more than by FEDOT's own API. The high-level Fedot class with fit, predict, and get_metrics has been stable across the 0.7.x line as far as the README shows, but the requirements file pins several libraries to ranges that will need attention on each upgrade: scipy below 1.13.0 and scikit-learn below 1.7.0 for Python versions before 3.13, and exact pins on hyperopt 0.2.7 and pyvis 0.2.1. An environment that shares dependencies with other ML tooling will feel those constraints.

The licence is BSD-3-Clause, declared in pyproject.toml and stated in the README. That is a permissive licence, which generally means you can use, modify, and redistribute the code including in commercial products, subject to the conditions in the licence text. This is not legal advice; read LICENSE.md and, if the distinction matters to your organisation, have counsel review it. Note that BSD-3-Clause covers FEDOT itself, not the third-party libraries it depends on, each of which carries its own terms.

Editorial conclusion

Use FEDOT when the pipeline structure itself is the unknown: mixed preprocessing plus several model families, time series with gaps, or hybrid models you want to wrap as custom operations. Skip it if you only need to tune hyperparameters of one gradient boosting model, since the evolutionary search adds cost without changing the structure. Before adopting it, install the pinned version and run the fit/predict example on your own data with a fixed timeout, then export the resulting pipeline to JSON and check that the graph matches operations you can actually serve in production. Note that the dependency set pins hyperopt to exactly 0.2.7 and caps scipy below 1.13.0 on Python before 3.13, so resolve those against your existing environment first.

Frequently asked questions

What is FEDOT?

FEDOT is an open-source framework for automated modeling and machine learning, distributed under the 3-Clause BSD licence. Its core is an evolutionary approach that designs machine learning pipelines as graphs and supports classification, regression, clustering, and time series prediction.

How do I install FEDOT?

The README gives pip install fedot as the simplest path, with pip install fedot[extra] adding optional dependencies for image and text processing and for deep neural networks. Docker images are also available through the docker directory in the repository.

Which Python versions does FEDOT support?

pyproject.toml declares requires-python as >=3.10,<3.15, and the classifiers list Python 3.10 through 3.14. Some dependencies differ by Python version, for example numpy 2.x is enabled only from Python 3.13 onward.

Can FEDOT export the resulting pipeline?

Yes. The README states that resulting pipelines can be exported separately as JSON or together with the input data as a ZIP archive, which supports experiment reproducibility. The documentation links to pages on pipeline import/export and project import/export.

Official sources

  1. aimclub/FEDOT on GitHub
  2. License: BSD-3-Clause
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/aimclub-fedot.svg)](https://hysenlabs.com/projects/aimclub-fedot)