gplearn: Symbolic Regression and Feature Engineering via Genetic Programming
Genetic Programming in Python, with a scikit-learn inspired API
At a glance
- What is it?
- gplearn is a Python library that implements genetic programming for symbolic regression, binary classification, and automated feature engineering, with a scikit-learn-compatible fit/predict API. It evolves populations of mathematical programs rather than fitting fixed model architectures, returning interpretable expressions instead of black-box weights.
- Who is it for?
- gplearn suits researchers and data scientists who need interpretable mathematical models from tabular data and are willing to trade training time for symbolic output. The main constraint is scope: the library is intentionally limited to symbolic regression, binary classification, and feature transformation, and the pyproject.toml lists it at Development Status Alpha.
- Can I use it commercially?
- Yes. BSD-3-Clause is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 46 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What Symbolic Regression Is and Why gplearn Exists
Most machine learning models fit parameters to a fixed architecture: a neural network has layers and weights, a random forest has split thresholds. Symbolic regression takes a different approach. It searches the space of mathematical expressions, building a population of candidate formulas and evolving them toward the one that best describes the relationship in the data. The output is a formula you can read, like a decision tree but for continuous relationships.
Genetic programming is the search mechanism: candidate programs are represented as trees, and each generation selects the fittest trees to undergo mutation and crossover. gplearn implements this process in Python and packages it as scikit-learn estimators, so the same fit/predict interface works for gplearn models as for any other scikit-learn model. The library was built to bring genetic programming into the scikit-learn ecosystem, accepting its constraints (focused scope, consistent API) rather than building a standalone framework.
Three Estimators: Regressor, Classifier, and Transformer
gplearn provides three estimators. SymbolicRegressor finds a mathematical expression that minimizes a regression loss on the training data. SymbolicClassifier handles binary classification problems. SymbolicTransformer does not produce a label or value directly; instead, it generates new features by evolving symbolic expressions of the input variables. The README describes SymbolicTransformer as "designed to support regression problems, but should also work for binary classification."
All three estimators plug into scikit-learn pipelines and the grid search module without modification. This means a gplearn estimator can be embedded in a Pipeline alongside scalers and preprocessors, and its hyperparameters can be tuned with GridSearchCV or RandomizedSearchCV the same way any other scikit-learn estimator's can. The README notes that there are many parameters to tweak and directs users to the documentation at gplearn.readthedocs.io for guidance on which are most relevant for a given problem.
Installation and Dependencies
gplearn is distributed as a Python package. The pyproject.toml in the repository root declares the following requirements:
- Python 3.11 or later - scikit-learn 1.8.0 or later - joblib 1.3.0 or later
Installation details are documented at the project homepage: http://gplearn.readthedocs.io/en/stable/installation.html. The README points there directly rather than listing installation commands inline. The build system is hatchling. The most recent release is version 0.4.3, published on 2026-01-07, and is available at pypi.python.org/pypi/gplearn.
The scikit-learn version floor (1.8.0) is a meaningful constraint. Projects pinned to an older scikit-learn release will need to upgrade before adding gplearn as a dependency, which may require testing for compatibility with other scikit-learn-dependent libraries in the same environment.
How the Genetic Programming Loop Works
The README describes the mechanism directly: gplearn "begins by building a population of naive random formulas to represent a relationship between known independent variables and their dependent variable targets in order to predict new data. Each successive generation of programs is then evolved from the one that came before it by selecting the fittest individuals from the population to undergo genetic operations."
The genetic operations are the standard tree-based ones: mutation replaces a subtree with a new random subtree, crossover swaps subtrees between two parent programs, and selection applies fitness pressure to prefer programs with lower loss. The function set (the set of mathematical operations available to the evolved programs) is configurable. The README labels the parameter governing this as "function set" and notes it as one of the key configuration choices. A function set that includes trigonometric functions will evolve different expressions than one restricted to arithmetic operators.
Parallelization is handled through joblib, which the pyproject.toml lists as a direct dependency. This means gplearn can parallelize the fitness evaluation of each population member across CPU cores using the same parallel backend that scikit-learn uses, without requiring a separate configuration step.
Where gplearn Falls Short
The library is intentionally constrained. The README states that gplearn is "purposefully constrained to solving symbolic regression problems," motivated by the scikit-learn ethos of powerful but straightforward estimators. Multi-class classification is not supported through SymbolicClassifier. Unsupervised tasks, time-series modeling, and survival analysis are outside the scope.
The pyproject.toml development status is Alpha. This signals that the API may still change across releases and that the library is not positioned as a production-hardened tool. The current release (0.4.3) followed 0.4.2 after nearly four years, which suggests a slow release cadence.
Symbolic regression is also computationally expensive relative to gradient-based methods. Evaluating a large population of tree programs across training data is slower than fitting a linear model or a gradient boosted tree, and the cost scales with population size, number of generations, and dataset size. For large datasets where interpretability is not a strict requirement, gradient boosted trees or neural networks will train much faster.
gplearn Versus DEAP for Genetic Programming in Python
DEAP (Distributed Evolutionary Algorithms in Python) is the most commonly referenced alternative for genetic programming in Python. The key difference is scope. DEAP is a general evolutionary computation framework that supports genetic algorithms, genetic programming, evolution strategies, and other evolutionary methods across a broad range of problem types. It provides the building blocks (populations, operators, fitness functions) but leaves the user to assemble them.
gplearn is narrower and more opinionated: it provides finished estimators for symbolic regression and classification, inherits the scikit-learn API, and handles the evolutionary loop internally. A user who wants to run symbolic regression with a familiar fit/predict interface will find gplearn easier to apply immediately. A user who wants to evolve programs for a custom problem type, or who needs evolutionary methods beyond GP, will need DEAP's greater generality. The related searches list for gplearn explicitly includes "gplearn vs deap," which confirms this is the comparison readers most often want to make.
Maintenance and License
The repository last received a push on 2026-08-14 and is not archived. Release 0.4.3 was published on 2026-01-07; the previous release, 0.4.2, was from 2022-05-03, a gap of roughly four years between minor versions. The gap between 0.4.2 and 0.4.3 coincides with the Python 3.11 requirement being formalized and the scikit-learn minimum being raised to 1.8.0, suggesting maintenance is focused on keeping the library compatible with the current Python and scikit-learn ecosystem rather than on adding new features.
The license is BSD-3-Clause, which permits use, modification, and redistribution with attribution. It is compatible with both commercial and academic projects. The project is listed under Development Status Alpha in the pyproject.toml classifiers, which should be weighed alongside the stable release history when assessing adoption risk.
Editorial conclusion
gplearn suits researchers and data scientists who need interpretable mathematical models from tabular data and are willing to trade training time for symbolic output. The main constraint is scope: the library is intentionally limited to symbolic regression, binary classification, and feature transformation, and the pyproject.toml lists it at Development Status Alpha. Teams that need multi-class classification, neural networks, or gradient boosting are looking at the wrong library. Before committing, verify that your scikit-learn version meets the 1.8.0 minimum requirement specified in the pyproject.toml, and check the readthedocs documentation for current parameter guidance, since the library exposes many configuration knobs.
Frequently asked questions
Is there a library for genetic algorithms in Python?
gplearn implements genetic programming (tree-based GP) in Python specifically for symbolic regression, binary classification, and feature engineering, with a scikit-learn-compatible API. DEAP is an alternative that covers a broader range of evolutionary computation methods beyond GP.
What is symbolic regression and how does it work?
Symbolic regression is a machine learning technique that searches the space of mathematical expressions to find one that best describes the relationship between variables. gplearn implements it by evolving a population of candidate formulas through genetic operations (mutation and crossover), selecting fitter programs each generation until a satisfactory expression is found.
Can gplearn handle multi-class classification?
No. gplearn's SymbolicClassifier supports binary classification only. The README does not document multi-class support, and the library's intentional scope is limited to symbolic regression, binary classification, and feature transformation.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/trevorstephens-gplearn)