PySR: symbolic regression in Python that hands you an equation, not a black box
High-Performance Symbolic Regression in Python and Julia
At a glance
- What is it?
- PySR wraps a Julia search engine behind a scikit-learn style Python API to fit interpretable expressions to low-dimensional data. Here is how it installs, how the search works, and where it stops being the right tool.
- Who is it for?
- Adopt PySR when you have a small number of input features and a reason to want the model written out as arithmetic: a physical law to recover, a neural network to distill, a simulator you want to compress. Do not adopt it as a general tabular regressor.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository received new commits within the last day.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem PySR solves, and who actually has it
Most regression libraries return a function you cannot read. A gradient-boosted tree with a few hundred splits will predict well and explain nothing. PySR targets the case where the fitted function itself is the deliverable: you want an expression made of operators you chose, short enough to write on a whiteboard, and accurate enough to be useful. The README frames this as symbolic regression, "a machine learning task where the goal is to find an interpretable symbolic expression that optimizes some objective".
The intended user is a scientist or engineer with a handful of input variables and a suspicion that a compact relationship exists. The README's own demo generates five features and a target built from cos, a square and a constant. That is the shape of problem PySR is built for. It is not a drop-in replacement for a tabular regressor on a 200-column dataset, and the README is explicit that symbolic regression works best on low-dimensional data.
A second use case gets less attention but is more interesting: symbolic distillation of neural networks. The README points to arXiv:2006.11287, where the technique converts a neural net into an analytic equation. If you have trained a network and want a human-readable surrogate, PySR is one route to it.
Inside the search: a Python wrapper over a Julia engine
PySR is two programs. The Python package you import is an interface layer; the search itself runs in SymbolicRegression.jl, which the README calls "the powerful search engine of PySR". Calling fit on a PySRRegressor launches a Julia process, and that process performs a multithreaded evolutionary search over candidate expressions.
The configuration you pass describes the search space rather than a model architecture. binary_operators and unary_operators list the functions the search may compose. maxsize caps expression size. niterations sets how long the search runs. The README describes the default 40 as a starting point and annotates it with a comment telling you to increase it for better results. Each iteration performs what the README calls "hundreds of thousands of mutations and equation evaluations", so the cost scales with the operator set you allow, not just with the number of rows.
The output is not one model but a Pareto front. After fitting, model.equations_ holds a table with columns for pick, score, equation, loss and complexity, ordered from simplest to most complex. You choose the row. That table is the real interface, and reading it is the skill PySR demands of you. The score column is what drives the automatic selection used by model.predict(X); passing an index, as in model.predict(X, 3), overrides that and evaluates the third equation instead.
Custom operators are declared in Julia syntax inside the Python call. The README's example defines inv(x) = 1/x as a unary operator and then supplies a matching extra_sympy_mappings entry so the Python side can still manipulate the resulting expression symbolically. The same Julia-syntax convention applies to elementwise_loss. This is a real seam in the API: you are writing Julia strings inside a Python constructor, and a typo surfaces at search time rather than at parse time.
Installing PySR and running a first fit
The documented install is a single pip command. Julia dependencies are not bundled; the README states they are installed at first import, so the first run of any PySR script is slower than later ones and needs network access.
pip install pysrA conda-forge package exists as well: conda install -c conda-forge pysr. For containers, the repository ships a Dockerfile with three targets: pysr-runtime (Python, Julia and PySR, without eager Julia precompilation), pysr (a dev image that installs extras and precompiles), and pysr-slurm. The README's Docker path is to clone the repo, build with docker build -t pysr ., then start a container with docker run -it --rm pysr ipython. On a cluster without root, the README offers Apptainer instead, built with apptainer build --notest pysr.sif Apptainer.def and launched with apptainer run pysr.sif.
Once installed, the quickstart is short. Generate data, build the regressor, fit it.
import numpy as np
from pysr import PySRRegressor
X = 2 * np.random.randn(100, 5)
y = 2.5382 * np.cos(X[:, 3]) + X[:, 0] ** 2 - 0.5
model = PySRRegressor(
maxsize=20,
niterations=40,
binary_operators=["+", "*"],
unary_operators=["cos", "exp", "sin", "inv(x) = 1/x"],
extra_sympy_mappings={"inv": lambda x: 1 / x},
)
model.fit(X, y)Equations print during training, and the README notes you can quit early by pressing q followed by enter. When the run finishes, print(model) shows the equations_ table, with each row giving a complexity and a loss. The README recommends IPython over Jupyter for this step because the printing is nicer, which is a small thing but accurate: the table is wide and the default notebook renderer does not do it justice.
Where PySR breaks down or is the wrong choice
The dimensionality ceiling is the first limit and the README states it plainly. Every additional feature multiplies the number of plausible expressions, and the operator set compounds that. If your dataset has dozens of columns, the search will spend its iterations on combinations that have no chance of being the answer. Feature selection before PySR is not optional at that point; it is the actual work.
The second limit is the Julia dependency. PySR is a Python package that cannot run without a Julia toolchain, and the README's troubleshooting section describes a failure that is worse than an error message: a hard crash at import reporting that GLIBCXX_... is not found. The cause is another Python dependency loading a mismatched libstdc++ library. The documented fix is to prepend the Julia libstdc++ directory to LD_LIBRARY_PATH, and the README warns that the path likely differs on your system. This is a configuration problem you can solve, but it is not a problem you would have with a pure-Python estimator, and on a locked-down environment you may not be able to change that variable at all.
The third limit is interpretive, not technical. A Pareto front is an answer only if you know what to do with it. The score column encodes a preference for accuracy against complexity, and that preference is a modelling decision PySR makes on your behalf. If the automatically selected equation disagrees with your domain knowledge, the table gives you no way to encode that knowledge except by restricting the operator set and rerunning. Nothing in the README describes a constraint mechanism for known-signed coefficients or monotonicity.
How PySR differs from a standard regressor and from a genetic-programming library
Against scikit-learn's GradientBoostingRegressor or a random forest, the difference is what comes out. Those return a fitted object whose predictions you can inspect and whose internals you mostly cannot. PySR returns a list of candidate formulas ranked by complexity, and the fit method deliberately mirrors the scikit-learn API so the switch costs little in code. It does cost something in runtime: a boosted model trains in seconds on the README's 100-row, 5-feature example, while PySR runs an evolutionary search over hundreds of thousands of candidate expressions per iteration. You are trading compute for legibility, and that trade is only worth making when legibility is the point.
Against general-purpose genetic programming frameworks, the difference is the search engine and the output format. PySR delegates to SymbolicRegression.jl and returns SymPy-compatible expressions, which means the result plugs into the Python symbolic stack rather than sitting in a bespoke tree structure. The extra_sympy_mappings parameter exists precisely to keep that bridge intact when you add custom operators. If your goal is to evolve arbitrary programs rather than to recover a compact equation, a general GP library is the better fit; PySR is narrow on purpose.
There is also a lighter-weight option worth naming: if you want the search engine without the Python layer, SymbolicRegression.jl is the Julia library PySR is built on, developed alongside it. Choosing between them is mostly a question of which language the rest of your pipeline lives in.
Maintenance, releases and what the licence means for you
The repository is not archived, and the last push was on 2026-09-09. Releases are frequent and versioned: v2.2.0 on 2026-09-02, v2.2.1 the same day, v2.3.0 on 2026-09-07. The pyproject.toml in the repository declares version 2.4.0, ahead of the latest tagged release listed, which is normal for a project that bumps the version file before tagging. The release configuration files (release-please-config.json and .release-please-manifest.json) indicate automated release management, and a CHANGELOG.md sits at the repository root.
Python support starts at 3.10 per requires-python. The dependency pins are conservative in places and loose in others: sympy is held below 2.0.0, numpy below 3.0.0, scikit_learn below 2.0.0, while juliacall is pinned to a narrow range between 0.9.29 and 0.9.36. That narrow juliacall pin is the one to watch. It is the bridge to Julia, and it means a future juliacall release will not be picked up until PySR widens the range.
The licence is Apache-2.0, declared both in the LICENSE file and in the pyproject.toml classifier. Apache-2.0 is permissive and includes an explicit patent grant, which matters if you are embedding PySR in a commercial product. It also carries attribution and notice requirements, so if you redistribute a modified copy you need to keep the notices intact. None of that is unusual, and none of it is legal advice: read the LICENSE file before you ship.
One more cost worth naming. The README asks that you cite arXiv:2305.01582 if you find PySR useful, and that you submit a pull request to the research showcase page when a project is finished. Those are requests, not licence terms, but in academic settings the citation expectation is real and you should plan for it.
Editorial conclusion
Adopt PySR when you have a small number of input features and a reason to want the model written out as arithmetic: a physical law to recover, a neural network to distill, a simulator you want to compress. Do not adopt it as a general tabular regressor. On wide datasets the search space explodes, and a gradient-boosted model will beat it on accuracy while PySR spends its budget wandering. The maintainers say so themselves: the README states that symbolic regression works best on low-dimensional datasets. Before committing, run the quickstart example on your own data with niterations raised well past the default 40 and check the equations_ table: if the complexity column climbs without the loss column falling, the search is not converging on your problem and no amount of tuning will fix the framing. Also verify the Julia toolchain resolves on your platform, since the first import triggers that install and a libstdc++ mismatch surfaces as a hard crash rather than a warning.
Frequently asked questions
What is a symbolic regression model?
It is a model whose output is a mathematical expression rather than a set of fitted weights. PySR's README defines symbolic regression as a machine learning task where the goal is to find an interpretable symbolic expression that optimizes some objective.
How can machine learning be used in astronomy?
PySR's README points to arXiv:2006.11287, which applies symbolic distillation of neural networks to N-body problems. The technique converts a trained neural network into an analytic equation, which makes the model inspectable rather than opaque.
What does the PySR quickstart example look like?
The README generates 100 datapoints with 5 features from 2.5382 * cos(x3) + x0^2 - 0.5, builds a PySRRegressor with maxsize=20 and niterations=40, then calls model.fit(X, y). After fitting, print(model) shows the equations_ table of candidate expressions with their loss and complexity.
How do I install PySR?
The documented install is pip install pysr, with Julia dependencies installed at first import. A conda-forge package is also available via conda install -c conda-forge pysr, and the repository ships a Dockerfile and an Apptainer.def for container use.
Why does importing PySR crash with a GLIBCXX error?
The README attributes this to another Python dependency loading an incorrect libstdc++ library, producing a hard crash at import. The documented fix is to prepend the Julia libstdc++ directory to LD_LIBRARY_PATH, and the README notes that the path likely differs on your system.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/astroautomata-pysr)