SymbolicRegression.jl: Distributed Equation Discovery in Julia
Distributed High-Performance Symbolic Regression in Julia
At a glance
- What is it?
- SymbolicRegression.jl searches for analytic expressions that fit data, using a genetic algorithm that can run across threads, processes and machines. It is aimed at people who need a readable formula rather than a black-box predictor, and who are willing to pay for that in setup complexity.
- Who is it for?
- Adopt SymbolicRegression.jl if you have tabular data, a small set of operators you believe are physically or structurally meaningful, and a need to read the resulting equation rather than call a model. Do not adopt it if you need a fixed, fast inference artifact, if your feature count is in the thousands, or if your team has no Julia in its stack and no appetite to add it.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 1 day ago.
- What is it written in?
- Mainly Julia, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
DEEP OPEN-SOURCE ANALYSIS
What SymbolicRegression.jl actually searches for
Most regression libraries return coefficients. You choose the functional form, and the fit adjusts numbers inside it. SymbolicRegression.jl inverts that: you supply operators, and the search returns the form itself. The README states the goal plainly, that the package "searches for symbolic expressions which optimize a particular objective." The objective is a fit to your target, and the search space is every expression you can build from the binary and unary operators you declare.
The audience follows from that. If you are doing scientific modelling and you want a formula you can write on a whiteboard, print in a paper, or hand to a domain expert for a sanity check, this is the tool's purpose. If you only need a number out the other end, gradient boosting or a neural network will get there with far less ceremony. The project also sits under an explainable-ai and equation-discovery framing in its own topics, which matches how it is used: as a discovery step, not a serving layer.
The cost is that the output is not a fixed artifact. You get a set of candidate expressions, and you choose among them. That choice is a modelling decision the library does not make for you.
Populations, operators and the Pareto front
The mechanism is a genetic algorithm over expression trees. You declare the operator set, the search evolves populations of candidate expressions, and each candidate is scored by a loss against your target. The README's low-level example sets populations=20, so the search runs twenty separate populations rather than one, which is the standard defence against premature convergence in this kind of search.
The important structural idea is the Pareto front. The README says you can view "the dominating Pareto front (best expression seen at each complexity)" using calculate_pareto_frontier. That means the search does not return a single winner. It returns, for each level of expression complexity, the best-fitting expression found at that complexity. A two-node expression and a twenty-node expression both appear, and you decide where the accuracy you need stops justifying the extra terms.
Expressions are represented as a Node type from DynamicExpressions.jl, wrapped by an Expression that carries operator and variable-name metadata. Expression objects are callable, so a discovered tree can be evaluated directly on new data. The README notes one behavioural detail worth knowing: the callable form "will automatically set all values to NaN if there were any Inf or NaN during evaluation," while the raw eval_tree_array returns an output plus a did_succeed flag. If you need to distinguish a failed evaluation from a genuine NaN, use the raw form.
Variable names come from your input container. Pass a NamedTuple or a Tables.jl-compatible table and the printed expressions use your column names; pass a plain matrix and you get x1, ..., xn. That naming is cosmetic to the search but matters a great deal when you are reading the output.
Installing SymbolicRegression.jl and running a first fit
The README gives the install as a standard Julia package add. From a Julia session:
using Pkg
Pkg.add("SymbolicRegression")The README's Quickstart builds a small synthetic problem with two named features, a target built from a cosine and a square, and additive noise. Note that the operators you pass are the entire vocabulary of the search, so the target must be reachable from them. Here the target uses cos, multiplication, subtraction and a power, but only +, -, * and cos are declared, which is exactly the kind of mismatch that decides whether a run succeeds.
import SymbolicRegression: SRRegressor, machine, fit!, predict, report
X = (a = rand(500), b = rand(500))
y = @. 2 * cos(X.a * 23.5) - X.b ^ 2
y = y .+ randn(500) .* 1e-3
model = SRRegressor(
niterations=50,
binary_operators=[+, -, *],
unary_operators=[cos],
)The machine wrapper binds the model to data, and fit! runs the search. After that, report(mach) prints the discovered expressions and predict(mach, X) evaluates the one selected by model.selection_method, which the README describes as "a mix of accuracy and complexity" by default. You can bypass that selection and pick from the Pareto front yourself by index:
mach = machine(model, X, y)
fit!(mach)
report(mach)
predict(mach, X)
predict(mach, (data=X, idx=2))Two details are easy to miss. First, the low-level equation_search interface expects column-major input of shape [features, rows], which is the opposite convention from the table-like inputs the high-level interface accepts. Second, for multiple outputs the README points to MultitargetSRRegressor, with an array of indices passed to idx for per-output equation selection. The same machine, fit!, predict and report functions are also exposed through MLJ if you want pipelines and tuning from that ecosystem.
Where the search breaks down
The first limitation is the operator set, and it is not a soft one. If the true relationship needs a function you did not declare, the search cannot express it, and it will happily return a more complex expression that approximates your data without capturing the mechanism. Adding operators widens the space, which makes the search slower and increases the chance of a spurious fit, so there is no free move here.
The second is dimensional cost. The search evaluates many candidate expressions against the full dataset across many populations and iterations. The README's own examples use a few hundred rows and a handful of features. Nothing in the documentation suggests the method scales to wide feature matrices, and the expression trees grow combinatorially with the number of variables. This is a tool for tens of features, not thousands.
The third is reproducibility of the output as an artifact. You get expressions, and the README shows exporting to SymbolicUtils.jl, but the package is not a serialization or serving format. If your deployment pipeline expects a frozen model object with a stable interface, you are building that layer yourself.
Finally, the default selection is a heuristic. The README describes selection_method as a mix of accuracy and complexity, which means the equation you get from predict is not necessarily the one you would have picked after looking at the Pareto front. For any result you intend to publish, read the front and choose explicitly.
PySR and the Python frontend question
The README points to PySR as "a Python frontend" to this package, and the discussions link in the README header goes to the PySR repository's discussions rather than a separate forum for the Julia package. That tells you where the project expects most of its user conversation to happen.
The difference is not algorithmic. PySR is a frontend, so the search itself is the Julia code. What changes is the surrounding workflow. If your data loading, preprocessing and downstream analysis already live in Python, PySR keeps you in one language and you pay a process boundary between the two runtimes. If you are already in Julia, or you want the low-level equation_search interface and direct manipulation of Expression and Node objects, the Julia package is the shorter path and gives you access to the internal types that a frontend would have to re-expose.
The trade-off is real in both directions. A Python user who installs the Julia package directly inherits a second language runtime and its package manager for no benefit. A Julia user who reaches for PySR adds a Python dependency to call code that is already available natively. Pick the side your data pipeline lives on.
Maintenance, licensing and the upgrade surface
The repository is not archived, and the last push was on 2026-09-10, the same day as the v2.4.1 release. Releases v2.3.0 and v2.4.0 landed on 2026-09-06, so the project is being cut frequently rather than sitting idle. That cadence is a maintenance cost as well as a signal: a CHANGELOG.md sits at the repository root, and release-please-config.json plus .release-please-manifest.json indicate automated release tooling, so the changelog is the place to check before upgrading rather than the commit history.
The licence is Apache-2.0. That is a permissive licence with an explicit patent grant, which matters more here than in a typical library because the output is a formula that may end up in a patent or a paper. Nothing in the repository indicates a copyleft obligation on the expressions you discover, but licence questions about generated output are fact-specific, and this is not legal advice.
On upgrade cost, the surface you depend on is the Options struct and the regressor constructors. The README notes that expressions are represented by the Node type from DynamicExpressions.jl, a separate package, so a version bump in that dependency can reach your code even when SymbolicRegression.jl itself is unchanged. If you build on the low-level interface, pin both.
Editorial conclusion
Adopt SymbolicRegression.jl if you have tabular data, a small set of operators you believe are physically or structurally meaningful, and a need to read the resulting equation rather than call a model. Do not adopt it if you need a fixed, fast inference artifact, if your feature count is in the thousands, or if your team has no Julia in its stack and no appetite to add it. Before committing, verify three things against your own data: that your operators close over the values the search will generate, that your target has a signal a short expression can capture, and that the number of rows you can afford to evaluate per iteration keeps the search within your time budget. Start from the Quickstart example in the README, run it once with niterations=50, and inspect report(mach) before you tune anything.
Frequently asked questions
What is symbolic regression?
It is a search for a symbolic expression, built from a set of operators you choose, that optimizes a fit to your target data. SymbolicRegression.jl performs this search and returns the best expression found at each level of complexity along a Pareto front.
Is symbolic regression NP hard?
The repository does not make any complexity claim, and the README describes the search as a genetic algorithm over populations rather than stating a theoretical bound. The practical consequence documented in the README is that you control cost through niterations and populations.
How can Julia be used in machine learning?
SymbolicRegression.jl is one example: it is a Julia package installed with Pkg.add, and its machine, fit!, predict and report functions are also available through MLJ for compatibility with the broader MLJ ecosystem. It targets equation discovery rather than general predictive modelling.
What is symbolic regression and how does it work in Python?
The README points to PySR as a Python frontend to this package, so the Python route runs the same Julia search behind a Python interface. The README's own examples are written in Julia, using SRRegressor or the lower-level equation_search.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/astroautomata-symbolicregression-jl)
Community notes