Library / SDK
abess-team/abess avatar
abess-team/abess

abess: best-subset selection for regression in Python and R

Fast Best-Subset Selection Library

472 stars43 forksC++NOASSERTION

At a glance

What is it?
abess implements a polynomial-time algorithm for best-subset selection across linear, logistic, Poisson, Cox and multi-task models. It is built for statisticians and ML engineers who want a sparse model with a chosen support size, not a penalty path.
Who is it for?
abess suits analysts who need an exact support of size k and can accept tuning over the support size. It is the wrong tool for pure prediction pipelines where a penalized path is enough.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Activity is slowing. The repository last received commits 6 months ago.
What is it written in?
Mainly C++, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem abess solves: choosing k predictors instead of shrinking all of them

Most sparse regression libraries in Python and R solve a penalized problem. You set a penalty strength, and the number of nonzero coefficients falls out of that choice. abess takes the opposite route. The README describes the goal as finding "a small subset of predictors such that the resulting model is expected to have the highest accuracy." The user controls the size of the subset, and the algorithm searches for the best one at that size.

That distinction matters when the size itself is the deliverable. A clinician reading a gene-expression model wants to know how many genes are in the panel before deciding whether the panel is practical. A penalty parameter is harder to justify to that audience than a fixed count. abess is aimed at that setting: scientific work and applied modelling where an interpretable, size-constrained subset is the output.

The library covers a wider family of models than the name suggests. The README lists linear regression, binary and multi-class classification, counting-response modelling, censored-response modelling, multi-response modelling, and Ising model estimation. It also supports group best-subset selection and nuisance penalized regression. For linear regression the README states that the time complexity of (group) best subset selection is certifiably polynomial, which is the central claim of the project.

How the algorithm and the two-language layout fit together

The repository is organized around a shared C++ core. The top-level entries include src/, include/, python/ and R-package/, with docs/ holding the tutorials and simulation scripts. The Python and R packages are interfaces over the same compiled implementation rather than two independent ports. That is why the two quick-start examples produce comparable behaviour on the same simulated data.

The workflow is a fit-then-inspect loop. You construct a model object, call fit on a feature matrix and a response, and read the selected support from the fitted object. The README's Python example generates data with make_glm_data(n = 300, p = 1000, k = 10, family = "gaussian") and then fits a LinearRegression model. The generator takes the true support size k, so the simulated problem has a known answer to compare against.

The generic framework is what lets one codebase serve so many response types. Each model family supplies its own likelihood and gradient, and the shared search procedure handles the combinatorial part. The trade-off is visible in the API: model families are separate classes under abess.linear (and the corresponding namespaces), so switching from Gaussian to logistic means switching class, not just a parameter.

Group selection and nuisance penalization are exposed as variants of the same interface rather than separate packages. That keeps the surface small, but it also means the documentation for advanced variants sits alongside the basic examples, and the README does not spell out how the nuisance and group options interact in a single call.

Installing abess and fitting a first linear model

The README gives two Python installation routes. From PyPI:

bash
pip install abess

Or from conda-forge:

bash
conda install abess

For R, the stable version is on CRAN:

r
install.packages("abess")

Once installed, the README's Python quick start builds a simulated problem with 300 samples, 1000 candidate predictors and a true support of 10, then fits a linear model:

python
from abess.linear import LinearRegression
from abess.datasets import make_glm_data
sim_dat = make_glm_data(n = 300, p = 1000, k = 10, family = "gaussian")
model = LinearRegression()
model.fit(sim_dat.x, sim_dat.y)

After fit returns, the fitted object holds the selected predictors and coefficients. The README does not print the attribute names for the support, so check the Python tutorials linked from the README before writing code that consumes the result. The R equivalent is shorter, because the top-level abess function wraps the fit:

r
library(abess)
sim_dat <- generate.data(n = 300, p = 1000)
abess(x = sim_dat[["x"]], y = sim_dat[["y"]])

Note the asymmetry: the R generator is called generate.data and takes n and p, while the Python generator is make_glm_data and takes n, p, k and family. If you work across both languages, do not assume the argument names carry over.

Where abess is the wrong tool

The clearest limitation is scope. abess targets best-subset selection under the model families the README enumerates. If your problem is a gradient-boosted tree, a neural network, or a survival model that is not a Cox proportional-hazards fit, the library does not address it. There is no claim in the README of a general-purpose estimator.

A second constraint is the search over support size. Best-subset selection at a fixed k is one problem; choosing k is another. The README does not describe a cross-validation helper or an automatic selection rule for the subset size. Users coming from glmnet-style workflows, where a single call returns a path over many penalty values, should expect to manage the size sweep themselves.

The polynomial-time guarantee is stated for linear regression, and specifically for (group) best subset selection in that setting. The README does not extend the complexity claim to the censored, counting or multi-task families. Treat the guarantee as scoped to what it says, and do not assume it transfers.

Finally, the README's runtime section compares abess against scikit-learn, glmnet, ncvreg and L0Learn on synthetic data. Those are the project's own measurements on its own simulation scripts. They are reproducible by running the scripts, but they are not an independent evaluation, and the README does not report memory behaviour at the scale of the p = 1000 example.

abess compared with glmnet and L0Learn

The README's own comparison set is the honest place to start. For R, it benchmarks against glmnet, ncvreg and L0Learn. For Python, it benchmarks against scikit-learn. Those are the alternatives the maintainers consider relevant.

The difference in approach is structural. glmnet solves a convex penalized problem along a regularization path, which makes it fast and stable but leaves the support size as an output rather than an input. L0Learn also targets sparse regression, but the README positions abess against it on runtime rather than on the shape of the solution. scikit-learn's linear models offer Lasso and related penalties, again path-based.

abess differs by making k the control. If you need a model with exactly ten features, abess expresses that directly. With a path-based method you either pick the penalty that yields ten features after the fact, or you refit. That is the practical trade: abess gives you the constraint you asked for, and in exchange you give up the single-call path that penalized methods hand you.

The README does not document how abess behaves when the true support is larger than the requested k, or how stable the selected set is across resamples. Those are the questions to answer with your own data before committing.

Maintenance, releases and the GPL v3 licence

The repository is not archived. The last push was on 2026-03-15. The most recent release listed is 0.4.11, dated 2026-03-04, which followed 0.4.8 in April 2024 and 0.4.7 in September 2023. The gap between 0.4.8 and 0.4.11 is roughly two years, so the release cadence has been uneven rather than steady.

The README links a GPL v3 badge, and the repository's LICENSE file is the place to confirm the exact terms. The GitHub metadata reports the licence as NOASSERTION, which means the automated classifier did not match a standard identifier. That mismatch is worth resolving before you embed the library in a product, because GPL v3 carries distribution obligations that permissive licences do not. This is a factual observation about the repository, not legal advice; if the terms matter to your distribution model, have someone qualified read the LICENSE file.

Upgrade cost is the usual one for a compiled extension. The Python and R packages both wrap the C++ core, so a version bump can change both the build requirements and the numerical results. The jump from 0.4.8 to 0.4.11 spans two years of changes, and the README's "What's New" section is the place to check what moved. Pin the version in your environment file and read the release notes before moving.

Editorial conclusion

abess suits analysts who need an exact support of size k and can accept tuning over the support size. It is the wrong tool for pure prediction pipelines where a penalized path is enough. Before adopting, verify that the model family you need is covered by the supported list and check the GPL v3 licence against your distribution constraints.

Frequently asked questions

How do I install abess in Python?

The README gives two routes: pip install abess from PyPI, or conda install abess from conda-forge. The R package installs from CRAN with install.packages("abess").

Which model families does abess support?

The README lists linear regression, binary and multi-class classification, counting-response modelling, censored-response modelling, multi-response modelling and Ising model estimation, plus group best-subset selection and nuisance penalized regression.

Is the abess algorithm guaranteed to run in polynomial time?

The README states that the time complexity of (group) best subset selection for linear regression is certifiably polynomial. It does not extend that guarantee to the other model families.

Official sources

  1. abess-team/abess on GitHub
  2. Issues
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/abess-team-abess.svg)](https://hysenlabs.com/projects/abess-team-abess)