Framework
rust-ml/linfa avatar
rust-ml/linfa

linfa: scikit-learn's shape, Rust's guarantees, eighteen algorithm crates

A Rust machine learning framework.

4,751 stars333 forksRustApache-2.0

At a glance

What is it?
linfa, Italian for sap, is a dual MIT and Apache-2.0 Rust machine learning framework aiming to provide a comprehensive toolkit, kin in spirit to Python's scikit-learn and focused on preprocessing and classical algorithms. Eighteen sub-packages span Naive Bayes through t-SNE, most marked tested and several benchmarked, with pluggable BLAS and LAPACK backends including Intel MKL and a wasm-bindgen feature for the browser.
Who is it for?
Use linfa when a Rust application needs classical machine learning, clustering, regression, classification or preprocessing, with memory safety and no interpreter runtime, and when scikit-learn's algorithm vocabulary maps to the problem. Look to Python for deep learning and the long tail of research methods, this is deliberately the everyday ML toolkit.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 40 days ago.
What is it written in?
Mainly Rust, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Sap: the name says the ambition

The project introduces itself with a dictionary entry, linfa, Italian, translating to sap in English, the vital circulating fluid of a plant, a name that frames the library as nourishment for a growing ecosystem rather than a finished product. The stated aim is to provide a comprehensive toolkit to build machine learning applications with Rust, and the reference point is named directly, kin in spirit to Python's scikit-learn, focusing on common preprocessing tasks and classical ML algorithms for everyday ML tasks. That positioning is a boundary, no deep learning frameworks competing with PyTorch, no exotic research methods, the algorithms a working engineer reaches for weekly, implemented in a language whose type system and memory safety change what an ML library can guarantee about its own correctness.

Eighteen crates, one status column

The algorithm table is the state of the union, eighteen sub-packages each with a purpose, a status, a category and notes. The supervised learning side includes bayes with Bernoulli, Gaussian and Multinomial Naive Bayes, ensemble with bagging, random forest and AdaBoost, elasticnet, lars with Least Angle Regression, linear with Ordinary Least Squares and Generalized Linear Models, logistic for two-class models, pls for Partial Least Squares, svm for classification and regression, and trees for linear decision trees. Unsupervised work covers clustering with K-Means, Gaussian Mixture Models, DBSCAN and OPTICS, hierarchical agglomerative clustering, ica with FastICA, and tsne with both the exact solution and the Barnes-Hut approximation. Every entry is marked Tested, and four carry Tested and Benchmarked, clustering, ftrl, nn and trees, the honest split between verified and measured. The notes column carries the practical detail a category cannot, for example that the clustering crate covers unlabeled data while the tsne crate distinguishes its exact and approximate solvers, distinctions that decide which crate a problem needs before any documentation is opened.

Preprocessing as first class crates

Three sub-packages serve the preprocessing category, and their presence in the table reflects the scikit-learn lineage. kernel maps feature vectors into higher dimensional space, the transformation that makes linear methods behave nonlinearly. nn provides nearest neighbours and distances, spatial index structures and distance functions that other algorithms consume, benchmarked alongside its tested status. preprocessing itself carries normalization and whitening plus count vectorization and tf-idf, the text pipeline basics, and reduction holds diffusion mapping, Principal Component Analysis and random projections for dimensionality reduction. A framework that ships preprocessing beside algorithms avoids the situation where users adopt the library for a model and reach for another dependency to feed it, the integration cost linfa's design tries to eliminate. Because these crates sit in the same repository and release together, their APIs are held to one consistent fit and transform shape, the scikit-learn convention that lets a preprocessing step chain into a model without adapter code between them.

The ftrl crate and partial fit

One row breaks the classification into a category of its own, ftrl, Follow The Regularized Leader proximal, categorized as partial fit rather than supervised learning. Its notes explain the distinction, it contains L1 and L2 regularization with possible incremental update, meaning the model learns from streams rather than requiring the full dataset in memory, the online learning mode that production systems feeding on continuous data need. Logistic regression is also marked partial fit, so the two crates cover the incremental training cases in the toolkit. The categorization itself is documentation, a user scanning for batch algorithms and stream algorithms finds the split encoded in the table rather than buried in per-crate docs.

BLAS backends, static or system

Some algorithm crates need an external library for linear algebra routines, and by default linfa uses a pure-Rust implementation, with the blas feature flag and a backend feature opting into external libraries. The choice spans openblas, netblas, that is Netlib, and intel-mkl, and the compatibility table is candid, OpenBLAS and Netlib support Linux only, while Intel MKL covers Linux, Windows and macOS. Each backend offers two features, linking the system library or statically building it, intel-mkl-static and intel-mkl-system being the documented example, and the Cargo manifest shows the same six features wired to ndarray-linalg's corresponding options. The pure Rust default keeps builds dependency free, and the MKL path buys performance at the cost of a proprietary library, with the static variant solving deployment and the system variant sharing memory. The are-we-learning-yet site that the current state section links tracks the wider Rust ML ecosystem, and linfa occupies its classical corner there, the reference point newcomers check before evaluating any of the crates individually.

WASM in the browser, explicitly

Browser support is a feature flag rather than an afterthought, for browser-style WASM on the wasm32-unknown-unknown target, enable linfa's wasm-bindgen feature. The manifest shows the mechanism, the feature pulls in getrandom with its js feature, since random number generation is the platform-sensitive dependency that differs between native and browser environments, and a separate serde feature enables serialization through ndarray's serde support. The combination means classical models, a trained clustering or a logistic regression, can run inside a web page without a server round trip, the use case where Rust's small footprint and WebAssembly's sandbox meet machine learning, and where linfa's classical rather than deep learning scope fits the browser's constraints.

A community argument, and the dataset crates

The README closes its state-of-the-union with an argument rather than a feature, we believe that only a significant community effort can nurture, build and sustain a machine learning ecosystem in Rust, there is no other way forward, with the roadmap issue linked and an invitation to get involved. The infrastructure around that belief is visible in the repository, a datasets crate providing winequality, iris, diabetes and generated datasets for examples and tests, a CHANGELOG tracking the release train through 0.8.1 on 2025-12-23 after 0.8.0 in September 2025, dual licensing files for MIT and Apache-2.0, and CI workflows named Codequality Lints and Run Tests. Four named authors, Luca Palmieri among them, maintain the core, with the community chat on Zulip and the website carrying the rendered documentation. The mascot.svg in the repository root and the rendered website give the project an identity beyond its crates, a small investment in approachability that classical ML libraries in younger language ecosystems usually skip.

Editorial conclusion

Use linfa when a Rust application needs classical machine learning, clustering, regression, classification or preprocessing, with memory safety and no interpreter runtime, and when scikit-learn's algorithm vocabulary maps to the problem. Look to Python for deep learning and the long tail of research methods, this is deliberately the everyday ML toolkit. Before adopting, check the algorithm table's status column for what is tested versus benchmarked, choose the BLAS story deliberately, pure Rust by default or OpenBLAS, Netlib or Intel MKL by feature with static and system variants, and enable the wasm-bindgen feature explicitly for browser targets.

Frequently asked questions

What is linfa?

linfa is a Rust machine learning framework, dual licensed under MIT and Apache-2.0, kin in spirit to Python's scikit-learn and focused on common preprocessing tasks and classical ML algorithms. It ships eighteen algorithm sub-packages including Naive Bayes, K-Means, SVM, decision trees, PCA and t-SNE, with pluggable BLAS backends and WASM support.

Which algorithms does linfa include?

The eighteen crates cover bayes, clustering with K-Means, Gaussian Mixture Models, DBSCAN and OPTICS, ensemble with bagging, random forest and AdaBoost, elasticnet, ftrl, hierarchical clustering, ica, kernel methods, lars, linear regression with OLS and GLM, logistic regression, nearest neighbours, partial least squares, preprocessing, dimensionality reduction with PCA and diffusion mapping, svm, trees and tsne with exact and Barnes-Hut variants.

Does linfa need BLAS?

Not by default, a pure-Rust implementation serves the linear algebra routines. Enabling the blas feature with openblas, netblas or intel-mkl switches to an external library, each with static and system variants, noting that OpenBLAS and Netlib support Linux only while Intel MKL covers Linux, Windows and macOS.

Official sources

  1. Issues
  2. License: Apache-2.0
  3. README
  4. Releases
  5. rust-ml/linfa on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/rust-ml-linfa.svg)](https://hysenlabs.com/projects/rust-ml-linfa)