Library / SDK
smartcorelib/smartcore avatar
smartcorelib/smartcore

smartcore: reading the manifest against the release tags and the benchmark alerts

A comprehensive library for machine learning and numerical computing. Apply Machine Learning with Rust leveraging first principles.

961 stars110 forksRustApache-2.0

At a glance

What is it?
The crate's own files disagree in small ways that decide what a build actually gets: a manifest version one tag ahead of the newest release, a feature listed and deprecated together, a serde flag that resolves differently by target, and two benchmark thresholds where one failure is advisory and the other is not.
Who is it for?
smartcore is a reasonable pick for a Rust codebase that needs classical supervised and unsupervised algorithms and wants a linear algebra abstraction of its own instead of binding to ndarray directly, with clustering, decomposition, linear models, forests, SVMs, neighbors, four Naive Bayes variants, preprocessing and model selection under one trait system. Check three things before pinning a version.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly Rust, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 3, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Cargo.toml reads 0.6.16 while the newest tag is v0.6.15

The manifest at the root of smartcorelib/smartcore declares `version = "0.6.16"`, and the published tag list stops one step behind at v0.6.15, tagged 2026-09-27. Below that sit v0.6.14 on 2026-08-26 and v0.6.13 on 2026-08-25, two patch releases one day apart, followed by a four week gap before v0.6.15. The last push recorded for the repository is 2026-10-02 and it is not archived, so the 0.6.16 in the tree is work in flight rather than something the tag history has caught up with. Two install lines sit side by side, and they hand you different things:

toml
[dependencies]
smartcore = "^0.6"
toml
[dependencies]
smartcore = { git = "https://github.com/smartcorelib/smartcore", branch = "main" }

The caret line resolves to the newest published 0.6 release, which is v0.6.15, and cannot reach 0.6.16 until a tag appears. The git line follows main, so it is the one that can carry the untagged 0.6.16 and whatever the 2026-10-02 push changed. Choosing between them is a choice about whether you want released code or current code, and the files give no signal about which of those is the safer default for you. Toolchain requirements are pinned in the same manifest: `edition = "2024"` alongside `rust-version = "1.85"`, so a build on an older compiler stops before it reaches any smartcore source. The change notes confirm the edition move as recent work, which makes 1.85 a floor that moved upward rather than a fixed target.

ndarray-bindings is offered and marked superseded on the same page

The install notes name three optional features and annotate the third in the same breath: datasets, serde, and `ndarray-bindings (deprecated in favor of ndarray-only support per recent changes)`. The manifest has not caught up with that annotation. It still declares the feature as a live switch, with a comment describing what it wires up:

toml
# Optional bindings for the ndarray crate (DenseMatrix <-> ndarray).
ndarray-bindings = ["dep:ndarray"]

The backing dependency `ndarray = { version = "0.17", optional = true }` is present in the dependency table too, where the optional surface is easy to bound: approx, cfg-if, num, num-traits and ordered-float are unconditional, and ndarray, serde and typetag are the only three entries marked optional. Elsewhere the same README says the crate moved to ndarray only, and that features such as nalgebra-bindings have been dropped in favor of ndarray-only paths. So a reader scanning the feature list finds one entry that is simultaneously offered as a way in and flagged as a way out, while the path it points toward is not spelled out as a feature name of its own. The install section closes by pointing readers at Cargo.toml for available features and compatibility notes, which is where the annotation and the declaration can be compared directly. Enabling the flag still works in the tree as it stands; what neither file settles is the intended end state, which release removes it, and whether code written against the bindings needs a different entry point afterwards.

The serde feature resolves to two different dependency sets by target

The feature is declared as `serde = ["dep:serde", "dep:typetag"]`, and its comment says it also enables typetag on non-wasm targets for trait object serialization. That qualifier is enforced by placement rather than by any conditional expression inside the feature list. `typetag` sits in a target specific dependency table:

toml
[target.'cfg(not(target_arch = "wasm32"))'.dependencies]
typetag = { version = "0.2", optional = true }

So a build for wasm32 resolves the same feature name without typetag appearing in the graph. The unconditional part of that table is short: `approx = "0.5.1"`, `cfg-if = "1.0.0"`, `num = "0.4"`, `num-traits = "0.2.12"` and `ordered-float = "5.1.0"`, with `rand` configured separately. A wasm32 build therefore carries numeric comparison, the ordered float wrapper and the allocation only generator, and nothing past that unless a feature asks for it. The portable half is deliberate: `default = []`, and the manifest explains that no features are enabled by default to keep the build WASM compatible and to avoid pulling in serde or rand unless a caller asks for them. The README frames the same decision as a WASM/WASI first posture, notes that some file system operations are restricted on wasm targets, and tells readers to enable serde selectively to minimize footprint. The practical consequence is that the trait object machinery the comment describes is present on a host build and absent from a wasm32 build of the same feature, so serialization code that relies on it compiles for one target and not the other.

Which random generator you get is a feature decision

The random number dependency is declared without most of its usual parts: `rand = { version = "0.10.1", default-features = false, features = ["alloc"] }`. What fills that gap depends on `std_rand`, whose comment says it uses the standard library RNG (StdRng / thread_rng) instead of SmallRng and is required for non-deterministic seeding without an explicit seed. Its own entry lists `rand/std_rng` and `rand/std`, followed by a third entry written as `rand/thread_` where the file stops. In a default build the small generator is therefore what backs the algorithms, and turning on `std_rand` is what makes the standard library generator available for a run that varies between executions rather than one driven by a seed you pass in. The change notes list seeds and deterministic controls across algorithms using RNG plumbing as recent work, so this layer is still moving and the feature boundary may shift. One coupling to watch: `datasets = ["std_rand", "serde"]`, so turning on bundled datasets switches both the generator and the serialization stack as a side effect.

The quick start predicts on the matrix it just fitted

The example the README opens with builds five two column rows, labels them, fits a KNN classifier with `Default::default()` and then predicts on the same matrix it trained on:

rust
use smartcore::linalg::basic::matrix::DenseMatrix;
use smartcore::neighbors::knn_classifier::KNNClassifier;

// Turn vector slices into a matrix
let x = DenseMatrix::from_2d_array(&[
    &[1., 2.],
    &[3., 4.],
    &[5., 6.],
    &[7., 8.],
    &[9., 10.],
]).unwrap();

// Class labels
let y = vec![2, 2, 2, 3, 3];

// Train classifier
let knn = KNNClassifier::fit(&x, &y, Default::default()).unwrap();

// Predict
let yhat = knn.predict(&x).unwrap();

Two things about that snippet are worth stating plainly. The prediction call reuses `x`, so the labels it returns are not held out estimates of anything; each row is scored against its own nearest neighbours inside the training set. And `Default::default()` supplies the neighbor count and the distance metric without either being written down, so anyone copying this does not know which `k` produced the output. The crate does ship the parts for a real split: CSV readers take a configurable delimiter and header rows, the generators make_blobs, make_circles and make_moons produce synthetic sets, and model selection provides K-fold and search parameters for the K-Means and SVM families.

Two benchmark pages with two different consequences for CI

Benchmark output is published by the benchmark-action/github-action-benchmark workflow onto the gh-pages branch through GitHub Pages, as two separate chart pages. The criterion page under smartcore-benches/dev plots wall-clock time per bench in nanoseconds per iteration across history, where lower is better, and its alert threshold of 200% is advisory: it posts a comment and does not fail CI. The iai-callgrind page under smartcore-benches/iai-dev plots retired instructions, labelled `Ir`, which the README describes as deterministic and machine independent, also lower is better, and its 120% threshold does fail the iai job and the status check. The gap between those two numbers is the part worth planning around: the timing suite can double and stay green, while the instruction count suite fails at twenty percent. Each page is a searchable line chart with one series per benchmark name, with `matmul/1024` and `iai_matmul::matmul::bench_matmul_256` given as examples, and `data.js` holds the raw history behind the chart. The chart section names those two series and then its closing line stops mid sentence.

The dataset lists in the README and the manifest do not match

Both files name built in sample data and they disagree. The README's data access section offers digits, diabetes, breast cancer and boston behind a feature gate, with serialization utilities to persist or refresh .xy bundles. The manifest's comment on the same feature names a different set: sample datasets (iris, boston, digits, ...) and dataset generators. Iris appears in the manifest comment and nowhere in the README's list. Diabetes and breast cancer appear in the README and nowhere in the comment. Boston is the single name both files carry. Neither file says which list is current, and the README answers the compatibility question by sending readers to Cargo.toml rather than resolving it. Enabling the feature also reaches past data access, because `datasets = ["std_rand", "serde"]`: the bundles ship serialized, so the flag turns on the standard library random generator and the serialization stack at once, a coupling the quick start example never surfaces.

Per-class metric bookkeeping was collapsed into one helper

The most specific entry in the README's change notes concerns classification metrics: multiclass macro F1 now averages per-class F-measures, which the file says matches sklearn, and Precision, Recall and F1 share a single per-class confusion-counts helper so the per class bookkeeping lives in one place. Both statements describe a change in what an existing metric returns rather than a new metric being added, which matters for anyone comparing numbers across versions. That sits inside a larger reorganization: a trait-system refactor with fewer structs and more object-safe traits, a move to the Rust 2024 edition with MSRV 1.85, and cleanup of duplicate code paths. Recent additions named alongside it include XGBoost-style regression, single-linkage clustering, Extra Trees, SVC and SVR with a kernel enum, multiclass SVM support, and a search parameter API for the K-Means and SVM families. The published package is narrower than the repository, because the manifest's exclude list drops `.github`, `.gitignore`, `smartcore.iml`, `smartcore.svg`, `tests/` and `AGENTS.md`, so the tests directory sitting at the repository root is not inside the crate a consumer downloads.

Editorial conclusion

smartcore is a reasonable pick for a Rust codebase that needs classical supervised and unsupervised algorithms and wants a linear algebra abstraction of its own instead of binding to ndarray directly, with clustering, decomposition, linear models, forests, SVMs, neighbors, four Naive Bayes variants, preprocessing and model selection under one trait system. Check three things before pinning a version. First, choose the source deliberately: the caret line stops at v0.6.15, while the git line on main carries the untagged 0.6.16 and the push recorded on 2026-10-02, and the two differ by exactly the work that has not been released. Second, settle the features from your target rather than convenience, because the default set is empty by design, serde gains typetag only off wasm32, datasets pulls in both serde and std_rand, and ndarray-bindings is still declared while the guide already calls it superseded. Third, do not read the quick start as an evaluation: it predicts on the matrix it fitted, with a defaulted neighbor count, so use the CSV readers, the dataset generators and the K-fold utilities to get an honest split. Leave it alone if you need deep learning, a GPU execution path, or a boosting implementation with its own tuning surface, none of which appear in this crate.

Frequently asked questions

Which smartcore version does the caret install line actually give me?

The line `smartcore = "^0.6"` resolves to the newest published 0.6 release, which is v0.6.15 from 2026-09-27. The Cargo.toml in the tree carries `version = "0.6.16"`, and that number has no tag yet, so it reaches you only through the git line pointing at the main branch.

Does smartcore build for WebAssembly without extra setup?

The feature table ships `default = []`, and the manifest comment says no features are enabled by default to keep the build WASM compatible and avoid pulling in serde or rand. The README notes that some file system operations are restricted on wasm targets, and the `serde` feature only gains typetag outside wasm32 because typetag is declared under a non-wasm32 target table.

Why does my smartcore build pick a different random generator than my colleague's?

The `rand` dependency is declared with `default-features = false` and only the alloc feature, and the `std_rand` feature is what switches from SmallRng to the standard library RNG (StdRng / thread_rng). That feature is also required for non-deterministic seeding without an explicit seed. Turning on `datasets` turns on `std_rand` and `serde` as well.

Do smartcore's tests and notebooks ship inside the published crate?

They do not. The Cargo.toml exclude list names `tests/` and `AGENTS.md`, so neither is packaged, even though both sit at the repository root. The notebooks live in a separate companion repository, smartcorelib/smartcore-jupyter, and the README recommends EVCXR to run Rust notebooks locally.

Is the smartcore ndarray-bindings feature safe to turn on now?

The feature is still declared in the manifest as `ndarray-bindings = ["dep:ndarray"]` over `ndarray = { version = "0.17", optional = true }`, so it resolves as written. The README's install list labels the same feature deprecated in favor of ndarray-only support, and the change notes say bindings such as nalgebra-bindings were dropped in favor of ndarray-only paths, so neither file states which release removes it.

Official sources

  1. License: Apache-2.0
  2. Project website
  3. README
  4. Releases
  5. smartcorelib/smartcore on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/smartcorelib-smartcore.svg)](https://hysenlabs.com/projects/smartcorelib-smartcore)