# Rumale: a scikit-learn style machine learning library for Ruby

> Rumale brings estimators, cross-validation and dataset loaders to Ruby with an API borrowed from scikit-learn. It is a good fit for Ruby shops that want to keep training code in the same language as the rest of their application.

**yoshoku/rumale** — Rumale is a machine learning library in Ruby

- Repository: https://github.com/yoshoku/rumale
- Website: https://rubygems.org/gems/rumale
- Stars: 918 · Forks: 34
- Language: Ruby
- License: BSD-3-Clause
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/yoshoku-rumale

## The gap Rumale fills for Ruby applications

Ruby has web frameworks, background job systems and a mature package manager, but its machine learning options are thin compared with Python. Rumale exists to close that gap without asking a Ruby team to stand up a Python service. The README describes it as a machine learning library in Ruby whose interfaces are similar to scikit-learn in Python, and the name is a contraction of Ruby machine learning.

The intended audience is a Ruby developer who needs to classify rows, cluster records or reduce dimensionality inside an existing application. The README lists a broad set of algorithms: Support Vector Machine, Logistic Regression, Ridge, Lasso, Multi-layer Perceptron, Naive Bayes, Decision Tree, Gradient Tree Boosting, Random Forest, K-Means, Gaussian Mixture Model, DBSCAN, Spectral Clustering, Multidimensional Scaling, t-SNE, Fisher Discriminant Analysis, Neighbourhood Component Analysis, Principal Component Analysis and Non-negative Matrix Factorization. That is a conventional supervised and unsupervised toolkit rather than a research library. Nothing in the README mentions deep learning frameworks, GPU execution or pretrained model hubs.

## How the estimator API and data flow work

The design follows the fit and transform pattern that scikit-learn popularised. An estimator object is constructed with hyperparameters, then fit with samples and labels, then reused for prediction or transformation. The README's first example chains two estimators: an RBF kernel approximation maps the input into a higher dimensional feature space, and a linear support vector classifier is trained on the transformed data. Chaining is explicit, so the reader can see exactly which array flows into which object.

Data enters through Rumale::Dataset.load_libsvm_file, which reads a file in LIBSVM format and returns samples and labels. Evaluation is handled by separate objects under Rumale::EvaluationMeasure, such as Accuracy, and model selection lives under Rumale::ModelSelection with splitters like StratifiedKFold and a CrossValidation runner. The repository layout mirrors this decomposition: the top level contains rumale-core, rumale-linear_model, rumale-ensemble, rumale-clustering, rumale-decomposition, rumale-model_selection, rumale-preprocessing and other directories, one per algorithm family. That layout suggests the gem is assembled from those component gems rather than being a single monolith.

One structural note matters for anyone upgrading: the README states that since v2.0.0 Rumale uses Numo::NArray Alternative instead of Numo::NArray as a dependency. Code written against the older array class will need attention.

## Installing the rumale gem and training a first classifier

Installation is a normal Ruby gem install. The README gives both the Bundler route and the direct route. Add the gem to your Gemfile and run bundle, or install it directly:

```bash
gem install rumale
```

The README's first worked example uses the pendigits dataset from the LIBSVM Data site. Download the training and test files first:

```bash
wget https://www.csie.ntu.edu.tw/~cjlin/libsvmtools/datasets/multiclass/pendigits
wget https://www.csie.ntu.edu.tw/~cjlin/libsvmtools/datasets/multiclass/pendigits.t
```

Training then loads the file, maps the samples into an RBF kernel feature space, fits a linear SVM and writes both objects to disk with Marshal:

```ruby
require 'rumale'

samples, labels = Rumale::Dataset.load_libsvm_file('pendigits')

transformer = Rumale::KernelApproximation::RBF.new(gamma: 0.0001, n_components: 1024, random_seed: 1)
transformed = transformer.fit_transform(samples)

classifier = Rumale::LinearModel::SVC.new(reg_param: 0.0001)
classifier.fit(transformed, labels)

File.open('transformer.dat', 'wb') { |f| f.write(Marshal.dump(transformer)) }
File.open('classifier.dat', 'wb') { |f| f.write(Marshal.dump(classifier)) }
```

At prediction time you reload both objects, transform the test samples with the same transformer, and score the classifier. The README shows the expected output of this pair of scripts as an accuracy of 98.5 percent on the pendigits test set. Treat that as the README's reported figure on that dataset, not as a general performance claim.

```ruby
require 'rumale'

samples, labels = Rumale::Dataset.load_libsvm_file('pendigits.t')

transformer = Marshal.load(File.binread('transformer.dat'))
classifier = Marshal.load(File.binread('classifier.dat'))

transformed = transformer.transform(samples)
puts("Accuracy: %.1f%%" % (100.0 * classifier.score(transformed, labels)))
```

Two optional gems change how fast this runs. Installing numo-linalg-alt lets matrix and vector products go through OpenBLAS, and the README says algorithms that compute those products frequently can be expected to speed up. Load it before Rumale:

```bash
gem install numo-linalg-alt
```

```ruby
require 'numo/linalg'
require 'rumale'
```

For parallel execution, install the parallel gem and load it the same way. Estimators that support parallelism take an n_jobs parameter, and the README states that a value of -1 uses all processors:

```ruby
estimator = Rumale::Ensemble::RandomForestClassifier.new(n_jobs: -1, random_seed: 1)
```

## Cross-validation and model selection in practice

The second README example shows the intended evaluation workflow. You construct an estimator, an evaluation measure, a splitting strategy and a CrossValidation object, then call perform on the samples and labels. The result is a report hash whose test_score key holds the per-fold scores; the README averages them by dividing by the number of splits.

The README's cross-validation example uses LogisticRegression with StratifiedKFold, five splits, shuffle enabled and random_seed set to 1. It reports a 5-CV mean accuracy of 95.5 percent on pendigits. The presence of random_seed on both the splitter and the estimators is worth noting: reproducibility is opt-in and threaded through each object rather than set globally.

```ruby
lr = Rumale::LinearModel::LogisticRegression.new
ev = Rumale::EvaluationMeasure::Accuracy.new
kf = Rumale::ModelSelection::StratifiedKFold.new(n_splits: 5, shuffle: true, random_seed: 1)
cv = Rumale::ModelSelection::CrossValidation.new(estimator: lr, splitter: kf, evaluator: ev)
report = cv.perform(samples, labels)
```

This is a plain, readable evaluation loop rather than a pipeline abstraction. If you are used to scikit-learn's Pipeline object or to a search API over parameter grids, the README does not describe either, so check the API documentation before assuming they exist.

## Where Rumale is the wrong tool

The serialisation approach is the first thing to weigh. Both examples persist models with Ruby's Marshal, which is a Ruby specific format tied to the class definitions that produced it. A model file written by one version of an estimator is not a portable artefact you can hand to a service written in another language, and the README does not document a versioning or migration story for those files. If your deployment target is a non-Ruby runtime, this is a real obstacle.

Second, the acceleration story depends on optional gems. Without numo-linalg-alt the README only says that speed can be expected to improve with it, which implies the default path is slower for matrix heavy algorithms. There is no benchmark table in the README to quantify the difference, so you would have to measure it on your own data.

Third, parallelism is described as partial. The README says several estimators support parallel processing through the parallel gem and that those estimators have an n_jobs parameter. It does not list which ones, so you cannot assume a given estimator scales across cores until you check its documentation.

Finally, the README documents no GPU support, no distributed training, no automatic hyperparameter search and no pretrained model distribution. Teams whose work depends on those capabilities should stay with a Python stack rather than port the effort into Ruby.

## Rumale compared with scikit-learn

The obvious alternative is scikit-learn in Python, and the comparison is more interesting than a simple feature count. Rumale deliberately copies the interface: constructors take hyperparameters, fit takes samples and labels, transform and predict produce arrays, and evaluation measures are separate objects. Anyone who knows scikit-learn can read the README examples without learning a new idiom.

The difference is ecosystem depth. Scikit-learn sits inside a Python environment with a much larger set of surrounding tools, and Rumale's answer to that is composition rather than imitation. The README points to two related projects: Rumale::SVM, which provides LIBSVM and LIBLINEAR support vector machine algorithms behind the Rumale interface, and Rumale::Torch, which provides learning and inference for neural networks defined in torch.rb behind the same interface. So the design intent is that Rumale supplies the common estimators and evaluation machinery, and specialised work is delegated to sibling gems that share the API.

That is a coherent strategy, but it also means the boundary of what Rumale does is defined by which sibling gems you are willing to add. If you need a support vector machine implementation that matches LIBSVM exactly, the README's own pointer is to install Rumale::SVM rather than to expect it from the core library.

## Maintenance, releases and licence

The repository is not archived, and the last push was on 2026-09-05. The release history shows v2.2.0 on 2026-07-05, v2.1.0 on 2026-02-08 and v2.0.2 on 2025-11-20, so releases arrive at a steady but unhurried pace rather than continuously. The CHANGELOG.md file at the repository root is where release notes live, and the README does not document a deprecation or rollback policy for the Marshal model files.

The v2.0.0 change of array dependency from Numo::NArray to Numo::NArray Alternative is the upgrade cost to plan for. Any code that constructs or inspects arrays directly may need to be revisited, and the README flags this change explicitly, which suggests it was a breaking one. The repository also carries a Gemfile.rubocop, a .rubocop.yml, an rspec.sh and a rumale-test.sh, so linting and tests are part of the normal working setup.

On licensing, the gem is released under the BSD-3-Clause License, and the LICENSE.txt file sits at the repository root. BSD-3-Clause is permissive and generally allows use in closed source products provided the copyright notice and licence text are retained, but the exact obligations depend on how you distribute the software. Read the licence text and, if your organisation has one, involve your legal team rather than relying on a summary.

## Conclusion

Adopt Rumale if your training and inference code already lives in Ruby and you want to avoid a Python service boundary. Do not adopt it if your team depends on the Python ecosystem for experiment tracking, GPU training or a wide catalogue of pretrained models, because the README lists none of that. Before committing, check that your Ruby version can build the numo-narray-alt dependency on your target platform, and confirm which estimators you need actually expose the n_jobs parameter, since the README says only that several estimators support it.

## FAQ

### How do I install Rumale?

Add gem 'rumale' to your Gemfile and run bundle, or install it directly with gem install rumale. The README also notes that since v2.0.0 the library depends on Numo::NArray Alternative rather than Numo::NArray.

### What algorithms does Rumale support?

The README lists Support Vector Machine, Logistic Regression, Ridge, Lasso, Multi-layer Perceptron, Naive Bayes, Decision Tree, Gradient Tree Boosting, Random Forest, K-Means, Gaussian Mixture Model, DBSCAN, Spectral Clustering, Multidimensional Scaling, t-SNE, Fisher Discriminant Analysis, Neighbourhood Component Analysis, Principal Component Analysis, Non-negative Matrix Factorization and others.

### How can I make Rumale run faster?

Install the numo-linalg-alt gem and require numo/linalg before rumale, which lets matrix and vector products use OpenBLAS. The README says algorithms that compute those products frequently can be expected to speed up.

### Does Rumale support parallel processing?

Several estimators do, implemented through the parallel gem, and those estimators expose an n_jobs parameter where -1 uses all processors. The README does not list which estimators support it.

## Sources

- [License: BSD-3-Clause](https://github.com/yoshoku/rumale/blob/main/LICENSE)
- [Project website](https://rubygems.org/gems/rumale)
- [README](https://github.com/yoshoku/rumale/blob/main/README.md)
- [Releases](https://github.com/yoshoku/rumale/releases)
- [yoshoku/rumale on GitHub](https://github.com/yoshoku/rumale)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/yoshoku-rumale
