# ankane/eps: machine learning for Ruby without leaving Ruby

> Eps trains regression and classification models from arrays of hashes and stores them as PMML, so a Rails app can serve models built in Ruby, Python or R. The API is small; the constraints are real.

**ankane/eps** — Machine learning for Ruby

- Repository: https://github.com/ankane/eps
- Stars: 693 · Forks: 15
- Language: Ruby
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/ankane-eps

## The gap Eps fills between Ruby apps and trained models

A Rails application that wants a prediction usually ends up with two runtimes: the Ruby app on one side, a Python training job and a scoring service on the other, plus a network hop and a second deployment to keep alive. Eps removes the second runtime. It trains models in Ruby from plain arrays of hashes and stores them as PMML, so the artifact is a file the Ruby process can read back. The README states the two goals plainly: build predictive models quickly, and serve models built in Ruby, Python, R, and more. That second goal matters more than the first. The gem is aimed at developers who already have data in ActiveRecord and would rather not add a service boundary just to get a regression or a classifier.

## How Eps trains: hashes in, PMML out

The input is an array of hashes, each hash one row, one key designated as the target. Eps infers the task from the target's type. A numeric target means regression; a categorical target means classification. Feature types are inferred the same way: numbers stay numeric, strings and booleans become categorical, and strings containing multiple words become bag-of-words text features. That inference is convenient and also the main source of surprises, because a numeric id passed as an integer will be treated as a continuous feature rather than a category. The README addresses this directly with the instruction to convert ids to strings so they are treated as categorical. Dates get the same treatment: you are expected to derive features such as day of week and month yourself rather than passing a timestamp.

Once the data is prepared, Eps splits it. If there are 30 or more data points, it holds out a validation set, trains on the rest, and reports performance in the summary: validation RMSE for regression, validation accuracy for classification. The split is random unless you name a time column, and the README recommends that whenever the data has a time associated with it. That recommendation is the difference between a model that looks good in validation and one that actually generalizes, since a random split on time-ordered rows lets the model see the future.

The trained model serializes to PMML, an XML standard for model storage. That choice is what makes cross-language serving possible, and it is also a ceiling: the format covers the model families Eps can express, not arbitrary ones.

## Installing Eps and training your first model

Installation is a Gemfile line. On macOS the README notes an extra step, because the underlying numerics need OpenMP, which is not part of the base system.

```ruby
gem "eps"
```

```sh
brew install libomp
```

After bundling, the smallest useful program is four rows of data and a target. The README's example uses house prices, which is a regression because the target is numeric.

```ruby
model = Eps::Model.new(data, target: :price)
puts model.summary
```

With fewer than 30 rows, as in that example, Eps does not hold out a validation set, so the summary is not a generalization estimate. Treat the printed summary as a sanity check until your dataset is larger than the threshold.

Predicting takes a hash with the same feature keys, minus the target. Passing an array of hashes returns multiple predictions in one call.

```ruby
model.predict(bedrooms: 2, bathrooms: 1)
```

To persist the model, write the PMML string to a file and load it back later. The README suggests checking the file into source control or storing it with a tool such as Trove.

```ruby
File.write("model.pmml", model.to_pmml)
model = Eps::Model.load_pmml(File.read("model.pmml"))
```

For a Rails app, the README recommends keeping model code in one file, for example under app/ml_models, subclassing Eps::Base and defining build and predict. It also notes that Spring must be restarted after creating that directory so the files are autoloaded.

```sh
bin/spring stop
```

## Feature engineering is where the accuracy actually comes from

The README says the best way to improve performance is feature engineering, and the API reflects that belief: there are no hyperparameter tuning loops to run, no grid search helper, no automatic feature generation beyond the type inference. What you get instead is control over how text becomes features. Text features accept a hash of options: min_occurrences, max_features, min_length, case_sensitive, tokenizer and stop_words. The tokenizer defaults to whitespace, so punctuation stays attached to words unless you supply a different pattern. Anyone coming from a Python pipeline should read that default carefully, because a regex-based tokenizer is something you have to bring yourself.

The categorical path has a similar shape. Strings and booleans become categories, which means high-cardinality string columns expand into many features. The README does not document cardinality limits or any encoding scheme beyond that, so a column with thousands of distinct values is a case where you should test the summary output rather than assume it behaves well.

If you need gradient boosting or a neural network, Eps is the wrong tool for training. Its role there is serving, not building.

## Serving models built in Python and R, and the consistency trap

This is the part of Eps that has no obvious Ruby equivalent. You train in Python with sklearn2pmml, or in R with the pmml package, export PMML, and load it in Ruby with Eps::Model.load_pmml. The README lists the supported families for serving: LightGBM, linear regression and naive Bayes. It points to ONNX Runtime and Scoruby for other model types, which is an admission that PMML coverage has edges.

The trap is feature consistency. A model trained in Python encodes the exact feature names and value transformations it saw. If the Ruby side computes month as a three-letter abbreviation while the Python side used a full month name, the model will not error; it will quietly produce wrong numbers. The README calls this out and recommends verifying it programmatically. That advice is the most important operational sentence in the document, and it is also the least specified: no verification helper is documented, so the check is on you.

## Monitoring drift with Eps.metrics

The README's monitoring section is short but concrete. Save predictions to the database alongside the eventual actual values, then compare the two arrays.

```ruby
actual = houses.map(&:price)
predicted = houses.map(&:predicted_price)
Eps.metrics(actual, predicted)
```

The guidance that follows is threshold-based. For RMSE and MAE, alert when they rise above a level you set. For ME, alert when it drifts too far from zero, since mean error near zero indicates the model is not systematically over- or under-predicting. For accuracy, alert when it drops below a threshold. Nothing here schedules the job or stores the metrics; Eps.metrics is a computation, and the surrounding infrastructure is your responsibility. That is a reasonable division for a library, but it means the monitoring story is only as good as the cron job you write around it.

## Licence, releases and what upgrading costs

Eps is MIT licensed, which permits commercial use and modification with the licence text retained; the repository carries LICENSE.txt at the top level. This is not legal advice, and if you redistribute the gem inside a product you should read the licence text itself rather than a summary.

On maintenance: the last push to the repository was on 2026-06-29, and the repository is not archived. The material retrieved lists no recent releases, so there is no changelog entry to reason about here; the repository does contain a CHANGELOG.md and a gemfiles directory, which is where version compatibility information would live. Practically, the upgrade surface is small. The public API in the README is a handful of methods: Eps::Model.new, predict, to_pmml, load_pmml, Eps::Base and Eps.metrics. The riskier upgrade path is not the gem version but the PMML files themselves: a model file written by an older version is an artifact you may want to retrain rather than port, since the README documents no migration path for stored models.

## Conclusion

Adopt Eps when you want a trained model inside an existing Ruby or Rails codebase and can accept PMML as the storage format, which also lets you serve LightGBM, linear regression or naive Bayes models built in Python or R. Do not adopt it if you need deep learning, GPU training, or an in-process Python bridge, because Eps trains and serves in Ruby and nothing in the README suggests otherwise. Before committing, verify three things: that your data reaches the 30-point threshold where Eps starts holding out a validation set, that your strongest features are numeric, categorical or bag-of-words text, and that whoever consumes the model file can read PMML. The last push was on 2026-06-29.

## FAQ

### Does ankane/eps require Python to be installed?

No. Eps trains and serves models in Ruby. Python is only involved if you choose to build a model in Python with sklearn2pmml and then serve the exported PMML file in Ruby.

### How does ankane/eps decide between regression and classification?

It looks at the target you pass to Eps::Model.new. A numeric target is treated as regression and reported with validation RMSE; a categorical target is treated as classification and reported with validation accuracy.

### Why does ankane/eps not report validation performance on my small dataset?

The README states that Eps splits the data into training and validation sets only when you have 30 or more data points. Below that threshold there is no held-out set, so the summary is not a generalization estimate.

## Sources

- [ankane/eps on GitHub](https://github.com/ankane/eps)
- [Issues](https://github.com/ankane/eps/issues)
- [License: MIT](https://github.com/ankane/eps/blob/master/LICENSE)
- [README](https://github.com/ankane/eps/blob/master/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/ankane-eps
