CLI tool
cardmagic/classifier avatar
cardmagic/classifier

cardmagic/classifier: Bayesian and LSI Text Classification in Ruby

A general classifier module to allow Bayesian and LSI classifications.

741 stars125 forksRubyNOASSERTION

At a glance

What is it?
A Ruby gem offering four classifiers plus TF-IDF vectorization, a native LSI extension and two command line tools. The README is detailed on training and scoring, but silent on accuracy numbers, so treat it as a toolkit rather than a turnkey model.
Who is it for?
Adopt cardmagic/classifier when you need text classification inside a Ruby process and want to choose between Bayesian, logistic regression, LSI and k-NN without leaving the language, or when you want a CLI that trains and scores from plain text files. Do not adopt it if your team works in Python, if you need a pre-trained general-purpose model, or if you require documented accuracy figures before committing; the README lists no benchmark results.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository last received commits 11 days ago.
What is it written in?
Mainly Ruby, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What cardmagic/classifier actually solves

The gem targets a specific gap: text classification that runs inside a Ruby process without shelling out to Python or standing up a model server. The README frames it as "Text classification in Ruby. Five algorithms, native performance, streaming support." That combination matters to a Rails application that already holds the text and wants a label back in the same request.

The audience is Ruby developers. A support tool that routes tickets by sentiment, a moderation filter that flags spam, a search feature that tags documents: all of these fit the shape of the library. The README's own examples are spam detection, IMDB sentiment and emotion detection, which is a fair summary of the intended scale. This is not a framework for training large neural models. It is a set of classical classifiers with a consistent Ruby API.

The comparison table in the README positions the gem against "Other Forks" and claims four classifiers plus TF-IDF vectorization against two classifiers, command line executables against none, incremental LSI, a native C extension and pluggable persistence. Those claims describe the repository's own feature surface. They are not accuracy claims, and the README offers no measured results to support a quality comparison.

Four classifiers and a TF-IDF vectorizer under one API

The mechanism differs by algorithm. Bayesian classification trains on labeled strings and returns a capitalized label. Logistic regression has a separate `fit` step that the README marks as "required before the first classify". LSI builds a latent semantic index from documents and classifies by similarity in that reduced space. k-Nearest Neighbors takes a `k` parameter and classifies by the labels of nearby training examples.

The README's quick start shows the shape of each:

ruby
classifier = Classifier::Bayes.new(:spam, :ham)
classifier.train(spam: "Buy viagra cheap pills now")
classifier.train(ham: ["Meeting tomorrow at 3pm", "Quarterly report attached"])
classifier.classify("Cheap pills!")  # => "Spam"
ruby
classifier = Classifier::LogisticRegression.new(:positive, :negative)
classifier.train(positive: "love amazing great wonderful")
classifier.fit                     # required before the first classify
classifier.classify("I love it!")  # => "Positive"
ruby
knn = Classifier::KNN.new(k: 3)
%w[laptop coding software developer programming].each { |w| knn.add(tech: w) }
%w[football basketball soccer goal team].each { |w| knn.add(sports: w) }
knn.classify("programming code")  # => "tech"

The TF-IDF class is separate from classification. It fits on an array of documents and transforms text into a term weight map: `tfidf.transform("Ruby programming")` returns `{rubi: 1.0}` in the README's example. Note the stem `rubi` rather than `ruby`. The README addresses this for the CLI, saying output "maps stems back to whole words, so a model built from `programming` prints `programming`, not `program`." The library-level example does not apply that mapping, which is worth knowing before you compare the two surfaces.

Installing the gem and running a first classification

The README gives two installation routes. For library use, add the gem to your bundle:

ruby
gem 'classifier'

For command line only, the README documents a Homebrew formula:

bash
brew install classifier

Once installed, the CLI can classify against pre-trained models without any Ruby code. The README lists sms-spam-filter, imdb-sentiment and emotion-detection as available models, and `classifier models` lists them.

bash
classifier -r sms-spam-filter "You won a free iPhone"
# => spam

Training your own model from files is the next step. The README's example trains on directories of positive and negative reviews, then classifies a new string:

bash
classifier train positive reviews/good/*.txt
classifier train negative reviews/bad/*.txt
classifier "Great product, highly recommend"
# => positive

After training, the README shows reading a label together with the terms that make the text distinctive, using two separate models in two separate formats. `classifier` takes `-f` and `keywords` takes `-m`, and neither reads the other's file:

bash
classifier -f reviews-model.json -p "Broken on arrival, awful quality"
# => positive:0.12 negative:0.88

The README is explicit that the terms are context, not an explanation of the label: "TF-IDF measures how well a term separates a document from its corpus, not how much it favors a category." That distinction is easy to miss and worth repeating to anyone who plans to show these terms to end users.

Incremental LSI and the persistence question

The feature the README argues hardest for is incremental LSI. Standard LSI rebuilds its decomposition when documents are added. The gem instead implements what the README calls "Brand's algorithm (no rebuild)". The usage pattern requires two constructor flags:

ruby
lsi = Classifier::LSI.new(incremental: true, auto_rebuild: false)

With `auto_rebuild` off, you add the starting corpus and then build once. The README's example adds five technology documents in a single `add` call. The trade-off is that incremental updates approximate the decomposition rather than recomputing it, and the README does not quantify how much accuracy drifts as the corpus grows. If your corpus is small enough to rebuild cheaply, the incremental path buys you nothing and costs you a configuration flag that is easy to set wrong.

Persistence is listed as "Pluggable (file, Redis, S3, SQL, Custom)" in the comparison table, and the repository layout includes a `lib/` directory and a `test/` directory, but the README does not document the interface for writing a custom store. If your deployment needs Redis or S3, plan to read the source or the docs site rather than the README. This is a real gap in the top-level documentation.

Where the README stops short

There are no accuracy numbers anywhere in the README. No precision, recall, F1, or comparison against a baseline. For a library whose entire purpose is producing labels, that is the most significant omission. The comparison table measures features, not results, and a feature count is not evidence that Bayesian classification will work on your corpus.

There is also no documented rollback or model versioning story. The README describes saving models to files and naming custom model files with `-m`, but it does not describe what happens when you retrain and want to revert. If you overwrite `reviews-model.json`, the previous model is gone unless you kept a copy.

The CLI error contract is one of the few operational details the README does specify: "A usage error exits 2 and any other error exits 1, so scripts can tell the two apart." That is a thoughtful touch for pipeline use and more than many CLI tools document.

Finally, the README does not state memory requirements for the streaming claim. It says the gem can "Train on multi-GB datasets" where other forks "Must load all data in memory", but it does not say what the streaming path holds in memory itself. Treat the claim as a design intent until you measure it on your own data.

How it compares to scikit-learn and other Ruby options

The obvious alternative is scikit-learn. It offers the same family of algorithms plus far more, and its documentation includes benchmark datasets and accuracy reporting as a matter of course. The difference in approach is architectural, not just numerical: scikit-learn is a Python library, so using it from a Ruby application means a separate process, an HTTP boundary, or a rewrite. cardmagic/classifier keeps the classifier in the same process as the application, which removes that boundary entirely.

Within Ruby, the README's comparison target is unnamed forks of the same project, which it says offer "2 classifiers only" and no executables, and require either pure Ruby or GSL for LSI. The gem's answer is a native C extension, which the README claims is "5-50x faster" than pure Ruby. That is a wide range and the README does not say what it was measured on, so treat it as an order-of-magnitude indication rather than a specification.

The practical choice: if your application is Ruby and the classification is one step in a larger request, the in-process option wins on integration cost. If classification is the main event and you need to iterate on model quality, scikit-learn's ecosystem and its habit of publishing accuracy make it the better place to do that work.

Maintenance, licence and upgrade cost

The repository is not archived, and the last push was on 2026-09-05. Three releases landed in the three months before that: v2.5.0 on 2026-06-01, v2.6.0 on 2026-06-25 and v2.7.0 on 2026-08-15. That release cadence is the strongest signal in the README and the repository metadata that the project is being worked on, and it is a better indicator than anything about star counts.

The licence is the part to read carefully. The README's badge links to LGPL 2.1, while the repository metadata reports the licence as NOASSERTION, meaning the automated detector could not classify the LICENSE file. Those two signals disagree. LGPL 2.1 is a copyleft licence with a linking exception, which matters if you distribute a modified version of the gem or statically link the native extension. This is not legal advice; if you ship the gem inside a distributed product, have someone read the LICENSE file and the gemspec rather than relying on a badge.

Upgrade cost depends on which surface you use. The Ruby API shown in the README has been stable across the quick-start examples. The CLI is where churn is most likely, since it gained the `keywords` command with its own model format and its own `-m` flag, separate from `classifier`'s `-f`. If you script against the CLI, pin the version and test after each bump. The README documents the exit code contract, which gives your scripts something stable to check.

Editorial conclusion

Adopt cardmagic/classifier when you need text classification inside a Ruby process and want to choose between Bayesian, logistic regression, LSI and k-NN without leaving the language, or when you want a CLI that trains and scores from plain text files. Do not adopt it if your team works in Python, if you need a pre-trained general-purpose model, or if you require documented accuracy figures before committing; the README lists no benchmark results. Verify first that the algorithm you intend to use behaves as expected on your own corpus, check the LGPL 2.1 licence against how you plan to distribute the gem, and confirm which backend you will use for persistence, since the README says file, Redis, S3, SQL and custom stores are all options but does not document each one's setup.

Frequently asked questions

How do I install cardmagic/classifier?

Add gem 'classifier' to your Gemfile for library use, or run brew install classifier if you only need the command line tools. The README documents both routes.

How do I use cardmagic/classifier to classify text?

Either call a classifier class in Ruby, such as Classifier::Bayes.new(:spam, :ham) followed by train and classify, or use the CLI, for example classifier -r sms-spam-filter "You won a free iPhone", which the README shows returning spam.

Which classification algorithms does cardmagic/classifier include?

The README lists Bayesian, logistic regression, LSI and k-Nearest Neighbors, plus a separate TF-IDF vectorizer. The quick start gives a code example for each.

Does cardmagic/classifier work from the command line without writing Ruby?

Yes. The README documents two executables, classifier and keywords. classifier can run pre-trained models such as sms-spam-filter, imdb-sentiment and emotion-detection, or train a new model from files.

What licence does cardmagic/classifier use?

The README's badge points to LGPL 2.1, while the repository metadata reports the licence as NOASSERTION. Check the LICENSE file and gemspec directly before distributing the gem.

Official sources

  1. cardmagic/classifier on GitHub
  2. Issues
  3. README
  4. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/cardmagic-classifier.svg)](https://hysenlabs.com/projects/cardmagic-classifier)