Library / SDK
oracle/tribuo avatar
oracle/tribuo

Tribuo: a Java prediction library with provenance built in

Tribuo - A Java machine learning library

1,417 stars196 forksJavaApache-2.0

At a glance

What is it?
Tribuo wraps Tribuo-native trainers, LibLinear, LibSVM, XGBoost and TensorFlow behind one Java interface, and attaches a serializable provenance record to every model and evaluation. The install path is Maven, and the trade-off is that native-backed interfaces only run where their native libraries exist.
Who is it for?
Adopt Tribuo if your models live inside a JVM service and you want a single Trainer interface across Tribuo-native algorithms, LibLinear, LibSVM, XGBoost and TensorFlow, with provenance attached to every model and evaluation. Do not adopt it if your pipeline is Python-first, if you need clustering algorithms beyond the two the README lists, or if you must run an ONNX Runtime, TensorFlow or XGBoost interface on a platform without that native library.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 145 days ago.
What is it written in?
Mainly Java, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What Tribuo is for, and who ends up using it

Tribuo is a machine learning library written in Java, led by Oracle Labs' Machine Learning Research Group, and it covers multi-class classification, regression, clustering, anomaly detection and multi-label classification. The interesting part is not the algorithm list. It is that Tribuo contains the code to load, featurise and transform data as well as the evaluation classes for every prediction type it supports, so a JVM service does not need a separate feature-engineering stack bolted on.

The audience is therefore narrow and specific. If your application is already a Java service and you want a model trained and evaluated in the same process, Tribuo fits without a Python sidecar. If your team lives in scikit-learn, this is the wrong library, because you would be rewriting existing work in a language your data scientists do not use.

One design decision stands out. Every model and evaluation carries a serializable provenance object recording the creation time, the identity of the data, any transformations applied, and the trainer hyperparameters. For evaluations it also records the model used. That is a deliberate bet that auditability matters as much as accuracy, and it shapes the API.

One Trainer interface over Tribuo, LibLinear, LibSVM, XGBoost and TensorFlow

Tribuo's architecture is a set of implementations plus wrappers. Some algorithms are implemented in the project itself: bagging, random forest, extra trees, K-NN, linear models trained with SGD, factorization machines, CART, SVM-SGD (an implementation of the Pegasos algorithm), Adaboost.SAMME, multinomial Naive Bayes, a linear chain CRF for sequence classification, Lasso via LARS, and Elastic Net via coordinate descent. Others are wrapped: regularised linear models from LibLinear, SVMs from LibSVM or LibLinear, gradient boosted decision trees from XGBoost, and neural networks from TensorFlow.

The unifying abstraction is the Trainer. Ensembles and K-NN are task agnostic and use a combiner to produce output; the combiners are task specific, with voting and averaging variants for multi-class classification, multi-label classification and regression. That separation is why bagging and Adaboost can take any Tribuo trainer as a base learner.

Configuration is handled by OLCUT, Oracle's configuration system. A trainer can be defined in an XML or JSON file and rebuilt repeatably, and example configurations for each supplied Trainer sit in the config folder of each package. Models and datasets serialize with Java serialization. Many models can also be exported to ONNX for deployment elsewhere.

The cost of this breadth is visible in the README itself: interfaces that need native code (ONNX Runtime, TensorFlow, XGBoost) are supported only where that native library exists. Tribuo is pure Java, but those wrappers are not.

Installing Tribuo and training a first model

The README does not print a dependency snippet. It states that tribuo-all is deployed to Maven Central, and that the model card and reproducibility packages are excluded from that deployment because they require Java 17. The README also gives the repository layout, in which each package has a config folder holding example trainer configurations.

bash
git clone https://github.com/oracle/tribuo.git
cd tribuo
mvn install

The clone and build follow the repository's own pom.xml at the top level and the package directories such as Classification/, Regression/, Clustering/, Core/ and Data/. Once the build finishes, the artifacts are in your local Maven repository and can be referenced from another project.

From there, a first model means building a dataset, choosing a trainer and calling train. The README's tutorials cover classification, clustering, regression, anomaly detection, TensorFlow, document classification and columnar data loading; they run on the IJava Jupyter kernel and need Java 10+, except the model card and reproducibility tutorials, which need Java 17. To convert tutorial code back to Java 8, the README says to replace var with the appropriate types.

The tutorials are the practical starting point because the README does not include a standalone quickstart snippet, and the Javadoc for 4.3 is published at tribuo.org. Expect to read the package overview before writing much: the split between Core, Data, Classification, Regression and the rest is what tells you which module a class lives in.

Provenance, redaction and the ONNX export path

Provenance is the feature that distinguishes Tribuo from a plain trainer collection. Each model and evaluation records when it was created, what data it saw, what transformations were applied, and which hyperparameters were used. The provenance object is serializable, and it can be extracted as JSON.

For production, the README describes a second mode: provenance can be redacted and replaced with a hash, so a deployment carries a model identifier that an external tracking system can resolve without embedding the training data description in the artifact. That is a sensible split, and it implies the intended workflow is training somewhere auditable and serving somewhere less permissive.

The export path is ONNX. Many Tribuo models can be exported to ONNX format for deployment in other languages, platforms or cloud services. This is the escape hatch for teams that train in Java but serve outside the JVM. The README does not document which models can and cannot be exported, so that has to be checked per model rather than assumed.

Where Tribuo is the wrong choice

The native-library constraint is the sharpest limitation. ONNX Runtime, TensorFlow and XGBoost interfaces require native code, and Tribuo tests on x86_64 on Windows 10, macOS and Linux (RHEL/OL/CentOS 7+). The README's own guidance for other platforms is to contact the developers of those libraries, which is a polite way of saying Tribuo cannot help you there. If your deployment target is an ARM server or an unusual container base image and you need XGBoost through Tribuo, verify the native library first.

Java version is a second constraint. Tribuo runs on Java 8+ and is tested on LTS versions plus the latest release, but the model card and reproducibility packages require Java 17 and are not in tribuo-all. A team on Java 8 that wants those packages has a problem the dependency declaration alone will not reveal.

Clustering is the thinnest area. The README states that Tribuo includes infrastructure for clustering and supplies two algorithm implementations, and that more are expected over time. If clustering is your primary task, two implementations is a small menu.

Finally, explainability is bounded. The LIME implementation can mix text and tabular data and use any sparse model as an explainer, such as regression trees or lasso, but it does not support images.

Tribuo compared with WEKA and Deeplearning4j

WEKA is the older Java option and the difference is in what each project treats as the centre of gravity. WEKA is an experimentation workbench with a GUI and a broad algorithm collection aimed at interactive use. Tribuo is a library: trainers are configured through OLCUT files and instantiated in code, models serialize with Java serialization, and provenance is attached automatically. If you want to explore a dataset by clicking, WEKA is the better fit. If you want a model inside a service with a record of how it was built, Tribuo's design points that way.

Deeplearning4j is the other frequent comparison, and the split is by model family. Deeplearning4j is built around neural networks and the ND4J tensor stack. Tribuo's neural network support comes through a TensorFlow wrapper, and its own implementations are classical: linear models, CART, random forest, extra trees, K-NN, Naive Bayes, Adaboost, a CRF. The README does not present Tribuo as a deep learning framework, and the algorithm tables bear that out. Choose Tribuo for tabular and sequence problems with a JVM deployment; choose a tensor-first framework when the model itself is the hard part.

Maintenance, licence and what an upgrade costs

The repository is not archived, and the last push was on 2026-05-07. The most recent release is v4.3.2, dated 2025-04-04, following v4.3.1 in 2022-12-23 and v4.2.2 in 2022-10-25. Release cadence has been uneven, so plan upgrades around specific versions rather than assuming a steady stream of patches.

Tribuo is licensed under Apache-2.0. That is a permissive licence, and it is compatible with commercial use, but the repository also ships a THIRD_PARTY_LICENSES.txt file because the wrapped libraries (LibLinear, LibSVM, XGBoost, TensorFlow, ONNX Runtime) carry their own terms. If you ship a model that uses an XGBoost or TensorFlow interface, the obligations attached to those dependencies travel with your build. This is not legal advice; read THIRD_PARTY_LICENSES.txt and the upstream licences before distribution.

The upgrade cost is dominated by the two structural constraints already noted. Native-backed interfaces tie you to platform support, and the Java 17 requirement for the model card and reproducibility packages creates a split between what tribuo-all gives you and what the full project offers. Both are worth confirming against a specific target environment before you build on them.

Editorial conclusion

Adopt Tribuo if your models live inside a JVM service and you want a single Trainer interface across Tribuo-native algorithms, LibLinear, LibSVM, XGBoost and TensorFlow, with provenance attached to every model and evaluation. Do not adopt it if your pipeline is Python-first, if you need clustering algorithms beyond the two the README lists, or if you must run an ONNX Runtime, TensorFlow or XGBoost interface on a platform without that native library. Before committing, verify which modules you actually need, because the model card and reproducibility packages require Java 17 and are not part of the tribuo-all Maven Central deployment, and check the config folder of each package for a Trainer configuration you can reuse instead of writing one from scratch.

Frequently asked questions

What Java version does Tribuo require?

Tribuo runs on Java 8+ and is tested on LTS versions of Java along with the latest release. The model card and reproducibility packages are the exception: they require Java 17 and are not part of the tribuo-all Maven Central deployment.

How do I add Tribuo to a Maven project?

The README states that tribuo-all is deployed to Maven Central, so it is added as a dependency from that repository. The model card and reproducibility packages are excluded from that deployment.

Which algorithms does Tribuo implement itself rather than wrap?

Tribuo implements bagging, random forest, extra trees, K-NN, linear models using SGD, factorization machines, CART, SVM-SGD, Adaboost.SAMME, multinomial Naive Bayes, a linear chain CRF, Lasso and Elastic Net. LibLinear, LibSVM, XGBoost and TensorFlow are wrapped instead.

Can Tribuo models be deployed outside Java?

Many Tribuo models can be exported in ONNX format for deployment in other languages, platforms or cloud services. The README does not list which models support export, so that needs checking per model.

Official sources

  1. License: Apache-2.0
  2. oracle/tribuo on GitHub
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/oracle-tribuo.svg)](https://hysenlabs.com/projects/oracle-tribuo)