Model or dataset
h2oai/h2o-3 avatar
h2oai/h2o-3

H2O-3: Distributed Machine Learning Platform with Python, R, and AutoML

H2O is an Open Source, Distributed, Fast & Scalable Machine Learning Platform: Deep Learning, Gradient Boosting (GBM) & XGBoost, Random Forest, Generalized Linear Modeling (GLM with Elastic Net), K-Means, PCA, Generalized Additive Models (GAM), RuleFit, Support Vector Machine (SVM), Stacked Ensembles, Automatic Machine Learning (AutoML), etc.

7,510 stars2,020 forksJupyter NotebookApache-2.0

At a glance

What is it?
H2O-3 is an in-memory distributed machine learning platform that runs on a single machine or across a Hadoop or Spark cluster. It provides Python, R, Scala, Java, and a web-based Flow UI for training GLM, GBM, XGBoost, Random Forest, Deep Neural Networks, and more, with a fully automated machine learning capability called H2O AutoML.
Who is it for?
Data scientists who need to train multiple machine learning algorithms at scale, with R or Python interfaces and the option to export models for fast production scoring, should evaluate H2O-3. Teams running small datasets on a single core and using only scikit-learn-compatible pipelines will find H2O-3's Java runtime and cluster overhead unnecessary.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 5 days ago.
What is it written in?
Mainly Jupyter Notebook, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What H2O-3 Is and the Problems It Addresses

H2O-3 is the third major version of the H2O platform. The README describes it as 'an in-memory platform for distributed, scalable machine learning.' It targets data scientists and machine learning engineers who need to train models on datasets that are too large for single-machine, single-core training tools, or who need to experiment with multiple algorithms quickly without switching tools.

The platform runs on a Java virtual machine and uses a distributed memory model that spans multiple nodes when deployed on Hadoop or Spark. The README notes that 'H2O uses familiar interfaces like R, Python, Scala, Java, JSON and the Flow notebook/web interface, and works seamlessly with big data technologies like Hadoop and Spark.' This makes it usable from the language a data scientist already works in, while the underlying distribution and parallelism is handled by the platform.

The predecessor H2O-2 is a separate archived repository. H2O-3 is the current and actively developed version. The top-level directory structure reflects the modular architecture: h2o-core/, h2o-algos/, h2o-automl/, h2o-genmodel/, h2o-parsers/, h2o-app/, and many other h2o-* subdirectories, each as a separate module in the Gradle multi-project build.

Algorithms and What H2O-3 Can Train

The README lists the algorithm implementations: Generalized Linear Models (GLM with Elastic Net), Gradient Boosting Machines (GBM), XGBoost, Random Forests, Deep Neural Networks, Stacked Ensembles, Naive Bayes, Generalized Additive Models (GAM), Cox Proportional Hazards, K-Means, PCA, Word2Vec, and H2O AutoML.

H2O AutoML is a fully automated machine learning mode. The README links to its documentation at docs.h2o.ai/h2o/latest-stable/h2o-docs/automl.html. AutoML trains and evaluates multiple algorithms and configurations automatically, returning a leaderboard of models ranked by performance on a held-out validation set.

The algorithm coverage spans tabular supervised learning (GLM, GBM, XGBoost, Random Forest), unsupervised learning (K-Means, PCA), survival analysis (Cox Proportional Hazards), text representations (Word2Vec), and deep learning. This breadth means a data scientist can run experiments across many algorithm families within a single platform without installing separate tools for each.

The stacked ensembles algorithm combines predictions from multiple base learners, which the AutoML feature uses to build final ensemble models. GAM support, listed in the repository description, extends the GLM family with nonlinear terms while retaining interpretability properties.

H2O-3 is extensible. The README states that 'H2O is extensible so that developers can add data transformations and custom algorithms of their choice and access them through all of those clients.' The h2o-extensions/ directory supports this.

Installing H2O-3 and Starting the Cluster

For Python users, install via pip:

bash
pip install h2o

For R users:

r
install.packages("h2o")

The README notes that Python users can also install through Anaconda and R users through CRAN. For the latest stable, nightly, Hadoop, or Spark releases, the README directs users to h2o.ai/download.

For teams that want to use H2O-3 as a Java library in a Gradle project, the README provides an example dependency configuration:

code
def h2oBranch = 'master'
def h2oBuildNumber = 'nnnn'
def h2oProjectVersion = "x.y.z.${h2oBuildNumber}"

The README notes that nightly builds publish to a build-specific S3-backed Maven repository, with the URL pattern https://s3.amazonaws.com/h2o-release/h2o-3/${h2oBranch}/${h2oBuildNumber}/maven/repo/.

Once installed, H2O starts by initializing a cluster in memory. From Python, this is typically done with h2o.init(), which launches a local H2O instance in the current process or connects to an existing one. The Flow web interface becomes available on a local port after initialization, giving access to a notebook-style UI for training and inspecting models without writing code.

Building from source requires Java and Gradle. The Makefile shows the standard build target as ./gradlew --parallel build -x test, skipping the shadow JAR assemblies. The Dockerfile uses Ubuntu 24.04 with OpenJDK 8 and Python 3.11 as the base.

The Flow Web UI and Multiple Language Bindings

H2O Flow is the web-based notebook interface described in the README as one of the ways to interact with H2O-3 alongside R, Python, Scala, Java, and JSON. Flow runs in a browser and provides a visual environment for uploading data, configuring model training, inspecting results, and building simple pipelines through a notebook-like cell interface.

Flow is particularly useful for users who prefer a GUI for exploratory work or for sharing H2O experiments with colleagues who do not write Python or R. It provides point-and-click model training across all supported algorithms, data summary views, and model diagnostics without requiring a notebook server like Jupyter.

The Python and R client libraries translate local function calls into HTTP requests to the H2O cluster's REST API. The JSON interface is the underlying protocol. This architecture means any client that can make HTTP requests can interact with H2O, and the Python and R packages are convenience wrappers around that API.

Scala bindings support Spark integration through the Sparkling Water project, referenced in the README as a separate repository at github.com/h2oai/sparkling-water. The h2o-hadoop-2/ and h2o-hadoop-3/ directories in the repository support deployment on Hadoop clusters.

MOJO and POJO: Exporting Models for Production

H2O-3 models can be exported for production scoring in two formats: POJO (Plain Old Java Object) and MOJO (Model Object, Optimized). The README describes these as allowing models to be 'exported into POJO or MOJO format for extremely fast scoring in production.' The h2o-genmodel/ directory holds the runtime library that enables POJO and MOJO scoring without a running H2O cluster.

POJO exports generate a standalone Java class that encodes the model. This class can be compiled and included in any Java application. MOJO is a newer format, described as optimized, that produces a binary file accompanied by the h2o-genmodel.jar runtime. MOJO scoring is faster than POJO for tree-based models because it avoids Java compilation at deployment time.

The production scoring path is a meaningful design advantage for teams that use H2O for training but deploy to Java-based backend services. The scoring library has no dependency on the H2O cluster, which means production scoring works without maintaining a running H2O instance.

The h2o-genmodel-extensions/ directory suggests extensions to the base scoring library for specific use cases. The README links to the full productionizing documentation at docs.h2o.ai/h2o/latest-stable/h2o-docs/productionizing.html for deployment details beyond what the README covers directly.

Where H2O-3 Is the Wrong Tool

H2O-3 requires Java. The h2o.init() call starts a JVM process. In environments where Java is unavailable or prohibited, such as certain containerized environments with locked-down runtimes or systems with strict binary whitelisting, H2O-3 will not run.

The in-memory constraint is real. H2O-3 requires all the data to fit in the combined RAM of the cluster nodes. For datasets larger than available memory, even distributed across multiple nodes, H2O-3 cannot load them. The README describes it as an 'in-memory platform,' and this is not a configuration limitation but an architectural one.

For small datasets on a single machine, scikit-learn is a more direct choice. scikit-learn has no Java dependency, integrates with the standard Python data science stack (NumPy, pandas, Jupyter), and uses a consistent API across all its algorithms. H2O-3's cluster overhead, the JVM startup time, and the separate Flow UI are not beneficial at small scale.

H2O-3's deep learning implementation is described in general terms in the README. Teams with specialized deep learning requirements, such as GPU-accelerated transformer training, should use PyTorch or TensorFlow directly rather than H2O-3's neural network implementation, which is designed for tabular data and does not provide GPU support in the open-source version based on the repository's documentation.

Maintenance, Build Infrastructure, and Apache-2.0 License

The last push to the repository was on 2026-09-25. The repository has no GitHub releases in the standard sense; H2O publishes versioned builds to its download server at h2o.ai/download. The README references a nightly build page at s3.amazonaws.com/h2o-release/h2o-3/master/latest.html for the most recent nightly artifact.

The build infrastructure is substantial. The repository includes h2o-docs-theme/, h2o-docs/, and the multi-module Gradle build (build.gradle, buildSrc/, gradle.properties). The Makefile shows three build targets: the default build (parallel, skipping some assemblies), a minimal build with PmainAssemblyName=minimal, and clean. The SECURITY.md, CONTRIBUTING.md, and DEVEL.md files document the development and contribution workflow.

Changes prior to version 3.28.0.1 are in the separate Changes-prior-3.28.0.1.md file. The main Changes.md tracks current release history.

The license is Apache-2.0. Commercial use, modification, and redistribution are permitted with attribution and license preservation. The CLAUDE.md at the repository root suggests the project uses AI coding assistance in its own development. The h2o-kaggle/ directory contains Kaggle competition examples.

For teams evaluating the upgrade cost: H2O's Python API is relatively stable between minor versions. Breaking changes tend to appear at major version boundaries. The Python package and R package each have their own versioning tied to the H2O build number.

Editorial conclusion

Data scientists who need to train multiple machine learning algorithms at scale, with R or Python interfaces and the option to export models for fast production scoring, should evaluate H2O-3. Teams running small datasets on a single core and using only scikit-learn-compatible pipelines will find H2O-3's Java runtime and cluster overhead unnecessary. Before adopting, verify that Java is available in the environment and check the h2o.ai/download page for the current stable release appropriate for your Python or R version.

Frequently asked questions

What is H2O-3 in machine learning?

H2O-3 is an open-source, in-memory distributed machine learning platform. It trains algorithms including GLM, GBM, XGBoost, Random Forest, Deep Neural Networks, and AutoML, and it runs on a single machine or on Hadoop and Spark clusters. Python and R are the primary interfaces.

What is the difference between H2O and H2O-3?

H2O-3 is the third major version of the H2O platform, succeeding H2O-2. It is the current, actively maintained version. H2O-2 is a separate archived repository. The README refers to H2O-3 as the current platform and links to the predecessor only for historical context.

Can H2O-3 export trained models for use without running H2O?

Yes. H2O-3 models can be exported as MOJO or POJO format. Both can be scored using the h2o-genmodel.jar runtime library without a running H2O cluster. The README describes this as enabling 'extremely fast scoring in production.'

Official sources

  1. h2oai/h2o-3 on GitHub
  2. Issues
  3. License: Apache-2.0
  4. Project website
  5. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/h2oai-h2o-3.svg)](https://hysenlabs.com/projects/h2oai-h2o-3)