Self-hosted service
Waikato/moa avatar
Waikato/moa

MOA (Massive Online Analysis): a Java framework for data stream mining

MOA is an open source framework for Big Data stream mining. It includes a collection of machine learning algorithms (classification, regression, clustering, outlier detection, concept drift detection and recommender systems) and tools for evaluation.

665 stars369 forksJavaGPL-3.0

At a glance

What is it?
MOA is a GPL-3.0 Java library and GUI for training classifiers, regressors, clusterers and drift detectors on data streams that never fit in memory. It suits researchers and engineers who need a benchmark suite, not a production streaming platform.
Who is it for?
Adopt MOA if you are reproducing stream mining research, teaching online learning, or need a reference implementation of algorithms such as Hoeffding Trees and ADWIN to compare against. Do not adopt it as the serving layer of a production pipeline: the README describes a benchmark suite for the stream mining community, and the last push was on 2026-08-30, so the project moves at an academic pace.
Can I use it commercially?
Yes, with conditions. GPL-3.0 is a copyleft licence: if you distribute software that includes it, you must release that software's source code under the same licence. Running it internally without distributing it does not trigger that obligation.
Is it still maintained?
Yes. The repository last received commits 32 days ago.
What is it written in?
Mainly Java, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What MOA solves, and who actually needs it

Batch machine learning assumes you can hold the training set, shuffle it, and pass over it many times. Stream mining assumes the opposite: examples arrive one at a time, the distribution can shift underneath you, and you cannot store the history. MOA is built for the second case. The README describes it as a collection of machine learning algorithms covering classification, regression, clustering, outlier detection, concept drift detection and recommender systems, plus tools for evaluation.

The audience is narrow and identifiable. It is researchers who need a shared benchmark so that a new drift detector or a new tree variant can be compared against published baselines. It is lecturers who want a GUI in which students can watch a classifier update as a stream generator produces data. It is engineers who have a specific streaming problem, such as scoring events on a feed where the label arrives late, and who want a reference implementation before writing their own. The README makes the intent explicit: the goal is to provide a benchmark suite for the stream mining community. That sentence should govern how you read everything else on this page.

How MOA is structured: tasks, streams and evaluators

The repository layout tells you most of the architecture. The moa/ directory holds the core Java sources; moa-kafka/ is a separate module for reading streams from Kafka; weka-package/ packages MOA for the WEKA ecosystem; docker/ contains container build material; and pom.xml at the top level makes the whole thing a Maven multi-module build. The README states the project is related to WEKA and also written in Java, while scaling to more demanding problems. That relationship is not cosmetic: WEKA's instance and attribute model is the substrate MOA's data streams are expressed in.

The mental model has three parts. A stream produces instances, either from a generator or from a file or a live source. A learner consumes those instances one at a time through a train-and-test loop rather than a fit-then-predict split. An evaluator measures the result, usually with a windowed or fading statistic, because a single accuracy number over an infinite stream is meaningless. Concept drift detection sits alongside the learner, watching the error signal rather than the data. The README states that MOA can be extended with new mining algorithms, new stream generators and new evaluation measures, which is the extension surface you would use if the built-in set does not cover your case.

Building MOA from source with Maven

The README points to a tutorial page titled Building MOA from the source, and the repository ships a top-level pom.xml, so Maven is the documented path. The README itself gives no build command, so the exact invocation has to come from that tutorial page. What the repository does tell you is that the build is a multi-module Maven project rooted at pom.xml, with the modules listed at the top level being moa, moa-kafka and weka-package.

The README also carries a Maven Central badge pointing at the nz.ac.waikato.cms.moa group, which means you can depend on a released artifact instead of building from source. The badge does not spell out the artifact coordinates in the README text, so check the linked Maven Central page for the exact artifactId that matches the module you want before writing it into a pom.xml. Until you have that, treat any coordinate you see quoted elsewhere as unverified.

Running MOA from Docker and a first task

The README links a DockerHub badge for the waikato/moa image, and the repository has a docker/ directory, so a container is the shortest route to a working environment without a JDK on your machine. The README does not reproduce the docker pull or docker run invocation, so take the exact image tag and entrypoint from the docker/ directory or the linked DockerHub page rather than guessing. Expect to inspect the container rather than assume a shell or a GUI appears.

Once you have a working MOA installation, the first real task is to evaluate a classifier on a synthetic stream. MOA ships a command line interface and a GUI, and the README's Getting Started page is the reference for both. The GUI's task panel is where you choose a stream, choose a learner, and choose an evaluator; the CLI exposes the same three choices as flags. Because the README does not reproduce the flag names, take them from the Getting Started page rather than from this article, and treat any command you find elsewhere as unverified until you have run it.

Where MOA is the wrong tool

MOA is a research framework, and the README's own framing says so. If you need a low-latency scoring service with an operational story around deployment, model versioning, rollback and monitoring, nothing in the README addresses those concerns. There is no mention of a model registry, a serving API, or a persistence format for trained models. The moa-kafka module shows that streaming ingestion from Kafka is contemplated, but ingestion is not serving.

The second limitation is the release cadence. The most recent release listed is 2024.07.0 from 2024-07-18, preceded by 2023.04.0 and 2021.07.0. That is roughly annual, and the gap between 2021.07.0 and 2023.04.0 was longer than a year and a half. The repository itself is not archived and the last push was on 2026-08-30, so development has not stopped, but a project whose tagged releases arrive yearly is not one to depend on for a fix you need this week. The third limitation is the licence, covered below. The fourth is scope: MOA is Java-first, so a Python or Go shop pays an integration cost that a native library would not impose.

MOA compared with River, a Python stream learning library

The closest thing to a like-for-like alternative is River, a Python library for online machine learning. The difference is not just language. River's API is built around Python objects that you compose and call incrementally, which fits naturally into a Python data stack and into notebooks. MOA's API is built around its own task, stream and evaluator abstractions, with a GUI and a CLI layered on top, and it interoperates with WEKA's instance model. If your team writes Python and wants incremental learning inside an existing pipeline, River removes the JVM from the picture entirely.

What you give up by leaving MOA is the benchmark suite. MOA's value is partly that its algorithms are the reference implementations that published stream mining papers compare against, and the README frames the project as serving exactly that community. Reproducing a MOA result in another library means reimplementing the generator, the evaluator and the drift detector, and any of those can differ in ways that make the numbers incomparable. Choose MOA when comparability with the literature matters. Choose the other when the deliverable is a working system rather than a comparable number.

Licence, maintenance and what an upgrade costs

MOA is licensed under GPL-3.0, as stated in the README badge and the LICENSE file at the top level. The practical consequence is that distributing a modified MOA, or a product that links against it, brings copyleft obligations that permissive licences do not. This is not legal advice; if you plan to ship MOA inside a closed product, get the question answered by someone qualified before you build on it. The GPL is also the reason some commercial streaming vendors reimplement algorithms rather than depend on this library.

Maintenance cost is shaped by the release cadence. With releases at 2024.07.0, 2023.04.0 and 2021.07.0, an upgrade is an infrequent, deliberate event rather than a routine one. The multi-module Maven layout means an upgrade touches the core module plus moa-kafka and weka-package if you use them, and the WEKA relationship means a WEKA version bump can ripple through. Budget for reading the release notes rather than assuming a drop-in swap, and pin the version you depend on so that an unreviewed build does not change your results between experiments.

Editorial conclusion

Adopt MOA if you are reproducing stream mining research, teaching online learning, or need a reference implementation of algorithms such as Hoeffding Trees and ADWIN to compare against. Do not adopt it as the serving layer of a production pipeline: the README describes a benchmark suite for the stream mining community, and the last push was on 2026-08-30, so the project moves at an academic pace. Before committing, verify that the Maven coordinates resolve for the modules you need, check the docker/ directory for the image you intend to run, and confirm the GPL-3.0 terms against how you plan to distribute your own code.

Frequently asked questions

What is MOA and what kinds of algorithms does it include?

MOA (Massive Online Analysis) is an open source Java framework for data stream mining. The README lists classification, regression, clustering, outlier detection, concept drift detection and recommender systems, plus tools for evaluation.

How do I install MOA or build it from source?

The README links a tutorial titled Building MOA from the source, and the repository has a top-level pom.xml, so Maven is the documented build path. A DockerHub badge for waikato/moa and a docker/ directory in the repository provide a container route as an alternative.

Is MOA related to WEKA?

Yes. The README states that MOA is related to the WEKA project, is also written in Java, and scales to more demanding problems. The repository also contains a weka-package/ directory.

What licence does MOA use?

MOA is licensed under GPL-3.0, per the licence badge in the README and the LICENSE file in the repository root. That is a copyleft licence, so distribution of modified or linked code carries obligations.

How often is MOA released?

The listed releases are 2024.07.0, 2023.04.0 and 2021.07.0, so tagged releases have arrived roughly annually. The repository is not archived and the last push was on 2026-08-30, so work continues between releases.

Official sources

  1. License: GPL-3.0
  2. Project website
  3. README
  4. Releases
  5. Waikato/moa on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/waikato-moa.svg)](https://hysenlabs.com/projects/waikato-moa)