# mleap: running a trained pipeline anywhere on the JVM without Spark, and the machinery that makes that safe

> This is a Scala project that takes a pipeline trained in Spark or scikit-learn, exports it to a portable bundle, and executes it on a lightweight runtime with no dependency on Spark, numpy or pandas. What makes it worth reading is not the export, it is the table of twenty tested dependency combinations and the cross-build test command that keeps that table honest.

**combust/mleap** — MLeap: Deploy ML Pipelines to Production

- Repository: https://github.com/combust/mleap
- Website: https://combust.github.io/mleap-docs/
- Stars: 1,546 · Forks: 313
- Language: Scala
- License: Apache-2.0
- Published: 2026-09-30 · Updated: 2026-09-30 · Language: en
- Canonical page: https://hysenlabs.com/projects/combust-mleap

## Twenty tested combinations, kept true by a cross-build command

The compatibility matrix in this readme is twenty rows long, and that is the most impressive thing in the document.

Each row pairs a release of this project with the versions of seven other things: the training framework, the language version, the Java version, the Python version, and the versions of two gradient boosting libraries and a deep learning framework. The oldest row is from the era when the training framework was at 2.4 and Java 8 was current. The newest is at 4.1 and Java 17. In between, the language version moves twice and the Python floor climbs from 3.6 to 3.13.

Most projects publish one tested combination, or a small number. Publishing twenty, spanning four major versions of the training framework and three Java versions, means somebody is checking all of them. And the build file shows how.

Every test target in the makefile invokes the Scala build tool with a cross-build flag before the task name. That flag does not mean cross-compile in the usual sense. It means run this task once for every version of the project built for every cross-target, and fail if any of them fails. So the executor tests, the benchmark tests, the boosting library tests, the root project's tests and the Python tests are each run across the whole matrix on every build.

That single character is the entire answer to how a table like that stays accurate for years. It is not maintained by hand and it is not maintained by hope. It is maintained by a test suite that is structurally incapable of going green while a combination is broken, and which nobody can merge without.

The test entry that proves it is one line in the makefile:

```
test_executor:
	$(SBT) "+ mleap-executor-tests/test"
```

The table also carries an honest caveat. Other combinations may work, and more recent Java versions in particular, but these are the ones that were tested. That phrasing matters: it tells you the matrix is a floor on your confidence rather than a wall, and it tells you which row to look at. If you are on a version in the table, you are on tested ground. If you are one minor above it, the project is telling you it probably works and has not promised.

Two of the three most recent releases are on the same afternoon, an hour apart, which suggests the patch releases are responsive to compatibility reports rather than scheduled. For a project whose value is compatibility, that is the right priority order.

## Parity is listed as a feature, which is the correct place for it

The feature list has six entries, and the one that matters most is the fifth.

Full parity tests between the training pipelines and the runtime pipelines. That is not a documentation claim, it is a test suite, and it exists because the entire premise of the project is falsifiable.

The premise is that you can train a pipeline somewhere heavy, export it, and get the same answers somewhere light. Everything else follows from that. If the answers match, you have decoupled serving from training, which is a real and valuable thing and the reason the project exists. If they do not match, the project is a slower, more awkward way of running the training framework with extra steps.

So the risk is concentrated in a single place: a transformer that behaves slightly differently in the two implementations. Not a crash, which is obvious and cheap, but a numerical difference. A different default for a missing value. A different tie-break in a sort. A different rounding at the boundary of a bucket. A different behaviour on an empty input. Each of those is a line of code, and none of them announces itself.

This is the classic failure mode of reimplementing something for portability, and it is why a reimplementation that ships a parity suite as a headline feature is a different proposition from one that does not. The parity suite is what makes the claim checkable, and it is a module in the repository rather than a script somebody runs.

The module list is telling about where the effort goes. There is a separate module for the executor's own tests and another for the training-framework side's tests, and a test kit for the training side. That is a test harness with two halves, which is what a cross-implementation parity suite requires: you need to be able to run a pipeline in both engines and compare.

It also explains something about the release cadence. When a training framework release lands, parity is the thing at risk, so the response is a release within hours, which is what the release history shows. A project whose only claim was speed would wait for a benchmark. A project whose claim is correctness patches on the same day.

What the feature list does not say, and what you should go and check, is how much of the surface is covered. Parity between two implementations of a large library is a spectrum, not a switch, and the number of transformers in the parity suite is the number you actually care about.

## A bundle is a zip file, and it will not run without a schema

The export step in the readme's example is short, and two details in it are the whole tutorial.

First, the output is a zip. The example writes to a path whose scheme is a file inside an archive, and the variable holding it is a bundle file type. So the portable format is an ordinary zip archive containing a manifest, the model parameters and whatever else the pipeline needs. That is a good choice and it is worth being explicit about why: a zip is something you can attach to a ticket, put in object storage, diff, checksum, and open with the tools everybody already has. A proprietary binary blob would be none of those things, and a directory you had to keep together would be worse.

Second, the example does not export the pipeline on its own. It constructs a bundle context and passes a data frame to it, and the data frame it passes is the training data already transformed by the pipeline. That looks redundant and is not. A model without its input schema is not runnable, because the runtime has to know what columns to feed it, what types they are, and what the model's features are called. The schema is part of what you are exporting, and the mechanism for capturing it is to run a row through the pipeline at export time.

That is a genuine subtlety, and it is the kind of thing that gets learned by getting it wrong. A bundle exported without a context either fails at load with an unhelpful message about a missing schema, or worse, loads and then mis-binds features by position. The readme teaches it by example rather than by warning, which is a legitimate documentation style and slightly unkind to the reader who does not yet know why the line is there.

The export call itself is worth noting for another reason. It is a method on the pipeline object itself, with the context supplied as an argument, rather than a separate export function. That means the training framework knows how to serialise itself, which is the right place for it: a transformer that is not aware it is being exported is a transformer whose state is easy to forget.

For a scikit-learn pipeline the same idea applies, and the readme says the runtime supports pipelines trained that way too, which means there are two export paths producing the same bundle format. That is the harder half of the project and the reason a parity suite matters even more there.

## Two serializations, two repositories, and a URI instead of a path

Three things in the module list together describe how a bundle moves through a system, and each is a small design decision with a consequence.

There are two serialisation formats, JSON and Protobuf, and they are separate modules rather than one module with a switch. The likely reason is that they are for different readers. A JSON bundle is a file a person can open and read when a model is producing the wrong answer, which is the situation where you most need to see what was saved. A Protobuf bundle is smaller and faster to load, which is what you want in a service that loads a model per request or per worker start. Shipping only one would mean either shipping something you cannot debug or shipping something wasteful in production, and most projects pick one and apologise.

Then there is the repository layer. Two modules provide backends for storing bundles, one for a distributed filesystem and one for object storage. So a bundle is addressed by a location rather than embedded in the application, and the location is a URI. That is a deployment decision disguised as a storage detail, and it is the decision that makes this project useful: the same bundle file can be promoted from a development bucket to a production bucket by changing a configuration value, with no rebuild, no redeploy, and no possibility of the two environments disagreeing about which model is live.

The alternative, a bundle compiled into the application, makes every model change a release. For a model retrained weekly that is untenable, and it is the reason people end up with model versions that are only known by asking whoever last deployed.

The object storage backend being a separate module rather than built in also means the runtime is not coupled to a particular cloud. For an on-premises deployment with a distributed filesystem and no object store, the two are genuinely different products and lumping them together would have forced one of them to be a dependency of the other.

The readme's own example writes to a local file, which is the simplest case of the general mechanism. A production deployment is the same code with a different URI, and that is the design goal stated in the first paragraph: deploying a pipeline should not be a time-consuming or difficult task.

## Three modules for one platform, one of them a fat jar

Scanning the module list, most of the entries follow a pattern: a core, a runtime, a training-side integration, a test kit, and tests. There are about four such groups, for the executor, the training framework, the boosting library runtime, and the boosting library's training side.

One group breaks the pattern badly, and it is the most interesting thing in the list.

There are three modules for a single hosted platform. One is the runtime for that platform. One is a test kit. And one is a fat variant of the runtime. Three modules, and the third exists only because a fat jar is a different artefact.

The reason is a hard constraint of the platform rather than a preference. A managed notebook service that runs your cluster ships its own libraries and controls the classpath, so you cannot hand it a thin jar and let it resolve a tree of dependencies at runtime. What you need is a single self-contained jar with everything inside it. That is a completely different build, with a different dependency strategy and a different failure mode, and it is why the project carries three modules where it could have carried one.

It is worth noticing what this tells you about the project's priorities. Serving a pipeline on a managed platform is not an afterthought here; it has its own runtime, its own build, and its own test kit. That is the shape of a project whose users actually deploy, rather than evaluate.

The same pattern appears once more, less visibly. There is a module for a popular Java framework integration and a separate module for a remote procedure call server. Both exist because a model runtime is not a library you call, it is a service you host, and hosting a service inside an application framework is a different packaging problem from hosting it standalone. The test kit next to the platform runtime is the giveaway again: a test kit exists to let other people's tests run against your implementation, which is what you build when you expect to be extended.

One thing this list does not contain is a module for a compiled language runtime other than the one the boosting library needs. The claim in the readme is portability across the JVM, not beyond it, and the module list confirms the boundary is respected.

## The Python side is a bridge, and the build script shows the seam

The feature list claims Spark, a Python variant of Spark, and scikit-learn support, and the last two deserve to be read carefully because they mean different things.

Spark support means a module that runs inside the training framework's own runtime and can see its types directly. The scikit-learn support means a training path for a library that has nothing to do with the JVM. The Python variant support is the third thing, and it is the one implemented as a bridge.

The makefile shows the seam. The Python test target does not run a Python test suite. It sources a shell script whose name says what it does, derives a Scala classpath, and then delegates to a make target inside the Python directory. In other words: the Python tests are real tests, and they work by making the compiled Scala classes visible to a Python process, then exercising the bindings over them.

That is the honest architecture. A Python package on a package index that exposes a JVM runtime is a set of bindings, not a native implementation, and every consequence follows. The JVM must be present and its version must match the one the classes were compiled for, which is why the compatibility matrix has both a Java and a Python column. The classpath has to be constructed, which is what the sourced script does, and which is a step a user will hit on their own. And the performance characteristics of a call from Python into the JVM are not the same as either side's native cost, which matters if you are scoring a large batch.

None of this is a criticism, because the alternative is two implementations of the same pipeline format, which is a much worse thing to maintain. A bridge is the right call. It is only worth knowing about if you were expecting the Python path to be independent.

The benchmark module and its test target deserve a mention for the same reason. In a project whose central claim is that the lightweight runtime performs comparably to the training framework, benchmarks are not decoration. They are the evidence for the claim that matters most after correctness, and having a benchmark module with its own test target means the numbers are regenerated rather than quoted from a blog post two years ago.

One small thing at the root: there are configuration files for two continuous integration systems, one of them a well-known older one. As with several projects of this vintage, the second was added and the first was never removed, and the badge in the readme points at the newer one.

## Conclusion

MLeap is the right tool if you train on a cluster and serve somewhere that should not need one, because the exported bundle is a versioned file you can promote by changing a path, and the parity test suite is the reason you can trust that it produces the same answers. It is a poor fit if you are on Python natively, because the Python package is a bridge into the JVM rather than an independent implementation, and a poor fit if you need a transformer the project has not implemented, since the runtime is a reimplementation rather than a wrapper and anything missing has to be ported. Before you commit, check your exact dependency combination against the published matrix rather than assuming a nearby version works, and read the parity test module to see which transformers are actually covered.

## FAQ

### What is MLeap and what problem does it solve?

It provides a portable execution engine and serialisation format for machine learning pipelines. You train a pipeline in Spark or scikit-learn, export it once, and execute it on a lightweight runtime that needs no dependency on Spark, scikit-learn, numpy or pandas, anywhere on the JVM.

### How does MLeap keep its dependency compatibility matrix accurate?

Every test target in the makefile invokes the Scala build tool with a cross-build flag, which runs the task once for every cross-built version of the project and fails if any fails. That is how twenty tested combinations spanning four major versions of the training framework and three Java versions are kept true without manual maintenance.

### What does MLeap's parity testing actually test?

That pipelines executed by the lightweight runtime give the same results as pipelines executed by the training framework. It is listed as a key feature rather than a claim, and it is implemented as a pair of test modules with a test kit, because the risk in a portable reimplementation is a small numerical difference in one transformer rather than a crash.

### What format is an MLeap bundle?

An ordinary zip archive written through a bundle file type, containing a manifest and the model. The export also requires a data frame transformed by the pipeline, because the input schema has to be captured and the runtime needs it in order to bind features correctly.

### Does MLeap's Python support mean a native Python implementation?

No. The Python tests run by sourcing a script that derives the Scala classpath and then exercising the bindings over the compiled classes, so the Python path is a bridge into the JVM. The practical consequences are that a matching JVM must be present, the classpath has to be constructed, and cross-boundary calls do not have native performance.

## Sources

- [combust/mleap on GitHub](https://github.com/combust/mleap)
- [License: Apache-2.0](https://github.com/combust/mleap/blob/master/LICENSE)
- [Project website](https://combust.github.io/mleap-docs/)
- [README](https://github.com/combust/mleap/blob/master/README.md)
- [Releases](https://github.com/combust/mleap/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/combust-mleap
