Self-hosted service
4paradigm/OpenMLDB avatar
4paradigm/OpenMLDB

OpenMLDB: a SQL feature platform where the training script becomes the serving script

OpenMLDB is an open-source machine learning database that provides a feature platform computing consistent features for training and inference.

1,716 stars332 forksC++Apache-2.0

At a glance

What is it?
OpenMLDB is an open source machine learning database from 4paradigm that runs the same SQL feature definitions offline in Spark and online in a purpose-built engine. It suits teams that already accept SQL for feature engineering, and it is a poor fit if you need a general-purpose OLTP store.
Who is it for?
Adopt OpenMLDB when your feature logic can be expressed in SQL and you are willing to run both a batch engine and a real-time engine to keep training and serving consistent. Do not adopt it as a general OLTP database or if your features live in Python libraries that SQL cannot express.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 20 days ago.
What is it written in?
Mainly C++, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem OpenMLDB targets: two toolchains for one feature

The README opens with a blunt claim about AI engineering: 95% of the time and effort is consumed by data related workloads. The specific pain it names is narrower and more useful. A data scientist writes a feature engineering script, usually in Python, and that script cannot go to production as written because it does not meet latency, throughput or availability requirements. An engineering team then rewrites it in a database or in C++. Two teams, two toolchains, one feature definition, and a consistency check between them that the README says costs significant time and human power.

OpenMLDB's answer is to remove the second rewrite. The README calls the goal Development as Deployment: you write the feature logic once in SQL, deploy that same SQL online, point the online engine at a real-time data source, and serve. The intended audience is therefore not the data scientist working alone in a notebook. It is the team that has already decided features must be computed the same way at training time and at request time, and is tired of maintaining two implementations. The README also notes the project originated from 4Paradigm's commercial product and was abstracted for community use in 2021, which explains the enterprise-shaped feature list: distributed storage, fault recovery, monitoring, heterogeneous memory.

Two SQL engines and one plan generator

The architecture in the README has four parts. SQL is the unified programming language. A real-time SQL engine serves online requests, and the README states it is built from scratch and optimized for time series data, with response times of a few milliseconds. A batch SQL engine handles offline work, and it is based on a tailored Spark distribution rather than upstream Spark. The fourth part is the one that matters most: a unified execution plan generator that bridges the batch and real-time engines to guarantee consistency.

That last component is the whole design argument. Consistency is not achieved by asking two engines to behave similarly; it is achieved by generating plans from one source so the offline and online paths agree. The README frames the benefit as hassle-free time travel without data leakage, meaning you can compute features for a past point in time for training without accidentally reading data that would not have been available then.

The SQL itself is extended for this domain. The README names LAST JOIN and WINDOW UNION as additions beyond standard SQL. If your feature logic depends on joining a label to the most recent preceding event, or on combining several time windows, those constructs are the reason to look here rather than at a generic SQL engine. If your logic does not, the extensions buy you nothing.

Installing OpenMLDB and running a first query

The README points to a Download and Install section and a QuickStart section, and the repository ships a demo directory with quickstart material for Python, Java and C++. The published artifacts named on the badges are a Docker image at 4pdosc/openmldb, a PyPI package named openmldb, and Maven artifacts com.4paradigm.openmldb:openmldb-batch and com.4paradigm.openmldb:openmldb-jdbc. The demo directory also contains setup_openmldb.sh and a standalone_dist.yml, which suggests a scripted standalone path, though the README excerpt here does not spell out the steps.

The Python route starts with the package:

bash
pip install openmldb

The container route uses the published image:

bash
docker pull 4pdosc/openmldb

What you should see after the pull is the OpenMLDB image on your machine; the README does not document which ports the container exposes, so check the image documentation before mapping ports. For a first real use, the repository's demo/python_quickstart and demo/quick_start directories are the concrete starting points. Building from source is a different commitment: the top-level Makefile sets NPROC to 1 by default with the comment that parallel builds can freeze the system, and it exposes flags such as SQL_PYSDK_ENABLE, SQL_JAVASDK_ENABLE, INSTALL_CXXSDK and TESTING_ENABLE. That default is a signal about build cost, not a formality.

Where OpenMLDB is the wrong tool

OpenMLDB is positioned as a feature platform, and the README's own FAQ says it is mainly positioned that way, while also noting it contains a time-series database used in finance and IoT. Read that as a boundary. If your workload is transactional, with many small updates and deletes and strict per-row consistency, a feature platform is the wrong shape: the value here comes from windowed aggregation over time-ordered data, not from row-level OLTP.

The second limit is the SQL surface itself. Everything in the consistency argument rests on features being expressible in SQL. A feature built from a Python library, a learned embedding computed outside the database, or arbitrary procedural code does not fit the model, and the README does not claim it does. You would be back to two implementations, which is the problem the project exists to remove.

The third is operational. The README lists distributed storage, fault recovery, high availability and scale-out as production features, which means the production story is a cluster, not a single process. The standalone deployment in the demo directory exists, but the README does not document its limits, and the excerpt here does not cover upgrade or rollback procedure at all. Treat the distributed path as the one you are signing up for.

OpenMLDB against a general feature store

The obvious comparison is a feature store such as Feast. The difference in approach is where the computation happens. A typical feature store is a registry plus materialization: you define features, compute them in your own pipeline, and the store handles storage, versioning and serving of already-computed values. OpenMLDB instead computes the features itself, in SQL, in two engines that share a plan generator. Consistency in the first model is a property of your pipeline discipline; in OpenMLDB it is a property of the engine.

That trade is real in both directions. OpenMLDB asks you to express features in its SQL dialect, including LAST JOIN and WINDOW UNION, and to run its engines. A registry-style store asks less of your feature code and more of your pipeline engineering. If you already have a reliable offline pipeline and only need a low-latency serving layer, the registry model is lighter. If the recurring bug in your organization is train-serve skew caused by two implementations of the same feature, the single-plan approach is the more direct fix.

A second comparison is a plain time-series database. OpenMLDB's README acknowledges it contains one, but a time-series database alone does not give you the offline and online feature consistency story, which is the part you are actually buying.

Licence, releases and the cost of staying current

OpenMLDB is Apache-2.0, which is permissive and carries the usual patent grant and attribution requirements; that is a description of the licence text, not legal advice, and you should have your own counsel review it if the deployment is commercially significant. The licence is not a barrier for most adopters.

The release cadence is the thing to look at. The most recent release listed is v0.9.3 from 2025-02-21, preceded by v0.9.2 on 2024-07-27 and v0.9.1 on 2024-07-18. The repository itself is not archived, and its last push was on 2026-09-10, so development activity on the main branch is recent even though the tagged release is older than that. That gap matters for planning: if you track the main branch you are following commits, and if you track releases you are on a version that has been stable for a while.

Upgrade cost is where the material is thin. The README excerpt does not document a rollback procedure, a compatibility policy between minor versions, or how a running cluster is upgraded. The README does list smooth upgrade among the production features, but without steps. For a stateful system holding feature data, that is the first question to raise with the project's Slack or GitHub Discussions before you put it behind production traffic. The demo directory has a standalone_dist.yml and a setup_openmldb.sh, which is where a local upgrade experiment would start.

Editorial conclusion

Adopt OpenMLDB when your feature logic can be expressed in SQL and you are willing to run both a batch engine and a real-time engine to keep training and serving consistent. Do not adopt it as a general OLTP database or if your features live in Python libraries that SQL cannot express. Before committing, verify three things yourself: that the SQL dialect covers your window and join patterns, that the version you install from PyPI or Docker matches the version of the documentation you are reading, and whether you need the standalone deployment or the distributed one, since the README separates them.

Frequently asked questions

What is OpenMLDB used for?

It is positioned as a feature platform for machine learning applications, with low-latency real-time features as its stated strength, and it provides Development as Deployment to cut the cost between offline training and online inference. The README also notes it contains a time-series database used in finance and IoT.

How do I install OpenMLDB?

The README points to its Download and Install and QuickStart sections. Published artifacts include the Docker image 4pdosc/openmldb, the PyPI package openmldb, and Maven artifacts for openmldb-batch and openmldb-jdbc; the repository also ships demo quickstarts for Python, Java and C++.

Does OpenMLDB guarantee the same features offline and online?

The README states that a unified execution plan generator bridges the batch and real-time SQL engines to guarantee consistency, producing consistent features for training and inference. The batch engine is a tailored Spark distribution and the real-time engine is separate, so the guarantee comes from the shared plan generation rather than from a single engine.

What SQL extensions does OpenMLDB add?

The README names LAST JOIN and WINDOW UNION as extended syntax for feature engineering beyond standard SQL. Those constructs are aimed at time-series feature patterns such as joining to the most recent preceding record and combining windows.

Official sources

  1. 4paradigm/OpenMLDB on GitHub
  2. License: Apache-2.0
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/4paradigm-openmldb.svg)](https://hysenlabs.com/projects/4paradigm-openmldb)