Open-source project
apache/cloudberry avatar
apache/cloudberry

Apache Cloudberry: a Greenplum fork that moved its PostgreSQL kernel forward

One advanced and mature open-source MPP (Massively Parallel Processing) database. Open source alternative to Greenplum Database.

1,407 stars248 forksCApache-2.0

At a glance

What is it?
An incubating MPP database written by the original Greenplum developers, carrying three build systems in one tree, a greenplum-prefixed directory layout, and an MCP server living next to the database source. Useful if you are already a Greenplum shop, and a poor fit if you are not.
Who is it for?
Adopt Cloudberry if your team already runs Greenplum and you have a specific reason to need a newer PostgreSQL kernel, because that is the trade the project exists to offer and the gp-prefixed layout means the operational knowledge transfers. Do not adopt it to replace a modern PostgreSQL with a data warehouse, since the design assumes a distributed MPP engine with its own storage layout and extension framework.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 11 days ago.
What is it written in?
Mainly C, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 20, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Two forks deep, with the greenplum prefixes still in the tree

The lineage is the first fact to establish, because it explains almost everything else. The project is described as created by the original developers of Greenplum Database, evolving from the open-source version of the Pivotal Greenplum Database, with a newer PostgreSQL kernel and more advanced enterprise capabilities. So the structure is Greenplum forked PostgreSQL to add MPP, and Cloudberry forked Greenplum to move the kernel forward. That second fork is a defensible decision and a costly one, and the cost is visible in the directory names. There is gpAux, gpMgmt and gpcontrib at the top level, three module groups carrying the prefix of a project that is no longer the upstream. The topics carry greenplum as well. For an adopter that cuts both ways. The operational knowledge transfers, since scripts, roles and extension expectations from a Greenplum estate mostly apply, which is the reason to choose this over starting fresh. The code is also not a thin patch, because a kernel upgrade touches the storage layer, the planner and the process model, so you should expect version-specific behaviour rather than drop-in equivalence. The claim of advanced and mature in the project description is its own; the test of it is your own workload.

Three build systems, and a Makefile that refuses to run before configure

Count the build entry points in the top-level listing and there are three families. There is a Makefile alongside GNUmakefile.in, with configure, configure.ac, aclocal.m4 and config/, which is the PostgreSQL autotools layout. There is meson.build with meson_options.txt, which is the newer build system PostgreSQL has been adopting. And there is pom.xml, which is Maven, for the Java side. That is a tree in the middle of a build migration, and it is why the build documentation lives off the repository, on the deployment guides linked from the README for Linux including RHEL and Rocky Linux and Ubuntu, and for macOS. The Makefile is the most revealing file, and it is inherited rather than written for this project. Its first job is to insist on GNU make, searching the path for gmake, gnumake or make and testing each one, and its second job is to refuse to proceed if configure has not been run, printing an error that points at an installation document. The message it prints is this:

bash
https://www.postgresql.org/docs/devel/installation.html

That is a PostgreSQL URL in a Greenplum-lineage database repository, and it is the clearest single illustration of what this project is. If you are building from source, expect to follow a PostgreSQL-shaped build process with Greenplum extensions in it, and expect the guides to be the authority rather than the tree.

Incubating, and a 2.x number reached across a major bump

The version history is short and worth reading exactly. Version 1.6.0 was released on 2024-09-03 with a plain tag. Then 2.0.0-incubating on 2025-09-01, and 2.1.0-incubating on 2026-04-15. So the project was already incubating at 1.6.0, took a major version bump while incubating, and is still incubating at 2.1.0. The repository is not archived and the last push was on 2026-09-18, so the work is current and nothing here suggests abandonment. What the sequence does tell you is that this is a young foundation project carrying a large inherited codebase, which is a specific and manageable risk rather than a general one. You are not choosing a young codebase, you are choosing a young project around an old and well-tested one, and those have different failure modes. The young-project risk is process: documentation that is still being written, deployment guides that change, and a graduation process that may reshape governance. The inherited-codebase risk is the one you would have with the upstream too, so it does not differentiate this from Greenplum. The files that make this visible are the usual foundation set, with a DISCLAIMER, a NOTICE, a CODE_OF_CONDUCT and a SECURITY.md, plus a security policy page for reporting a vulnerability. The Python side keeps its dependencies in a single file, python-dependencies.txt, and version handling goes through a getversion script and a putversion script, which is the inherited convention rather than a modern one.

An MCP server in the database tree, beside a template for coding agents

The top-level listing contains a directory called mcp-server, and it sits in the same repository as the query engine. That is an unusual placement and it deserves a decision rather than a shrug. A Model Context Protocol server is a surface for tool use: it exposes operations that a model can call. Inside a data warehouse, the operations that matter are queries, so you are placing an agent-facing interface on the same tree as the storage layer, with the same review, the same release and the same version. That is convenient, because an agent can inspect the actual schema rather than a description of it, and it is a widening of the trust boundary, because whatever the MCP server can do is now reachable by a model. The project has also written down its position on automated contribution, which is unusually far ahead. There is an AI_GUIDELINE.md and an AGENTS.md.template, so contributors are told how to use AI assistance when writing code, and a template means the repository expects a per-contributor agent configuration file. Combined with a mcp-server directory, that is a coherent stance: the project is preparing for agents both as consumers of the database and as participants in its development. Whether the boundary is in the right place is a question worth asking before you deploy it, and the presence of sonar-project.properties and an ABI check directory suggests the maintainers are already measuring change carefully.

The product is five repositories, and PXF is the one you may need

The README lists ecosystem repositories beside the main one, and the division is meaningful. There is a site repository holding the website and documentation sources, which is why the documentation is versioned separately from the database and can describe a version you are not running. There is a backup utility, which tells you that taking a backup is not part of the core distribution and is a component you have to select. There is a Go libraries repository, which implies first-class Go clients maintained alongside the database. And there is the Platform Extension Framework, PXF, in its own repository. PXF is the one to look at carefully, because it is the mechanism for reaching data that is not in the database, which for an MPP warehouse used as a query layer over object storage is not a detail but the main event. It is also inherited from Greenplum, so it is one of the pieces that came with the fork. The practical consequence of the split is version management. Your database, your PXF and your backup utility are three releases that have to be compatible, and the main repository is not where that compatibility matrix lives. The repository is the engine; the rest of the surface is somewhere else.

ABI checks, Coverity, Sonar and a licence audit in the CI surface

The badges and directories describe how a fork of a large C database gets maintained, and the list is unusually specific. There is a .abi-check directory, which matters more than its size suggests. A database with a C extension interface has a binary compatibility contract with anything compiled against it, and an ABI check in the build means a change that breaks that contract fails before release rather than after a user loads an extension. There is Coverity Scan for static analysis and a SonarQube Cloud badge, so two independent analysers rather than one. There is a .clang-tidy file for C style, a sonar-project.properties file that configures the analyser, and an Apache RAT audit workflow, which checks that every file carries an acceptable licence header. That last one is the standard foundation guard against accidentally committing code with the wrong terms, and in a project with a vendored licences directory and a long inherited history it is not ceremonial. The build itself is a GitHub Actions workflow named build-cloudberry, so unlike some foundation projects the visible build history is on GitHub. There is also a .gitmodules file, so some dependencies arrive as submodules, and a .git-blame-ignore-revs file, which is the standard remedy for a history rewritten in bulk, common in a project that arrives through a donation or a rename.

Try the sandbox before you read a single build guide

For evaluation, the low-risk path is documented and it is not a from-source build. The README points to a Docker-based sandbox in devops/sandbox, described as tailored to help you gain a basic understanding of Cloudberry's capabilities and features. That directory is in this repository, so the sandbox is versioned with the code rather than published separately, which means the sandbox you try matches the branch you read. Spend the first hour there. You will learn how the MPP engine is deployed, how the coordinator and the segment hosts are arranged, and what the administrative surface looks like, without a from-source build that requires configure, GNU make and three build systems. The wrong-tool boundary is then easy to state. Cloudberry is a choice against PostgreSQL, not an addition to it, because an MPP engine has its own storage layout, its own distribution and its own extension model, and PXF is how you reach data outside it. If your requirement is a single-node analytical database with extensions you can install from a package manager, this is a much larger machine. If your requirement is distributed scans and concurrent analytics over a large fact set, and your team already knows Greenplum, then the PostgreSQL kernel upgrade is the whole argument, and the sandbox is how you check it against your own data.

Editorial conclusion

Adopt Cloudberry if your team already runs Greenplum and you have a specific reason to need a newer PostgreSQL kernel, because that is the trade the project exists to offer and the gp-prefixed layout means the operational knowledge transfers. Do not adopt it to replace a modern PostgreSQL with a data warehouse, since the design assumes a distributed MPP engine with its own storage layout and extension framework. Three things to check first. Try the Docker sandbox in devops/sandbox before reading any build guide, because a from-source build runs configure, needs GNU make and lives across three build systems with no single canonical entry point in the tree. Confirm where the extension layer lives, because PXF is a separate repository along with the backup utility and the Go libraries. And read the version line honestly: 1.6.0, then 2.0.0-incubating in 2025, then 2.1.0-incubating in 2026, so the project crossed a major version under incubation and is not yet graduated.

Frequently asked questions

What is Apache Cloudberry and how is it related to Greenplum?

It is described as created by the original developers of Greenplum Database, evolving from the open-source version of the Pivotal Greenplum Database with a newer PostgreSQL kernel. The inherited layout is still visible, with gpAux, gpMgmt and gpcontrib directories at the top level.

How do I try Apache Cloudberry without building it from source?

Build the Docker-based sandbox in the devops/sandbox directory, which the README describes as a way to gain a basic understanding of Cloudberry's capabilities. Build-from-source instructions are separate deployment guides covering Linux, including RHEL and Rocky Linux and Ubuntu, and macOS.

What is the current Apache Cloudberry release?

2.1.0-incubating, published on 2026-04-15, following 2.0.0-incubating in September 2025 and 1.6.0 in September 2024. The project is still incubating and the last push to the repository was on 2026-09-18.

Where does Apache Cloudberry keep its extension framework?

In a separate repository, the Platform Extension Framework under apache/cloudberry-pxf, alongside a backup utility repository, a Go libraries repository and a site repository holding the documentation sources. The main repository is the database engine.

What build systems does the Apache Cloudberry repository contain?

Three families: the PostgreSQL-style autotools layout with a Makefile, GNUmakefile.in, configure and config; a meson.build with meson_options.txt; and a Maven pom.xml. The inherited Makefile searches for GNU make and refuses to build before configure has been run.

Official sources

  1. apache/cloudberry on GitHub
  2. License: Apache-2.0
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/apache-cloudberry.svg)](https://hysenlabs.com/projects/apache-cloudberry)