Open-source project
alldatacenter/alldata avatar
alldatacenter/alldata

AllData: a GPL-3.0 data platform that assembles twenty open source data tools

🔥🔥 AllData可定义数据中台,以数据平台为底座,以数据中台为桥梁,以机器学习平台为工厂,以大模型应用为上游产品,提供全链路数字化解决方案。产品正式演示体验、社群咨询、商务采购:https://docs.qq.com/doc/DVHlkSEtvVXVCdEFo

3,102 stars978 forksJavaGPL-3.0

At a glance

What is it?
AllData数据中台 is a Chinese data-platform distribution that wires together separate engines for integration, sync, metadata, scheduling and BI. The README is mostly screenshots, so the real question is what the repository actually ships.
Who is it for?
Adopt AllData if you are building a Chinese-language data platform and would rather assemble DataX, SeaTunnel, DolphinScheduler, DataVines and the rest yourself than operate twenty separate deployments, and if you can read the product handbook at yuque.com/aolingdata/product because the README will not tell you how it is wired.
Can I use it commercially?
Yes, with conditions. GPL-3.0 is a copyleft licence: if you distribute software that includes it, you must release that software's source code under the same licence. Running it internally without distributing it does not trigger that obligation.
Is it still maintained?
Yes. The repository last received commits 18 days ago.
What is it written in?
Mainly Java, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What problem AllData数据中台 is trying to solve

A mid-sized company that wants a data platform ends up operating a stack, not a product. Ingestion, database sync, metadata, scheduling, data quality and dashboards each arrive as a separate open source project with its own deployment model, its own UI and its own upgrade cadence. The repository topics list reads like that stack written out: DataX and SeaTunnel for movement, DolphinScheduler and StreamPark for orchestration, DataVines for quality checks, OpenMetadata and Gravitino for catalog and metadata, Paimon for lakehouse storage, DataRT and Supersonic for visualisation, plus Dinky, DBSwitch, Chat2DB, CloudEon, SQLREST, TIS, Amoro, Datasophon, BiSheng and cube-studio. AllData's claim is that it puts one shell over that collection. The README describes the shape in one line: a data platform as the base, a data middle platform as the bridge, a machine learning platform as the factory, and large-model applications as the upstream product. Who it is for is narrower than that sentence suggests. The documentation links point to Yuque, to a Tencent Docs page for community and business contact, and to a Read the Docs site, all in Chinese. If your team cannot work from Chinese documentation, the integration value is hard to reach.

How the pieces relate: base, bridge, factory, upstream

The README gives an architecture diagram, a product function flow diagram, and then a long sequence of screenshots grouped by function: 数据源平台, 数据库同步平台, 数据中枢平台, 数据汇聚平台 split into 数据集成管理, 数据集成平台 and 数据同步平台, and 数据存储平台. What that tells you is that the integration is presented as a set of modules rather than as one runtime. There is no single data flow described in text. The repository layout is more informative than the README prose: a top-level pom.xml makes this a Maven multi-module Java build, and there are separate docker/, install/ and quickstart/ directories, plus moat/ and moat_ui/ modules. The naming suggests the platform is assembled from components that keep their own identities, which is the honest reading of a project whose topics are other projects' names. The trade-off is visible in the same place. A distribution that wraps many engines inherits their version constraints. The 0.7.0 release notes entry is a single line, so the README does not document which upstream versions are pinned, and it does not describe rollback. If you need to know whether a DolphinScheduler upgrade will break the AllData shell, the README will not answer that; the product handbook at yuque.com/aolingdata/product is where the project points.

Installing AllData and running a first quickstart

The README itself gives no install commands. It points to documentation sites and to a quickstart/ directory in the repository, which contains quickstart_bi.md, quickstart_dts.md and quickstart_studio.md. Those three filenames map to the BI, data-sync and studio modules, so the intended first contact is per module rather than a single platform boot. Because the README does not print a command sequence, the honest instruction is to read the quickstart file for the module you want before touching docker/ or install/. The repository is a Maven build, so a local build starts from the parent POM rather than from a container image:

bash
git clone https://github.com/alldatacenter/alldata.git
cd alldata
mvn -f pom.xml clean package -DskipTests

What you should see is Maven resolving the modules declared in pom.xml and writing artifacts into each module's target directory. If your environment cannot reach the dependencies, the build fails at resolution rather than at compile, and the README does not list a mirror configuration. For a container path, the docker/ directory is the place to look; the README does not state which ports the platform binds, so do not assume a default port from this article. Treat quickstart_bi.md as the first document to open if your goal is dashboards, and quickstart_dts.md if your goal is moving data between databases.

The documentation gap is the main limitation

The README is a product brochure. It carries a title, five documentation links, four diagrams, a star-history embed, and then dozens of screenshots with captions like 登录页 and 首页. There is no installation section, no configuration reference, no list of supported data sources in text, no API description and no upgrade guide. For an engineer deciding whether to adopt, that is the decisive fact. You cannot evaluate AllData from the repository alone; you evaluate it from the Yuque product handbook, which is a separate site and is not mirrored in the repository. The second limitation follows from the first. Because the project bundles DataX, SeaTunnel, DolphinScheduler, DataVines, OpenMetadata, Gravitino, Paimon and others, its failure modes are their failure modes plus the shell's own. A connector bug in SeaTunnel becomes an AllData bug from your operators' point of view, and the README does not describe how the project tracks upstream fixes. A third case where it is the wrong tool: if you need exactly one capability, say database sync, installing a whole data middle platform to get DBSwitch or DataX is more surface area than the job requires. Take the component directly.

Where AllData differs from a single-vendor platform

The obvious alternative is a commercial integrated platform, where one vendor owns ingestion, catalog, scheduling and BI and sells them as one licence with one support contract. The difference in approach is not features; it is who resolves version conflicts. In a commercial suite the vendor does it and you pay for that. AllData does the opposite: it keeps the upstream projects as the units, which means you can replace a component, read its source, and hire people who already know DolphinScheduler or SeaTunnel. The cost is that nobody contractually owns the seams. A second alternative is to skip the distribution entirely and deploy the same open source projects yourself with Helm charts or Compose files. That gives you full control over versions and no dependency on AllData's release cadence, which matters because the README does not publish an upgrade path. What you lose is the unified UI and the single login. AllData's value proposition is precisely that unified shell, so the build-your-own route is a real competitor, not a straw man.

Maintenance, releases and the GPL-3.0 question

The repository is not archived, and the last push was on 2026-09-12, the same day as the 0.7.0 release. That is a recent push, so the project is being worked on, but the release history shown here contains a single entry, which means there is no visible track record of how often versions land or how breaking changes are handled. The README does not document rollback or a downgrade procedure. Plan for that gap: before upgrading a production deployment, check the quickstart files and the Yuque handbook for the version you are moving to, and keep your own notes, because the repository will not provide them. On licensing, the repository is GPL-3.0 and also carries LICENSING.md and THIRD_PARTY_NOTICE at the top level. That combination is worth reading before you build anything commercial on top. GPL-3.0 is a copyleft licence, and a platform that aggregates many upstream projects may carry additional terms from those projects recorded in THIRD_PARTY_NOTICE. This is a description of what the repository contains, not legal advice; if you intend to redistribute AllData or embed it in a product, have your own counsel read LICENSING.md and THIRD_PARTY_NOTICE rather than relying on the licence label alone.

Editorial conclusion

Adopt AllData if you are building a Chinese-language data platform and would rather assemble DataX, SeaTunnel, DolphinScheduler, DataVines and the rest yourself than operate twenty separate deployments, and if you can read the product handbook at yuque.com/aolingdata/product because the README will not tell you how it is wired. Do not adopt it if you need English documentation, a published upgrade path, or a licence you can embed in a closed product: the repository is GPL-3.0 and carries a separate LICENSING.md and THIRD_PARTY_NOTICE that you should read before anything else. Verify first that the components you need are actually present in the release you pull, because the README lists capabilities as screenshots rather than as a manifest.

Frequently asked questions

What is AllData数据中台?

It is a Java data platform distribution from alldatacenter that bundles open source data tools such as DataX, SeaTunnel, DolphinScheduler, DataVines, OpenMetadata, Gravitino and Paimon behind a single interface. The README describes it as a data platform base with a data middle platform as the bridge, a machine learning platform as the factory, and large-model applications as the upstream product.

How do I install AllData?

The README does not give install commands. The repository contains docker/, install/ and quickstart/ directories, and the quickstart folder holds quickstart_bi.md, quickstart_dts.md and quickstart_studio.md. The parent pom.xml makes it a Maven multi-module Java build, so a local build runs from that POM.

Is AllData free to use?

The repository is licensed GPL-3.0, so the source is available under that licence. The repository also carries LICENSING.md and THIRD_PARTY_NOTICE, which record terms from the bundled upstream projects, so the single licence label does not tell the whole story for commercial use.

Which open source projects does AllData include?

The repository topics list Amoro, BiSheng, Chat2DB, CloudEon, cube-studio, DataRT, Datasophon, DataVines, DataX, DBSwitch, Dinky, DolphinScheduler, Gravitino, OpenMetadata, Paimon, SeaTunnel, SQLREST, StreamPark, Supersonic and TIS. The README presents their capabilities as screenshots grouped by module rather than as a written component manifest.

Official sources

  1. alldatacenter/alldata on GitHub
  2. License: GPL-3.0
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/alldatacenter-alldata.svg)](https://hysenlabs.com/projects/alldatacenter-alldata)