Self-hosted service
spitfireuptown/datalinkx avatar
spitfireuptown/datalinkx

DatalinkX: a self-hosted console for moving data between HTTP, Oracle, MySQL and Elasticsearch

🔥🔥DatalinkX异构数据源之间的数据同步系统,支持海量数据的增量或全量同步,同时支持HTTP、Oracle、MySQL、ES等数据源之间的数据流转,支持中间transform算子如SQL算子、大模型算子,底层依赖Flink、Seatunnel引擎,提供流转任务管理、任务级联配置、任务日志采集等功能🔥🔥

382 stars51 forksJavaNOASSERTION

At a glance

What is it?
DatalinkX wraps Flink, ChunJun and SeaTunnel in a Spring Boot web console so teams can define, schedule and chain cross-database sync jobs from a browser. The trade-off is a stack of engines you now have to operate.
Who is it for?
Adopt DatalinkX if you already run MySQL 8, Redis and xxl-job and you want scheduled cross-database syncs defined in a browser instead of hand-written Flink jobs. Do not adopt it if you need per-user permissions, real-time sources beyond Kafka, or a documented migration path between releases.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository last received commits 93 days ago.
What is it written in?
Mainly Java, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem DatalinkX addresses: sync jobs scattered across teams

The README states the case plainly: once a company is large enough that departments exchange data, someone needs a shared place to define and watch those transfers. Its own example is a crawler team pushing scraped data on a schedule into a warehouse team's database. Without a shared service, each team writes its own script, keeps its own logs, and nobody can answer which job last ran or what it wrote.

DatalinkX is aimed at that internal platform role rather than at individual developers. It is a Java and Spring Boot application with a Vue 2 and Ant Design front end, so the people who use it are data engineers and platform engineers who want a UI for creating data sources and sync tasks. The project description lists HTTP, Oracle, MySQL and Elasticsearch as sources it moves data between, and the README notes batch jobs, real-time jobs, and compute jobs with transform operators. Real-time jobs, per the README, only support Kafka.

How a DatalinkX job actually runs: job_graph, engines and scheduling

The unit of work is a job_graph. In a batch task you configure from_db and to_db, and the README describes that pair as what constructs the graph. The compute task type goes further: you configure canvas information, and the graph can include transform operators such as an SQL operator or a large-model operator. That canvas is the project's answer to the fact that a sync is rarely a straight copy.

Execution is delegated. The README lists Flink 1.10.3 and SeaTunnel 2.3.8 as distributed compute engines, with ChunJun (formerly FlinkX) 1.10_release named as the sync framework. The top-level repository layout reflects this: there are flinkx/ and apache-seatunnel-2.3.8/ directories alongside the Java modules. So DatalinkX is not itself a data movement engine. It compiles your UI configuration into work for one of those engines and then reports on it.

Scheduling is handled by xxl-job 2.3.0, and the README says you set a cron expression to trigger a sync task. Redis 5.0 or later serves as both cache and message middleware, using Redis Stream. OpenFeign 3.1.9 handles RPC between services, and the repository splits into datalinkx-server, datalinkx-job, datalinkx-connector, datalinkx-feign, datalinkx-messagehub, datalinkx-sse, datalinkx-copilot and datalinkx-common. The connector module is where the plugin loading lives: the README says you can develop a custom driver to a fixed rule, drop it into driver-dist, and it becomes usable.

Installing DatalinkX and creating a first sync task

The README does not publish a step-by-step install guide. It points to a separate documentation set (described as 92 documents) hosted on a Youdao note, and it names Docker and Docker Compose as the deployment route, with a docker/ directory in the repository. Since no install commands appear in the README, the honest starting point is to read that linked documentation and inspect docker/ rather than guess at a compose invocation.

The usage flow itself is documented, and that is what you can follow once the stack is up. You log in with the default credentials admin / admin. The README states there is no permission control, so anyone who can reach the login page is effectively an administrator.

After login, the order is data source, then task. A batch task is defined by choosing from_db and to_db, and the README describes that pair as constructing the job_graph. For a compute task you configure canvas information and attach transform operators, and for a real-time task you again set from_db and to_db, with Kafka as the only supported real-time source. Scheduling is then a cron expression in xxl-job, which the README shows as a screenshot rather than as text. Cascading and lineage are separate screens: one chains tasks together, the other draws the relationship between them.

Where DatalinkX is the wrong tool

The permission model is the first hard limit. The README says outright that login has no permission control. If your data sources include anything regulated, or if you need to prove who changed a job, this build does not give you that. The README also advertises an enterprise edition, DatalinkX-Pro, whose stated differences are a more complete user permission system, more database plugins, 5x24 support and custom development. That is a clear signal about where the open project stops.

Real-time scope is the second limit. The README says real-time tasks support Kafka only. If your change stream comes from a database binlog or another broker, this is not the tool for that path in its documented form.

The third issue is version weight. Flink 1.10.3, ChunJun 1.10_release and SeaTunnel 2.3.8 span three different generations of the same ecosystem. You are not just running DatalinkX; you are running all of those runtimes, plus MySQL 8, Redis 5 and xxl-job. For a team that only needs a nightly table copy between two MySQL instances, that is a large amount of infrastructure to keep patched for a small amount of movement. A single scheduled script or a lighter tool would be the better call.

Finally, the README does not document rollback, retry semantics, or what happens to an in-flight Flink job when a task is edited. Those are the questions to ask before trusting it with production pipelines.

How DatalinkX differs from running SeaTunnel or Flink directly

The obvious alternative is to use SeaTunnel or Flink on their own, since DatalinkX sits on top of both. The difference is where the definition of a job lives. With SeaTunnel you write a config file per job and manage those files, their secrets and their scheduling yourself. DatalinkX turns that into rows in a MySQL-backed UI: data sources are registered once, jobs are built from them, xxl-job triggers them, and logs are collected centrally. The README's stated benefit is exactly this consolidation of task management and logs.

That convenience has a cost in transparency. When a DatalinkX job fails, you are debugging a generated job on a Flink 1.10.3 cluster, and the README does not describe how the generated job maps back to the UI fields. With hand-written SeaTunnel configs you own every line. ChunJun is the other comparison point and it is already inside the project as the sync framework, so choosing DatalinkX over ChunJun alone is choosing a management layer, not a different engine. If your team is three people and your pipelines are stable, the management layer is the part you can skip.

Maintenance, licensing and the upgrade question

The repository is not archived, and the last push was on 2026-06-29. That is recent enough that the project is not dormant, but the README records no releases at all, so there is no versioned artifact, no changelog and no stated compatibility contract between builds. Upgrades therefore mean tracking the main branch and re-checking that your custom connector drivers, which the README says live in driver-dist, still load against a changed connector interface. The bundled flinkx/ and apache-seatunnel-2.3.8/ directories also mean engine upgrades arrive as repository changes rather than as dependency bumps you can pin independently.

The licence is the other thing to settle before you build on it. The repository metadata reports NOASSERTION, which means the licence could not be identified automatically, and the README says nothing about licensing terms. Read the LICENSE file in the repository root and decide from the actual text whether your use is covered. This is not a formality for a component that sits between production databases. Note also that the README links an enterprise edition with features the open project does not have, so check whether the parts you depend on are in the open tree or only in that edition.

Editorial conclusion

Adopt DatalinkX if you already run MySQL 8, Redis and xxl-job and you want scheduled cross-database syncs defined in a browser instead of hand-written Flink jobs. Do not adopt it if you need per-user permissions, real-time sources beyond Kafka, or a documented migration path between releases. Before committing, verify the licence text in the LICENSE file, check which connector drivers ship under datalinkx-connector, and confirm that the Flink 1.10.3 and SeaTunnel 2.3.8 versions in the README are versions you are willing to run in production.

Frequently asked questions

What is DatalinkX used for?

It is a heterogeneous data source sync system that moves data between sources such as HTTP, Oracle, MySQL and Elasticsearch, and manages those sync tasks. The README frames it as a shared internal service for teams that exchange data between departments.

Which compute engines does DatalinkX use?

The README lists Flink 1.10.3 and SeaTunnel 2.3.8 as the distributed compute engines, with ChunJun (formerly FlinkX) 1.10_release as the sync framework. Scheduling runs on xxl-job 2.3.0.

Does DatalinkX support real-time synchronization?

It has a real-time task type, but the README states that real-time tasks only support Kafka. Other change sources are not documented.

How do I install DatalinkX?

The README names Docker and Docker Compose as the deployment method and the repository contains a docker/ directory, but it does not list install commands. It points to a separate documentation set for setup details.

Does DatalinkX have user permissions?

The README says login uses the default admin / admin credentials and that there is no permission control. A more complete user permission system is listed as a feature of the separate enterprise edition.

Official sources

  1. Issues
  2. README
  3. spitfireuptown/datalinkx on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/spitfireuptown-datalinkx.svg)](https://hysenlabs.com/projects/spitfireuptown-datalinkx)