# Apache Linkis: A Computation Middleware Layer Between Applications and Data Engines

> Apache Linkis puts a REST, WebSocket and JDBC gateway in front of Spark, Hive, Presto and Flink so applications stop talking to engines directly. It is a Java-based Apache project under Apache-2.0, and the trade-off is operational weight.

**apache/linkis** — Apache Linkis builds a computation middleware layer to facilitate connection, governance and orchestration between the upper applications and the underlying data engines.

- Repository: https://github.com/apache/linkis
- Website: https://linkis.apache.org/
- Stars: 3,412 · Forks: 1,168
- Language: Java
- License: Apache-2.0
- Published: 2026-09-23 · Updated: 2026-09-23 · Language: en
- Canonical page: https://hysenlabs.com/projects/apache-linkis

## The problem Linkis solves: applications talking to too many engines

A big data platform usually accumulates engines. Spark for batch and interactive SQL, Hive for the warehouse, Presto or Trino for federated queries, Flink for streaming, plus JDBC sources. Each has its own client, its own session model and its own way of handling a user's JARs, UDFs and variables. The application layer ends up encoding all of that.

Linkis inserts a computation middleware layer between the two. The README describes it as a layer that facilitates connection, governance and orchestration between upper applications and underlying data engines, and says the point is to decouple the application layer from the engine layer and simplify the network call relationship. The audience is platform and data infrastructure teams, not individual analysts.

The README claims more than 700 trial companies and 1000+ sandbox trial users since the first release in 2019, spanning finance, banking, telecom and manufacturing. Those are adoption claims from the project itself, not independently verified numbers, and they say nothing about whether the deployment was small or large.

## How the middleware layer works: interfaces, engines and shared context

Upper applications reach Linkis through standard interfaces: REST, WebSocket and JDBC. Those are the only entry points an application needs to know. Behind them, Linkis manages the engines.

The README lists the engines it can front: Spark, Hive, Python, Shell, Flink, JDBC, Pipeline, Sqoop, OpenLooKeng, Presto, ElasticSearch, Trino and SeaTunnel. Languages include SparkSQL, HiveSQL, Python, Shell, Pyspark, Scala, JSON and Java. The repository layout shows how this is packaged: linkis-engineconn-plugins holds the engine connectors, linkis-computation-governance covers routing and governance, linkis-orchestrator handles task orchestration, and linkis-spring-cloud-services plus linkis-public-enhancements provide the service layer. A separate linkis-web and linkis-web-next hold the front end.

The part that distinguishes Linkis from a plain query gateway is the context service. The README states it supports cross-user, system and computing engine association and management of resource files (JAR, ZIP, Properties), result sets, parameter variables, functions and UDFs, summarized as one setting, automatic reference everywhere. So a UDF registered once is available to a Spark job and a Hive query without re-uploading it. Governance is label-based: the README lists task routing, load balancing, multi-tenant, traffic control and resource control based on multi-level labels. There is also a unified data source management layer for Hive, ElasticSearch, MySQL, Kafka and MongoDB, with version control, connection testing and metadata queries, plus an error code catalog for common task failures.

One thing the README does not document is the internal scheduler behaviour under contention, or how routing decisions are made when labels conflict. The README describes the capability, not the algorithm.

## Installing Apache Linkis and running a first engine task

The README points to the documentation site at linkis.apache.org for installation, not to inline steps. It also shows a quick-build.sh script at the repository root, alongside quick-build.cmd and quick-build.ps1 for Windows, and a Maven wrapper (mvnw, mvnw.cmd) with a pom.xml. The JDK badge in the README states JDK 8, so that is the baseline to match before building.

A build from a checkout therefore starts like this. The Maven wrapper avoids depending on a locally installed Maven, and the script wraps the multi-module build.

```bash
./mvnw clean install -DskipTests
./quick-build.sh
```

Expect a long build: the repository is split into linkis-commons, linkis-computation-governance, linkis-engineconn-plugins, linkis-orchestrator, linkis-public-enhancements, linkis-spring-cloud-services and linkis-dist, and the engine connectors pull in engine-specific dependencies. The README's Engine Type table notes that Spark and Hive connectors are included in the release package by default, with default dependency versions of Apache Spark 3.2.1 and Apache Hive 3.1.3 respectively, and that both require Linkis 1.0.3 or later. Because the README does not list deployment steps, the exact service startup order has to come from the documentation site rather than from this file.

Once a deployment is up, the surface an application uses is the REST interface. The README does not include a request example, so the practical first step is to read the interface documentation on linkis.apache.org and issue a submission against a Hive or Spark engine, then confirm the task appears in the Linkis web front end.

## Where Linkis is the wrong tool

Linkis is middleware, and middleware has to run. A team with a single Spark cluster and one application does not need a layer between them; the application can submit to Spark directly and skip an entire set of services to operate. The README's own framing, decoupling the application layer from the engine layer, only pays off when there are multiple engines or multiple applications to decouple.

Version coupling is the second constraint. The Engine Type table ties each connector to a supported component range and a minimum Linkis version. Spark is listed as Apache >= 2.0.0 and CDH >= 5.4.0 with a default of Apache Spark 3.2.1, and Hive as Apache >= 1.0.0 and CDH >= 5.4.0 with a default of Apache Hive 3.1.3. If your cluster runs a version outside those ranges, the connector is not promised to work, and the README does not describe a fallback path.

The third is the JDK 8 badge. The README presents JDK 8 as the supported Java version and does not state a newer baseline, so a platform standardised on a later JDK has to reconcile that before adopting.

Finally, the context service is a shared namespace. The README describes sharing materials across users and systems as a feature. In a multi-tenant deployment that same mechanism is a boundary you have to configure, and the README does not document the isolation model in detail.

## Linkis compared with Apache Livy and a plain JDBC gateway

Apache Livy is the closest comparison and appears in the repository topics. Livy is a REST interface to Spark: submit a session or a batch, get back a job. It is narrower by design. Linkis covers Spark but also Hive, Presto, Flink, Trino, SeaTunnel and JDBC sources behind one interface, and adds the context service, label-based governance and data source management that Livy does not attempt.

The difference in approach matters operationally. With Livy you run one service next to your Spark cluster; with Linkis you run a set of services (computation governance, orchestrator, public enhancements, Spring Cloud services, plus the web front end) and gain cross-engine resource sharing in return. If your only engine is Spark, Livy is the smaller thing that does the job. If you are already juggling Hive, Presto and Spark and re-uploading the same UDF to each, the Linkis context service is the feature that Livy has no answer for.

A plain JDBC gateway is the other end of the spectrum. It forwards SQL and nothing else. Linkis accepts SparkSQL, HiveSQL, Python, Shell, Pyspark, Scala, JSON and Java through REST, WS and JDBC, so it is a task submission layer rather than a query proxy.

## Maintenance, releases and the Apache-2.0 licence

The repository is not archived and the last push was on 2026-09-13, so the codebase is being changed. Release cadence is visible in the tags: release-1.8.0 on 2025-10-17, 1.7.0 on 2025-01-09 and 1.6.0 on 2024-07-12. That is roughly one release per nine to twelve months, which means engine connector updates for a new Spark or Hive version arrive on that schedule rather than immediately.

Upgrade cost is dominated by that coupling. Because connectors declare supported component ranges and minimum Linkis versions, moving to a new engine version can mean moving Linkis too. The README does not document a rollback procedure or a compatibility matrix beyond the Engine Type table, so an upgrade plan should be built from the release notes and the documentation site rather than from the README.

Licensing is Apache-2.0, stated in the README badge and in the LICENSE file at the repository root. Apache-2.0 permits commercial use and modification and includes a patent grant. The repository also carries a NOTICE file and a licenses/ directory, which is where bundled third-party dependency licences are recorded; anyone redistributing a built distribution should read those rather than assume the top-level licence covers every bundled jar. This is a description of the files present, not legal advice.

## Conclusion

Adopt Linkis if you already run several engines and need one entry point with shared variables, UDFs and resource files across Spark, Hive, Presto or Flink. Do not adopt it for a single engine, or if you cannot run the full set of Linkis services and their dependencies. Before committing, verify the engine version you actually run against the Engine Type table in the README, and confirm the JDK 8 requirement, since the README badge states JDK 8 and does not document a newer baseline.

## FAQ

### What is Apache Linkis used for?

It is a computation middleware layer that sits between upper applications and underlying data engines. Applications use its REST, WS or JDBC interfaces to reach engines such as Spark, Hive, Presto and Flink, and to share variables, scripts, UDFs and resource files across them.

### Which engines does Apache Linkis support?

The README lists Spark, Hive, Python, Shell, Flink, JDBC, Pipeline, Sqoop, OpenLooKeng, Presto, ElasticSearch, Trino and SeaTunnel. Spark and Hive connectors are included in the release package by default, with default dependency versions of Apache Spark 3.2.1 and Apache Hive 3.1.3.

### What Java version does Apache Linkis require?

The README carries a JDK 8 badge, and it does not state a newer Java baseline. The repository also ships a Maven wrapper (mvnw) and a quick-build.sh script for building from source.

### Is Apache Linkis actively maintained?

The repository is not archived and its last push was on 2026-09-13. Recent releases are release-1.8.0 on 2025-10-17, 1.7.0 on 2025-01-09 and 1.6.0 on 2024-07-12.

## Sources

- [apache/linkis on GitHub](https://github.com/apache/linkis)
- [License: Apache-2.0](https://github.com/apache/linkis/blob/master/LICENSE)
- [Project website](https://linkis.apache.org/)
- [README](https://github.com/apache/linkis/blob/master/README.md)
- [Releases](https://github.com/apache/linkis/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/apache-linkis
