Apache Kyuubi: a multi-tenant SQL gateway in front of Spark
Apache Kyuubi is a distributed and multi-tenant gateway to provide serverless SQL on data warehouses and lakehouses.
At a glance
- What is it?
- Kyuubi puts a HiveServer2-style Thrift JDBC/ODBC endpoint in front of serverless Spark engines so that administrators can isolate resources and end users can stay in SQL. This review covers the mechanism, the install path, and where the design stops helping.
- Who is it for?
- Adopt Kyuubi when you already run Spark on YARN or Kubernetes, you have a small team that can own deployment and tuning, and your end users should never touch a Spark client. Skip it when a single Spark Thrift Server already satisfies your concurrency and isolation needs, or when nobody on the team can operate the server and engine lifecycle.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 2 days ago.
- What is it written in?
- Mainly Scala, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 28, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem Kyuubi solves: Spark Thrift Server cannot be multi-tenant
The README frames the whole project as a response to one architectural fact about Spark Thrift Server. STS is a single Spark application, and the user and queue it belongs to are fixed at startup. That means it cannot use YARN or Kubernetes to isolate and share resources per caller, and it cannot control access by the single user inside the system. The Thrift Server also lives inside the Spark driver JVM, which the README describes as a coupled architecture that puts high risk on server stability and makes high client concurrency or load balancing hard because the process is stateful.
Kyuubi's answer is to split the server from the engine. The project describes itself as a distributed and multi-tenant gateway that provides serverless SQL on data warehouses and lakehouses, and it exposes a pure SQL gateway through a Thrift JDBC/ODBC interface. The target audience is explicit: system administrators, a small group of Spark experts responsible for deployment, configuration and tuning, and end users who focus on their own business data rather than where it is stored or how it is computed. Anyone who can write SQL is in scope, and the README notes that SQL skills are not even necessary when Kyuubi is paired with Apache Superset for dashboards.
Server and engine separation is the whole design
The mechanism is a loosely coupled two-tier architecture. A Kyuubi server accepts client connections over the Thrift JDBC/ODBC protocol, which the README describes as a HiveServer2-like API, and it hands work to Spark SQL engines that it launches and manages on a cluster manager. Because the engines are separate processes rather than threads inside the server, the server itself is not the Spark driver, and the README states that this separation improves client concurrency and service stability.
The repository layout matches that story. The top level contains kyuubi-server for the gateway, kyuubi-ha for high availability, kyuubi-zookeeper for coordination, kyuubi-metrics and kyuubi-events for telemetry, kyuubi-ctl for command-line control, kyuubi-hive-jdbc and kyuubi-hive-jdbc-shaded for the client driver, plus kyuubi-rest-client, kyuubi-common, kyuubi-util and kyuubi-util-scala as shared modules. There is a charts directory and a docker directory, which implies Helm-based and container-based deployment paths even though the README does not spell them out. The multi-tenancy is not just a label: the README says the server and engines' multi-tenant architecture gives administrators a way to achieve computing resource isolation, data security, high availability and high client concurrency.
On the data side, the project positions itself as a single portal for different workload types. ETL processing and BI analytics are meant to run on one copy of data behind one SQL interface. The README lists logical view support through the Kyuubi DataLake Metadata APIs, multiple catalogs, and SQL standard authorization for the data lake as coming. Note the tense: authorization is described as coming, not shipped, so anyone evaluating Kyuubi as a security boundary should treat that item as unfinished.
Installing Kyuubi and running a first query
The README does not contain install steps. It points to the online documentation and to a Getting Started page under the quick_start path on kyuubi.readthedocs.io, and the project publishes Docker images under the apache/kyuubi repository on Docker Hub, which the README references through its badge. The repository also ships a docker directory and a charts directory for Kubernetes. Because the README gives no concrete commands, the only honest instruction is to follow the quick start guide for the release you intend to run, and to use the container or chart path if you do not want to build from source.
What the README does give is the shape of the client interaction. Kyuubi speaks Thrift JDBC/ODBC, so from an application's point of view you connect the way you would connect to HiveServer2. The repository ships kyuubi-hive-jdbc and a shaded variant, which is the driver you would place on a client classpath. The README does not print a connection string, a default port or any client command, so there is nothing to copy here: take the JDBC URL, driver class and port from the documentation for your release rather than from this page. The one thing the README makes clear is the endpoint type. You connect to the Kyuubi server over Thrift JDBC/ODBC, and after connecting you are issuing ordinary SQL against a Spark SQL engine that the server provisioned for you. That is the point of the product: the user never starts a Spark application.
For deployment, the repository's charts directory suggests a Helm-based install for Kubernetes, and the docker directory plus the Docker Hub image suggest a container path. The README does not document either in detail, so treat the chart values and the image tag as things to read in the repository rather than things described in the project's front page.
Where Kyuubi is the wrong tool
Kyuubi is infrastructure, and it is priced accordingly in operational attention. The README is direct that a small group of Spark experts is expected to own deployment, configuration and tuning. If your organization has no one who can operate a Spark cluster, adding Kyuubi does not remove that requirement; it adds a stateful gateway, a high-availability layer, a ZooKeeper dependency and a metrics pipeline on top of it. The module list is a fair proxy for the surface area: kyuubi-ha, kyuubi-zookeeper, kyuubi-metrics and kyuubi-events are separate components that each need configuration.
There is also a class of workload where the gateway is pure overhead. If you have one team, one queue and modest concurrency, a single Spark Thrift Server already gives you a JDBC endpoint and no extra hop. Kyuubi's value comes from multi-tenancy, isolation and concurrency, and none of those are problems you have. The README's own comparison is worth reading literally: it criticizes STS for being unable to control access by the single user inside the system, which only matters when there is more than one user.
The third boundary is security expectations. The README lists SQL standard authorization for the data lake as coming, not delivered. An administrator who reads multi-tenant and assumes row-level or table-level policy enforcement out of the box is reading ahead of the project. Resource isolation through the cluster manager and authentication of callers are different things from fine-grained authorization inside the SQL layer, and the README does not claim the latter is finished.
Kyuubi against Apache Livy and plain Spark Thrift Server
The closest conceptual alternative is Apache Livy, which also sits between clients and Spark and also targets cluster managers. The difference in approach is the protocol and the unit of tenancy. Livy exposes a REST interface and manages Spark sessions and batches as resources you create through HTTP calls, which suits programmatic submission from notebooks and pipelines. Kyuubi exposes a Thrift JDBC/ODBC interface that is HiveServer2-like, so existing BI tools, JDBC drivers and SQL clients connect without a rewrite. If your consumers are SQL clients and dashboards, that protocol choice is the deciding factor, and it is also why the README can claim that SQL skills are unnecessary when Superset is in front.
The other alternative is the thing Kyuubi replaces: Spark Thrift Server. The README's argument is structural rather than feature-based. STS is one Spark application with a fixed user and queue, with the Thrift Server coupled into the driver JVM. Kyuubi decouples the server from the engines and uses multi-tenancy to interact with cluster managers for resource sharing and isolation. If you do not need per-caller isolation, the simpler component wins on operational cost. If you do, the coupling is not something you can configure away.
Maintenance, releases and the Apache-2.0 licence
The repository is not archived, and the last push was on 2026-09-28. Releases are frequent enough to plan around: v1.12.0 on 2026-07-28, v1.11.1 on 2026-04-13 and v1.10.3 on 2025-12-26. A cadence with a minor release roughly every quarter and patch releases in between means upgrade work is a recurring line item, not a one-time cost. The repository carries a LICENSE, LICENSE-binary, NOTICE, NOTICE-binary and licenses and licenses-binary directories, which is the standard Apache layout for a project that redistributes bundled dependencies, and the README's own header is the Apache License, Version 2.0.
For adopters, the licence is permissive and the practical question is not legal but operational: which Spark versions a given Kyuubi release supports, and whether your cluster manager is covered by the engine deployment modes that release documents. The README discusses YARN and Kubernetes as cluster managers for engines, and the repository's charts and docker directories point at Kubernetes and container deployments. It does not enumerate supported Spark versions, so that check belongs in the release notes and the documentation for the version you pick. Nothing here is legal advice; if you redistribute Kyuubi inside a product, read the NOTICE and LICENSE-binary files yourself.
Editorial conclusion
Adopt Kyuubi when you already run Spark on YARN or Kubernetes, you have a small team that can own deployment and tuning, and your end users should never touch a Spark client. Skip it when a single Spark Thrift Server already satisfies your concurrency and isolation needs, or when nobody on the team can operate the server and engine lifecycle. Before rolling it out, verify that the current release's engine deployment mode matches your cluster manager, and read the configuration reference for the specific keys you intend to change, because the README itself documents none of them.
Frequently asked questions
What is Apache Kyuubi?
It is a distributed and multi-tenant gateway that provides serverless SQL on data warehouses and lakehouses, exposing a Thrift JDBC/ODBC interface that the README describes as a HiveServer2-like API. It runs Spark SQL engines on your behalf so that end users only need SQL.
How does Kyuubi relate to Spark Thrift Server?
The README argues that Spark Thrift Server is a single Spark application with a fixed user and queue, and that its Thrift Server is coupled into the driver JVM, which limits concurrency and high availability. Kyuubi decouples the server from the engines and uses multi-tenancy to interact with cluster managers for resource sharing and isolation.
Which cluster managers can Kyuubi engines run on?
The README states that Kyuubi can deploy its engines on different kinds of cluster managers, naming Hadoop YARN and Kubernetes.
Does Kyuubi provide fine-grained SQL authorization?
The README lists SQL standard authorization for the data lake as coming, so it is not presented as a delivered capability. Resource isolation through the cluster manager and caller authentication should not be read as the same thing.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/apache-kyuubi)