# Apache Polaris: an Iceberg REST catalog for multi-engine setups

> Apache Polaris implements the Iceberg REST catalog API so Spark, Trino, Flink and other engines can share one metadata layer. Here is how it is built, how to run it locally, and where it stops being the right answer.

**apache/polaris** — Apache Polaris is an interoperable metadata catalog for Apache Iceberg tables designed for cross-engine compatibility.

- Repository: https://github.com/apache/polaris
- Website: https://polaris.apache.org/
- Stars: 2,067 · Forks: 532
- Language: Java
- License: Apache-2.0
- Published: 2026-08-04 · Updated: 2026-08-18 · Language: en
- Canonical page: https://hysenlabs.com/projects/apache-polaris

## The problem Apache Polaris solves: one Iceberg catalog, many engines

Iceberg tables store their metadata in the table format itself, but engines still need somewhere to resolve a table name to a metadata location, and to coordinate commits. Without a shared catalog, each engine keeps its own view: Spark writes a table through a Hive metastore, Trino reads it through a different configuration, and the two drift. Apache Polaris is an open-source catalog for Apache Iceberg that implements Iceberg's REST API, so engines that speak that protocol point at one service instead of each maintaining its own catalog.

The README names the engines it targets: Apache Doris, Apache Flink, Apache Spark, Dremio OSS, StarRocks and Trino. That list is the audience. If your stack is one of those engines plus a second one that also needs the same tables, Polaris is aimed at you. If you have a single engine, the catalog is one more process to run for a capability you already have.

The project is not a query engine and not a storage layer. It holds catalog state (namespaces, tables, credentials, access control) and hands engines the metadata they need. Data still lives in your object store or filesystem.

## How the Polaris server is put together

The repository is a Gradle multi-module build, and the module names describe the architecture better than any diagram would. Polaris-core holds entity definitions and core business logic. The API layer is generated from OpenAPI specifications and split into four modules: polaris-api-management-model and polaris-api-management-service for the Polaris Management API, polaris-api-iceberg-service for the Iceberg REST service, and polaris-api-catalog-service for the Polaris Catalog API. Two browsable specifications sit in the spec directory, and the README links to Swagger editors for both.

The runtime side is where the choices get concrete. polaris-server is a Quarkus-based server. polaris-admin is described as an admin tool, mainly for bootstrapping persistence. polaris-runtime-service is the service package, polaris-runtime-defaults holds default runtime configuration, and polaris-distribution handles packaging. Persistence is pluggable: polaris-relational-jdbc is a JDBC implementation of BasePersistence, which implies other implementations can exist behind the same interface.

Extensions cover two areas. Catalog federation has modules for Hive, Hadoop and BigQuery, so Polaris can expose catalogs that live elsewhere. External authorization has modules for OPA and Ranger. Both are optional at the module level, which matters: a default deployment does not need either.

The build requires Java 21+ and Docker 27+, and the README notes that integration tests depend on Docker, so a build without a running container runtime will fail at the test stage rather than at compile time.

## Running Polaris locally and creating your first table

The README's path to a local instance is Gradle. The server is reachable at localhost:8181, and the build expects Java 21+ and Docker 27+.

```bash
./gradlew build
./gradlew run
```

The first command builds and runs tests; the README warns that Docker must be running because integration tests depend on it. If you want to skip tests, the README gives ./gradlew assemble. If you want checks without integration tests, it gives ./gradlew check -PnoIntegrationTests.

Credentials are set with a system property. The README's example is:

```bash
./gradlew run -Dpolaris.bootstrap.credentials=POLARIS,root,secret
```

In that value, POLARIS is the realm, root is the CLIENT_ID and secret is the CLIENT_SECRET. The README states that if credentials are not set, preset credentials POLARIS,root,s3cr3t are used. That default is convenient for a local run and is exactly the kind of thing you should not carry into anything reachable from a network.

The README's first real use is Spark SQL through the regression test script, ./regtests/run_spark_sql.sh. The example commands it gives are:

```sql
create database db1;
show databases;
create table db1.table1 (id int, name string);
insert into db1.table1 values (1, 'a');
select * from db1.table1;
```

What you should see is the database listed by show databases, then the row inserted and returned by the final select. Regression tests run locally with env POLARIS_HOST=localhost ./regtests/run.sh, and the regtests README documents further options. Beyond Gradle, the repository also carries Helm charts under helm/ and a Makefile with targets such as make build-server and make build-admin that build components and a container.

## Where Polaris is the wrong tool

The clearest limitation is the engine list. Polaris implements the Iceberg REST API, so it is useful to engines that speak that protocol. An engine that reads Iceberg through a different integration does not become multi-engine compatible because Polaris exists. The README names Doris, Flink, Spark, Dremio OSS, StarRocks and Trino; it does not claim universal coverage, and you should treat anything outside that list as unverified for your case.

The second limitation is operational. Polaris is a server with a persistence backend, not a library you embed. The repository ships polaris-relational-jdbc as the JDBC implementation of BasePersistence, and polaris-admin exists mainly for bootstrapping persistence. Someone has to run that database, back it up, and decide what happens when it is unavailable. A team that picked Iceberg specifically to avoid running catalog infrastructure may find this the wrong trade.

The third is the credential model. Bootstrap credentials are supplied as a system property, and the README documents a preset default of POLARIS,root,s3cr3t when none is given. That is fine for a local run and for regression tests. It is not a secret-management story, and the README does not present it as one. If you need to know how credentials are rotated, how they are scoped per principal, or how they are revoked, the README is not the place that answers it; the site documentation and the Management API specification are. The repository does include a SECURITY-THREAT-MODEL.md and a SECURITY.md, which is a signal that the project has thought about this, but the README itself stops at the system property.

## Polaris compared with a Hive metastore as the shared catalog

The realistic alternative for many teams is a Hive metastore, which is why the federation extension for Hive is worth noting. The difference in approach is the protocol. A Hive metastore exposes the Thrift metastore API, and engines integrate with it through Hive-compatible clients. Polaris exposes the Iceberg REST API, and engines integrate with it through their REST catalog support. Both are shared catalogs; they are shared over different contracts.

That distinction has practical consequences. A REST catalog is a plain HTTP service, so clients do not need Thrift libraries or a metastore client on the classpath, and the service boundary is easier to reason about from outside. A metastore deployment is a familiar shape for teams that have run one for years, and the surrounding tooling (backup, monitoring, migration) is well trodden in a way a newer REST service is not.

Polaris also has a federation module for Hive, which suggests the intended pattern is not necessarily replacement. You can keep a Hive catalog and expose it through Polaris alongside native Polaris catalogs, which is a migration path rather than a fork in the road. The README does not document that module's behaviour in detail; the module has its own README under extensions/federation/hive.

## Maintenance, releases and what the Apache licence means here

The repository is not archived. Its most recent push was on 2026-08-02, which is the same date as the apache-polaris-1.7.0 release. Before that, 1.6.0 was released on 2026-07-09 and 1.5.0 on 2026-05-18. Three releases in roughly three months is a cadence worth noting if you plan to pin a version, because upgrade work arrives on a schedule rather than as an occasional event.

The upgrade cost is not only the server artifact. The API modules are generated from OpenAPI specifications, and the persistence layer is a pluggable interface with a JDBC implementation. A version bump can move the API surface, the persistence expectations, or both. The repository carries a CHANGELOG.md, which is the file to read before bumping rather than after.

Licensing is Apache-2.0, which is the standard permissive Apache licence. It permits commercial use and modification, and it includes a patent grant. It also requires that you preserve the licence and NOTICE files. The repository ships a NOTICE file and an aggregated-license-report module, which is where you would look at what the distribution pulls in. That is a description of the licence text, not legal advice; if your organisation has a policy on bundled dependencies, the aggregated report is the artifact to hand to whoever reviews it.

One more maintenance point: the repository points at a separate apache/polaris-tools repository for additional tooling, and the README links a third-party auto-generated code wiki while stating that the source tree remains the authoritative reference. Treat the wiki as orientation only.

## Conclusion

Adopt Polaris if you already run more than one engine against the same Iceberg tables and want one catalog process to own the metadata, rather than a per-engine catalog wired into each tool. Do not adopt it if your tables live in a single engine and nothing else needs to see them, because you take on a server, a persistence backend and a credential model for no gain. Before committing, verify three things against your own environment: that your engines are on the list the README names, that the persistence module you intend to use matches the one you can operate, and that the credential bootstrapping path (the bootstrap credentials system property or the admin tool) fits how you provision secrets. The repository is Apache-2.0 and the last push was on 2026-08-02.

## FAQ

### What is Apache Polaris?

It is an open-source catalog for Apache Iceberg that implements Iceberg's REST API, so multiple engines can share one metadata layer. The README lists Apache Doris, Apache Flink, Apache Spark, Dremio OSS, StarRocks and Trino among the supported platforms.

### How do you install Apache Polaris and run it locally?

The README builds it with Gradle using Java 21+ and Docker 27+, then runs the server with ./gradlew run, which makes it reachable at localhost:8181. Bootstrap credentials are passed with the polaris.bootstrap.credentials system property, and the README notes a preset default of POLARIS,root,s3cr3t when none is set.

### How do you use Apache Polaris with Spark?

The README points to ./regtests/run_spark_sql.sh to connect from Spark SQL, and gives example statements such as create database db1, create table db1.table1 (id int, name string) and a select against it. Regression tests run locally with env POLARIS_HOST=localhost ./regtests/run.sh.

## Sources

- [Official documentation](https://polaris.apache.org/)
- [Official README](https://github.com/apache/polaris#readme)
- [Project repository](https://github.com/apache/polaris)
- [Release notes](https://github.com/apache/polaris/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/apache-polaris
