Unity Catalog: the open source catalog behind tables, files, functions and models
Open, Multi-modal Catalog for Data & AI
At a glance
- What is it?
- Unity Catalog is an Apache-2.0 catalog server with an OpenAPI spec and an OSS implementation, compatible with the Hive metastore and Iceberg REST APIs. It is a sandbox project at LF AI & Data, and its APIs are still evolving.
- Who is it for?
- Adopt it if you need one catalog covering tables, volumes, functions and models, and you can run a JVM service. Do not adopt it if you need a frozen API surface or a catalog with no server to operate: the README states the APIs are evolving and should not be assumed stable.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 3 days ago.
- What is it written in?
- Mainly Java, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 28, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What problem Unity Catalog solves, and who ends up running it
Most data platforms accumulate more than one kind of asset. Tabular data sits in one place, files in another, and the functions and models produced by AI work sit somewhere else again, each with its own registry. Unity Catalog's stated goal is to put those under a single interface: the README describes it as a universal catalog for data and AI, and lists tables, files, functions and AI models as the asset types it covers. The governance pitch follows from that: one place to describe and secure all of them.
The intended user is not an analyst clicking through a UI. The README's own quickstart pairs a server process with a CLI, and the repository layout carries separate directories for the server, the API, the clients, the connectors and the UI. That is the shape of a piece of infrastructure that platform engineers deploy and then point other systems at. If you are choosing where table metadata lives for a lakehouse that several engines read, that is the decision this project is aimed at.
One caveat belongs here rather than at the end. The README states that Unity Catalog is currently a sandbox project with LF AI & Data Foundation, part of the Linux Foundation. Sandbox is an early stage in foundation terms, and the README says so plainly instead of burying it.
How the catalog is structured: server, spec, clients and connectors
The architecture visible in the repository is a client-server catalog. The server module holds the implementation, the api module holds the interface, and the clients directory holds code that talks to it. A specification lives under spec, and the README points to a published OpenAPI document for the REST API. The practical consequence is that the server is not tied to one engine: anything that can speak the API, or one of the compatible protocols, can read what is cataloged.
Compatibility is the part that matters most for adoption. The README states that Unity Catalog is compatible with Apache Hive's metastore API and with Apache Iceberg's REST catalog API. Those two interfaces are how a large amount of existing tooling already talks to a catalog, so the compatibility is what lets an engine connect without a bespoke integration. The README also lists multi-format support, naming Delta Lake, Apache Iceberg and Apache Hudi via UniForm, Apache Parquet, JSON and CSV among the formats it handles.
The client side is split by language and purpose. There is a CLI under bin, a UI under ui, and connectors under connectors, including a Spark integration module that the build instructions mention explicitly. The Dockerfile builds the server and the example project together, and the compose file wires a server container on port 8080 to a UI container on port 3000. That split, one process serving metadata and another serving a browser interface, tells you the UI is optional and the server is not.
Installing Unity Catalog and running a first query
The README gives two paths. The direct one requires a clone of the repository and JDK 17, with JAVA_HOME pointed at it, then a compile step. The Docker path is a single command, and the compose file in the repository defines both services.
docker compose upThat starts the server on port 8080 and the UI on port 3000, according to compose.yaml. The compose file mounts ./etc/conf into the container and keeps data in a named volume called unitycatalog_data, so state survives a restart.
For the non-Docker path, the README's prerequisites are a clone, JDK 17 in JAVA_HOME, and this build command:
build/sbt packageThen the server starts from the repository root:
bin/start-uc-serverLeave that terminal running and open another. The CLI is the fastest way to confirm the server is alive. Listing tables takes a catalog and schema name:
bin/uc table list --catalog unity --schema defaultThe README says you should see a few tables, with nested details truncated unless you add --output jsonPretty. Fetching one table's metadata uses the full three-part name:
bin/uc table get --full_name unity.default.numbersThe README notes the output identifies it as a Delta table. For Delta tables specifically, the CLI can print a snippet of the table contents, which the README attributes to the Delta Kernel Java project:
bin/uc table read --full_name unity.default.numbersIf you would rather query from SQL, the README walks through DuckDB 1.0. Inside the DuckDB shell you install and load two extensions, then register an endpoint as a secret. The token in the README's example is the literal string 'not-used', which is worth noticing: the default server is not asking you to authenticate.
CREATE SECRET (
TYPE UC,
TOKEN 'not-used',
ENDPOINT 'http://127.0.0.1:8080',
AWS_REGION 'us-east-2'
);
ATTACH 'unity' AS unity (TYPE UC_CATALOG);After attaching, SHOW ALL TABLES and a SELECT against unity.default.numbers should return the same data the CLI printed. The UI is a separate process: the README requires Node and Bun, then bun install and bun run start from the ui directory.
Where Unity Catalog is the wrong choice
The API stability statement is the first thing to weigh. The README says the APIs are currently evolving and should not be assumed to be stable. For a catalog that other systems depend on, that is a real constraint: an upgrade can change the surface your connectors were written against. A team that needs a frozen interface for a multi-year platform commitment is looking at the wrong project stage, not the wrong feature set.
The second limitation is operational. This is a server you run, with a JVM, a configuration directory, and a data volume. The Dockerfile shows a two-stage build that compiles with sbt and then copies artifacts and the coursier cache into a runtime image, and the comments in that file explain why the cache path has to resolve identically in both stages. That is the kind of detail that matters if you build your own image: get it wrong and the server starts with classpath entries missing. A project that embeds a catalog library in-process, or a managed service where someone else carries the JVM, avoids this entirely.
The third is scope. The README claims broad asset coverage, but the tutorial it provides is narrow: list tables, read a Delta table, query through DuckDB. Someone whose catalog needs are entirely about, say, lineage graphs or a semantic layer will find that the README does not demonstrate those paths, and the repository's roadmap file is where that has to be checked rather than assumed.
Unity Catalog compared with Hive Metastore and OpenMetadata
The Hive metastore comparison is the one the project answers directly, by being compatible with it. Hive's metastore API is the older interface that a great deal of engine tooling already speaks, and Unity Catalog implements compatibility with it so those engines can connect. The difference in approach is what sits behind the interface: the README describes Unity Catalog as multi-format and multimodal, covering tables, files, functions and models, while a Hive metastore is a table metadata store that engines reach through the Thrift interface. If your engines already speak Hive and your assets are all tables, the compatibility layer is the useful part and the extra asset types may be irrelevant to you.
OpenMetadata sits at a different layer. It is a metadata catalog in the governance and discovery sense: it collects metadata about systems and presents it for people to browse and reason about. Unity Catalog is positioned as the catalog those systems read at query time, with the README framing governance as a single interface over tabular data, unstructured assets and AI assets. The practical difference is who consumes it. An engine consumes Unity Catalog to resolve a table name into storage locations; a person consumes a discovery catalog to find out that the table exists. Those are not substitutes, and a platform can plausibly run both.
The same distinction applies to Purview, which is a governance service rather than a query-time catalog. The honest summary is that Unity Catalog competes with metastore-style catalogs on the read path, and with governance products only on the overlap of describing and securing assets.
Maintenance, releases and what the licence allows
The repository is not archived, and the last push was on 2026-09-20. Releases have been coming at a steady cadence: v0.5.0 in June 2026, v0.5.1 in July, and v0.6.0 in August. Note that these are 0.x versions, which is consistent with the README's own statement that the APIs are evolving. An upgrade from 0.5 to 0.6 should be treated as a compatibility question to check against the release notes, not a patch you apply without reading.
The licence is Apache-2.0, stated in the README and present as a LICENSE file at the repository root, alongside a NOTICE file. Apache-2.0 is a permissive licence with an explicit patent grant and a requirement to preserve notices, which is the usual reason a company can embed this kind of component without a legal review turning into a project. That is a description of the licence text, not legal advice; the NOTICE file and any bundled third-party dependencies are what your own review would look at.
Upgrade cost is dominated by the API surface rather than the build. The build itself is sbt, and the README documents both a full build with publishLocal and a createTarball target that produces a deployable archive in the target directory. If you deploy from the published container image, the compose file pins nothing beyond the unitycatalog/unitycatalog:latest tag, so a reproducible deployment means pinning a digest yourself.
What the Unity Catalog UI and CLI actually cover
The CLI is the broader of the two interfaces in the README's own framing. It is described as the way to create and manage catalogs, schemas and tables, and to operate on volumes and functions, with a pointer to docs/usage/cli.md for the full surface. The quickstart only demonstrates the read side, listing and getting and reading tables, so the create and manage operations are documented elsewhere rather than in the main README.
The UI is deliberately thin in the instructions. The README shows a screenshot and gives prerequisites of Node and Bun, then two commands from the ui directory: bun install and bun run start. It assumes the server is already running, and in the compose file the UI service declares a dependency on the server service. Nothing in the README describes authentication for either interface, and the DuckDB example uses a placeholder token, which suggests the default configuration does not enforce one. That is a deployment question to settle before exposing the server beyond localhost.
The example project under examples/cli is described as a demonstration of the SDK for various assets and as a way to explore any Unity Catalog server implementation. That last phrase is the useful part: the CLI is not tied to this particular server, so it can be pointed at another implementation of the same API.
Editorial conclusion
Adopt it if you need one catalog covering tables, volumes, functions and models, and you can run a JVM service. Do not adopt it if you need a frozen API surface or a catalog with no server to operate: the README states the APIs are evolving and should not be assumed stable. Before committing, verify that your query engine can speak to it through the Hive metastore or Iceberg REST interface, and check the roadmap for the governance features you actually depend on.
Frequently asked questions
What is the difference between the Hive Metastore and Unity Catalog?
Unity Catalog is compatible with Apache Hive's metastore API, so engines that speak Hive can read from it. The difference is scope: the README describes Unity Catalog as multimodal, covering tables, files, functions and AI models, and as supporting Delta Lake, Iceberg and Hudi via UniForm along with Parquet, JSON and CSV.
Is Unity Catalog a Databricks feature?
The project covered here is the open source implementation, released under Apache-2.0 and described in the README as a sandbox project with LF AI & Data Foundation, part of the Linux Foundation. It is a separate repository with its own server, CLI and OpenAPI specification.
How do you install Unity Catalog?
The README gives two routes. You can clone the repository, point JAVA_HOME at JDK 17 and run build/sbt package, or you can run docker compose up, which starts the server on port 8080 and the UI on port 3000.
What are the benefits of Unity Catalog?
The README lists a multimodal interface that covers any format, engine and asset, an open source API and implementation under Apache-2.0, and unified governance for tabular data, unstructured assets and AI assets through a single interface.
What is Unity Catalog used for?
It is a catalog that holds metadata for tables, files, functions and AI models, and the README frames it as a single interface for governing and securing those assets. The quickstart shows it being used to list tables, read a Delta table through the CLI, and query the same data from DuckDB.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/unitycatalog-unitycatalog)