Open-source project
open-metadata/OpenMetadata avatar
open-metadata/OpenMetadata

OpenMetadata: the open context layer for data and AI agents

The Open Context Layer for Data and AI , OpenMetadata is the open platform for building trusted data context and business semantics for humans, AI assistants, and agents.

15,334 stars2,410 forksTypeScriptApache-2.0

At a glance

What is it?
OpenMetadata connects technical metadata, lineage, quality signals and business semantics into one graph, then exposes it over APIs, SDKs and an MCP server. The pitch is context for AI; the cost is a heavyweight service stack.
Who is it for?
Adopt OpenMetadata if you already run a warehouse plus several BI and pipeline tools and need one governed place where lineage, ownership, quality and business terms meet, and if you are prepared to operate its service stack and Postgres or MySQL backend. Do not adopt it as a lightweight column dictionary for a single database, and do not treat the MCP server as a substitute for access control at the warehouse.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 4 days ago.
What is it written in?
Mainly TypeScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What OpenMetadata solves, and for whom

A warehouse connection tells an assistant which tables exist. It does not tell the assistant that a column holds certified customer revenue, that a contract governs its freshness, or that a pipeline three hops downstream breaks if the column is renamed. OpenMetadata exists to hold that second layer. The README frames the argument bluntly: AI does not need another raw database connector, AI needs context and memory.

The people this is built for are data platform engineers, stewards and analytics leads inside organizations that already run several tools at once. A single Postgres instance does not generate enough ambiguity to justify a metadata platform. A stack of two warehouses, a BI tool, an orchestrator and a streaming layer does. That is where ownership, lineage and glossary terms stop fitting in anyone's head, and where the README's list of questions (which datasets power this dashboard, who owns this data product, which columns contain sensitive customer information) becomes a real support burden.

The project describes itself as the open context layer for data and AI, and the repository reflects that ambition in its layout: separate modules for the service, ingestion, the UI, the spec, the SDK, the MCP server and a Kubernetes operator. That breadth is the product. It is also the reason adoption is a project rather than an afternoon.

How the metadata graph is assembled

The README describes six stages: collect, normalize, connect, preserve memory, govern, activate. Read that as a pipeline rather than marketing. Connectors and ingestion APIs pull technical metadata from warehouses, lakes, BI tools, pipelines, ML platforms, messaging and storage systems. That raw material is normalized against open schemas so an asset, a lineage edge, a policy and a memory all share one representation. The normalized records are then joined into a single graph, which is the thing queries actually hit.

The schema-first choice matters more than it sounds. The repository carries an openmetadata-spec module and an openspec directory, and the README lists DCAT, DPROD, PROV-O, OpenLineage, ODCS, RDF/OWL, JSON-LD, SHACL and JSON Schema as the standards the platform works with. Because the model is declared rather than implied, the same graph can be read by the UI, by the Python SDK, by the REST API and by the MCP server without four separate translation layers. That is the real architectural bet, and it is the reason the connector count can grow without the core model fragmenting.

The memory layer is the part that departs from a conventional catalog. Conversations, AI threads, decisions, assumptions and remediation notes are stored as governed objects attached to assets, users and data products. The README is explicit that memory is part of the architecture and not a side channel. Whether that is genuinely useful or merely a place for stale notes to accumulate depends entirely on whether your teams write them down, which no architecture can enforce.

Installing OpenMetadata and running a first ingestion

The README does not carry a step-by-step install section; it points to the documentation and the project homepage at open-metadata.org. What the repository itself shows is a Docker deployment directory and a Kubernetes operator module, plus a Makefile that drives local development. The Makefile exposes a single setup target for macOS and Linux, with flags passed through an ARGS variable, and a companion target that reports on the environment without changing it.

bash
make prerequisites
make dev_setup
make dev_check

The first target runs a prerequisite check script. The second performs the one-call dev environment setup, and the third diagnoses the environment without modifying anything, which is the one to run first if a previous setup attempt failed halfway.

For the user interface, the Makefile installs dependencies with Yarn from the UI directory and then starts the development server. Note the engine constraint in package.json: Node 10 or later and Yarn 1.22 or later.

bash
make yarn_install_cache
make yarn_start_dev_ui

The install step uses a frozen lockfile, so it will fail rather than silently resolve a different dependency tree. The second command serves the UI locally from the openmetadata-ui module.

Ingestion is a separate Python module. The Makefile shows how the ingestion package is installed with an extras group, and the repository ships an examples/python-sdk directory as a starting point.

bash
python -m pip install "ingestion[e2e_test]/"

That command installs the ingestion module with end-to-end test dependencies, including Playwright, and the Makefile pairs it with a Playwright browser install. For the SDK path specifically, read examples/python-sdk rather than guessing at client construction; the README does not document the client signature.

Where OpenMetadata strains

The honest limitation is operational weight. The repository contains a Java service module, a TypeScript UI module, a Python ingestion module, a spec module, shaded dependencies, an Airflow API module and a Kubernetes operator. Running that locally means running several of them, and the Makefile's own targets split cleanly along those lines. A team that wants a searchable list of tables and column descriptions will spend more time on the platform than the platform saves them.

The second limitation is connector depth versus connector breadth. The README states 130+ connectors, and that number is a count, not a quality signal. Connectors vary in how much they extract: some surface table and column metadata, others also emit lineage, profiling or quality results. The README does not document per-connector coverage, so the only reliable way to know whether your source is fully supported is to check that connector's documentation before you plan around it.

The third is that context does not create itself. Lineage is only as complete as the tools feeding it, and the README leans on OpenLineage events for pipeline lineage. If your orchestrator does not emit them and no connector covers it, that part of the graph stays empty. Similarly, the memory layer depends on humans and agents writing things down. An empty memory store is indistinguishable from a feature that does not work.

Finally, the MCP server gives agents a governed view of metadata. It does not give them access to the underlying rows, and it is not a substitute for row-level or column-level controls at the warehouse. Treating metadata governance as data governance is a category error the README does not make, but a hurried deployment might.

OpenMetadata against DataHub and OpenLineage

The comparison people reach for is DataHub, and the difference is one of center of gravity. DataHub is built around a metadata event stream and a push-based ingestion model, where emitters publish change events into a broker and the catalog consumes them. OpenMetadata's README describes a pull-oriented collector model, with 130+ connectors and ingestion APIs doing the extraction, then normalization against open schemas before records land in the graph. Push versus pull changes what you operate: an event-stream catalog asks you to instrument producers, a connector catalog asks you to schedule and monitor ingestion jobs.

OpenLineage is not a competitor at all, and the search traffic that pairs the two is a category confusion worth clearing up. OpenLineage is a specification for lineage events, and OpenMetadata is a consumer of it. The README lists OpenLineage among the interoperability standards the platform supports, alongside DCAT, PROV-O and ODCS. If you already emit OpenLineage from your orchestrator, that is an input to OpenMetadata, not an alternative to it.

The more interesting contrast is with a hosted commercial catalog. A managed service removes the operating burden described above and typically bundles support. What it does not give you is the schema and the connector code sitting in your own repository under Apache-2.0, which is the actual reason teams pick the open path: the ability to read the ingestion code when a connector behaves unexpectedly, and to extend it without a vendor ticket.

Licence, maintenance and upgrade cost

OpenMetadata is licensed under Apache-2.0, and the repository carries both a LICENSE and a NOTICE file. The NOTICE file is the one to read before redistribution, because it records bundled third-party components and their attribution requirements. Apache-2.0 permits commercial use and modification, and it includes a patent grant, but it also requires that you preserve licence and attribution notices. That is a description of the licence text, not legal advice; if you are redistributing OpenMetadata inside a product, have counsel review the NOTICE contents.

The release history shows a fast cadence with parallel lines: 2.0.0-release on 2026-08-24, 1.13.4-release on 2026-08-21, and 2.0.0-rc2-release on 2026-08-13. The last push to the default branch was on 2026-08-24. Two maintained lines means you must decide deliberately which one you track; a patch on 1.13.x will not carry the 2.0 changes, and vice versa. The repository is not archived.

Upgrade cost sits mostly in the schema. Because the platform is schema-first, a version bump can change entity definitions, and the ingestion module, the SDK and the MCP server all read those definitions. The practical implication is that the service, the ingestion package and any client you have written should move together rather than independently. The repository's integration tests and the Airflow API module exist for exactly this reason; the README does not document a rollback procedure, so plan your own before upgrading.

Editorial conclusion

Adopt OpenMetadata if you already run a warehouse plus several BI and pipeline tools and need one governed place where lineage, ownership, quality and business terms meet, and if you are prepared to operate its service stack and Postgres or MySQL backend. Do not adopt it as a lightweight column dictionary for a single database, and do not treat the MCP server as a substitute for access control at the warehouse. Before committing, verify which of the 130+ connectors covers your specific sources, check the Apache-2.0 LICENSE and NOTICE files for bundled third-party components, and confirm your deployment path against the docker/ and openmetadata-k8s-operator/ directories in the repository.

Frequently asked questions

What does OpenMetadata do?

It collects technical metadata, quality signals, lineage, ownership, usage, policies, glossaries and conversations from across a data stack and connects them into a unified metadata knowledge graph. That graph is then exposed through semantic search, APIs, SDKs and an MCP server.

Is OpenMetadata free to use?

The repository is licensed under Apache-2.0, which permits commercial use and modification. The NOTICE file records bundled third-party components and their attribution requirements, so check it before redistributing.

Is OpenMetadata a data catalog?

It covers cataloging, but the README positions it more broadly as an open context layer that also holds semantics, a knowledge graph, memory and activation surfaces. Data cataloging is one of the capabilities it lists, not the whole of it.

How do I install OpenMetadata?

The README does not carry install steps and points to the documentation and open-metadata.org. The repository provides a docker directory, a Kubernetes operator module, and a Makefile with a dev_setup target for macOS and Linux plus a dev_check target that only reports on the environment.

How do I use OpenMetadata?

The documented path is to ingest metadata through the 130+ connectors, ingestion APIs or SDKs, then read the resulting graph through semantic search, the REST API, the Python SDK or the MCP server. The repository ships an examples/python-sdk directory as a starting point.

What is the OpenMetadata Standard?

The README describes OpenMetadata as built around open metadata standards and lists DCAT, DPROD, PROV-O, OpenLineage, ODCS, RDF/OWL, JSON-LD, SHACL and JSON Schema among them. The repository keeps the entity definitions in the openmetadata-spec module.

Official sources

  1. Official documentation
  2. Official README
  3. Project repository
  4. Release notes
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/open-metadata-openmetadata.svg)](https://hysenlabs.com/projects/open-metadata-openmetadata)