Self-hosted service
datahub-project/datahub avatar
datahub-project/datahub

DataHub: the open source metadata platform behind your data catalog

The Context Platform for your Data and AI Stack

12,778 stars3,705 forksPythonApache-2.0

At a glance

What is it?
DataHub is a metadata platform for discovery, governance and observability across warehouses, BI tools and AI agents. This review covers how ingestion works, how to install the CLI, and where the project stops short.
Who is it for?
Adopt DataHub if you have a fragmented stack and can run its services plus an ingestion pipeline, or if you only need the Python CLI and a handful of connectors. Do not adopt it if you want a single binary with no backend, or if your metadata fits in a spreadsheet.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem DataHub solves: metadata scattered across dozens of tools

A modern data stack is not one system. It is a warehouse, a lake, a BI layer, an orchestrator, a feature store, and now a set of AI agents querying all of it. Each of those tools knows something about the data it holds, and none of them knows what the others know. The README describes this directly: finding the right data and understanding its lineage is, in its words, "like searching through a maze blindfolded."

DataHub's answer is to collect metadata from every one of those systems into a single graph, and then serve that graph back to humans through search and lineage views, and to machines through APIs. The repository topics list the intended scope: data-catalog, data-discovery, data-governance, data-observability, context-management, agent-platform.

The audience is not a solo analyst. It is a platform or data engineering team that already operates several of these systems and has been asked, repeatedly, where a given column came from and who is allowed to see it. If your entire data estate is one Postgres instance and a BI tool, the graph you would build is smaller than the effort of building it.

How the metadata graph is built: connectors, events and the GMS service

DataHub separates metadata collection from metadata serving. The collection side lives in metadata-ingestion, a Python package published as acryl-datahub. Connectors pull or receive metadata from source systems: the README claims 80+ production-grade connectors extracting column lineage, usage statistics, profiling results and quality metrics. That extraction is the expensive part, and it is why the README calls the push/pull ingestion framework its pioneering architectural choice.

On the serving side, the repository layout shows the split clearly. metadata-service holds the backend, and the Docker images referenced in the README include linkedin/datahub-gms, the generalized metadata service. metadata-models and entity-registry define what a metadata entity looks like and which aspects are attached to it. metadata-events carries the change stream, which is how the README's claim of metadata updates "in seconds, not hours or days" is implemented: ingestion emits events, and the graph is updated incrementally rather than rebuilt.

Around that core sit datahub-frontend and datahub-web-react for the UI, datahub-graphql-core for the GraphQL API, datahub-actions for reacting to events, and ingestion-scheduler for recurring runs. The README also points to a Model Context Protocol server, @acryldata/mcp-server-datahub, as the path for Cursor, Claude Desktop or Cline to query the catalog. That is the context-management angle: the same graph that answers a human's search also feeds an agent's prompt.

One design consequence is worth stating plainly. Because entities and aspects are defined in metadata-models and validated by metadata-models-validator, extending DataHub with a custom entity type is a schema change, not a config toggle. The repository ships metadata-models-custom precisely because teams do this, but it is a build-time exercise.

Installing DataHub: the Python CLI and a first ingestion run

The README does not walk through a server install; it links to a quickstart at docs.datahub.com/docs/quickstart and a live demo at demo.datahub.com. What the README does show directly is the Python package, acryl-datahub, whose PyPI badge appears at the top of the file. The CLI that ships in that package is the practical entry point for a first real use, because it lets you verify a connector against a source before you commit to running the full platform.

For the analytics agent, the README gives a short path in a separate repository, and this block is copied verbatim from the README:

bash
git clone https://github.com/datahub-project/analytics-agent.git
cd analytics-agent && bash quickstart.sh

That agent is Apache 2.0 and brings its own LLM, per the README. Note that it is a distinct repository from datahub-project/datahub, and its quickstart is not the same thing as installing the catalog.

The README also gives one command for connecting AI coding assistants to DataHub over the Model Context Protocol:

bash
npx -y @acryldata/mcp-server-datahub init

That is the entire setup the README shows for the MCP path. Everything else about configuring a source and a sink lives in the docs, not in this file.

Where DataHub is the wrong tool

The honest limitation is operational weight. DataHub is not a library you import; it is a set of services. The repository contains datahub-frontend, datahub-gms, datahub-upgrade, datahub-actions and an ingestion scheduler, plus a Kubernetes chart under datahub-kubernetes and a docker directory of compose files. Running the full platform means running a message backbone for metadata-events and storage behind metadata-service. The README advertises a free cloud trial and a live demo, which is a reasonable signal that self-hosting the whole thing is not a five-minute task.

A second boundary is connector coverage. "80+ connectors" is a count, not a guarantee that yours is among them. If your source system is niche or internal, you will be writing a connector against the metadata-ingestion framework, and that work sits in Python alongside a Java-based backend. Teams without both skill sets should expect friction.

A third is the mismatch between a catalog and a data quality tool. DataHub records profiling and quality metrics as metadata, and the topics include data-observability. It is not a test runner. If what you want is assertions that fail a pipeline when a table goes stale, DataHub will store the result of that assertion, not run it for you.

Finally, the README itself spends a paragraph warning that this project is not datahub.io, the separate public dataset hosting service, and that the old datahubproject.io domain now redirects to datahub.com. That confusion is common enough to be documented in the project's own file, and it means any evaluation should start by confirming you are reading the right project's docs.

DataHub compared with OpenMetadata

The comparison people search for is DataHub versus OpenMetadata, and the difference that matters is architectural lineage rather than a feature checklist. DataHub grew out of LinkedIn's internal metadata work, and the repository still reflects that: a Java metadata service, a Python ingestion framework, a GraphQL API, and an entity model defined in metadata-models with its own registry and validator. The README describes the ingestion framework as having been "widely adopted by other catalogs," which is a claim about influence, not about parity.

OpenMetadata takes a different route. It is built around a single consolidated backend and a schema-first API, which tends to mean fewer moving services to operate and a more uniform extension story. DataHub's split between a Python collection layer and a Java serving layer is more flexible at the edges, and more work in the middle: you can write a connector in Python without touching the backend, but you cannot add an entity type without touching metadata-models.

The practical test is not which one has more connectors listed. It is which one your team can extend on a bad day. If your engineers are stronger in Python and want to add a source, DataHub's ingestion framework is the gentler path. If you want one process to reason about and a schema you edit in one place, the consolidated design is easier to hold in your head. Neither README settles this, and a proof of concept against your two ugliest sources will.

Maintenance cost, release cadence and the Apache-2.0 licence

The repository is not archived, and its last push was on 2026-09-20. Recent releases are all release candidates: v1.8.0rc1 on 2026-09-04, v1.8.0rc2 and v1.8.0rc3 on 2026-09-07. That pattern tells you the project ships frequently and that the newest line is still stabilising. Pinning to a stable release rather than an rc is the lower-risk default for a production catalog, and the presence of datahub-upgrade as a top-level module confirms that upgrades are a first-class operation with their own tooling rather than a redeploy.

Upgrade cost scales with how far you have customised. Stock deployments follow the documented path. Deployments that added custom entity types through metadata-models-custom, or wrote their own connectors in metadata-ingestion, carry those changes forward across releases and should expect to re-validate them. The e2e-test and perf-test directories exist because the maintainers run both, but they do not run your extensions.

The licence is Apache-2.0, stated in the README badge and in the LICENSE file. That is a permissive licence with a patent grant, and it is compatible with commercial use and with distributing modified versions, provided you keep the notices. It is not a copyleft licence, so it does not force you to publish your connectors. Read the NOTICE file alongside LICENSE, since that is where attribution requirements for bundled work typically live. This is a description of the licence, not legal advice; your counsel should review your specific distribution.

Editorial conclusion

Adopt DataHub if you have a fragmented stack and can run its services plus an ingestion pipeline, or if you only need the Python CLI and a handful of connectors. Do not adopt it if you want a single binary with no backend, or if your metadata fits in a spreadsheet. Before committing, verify which connector covers your sources, whether your deployment path is the Docker quickstart or the Kubernetes chart, and how you will handle schema evolution in metadata-models.

Frequently asked questions

What is DataHub used for?

DataHub is an open source metadata platform for discovery, governance and observability across a data ecosystem. It collects metadata from warehouses, BI platforms, ML systems and other tools into a unified graph, then serves it to people through search and lineage views and to AI agents through APIs.

How do I install the DataHub CLI?

The Python package is published as acryl-datahub, so it installs with pip. The README shows the package badge but does not give the install command, and it links to docs.datahub.com/docs/quickstart for the server setup rather than documenting it inline.

How much is DataHub?

The project is licensed under Apache-2.0, so the software itself carries no licence fee. The README also links to a free cloud trial at datahub.com, which is a separate hosted offering from the open source repository.

Who owns DataHub?

The repository lives under the datahub-project organization on GitHub and the README says it was originally built at LinkedIn, with the footer crediting DataHub and LinkedIn engineering. The README also notes the project was previously hosted at datahubproject.io, which now redirects to datahub.com.

What is DataHub GMS?

GMS stands for generalized metadata service, the backend that serves the metadata graph. It appears in the repository as metadata-service and as the Docker image linkedin/datahub-gms, which the README references in its badge list.

Official sources

  1. datahub-project/datahub on GitHub
  2. License: Apache-2.0
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/datahub-project-datahub.svg)](https://hysenlabs.com/projects/datahub-project-datahub)