Self-hosted service
apache/atlas avatar
apache/atlas

Apache Atlas: a metadata catalogue whose integration surface is eight hook tarballs

Apache Atlas - Open Metadata Management and Governance capabilities across the Hadoop platform and beyond

2,148 stars907 forksJavaApache-2.0

At a glance

What is it?
The governance and metadata framework for Hadoop, where the build produces one server bundle plus a separate hook bundle per ecosystem, and the list of those ecosystems is the clearest statement of what Atlas actually integrates with. Two dashboards in the tree, two review systems, and a Python package on PyPI.
Who is it for?
Adopt Atlas if your metadata estate is Hadoop, Hive, HBase, Impala, Kafka, Sqoop or the rest of that list, because the hook bundles are already built and maintained and rebuilding those integrations against the common store is the work you would otherwise do.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 1 day ago.
What is it written in?
Mainly Java, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Eight hook tarballs are the real compatibility statement

The build output is the most informative thing in this README, and it is a list of eleven archives. Two are conventional: a binary distribution and a sources archive, plus one dedicated to the server. The other eight are named for an ecosystem and end in hook: hbase, hive, impala, kafka, sqoop, storm, falcon and couchbase. That naming is the architecture in one word. Atlas is not a monolithic service that connects to everything through a single adapter layer; it ships a discrete integration artefact per system, so the systems it supports are precisely the systems on that list, and adding one is a separate build output. Read the list as a snapshot of a specific moment. Storm and Falcon are Hadoop-generation projects that most new clusters will not run. Couchbase is the one non-Hadoop entry and is a genuinely different database, which suggests someone had a use case for a document store. Kafka is the one entry that has stayed relevant. What is absent is as informative as what is present: there is no hook for a modern table format, no stream processor from the last five years, no transformation tool, and nothing for an orchestrator. If your platform is built on those, the integrations you would need are yours to write, and each one is a separate artefact in this build.

Java 17 needs four add-opens flags and Java 8 does not

Atlas builds on Java 8, Java 11 and Java 17, and the README adjusts the build per version, which is a more honest statement about a large Java codebase than a support matrix. The clone step is the ordinary one:

bash
git clone https://github.com/apache/atlas.git
cd atlas

The version-sensitive part is the Maven heap and module flags. For Java 8 and Java 11 it is two lines of memory configuration and nothing else:

bash
export MAVEN_OPTS="-Xms2g -Xmx2g"

For Java 17 it becomes five, because the build has to reopen parts of the base module:

bash
export MAVEN_OPTS="--add-opens=java.base/java.lang=ALL-UNNAMED \
--add-opens=java.base/java.lang.reflect=ALL-UNNAMED \
--add-opens=java.base/java.nio=ALL-UNNAMED \
--add-opens=java.base/java.net=ALL-UNNAMED \
-Xms2g -Xmx2g"

Those four flags are the fingerprint of code that still reaches into JDK internals by reflection, which is what you would expect from a project whose lineage goes back to the Hadoop 2 era. They are not a warning so much as an instruction: pick a Java version before you start, and if you are on 17, copy the longer block rather than assuming the short one works. The two build commands that follow are the same in both cases, a normal install followed by a distribution profile that produces the archives:

bash
mvn clean install
mvn clean package -Pdist

The 2g floor on both ends is also a hint about the memory the build needs before you get an out-of-memory failure.

dashboard and dashboardv2 both exist, so the interface is mid-migration

Two web front-end directories sit side by side at the top level: dashboard and dashboardv2, alongside webapp and rest-notification-webapp. Two dashboards with a version suffix is the shape of an interface that is being replaced rather than extended, and nobody has deleted the old one. That has direct consequences for anyone building on Atlas. If you script against the interface, a web test, or a screen you have automated, you need to know which dashboard a given build ships, and the repository will not tell you, because there are no GitHub releases to compare tags against. It also means the REST surface underneath may be stable while the pages above it are not, which is a useful thing to establish early: integration through the API is likelier to survive an interface migration than integration through the page. The notification path is a separate application, rest-notification-webapp, which fits a design where a hook fires an event and a service delivers it. The addons and tools directories suggest further extension points, and plugin-classloader is the one that deserves a note, since isolating third-party code behind a class loader is what keeps a plugin from colliding with the server's own dependencies.

Ranger enforces at runtime, Atlas only describes

The security paragraph is short and it is the most important design statement in the project. The metadata veracity is maintained by using Apache Ranger to prevent unauthorised access paths to data at runtime, and security is both role based and attribute based. So the division of labour is explicit: Atlas holds the catalogue and the lineage, and Ranger is the component that actually stops a read. That matters for how you scope a governance project, because a rollout that installs Atlas and skips Ranger delivers a searchable inventory and no enforcement. Anyone auditing access to data through Atlas metadata should follow it into Ranger, and the attribute-based half of the model is the part most deployments leave unconfigured, since role-based rules are the conventional starting point. The other two modes Atlas operates in are worth separating, because they answer different questions. The prescriptive model is what you intend the estate to look like, and the forensic model is what an investigation reconstructs from what was recorded. Lineage enriched by business taxonomical metadata is the join between the two, because raw lineage says a job wrote a table and the taxonomy says that table is customer data, which is what makes the forensic answer useful to a compliance question.

One store, so consumers do not need interfaces to each other

The architectural argument for a metadata layer is stated once and it is about integration cost rather than about any single feature. The framework is described as an extensible set of core foundational governance services for meeting compliance requirements within Hadoop, and the claim is that it enables any metadata consumer to work interoperably without discrete interfaces to each other, because the metadata store is common. That is the whole design in one sentence. The alternative model is the one most estates actually have, where a catalogue, a lineage tool, a data quality tool, an access tool and a scheduler each keep their own copy of what a table is, and each pair needs an integration. The number of interfaces in that arrangement grows quadratically, and every one of them is something to maintain and something to break. A shared store inverts that: consumers read one thing. The module list supports the claim, with graphdb as the store implementation, repository and server-api as the interface layers, common and server-common as shared code, and client for consumers. The cost of the shared-store model is that it is a critical dependency for everything downstream, which is the same property that makes it valuable when it works and expensive when it is unavailable.

Two review systems, a mandatory JIRA, and no GitHub releases

The contribution section asks for more than a pull request. Contributions are accepted as pull requests on GitHub, or through the Review Board at reviews.apache.org, and in both cases you are asked to create an Atlas JIRA issue and mention it in the pull request or review. So the process is two systems plus a ticket, and the ticket is not optional even when the code change is. That is standard foundation practice, and it is a real cost for an outside contributor who would rather open a patch and a question. Two other foundation signals in the tree are worth noting. There is a 3party-licenses directory, which is where bundled dependencies and their terms are recorded, and for a project that ships shaded clients that matters for anyone reviewing its supply chain. There is also a plugin-classloader module, and release-build.xml with a release-log.txt at the root, which together describe a formal release process run outside GitHub. That is why there are no GitHub releases, and it has a practical effect: the build output filenames carry a version placeholder, so the version you deployed is the one in the archive name after you build, and the tag list is the only version history you can browse. The last push was on 2026-09-28, so the work is current even though the release surfaces are elsewhere.

Where Atlas is the wrong shape of tool

Three boundaries, in order of how often they come up. First, platform. Everything in the project is anchored to Hadoop, and the hook list is the honest measure of that. If your data lives in a modern table format on object storage, there is no integration artefact for it here, and the lineage you get will stop at the boundary of whatever Atlas does have a hook for. Second, enforcement. If what you need is a control that blocks an unauthorised read, Atlas is the wrong project and Ranger is the right one; they are designed to be used together, and confusing them produces a governance programme with no enforcement in it. Third, interfaces. With two dashboards in the tree, the web layer is not a stable contract, so anything you automate against pages will need revisiting. The case for adoption is narrow and clear: a Hadoop estate, a compliance requirement that needs a catalogue and lineage, and a willingness to run Ranger alongside it. The store is common, so once it exists, other consumers can be added without new interfaces, and that compounding benefit is the reason to run one at all.

Editorial conclusion

Adopt Atlas if your metadata estate is Hadoop, Hive, HBase, Impala, Kafka, Sqoop or the rest of that list, because the hook bundles are already built and maintained and rebuilding those integrations against the common store is the work you would otherwise do. Do not adopt it for a modern lakehouse, since no hook is listed for a table format or a stream processor from the current era, and the governance enforcement it describes sits in Apache Ranger rather than in Atlas itself. Three things to check before committing. Confirm the hook list still matches your estate, and treat the absence of anything for your format as the deciding fact. Pick your Java version deliberately, since Java 17 needs four add-opens flags that Java 8 and 11 do not. And read the dashboard situation before writing UI automation, because the tree carries both dashboard and dashboardv2 while there are no GitHub releases to tell you which one a given tag ships.

Frequently asked questions

How do I build Apache Atlas from source?

Clone the repository, set JAVA_HOME for Java 8, 11 or 17, export the MAVEN_OPTS for that version, then run mvn clean install followed by mvn clean package -Pdist. The distribution archives are produced under distro/target.

Why does building Apache Atlas with Java 17 need extra MAVEN_OPTS?

The Java 17 setting adds four --add-opens flags for java.base packages including java.lang, java.lang.reflect, java.nio and java.net, on top of the 2g heap settings used for Java 8 and Java 11. The README gives the longer block specifically for Java 17.

Which systems does Apache Atlas have integrations for?

The build produces a separate hook archive per ecosystem, listed as hbase, hive, impala, kafka, sqoop, storm, falcon and couchbase. No hook is listed for a modern table format or a recent stream processor, so that list is the practical compatibility boundary.

Does Apache Atlas enforce access control on data?

No. The README states that metadata veracity is maintained by using Apache Ranger to prevent unauthorised access paths to data at runtime, and that security is both role based and attribute based. Atlas is the catalogue and Ranger is the enforcement point.

How do I contribute a change to Apache Atlas?

Through a pull request on GitHub or the Review Board at reviews.apache.org, and in both cases you are asked to create an Atlas JIRA issue and mention it in the change. The wiki is at cwiki.apache.org/confluence/display/ATLAS.

Official sources

  1. apache/atlas on GitHub
  2. Issues
  3. License: Apache-2.0
  4. Project website
  5. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/apache-atlas.svg)](https://hysenlabs.com/projects/apache-atlas)