# Apache Ambari: a Hadoop control plane whose metrics directory was repurposed

> RESTful APIs plus a browser interface for provisioning Hadoop clusters, retrofitted onto Prometheus-compatible metrics with a bundled VictoriaMetrics store, and an agent written in Python with its crypto library pinned to an exact version. The README is four links to a wiki.

**apache/ambari** — Apache Ambari simplifies provisioning, managing, and monitoring of Apache Hadoop clusters.

- Repository: https://github.com/apache/ambari
- Website: https://ambari.apache.org
- Stars: 2,314 · Forks: 1,742
- Language: Java
- License: Apache-2.0
- Published: 2026-09-30 · Updated: 2026-09-30 · Language: en
- Canonical page: https://hysenlabs.com/projects/apache-ambari

## Trunk, a Jenkins job, and two announcements seven days apart

The build badge points at builds.apache.org and a job named Ambari-trunk-Commit, the default branch is called trunk rather than main, and the top of the tree has a Jenkinsfile. So this is a Maven project built by Jenkins on Apache infrastructure, which is why the build history is not on GitHub. The release story needs reading carefully because the two most recent announcements are almost a month old and out of order. release-3.0.0 was announced on 2025-04-10, and release-2.7.9 was announced on 2025-04-17, a week later. That is not a mistake, it is the normal closing of a line: 2.7.9 was the last release of the 2.x series, cut after 3.0.0 existed so that operators on 2.x had a final supported destination. The state of trunk is the useful signal. The Python package declares a version of 3.1.0.0.dev0, so development has moved on to a 3.1 line that has not been announced, and the last push was on 2026-09-24. Read together, the project is being worked on daily and its release line stopped in April 2025. For an operator that split matters: you deploy 3.0.0, you do not deploy trunk, and the fixes you want are on trunk until the next announcement.

## The ambari-metrics directory no longer means what it used to

This is the single most confusing thing about the project for anyone arriving from Ambari 2.x, and the README spends two sentences on it because it needs to. The optional Ambari Metrics RPM, described in ambari-metrics/README.md, packages the pinned VictoriaMetrics release used by the default deployment. It replaces the legacy Ambari Metrics System and the Ganglia integrations. And it is explicitly not the former external apache/ambari-metrics project. So a directory named ambari-metrics now contains a package whose job is to ship a time-series database, not a metrics collection system, and the same words refer to a different thing entirely. That is a name collision with history, and the consequences are concrete. Anyone planning a migration who searches for the old project by directory name lands in the wrong place. Anyone reading a 2.x-era document about Ambari Metrics System is reading about software that no longer has a role here. And anyone who had Ganglia dashboards built against collected metrics is not getting an upgrade path, they are getting a different backend with a different query language. The clarifying sentence in the README is doing a lot of work, and if you are reading this to decide whether to upgrade, treat the metrics change as a replacement rather than an addition.

## Prometheus-compatible endpoints, with a store you can decline

The monitoring architecture is now shaped by the Prometheus ecosystem rather than by a proprietary stack. The monitoring stack is described as using Prometheus-compatible APIs, Ambari Agent exporters, and a bundled VictoriaMetrics storage provider. Read that as three layers. The agent runs on each node and exposes metrics. Those endpoints speak the Prometheus exposition format, which is why the agent is described as an exporter rather than as a collector. And the storage is a bundled provider, which is the part you can choose about, because the RPM is described as optional. So there are two supported architectures, not one. The default is self-contained: install the optional package and the deployment brings its own time-series database, with the VictoriaMetrics version pinned so the tested combination is reproducible. The alternative is to leave the RPM out and point the agent exporters at storage you already run, which is the right choice if you have a Prometheus-compatible platform and would rather not operate another stateful service inside a Hadoop cluster. The README does not describe the exporter configuration in any detail, so the decision between those two paths is made in the Confluence documentation and in whatever your platform expects, not from the repository.

## The agent is Python, and its crypto dependency is pinned exactly

The Java in this project is the server. The agent, which is the thing that runs on every host in the cluster, is Python, and that is visible in the packaging rather than hidden. The distribution name is ambari-python, it needs Python 3.9.2 or newer, and the package root is the source tree inside ambari-common, which the build script declares: 

```python
AMBARI_COMMON_PYTHON_FOLDER = "ambari-common/src/main/python"
```

The core dependencies are pinned to exact versions, cryptography at 50.0.1, distro at 1.9.0 and Jinja2 at 3.1.6, and the optional extras are split by role. The agent extra adds APScheduler, stomp.py and websocket-client, which describes what the agent does: schedule work, talk to a message broker, and hold a websocket open. The server extra adds javaproperties and PyYAML. A tooling extra pins ruff and pulls in tomli for Python below 3.11. The exact pins are the operationally significant part. A library pinned to a patch version cannot take a security fix without a release, so a cryptography advisory means the Ambari Python package is rebuilt, and the agent is redeployed across every node in the fleet. That is a slower path than a floating range, and for a fleet-wide rollout it is a real consideration. The version handling is also worth one look, because the build script converts a Maven-style SNAPSHOT suffix into a dev suffix and refuses anything it does not recognise: 

```python
if version.endswith("-SNAPSHOT"):
    version = version[: -len("-SNAPSHOT")] + ".dev0"
```

That is how a Maven version becomes a Python package version in a project where both build systems run side by side.

## Twelve ambari modules, and the SPI is the extension point

The top-level listing is the architecture diagram, because the module names describe the responsibilities without needing documentation. There is ambari-server and ambari-server-spi, and the separation is the interesting part: an SPI module is a published interface that exists so other code can plug into the server, so third-party extension goes through ambari-server-spi rather than by patching the server. Then there is ambari-agent for the per-host process, ambari-common for shared code including the Python package, ambari-admin and ambari-web for the interface, ambari-metrics for the bundled store, and ambari-views, which is the pluggable view layer for the web interface. ambari-serviceadvisor is a pre-install recommendation engine, so it can check a host and tell the installer what to adjust before anything is configured. ambari-project holds the parent POM, ambari-utility is shared helpers, and ambari-funtest is the functional test suite. Two sub-projects live in separate repositories with both GitHub and GitBox links: Ambari Log Search for log aggregation and Ambari Infra for the platform monitoring services. If you need either, you install something else, which is the normal Apache arrangement and worth knowing before you plan a deployment.

## The README is four wiki links, and the wiki is not versioned

Count the content in the README. The description, the monitoring paragraph, the metrics clarification, the two sub-projects, and then four bare URLs: a Quick Start Guide, a Technology Stack page, a How to Contribute page, and the licence page. All four point at cwiki.apache.org. So the repository contains no installation instructions, no configuration reference, no architecture description and no API documentation. For a tool whose entire purpose is provisioning clusters, that is a striking amount of omission, and it is a deliberate Apache pattern rather than neglect: the code lives in the repository and the documentation lives in the project's Confluence space. The practical consequence is a versioning mismatch. The wiki describes what is current, not what is in the tag you checked out, so a 3.0.0 deployment is documented by a wiki that may have been edited for trunk. The two documentation links worth reading first are the Quick Start Guide, which is where the install sequence lives, and the Technology Stack page, which is where the component versions are pinned for a release. The Getting Started link in full is https://cwiki.apache.org/confluence/display/AMBARI/Quick+Start+Guide. The rest of the repository is more disciplined than the README: there is an AGENTS.md at the root, a KEYS file, an .asf.yaml for the foundation metadata, NOTICE.txt and LICENSE.txt, and three separate Python dependency files for the build, the lock and the tooling.

## Where Ambari is the wrong shape for the job

The honest boundary is the deployment model. Ambari provisions, manages and monitors a Hadoop cluster whose hosts have addresses, names, SSH access and a known operating system, and its value is that it holds that inventory and applies changes consistently across it. Every one of those assumptions is an argument against it if your data platform runs somewhere else. On a container orchestrator, workload identity, scheduling and configuration are already handled by the platform, and what you actually need from a data tool is a way to submit jobs and read their results, which is a much smaller problem. So the case for Ambari is a Hadoop installation you intend to keep running on bare hosts, where the REST APIs give you something to automate and the agent gives you a consistent metrics surface. The case against is anything ephemeral, anything where the platform already owns node configuration, and anything where the release gap is a problem, because you would be standardising on 3.0.0 from April 2025 with fixes sitting unreleased on trunk. A second, smaller boundary: if all you want is metrics, the Prometheus-compatible agent exporters are the part to evaluate, and the bundled VictoriaMetrics package is the part you can decline.

## Conclusion

Adopt Ambari if you run Hadoop on hosts you manage yourself, because the provisioning model, the REST APIs and the agent-based exporters are built for a cluster of machines with names and addresses, and no amount of configuration makes that the right shape for a container orchestrator. Do not adopt it expecting the 2.x monitoring stack, because the ambari-metrics path now packages a pinned VictoriaMetrics and replaces both the Ambari Metrics System and Ganglia, which is a rename in place rather than a compatibility layer. Three things to check before you plan an upgrade. The agent is Python and its cryptography dependency is pinned to an exact version, so a security update means a package release and a rollout to every node. The last announced release is 3.0.0 from 2025-04-10 while trunk is at a 3.1.0.0.dev0 version, so fixes land before releases. And the documentation you need is on an Apache Confluence wiki that is not versioned with the code, so read it against the tag you actually deploy.

## FAQ

### What does the ambari-metrics component of Apache Ambari contain now?

The optional Ambari Metrics RPM in ambari-metrics packages the pinned VictoriaMetrics release used by the default deployment. It replaces the legacy Ambari Metrics System and the Ganglia integrations, and the README states it is not the former external apache/ambari-metrics project.

### Which monitoring stack does Apache Ambari 3 use?

Prometheus-compatible APIs, Ambari Agent exporters, and a bundled VictoriaMetrics storage provider. The VictoriaMetrics package is optional, so the agent exporters can instead feed storage you already operate.

### What is written in Python in Apache Ambari?

The agent. The Python distribution is named ambari-python, requires Python 3.9.2 or newer, and is built from ambari-common/src/main/python. Its core dependencies are pinned exactly, including cryptography and Jinja2, with separate optional extras for the agent, the server and tooling.

### How do I install Apache Ambari?

The repository does not document installation. The README links a Quick Start Guide on the Apache Confluence wiki, along with a Technology Stack page and a contribution guide, and the two sub-projects Ambari Log Search and Ambari Infra live in their own repositories.

### What is the latest Apache Ambari release?

release-3.0.0, announced on 2025-04-10, with release-2.7.9 announced a week later on 2025-04-17 as the final release of the 2.x line. Development has moved to a 3.1 line, and the Python package on trunk declares version 3.1.0.0.dev0.

## Sources

- [apache/ambari on GitHub](https://github.com/apache/ambari)
- [License: Apache-2.0](https://github.com/apache/ambari/blob/trunk/LICENSE)
- [Project website](https://ambari.apache.org)
- [README](https://github.com/apache/ambari/blob/trunk/README.md)
- [Releases](https://github.com/apache/ambari/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/apache-ambari
