# Apache DevLake: a dev data platform that turns tool sprawl into queryable tables

> An Apache project in Go that pulls commits, issues and deployments out of GitHub, Jira, Jenkins and SonarQube into one schema, then serves Grafana dashboards over it. The value is the normalized schema, and so is the cost of accepting it.

**apache/devlake** — Apache DevLake is an open-source dev data platform to ingest, analyze, and visualize the fragmented data from DevOps tools, extracting insights for engineering excellence, developer experience, and community growth.

- Repository: https://github.com/apache/devlake
- Website: https://devlake.apache.org/
- Stars: 3,162 · Forks: 819
- Language: Go
- License: Apache-2.0
- Published: 2026-10-06 · Updated: 2026-10-06 · Language: en
- Canonical page: https://hysenlabs.com/projects/apache-devlake

## The normalized schema is the product, not the dashboards

Most teams trying to measure engineering productivity end up with a folder of exports and a spreadsheet that only one person understands. DevLake's bet is that the hard part is not the charts, it is agreeing on what a deployment, an issue and a pull request each mean across five different vendors. So the interesting work lives in `backend/`, which defines a canonical domain layer that every collector writes into.

That is a design decision with real consequences. Once your data lands in a common schema, a query about deployment frequency does not need to know whether the deployment came from Jenkins, Argo or a GitHub Actions run. The cost is that the schema forces a single opinion about each entity. A tracker with unusual concepts will either get bent into the canonical shape or need its own fields, and the README does not promise that every vendor quirk survives the trip. Topics on the repository include `domain-layer`, `etl` and `data-transfers`, which tells you the project thinks of itself as a pipeline first and a dashboard product second.

The repository layout backs that up: `backend/` holds the server and collectors, `config-ui/` is the separate React configuration app, `grafana/` holds the dashboard assets, `e2e/` holds end to end tests, and `tools/` holds developer utilities. Three deployable pieces, not one, which is why the build produces three images.

## Getting a first dataset in with Docker Compose or Helm

The README lists three install paths and points at step by step documentation for two of them: Docker Compose and Helm. Both end up with the same three services running, a relational database for the collected data, a configuration UI and a Grafana instance with provisioned dashboards.

The repository keeps development variants of the Compose setup at the top level, which is a decent hint about what the production deployment looks like. The three files to read first are the MySQL and PostgreSQL development Compose files and the datasource overlay.

There are separate Compose files for MySQL and PostgreSQL development, plus an `env.example` at the root that lists the environment variables the images expect. If you are evaluating DevLake before installing anything, reading `env.example` and the two Compose files tells you more about the operational surface than the README does, because it names the databases and the environment variables by name.

The third option is worth noting because it is unusual for an Apache project: `gh-devlake`, a GitHub CLI extension that deploys, configures and monitors DevLake from a terminal. The README describes it as supporting local Docker and Azure deployments. A CLI wrapper is convenient for trying the platform without hand-editing Compose files, though the documentation it points to lives in a separate repository rather than here.

## Blueprints are the configuration object, and the UI is where you build one

A Blueprint is the thing you create after installing: it names the data connections, the scope of data to pull, the transformations to apply and how often to sync. The configuration UI walks you through creating one, and then the usage sequence is short enough to state in full. Set up DevLake, create a Blueprint, track its progress, view the prebuilt dashboards once the first run completes, and customize with SQL when the panels do not fit.

That fifth step is the one that decides how much of DevLake you actually use. Every metric in the prebuilt dashboards is a query over tables you own, so you can edit the panel, write a new query, or build a dashboard from scratch. You are not locked into the shipped definitions. The cost of that freedom is that the semantic definitions are yours to maintain, because a hand-written DORA query that nobody reviews will quietly disagree with the team standard.

Release notes for v1.0.3-beta16 show the configuration surface getting real attention: explicit initial values for plugins that ship without them, and recovery of the create a new connection action in the UI. That is the kind of change that only matters if you are running the platform, which is a decent sign about where the effort is going.

## Writing SQL over collected data instead of trusting a metric definition

The Grafana instance is provisioned by DevLake rather than bolted on, and the repository ships live demos you can open in a browser: a DORA dashboard plus separate sets of dashboards for engineering leads and for open source maintainers. Those three audiences are a useful hint about what the shipped panels assume.

Grafana 13 support landed in v1.0.3-beta16, which bumped both the dashboards and their provisioning. The same release moved the bundled databases forward, pinning PostgreSQL to 18.1 and MySQL to 8.4.10. For a self-hosted analytics platform this is routine and still worth noting, since the provisioning files and the panel queries both change when Grafana changes major versions.

A concrete limit shows up in the same release: a fix for Grafana variable interpolation with string team and user IDs in the MySQL dashboards. That is the sort of bug that only appears once real deployments with real identifiers start rendering panels. It is also a reminder that DevLake supports both MySQL and PostgreSQL as backing stores, which doubles the surface you have to reason about when a panel renders empty.

## A collector ecosystem measured in plugins, not in a unified table

DevLake supports connections to what the README calls popular development tools, naming GitHub, GitLab, Jenkins, Jira and Sonarqube, with a documentation page listing every data source along with its scope and supported versions. Plugin work is an explicit contribution path: the README links to issues labelled for adding a plugin and to a plugin development manual under `backend/DevelopmentManual`.

Recent releases show the collector list growing faster than anything else. v1.0.3-beta16 added a ClickUp data source with folder-scoped boards, sprints and DORA metrics. v1.0.3-beta15 added historical SonarQube project metrics with Grafana trends and Copilot team adoption dashboards, and fixed hyphenated enterprise slugs for GitHub Copilot. v1.0.3-beta14 added Claude Console organization support through a usage report endpoint.

The fixes are more informative than the features. Several releases are dedicated to vendor API changes that broke collectors: BitBucket moving off cross-workspace APIs, CircleCI returning HTTP 500 that the collector had to tolerate, Azure DevOps returning 204 with no content on a timeline call, and SonarQube databases arriving without the expected migration indexes. Anyone planning a long deployment should assume that a collector will need attention when the upstream tool changes shape, and should check which collectors their stack actually depends on before assuming coverage.

## Building the three images yourself and what the release cadence tells you

The top level `Makefile` is the clearest statement of how the project is packaged, because it builds three separate Docker images rather than one bundle:

```bash
build-config-ui-image:
	cd config-ui; docker build -t $(IMAGE_REPO)/devlake-config-ui:$(TAG) --file ./Dockerfile .
```

The same Makefile has targets for the server image, which delegates into `backend/`, and for the Grafana dashboard image from the `grafana/` directory. `build-images` chains all three, and matching `push-*` targets push each one. The version string is assembled from the current git tag and short SHA, which means images are tied to a commit rather than to a floating `latest`.

The repository is licensed under Apache 2.0 and is not archived, with the last push on 2026-09-23. The version line is still a beta: v1.0.3-beta16 published on 2026-08-27, preceded by beta15 on 2026-07-19 and beta14 on 2026-07-10. Roughly monthly beta tags with long change lists are fine for an Apache incubation project, but they also mean you should pin a version rather than track a branch, and you should read a release before upgrading, since Grafana provisioning and database pins move between tags.

## Conclusion

DevLake makes sense for an engineering organisation that already lives in GitHub plus a tracker plus a CI server and keeps re-implementing the same joins in spreadsheet reports. It is a poor fit for a single toolchain with a handful of developers, because the setup cost is real and the payoff is thin. Verify three things before committing: how long the first blueprint takes on your own data volumes, whether the prebuilt panels cover the metrics your leads actually review, and whether the plugin you need for your tracker is already in the backend tree. Start from the Docker Compose setup on devlake.apache.org, look at the schema the collector produced, and only then decide whether to write your own SQL or add a data source plugin.

## FAQ

### Who is behind Apache DevLake?

DevLake is developed under the Apache Software Foundation, which is why the repository lives in the `apache` organization, the license is Apache 2.0, and the README opens with the standard ASF contributor license header. Governance and contribution run through the ASF process, including an issue tracker, a mailing list and a community site at devlake.apache.org.

### Is Apache DevLake free to use?

Yes. The project is licensed under Apache 2.0, which permits commercial use, modification and redistribution with notice and preservation of the license terms. There is no paid edition described in the repository; you pay for the infrastructure you run it on.

### Is Apache DevLake open source?

Yes, and the whole stack is public: the Go backend, the configuration UI and the Grafana dashboard assets are all in the repository, along with the Makefile targets that build each container image. The repository also carries an `AGENTS.md` file and contribution guides for code and plugins.

## Sources

- [apache/devlake on GitHub](https://github.com/apache/devlake)
- [License: Apache-2.0](https://github.com/apache/devlake/blob/main/LICENSE)
- [Project website](https://devlake.apache.org/)
- [README](https://github.com/apache/devlake/blob/main/README.md)
- [Releases](https://github.com/apache/devlake/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/apache-devlake
