Open-source project
docker-library/repo-info avatar
docker-library/repo-info

docker-library/repo-info: extended metadata for Official Images, generated from the registry

Extended information (especially license and layer details) about the published Official Images

607 stars358 forksPerlApache-2.0

At a glance

What is it?
A Perl script set that pulls license and layer details for the published Official Images and writes them into the repos/ tree. The README calls it a work in progress, and the repository layout shows why: it is a data pipeline, not a library.
Who is it for?
Adopt docker-library/repo-info when you need license and layer facts for Official Images at a scale where reading each image's own repository by hand is not practical, and you accept that the README describes the project as a Work In Progress. Do not adopt it expecting a supported library, a stable API, or a packaging of the data for offline use; the repository ships scripts and generated files, and nothing in the repository promises a schema contract.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly Perl, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What docker-library/repo-info actually produces

Docker Official Images are published from a set of source repositories, and each of those repositories documents its own image in its own way. The problem this project addresses is that the published image and the repository that builds it are two different artifacts: the repository holds a Dockerfile and a README, while the registry holds layers, labels and an installed package set that no source file lists completely. The README describes the repository as "a set of scripts for generating reports of extended information about published official image repositories," with license and layer details called out as the emphasis.

That framing sets the audience. This is not for someone who wants to inspect a single image once; docker inspect and a look at the upstream repository cover that. It is for people who need the same facts across many Official Images at once, in a form that can be committed, diffed and read without a registry client. The output lives in the repos/ directory at the top level of the repository, which means the data is versioned the same way the scripts are. You get history, not an API.

The remote and local halves of the pipeline

The repository splits the work into two scripts. update-remote.sh is the remote half, paired with Dockerfile.remote and .remote.pl. scan-local.sh is the local half, and it is the one the README marks as running "one job per repo" on the automated side. The distinction matters because the two halves answer different questions: what the registry publishes, and what a container built from a given base image actually contains once packages are installed.

The local half carries three package-manager Dockerfiles: Dockerfile.local-apk, Dockerfile.local-dpkg and Dockerfile.local-rpm. That trio is the clearest architectural signal in the repository. Package metadata, and therefore license metadata, is not uniform across distributions. Alpine records installed packages through apk, Debian and Ubuntu through dpkg, and RPM-based images through rpm. Rather than parse one universal manifest, the project scans each image with the tool that matches its package manager. The cost is three code paths to keep working; the benefit is that the license report reflects what the package database says rather than what a naming convention suggests.

Perl is the implementation language, and .remote.pl is the visible entry point for the remote side. Perl is an unusual choice for new tooling in 2026, but for text-heavy extraction and report generation it is not a surprising one. The practical consequence for a reader is that extending the pipeline means reading Perl, not Python or Go.

Installing docker-library/repo-info and running a first scan

There is no published package and the README gives no install command, so the only supported route is cloning the repository and running the shell entry points in place. The repository layout lists update-remote.sh, scan-local.sh and generate-readme.sh at the top level alongside the Dockerfiles they build from.

Start by cloning the repository and moving into it. Nothing here is installed system-wide; the scripts run from the checkout.

bash
git clone https://github.com/docker-library/repo-info.git
cd repo-info

The remote half refreshes what the registry publishes. The README links an automated Jenkins job named update-remote.sh, which is the same script you would call locally.

bash
./update-remote.sh

Expect this to take a while and to need network access to the registry. The results land under repos/, where the generated files are committed rather than written to a cache directory.

The local half is the package-level scan. The README notes that its automated job runs one job per repository, which is a hint about runtime: a single invocation is not a sweep of every Official Image, and you should point it at the repositories you care about rather than treating it as a full-corpus command.

bash
./scan-local.sh

After either script runs, the useful artifact is the diff under repos/. Reading that diff is the first real use of the project: it tells you which images changed license metadata or layer composition since the last commit, without you querying the registry yourself.

Where the pipeline is likely to break

The README states plainly that the project "is still a firm Work In Progress," and asks for "concrete suggestions for improvement in gathering or presentation." That is an honest description of the state of things, and it should shape how you depend on the output. There is no schema document in the repository, no versioning convention for the generated files, and no stated compatibility promise. If a downstream tool parses repos/, it is parsing a format that the maintainers have explicitly reserved the right to change.

The three local Dockerfiles also imply a coverage boundary. An image whose base distribution is not Alpine, Debian/Ubuntu or RPM-based has no obvious scanner here, and the repository does not describe a fallback. In that case the local scan either skips the image or produces an incomplete package list, and a license report built on it would understate what is installed. The remote half is not a substitute, because layer and label data from the registry does not enumerate installed packages.

A second boundary is freshness. The README points to two separate automated jobs, one for the remote script and one per repository for the local script. Nothing in the repository describes a single end-to-end schedule that keeps both halves current for every image, so the age of a given file under repos/ is a property of that file, not of the repository as a whole. Check the commit that touched the file you are reading before treating it as current.

How this differs from querying the registry or the upstream repos

The obvious alternative is to skip the project and work against the sources directly: pull the image, run the distribution's package query inside it, and read the license fields yourself. That approach has the advantage of being current by construction, since you are measuring the image that exists right now rather than a committed report. It also avoids Perl entirely and fits naturally into a CI job written in whatever language your team already uses.

The difference in approach is where the work happens. Running the query yourself means executing a container per image per run, which requires a Docker daemon, network access and time proportional to the number of images you cover. docker-library/repo-info moves that cost into a scheduled job and commits the result, so consumers read files instead of running containers. That trade favors consumers who need breadth and history over consumers who need the state of one image at this second.

The second alternative is reading each Official Image's own repository, which is where the Dockerfiles and READMEs live. Those repositories are authoritative about how an image is built, but they do not list the installed package set of the published artifact, and license text in a README is a claim rather than a measurement. The project's value is that it records the measured side. Neither source replaces the other, and a license review that reads only one of them is incomplete.

Maintenance cost, licensing and what the repository does not cover

The repository itself is licensed under Apache-2.0, and the LICENSE file sits at the top level. That covers the scripts. It does not cover the data under repos/, which describes third-party images and their installed packages; the licenses of those packages are what the reports are about, and the repository states no additional terms for the generated output. Treat the Apache-2.0 grant as applying to the tooling and check the upstream image licenses separately before redistributing anything derived from a report.

Upgrade cost is low in the usual sense and high in an unusual one. There are no releases to track, so there is nothing to pin and nothing to break on a version bump; you take the branch as it is. The unusual cost is that the generated files change as upstream images change, which means a pull can produce a large diff even when no script was touched. If you vendor this repository, budget for reviewing data diffs, not code diffs.

The repository does not describe a rollback path, a pinned snapshot of repos/, or a tagged release of the data. If you need a frozen license report for an audit, the repository gives you commit hashes and nothing more structured than that.

Editorial conclusion

Adopt docker-library/repo-info when you need license and layer facts for Official Images at a scale where reading each image's own repository by hand is not practical, and you accept that the README describes the project as a Work In Progress. Do not adopt it expecting a supported library, a stable API, or a packaging of the data for offline use; the repository ships scripts and generated files, and nothing in the repository promises a schema contract. Before relying on it, verify two things: which of the three local package scanners (apk, dpkg, rpm) covers the base images you care about, and whether the branch of repos/ you intend to read has been refreshed recently, since the README points at two separate Jenkins jobs rather than one schedule.

Frequently asked questions

What is docker-library/repo-info used for?

It generates reports of extended information about published Official Images, with license and layer details named as the emphasis in the README. The generated files are committed under repos/ in the same repository as the scripts.

How do I install docker-library/repo-info?

There is no package and the README gives no install steps, so the route is cloning the repository and running update-remote.sh or scan-local.sh from the checkout. The README links automated jobs for both scripts rather than a distribution channel.

Which package managers does the local scan support?

The repository ships Dockerfile.local-apk, Dockerfile.local-dpkg and Dockerfile.local-rpm, covering apk, dpkg and rpm. The repository does not describe a fallback for images whose base distribution uses none of those.

Is docker-library/repo-info stable enough to build on?

The README calls the project a firm Work In Progress and invites suggestions for improvement in gathering or presentation. No schema, versioning convention or compatibility promise for the generated files appears in the repository.

What licence applies to docker-library/repo-info?

The repository is Apache-2.0, with the LICENSE file at the top level. That covers the scripts; the reports describe third-party images and their packages, whose licences are separate.

Official sources

  1. Official README
  2. Project repository