Open-source project
openaddresses/openaddresses avatar
openaddresses/openaddresses

OpenAddresses: A Global Register of Address, Parcel and Building Data Sources

A global repository of open address, building, and parcel data.

3,269 stars878 forksPythonBSD-3-Clause

At a glance

What is it?
OpenAddresses is a repository of references to open address, cadastral parcel, building footprint and street centerline sources, plus the pipeline that turns them into downloadable data. It is a source register and batch pipeline, not a geocoder you can query.
Who is it for?
Adopt OpenAddresses if you need bulk address or parcel data for a region and can work with per-source licences; the sources/ directory is the thing you actually edit. Do not adopt it if you need a live geocoding API, since the project is a source register and batch pipeline, not a query service.
Can I use it commercially?
Yes. BSD-3-Clause is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What OpenAddresses actually collects, and who needs it

The README is explicit that the repository is "a collection of references to address, cadastral parcel, building footprint and street centerline data sources." That sentence is the whole product description. The repository does not host address points itself; it holds JSON files that describe where address data lives, how to fetch it, and what licence covers it. The processed output is published elsewhere, at openaddresses.io and batch.openaddresses.io.

The audience follows from that. If you are building a geocoder, a routing graph, a delivery-zone model or a civic dataset and you need street names, house numbers and coordinates in bulk, you want a register of upstream sources you can re-download on a schedule. OpenAddresses is aimed at that job. It is not aimed at someone who wants to type an address into a box and get coordinates back. The README frames the motivation in infrastructure terms: street names, house numbers and post codes combined with geographic coordinates connect digital records to physical places.

The scope is broader than the name suggests. Parcels, building footprints and street centerlines sit alongside addresses in the same register, which means one schema and one pipeline cover four related data types.

How the sources/ directory and the batch pipeline fit together

The unit of work is a source JSON file. Each one lives under sources/, organised by country and region, and points at an upstream dataset. The README points to the San Francisco source as an example and notes that its JSON contains a link to the county's open data licence. The repository also carries a schema/ directory, and package.json depends on ajv, ajv-formats, @apidevtools/json-schema-ref-parser and glob, which is consistent with validating those source files against a JSON schema before anything is fetched.

The pipeline itself is documented in DEVELOPMENT.md, which the README describes as covering how the pipeline works end-to-end, how to test sources locally, how pull request CI works, and how the weekly batch run is structured. That weekly cadence matters for planning: the register is not a live feed. A source that breaks upstream will show up as a failure in a subsequent batch run, not instantly.

Around the pipeline sit apply-us-data.py, the us-data/ and ci/ directories, and a test/ directory driven by tape and tap-dot through npm test. The scripts/ directory holds supporting tooling. So the data flow is: a contributor writes or edits a source JSON, schema validation and tests run in CI, and the batch job fetches the upstream data and produces the downloadable output.

Installing the tooling and validating a source file

The repository is a Node project for its tooling, with Python in the pipeline scripts. package.json declares "name": "oa", "version": "5.15.0", "type": "module", and an engines field of node >=10.x.x. There are no published install instructions in the README beyond cloning the repository, so the realistic path is a clone plus npm install.

bash
git clone https://github.com/openaddresses/openaddresses.git
cd openaddresses
npm install

After that, the two scripts declared in package.json are the ones you will use most. The test script runs the tape suite and pipes it through tap-dot; the lint script runs eslint over test/**.js.

bash
npm test
npm run lint

If you are adding a source rather than fixing tooling, the README gives two routes: open an issue with a link to the data and a description of the coverage area, or create a pull request against the sources/ directory. CONTRIBUTING.md holds the detail. Before opening the pull request, run the test script, because CI validates source files against the schema in schema/ before the batch job will touch them. A source that fails schema validation will not reach the download stage.

The licence situation is per-source, and that is the real constraint

This is the part that catches people. The README states plainly that the data produced by the processing pipeline is not relicensed from the original sources, and that individual sources will have their own licences. The OpenAddresses team summarises those licences in each source JSON, but a summary is not a grant. If you redistribute a merged dataset, you are redistributing data under a patchwork of terms.

The repository's own licensing splits in two. The source JSON under sources/ is licensed under CC0 1.0 Universal, as described in sources/LICENSE. The rest of the repository is BSD 3-Clause, matching the license field in package.json and the LICENSE file at the top level. So the code and the register are permissively licensed while the data they point at is not uniformly so.

Practically, this means licence review is per source, not per project. A global extract that merges thousands of sources inherits thousands of licence questions, and the source JSON is where you answer them. The README does not describe any tooling that aggregates licences into a single report, so that aggregation is work you do yourself.

Where OpenAddresses is the wrong tool

The clearest failure mode is expecting a query endpoint. Nothing in the README describes an API for looking up a single address. The homepage is described as a place for a data download, and the batch site is described as where pipeline output lives. If your application needs sub-second forward or reverse geocoding, this register does not provide it; you would load the bulk data into your own index.

A second limitation is freshness. The README describes the batch run as weekly. Upstream sources change on their own schedules, and some disappear. A source whose URL stops resolving produces a failure in the batch output rather than a correction, and the register will keep the stale reference until someone updates the JSON.

A third is coverage asymmetry. The register is global in ambition, and the README says the project is "just getting started," but the depth of coverage varies by country and region. For a country with no contributed source, there is nothing to download no matter how complete your tooling is. Checking the sources/ tree for your target region is the first step, not the last.

How it differs from OpenStreetMap and Who's on First

OpenStreetMap is the obvious comparison, and the difference is structural. OSM is a single editable geospatial database with its own tagging conventions, where addresses are contributed into the map itself. The search data suggests people conflate the two, since one of the common questions is how to add an address to OpenStreetMap. OpenAddresses takes the opposite approach: it does not hold the data, it holds pointers to authoritative sources, often government open data portals. If you need a source that a public agency maintains and updates, that model is a better fit than a crowdsourced map. If you need one schema across the whole planet, it is worse.

Who's on First is a gazetteer of places and administrative geography rather than a register of address point sources. The two are complementary: a gazetteer tells you what a place is, an address register tells you where the houses are. GeoNames is a geographic name database, useful for place lookup, but it is not a per-source pipeline that fetches and normalises address points from municipal portals. Choosing between them comes down to whether your problem is naming places or locating addresses.

Maintenance cost and what the repository tells you about it

The last push to the default branch was on 2026-09-19, and the repository is not archived. That is recent, but it says nothing about how quickly individual sources are fixed. The maintenance burden you take on is proportional to how many sources you depend on: each one is an upstream URL that can change format, move, or vanish, and each one carries its own licence summary that may need revisiting.

On the tooling side the cost is modest. package.json pins a small dependency set (ajv, ajv-formats, json-schema-ref-parser, glob) and a dev set built around tape and eslint, with Node >=10.x.x as the floor. The repository also carries .pre-commit-config.yaml and a pre-commit CI badge, so formatting and lint checks run before code lands. Upgrading means tracking those dependencies and the schema in schema/, which is where a breaking change to the source format would surface.

If you contribute sources rather than consume them, the weekly batch cadence sets your feedback loop. A pull request that passes CI still waits for a batch run before its output appears, so plan for days rather than minutes between merge and downloadable data.

Editorial conclusion

Adopt OpenAddresses if you need bulk address or parcel data for a region and can work with per-source licences; the sources/ directory is the thing you actually edit. Do not adopt it if you need a live geocoding API, since the project is a source register and batch pipeline, not a query service. Before you build on it, open the source JSON for your country under sources/, read its licence and coverage fields, and check whether the upstream URL still resolves.

Frequently asked questions

What is OpenAddresses?

It is a global collection of references to address, cadastral parcel, building footprint and street centerline data sources, plus the pipeline that processes them. The repository holds source JSON files under sources/ rather than the address data itself; downloads are published at openaddresses.io and batch.openaddresses.io.

Is OpenAddresses a free address database?

The README describes the collection as open and free to use, and data downloads are available through openaddresses.io. The pipeline output is not relicensed from the original sources, so each source carries its own licence, summarised in its source JSON.

How do I contribute an address source to OpenAddresses?

The README gives two routes: open an issue with a link to the data and a description of the coverage area, or create a pull request against the sources/ directory. CONTRIBUTING.md holds the detail, and DEVELOPMENT.md explains how to test sources locally.

Does OpenAddresses provide an API for looking up a single address?

The README does not describe a query API. It describes a data download at openaddresses.io and pipeline output at batch.openaddresses.io, so single-address lookup would require loading the bulk data into your own index.

What licence covers the OpenAddresses source files?

The source JSON under sources/ is licensed under CC0 1.0 Universal as described in sources/LICENSE, and the rest of the repository is BSD 3-Clause. The data produced by the pipeline is not relicensed from the original sources, which keep their own licences.

Official sources

  1. Issues
  2. License: BSD-3-Clause
  3. openaddresses/openaddresses on GitHub
  4. Project website
  5. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/openaddresses-openaddresses.svg)](https://hysenlabs.com/projects/openaddresses-openaddresses)