Model or dataset
zhu-xlab/GlobalBuildingAtlas avatar
zhu-xlab/GlobalBuildingAtlas

GlobalBuildingAtlas: global building geometry split three ways by license

GlobalBuildingAtlas: an open global and complete dataset of building polygons, heights and LoD1 3D models

2,241 stars219 forksPythonNOASSERTION

At a glance

What is it?
A research dataset that had to be cut in half to stay legal. Three separately licensed parts, an ODbL conflict nobody could resolve, and a set of practical steps to actually get the tiles you need.
Who is it for?
The licensing problem in this repository is not a footnote, it is the reason the dataset is shaped the way it is. OpenStreetMap footprints are ODbL, machine-derived heights from PLANET imagery are not, and combining them creates a licensing obligation the authors could not satisfy, so the data ships as three parts and you assemble the final LoD1 GeoJSON yourself with the provided script.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository last received commits 95 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 8, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Why one dataset became three

The FAQ explains the split better than any design document could. The project used building footprints from ODbL-licensed sources, naming OpenStreetMap and the Microsoft Global ML Building Footprints, alongside footprints derived from PLANET imagery. According to ODbL, derivatives must also be ODbL-licensed, which conflicts with the BY-NC data derived from their own imagery.

That is a genuine legal dead end rather than a licensing oversight. If you mix an ODbL polygon with a non-ODbL property into one derived product, the ODbL share taints the result, and the authors could not retroactively relicense PLANET-derived content. So the data is split at exactly the line where the obligations diverge.

Part I, `GBA.ODbLPolygon`, contains only building polygons derived from ODbL-licensed sources and lives on HuggingFace. Part II, `GBA.LoD1`, also on HuggingFace, contains additional building footprints from other sources plus LoD1 JSON files that link all the polygon features from both parts. The height maps, `GBA.Height`, are on mediaTUM.

The consequence for a user is that there is no single download that gives you a finished LoD1 file. The README is explicit that the final LoD1 GeoJSON files are no longer distributed as such, and that you derive them. That is a step most dataset consumers will not expect from a project that presents itself as a finished global dataset.

The license notice is the most important paragraph in the README

The notice is placed immediately after the introduction, before anything else, and it is written in four warning-flagged lines. The dataset is provided in three parts: ODbL-licensed polygons in `GBA.ODbLPolygon`, CC BY-NC 4.0 polygons and LoD1 building models in `GBA.Polygon` and `GBA.LoD1`, and CC BY-NC 4.0 height maps in `GBA.Height`.

The notice then says users may combine these datasets for analysis or downstream applications, but doing so may create license implications, that it is the responsibility of each user to ensure their use complies, and that the repository does not provide legal advice.

This is worth pausing on, because the non-commercial clause on two of the three parts has a scope question in it. CC BY-NC 4.0 restricts commercial use, and the authors explicitly say combining parts may create implications. Their statement is not that combination is prohibited, but that the obligation is yours to work out. A commercial project that wants global building geometry has to read both licenses itself before touching the data.

GitHub reports the license as NOASSERTION for the repository, which is the same thing in a different form. There is no single license to point at because there are three, and the code underneath carries a different one again: the Code License section describes it as MIT with Commons Clause, which permits use but forbids selling the code as such.

So there are four licensing regimes in one repository: ODbL for one part, CC BY-NC 4.0 for two, a Commons Clause restriction on the code, and no single repository-level grant. Anyone shipping a product built on this data should establish which of those applies before the data reaches a build pipeline.

Coordinate reference systems, and one instruction to ignore the file

The CRS question has its own FAQ entry, and the answer is unusual enough to repeat exactly. All building polygons are recorded in EPSG:3857. Some files in `GBA.ODbLPolygon` on HuggingFace may appear in EPSG:4326, and the instruction is to treat them as EPSG:3857.

That is a directive to ignore metadata rather than interpret it, which means any tool that reads a `.geojson` or `.prj` file and trusts it will silently produce wrong results for the affected files. EPSG:3857 is Web Mercator, the projection web maps use, so treating a 4326 file as 3857 will misplace geometry rather than fail loudly. The only safe approach is to assume 3857 for everything and verify against a tile index you already trust.

This is a real operational hazard and it is the kind of thing that makes a dataset dangerous to use casually. A coordinate system error in building data produces plausible-looking geometry in the wrong place, which is far harder to catch than an error that throws.

The FAQ also handles coverage and completeness questions directly. The dataset aims at global coverage so all countries, territories and cities should be included, but due to data quality limitations some areas may be absent or may not have height attributes. For availability questions the answer is to check the web viewer.

The five step procedure for getting the data you need

The README's usage section is a numbered procedure, and it is the part you should follow literally. Step one is to access a representative dataset, either from the `representative/` folder on HuggingFace or via mediaTUM. Doing this first is the advice the README implies rather than states, and it costs you nothing.

Step two is to identify which tiles overlap your region. For polygons and LoD1 models you use `lod1.geojson` to find intersecting tiles in `GBA.Polygon` or `GBA.LoD1`. For heights you use `height_zip.geojson` and `height_tif.geojson` to find intersecting tiles in `GBA.Height`. So there are index files, and the dataset is tile-based, which means you download only what you need rather than the whole planet.

Step three is downloading into specific directories, and the required paths are literal:

code
./ODbLPolygon
./Polygon
./LoD1

Step four is running the enrichment script:

bash
python produce_lod1.py

Or with explicit paths, which is what you want for anything reproducible:

bash
python produce_lod1.py \
    --odbl_root /path/to/odbl \
    --polygon_root /path/to/polygon \
    --json_root /path/to/json \
    --output_root /path/to/output

Step five is the output, written under `./LoD1_GeoJSON` unless you specified otherwise. The script reads the two GeoJSON folders and the JSON folder, merges the properties, and adds height and var fields. If you only want `GBA.Polygon` you can ignore the height and var fields entirely.

A machine learning product, with the error rate stated rather than hidden

The FAQ question about incorrect heights is answered in one sentence: this is a machine learning derived product, errors may occur, and the reader should refer to the publication for validation results. The README does not quote a number. That restraint is deliberate and worth respecting, because a dataset that published a headline accuracy figure would invite exactly the misuse a global geometry product attracts.

What the repository does give you is the pipeline, organised as six directories matching sections of the paper. `./im2bf` holds the code for building map extraction, regularization, polygonization and simplification, which the README ties to sections 4.3.2 through 4.3.4. `./im2bh` is the monocular height estimation using HTC-DC Net from section 4.4.2. `./infer_height` covers global inference and uncertainty quantification from section 4.4.3.

`./fuse_bf` handles quality-guided building polygon fusion, section 4.5.1, and `./make_lod1` the LoD1 model generation, section 4.5.2. `./make_plots` reproduces the figures in the manuscript. There is also a `.gitmodules` entry, which means at least one of these directories is a submodule rather than vendored source, so a shallow clone will not give you everything.

The presence of an `infer_height` directory devoted to uncertainty quantification is the most interesting thing in that list. It suggests the authors treated height confidence as something you can model and therefore something you should be able to query, rather than as a single number attached to each building.

The web viewer is a viewer, not an API

There is a web interface at a TUM-hosted address, and the README warns about how to use it in unusually direct language. The webviewer is intended for interactive visualization only. Please do not use the WFS service for feature streaming, automated querying or bulk data extraction. For data access, use the dataset release instead.

That warning is worth reading as a design fact rather than as etiquette. A WFS service is a tempting thing to build a pipeline on, because it looks like an API, and the project has decided that traffic to the visualization server is not a distribution channel. The FAQ adds that high traffic can occasionally affect the web viewer and that the server is restarted as needed to maintain access, which explains why relying on it for anything automated would be fragile even if it were permitted.

Downloading also carries a condition. The README says that by downloading the data you agree to the Terms of Use hosted alongside it and acknowledge the License Notice. Both are on the same server as the viewer, so the terms are reachable before you pull anything.

The version story is thin. There is a single release, v1.0.0 from 2025-11-06, and the last push was on 2026-07-06. For a dataset that will be consumed by scripts, the absence of a changelog and a version history means you should pin your download by date and record the terms version you accepted.

Editorial conclusion

The licensing problem in this repository is not a footnote, it is the reason the dataset is shaped the way it is. OpenStreetMap footprints are ODbL, machine-derived heights from PLANET imagery are not, and combining them creates a licensing obligation the authors could not satisfy, so the data ships as three parts and you assemble the final LoD1 GeoJSON yourself with the provided script. That script is the actual entry point of this project. Run it on the representative sample first, treat the CRS as EPSG:3857 even where files claim otherwise, and read the validation results in the publication before you rely on a height value for anything that matters.

Frequently asked questions

Where can I find free 3D models of buildings?

This dataset is one source, with the caveat that it ships as three separately licensed parts rather than one finished file. GlobalBuildingAtlas provides building polygons, machine-derived heights and LoD1 models at global coverage, hosted on HuggingFace and mediaTUM. You assemble the final LoD1 GeoJSON yourself using a provided script rather than downloading a single archive.

Why is GlobalBuildingAtlas split into three separate datasets?

Because OpenStreetMap and Microsoft footprint data are ODbL-licensed, and ODbL requires derivatives to be ODbL too, which conflicts with the non-ODbL footprints and heights derived from their own imagery. Rather than taint the derived data or drop the ODbL sources, the project separated them by license and made you merge them.

What coordinate system does GlobalBuildingAtlas use?

EPSG:3857 for all building polygons. The README warns that some files in the ODbLPolygon part on HuggingFace may appear to be EPSG:4326 and instructs you to treat them as EPSG:3857 anyway, so do not trust the coordinate declaration in the files themselves.

How do I get the finished LoD1 GeoJSON files?

You derive them. Download the representative dataset, use lod1.geojson to find the tiles covering your region, download those tiles into the expected folders, then run produce_lod1.py, optionally with explicit root paths. The output is written to ./LoD1_GeoJSON or wherever you direct it.

Can I use GlobalBuildingAtlas for a commercial project?

Check before you rely on it. Two of the three parts, GBA.Polygon and GBA.LoD1, plus the height maps, are CC BY-NC 4.0, which is non-commercial, and only GBA.ODbLPolygon is under ODbL. The README states that combining them may create license implications and that determining compliance is the user's responsibility, not the repository's.

Official sources

  1. Issues
  2. README
  3. Releases
  4. zhu-xlab/GlobalBuildingAtlas on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/zhu-xlab-globalbuildingatlas.svg)](https://hysenlabs.com/projects/zhu-xlab-globalbuildingatlas)