Open-source project
geopandas/geopandas avatar
geopandas/geopandas

GeoPandas: geographic operations on pandas DataFrames, and where it stops

Python tools for geographic data

5,262 stars1,061 forksPythonBSD-3-Clause

At a glance

What is it?
GeoPandas adds GeoSeries and GeoDataFrame to pandas so vector data can be read, reprojected, joined and plotted in one process. It is a solid fit for tabular pipelines that already live in pandas, and a poor fit for raster work or datasets that no longer fit in memory.
Who is it for?
Adopt GeoPandas when your vector data already flows through pandas, when you need spatial joins, dissolves or reprojection inside a Python process, and when the dataset fits in memory. Do not adopt it for raster analysis, for streaming ingestion, or for very large datasets where a database with spatial indexes is the better home.
Can I use it commercially?
Yes. BSD-3-Clause is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 6 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What GeoPandas actually adds to pandas

The project describes itself as adding support for geographic data to pandas objects. Concretely, it implements two types, GeoSeries and GeoDataFrame, which subclass pandas.Series and pandas.DataFrame. That subclassing is the whole design bet: anything you already do with pandas indexing, grouping, joining and plotting keeps working, and geometry rides along as one more column.

The target user is someone doing vector analysis in Python who does not want to leave the DataFrame model. A typical task is taking a file of administrative boundaries, filtering rows by attribute, reprojecting, and joining points to polygons. GeoPandas handles that in a few lines because the geometry column is a first-class citizen rather than an opaque blob.

What it is not: a raster toolkit, a spatial database, or a tile server. The README is explicit that geometry operations are cartesian. There is no enforcement of like coordinates for operations, though the README says that may change. In practice that means a buffer in degrees and a buffer in metres look identical in code and differ wildly in result. The crs attribute is metadata you are responsible for keeping honest.

How the GeoDataFrame stores geometry and CRS

A GeoDataFrame is a pandas DataFrame with at least one column holding shapely geometry objects, plus a crs attribute describing the coordinate reference system for that column. When you load a file, the crs is set automatically from the file's own metadata. When you build geometry by hand, it is not set at all until you assign it.

The dependency list in pyproject.toml shows the split of labour. numpy and pandas carry the tabular layer. shapely provides the geometry objects and the operations on them. pyproj handles coordinate transformations, which is what to_crs() calls into. pyogrio is the I/O engine: read_file and the alternate constructors accept any format pyogrio recognizes. packaging is there for version checks.

Operations fall into two categories, and the README demonstrates both. Attribute-style operations such as area return plain pandas objects: g.area gives a pandas.Series of floats. Geometry-producing operations such as g.buffer(0.5) return GeoPandas objects, so the crs and the geometry dtype survive. That distinction matters when you chain calls, because a returned pandas.Series has lost the spatial context and cannot be plotted as a map directly.

Installing GeoPandas with conda or pip

The README points to the installation docs for full detail and states the runtime dependencies: pandas, shapely, pyogrio, pyproj and packaging. matplotlib is optional and only needed for plotting. The README also warns that those packages depend on several low-level libraries for geospatial analysis, which can be a challenge to install, and therefore recommends conda.

The README's install section gives the package name and the dependency list rather than a literal command, so the conda route is the one the project directs you to. Note the version floor in pyproject.toml: requires-python is >=3.11, and the dependencies pin numpy >= 2, pandas >= 2.2.0, pyproj >= 3.7.0, shapely >= 2.1.0 and pyogrio >= 0.8. Those floors are what to check against your environment before you start, because the compiled dependencies are the part that fails.

On platforms where wheels for those compiled dependencies are unavailable, pip will try to build from source and fail. That is the failure the README is warning about, and it is why conda is the documented default. The installation docs at geopandas.readthedocs.io carry the exact commands for each platform.

Reading a shapefile and running a spatial join

The README's own example loads the New York City borough boundaries through the geodatasets package, which fetches the file and returns a path. read_file then parses it into a GeoDataFrame with a geometry column and the crs taken from the shapefile.

python
import geopandas
import geodatasets

nybb_path = geodatasets.get_path('nybb')
boros = geopandas.read_file(nybb_path)
boros = boros.set_index('BoroCode').sort_index()

The result is a GeoDataFrame with BoroName, Shape_Leng, Shape_Area and geometry columns. You should see five rows, one per borough, and the geometry column holding MULTIPOLYGON values.

From there, geometry operations chain like pandas methods. The README shows `boros['geometry'].convex_hull` returning a GeoSeries of POLYGON values, and `g.plot()` producing a matplotlib figure. Reprojection uses to_crs(), and overlay and sjoin are the two operations for combining layers: sjoin attaches attributes from one layer to another based on a spatial predicate, while overlay computes the geometric intersection or union of two layers and keeps attributes from both. Neither is shown in the README example, so check the documentation for the exact predicate and how column names are suffixed when both layers share a name.

The cartesian geometry assumption and its consequences

The single most important limitation is stated plainly in the README: GeoPandas geometry operations are cartesian, and there is currently no enforcement of like coordinates for operations. Nothing stops you from computing the area of polygons stored in EPSG:4326, and the number you get back will be in square degrees. It will be a plausible-looking float. It will be meaningless as a measurement.

The same applies to buffer distances. A 0.5-unit buffer on a geographic CRS is roughly 55 kilometres at the equator and a different distance everywhere else. The README's own buffer example uses the synthetic polygons from the introduction, which have no crs set, so the units are arbitrary by construction.

The practical rule is to reproject to a projected CRS before any operation whose result carries a unit. That is a discipline the library does not enforce, and the README says enforcement may come in the future. Until then, the burden is on the caller.

The second limitation is memory. GeoPandas is pandas underneath, so the whole dataset is materialized in the process. A shapefile with a few hundred thousand polygons is fine. A national parcel dataset with tens of millions is not, and the failure mode is a MemoryError rather than a graceful degradation. At that scale a spatial database with an index is the right tool, not a DataFrame.

Where PostGIS and DuckDB spatial differ

PostGIS is the natural comparison for anyone who has outgrown a single process. The difference is architectural rather than a matter of features. PostGIS stores geometry in a server, builds spatial indexes on disk, and pushes predicates into the query planner, so a spatial join over a table that does not fit in RAM still completes. GeoPandas loads both sides into memory and does the work in Python. For datasets that fit, GeoPandas is simpler because there is no server to run. For datasets that do not, PostGIS wins by default.

DuckDB with its spatial extension is the closer alternative for analytical work. It keeps the single-process, embedded model like GeoPandas but executes vectorized queries over data that can spill to disk, and it can read GeoParquet directly. The trade-off is the API: you write SQL, and the pandas-style indexing, groupby and plotting that GeoPandas inherits are not there. If your analysis is a sequence of DataFrame transformations and a matplotlib figure, GeoPandas is the shorter path. If it is a large aggregation over Parquet files, DuckDB is.

Note that GeoPandas itself can read and write GeoParquet through pyarrow, which is listed in the all extra. That makes the two tools complementary rather than exclusive: DuckDB for the heavy aggregation, GeoPandas for the final DataFrame-shaped step.

Maintenance, licence and upgrade cost

The repository is not archived, and the last push was on 2026-09-20. Recent releases are v1.1.4 on 2026-06-26, v1.1.3 on 2026-03-10 and v1.1.2 on 2025-12-22, so the release cadence is a few per year rather than continuous. The project states it uses an open governance model and is fiscally sponsored by NumFOCUS. pyproject.toml classifies it as Production/Stable.

The licence is BSD-3-Clause, declared both in the repository metadata and in the pyproject.toml license field. That is a permissive licence, which generally means you can use, modify and redistribute it, including in proprietary products, provided the copyright notice and licence text are retained. Nothing here is legal advice; read LICENSE.txt before you ship.

The upgrade cost is dominated by the dependency floors, not by GeoPandas itself. requires-python is >=3.11, and the pins require numpy >= 2, pandas >= 2.2.0, shapely >= 2.1.0 and pyproj >= 3.7.0. If your environment is still on numpy 1.x, installing a current GeoPandas means upgrading numpy, which can break other compiled packages in the same environment. That is the real migration risk, and it is worth checking before you pin a version. The optional extras in the all group (pyarrow, matplotlib, mapclassify, folium, SQLAlchemy, GeoAlchemy2 and others) are only needed for the specific features that use them, so you can keep the install lean.

Editorial conclusion

Adopt GeoPandas when your vector data already flows through pandas, when you need spatial joins, dissolves or reprojection inside a Python process, and when the dataset fits in memory. Do not adopt it for raster analysis, for streaming ingestion, or for very large datasets where a database with spatial indexes is the better home. Before committing, verify your pandas, shapely and pyproj versions against the dependency floor in pyproject.toml, and confirm that the conda or pip route you pick actually resolves the low-level geospatial libraries on your platform.

Frequently asked questions

What is GeoPandas used for?

It adds support for geographic data to pandas objects, implementing GeoSeries and GeoDataFrame as subclasses of pandas.Series and pandas.DataFrame. It is used for reading vector files, reprojecting with to_crs(), running geometry operations like buffer and convex_hull, and plotting with matplotlib.

How do I install GeoPandas in Python?

The README recommends conda because the underlying geospatial libraries can be hard to install, and points to the installation docs for details. The runtime dependencies are pandas, shapely, pyogrio, pyproj and packaging, with matplotlib optional for plotting. A pip install is possible but may fail where compiled wheels are unavailable.

How can I read a shapefile using GeoPandas?

Use geopandas.read_file() with the path to the file. The README example fetches the New York City borough boundaries through the geodatasets package and passes the returned path to read_file, which produces a GeoDataFrame with a geometry column and the crs set automatically from the file.

How do I install GeoPandas with pip?

The package is on PyPI, but the README warns that the compiled low-level geospatial libraries can be a challenge to install and recommends conda instead. The pyproject.toml requires Python 3.11 or newer and pins numpy >= 2, pandas >= 2.2.0, pyproj >= 3.7.0, shapely >= 2.1.0 and pyogrio >= 0.8.

What are some good Python libraries for geospatial analysis?

GeoPandas builds directly on several of them: pandas and numpy for the tabular layer, shapely for geometry objects, pyproj for coordinate transformations and pyogrio for reading and writing formats. matplotlib is the optional dependency used for plotting GeoSeries and GeoDataFrames.

Official sources

  1. geopandas/geopandas on GitHub
  2. License: BSD-3-Clause
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/geopandas-geopandas.svg)](https://hysenlabs.com/projects/geopandas-geopandas)