Library / SDK
apache/sedona avatar
apache/sedona

Apache Sedona: spatial SQL on Spark and Flink

A cluster computing framework for processing large-scale geospatial data.

2,414 stars787 forksJavaApache-2.0

At a glance

What is it?
Apache Sedona adds spatial data types, indexes and query operators to distributed engines. It is for teams whose geometry data no longer fits comfortably in one machine, and it assumes you already run Spark or Flink.
Who is it for?
Adopt Apache Sedona if your geometry lives in object storage or a warehouse and you already operate Spark or Flink; the spatial functions are exposed as SQL and as DataFrames, so the migration is mostly about data layout and partitioning. Do not adopt it for a few million points on one laptop: SedonaDB, the subproject the README presents for single-node analytics, is the better fit there, and a plain PostGIS instance will be simpler still.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 4 days ago.
What is it written in?
Mainly Java, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 26, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What Apache Sedona actually adds to a Spark or Flink job

Spark and Flink know nothing about geometry. You can store a polygon as a string or a byte array, but every distance check, containment test and intersection becomes hand-written code inside a map function, and the planner cannot help you. Apache Sedona supplies the missing layer: spatial types, spatial predicates and constructors expressed as SQL functions and as DataFrame operators, plus spatial indexing so that a join between two large geometry tables does not degrade into comparing every row with every other row.

The intended user is not a GIS analyst working in a desktop tool. It is a data engineer who already runs a cluster and has been asked to answer questions like which taxi trips started inside which zone, or which parcels intersect a flood boundary. The README's own example follows that shape: load NYC taxi trips and taxi zone boundaries from CSV files stored on AWS S3, run a spatial SQL query to return only trips in Manhattan, then spatially join the trips DataFrame against the zones DataFrame. The output is an ordinary DataFrame that can be plotted with GeoPandas. Nothing in that pipeline requires a GIS server.

How the spatial layer is organised inside the repository

The top level of the repository is a monorepo, and the directory names tell you where the engine-specific code lives. There are separate spark, flink, snowflake, common and python trees, plus R, zeppelin, docker and examples. The common directory holds code shared across engines; the shaded directories exist because spatial libraries pull in dependencies that would otherwise collide with a user's own classpath. That shading is a deliberate trade-off: it avoids version conflicts at the cost of larger artifacts and a more involved build.

The Python package is defined in python/pyproject.toml, not in the root pyproject.toml. The root file states plainly that it is an internal meta project for documentation and development tooling and is not published. If you are looking for the installable Python distribution, the root file is the wrong place to look, and that is a real source of confusion for newcomers who read the repository from the top down.

Examples are grouped by engine under examples/: flink-sql, java-spark-sql and spark-sql. That grouping matches the supported entry points. You write spatial SQL, or you write Java against the Spark SQL API, or you write Flink SQL, and the same function names appear in each.

Installing Apache Sedona and running a first spatial query

The README points to the project documentation for installation, and the repository publishes artifacts to Maven, PyPI, conda-forge, CRAN and DockerHub. For a Python workflow the package name is apache-sedona on PyPI and conda-forge. Install it alongside the engine you intend to use:

bash
pip install apache-sedona

After installation, you register the spatial functions with your Spark session. The documentation gives the SedonaContext entry point for this; the exact call differs between the PySpark and Scala APIs, so follow the version-specific page rather than copying a snippet from an older release.

A first real query follows the README's taxi example. Load two spatial DataFrames, one of points and one of polygons, and then run a spatial SQL query that returns only the trips inside Manhattan, followed by a spatial join between the trip DataFrame and the zone DataFrame. The README shows this as a SQL query and as a DataFrame join; the function names are documented per engine, and the join predicate is what determines whether a spatial index is used.

What you should see is a DataFrame whose row count is smaller than the trip table and whose rows all satisfy the containment predicate. If the count is close to the full trip table, the predicate is probably not being pushed into a spatial index, and you should inspect the query plan before scaling the job.

Where Sedona is the wrong tool

The README is direct about the boundary. SedonaDB is described as a single-node analytical database engine with geospatial as a first-class citizen, aimed at developers who want spatial analytics without distributed system complexity. That sentence is also a description of when the cluster framework is overkill. If your dataset fits on one machine, running Spark to process it means paying for a scheduler, a shuffle and serialisation overhead to solve a problem that a local engine handles directly.

The second limitation is operational rather than algorithmic. Sedona is a library that runs inside Spark, Flink or Snowflake; it does not remove the need to size and tune those systems. A spatial join between two large tables still triggers a shuffle, and the partitioning strategy you choose determines whether that shuffle is balanced. The repository does not document an automatic partitioner that inspects your data distribution, so the responsibility sits with the person writing the job.

The third is version coupling. The repository carries separate trees for Spark, Flink and Snowflake, and the release notes for 1.9.1, 1.9.0 and 1.8.1 are the record of what changed. Sedona's spatial functions are only available on the engine versions each release supports, so an upgrade of Spark can force an upgrade of Sedona, and the release notes are the place to confirm the pairing before you plan the change.

How it compares with PostGIS

PostGIS is the obvious alternative and the difference is architectural, not a matter of feature lists. PostGIS is an extension inside a single PostgreSQL instance. Your geometry lives in that database, the spatial index is a GiST index maintained by the server, and the planner decides how to execute a spatial join using statistics it collects itself. Everything is transactional and consistent by default.

Sedona inverts that. The data lives in files or a warehouse, the compute is separate, and the index is built as part of the job rather than maintained by a server. You gain horizontal scale and the ability to process data that would never fit in one PostgreSQL instance. You lose transactional guarantees over the geometry, and you take on the job of deciding how the data is partitioned. A team that already has a well-tuned PostGIS instance and data that fits should not move; a team whose geometry has outgrown that instance has a reason to.

The repository also lists Snowflake as a target, which matters if your warehouse is already there. In that case Sedona runs where the data sits and you avoid exporting geometry to a separate cluster.

Licence, releases and what an upgrade costs

Apache Sedona is released under the Apache License 2.0, and the repository carries both LICENSE and NOTICE files, which is the standard arrangement for an Apache Software Foundation project. For most users this means the licence permits commercial use and modification; the NOTICE file must be preserved when you redistribute. This is a description of the licence text, not legal advice, and if you are redistributing a modified build you should read the licence yourself.

The release cadence visible in the repository is roughly quarterly: 1.8.1 in January 2026, 1.9.0 in April 2026 and 1.9.1 in August 2026. The last push to the default branch was on 2026-08-05, the same day as the 1.9.1 release. That is a normal pattern for a project that tags a release and then goes quiet on master for a while.

The upgrade cost is not in the library itself but in the engine underneath it. Because the spark, flink and snowflake trees are maintained separately, a Sedona version bump may also require a compatible engine version, and the shaded artifacts mean you cannot always mix a newer Sedona with an older engine. Budget for testing the spatial functions you actually call, not for a drop-in version change. The Makefile shows the project's own checks run through pre-commit via uv, which tells you the maintainers gate changes on linting and formatting, but it says nothing about the compatibility matrix you will face.

What the repository does not document

Several things a prospective adopter would want are simply absent. There is no documented rollback procedure if a Sedona upgrade breaks a job, and no published compatibility table in the README mapping each Sedona release to specific Spark and Flink versions; the release notes are the only record and they are per-release. There is no guidance on choosing a partition count for a spatial join, which is the single decision most likely to determine whether your job finishes or thrashes.

The README also does not document how spatial indexes are built or persisted across a session. Whether an index is created implicitly by the join operator or must be requested explicitly is a question the documentation answers per engine, and the answer affects both performance and memory. Check that page for your engine before you size a cluster, because an in-memory index over a large polygon table is a different resource profile from a shuffle-based join.

Editorial conclusion

Adopt Apache Sedona if your geometry lives in object storage or a warehouse and you already operate Spark or Flink; the spatial functions are exposed as SQL and as DataFrames, so the migration is mostly about data layout and partitioning. Do not adopt it for a few million points on one laptop: SedonaDB, the subproject the README presents for single-node analytics, is the better fit there, and a plain PostGIS instance will be simpler still. Before committing, verify which engine version your release line supports, check whether your geometry arrives as WKT, WKB or GeoParquet because that decides your loader, and confirm that the spatial join in your query plan uses a spatial index rather than a broadcast nested loop.

Frequently asked questions

How do I install Apache Sedona?

The README points to the project documentation for installation and lists distribution through Maven, PyPI, conda-forge, CRAN and DockerHub. For Python, the package is apache-sedona on PyPI and conda-forge, installed alongside the Spark or Flink version your release supports.

How do I use Apache Sedona?

You register the spatial functions with your engine session, then write spatial SQL or DataFrame operations against geometry columns. The README's example loads NYC taxi trips and taxi zones from CSV files on S3, filters trips to Manhattan with a spatial predicate, and spatially joins the two DataFrames.

What is Apache Sedona and who is it for?

It is a cluster computing framework for processing large-scale geospatial data, adding spatial types, predicates and indexes to Spark, Flink and Snowflake. It suits data engineers who already operate one of those engines and whose geometry has outgrown a single machine.

Does Apache Sedona work without a cluster?

The README presents SedonaDB, a subproject described as a single-node analytical database engine with geospatial as a first-class citizen, for developers who want spatial analytics without distributed system complexity. That is the path for data that fits on one machine.

What licence is Apache Sedona released under?

Apache Sedona is released under the Apache License 2.0, and the repository includes both LICENSE and NOTICE files. Redistribution of a modified build requires preserving the NOTICE file.

Official sources

  1. Official documentation
  2. Official README
  3. Project repository
  4. Release notes
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/apache-sedona.svg)](https://hysenlabs.com/projects/apache-sedona)