Sail: a Rust-native Spark Connect server reviewed for PySpark teams
Drop-in Apache Spark replacement written in Rust, unifying batch processing, stream processing, and compute-intensive AI workloads.
At a glance
- What is it?
- Sail from LakeHQ replaces the JVM Spark engine with a Rust server that speaks the Spark Connect protocol, so existing PySpark code connects unchanged. The trade-off is behavioural parity, not API surface.
- Who is it for?
- Sail is worth adopting for teams whose PySpark workloads are SQL and DataFrame heavy, who already tolerate Spark Connect as the client boundary, and who want to remove JVM tuning from the operational picture. It is the wrong choice if your pipeline depends on Spark SQL strings, third-party JVM libraries, or exact behavioural parity that the compatibility script explicitly does not check.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository received new commits within the last day.
- What is it written in?
- Mainly Rust, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem Sail targets: Spark's JVM foundation
Spark has been the default distributed processing engine for well over a decade, and the README frames its age as the problem rather than its feature set. The JVM foundation that made Spark portable across operating systems is now the source of garbage collection pauses, memory tuning work, and slow process startup. Sail's answer is to keep the client contract and replace the engine. It is a drop-in replacement for Spark SQL and the Spark DataFrame API, implemented in Rust, and it speaks the Spark Connect protocol so existing PySpark code runs once the session points at a Sail server. The intended audience is narrow and identifiable: teams with working PySpark pipelines who are paying for cluster memory they spend on JVM overhead rather than on data. The README also positions Sail beyond batch, naming stream processing and compute-intensive AI workloads as part of the same engine, but the concrete compatibility claims in the README are about Spark SQL and DataFrames.
How the Spark Connect boundary actually works
Sail does not reimplement the PySpark client library. It implements the server side of Spark Connect, so the client you already use sends the same protocol messages to a different process. That is the whole architecture in one sentence: PySpark builds a logical plan, Spark Connect ships it over gRPC, and Sail parses and executes it. The README describes a custom Rust parser built with parser combinators and procedural macros that covers Spark SQL syntax, and execution built on Apache Arrow for the columnar in-memory format and Apache DataFusion for the query engine. Two details matter for performance claims. First, workers exchange Arrow columnar data directly during shuffles, which is where joins and aggregations spend most of their time. Second, Python UDFs run inside Sail with Arrow array pointers shared zero-copy, so the README claims no serialization overhead for Python, Pandas, and Arrow UDFs. The README states Sail is roughly 10x faster than Spark and 98% cheaper on infrastructure, derived from TPC-H benchmarks it links to. Those are the project's own numbers on its own benchmark; treat them as a starting hypothesis for your workload, not a guarantee.
Installing pysail and connecting a first session
Sail ships as a Python package on PyPI named pysail, and the README pairs it with pyspark-client. Both install from pip. The Python requirement in pyproject.toml is >=3.10,<3.15, and the package declares no runtime dependencies of its own, so pip will not pull a large tree behind it.
pip install pysail
pip install "pyspark-client"Start the local server on the command line. The port flag is the one the README uses throughout, and the server is the process your PySpark client will talk to.
sail spark server --port 50051There is also a Python entry point that starts the same server in-process, which is convenient for notebooks where you do not want a separate terminal.
from pysail.spark import SparkConnectServer
server = SparkConnectServer(port=50051)
server.start(background=False)Connect from PySpark with the remote builder pointed at sc://localhost:50051. The README states no changes are needed in your PySpark code beyond the connection string, so this is the only line that differs from a normal local Spark session.
from pyspark.sql import SparkSession
spark = SparkSession.builder.remote("sc://localhost:50051").getOrCreate()
spark.sql("SELECT 1 + 1").show()If that returns 2, the protocol path works. For cluster deployment the README points to a Kubernetes guide and shows applying a manifest and port-forwarding the service, but the manifest itself is not reproduced in the README, so read the deployment guide before assuming a working YAML exists in the repository.
Where the compatibility check stops short
Sail ships an experimental script that scans a PySpark codebase and reports the support status of the functions it finds. The README is unusually direct about its limits, and those limits are the most useful thing in the document. The script checks whether referenced PySpark functions are implemented in Sail. It does not verify behavioural parity. It looks at functions used in DataFrame operations and does not cover Spark SQL strings at all. So a clean report means your function names exist, not that your results will match. For a team whose pipelines are mostly DataFrame chains, that is a reasonable first pass. For a team with a large body of embedded SQL, the script is close to silent, and the only honest verification is running the queries on both engines and comparing output. The README states Sail is designed to be compatible with Spark 3.5.x, Spark 4.x, and later versions, which is a forward-looking statement rather than a tested matrix in the repository documentation.
The JVM dependency you cannot drop
The most consequential limitation is structural. Sail is a Spark replacement that runs on the Spark Connect protocol, which means the client side remains Spark's. If your code depends on JVM libraries loaded through Spark's classpath, or on a Spark feature that lives in the driver rather than the protocol, Sail cannot serve it regardless of how many functions the compatibility script marks as implemented. The same applies to anything outside Spark SQL and the DataFrame API. The README lists Python UDF, UDAF, UDWF, and UDTF support with Python, Pandas, and Arrow variants following Spark conventions, which covers a lot of custom logic, but it does not claim a general escape hatch for arbitrary JVM code. A second limitation is maturity signalling. The package metadata in pyproject.toml declares Development Status 4 - Beta, and the compatibility script is labelled experimental in the README. That is a project telling you it expects rough edges. Teams with strict change-control or audit requirements should weigh that against the operational savings.
Sail against DataFusion and against plain Spark
The obvious alternative is staying on Apache Spark. The difference is not features but the runtime you operate. Spark gives you a JVM, a mature ecosystem of connectors and libraries, and a decade of behavioural documentation. Sail gives you a Rust binary with no garbage collector, faster startup, and a smaller memory footprint, at the cost of a compatibility surface that the project itself describes as partial. Choosing between them is really choosing which risk you prefer: JVM tuning and resource cost, or unverified behavioural parity. A second alternative is using Apache DataFusion directly, since Sail is built on it. DataFusion is a query engine library, not a Spark-compatible server, so adopting it means rewriting pipelines against its own API rather than pointing SparkSession at a new host. Sail exists precisely to avoid that rewrite. If your team is comfortable writing Rust or a Python binding against a query engine, DataFusion removes a layer. If you have working PySpark and no appetite for a port, Sail's Connect compatibility is the entire value proposition.
Licence, releases and the cost of upgrading
Sail is Apache-2.0, and the licence identifier appears in both Cargo.toml and the Python package classifiers, so the Rust workspace and the Python distribution carry the same terms. That is a permissive licence with the usual patent grant, and it is the same licence Spark uses, which removes licence compatibility as a migration question. This is not legal advice; review the LICENSE file for the terms that bind you. On maintenance, the last push to main was on 2026-09-09, and the most recent release is v0.7.1 from 2026-08-24, following v0.7.0 on 2026-08-03 and v0.6.6 on 2026-07-07. That cadence means upgrades arrive often, and each one can move the compatibility boundary in either direction. The practical upgrade cost is not the pip install; it is re-running your verification against the new version, because a function that behaved one way in v0.6.6 may not behave identically in v0.7.1. Pin the version in production and treat upgrades as a tested change rather than a routine patch.
Editorial conclusion
Sail is worth adopting for teams whose PySpark workloads are SQL and DataFrame heavy, who already tolerate Spark Connect as the client boundary, and who want to remove JVM tuning from the operational picture. It is the wrong choice if your pipeline depends on Spark SQL strings, third-party JVM libraries, or exact behavioural parity that the compatibility script explicitly does not check. Before committing, run python -m pysail.examples.spark.compatibility_check against your codebase, then diff the results of your slowest queries between Sail and your current Spark cluster, because the script only reports whether functions are implemented.
Frequently asked questions
What are the downsides of Apache Spark?
The README frames the JVM foundation as Spark's main drawback: garbage collection pauses, JVM memory tuning, and slow process startup. Sail's pitch is that removing the JVM removes those costs while keeping the Spark SQL and DataFrame API.
Is Apache Spark still relevant?
Sail's README states Spark has been the default engine for distributed data processing for over 15 years and powers ETL, machine learning, and analytics across nearly every industry. The project's premise is that the engine should change, not that Spark's role disappeared.
What language is used in Apache Spark?
Sail's README describes Spark as built on a JVM foundation, which is what Sail replaces with a Rust-native engine. Sail itself is written in Rust, and its Spark SQL parser is a custom Rust parser built with parser combinators and procedural macros.
Is Spark 100x faster than MapReduce?
The repository documentation does not address MapReduce performance comparisons. It does state that Sail is roughly 10x faster than Spark based on derived TPC-H benchmarks, which is a different comparison entirely.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/lakehq-sail)