Fugue: Ray is excluded on Python 3.14 that the classifiers claim, and the docs target deletes twice
A unified interface for distributed computing. Fugue executes SQL, Python, Pandas, and Polars code on Spark, Dask and Ray without any rewrites.
At a glance
- What is it?
- Fugue is an abstraction layer that lets a Pandas function run on Spark, Dask or Ray without being rewritten, plus an SQL dialect that can invoke Python. Reading its manifest and build file closely shows one backend dependency conditioned away on a Python version the package simultaneously advertises support for, a feature that moved out of core dependencies at a version bump without a major version, a build file that declares six phony targets and defines around twenty, and a documentation target that removes the same directory twice.
- Who is it for?
- Fugue suits a data team with an existing Pandas codebase that occasionally needs to run at scale on Spark, Dask or Ray, and that wants one SQL dialect over both, because the engine switch is a context manager and the same query returns a different dataframe type per backend. It suits less well a team that needs every backend on the newest Python, since the Ray extra is conditioned away on 3.14 even though the package advertises that version.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 139 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 5, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Ray is conditioned away on the Python version the classifiers claim
The Ray extra carries an environment marker that excludes one interpreter:
ray = [
"ray[data]>=2.30.0; python_version<'3.14'",
"duckdb>=0.5.0",
"pyarrow>=7.0.0",
"pandas",
]The requirement list that a contributor sees is therefore four packages wide on any interpreter below 3.14 and three packages wide on 3.14 and above, because the Ray line is skipped. Now put that next to the classifier list, which declares support for Python 3.10, 3.11, 3.12, 3.13 and 3.14, and the floor is 3.10. So a user on 3.14 who runs the documented install for the Ray backend gets a successful resolution and a working install of three of the four dependencies, and no Ray at all. Nothing in the failure is loud, because the marker is a supported mechanism rather than a mistake, and a package manager is doing exactly what it was told. The reason is visible in the marker itself and is almost certainly upstream: the Ray distribution publishes its own lower bound per interpreter, and pinning this project against a floor it cannot satisfy on the newest Python is a reasonable local decision. The cost is that the extra's name no longer guarantees its contents on that interpreter. DuckDB, by contrast, appears unconditionally in the same block, so it has the opposite property.
The SQL feature moved out of the core install at 0.9.0
The extras list explains one entry in unusually candid terms. Without the SQL extra the non-SQL part still works. Before version 0.9.0 the extra was included in the core dependencies, so nobody had to install it explicitly. From 0.9.0 onward it becomes required if you want the SQL dialect. That is a dependency relocation at a minor version bump in a zero-major package, and the warning that announces it is written with a comma in place of a decimal point, reading as 0,9.0 rather than 0.9.0. Small as the typo is, it sits inside a bolded caution, which is the one place on the page where precision matters most. The install section gives both commands plainly:
pip install fuguepip install fugue[sql]and states that the SQL extra is strongly recommended for the dialect. What a reader cannot learn from the page is how to detect the condition in an existing environment. A project that upgraded across the boundary and never used the dialect will see no change, and one that did use it will see an import failure, and neither symptom points back at this paragraph. Worth noting alongside it: the core dependency list is only three packages, two of which are the project's own and one of which is Pandas with an upper bound, so the footprint is genuinely small and the SQL dialect is the expensive part.
The build file declares six phony targets and defines about twenty
The Makefile opens with a single declaration line naming the targets it considers phony:
.PHONY: help clean dev docs package testSix names. The file then goes on to define a target for setting up a development environment with the package manager, one for initialising a Codespace, a clean target, documentation generation, a notebook server, twelve separate test targets, and two for Spark Connect. That is roughly twenty definitions against six declarations. A phony target is one make never treats as a file, so a declared one rebuilds when you ask for it, and an undeclared one does not: if a directory or file of the same name ever appears, make considers the target up to date and does nothing. Twelve of the undeclared targets are test targets, one per backend, which is exactly the set where a stale result would be most misleading, since a test target that silently skips is indistinguishable from a passing one. There is also a `dev` name in the declaration line with no visible `dev` target beside it, while the environment setup target is called `devenv`, so that pair looks like a rename that did not propagate to the top of the file.
The Codespace target curls a remote install script before it pulls the code
One recipe in that file is worth reading on its own terms. Initialising a Codespace environment does three things, and the order is the point. First it runs:
curl -fsSL https://claude.ai/install.sh | bashThen it runs a git pull, with the failure tolerated, and then it synchronises all dependencies in frozen mode. So a developer environment is created by piping a script from an external host into a shell, and that happens before a single line of the repository's own code has been fetched. The command is not hidden, it is not wrapped in a target most people would run by accident, and the trailing flag means a failing download stops the recipe rather than continuing. It is still the pattern that supply chain reviewers ask about, and it is the first thing the environment setup does. The other environment recipe in the same file takes a different approach: it synchronises dependencies in either upgrade or frozen mode depending on a make variable, freezes the resolved set to output, and installs the pre-commit hooks through the environment runner. One target reaches the network for a tool installer, the other only resolves packages.
The documentation target removes the same directory twice
Generating the API documentation starts by clearing what a previous run produced, and the recipe reads:
docs:
rm -rf docs/api
rm -rf docs/api_sql
rm -rf docs/api_spark
rm -rf docs/apiThe first line and the fourth line are the same command. So `docs/api` is deleted, then `docs/api_sql`, then `docs/api_spark`, then `docs/api` again, and the last removal has nothing left to remove. Two things follow. The duplication is a copy-paste artifact, which is harmless in effect but tells you the recipe was extended by appending rather than maintained. More usefully, the four names are the visible evidence of how many API surfaces exist: one base, one for the SQL dialect, one for the Spark integration, and a fourth that never got a line. Since the repository carries separate packages for each engine, the naming convention here implies an `api_duckdb`, an `api_dask` and an `api_ray` somewhere, either in the part of the recipe that is not shown or in a target that does not exist. A stale build directory from a renamed surface would then survive a docs build, which is the kind of thing that produces documentation for an API you have removed.
The PySpark equivalent it compares against stops mid-expression
The page justifies the library by showing what you would otherwise write, inside a collapsed block. The wrapper it shows takes an iterator of dataframes and yields the transformed result, and the runner accepts either a Spark dataframe or a Pandas one, converting by creating a Spark dataframe from a copy in the second case. It then builds a struct type from the input schema fields, and the block ends on that line, mid-expression, with the parenthesis unclosed. So the alternative the page argues against is never finished on the page. That matters because the claim being made is not that Fugue is faster, which would be a performance claim, but that the syntax is simpler, cleaner and more maintainable than the native route, which is a readability claim, and a readability claim is only credible if the other version is visible in full. The collapsed block is also the one place where a reader could judge the trade for themselves, and it is the one piece of the demonstration that does not render completely. The functional call it wraps is four lines long, and the abbreviated version is three.
Four descriptions and three different engine lists
The project describes itself four different ways across three files. The hosting summary calls it a unified interface for distributed computing that executes SQL, Python, Pandas and Polars code on Spark, Dask and Ray without rewrites. The manifest description is an abstraction layer for distributed computing, naming no languages and no engines at all. The bolded line on the page calls it a unified interface for distributed computing that lets users execute Python, Pandas and SQL on Spark, Dask and Ray with minimal rewrites. So Polars appears in the summary and not in the page's own summary, and the manifest omits the detail entirely. The engine list is unstable in a second way, visible only in the tree. Alongside the core package there are separate directories for Dask, DuckDB, Ibis, Polars, Ray, Spark, SQL, notebooks, contributed extensions and the test harness. DuckDB and Ibis are therefore first-class backends with their own packages, their own extras and their own make targets, and neither is named in any of the three summaries. The install extras list on the page ends with an empty bullet, which suggests the list continues past where the page is given.
Calling Python from SQL works by templating JSON into the query
The SQL dialect reaches Python by string substitution, and the example shows the whole mechanism:
query = """
SELECT id, value
FROM input_df
TRANSFORM USING map_letter_to_food(mapping={{mapping}}) SCHEMA *
"""
map_dict_str = json.dumps(map_dict)
# returns Pandas DataFrame
fugue_sql(query,mapping=map_dict_str)
# returns Spark DataFrame
fugue_sql(query, mapping=map_dict_str, engine="spark")Three mechanisms are visible in six lines. The dialect adds a `TRANSFORM USING` clause to a standard select, naming a Python function and binding a parameter. The double-brace syntax is a template marker, which is why a templating engine is a dependency of the SQL extra, and the bound value is a Python dictionary serialised to JSON before it reaches the query, so the Python function receives a string and parses it. And the same query text returns a Pandas dataframe or a Spark dataframe depending only on the engine argument. That last line is the actual claim of the library in a single call, and it is worth reading alongside the one addition the abstraction demands of your code: the output schema, declared as `SCHEMA *` here to mean every input column, which the page explains is required because distributed frameworks demand it.
Editorial conclusion
Fugue suits a data team with an existing Pandas codebase that occasionally needs to run at scale on Spark, Dask or Ray, and that wants one SQL dialect over both, because the engine switch is a context manager and the same query returns a different dataframe type per backend. It suits less well a team that needs every backend on the newest Python, since the Ray extra is conditioned away on 3.14 even though the package advertises that version. Before upgrading across the 0.9 boundary, check whether your code uses the SQL dialect, because its dependencies moved out of the core install at exactly that version. And if you are adopting the build file, note that its phony declaration lists a handful of targets while the file defines twenty, so a target missing from that line is not phony and will not rebuild when its inputs change.
Frequently asked questions
How do I install Fugue with SQL support?
The base install is pip install fugue, and the SQL dialect additionally needs the sql extra: pip install fugue[sql]. The page warns that before version 0.9.0 the SQL dependencies were part of the core install, and that from 0.9.0 onward the extra is required if you want the SQL dialect, so upgrading across that boundary can break existing dialect usage.
Which Python versions does Fugue support?
The manifest requires Python 3.10 or newer and classifies 3.10, 3.11, 3.12, 3.13 and 3.14. One exception sits inside the extras: the Ray backend's dependency carries an environment marker that excludes Python 3.14, so installing the Ray extra on 3.14 resolves without Ray in it.
How do I run the same Fugue code on Spark instead of Pandas?
Wrap the calls in an engine context and pass the backend name, so the same body runs on the default local engine, on Spark or on Dask. The SQL dialect works the same way: the identical query text returns a Pandas DataFrame or a Spark DataFrame depending on the engine argument alone.
Does Fugue require rewriting my Pandas function to run on Spark?
No edits are made to the original Pandas function. You pass it to transform together with an output schema, which is needed because distributed frameworks require it, and a schema of star means all input columns pass through. The same function is also called directly on Pandas dataframes.
What are the latest Fugue releases?
Version 0.9.7, followed by 0.9.6 and 0.9.5, with the manifest also reading 0.9.7. Tag naming is inconsistent, since the newest tag carries a v prefix and the two before it do not. The last commit to the default branch is dated 2026-05-19, several months after the most recent release.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/fugue-project-fugue)