# xorq: a git catalog of hashed dataframe builds, with two entry points to the same idea

> Xorq turns lazy dataframe expressions into content-addressed build directories that a git repository catalogs, so agents can discover and reuse pipelines instead of leaving folders of one-off scripts. Its costs are the ones git imposes, plus a packaging workaround for environment templates that a .gitignore rule nearly erased.

**xorq-labs/xorq** — Git-native executable catalog for agentic data work.

- Repository: https://github.com/xorq-labs/xorq
- Website: https://docs.xorq.dev
- Stars: 553 · Forks: 37
- Language: Python
- License: Apache-2.0
- Published: 2026-09-15 · Updated: 2026-09-15 · Language: en
- Canonical page: https://hysenlabs.com/projects/xorq-labs-xorq

## A build is four files, and the wheel name keeps a placeholder

The unit of work is a build directory whose name is a hash, and the documented shape of it is four entries:

```text
builds/fa2122f6a9e9/
├── expr.yaml
├── expr_metadata.json
├── requirements.txt
└── xorq-<version>-py3-none-any.whl
```

Two of those four files are the specification, `expr.yaml` and the matching `expr_metadata.json`, and the README calls the pair the content-addressed specification that captures the computation with its operations, schema, lineage and inputs. `requirements.txt` and the wheel are the environment half. The wheel filename is written as `xorq-<version>-py3-none-any.whl`, so the literal tree in the documentation is a template rather than a real directory listing: the version slot is unfilled and the hash is the twelve-character example from the surrounding commands.

Those commands are the whole manual path. You save the expression as `expr.py`, run

```bash
xorq uv build expr.py
```

then register the directory in the catalog:

```bash
xorq catalog add builds/fa2122f6a9e9/ -a penguins-agg
```

The catalog is a git repository of entries, and an entry is discoverable by its hash or by a human-readable alias, which is what `-a penguins-agg` sets. After that another person or another agent can discover, reproduce, run or compose the pipeline without reconstructing the original from scripts, environment setup and chat history.

## The catalog is a git repo of builds with git-annex behind it

Storage is not a service. The design table names Git as the choice for state and storage, and spells out what that means: a git repo of entries with git-annex support for large files. So there is no metastore to run, no API to authenticate against, and no control plane, which the comparison table states as the operating model being a git repo of build artifacts on disk with no service to call. The cost is that whatever git can express is the whole feature set: branches, review, diff and blame on top of entries that are directories of files.

That also decides what a lookup gives you. A hit returns an Arrow stream produced by running the entry, not a governed table you query with external code, and lineage is read off the manifest, which the table calls exact by construction rather than harvested from executed queries on a best-effort basis. Lineage that cannot drift is a direct consequence of cataloging the computation instead of the query log, and it is the strongest claim in the comparison.

Git-annex is the escape hatch for the size problem. A build that embeds a wheel and a pinned requirements file is fine in git, and an entry that also holds a model or a large table is not, so large files are pushed to annex rather than into the object database. The README does not spell out the annex commands, which means the large-file path is a documented intention rather than a documented workflow.

## Twelve engines are named and only four are embedded

Expressions are Ibis expressions, and the engine table splits the backends by category. Embedded: DataFusion, DuckDB, SQLite, pandas. Warehouses: Snowflake, Databricks, Trino, Postgres. Lakehouse: PyIceberg. Arrow Flight: GizmoSQL. `into_backend` moves data between them, and the design table credits DataFusion with in-process SQL and UDF execution while Arrow carries IPC and the network, so operators exchange Arrow RecordBatches rather than Python objects.

The portability claim is narrow in a way the table makes clear. The same expression runs on DataFusion, DuckDB, Snowflake and Databricks; the comparison row puts the alternative platforms as bound to the host platform. Only the embedded four need nothing installed beyond the wheel, which is also why DataFusion is the default compute rather than an option among equals.

The remaining eight are reachable only if you bring the connection. Nothing in the dependency list installs a Snowflake, Databricks, Trino, Postgres, Iceberg or GizmoSQL client, so testing a warehouse backend is a configuration job on top of the library rather than an install. The compose file shows how seriously the project takes that surface: it stands up a postgres built from `./docker/postgres` and tagged `ibis-postgres` on port 5432, a Hive metastore database on postgres 17.9-alpine listening on port 23456, a minio object store, and the Hive metastore itself on `starburstdata/hive:3.1.3-e.4`.

## Lineage is read off the manifest instead of harvested from query logs

The problem statement is written about coding agents specifically. Ask one to build a dashboard and the stated expectation is a folder of one-off Python scripts importing each other in non-obvious ways, an embedded JSON holding intermediate state, and a requirements.txt last regenerated two sessions ago. It may run end to end on your laptop, and reproducing it on another machine means rewriting some of it.

The pain table names four failures. Imperative, stateful artifacts leave a folder of `.py`, `.json` and `.html` files with no declarative spec to re-run in order. There is no discoverable shared index, and the README points at `~/.claude/memory/*.md` with a `MEMORY.md` of one-line notes as what team memory looks like today. There is no lineage graph, so renaming a column upstream breaks a downstream model at runtime because the dependency lived in chat history. And there is no portable environment, so a pipeline from one agent session has no path to another sandbox, your machine, or production.

The answer is the expression. Instead of a sequence of scripts that must be executed in the right order,

```python
import xorq.api as xo

penguins = xo.examples.penguins.fetch()
penguins_agg = (
    penguins
    .filter(xo._.species.notnull())
    .group_by("species")
    .agg(avg_bill_length=xo._.bill_length_mm.mean())
)
```

gives a structured representation of the computation, which is what can be hashed, versioned and diffed. The README summarises the design as giving a dataframe pipeline a reproducible identity, and the tradeoff it accepts is that you write a lazy expression instead of imperative code.

## Two starting points, and the agent one needs a marketplace round trip

With an agent, you install a Claude Code plugin from a separate repository:

```
/plugin marketplace add xorq-labs/claude-plugins
/plugin install xorq@xorq-plugins
```

The plugin then contributes four slash commands. `/xorq:init` loads CSV or Parquet files as catalog entries, `/xorq:catalog-explore` browses what is already in a catalog, `/xorq:composer` combines entries into new joined or aliased entries, and `/xorq:builder` assembles ML pipelines and semantic-layer entries. The division of labour is stated plainly: the agent does the building, you keep the catalog.

Manually, you install the library and write expressions:

```bash
pip install xorq[examples]
xorq init -t penguins
```

The example extra is what makes `xo.examples.penguins.fetch()` work, since the examples module is not part of the base install. Documentation for the CLI and the plugin lives outside the repository on docs.xorq.dev, with a separate website at www.xorq.dev and the plugin source in its own repo, xorq-labs/claude-plugins.

The two routes are not equivalent in surface area. The manual route gets you `xorq init` and then the Python API; the plugin route gets you four commands that assume an agent is sitting there to answer questions while it works.

## env_templates is force-included because .gitignore would eat it

The wheel is built by hatchling with `sources = ["python"]`, so the importable package lives under python/ in the repository and becomes xorq/ in the wheel, and the sdist narrows to `only-include = ["python"]` plus a force-include of `uv.lock`. The interesting part is a packaging workaround with a three-paragraph comment explaining it.

The `python/xorq/env_templates` tree is excluded from the normal VCS scan, and the comment says why: `git ls-files` picks those files up inside a git repository, but the force-include below also adds them, which causes duplicates in the wheel ZIP. The second comment is the real hazard: `.gitignore` contains `.env.*`, and hatchling treats that as an exclusion pattern in non-VCS contexts such as sdist rebuilds, staging directories and `uv tool run`. So the directory holding environment templates is invisible to both the VCS scan and the ignore rules at different moments, and force-include is what guarantees it lands in the wheel anyway.

That tells you something about what the feature does. Environment templates travel inside the wheel, which is what lets a built entry carry the environment needed to reproduce and execute it rather than only naming one. It also means the templates are a distribution concern with a git interaction attached, and anyone vendoring or mirroring the repository has to keep `.gitignore` from taking those files with it.

## Dependency floors exist per interpreter, and 3.14 is deliberately excluded

The dependency list is mostly ordinary: attrs, pyarrow, structlog, pandas, atpublic, parsy, packaging, and then a numpy row preceded by the longest comment in the file. Every major dependency carries a `python_version >= '3.10' and python_version < '4.0'` marker, so the constraints do not vary per interpreter for the ordinary packages.

numpy does. The comment explains the per-interpreter floors exist so `--resolution lowest-direct` gets a wheel rather than a source build, and that each floor is the first release with wheels for that Python, with one exception: numpy on 3.10, where 1.22.4 is pandas own floor because 1.21.3 was the first with cp310 wheels. The last sentence is the constraint that matters: the `>= '3.13'` rows are open upward and only safe while requires-python caps below 3.14, so raising that cap needs a new row per interpreter or lowest-direct picks a release with no wheel for it.

So the project is testing its dependency floors against every interpreter it claims, and has deliberately capped the supported range to keep that test honest. On a 3.14 install the markers exclude the marked packages rather than resolving something untested. This is the sort of bookkeeping that produces a working install on old interpreters and a clear failure on a new one.

## Conclusion

xorq fits teams where agent-written dataframe work has to survive the session that produced it, where git review of a pipeline is acceptable as review of code, and where the engines in use are the embedded ones or a connection you already control. It does not fit a group that needs a metastore with permissions and a UI, a Python newer than 3.13, or large catalog entries without reading up on git-annex first. Before adopting it, look at what the hashed build directory actually contains for one of your pipelines, decide whether git with annex is the storage you want, and check the engine backends you need against the four embedded ones and the eight that need a client you supply.

## FAQ

### What does xorq produce when it builds an expression?

A hashed build directory holding expr.yaml, expr_metadata.json, requirements.txt and a wheel named xorq-<version>-py3-none-any.whl. The manifest pair is the content-addressed specification carrying the operations, schema, lineage and inputs, and the build packages it with the environment needed to reproduce it.

### Which Python versions does xorq support?

Markers in pyproject.toml gate the major dependencies on python_version >= 3.10 and < 4.0, and a comment on the numpy rows states the rows open upward are only safe while requires-python caps below 3.14. A 3.14 install therefore excludes the marked packages rather than resolving untested floors.

### Which query backends can xorq expressions run against?

Embedded ones are DataFusion, DuckDB, SQLite and pandas. The table also lists Snowflake, Databricks, Trino and Postgres, PyIceberg for lakehouse work, and GizmoSQL over Arrow Flight, with into_backend moving data between them. The README states plainly that it is complementary to Atlan, Collibra and Unity Catalog rather than a replacement.

### How do I start using xorq with Claude Code?

Add the marketplace and install the plugin from the separate xorq-labs/claude-plugins repository. That contributes four slash commands: /xorq:init, /xorq:catalog-explore, /xorq:composer and /xorq:builder. The manual route instead is pip install xorq[examples] followed by xorq init -t penguins.

### Does xorq need a catalog server to run?

No. The catalog is a git repository of entries on disk with git-annex support for large files, and the comparison table describes the operating model as no service to call. A lookup runs the entry and returns an Arrow stream rather than handing you a governed table to query.

## Sources

- [License: Apache-2.0](https://github.com/xorq-labs/xorq/blob/main/LICENSE)
- [Project website](https://docs.xorq.dev)
- [README](https://github.com/xorq-labs/xorq/blob/main/README.md)
- [Releases](https://github.com/xorq-labs/xorq/releases)
- [xorq-labs/xorq on GitHub](https://github.com/xorq-labs/xorq)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/xorq-labs-xorq
