Model or dataset
rocky-data/rocky avatar
rocky-data/rocky

Rocky: compile-time type checking for SQL pipelines across Databricks, Snowflake, BigQuery and DuckDB

A SQL transformation engine that type-checks your whole pipeline and catches breaking changes before they run, branches, replay, column-level lineage, compile-time contracts, per-model cost. Adapters: Databricks, Snowflake, BigQuery, DuckDB. Single static Rust binary. Apache 2.0.

300 stars18 forksRustApache-2.0

At a glance

What is it?
Rocky is a Rust SQL transformation engine that checks types, refs and contracts across a whole pipeline before anything runs. It is aimed at data engineers on Databricks first, and it ships as a single static binary under Apache 2.0.
Who is it for?
Adopt Rocky if you run SQL transformations on Databricks, Snowflake, BigQuery, Trino or DuckDB and you want type errors, broken refs and contract violations to surface at compile time rather than in production. Do not adopt it if you need a scheduler: the README points at Dagster for that, and the repository ships a Dagster integration rather than a built-in one.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 5 days ago.
What is it written in?
Mainly Rust, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The quiet failure Rocky is built to catch

The README frames the problem in one line: "The failures that cost the most are the quiet ones." A source column changes type. Someone renames a column and three downstream models stop working. A query passes in dev and fails in production. None of these throw an error at the moment the edit is made, because SQL engines generally do not resolve types until the query is submitted to the warehouse.

Rocky moves that resolution earlier. It reads your model files, resolves column types and references across the whole dependency graph, and reports problems before a row is written. The audience is stated plainly: data engineers on Databricks, where the README argues a silent failure costs the most and where Dagster usually runs the schedule. Snowflake, BigQuery, Trino and DuckDB are also supported, with the README pointing readers to the Adapters section for what each one does today. That pointer is worth taking literally, because the README does not enumerate per-adapter capability in the text it provides.

How the compile and run stages differ

The mechanism is a two-stage split. You edit SQL model files, then `rocky compile` checks the whole pipeline (types, refs, contracts) without touching the warehouse, and then `rocky run` performs the writes. The README's own diagram shows a problem being found at the compile stage, before the arrow reaches the warehouse box.

The compile output is a per-model column count and an error and warning tally. The run output is a receipt: model names, the schema they landed in, and the refresh mode. In the README's own example, three models compile with zero errors and zero warnings, and the run reports `3 model(s) executed in 20ms` against the `playground.main` schema. That 20ms figure comes from the README's illustration on local DuckDB, not from a production warehouse, so treat it as a demonstration of the shape of the output rather than a performance claim.

Two deployment paths exist. `rocky plan` saves what will change and `rocky apply <plan-id>` runs it, which is the production route. `rocky run` collapses both steps, and the README recommends it for local work and automation.

Installing Rocky and running the DuckDB playground

The README gives install commands for macOS, Linux and Windows. The macOS and Linux path pipes a shell script from the repository into bash; the Windows path does the equivalent in PowerShell. Both fetch from the `main` branch of the GitHub repository, so you are installing from source-controlled script rather than a package manager.

bash
curl -fsSL https://raw.githubusercontent.com/rocky-data/rocky/main/engine/install.sh | bash
bash
irm https://raw.githubusercontent.com/rocky-data/rocky/main/engine/install.ps1 | iex

The README then describes a playground that needs no credentials, because it runs on local DuckDB. The command is `rocky playground my-first-project`, which scaffolds a directory you then change into.

bash
rocky playground my-first-project
cd my-first-project
rocky compile && rocky test && rocky run

What you should see: `rocky compile` prints a checkmark and a column count for each model, followed by a summary line such as `Compiled: 3 models, 0 errors, 0 warnings`. `rocky test` runs the model tests, and `rocky run` reports the executed models with their schema and refresh mode. If you are building from source instead, the `justfile` at the repository root defines `build-engine` as `cd engine && cargo build --release`, and `build-ui` as an npm build inside `engine/ui` that plain `cargo build` does not need.

Lineage diff, branches and replay as review tools

The most interesting command in the README is `rocky lineage-diff main`, which compares two versions of a project and emits Markdown you can paste into a pull request comment. The output is a table of column changes with a downstream consumers column. In the README's abridged example, adding `amount_usd` and removing `amount` in `stg_orders` shows that `fct_revenue.total_revenue` reads the removed column. That is the useful part of the report: not that a column changed, but who was reading it.

The README notes one honest gap in the same output. A removed column is marked `(removed; not traceable)`, meaning the tool can tell you the column is gone but cannot always tell you where it went. Renaming is inferred from the pairing of an added and a removed column in the same model, not from an explicit rename declaration.

Branches and replay get their own demo. The description is that you run against an isolated copy, inspect it, then drop or promote it. That is a workflow for testing a change against real data shape without writing to the production tables. The README does not document what happens to a branch if the process is interrupted between creation and promotion, so the cleanup path is worth confirming against the demo scripts in `examples/playground/pocs/00-foundations/06-branches-replay-lineage/`.

Where Rocky is the wrong tool

Rocky is not a scheduler. The README states that Dagster usually runs the schedule for its primary audience, and the repository ships an `integrations/dagster` directory with its own build and test targets in the `justfile`. If you are looking for an orchestrator with backfills, sensors and retries, Rocky is the layer underneath that, not a replacement for it.

The adapter story is also uneven by the project's own account. The README lists Databricks, Snowflake, BigQuery, Trino and DuckDB as targets, but explicitly defers per-adapter capability to a separate Adapters section rather than summarizing it. The BigQuery cost demo is flagged as needing credentials, which means the byte-accurate cost reporting cannot be evaluated in the credentials-free playground. If cost prediction is your reason for adopting, that is the piece you have to test with a real account.

Finally, the install path is a piped remote script with no package-manager option documented in the README. Teams with restrictions on `curl | bash` will need to build from source via the `justfile` targets or vendor the script.

How Rocky differs from dbt

The obvious comparison is dbt, and the difference is where the checking happens. dbt compiles Jinja into SQL and hands the result to the warehouse, which resolves types when the query executes. Rocky resolves types across the pipeline itself, before the warehouse is involved, which is why `rocky compile` can report a column count and a type for each model without credentials.

That design choice has a cost. Rocky has to model the type system of each adapter, and the README's deferral of adapter detail to a separate section suggests that coverage varies. dbt's approach inherits whatever the warehouse supports, at the price of finding out later.

The second difference is the plan and apply split. `rocky plan` produces a saved plan that `rocky apply <plan-id>` executes, which is closer to Terraform's model than to a transformation runner that executes on invocation. For teams that want a reviewable artifact between "this is what will change" and "this is what changed", that is the meaningful distinction. For teams that want one command, `rocky run` is still there.

Agent policy, licence and the upgrade surface

The README devotes a section to AI agents writing pipeline changes, and the framing is a permissions problem rather than a generation problem. Rocky type-checks what an agent writes, the agent produces a plan, and a plan never applies itself: it must pass rules you wrote, and every decision lands in a ledger you can query. There is a demo at `examples/playground/pocs/03-ai/07-policy/` where CI catches a rule that was loosened by accident.

The repository layout supports the claim that this is a first-class concern rather than a marketing section. There are `AGENTS.md`, `AGENT_REVIEW.md`, `CLAUDE.md` and `CODEX_REVIEW.md` at the root, plus a `justfile` target `evals` that runs an agent conformance suite requiring a `rocky` build, `$ANTHROPIC_API_KEY` and the claude and duckdb CLIs, and skips cleanly without them. A separate `evals-selftest` target runs the harness plumbing without a model key, which is a sensible split for CI.

Licensing is Apache 2.0, confirmed by the LICENSE file at the repository root. That permits commercial use and modification, and it includes a patent grant. It does not answer whether the bundled VS Code extension or the Python SDK carry different terms; the README links to the Marketplace listing for the extension but does not discuss its licence. Check the `editors/vscode` and `sdk/python` directories before assuming a single licence covers the monorepo.

On maintenance: the repository is not archived, and the last push was on 2026-08-26, with `engine-v1.72.0`, `sdk-v0.13.0` and `vscode-v1.39.0` released the same day. Upgrade cost is dominated by the engine, which is a single static Rust binary, so the practical question is whether your model SQL depends on behaviour that changed between engine versions. The README does not document a deprecation policy or a migration guide, so pin the engine version in CI and read the release notes before moving.

Editorial conclusion

Adopt Rocky if you run SQL transformations on Databricks, Snowflake, BigQuery, Trino or DuckDB and you want type errors, broken refs and contract violations to surface at compile time rather than in production. Do not adopt it if you need a scheduler: the README points at Dagster for that, and the repository ships a Dagster integration rather than a built-in one. Before you commit, verify what each adapter actually supports today, since the README directs you to the Adapters section for that, and confirm the Apache 2.0 terms in the LICENSE file against your own distribution model.

Frequently asked questions

What is Rocky and who is it for?

Rocky is a SQL transformation engine that checks a whole pipeline (types, refs, contracts) before it runs. The README says it is built first for data engineers on Databricks, and it also runs on Snowflake, BigQuery, Trino and DuckDB.

How do I install Rocky?

The README gives a shell install for macOS and Linux that pipes engine/install.sh from the main branch into bash, and a PowerShell equivalent using engine/install.ps1. There is no package-manager install documented in the README.

Can I try Rocky without warehouse credentials?

Yes. The README states the playground needs no credentials because it runs on local DuckDB, and the sequence is rocky playground my-first-project followed by rocky compile && rocky test && rocky run. The BigQuery cost demo, by contrast, is flagged as needing credentials.

Does Rocky replace an orchestrator like Dagster?

No. The README says Dagster usually runs the schedule for Rocky's primary audience, and the repository ships a separate integrations/dagster directory with its own build and test targets. Rocky handles compile-time checking and execution, not scheduling.

What licence does Rocky use?

Apache 2.0, per the LICENSE file at the repository root. The README does not state whether the VS Code extension or the Python SDK are covered by the same terms, so check those directories if it matters to you.

What does rocky lineage-diff output?

It compares two versions of a project and writes Markdown listing the downstream tables and columns each change affects, intended to be pasted into a pull request comment. The README notes that removed columns can appear as not traceable.

Official sources

  1. Official documentation
  2. Official README
  3. Project repository
  4. Release notes
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/rocky-data-rocky.svg)](https://hysenlabs.com/projects/rocky-data-rocky)