Duckle: Open-Source ETL on DuckDB with a Visual Canvas and In-Process AI
Open-source ETL/ELT on DuckDB. Write, wire, or draw one pipeline: 350+ components, 160+ connectors, dbt, CDC, data quality, a Python API, and MCP for AI agents. Runs anywhere, no lock-in.
At a glance
- What is it?
- Duckle is an open-source ETL platform from SlothFlowLabs that compiles pipelines to SQL and executes them on DuckDB. It offers a visual canvas, a Python API, and a SQL editor as three authoring modes, ships as a 73 to 110 MB self-contained binary, and includes an in-process AI assistant that requires no external API key.
- Who is it for?
- Duckle suits data engineers and small teams who want a self-hosted ETL tool that runs on their own machines or servers without per-row billing or a vendor cloud dependency. The DuckDB execution model is efficient for batch workloads on a single machine, and the visual canvas plus git-friendly file format make pipelines auditable by non-engineers.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 5 days ago.
- What is it written in?
- Mainly Rust, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What Problem Duckle Solves and Who It Targets
Hosted ETL platforms such as Fivetran and Airbyte charge per row or per connector, which makes them expensive for teams with high data volumes. Self-hosting those same tools adds operational overhead without eliminating the dependency on their vendor infrastructure. Duckle takes a different position: it is a single binary that the user downloads and runs on their own server or laptop, with no per-row pricing and no mandatory cloud component.
The README describes it as "a free, open-source, single-engine alternative to hosted, per-row-priced ETL platforms." The target user is a data engineer or small team that runs batch pipelines on a single machine, wants to version-control pipelines in git, and prefers to understand every step of what the tool is doing. The visual canvas and generated SQL are both visible during pipeline construction, so there is no hidden transformation logic.
Duckle is an independent project by SlothFlowLabs, built on DuckDB but not affiliated with DuckDB Labs or MotherDuck. The README states this explicitly.
Three Ways to Build a Pipeline: Canvas, Python, and SQL
Duckle provides three authoring modes for the same pipeline format. The visual canvas allows dragging sources, transforms, validators, and sinks onto a workspace and wiring them together. The README describes the canvas as compiling the graph to SQL with live previews and the generated SQL visible on every node, so engineers can inspect what the pipeline will actually execute at each step.
The Python API provides a programmatic interface for teams that prefer code over a graphical tool or that need to generate pipelines dynamically. The SQL editor is the third option for teams already comfortable with SQL-based transformation.
All three modes produce the same pipeline file format. The README states that every pipeline is one file in git, meaning a pipeline built in the canvas and one written in Python produce the same artifact. This portability matters for team workflows: a data analyst can prototype a pipeline in the canvas and a data engineer can review and modify the resulting file in a code editor.
DuckDB as the Only Execution Engine
Duckle compiles every pipeline to SQL and executes it on DuckDB, which provides vectorized columnar processing using all CPU cores on the host machine. The README states that scaling means giving the process more hardware: "a bigger instance is a faster pipeline." The performance claim in the README is specific: 96 million rows from Postgres to Parquet in 39.9 seconds.
The Cargo workspace reveals the internal architecture. The duckle-duckdb-engine crate is the DuckDB integration. Separate crates handle streaming (duckle-stream-engine), transformations (duckle-transform-engine), workflow orchestration (duckle-workflow-engine), and the runner (duckle-runner). The duckle-mcp crate packages the MCP server for AI agent integration.
DuckDB downloads on first launch through a guided step rather than being bundled inside the binary. This keeps the main binary at 73 to 110 MB depending on platform while allowing the DuckDB version to be updated independently. The AI engine for Duckie is opt-in and separate from the base install.
Deploying Pipelines with duckle-runner
The README describes the deployment workflow as authoring on a laptop and then shipping the same file to a server. The duckle-runner binary handles the server side. Launching it with the serve subcommand starts a headless execution environment with scheduling, a web console, role-based access control, and an audit log:
duckle-runner serveThe runner can execute in Docker or directly on a server the team controls. Pipelines run on a schedule that the console manages, and the audit log records execution history. The README emphasizes that nothing is rewritten between development and deployment: the same pipeline file runs locally and in production without transformation.
For version control, pipelines, connections, context variables, and routines persist as plain files in a directory the user chooses. The README describes this as diff-friendly: the files can be reviewed in a code review, branched, and merged using standard git workflows. In-app git integration with GitHub and GitLab is listed in the README's feature table.
Duckie: The Built-In AI Assistant
Duckle ships with an AI assistant called Duckie that accepts natural-language descriptions of a desired pipeline and writes the JSON configuration, placing it on the canvas. The README describes three properties of this assistant that distinguish it from external AI integrations: the model runs in-process by default without requiring an API key, prompts and data stay inside the user's infrastructure, and the user can point Duckie at their own OpenAI-compatible endpoint if they prefer a different model.
This is a practical difference from tools where AI assistance requires sending data or schema information to a third-party API. For teams with data governance requirements that restrict sending schema details to external services, the in-process option removes that concern.
The in-process AI engine is opt-in and not bundled in the base binary download. The README lists it as a separate install step alongside the DuckDB engine. The MCP server in the duckle-mcp crate exposes Duckle's pipeline management to external AI agents running in tools like Claude or Cursor.
Limitations: Single-Machine Scope and the Connector Coverage Question
The README is explicit about Duckle's scope: it is single-machine and embedded by design. The stated purpose is making local and small-team data work fast, not replacing a distributed data warehouse. Teams that need to distribute processing across a cluster for petabyte-scale workloads should look at Apache Spark or Databricks instead. Spark is a distributed processing engine that runs across multiple machines and is designed for horizontal scaling, while Duckle targets single-node efficiency on DuckDB.
The 400+ connector count covers files, SQL databases, warehouses, lakehouses, object stores, SaaS REST APIs, NoSQL databases, streaming brokers, vector databases, FTP, IMAP, and SMTP according to the README. The gap is that very niche or enterprise-specific connectors may not be present. Before adopting Duckle, engineers should verify their specific sources and sinks are in the connector list, since discovering a missing connector after a migration decision is costly.
The dbt integration and CDC support are listed in the repository description but the README does not document their specific configurations or limitations.
Editorial conclusion
Duckle suits data engineers and small teams who want a self-hosted ETL tool that runs on their own machines or servers without per-row billing or a vendor cloud dependency. The DuckDB execution model is efficient for batch workloads on a single machine, and the visual canvas plus git-friendly file format make pipelines auditable by non-engineers. The README explicitly positions it as a single-machine and embedded tool: teams that need distributed execution across multiple nodes for very large-scale workloads will need something else. Before adopting it, verify that your data sources are covered by the connector list on the Duckle website, since 400+ connectors is a large set but not exhaustive, and check that the dbt and CDC features meet your specific version requirements.
Frequently asked questions
What is Duckle?
Duckle is an open-source ETL platform built on DuckDB. It provides a visual canvas, Python API, and SQL editor for building data pipelines that compile to SQL and execute locally on DuckDB. It ships as a self-contained binary and includes a headless runner for scheduled server deployment.
Does Duckle require an API key for its AI assistant?
The built-in Duckie assistant runs in-process by default and requires no external API key. Users can optionally configure an OpenAI-compatible endpoint to use a different model, in which case the requests go to that endpoint rather than an in-process model.
Can Duckle pipelines be managed in git?
Yes. The README states that every pipeline is one file in git. Pipelines, connections, context variables, and routines persist as plain files in a directory the user chooses, making them compatible with standard git diff, branch, and merge workflows. In-app GitHub and GitLab integration is also listed in the feature documentation.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/slothflowlabs-duckle)