Databend rebuilds the cloud warehouse as an agent runtime on your own object storage
Data Agent Ready Warehouse : One for Analytics, Search, AI, Python Sandbox. — rebuilt from scratch. Unified architecture on your S3.
At a glance
- What is it?
- A Rust data warehouse unifying analytics, vector search and full-text search on S3-compatible storage, now adding sandboxed Python UDFs and SQL orchestration aimed at AI agents.
- Who is it for?
- Databend fits data teams that want warehouse-scale analytics plus retrieval workloads on their own object storage, and that are ready to run agent logic inside sandboxed SQL functions with branching as the safety net. Skip it if a fully managed warehouse matches your staffing reality, if you already run a dedicated vector store you like, or if the Elastic License's service-provider restrictions touch your business model.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository last received commits 2 days ago.
- What is it written in?
- Mainly Rust, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 28, 2026, and from our analysis. They are not legal advice.
Editorial analysis
A warehouse rebuilt for agents, not just dashboards
Databend has years of history as an open-source, cloud-native analytics warehouse, and its README now leads with a repositioning: an enterprise data warehouse for AI agents, rebuilt from scratch, with the description calling it a Data Agent Ready Warehouse. The engine is written in Rust, stores data on your own object storage, S3, Azure or GCS, and unifies four jobs that usually require four products: SQL analytics, vector search, full-text search and auto schema evolution, with transactions underneath.
The agent framing is not a sticker on the website. It corresponds to concrete machinery: sandboxed Python UDFs that can hold agent logic, SQL as the orchestration layer that calls that logic at scale, transactions for reliability, and git-like branching so an agent can experiment on a production snapshot without touching production. That combination is the actual product thesis, and the README documents each piece.
Quick start: a pip one-liner and a Docker container
Three entry paths are documented, in descending order of hand-holding. Databend Cloud is the recommended route for production, described as ready in sixty seconds. For local development, a Python driver runs an embedded instance; the README specifies Python 3.12 or 3.13 and databend-driver 0.34.0 or later:
pip install "databend-driver[local]>=0.34.0"and then two lines of Python get you a query:
from databend_driver import connect
conn = connect("databend+local:///./local-state")
print(conn.query_row("SELECT 'Hello, Databend!'").values())The full warehouse runs locally in one Docker command:
docker run -p 8000:8000 datafuselabs/databendThe local driver writes state into a directory you choose, which makes it a realistic option for tests and prototypes rather than a toy demo mode.
The Sandbox UDF: agent logic as a SQL function
The agent architecture has three named planes. A control plane handles resource scheduling, permission validation and sandbox lifecycle. The execution plane is Databend itself, orchestrating through SQL and issuing work over Arrow Flight. The compute plane consists of isolated sandbox workers running your agent code.
The SQL shape is the part worth staring at, because it is unusual:
CREATE FUNCTION my_agent(input STRING) RETURNS STRING
LANGUAGE python HANDLER = 'run'
AS $$
def run(input):
# Your agent logic: LLM calls, tool use, reasoning...
return response
$$;
SELECT my_agent(question) FROM tasks;A Python function becomes a SQL function; SQL then fans it out across a table of inputs. The README's example comment lists LLM calls and tool use as the intended body, which means warehouse rows can drive agent invocations with the engine handling scheduling and isolation. For data teams, that moves agent workloads into the same governance and audit perimeter as every other query.
Branching and transactions: the safety story
Two older warehouse virtues are repositioned as agent safety features, and the reframe holds up. Branching is described as git-like data versioning: an agent that needs to try an aggressive transformation can work on a branch, compare results, and be discarded without contaminating production data. Transactions mean multi-step operations either land completely or not at all, which is the difference between an agent that corrupts a table halfway through a plan and one that simply fails.
The README's use-case table makes the pairing explicit: sandbox UDFs give agents somewhere to run, SQL orchestration gives them scale, and branching gives them a safety net on production snapshots. For anyone who has watched an autonomous script mutate the wrong table, that trio, not the vector search, is the reason to take the agent positioning seriously.
One engine, four workloads
The unification claim deserves its own accounting. Analytics is the heritage: large-scale SQL with elastic compute over object storage. Vector search and full-text search cover the retrieval half of RAG pipelines, so embeddings and keyword indexes live beside the structured tables they join against, instead of in a separate specialty store. Auto schema evolution reduces the friction of semi-structured sources that change shape under you.
The practical consequence is architectural: a RAG pipeline can ingest documents, chunk and embed them, index them for both retrieval modes, and join retrieved chunks against operational tables in one system, governed by one permission model. The README's search and RAG use case links to documentation for exactly this pattern. Whether one engine's vector performance matches a dedicated vector database at your scale is a benchmark question the README wisely does not answer with marketing numbers; it is answerable on your own data.
Licence: Apache 2.0 plus Elastic 2.0
The licensing is dual: Apache 2.0 alongside Elastic License 2.0, with both texts in the repository and a licensing FAQ in the documentation. GitHub's own detector reports the repository as NOASSERTION, which is what happens when a repo is dual-licensed this way, and the distinction matters for anyone planning to build on it.
The short version, without legal advice: Apache 2.0 is the permissive licence most users expect, while the Elastic License restricts offering the software itself as a managed service, which is the standard open-core boundary for infrastructure companies. If you are a data team adopting Databend internally, the question barely arises. If you are a cloud provider or startup planning to sell it as a service, read the Elastic text and the FAQ before writing a business plan around it. The repository also carries an AI_POLICY.md and its own AGENTS.md and CLAUDE.md files, a small sign the maintainers run AI tooling against this codebase under a stated policy.
Against Snowflake, ClickHouse and DuckDB
The comparisons in the search data are with Snowflake, and the structural difference is control: Snowflake is a managed commercial service where the vendor runs everything; Databend runs on your own S3-compatible storage, including on-prem object stores, with an open codebase you can read and an embedded local mode for development. Teams with data-residency requirements or existing storage contracts weigh that differently than teams that just want zero operations.
Against column-oriented engines like ClickHouse, the difference is scope rather than speed: ClickHouse remains a specialist at high-throughput analytics, while Databend bundles vector and full-text search, branching and the Python sandbox into the same engine. DuckDB occupies the embedded analytics slot and is a more natural comparison for the pip-install local mode than for the warehouse. Databend's defensible middle is the combination: warehouse scale, retrieval workloads, and agent execution in one governed place, which none of the three alternatives offers as a single system.
Editorial conclusion
Databend fits data teams that want warehouse-scale analytics plus retrieval workloads on their own object storage, and that are ready to run agent logic inside sandboxed SQL functions with branching as the safety net. Skip it if a fully managed warehouse matches your staffing reality, if you already run a dedicated vector store you like, or if the Elastic License's service-provider restrictions touch your business model. Verify the claims on your data, not the README's: run the Docker one-liner, load a real table, try the vector and full-text paths, and only then measure what a sandboxed UDF agent would actually cost you in compute.
Frequently asked questions
Can Databend run locally?
Yes. For development, pip install databend-driver[local] version 0.34.0 or later runs an embedded instance on Python 3.12 or 3.13, and docker run -p 8000:8000 datafuselabs/databend runs the full warehouse locally.
What licence does Databend use?
It is dual-licensed under Apache 2.0 and Elastic License 2.0, with both texts in the repository and a licensing FAQ in the docs. GitHub's detector reports NOASSERTION because of the dual licensing.
What makes Databend agent-ready?
Sandboxed Python UDFs hold agent logic as callable SQL functions, SQL orchestrates them at scale over Arrow Flight, transactions guarantee atomic multi-step operations, and git-like branching lets agents work safely on production snapshots.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/databendlabs-databend)