Datus: a data agent whose requirements file is hand-copied
The Future of Data Engineering — A CLI SQL client for the modern data stack, enabling AI-native context engineering for data.
At a glance
- What is it?
- Datus is a Python data engineering agent that connects a warehouse, catalog, semantic layer and BI, ships as a CLI, a web chat and a REST and MCP API, and is built from five separately published first-party packages. Its most revealing file is requirements.txt, a hand-maintained duplicate of the pyproject dependency list with a CI script that fails the release if the two drift apart.
- Who is it for?
- Evaluate Datus if your problem is that your warehouse has no trustworthy semantic layer, because automated semantic modelling from schema plus SQL history and a Dosi execution engine that compiles one model into thirteen dialects is a more durable asset than another text-to-SQL prompt, and the context engine that writes every correction back is the mechanism that makes it improve.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository last received commits 1 day ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Three descriptions of one product, and the word Compiled
The same tool describes itself three different ways in three different places, and the divergence is worth reading because each version is aimed at a different audience.
The repository description is the marketing one: The Future of Data Engineering, a CLI SQL client for the modern data stack, enabling AI-native context engineering for data.
The README's own first line is the engineering one: Datus is the open-source data engineering agent for the modern data stack, one agent that connects your warehouse, catalog, semantic layer and BI, grounded in an evolvable context engine your team owns.
And the PyPI metadata, which is what a person sees when they type pip install datus-agent, says: AI-powered SQL Agent for data engineering, followed by the two words Compiled Version in parentheses.
That parenthetical is the odd one out. Nothing else in the repository says anything about compilation. The Makefile builds a Python distribution through a script called build_pypi_package.py, the package is a pure-Python project with a pyproject.toml and a requires-python of 3.12 or newer, and the README describes a system of agents, a context engine and adapters rather than a binary. A parenthesised Compiled Version in the package description suggests either an older product identity that was not updated, or a plan to distribute something that is not a pure-Python wheel, or simply a stray phrase.
The three descriptions also disagree on the category. The repository description calls it a CLI SQL client. The README calls it an agent with three entry points: a CLI for data engineers, a chat interface for analysts on the web, in Slack or Feishu, or in VS Code, and a REST and MCP API for other agents and applications. A CLI SQL client and a three-surface agent platform are different products, and a reader arriving from the PyPI listing and a reader arriving from the GitHub repository form different expectations about the same install command.
None of this is a correctness problem. The three descriptions are compatible once you know that the product has grown from a command line tool into a platform, and the repository description is simply the shortest of the three. It is worth naming because package metadata is read by people deciding whether to install something, and because a phrase like Compiled Version in a wheel's description is the kind of thing that makes a reader wonder what they are about to get.
requirements.txt is a hand-kept copy, and CI fails when it drifts
The first thing in requirements.txt is a comment explaining why the file exists at all, and it is the most interesting engineering decision visible in this repository.
It reads: runtime dependencies, mirroring project.dependencies in pyproject.toml. Then: ci/check_release_readiness.py compares the two and fails on any drift. And then the consequence: every edit here needs the matching edit there, and vice versa.
So this project maintains its dependency list twice, on purpose, and has written a release gate that refuses to proceed if the two copies disagree.
That is unusual enough to be worth asking why. A modern Python package declares its dependencies in pyproject.toml and does not ship a requirements.txt at all; pip resolves from the metadata. Several reasons could make a maintainer want a flat list anyway, and the file gives hints about which applies here.
The most likely reason is that the flat list is for a deployment environment that does not use the package. A team shipping this into a container, a VM or an air-gapped industrial network frequently wants a single text file to feed to a provisioning tool, to review in a security scan, or to install into a locked-down image builder. In that case the duplication is not an accident but a deployment interface, and the CI check is what keeps an interface from rotting.
The file also looks like a security-review artefact. Almost every entry carries an exact pin or a tight range, several carry a comment explaining why, and the set includes a broad and deliberate spread: HTTP clients, a SQL toolkit, a data frame library, two search engines, an embedded analytics database, a spreadsheet writer and two spreadsheet readers, an image library, an LLM gateway, two model provider SDKs, an agent framework, tracing, a terminal UI stack, a web framework, an ASGI server, the Model Context Protocol client, GitHub integration, HTML parsing, and two XML libraries one of which is defusedxml. That last detail is a small good sign: defusedxml is the hardened XML parser, and having it in a project that also has lxml and an HTML parser suggests someone thought about XML parsing rather than reaching for the standard library.
A few of the pins come with their reasoning inline, which is a practice worth copying. The comment on the openai pin explains that OpenAI 2.45 makes cache_write_tokens a first-class field in the input token details, so the session store persists it without a patch on the project side. The comment on openpyxl explains that DuckDB's Excel extension can read sheet contents but cannot enumerate sheet names or report the used range, both of which are needed to build a correct view, and that xlrd covers legacy .xls formats DuckDB does not support at all. Those are dependency decisions with the reasoning attached, which is what you want from a project that pins exactly.
The cost of the duplication is the cost the CI check pays for. Two files to edit, a script to keep working, and a class of failure where a developer's local install differs from a release. The project has made that cost explicit and paid it deliberately, and the comment tells you exactly what to do when you add a dependency. That is better than the alternative, which is a requirements.txt that drifts for two years and nobody notices.
Two search stacks in the dependencies, and one the README does not mention
The context engine is described in the README in one sentence, and the dependency list describes it in a way that adds something.
The README says the context engine holds metadata, metrics, reference SQL, knowledge and local files, retrieved through business-domain trees plus vector search, with storage on embedded LanceDB and SQLite and PostgreSQL for teams that share context.
So the documented retrieval mechanisms are two: a tree structure organised by business domain, and vector search. The storage is named as LanceDB and SQLite embedded, with PostgreSQL as the shared option.
The requirements file names three more libraries that belong to that layer. lancedb is pinned at 0.34.0 and is the vector store the README mentions. fastembed is pinned at 0.4.2 and is an embedding library, which is what produces the vectors; the README does not mention it, which is reasonable since embedding is an implementation detail of vector search. And tantivy, at 0.22.2 or newer, is a full-text search engine written in Rust and best known as the search core behind Meilisearch. There is no full-text search component in the README's description of the context engine.
So the context engine appears to do lexical search as well as vector search, and the lexical half is undocumented. That matters for two reasons. A business-domain tree is a hierarchical filter, and a tree plus vector search is a common and effective combination, because the tree narrows the candidate set and the vectors rank within it. Adding a full-text engine on top is a third mechanism, and lexical search is what reliably finds an exact identifier: a table name, a column name, a specific metric string. Vector search on identifiers is unreliable in a way that surprises people, and a full-text engine fixes it.
So the dependency list suggests a better retrieval design than the documentation describes, which is a pleasant kind of finding: the implementation is ahead of the prose. It also means there is an unstated third thing to evaluate, because the quality of a context engine for a data team depends heavily on whether it can find the exact table the analyst meant, and the answer here appears to be yes, via a dependency the README does not mention.
The broader pattern is worth naming. Across fifty-odd runtime dependencies there are at least four search or retrieval systems: LanceDB for vectors, Tantivy for full text, and the business-domain tree implemented in the project itself, plus whatever the Dosi semantic layer does to resolve a metric name to a query. Choosing to run three retrieval mechanisms in one process is a real operational decision, with three indexes to build, three to keep in sync with the warehouse, and three sets of failure modes. It may well be the right call for a data agent whose failure mode is confidently answering from the wrong table, and it is a call the documentation does not surface.
Five separately published packages, so the monorepo boundary is the API
The last five lines of requirements.txt are first-party dependencies, and they are the most important entries in the file for anyone deciding how to depend on this project.
They are datus-storage-base at 0.1.5 or newer and below 0.2.0, datus-db-core at 0.1.7 or newer, datus-semantic-core at 0.2.4 or newer, datus-bi-core at 0.1.2 or newer, and datus-scheduler-core at 0.1.1 or newer.
So Datus is not one package. It is five, published to PyPI, split by domain: database access, storage, the semantic layer, business intelligence, and scheduling. The main datus-agent package depends on all of them.
That is a deliberate and conventional way to structure a project of this size, and it has a specific consequence for consumers. Depending on datus-agent means depending on five version lines, each with its own release cadence, and the upper bound on datus-storage-base at below 0.2.0 tells you the project is using a major-version scheme to signal breaking changes in the internal libraries. All five are at 0.1.x or 0.2.x, so none of them has reached 1.0 either.
The domain split also explains the architecture section. The connected-systems list names LLM providers, data warehouses, the Dosi semantic layer, job schedulers, BI tools and MCP servers and clients, and the five packages map onto four of those six categories, with the LLM providers and MCP living in the main package. So the package boundaries follow the integration boundaries rather than the code layout, which is the right way round and is visible from the dependency list alone.
For an evaluator, the useful reading is that the internal structure is already a published API. If you want only the semantic layer, you take datus-semantic-core and skip the agent. If you want the storage layer, you take datus-storage-base. That is a much better adoption surface than a monolith, and it means the agent is a convenience layer over components you can use individually.
The counter-argument is the one the pin ranges make. Five pre-1.0 packages with a coupled release process is five things that can be incompatible with each other, and the fact that requirements.txt exists as a deployment artefact suggests the vendor expects teams to install the whole set together rather than to compose them. So the modularity is real but the coupling is real too, and the version numbers on the internal packages are the thing to watch as the project matures.
Alpha, at 0.4.1, with three releases in six weeks
The release history and the package classifier tell two different stories, and the tension between them is the honest summary of this project's maturity.
The three most recent releases are v0.3.9 on 2026-08-02, v0.4.0 on 2026-08-27, and v0.4.1 on 2026-09-10. The last push to the repository was on 2026-09-19. So three releases in six weeks, a minor version bump in the middle, and nine days of further commits after the newest tag. That is an actively developed project on a monthly-ish minor cadence with patch releases behind each minor.
The classifier in pyproject.toml says Development Status :: 3 - Alpha. The version is 0.4.1. The project is self-declared as alpha software.
Read together, the two statements are consistent in the narrow sense and inconsistent in the wider one. Narrowly, the project is telling you the API is not stable, which a 0.x version number already implied, and it is using the classifier to make that explicit to tooling. Widely, the README describes itself in the language of a finished platform: one agent that connects your warehouse, catalog, semantic layer and BI, grounded in an evolvable context engine your team owns, with the whole stack staying open and flexible. That is confident framing for something declaring itself alpha, and a reader could reasonably end up with a different sense of the risk than the classifier conveys.
The more useful signal is in the numbers rather than the words. Roughly twenty of the fifty-odd runtime dependencies carry an exact pin using the double-equals form, and a substantial number of the rest carry a floor with no ceiling. For a library that others install, exact pins are a hard-nosed choice with two consequences that pull in opposite directions. On the good side, a Datus release resolves to one known set of versions, so a conflict with your own dependency on, say, the data frame library or the model provider SDK cannot happen. On the bad side, you cannot take a patch release of any pinned dependency without waiting for a Datus release, and the internal five packages carry the same risk because they pin their own trees.
For a data engineering agent whose output quality depends on the model provider and the retrieval libraries behaving as expected, exact pinning is defensible. For a security response, it is a problem: when a pinned dependency has a vulnerability, your remediation path is a Datus release, not a resolution change. That is a real operational constraint and it is worth raising with the project before you adopt rather than after.
The other pinned detail worth noting is the Python floor. The package requires 3.12 or newer, so it is a modern-only package, while several of its pins are from earlier Python generations. Nothing inconsistent there, but it does mean the project cannot be installed on the older interpreters that a lot of data infrastructure still runs, which narrows where you can deploy it.
The Makefile is one Python script wearing seven aliases
The Makefile in this repository contains no logic at all, which is a deliberate choice worth examining because it is unusual to see done this consistently.
Every single target is a call into the same Python script:
build: ## Build the package
python build_scripts/build_pypi_package.py build
install: ## Install package locally (editable mode)
python build_scripts/build_pypi_package.py install
test: ## Test the installation
python build_scripts/build_pypi_package.py test
check: ## Check package before upload
python build_scripts/build_pypi_package.py check
upload: ## Upload to PyPI
python build_scripts/build_pypi_package.py uploadThe targets are clean, install, install-dist, test, check, upload-test, upload, all and publish, and clean, plus dev-install and setup-dev which use uv directly. The composite targets are compositions: all is clean plus build plus check plus test, and publish is clean plus build plus check plus upload. And there are convenience targets that chain the composites, with quick-build, quick-test and quick-publish.
So the Makefile exists to give the project the command names its users and contributors expect, and to document them through its own help target, which greps the target list for a double-hash comment and prints the result with the target name padded to twenty columns. That is a functioning help system, and it is why every target in the file carries a comment.
The design has a clear rationale and a clear cost. The rationale is that build logic in a Python project belongs in Python, where it can be tested, where it can use the same packaging libraries the build produces, and where it does not need to be reimplemented in shell. The cost is that a Makefile which used to be a transparent description of what happens is now an indirection, and a reader who wants to know what the release does has to open build_scripts/build_pypi_package.py instead.
There is one asymmetry. The dev-install and setup-dev targets bypass the script and call uv directly, installing the project in editable mode with the dev extra and then installing build and twine. So the project uses the script for the packaging lifecycle and uv for development setup. That is a sensible split, and it also means the toolchain has two build paths, which is a mild version of the requirements.txt duplication discussed earlier: two places to keep in sync, guarded in one case by CI and not in the other.
The rest of the repository's tooling is more conventional. There is a uv.lock committed, so development environments are reproducible. There is a pre-commit configuration, a pytest configuration, a separate requirements-test.txt and requirements-docs.txt, a Makefile-adjacent install.sh and install-dev.sh, a conf/ directory for configuration, and a ci/ directory which is where the release-readiness script that guards the dependency duplication lives.
Apache-2.0 everywhere except the metadata, and a benchmark nobody mentions
Two loose ends in this repository, one of which is a discrepancy you should resolve before relying on the project and one of which is a question you should ask.
The licence first. The repository's licence field reads as a custom licence that GitHub could not classify, with a note to see the LICENSE file. Everything else in the repository says Apache 2.0. The pyproject.toml sets license to Apache-2.0 and declares license-files as LICENSE*, the README carries an Apache 2.0 badge linking to the Apache licence page, and there is a LICENSE file at the root. The package description on PyPI is built from that same metadata.
So the position is that the intended licence is unambiguously Apache 2.0 and the repository-level field is the odd one out. A classifier that cannot categorise a licence usually means the file has something appended, such as a notice, or a bundled third-party section. The pyproject declaration using the modern SPDX string form and the license-files glob both suggest the project has been attentive to the packaging-standard way of expressing licensing, so the metadata field is more likely stale than meaningful. Read the LICENSE file, confirm it is unmodified Apache 2.0, and the question closes.
The benchmark directory is the more interesting loose end. There is a top-level benchmark/ directory in the repository, alongside build_scripts/, ci/, conf/, datus/, docs/, scripts/ and tests/. For a data engineering agent, a benchmark directory is the single most informative thing that could be in a repository, because the central question about an agent is not whether it runs but whether its answers are right, and that is a measurement problem.
And the README never mentions it. The Features section describes automated semantic modelling, the Dosi execution engine compiling models for thirteen-plus dialects, metric question answering and dimension attribution, built-in subagents for cross-database migration, ETL generation and wide-table builds, report and dashboard generation, nineteen database adapters, ten or more LLM providers and an MCP server and client. The Architecture section describes the three entry points, the subagents and skills, the context engine and the plugins. Not one of them refers to a benchmark, a leaderboard, an accuracy figure or an evaluation harness.
There are ways to read that absence that are not damning. The benchmark may be an internal tool the team runs internally, which is common in agent projects because a public leaderboard invites gaming. The accuracy claim in the README is phrased qualitatively, saying the context engine gathers schemas, reference SQL and business rules and writes every correction back so later answers keep getting more accurate, which is a claim about a mechanism rather than a number. And the package is at 0.4.1 with an Alpha classifier, so a public benchmark may simply be premature.
But for an evaluator it is the first question to ask, and the answers matter. What does the benchmark measure, on which datasets, with which reference answers, and is the result stable across the ten LLM providers the project supports? A data agent whose accuracy depends on which model you point it at, with no published measurement, is a tool you have to evaluate yourself, and the cost of that evaluation is the thing the benchmark directory should have saved you.
Editorial conclusion
Evaluate Datus if your problem is that your warehouse has no trustworthy semantic layer, because automated semantic modelling from schema plus SQL history and a Dosi execution engine that compiles one model into thirteen dialects is a more durable asset than another text-to-SQL prompt, and the context engine that writes every correction back is the mechanism that makes it improve. Do not adopt it as a stable dependency yet, because the package declares itself Alpha at version 0.4.1 with exact pins on roughly twenty of fifty runtime dependencies, so a security patch upstream will not reach you without a release. Do not expect the default install to be one package, because the first-party components are published separately as datus-db-core, datus-storage-base, datus-semantic-core, datus-bi-core and datus-scheduler-core, and you will be pinning five version lines rather than one. Verify five things. Whether the licence is what you think, since the repository metadata says a custom licence that GitHub could not classify while pyproject, the badge and the LICENSE file all say Apache-2.0. What the benchmark/ directory measures, because it is a top-level directory that the README never mentions and for an agent that is the most interesting thing in the repository. Whether LanceDB and Tantivy are both doing retrieval work, since the requirements show two search stacks and the README only describes vector search. How the LLM provider floors interact with your own, given litellm, anthropic and openai-agents are all pinned exactly. And whether the Alpha classifier is a statement about the release process or about the API. The deciding fact is that the engineering discipline around the dependency list is stronger than the discipline around the version promise, and the dependency list is the part you will live with.
Frequently asked questions
What is the Datus data engineering agent?
It is a Python agent that connects a warehouse, catalog, semantic layer and BI through one interface, with three entry points: a CLI for data engineers, a chat interface for analysts on the web, in Slack or Feishu, or in VS Code, and a REST and MCP API for other agents and applications. It handles SQL authoring and validation, semantic model and metric construction, and the generation of pipelines, reports and dashboards, and every run and correction settles into its context.
How does the Datus context engine store and retrieve context?
The README describes metadata, metrics, reference SQL, knowledge and local files, retrieved through business-domain trees plus vector search, with storage on embedded LanceDB and SQLite, and PostgreSQL for teams that share context. The requirements file adds a third mechanism the README does not mention: Tantivy, a full-text search engine, alongside LanceDB for vectors and fastembed for the embeddings themselves.
What is Dosi in the Datus project?
Dosi is a separate semantic layer with its own site, and its execution engine compiles one semantic model into SQL for thirteen or more database dialects. It ships as an independent program that can also be run as a CLI, a REST server or an MCP server, and it defines the OSI semantic model format that Datus's automated semantic modelling generates, with no hand-written YAML.
How many packages make up Datus?
Five first-party ones, all published to PyPI and all pre-1.0: datus-db-core at 0.1.7 or newer, datus-storage-base at 0.1.5 or newer and below 0.2.0, datus-semantic-core at 0.2.4 or newer, datus-bi-core at 0.1.2 or newer and datus-scheduler-core at 0.1.1 or newer. The split follows the integration boundaries, so the semantic layer or the storage layer can be taken without the agent.
What does Datus's requirements.txt do that pyproject.toml does not?
It duplicates the dependency list. The header comment says the file mirrors project.dependencies in pyproject.toml, and that ci/check_release_readiness.py compares the two and fails on any drift, so every edit in one needs the matching edit in the other. The flat list is presumably a deployment and security-review artefact, and it is also where the inline reasoning for several pins lives.
What licence is Datus released under?
Apache 2.0 according to the pyproject.toml, the README badge and the LICENSE file at the root, all of which agree. The repository-level licence metadata is the exception, reading as a custom licence GitHub could not classify, so read the LICENSE file directly to confirm it is unmodified Apache 2.0. The package requires Python 3.12 or newer and declares itself Development Status 3 - Alpha at version 0.4.1.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/datus-ai-datus-agent)