Open-source project
datagallery-ai/dataagent avatar
datagallery-ai/dataagent

DataFoundry ships 28 agent limit knobs, no Docker, and one release named for the old name

DataFoundry is an open-source AI workbench for data analysis, unifying data sources, knowledge, tools, and agent runtime into a governed workspace for interactive analytics.

780 stars113 forksTypeScriptApache-2.0

At a glance

What is it?
A self-hosted TypeScript workbench that puts an agent inside a semantic layer and a wall of integer caps. The identity is muddled: the repository is dataagent on a branch named datafoundry, every clone command points at a different owner, and the only published release is DataAgent v0.1.0 while the manifest says 0.2.0.
Who is it for?
The engineering judgment here is better than the packaging story. A read-only boundary, credential isolation, field masking, row limits and timeouts as defaults, with every SQL statement and tool call persisted for replay, is the right set of answers to the four questions the project says actually matter.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 5 days ago.
What is it written in?
Mainly TypeScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 4, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Every clone command targets a different owner than the repository holding it

Follow the Quick Start literally and you leave this project. The clone URL in both install paths is github.com/datagallery-lab/datafoundry.git, while this repository is datagallery-ai/dataagent, on a default branch named datafoundry. The manifest agrees with the clone URL rather than with the repository: its name field is datafoundry, and its repository, homepage and bugs URLs all point at datagallery-lab. The documentation link in the header goes to datagallery-lab.github.io, and the recorded homepage for this repository is a third address, datagallery-ai.github.io/dataagent. So four addresses describe what is arguably one product, and the two a reader is bound to act on, the clone URL and the docs link, both leave the repository they are printed in.

bash
git clone https://github.com/datagallery-lab/datafoundry.git
cd datafoundry
./deploy.sh

The second install path, for native Windows and macOS, opens with the same two lines and continues with npm install and copies of the two example environment files.

The single release is DataAgent v0.1.0 while the manifest and the changelog say 0.2.0

There is one GitHub release, v0.1.0, dated 2026-08-29 and titled DataAgent v0.1.0. The root package.json says 0.2.0, and the README carries a section headed with what is new in v0.2.0, listing branchable concurrent analysis, evidence-first follow-ups referencing a selected table or text region, a semantic trace DAG, an external DataLink integration, reusable outputs and workspace assets, and a production-facing web layer with built-in password authentication. The version, the product name in the release title, and the branch name all moved at different times. There is no 0.2.0 release, so the work described in that section exists as a branch, a docs page at docs/en/releases/v0.2.0.md, and a manifest version number. Anyone installing a tagged release gets a month-old 0.1.0 under the older name.

Formal deploy means do not run npm run dev, and there is no container

The deployment section opens with two constraints that shape everything after it: formal mode has two paths, and npm run dev is not one of them, and Docker and Compose are not provided in this release. That is a notable omission for a product sold on self-hosting, and it is stated rather than hidden. The recommended path is a single script on Ubuntu or Debian, which generates configuration, installs dependencies, builds the web app, the API and the TUI, then starts web plus API as a detached background process. The TUI is built but is a foreground client that does not stay up with the stack, DataLink is external and is not started by the script, and no model key is needed during deploy because a model profile is created in the web UI after login. Native Windows and macOS are not supported by that path.

The management commands assume a process group, and reconfiguration stops the stack

The script exposes a small control surface:

bash
./deploy.sh status
./deploy.sh start
./deploy.sh stop
./deploy.sh logs
./deploy.sh doctor
./deploy.sh tui      # optional: foreground TUI client (API must already be up)

Two operational details are easy to skim past. Re-running the deploy uses a maintenance window, stopping the managed process group before install and build, so a second deploy is not a zero-downtime operation. And changing ports or the public URL goes through a reconfigure flag that preserves secrets and backs up the environment file first, which is the right order of operations. The script also skips its configuration questions entirely when a complete environment file already exists, so a half-configured file from an earlier attempt is the state to check before wondering why the prompts are missing.

The published package omits deploy.sh, deploy/ and services/

The files array in the root manifest is short and deliberate. It ships the environment example, CONTRIBUTING, LICENSE, both READMEs, apps, docs, mkdocs.yml, packages, the two requirements files, scripts, and the two TypeScript configs. Three paths that exist at the root are absent from it: deploy.sh, deploy/ and services/. The consequence is concrete. The one-click deploy story, the management commands above, and anything living in the services directory are reachable only from a git clone, not from the registry. Anyone wiring DataFoundry into an internal builder that installs from npm will find that the deployment experience described in the README simply is not in the tarball they fetched.

npm install executes a postinstall script, and dev has its own pre-hook

The recommended manual path on Windows and macOS is a git clone, npm install, then two copies of example environment files. That npm install runs a postinstall hook, node scripts/postinstall.mjs, before you have read anything. The rest of the script surface is small and worth knowing by name: dev runs node scripts/dev.mjs behind a predev:api hook that runs node scripts/ensure-dev-environment.mjs, start runs node scripts/start.mjs, build and typecheck both drive tsc against tsconfig.build.json, and the deploy tests run node --test over scripts/deploy/*.test.mjs plus scripts/stack-runtime-config.test.mjs. The workspaces are apps/* and packages/*, and the manifest requires node 22 or newer, so an older runtime in your path fails before anything else does.

Safe by default is 28 commented integer ceilings, and two of them bind first

The environment example carries twenty-eight commented caps, and they are the whole safety story in numeric form. The interesting ones: 80 agent steps, 60 SQL executions, 100 protocol actions by default, 500 for data analysis, 100 for a general task, 3 completion rejections, 8 commit retries, 10 for automatic action depth, 20 tables and 50 columns per table for schema inspection, 20 model rows and 20 activity rows for SQL results, 4000 characters of SQL, 500 characters per cell, 500 characters per tool error message, a knowledge top-K of 20, 16 claims and 32 output fields for a requirement commit, 1 step for the model helper, 512 output tokens for the protocol classifier, 4096 for the requirement extractor, 8192 across 2 attempts for the contract grounder, and 12000 characters of tool observation. Context is capped twice, at 32000 tokens and at 32000 characters, so which limit binds depends on the mix of prose and result rows in the context. The file says the authoritative defaults and safe ranges live in packages/agent-runtime/src/config/agent-runtime-limits.ts.

The shipped defaults point at one model vendor and one embedding vendor

The pitch is that any OpenAI-compatible provider works, naming Qwen, DeepSeek and GPT. The example configuration encodes something narrower. LLM_PROVIDER is openai-compatible, LLM_MODEL is qwen-plus, and LLM_BASE_URL is the DashScope compatible-mode endpoint, with deepseek-chat and gpt-4o offered in a trailing comment. On the embedding side the defaults are EMBEDDING_PROVIDER set to bailian, EMBEDDING_MODEL set to text-embedding-v4, and EMBEDDING_DIM set to 1024, with the file cutting off mid-line right after that. A fixed embedding dimension is the kind of value that has to match an existing index, so anyone switching providers should expect to rebuild knowledge rather than just swap a key. The repository also carries a Python requirements file listing numpy, pandas, matplotlib and scikit-learn inside an otherwise TypeScript monorepo.

Editorial conclusion

The engineering judgment here is better than the packaging story. A read-only boundary, credential isolation, field masking, row limits and timeouts as defaults, with every SQL statement and tool call persisted for replay, is the right set of answers to the four questions the project says actually matter. The ceilings are unusually legible too, 28 documented integers you can raise deliberately rather than a vague safety claim. What to settle before deploying is identity and packaging. Three repository names are in play, the clone URL does not match the repository you are reading, the one-click path is a shell script that is excluded from the published package, there is no container image, and 0.2.0 exists only as a branch and a release-notes page. Read the deploy script before running it on a server, decide whether 20 model rows and 4000 SQL characters fit your questions, and point LLM_BASE_URL and EMBEDDING_PROVIDER at your own endpoints rather than the defaults, because the example configuration targets one vendor end to end.

Frequently asked questions

What is DataFoundry built on?

A TypeScript monorepo with apps/* and packages/* workspaces, ES modules, and a node 22 or newer requirement. The root package is named datafoundry at version 0.2.0, and the badges point at Mastra, the AG-UI protocol and Ink.

How do I install DataFoundry?

On Ubuntu or Debian, ./deploy.sh configures, installs, builds web, API and TUI, and starts web plus API detached. Elsewhere, manual npm: git clone, npm install, then copy .env.example to .env and apps/web/.env.example to apps/web/.env.local.

Does DataFoundry provide Docker or Compose files?

Not in this release. The formal deploy section states that Docker and Compose are not provided, and warns against running npm run dev for a formal deployment.

What limits does the DataFoundry environment example expose?

Twenty-eight commented agent runtime caps, including 80 agent steps, 60 SQL executions, 20 model rows, 4000 SQL characters, 500 characters per cell, 20 tables and 50 columns per table, and a 32000 context ceiling in both tokens and characters.

Which model provider does DataFoundry use by default?

An OpenAI-compatible endpoint with LLM_MODEL set to qwen-plus and LLM_BASE_URL pointing at the DashScope compatible-mode endpoint, with deepseek-chat and gpt-4o named as alternatives. Embeddings default to provider bailian, model text-embedding-v4 at 1024 dimensions.

What is DataLink in DataFoundry?

An external component the deploy script does not start. It can be connected later through MCP in the web interface, mapping tables and columns to business concepts, entities, joinable paths and confidence-scored relationships.

Official sources

  1. datagallery-ai/dataagent on GitHub
  2. License: Apache-2.0
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/datagallery-ai-dataagent.svg)](https://hysenlabs.com/projects/datagallery-ai-dataagent)