Self-hosted service
ModelEngine-Group/DataMate avatar
ModelEngine-Group/DataMate

DataMate: a self-hosted data work platform for fine-tuning and RAG pipelines

DataMate is an enterprise-level data processing platform designed for model fine-tuning and RAG retrieval.

368 stars46 forksTypeScriptMIT

At a glance

What is it?
DataMate bundles collection, cleaning, synthesis, annotation, evaluation and knowledge generation behind a drag-and-drop workflow builder. It installs as a Docker Compose or Helm stack, and the interesting question is whether that breadth fits your team.
Who is it for?
Adopt DataMate if you already run Kubernetes or Docker Compose and want collection, cleaning, synthesis, annotation, evaluation and knowledge generation in one self-hosted place, with Label Studio and Mineru available as optional add-ons through the Makefile. Do not adopt it if you need a hosted service with published pricing, or if you cannot operate a multi-service stack.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 34 days ago.
What is it written in?
Mainly TypeScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What DataMate is for, and who ends up running it

The README describes DataMate as an enterprise-level data processing platform for model fine-tuning and RAG retrieval. The list of core modules is the clearest statement of intent: data collection, data management, an operator marketplace, data cleaning, data synthesis, data annotation, data evaluation and knowledge generation. Those are the stages a team walks through between raw documents and a training set or a retrieval index, and DataMate puts them in one application instead of leaving them spread across notebooks and shell scripts.

The audience is implied by the deployment options rather than stated outright. Docker Compose is offered for a single machine, and Kubernetes with Helm is offered for a cluster, with Sealed Secrets required on the Kubernetes path. That is a platform team's toolkit, not a laptop's. If your workflow is one person and one CSV file, the operator marketplace and the annotation module are weight you will not use. If three teams are each building their own cleaning scripts and nobody can reproduce last quarter's dataset, the shared surface starts to pay for itself.

How the pieces fit: a TypeScript backend, a Python runtime, and Ray

The repository layout separates concerns in a way worth understanding before you deploy. The primary language is TypeScript, and backend/ holds an api-gateway plus a main-application service, with shared libraries split into domain-common and security-common. The runtime/ directory is a separate world: a FastAPI service under runtime/datamate-python, a Ray executor framework under runtime/python-executor, the operator ecosystem under runtime/ops, and something called a DataX data frame under runtime/datax.

So the flow the layout implies is: the frontend talks to the API gateway, the gateway routes to the main application, and execution of actual data work is handed to the Python runtime, where Ray distributes it across workers. Operators are the unit of work in that runtime, which is why the marketplace exists as a first-class concept rather than a plugin folder. The practical consequence is that a custom operator is a Python-side artifact, even though the platform's own services are TypeScript. If your team is Java-only or Go-only, the extension path runs through a language you may not have.

The README does not document the wire format between gateway and runtime, nor the scheduling policy Ray uses for a given operator. Those are things you would read from backend/api-gateway/README.md and runtime/python-executor/README.md in the repository.

Installing DataMate with Docker Compose and opening the first workflow

The prerequisites are Git, Make, Docker and Docker-Compose; add Kubernetes and Helm only for the cluster path. The README also offers a one-liner that skips cloning entirely, pulling the Compose file straight from the repository and starting the stack against the published images. The REGISTRY variable selects ghcr.io/modelengine-group/ as the image source.

bash
wget -qO docker-compose.yml https://raw.githubusercontent.com/ModelEngine-Group/DataMate/refs/heads/main/deployment/docker/datamate/docker-compose.yml \
 && REGISTRY=ghcr.io/modelengine-group/ docker compose up -d

If you prefer the source tree, clone it and run the installer. The Makefile prompts for the deployment method and expects a number: 1 for Docker/Docker-Compose, 2 for Kubernetes/Helm.

bash
git clone [email protected]:ModelEngine-Group/DataMate.git
cd DataMate
make install

The README notes that if Make is unavailable, the equivalent Compose invocation uses the milvus profile explicitly, which is what brings up the vector store alongside the application.

bash
REGISTRY=ghcr.io/modelengine-group/ docker compose -f deployment/docker/datamate/docker-compose.yml --profile milvus up -d

Once the container is running, the frontend is served at http://localhost:30000. That is the point at which you would build your first drag-and-drop workflow. Two optional add-ons install through their own targets: make install-label-studio brings up the annotation tool, and make build-mineru followed by make install-mineru adds enhanced PDF processing. To see every target and flag the Makefile accepts, run make help; the README points there rather than enumerating them.

The Kubernetes path and why Sealed Secrets is not optional there

On Kubernetes, DataMate takes a firm position on secrets. Database passwords and JWT secrets are stored encrypted in Git under deployment/kubernetes/sealed-secrets/ and decrypted in-cluster by the Bitnami Sealed Secrets Controller at deploy time. The README states plainly that the controller is required for this path. That is a real operational dependency: you cannot install the Helm chart on a bare cluster and fill in secrets afterwards through a values file.

bash
helm repo add sealed-secrets https://bitnami-labs.github.io/sealed-secrets
helm install sealed-secrets sealed-secrets/sealed-secrets -n kube-system

For air-gapped environments the README walks through pulling bitnami/sealed-secrets-controller:latest, saving it with docker save to a tar file, and installing kubeseal separately (brew on macOS, a wget of the linux-amd64 binary on Linux). Rotating a password means re-encrypting with kubeseal against the namespace and scope the secret was sealed for.

bash
echo -n "new-password" | kubeseal --raw --name datamate-conf --namespace datamate --scope namespace-wide

The Docker path is deliberately different: secrets live in a .env file that .gitignore excludes. That is a lower bar to entry and a weaker guarantee, since nothing stops a .env from being copied into a ticket. The README does not describe a migration path between the two secret models, so a team that starts on Compose and later moves to Kubernetes should expect to reseal every secret by hand.

Where DataMate is the wrong tool

The Makefile's own help text is the most honest limitation. It documents that make install INSTALLER=k8s requires the Sealed Secrets Controller, which means the cluster path has a prerequisite that a plain helm install does not. A team without cluster admin rights cannot complete that path at all.

The second limitation is surface area. A deployment brings up the DataMate application, Milvus, and optionally Label Studio and Mineru. The uninstall target confirms the shape of the stack: make uninstall prompts once whether to delete volumes, applies that single answer to every component, and tears down in the order milvus, then label-studio, then datamate, so the datamate network is removed after the services using it have stopped. A single volume-retention answer for the whole stack is convenient until you want to keep your Milvus index and discard annotation drafts, or the reverse.

Third, the project is a platform, not a library. If your problem is one cleaning pass over a few thousand rows, a script and a dataframe will finish before DataMate's containers are healthy. The operator marketplace and the visual workflow builder are worth their cost when several people share the same pipeline definition, and not before. The README also does not document rollback of a deployment or a dataset version, so treat reproducibility as something you build on top rather than something the platform hands you.

DataMate against a plain script-and-store setup

The obvious alternative is not another platform but the absence of one: pandas or Spark for cleaning, a labeling vendor or an in-house tool for annotation, MLflow or a folder of JSONL files for tracking, and a vector database such as Milvus on its own. That stack is smaller and every piece is independently replaceable. Its weakness is the seams. The cleaning logic lives in one repository, the annotation export in another, and the evaluation results in a third, and nobody can point to the single definition that produced the training set.

DataMate's answer is to make the pipeline itself the artifact, with operators as the reusable unit and a visual editor as the authoring surface. You give up the freedom to swap any single component for an unrelated one, since the operator ecosystem and the Ray executor are the execution model. In exchange, the cleaning, synthesis, annotation and evaluation stages share a data-management layer and a common workflow definition. Whether that trade is good depends on how many people need to read the pipeline. One engineer does not need it. Four engineers across two teams probably do.

Maintenance, upgrades and the MIT licence

The last push to the default branch was on 2026-08-27, and the repository is not archived. The most recent tagged release is v1.0.2 from 2026-04-29, preceded by v1.0.1 on 2026-04-23 and v1.0.0 on 2026-04-01. The gap between the April release and the August push means main carries changes that are not in a tag, so pinning to v1.0.2 and pinning to main are different bets. The Makefile accepts VERSION=<version> with a default of latest, which is the knob that decides which bet you are making; leaving it at latest means every install can pull a different image.

Upgrades are image swaps on both paths, with the caveat that the Kubernetes route re-reads sealed secrets at deploy time, so a secret rotation and an application upgrade are separate operations. The README does not document a database migration procedure or a downgrade path, which is the thing to test on a staging cluster before touching production.

The licence is MIT. That permits commercial use and modification, and it comes with no warranty. The README does not discuss the licences of the bundled components, and the deployment pulls Milvus, Label Studio and Mineru images alongside the DataMate images. If your organization reviews third-party licences, that list is the place to look, not the DataMate LICENSE file alone. Nothing here is legal advice.

Editorial conclusion

Adopt DataMate if you already run Kubernetes or Docker Compose and want collection, cleaning, synthesis, annotation, evaluation and knowledge generation in one self-hosted place, with Label Studio and Mineru available as optional add-ons through the Makefile. Do not adopt it if you need a hosted service with published pricing, or if you cannot operate a multi-service stack. Before committing, run make help to confirm the available targets, check that the ghcr.io/modelengine-group/ images resolve for your platform, and confirm that port 30000 is free on the host.

Frequently asked questions

What is DataMate used for?

The README describes it as an enterprise-level data processing platform for model fine-tuning and RAG retrieval. Its core modules cover data collection, data management, an operator marketplace, data cleaning, data synthesis, data annotation, data evaluation and knowledge generation.

What kind of reports can DataMate generate?

The README lists data evaluation as a core module but does not describe report formats or outputs. The repository's docs/ directory and the runtime documentation are where that detail would live.

What industries use DataMate?

The README does not name any industries. It positions the project for teams doing model fine-tuning and RAG retrieval, and the deployment options assume Docker Compose or a Kubernetes cluster rather than a particular sector.

What are DataMate's pricing plans?

There are none to describe. DataMate is released under the MIT licence and deployed from source or from published container images, so there is no pricing page in the repository.

What is DataMate?

It is a self-hosted, MIT-licensed data processing platform from ModelEngine-Group, written primarily in TypeScript with a Python runtime. The README frames it around model fine-tuning and RAG retrieval, with a drag-and-drop workflow designer and an operator marketplace.

Official sources

  1. License: MIT
  2. ModelEngine-Group/DataMate on GitHub
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/modelengine-group-datamate.svg)](https://hysenlabs.com/projects/modelengine-group-datamate)