# Diffgram: Self-Hosted AI Data Labeling and Annotation Platform

> Diffgram is a self-hosted platform that combines data labeling for multiple media types, an AI datastore for schemas and predictions, and a human supervision workflow. It deploys via Docker Compose and uses a custom commercial open-source license rather than a permissive one.

**diffgram/diffgram** — The AI Datastore for Schemas, BLOBs, and Predictions. Use with your apps or integrate built-in Human Supervision, Data Workflow, and UI Catalog to get the most value out of your AI Data.

- Repository: https://github.com/diffgram/diffgram
- Website: https://diffgram.com
- Stars: 1,910 · Forks: 132
- Language: Python
- License: NOASSERTION
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/diffgram-diffgram

## What Diffgram is and who it targets

Diffgram occupies the space between a pure annotation tool and a full AI data management platform. The README describes it as an AI Datastore for Schemas, BLOBs, and Predictions. A team can use it to label raw data through its UI, store model predictions alongside those labels, and manage data flow between annotation tasks and AI training pipelines, all within one self-hosted system.

The platform is installed and operated by the team using it. Diffgram does not run as a hosted SaaS by default; you own the deployment and the data that goes into it. That design choice makes it relevant to teams with data residency or compliance requirements that prevent sending labeled data to third-party cloud annotation services.

Primary users are teams building AI applications that require large volumes of labeled training data across different media types, ML engineers who need to manage human labeling workflows at scale, and organizations that need an audit trail for their annotation process.

## Media types supported for annotation

The README lists the following annotation types:

- Grid and Multi-Modal
- Conversational and LLM (listed as Preview)
- Image
- Video
- 3D
- Text
- Audio
- GeoSpatial
- Document (listed as Roadmap)
- HTML (listed as Roadmap)
- DICOM (listed as Roadmap)
- Custom

The Conversational and LLM type is noted as a preview, meaning it is not yet production-stable. Document, HTML, and DICOM are listed as roadmap items and are not currently available. The coverage across image, video, 3D, text, and audio is the production-ready set.

This breadth distinguishes Diffgram from tools that focus on a single modality. A team that needs to label both images and audio recordings for a multimodal model does not have to operate two separate annotation systems.

## Deployment: a multi-service Docker Compose stack

Diffgram does not install as a single binary or a pip package. It deploys as a set of Docker containers orchestrated by a docker-compose.yaml. The compose file defines the following services:

- dispatcher: routes traffic, exposes port 8085
- frontend: the web UI, serves on port 8081
- default: the main API service, with a health check on /api/status
- walrus: a background processing service on port 8082
- eventhandlers: an async event processor on port 8086
- db: the database, required to pass a health check before dependent services start
- rabbitmq: the message queue, also health-checked before eventhandlers and walrus start
- minio: object storage

The docker-compose.yaml also mounts a GCP service account file from a path set by the GCP_SERVICE_ACCOUNT_FILE_PATH environment variable. Running without a GCP service account file requires adjusting that volume mapping.

To develop locally, the Makefile provides targets for setting up a Python virtualenv, starting background services, and running each service process separately:

```bash
make run-bg-services
make run-default
make run-walrus
```

The Makefile also covers frontend dependencies and end-to-end tests through Cypress.

## The Diffgram license and what it means in practice

Diffgram does not use a standard open-source license. The repository reports NOASSERTION as its SPDX license identifier, and the README links to the Diffgram license version 2 (DLv2) published in September 2023. The README also mentions a contributor license (CL) available at no financial cost for contributions, and notes that MSA (Master Services Agreement) customers receive a financial credit for contributions.

This is a custom commercial open-source license. The terms differ from MIT, Apache 2.0, or GPL, which are the common baseline licenses in the annotation tool space. Anyone planning to use Diffgram in a commercial product or deploy it for customers should read the full DLv2 text at the link in the README rather than assuming standard open-source permissions apply.

The README also states that commercial firms have been using Diffgram since 2018, which indicates the project has a production track record, but the licensing model means the terms of that use are governed by DLv2 rather than a permissive open-source license.

## Limitations and operational considerations

The multi-service architecture is a significant operational burden compared to tools that run as a single process. Running Diffgram in production requires managing a database, a message broker, an object storage service, and multiple application services, each with its own configuration, health checks, and update cycle. The Makefile and docker-compose.yaml provide the scaffolding, but ongoing maintenance falls on the team.

The most recent GitHub release is 1.25.3, dated October 2024. The repository's last push was on 2026-06-22, which indicates active development beyond what the release tags reflect. This gap means there is no straightforward way to know the current state of unreleased changes without reading commit history.

The Conversational and LLM annotation type is in preview. Document, HTML, and DICOM annotation are on the roadmap. Teams that specifically need those types cannot rely on Diffgram today for them.

The requirements.txt pins specific versions for boto3, google-cloud-storage, and azure-storage-blob. Pinned versions in a general requirements file can cause dependency conflicts in environments that also need other packages.

## Comparison with CVAT

CVAT (Computer Vision Annotation Tool) is an open-source annotation tool originally developed at Intel and now maintained independently. It is licensed under MIT and focuses primarily on image and video annotation for computer vision tasks, with support for bounding boxes, polygons, keypoints, and similar shapes.

Diffgram takes a broader scope: it includes annotation across more media types including audio, text, and 3D, and adds datastore and prediction management features that CVAT does not provide. CVAT is the simpler choice for teams that only need image and video annotation under a permissive license. Diffgram is the more complete choice when annotation is one part of a larger AI data pipeline that also needs schema management, prediction storage, and workflow control. The licensing cost and operational overhead are correspondingly higher with Diffgram.

## Maintenance and related projects

The repository is not archived. The last push was on 2026-06-22. The README includes news entries for two new open-source projects from the team behind Diffgram: dos-kernel, a trust verification tool for AI agent outputs, and fak, a capability gate for AI agents. Neither is part of Diffgram itself; they are separate repositories noted in the README as related work.

The license is Diffgram license version 2 (DLv2), not a standard open-source license. The full text is linked from the README and from the repository's LICENSE.md file.

## Conclusion

Diffgram suits teams that need a self-hosted, end-to-end platform covering annotation across multiple media types, workflow management, and an AI datastore in one system. It is a poor fit for individual developers or small teams that need a quick open-source annotation tool under Apache or MIT, because the Diffgram license version 2 (DLv2) is a custom commercial license and should be read carefully before production use. Before deploying, verify that all services in the docker-compose.yaml, including the database, RabbitMQ, and MinIO, can run in your infrastructure, and confirm which version tag you are deploying since the latest GitHub release is 1.25.3 from October 2024 while the repository has continued to receive commits since then.

## FAQ

### What is Diffgram?

Diffgram is a self-hosted platform for AI data labeling and management. It covers annotation of image, video, 3D, text, and audio data, and includes a datastore for schemas, binary data, and model predictions alongside a human supervision workflow.

### Does Diffgram support audio and 3D annotation?

Yes. The README lists audio and 3D annotation as supported media types alongside image, video, text, and geospatial data. Multi-modal and conversational annotation are also listed, with the conversational type noted as a preview.

### What license does Diffgram use?

Diffgram uses the Diffgram license version 2 (DLv2), a custom commercial open-source license. It is not MIT, Apache 2.0, or GPL. The README links to the full license text and a separate contributor license for those who contribute to the project.

## Sources

- [diffgram/diffgram on GitHub](https://github.com/diffgram/diffgram)
- [Issues](https://github.com/diffgram/diffgram/issues)
- [Project website](https://diffgram.com)
- [README](https://github.com/diffgram/diffgram/blob/master/README.md)
- [Releases](https://github.com/diffgram/diffgram/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/diffgram-diffgram
