Open-source project
opendatalab/labelU avatar
opendatalab/labelU

LabelU: multimodal annotation where the AI assist is a server you host

Open-source multimodal data annotation platform with AI auto-annotation support.

1,683 stars187 forksPythonApache-2.0

At a glance

What is it?
OpenDataLab's Python annotation platform covers image, video and audio, and now bolts on Florence-2, GroundingDINO and SAM through a separate model server you run yourself.
Who is it for?
LabelU is a solid choice when your annotation needs span images, video and audio under one interface and you would rather not send data to a hosted service. The Python side is unremarkable in the best way: FastAPI, SQLAlchemy with Alembic migrations, SQLite by default and MySQL when you outgrow it, and a real exporter that writes the formats model training actually consumes.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 73 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 8, 2026, and from our analysis. They are not legal advice.

Editorial analysis

One interface for three modalities

LabelU is a data annotation platform from OpenDataLab, written in Python, Apache-2.0 licensed, and sitting at roughly 1,675 stars with 182 forks and 32 open issues. The repository was last pushed on 2026-07-28, which is recent enough to call active work, and it is not archived.

The scope is broader than most annotation tools. For images it provides two-dimensional bounding boxes, semantic segmentation, polylines and keypoints. For video it supports segmentation, classification and information extraction, with the claim that long-duration footage is handled. For audio it offers segmentation, classification and information extraction, visualising the signal rather than treating audio as a file with a start and end. The stated use cases are object detection, scene analysis and action recognition, which are the standard reasons to buy annotation software in the first place.

What the repository does not do is train anything. Its output targets are data formats rather than models, and the export list in the README is JSON, COCO and MASK. That boundary is what makes the tool interesting to evaluate, since it can sit in front of a training pipeline without being coupled to one.

One structural detail worth knowing up front: the annotation user interface does not live in this repository. LabelU is built on a separate JavaScript kit, LabelU-Kit, and this repository is the Python backend that serves it. When you clone for local development, a script pulls the frontend statics in from that other repository.

AI annotation is a server you run

The feature that separates this from an ordinary annotation tool is automatic annotation of image data, triggered from a button on the annotation page, with batch processing for a whole task and progress tracking as it goes. The README describes three reference model servers as provided out of the box, and then points at a separate `model_server/README.md` for setup instructions.

Those two statements sit in mild tension, and the sensible reading is that the server code is included while the models still have to be configured and started. What is genuinely useful is the specification of what each option costs. Florence-2 is described as lightweight and CPU-friendly at roughly 4GB of VRAM. GroundingDINO paired with SAM ViT-B targets higher quality detection and segmentation at roughly the same 4GB. SAM 3 is described as a state-of-the-art unified model needing roughly 8GB of VRAM and a high-end GPU.

The three options form a real decision rather than a menu of equivalents. The first two cost about the same memory and differ in quality, while the third roughly doubles the requirement. If you are annotating on a workstation without a serious GPU, Florence-2 or the GroundingDINO plus SAM combination are the only realistic starting points, and you should measure throughput on your own hardware before committing, because the README gives memory figures and not speed figures.

There is also a separate feature from auto-annotation: one-click loading of pre-annotated data that you then adjust. That is the more valuable one for a team with existing labels, because importing and correcting is a different and much safer workflow than generating from nothing.

Storage, import and the formats it exports

Annotation data can be imported directly from S3-compatible object storage, which the README names as AWS S3 and MinIO. You configure data source connections in the task settings, browse and preview files, then import either selected files or everything under a path in one click. The dependency list confirms this is real rather than aspirational, since boto3 is a direct dependency of the package.

That matters more than it first appears. A labeling task is usually an ongoing arrangement where raw media arrives continuously from a capture pipeline, and pointing the tool at object storage instead of uploading files one at a time removes a step that would otherwise dominate the workflow. It also means the annotation tool does not need to be the system of record for your media.

The export side is the other half of the contract, and it is where a data tool earns its keep. JSON, COCO and MASK cover detection, generic structured annotation and masks. The dependency list adds tfrecord, which is a TensorFlow-native format, so there is a path to TensorFlow pipelines without an external converter. The documentation site is the reference for the actual schema, and the README links to the format documentation directly.

Websockets appear as a dependency, which is consistent with a UI that receives progress updates from long-running batch operations such as batch annotation rather than blocking on a request.

Version lines that do not line up

There are three published releases and they describe a project with more than one version line running in parallel, which is worth understanding before you pin anything.

The newest is v1.5.6, published on 2026-07-28. Its release note is a single bug fix, updating a version to v5.12.0 against issue 289. That number is not a LabelU backend version: the package version in `pyproject.toml` is 1.5.6, matching the tag. A jump to v5.12.0 in a 1.5.6 release almost certainly refers to the separate frontend kit, which has its own much higher version series. So there are two independent version lines in play, the Python backend and the JavaScript kit, and the release notes only mention the one that changed.

The other two releases are v1.5.5 from 2026-07-16 and v1.3.0-alpha.20 from 2026-07-15, and both cite the same issue, 284. Two different version lines closing the same issue on consecutive days means the stable branch and an alpha line were both carrying fixes for one report. That is normal for a project with a stable track and a preview track, but it means the alpha tag is not simply an older stable release and should not be treated as one.

The root also carries a `.VERSION` file, which is a fourth place version information can live, and a `.releaserc.json` indicating releases are automated rather than hand-tagged. When the changelog is this thin, reading the commit history is more useful than reading the tags.

Installing it, and the database question

Installation is a Conda environment, one pip command and a launcher. Python 3.11 is the floor: the project metadata declares a minimum of 3.11, and the container image builds on a 3.11 slim base.

bash
conda create -n labelu python=3.11

After activating that environment, the package installs from PyPI and starts with a single command:

bash
pip install labelu

MySQL support is an optional extra rather than a separate install, which is the cleaner arrangement:

bash
pip install labelu[mysql]

The local development path is more modern and worth copying if you plan to contribute. It clones the repository, syncs dependencies with uv, copies the example environment file, and runs the app under uvicorn with reload. One step in that sequence is unusual and important: a shell script downloads the frontend statics from the LabelU-Kit repository, because the UI is not in this repo. A developer who skips it will get a running API and no interface.

The database story is simple at the start and has one sharp edge. By default the application uses SQLite through a connection string pointing at a local file, which needs no server. If you are upgrading from the 1.x line and had moved to MySQL, the README gives a dedicated migration command that copies data from the built-in SQLite database into MySQL, with the connection URL passed in as an environment variable. What is not documented is a general upgrade path for later versions, so if you are planning a multi-step upgrade with a populated database, confirm the current procedure with the maintainers rather than assuming the 1.x command still applies.

The container image has a detail worth knowing

The Dockerfile is four lines of real configuration inside a slim Python 3.11 base, and one of them deserves a second look:

code
CMD ["sh", "-c", "labelu --host=0.0.0.0 --media-host=$MEDIA_HOST"]

The command binds to all interfaces and takes the media host from an environment variable that the image sets to a localhost URL by default. In a container, a localhost reference points back at the container itself, so running the image as-is serves media that your browser cannot reach. You have to override that variable with a URL the browser can resolve, which is normal for an image that ships a localhost default, but it is the first thing that will bite you on a first deployment.

The install step is the second detail. The pip invocation points at the PyPI test index first and adds the production index as an extra index. In other words the image resolves packages preferring pre-release artifacts unless they are absent from production PyPI. For an internal deployment you may prefer to pin known versions, and at minimum you should be aware that the image is not built to resolve only stable releases.

The application itself is a conventional stack and that conventionality is a positive. FastAPI serves it, Uvicorn runs it, SQLAlchemy handles persistence with Alembic for migrations, Pydantic and pydantic-settings handle configuration, python-jose issues JSON web tokens, and typer provides the command line interface through a console script entry point. bcrypt is pinned to an exact version rather than a range, which is the kind of detail that reflects a real dependency conflict having been worked around.

Where it fits

The strongest case for LabelU is a team annotating mixed media under data residency constraints. Self-hosting the backend and the model server means images, video and audio never leave your infrastructure, and the S3 import path means you can point the tool at existing object storage instead of migrating data into an annotation vendor. For teams that also want the convenience of a hosted trial, the README links to an online instance and to the standalone annotation toolkit.

The weaker case is a small image-only project. Bounding boxes and segmentation are widely available, the model requirements are real, and the operation of batch annotation plus human correction plus export to the exact schema your pipeline expects is work that a purpose-built tool for your modality may do better. The export formats are documented rather than guaranteed, so verify the schema against your training code before annotating at volume.

The AI feature deserves a specific caution rather than a general one. Automatic detection and segmentation are excellent at producing a first pass over many thousands of images and unreliable at producing labels you can ship without review. Budget the review, sample the output, and treat the annotation volume you can verify rather than the volume the model produced as your throughput. The same applies to the pre-annotated import path, which is safer because a human produced the original labels.

For a final check on project state: 32 open issues on a repository this size is modest, the release notes for the last three versions are one-line fixes rather than feature announcements, and the citation in the README points at the OpenDataLab paper rather than at LabelU itself. LabelU is a maintained, unpretentious tool in a well-supported ecosystem. Check the documentation site for the schema before you commit annotators to it.

Editorial conclusion

LabelU is a solid choice when your annotation needs span images, video and audio under one interface and you would rather not send data to a hosted service. The Python side is unremarkable in the best way: FastAPI, SQLAlchemy with Alembic migrations, SQLite by default and MySQL when you outgrow it, and a real exporter that writes the formats model training actually consumes. The AI auto-annotation is the interesting part and also the part that needs care, because the three reference models ship as a separate server you configure yourself and the VRAM figures range from a CPU-friendly four gigabytes to eight gigabytes for SAM 3. My practical advice is to treat AI annotation as a first-pass generator rather than a labeler, budget human review time for every batch it produces, and pin the version you install. It was last pushed on 2026-07-28 and the published release notes are thin, so the commit history is more informative than the changelog.

Frequently asked questions

What is LabelU?

It is an open-source multimodal data annotation platform from OpenDataLab, written in Python and licensed under Apache-2.0. It covers image annotation with bounding boxes, semantic segmentation, polylines and keypoints, video segmentation and classification, and audio segmentation and classification, and it exports to JSON, COCO and MASK.

How do I install and run LabelU locally?

Create a Python 3.11 environment with Conda, activate it, then run `pip install labelu` and start the app with `labelu`. The interface is served at localhost on port 8000. MySQL support is an optional extra installed with `pip install labelu[mysql]`, and the default database is SQLite so no database server is needed to start.

What does LabelU need for AI auto-annotation?

It connects to model servers you configure yourself through the `model_server` directory in the repository. Three reference options are described: Florence-2 at roughly 4GB of VRAM, GroundingDINO with SAM ViT-B at roughly 4GB, and SAM 3 at roughly 8GB on a high-end GPU. Setup instructions live in `model_server/README.md` rather than in the main README.

Can LabelU import data from object storage?

Yes. It imports annotation data directly from S3-compatible storage such as AWS S3 or MinIO. You configure a data source connection in the task settings, browse and preview files, then import selected files or everything under a path in one click.

Why does the LabelU repository not contain the web interface?

The annotation user interface lives in a separate repository, LabelU-Kit, which is the JavaScript kit the Python backend serves. For local development a script downloads the frontend statics from that repository, which is why a fresh clone needs that step before the interface appears.

Official sources

  1. License: Apache-2.0
  2. opendatalab/labelU on GitHub
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/opendatalab-labelu.svg)](https://hysenlabs.com/projects/opendatalab-labelu)