Model or dataset
CVHub520/X-AnyLabeling avatar
CVHub520/X-AnyLabeling

X-AnyLabeling: the extra you install decides your ONNX Runtime build

X-AnyLabeling: A lightweight, efficient, and unified cross-platform desktop application for annotating text, image, video, and multimodal data, combining versatile built-in tools with state-of-the-art AI models and flexible multi-format export.

10,592 stars1,172 forksPythonGPL-3.0

At a glance

What is it?
X-AnyLabeling is a GPL-3.0 desktop annotation app for text, image, video, point cloud and multimodal data, backed by a long model table and a flat export format list. What shapes your install is the extra you pick, and what it cannot do is reach a remote model without a second repository.
Who is it for?
X-AnyLabeling fits a labeling team that annotates several modalities on one desktop, wants AI-assisted shapes from a broad model table, and is willing to choose its CUDA generation at install time. It does not fit a headless pipeline that needs remote models or a server component, because that lives in X-AnyLabeling-Server.
Can I use it commercially?
Yes, with conditions. GPL-3.0 is a copyleft licence: if you distribute software that includes it, you must release that software's source code under the same licence. Running it internally without distributing it does not trigger that obligation.
Is it still maintained?
Yes. The repository last received commits 3 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 3, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Four extras pick the runtime, and two ONNX packages must not coexist

The build metadata is where the first real decision sits. The build system requires setuptools>=77.0.0 and wheel with the setuptools.build_meta backend, and the distribution is named x-anylabeling-cvhub with a dynamic version. The plain install line is:

bash
pip install "x-anylabeling-cvhub[xxx]"

The bracket is not decoration, because the runtime has to be named there. Four development forms cover the four cases, and they differ only in the extra name:

- CPU is `pip install -e ".[cpu,dev]"` - CUDA 12.x is `pip install -e ".[gpu,dev]"` - CUDA 11.x is `pip install -e ".[gpu-cu11,dev]"` - CUDA 13.x is `pip install -e ".[gpu-cu13,dev]"`

The editable form is the one intended for working on the code, because it allows real-time code modifications without the need for re-installation. What the extras cannot do is detect your hardware for you, so a plain install with no accelerator extra leaves the choice to you. One constraint is stated flatly in the same file: onnxruntime and onnxruntime-gpu must not be installed simultaneously.

The model table names weights, not the engine that runs them

The feature list claims local and remote inference through engines and serving frameworks including ONNX Runtime, TensorRT, OpenCV DNN, vLLM and SGLang. That sentence is the whole story about runtime coverage, and it sits apart from a model table that runs to roughly twenty rows: image classification with YOLOv5-Cls, YOLOv8-Cls, YOLO11-Cls, InternImage and PULC; object detection across YOLOv5/6/7/8/9/10, YOLO11/12/26, YOLOX, YOLO-NAS, D-FINE, DAMO-YOLO, Gold_YOLO, RT-DETR, RF-DETR and DEIMv2; plus rows for instance segmentation, semantic segmentation with U-Net, pose, face, tracking, rotated detection, depth, matting with RMBG 1.4/2.0, tagging with RAM and RAM++, OCR with PP-OCRv4 through PP-OCRv6, document parsing with PaddleOCR-VL, lane detection with CLRNet and counting with CountGD, GeCO and GeCo2.

The table gives no pairing between a model and an engine, and no weights filename. Custom work has its own page at docs/en/custom_model.md, and a wider list lives at docs/en/model_zoo.md, so the engine and weight pairing has to be resolved per model rather than read off the table.

SAM 3 sits in three different rows of the same table

One model name recurring across task categories is where the table gets confusing. The Segment Anything row lists SAM 1/2/3 alongside SAM-HQ, SAM-Med2D, EdgeSAM, EfficientViT-SAM and MobileSAM. The Tracking row lists TrackTrack, Bot-SORT, ByteTrack and SAM2/3-Video. The Grounding row lists Grounding DINO, YOLO-World, YOLOE, SAM 3 and LocateAnything.

So the string SAM 3 can mean a prompted mask, a video tracking component, or a text grounded proposal depending on which row you read it from, and each implies a different label type on the canvas: a mask shape, a track id, or a text query tied to a box. The same pattern shows up elsewhere in the table, where YOLO appears in classification, detection, segmentation, pose, and rotated rows under different suffixes. The consequence for a user is that searching the table by model name is not enough; you have to know the task row first, because the task decides what the output becomes and which shape tool receives it.

Ten formats share one list while examples are split by task

Import and export are covered by a single line: COCO, VOC, YOLO, DOTA, MOT, MASK, PPOCR, MMGD, VLM-R1 and ShareGPT. That list mixes families that belong to different tasks, since PPOCR sits alongside detection formats and VLM-R1 and ShareGPT belong to the vision language side of the app. No per-task mapping is given with the list, so choosing an export format is a decision you make per project rather than a setting the app infers.

The examples tree is organized the opposite way, grouped by task: classification with image-level and shape-level subfolders, detection split into HBB and OBB, segmentation split into instance, binary semantic and multiclass semantic, description split into tagging and captioning, estimation split into face, pose and depth, optical character recognition split into text recognition and key information extraction, then multiple object tracking, counting, grounding, matting, interactive video object segmentation, training and vision language. That layout is the practical answer to the format question, because each task folder shows the shape of the data you would be exporting.

3D point cloud landed the same day as its release

The dated change list records 3D point cloud annotation on 2026-10-01, covering 3D point cloud detection and segmentation, and pointing at docs/en/point_cloud.md. The release tag v4.1.0 carries the same date, 2026-10-01, and the repository was last pushed on 2026-10-01 as well. Point cloud shapes belong to the same toolbox as the rest of the annotation work, with cuboids and points listed among the available shape types.

What is missing is example material. The examples tree listed for the repository covers classification, counting, description, detection, estimation, grounding, interactive video object segmentation, matting, multiple object tracking, optical character_recognition, segmentation, training and vision language, with no point cloud directory among them. So the newest capability is also the one without a worked example, and for a reader deciding whether to adopt it that gap matters: the docs page is the only place described that shows what a finished point cloud annotation looks like.

The Linux desktop entry sits next to packaging and tool directories

The repository root carries a Linux desktop entry file named x-anylabeling.desktop, alongside MANIFEST.in, a packaging directory, a scripts directory, a tools directory and a tests directory. The app is described as running on Windows, Linux and macOS, with interfaces available in English, Simplified Chinese, Japanese and Korean, and the root holds both README.md and README_zh-CN.md.

Contributor tooling sits next to that. The tree includes .flake8, .pre-commit-config.yaml, .vscode/, CITATION.cff, CLA.md, CONTRIBUTING.md, SECURITY.md and a CHANGELOG.md, and the tests directory is the target of the test script defined in the root package metadata of a project of this shape. What the root does not do is tell you how any of this becomes a per-platform build: no packaging commands appear there, and the installation steps are deferred to the Installation & Quickstart page at docs/en/get_started.md. Desktop entry file, packaging directory and a missing set of documented build commands means platform packaging is real work you schedule, not a one line install.

Releases land weekly, so the version you pin sets your bug surface

Three releases appear close together: v4.0.5 on 2026-08-28, v4.0.6 on 2026-09-05 and v4.1.0 on 2026-10-01. The change list moves at the same pace. Image tagging with creation, editing, reordering and batch deletion arrived on 2026-08-19, D-FINE-seg instance segmentation on 2026-08-12, and on 2026-08-08 both the RT-DETRv2-OBB rotated object detection model and a Magic Wand tool for creating polygons from contiguous color regions, with v4.0.0 released on 2026-08-05.

The consequence for a reader is about pinning rather than about features. A repository that pushes and tags weekly keeps a gap between what main holds and what your installed build holds, and the change list itself defers the detail elsewhere, telling readers to refer to the CHANGELOG for more. So a fix or a model entry you read in the tree may sit in a commit that no tag contains yet, and a behavior you rely on may change under you on the next monthly tag. Track the CHANGELOG alongside the version rather than treating either alone as the state of the app.

Remote inference is a separate repository, not a switch

Remote work is delegated rather than included. For remote inference, X-AnyLabeling-Server is named as a lightweight, extensible backend for connecting custom models and compute resources, and it is a different repository with its own documentation entry, listed first in the docs index before Installation & Quickstart and Usage. The desktop app is the annotator; the server is where custom models and compute resources get connected.

That division has a direct consequence. The desktop application cannot reach a remote model on its own, so a team whose models live behind a serving stack has two things to deploy and two things to keep aligned, even though vLLM and SGLang are named among the serving frameworks the app supports. It also means the plugin story splits: local model integration runs through docs/en/custom_model.md inside this repository, while anything routed through remote compute goes through the server project instead.

Editorial conclusion

X-AnyLabeling fits a labeling team that annotates several modalities on one desktop, wants AI-assisted shapes from a broad model table, and is willing to choose its CUDA generation at install time. It does not fit a headless pipeline that needs remote models or a server component, because that lives in X-AnyLabeling-Server. Before committing, confirm three things: the extra matches your CUDA version, onnxruntime and onnxruntime-gpu are not both installed, and the engine backing your chosen model is listed among ONNX Runtime, TensorRT, OpenCV DNN, vLLM and SGLang.

Frequently asked questions

what is x anylabeling

X-AnyLabeling is a lightweight, unified cross-platform desktop application for annotating text, image, video, point cloud and multimodal data, combining built-in tools with deep learning models and multi-format import and export. Its project page is xanylabeling.com.

how to install x anylabeling

The distribution is x-anylabeling-cvhub and the install line is `pip install "x-anylabeling-cvhub[xxx]"`, where the extra picks the runtime: cpu, gpu for CUDA 12.x, gpu-cu11 for CUDA 11.x, or gpu-cu13 for CUDA 13.x. Development installs add the dev extra, and onnxruntime and onnxruntime-gpu must not be installed at the same time.

x anylabeling how to use

The docs hold a Usage page at docs/en/user_guide.md and a Command Line Interface page at docs/en/cli.md, with separate pages for the Chatbot, VQA, Image Classifier, Video Classifier, document parsing and 3D point cloud annotation. Installation and quickstart live at docs/en/get_started.md, and each task has an examples folder.

x anylabeling vs cvat

The README does not make that comparison. It describes X-AnyLabeling as a cross-platform desktop application with interfaces in English, Simplified Chinese, Japanese and Korean, and for remote inference it points to X-AnyLabeling-Server as a separate backend.

Official sources

  1. Official documentation
  2. Official README
  3. Project repository
  4. Release notes
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/cvhub520-x-anylabeling.svg)](https://hysenlabs.com/projects/cvhub520-x-anylabeling)