# YOLOv12 is a patched copy of Ultralytics, and the README says so first

> The attention-centric detector from the NeurIPS 2025 paper ships as a vendored fork of the Ultralytics package, which tells you a great deal about what it will cost you to run.

**sunsmarterjie/yolov12** — [NeurIPS 2025] YOLOv12: Attention-Centric Real-Time Object Detectors

- Repository: https://github.com/sunsmarterjie/yolov12
- Stars: 2,961 · Forks: 435
- Language: Python
- License: AGPL-3.0
- Published: 2026-10-07 · Updated: 2026-10-07 · Language: en
- Canonical page: https://hysenlabs.com/projects/sunsmarterjie-yolov12

## The tree is a vendored Ultralytics, not a fresh package

The clearest signal is in the packaging metadata. `pyproject.toml` declares `name = "ultralytics"`, describes itself as Ultralytics YOLO for SOTA object detection, multi-object tracking, instance segmentation, pose estimation and image classification, and carries `license = { "text" = "AGPL-3.0" }`. None of that is boilerplate copied from somewhere else. This repository builds a distribution called ultralytics, from a fork of ultralytics, and the tree confirms it with a top-level `ultralytics/` directory sitting next to `app.py`, `tests/`, `examples/`, `docker/`, `assets/`, `logs/` and an `mkdocs.yml`.

That structure is the single most useful fact for anyone considering the project, and the README states the reason plainly in its first update entry. Dated 2025/06/17, it says to use this repository for YOLOv12 instead of the implementation in the ultralytics package, because that implementation is inefficient, needs more memory, and has unstable training, all of which are described as fixed here. In other words, the architectural contribution and the packaging contribution are the same artifact: you get the paper's model by adopting the authors' fork.

For research reproduction this is a reasonable trade. For a product that already depends on the official Ultralytics package, it is a fork to track, and a conflict waiting to happen the next time you upgrade on the other side. GitHub reports 2961 stars, 435 forks and 125 open issues, with the last push on 2026-09-15.

## Attention-centric means the speed claim is the whole argument

The paper's framing is that the YOLO line has improved through CNN changes while attention mechanisms model better but could not match CNN speed on real-time hardware. YOLOv12's claim is to close that gap: attention modelling at CNN-like throughput. The abstract reports that YOLOv12-N reaches 40.6% mAP at 1.64 ms of inference latency on a T4 GPU, beating YOLOv10-N and YOLOv11-N by 2.1% and 1.2% mAP at comparable speed, and that YOLOv12-S beats RT-DETR-R18 and RT-DETRv2-R18 while running 42% faster, using 36% of the computation and 45% of the parameters.

Read those latency figures as TensorRT figures. The table header names T4 TensorRT10 as the measurement stack, which is a specific combination of an older GPU, a specific TensorRT major version, and an export path. A number obtained that way is a floor for what the architecture can do after compilation, not a prediction for what you will see in eager PyTorch on your own hardware.

The authors are Yunjie Tian and David Doermann at the University at Buffalo, SUNY, with Qixiang Ye at the University of Chinese Academy of Sciences. The paper is arXiv 2502.12524 and was public on 2025/02/19, with Hugging Face Spaces demos listed alongside it.

## Two sets of headline numbers, and they are not the same measurement

Here is the discrepancy worth knowing about before you compare this project against another detector. The abstract gives YOLOv12-N as 40.6% mAP and 1.64 ms. The results table in the README, under the heading Turbo (default), gives YOLO12n as 40.4 mAP and 1.60 ms, at 640 pixels, 2.5 million parameters and 6.0 GFLOPs.

Both are in the repository and neither is marked wrong. The likeliest explanation is that the abstract carries the paper's figures for one variant while the table reports the turbo variant, which was released later, on 2025/03/09 as a faster version of YOLOv12. The table also points to a separate V1.0 branch for the earlier weights, so the two number sets belong to two different artifacts rather than one measurement quoted twice.

What a reader should take from this: do not treat a single mAP number from this README as the specification. Match the variant, the input size and the runtime before comparing anything. The turbo table covers the full ladder, and the download links go straight to the corresponding `.pt` weights in the release assets:

```text
torch==2.2.2
torchvision==0.17.2
timm==1.0.14
albumentations==2.0.4
onnx==1.14.0
```

Those lines come from `requirements.txt`, and the numbers in that file matter at least as much as the table does.

## The dependency pins decide whether your machine qualifies

Requirements are pinned hard, and one of them is not a version at all. Alongside torch 2.2.2, torchvision 0.17.2, timm 1.0.14, albumentations 2.0.4, onnx 1.14.0, onnxruntime 1.15.1, onnxruntime-gpu 1.18.0, gradio 4.44.1, opencv-python 4.9.0.80, numpy 1.26.4, pycocotools 2.0.7, safetensors 0.4.3 and supervision 0.22.0, the file contains a single line with a complete wheel filename:

```text
flash_attn-2.7.3+cu11torch2.2cxx11abiFALSE-cp311-cp311-linux_x86_64.whl
```

Read that as a constraint list rather than a download. It encodes CUDA 11, torch 2.2, the non-CXX11-ABI torch build, CPython 3.11, and Linux on x86-64, all at once. Installing it unmodified on Apple Silicon, on a newer CUDA runtime, or on Python 3.12 does not resolve to a working attention kernel, and `requirements.txt` gives you no marker or environment-variable escape hatch for that case.

The rest of the packaging is more forgiving. `requires-python = ">=3.8"` and a setuptools backend with `setuptools>=70.0.0` mean the Python floor is not the obstacle. The obstacle is the CUDA-flavored wheel. Anyone planning a deployment should treat editing that one line as a required step and budget for it.

## Detection first, then segmentation and classification

The releases show a project that broadened its task coverage over about four months. There are three, and they are terse: YOLOv12-turbo models on 2025-03-08, instance segmentation models on 2025-06-04, and classification models on 2025-07-01. The README update log matches those dates and points segmentation code at a `Seg` directory and classification code at a `Cls` directory, each on its own branch rather than on `main`.

That branch layout is worth pausing on. The default branch is `main` and it carries the detection work, so cloning the repository gives you detection only. Anyone arriving for segmentation or classification has to switch branches, which is not obvious from the tree and is the kind of thing that turns into a support question.

The `docker/` directory in the tree suggests a containerized path for inference and training, and `mkdocs.yml` indicates a documentation site is configured. The README itself is mostly a paper abstract, a results table and a long list of community ports, so the operational detail for training lives in Ultralytics-style configuration and in the documentation rather than on the repository front page.

## The deployment story was built by other people

A striking share of what makes this model usable sits in the README's update log and belongs to third parties. The TensorRT C++ inference repository and Colab notebook from mohamedsamirx, the ncnn-based Android deployment from mpj1234, TensorRT-YOLO from laugh12321, and an ONNX C++ port are all listed as community work rather than first-party support.

The hosted and notebook paths are similarly third-party or partner-led: Roboflow inference and training posts, a Colab notebook, a Kaggle notebook, LightlyTrain support with its own Colab tutorial, an OpenBayes tutorial, and three Hugging Face Spaces demos. Ultralytics and LearnOpenCV both published explanatory posts on 2025/02/24.

Taken together this is the real adoption profile of the model. The paper and the training code are here; the fast inference paths are assembled by the community around exported weights. If your deployment target is not on that list, you are the one writing the port, and the practical starting point is the turbo `.pt` weights and an ONNX export rather than anything in `src/`.

The licensing question deserves its own attention because it outlives all of this. The package is AGPL-3.0, which means network use counts as distribution in the way that MIT or Apache-2.0 does not. For an internal service, read the license before you build on it; for anything a paying customer touches, that decision belongs to whoever handles licensing at your company.

## Conclusion

YOLOv12 is best understood as a set of architecture changes delivered as a patched fork rather than as a standalone library, and the README opens by telling you to prefer it over the official Ultralytics implementation of the same model. Practical evaluation means three checks. Confirm the numbers you were quoted against the table you will actually download, since the paper's figures and the turbo table do not match line for line. Read `requirements.txt` before you build an environment, because the pinned torch version and the hardcoded flash-attention wheel will decide whether your machine is supported. Then settle the AGPL-3.0 question, because that license is the constraint most likely to matter to a commercial deployment. For a single-scale detector on 640-pixel input, the turbo table is the honest starting point.

## FAQ

### What is the YOLO model used for?

Real-time object detection, which is the task YOLOv12 targets in the README results table, where five detection sizes are scored at 640-pixel input on mAP 50-95 and latency. The same Ultralytics codebase this project forks covers adjacent tasks too, including multi-object tracking, instance segmentation, pose estimation and image classification, and YOLOv12 released its own segmentation and classification checkpoints in June and July 2025.

### Which is better, YOLO or SSD?

The repository does not benchmark against SSD at all, so it cannot answer that. Its comparisons are with the generation before it, YOLOv10 and YOLO11, and with the DETR line in the form of RT-DETR and RT-DETRv2, where the paper reports YOLOv12-S running 42% faster with 36% of the computation and 45% of the parameters. Deciding against SSD needs numbers from another source, matched to the same hardware and runtime.

### Why does the README tell me not to use the Ultralytics implementation of YOLOv12?

Because it is a different code path for the same model, and the authors consider it worse. The 2025/06/17 update says the ultralytics package version is inefficient, requires more memory, and has unstable training, all of which they say are fixed in this repository. Practically, adopting this project means adopting their fork of the whole Ultralytics library, not installing a small add-on.

## Sources

- [Issues](https://github.com/sunsmarterjie/yolov12/issues)
- [License: AGPL-3.0](https://github.com/sunsmarterjie/yolov12/blob/main/LICENSE)
- [README](https://github.com/sunsmarterjie/yolov12/blob/main/README.md)
- [Releases](https://github.com/sunsmarterjie/yolov12/releases)
- [sunsmarterjie/yolov12 on GitHub](https://github.com/sunsmarterjie/yolov12)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/sunsmarterjie-yolov12
