Open-source project
Peterande/D-FINE avatar
Peterande/D-FINE

D-FINE: a real-time DETR that turns box regression into distribution refinement

D-FINE: Redefine Regression Task of DETRs as Fine-grained Distribution Refinement [ICLR 2025 Spotlight]

3,348 stars326 forksPythonApache-2.0

At a glance

What is it?
D-FINE is the official implementation of the ICLR 2025 spotlight paper on Fine-grained Distribution Refinement, aimed at engineers who need a COCO-trained, real-time object detector they can fine-tune on their own data. It is a research codebase with a Python install path, not a packaged library.
Who is it for?
Adopt D-FINE if you need a COCO-pretrained real-time detector whose regression head you can read and modify, and you are comfortable running a research repository from source. Skip it if you need a pip-installable inference library with a stable API, or if your deployment target cannot run the PyTorch versions in requirements.txt.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 45 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 3, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What D-FINE changes about DETR regression

DETRs predict boxes by regressing four coordinates directly. D-FINE reframes that step as Fine-grained Distribution Refinement (FDR), and adds Global Optimal Localization Self-Distillation (GO-LSD). The README states the two techniques achieve their results "without introducing additional inference and training costs", which matters because most accuracy gains in detection come with a heavier backbone or a longer schedule. The repository is the official implementation of the paper by Yansong Peng, Hebei Li, Peixi Wu, Yueyi Zhang, Xiaoyan Sun and Feng Wu at the University of Science and Technology of China, and it was accepted as an ICLR 2025 Spotlight.

The intended audience is narrow. This is for someone who already works with detection training loops and wants a detector they can retrain, not for an application developer who wants a one-line predict call. The model zoo covers five sizes from D-FINE-N at 4M parameters to D-FINE-X at 62M, each with a COCO config, a checkpoint and a training log, plus Objects365+COCO variants for the smaller sizes. If your interest is the regression mechanism itself, the FDR head is the part worth reading.

How FDR and GO-LSD fit into the training loop

The repository layout tells you where each piece lives. configs/ holds the YAML definitions, src/ holds the model and solver code, tools/ holds the entry points, and train.py sits at the top level as the training entry. The README points at a blog under src/zoo/dfine/blog.md for the English write-up of the method, so the conceptual explanation is kept next to the model code rather than in the top-level README.

What the README does not give you is a diagram of the data flow or a description of the loss terms. It gives you the config and checkpoint pairs, and the numbers those pairs produce. That is a deliberate choice for a paper release, but it means the practical way to understand FDR is to open the config for the size you care about, follow which modules it instantiates, and read the corresponding source files. The Objects365+COCO configs under configs/dfine/objects365/ show a two-stage data recipe where a model is pretrained on Objects365 and then taken to COCO, which is the pattern to copy if you have a large weakly labelled corpus and a smaller clean set.

Installing D-FINE and running a first COCO evaluation

There is no pip package. The README describes the source route, and requirements.txt lists the dependency set: torch>=2.0.1, torchvision>=0.15.2, faster-coco-eval>=1.6.6, PyYAML, tensorboard, scipy, calflops, transformers and loguru. Install that file into an environment of your choosing; the repository does not dictate how the environment is created.

bash
pip install -r requirements.txt

The Dockerfile is the alternative. It builds from a prebuilt base image on a registry mirror, and the commented history in the file shows an Ubuntu 18.04 CUDA 12.0.1 image with Miniconda installed on top. If you use it, note that the base image reference is a mirror host, so a build in a network that cannot reach that host will fail before any D-FINE code runs.

The README does not spell out a single evaluation command in the excerpt available, so the config and checkpoint columns in the model zoo are the entry point: pick a row, take its yml path and its .pth URL, and drive them through the tools/ scripts. For a first run, D-FINE-N is the cheapest row to validate your environment, at 4M parameters and 7 GFLOPs. If the checkpoint downloads and the config loads, your stack is sound and you can move up to D-FINE-S or D-FINE-M.

Fine-tuning on a custom dataset

The updates list records that custom dataset finetuning configs were added on 2024-10-25 in response to issue #7. Those configs are the intended path for your own data, and they are the reason to prefer this repository over reimplementing the paper. You start from a COCO checkpoint rather than from random weights, which is what makes a small labelled set viable.

The constraint is that your annotations have to be converted into whatever format those configs expect. The README does not document that conversion in the available text, and there is no dataset preparation section in it. Budget time for reading the custom config and the dataset loader in src/ before you label anything, because the annotation schema drives the whole pipeline. A second constraint: the reported latencies are measured on a T4, so a desktop GPU or a CPU-only box will not reproduce them, and the README does not offer latency numbers for other hardware.

Where D-FINE is the wrong tool

D-FINE is a detector. If your task is segmentation, pose estimation or tracking, the model zoo gives you detection checkpoints only, and there is no indication in the README that the heads extend to those tasks. You would be adapting the codebase, not using it.

The second mismatch is deployment. The repository ships a training-oriented Dockerfile and a requirements.txt aimed at a GPU training environment. Nothing in the README describes an export path to ONNX, TensorRT or a mobile runtime. If your target is an edge device with a fixed inference engine, you are on your own for the conversion, and that work is often larger than the fine-tune itself.

The third is API stability. This is research code tied to a paper. The updates log shows model releases and a pretrained weight revision between October and November 2024, including a D-FINE-L update that changed performance by 2.0 percent. Treat the config schema as something that can move between revisions of the repository, and pin the commit you validated.

D-FINE against YOLO and RF-DETR

The README itself frames the comparison against YOLO. It describes a street scene video where D-FINE-X detects nearly all targets under backlighting, motion blur and dense crowds, including small objects such as backpacks, bicycles and traffic lights, and states that its confidence scores and localization precision on blurred edges are higher than YOLO11's. That is the project's own demonstration, not an independent benchmark, and it is a qualitative video comparison rather than a table.

The architectural difference is the one that matters for adoption. YOLO-family detectors regress boxes directly through a convolutional head and are generally distributed with export tooling and a stable inference API. D-FINE keeps the DETR formulation and changes the regression target to a distribution that is refined across decoder layers, with self-distillation guiding localization. That is what the AP column in the model zoo is buying: D-FINE-M reaches 52.3 AP at 19M parameters and 5.62ms on a T4, and D-FINE-X reaches 55.8 AP at 62M and 12.89ms. Against RF-DETR, the other DETR-family alternative people search for, the shared premise is transformer detection; the difference is that D-FINE's contribution is specifically in the regression head and its distillation signal, so if you are choosing between them, compare what each does to the box refinement stage rather than the backbone.

Licence, maintenance and the cost of upgrading

D-FINE is Apache-2.0. That is a permissive licence, and the repository carries the LICENSE file at the top level. Apache-2.0 includes an explicit patent grant and requires that you preserve notices; it does not require you to open your own code. This is not legal advice, and if you ship a product built on these weights you should read the licence text yourself, particularly since the checkpoints are hosted on a separate GitHub releases page under the same author's storage repository rather than in the main tree.

The last push to the repository was on 2026-08-19, so there is recent activity. That is not the same as a support commitment. There are no retrieved releases, so versioning happens through commits and through the model zoo's checkpoint URLs. The practical upgrade cost is the config and checkpoint pairing: a new checkpoint may expect a config that differs from the one you tuned against, and the D-FINE-L revision in the updates log is the precedent. Keep your fine-tuned weights and the exact commit hash together, and re-run your evaluation set whenever you move either one.

Editorial conclusion

Adopt D-FINE if you need a COCO-pretrained real-time detector whose regression head you can read and modify, and you are comfortable running a research repository from source. Skip it if you need a pip-installable inference library with a stable API, or if your deployment target cannot run the PyTorch versions in requirements.txt. Before committing, verify that the checkpoint for the size you picked loads through the config in configs/dfine/ and that your own dataset converts into the format the custom finetuning configs expect, because the README does not document a rollback path if a fine-tune diverges.

Frequently asked questions

What does D-FINE do?

It is a real-time object detector that redefines the bounding box regression task in DETRs as Fine-grained Distribution Refinement, with Global Optimal Localization Self-Distillation added on top. The README states both techniques work without adding inference or training cost.

How does D-FINE compare with YOLO?

The README presents a street scene video comparison in which D-FINE-X detects nearly all targets under backlighting, motion blur and dense crowds, with higher confidence and better localization precision on blurred edges than YOLO11. That is the project's own demonstration rather than an independent benchmark.

What is the difference between D-FINE and RF-DETR?

Both are DETR-family detectors, but D-FINE's contribution is in the regression stage, reframing box regression as distribution refinement and adding localization self-distillation. The README does not contain a direct comparison with RF-DETR.

How does D-FINE compare with RT-DETR?

Both are real-time DETR-family detectors. D-FINE's stated contribution is reframing bounding box regression as Fine-grained Distribution Refinement and adding Global Optimal Localization Self-Distillation, which the README says comes without extra inference or training cost. The README does not include a direct comparison with RT-DETR.

How does D-FINE compare with YOLOv11?

The README's video comparison states that D-FINE-X detects nearly all targets in a crowded street scene and shows higher confidence scores and better localization precision on blurred edges than YOLO11. The comparison is the project's own and is qualitative.

What are the alternatives to D-FINE?

The README positions D-FINE against YOLO11 in its video comparison, and the model zoo lists five COCO sizes plus Objects365+COCO variants for the smaller ones. Other DETR-family detectors such as RF-DETR and RT-DETR share the transformer detection premise but differ in how they handle the regression stage.

Official sources

  1. Issues
  2. License: Apache-2.0
  3. Peterande/D-FINE on GitHub
  4. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/peterande-d-fine.svg)](https://hysenlabs.com/projects/peterande-d-fine)