DEIMv2: DINOv3 features pushed down to a 0.5M parameter detector
[DEIMv2] Real Time Object Detection Meets DINOv3
At a glance
- What is it?
- The successor to DEIM adds mobile-sized real-time detectors that clear 50 AP on COCO, plus an Intel Geti integration and a TensorRT fix for FP16.
- Who is it for?
- DEIMv2 is a better starting point than its predecessor for almost anyone starting fresh, because the model zoo spans 0.5M to the large variants and the Atto and Femto entries were designed for constrained hardware rather than trimmed down afterwards. The caveat is maturity.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository last received commits 43 days ago.
- What is it written in?
- Mainly Jupyter Notebook, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 20, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Same lab, different bet: DINOv3 features instead of matching tricks
DEIMv2 is a rewrite rather than a version bump. Its description is Real Time Object Detection Meets DINOv3, and the README states the design directly: DEIMv2 is an evolution of the DEIM framework while leveraging the rich features from DINOv3. Where DEIM attacked convergence speed in the training loop, DEIMv2 goes after the backbone, pulling in features from a self-supervised vision transformer.
That changes what you inherit. A DEIM-trained model costs the same at inference as the baseline it was trained on, because nothing about the architecture moved. DEIMv2 models do not have that property, since the DINOv3 feature extraction is part of the model.
The size range is the headline claim. The README says the method is designed with various model sizes, from an ultra-light version up to S, M, L and X, and that the S-sized model notably surpasses 50 AP on the challenging COCO benchmark. Those are the project's own numbers measured against its own baselines, so treat them as a starting point for comparison rather than a settled claim.
Authorship changed from DEIM. This paper lists Shihua Huang, Yongjie Hou, Longfei Liu, Xuanlong Yu and Xi Shen, with Intellindust AI Lab and Xiamen University as the two institutions. Longfei Liu appears here but not on the DEIM paper, and Xuanlong Yu is new to the list. Xi Shen is the corresponding author again, and the badge row still points at [email protected].
Eight model sizes, from 23.8 AP at 0.5M parameters
The model zoo is the substance of the repository, and the small end is where it differs most from its predecessor. Every entry carries a COCO AP figure, a parameter count, GFLOPs, a latency measurement, a YAML config path, a Hugging Face link and Google Drive and Quark links for both the checkpoint and the training log.
The three ultra-light entries are the ones to look at first. Atto reaches 23.8 AP on COCO with 0.5M parameters, 0.8 GFLOPs and 1.10ms latency, configured at `configs/deimv2/deimv2_hgnetv2_atto_coco.yml`. Femto gets 31.0 AP at 1.0M parameters, 1.7 GFLOPs and 1.45ms. Pico reaches 38.5 AP at 1.5M parameters, 5.2 GFLOPs and 2.13ms. The gap between Femto and Pico in FLOPs, from 1.7 to 5.2, for 7.5 AP is the steepest trade in the table.
Above that, N is listed at 43.0 AP with 3.6M parameters, 6.8 GFLOPs and 2.32ms. The parameter counts are worth comparing against DEIM's naming: DEIM's N was 4M at 2.12ms, and DEIMv2's N is 3.6M at 2.32ms. Roughly the same size and the same COCO score, which is the clearest evidence in the two tables that the DINOv3 features buy efficiency rather than raw accuracy at this tier.
The config filenames follow `deimv2_hgnetv2_<size>_coco.yml` across `configs/deimv2/`, so the backbone is HGNetv2 throughout, the same family DEIM used. The Hugging Face organisation is `Intellindust`, and the README credits NielsRogge for the upload.
What changed after release, and what broke
The Updates list is dated and covers the period since the 2025-09-26 release, and two entries there are about deployment rather than accuracy.
The first is a fix, dated 2026-01-07: for FP16 inference, use TensorRT 10.6 or later to ensure stable execution and correct detection results. That is an unusually blunt warning to find in a research README, and it tells you that running these models in half precision on an older TensorRT gives you silently wrong output rather than a crash. If you are deploying, treat that as a hard floor.
The second is an optimization dated 2025-10-28: the attention module in ViT-Tiny was optimized, reducing memory usage by half for the S and M models. Read that carefully, because it applies to the ViT-Tiny variant and specifically calls out S and M. The larger variants and the ultra-light ones are not covered by that saving.
The adoption entries are the more interesting signal. On 2025-10-02 DEIMv2 was integrated into X-AnyLabeling, which puts it in front of people doing annotation work rather than model research. On the same 2026-01-07 date, STA, the technique introduced in DEIMv2, was integrated into the LightlyTrain distillation library, which is the clearest evidence that something here travelled outside the lab. On 2026-08-13 DEIMv2 was integrated into Intel Geti, where you can fine-tune the DINOv3 S, M and L variants on your own data with no code, train on Intel dGPU, iGPU or CPU, export to OpenVINO IR and optimize for INT8.
An inference toolkit that lives in tools/
The repository carries a Jupyter notebook, `hf_models.ipynb`, at the root alongside `train.py`, and a `tools/` directory that the table of contents describes as one of seven sections. Together with the tensorboard and scipy dependencies, that points at a project set up for training runs you inspect afterwards rather than a packaged library.
The dependency list is tighter than DEIM's and pins PyTorch exactly, which is the opposite choice from the predecessor:
torch==2.5.1
torchvision==0.20.1
faster-coco-eval>=1.6.7
PyYAML
tensorboard
scipy
calflops
transformersExact pins on torch and torchvision mean a newer release will not be picked up automatically, and `faster-coco-eval` moves its floor from 1.6.5 to 1.6.7. Whether the stricter pins are deliberate or a side effect of a working environment, you should assume reproducing the reported numbers means reproducing these versions.
The tree is otherwise identical in shape to DEIM: `configs/`, `engine/`, `figures/`, `tools/`, `train.py`, `requirements.txt`, plus `LICENSE.md`, a `README.md` and an `example.jpg` that the repository presumably uses for a detection demo image. The one file with no counterpart in DEIM is the notebook.
The README's Deployment instructions are linked from the FP16 entry and anchored to section 4, Tools. That is the place to look for the export and inference path, and it is also the section most likely to have changed with the 2026-08-13 Geti integration.
Sizing it up before you commit
The repository has 2,037 stars, 216 forks and 44 open issues, with no tagged releases and no `CHANGELOG`. The last push to the default branch was 2026-08-24, which is recent enough that the EdgeCrafter work announced in March and accepted to TMLR in August is plausibly still landing here.
Two gaps are worth naming. The README says the S, M, L and X variants leverage DINOv3 features, distilled or pretrained, but it does not break down which of those four apply to each individual size. For the Atto, Femto and Pico entries that question is left open, and those are exactly the sizes where you would want to know whether you are shipping a DINOv3 backbone or a compact one. There is also no inference script named at the root, so the tooling path is inside `tools/` and you will be reading code to map it.
On licensing, the same mismatch appears as in DEIM. No recognised license identifier is attached to the repository, the README badge says DEIMv2 License rather than naming a standard one, and the tree contains a `LICENSE.md` file. Read that file rather than relying on the badge.
The fair comparison for anyone choosing between the two repositories: DEIMv2 costs more to train because the DINOv3 features have to be produced, and the README says the S, M, L and X variants use them. DEIM's claim is that it converges faster with identical inference cost. Those are different tradeoffs and the numbers do not settle which wins for your hardware, so benchmark the variant you intend to ship rather than the one that reads best in a table.
Editorial conclusion
DEIMv2 is a better starting point than its predecessor for almost anyone starting fresh, because the model zoo spans 0.5M to the large variants and the Atto and Femto entries were designed for constrained hardware rather than trimmed down afterwards. The caveat is maturity. The series was released on 2025-09-26, there are no tagged releases to pin against, and the last push was 2026-08-24, so you should expect to run from a clone. Verify three things before committing: which DINOv3 variant a given size uses, since the README says S, M, L and X leverage DINOv3 features while leaving the smaller sizes ambiguous, whether you need TensorRT 10.6 or later to avoid broken FP16 results, and whether the Intel Geti integration covers the size you actually want. For anyone reproducing the original CVPR work, DEIM is still the right repository; for deployment, this is it.
Frequently asked questions
What is DEIMv2 and how is it different from DEIM?
DEIMv2 is the successor to DEIM from Intellindust AI Lab, built on the description Real Time Object Detection Meets DINOv3. DEIM was a training technique that changed DETR's matching so convergence was faster with no inference cost. DEIMv2 instead builds on features from the DINOv3 self-supervised vision transformer and expands the model zoo down to a 0.5M parameter Atto model.
What accuracy does DEIMv2 reach on COCO?
The S-sized model surpasses 50 AP on COCO by the project's own account. The smaller entries are Atto at 23.8 AP with 0.5M parameters, Femto at 31.0 AP with 1.0M parameters, Pico at 38.5 AP with 1.5M parameters, and N at 43.0 AP with 3.6M parameters. Latency for Atto is listed at 1.10ms and for N at 2.32ms.
How do I run DEIMv2 models in FP16 without breaking?
Use TensorRT 10.6 or later. The README's 2026-01-07 update states that FP16 inference needs TensorRT at that version or above to ensure stable execution and correct detection results, which implies older versions produce wrong output rather than an error. Deployment instructions are in section 4, Tools. A separate 2025-10-28 update reduced attention module memory usage by half in the ViT-Tiny variant for the S and M models.
Where can I download DEIMv2 checkpoints?
Three ways per model. Each row of the model zoo links a Hugging Face repository under the `Intellindust` organisation, a Google Drive checkpoint, and a Quark mirror, plus Google Drive and Quark links for the training log. Configs are in `configs/deimv2/` with names following `deimv2_hgnetv2_<size>_coco.yml`, and the Hugging Face upload arrived on 2025-11-03 with credit to NielsRogge.
What is STA in DEIMv2 and where else is it used?
STA is a technique introduced in DEIMv2, and on 2026-01-07 it was integrated into the LightlyTrain distillation library, which the README presents as evidence of value in real-world training pipelines. That date also carries the FP16 inference fix. Separate integrations include X-AnyLabeling from 2025-10-02 and Intel Geti from 2026-08-13, where you can fine-tune the DINOv3 S, M and L variants without code and export to OpenVINO IR with INT8 optimization.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/intellindust-ai-lab-deimv2)