Library / SDK
facebookresearch/dinov3 avatar
facebookresearch/dinov3

DINOv3: Meta's Self-Supervised Vision Backbones and What It Takes to Run Them

Reference PyTorch implementation and models for DINOv3

11,468 stars968 forksJupyter NotebookNOASSERTION

At a glance

What is it?
DINOv3 is a family of pretrained vision foundation models from Meta AI Research, released as a reference PyTorch implementation with gated weights. It is strongest as a frozen feature extractor for dense tasks, and the access process is the first thing you have to plan for.
Who is it for?
Adopt DINOv3 if you need frozen dense features for segmentation, depth or detection and can accept gated weight access plus the licence terms. Do not adopt it if you need a permissively licensed checkpoint you can redistribute or an offline install with no approval step; the weights are not on PyPI.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository last received commits 76 days ago.
What is it written in?
Mainly Jupyter Notebook, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

DEEP OPEN-SOURCE ANALYSIS

What DINOv3 actually solves, and who it is for

Training a vision backbone from scratch is expensive, and fine-tuning a large one per task is only slightly cheaper. DINOv3 is positioned against both problems. The README describes it as "An extended family of versatile vision foundation models producing high-quality dense features and achieving outstanding performance on various vision tasks including outperforming the specialized state of the art across a broad range of settings, without fine-tuning." The operative phrase is without fine-tuning. The intended workflow is to load a frozen backbone, extract per-patch features, and attach a small task head on top.

The audience is therefore narrower than the model family's breadth suggests. It fits research and applied teams working on dense prediction: semantic segmentation, monocular depth, detection, and remote sensing. The repository ships linear probing code for ADE20K segmentation and NYUv2-Depth estimation, and a Canopy Height Maps v2 (CHMv2) model for canopy height estimation from satellite imagery. It does not fit teams that want a single packaged classifier with a stable Python API and no weight request form.

The backbone and adapter split in the DINOv3 repository

The model listing separates two kinds of artifact. Backbones are the ViT and ConvNeXt networks pretrained on LVD-1689M, ranging from ViT-S/16 distilled at 21M parameters to ViT-7B/16 at 6,716M parameters. Adapters are the task-specific heads, such as the DPT head used by CHMv2. The README states that the download email contains URLs for "all the available model weights (both backbones and adapters)", and that those URLs can either be pointed at local files via the `weights` or `backbone_weights` parameters, or passed directly to `torch.hub.load()` for download and load.

That split matters operationally. A backbone URL is reusable across tasks; an adapter URL is tied to the head it was trained for. Loading code has to know which of the two parameters it is filling. The repository layout reinforces the separation: a `dinov3/` package, a top-level `hubconf.py` for torch.hub entry points, `notebooks/` for worked examples, and `MODEL_CARD.md` alongside `DATASETS.md` for provenance. Distillation code for ConvNeXt backbones was released separately, which is why the ConvNeXt entries exist alongside the ViT ones.

Installing DINOv3 and loading a first backbone

The package declares `REQUIRES_PYTHON = ">=3.11"` in `setup.py`, and the importable package name is `dinov3`. Runtime dependencies in `requirements.txt` are short and conventional: `torch`, `torchvision`, `torchmetrics`, `scikit-learn`, `omegaconf`, `submitit`, `ftfy`, `regex` and `termcolor`. A conda environment file is provided at `conda.yaml`.

The first hard constraint is not in the install command. Weights are gated. The README instructs you to follow the download link for access, after which an email arrives with the full list of URLs. It also warns to use `wget` rather than a web browser for the download. Until that email arrives, there is nothing to load.

Once you have a URL, the load path is short. This is the shape the README describes, with the URL supplied by the access email:

python
import torch

model = torch.hub.load(
    repo_or_dir="facebookresearch/dinov3",
    model="dinov3_vits16",
    weights="<URL from the access email>",
)

For an adapter, the same call takes `backbone_weights` instead. If you would rather keep the checkpoint on disk, download it first and pass the local path in the same parameter. The README does not document a checksum or a signature check on the downloaded file, so verify the file yourself against whatever the email provides.

Where DINOv3 is the wrong tool

Three cases stand out. The first is redistribution. The licence is listed as NOASSERTION, and `setup.py` carries the line "This software may be used and distributed in accordance with the terms of the DINOv3 License Agreement." That is a pointer to a separate agreement, not an open source identifier. If your product needs to ship a checkpoint, or your legal review requires an OSI-approved licence, read `LICENSE.md` before you build anything on top.

The second is a zero-friction install. There is no published package for the weights, and the access step is a human approval. Teams that need reproducible CI without an outbound request to Meta will find this awkward; a mirror of the weights inside your own artifact store is the usual workaround, and the licence governs whether you may do that.

The third is small-data fine-tuning. DINOv3's pitch is frozen features plus a linear or lightweight head. If your task is far from natural images and you have thousands of labelled examples, a smaller supervised backbone trained end to end may beat a frozen 7B model on both accuracy and cost. The repository's own linear probing recipes are the honest baseline to compare against before assuming the larger backbone helps.

DINOv3 versus timm and Hugging Face Transformers

The most practical alternative is not another self-supervised method, it is loading the same weights through a different library. The README records that DINOv3 backbones are supported by PyTorch Image Models (timm) starting with version 1.0.20, and by Hugging Face Transformers starting with version 4.56.0, with weights on the Hugging Face Hub. The CHMv2 model is also on the Hub and documented in the Transformers model docs.

The difference in approach is real. This repository gives you the reference implementation: training and distillation code, linear probing recipes for ADE20K and NYUv2-Depth, notebook examples, and the exact configuration used in the paper. The Transformers and timm paths give you the backbone as a standard model class inside an ecosystem you may already depend on, with the surrounding tooling that implies. If you only need inference on a frozen backbone, the ecosystem route is less code. If you need to reproduce a probing result, run distillation, or adapt the training loop, this repository is the one that carries that machinery. Choosing between them is mostly a question of whether you are consuming the model or studying it.

Maintenance, upgrade cost and the licence question

The last push to the default branch was on 2026-07-15, and the repository is not archived. The release history visible in the README is steady rather than fast: semantic segmentation and depth probing code in October 2025, ConvNeXt distillation in November 2025, CHMv2 in March 2026, and metadata-guided training recipes on the `FINO` branch in June 2026. That cadence suggests the surface area keeps growing, which raises the cost of pinning.

Upgrades are not drop-in. The README states a minimum Python of 3.11, and the package version is read from `dinov3/__init__.py`, so a version bump in that file is the signal to watch. Backbone weights are addressed by URL, and adapter weights are tied to specific heads, so a new adapter does not invalidate an old one but a changed backbone does invalidate anything trained on its features. Budget for re-running linear probes when you move between backbones.

On licensing: the repository ships `LICENSE.md`, and `setup.py` refers to the DINOv3 License Agreement. The GitHub licence field reads NOASSERTION, which means GitHub could not map the file to a known identifier. That is a signal to read the agreement itself, not a statement about what it permits. Nothing here is legal advice; the only concrete step is to open `LICENSE.md` and `MODEL_CARD.md` before you decide how the weights will be used or stored.

Editorial conclusion

Adopt DINOv3 if you need frozen dense features for segmentation, depth or detection and can accept gated weight access plus the licence terms. Do not adopt it if you need a permissively licensed checkpoint you can redistribute or an offline install with no approval step; the weights are not on PyPI. Verify first that your request for the download links has been accepted, that you can run the PyTorch version your environment pins, and that your intended adapter (segmentation, depth, or a downstream head) is among the released ones.

Frequently asked questions

Is DINOv3 a transformer?

The family includes ViT backbones such as ViT-S/16, ViT-B/16, ViT-L/16, ViT-H+/16 and ViT-7B/16, so yes, the main line is transformer-based. The repository also ships ConvNeXt backbones, with distillation code and configurations released for them.

How many images was DINOv3 trained on?

The pretraining dataset is listed as LVD-1689M for both the ViT and ConvNeXt model tables. The README does not break that figure down further.

How do I get access to the DINOv3 weights?

The README says to follow the download link to request access, and that once accepted an email is sent with the complete list of URLs for all available model weights, both backbones and adapters. Those URLs can be passed to torch.hub.load() or used to download the weights locally.

How do I install DINOv3?

The package requires Python 3.11 or later and depends on torch, torchvision, torchmetrics, scikit-learn, omegaconf, submitit, ftfy, regex and termcolor. A conda environment file is provided at conda.yaml, and setup.py installs the dinov3 package.

What is DINOv3 used for?

The README describes it as a family of vision foundation models producing high-quality dense features, evaluated without fine-tuning. The repository ships linear probing code for ADE20K semantic segmentation and NYUv2-Depth monocular depth estimation, plus a Canopy Height Maps v2 model for satellite imagery.

Official sources

  1. facebookresearch/dinov3 on GitHub
  2. Issues
  3. README
For maintainers

Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/facebookresearch-dinov3.svg)](https://hysenlabs.com/projects/facebookresearch-dinov3)
Community notes

Community notes