# facebookresearch/dinov2: Loading the Backbones and What the Repo Actually Ships

> DINOv2 is a family of self-supervised vision transformers that produce features usable by a linear classifier. The repository is easiest to consume through torch.hub, and the harder part is everything around the training code.

**facebookresearch/dinov2** — PyTorch code and models for the DINOv2 self-supervised learning method.

- Repository: https://github.com/facebookresearch/dinov2
- Stars: 13,380 · Forks: 1,270
- Language: Jupyter Notebook
- License: Apache-2.0
- Published: 2026-09-21 · Updated: 2026-09-21 · Language: en
- Canonical page: https://hysenlabs.com/projects/facebookresearch-dinov2

## What DINOv2 gives you that a supervised ImageNet backbone does not

The pitch is narrow and specific. DINOv2 models produce visual features that the README says can be used directly with classifiers as simple as linear layers across a range of computer vision tasks, without fine-tuning. The models were pretrained on 142 million images with no labels or annotations. That matters when your labelled set is small, because the cost of adapting to a new domain drops to fitting a linear probe or a small head rather than retraining a backbone.

The intended user is someone who has images and a downstream task but no large annotated dataset. If you already own a labelled corpus of hundreds of thousands of images in your exact domain, a supervised backbone trained on that corpus is a legitimate competitor and may win. DINOv2 is a general-purpose feature extractor, not a task-specific model. It does not ship a classifier, a detector, or a segmentation head as a finished product.

The repository also carries research extensions that go beyond the original release: dino.txt inference code added in June 2025, Channel-Adaptive DINO and Cell-DINO added in December 2025, and XRay-DINO backbone loading added the same month. Those are separate papers with separate code paths and, in two cases, separate licence files.

## ViT backbones, registers and the patch-feature output

The architecture is a Vision Transformer, and the README lists four sizes: ViT-S/14 at 21M parameters, ViT-B/14 at 86M, ViT-L/14 at 300M, and ViT-g/14 at 1,100M. The /14 suffix is the patch size. Each size exists in two variants, one plain and one with registers, following the paper Vision Transformers Need Registers. The register variants add tokens that absorb the high-norm outlier patches the paper identifies; the README's own tables show the accuracy difference is small, for example ViT-g/14 goes from 83.5% to 83.7% k-NN and 86.5% to 87.1% linear.

What you get out of the model is not a class label. It is a set of patch features, one per patch, plus a global representation. The repository's own visualization maps the first three principal components of the patch features to RGB, which is a good mental model: each patch is an embedding, and downstream tasks either pool them or attach a head to the grid.

The distilled models are worth noting. The small and base entries in the README are labelled distilled, which is why a 21M parameter model reaches 79.0% k-NN. The g/14 is not distilled. If you are choosing on memory budget, the plain ViT-S/14 at 21M is the entry point, and the register variant of the same size is a drop-in swap.

## Installing DINOv2 and running a first forward pass

The README states that PyTorch is the only required dependency for loading the model, and points to the PyTorch install instructions, with CUDA support strongly recommended. The repository also ships a requirements.txt for the full codebase, which pins torch==2.0.0 and torchvision==0.15.0 against the CUDA 11.7 wheel index, plus omegaconf, torchmetrics==0.10.3, fvcore, iopath, xformers==0.0.18, submitit, and cuml-cu11 from the NVIDIA index. That file is for training and evaluation, not for inference. If you only want features, installing PyTorch alone is enough.

The fastest path is torch.hub, which pulls the repository and the checkpoint together. This snippet is the one the README gives:

```python
import torch

dinov2_vits14 = torch.hub.load('facebookresearch/dinov2', 'dinov2_vits14')
dinov2_vitb14 = torch.hub.load('facebookresearch/dinov2', 'dinov2_vitb14')
dinov2_vitl14 = torch.hub.load('facebookresearch/dinov2', 'dinov2_vitl14')
dinov2_vitg14 = torch.hub.load('facebookresearch/dinov2', 'dinov2_vitg14')
```

Running that downloads the backbone weights and returns a module. The register variants use the same pattern with a _reg suffix, for example dinov2_vits14_reg, dinov2_vitb14_reg, dinov2_vitl14_reg and dinov2_vitg14_reg.

For a full checkout rather than hub loading, the repository has a conda.yaml in the root and a setup.py exposing the package as dinov2, with REQUIRES_PYTHON set to >=3.9.0. The XRay-DINO backbone is the exception to the simple path: the README says you request it through a form at ai.meta.com, receive an email with a temporary link, and then either download with wget and point at the local checkpoint path or pass the URL from the email into the loading code. That is a gated download, not a public one.

## Where DINOv2 stops being the right tool

The repository is research code, and the maintenance signal is mixed. The last push was on 2026-06-03, and the README itself points readers to DINOv3 as a more recent effort continuing the line of work. That is a direct statement from the project that a successor exists. Anyone starting a new project and expecting DINOv2 to be the growth path should read that line carefully.

The dependency pinning is a real constraint. torch==2.0.0 with xformers==0.0.18 and cuml-cu11 is a snapshot of a particular CUDA 11.7 environment. If your infrastructure is on a newer PyTorch or a different CUDA major version, the training and evaluation path in requirements.txt is not a drop-in. Inference through torch.hub avoids most of this because it does not need the extras.

Feature extraction is also not free at the large end. ViT-g/14 is 1,100M parameters, and the README does not document memory requirements or throughput for any variant. If your target is an edge device or a high-throughput service, the 21M ViT-S/14 is the only entry that plausibly fits, and even then you are running a transformer over patches rather than a small CNN. Nothing in the repository is a deployment format: there is no ONNX export, no quantized build, and no serving wrapper documented in the README.

## DINOv2 against CLIP-style image-text models

The obvious alternative for a frozen-feature workflow is a contrastive image-text model such as CLIP, which is trained to place images and captions in one embedding space. The difference in approach is the supervision signal. CLIP learns from paired image and text, so it can do zero-shot classification by comparing an image embedding against text embeddings for candidate labels. DINOv2 is trained without labels or text, so it has no text encoder and cannot classify by prompt. You need labelled examples to fit even a linear head.

That trade cuts both ways. Because DINOv2 is not tied to captions, its patch features are not shaped by language supervision, which is why the repository can show patch-level visualizations and why the features work for dense tasks like segmentation. The README's own extension, dino.txt, exists precisely because the base method has no text alignment; the paper is titled DINOv2 Meets Text, and the inference code was added in June 2025. If you need text-conditioned retrieval, you are looking at that add-on, not the base model.

A second alternative is simply a supervised backbone from torchvision, fine-tuned on your data. It will beat DINOv2 when you have enough labels, and it comes with a simpler dependency story. DINOv2's advantage is concentrated in the low-label regime.

## Licensing, the extra model files, and what to check before shipping

The repository is Apache-2.0, and the LICENSE file sits in the root. That does not cover everything in the tree. The top-level entries include LICENSE_CELL_DINO_CODE, LICENSE_CELL_DINO_MODELS, and LICENSE_XRAY_DINO_MODEL, which means the Cell-DINO and XRay-DINO additions carry their own terms, and in the Cell-DINO case the code and the models are licensed separately. The README also notes a model card, MODEL_CARD.md, is included in the repository. If you plan to use only the standard ViT backbones, the Apache-2.0 terms are the relevant ones; if you pull in XRay-DINO or Cell-DINO, read the corresponding file rather than assuming the root licence applies. This is a description of what the repository contains, not legal advice.

Upgrade cost is low if you stay on torch.hub, because you are consuming checkpoints rather than the training stack. It is high if you depend on the training and evaluation code, because requirements.txt pins exact versions of torch, torchvision, torchmetrics and xformers, and moving any of them is an untested combination as far as the repository documents. There are no retrieved releases, so there is no changelog to diff between versions; the README's dated entries are the closest thing to a release history.

## Conclusion

Adopt DINOv2 if you need frozen visual features for classification, retrieval or segmentation without training a backbone yourself, and start by loading a single checkpoint through torch.hub before touching the training stack. Do not adopt it if you need image-text alignment out of the box or if you cannot run PyTorch with CUDA, since the pinned requirements target CUDA 11.7 wheels. Before committing, verify which checkpoint variant you need (with or without registers), confirm the licence file that applies to the specific weights you download, and check the XRay-DINO and Cell-DINO licence files separately if you plan to use those backbones.

## FAQ

### What are the uses of DINOv2?

The README says the models produce visual features that can be used directly with classifiers as simple as linear layers across a variety of computer vision tasks, without fine-tuning. The repository also includes inference code for dino.txt and backbones for XRay-DINO and Cell-DINO.

### What is a DINOv2 patch feature?

The model is a Vision Transformer with a patch size of 14, and it emits one embedding per patch alongside a global representation. The repository visualizes the first three principal components of the patch features mapped to RGB.

### How do I install DINOv2?

For loading a model, PyTorch is the only required dependency, and the README points to the PyTorch install instructions with CUDA strongly recommended. The full training codebase uses requirements.txt, which pins torch==2.0.0 and torchvision==0.15.0 against the CUDA 11.7 wheel index.

### What is DINOv2 trained on?

The README states the models were pretrained on a dataset of 142 million images without using any labels or annotations.

### What type of model is DINOv2?

It is a self-supervised Vision Transformer, available in ViT-S/14, ViT-B/14, ViT-L/14 and ViT-g/14 sizes, each with an optional register variant following the paper Vision Transformers Need Registers.

### What is the output of DINOv2?

The README does not describe the output tensor in detail, but it does show patch features being reduced to principal components for visualization, which indicates the model returns per-patch embeddings rather than class predictions.

## Sources

- [facebookresearch/dinov2 on GitHub](https://github.com/facebookresearch/dinov2)
- [Issues](https://github.com/facebookresearch/dinov2/issues)
- [License: Apache-2.0](https://github.com/facebookresearch/dinov2/blob/main/LICENSE)
- [README](https://github.com/facebookresearch/dinov2/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/facebookresearch-dinov2
