Library / SDK
apple-aiml-research/ml-cvnets avatar
apple-aiml-research/ml-cvnets

CVNets: Apple's Training Toolkit for Mobile and Non-Mobile Vision Models

CVNets: A library for training computer vision networks

1,991 stars258 forksPythonNOASSERTION

At a glance

What is it?
CVNets is a PyTorch library from Apple for training classification, detection, segmentation and CLIP-style foundation models, with a strong focus on mobile architectures. It installs from source with pip and ships console entry points for training, evaluation and CoreML conversion.
Who is it for?
Adopt CVNets if you are training or reproducing mobile-oriented vision architectures such as MobileViT, MobileNet or EfficientNet, and you want the same codebase used in the papers listed in the README. Skip it if you need a stable pip release with semantic versioning, or if you want a framework that abstracts away the training loop: CVNets expects you to work with its config files and entry points directly.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository last received commits 19 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The gap CVNets fills for mobile vision work

Most PyTorch training code you find in a research repository is written for one paper and one dataset. Reproducing a second architecture means rewriting the data pipeline, the optimizer setup, the augmentation stack and the evaluation loop. CVNets packages those parts once and exposes them through configuration files, so a MobileNetV3 run and a Swin Transformer run share the same engine. The README frames the scope explicitly: standard and novel mobile and non-mobile models, covering object classification, object detection, semantic segmentation and foundation models such as CLIP.

The audience is narrow but clear. This is for researchers and engineers who need to train these specific families of models, or who want to reproduce the Apple publications the README lists, including MobileViT, MobileViTv2, RangeAugment and ByteFormer. If your work is on large language models, tabular data or audio, nothing here applies. The library assumes images, and its model zoo is organized around ImageNet, MSCOCO, ADE20K and Pascal VOC style benchmarks.

How the training engine is put together

The repository layout tells you most of the architecture. Top-level directories separate concerns: cvnets/ holds models, layers and transforms; data/ holds dataset readers; loss_fn/, metrics/ and optim/ hold the training components; engine/ holds the loop that ties them together; options/ parses configuration. The entry points are thin wrappers over that engine, which is why main_train.py, main_eval.py and main_conversion.py sit at the repository root rather than inside a package.

Configuration is the control surface. Options are parsed from config files, and the examples/ directory is organized by feature: examples/byteformer/, examples/range_augment/ and examples/vit/. A run is defined by pointing an entry point at a config and letting the option parser fill in the rest. That design is what makes swapping a backbone or a dataset a config edit rather than a code change, and it is also the main thing you have to learn before the library becomes useful. The README does not document a programmatic Python API for building a model without going through the config path, so treat configuration as the intended interface.

Installing CVNets and running a first training job

The README recommends Python 3.10+ and PyTorch version >= v1.12.0, and gives Conda-based instructions. The first step clones the repository and creates an isolated environment. Note that the README's clone command uses the SSH form of the URL, so you need a GitHub SSH key configured, or you substitute the HTTPS URL yourself.

bash
git clone [email protected]:apple/ml-cvnets.git
cd ml-cvnets
conda create -n cvnets python=3.10.8
conda activate cvnets

After activation, the README installs dependencies with a constraints file and then installs the package itself in editable mode. The constraints file is what pins the dependency graph; without it, pip resolves versions freely and you may get a combination the maintainers did not intend.

bash
pip install -r requirements.txt -c constraints.txt
pip install --editable .

The editable install is what registers the console scripts declared in setup.py. After it completes, cvnets-train, cvnets-eval, cvnets-eval-seg, cvnets-eval-det, cvnets-convert and cvnets-loss-landscape should be on your PATH. The README points to docs/source/en/models and the examples/ folder for concrete training and evaluation commands rather than inlining them, so the next step is to open a config under examples/ and run cvnets-train against it. The README does not document what a successful first run prints, so check the docs for the expected output.

Where CVNets gets in your way

The packaging is the first friction point. setup.py declares VERSION = 0.3, while the README's What's new section describes version 0.4 of the library. That mismatch between the packaged version string and the documented release is the kind of detail that matters if you are pinning dependencies in a downstream project. There are no retrieved releases on the repository, so you are installing from a branch, not from a tagged artifact.

The second constraint is the dependency surface. requirements.txt pulls torch, torchvision, torchtext, torchaudio and torchdata together, plus coremltools, pycocotools, cityscapesscripts, pytorchvideo, av and fvcore. Several of those are heavy and platform-sensitive; coremltools in particular is aimed at Apple platforms. A Linux CI image that only needs ImageNet classification still resolves the full set unless you curate the file yourself. The README does not describe a minimal install profile.

Finally, this is a research toolkit, not a deployment framework. Converting a trained model to CoreML is documented as a separate workflow with its own entry point, and the README points to a dedicated guide for it. If your goal is an inference runtime rather than a training pipeline, you are using the wrong half of the project.

CVNets against timm and plain torchvision training loops

The closest comparison in spirit is timm, which also collects vision architectures behind a shared training interface. The difference is emphasis. timm's center of gravity is a broad, quickly growing catalog of pretrained backbones with a consistent model-creation API, and its training scripts are comparatively self-contained. CVNets is organized around a configuration-driven engine with explicit modules for losses, metrics, optimizers and augmentation, and its catalog is deliberately narrower: MobileNet variants, EfficientNet, ResNet, RegNet, ViT, MobileViT v1 and v2, Swin, SSD, Mask R-CNN, DeepLabv3, PSPNet, CLIP and ByteFormer.

If you want to fine-tune an arbitrary backbone from a checkpoint, timm is the shorter path. If you want to reproduce a MobileViT training recipe, or run distillation with the soft and hard modes the README lists, CVNets is the implementation those recipes came from. The distinction is not quality; it is whether the config engine is an asset or overhead for your specific job.

Maintenance, licence and what an upgrade costs

The repository is not archived, and the last push was on 2026-09-11, so it is being touched. That said, the README's own What's new section stops at July 2023 and describes version 0.4, while setup.py still declares 0.3. The maintainers listed are Sachin, Maxwell Horton, Mohammad Sekhavat and Yanzi Jin, with Farzad named as a previous maintainer. Treat the release cadence as slow and the changelog as lagging the code.

Upgrade cost is driven by the configuration schema. Because options are parsed from config files, a change to an option name or default breaks existing configs rather than raising an obvious Python error. If you keep your own configs outside the repository, budget time for diffing them against examples/ after any pull. Pinning to a commit hash is more predictable than tracking main.

The licence is the open question. The repository's licence identifier is NOASSERTION, the README says only "For license details, see LICENSE", and every source file carries a header pointing at that same file. Nothing in the README states the terms. Read LICENSE before you plan around this code, and if the terms are unclear for your use case, ask someone qualified rather than inferring from the fact that the repository is public.

Editorial conclusion

Adopt CVNets if you are training or reproducing mobile-oriented vision architectures such as MobileViT, MobileNet or EfficientNet, and you want the same codebase used in the papers listed in the README. Skip it if you need a stable pip release with semantic versioning, or if you want a framework that abstracts away the training loop: CVNets expects you to work with its config files and entry points directly. Before committing, verify the LICENSE file for the terms that apply, check whether the console scripts appear after pip install --editable ., and confirm that your PyTorch build matches what requirements.txt and constraints.txt resolve to on your platform.

Frequently asked questions

Is computer vision part of machine learning?

CVNets sits in that overlap: the README describes it as a computer vision toolkit for training models, and the training stack is built on PyTorch with optimizers, losses and metrics under optim/, loss_fn/ and metrics/.

What does CV mean in AI?

In this project's context it means computer vision. CVNets covers object classification, object detection, semantic segmentation and foundation models such as CLIP, with benchmark ties to ImageNet, MSCOCO, ADE20K and Pascal VOC.

Is OpenCV still relevant?

CVNets is a separate project and the README does not mention OpenCV. CVNets handles model training and evaluation in PyTorch, while OpenCV is an image processing library, so they occupy different parts of a vision pipeline.

What is OpenCV in machine learning?

The README does not describe OpenCV or any relationship to it. CVNets is Apple's PyTorch library for training vision networks, installed from source with pip install --editable . after the requirements and constraints files.

Official sources

  1. apple-aiml-research/ml-cvnets on GitHub
  2. Issues
  3. Project website
  4. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/apple-aiml-research-ml-cvnets.svg)](https://hysenlabs.com/projects/apple-aiml-research-ml-cvnets)