# FastViT: Apple's ICCV 2023 hybrid transformer with structural reparameterization

> FastViT is a hybrid vision transformer from Apple research that uses structural reparameterization to maintain multi-branch training blocks while fusing them into single convolutions at inference. Seven ImageNet-1K pretrained variants are available as PyTorch checkpoints and CoreML models, benchmarked on iPhone 12 Pro using the ModelBench app.

**apple-aiml-research/ml-fastvit** — This repository contains the official implementation of the research paper, "FastViT: A Fast Hybrid Vision Transformer using Structural Reparameterization" ICCV 2023

- Repository: https://github.com/apple-aiml-research/ml-fastvit
- Stars: 2,039 · Forks: 130
- Language: Python
- License: NOASSERTION
- Published: 2026-10-09 · Updated: 2026-10-09 · Language: en
- Canonical page: https://hysenlabs.com/projects/apple-aiml-research-ml-fastvit

## The structural reparameterization trick that separates training from inference

FastViT inherits structural reparameterization from MobileOne, the Apple research model the repository cites through its ModelBench app benchmark. During training the model maintains multi-branch blocks whose parallel paths give gradients a wider set of routes during backpropagation. At inference those branches are algebraically merged into a single convolution, which reduces memory operations per forward pass without changing the model output.

The repository exposes this through reparameterize_model in models/modules/mobileone. After training, passing a model instance through that function returns a new object in the fused state. This conversion is one-way for a given set of weights: a fused checkpoint cannot be unfolded back into training branches. That is why every architecture in the model zoo ships as two separate files. An unfused file, using the architecture name alone, is the one to load before fine-tuning or downstream task training such as detection or segmentation. A fused file, whose name ends in _reparam, is for inference and latency measurement.

Skipping reparameterize_model before benchmarking the unfused model produces numbers that correspond to no figure in the README. The latency table reflects fused weights on iPhone 12 Pro, and comparisons using the unfused model are not comparable to those numbers.

## Python 3.9 and PyTorch 1.11.0 pin the environment to a March 2022 stack

The setup block in the README specifies a Conda environment built on Python 3.9, PyTorch 1.11.0, torchvision 0.12.0, torchaudio 0.11.0 and CUDA toolkit 11.3. These package versions predate the current PyTorch and CUDA release lines by several major versions. On a system with a recent NVIDIA driver or a newer GPU architecture, installing this exact combination under Conda may require additional configuration and may not succeed without tweaking channel priorities or accepting partial downgrades of system libraries.

requirements.txt adds only two packages: timm and scipy. The timm library handles the create_model call that constructs FastViT variants by name, so the installed version of timm determines which architecture names are registered. The README does not pin a timm version. A fresh pip install -r requirements.txt installs the current timm release, which may introduce API changes or register new defaults that differ from what the authors used.

No Dockerfile appears in the top-level repository entries, and no container-based path is described. The Conda instructions are the only documented environment method. Anyone reproducing the setup on hardware that does not accept the pinned CUDA toolkit needs to resolve compatible alternatives independently, starting from the PyTorch compatibility matrix for their driver version.

## Setting up the conda environment and loading a pretrained checkpoint

The README gives four commands to create and populate the environment:

```bash
conda create -n fastvit python=3.9
conda activate fastvit
conda install pytorch==1.11.0 torchvision==0.12.0 torchaudio==0.11.0 cudatoolkit=11.3 -c pytorch
pip install -r requirements.txt
```

Once the environment is active, the README provides a usage block covering the three main phases: creating a model, loading a checkpoint for fine-tuning, and preparing for inference:

```python
import torch
import models
from timm.models import create_model
from models.modules.mobileone import reparameterize_model

# To Train from scratch/fine-tuning
model = create_model("fastvit_t8")

# Load unfused pre-trained checkpoint for fine-tuning
checkpoint = torch.load('/path/to/unfused_checkpoint.pth.tar')
model.load_state_dict(checkpoint['state_dict'])

# For inference
model.eval()      
model_inf = reparameterize_model(model)
# Use model_inf at test-time
```

The checkpoint key is state_dict. The unfused T8 file URL follows the convention in the model zoo table, ending in fastvit_t8.pth.tar. The fused variant for inference ends in fastvit_t8_reparam.pth.tar. For training from scratch, create_model alone is sufficient and no checkpoint is required.

Three additional scripts sit at the repository root: train.py, validate.py and export_model.py. validate.py handles ImageNet-1K evaluation, and train.py handles full training runs. Neither script's arguments or expected inputs are described in the README, so reading the files directly is necessary before using them.

## Seven model variants from T8 at 0.8 ms to MA36 at 4.6 ms on iPhone 12 Pro

The model zoo lists seven architectures for image classification, all trained on ImageNet-1K and benchmarked on iPhone 12 Pro with the ModelBench app. The accuracy and latency figures for the standard trained variants are: FastViT-T8 at 76.2% Top-1 and 0.8 ms; FastViT-T12 at 79.3% and 1.2 ms; FastViT-S12 at 79.9% and 1.4 ms; FastViT-SA12 at 80.9% and 1.6 ms; FastViT-SA24 at 82.7% and 2.6 ms; FastViT-SA36 at 83.6% and 3.5 ms; FastViT-MA36 at 83.9% and 4.6 ms.

A parallel table covers the same seven architectures trained with knowledge distillation. Distillation shifts Top-1 accuracy upward at identical latency. FastViT-T8 goes from 76.2% to 77.2%, FastViT-T12 from 79.3% to 80.3%, and FastViT-S12 from 79.9% to 81.1%. The T suffix variants are the smallest; the MA suffix at MA36 is the largest. Distilled checkpoints and standard checkpoints share the same latency numbers at the same architecture name.

All latency figures are specific to iPhone 12 Pro using the ModelBench app. No measurements appear for other phones, server GPUs or desktop processors. The accuracy figures come from ImageNet-1K, not from any other dataset or task.

## CoreML packages target Apple Neural Engine, not a general deployment path

Every model in the zoo comes with a CoreML file alongside the PyTorch checkpoint. Each CoreML package is distributed as a .mlpackage.zip archive from Apple's developer documentation CDN. File names follow the same pattern as the PyTorch checkpoints, with _reparam indicating fused weights, so fastvit_t8_reparam.pth.mlpackage.zip corresponds to the fused inference model.

CoreML runs on Apple Neural Engine, which is available on iPhone, iPad and Mac devices with Apple silicon. A project targeting Android phones, Linux servers or Windows workstations cannot use these CoreML files. For those targets, the PyTorch checkpoints are the only documented option.

The repository includes export_model.py at its root. This file likely handles conversion from PyTorch to CoreML, but the README does not describe its arguments, inputs or outputs. Building a deployment pipeline for a non-iPhone target requires reading that file before making any assumptions about supported formats. No ONNX or TFLite export path is documented anywhere in the repository files listed, and the README does not mention Android or server-side inference as intended use cases.

## When FastViT is the wrong choice: dataset mismatch and licensing limits

All pretrained models were trained on ImageNet-1K. Applying those weights to domains that differ significantly from ImageNet photographs, such as medical scans, satellite imagery or industrial inspection, requires fine-tuning from an unfused checkpoint. The README's usage block shows how to load an unfused checkpoint for that purpose, but the README says nothing about the amount of data or the number of epochs that would be needed for a given distribution shift. The quality of fine-tuned models on out-of-distribution data is not discussed.

The license is the other limit to plan around before adoption. GitHub classifies it as Other because it does not match a recognized standard license. Apple's research repositories frequently use a custom sample code license, which can restrict commercial use and redistribution more tightly than MIT or Apache 2.0 would. The LICENSE file at the repository root is the authoritative text, and the ACKNOWLEDGEMENTS file also sits there. Reading both before building a commercial product, distributing fine-tuned weights or incorporating the architecture into a larger codebase is the only way to know what the terms permit.

## No GitHub releases and no standard license make version tracking manual

The repository has no GitHub releases. There is no tagged version to reference in a paper citation, a requirements lock file or a dependency specification. The last push was on 2026-09-11, and the repository is not archived.

The five named authors are Pavan Kumar Anasosalu Vasu, James Gabriel, Jeff Zhu, Oncel Tuzel and Anurag Ranjan. The paper appeared at ICCV 2023 and is available at arxiv.org/abs/2303.14189. No CITATION.cff file appears in the top-level repository entries, so the arxiv preprint is the reference to use in academic citations.

Without versioned releases, tracking exactly what the codebase contains on any given day requires pinning to a specific commit hash. A clone at HEAD will pick up future changes to train.py, validate.py or the model definitions without any signal in a changelog. The only top-level files that document the project's history are README.md and ACKNOWLEDGEMENTS; neither functions as a changelog. For production use where reproducibility matters, record the commit hash at the time of cloning alongside the checkpoint filename.

## Conclusion

FastViT is worth considering when your target device is iPhone and you want an ICCV 2023 published architecture with pretrained weights you can fine-tune through the timm API. It is not the right starting point for server-side deployment, non-Apple mobile hardware, or any project where a standard open source license is a hard requirement. Before building a pipeline on it, read the LICENSE file to understand the custom Apple terms, confirm that PyTorch 1.11.0 and CUDA 11.3 are compatible with your GPU drivers, and run validate.py against the downloaded checkpoint to confirm you have the correct fused or unfused variant for your intended use.

## FAQ

### What is FastViT?

FastViT is a hybrid vision transformer from Apple research, presented at ICCV 2023, that uses structural reparameterization to separate training and inference graph structures. Seven variants are available as pretrained PyTorch checkpoints and CoreML models, all trained on ImageNet-1K.

### How do I install FastViT?

The README specifies a Conda environment with Python 3.9, PyTorch 1.11.0, torchvision 0.12.0, torchaudio 0.11.0 and cudatoolkit 11.3, followed by pip install -r requirements.txt to add timm and scipy. No Docker or pip-only path is documented.

### What is the difference between fused and unfused FastViT checkpoints?

An unfused checkpoint stores the multi-branch training structure and is the file to load before fine-tuning or downstream task training. A fused checkpoint has had those branches merged into single convolutions via reparameterize_model and is the one used for inference and latency benchmarking.

### Can FastViT run on hardware other than iPhone?

The PyTorch checkpoints can be loaded on any hardware that PyTorch supports. The CoreML model files require Apple hardware with Neural Engine. The README's latency figures apply only to iPhone 12 Pro measured with the ModelBench app, and no figures for other devices are provided.

### What license does FastViT use?

GitHub classifies the license as Other because it does not match a recognized standard license. The LICENSE file at the repository root contains the full terms, which should be read before any commercial use or redistribution of the code or checkpoints.

## Sources

- [apple-aiml-research/ml-fastvit on GitHub](https://github.com/apple-aiml-research/ml-fastvit)
- [Issues](https://github.com/apple-aiml-research/ml-fastvit/issues)
- [README](https://github.com/apple-aiml-research/ml-fastvit/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/apple-aiml-research-ml-fastvit
