FastVLM: Apple's FastViTHD Vision Encoder and What the ml-fastvlm Repository Actually Ships
This repository contains the official implementation of "FastVLM: Efficient Vision Encoding for Vision Language Models" - CVPR 2025
At a glance
- What is it?
- Apple's CVPR 2025 vision-language model cuts time-to-first-token by shrinking the vision encoder's token output. The repository gives you PyTorch inference and a CoreML export path, but training is left to LLaVA.
- Who is it for?
- Adopt FastVLM if you need a vision-language model with a small encoder and fast time-to-first-token, and you are comfortable driving inference through predict.py or exporting to CoreML for Apple Silicon. Do not adopt it if you need a training recipe: the README points to the LLaVA codebase for that, and this repository only ships inference instructions.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository last received commits 19 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The Problem FastVLM Solves: Vision Encoding Is the Bottleneck
Vision-language models spend a disproportionate amount of their latency budget turning an image into tokens. A high-resolution image produces a large grid of visual tokens, the language model has to attend over all of them, and time-to-first-token grows with the square of that token count in the attention layers. FastVLM attacks the encoder side of that pipeline rather than the language model side.
The repository introduces FastViTHD, described in the README as a hybrid vision encoder designed to output fewer tokens and reduce encoding time for high-resolution images. The audience is narrow and specific: engineers who want to run a vision-language model where first-token latency matters, and who are willing to accept a fixed set of released checkpoints. The README claims the smallest variant outperforms LLaVA-OneVision-0.5B with 85x faster time-to-first-token and a 3.4x smaller vision encoder, and that the larger Qwen2-7B variants outperform Cambrian-1-8B with 7.9x faster time-to-first-token. Those are the authors' numbers from the paper, not independent measurements.
The project is not a general-purpose multimodal framework. It is a paper implementation with three model sizes, a PyTorch inference script, an export path, and a demo app.
How FastViTHD Changes the Vision-to-LLM Token Path
The architecture visible in the repository is a two-part stack. A FastViTHD vision encoder consumes the image and emits visual tokens; a language model consumes those tokens alongside the text prompt. The repository layout shows this split directly: the llava/ directory holds the modelling code inherited from the LLaVA codebase, model_export/ holds the conversion tooling, and predict.py is the entry point that wires the two halves together.
The design bet is that fewer visual tokens per image means less work in the language model's prefill pass, which is where time-to-first-token is spent. The README's comparison numbers are all framed as time-to-first-token ratios, which is consistent with that bet. What the README does not document is the token count per image resolution, the encoder's patch schedule, or how the hybrid convolution and attention layers are arranged. For those details you have to read the paper.
Checkpoints come in two stages. Stage 2 and stage 3 weights are published separately for each of the 0.5B, 1.5B, and 7B variants. The README does not explain what changed between stages, so treat the stage number as an opaque label until you check the paper.
Installing FastVLM and Running Your First Prediction
Setup is a conda environment plus an editable install. The README specifies Python 3.10, and pyproject.toml pins the heavy dependencies, including torch 2.6.0, torchvision 0.21.0, transformers 4.48.3, and coremltools 8.2. The install name in pyproject.toml is llava, version 1.2.2.post1, so the package you install is not called fastvlm.
conda create -n fastvlm python=3.10
conda activate fastvlm
pip install -e .Weights are not bundled. The repository ships a shell script that downloads every checkpoint into a checkpoints directory, and the README warns this can take a while depending on your connection.
bash get_models.sh # Files will be downloaded to `checkpoints` directory.Inference itself is a single script invocation with three arguments: the checkpoint directory, an image file, and a prompt. The README gives this exact example.
python predict.py --model-path /path/to/checkpoint-dir \
--image-file /path/to/image.png \
--prompt "Describe the image."Expect a text completion describing the image. If the model path points at a stage 2 checkpoint rather than stage 3, you are running the intermediate model, and the README gives no guidance on which stage is appropriate for which task.
Exporting FastVLM to CoreML for Apple Silicon and iOS
The PyTorch checkpoints do not run on Apple Silicon as-is. The README states that checkpoints have to be exported to a format suitable for Apple Silicon, and points to the model_export/ subfolder for the instructions and code. That subfolder has its own README, which is where the actual conversion steps live; the top-level README does not reproduce them.
Apple publishes three pre-exported models so you can skip the conversion on a first pass: fastvlm_0.5b_stage3 in fp16, fastvlm_1.5b_stage3 in int8, and fastvlm_7b_stage3 in int4. Note the quantisation pattern: the larger the model, the more aggressive the precision reduction. That is a sensible memory trade, but the README does not report accuracy deltas for the quantised versions, so you cannot tell from this repository what int4 costs you on the 7B model.
The README also encourages exporting your own model at a quantisation level of your choosing. For on-device use, the app/ subfolder covers running inference on iPhone, iPad, or Mac, and the README describes a demo iOS app built to show the model's performance on a mobile device. The app is a demonstration, not a product; the README does not describe an App Store distribution path or a stable API surface.
Where FastVLM Is the Wrong Tool
The most concrete limitation is training. The README is explicit that the codebase used to train FastVLM variants is LLaVA, and that anyone wanting to train or finetune their own variants should follow the instructions in the LLaVA repository. This repository provides inference instructions only. If your goal is to adapt the model to a domain-specific dataset, you are not working in this repository at all.
The second limitation is platform coupling. The fastest path the project documents is the CoreML export route for Apple Silicon, and the pre-exported checkpoints exist for that route. Nothing in the README describes a CUDA-optimised deployment, a TensorRT path, a server with batching, or a quantised format for non-Apple accelerators. If your inference fleet is NVIDIA hardware, the PyTorch path is what you get, and the README makes no latency claims about it.
The third is the licence. The repository's licence identifier is NOASSERTION, and the README directs you to two separate files: LICENSE for the code and LICENSE_MODEL for the released models. Two files means two sets of terms. Anyone planning commercial deployment needs to read both before building on the weights.
FastVLM Versus LLaVA-OneVision and Cambrian-1
The README's own comparisons name LLaVA-OneVision-0.5B and Cambrian-1-8B as baselines, and the difference in approach is worth stating plainly. LLaVA-OneVision and Cambrian-1 scale up the visual side: more tokens per image, multiple image encoders in Cambrian-1's case, and correspondingly heavier prefill. FastVLM goes the other direction, using a single encoder that emits fewer tokens.
That is a real architectural disagreement, not a packaging difference. Cambrian-1-8B combines several vision encoders to cover different visual competencies; FastVLM's 7B variant uses one encoder and claims a 7.9x faster time-to-first-token against it. Whether the single-encoder design holds up on the tasks where Cambrian-1's ensemble helps is not something this repository answers.
If your constraint is throughput on a fixed accelerator budget, the token-count reduction is the relevant axis. If your constraint is maximum accuracy on a benchmark where ensemble encoders win, the FastVLM trade may not pay off. The README reports latency ratios, not the accuracy numbers behind them; those are in the paper.
Maintenance, Upgrade Cost and Licence Exposure
The last push to the default branch was on 2026-09-11. There are no retrieved releases, so there is no versioned artefact to pin against; you track the main branch or you pin a commit yourself. The package version in pyproject.toml is llava 1.2.2.post1, which is the LLaVA codebase's version number rather than a FastVLM-specific one, so it tells you little about the state of the FastVLM additions.
Upgrade cost is dominated by the pinned dependency set. torch, torchvision, transformers, tokenizers, and coremltools are all pinned to exact versions, and peft is constrained to >=0.10.0,<0.14.0. Moving any of those forward means testing the export path as well as the PyTorch path, because coremltools conversions are sensitive to the model graph they receive. The optional train extra pulls deepspeed 0.13.1, ninja, and wandb, but the README routes training to LLaVA, so that extra is inherited rather than exercised here.
On licensing, this is not legal advice. The repository identifier is NOASSERTION, the code licence lives in LICENSE, the model licence lives in LICENSE_MODEL, and the README asks you to check both before using the provided code and the released models. The practical implication is that you cannot assume the Apache 2.0 classifier in pyproject.toml covers the weights.
Editorial conclusion
Adopt FastVLM if you need a vision-language model with a small encoder and fast time-to-first-token, and you are comfortable driving inference through predict.py or exporting to CoreML for Apple Silicon. Do not adopt it if you need a training recipe: the README points to the LLaVA codebase for that, and this repository only ships inference instructions. Before you commit, verify the LICENSE and LICENSE_MODEL files, since the repository carries an NOASSERTION licence identifier and the two files may impose different terms on code and weights.
Frequently asked questions
Which FastVLM model is the fastest?
The README frames its headline claim around the smallest variant, FastVLM-0.5B, which it says outperforms LLaVA-OneVision-0.5B with 85x faster time-to-first-token. The 7B variants are compared at 7.9x faster time-to-first-token against Cambrian-1-8B. Those figures come from the paper and are not independently verified in the repository.
How do I install and run FastVLM?
The README creates a conda environment with Python 3.10, runs pip install -e . from the repository root, downloads checkpoints with bash get_models.sh, and then runs predict.py with --model-path, --image-file, and --prompt. The package installs under the name llava, not fastvlm.
Does FastVLM work on Apple Silicon and iPhone?
PyTorch checkpoints must first be exported to an Apple Silicon compatible format using the code in the model_export/ subfolder, which has its own README. Apple also publishes three pre-exported models at fp16, int8, and int4. Running on iPhone, iPad, or Mac is covered by the app/ subfolder.
Can I train or finetune FastVLM in this repository?
No. The README states that LLaVA is used to train FastVLM variants and directs anyone wanting to train or finetune their own variants to follow the instructions in the LLaVA codebase. This repository provides inference instructions only.
What licence applies to FastVLM models and code?
The repository's licence identifier is NOASSERTION, and the README points to LICENSE for the code and LICENSE_MODEL for the released models. The README asks you to check both before using the provided code and the released models, so the two may not carry identical terms.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/apple-aiml-research-ml-fastvlm)