Model or dataset
apple-aiml-research/ml-aim avatar
apple-aiml-research/ml-aim

apple-aiml-research/ml-aim: one AIMv2 checkpoint across torch, MLX and JAX, with three return shapes

This repository provides the code and model checkpoints for AIMv1 and AIMv2 research projects.

1,424 stars73 forksPythonNOASSERTION

At a glance

What is it?
The code and checkpoints for a family of autoregressive vision encoders, published with three inference backends behind one loader function whose return value differs per backend, an AIMv2 example that imports its transforms from AIMv1, and a recommended Hugging Face path that turns on remote code execution.
Who is it for?
Adopt this if you want an autoregressive vision encoder as a frozen feature extractor and you are deciding between frameworks late, because one checkpoint loading through PyTorch, MLX and JAX is genuinely useful when a model has to run on a laptop and in a serving stack.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository last received commits 18 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 20, 2026, and from our analysis. They are not legal advice.

Editorial analysis

One loader, three backends, and a return value that changes shape

The model gallery in the README carries four badges, and three of them are inference frameworks. There is a PyTorch path, a JAX path and an MLX path, plus a Hugging Face collection. The interesting part is that the three examples are almost identical and diverge in exactly two places. The PyTorch and MLX examples both call the loader and assign a single object, passing a backend argument naming the framework. The JAX example calls the same function with the same model identifier and receives two objects, a model and a set of parameters, and then has to invoke the model by applying it to a parameters dictionary. That difference is not a documentation shortcut, it is functional programming. A JAX model is a pure function of parameters and inputs, so the parameters are not attached to it and you supply them at apply time. Two consequences follow for anyone writing code. First, a backend-agnostic wrapper has to special-case JAX, because the return arity changes. Second, the JAX path is the one that fits a serving stack where parameters are sharded across devices, while the other two give you a module you call directly. The install route reflects the same split, since the package is installed per subdirectory and MLX support is an extra install rather than part of the default:

bash
pip install mlx

The three-way split is unusual for a research release and it is the most reusable thing in the repository, since it lets the same weights serve an Apple-silicon experiment and a server-side inference path without a conversion step.

An AIMv2 example imports its transforms from AIMv1

Read the import lines in the examples and you find something that looks like a mistake. The AIMv2 model gallery example, whichever backend it uses, imports its validation transform from the first generation: the import is from the v1 package's torch data module, and the checkpoint it loads is named for version 2. The preprocessing pipeline for a v2 checkpoint therefore lives in the v1 subpackage. That is a real coupling and it has practical consequences. It means installing the v2 subdirectory alone is not enough for the documented example to run, because the transform import reaches across into v1, so in practice you install both. It also means the two generations share a data pipeline by design, which is sensible while the input pipeline is unchanged, and which will become a liability if v2's preprocessing ever diverges from v1's. If you are reading this to understand the dependency footprint, note that the two install commands install two subdirectories of one repository, and the example code needs the code from both. The repository structure is the simplest possible, with the top level containing only the two subdirectories plus a licence, an acknowledgements file, a readme, an editor configuration and a pre-commit configuration. There is no package index, no build configuration and no release process in the tree, so the installation story is entirely the two pip commands.

The install command points at a different organisation path

Compare two strings that should be the same and are not. This repository is published under the organisation named apple-aiml-research. Every installation command in the README installs from a different path, one naming the apple organisation directly:

bash
pip install 'git+https://github.com/apple/ml-aim.git#subdirectory=aim-v1'
pip install 'git+https://github.com/apple/ml-aim.git#subdirectory=aim-v2'

The most likely explanation is an organisation rename, since an enterprise account that changes its name keeps its repositories reachable at a redirecting path, and the README was written before or during the change. The practical consequences are small but real. The commands work today, which is what a reader needs, but a script that pins the git path is pinned to a name that is no longer the canonical one, and anyone auditing where their dependencies come from will see two names for the same repository. The larger point is the shape of the install. There is no package on an index. You are asking pip to clone a repository and take one subdirectory, which means your build now depends on the current state of the default branch rather than on a released version, because there are no GitHub releases in this project and no version to pin. The instructions do say to install PyTorch from its own official instructions first, which is the right dependency to take from its own installer. Everything else you need, an image library and a transformer library depending on the route, you install yourself.

The recommended Hugging Face path runs code from the model repository

The checkpoints are also published as a collection on the Hugging Face Hub, and the README shows the shortest possible way to use them: an automatic image processor and an automatic model, both loaded by a name, then a processor call and a model call. That route is two imports and four lines, and it is the one most people will take. It also contains a flag you should read twice. The model is loaded with trust_remote_code set to true. That flag tells the transformers library to download and execute Python code from the model repository in order to build the model class, because AIMv2's architecture is not one the library implements natively. It is a genuine convenience, since it means the model works without a library release, and it is a genuine trust decision, because the code that runs is whatever is in that repository at that moment. Before you put that line in a production service, read the code in the model repository and understand what it does at import time. The contrast with the pip route is instructive, because there you install code from the source repository you have already reviewed, and the architecture lives in the aim package. The two routes have different trust models and the same weights, which is worth remembering when a team standardises on one of them.

The name says 3B and the table says 2.7B

The checkpoint tables are the most useful part of the README and they carry one detail that matters operationally. Each row gives a model identifier, a parameter count, a top-1 accuracy figure, links to the model and to the weights file, and the backbone. Reading down the 224-pixel table, the identifiers are named by capacity tier and the parameter counts do not match the numbers in the names. The large checkpoint at this resolution has 0.3B parameters. The huge one has 0.6B. The one named 1B has 1.2B. The one named 3B has 2.7B. So the size in the identifier is a tier name, not a count, and if you are sizing a machine, planning a download or writing a memory budget, the parameter column is the number that matters and the name is not. The accuracy figures rise with size across those four rows, from 86.6 to 87.5 to 88.1 to 88.5, which is the shape you would expect, and it also shows diminishing return: going from 0.3B to 0.6B is worth nearly a point, while going from 1.2B to 2.7B is worth four tenths. One number in the overview needs its own caveat. The headline says the 3B model achieves 89.5 on ImageNet using a frozen trunk, which is a higher figure than the table entry for that checkpoint, because the headline describes a frozen-trunk evaluation and the table column is not described that way. Do not treat those two numbers as the same measurement.

Six checkpoint families, and one of them is the recommended one

The gallery is not a list of sizes, it is a list of purposes, and reading the six headings tells you how the maintainers expect these models to be used. There are three resolution variants, at 224 pixels, 336 pixels and 448 pixels, plus a variant at native resolution, which is the one that matters if your input aspect ratio is unusual. Then there are two that are not about resolution at all. A distilled ViT-Large variant is marked as recommended for multimodal applications, which is the most actionable line in the README, because it tells you that for a multimodal task the distilled model is the one to start with rather than the largest checkpoint. And a zero-shot adapted variant exists, which means a checkpoint tuned for a task without labelled data. The recommendation matters because the two axes, size and adaptation, are independent, and picking the biggest model is the most common mistake. A frozen trunk changes the arithmetic entirely, since the encoder is what you are shipping and the head is what you train, so the size that reaches production is the trunk and the accuracy figure that matters is the one measured with it frozen. The resolution axis is the other half of the decision, and the example in the README uses a large checkpoint at 336 pixels, which is a middle choice on both axes and a reasonable starting point for a feature extraction experiment.

Two papers in one repository, and a licence to read before you ship

The repository holds two generations of work, and the README is explicit about how to tell them apart. AIMv1 is described as scalable pre-training of large autoregressive image models, published at ICML 2024, and the README adds a note that if you are looking for the original AIM model you should refer to a readme inside the v1 subdirectory. AIMv2 is described as multimodal autoregressive pre-training of large vision encoders, published at CVPR 2025 where it is marked as a highlight, and the two generations have separate arXiv entries, separate badges and separate readmes. That structure is why the install is per subdirectory, and it is a sensible layout for a research group that iterates on a model family: the older generation stays reproducible while the new one moves. Two things to note about the project rather than the model. There are no GitHub releases, so the checkpoints are versioned by name rather than by tag, and the code is installed from the default branch, so there is no code version to record alongside a checkpoint. And the licence is one GitHub cannot classify automatically, which is expected for a commercial research organisation publishing model weights and is not a warning sign in itself. It does mean the terms are not the ones you can guess from a standard identifier, and for weights rather than code that matters: read the file before a checkpoint goes anywhere near a product.

Editorial conclusion

Adopt this if you want an autoregressive vision encoder as a frozen feature extractor and you are deciding between frameworks late, because one checkpoint loading through PyTorch, MLX and JAX is genuinely useful when a model has to run on a laptop and in a serving stack. Do not adopt it expecting a drop-in replacement for a contrastive encoder, since the training objective is different and the reported comparisons against CLIP, SigLIP and DINOv2 are the paper's own claims on its own benchmarks. Three things to check first. Read the parameter count column rather than the checkpoint name, since a name reading 3B corresponds to 2.7B parameters in the published table. Decide which installation route you want, because the pip route installs from one organisation path while the repository sits under another, and the Hugging Face route pulls in trust_remote_code and therefore executes code from the model repository. And note the licence is a custom one that GitHub cannot classify, so read it before you use a checkpoint in a product.

Frequently asked questions

How do I install the AIM vision model code?

Install PyTorch from its own installation instructions first, then install from a subdirectory of the repository: pip install with a git URL and a subdirectory parameter for aim-v1, and the same for aim-v2. MLX support on Apple silicon is a separate pip install mlx.

Which inference backends does ml-aim support?

PyTorch, JAX and MLX, all through one loader function that takes a backend argument and a checkpoint name. The JAX path returns a model and a parameters object that you apply, while the torch and mlx paths return a single callable object.

How many parameters does the aimv2-3B checkpoint have?

The published table gives 2.7B parameters for the checkpoint named aimv2-3B-patch14-224, so the size in the checkpoint name is a capacity tier rather than a parameter count. The table is the place to read actual counts.

Which AIMv2 checkpoint does the README recommend for multimodal work?

The distilled ViT-Large variant is marked as recommended for multimodal applications. The gallery also offers 224, 336, 448 and native resolution variants, plus a zero-shot adapted variant.

What does trust_remote_code do in the Hugging Face example for AIMv2?

It tells the transformers library to download and execute Python code from the model repository in order to build the model class, which is needed because the architecture is not natively implemented. That is a trust decision about code from the model repository, as opposed to the pip route which installs code from the source repository you have read.

Official sources

  1. apple-aiml-research/ml-aim on GitHub
  2. Issues
  3. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/apple-aiml-research-ml-aim.svg)](https://hysenlabs.com/projects/apple-aiml-research-ml-aim)