Open-source project
apple-aiml-research/ml-ane-transformers avatar
apple-aiml-research/ml-ane-transformers

ane_transformers: Apple's Reference Transformer for the Neural Engine

Reference implementation of the Transformer architecture optimized for Apple Neural Engine (ANE)

2,743 stars93 forksPythonNOASSERTION

At a glance

What is it?
A PyTorch reference implementation that restructures Transformer blocks so they map cleanly onto the Apple Neural Engine, plus an optimized Hugging Face distilbert class. Useful if you ship Core ML models to A14 or M1 hardware and your baseline conversion is leaving the ANE idle.
Who is it for?
Adopt ane_transformers if you are converting a Transformer to Core ML for A14 or newer and M1 or newer hardware and you want Apple's own reference for how the block should be shaped before conversion. Do not adopt it if you are targeting older devices, if you need a maintained general-purpose conversion library, or if your model is not a Transformer.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository last received commits 19 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 24, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem ane_transformers actually solves

A Transformer written the ordinary PyTorch way does not map onto the Apple Neural Engine. The ANE is a fixed-function accelerator, and the README's framing is that baseline implementations leave performance on the table: Apple claims up to 10 times faster execution and 14 times lower peak memory consumption compared to baseline implementations, on A14 or newer and M1 or newer chips. Those numbers come from Apple's own research article, not from independent measurement, and the README is explicit that the tutorial's own configuration (sequence length 128, batch size 1) improves latency by a factor of 2.84 rather than 10. The larger figures apply at sequence lengths such as 512 and batch sizes such as 8.

The audience is narrow and worth stating plainly. This is for engineers who already intend to ship a Transformer inside an iOS or macOS app through Core ML, and who are willing to restructure the model before conversion. It is a reference implementation, in Apple's words, not a training framework and not a runtime. If you never touch Core ML, the package has nothing to offer you.

What the optimized model changes, and why the operation count grows

The package is split into two parts. ane_transformers.reference is a standalone reference implementation of the architecture. ane_transformers.huggingface holds optimized versions of Hugging Face model classes, with distilbert as the worked example. The optimized model is mathematically equivalent to the baseline, which is what makes the tutorial's load_state_dict step possible: you build the baseline, build the optimized model from the same config, then copy parameters across.

The trade-off is visible in the numbers the README gives. The optimized model contains 606 operations, and 4 of them still execute on the CPU because they are embedding lookup operations that are more efficient there for this configuration. The README also notes that load and compilation times increase because the operation count grows. Apple's position is that these are one-time costs and that loading the model asynchronously keeps them away from the user. That is a reasonable position for a model loaded once at app start, and a poor one for anything that reloads models frequently.

One detail worth internalizing: the shape of the win depends on your workload. The 2.84x figure for the tutorial's 128-token, batch-1 case versus the up-to-10x figure at 512 tokens and batch 8 means short-sequence, single-item inference is the least favorable case for this approach.

Installing ane_transformers and converting distilbert

The README gives two install paths. The fastest is a plain pip install; the alternative is an editable install from a clone. The pinned dependency range matters here: setup.py requires torch>=1.10.0,<=1.11.0 alongside coremltools>=5.2.0, transformers>=4.18.0 and protobuf>=3.1.0,<=3.20.1. Installing into an environment with a newer torch will conflict.

bash
pip install ane_transformers

If the wheel build fails with an error about a missing Rust compiler while building tokenizers, the README points to a Hugging Face issue thread rather than offering a fix of its own. For a local checkout, the editable install is:

bash
pip install -e .

The tutorial then builds the baseline model and the optimized counterpart, copying weights between them. Note return_dict=False and torchscript=True on the baseline, and that the optimized class is constructed from the baseline's config:

python
import transformers
from ane_transformers.huggingface import distilbert as ane_distilbert

model_name = "distilbert-base-uncased-finetuned-sst-2-english"
baseline_model = transformers.AutoModelForSequenceClassification.from_pretrained(
    model_name, return_dict=False, torchscript=True
).eval()
optimized_model = ane_distilbert.DistilBertForSequenceClassification(
    baseline_model.config
).eval()
optimized_model.load_state_dict(baseline_model.state_dict())

Tokenization uses max_length=128 and padding="max_length", which fixes the traced input shape. The model is then traced with torch.jit.trace and passed to coremltools with convert_to="mlprogram" and compute_units=ct.ComputeUnit.ALL. The saved artifact is named HuggingFace_ane_transformers_distilbert_seqLen128_batchSize1.mlpackage. To check performance, the README says to add the package as a resource in an Xcode project and use the Performance tab to generate a report on a locally available device. There is no command-line benchmark in the repository; Xcode is the measurement path.

Where the reference implementation stops being the right tool

The pinned torch range is the first hard boundary. torch>=1.10.0,<=1.11.0 is a narrow window, and anyone on a current PyTorch release cannot install this package without either a separate environment or a dependency override that the README does not describe. The last tagged release is v0.1.3 from 2022-08-09, and setup.py still classifies the project as Development Status :: 4 - Beta with Python classifiers running only to 3.9. The repository's last push was on 2026-09-11, so it is not abandoned, but the dependency pins have not moved with the ecosystem.

Second, the device floor is real. The README states the spec as M1 or newer for Mac and A14 or newer for iPhone and iPad, and notes that the unit tests print a warning when run on devices outside that spec. It also says a model generated on an out-of-spec Mac should still work on in-spec devices, which is a useful escape hatch but not a performance guarantee.

Third, and most important: this is a reference, not a product. The huggingface module covers distilbert as a demonstration of the optimization principles. If your model is not distilbert, you are reading the reference implementation to learn the approach and applying it yourself. Nothing in the README suggests a general-purpose converter that will restructure an arbitrary Transformer for you.

ane_transformers compared with converting a stock model through coremltools

The obvious alternative is to skip this package entirely: take a stock Hugging Face model, trace it, and hand it to coremltools with the same convert_to="mlprogram" and compute_units=ct.ComputeUnit.ALL settings. That path works, requires no extra dependency, and is not bound to torch 1.11. What you give up is the restructuring. The whole point of the reference implementation is that the baseline block shape does not partition well onto the ANE, and Apple's claim of up to 10x latency and 14x peak memory improvement is measured against baseline implementations, not against a differently optimized one.

So the difference is not "wrapper versus no wrapper". It is whether you are willing to change the model's internal structure before conversion. If you are, the reference implementation shows you the target shape and the distilbert class shows a working example. If you are not, coremltools alone is the more honest choice, and you should measure your own latency rather than assuming the published multipliers apply to your model.

Versions, licence and what upgrading costs

The release history is short: v0.1.1 on 2022-06-07, v0.1.2 on 2022-07-30, v0.1.3 on 2022-08-09. There has been no tagged release since, although the repository has received pushes as recently as 2026-09-11. That combination means you should treat the pinned dependency range in setup.py as the authoritative statement of what this code was built against, and check the repository directly rather than assuming a newer release exists.

The licence file is LICENSE.md and the repository metadata reports the licence as NOASSERTION, meaning GitHub could not match it to a recognized SPDX identifier. That is a signal to read LICENSE.md yourself before shipping anything derived from this code in a commercial app, particularly since the package is published by Apple Inc. and the optimized model classes are derived from Hugging Face model definitions, which carry their own licence. Nothing here is legal advice; the practical point is that the licence is not a one-word answer you can copy from a badge.

Upgrade cost is dominated by the torch pin. Moving to a newer PyTorch means either verifying that the reference implementation still behaves identically on the newer version or maintaining a fork, and the README offers no migration guidance. The Makefile's test target runs ane_transformers/reference/test_transformer.py and ane_transformers/huggingface/test_distilbert.py, which is the check to run first if you do move versions.

Editorial conclusion

Adopt ane_transformers if you are converting a Transformer to Core ML for A14 or newer and M1 or newer hardware and you want Apple's own reference for how the block should be shaped before conversion. Do not adopt it if you are targeting older devices, if you need a maintained general-purpose conversion library, or if your model is not a Transformer. Before committing, verify that your torch version falls inside the pinned range in setup.py, that your target sequence length and batch size are the ones the README quotes speedups for, and that you can reproduce the tutorial's .mlpackage on your own device before rewriting any model code around it.

Frequently asked questions

Is an LLM just a transformer?

The README frames ane_transformers as a reference PyTorch implementation of the Transformer architecture optimized for the Apple Neural Engine, and its worked example is a distilbert sequence classification model rather than a generative language model. It does not make claims about how large language models relate to the Transformer architecture in general.

Why is GPT called a transformer?

The repository does not discuss GPT or explain the naming of that model family. What it does describe is the Transformer architecture as implemented in ane_transformers.reference and the optimized Hugging Face model classes in ane_transformers.huggingface.

What are the main 3 types of ML models?

This question is not addressed anywhere in the repository. The README covers a Transformer reference implementation for ANE deployment, the distilbert tutorial, unit tests and installation troubleshooting, and does not categorize machine learning models.

Is transformers a part of NLP?

The repository treats transformers as a Python package dependency: setup.py requires transformers>=4.18.0, and the tutorial imports it to load the distilbert tokenizer and baseline model. It does not discuss natural language processing as a field.

Official sources

  1. apple-aiml-research/ml-ane-transformers on GitHub
  2. Issues
  3. README
  4. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/apple-aiml-research-ml-ane-transformers.svg)](https://hysenlabs.com/projects/apple-aiml-research-ml-ane-transformers)