ANE Training: Running Backpropagation on Apple's Neural Engine
Training neural networks on Apple Neural Engine via reverse-engineered private APIs
At a glance
- What is it?
- maderix/ANE is a research proof-of-concept that trains transformer models directly on the Apple Neural Engine by reverse-engineering Apple's undocumented _ANEClient and _ANECompiler APIs. The author is direct: ANE utilization currently sits at 5 to 9 percent of theoretical peak, and this is not a production framework.
- Who is it for?
- maderix/ANE is the right choice for hardware researchers and compiler engineers who want to study how the ANE handles gradient computation at the firmware level. It is not appropriate as a dependency in any production system because the private APIs it uses can change or break without notice in any macOS update.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Activity is slowing. The repository last received commits 6 months ago.
- What is it written in?
- Mainly Objective-C, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What ANE Training Demonstrates and Who It Is For
Apple's Neural Engine on M4-series chips delivers 15.8 TFLOPS of FP16 compute. CoreML, Apple's public framework, uses the ANE for inference but exposes no training API. The maderix/ANE project exists to show that this restriction is a software decision, not a hardware limitation.
The project reverse-engineers the undocumented _ANEClient and _ANECompiler private APIs to run full forward and backward passes for transformer models directly on the ANE. No CoreML, no Metal, no GPU is involved in the compute path. The README describes the project as a proof of concept for ANE training via these private APIs and as a reference for anyone exploring direct ANE access outside CoreML.
The primary audience is compiler researchers and hardware engineers who study NPU compute graphs, memory hierarchy, and kernel scheduling at the firmware level. It may also interest engineers working on edge AI optimization for Apple Silicon. It is not a library you import into a product, and the author states this plainly in the README. The project also documents the performance ceiling honestly, which is rarer than it should be for this kind of research work.
The Private API Bridge: _ANEClient, _ANECompiler, and MIL Programs
Two undocumented Apple-internal Objective-C classes form the foundation of the project. _ANEClient dispatches compute graphs to the hardware. _ANECompiler takes a program in MIL (Model Intermediate Language) format and compiles it into an ANE-native artifact. Apple uses both internally when CoreML prepares models for ANE execution; neither appears in any public SDK or developer documentation.
MIL is the intermediate representation Apple's own tools use when converting models for deployment. The maderix/ANE project constructs MIL programs at runtime, representing each transformer component as a graph of MIL operations. The QKV projection, scaled dot-product attention, and output projection map to one kernel. The SwiGLU feed-forward block maps to another. Each backward pass gets dedicated kernels sized to fit within the ANE's L2 SRAM budget.
A non-obvious operating system constraint shapes how the project must run: each process can compile at most approximately 119 ANE models before hitting a per-process ceiling that is not documented in any Apple reference. The project works around this by calling exec() to restart the process, which resets the counter while checkpoint and resume state persists across the restart. This is a property of the OS runtime, not a design choice the project made.
Cloning the Repository and Examining the Source Files
The repository is Objective-C research code. There is no package manager entry point, no Swift Package Manager configuration, and no precompiled binary. To work with it, clone the repository:
git clone https://github.com/maderix/ANE
cd ANEThe root-level files each serve a specific purpose. api_exploration.m holds the initial API discovery work. inmem_basic.m provides an in-memory MIL compilation proof-of-concept in minimal form. inmem_bench.m measures ANE dispatch latency. inmem_peak.m measures peak throughput using 2048x2048 matrix multiplications. ane_int8_bench.m compares INT8 W8A8 against FP16 throughput. sram_bench.m and sram_probe.m probe the ANE's L2 SRAM bandwidth and spatial layout.
The full transformer training implementations are in the training/ directory. The README directs readers to three Substack articles the author published covering reverse engineering, benchmarks, and training in sequence. The source files are designed to be read alongside those articles, not as self-contained documentation. Each file is narrow in scope, which makes the codebase approachable for study even though there is no setup guide beyond the README.
Dynamic Pipeline Architecture and Per-Layer Kernel Design
The Stories110M model (109 million parameters, 12 layers, dimension 768, multi-head attention with 12 heads) uses 6 ANE kernels per transformer layer. The sdpaFwd kernel handles QKV projection, scaled dot-product attention, and the output projection in one dispatch. The ffnFused kernel handles the SwiGLU FFN block, which combines W1, W3, a SiLU activation, and W2 in sequence. Backward pass work splits across ffnBwdW2t, ffnBwdW13t, sdpaBwd1, and sdpaBwd2: the splits reflect SRAM capacity limits, not a preference for small kernels.
The Qwen3-0.6B model (596 million parameters, grouped-query attention) requires 10 kernels per layer. The increase follows from GQA's structure: the query dimension differs from the model dimension, which requires separate woFwd, qBwd, and kvBwd kernels for the output projection and the key-value backward passes.
Weights are stored in the channel-first CPU layout that matches the ANE's IOSurface [1, C, 1, S] format. This decision eliminates transpose overhead at every kernel dispatch. dW gradient accumulation runs via cblas_sgemm on the Accelerate framework, scheduled on a serial GCD dispatch queue to overlap with ANE forward passes. Each cblas wait is deferred to the next step's forward pass, squeezing additional throughput.
INT8 W8A8 quantization halves activation bandwidth between tiles. The README reports 1.88x throughput on M4 H16G: 35.1 TOPS in INT8 versus 18.6 TOPS in FP16 for a 128x convolution with 512 channels at 64x64 spatial dimensions. Weights use constexpr_affine_dequantize so they remain int8 on device and convert to FP16 at compile time.
Measured Step Times and the Real Limits of the Current Implementation
The README reports 91 ms per step for Stories110M and 412 ms per step for Qwen3-0.6B using the dynamic pipeline, which avoids recompilation when weights change between steps. All forward passes and backward dx passes run on the ANE. Weight gradient accumulation for the dW terms runs on CPU via Accelerate's cblas_sgemm because the ANE's efficiency on those operations does not justify the dispatch overhead.
The author is precise about what is not working well. ANE utilization sits at 5 to 9 percent of the hardware's theoretical peak. Many element-wise operations fall back to the CPU because the ANE handles them at higher latency than it handles convolution and matrix workloads. The Substack articles published alongside the repository document these gaps in technical detail.
The deeper structural problem is private API dependency. Apple can change or remove _ANEClient and _ANECompiler behavior in any macOS or iOS release without any public notice. There is no compatibility table in the repository. There is no test suite for OS version differences. Code that runs correctly on one macOS version may fail silently or crash on the next.
When CoreML, MLX, or llama.cpp Are the Right Choice Instead
For production inference on Apple Silicon, CoreML is the correct tool. Apple exposes CoreML publicly, maintains backward compatibility across OS versions, and provides conversion tooling for most standard model formats. It schedules ANE use transparently during inference without requiring any private API access.
Apple's MLX framework is the most relevant comparison point for training research on Apple Silicon. MLX runs on the GPU and CPU using documented public APIs, provides a NumPy-like Python interface with autograd, and is maintained by Apple. It covers the same transformer training use case as maderix/ANE but targets the GPU and CPU rather than the ANE directly. The difference in risk profile is significant: MLX is a supported research framework with stable APIs; maderix/ANE is an ANE-specific experiment built on undocumented internals.
llama.cpp focuses on inference for large language models via Metal and CPU backends on Apple hardware. It does not attempt to use the ANE and is not trying to. For anyone whose goal is deploying a language model locally, llama.cpp provides a stable, maintained path with no private API exposure.
Maintenance Status and MIT License
The last push to the repository was on 2026-03-10. The project is not archived. The README describes the author's intent: updates happen when they discover something interesting, bug fix and benchmark contributions are welcome, feature requests will likely go unaddressed, and pull requests will be merged slowly to avoid becoming a bottleneck. The author explicitly invites forking and independent development under the MIT license.
MIT permits forking, modification, and redistribution without restriction. This aligns with the author's stated goal of making the proof of concept available for the community to extend. The license does not address API stability or forward compatibility with future OS releases.
Anyone who forks this code accepts the private API risk independently. Apple has changed the internal ANE interfaces before without public notice, and the README gives no indication of testing across macOS versions. The project is valuable for what it has already documented: that training on the ANE is technically possible at the firmware level.
Editorial conclusion
maderix/ANE is the right choice for hardware researchers and compiler engineers who want to study how the ANE handles gradient computation at the firmware level. It is not appropriate as a dependency in any production system because the private APIs it uses can change or break without notice in any macOS update. Before working with the code, verify that your chip generation and macOS version match the M4 hardware the benchmarks were gathered on: the README does not document a compatibility matrix across chip generations.
Frequently asked questions
What is Apple's Neural Engine used for?
According to the maderix/ANE README, Apple restricts the Neural Engine to inference through CoreML, and it delivers 15.8 TFLOPS of FP16 compute on M4-series chips. The maderix/ANE project shows that training is also possible when the undocumented _ANEClient and _ANECompiler APIs are accessed directly, though utilization in the current implementation sits at 5 to 9 percent of peak.
Is the Neural Engine the same as a GPU?
The Apple Neural Engine and the GPU are separate processors on Apple Silicon. The maderix/ANE project runs its compute path exclusively on the ANE with no Metal or GPU involvement, though one pipeline configuration uses IOSurface shared memory for a zero-copy hand-off from GPU prefill to ANE decode.
What private APIs does maderix/ANE use to access the Apple Neural Engine?
The project uses two undocumented Objective-C classes: _ANEClient, which dispatches compute graphs to the hardware, and _ANECompiler, which compiles MIL programs into ANE-native executables. Both are Apple-internal classes not documented in any public SDK.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/maderix-ane)