# MLX: Apple's Array Framework for Machine Learning on Apple Silicon

> MLX is an array framework from Apple machine learning research with a NumPy-style Python API, composable function transformations and a unified memory model. It fits Mac-based research and small-scale training, and is the wrong tool if you need CUDA-only kernels.

**ml-explore/mlx** — MLX: An array framework for Apple silicon. Composable function transformations: MLX supports composable function transformations for automatic differentiation, automatic vectorization, and computation graph optimization.

- Repository: https://github.com/ml-explore/mlx
- Website: https://ml-explore.github.io/mlx/
- Stars: 28,552 · Forks: 2,277
- Language: C++
- License: MIT
- Published: 2026-08-08 · Updated: 2026-08-18 · Language: en
- Canonical page: https://hysenlabs.com/projects/ml-explore-mlx

## What MLX solves and who it is actually for

MLX is an array framework for machine learning on Apple silicon, released by Apple machine learning research. The README is explicit about the audience: it is "designed by machine learning researchers for machine learning researchers," and the stated goal is to make the framework easy for researchers to extend rather than to serve as a production serving stack.

The concrete problem it addresses is friction. If you write NumPy code on a Mac, you get CPU execution and explicit data movement when you want the GPU. If you write PyTorch or JAX code, you get GPU execution but you inherit a framework whose abstractions were shaped by discrete accelerator memory. MLX targets the middle ground: a NumPy-shaped API that runs on the GPU without you rewriting array code.

The secondary audience is people who want to read and modify the framework itself. The repository ships C++, C and Swift APIs that mirror the Python one, plus higher-level packages named mlx.nn and mlx.optimizers whose APIs follow PyTorch. That layering matters if you plan to write a custom layer or optimizer, because you are not forced into a Python-only extension path.

## Unified memory, lazy evaluation and dynamic graphs

The design decision that separates MLX from most array frameworks is the unified memory model. The README states that arrays live in shared memory and that operations on MLX arrays can run on any supported device without transferring data. On Apple silicon the CPU and GPU address the same physical memory, so a framework built for that hardware does not need an explicit device-placement step. This is the most consequential difference from PyTorch, where device placement is a first-class part of the programming model.

Computation is lazy. Arrays are only materialized when needed, which means a chain of operations builds a graph and the framework decides when to evaluate it. Dynamic graph construction is the companion property: changing the shapes of function arguments does not trigger slow recompilations, and the README says debugging stays simple as a result. That contrasts with ahead-of-time tracing approaches where a shape change is a recompile.

Function transformations are composable, covering automatic differentiation, automatic vectorization and computation graph optimization. The README names NumPy, PyTorch, JAX and ArrayFire as inspirations, and the transformation story is the JAX-like part of that inheritance. Supported devices are currently the CPU and the GPU. The README does not describe a distributed or multi-node execution path, so treat single-machine use as the documented scope.

## Installing MLX and running a first array operation

MLX is published on PyPI. On macOS the README gives a single command:

```bash
pip install mlx
```

After installing, the README points to the quick start guide in the documentation for usage. The Python API follows NumPy closely, and the README states that MLX also has higher-level packages like mlx.nn and mlx.optimizers with APIs that closely follow PyTorch. The README does not reproduce an array-construction snippet, so the quick start guide is the place to look for the first working example.

Linux users have two documented options. For the CUDA backend, the README gives:

```bash
pip install mlx[cuda]
```

For a CPU-only Linux package, it gives:

```bash
pip install mlx[cpu]
```

The distinction matters because the base mlx package targets macOS. If you install the plain package on Linux and expect GPU execution, you have installed the wrong extra. The README points to the documentation for building the C++ and Python APIs from source, and the repository carries a top-level CMakeLists.txt, a cmake directory and an examples/cmake_project directory for that path.

## The examples repository is where the real workloads live

The core repository is the framework, not a model zoo. The README directs readers to a separate MLX examples repository for actual workloads, and lists four: transformer language model training, large-scale text generation with LLaMA plus finetuning with LoRA, image generation with Stable Diffusion, and speech recognition with OpenAI's Whisper.

That split is worth understanding before you evaluate MLX. Installing the framework gives you arrays, transformations and neural network layers. It does not give you a tokenizer, a checkpoint loader or a training loop. If your goal is to run a quantized language model locally, the framework is a dependency of that workflow rather than the workflow itself.

The in-repository examples directory is narrower in scope. It contains examples/cmake_project, examples/cpp, examples/export, examples/extensions and examples/python. These are integration examples: how to consume MLX from C++ via CMake, how to write an extension, how to export. They are useful when you are embedding MLX in a larger C++ application, and less useful when you are trying to train a model.

## Where MLX is the wrong choice

The most obvious limitation is hardware. The README frames MLX as a framework for Apple silicon, and the supported device list is CPU and GPU. Linux is addressed through CUDA and CPU-only extras, but the project's identity and documentation are built around the Mac. If your team's training happens on a rented multi-GPU Linux cluster, MLX is not the framework that work should be written in.

The second limitation is ecosystem depth. PyTorch has years of third-party libraries, pretrained checkpoint conventions and deployment tooling. MLX has a NumPy-style API, PyTorch-style mlx.nn and mlx.optimizers, and a set of examples. The README does not claim parity, and the API resemblance is a migration convenience rather than a guarantee that arbitrary PyTorch code will port.

The third is operational. The README documents installation and points at source builds, but it does not document a serving story, a distributed training story, or a rollback procedure for a bad release. If you need those guarantees written down before you adopt a framework, this README will not supply them. Version pinning is your responsibility, and the release history shows a steady cadence of patch releases, so pinning is not optional in a reproducible environment.

## How MLX differs from JAX and PyTorch

The README itself names the inspirations, which makes the comparison fair game. Against JAX, the overlap is the transformation model: composable automatic differentiation, automatic vectorization and graph optimization. The difference is the execution target and the memory model. JAX compiles traced functions for accelerators with discrete memory and is typically deployed on Linux hardware. MLX evaluates lazily on unified memory and treats the Mac as the primary platform. If you like JAX's programming model but work on a laptop, MLX is the closer fit; if you need XLA's backend coverage, it is not.

Against PyTorch, the difference is device management. PyTorch makes device placement explicit and has a mature distributed story. MLX removes the placement step because arrays are shared, and the README does not describe an equivalent distributed layer. PyTorch's surface is broader and better documented, and its ecosystem is far larger. The trade is that MLX code on a Mac avoids the host-to-device copies that PyTorch pays for on the same hardware.

Against NumPy, the difference is that NumPy is CPU-only and eager. MLX keeps the array semantics but adds GPU execution, laziness and transformations. If your code is pure NumPy and fast enough, switching buys you little.

## Licence, maintenance and upgrade cost

MLX is released under the MIT licence, which is permissive and places few restrictions on redistribution or modification. The repository includes a LICENSE file at the top level and a CITATION.cff, and the README provides a BibTeX entry for academic citation. The MIT terms apply to the framework code; the examples repository and any model weights you load through it may carry separate terms, and the README does not address that.

On maintenance, the most recent push to the default branch was on 2026-08-25, and the release listed for that same date is v0.32.2, following v0.32.1 on 2026-08-18 and v0.32.0 on 2026-07-07. The project is not archived. The patch releases arriving roughly a week apart within a minor line suggest active bug-fix work, but that is an observation about release timing, not a statement about long-term support commitments, which the README does not make.

Upgrade cost is the practical concern. The version string is generated at build time from mlx/version.h, and the setup.py logic appends a dev date stamp and a short git hash unless the PYPI_RELEASE environment variable is set. That means a locally built wheel and a PyPI wheel can carry different version strings for the same source. Pin exact versions in your environment files and read the release notes before moving a minor version, because the README does not document a deprecation policy.

## Conclusion

Adopt MLX if you are doing research or small-scale training on a Mac and want NumPy and PyTorch style APIs with lazy evaluation and no host-to-device copies. Do not adopt it if your workflow depends on CUDA-only custom kernels or a mature multi-node distributed training stack, because the README lists only CPU and GPU as supported devices and says nothing about distributed execution. Before committing, verify that the Python API surface you need exists in the version you install, confirm the mlx[cuda] or mlx[cpu] extra matches your Linux toolchain, and check that the examples repository already covers your workload type.

## FAQ

### What is MLX in machine learning?

MLX is an array framework for machine learning on Apple silicon, from Apple machine learning research. It offers a Python API that follows NumPy, plus C++, C and Swift APIs, and higher-level mlx.nn and mlx.optimizers packages that follow PyTorch.

### Is MLX made by Apple?

Yes. The README states that MLX is brought to you by Apple machine learning research, and the citation entry credits Awni Hannun, Jagrit Digani, Angelos Katharopoulos and Ronan Collobert as equal contributors.

### What does "mlx" mean?

The README does not expand the name into a phrase. It is used throughout as the project name for the array framework, with related packages named mlx.nn, mlx.optimizers and mlx.core in the Python API.

### What are the differences between MLX-lm and MLX?

The README describes MLX as the array framework itself, and points to a separate MLX examples repository for language model work such as LLaMA text generation and LoRA finetuning. The README does not document an MLX-lm package, so the distinction between the framework and the language model tooling is not spelled out there.

## Sources

- [Official documentation](https://ml-explore.github.io/mlx/)
- [Official README](https://github.com/ml-explore/mlx#readme)
- [Project repository](https://github.com/ml-explore/mlx)
- [Release notes](https://github.com/ml-explore/mlx/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/ml-explore-mlx
