Model or dataset
lablup/mlxcel avatar
lablup/mlxcel

mlxcel: one Rust binary for MLX checkpoints, three GPU backends, one of them source-only

High-performance LLM, VLM, embedding, reranking, and audio inference for Apple Silicon and NVIDIA CUDA systems.

472 stars56 forksRustApache-2.0

At a glance

What is it?
lablup/mlxcel is a Rust CLI and server that reads MLX SafeTensors checkpoints directly, with Apple Silicon first and Linux/CUDA second. The AMD path exists only as a source build with kernel fallbacks, the serving surface is classified rather than proven against a frozen llama-server manifest, and the branch described in the README sits ahead of the v0.7.0 tag.
Who is it for?
mlxcel suits an Apple Silicon shop that wants OpenAI, Anthropic, and llama-server routes out of one native binary and is willing to track a branch instead of a tag. It does not suit a team that needs AMD GPUs in production, needs speculative decoding outside greedy sampling, or expects the 376 classified llama-server options to be interchangeable with the reference server.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 1 day ago.
What is it written in?
Mainly Rust, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 3, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The workspace root is the binary, and the gate skipped four crates

mlxcel's Cargo.toml opens with a long comment defending one absent line. There is deliberately no `default-members`, because the workspace root is itself the `mlxcel` package: a bare cargo command resolves to `-p mlxcel`, and adding `default-members` would re-point every invocation in the tree at once, including the release workflow's build, which would start compiling the default-off `mlxcel-xla` into every release build. The cost is that a command has to say `--workspace` to see all five members, which are the root, `src/lib/mlxcel-core`, `src/lib/mlxcel-mlx-pin`, `src/lib/mlxcel-surgery`, and `src/lib/mlxcel-xla`, on resolver 2.

That gap was not theoretical. Until issue 1007, the quality gate covered the root package only, leaving 1754 tests across the other members ungated, 1354 of them inside mlxcel-core, along with their whole lint surface. Both targets now pass `--workspace` explicitly.

bash
cargo build --release --target aarch64-apple-darwin --locked
make verify-test
make verify-clippy

The first line is the release build the comment warns about. One member, `mlxcel-mlx-pin`, has no production role at all: it hosts the unit tests for the MLX-pin logic in `mlxcel-core/build_support/mlx_pin.rs` so that logic can be exercised without compiling MLX C++.

Two one-liners describe the same crate and they disagree. The manifest says `High-performance LLM/VLM/VLA inference on Apple Silicon and CUDA GPUs`; the repository description says `High-performance LLM, VLM, embedding, reranking, and audio inference for Apple Silicon and NVIDIA CUDA systems`. VLA appears in one and not the other, while embedding, reranking, and audio appear in the other and not the manifest.

The same skew shows up in versioning. The manifest version is 0.7.0 and the latest tag is v0.7.0 from 2026-09-09, but the branch is described as v0.7.0 plus unreleased work, and the highlights on it cover backends and architectures that have not reached a tag. Anything written about this project from its README describes main, not the build you would get from the releases link.

A six-job build cap measured on a host shared with three servers

The Makefile's most heavily commented line is a job cap. `ROCM_JOBS ?= -j 6` is written as a whole flag so that setting `ROCM_JOBS=` to empty disables it, and the comment above it explains the measurement behind it. The gfx1151 validation host has 32 cores against 30 GiB of RAM shared with the GPU services it also runs. On 2026-09-28 three `llama-server` processes held 25.5 GiB there, leaving roughly 2 GiB for a build, and cargo's default of one rustc per core cannot fit in that. The failure mode is an OOM kill partway through compiling, which reads as an unrelated crash rather than as memory pressure.

So the number is a function of what else is resident, not of the core count. Six jobs builds this workspace when the GPU services are idle and is not enough when they are not. With the services stopped, a full `verify-rocm` never dropped below 14.8 GiB available, which is why `verify-rocm-smoke` prints available memory before it starts.

bash
ROCM_JOBS ?= -j 6
RUSTFLAGS := RUSTFLAGS="-C target-cpu=native"

That second line deserves attention before anyone packages a release. Anything built through this Makefile is tuned for the CPU of the machine that built it, so the `mlxcel` and `mlxcel-server` executables the project calls simple deployment artifacts are not the same binary on every host. The same file still expects a Python interpreter on the build machine: `WEBUI_CONTRACT_PY`, `WEBUI_BUNDLE_PY`, and `WEBUI_EVIDENCE_PY` all default to `python3`, and the file as published stops partway down a run of empty `WEBUI_*` variables. The promise of no Python in the request path holds for the request path, not for the build.

ROCm is a source build whose fused kernels fall back

Apple Silicon is the primary target, Linux with CUDA is the secondary one, and AMD is described as experimental and source-build-only. `rocm` is a cargo feature, so reaching it means compiling: the backend arrives inside the source tree as the ROCm part of mlxcelverse and is applied on the same pinned MLX commit as the Metal and CUDA builds. That single commit serving three backends is deliberate, and it is why the pin logic has its own test-only crate.

What runs there is spelled out narrowly. Affine, mxfp8, and mxfp4 checkpoints, including gpt-oss MoE experts, run on RDNA 3.5, listed as `gfx1151`. The rest arrives with conditions attached: fused kernels fall back to MLX graphs, and affine MoE models need `MLXCEL_FUSED_MOE=0` for now. Readers are sent to `docs/installation.md` for the build and the open gaps rather than being told the target is finished.

A fourth path sits beside them, an opt-in OpenXLA and IREE backend offered as an alpha development path and kept out of the default build. That is the same reason the Cargo.toml refuses to add `default-members`: `mlxcel-xla` is one of the five workspace members and stays off unless its feature is enabled. A team planning an AMD or XLA deployment should treat the result as something it compiles and validates itself, starting from the installation page rather than from a release artifact.

376 llama-server options classified rather than proven equal

Compatibility with `llama-server` is the claim the project backs with a count. A frozen manifest at b10621 classifies all 376 pinned options, routes, and native request fields, with no deferred entries, and the stated point of the classification is that nothing is quietly ignored. Native completion, embedding, tokenization, template, infill, props, slots, metrics, resumable-stream, router, and LoRA surfaces are each either implemented or given a label.

The labels are where the care sits. `mlxcel-server` accepts a broad `llama-server` flag surface and the `LLAMA_ARG_*` environment variables, and the checked-in manifest records every case as supported, aliased, not-applicable, or intentional-difference. Those last two categories are the reason to read it rather than skim it, because they are exactly where a drop-in migration stops being a drop-in.

The same binary also answers OpenAI, Anthropic, and Vertex-shaped requests. Chat Completions, Completions, Responses, Embeddings, Reranking, Audio, and Anthropic Messages sit alongside the native routes, with optional Vertex AI custom-container routing driven by the standard `AIP_*` variables. Continuous batching, prompt-prefix caching, and automatic prefix caching are on by default, while speculative decoding, KV-cache compression, router mode, live LoRA, and distributed modes appear where the model and backend support them. `docs/llama-server-compat.md` holds the exact boundary and `docs/server-features.md` the route and deployment map.

Speculative decoding probes greedy exactness before it engages

Speculative decoding here is conditional, and the project states what the condition is. The Gemma 4, Qwen, and GLM-4.7-Flash MTP paths, the LFM2 and LFM2.5 DSpark path, and the Muse Glimmer assistant path all probe greedy exactness before enabling the fast path, and the server publishes its adaptive decision at `GET /v1/internal/mtp-policy`. That endpoint is the difference between a policy and a guess, since a deployment can read the current decision instead of inferring it from token timings.

The pairings narrow the scope. LFM2.5 targets pair with LiquidAI's published DSpark drafters on `mlxcel-server` through `--draft-model`, and that pairing is greedy decoding only. A deployment sampling above zero temperature gives up the draft acceptance path for those checkpoints, whatever the policy endpoint reports. Muse Glimmer pairs with `meta-models/Muse-Glimmer-30B-assistant` on text-only requests, and its round loop measures the verify width instead of fixing it, so the accept width is a runtime measurement rather than a number in the config.

bash
mlxcel split-mtp

That command extracts the GLM-4.7-Flash drafter from the raw checkpoint, which is the step that turns one published checkpoint into a pair. Note the division of labour: the acceptance policy is measured inside the server, while the draft model itself is produced by a separate offline pass over the weights.

Three audio containers at the boundary, WAV only for most families

Audio support is where coverage stops being uniform. The compatible transcription boundary recognizes WAV, MP3, and FLAC by content rather than by extension, which is the right way to handle uploads. Decoding is narrower: Phi-4 Multimodal and Gemma 3n handle all three containers, while other current audio families and dedicated Whisper remain WAV-only. A deployment that hands MP3 or FLAC to a Whisper checkpoint through the compatible route has to convert before the request, whatever the boundary will accept.

Streaming differs too. Chat-model transcription streams one ASR delta per decoded token, tying the audio output to the decode loop rather than to a separate segmenter. On the far side of the same boundary, Kokoro provides text-to-speech.

The architecture list is wider than the audio list. Dense transformers, sparse MoE, hybrid SSM, VLM and OCR, block diffusion, embeddings, rerankers, ASR, TTS, and full-duplex speech-to-speech in the form of Nemotron VoiceChat are all represented. The retrieval and multimodal side carries its own enumeration: BERT, ModernBERT, SigLIP, Qwen3 and Qwen3-VL, Llama and Nemotron, LFM2.5, ColBERT-style models, cross-encoders, and generative rerankers, with Qwen-VL video input and Responses-native image parts.

It is also where the coverage bullet on architecture ends mid-sentence, stopping on the words `for checkpoint` after pointing at `mlxcel arch` for the binary's own catalog and at `docs/supported-models.md`. The catalog the binary prints is the list to trust, not the prose that introduces it.

Half-precision states stopped widening to f32

One entry in the highlights list describes a dtype bug and its cost together, which makes it the most useful thing in that list. Activation helpers, the multimodal rotary, the Gemma3n load policy, and the GraniteMoeHybrid gated norm used to widen a half-precision hidden state to f32, and because the state was f32 from that point on, every later matmul promoted its own weight. The project states that the affected checkpoints move by up to 3.26x once the widening is gone.

The second half matters more than the ratio. The bridge's reductions now accumulate in f32 to close a bfloat16 `max` that returned NaN for a finite input. A reduction producing NaN from finite operands does not crash a serving process; it returns bad text, which is far harder to trace back than an exception would be. Keeping the reduction in f32 while the stored state keeps its original dtype is the shape of fix that addresses both halves at once.

Read as a pattern rather than a changelog entry, it says something about how the runtime is organized: half-precision paths are expected to hold their dtype, and the places where they used to widen are enumerated by name. Anyone profiling this on an M-series host has a short list of checkpoints to compare against first.

Weight surgery runs at load time from a YAML file

Two capabilities change what a checkpoint contains at load time rather than at training time, and neither needs a conversion step. YAML load-time weight edits ship in the default build, reached through `--surgery` or `MLXCEL_SURGERY`, with `scale`, `add`, `prune`, `replace`, and `interpolate` as the operations. An `examples/surgery/` directory sits in the tree beside twenty-two example files, eighteen of which carry bench, benchmark, microbench, or profile in the name, so the surgery path has examples of its own while the rest of that directory is measurement code.

Router mode is the other. It discovers checkpoints from the model store, a model directory, or INI presets, loads them on demand, and bounds the resident set with LRU eviction. LoRA adapters can stay unfused, which keeps per-request and live scale changes available, or be fused for zero decode overhead, and that is a genuine trade rather than a quality setting.

bash
mlxcel arch

That is the command to run first on a new host, since it prints the architecture catalog the binary itself knows about. Checkpoints are read as MLX SafeTensors straight from HuggingFace, including the `mlx-community` conversions, and many standard SafeTensors embedding checkpoints load without conversion. Loading, scheduling, and inference stay inside one process, and platform runtime libraries are still required once the executables are packaged, so the deployment story is two native binaries plus whatever the host platform insists on.

Editorial conclusion

mlxcel suits an Apple Silicon shop that wants OpenAI, Anthropic, and llama-server routes out of one native binary and is willing to track a branch instead of a tag. It does not suit a team that needs AMD GPUs in production, needs speculative decoding outside greedy sampling, or expects the 376 classified llama-server options to be interchangeable with the reference server. The last push landed on 2026-09-30, three weeks after the v0.7.0 tag on 2026-09-09, so before adopting, read CHANGELOG.md rather than the tag page, check which MLX commit your build pins, and confirm whether your workload needs the fused ROCm kernels that currently fall back to MLX graphs.

Frequently asked questions

Does mlxcel run on AMD GPUs, or is it limited to Apple Silicon and NVIDIA cards?

AMD is an experimental, source-build-only target. The `rocm` cargo feature builds on Linux against an MLX ROCm backend vendored into the source tree as the ROCm part of mlxcelverse, and the notes say fused kernels fall back to MLX graphs while affine MoE models need `MLXCEL_FUSED_MOE=0` for now. Apple Silicon is the primary target and Linux with CUDA the secondary one.

Can mlxcel serve an OpenAI-compatible API, or only llama-server routes?

Both. Chat Completions, Completions, Responses, Embeddings, Reranking, Audio, and Anthropic Messages are served alongside the native `llama-server` routes, and a frozen b10621 manifest classifies all 376 pinned options, routes, and native request fields with no deferred entries, marking each as supported, aliased, not-applicable, or an intentional difference.

What does mlxcel need to build from source, and does it need Python?

The pinned MLX commit applied per backend, including the ROCm part of mlxcelverse that ships as source. No Python sits in the request path, but the Makefile still defaults `WEBUI_CONTRACT_PY`, `WEBUI_BUNDLE_PY`, and `WEBUI_EVIDENCE_PY` to `python3`, and platform runtime libraries are required at run time even after the binaries are packaged.

Which mlxcel version does the README describe, and which one would I install?

The README describes the `main` branch, which is v0.7.0 plus unreleased work. The latest tag is v0.7.0 from 2026-09-09, preceded by v0.7.0-beta.1 on 2026-09-04 and v0.6.0 on 2026-08-22, while the last push to the branch landed on 2026-09-30. Reading the changelog is how you tell those apart.

Official sources

  1. Issues
  2. lablup/mlxcel on GitHub
  3. License: Apache-2.0
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/lablup-mlxcel.svg)](https://hysenlabs.com/projects/lablup-mlxcel)