Library / SDK
mizorewww/laya-mlx avatar
mizorewww/laya-mlx

laya-mlx: typed decisions on Apple Silicon without a token decoder

Native MLX runtime for Laya typed decision models — 7–14 ms short decisions on M3 Max. No text generation, PyTorch, or cloud API.

6,654 stars523 forksPythonApache-2.0

At a glance

What is it?
An independent MLX port of the Laya decision models that returns probabilities over options, rubric levels or a truth value instead of generating text. It runs on Apple Silicon only, and the README's latency figures come from a one-question API benchmark rather than the Snake demo loop.
Who is it for?
Adopt laya-mlx if you need a bounded decision (a department, a rubric level, a yes/no probability) on an Apple Silicon machine and you want it local, without a PyTorch or Transformers runtime. Skip it if you need generated text, if you are not on arm64 macOS, or if you need to fine-tune: the README states that RLCD training and fine-tuning remain in the upstream project, and the port covers inference and conversion only.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 9 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem laya-mlx solves: a choice, not a paragraph

Most local model stacks are built to produce text. A routing rule, a rubric score or a yes/no gate does not need text. It needs a distribution over a small set of named outcomes, and it needs that distribution fast enough to sit inside a request handler. laya-mlx targets exactly that gap. The README frames the project as answering "constrained questions in a bidirectional forward pass, without token-by-token decoding or generated JSON," and reports 0 output tokens.

The audience is narrow on purpose. You need an Apple Silicon Mac, Python 3.11 or newer, and macOS 14 or newer. If your deployment target is a Linux container, this project is not for you, and the dependency list says so directly: the mlx requirement is marked with sys_platform == 'darwin' and platform_machine == 'arm64'. That is not a soft preference. It is the runtime.

How the bidirectional encoder and decision heads fit together

The pipeline the README gives is short: state plus typed question goes into a bidirectional encoder, then decision heads, then probabilities. There is no autoregressive loop, so there is no decoding step to tune and no JSON to parse out of a string.

Three question types are exposed. choice returns probabilities over named options. score returns probabilities over ordered rubric levels plus their expected score. noul returns P(true) for a proposition. The encoder, the decision Transformer, the scoring head and the action head all run in MLX. Tokenization is delegated to Hugging Face's Rust tokenizer.

One design point is stated plainly and is worth reading twice: question rows are batched independently, and their encoder representations depend on both state and question. The README explicitly says the runtime "does not claim to encode the state once and reuse its hidden states across arbitrary questions." If you were hoping to amortize a long shared context across many questions, that is not what this does. The Snake optimization path does use prefix reuse, but that is a separate documented mode, not the general API contract.

Three checkpoints are supported. convaiinnovations/laya is a 421M ModernBERT-large encoder with a 512-token context, for English. convaiinnovations/laya-multilingual is a 322M mmBERT-base encoder with a 1,024-token context. convaiinnovations/laya-typed-decisions is another 421M ModernBERT-large with a 1,024-token context, aimed at upstream typed-decisions workflows. Context covers instructions, options and state together, so a long state string eats into the room left for the question.

Installing laya-mlx and running a first typed decision

The package is on PyPI. The README's quick start is a single pip install, and the first load pulls the checkpoint from Hugging Face; after that, inference is local.

bash
pip install laya-mlx

The example below asks a support message to be routed. It loads the pre-converted FP16 checkpoint aac6fef/laya-mlx, then calls predict with the state string and a dictionary describing one choice question with three criteria. The result is read back through result["answers"]["department"], which holds the selected option.

python
import laya_mlx as laya

agent = laya.load("aac6fef/laya-mlx")
result = agent.predict(
    "I was billed twice. Please refund the duplicate.",
    {
        "department": {
            "type": "choice",
            "instructions": "Who should handle this?",
            "criteria": ["billing", "technical", "sales"],
        }
    },
)
print(result["answers"]["department"])

For a terminal demo, the README installs the demo extra, downloads the multilingual checkpoint in advance, and runs the Snake CLI. The download step matters because the demo is meant to run offline afterwards.

bash
pip install 'laya-mlx[demo]'
hf download aac6fef/laya-multilingual-mlx
laya-snake

Use a terminal of at least 104 x 35 cells. Space pauses, the up and down arrows change speed, R resets and Q quits. The README also documents laya-snake --max-speed, which makes a fresh decision for every move without pacing, and laya-snake --optimize --max-speed, which enables the tested compilation and prefix-reuse path. For a development install from source, the repository uses uv: gh repo clone mizorewww/laya-mlx, then cd laya-mlx, uv sync --extra demo, and uv run --extra demo laya-snake. The package also exposes a laya-mlx console script entry point.

Latency numbers and what they do not cover

The headline figures are 13.42 ms P50 for one short question on the 421M model and 7.39 ms for the 322M multilingual model, both FP16 end to end on an M3 Max with 40 GPU cores and 128 GiB of memory. P95 is 13.92 ms and 7.79 ms respectively. Throughput at 50 questions is 146.8 q/s and 395.0 q/s. Peak MLX allocation for a single short question is 943.6 MiB and 687.6 MiB.

Read the fine print. Timing includes prompt preparation, tokenization, tensors, synchronized inference, calibration and result formatting, but excludes model loading. The 50-question measurement uses batch_size=64 while the API defaults to 16, so the throughput row is not what a default caller gets. And the README states outright that the latency figures are the one-question API benchmark, not the frame time of the three-question Snake loop. Quoting 13.4 ms as your per-decision budget in a multi-question loop would be a misreading of the source.

The environment was macOS 27.2, Python 3.12.13 and MLX 0.32.2. That MLX release ships macOS 14, 15 and 26 wheels, and the README notes the local installer selected the 26 wheel, with older supported macOS versions untested on that machine. If you are on macOS 14 or 15, you are inside the stated support range but outside the measured one.

On correctness, the README reports that all three checkpoints matched the upstream selected answer on 63/63 validation questions in both FP32 and FP16, 378 comparisons in total, and that each configuration passed 100 repeated finite, deterministic calls with zero measured active-memory growth. The README itself scopes this: it measures fidelity on those fixtures, not accuracy on every possible question. Treat it as a port-fidelity claim, not a quality claim.

Where laya-mlx is the wrong tool

The clearest boundary is platform. There is no CUDA path, no CPU fallback described, and no Linux wheel. The mlx dependency is conditioned on darwin and arm64, so a pip install on anything else does not give you a working runtime.

The second boundary is task shape. This is not a text generator. If your output is a summary, a reply draft or a tool call in JSON, you are outside the three question types and outside the project's stated purpose. The README's own framing, "0 output tokens," is a feature for routing and a disqualifier for generation.

The third is training. The README states that the repository provides inference and conversion, and that RLCD training and fine-tuning remain in the upstream project. There is also no mention of GPU memory tuning knobs beyond dtype and batch size, and no documented rollback path if a checkpoint revision changes behaviour. The published checkpoints do carry pinned revisions and weight hashes in hub-publication.json, which is the closest thing to a reproducibility handle here, but the README does not describe a downgrade procedure.

Finally, the project is an independent port and says so: it is not an official Convai Innovations release. If you need vendor support, that is a real consideration.

laya-mlx versus keeping the PyTorch reference stack

The natural alternative is the upstream Laya project with its PyTorch and Transformers reference path. The pyproject.toml makes the split concrete: there is a reference extra pulling torch>=2.14,<3, transformers>=5.17,<6 and safetensors>=0.6, and it is optional. The default dependency set is mlx, numpy, huggingface-hub and tokenizers.

The difference is not accuracy on the fixtures, since the port matched upstream on all 63 validation questions in both precisions. The difference is the runtime. The reference stack gives you the training and fine-tuning workflows that this repository explicitly leaves upstream, plus portability off Apple Silicon. laya-mlx gives you a smaller dependency surface on a Mac and a path that the README measures in single-digit to low-double-digit milliseconds per short question.

If you are prototyping on a Mac and will later train or fine-tune, the reference extra is the more honest starting point, because you will not have to switch stacks when you reach the part this port does not cover. If you only ever call the model for inference on Apple hardware, the MLX path removes an entire framework from your install.

Licence, maintenance and what an upgrade costs you

The project is Apache-2.0, declared both in the repository metadata and in pyproject.toml, with a LICENSE and a NOTICE file at the top level. Apache-2.0 includes an explicit patent grant and requires that you preserve the NOTICE file and state changes you make. That is a description of the licence text, not legal advice; if you are redistributing the package or the converted weights, have someone check your obligations.

The licence of the package is one question. The licence of the model weights is another. The README says each published checkpoint includes its own model card, provenance and license, so the weights carry terms separate from the Apache-2.0 code, and you need to read the model card for the checkpoint you actually deploy. The README also states the port is independent and not an official Convai Innovations release.

On cadence: the last push to the repository was on 2026-09-22. Version 0.2.0 is classified as Development Status 4 - Beta, so expect interface movement. The upgrade surface is small in one respect and awkward in another. Small, because the dependency list is four packages and the MLX pin is a single minor range, mlx>=0.32.2,<0.33, which means an MLX 0.33 release will require a version bump here. Awkward, because checkpoint behaviour is tied to pinned revisions and weight hashes recorded in benchmarks/results/hub-publication.json, so a checkpoint update is a behavioural change you should treat like a code change. The README does not document a rollback procedure for that case.

Editorial conclusion

Adopt laya-mlx if you need a bounded decision (a department, a rubric level, a yes/no probability) on an Apple Silicon machine and you want it local, without a PyTorch or Transformers runtime. Skip it if you need generated text, if you are not on arm64 macOS, or if you need to fine-tune: the README states that RLCD training and fine-tuning remain in the upstream project, and the port covers inference and conversion only. Before wiring it into anything, verify two things yourself: that your macOS version has an MLX wheel, since the README notes older supported macOS versions were not tested on the measuring machine, and that your questions fit the context limit of the checkpoint you pick, 512 tokens for convaiinnovations/laya and 1,024 for the other two.

Frequently asked questions

What does laya-mlx need to run?

An Apple Silicon Mac with Python 3.11 or newer and macOS 14 or newer. The mlx dependency is conditioned on darwin and arm64, so there is no Linux or CUDA path described. The first load downloads the checkpoint and later inference is fully local.

Does laya-mlx generate text?

No. The README describes it as answering constrained questions in a bidirectional forward pass and reports 0 output tokens. It returns probabilities over named options for choice, over ordered rubric levels for score, and P(true) for noul.

Can I fine-tune a Laya model with laya-mlx?

No. The README states that this repository provides inference and conversion, and that RLCD training and fine-tuning remain in the upstream project. The optional reference extra pulls torch and transformers for the upstream path.

Is laya-mlx an official Convai Innovations release?

It is not. The README describes it as an independent MLX port that retains the original pretrained weights, question formatting, calibration and output schema, but it is not an official Convai Innovations release.

Official sources

  1. Issues
  2. License: Apache-2.0
  3. mizorewww/laya-mlx on GitHub
  4. Project website
  5. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/mizorewww-laya-mlx.svg)](https://hysenlabs.com/projects/mizorewww-laya-mlx)