PMetal, twenty crates and a desktop app deliberately outside the workspace
PMetal: high-performance Apple Silicon framework for local LLM inference, LoRA/QLoRA fine-tuning, serving, quantization, and MLX/Metal acceleration.
At a glance
- What is it?
- PMetal is a machine learning SDK written in Rust for Apple Silicon, covering Metal kernels, Neural Engine integration, training, quantization, serving, a twenty tab terminal interface, and a desktop application. The parts that reveal the engineering are the build files: a preflight recipe that mirrors continuous integration step for step, a minimum feature build nobody else compiles, and a GUI backend excluded from the workspace so it needs its own formatting pass.
- Who is it for?
- PMetal is worth trying if you are training or serving on a Mac and want everything behind one binary, since the same install gives you a command line, a terminal interface, and a desktop app. Check four things first.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository last received commits 1 day ago.
- What is it written in?
- Mainly Rust, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 3, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Twenty workspace crates and one desktop app outside it
The workspace manifest lists twenty members and excludes two directories.
The members run from a bridge crate through core, a derive macro crate, the Metal and MLX bindings, model definitions, adapters, data, the trainer, a hub client, the GGUF format, a vocoder, merging, distillation, distribution, a mixed crate, serving, an MCP server, the main binary, and a Python binding. The dependency table then repeats each of those as a path dependency with an explicit version, so every internal crate carries the workspace number rather than a wildcard.
The exclusions are the interesting part. A fuzzing directory is excluded, and so is the desktop application's own Tauri backend.
The build file explains the consequence in a comment rather than leaving you to discover it. Because that backend is outside the workspace, a formatting command scoped to all members does not reach it, so it needs its own invocation or it drifts out of format over time. The same file does the same thing for linting, and the check recipe runs both.
That is a small piece of maintenance discipline that most projects only notice when it has already gone wrong.
Nineteen pages in the desktop app and twenty tabs in the terminal
The two full interfaces do not have the same feature list, and the difference is one page.
The desktop application has nineteen pages: dashboard, training, the group-relative training page, distillation, pretraining, inference, speculative decoding, models, datasets, merging, quantization, embedding training, reinforcement learning with distillation, Ollama integration, serving, benchmarking, evaluation, jobs, and settings. The terminal interface has twenty tabs, the extra one being a device page that reports GPU and Neural Engine information, detects Metal features, shows a memory gauge, offers kernel tuning, and describes a topology.
There is a second structural difference that matters more than the count. Training, inference, distillation, and the group-relative training run inside the application's own process with live progress. Every other page shells out to the command line binary as a subprocess, and the application bundles that binary with itself.
So the desktop app is a hybrid: a few heavy paths are linked directly for responsiveness, and the long tail is a front end over the same command line everyone else uses. That is a pragmatic split, and it means the bundled binary and the installed binary can drift if you update one and not the other.
Six of thirty-five commands exist only to measure
The command table has thirty-five entries, and six of them are benchmarks.
There is a general training benchmark, a generation loop timing benchmark, an interface-call overhead benchmark, a real cached workload benchmark, a structured kernel benchmark with JSON reporting, and a benchmark for one specific architecture's backends on real layer shapes. There is also a cluster mode with its own all-reduce throughput and pipeline transport benchmarks.
Six of thirty-five is a ratio that tells you something about the project. Most command line tools have one benchmark command. Having a dedicated one for foreign-call overhead and another for real cached workloads suggests the authors have been chasing specific bottlenecks rather than reporting a single number.
The rest of the table is unusually broad for one binary. There are five separate training entry points covering adapter fine-tuning, a multi-token-prediction predictor, a draft model for speculative decoding, a block-diffusion model, and full-parameter pretraining. Merging supports twelve strategies. Quantization supports twenty-four methods across two formats. There is a separate command to fuse adapter weights into base weights, which is a different operation from merging two models, and another to pack expert weights for offloaded inference.
There is also a server, an integration path for another inference runtime's model format, and an MCP server exposing fifty-one tools over standard input and output.
The fine-tuning entry point takes its arguments the way you would expect, and sequence packing is the default rather than an opt-in:
# LoRA fine-tuning with sequence packing (default)
pmetal train \
--model Qwen/Qwen3-0.6B \
--dataset train.jsonl \
--output ./output \
--lora-r 16 --batch-size 4 --learning-rate 2e-4The learning rate is adjustable while a run is going
One keybinding in the terminal interface is worth pulling out of the list.
Alongside the usual navigation keys, there is a single key to adjust the learning rate mid-run. There is also a jump-to-any-tab command that filters as you type, numeric shortcuts for the first nine tabs, a contextual help key, and a quit key.
Changing a learning rate during a run is the kind of feature that gets built by someone who has watched a schedule go wrong and did not want to kill the run. It is also the kind of feature that quietly changes what an experiment means: a run whose learning rate moved at step four thousand is not the run your configuration file describes, and nothing in the terminal interface appears to record that it happened.
That makes the jobs tab worth noting. It holds training run history with a log viewer, status tracking, and metadata. Whether the mid-run change is captured in that metadata is not something the README says, and it is exactly the kind of detail that determines whether a resumed run matches the original.
The dashboard tab goes further in the same direction, with loss curves drawn in braille characters, a schedule view, throughput sparklines, and timing breakdowns rendered as gauges.
The cluster picks Thunderbolt because it measured the link
The multi-machine mode is the feature that has no equivalent anywhere else in the table.
Two or more Apple Silicon machines are connected into what the project calls a home cluster. It enumerates every network interface on the machine, which includes Thunderbolt bridge, Ethernet, and Wi-Fi, advertises them over multicast name resolution, and forms a ring biased toward the fastest fabric. Thunderbolt is preferred over Ethernet, and Ethernet over Wi-Fi, with the README noting that this happens without configuration.
The commands are laid out as a procedure rather than a reference. Connect the machines with a Thunderbolt cable, then on every machine check status and join the cluster, then on every machine at the same time run the all-reduce throughput benchmark for each fabric and the pipeline activation transport benchmark, and only then start distributed training.
The ordering matters. Measuring the fabrics before training is the point of a ring that adapts itself, since the choice of interconnect is the thing that decides whether distributed training is worth doing at all.
A caveat the README does not hide: Thunderbolt four or five cables are recommended, and Ethernet is described as also working.
The workspace says 0.6.0 and the newest tag is 0.5.0
The manifest and the release history are a version apart, which is normal and worth noting for anyone pinning.
The workspace package block declares version 0.6.0, edition 2024, and a minimum Rust version of 1.91. The licence field is the dual form, offering either the MIT licence or the Apache licence, and both licence files are present at the repository root alongside a third-party notices file.
The three most recent releases are 0.3.7 in March, 0.4.0 a week later, and 0.5.0 in May. So the tree is one minor version ahead of anything you can install from the registry, which is the normal state of a project between releases.
The maintenance signal is strong. The last push is dated 1 October 2026, the same day this was checked, and the repository is not archived.
The remaining top-level entries describe how the project is operated: a pinned toolchain file, a build recipe driver, a dependency audit configuration, a separate bench directory, a fuzzing directory that is present but excluded from the workspace, a supply-chain directory, and security and changelog files.
A twelve step preflight that claims to be what CI runs
The build recipe file has a default target that lists the others, and one target that lists everything.
The recipe driver sets its shell to a login interactive one rather than the usual non-interactive default, and derives the project version by reading it out of the manifest with a shell pipeline, so there is one source of truth for the number.
The preflight target runs twelve things: a format check, the default lint, the all-features lint, the minimum-features lint, tests, release tests, a GUI check, a GUI lint, a version check, a lockfile check, and the dependency audit. The comment above it says it runs every check that continuous integration performs, and the success message says if it passes, publishing is safe. That claim is auditable by reading the recipe against the workflow.
The three lint variants are the interesting part. The default lint excludes the Python binding crate. The all-features lint exists to catch field mismatches hidden behind conditional compilation. And the minimum-features lint builds the smallest command line configuration with no trainer, no adapters, and no Neural Engine, because the comment says nothing else compiles those paths and so unused-code warnings there went unseen.
Building the smallest configuration to catch dead code is a practice more projects should copy and few do.
Editorial conclusion
PMetal is worth trying if you are training or serving on a Mac and want everything behind one binary, since the same install gives you a command line, a terminal interface, and a desktop app. Check four things first. Which model families you actually need, because support is named per family rather than general. Whether you have more than one Mac and a Thunderbolt cable, since that is the difference between the cluster feature working well and not at all. Which interface you will use, since the desktop app and the terminal interface do not have the same feature lists. And whether you can wait for the next tag, because the workspace version is ahead of the newest release.
Frequently asked questions
What is PMetal?
It is a machine learning SDK, framework, and application suite for Apple Silicon written in Rust, published on the Rust registry as pmetal. It covers low-level Metal GPU kernels, Apple Neural Engine integration, high-level training APIs, a terminal interface with twenty tabs, and a desktop application.
Does pmetal support LoRA and QLoRA fine-tuning?
Yes. The main training command fine-tunes with LoRA or QLoRA using supervised fine-tuning, sequence packing is the default, and a separate command fuses adapter weights into the base model. There are also entry points for a multi-token-prediction predictor, a speculative decoding draft model, and a block-diffusion model.
What does pmetal do with multiple Macs?
It can form a home cluster of two or more Apple Silicon machines. It detects every network interface, advertises them over multicast name resolution, and forms a ring biased toward the fastest fabric, preferring Thunderbolt over Ethernet over Wi-Fi without configuration.
How does pmetal expose its tools to an AI assistant?
There is a command that starts an MCP server over standard input and output, exposing fifty-one tools for use with Claude Desktop and Claude Code. It sits alongside a command line server that speaks both OpenAI and Anthropic compatible APIs.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/epistates-pmetal)