Model or dataset
ashhart/TensorFold avatar
ashhart/TensorFold

TensorFold: exact speculative decoding on Apple Silicon and NVIDIA, one backend short of its own pitch

LLM Inference Engine for Metal, CUDA and Vulkan.

1,150 stars199 forksPythonApache-2.0

At a glance

What is it?
An OpenAI-compatible inference server whose real documentation is a matrix of weight formats, where a draft token is accepted only if it matches serial decoding. The repository description promises three backends and the README documents two.
Who is it for?
TensorFold is unusually precise about the things that usually get hand-waved. The exactness rule for drafts is stated as a condition rather than a benchmark, the per-model weight format requirements are documented rather than discovered, and unsupported hardware is refused at startup instead of producing wrong numbers.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 9, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Two backends in the docs, three in the repository description

The repository description reads LLM Inference Engine for Metal, CUDA and Vulkan. The README describes something narrower and different: it serves language models on Apple Silicon and NVIDIA GPUs through an OpenAI-compatible API, with each model family supplying its own kernels and draft verification. Metal and Vulkan appear nowhere in the README, and neither does the word Vulkan anywhere in the file. The `pyproject.toml` description agrees with the README, saying fast, exact LLM decoding on Apple Silicon through MLX and NVIDIA GPUs through CUDA.

So two of the three claimed backends are documented and one is not, and the naming shifts as well. MLX is what the project actually calls the Apple path, and the dependency list is explicit about how it is used: MLX and mlx-lm are installed only on Darwin, with mlx-lm deliberately held below its next minor release. The CUDA path takes a different shape, since a comment in the dependencies explains that the CUDA build uses the container's torch and triton rather than pinning them.

The repository topics tell the same story as the README, listing `ai`, `llm`, `llm-inference` and `llm-tools` with no hardware terms at all, and the package keywords are specific about MLX, Apple Silicon, CUDA and DGX Spark. The classifiers cover MacOS, Linux and NVIDIA CUDA environments, and nothing else, so Windows is not a target.

The honest reading is that the description overstates the surface area. Nothing in the documentation suggests a Vulkan path exists, and given how much work each backend's kernels require, a description promising three of them when two are documented is the first thing to reconcile before you plan around it.

Installation runs from a git URL, and the manifest adds detail the README omits

There is no package index in the README. Installation is a pip install straight from the repository, and Homebrew carries a formula for the Mac path:

bash
python -m pip install git+https://github.com/ashhart/TensorFold.git
tensorfold serve TensorFold/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-MLX-4bit

The manifest adds requirements the README states only in prose. Python 3.11 or newer is required. On Darwin, MLX is pinned to a window of `>=0.32.2,<0.32.4` and mlx-lm to `>=0.31.3,<0.33`, with a comment explaining that the prompt kernels and SSD streaming were tested on MLX 0.32.2 and 0.32.3. Everywhere except Darwin, the package pulls in tokenizers, safetensors and jinja2 separately, which is consistent with the Mac path getting those from MLX's own stack.

That narrow MLX pin is the operational detail to plan around. Holding a two-version window means a new MLX minor will not install until this project ships, so an MLX release and a TensorFold release arrive together. The same kind of coupling appears in an optional extra that builds a small MLX extension on first use through MLX's nanobind, with the comment noting which nanobind version MLX 0.32.2 takes.

The README routes installation details to RUNBOOK.md and the endpoint contract to `docs/api.md`. Both backends serve chat completions, plain completions, OpenAI Responses at `/v1/responses` and Anthropic Messages at `/v1/messages`, and the client base URL is `http://127.0.0.1:8080/v1` with model IDs read from `/v1/models`. Serving four API dialects from one process is a genuine convenience if you are migrating an application between providers, and it is also four contracts to keep correct.

A draft is accepted only when serial decoding would have produced it

The exact decoding section is the heart of the project and it is short. A draft token is accepted only when it equals the token the same engine would produce serially. Sampling depends on the prompt or an explicit seed, absolute position and token ID, so the verification has enough context to be a genuine equality check rather than a heuristic.

That condition is what makes the speedups claimable rather than merely plausible. Speculative decoding normally trades output distribution for throughput, and a drafter that diverges means your benchmark numbers describe a different model than the one you deployed. TensorFold's answer is to keep the output bit-identical and move the cost into verification, with each round's depth chosen from measured per-device costs and a round staying plain where drafting would not pay.

The reported numbers are specific enough to be checkable. On an M3 Ultra, greedy, against a no-drafts baseline, short prose runs 1.46 times faster at 173 against 119 tokens per second, short code 2.05 times at 245 against 119, thinking 1.65 times, and a 28,400-token prompt 1.34 times with the same time to first token. Against mlx-vlm 0.7.4 with its own MTP drafter, the comparison is 1.45 to 1.73 times. Every reply's tokens are stated to have matched plain decoding, greedy and sampled, which is the claim that matters and the one an evaluator should try to reproduce.

Nemotron works the same way on CUDA, where the MTP head drafts up to fifteen levels and a chain stops where the next verify row would cost more time than its draft is expected to save, with window and level costs measured at startup on each GPU. The consequence is that the same binary drafts deeper on a faster card, which is a design choice with a tradeoff: throughput adapts to hardware, and so does the profile of what you are measuring.

Kernels compiled on Blackwell for cards nobody has run them on

The hardware section is where the README is most candid and most awkward at once. The CUDA kernels need compute capability 8.9 or newer, covering Ada, Hopper and Blackwell, including the DGX Spark's GB10 and the RTX 50 series. RTX 30 cards at 8.6 are unsupported, and the server refuses a GPU below 8.9 at startup rather than limping along.

Then comes the sentence that deserves attention: the RTX 40, Hopper and B200 builds are compiled and bit-checked on Blackwell but not yet run on those cards. So three of the named hardware families have had their kernels produced and numerically compared against a reference, on different hardware, without the binaries ever executing on the hardware they target. That is a reasonable way to catch kernel generation errors and it is not the same as running them, and the project says so rather than implying coverage it does not have.

Quantization support follows the same shape. NVFP4 and FP8 checkpoints run from capability 8.9, using their own math where the GPU has the relevant matrix-multiply instruction, FP4 on 12.x and FP8 from 8.9, falling back to W4A16 elsewhere. The README also flags the EXL3 rows in the model table as experimental, and one of them branches across 1 to 8 bits per weight with any codebook, which is a wide surface to call experimental.

One model in the table carries a licensing condition worth catching early. GLM's optional DFlash2 drafter checkpoint is noted as having non-commercial license terms, documented in THIRD_PARTY_NOTICES.md. The project itself is Apache-2.0 with a LICENSE, a NOTICE and a LICENSES directory, so the core is clean, but an optional accelerator with different terms is exactly the kind of thing that matters if you are deploying commercially.

Patch releases that each carry one feature

The release train is fast and unusually granular. Three releases land between 2026-10-03 and 2026-10-06, and version 0.6.6 is titled for a single flag, name-priority on the CUDA server.

That flag has a good story behind it. A request naming the served background model ID and sending no priority of its own is served at background priority, so it yields to foreground traffic, while a request's own priority field always wins. The measured effect is dramatic and narrow: on one DGX Spark, a foreground request queued behind four background ones got its first token in 0.2 seconds instead of 21. The stated motivation is that some clients, batch extractors among them, can choose a model ID but cannot add a field to the request.

The other two releases are broader. 0.6.5 adds Qwen3.6-35B-A3B drafting with its own MTP layer on Macs, chains of up to four drafts verified in lane rounds with per-round depth from measured Mac costs, and Nemotron on CUDA drafting as deep as its rows pay for, reported as roughly 6% faster on one and two DGX Sparks and 7% on an RTX PRO 6000. It also adds API keys. Version 0.6.4 makes Flash Next serve concurrent requests on two DGX Sparks with a parallel flag, unifying the communicator interface for two-rank engines and adding a plan subcommand.

There is a continuity detail buried in the checkpoint ids that suggests this project was renamed rather than started. The README notes that the `TensorFold/...` identifiers moved from the Vontra organization on Hugging Face on 2 October 2026, with the old names redirecting. Anyone with scripts pinned to the earlier namespace gets redirects rather than breakage, which is a courtesy worth noting.

The weight format matrix is the actual documentation

Skip past the benchmark numbers and the most useful part of the README is a paragraph per model describing exactly which weight layout will load. Qwen3.8-27B reads MLX affine checkpoints at 2, 3, 4, 5, 6 and 8 bits including mixed layer formats, and reads packed rows in groups of 32, 64 and 128 on both Apple Silicon and CUDA, with hardware qualification still pending for the newer paths. Flash Next requires 4-bit with group-32 weights and refuses other layouts before downloading anything. Nemotron on CUDA requires 4-bit group-64 and an MTP head unless drafting is disabled. GLM on MLX reads 4-bit group-64 plus mlx-lm's mixed-bit conversions, whose 5, 6 and 8-bit tensors take their own row kernels.

Gemma 4 is stricter still, reading 4-bit weights in groups of 32 or 64 with an 8-bit router exactly as the mlx-community conversion stores them, and refusing other layouts before download. DeepSeek-V4-Flash reads the mlx-community conversion with affine 4-bit group-64 weights and mxfp4 routed experts.

That refusal behavior is a design decision worth more than it first appears. Refusing before downloading means a user with the wrong checkpoint learns in seconds rather than after a multi-gigabyte transfer, and it converts a class of confusing kernel errors into a clear message.

The tooling around it is small and sensible. `tensorfold models` lists families and checkpoints, `tensorfold info MODEL` checks configuration without fetching weights, `serve` downloads a missing checkpoint and `pull` fetches it in advance. Pulling a drafter and its target together is the documented pattern:

bash
tensorfold pull TensorFold/Qwen3.8-27B-MLX-4bit z-lab/Qwen3.8-27B-DFlash2
tensorfold serve TensorFold/Qwen3.8-27B-MLX-4bit

Two practical limits are visible in the model table without being argued about. GLM-5.3-Flash and DeepSeek-V4-Flash on MLX both require a 256 GB Mac, which in practice means the largest unified-memory configuration available. And vision input needs the optional extra plus a supported dense checkpoint started with a vision flag, with GLM images running on MLX and dense Qwen images on either backend, so image support is a per-model capability rather than a server-wide one.

Editorial conclusion

TensorFold is unusually precise about the things that usually get hand-waved. The exactness rule for drafts is stated as a condition rather than a benchmark, the per-model weight format requirements are documented rather than discovered, and unsupported hardware is refused at startup instead of producing wrong numbers. Those are the reasons to look at it if you serve models yourself on a Mac or an NVIDIA card. What the README does not give you is a Vulkan or Metal backend, and the pyproject agrees with the README rather than the repository description, so plan for MLX or CUDA only. Expect the MLX pin to be the operational constraint: the dependency is held between 0.32.2 and 0.32.4 because kernels were tested on exactly those builds, which means a TensorFold release is likely whenever a new MLX ships. Read RUNBOOK.md first, then the recipe for your specific model.

Frequently asked questions

Does TensorFold support Vulkan or a Metal backend?

The documentation describes two backends: MLX on Apple Silicon and CUDA on NVIDIA. Neither the README nor the pyproject description mentions Vulkan or a Metal path, and the repository topics carry no hardware terms. The CUDA side additionally expects the container's torch and triton rather than pinning them, so treat the description's three-backend claim as aspirational until a release says otherwise.

Does speculative decoding in TensorFold change the output?

That is the design constraint rather than a hoped-for property. A draft token is accepted only when it matches what the same engine would have produced serially, and sampling depends on the prompt or an explicit seed together with absolute position and token ID. The reported comparisons state that every reply's tokens matched plain decoding, greedy and sampled, so the speedups are meant to be free of distribution drift.

What hardware does TensorFold need for CUDA?

Compute capability 8.9 or newer, covering Ada, Hopper and Blackwell including the DGX Spark GB10 and the RTX 50 series. RTX 30 cards at 8.6 are unsupported and the server refuses anything lower at startup. The README also notes that the RTX 40, Hopper and B200 builds are compiled and bit-checked on Blackwell but not yet run on those cards.

How do I install TensorFold?

From the repository rather than a package index, with python -m pip install pointed at the git URL, or through Homebrew on a Mac. Python 3.11 or newer is required and the Mac path needs MLX 0.32.2 or newer, which pip installs for you. Serve a model with tensorfold serve and a checkpoint id, then point your client at http://127.0.0.1:8080/v1 and read the model id from the models endpoint.

Official sources

  1. ashhart/TensorFold on GitHub
  2. License: Apache-2.0
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/ashhart-tensorfold.svg)](https://hysenlabs.com/projects/ashhart-tensorfold)