# ZINC leaves its own speculative decoding out of the chart on purpose

> A single Zig binary that runs open-weight models on the GPU you already have, with a browser chat and an OpenAI-compatible server on one port. Its benchmark page is more interesting than the results: it measures against the same engine on the same card, keeps the rows where the other engine wins, and deliberately excludes the one number that would have flattered it.

**zolotukhin/zinc** — Zig INferenCe Engine — Local LLM inference on AMD GPUs and Apple Silicon

- Repository: https://github.com/zolotukhin/zinc
- Website: https://zolotukhin.ai/zinc/
- Stars: 520 · Forks: 22
- Language: Zig
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/zolotukhin-zinc

## One binary, and the first thing it does is report the hardware

The pitch is deliberately narrow: a single Zig binary, no Python, and no assumption that the machine has a CUDA-capable device. It loads GGUF files and gives you a command line, a browser chat, a model manager and an OpenAI-compatible API.

The getting started sequence is four steps and the second one is unusual. Before anything else you build and then run a check flag, described as answering the question of which GPU it found. That is the right first command for an engine whose whole reason to exist is that hardware support varies by vendor.

```bash
git clone https://github.com/zolotukhin/zinc.git
cd zinc
zig build -Doptimize=ReleaseFast

./zig-out/bin/zinc --check                      # what GPU did it find?
./zig-out/bin/zinc model pull qwen35-9b-q4k-m   # fetch a model
./zig-out/bin/zinc --model-id qwen35-9b-q4k-m --prompt "Hello" --chat
```

The prerequisites are versioned. The compiler must be a specific minimum version or newer, Linux builds using the portable backend additionally need a shader compiler and a Vulkan loader, and the other backend needs a working vendor runtime installation. Two runtime variables appear in the build and run commands: one points at the vendor runtime installation directory, and one selects the visible device.

## The best number on the page is deliberately not on the chart

One model in the tuning list ships an extra prediction block, and the engine uses it to draft tokens that the full model then verifies in a single batched pass. The claim is identical output with fewer passes and 1.9 times the generation speed.

The benchmark page does not show that number.

The reason is given at length, and it is the most careful passage in the file. The competing engine can speculate with the same block, but its converter exports it as a separate draft model, and the runs behind the chart do not give it one. Comparing 1.9 times against an engine that was not allowed to speculate would not be a fair comparison, so the chart shows both engines without speculation, and a switch on the benchmark page reveals the speculative figure separately.

The same restraint shows up in the row counts. Prompt processing is reported as ahead in all 24 rows, with the range given for the short chat prompt shown in the readme and a narrower range for the longest prompts, which is the opposite of how a benchmark is usually framed. Token generation is also ahead in all 24 rows, with the slowest and fastest cases named by model.

And the scope is stated rather than implied: one card and six models, with other cards measured separately, and rows where the other engine is ahead left visible on the page.

## The chart is regenerated from a data file, never edited by hand

The provenance of the numbers is specified in the same section that presents them.

Every figure comes from one named script and lands in one named data file inside the site directory. That file records the prompts, the raw samples, the build revisions, and the commit of the competing engine used for the comparison. The chart above it is regenerated from that file and, in the documentation's words, never hand-edited.

That is a stronger claim than most projects make about their benchmarks, because it means the chart cannot drift from the measurements. It also means the comparison is reproducible in a specific sense: given the recorded prompts and the recorded engine commit, someone can regenerate the same chart. The benchmarking guide is pointed at for anyone who wants to reproduce or extend the results, and it is said to document how to publish a pair of runs with speculation on and off.

The measurement setup is described just as carefully. Both engines run on the same card, through their respective vendor backends, with the competing engine built from the same commit as the project's other comparison. Both load the same model file, and the harness checks that the prompt token counts match, so a difference in speed cannot be explained by one side having read more input. Both run in a reusable server with identical warmups and repeat counts.

That is the difference between a benchmark and a screenshot, and it is why the exclusions in the previous section are credible.

## Five hardware rows, and no backend inherits another's numbers

The support table covers five combinations, and the status column is where the honesty lives.

Two rows cover AMD cards on two generations, one through the vendor compute stack and one through the portable graphics backend, with the first marked as the fastest for prompt processing on those cards and the second as portable and needing no vendor runtime installation. Intel's discrete cards are listed as supported on the portable backend. Apple hardware is listed as supported on its own backend. And an NVIDIA row exists with the status given as experimental.

One NVIDIA row marked experimental in a project whose pitch is that it makes no CUDA-only assumption is a notable admission, and it is consistent with the rest of the file.

Underneath the table is the rule that keeps the numbers meaningful: each backend has its own kernels and is benchmarked separately, so no backend inherits another's results. That matters because the numbers in the readme come from one specific backend on one specific card.

The guidance that follows is a recommendation to pick the backend for the job. The compute-stack backend reads prompts considerably faster than the portable one, while for generation the two are close and the portable backend is ahead on some models, with one named comparison given as 105 tokens per second against 97. So neither backend wins outright, and the documentation says so rather than optimising for the flattering half.

## The device variable in the example is spelled unusually

The build and run commands for the vendor backend set two environment variables, and one of them does not match the conventional spelling.

```bash
ROCM_PATH=/opt/rocm zig build -Dbackend=rocm -Doptimize=ReleaseFast
ROCR_VISIBLE_DEVICES=0 ./zig-out/bin/zinc --check
```

The first is the runtime installation path, which is straightforward. The second selects which device the engine sees, and its name has an extra letter inserted into the middle of the vendor prefix compared with the name that vendor's documentation uses. Anyone copying that line from the readme into a shell will set an environment variable the runtime does not read, and the engine will fall back to seeing every device rather than the one you asked for.

It is a small thing in a readme that is otherwise unusually careful, and it is the kind of detail that only matters on a multi-card machine, which is exactly where you would use the variable.

The other thing to notice about those two commands is that the backend is chosen at build time rather than at run time, as a build option, and the resulting binary is a different one from the portable default. There is no flag to switch backends in a running binary, so changing your mind means another build.

## The browser chat and the API are the same server on one port

One command starts both halves of the user facing product:

```bash
./zig-out/bin/zinc chat --model-id qwen35-9b-q4k-m
```

That starts the browser chat and an OpenAI-compatible API on the same port, which is stated as the reason existing clients and SDKs work unchanged. The endpoints named are a liveness path and the two standard paths for listing models and for chat completions, and a separate guide holds command line and SDK examples for each.

That design decision explains most of the rest of the repository layout. If the product is a server, then the benchmark harness is a client, and the client needs an SDK. The JavaScript manifest in the repository root carries exactly one dependency, and it is an SDK for the competing service, used to prove the endpoint behaves like an endpoint.

The manifest is private and fixed at version 0.1.0, and the repository has no GitHub releases, so neither number means anything to an outside user. It is a tool manifest for the project's own scripts.

Model handling has two paths. You can point the engine at a GGUF file already on disk, or you can use the hub flag with a repository and a quantisation to fetch one, and there is also a managed catalog you pull from by name. The documentation is specific about which checkpoint its own measurements used, linking the exact file published by the model's author, so a benchmark run can be reproduced against the same weights rather than a similar quantisation.

## Conclusion

ZINC is worth trying if you have an AMD card and no appetite for a Python stack, since a single binary that serves an OpenAI-compatible API is an easy drop-in for existing clients. Two things to check. Pick your backend deliberately, because prompt processing and token generation peak on different backends and the documentation says so rather than hiding it. And read the benchmarking guide before quoting any figure, because the headline numbers come from one card and six models with the project's own best result excluded.

## FAQ

### What does ZINC need to build and run?

A specific minimum version of the Zig compiler or newer. Linux builds on the portable graphics backend also need a shader compiler and a Vulkan loader, and builds on the vendor compute backend need a working vendor runtime installation.

### Why is ZINC's speculative decoding not in its benchmark chart?

Because the competing engine could speculate with the same prediction block but was not given a draft model in those runs, so the comparison would have been unfair. The chart shows both engines without speculation, and a switch on the benchmark page reveals the speculative figure.

### Where do the ZINC benchmark numbers come from?

From one performance script, which records prompts, raw samples, build revisions and the competing engine's commit into a JSON data file in the site directory. The chart is regenerated from that file and never hand-edited.

### Which GPUs does ZINC support?

AMD cards on two recent generations through both the vendor compute stack and the portable graphics backend, Intel discrete cards on the portable backend, and Apple hardware on its own backend. An NVIDIA row exists with the status given as experimental.

### Can I use ZINC with an OpenAI client?

Yes. The chat command starts the browser interface and an OpenAI-compatible API on the same port, with liveness, model listing and chat completion endpoints, so existing clients and SDKs work unchanged.

### How do I point ZINC at a model?

Either with a path to a GGUF file you already have, or with a hub flag taking a repository and a quantisation to fetch one. There is also a managed catalog you pull by name.

## Sources

- [Issues](https://github.com/zolotukhin/zinc/issues)
- [License: MIT](https://github.com/zolotukhin/zinc/blob/main/LICENSE)
- [Project website](https://zolotukhin.ai/zinc/)
- [README](https://github.com/zolotukhin/zinc/blob/main/README.md)
- [zolotukhin/zinc on GitHub](https://github.com/zolotukhin/zinc)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/zolotukhin-zinc
