Model or dataset
guoqingbao/xinfer avatar
guoqingbao/xinfer

xinfer: a mistyped manifest at the root, a tensor library pinned to one commit of a fork, and two ports with no auth

Blazing-fast LLM inference in pure Rust. No PyTorch and Python runtime.

330 stars46 forksRustMIT

At a glance

What is it?
A large language model inference engine written in Rust with no Python runtime, shipping OpenAI and Anthropic compatible endpoints, a chat interface, MCP tool calling and an aggressive key value cache compression path. The engineering claims are specific and checkable, and several of the numbers do not line up with each other.
Who is it for?
xinfer is worth an afternoon if you are weighing whether you can drop a Python runtime from your serving stack, because the performance table is per model, per quantisation and per card, which is more honesty than most inference projects manage. Three things to settle first.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 24 days ago.
What is it written in?
Mainly Rust, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 3, 2026, and from our analysis. They are not legal advice.

Editorial analysis

There is a myproject.toml sitting next to the real manifest

The repository root has twenty entries and two of them are meant to be the same file.

One is the Rust package manifest, and it is complete: a package name, a version matching the newest release, an edition, a default run target, a description, repository and homepage fields, keywords, categories, and a licence.

The other is named with a misspelled first word. It is not a valid Rust package manifest filename, so the build system will never read it. It is committed to the repository and appears in the root listing next to the manifest that is actually used.

Two details inside the real manifest are worth noting. The licence is written as a plain string rather than the modern licence expression form, and the readme field points at a file whose name is capitalised oddly and carries a hyphen, which is the English readme while a second file holds the Chinese one.

The version is the useful part. It reads zero point fourteen point five, and the newest release tag is the same number on the same day as the last commit. So unlike most of the projects in this batch, the manifest, the tag and the commit history agree, and an installed binary can be traced back to a version.

The tensor library is a fork on the maintainer's own account at one commit

Two of the dependencies are not from a registry.

Both are declared against the same git repository, under the account that owns this project, at a fixed commit identifier and a fixed upstream version. They are the tensor core and the neural network layer, which for an inference engine are the two libraries that matter most.

Everything else in the manifest comes from a registry with a version range. The tensor core and the network layer come from a branch of somebody's working copy.

The stated reason for a fork is visible in the dependency list and reinforced by the feature table: the engine advertises native attention kernels, a graph-capture path, and a four-bit numeric format that works on cards without hardware support for it. Those are exactly the changes that have to happen below the tensor library, which is a reasonable motivation. But the cost is on the reader: there is no published release of the fork to diff against, no upstream relationship declared, and a commit identifier as the only thing pinning your build to a version that has no changelog.

The rest of the dependency list is conventional and well chosen. A tokenizer library with the HTTP feature for fetching from a model hub, an async runtime with only the synchronisation feature, a schema validation library with the JSON schema feature enabled, a template engine with a Python-compatibility module, an interprocess communication crate, a big array codec for tensor serialisation, and a readline library for the terminal interface.

Three names, two install routes, and one of them is a script piped from a static site host

The project has three identifiers. The repository, the crate and the command are all the same six letters. The package on the language package registry is not: it has a suffix, and it is the only one of the four that differs.

There are two documented install routes. One is a shell script piped from the project's documentation site into an interpreter. The other is a global install of the package-registry name.

The first has no version argument, no checksum and no way to read the script before running it, and the host is a static site on a personal pages domain rather than a release endpoint. The second is a versioned package, which is the safer of the two by default.

The repository record lists the documentation site over plain HTTP rather than HTTPS, while the install command in the readme uses HTTPS against the same host. Both can be true of the same domain, but the recorded homepage is the one a person will paste into a browser.

There is also a third path implied by the packaging directory at the root and by the optional Python entry point, which means a Python wheel exists for people who want one. So a project whose pitch is that it needs no Python has three install routes and one of them installs Python bindings.

Two ports, both plain HTTP, and no bind address or authentication anywhere in the file

One tip line tells you everything about the network surface. Open one address for the built-in chat interface, and use another address with a versioned path suffix as the base URL for the API.

Two ports, the interface on the higher one and the API on the lower. Both are written as plain HTTP with a literal address placeholder in place of the host.

The three run commands in the quick start differ only in how the model is named and which device list is passed:

bash
xinfer --m Qwen/Qwen3.6-27B-FP8 --kvcache-dtype turbo4 --ui-server
bash
xinfer --m /home/Qwen3.6-35B-A3B --d 0,1 --ui-server
bash
xinfer --m Qwen/Qwen3.5-35B-A3B --d 0,1 --ui-server --num-speculative-tokens 3

Three things are missing, and all three matter for anything beyond a laptop.

There is no bind address. The placeholder is an address, not the loopback interface, which means the default is to listen on every interface unless a flag says otherwise, and no such flag is documented in the visible file.

There is no authentication. The readme advertises OpenAI and Anthropic compatible endpoints, which is a compatibility promise about request and response shapes, not a security model. An inference server with an OpenAI compatible shape and no key is the shape most likely to be picked up by a scanner.

There is no TLS. Both addresses are HTTP.

The feature table is otherwise unusually complete about what is served: compatible chat endpoints, a web interface in the style of a consumer chat product, MCP tool calling, structured outputs, separate embedding and tokenizer endpoints, and multi-token prediction.

The cache table puts a throughput figure in a column headed quality

The key value cache table has five rows and four columns, and the compression figures are exact.

The default is uncompressed. Then four compressed modes with compression factors of two times, two point six times, three point seven times and four point seven times. Two of them state a compute-capability floor for Nvidia cards, and the first two also list Apple hardware.

The third column is the one with the error. For the two point six times row it reads as a percentage range, and that range describes throughput, not quality. The table's own quality column for the other rows contains words like near-lossless, best balance and max compression. One cell answers a different question from the rest of its column.

The compression figures also do not reconcile with the headline. The feature table says the technique extends context up to four point three times. The table's best balance row is three point seven times, and the maximum compression row is four point seven times, which is above the headline. A separate table of context budgets gives gains of three point eight and three point nine times.

So four different ratios appear for what a reader will reasonably assume is one number: the headline extension, the compression factors, and the measured gains. None is wrong on its own terms, and none is labelled clearly enough to tell which one to quote.

The fastest number in the file is a batched aggregate, not a decode rate

The main table has one column, decoding speed without speculative decoding, and it is the right column for the numbers in it.

The Apple section underneath does not use the same column. Its columns are batch size, output tokens, time and throughput. And the numbers there are the fastest in the entire file by an order of magnitude: a small model at batch size one hundred and twenty eight reports over seven hundred tokens per second, while the same model at batch size one reports under fifty.

That is not an inconsistency, it is two different measurements, and the batch column is what tells you so. Throughput at a large batch is aggregate tokens per second across concurrent sequences. Decode speed is tokens per second for one sequence. A reader who takes the seven hundred and compares it to the ninety on a large mixture-of-experts model on a datacentre card is comparing a throughput number to a latency number.

The Apple rows also cover quantisation levels from sixteen bit down to two bit, and the two bit row is the slowest in the table, which is the expected direction and is worth noticing because it is the only place in the file where the cost of extreme quantisation is visible in a time column.

Three of the sixteen benchmark rows are four-bit formats emulated in software

The feature table makes a specific claim: four-bit numeric formats with a low-bit cache on an older datacentre card, with no hardware four-bit support needed and coherent output.

The performance table is where to check it, because three rows carry a software annotation in parentheses.

One mixture-of-experts model appears twice. On a consumer card it runs at just under two hundred tokens per second. On the older datacentre card, with the four-bit format emulated in software, the same model at the same quantisation runs at just over seventy. That is a factor of nearly three, and it is the clearest statement in the file of what the software path costs.

Two more rows carry the same annotation, both on a newer datacentre card with the format emulated. One is a mixture-of-experts model at just under eighty tokens per second. The other is a two hundred and twenty-nine billion parameter mixture-of-experts model at just over sixty, and it is the only row in the table that needs two devices, with tensor parallelism set to two.

So a third of the benchmark evidence is about emulating a format the hardware does not have. That is worth knowing precisely because the feature table sells the emulation as the advantage.

The default container build compiles every optional path and resolves one version at build time

The container build declares five features as its default set: the CUDA path, the collective communication library, the Python bindings, a fast inference kernel library, and a kernel library. That is every heavy path in the project compiled into the image whether you use them or not.

It also declares a default compute capability of 8.0 and a thread pool size of thirty-two, and the base image is chosen by three build arguments, a CUDA version, a distribution version and a flavour.

One line in that build deserves attention. The collective communication library and its runtime are installed at whatever version the package manager's metadata offers for the major and minor CUDA version being requested, extracted with an inline script at build time. Nothing pins the result.

So the image is not reproducible from the repository. Two builds a month apart with the same arguments can carry different versions of a library that talks directly to the GPUs.

The last build argument is the interesting one. When it is set, the build rewrites the Rust toolchain's update host and its distribution server to two regional mirrors, and then replaces the crates registry with a third mirror by writing a source replacement into the toolchain's global configuration. That is a sensible accommodation, and it also means the same commit builds against a different registry depending on one flag, with the flag defaulting to off.

Editorial conclusion

xinfer is worth an afternoon if you are weighing whether you can drop a Python runtime from your serving stack, because the performance table is per model, per quantisation and per card, which is more honesty than most inference projects manage. Three things to settle first. The two crates everything depends on come from a fork under the maintainer's own account at a single commit, so you are auditing that commit rather than a published release. The container build resolves a communication library version at build time, so it is not reproducible. And the served endpoints are described with a bare address and no authentication, so do not put one on a network you do not control.

Frequently asked questions

Does xinfer need Python or PyTorch to run?

No. It is written in Rust with no PyTorch and no Python runtime for the inference path. There is an optional Python binding wheel when you want a Python entry point, and the container build includes the Python feature by default, so the runtime is optional but the bindings exist.

How do I install xinfer?

Two documented routes: a shell script piped from the project documentation site into an interpreter, or a global install of the package-registry name, which carries a suffix the repository and crate names do not. A Python wheel also exists as an optional entry point. The piped script takes no version argument and offers no checksum.

Which ports does xinfer serve and does it need a key?

Two. The built-in chat interface is on the higher port and the versioned API base URL is on the lower one. Both are documented as plain HTTP with an address placeholder rather than a loopback address, and the file describes no authentication, no bind flag and no transport security for either.

How much context can xinfer hold with the compressed cache?

The readme gives a budget table rather than one figure. On a mixture-of-experts model at four-bit weights, a seven gigabyte cache budget goes from seven hundred thousand tokens uncompressed to two point seven million with the best balance mode, a gain of three point nine times. The headline in the feature table quotes four point three times, which appears in neither table.

Can xinfer run four-bit models on a V100?

The feature table says yes, with the format emulated rather than supported by the hardware. The performance table is specific about the cost: the same thirty-billion-parameter mixture-of-experts model runs at just under two hundred tokens per second on a consumer card and just over seventy on the older datacentre card with software emulation. Three of the sixteen benchmark rows carry that annotation.

Official sources

  1. guoqingbao/xinfer on GitHub
  2. License: MIT
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/guoqingbao-xinfer.svg)](https://hysenlabs.com/projects/guoqingbao-xinfer)