# Xinference: one API for local LLMs, speech and multimodal models

> Xorbits Inference (Xinference) serves language, speech recognition and multimodal models behind a single API, installed with pip or Docker. Its value is model breadth behind one endpoint; its cost is that the model catalogue, not the API, decides what you can run.

**xorbitsai/inference** — Swap GPT for any LLM by changing a single line of code. Xinference lets you run open-source, speech, and multimodal models on cloud, on-prem, or your laptop — all through one unified, production-ready inference API.

- Repository: https://github.com/xorbitsai/inference
- Website: https://inference.readthedocs.io
- Stars: 9,576 · Forks: 866
- Language: Python
- License: Apache-2.0
- Published: 2026-09-09 · Updated: 2026-09-09 · Language: en
- Canonical page: https://hysenlabs.com/projects/xorbitsai-inference

## The problem Xinference solves: one API in front of many open models

Running open models usually means running several tools. vLLM for one family, llama.cpp for quantized weights, a separate diffusers script for image models, another process for Whisper. Each has its own launch flags, its own HTTP surface, its own way of reporting readiness. The README positions Xinference against exactly that spread: it is "designed to serve language, speech recognition, and multimodal models" and lets you deploy built-in models "using just a single command."

Who this is for: teams that want an OpenAI-shaped endpoint on their own hardware and do not want to write glue code per model family. The README names researchers, developers and data scientists. The repository also ships a frontend directory and a monitor directory, which suggests a bundled web UI and a monitoring component rather than a library-only package. That matters if you plan to hand the endpoint to people who will not use curl.

Where it is the wrong fit is worth stating early. If you have standardized on one engine because you have tuned it, Xinference sits between you and that engine. It is a serving layer with opinions about which backends to use, not a thin proxy.

## How the serving layer is put together

The dependency list in pyproject.toml is the clearest picture of the architecture. xoscar appears first, and xoscar is the actor framework from the same organization, so workers are actors rather than plain processes. The README confirms this direction with a "Distributed inference: running models across workers" entry. A supervisor process launches workers, and a model is placed on a worker with enough resources.

Above the workers sits a FastAPI application (fastapi is pinned to >=0.110.3,<0.137, with uvicorn as the server and sse_starlette for streaming responses). JSON responses go through orjson. The openai package is a declared dependency, and the README describes the surface as a "unified, production-ready inference API," which matches the OpenAI-compatible endpoint the project is known for. An OpenAI-compatible route means existing clients that speak that protocol can point at a local base URL instead of a hosted one.

Backend selection is the other half. The topics list names vllm, sglang, llama-cpp, transformers and diffusers, so the same server can put a request through different engines depending on the model. The README also points to Xllamacpp, a llama.cpp Python binding maintained by the Xinference team that "supports continuous batching." That is a real design commitment: rather than depend on an upstream binding, the project keeps its own.

The cost of this breadth is dependency weight. torch is a top-level dependency, not an extra, and the Python floor is 3.10 with classifiers up to 3.14. A base install is not small, and the engine you actually want may pull more.

## Install Xinference and serve a first model

The README links to a self-hosting installation page rather than inlining the pip command, so treat the exact package extras as something to confirm on that page. The distribution name is xinference, which is what the PyPI badge and the [project] table in pyproject.toml both use. The PyPI page for the package is the authoritative source for the install command, and the README's installation link points there.

The project also publishes a Docker image. The README's Docker badge points at the xprobe/xinference repository on Docker Hub, so that is the image name to pull.

Once installed, the entry point is a single command. The README's headline claim is that you deploy models "using just a single command," and the CLI is built on click (pinned to <8.2.0 in pyproject.toml). The README does not print the exact launch invocation, so read the getting-started page for the current syntax rather than copying a flag from a blog post.

After launch, the server exposes an HTTP API. Because the surface is OpenAI-compatible, a client that already speaks that protocol can be pointed at the local endpoint by changing the base URL. The README frames the whole project this way: swap GPT for any LLM by changing a single line of code.

## The catalogue is the ceiling

The most consequential limitation is not performance, it is coverage. Xinference serves built-in models. The README's new-model entries are a long list of specific checkpoints with pull-request links: Kimi-K3, GLM-5.2, GLM-Image, DeepSeek-V4-Flash-0731, Krea 2, ACE-Step 1.5, Ornith 1.5, WeMM-Embedding, NaviDC-OCR, and several world models. Each of those is a deliberate integration.

The flip side: a checkpoint that is not on that list is not served by pointing Xinference at a directory. You either wait for an integration or you run the model through its own engine and lose the unified API. For teams that track new releases weekly, that gap is the practical constraint, not throughput.

There is a second, quieter cost. Version 3.0.0 shipped with "migration notes and breaking changes," per the README's Hot Topics section. A serving layer that owns the request path, the model catalogue and the worker protocol will have breaking changes across major versions, and 3.1.0, 3.2.0 and 3.3.0 followed within roughly two months of each other according to the release list. If your deployment pins versions and upgrades on a quarterly cadence, budget for reading migration notes each time.

The README does not document rollback procedures for a failed model launch or a failed upgrade. That silence is worth noting before you put it in front of production traffic.

## How Xinference differs from running vLLM or llama.cpp directly

The honest alternative is not another model server. It is running the engine yourself. vLLM and llama.cpp are both listed in the repository topics, and Xinference can use them as backends. So the real question is what the extra layer buys.

Running vLLM directly gives you one engine, one configuration surface, and a direct line to that project's tuning knobs. You get its scheduling behavior without an intermediary deciding how to invoke it. What you do not get is a second model family in the same process, or a speech model next to your chat model, or a web UI that lists what is loaded. Xinference's distributed inference work and its shared KV cache across replicas (referenced in the README's Hot Topics) are features of the serving layer, not of any single engine.

llama.cpp is the closer comparison for laptop and CPU-bound use. It is designed for quantized weights and modest hardware, and Xinference's own Xllamacpp binding exists because the team wanted continuous batching there. If your entire workload is one quantized model on one machine, Xinference adds a supervisor, an HTTP layer and a dependency tree around something you could start with a single binary.

The dividing line: pick Xinference when the variety of models you must serve is the problem. Pick a single engine when the depth of tuning on one model is the problem.

## Licence, maintenance and what an upgrade costs

Xinference is Apache-2.0, declared in pyproject.toml as license = "Apache-2.0" with license-files = ["LICENSE"]. Apache-2.0 is a permissive licence with an explicit patent grant and a requirement to preserve notices. That is a statement about the project's own code. It says nothing about the licences of the model weights you load through it, and those vary widely across the families listed in the README. Checking the model card for each checkpoint you serve is your responsibility, not something the server resolves for you. This is not legal advice.

The repository is not archived, and the last push was on 2026-09-09, so the project is being worked on. The release cadence supports that: v3.1.0 on 2026-07-31, v3.2.0 on 2026-08-15, v3.3.0 on 2026-08-30. Roughly two-week intervals.

Upgrade cost is the part teams underestimate. A minor version bump can add models, which is free, or change the request path, which is not. The 3.0.0 notes are described as containing breaking changes, so the 3.x line is where you should expect to read before you deploy. The practical sequence is to pin the version in your environment, read the release notes for the target version, and launch one model on a staging worker before rolling the supervisor.

## Conclusion

Adopt Xinference if you need one endpoint in front of many open model families (LLM, speech, multimodal) and you are willing to let the built-in catalogue decide which checkpoints you can serve. Skip it if you need a single engine tuned to one model family, or if every dependency in your stack must be pinned for years. Before committing, check the built-in model list for the exact checkpoint you need, and read the v3.0.0 migration notes if you are upgrading from a 2.x deployment, since the release notes describe breaking changes.

## FAQ

### How do I install Xinference?

The README points to the self-hosting installation page in the documentation, and the distribution is published on PyPI under the name xinference. The project also publishes a Docker image at xprobe/xinference on Docker Hub, per the README's Docker badge.

### What is Xinference used for?

It serves language, speech recognition and multimodal models behind one API. The README describes deploying built-in models with a single command and lists support for LLM, image, embedding, OCR and world-model families.

### Can Xinference serve any model, or only specific ones?

Only built-in models. The README's new-model entries are individual integrations such as Kimi-K3, GLM-5.2 and GLM-Image, each added through a pull request, so a checkpoint outside that catalogue is not served by pointing the server at a directory.

### What are the breaking changes in Xinference 3.0.0?

The README links to release notes at xinference.co for v3.0.0 and states that the release includes migration notes and breaking changes. The README itself does not enumerate them, so the linked release notes are the place to read before upgrading.

### What licence does Xinference use?

Apache-2.0. pyproject.toml declares license = "Apache-2.0" and license-files = ["LICENSE"]. That covers the project's code, not the weights of the models you load through it.

## Sources

- [License: Apache-2.0](https://github.com/xorbitsai/inference/blob/main/LICENSE)
- [Project website](https://inference.readthedocs.io)
- [README](https://github.com/xorbitsai/inference/blob/main/README.md)
- [Releases](https://github.com/xorbitsai/inference/releases)
- [xorbitsai/inference on GitHub](https://github.com/xorbitsai/inference)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/xorbitsai-inference
