# OpenLLM: self-hosting Llama, Qwen and DeepSeek behind an OpenAI-compatible API

> BentoML's OpenLLM turns a model name into a running OpenAI-compatible endpoint on port 3000, with a chat UI and a CLI chat mode. It is a serving layer for teams that already have GPUs, not a desktop model runner.

**bentoml/OpenLLM** — Run any open-source LLMs, such as DeepSeek and Llama, as OpenAI compatible API endpoint in the cloud.

- Repository: https://github.com/bentoml/OpenLLM
- Website: https://bentoml.com
- Stars: 12,535 · Forks: 841
- Language: Python
- License: Apache-2.0
- Published: 2026-09-09 · Updated: 2026-09-09 · Language: en
- Canonical page: https://hysenlabs.com/projects/bentoml-openllm

## What OpenLLM actually solves for a team with GPUs

The gap OpenLLM targets is the distance between a model checkpoint on Hugging Face and an HTTP endpoint your application can call. Without a serving layer you write model loading, batching, a request schema and a streaming response format yourself, and then every client that talks to your service needs custom code. OpenLLM collapses that into one command that produces an OpenAI-compatible API, so existing OpenAI SDKs and frameworks point at your machine instead. The README states the project runs any open-source LLMs or custom models as OpenAI-compatible APIs with a single command, and lists a chat UI, inference backends and a path to Docker and Kubernetes deployment as the surrounding pieces. The audience is therefore infrastructure and platform engineers inside companies that already own GPU capacity. It is not aimed at someone who wants a model on a laptop for personal use, and the required-GPU column in the model table makes that explicit: phi4:14b is listed at 80G, llama3.1:8b at 24G. Those are deployment targets, not suggestions.

## The model repository is the mechanism, not the CLI

The interesting design decision is that OpenLLM separates the model catalog from the serving binary. A model repository is a catalog of available LLMs, and OpenLLM ships a default one hosted in the bentoml/openllm-models GitHub repository. Commands like openllm model list read that catalog, openllm repo update synchronizes your local list with connected repositories, and openllm model get inspects a single entry. When you run openllm serve llama3.2:1b, the tag is resolved against the catalog, which carries the metadata the server needs, including the GPU requirement shown in the README table. That indirection is why adding a custom model means adding a model repository rather than patching the server. It also means the catalog is a dependency you inherit: a tag that exists today is defined by an external repository, and the README documents a repo update command precisely because the local list can drift from upstream. The practical consequence is that pinning a tag in a deployment manifest is not the same as pinning the model definition. If you care about reproducibility, record which repository and revision the tag resolved to when you last ran the update.

## Installing OpenLLM and calling your first endpoint

Installation is a single pip package, and the README pairs it with an interactive command meant to show you what the tool does before you commit a GPU to it. The pip line installs the openllm package; openllm hello then walks you through the project interactively. Run both, and expect the hello command to be exploratory rather than a server.

```bash
pip install openllm  # or pip3 install openllm
openllm hello
```

Before starting a server for a gated checkpoint such as meta-llama/Llama-3.2-1B-Instruct, the README says OpenLLM does not store model weights and a Hugging Face token is required for gated models. The documented sequence is to create a token, request access to the gated model, and export the token as an environment variable. The variable name is HF_TOKEN.

```bash
export HF_TOKEN=<your token>
```

With the token in place, one command starts the server. The README states the server is accessible at http://localhost:3000 and exposes OpenAI-compatible APIs.

```bash
openllm serve llama3.2:1b
```

To confirm the endpoint answers, point the OpenAI Python client at the base URL with a placeholder API key. Note that the model string in the request is the full checkpoint name, not the short tag you served. The README's example uses meta-llama/Llama-3.2-1B-Instruct and streams the response.

```python
from openai import OpenAI

client = OpenAI(base_url='http://localhost:3000/v1', api_key='na')

chat_completion = client.chat.completions.create(
    model="meta-llama/Llama-3.2-1B-Instruct",
    messages=[{"role": "user", "content": "Explain superconductors like I'm five years old"}],
    stream=True,
)
for chunk in chat_completion:
    print(chunk.choices[0].delta.content or "", end="")
```

If you would rather not write client code first, two shortcuts exist. The chat UI lives at the /chat endpoint on the same host, so http://localhost:3000/chat gives you a browser interface, and openllm run llama3:8b starts a chat conversation in the terminal.

## Where OpenLLM is the wrong tool

The README does not document rollback, and it does not describe what happens when a model tag is removed or renamed in the upstream repository. That silence matters more than it first appears, because the tag is the unit of deployment here. A service defined as openllm serve llama3.2:1b depends on a catalog entry that lives outside your repository, and the documented remedy for drift is to run openllm repo update, which pulls the latest list rather than freezing it. Teams that need a frozen artifact should treat the catalog as something to snapshot themselves. The second limitation is hardware. The supported-model table attaches a required GPU to every row, from 12G for gemma2:2b up to 80Gx16 for deepseek:r1-671b, and the project classifiers list NVIDIA CUDA 11.7, 11.8 and 12 environments. There is no CPU path described in the README, so a machine without a supported NVIDIA GPU is outside the documented envelope. Third, OpenLLM is a server, not a client library. If your goal is to route requests across hosted providers, or to call a model that runs somewhere else, the serving machinery is weight you carry for nothing. And if you only need one model for one script, the install plus a multi-gigabyte weight download is a heavier path than calling a hosted API.

## OpenLLM compared with Ollama and vLLM

The README links a design-philosophy post titled from Ollama to OpenLLM running LLMs in the cloud, which frames the contrast directly. Ollama is the tool people reach for when they want a model running on their own machine with minimal setup, and the OpenLLM documentation positions itself on the other side of that line: cloud deployment, Docker and Kubernetes, enterprise-grade serving. If your constraint is a workstation and a single user, the OpenLLM model table is a poor fit before you write any code. Against vLLM the difference is scope rather than raw inference. vLLM is an inference engine you embed in a service you build; OpenLLM is the service, with the catalog, the OpenAI-compatible surface, the chat UI and the deployment story already assembled. Choosing OpenLLM means accepting its catalog format and its tag resolution in exchange for not writing the serving layer. Choosing vLLM means writing that layer but keeping full control over how models are identified and loaded. Neither is a strict upgrade; the decision is whether the catalog abstraction earns its keep in your environment.

## Licence, maintenance and what an upgrade costs

OpenLLM is Apache-2.0, and pyproject.toml carries the Apache Software License classifier alongside Development Status 5 - Production/Stable. Apache-2.0 permits commercial use and modification and includes a patent grant, which is why it is a common choice for serving infrastructure; it also means you take on the obligation to preserve notices and state changes. That is a summary of the identifier, not legal advice, and the LICENSE file in the repository is the text that governs. The pinned dependencies are worth reading before an upgrade: pyproject.toml pins bentoml==1.4.23 and openai==1.90.0 exactly, while typing-extensions carries a floor of >=4.12.2. An exact pin on BentoML means an OpenLLM release is coupled to one BentoML version, so upgrading OpenLLM can move your BentoML version underneath you. The last push to the repository was on 2026-09-07, and the most recent release listed is v0.6.30 from 2025-04-21, with v0.6.29 and v0.6.28 both dated 2025-04-16. Read that gap as it is: commits continue, tagged releases are less frequent. Before upgrading, check whether the model tags you depend on still resolve and whether a BentoML version change affects anything else in the same environment.

## Conclusion

Adopt OpenLLM if you have the GPUs the model table lists and you want an OpenAI-compatible endpoint plus a chat UI without writing serving code; the deepseek r1-671b row asking for 80Gx16 is a fair picture of how literal those requirements are. Skip it if you are on a laptop without a supported GPU, or if you only need to call hosted models, since OpenLLM does not store weights itself and expects a Hugging Face token for gated checkpoints. Before committing, run openllm model get on the exact tag you plan to deploy and confirm the required GPU line matches your hardware, then check whether the checkpoint you want is gated and whether you already hold access to it.

## FAQ

### What is OpenLLM?

OpenLLM is a Python tool from BentoML that runs open-source LLMs as OpenAI-compatible API endpoints. The README states it can serve models such as Llama 3.3, Qwen2.5 and Phi3 with a single command, and it includes a chat UI plus a workflow for Docker and Kubernetes deployment.

### How to install OpenLLM?

The README gives one install command, pip install openllm, followed by openllm hello for an interactive walkthrough. After that, openllm serve llama3.2:1b starts a server on http://localhost:3000.

### How to use OpenLLM?

Start a server with openllm serve and a model tag, then call http://localhost:3000/v1 with any OpenAI-compatible client. The README also documents a browser chat UI at the /chat endpoint and a terminal chat mode via openllm run llama3:8b.

### OpenLLM vs vLLM: what is the difference?

The README does not compare the two, so the difference has to be read from scope. vLLM is an inference engine you embed in a service you write; OpenLLM is the assembled service, with a model catalog, an OpenAI-compatible API and a chat UI already in place.

### Is there a library for LLMs in Python?

OpenLLM is one, distributed on PyPI as the openllm package and installed with pip install openllm. It is a serving library rather than a training or fine-tuning one: the README's workflow is installing it, starting a server with openllm serve, and calling the resulting endpoint.

## Sources

- [bentoml/OpenLLM on GitHub](https://github.com/bentoml/OpenLLM)
- [License: Apache-2.0](https://github.com/bentoml/OpenLLM/blob/main/LICENSE)
- [Project website](https://bentoml.com)
- [README](https://github.com/bentoml/OpenLLM/blob/main/README.md)
- [Releases](https://github.com/bentoml/OpenLLM/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/bentoml-openllm
