# BentoML: building inference APIs in Python and shipping them as Docker images

> BentoML turns a Python inference script into an HTTP service with a decorator and a type hint, then packages the code, dependencies and model into a reproducible Docker image. It is Apache-2.0, requires Python 3.9 or newer, and the last push to main was on 2026-09-07.

**bentoml/BentoML** — The easiest way to serve AI apps and models - Build Model Inference APIs, Job queues, LLM apps, Multi-model pipelines, and more!

- Repository: https://github.com/bentoml/BentoML
- Website: https://bentoml.com
- Stars: 8,869 · Forks: 1,038
- Language: Python
- License: Apache-2.0
- Published: 2026-09-09 · Updated: 2026-09-09 · Language: en
- Canonical page: https://hysenlabs.com/projects/bentoml-bentoml

## The gap BentoML fills between a working script and a running endpoint

Most model code starts as a script: load weights, run a forward pass, print a result. Turning that into something another service can call means adding an HTTP framework, deciding how requests map to function arguments, handling concurrency, and then reproducing the environment somewhere else. BentoML is aimed at that second step. The README describes it as "a Python library for building online serving systems optimized for AI apps and model inference," and the pyproject metadata lists Python 3.9 through 3.12 with the classifiers for Scientific/Engineering and Artificial Intelligence.

The intended user is a Python developer who already has the model code and does not want to own the serving layer. The README's own framing is that a model inference script becomes a REST API server "with just a few lines of code and standard Python type hints." That is the whole pitch: the function signature is the API contract. If you are the person who writes the inference function, you are the person who writes the endpoint, and there is no separate config file mapping routes to handlers.

What it is not is a model registry or an experiment tracker. BentoML sits after training. The dependency list in pyproject.toml is all serving infrastructure: aiohttp, httpx, attrs, cattrs, click, Jinja2, PyYAML, prometheus-client, nvidia-ml-py, and a set of opentelemetry packages. There is nothing in there for training loops or dataset versioning.

## How a service class becomes an HTTP server

The mechanism is a decorated Python class. A @bentoml.service decorator marks the class, and each @bentoml.api method becomes a route. In the README example, a Summarization class builds a transformers pipeline in __init__ and exposes summarize(texts: list[str]) -> list[str]. The type hints are not decoration: they describe the request and response shape, which is why the same function can be called from the generated HTTP client without a hand-written schema.

The api decorator also takes serving options. The README example passes batchable=True, which is the switch that lets the server group concurrent calls into one invocation of the underlying pipeline. That is the dynamic batching feature the README lists under serving optimization, alongside model parallelism, multi-stage pipelines and multi-model inference-graph orchestration. The practical consequence is that batching is declared on the function rather than implemented by the caller, and the function has to accept a list and return a list for it to work.

Environment and model packaging live in the same file. The service decorator in the example carries image=bentoml.images.Image(python_version="3.11").python_packages("torch", "transformers"). That declaration is what the build step reads later. So the service file is three things at once: the API definition, the dependency manifest, and the input to the image builder. The README calls the resulting artifact a Bento, described as "the standardized deployable artifact in BentoML."

The runtime is not the development server you might expect from a Python framework. When you run bentoml serve, the README shows the log line \"Starting production HTTP BentoServer from \"service:Summarization\" listening on http://localhost:3000\". Port 3000 is the default, not 8000, which trips people who assume the usual Python web convention.

## Installing BentoML and serving a first model

Installation is a single pip command. The README notes it requires Python 3.9 or newer, which matches the requires-python field in pyproject.toml.

```bash
pip install -U bentoml
```

Create a file named service.py. The README's example wraps a summarization pipeline and exposes one batched endpoint. The image declaration inside the decorator is what the later build step will use, so it is worth writing even for a local run.

```python
import bentoml

@bentoml.service(
    image=bentoml.images.Image(python_version="3.11").python_packages("torch", "transformers"),
)
class Summarization:
    def __init__(self) -> None:
        import torch
        from transformers import pipeline

        device = "cuda" if torch.cuda.is_available() else "cpu"
        self.pipeline = pipeline('summarization', device=device)

    @bentoml.api(batchable=True)
    def summarize(self, texts: list[str]) -> list[str]:
        results = self.pipeline(texts)
        return [item['summary_text'] for item in results]
```

The service imports torch and transformers at call time, so install them into the same virtual environment before running. The README calls these "additional dependencies for local run."

```bash
pip install torch transformers
```

Start the server from the directory containing service.py. The README gives bentoml serve with no arguments, which discovers the service file.

```bash
bentoml serve
```

The expected output names the service and the address: a line starting with Starting production HTTP BentoServer and listening on http://localhost:3000. You can then open that URL in a browser, or call it from Python with the bundled client, which mirrors the method names on the service class.

```python
import bentoml

with bentoml.SyncHTTPClient('http://localhost:3000') as client:
    summarized_text: str = client.summarize([bentoml.__doc__])[0]
    print(f"Result: {summarized_text}")
```

Moving to a container is two more commands. bentoml build collects code, models and dependency configuration into the Bento, and bentoml containerize turns that into an image. The README's example tags it with the service name and latest.

```bash
bentoml build
bentoml containerize summarization:latest
docker run --rm -p 3000:3000 summarization:latest
```

Docker has to be running for the containerize step, and the README says so explicitly. There is a third path, bentoml cloud login followed by bentoml deploy, which targets BentoCloud and requires a signup. It is optional; the local serve and build flow does not need it.

## Where the abstraction costs you

The service class owns the process lifecycle, and that is the main constraint. Anything that needs to happen before the first request, such as downloading weights or warming a GPU, has to fit inside __init__, because that is the only hook the README shows. There is no documented pre-start hook, and no documented rollback for a deploy, so a bad release is a redeploy rather than an undo.

The batching flag has a cost that is easy to miss. Marking a method batchable=True tells the server to group concurrent requests, which means a caller's latency now depends on how many other callers arrived in the same window. Under light traffic, a request can wait for a batch that never fills. The README presents dynamic batching as a utilization feature; it is also a latency knob, and the documentation excerpt does not give a tuning parameter for the wait window.

Dependency resolution is delegated to the image builder, which reads the python_packages list from the service decorator. That is convenient until a framework needs a system library or a pinned CUDA build, at which point the declaration in the decorator may not be expressive enough and you are back to writing Dockerfile-level detail. The README does not document how to inject arbitrary system packages into the generated image.

Finally, the project is a serving framework, not a model server. There is no bundled inference engine for large language models in the core dependency list. The README points to separate example repositories such as BentoVLLM for Llama and Mistral, which means the LLM path is a composition of BentoML plus another runtime, not a single install.

## BentoML against FastAPI and against dedicated model servers

The most direct comparison is FastAPI, because BentoML uses the same underlying idea of Python functions as endpoints. The difference is scope. With FastAPI you write the route decorators, the Pydantic models, the Dockerfile and the dependency pinning yourself, and you get a general web framework that can do anything. With BentoML the route decorator is @bentoml.api, the schema comes from type hints, and the Dockerfile is generated from the image declaration in the service decorator. You trade control over the packaging layer for not having to maintain it. If your service needs middleware, background tasks or a database session per request, FastAPI's surface is wider and the BentoML docs excerpt does not describe equivalents.

Against a dedicated inference server such as vLLM or Triton, the split is different again. Those tools specialize in running one class of model efficiently, with their own scheduling and memory management. BentoML's selling point is composition: multiple models, multiple frameworks and custom Python logic in one service, with the README listing model composition and workers as advanced topics. If your entire workload is one transformer served at high throughput, a specialized server is the narrower and more direct answer. If your workload is a retrieval step, a reranker and a generator behind one endpoint, the composition story is why you would pick BentoML over either.

Against MLflow, the comparison is mostly a category error, but people search for it, so it is worth stating: MLflow tracks experiments and registers models, while BentoML serves them. They can sit in the same stack without overlapping. The same applies to Ray Serve and KServe, which operate at the orchestration and cluster layer rather than the Python service definition layer.

## Licence, release cadence and what upgrading involves

BentoML is Apache-2.0, declared in both the README badge and the pyproject license field. For most teams that means permissive use, modification and redistribution, with the usual obligations around notices and patent terms. This is not legal advice; if you redistribute a modified version or embed it in a product, read the licence text in the repository rather than a summary.

The release history shows a steady patch cadence: v1.4.37 on 2026-03-25, v1.4.38 on 2026-04-02, and v1.4.39 on 2026-05-07. The last push to main was on 2026-09-07, so the repository is being worked on between releases. The version numbering suggests a stable 1.4 line with incremental changes rather than breaking releases, but the README does not include an upgrade guide, so the practical way to judge the cost of moving from one patch to the next is the release notes for the version you are installing.

Upgrade cost is dominated by the generated image, not the library. Because the service decorator pins python_version and a python_packages list, a bump in BentoML can change how that list is resolved, and the resulting image is what your deployment actually runs. Rebuilding and comparing the image is the check that matters. The opentelemetry and prometheus-client dependencies in pyproject.toml also mean observability wiring ships with the library, so a version bump can move instrumentation behavior even when your service code is untouched.

## Conclusion

Adopt BentoML if you already write Python inference code and want an HTTP endpoint plus a Docker image from the same file, without hand-writing a FastAPI layer or a Dockerfile. Skip it if you only need to serve one transformer model and want a dedicated LLM server, or if you cannot run Docker in your build environment, since bentoml containerize depends on it. Before committing, check the changelog for the release you install, confirm that the image builder can resolve your framework's wheels for your Python version, and decide whether you need BentoCloud at all, because the local serve and build path is complete on its own.

## FAQ

### Is BentoML free?

Yes. The project is licensed under Apache-2.0, and the README badge and pyproject license field both state that. BentoCloud, the hosted deployment target the README links to, is a separate commercial service that requires a signup.

### How to install BentoML?

Install it with pip, which requires Python 3.9 or newer according to the README. The single command is pip install -U bentoml; framework packages such as torch and transformers are installed separately for a local run.

### What is BentoML used for?

It is used to turn Python model inference code into an HTTP API server. The README describes it as a library for building online serving systems, and the same service file also declares the dependencies used to generate a Docker image.

### Is BentoML open source?

Yes. The source is on GitHub under the Apache-2.0 licence, and the repository contains the full source tree under src/ along with tests, examples and documentation.

### What is OpenLLM?

The README does not describe OpenLLM. It lists separate example repositories for large language models, such as BentoVLLM for Llama and Mistral, but does not explain what OpenLLM is.

### How does BentoML compare with vLLM?

The README does not make that comparison. It points to BentoVLLM as an example repository for serving Llama and Mistral, which indicates the two are used together rather than as substitutes, but it gives no benchmark or feature comparison.

## Sources

- [bentoml/BentoML on GitHub](https://github.com/bentoml/BentoML)
- [License: Apache-2.0](https://github.com/bentoml/BentoML/blob/main/LICENSE)
- [Project website](https://bentoml.com)
- [README](https://github.com/bentoml/BentoML/blob/main/README.md)
- [Releases](https://github.com/bentoml/BentoML/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/bentoml-bentoml
