LitServe: a pure-Python inference server you write yourself
A minimal Python framework for building custom AI inference servers with full control over logic, batching, and scaling.
At a glance
- What is it?
- LitServe hands request handling, batching and routing back to your own Python class and keeps only the serving machinery. It is a good fit for agents, RAG pipelines and multi-model endpoints, and the wrong fit if you want a finished model server.
- Who is it for?
- Adopt LitServe when the inference logic is the product: multi-model pipelines, agents, RAG, or anything vLLM's request schema cannot express, and when you are willing to own the model loading code.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 8 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem LitServe solves: serving logic that does not fit a fixed schema
Most serving tools assume one model behind one request shape. The README states the case plainly: most serving tools "are built for a single model type and enforce rigid abstractions," and they work until you need custom logic, multiple models, agents, or non-standard pipelines. LitServe's answer is to let you write the inference engine in Python. You define how requests are handled, how models are loaded, how batching and routing work, and how outputs are produced. The framework keeps concurrency, scaling and deployment.
The intended audience is a Python developer who already has a working model or pipeline and now needs an HTTP surface for it. The README lists inference APIs, agents, chatbots, RAG systems, MCP servers and multi-model pipelines as targets. That is a wide net, and the shared trait is that the steps between the HTTP request and the model call are yours, not the framework's.
How LitAPI and LitServer split the work
The mechanism is two classes. You subclass ls.LitAPI and implement setup(self, device) to load models, plus predict(self, request) to answer a request. LitServer wraps that class, takes an accelerator argument, and exposes the HTTP endpoint. In the README's inference engine example, setup assigns two lambdas to self.text_model and self.vision_model, and predict reads request["input"], calls both, adds the results and returns {"output": c}. Nothing about that is model-specific; it is ordinary Python running inside a request handler.
The dependency list in pyproject.toml shows what sits underneath: fastapi, pyzmq and uvicorn[standard]. FastAPI and uvicorn handle HTTP; ZeroMQ is the transport between the web process and the worker processes. That is where the worker model comes from, and it is why batching can be configured per server rather than per route. The README also advertises autoscaling, GPU support and streaming, with batching enabled through constructor arguments such as max_batch_size. The README does not document the full set of constructor options; it points to the LitServer API reference for that.
Install LitServe and send your first request
The README gives one install command. It does not pin a version, and the install page is linked for other options such as accelerator-specific extras.
pip install litserveWrite a server file. This is the README's inference engine example in full, including the port argument.
import litserve as ls
class InferenceEngine(ls.LitAPI):
def setup(self, device):
self.text_model = lambda x: x**2
self.vision_model = lambda x: x**3
def predict(self, request):
x = request["input"]
a = self.text_model(x)
b = self.vision_model(x)
return {"output": a + b}
if __name__ == "__main__":
server = ls.LitServer(InferenceEngine(max_batch_size=1), accelerator="auto")
server.run(port=8000)Run it with python server.py, or with the CLI the README shows: lightning deploy server.py for a local run, and lightning deploy server.py --cloud for Lightning's hosted option. Then post to /predict. The README's curl returns the computed value for input 4.0.
curl -X POST http://127.0.0.1:8000/predict -H "Content-Type: application/json" -d '{"input": 4.0}'Two things to note before adapting this. The endpoint is /predict, and the request body keys are whatever predict reads; LitServe does not generate a schema from your class. The second is that setup receives a device argument, so model placement is your decision inside that method.
Where LitServe gets in the way
The freedom is the cost. Because predict is yours, so is every failure inside it: a slow model call blocks that worker, and the README does not describe a timeout or cancellation policy for in-flight requests. If your pipeline calls an external API, as the news agent example does with openai.OpenAI, the latency of that call is now part of your server's latency, and the framework has no opinion about it.
Batching is opt-in and configured at construction. The quick start passes max_batch_size=1, which means no batching at all; raising it changes how requests are grouped, and the README does not state the wait window or how a batch is assembled. Anyone tuning throughput has to read the LitServer reference rather than the README.
The project also labels itself Beta in pyproject.toml (Development Status :: 4 - Beta) and the release history is 0.2.x, with v0.2.19 on 2026-09-09. That is fine for a service you control, and less fine if you need a frozen API surface across a long release cycle. Finally, LitServe is not a model server. It does not ship weights, tokenizers or a chat completions schema. If your requirement is "serve this checkpoint over an OpenAI-compatible endpoint," you are writing that layer yourself.
LitServe compared with FastAPI and Ray Serve
The README's own comparison is against FastAPI, and it claims LitServe is 2x faster. Treat that as the project's claim, not a measured result here; the number appears in the feature list with no stated workload or hardware. The structural difference is more useful than the multiplier. A plain FastAPI app gives you routing and validation and leaves worker processes, model loading and batching to you. LitServe supplies those pieces and asks you to fill in setup and predict. If your endpoint already has a clean FastAPI implementation with no batching needs, LitServe adds a layer without removing much.
Ray Serve sits at the other end. It is a general serving layer over Ray's distributed runtime, so scaling across a cluster and composing multiple deployments are first-class, at the cost of adopting Ray's programming model and operating a Ray cluster. LitServe's scaling story, as described in the README, is accelerator selection plus autoscaling on Lightning's cloud or your own host. If your problem is distributing one model across many nodes, Ray Serve is the more direct tool; if your problem is expressing an unusual pipeline on one machine or one GPU box, LitServe keeps the code smaller. The README does not document a migration path from either tool.
Licence, releases and what upgrades cost you
LitServe is Apache-2.0, declared in pyproject.toml and in the LICENSE file at the repository root. That is a permissive licence with an explicit patent grant, and it does not require you to publish your own server code. It also means no warranty and no support obligation from the maintainers. The README points users to a Discord for help, which is where community support lives.
Upgrade cost is tied to the release cadence. v0.2.17 landed on 2025-12-23, v0.2.18 on 2026-08-12 and v0.2.19 on 2026-09-09, with the last push to main on 2026-09-10. The gap between the December and August releases suggests the cadence is uneven, so pinning a version in your own requirements and reading the release notes before moving is cheaper than tracking main. The repository also carries GOVERNANCE.md and CONTRIBUTING.md, so the project has stated processes, though neither is summarised in the README.
One practical detail from pyproject.toml: requires-python is >=3.10. If you are still on 3.9, LitServe is not installable without an interpreter upgrade.
Editorial conclusion
Adopt LitServe when the inference logic is the product: multi-model pipelines, agents, RAG, or anything vLLM's request schema cannot express, and when you are willing to own the model loading code. Skip it if you want a finished server for a single Hugging Face checkpoint, or if you cannot maintain a Python service that tracks a project at 0.2.x. Before committing, run the quick start locally, confirm the batching defaults against your own latency budget, and check the install page for the extras your accelerator needs.
Frequently asked questions
What is an AI server?
In LitServe's framing it is a process that accepts requests over HTTP and runs your inference logic for each one. You supply that logic by subclassing ls.LitAPI and implementing setup and predict; LitServer handles the HTTP layer and the worker transport.
What are the alternatives to LitServe?
The README contrasts LitServe with vLLM, which it describes as built for a single model type with rigid abstractions, and with FastAPI, which it claims to be 2x faster than. Ray Serve is a general serving layer over Ray's distributed runtime for cluster-scale deployments.
How do I install LitServe?
The README gives a single pip command: pip install litserve. The install page linked from the README covers other options, and pyproject.toml sets requires-python to >=3.10.
Does LitServe support batching and streaming?
The README lists batching and streaming among the features and shows max_batch_size passed to the LitAPI subclass in the quick start, where it is set to 1. The README does not document the batching wait window or the full set of constructor options; those are in the LitServer API reference.
What licence does LitServe use?
Apache-2.0, declared in pyproject.toml and in the LICENSE file at the repository root. The README also carries an Apache 2.0 licence badge.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/lightning-ai-litserve)