Model or dataset
mosecorg/mosec avatar
mosecorg/mosec

Mosec: a Rust web layer in front of your Python model

A high-performance ML model serving framework, offers dynamic batching and CPU/GPU pipelines to fully exploit your compute machine

902 stars74 forksPythonApache-2.0

At a glance

What is it?
Mosec wraps a Python Worker class in a Rust HTTP server that does dynamic batching and pipelined multi-process stages. It fits teams that already have a forward function and want it behind an API without writing a serving stack.
Who is it for?
Adopt Mosec if you have a Python forward function and want batching plus multi-stage pipelining without building the HTTP and coordination layers yourself; the Rust side handles that and the Python side stays testable offline. Do not adopt it if you need a model registry, autoscaling or a workflow DAG, because the project states it focuses on the online serving part only.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 1 day ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The gap Mosec fills between a trained model and an HTTP endpoint

Most serving stacks ask you to adopt a config format, a model registry or a graph definition before you can answer a request. Mosec inverts that. You write a Python class that inherits mosec.Worker, put your model in __init__, and put your inference code in a forward method. The framework supplies the HTTP server, the request queue, batching and process management around it.

The project describes itself as bridging "the gap between any machine learning models you just trained and the efficient online service API." The intended user is a Python engineer who already has working inference code and does not want to rewrite it in a serving DSL. The README's stable diffusion example keeps the diffusers call inside forward, unchanged from what you would run offline.

The framework is framework-agnostic on the model side. The topics list covers PyTorch, TensorFlow, JAX and MXNet, and the examples directory includes jax_single_layer, distil_bert_server_pytorch and resnet50_msgpack. Nothing in the Worker interface references a specific tensor library.

Where it is not aimed: Mosec does not manage model versions, does not schedule training jobs, and does not give you a DAG editor. The README states the design goal as "Do one thing well": serving is the scope, and model optimization and business logic stay with you.

How the Rust coordinator and Python workers split the work

The split is the core design decision. The web layer and task coordination are written in Rust, per the README, and the user interface is purely Python. Cargo.toml shows the concrete pieces on the Rust side: axum for HTTP, tokio for the async runtime, async-channel for the message passing, prometheus-client for metrics, and utoipa with utoipa-swagger-ui for the generated API documentation.

A request arrives at the Rust server, which places it on a queue. The coordinator aggregates requests from different users into a batch and hands that batch to a worker process. The worker calls your forward with a list when dynamic batching is configured, and the results are distributed back to the individual waiting requests. That is what the README means by "aggregate requests from different users for batched inference and distribute results back."

For multi-stage services, you chain workers. The README describes the arrangement: the first worker takes the request, the last worker sends the response, and data moves between workers over inter-process communication. Stages can be assigned to different processes so a CPU-bound preprocessing step and a GPU-bound inference step do not block each other. The repository includes examples/shm_ipc and examples/segment, which are the two shapes this usually takes.

The Python side stays isolated in its own process. That matters for a practical reason: a segfault or a CUDA out-of-memory in your model code should not take down the HTTP listener. It also means your model object is not shared across threads, so you do not have to reason about Python's GIL during inference.

Installing Mosec and serving a first model

Mosec ships as a PyPI package for Linux and macOS. The README lists three install paths, and the conda and pixi variants resolve the same wheel:

bash
pip install -U mosec
# or install with conda
conda install conda-forge::mosec
# or install with pixi
pixi add mosec

Building from source requires Rust and the Makefile target `make package`, which runs maturin and writes a wheel into the dist folder:

bash
make package

For a first service, the README's stable diffusion example is the reference shape. You need the diffusers and transformers packages first:

bash
pip install --upgrade diffusers[torch] transformers

The server itself is a class. Note the example attribute, which is the warmup input, and the forward signature, which takes a list when batching is on:

python
from mosec import Server, Worker, get_logger
from mosec.mixin import MsgpackMixin

class StableDiffusion(MsgpackMixin, Worker):
    def __init__(self):
        self.pipe = StableDiffusionPipeline.from_pretrained(
            "sd-legacy/stable-diffusion-v1-5", torch_dtype=torch.float16
        )
        self.example = ["useless example prompt"] * 4  # warmup (batch_size=4)

    def forward(self, data: List[str]) -> List[memoryview]:
        res = self.pipe(data)
        return [to_jpeg_bytes(img) for img in res[0]]

The README is explicit about what the warmup example does: if you set it, the service is not marked ready until that example has passed through the handler. If you leave it unset, the first real request pays the longer latency. That is a real behaviour to plan around, because a readiness probe that passes before the GPU is allocated will send traffic into a cold process.

Serialization is a second choice you make at class definition time. Without a mixin, Mosec uses JSON. Inheriting MsgpackMixin switches to msgpack, which the README recommends when you return binary data, since JSON would force base64 and inflate the payload. The mixin module also lists other formats if neither fits.

Finally, the Server object wraps the worker and starts the process. The README's full example constructs `Server()` and registers the worker class with it; the documentation site at mosecorg.github.io/mosec carries the complete runnable version and the CLI invocation.

Dynamic batching is the feature that decides your throughput

Batching is where the framework earns its place. Static batching means you fix a batch size and wait until it fills, which adds latency to every request. Mosec's dynamic batching collects whatever requests are in flight and forms a batch from them, so a quiet period produces small batches and a burst produces large ones. The README frames this as aggregating requests from different users.

The consequence for your code is the forward signature. The README states the contract directly: `forward(self, data: Any | List[Any]) -> Any | List[Any]`, and whether you receive a single item or a tuple depends on whether dynamic batching is configured. This is the single most common source of confusion, because the same forward method behaves differently under two configurations. If you write `forward(self, data: List[str])` and then disable batching, you will index into a string.

The batch size is bounded by configuration, and the README points at its Configuration section for the limits. Those limits are what keep a large burst from allocating a tensor that does not fit in GPU memory. The trade-off is that a batch limit plus a wait window is a latency decision, and Mosec does not choose it for you. A text-to-image model with a batch of four has very different tail latency from a four-item embedding batch, and the same configuration value serves both.

There is a second subtlety with warmup. The example in the README is `["useless example prompt"] * 4`, an explicit comment that the warmup batch size is four. If your configured maximum batch differs, the warmup exercises a different code path than production traffic.

Operational surface: warmup, shutdown and Prometheus metrics

The README lists model warmup, graceful shutdown and Prometheus monitoring metrics as the cloud-facing features. The Rust dependencies confirm the metrics part: prometheus-client is a direct dependency, and the repository has an examples/monitor directory.

Metrics alone do not tell you what to alert on. The interesting numbers for a batched server are the batch size distribution and the queue wait time, because those are what dynamic batching trades against latency. Mosec exposes Prometheus-format metrics, but the documentation does not enumerate every metric name in the pages available here, so plan to scrape the endpoint once and read what is actually there before writing alert rules.

Graceful shutdown matters more than it sounds. With multiple worker processes holding GPU memory, a hard kill can leave the device in a state that the next container start has to recover from. The README states graceful shutdown is supported, and the Rust side includes tokio's signal feature, which is the mechanism for it.

Deployment is left to you. The README says the service is "easily managed by Kubernetes or any container orchestration systems," and the repository ships a Dockerfile built from an NVIDIA CUDA base image with Miniconda installed. That Dockerfile is a starting point for a GPU image, not a general-purpose one: the default base argument is a CUDA runtime image, so a CPU-only deployment should override it.

Where Mosec is the wrong tool

Mosec has no model registry and no versioning. If you need to promote a model from staging to production by tag, roll back to a previous artifact, or run A/B splits across two model versions, the framework does not provide that. You would build it around the service, or use a different layer entirely.

The same applies to autoscaling and routing. Mosec serves; Kubernetes or your orchestrator scales. There is no built-in queue depth-based scaling signal documented in the README, only Prometheus metrics you could build one from.

Multi-stage pipelines in Mosec are linear. Workers are chained, the first takes the request and the last returns the response. There is no branching, no conditional routing between stages, and no fan-out to parallel workers with a join. If your inference graph looks like that, a workflow engine is a better fit and Mosec would be a poor substitute.

One more boundary worth stating plainly: the Python interface is the only interface. Cargo.toml publishes the crate with documentation at docs.rs, but the README positions the Rust layer as internal machinery, and there is no documented path for writing a worker in Rust. Teams that want a single-language Rust service should look elsewhere.

Finally, the model code runs in a separate process from the HTTP layer, which is good for isolation but means the request and response payloads cross a process boundary. For very small, very high-frequency requests, that serialization and IPC cost is a fixed tax on every call, and a simpler in-process server may beat it.

How it compares to a Python-native serving framework

The natural comparison is BentoML, which also wraps Python inference code in a service. The difference is where the work happens. BentoML builds a Python service object and packages it into a bento artifact with its own dependency and model management, then runs it through a runner architecture. Mosec keeps the surface to a Worker subclass and puts the HTTP server, queueing and batching in Rust instead of Python.

That difference shows up in what you get for free. BentoML gives you model artifact management and a packaging story; Mosec does not, and its README says so by scoping itself to the online serving part. Mosec gives you a Rust HTTP layer and process-level pipelining as the default rather than as an option.

A second comparison is TorchServe, which is tied to PyTorch and uses a handler class with a .pt model archive. Mosec's Worker does not reference a tensor framework, so a JAX or scikit-learn model fits the same interface. The trade-off runs the other way for TorchServe users who want the model archive format and the PyTorch-specific tooling.

For a single model behind one endpoint, a plain FastAPI app with a manual batch loop is less machinery. Mosec's advantage appears when you need the batching, the multi-process pipeline and the metrics without writing them, and when the Rust layer's async I/O is doing work your Python server would have to do with threads.

Editorial conclusion

Adopt Mosec if you have a Python forward function and want batching plus multi-stage pipelining without building the HTTP and coordination layers yourself; the Rust side handles that and the Python side stays testable offline. Do not adopt it if you need a model registry, autoscaling or a workflow DAG, because the project states it focuses on the online serving part only. Verify two things first: that the warmup example you set matches the batch shape your forward expects, and that your serialization choice (JSON by default, msgpack via MsgpackMixin) can carry your payload without base64 inflation.

Frequently asked questions

How do I install Mosec?

The README gives three options: pip install -U mosec, conda install conda-forge::mosec, or pixi add mosec, on Linux or macOS. Building from source requires Rust and the make package target, which produces a wheel in the dist folder.

Does Mosec do dynamic batching automatically?

It aggregates requests from different users into a batch and distributes results back, but whether your forward receives a single item or a list depends on whether dynamic batching is configured. The README states the signature as forward(self, data: Any | List[Any]) -> Any | List[Any] for exactly that reason.

What serialization does Mosec use for requests and responses?

JSON is the default. Inheriting mosec.mixin.MsgpackMixin switches the protocol to msgpack, which the README recommends when returning binary data, since JSON cannot carry it without base64 encoding and a larger payload.

What happens if I do not set a warmup example in Mosec?

The README says the service is only ready after the example has been forwarded through the handler when one is set. Without an example, the first request's latency is expected to be longer, typically because GPU memory has not been allocated yet.

Can Mosec run multiple stages of a pipeline?

Yes. Workers can be chained so the first takes the request and the last sends the response, with data moving between them over inter-process communication. The README notes this lets you mix CPU, GPU and IO bound stages in one service.

Official sources

  1. License: Apache-2.0
  2. mosecorg/mosec on GitHub
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/mosecorg-mosec.svg)](https://hysenlabs.com/projects/mosecorg-mosec)