Model or dataset
intentee/paddler avatar
intentee/paddler

Paddler review: an LLM load balancer built around llama.cpp and two binaries

Open-source LLM/VLM load balancer and serving platform for self-hosting LLMs (and VLMs) at scale 🏓🦙 Alternative to projects like llm-d, Docker Model Runner, etc but with less moving parts and simple deployments built around ggml ecosystem. Runs on CPU and GPU.

1,676 stars99 forksRustApache-2.0

At a glance

What is it?
Paddler is an Apache-2.0 load balancer and serving platform for self-hosting LLMs and VLMs, shipped as a single binary with a balancer and agents. It trades the multi-component Kubernetes stack for a ggml-based design, and the trade-offs are visible in the repository layout.
Who is it for?
Adopt Paddler if you want LLM inference on your own hardware with a small number of moving parts and you are comfortable with the ggml/llama.cpp model format. Do not adopt it if you need a documented rollback path for model swaps or a deployment story that is not a Rust build.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly Rust, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem Paddler targets: per-token pricing and closed model providers

The README opens with a stated motivation rather than a feature list: digital products need privacy, reliability, cost control, and independence from closed-source model providers. Paddler is the project's answer to that, an open-source LLM load balancer and serving platform that runs inference on infrastructure you control.

The audience is named explicitly. Product teams that need LLM inference and embeddings inside their features. DevOps and LLMOps teams running models at scale. Organizations with compliance and privacy requirements in medical or financial settings. Teams that want predictable costs instead of exposure to per-token pricing. That list matters because it tells you what the project optimizes for: operational simplicity on hardware you already own, not squeezing the last percent of throughput out of a rented GPU.

There is a second audience the README is honest about. Paddler ships a desktop application in beta for casual use cases, such as pooling multiple laptops and PCs into a local cluster or building an office-wide second brain without touching a console. The README suggests mixing the two modes: a balancer on a server rack, plus a colleague's RTX 5090 joining ad hoc as an agent when it is idle. That is an unusual deployment story, and it is only plausible because agents can be added dynamically.

Balancer, agents, slots: the three layers of a Paddler cluster

Paddler has exactly two deployable components. The balancer distributes incoming requests. Agents generate tokens and embeddings through slots. Everything else in the repository is a library or a test harness.

The balancer exposes three services. An inference service that applications connect to for tokens or embeddings. A management service used internally to coordinate the fleet. And an optional web admin panel for viewing and testing the setup. Agents are usually deployed on separate instances and further distribute requests to slots, which are the units that actually generate tokens and embeddings.

This is where the project diverges from wrapping llama.cpp's server. Paddler uses a built-in llama.cpp engine for inference, but implements its own version of llama.cpp slots, and the README states that these slots keep their own context and KV cache. That design choice is what makes LLM-specific load balancing possible: the balancer can route against slot state rather than treating each backend as an opaque HTTP endpoint. It also means the project owns code that a wrapper would inherit from upstream.

The Cargo workspace confirms the shape of the system. Members include paddler_balancer, paddler_agent, paddler_cli, paddler_client, paddler_messaging, paddler_download_manager, paddler_gui, and paddler_state_conversion, alongside test crates such as paddler_test_cluster_harness and paddler_opencode_tests. The presence of a dedicated messaging crate and a state conversion crate suggests the balancer and agents exchange typed state rather than only proxying bytes. There is also a JavaScript client workspace, paddler_client_javascript, and a Python client directory, so the intended integration surface is not limited to raw HTTP.

Installing Paddler and running a first balancer with one agent

Paddler is self-contained in a single binary. The README gives two ways to obtain it: download the latest release from GitHub releases, or build from source. The stated MSRV is 1.88.0, and the toolchain pin lives in rust-toolchain.toml. If you build from source, the Makefile is the entry point; it builds the paddler_cli package and, for GPU variants, passes feature flags such as cuda or metal together with web_admin_panel.

Once the binary is on your PATH, the whole surface is the paddler command. Running paddler --help lists the available commands, which is the fastest way to confirm the build is intact.

Start the balancer first. The README's example binds inference, management, and the web admin panel to loopback:

bash
paddler balancer --inference-addr 127.0.0.1:8061 --management-addr 127.0.0.1:8060 --web-admin-panel-addr 127.0.0.1:8062

The inference address is what your application will call. The management address is what agents dial. The web admin panel flag is optional, but without it you lose the dashboard and the GUI test page.

With the balancer running, start an agent and point it at the management address. The README uses four slots for the example:

bash
paddler agent --management-addr 127.0.0.1:8060 --slots 4

Four slots means four concurrent generation units on that host, each with its own context and KV cache. That number is a memory decision as much as a concurrency decision, and the README does not offer a formula for picking it. After the agent registers, open the web admin panel at the address you passed to --web-admin-panel-addr. The dashboard shows the fleet; the model section is where you add or update a model and customize the chat template and inference parameters; the prompt section is a GUI for testing inference. The README links a longer walkthrough for setting up a basic LLM cluster, and a Dockerfile exists in the repository if you prefer to build the binary inside a container rather than on the host.

Request buffering, dynamic model swapping, and scaling from zero

Two features in the README deserve separate treatment because they determine whether Paddler fits your workload.

Request buffering, described as enabling scaling from zero hosts, means the balancer holds incoming requests while no agent is available. Combined with dynamically added agents, this is what makes the project compatible with autoscaling tools: you can let compute drop to nothing and accept that the first request after a scale-up waits. The README does not state a buffer limit, a timeout, or what happens to buffered requests when the buffer is exhausted. For interactive chat that is a detail you would want to measure before committing.

Dynamic model swapping is the other one. Because slots own their context and KV cache, changing the model is a state transition the system manages rather than a process restart. The README does not document rollback for a swap, nor does it describe what happens to in-flight requests during one. Treat swapping as a capability you should exercise in a staging cluster first, with the web admin panel open, so you can see the fleet state change rather than infer it.

The repository also contains paddler_download_manager and a dependency on hf-hub, which points at Hugging Face as a model source. The README does not spell out the download workflow, so the model section of the web admin panel is the place to look. Whatever you download, the model's own license is a separate question from Paddler's.

Where Paddler is the wrong tool

The clearest limitation is the engine. Paddler's inference path is built on llama.cpp and the ggml ecosystem, with llama-cpp-bindings pinned in the workspace. If your models are not in a ggml-compatible format, or if your organization has standardized on a different serving runtime, Paddler is not a drop-in layer you can put in front of it. The load balancing logic is tied to Paddler's own slot implementation, so you cannot point the balancer at arbitrary OpenAI-compatible servers and get the same behavior.

The second limitation is deployment machinery. Paddler is a Rust workspace with a Node-based frontend build, a Makefile that drives both, and GPU feature flags chosen at compile time. The Dockerfile reflects that: it starts from node:latest, installs a Rust toolchain and a C++ build chain, clones the repository, and runs make. That is a build-time story, not a pull-a-published-image story. Teams that expect a pinned container image from a registry will need to produce one themselves.

Third, the desktop application is labeled beta in the README. The ad hoc laptop-as-agent scenario is appealing, but it is the least proven part of the surface. And for anything requiring strict multi-tenant isolation or per-request accounting across untrusted users, the README describes no such mechanism; it describes a fleet you operate.

How Paddler differs from llm-d and Docker Model Runner

The README positions Paddler as an alternative to projects like llm-d and Docker Model Runner, with the stated difference being fewer moving parts and simple deployments built around the ggml ecosystem.

That framing is accurate as far as it goes. The practical difference is where the complexity lives. A Kubernetes-native serving stack spreads configuration across custom resources, controllers, and a scheduler, and gains elasticity and scheduling policy in return. Paddler collapses that into a balancer process and agent processes, with the web admin panel as the control surface. You give up the scheduling vocabulary and the ecosystem of operators; you gain a deployment you can describe in two commands.

The comparison to Docker Model Runner is a different axis. Paddler's agents are designed to be added dynamically and can run on separate instances, including machines that are not part of a cluster. That is the multi-device story the README leans on, and it is not what a local model runner is built for. If your goal is a single developer running a model on one laptop, neither Paddler's balancer nor its agent layer earns its keep; a plain llama.cpp invocation is closer to the job. Paddler starts to matter when there is more than one host, or when you need the fleet to shrink to zero and grow again.

Maintenance, licensing, and what the repository tells you about cost

Paddler is not archived. The last push to the default branch was on 2026-07-19, the same day as the v4.1.0 release, which the release notes label OpenCode support. The preceding releases were v4.1.0-rc1 on 2026-07-11 and v4.0.1 on 2026-07-07. That is a release cadence measured in weeks across that window, and the presence of a release candidate before the final tag suggests a staged process. The repository carries CODEOWNERS, SECURITY.md, clippy.toml, rustfmt.toml, and a coverage check as a dev dependency, which indicates the project invests in its own tooling. None of that is a promise about the next release.

Licensing is Apache-2.0 for the project, declared in both Cargo.toml and package.json. Apache-2.0 includes an express patent grant and requires preservation of notices, which is generally friendlier for commercial embedding than a copyleft license. This is not legal advice, and it covers Paddler's code only. Models you download have their own terms, and the README does not discuss model licensing at all, so that review is on you.

The upgrade cost is dominated by the build. Because GPU support is selected through Cargo features and the Makefile maintains separate target directories per device, moving between CPU, CUDA, and Metal builds means separate artifacts rather than one binary that adapts. The workspace pins several dependencies to exact versions, including llama-cpp-bindings at 0.12.0 and actix-web at 4.13.0, so upgrades will occasionally require coordinated bumps. Budget for a build pipeline rather than a version bump.

Editorial conclusion

Adopt Paddler if you want LLM inference on your own hardware with a small number of moving parts and you are comfortable with the ggml/llama.cpp model format. Do not adopt it if you need a documented rollback path for model swaps or a deployment story that is not a Rust build. Verify first that your target hardware is covered by the build features in the Makefile, and confirm the license terms of every model you download through the built-in download manager.

Frequently asked questions

What is Paddler?

Paddler is an open-source LLM load balancer and serving platform for running inference, deploying, and scaling LLMs on your own infrastructure. It ships as a single binary with two deployable components, a balancer and agents, and uses a built-in llama.cpp engine for inference.

What does Paddler mean in this project?

In this project, Paddler is the name of the software itself, an LLM and VLM load balancer from Intentee. The README does not explain the origin of the name.

How do I install Paddler?

Download the latest release binary from GitHub releases, or build from source with an MSRV of 1.88.0. The README states the project is self-contained in a single binary, so installation means making the paddler binary available on your system.

Does Paddler run on CPU only, or does it need a GPU?

The repository description states it runs on CPU and GPU. The Makefile builds device-specific variants using Cargo features such as cuda and metal, with a TEST_DEVICE variable that defaults to cpu.

What license does Paddler use?

Apache-2.0, declared in both Cargo.toml and package.json. That covers the project's code; models you download through the built-in download manager carry their own licenses, which the README does not address.

Official sources

  1. intentee/paddler on GitHub
  2. License: Apache-2.0
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/intentee-paddler.svg)](https://hysenlabs.com/projects/intentee-paddler)