lollms_hub: a proxy and admin console for multiple Ollama instances
A proxy server for multiple ollama instances with Key security
At a glance
- What is it?
- lollms_hub puts authentication, rate limiting and model orchestration in front of one or many Ollama, vLLM or OpenAI-compatible backends. It is worth a look if you already run more than one inference server, and overkill if you run one.
- Who is it for?
- Adopt lollms_hub if you already run two or more Ollama or vLLM backends and want one authenticated endpoint, per-user keys and a UI for model management instead of SSH sessions.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 161 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem lollms_hub addresses: many Ollama servers, no shared front door
Ollama exposes an HTTP API on a port and nothing else. Once you have more than one machine serving models, or more than one person using them, you are left assembling the missing pieces yourself: who is allowed to call which model, how requests spread across GPUs, what happens when the first backend is busy. lollms_hub is built for exactly that gap. The pyproject.toml describes it as "a secure, high-performance universal AI gateway and load balancer for Ollama, vLLM, and llama.cpp", and the README frames the target user as someone with several compute nodes rather than a single laptop.
The intended audience is narrow and specific. If you have one Ollama install and one user, this adds a Python service, a database, a session layer and an admin UI between you and a working endpoint. If you have a Windows gaming PC with two GPUs, a Linux box and a cloud API key, the project is aimed at you. The README's cluster diagram shows a master hub in front of worker machines, each running either a local slave hub or a raw Ollama service, with a cloud vLLM or OpenAI endpoint as a third kind of node.
The slave hub exists to keep the master simple. A worker running its own hub can spawn and manage local processes, so the master holds one remote server entry rather than a list of raw instances. It also changes the security boundary: only the master needs public reachability, and the README states that slaves can stay on a private VPN or VLAN. That is a real architectural decision, not a cosmetic one, because it means the master is a single point of failure for every client regardless of how many workers sit behind it.
How the master hub, slave hubs and routers fit together
The architecture is a tree. A master node holds the public-facing API, and because that API is stated to be compatible with both Ollama and OpenAI, other hubs can be chained behind it as ordinary backends. A worker machine can run a slave hub that manages its own local Ollama processes, so the master sees one remote server entry instead of several raw instances. The README gives three reasons for this indirection: local process management on the worker, a simpler master configuration, and the ability to keep slaves on a private VPN or VLAN with only the master exposed.
Routing sits above the backends. A Smart Router is a named virtual model that evaluates rules top to bottom. Fast rules cover keywords, regex patterns, message length, image detection and specific users; semantic rules hand the request to a small LLM for intent classification when pattern matching is not enough. Three strategies decide which backend actually serves the call: priority, which tries the first model and falls back to the second if it is unavailable; random, for even spread across a large cluster; and least loaded, which picks the backend with the lowest active request count. That last one is the only strategy the README ties to a workload, describing it as best for high-TPS applications.
The vision router example is the clearest illustration of the design. A request containing an image is matched by a has_images rule and sent to a vision model such as gemma3:27b; the resulting description is then passed, together with the original prompt, to a text model such as llama3.1:70b. The two-model hop is the point: the text model never sees the image, only a description of it, so image understanding is bounded by what the vision model chose to write down. Ensembles work differently. Several agents run in parallel on the same query and a master model synthesises their outputs into one answer, which trades latency and GPU capacity for breadth of reasoning. Both features are configuration objects created in the UI, not code.
Installing lollms_hub and adding a first backend
The README gives two paths. The scripted path clones the repository and runs the platform-specific launcher, which on first run presents a setup wizard for the admin credentials:
git clone https://github.com/ParisNeo/lollms_hub.git
cd lollms_hub
chmod +x run.sh
./run.shOn Windows the README says to double-click run_windows.bat instead. Stopping the hub means closing the terminal or pressing Ctrl+C.
The Dockerfile is the second path. It builds in two stages on python:3.13-slim, installs dependencies with Poetry in the builder stage, copies the app and gunicorn_conf.py into the runtime stage, and runs as a non-root app user. The exposed port is 8080 and the start command is gunicorn with gunicorn_conf.py:
docker build -t lollms-hub .
docker run -p 8080:8080 lollms-hubThe container expects ./app/static/uploads, ./.ssl and ./benchmarks to exist, and the Dockerfile creates them with ownership for the app user. After the hub is up, the README's workflow is UI-driven: log in with the credentials from the wizard, open the server management page, add your Ollama instances, then pull, update or delete models on any of them from the browser. The repository also ships a reset_admin_password.py script plus .bat and .sh wrappers, which is the recovery path if the admin password is lost. A .env.example file sits at the top level for environment-based configuration, and gunicorn_conf.py is where the server binds when you run the container.
Where lollms_hub gets in your way
The dependency list is heavier than the job sounds. pyproject.toml pulls in SQLAlchemy and alembic for persistence, redis, bcrypt and passlib for auth, scikit-learn and numpy, pandas and matplotlib, plus xhtml2pdf and pipmaster. That is a lot of surface area for a proxy, and it means the install is not a single small binary. The requirements.txt and pyproject.toml also disagree in places: requirements.txt pins cryptography==42.0.8 and numpy==1.26.4 while pyproject.toml asks for cryptography ^46.0.5 and numpy ^2.3.4. Whichever file your deployment actually uses matters, and the README does not say which one is authoritative.
The README is also uneven. It documents HTTPS configuration in detail, including the two options in Settings, HTTPS/SSL: upload key.pem and cert.pem through the UI, or give full file paths for certificates managed elsewhere, with a server restart required to apply changes. What it does not document is what happens when a backend fails mid-request, how the least-loaded counter is maintained across restarts, or how to roll back a configuration change. There is a docs/ directory and a DEVELOPMENT.md at the top level, so some of this may live there rather than in the README, but a reader deciding from the README alone is working with an incomplete picture. Two identical "Step 6" headings appear in the visual walkthrough, which suggests the README has not been proofread recently.
The wrong-tool case is straightforward. If your Ollama instance is bound to localhost and used by one person on one machine, a reverse proxy with basic auth, or Ollama's own configuration, covers the requirement with far less running software. lollms_hub earns its footprint when the number of backends, users or models crosses the point where manual coordination stops working. Note also that the hub does not make a slow backend fast. Least-loaded routing spreads work, but if every backend is saturated, the queue simply moves.
lollms_hub compared with a plain nginx reverse proxy
The obvious alternative is nginx in front of Ollama. Nginx terminates TLS, can distribute requests across upstream servers, and adds basic authentication. What it does not do is understand models. It cannot route a request to a different backend because the payload contains an image, it cannot fall back to a second model because the first is unavailable, and it has no concept of a named virtual model that resolves to different backends over time.
lollms_hub's routing rules operate on message content and user identity, which is a different layer of the stack. The trade-off is that you are running a Python application with a database, a session layer and a web UI, and you inherit its bugs and its upgrade cycle. Nginx is a smaller, better-understood component. The honest split is that nginx solves transport and lollms_hub solves dispatch. If your problem is only TLS and a shared port, nginx is the cheaper answer. If your problem is that different requests should reach different models, or that one user should not exhaust a GPU another user needs, nginx has no mechanism for it.
There is a middle option worth naming: a small FastAPI or Flask shim that holds an API key and forwards to one Ollama host. That covers authentication and nothing else, and it is perhaps fifty lines. The reason to prefer lollms_hub over such a shim is the parts you would otherwise write yourself and maintain: user records, usage analytics, per-model access, and a UI for pulling models on remote machines. The reason to prefer the shim is that you can read all of it in one sitting.
Maintenance, licensing and what an upgrade actually costs
The repository is not archived, and the last push was on 2026-04-23. That is roughly five months before the date of this article, so it is not fair to describe the project as actively developed on the strength of that alone. The release history is unusual: v9.0.0 is dated 2025-10-31, v8.0.0 is dated 2025-09-06, and v17.1.0 is dated 2025-09-04 with the note "Last stable version before full revamp". A 17.x release appearing before a 9.0.0 release in the timeline, and pyproject.toml declaring version 10.0.0, means the version numbers do not form a single monotonic line. Anyone pinning a version should read the release notes rather than assume higher means newer.
The licence is Apache-2.0, declared in both the LICENSE file and pyproject.toml. That permits commercial use and modification, and it includes an explicit patent grant, which matters if you are embedding the hub in a product. It also requires that you preserve copyright and licence notices and state significant changes. This is a description of the licence text, not legal advice; if you plan to redistribute a modified hub, read the LICENSE file and talk to someone qualified.
Upgrade cost is the real question. Alembic is present, so schema migrations exist, but the README does not describe a migration procedure or a downgrade path. The reset_admin_password.py script exists for credential recovery, and reset.sh and reset.bat exist at the top level, which suggests recovery is handled by resetting state rather than by rolling back. Before upgrading a hub that holds user accounts and usage history, check the docs/ directory and the release notes for the migration steps, because the README will not tell you. The .gitlab-ci.yml and .github/ directories both exist, so there is CI, but the README does not state what the pipeline runs.
Editorial conclusion
Adopt lollms_hub if you already run two or more Ollama or vLLM backends and want one authenticated endpoint, per-user keys and a UI for model management instead of SSH sessions. Skip it if you run a single Ollama instance on localhost, or if you need a proxy whose behaviour is fully documented: the README covers routing strategies at a high level but does not document rollback, failure handling when a backend dies mid-request, or how the master hub behaves when a slave hub is unreachable. Verify first that the routing strategy you need (priority, random or least loaded) matches the traffic shape you have, and check the docs/ directory for configuration keys the README does not list.
Frequently asked questions
What is lollms_hub and what does it do with my Ollama instances?
It is a proxy and admin layer that sits in front of one or more Ollama, vLLM or OpenAI-compatible backends, exposing a single API that the README states is compatible with both Ollama and OpenAI. It adds authentication, rate limiting, user management and model routing, and lets you pull or delete models on any registered server from a web UI.
How do I install lollms_hub?
Clone the repository and run ./run.sh on macOS or Linux, or double-click run_windows.bat on Windows; the first run presents a setup wizard for admin credentials. A Dockerfile is also provided and the container exposes port 8080, running the app through gunicorn with gunicorn_conf.py.
How do I secure Ollama with lollms_hub?
The hub handles authentication and user management itself, and the README documents HTTPS under Settings, HTTPS/SSL, where you either upload key.pem and cert.pem through the UI or supply full file paths to certificates managed elsewhere. A server restart is required for the change to take effect.
Can lollms_hub route requests to different models automatically?
Yes, through Smart Routers, which evaluate rules top to bottom. Fast rules match keywords, regex patterns, message length, image detection and specific users, and semantic rules use a small LLM to classify intent; the backend is then chosen by priority, random or least-loaded strategy.
What if I forget the lollms_hub admin password?
The repository includes reset_admin_password.py along with reset_admin_password.bat and reset_admin_password.sh wrappers. The README does not describe the script's options, so read the file before running it.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/parisneo-lollms-hub)