Model or dataset
predibase/lorax avatar
predibase/lorax

LoRAX: serving thousands of LoRA adapters from one GPU

Multi-LoRA inference server that scales to 1000s of fine-tuned LLMs

3,834 stars326 forksPythonApache-2.0

At a glance

What is it?
LoRAX is Predibase's multi-LoRA inference server, built on a Rust router and a Python server with custom CUDA kernels. It is a strong fit when you host many fine-tuned adapters over one shared base model, and the wrong tool when you need one huge unquantized model on a single card.
Who is it for?
Adopt LoRAX if you already host a supported base model and your problem is many PEFT adapters sharing one GPU, since the adapter_id parameter is the whole integration surface. Do not adopt it if you need a base architecture outside the supported list, a non-NVIDIA or pre-Ampere GPU, or a serving stack you can debug without CUDA.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 124 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

DEEP OPEN-SOURCE ANALYSIS

The problem LoRAX targets: many adapters, one shared base model

The README states the goal directly: serve thousands of fine-tuned models on a single GPU without compromising on throughput or latency. The unit of deployment is not the model. It is the pair of a base model, shared across every adapter, and an adapter, the task-specific weights loaded per request. A team that fine-tunes one adapter per customer, per language, or per task normally ends up with one process per adapter, and the base weights are duplicated in every process. LoRAX collapses that into one server holding the base model once and swapping adapters in and out. The audience is therefore narrow and specific: platform teams running a shared inference endpoint for many small fine-tunes, not someone serving a single model behind an API. The supported adapters are LoRA adapters trained with PEFT or Ludwig, and the README says any of the model's linear layers can be adapted. That last sentence matters more than it looks. It means the adapter has to be built against the same base architecture the server loaded, and a mismatch is a load-time failure, not a silent quality drop.

Heterogeneous batching and adapter exchange scheduling

The mechanism has two named parts. Heterogeneous continuous batching packs requests for different adapters into the same batch, which is what keeps latency and throughput roughly constant as the number of concurrent adapters grows. Adapter exchange scheduling prefetches and offloads adapter weights between GPU and CPU memory asynchronously, so a request for a cold adapter does not block the requests already in flight. Together these explain why the server can accept an adapter_id at request time: the router receives the HTTP call, the Python server decides whether those weights are resident, and scheduling decides whether to fetch them now or batch around them. The inference path itself uses tensor parallelism, paged attention, flash-attention and SGMV kernels, with fp16 or quantized base weights via bitsandbytes, GPT-Q or AWQ. One design consequence is worth stating plainly: adapter weights live in CPU memory when they are not resident, so the realistic ceiling is set by host RAM and PCIe bandwidth as much as by VRAM. The README does not publish a figure for how many adapters fit, and it should not be read as a promise that a thousand adapters are simultaneously resident.

Installing LoRAX with Docker and sending a first request

The README recommends the prebuilt image to avoid compiling custom CUDA kernels. The stated requirements are an Nvidia GPU of Ampere generation or above, CUDA 11.8 compatible drivers, Linux, and Docker. Install nvidia-container-toolkit first, then reload and restart the Docker daemon, or the container will not see the GPU.

bash
sudo systemctl daemon-reload
sudo systemctl restart docker

The launch command pulls the image, mounts a local data directory, and passes the base model id. Port 80 inside the container is published on 8080 on the host, so every later request goes to 127.0.0.1:8080. Expect a long first start: the base weights have to download before the server answers.

shell
model=mistralai/Mistral-7B-Instruct-v0.1
volume=$PWD/data

docker run --gpus all --shm-size 1g -p 8080:80 -v $volume:/data \
    ghcr.io/predibase/lorax:main --model-id $model

A first call against the base model uses the /generate endpoint. The only parameter here is max_new_tokens.

shell
curl 127.0.0.1:8080/generate \
    -X POST \
    -d '{
        "inputs": "[INST] Natalia sold clips to 48 of her friends in April, and then she sold half as many clips in May. How many clips did Natalia sell altogether in April and May? [/INST]",
        "parameters": {
            "max_new_tokens": 64
        }
    }' \
    -H 'Content-Type: application/json'

Adding an adapter is the same request with one extra key, adapter_id. The README's example points at a public Hub adapter, and the first call that names it pays the load cost.

json
{
  "inputs": "[INST] Natalia sold clips to 48 of her friends in April, and then she sold half as many clips in May. How many clips did Natalia sell altogether in April and May? [/INST]",
  "parameters": {
    "max_new_tokens": 64,
    "adapter_id": "vineetsharma/qlora-adapter-Mistral-7B-Instruct-v0.1-gsm8k"
  }
}

The Python client is a separate package. It wraps the same endpoint, so the adapter is again just a keyword argument.

bash
pip install lorax-client
python
from lorax import Client

client = Client("http://127.0.0.1:8080")

prompt = "[INST] Natalia sold clips to 48 of her friends in April, and then she sold half as many clips in May. How many clips did Natalia sell altogether in April and May? [/INST]"
print(client.generate(prompt, max_new_tokens=64).generated_text)

adapter_id = "vineetsharma/qlora-adapter-Mistral-7B-Instruct-v0.1-gsm8k"
print(client.generate(prompt, max_new_tokens=64, adapter_id=adapter_id).generated_text)

The repository also carries a Makefile target that runs the server directly on the host, which is useful when you want to change server code without rebuilding the image.

bash
make run-mistral-7b-instruct

Where LoRAX is the wrong tool

The hardware floor is the first constraint, and it is not negotiable: an Nvidia GPU of Ampere generation or above, CUDA 11.8 compatible drivers, and Linux. A team on Apple silicon, on an older card, or on Windows cannot run the documented path at all. The second constraint is the base model. The README lists Llama, CodeLlama, Mistral, Zephyr and Qwen, and points to a supported-architectures page for the complete list. If your fine-tune sits on an architecture outside that list, LoRAX has nothing to load. The third is the adapter format. PEFT and Ludwig LoRA adapters are supported; a fully fine-tuned checkpoint is not, because there is no base model to share. The fourth is operational. The repository is a Rust router plus a Python server plus custom CUDA kernels, and the Dockerfile builds both toolchains. When something goes wrong below the HTTP layer, the debugging surface is larger than a pure-Python server. Maintenance is also worth checking before you commit: the last push to the default branch was on 2026-05-28, and the most recent release in the list is lorax-0.4.0 from 2025-01-13, so the release cadence is not what it was in late 2024.

LoRAX against vLLM's multi-LoRA mode

vLLM is the obvious comparison because it also serves LoRA adapters alongside a base model, and the difference is one of emphasis rather than capability. vLLM's multi-LoRA support is a mode of a general-purpose serving engine whose primary job is high-throughput serving of whole models, quantized or not, with a large supported-architecture list. LoRAX inverts that: the adapter is the first-class object, the request carries adapter_id, and the scheduling machinery exists to keep adapter swapping from wrecking the batch. The README even describes merging adapters per request to build ensembles, which is an adapter-centric operation rather than a model-serving one. The practical split follows from that. If your workload is many small fine-tunes over one or two base models, LoRAX's scheduler is built for exactly that shape. If your workload is a handful of large models, or you need an architecture the LoRAX supported list does not cover, a general serving engine is the less constrained choice. Neither is a drop-in for the other, and the API surfaces differ enough that a migration is real work.

Licence, packaging and what an upgrade actually costs

LoRAX is Apache-2.0, which the README calls out as free for commercial use. That is a permissive licence with a patent grant, and it imposes no copyleft obligation on your own code, but it says nothing about the licences of the base models or adapters you load. Those come from their own publishers, and a Mistral or Llama base model carries terms that Apache-2.0 does not override. Check the model card, not the server licence, before shipping. On upgrades, the repository ships prebuilt images on ghcr.io and Helm charts under charts/, so the normal path is to pin an image tag rather than rebuild. Note that the README's example uses the main tag, which moves; pinning to a release tag is the safer default for anything you intend to keep running. The Makefile shows the from-source path, and it is heavier than most Python projects: install-server, install-router, install-launcher, and install-custom-kernels, the last of which requires BUILD_EXTENSIONS=True and warns that kernels might not work on all hardware. Upgrading a source build therefore means a Rust toolchain, a CUDA toolchain, and a rebuild of the kernels, not a pip upgrade.

Editorial conclusion

Adopt LoRAX if you already host a supported base model and your problem is many PEFT adapters sharing one GPU, since the adapter_id parameter is the whole integration surface. Do not adopt it if you need a base architecture outside the supported list, a non-NVIDIA or pre-Ampere GPU, or a serving stack you can debug without CUDA. Before committing, verify that your adapter's target linear layers match the base model, measure adapter load latency under your own concurrency, and read the reference for the exact server flags your version exposes.

Frequently asked questions

What does LoRAX mean?

The README expands the name as LoRA eXchange. It is a multi-LoRA inference server that serves many fine-tuned adapters over one shared base model.

What is LoRAX?

It is a framework for serving thousands of fine-tuned models on a single GPU. Each request names a base model plus an optional LoRA adapter, and the server loads that adapter just in time.

How do I install and run LoRAX?

The README recommends the prebuilt Docker image. After installing nvidia-container-toolkit and restarting Docker, run the ghcr.io/predibase/lorax:main image with --gpus all and a --model-id, publishing container port 80 to a host port.

Does LoRAX require an Nvidia GPU?

Yes. The stated minimum requirements are an Nvidia GPU of Ampere generation or above, CUDA 11.8 compatible drivers, a Linux OS, and Docker for the documented guide.

Which base models and adapters does LoRAX support?

The README lists Llama, CodeLlama, Mistral, Zephyr and Qwen as base models, with a supported-architectures page for the full list. Adapters must be LoRA adapters trained with PEFT or Ludwig.

Official sources

  1. License: Apache-2.0
  2. predibase/lorax on GitHub
  3. Project website
  4. README
  5. Releases
For maintainers

Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/predibase-lorax.svg)](https://hysenlabs.com/projects/predibase-lorax)
Community notes

Community notes