# OptiLLM: an OpenAI-compatible proxy that adds inference-time reasoning techniques

> OptiLLM sits between your client and any OpenAI-compatible endpoint and applies one of 20+ reasoning techniques selected by a model-name prefix. It is for teams that want better reasoning output without fine-tuning, and it costs extra tokens and latency on every call.

**algorithmicsuperintelligence/optillm** — Optimizing inference proxy for LLMs

- Repository: https://github.com/algorithmicsuperintelligence/optillm
- Stars: 4,304 · Forks: 388
- Language: Python
- License: Apache-2.0
- Published: 2026-09-09 · Updated: 2026-09-09 · Language: en
- Canonical page: https://hysenlabs.com/projects/algorithmicsuperintelligence-optillm

## The problem OptiLLM targets: reasoning quality without fine-tuning

Most teams that want better answers from a language model face two options. Fine-tune, which needs labelled data, GPU time, and a retraining loop every time the base model changes. Or prompt harder by hand, which does not scale past a handful of prompts. OptiLLM takes a third route: keep the model frozen and spend more compute at inference time. The README frames this as beating frontier models "by doing additional compute at inference time", and cites the CePO approach from Cerebras as an example of combining techniques.

The intended user is an engineer who already has working code against an OpenAI-compatible API. The proxy is a drop-in replacement: you change the base_url, and you change the model name. The README's example uses the prefix moa- on gpt-4o-mini to route the request through Mixture of Agents. Nothing else in the client changes. That is the whole pitch, and it is a good one for evaluation work, because you can A/B a technique against a baseline by editing one string.

## How the proxy picks a technique from the model name

The routing mechanism is prefix matching on the model field. A request for moa-gpt-4o-mini is interpreted as: apply the moa approach, then call the underlying model gpt-4o-mini. The README lists slugs for each technique, including mars, cepo, cot_reflection, plansearch, and re2. The default approach is auto, which the startup log line confirms: "Starting server with approach: auto".

Under the hood the proxy is a Flask application started by optillm.py, with dependencies that hint at what the techniques need. networkx appears for graph-based search, z3-solver for constraint solving, numpy and scikit-learn for scoring, litellm for talking to many providers, and presidio_analyzer plus presidio_anonymizer for the privacy plugin. The startup log also shows a memory plugin loading.

The consequence of this design is that the technique is a property of the request, not of the deployment. You can run one proxy process and send cot_reflection requests next to plansearch requests. The trade-off is that the model name is now overloaded: it carries both the routing decision and the upstream model identifier, so any client-side logic that parses model names needs to know about the prefixes.

## Installing OptiLLM with pip and running a first request

The README gives pip as the shortest path. Install the package, export a provider key, and start the server. The startup output the README shows includes lines for the privacy and memory plugins and the approach selection.

```bash
pip install optillm
export OPENAI_API_KEY="your-key-here"
optillm
```

With the server running on port 8000, point any OpenAI client at it. The example below is adapted from the README's quick start: the only difference from a normal call is the moa- prefix on the model name.

```python
from openai import OpenAI

client = OpenAI(base_url="http://localhost:8000/v1")

response = client.chat.completions.create(
    model="moa-gpt-4o-mini",
    messages=[{"role": "user", "content": "Solve: If 2x + 3 = 7, what is x?"}]
)
print(response.choices[0].message.content)
```

The README contrasts a bare answer ("x = 1") with a worked answer that shows the steps and arrives at x = 2. Expect the response to contain reasoning text rather than a single token, and expect it to take longer than a direct call, because the technique fans out to more than one model invocation.

Docker is the alternative path. The README documents three image variants: latest for the full image with local inference and plugins, latest-proxy for a lightweight proxy without local inference, and latest-offline which bundles pre-downloaded spaCy models for air-gapped use.

```bash
docker pull ghcr.io/algorithmicsuperintelligence/optillm:latest
docker run -p 8000:8000 ghcr.io/algorithmicsuperintelligence/optillm:latest
```

The compose file in the repository wires the same service with an env_file, exposes port 8000 by default through OPTILLM_PORT, and defines a healthcheck against /health. It also lists commented environment variables for per-technique parameters, including OPTILLM_APPROACH, OPTILLM_MODEL, OPTILLM_BEST_OF_N, OPTILLM_SIMULATIONS, OPTILLM_EXPLORATION, OPTILLM_DEPTH, and the RSTAR knobs OPTILLM_RSTAR_MAX_DEPTH, OPTILLM_RSTAR_NUM_ROLLOUTS and OPTILLM_RSTAR_C. Those names are the ones to check first when a technique behaves differently from what you expect.

## The cost model: more accuracy means more inference calls

Every technique here buys accuracy with compute. Best-of-N generates N candidates and picks one. MCTS explores a tree of partial solutions. CePO combines best-of-N with chain-of-thought, self-reflection and self-improvement. None of that is free. The compose file's default OPTILLM_BEST_OF_N is 3, which is a useful hint about the multiplier on a single user request.

This is the main limitation and it is not a bug. If your workload is a latency-sensitive autocomplete, a high-volume classification job, or anything with a hard per-request budget, the proxy is the wrong layer to add. The README does not document a per-request token cap, a caching layer for repeated prompts, or a fallback that returns the base model's answer when a technique times out. Those may exist in the code, but the README is silent on them, and silence is what you should assume when sizing a deployment.

There is a second constraint worth naming. The techniques are tuned for reasoning tasks: math, code, logic. The README's benchmark table is entirely reasoning benchmarks (AIME 2025, Math-L5, GPQA-Diamond, InfiniteBench, LiveCodeBench, Arena-Hard-Auto). For a summarisation or extraction task, the extra calls are likely to buy you little, and you would be paying the multiplier for nothing.

## OptiLLM compared with calling a stronger model directly

The obvious alternative is to skip the proxy and call a larger model. The README itself makes this comparison: MOA on GPT-4o-mini is described as matching GPT-4 on Arena-Hard-Auto, and the quick-start comment says the moa- prefix "gives you GPT-4o performance from GPT-4o-mini".

The difference in approach is where the compute goes. Calling a stronger model sends one request to a more expensive endpoint. OptiLLM sends multiple requests to a cheaper endpoint and combines the results. Which wins depends on the price gap between the small and large model, and on whether your provider charges per token or per request. If the small model is an order of magnitude cheaper, the fan-out can be cheaper overall. If not, the stronger model is simpler and has one fewer moving part in your stack.

A second alternative is to implement one technique yourself, for example a best-of-N loop in your application code. That is maybe twenty lines. OptiLLM's value is that it packages 20+ techniques behind one interface and one set of slugs, so you can compare them without writing each one. If you already know which technique you want and only need that one, the proxy is overhead.

## Maintenance, releases and the dependency surface

The repository is not archived. The last push was on 2026-07-18, and v0.3.22 was released the same day, following v0.3.21 and v0.3.20 on 2026-07-05. That is a recent, frequent release cadence.

The dependency list is the real maintenance cost. It includes torch, litellm, transformers, peft, bitsandbytes, spacy, presidio, gradio, outlines, selenium, webdriver-manager and a pinned z3-solver. Several of those are heavy. The requirements file also carries an explicit version cap: transformers is pinned to >=5.0.0,<5.13.0, with a comment explaining that transformers 5.13.0 tightened AutoTokenizer.register() to require a class, which breaks mlx-lm because it registers a tokenizer by string name. The comment says the cap can be lifted once mlx-lm ships a compatible release. That is a concrete example of the kind of coupling you inherit: an upstream change in one library blocks an upgrade in another.

On licensing, the project is Apache-2.0, and the Dockerfile labels the image with that licence. Apache-2.0 permits commercial use and modification and includes a patent grant. It also requires that you preserve notices and state changes. What it does not do is resolve the licences of the dependencies you pull in, several of which have their own terms. Check those separately for your distribution model; this is a factual note about the licence identifier, not legal advice.

## Where the documentation is thin

Three gaps stand out from the README alone. First, the benchmark table reports improvements but does not say how to reproduce them; the eval extra dependency group (tabulate, accelerate, huggingface_hub, httpx, tqdm, pandas) suggests there is a harness, but the README does not walk through running it.

Second, the SSL section documents --no-ssl-verify and --ssl-cert-path, and the OPTILLM_SSL_VERIFY and OPTILLM_SSL_CERT_PATH environment variables, with a clear warning that disabling verification is insecure. What it does not document is what happens to in-flight requests when a certificate check fails mid-run.

Third, and most important for operations, the README does not describe rollback. If a technique produces worse output after an upgrade, there is no documented procedure for pinning to the previous behaviour beyond installing an older package version. The release history gives you version numbers to pin to, which is something, but the README does not present it as a rollback path.

## Conclusion

Adopt OptiLLM if you already call an OpenAI-compatible endpoint and want to compare reasoning techniques without changing client code; the model-name prefix is the whole integration surface. Do not adopt it if you are latency- or cost-sensitive on every request, since techniques such as best-of-N and MCTS multiply inference calls, or if you need a documented rollback path, which the README does not provide. Before committing, verify three things in your own environment: that the technique you pick actually improves your task, that your provider's rate limits survive the extra calls, and that the Apache-2.0 licence and the dependency tree (torch, litellm, presidio, spacy, and others) fit your deployment constraints. The repository's last push was on 2026-07-18, with v0.3.22 released the same day.

## FAQ

### What is OptiLLM?

OptiLLM is an OpenAI API-compatible proxy that applies one of 20+ inference-time reasoning techniques to a request, selected by a prefix on the model name. It requires no model training or fine-tuning; you install it, point your client at it, and change the model string.

### What is the OptiLLM proxy?

It is a Flask server started by optillm.py that listens on port 8000 by default, accepts OpenAI-compatible chat completion requests, and routes each one through a technique such as moa, cepo or plansearch before calling the upstream model. The technique is encoded in the model name, for example moa-gpt-4o-mini.

### How do I install OptiLLM?

The README gives pip install optillm, then export OPENAI_API_KEY and run optillm. Docker is the alternative: docker pull ghcr.io/algorithmicsuperintelligence/optillm:latest and docker run -p 8000:8000 with that image.

### Does OptiLLM need a GPU or model training?

The README states the techniques require zero training or fine-tuning. The full Docker image includes dependencies for local inference, and the proxy-only variant is described as lightweight without local inference capabilities, so a GPU is only relevant if you use the local inference path.

### Which providers does OptiLLM support?

The README lists OpenAI, Anthropic, Google and Cerebras, and says 100+ models are reachable via LiteLLM. The compose file also notes that OPTILLM_BASE_URL can be set to any OpenAI API-compatible endpoint.

## Sources

- [algorithmicsuperintelligence/optillm on GitHub](https://github.com/algorithmicsuperintelligence/optillm)
- [Issues](https://github.com/algorithmicsuperintelligence/optillm/issues)
- [License: Apache-2.0](https://github.com/algorithmicsuperintelligence/optillm/blob/main/LICENSE)
- [README](https://github.com/algorithmicsuperintelligence/optillm/blob/main/README.md)
- [Releases](https://github.com/algorithmicsuperintelligence/optillm/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/algorithmicsuperintelligence-optillm
