# NVIDIA NeMo Switchyard: an OpenAI and Anthropic compatible LLM routing proxy

> Switchyard is a Rust proxy and library that keeps clients speaking OpenAI or Anthropic formats while routing requests to vLLM, NVIDIA NIM, Ollama or any OpenAI-compatible endpoint. It is pre-alpha, and its own README labels the server a demo.

**NVIDIA-NeMo/Switchyard** — Switchyard lets LLM applications route traffic across models and providers while preserving native OpenAI and Anthropic API compatibility - enabling flexible model selection, benchmarking, and cost/performance optimization.

- Repository: https://github.com/NVIDIA-NeMo/Switchyard
- Stars: 3,253 · Forks: 317
- Language: Rust
- License: Apache-2.0
- Published: 2026-08-17 · Updated: 2026-08-18 · Language: en
- Canonical page: https://hysenlabs.com/projects/nvidia-nemo-switchyard

## The problem: agents that only speak one API

A coding agent such as Claude Code or Codex is written against one wire format. If you want that agent served by an open-source model on vLLM, NVIDIA NIM or Ollama, either the agent needs a patch or something in front of it has to translate. Switchyard is that something. It accepts OpenAI Chat Completions, OpenAI Responses and Anthropic Messages on one side, picks a configured backend, forwards the request in that backend's own format, and translates the response back into the shape the client expects.

The audience is narrow and specific. First, teams running self-hosted inference who want existing agent tooling to work unchanged. Second, teams doing model comparison: the random route gives a fixed traffic split for A/B tests and cost experiments, so the same client request can land on different models without client-side changes. Third, Rust teams that want routing logic as a library rather than a process. The README is explicit that this is a young project showcasing active research, and the Python package metadata classifies it as Development Status 3 - Alpha.

## How routing and translation actually fit together

The architecture is a single hop. Clients talk to Switchyard in OpenAI or Anthropic format; Switchyard talks to backends in provider-native format. Each configured LLM client selects one upstream format, so the translation layer knows what it is converting from and to. The server accepts OpenAI Chat Completions, OpenAI Responses and Anthropic Messages.

Routing is a separate concern from translation. A route has a type, and the README lists six: llm_classifier, stage_router, escalation (which is llm_classifier with mode = "escalation"), composite, random, and passthrough. The differences matter. The LLM classifier spends an extra model call to decide whether a turn needs the weak or strong tier. The stage router avoids that call by reading signals already in the conversation, such as tool results and errors. The escalation router runs every turn on the weak tier first, then a judge reads that answer to decide whether to resend the same request to the strong tier. Composite lets one algorithm set the configuration of another before handing off; today the documented case is an LLM classifier setting the tier a stage router falls open to. Random ignores content entirely and gives you a fixed split. Passthrough registers one target under one model ID with no routing decision at all.

The library path is the more interesting design choice. switchyard-libsy embeds the algorithms and never calls a model itself. An algorithm decides which target to use and hands every model call back to the caller, so it drops into an existing proxy, gateway or agent runtime without owning an HTTP stack. Pair it with switchyard-llm-client when you want the calls made for you. That split is deliberate and it is the reason libsy is the most mature component in the tree.

## Install and first request with switchyard-server

The README gives a server path and a library path. For the server, install Rust with Cargo, then install the published binary. Cargo builds the release binary and installs it into ~/.cargo/bin by default.

```bash
cargo install --locked switchyard-server
switchyard-server --help
```

You then need a routes.toml. The README points at docs/getting_started.md for a complete configuration rather than inlining one, so treat that file as the source of truth for the route shape. Once you have it, validate before starting: the --dry-run flag checks the config without serving traffic. The example below also sets the OpenRouter key the README uses in its own snippet.

```bash
export OPENROUTER_API_KEY="your-openrouter-key"
switchyard-server --config routes.toml --dry-run
switchyard-server --config routes.toml --host 127.0.0.1 --port 4000
```

With the server running, the README's verification step is a health check on port 4000. A 200 response means the process is up; it says nothing about whether your routes are correct.

```bash
curl http://localhost:4000/health
```

For the library path, add the crates as git dependencies pinned to the v0.2.0 tag. Note that the README pins by tag, not by a crates.io version.

```toml
[dependencies]
switchyard-libsy = { git = "https://github.com/NVIDIA-NeMo/Switchyard.git", tag = "v0.2.0" }
switchyard-protocol = { git = "https://github.com/NVIDIA-NeMo/Switchyard.git", tag = "v0.2.0" }
```

There is also a Dockerfile in the repository root that builds switchyard-server on rust:1.96.1-bookworm and runs it on debian:bookworm-slim as user 1000:1000, exposing port 4000 with ENTRYPOINT ["switchyard-server"]. The image sets HOME=/tmp, which matters because the runtime user has no writable home directory otherwise.

## Where Switchyard is the wrong choice

The maturity warning is the first thing to read, not the last. The README states that switchyard-server is a demo server, not for production use. If you are looking for a gateway to put in front of paying traffic, that sentence disqualifies the server path today, regardless of how good the routing algorithms are. The other components carry their own labels: libsy is Beta and ready for trial integration, switchyard-llm-client and switchyard-runner are Alpha, and the API and algorithms are expected to change significantly before v1.0.

The second limitation is scope. Switchyard routes and translates. It is not a caching layer, not a rate limiter, and not a cost accounting system; the README's feature list is protocol translation, multi-backend routing and Prometheus metrics covering requests, errors, latency, tokens and routing overhead. If you need quota enforcement per key, that is not in the described feature set.

The third is that two of the more interesting strategies cost you something. The LLM classifier adds a model call per routing decision, which is exactly the overhead a latency-sensitive path may not want; the stage router exists to avoid that, but it only works when the conversation already carries usable signals such as tool results and errors. In a single-turn request with no prior tool output, there is nothing for a stage router to read.

Finally, the library path requires Rust. The Python package nemo-switchyard exists and requires Python 3.10 or later, but the README's own getting-started instructions are Cargo and TOML, not pip. A Python team expecting a drop-in Python gateway should check what the Python binding actually exposes before assuming parity.

## How it differs from LiteLLM

The repository ships examples/litellm/, so the comparison is one the project itself invites. The difference in approach is where the routing decision lives. LiteLLM is a Python proxy: you configure it, run it as a service, and routing policy is expressed in its configuration. Switchyard splits into a proxy binary and a Rust library, and the library deliberately does not make model calls. libsy picks a target and hands the call back to you, which means you can embed the routing decision inside an existing gateway or agent runtime that already owns its HTTP stack and its credentials. The README states this explicitly: it drops into an existing proxy, gateway or agent runtime without owning an HTTP stack.

That is a real architectural difference, not a packaging one. It also means the two are not mutually exclusive: a Python service could call into the Switchyard Python package for routing while keeping its own transport. The trade-off is that the embedded path gives you less out of the box. You get the algorithm and the protocol types, and you write the plumbing. The server path gives you the plumbing, but the README calls that server a demo.

## Licence, versioning and what upgrades cost

Switchyard is Apache-2.0, copyright NVIDIA Corporation, with a NOTICE file in the repository root. Apache-2.0 permits commercial use and modification and includes a patent grant; it also requires that you preserve the licence and NOTICE files and state significant changes. That is the general shape of the licence, not advice about your situation.

The upgrade cost is the part that should shape a decision. The workspace version is 0.2.0, and the release history is short: v0.0.1 and v0.1.0 both landed on 2026-06-30, and v0.2.0 on 2026-08-10. The README warns that the API and algorithms are expected to change significantly before v1.0. The library dependency in the README is pinned to a git tag rather than a published crate version, which means an upgrade is a deliberate edit of the tag in Cargo.toml, not a cargo update. Cargo.toml also sets rust-version = 1.96.1 and edition 2024, and the Dockerfile comments that RUST_VERSION must be kept in sync with rust-toolchain.toml. If your toolchain is older than 1.96.1, you cannot build the workspace at all.

For the Python side, dev tooling lives in a PEP 735 dependency group rather than an optional extra, so pytest, ruff and mypy are not installed by pip install nemo-switchyard and do not appear in the published wheel's metadata. Install them with uv sync or uv sync --group dev from a checkout.

## Conclusion

Adopt Switchyard if you are a Rust team that wants to point a coding agent at a self-hosted model, or you need a fixed traffic split for A/B benchmarking, and you can accept pre-alpha churn. Do not adopt it if you need a production gateway today: the README calls switchyard-server a demo, not for production use. Before committing, read docs/getting_started.md, run switchyard-server --config routes.toml --dry-run against your own routes.toml, and check whether the route types you need exist in the v0.2.0 tag rather than on main.

## FAQ

### What is NVIDIA NeMo Switchyard?

It is a Rust proxy and library for LLM traffic that routes requests across providers and translates between the OpenAI Chat, Anthropic Messages and OpenAI Responses formats. The README describes it as pre-alpha software that is evolving rapidly.

### How do I install Switchyard?

The README's server path is cargo install --locked switchyard-server, which builds the release binary and installs it into ~/.cargo/bin by default. The library path instead adds switchyard-libsy and switchyard-protocol as git dependencies pinned to the v0.2.0 tag.

### Is Switchyard ready for production?

The README labels switchyard-server a demo server, not for production use. libsy is listed as Beta and ready for trial integration, while switchyard-llm-client and switchyard-runner are Alpha.

## Sources

- [Official README](https://github.com/NVIDIA-NeMo/Switchyard#readme)
- [Project repository](https://github.com/NVIDIA-NeMo/Switchyard)
- [Release notes](https://github.com/NVIDIA-NeMo/Switchyard/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/nvidia-nemo-switchyard
