Shepherd Model Gateway (SMG): a Rust router in front of vLLM, SGLang and hosted LLM APIs
Engine-agnostic LLM gateway in Rust. Full OpenAI & Anthropic API compatibility across vLLM, TRT-LLM, TokenSpeed, SGLang, OpenAI, Gemini & more. Industry-first gRPC pipeline, KV cache-aware routing, chat history, tokenization caching, Responses API, embeddings, WASM plugins, MCP, and multi-tenant auth.
At a glance
- What is it?
- SMG is an engine-agnostic gateway that puts one OpenAI- and Anthropic-compatible endpoint in front of self-hosted inference engines and cloud providers. The interesting part is KV cache-aware routing; the hard part is that you now operate a router.
- Who is it for?
- Adopt SMG if you already run several inference engines or mix self-hosted GPUs with hosted APIs and want one endpoint, one auth layer and one metrics surface in front of them. Skip it if you have a single worker behind a plain proxy, because a Rust router with a mesh mode, WASM plugins and a control plane is more moving parts than that setup needs.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 1 day ago.
- What is it written in?
- Mainly Rust, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem: many engines, many APIs, one client
If you serve models yourself, you probably run more than one engine. vLLM for one workload, SGLang for another, maybe TensorRT-LLM or TokenSpeed where latency matters, MLX on a Mac for local work. Each exposes its own HTTP surface, its own health endpoint, its own notion of a worker pool. Add a hosted provider for overflow or for a model you cannot host, and the client side of your stack starts carrying provider-specific branches.
SMG targets that situation. The README describes it as an engine-agnostic model-routing gateway that centralizes worker lifecycle management, balances traffic across self-hosted engines and cloud providers, and puts multi-tenancy, chat-history storage, MCP tooling and observability behind one unified endpoint. The audience is a platform or inference team that already has GPUs in production and wants the routing layer to be a separate concern from the engines.
It is not a serving engine. It does not load weights or run kernels. Every request it handles ends up at a worker or a provider that does.
How SMG routes: radix trees, gRPC and prefill/decode split
The mechanism that separates SMG from a generic reverse proxy is cache-aware routing. The README states that it tracks each worker's KV-cache state in radix trees so that prefixes can be reused across SGLang, vLLM, TensorRT-LLM, TokenSpeed and MLX, and that load modeling accounts for queued token work and KV pressure. In practice this means the router keeps a per-worker view of which token prefixes are already cached, and prefers the worker that can skip the most prefill. The repository layout backs this up: there is a dedicated crates/radix_tree member and a crates/kv_index member in the Cargo workspace.
Routing policy is a launch parameter. The README lists ten: cache_aware, least_load, power_of_two, consistent_hashing, prefix_hash, bucket, round_robin, random, manual and passthrough. That is a wide menu, and it is worth reading the routing docs before picking one, because cache_aware and least_load optimize different things and will produce different latency distributions on the same cluster.
The transport layer is the other half. SMG ships a native streaming gRPC pipeline to the engines, with prefill/decode disaggregation and a separate encode stage for vision workloads, plus DP-aware routing for data-parallel engines. Workers that only speak HTTP are supported too, so the gRPC path is an option rather than a requirement.
Around that core sit the pieces you would otherwise write yourself: 21 tool-call parsers and 16 reasoning parsers with automatic model detection, MCP tool discovery and execution over stdio, SSE and streamable HTTP with approval policies, a mesh mode using SWIM gossip and CRDT-replicated state for multi-node deployments, and Kubernetes pod watchers with label selectors and per-role prefill/decode/encode selectors.
Installing SMG and sending the first request
The README gives four install paths. Docker, Helm, pip and cargo. Docker is the shortest route if you just want to try it against an existing worker:
docker pull lightseekorg/smg:latestFor a Kubernetes cluster the README shows a Helm chart published as an OCI artifact:
helm install smg oci://ghcr.io/smg-project/charts/smgThe Python package is on PyPI and the Rust binary is on crates.io. The README notes that the cargo install needs protoc available, which follows from the gRPC pipeline being part of the build:
pip install smg
cargo install smgOnce installed, point the gateway at your workers. A single worker is the simplest case, and the cache-aware policy is the one the README uses for a multi-worker example:
smg launch --worker-urls http://localhost:8000
smg launch --worker-urls http://gpu1:8000 http://gpu2:8000 --policy cache_awareRequests then go to the gateway's own port, not the worker's. The README's example uses port 30000 and the OpenAI chat completions path:
curl http://localhost:30000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model": "llama3", "messages": [{"role": "user", "content": "Hello!"}]}'If your worker is healthy and registered, you get a normal OpenAI-shaped response back. If nothing answers, the first thing to check is whether the worker URL is reachable from inside the gateway process, since the gateway is now the component holding the connection.
The cost of adding a router you now have to run
SMG is a distributed system you operate. The high-availability mode uses mesh networking with SWIM gossip and CRDT-replicated state, and the launch example for it takes three flags: --enable-mesh, --mesh-advertise-host and --mesh-peer-urls. That is real configuration surface, and gossip-based membership has its own failure modes under network partitions that a single-process proxy does not have.
The build itself is not trivial either. The Cargo workspace lists more than twenty member crates, including separate crates for auth, MCP, mesh, WASM plugins, multimodal handling, RDMA, a radix tree, a KV index and a ZMQ client. The Makefile caps parallel jobs at 16 and prints a hint to install sccache if it is not found, which tells you the maintainers expect compilation to be slow enough that a compiler cache matters.
There is also a versioning shape worth noticing. The workspace dependencies pin internal crates at their own versions: openai-protocol at 1.13.0, smg-mcp at 2.3.3, smg-auth at 1.2.3, llm-tokenizer at 1.7.0. Those do not move in lockstep with the gateway release, so a bug you hit may live in a crate with a different version number than the binary you installed.
And the wrong-tool case is simple: one worker, one model, no hosted fallback. A single-process proxy or the engine's own frontend handles that with far less to maintain.
SMG compared with LiteLLM and with a plain load balancer
The closest comparison in this space is LiteLLM, which also presents one OpenAI-compatible surface in front of many providers. The difference in approach is where the intelligence sits. LiteLLM's centre of gravity is provider translation and budget or key management across hosted APIs. SMG's centre of gravity is the GPU side: KV cache state tracked per worker in radix trees, load modeling that accounts for queued token work, prefill/decode disaggregation over gRPC, and DP-aware routing. If your problem is "we call six hosted providers and want one SDK", SMG is aimed somewhere else. If your problem is "our prefixes keep landing on the wrong worker", that is the case SMG was built for.
Against a plain L7 load balancer the difference is even sharper. A load balancer distributes connections; it has no idea which worker already holds the KV cache for a prompt prefix, no notion of queued token work, and no tool-call or reasoning parser layer. It also cannot split prefill from decode. What it does have is a much smaller operational footprint and a decade of tooling around it. The trade is deliberate: SMG buys prefix reuse and engine awareness at the price of running a Rust service with its own config, mesh and metrics.
Licence, releases and what upgrading involves
SMG is Apache-2.0, stated in the README badge and present as a LICENSE file at the repository root. Apache-2.0 is permissive and includes an explicit patent grant, which matters if you are embedding the gateway in a commercial product. It also means you carry the usual obligation to preserve notices. This is a description of the licence text, not legal advice; check the terms with your own counsel if the gateway becomes part of something you ship.
The release cadence visible in the repository is steady: v1.9.0 on 2026-07-30, v1.10.0 on 2026-08-25, v1.10.1 on 2026-08-27. The last push to the default branch was on 2026-09-10, and the repository is not archived. Upgrade cost depends on which install path you chose. Docker and Helm users pull a new tag or chart version. Python users get a wheel. Rust users rebuild, and that rebuild pulls the workspace crate versions listed in Cargo.toml, which as noted move independently of the gateway version. If you run the mesh mode, the CRDT-replicated state is the part to think about during a rolling upgrade, because peers exchange membership and state rather than just forwarding requests.
Editorial conclusion
Adopt SMG if you already run several inference engines or mix self-hosted GPUs with hosted APIs and want one endpoint, one auth layer and one metrics surface in front of them. Skip it if you have a single worker behind a plain proxy, because a Rust router with a mesh mode, WASM plugins and a control plane is more moving parts than that setup needs. Before committing, check the routing documentation for the exact behaviour of the policy you intend to use, and verify that your engine version speaks the gRPC path if you plan to enable it.
Frequently asked questions
What is SMG (Shepherd Model Gateway)?
It is an engine-agnostic LLM gateway written in Rust that centralizes worker lifecycle management and balances traffic across self-hosted engines and cloud providers behind one unified endpoint. The README describes it as a model-routing gateway rather than a serving engine.
Which inference engines and providers can SMG route to?
The README lists vLLM, SGLang, TokenSpeed, TensorRT-LLM and MLX as self-hosted engines, plus any OpenAI-compatible server such as Ollama. On the hosted side it names OpenAI, Anthropic, Google Gemini, xAI, OCI Generative AI, AWS Bedrock and Azure OpenAI, along with any OpenAI-compatible provider.
How do I install SMG?
The README gives four options: a Docker image at lightseekorg/smg, a Helm chart at oci://ghcr.io/smg-project/charts/smg, the smg package on PyPI, and cargo install smg. The README notes that the Rust build needs protoc.
What port does SMG listen on by default?
The README's request example sends a curl to http://localhost:30000/v1/chat/completions, so that is the port used in the documented quick start. The README does not state whether the port is configurable or what the flag for changing it is.
What routing policies does SMG support?
The README lists ten: cache_aware, least_load, power_of_two, consistent_hashing, prefix_hash, bucket, round_robin, random, manual and passthrough. The policy is selected at launch with the --policy flag.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/smg-project-smg)