weave-os/router: a Go model router for agentic systems
Model router for agentic systems. Routes every prompt to the right model in <50ms. Cut costs 40-70% with just an endpoint change.
At a glance
- What is it?
- weave-os/router is an OpenAI-compatible endpoint that picks a model per prompt using an in-process cluster scorer. Here is what the repository shows about how it works, how to run it locally, and where it stops being the right tool.
- Who is it for?
- Adopt weave-os/router if you already run agentic traffic through an OpenAI-compatible client and want the routing decision to happen inside your own infrastructure, with Postgres as the only stateful dependency. Do not adopt it if you need a documented rollback path, a published cost model, or a licence you can classify without asking.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository received new commits within the last day.
- What is it written in?
- Mainly Go, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem weave-os/router targets: per-prompt model choice inside an agent loop
Agentic systems send many prompts per task. Some are cheap classification steps, some are long tool-call reasoning turns, and the cost difference between sending all of them to a frontier model and sending the easy ones to a small OSS model is large. The repository describes the project as a "Model router for agentic systems" that "Routes every prompt to the right model in <50ms" and claims it can "Cut costs 40-70% with just an endpoint change". Those numbers are the project's own claims, not measurements published with a methodology in the README.
The audience is narrow and identifiable. The topics list agentic-coding, ai-gateway, anthropic, claude-code, codex, model-router and openai-compatible. That is a tool for teams running coding agents or similar tool-calling loops who already speak the OpenAI chat-completions dialect and do not want to rewrite their client. The pitch is that routing is a configuration change at the base URL, not a code change. Whether that holds depends on how much of your prompt shape the scorer can see, which the README does not spell out.
How the routing decision is made: ONNX embedders, a cluster scorer and a frozen HMM
The mechanism visible in the repository is a two-stage scorer. The Dockerfile states that the router "uses hugot + onnxruntime_go (CGO; dynamic-links against libonnxruntime.so) to run the cluster scorer's embedders (Jina v2, Qwen3-0.6B) in-process". So the embedding step is not a network call to a hosted embedding API. It runs inside the router process, which is why the image needs a glibc base: the Dockerfile notes that "Alpine/musl can't load the library out of the box", with a bookworm-glibc builder and a distroless/cc-debian12 runtime.
The embedders are pulled from Hugging Face at build time. HF_MODEL_REPO defaults to jinaai/jina-embeddings-v2-base-code, pinned by HF_MODEL_REVISION to a specific commit, and the Dockerfile comments explain the pin as a guard against "the build silently picked up new weights". A second embedder, Qwen3-Embedding-0.6B, is optional: setting HF_QWEN_REPO to empty skips the pull, and the Dockerfile notes the runtime constructs embedders lazily. The go.mod confirms the runtime dependency set: hugot for ONNX inference, gin for HTTP, pgx for Postgres, and a large OpenTelemetry block for traces and metrics.
On top of the embeddings sits a frozen model. The only release listed is hmm-model-v1, dated 2026-07-13, described as "HMM frozen model v1". The name suggests a hidden Markov model, and the word frozen suggests it is not retrained at runtime. The README does not document the training data, the label set, or how a prompt is mapped to a model choice, so the internal scoring logic is the least transparent part of the project. What can be said is that the decision is local, deterministic given the same weights, and does not require an outbound call to a routing service.
State lives in Postgres. The docker-compose file brings up postgres:15-alpine, applies migrations, and mentions a Pub/Sub emulator "used for cross-replica cache invalidation", with GCP Pub/Sub in production. That is the one architectural detail worth pausing on: if you run more than one replica, cache invalidation crosses a message bus, and the local stack substitutes an emulator for it.
Running weave-os/router locally with docker compose
The repository ships a compose file explicitly labelled "DEV ONLY". Its header says it "Brings up Postgres, applies migrations, and starts the server end-to-end" and warns not to deploy it "to any environment exposed beyond localhost". Start there.
docker compose up --buildThe compose file documents two health checks. The first should return a success response; the second should return 401 because the key is deliberately invalid.
curl http://localhost:8080/health
curl -H 'Authorization: Bearer rk_invalid' http://localhost:8080/validateTo get a usable API key, the compose header gives a seed command that creates an installation and a key.
docker compose run --rm seedOne port detail matters if you already run Postgres locally. The compose file maps the database to host port 5433 rather than 5432, and the comment explains why: to avoid colliding with the Weave backend's local Postgres. Inside the compose network the router still talks on the internal 5432. If you connect with psql from the host, use 5433.
For a non-Docker path, .env.example is the template. It says to copy it to .env.local, that "docker-compose loads .env.local automatically", and that "make dev reads it via godotenv". Anything marked REQUIRED must be set for the router to boot. DATABASE_URL takes precedence over the individual POSTGRES_* variables, and POSTGRES_SSLMODE should be set to disable for local Docker. The same file notes that at least one provider API key is recommended, that OpenRouter "is the suggested baseline" because it "unlocks the OSS-model pool the cluster scorer is trained against", and that Anthropic is special: when ANTHROPIC_API_KEY is unset, the router forwards client Anthropic auth headers upstream. That last behaviour is a design decision worth reading twice before you rely on it in a shared deployment.
Where the repository goes quiet: licence, rollback and the cost claim
The licence field is NOASSERTION. That is not a licence. It means the automated classifier could not identify one from the LICENSE file, and it is the first thing to resolve before this touches production traffic. The repository does contain a LICENSE file at the top level, so the answer exists; it just is not summarised anywhere in what is published. Treat any legal question about redistribution or modification as unanswered until someone reads that file.
The README does not document rollback. There is no described procedure for reverting a routing decision, disabling the scorer, or pinning all traffic to a single model if the scorer misbehaves. The closest thing to an escape hatch visible in the published files is the Anthropic auth passthrough, which is about credentials rather than routing. If your agent loop cannot tolerate a wrong model choice on a critical path, you are relying on an undocumented fallback.
The 40-70% cost reduction is stated without a baseline. It does not say which model mix produces that range, what the traffic distribution was, or whether the figure includes the embedding compute the router adds on every prompt. Running two ONNX embedders in-process is not free; it is CPU and memory you now pay for on your own hardware. The claim may well be true for the traffic the authors measured, but the repository does not let you check it against yours.
Finally, the build is heavier than the Go module list suggests. CGO, a glibc base image, a pinned ONNX Runtime version, and a build-time Hugging Face pull with an optional HF_TOKEN secret all mean your CI needs network access to Hugging Face and a Docker builder that can handle CGO. That is a real operational cost the "just an endpoint change" framing understates.
weave-os/router versus a fixed-model gateway
The obvious alternative is a plain OpenAI-compatible gateway that forwards every request to one model, or to a model you select per client. The difference is where the decision happens. A fixed gateway makes the choice at configuration time, in your code or your environment variables. weave-os/router makes it per prompt, at request time, using a scorer trained against a particular pool of models.
That is a genuine trade-off, not a strict upgrade. A fixed gateway is trivially debuggable: the same prompt always goes to the same model, so a regression in output quality has one candidate cause. With a router, the same prompt can go to different models if the scorer's inputs change, and the README does not document a way to inspect or override the choice. You gain cost control and lose reproducibility unless the project provides an audit surface that the published files do not describe.
A second alternative is to keep the decision in application code: classify the prompt yourself and call whichever provider you want. That is more work and it puts routing logic in every client, but it is fully visible and testable with your existing tooling. weave-os/router is the better fit when you have many clients and want one place to change the policy, and the worse fit when you have one client and need to explain every routing decision to someone.
Maintenance, upgrades and what the release history shows
The repository is not archived, and the last push was on 2026-09-20. On that evidence alone it is current. The release history is thin: a single listed release, hmm-model-v1, dated 2026-07-13. That release name is the one to watch, because it is the frozen scoring model. If a future release changes the HMM, routing behaviour changes without any change to your client code, and the README does not describe a versioning or pinning scheme for the model itself.
The Dockerfile does show deliberate pinning discipline for the parts the authors control: ONNX Runtime version, tokenizers version, and Hugging Face revisions for both embedders, with comments telling the reader to bump them deliberately. That is a good sign for build reproducibility. It does not extend to the scoring model, which is the component most likely to alter your costs and output quality.
Upgrade cost therefore splits in two. Rebuilding the image is routine if your CI can reach Hugging Face and build CGO. Changing the frozen model is the upgrade that needs a test plan, and the repository does not provide one. The go.mod pins Go 1.25.0 with toolchain go1.25.9, so a Go toolchain upgrade on your side is a prerequisite rather than something the project manages for you.
Editorial conclusion
Adopt weave-os/router if you already run agentic traffic through an OpenAI-compatible client and want the routing decision to happen inside your own infrastructure, with Postgres as the only stateful dependency. Do not adopt it if you need a documented rollback path, a published cost model, or a licence you can classify without asking. Before anything else, check the LICENSE file, run docker compose up --build and confirm /health responds, then point one non-critical client at the endpoint and compare its routing decisions against your current fixed-model setup.
Frequently asked questions
What is weave-os/router used for?
It is a model router for agentic systems. According to the repository description it routes every prompt to a model in under 50ms, and it presents itself as an OpenAI-compatible endpoint so clients can point at it with an endpoint change.
How do I install weave-os/router locally?
The repository ships a docker-compose.yml that the file itself labels DEV ONLY. Running docker compose up --build brings up Postgres, applies migrations and starts the server; the compose header lists curl http://localhost:8080/health as the health check and docker compose run --rm seed to create an installation and API key.
Does weave-os/router run models itself?
No. The Dockerfile states that it runs the cluster scorer's embedders, Jina v2 and Qwen3-0.6B, in-process via hugot and onnxruntime_go. The generation models are reached through provider API keys configured in .env.example, with OpenRouter described as the suggested baseline.
What licence is weave-os/router under?
The repository's licence field is NOASSERTION, which means the licence could not be identified automatically. A LICENSE file exists at the top level, so the terms have to be read from that file rather than from any summary.
Can I run weave-os/router without Docker?
.env.example describes a non-Docker path: copy the template to .env.local, set the REQUIRED values, and note that make dev reads it via godotenv. DATABASE_URL takes precedence over the individual POSTGRES_* variables, and POSTGRES_SSLMODE should be set to disable for local use.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/weave-os-router)