Model or dataset
katanemo/plano avatar
katanemo/plano

katanemo/plano: an Envoy-based LLM gateway and agent router written in Rust

Plano is an AI-native proxy server and data plane for agentic apps. Smart LLM routing, observability, agent orchestration, and guardrails so you stay focused on your agents core logic.

7,070 stars484 forksRustApache-2.0

At a glance

What is it?
Plano is an Apache-2.0 proxy server that moves agent routing, model selection and tracing out of your application code and into a config file. It is a good fit if you already run more than one agent or more than one model provider, and the wrong tool if you want a library you can import.
Who is it for?
Adopt Plano if your agent code has grown routing branches, provider adapters or hand-rolled tracing, and you are willing to run another process in front of your services. Skip it if you have a single agent calling a single model, or if you cannot accept a hosted routing model in the request path.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 1 day ago.
What is it written in?
Mainly Rust, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The hidden middleware Plano is trying to delete

The README opens with a claim worth taking seriously: building an agentic demo is easy, shipping one is not. The work that appears after the demo is what the project calls hidden middleware. Routing logic that decides which agent handles a request. Guardrail hooks for moderation. Evaluation and tracing glue. Provider quirks spread across frameworks and application code. Plano's answer is to move those concerns out of process, into a data plane that sits in front of your agents and your model providers.

The target user is a team that already has agents running as HTTP services and is now maintaining the connective tissue between them. The README is explicit that agents stay ordinary HTTP servers implementing an OpenAI-compatible chat completions endpoint, in any language or framework. Plano does not ask you to adopt an SDK or rewrite agents against a new interface. That is the main architectural bet, and it is also what makes the project easy to evaluate: if your agents already speak the OpenAI chat completions shape, the integration surface is a URL and a YAML file.

How the data plane is put together: Envoy, WASM filters and a Rust binary

The repository layout and Dockerfile show the mechanism. Plano is not a from-scratch proxy. The Dockerfile pins ENVOY_VERSION=v1.37.0 and comments that the value must be kept in sync with cli/planoai/consts.py ENVOY_VERSION. On top of Envoy, two Rust crates are compiled to WebAssembly with the wasm32-wasip1 target: prompt_gateway and llm_gateway. A third crate, brightstaff, is built as a normal release binary rather than a WASM module. The README describes Plano as built on Envoy by its core contributors, which matches what the build files do.

That split matters when you reason about failure. Envoy owns listeners and connection handling. The WASM filters own request-level logic such as prompt handling and LLM routing. brightstaff is a separate process, which is consistent with the idea that orchestration decisions are not made inline in the proxy filter. A request arrives at a listener, the filter chain runs, and the request is forwarded either to an agent URL or to a model provider depending on the listener type. Configuration is declarative: agents, model_providers, listeners and tracing are all top-level keys in one YAML file. The README does not document the internal protocol between the proxy and brightstaff, so treat that boundary as opaque unless you read the crates.

Installing Plano and routing a first request between two agents

The README points to the quickstart guide at docs.planoai.dev/get_started/quickstart.html for prerequisites and installation, and shows planoai up config.yaml as the command that starts the server. The CLI lives in the cli/ directory of the repository. Because the README defers installation details to the quickstart rather than listing package manager commands, check that page for the current install path instead of guessing a package name.

The configuration below is the example from the README, a travel assistant that fronts two agents. It declares the agent URLs, two model providers with access keys read from environment variables, and a listener of type agent on port 8001 using the router plano_orchestrator_v1. The agents block inside the listener carries natural language descriptions, which is what the router uses to pick a destination.

yaml
version: v0.3.0

agents:
  - id: weather_agent
    url: http://localhost:10510
  - id: flight_agent
    url: http://localhost:10520

model_providers:
  - model: openai/gpt-4o
    access_key: $OPENAI_API_KEY
    default: true
  - model: anthropic/claude-3-5-sonnet
    access_key: $ANTHROPIC_API_KEY

listeners:
  - type: agent
    name: travel_assistant
    port: 8001
    router: plano_orchestrator_v1
    agents:
      - id: weather_agent
        description: |
          Gets real-time weather and forecasts for any city worldwide.
      - id: flight_agent
        description: |
          Searches flights between airports with live status and schedules.

tracing:
  random_sampling: 100

Start the server with the CLI, then send an OpenAI-shaped request to the listener port rather than to the agents directly. The README's third step describes querying Plano so that it routes to both agents in a single conversation.

bash
planoai up config.yaml

Agent code stays small. The README's weather agent example is a FastAPI app that points an AsyncOpenAI client at http://localhost:12001/v1 with api_key="EMPTY", which is Plano's LLM gateway rather than the provider directly. The agent fetches its own data and streams a completion back through that base URL, so model selection and provider credentials never appear in agent code.

python
from openai import AsyncOpenAI

llm = AsyncOpenAI(base_url="http://localhost:12001/v1", api_key="EMPTY")

stream = await llm.chat.completions.create(
    model="openai/gpt-4o",
    messages=[{"role": "system", "content": f"Weather: {weather_data}"}, *messages],
    stream=True,
)

The two ports worth remembering are 8001 for the agent listener and 12001 for the LLM gateway, both taken from the README example. If your first request returns a connection error rather than a model response, the listener port is the first thing to check.

The orchestrator model is the part to interrogate before production

The listener config sets router: plano_orchestrator_v1, and the README annotates it as powered by a 4B-parameter routing model, adding that you can change it to different models. A routing decision therefore involves a model call, not just a lookup table. That is the design's most interesting trade-off. Natural language agent descriptions mean you add an agent by editing YAML instead of writing an intent classifier, which is genuinely less code. It also means intent classification is a learned behaviour with the usual failure modes: overlapping descriptions, ambiguous queries, and a routing step that can be wrong in ways a deterministic rule would not be.

The README states that Plano and the Plano family of LLMs, including Plano-Orchestrator, are hosted free of charge in the US-central region for a first-run developer experience, and that to scale in production you can run these LLMs locally or contact the team on Discord for API keys. Read that literally. The convenient default path sends routing traffic to a hosted endpoint in one region, and the documented production path is to self-host the model or obtain keys. Neither option is a one-line change, and the README does not describe the resource requirements for running the orchestrator locally. If your agents handle data that cannot leave a region, or if a hosted dependency in the request path is unacceptable, this is the decision that gates everything else.

Filter chains, signals and what the README leaves unspecified

Two capabilities have documentation pages linked from the README but no detail in the README itself. Filter Chains are described as the mechanism for jailbreak protection, moderation policies and memory, and Agentic Signals are described as zero-code capture of signals plus OTEL traces across every agent. The tracing block in the example sets random_sampling: 100 with the comment that traces are captured automatically for evaluation. What the README does not say is where traces are exported, what the sampling value's scale is, or what storage the signals land in. Those are exactly the questions an operations team asks first, and they are answered, if at all, on docs.planoai.dev rather than in the repository front page.

The same gap applies to failure behaviour. The README does not document what a listener does when an agent URL is unreachable, whether requests are retried, or how model fallbacks are ordered when several providers are configured. The comment in the YAML says you do not write model fallbacks yourself, which implies Plano handles them, but the policy is not stated. Treat the routing and fallback behaviour as something to establish by reading the crates or testing directly, not as something the README settles.

Where a plain OpenAI-compatible gateway is the better choice

Plano's differentiator is the agent listener: routing between your own agents by description, in addition to routing between model providers. If you only need the second half, a conventional LLM gateway is a smaller thing to operate. LiteLLM is the obvious comparison. It presents itself as a proxy that exposes an OpenAI-compatible interface in front of many providers, with spend tracking and key management, and it is distributed as a Python package you can run in-process or as a server. The difference in approach is the unit of routing. LiteLLM routes between models and providers; Plano routes between agents and models, and the agent routing step is what pulls in the orchestrator model and the hosted-or-self-hosted question above.

That distinction should drive the decision rather than a feature checklist. A team with one agent and three providers gets most of the value from the simpler gateway and takes on less. A team with several agents, each owning a slice of a conversation, is the case Plano was built for, and the YAML in the README is a fair picture of how much configuration that costs.

Licence, release cadence and the cost of running another process

Plano is Apache-2.0, which permits commercial use and modification, and the LICENSE file sits at the repository root. The practical implication to note is the hosted orchestrator: the code is permissively licensed, but the default routing model is a service the project operates, and the README describes production use as running those LLMs locally or obtaining API keys. Licence terms on the source do not settle the terms of that service, so read them separately if routing traffic would leave your infrastructure. This is a description of what the repository states, not legal advice.

On maintenance, the last push to the default branch was on 2026-08-19, and releases 0.4.34 through 0.4.36 landed on 2026-08-17, 2026-08-18 and 2026-08-19. The repository is not archived. The version string in the example config is v0.3.0 while the releases are on 0.4.x, so expect config schema drift between the README example and the current CLI; check the quickstart for the version your install expects. Upgrade cost is dominated by two moving parts: the pinned Envoy version, which the Dockerfile ties to a constant in the Python CLI, and the WASM filters, which are rebuilt from the Rust crates. If you build from source rather than pulling an image, you are tracking both.

Editorial conclusion

Adopt Plano if your agent code has grown routing branches, provider adapters or hand-rolled tracing, and you are willing to run another process in front of your services. Skip it if you have a single agent calling a single model, or if you cannot accept a hosted routing model in the request path. Before committing, verify three things in your own environment: whether plano_orchestrator_v1 can be replaced with a locally hosted model for your traffic, what the listener does when an agent URL is unreachable, and how the tracing sampling setting behaves under load.

Frequently asked questions

Does Plano work with agents written in languages other than Python?

Yes, according to the README. Agents are HTTP servers that implement the OpenAI-compatible chat completions endpoint, and the README states you can use any language or AI framework. The Python example is one illustration, not a requirement.

What is the difference between Plano's agent listener and its LLM gateway?

A listener of type agent routes a request to one of your agents based on the natural language descriptions in the config, while the LLM gateway on port 12001 handles model selection and provider credentials for the agents. The README's example config uses both together, and it states you can configure only one of them.

Can Plano run in production without the hosted orchestrator?

The README states that Plano and the Plano family of LLMs are hosted free of charge in the US-central region for a first-run developer experience, and that to scale and run in production you can either run these LLMs locally or contact the team on Discord for API keys. The README does not document the resources needed to run the routing model locally.

Official sources

  1. katanemo/plano on GitHub
  2. License: Apache-2.0
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/katanemo-plano.svg)](https://hysenlabs.com/projects/katanemo-plano)