# Portkey AI Gateway: a self-hosted router for 1,600+ models with guardrails

> Portkey's AI Gateway is an MIT-licensed TypeScript service that puts retries, fallbacks, load balancing and output guardrails in front of any LLM provider. The routing config is the interesting part; the maintenance picture is not as current as the README implies.

**Portkey-AI/gateway** — A blazing fast AI Gateway with integrated guardrails. Route to 1,600+ LLMs, 50+ AI Guardrails with 1 fast & friendly API.

- Repository: https://github.com/Portkey-AI/gateway
- Website: https://portkey.ai/features/ai-gateway
- Stars: 13,022 · Forks: 1,305
- Language: TypeScript
- License: MIT
- Published: 2026-09-09 · Updated: 2026-09-09 · Language: en
- Canonical page: https://hysenlabs.com/projects/portkey-ai-gateway

## The problem: provider sprawl in application code

Calling one model provider is easy. Calling four is not, because each one brings its own client, its own authentication header, its own error shapes and its own outage pattern. The usual result is provider-specific code scattered through the application, with retry logic duplicated next to each call site and no single place to see what was sent where.

Portkey's AI Gateway is aimed at that layer. It is a TypeScript service that sits between your application and the provider, exposes an OpenAI-compatible endpoint, and holds the routing policy in a config object rather than in your application code. The README describes it as "designed for fast, reliable & secure routing to 1600+ language, vision, audio, and image models" and claims a 122kb footprint. Those numbers come from the project's own README, not from independent measurement.

The audience is teams that have outgrown a single provider but do not want to build a routing layer themselves: platform engineers supporting several product teams, and application developers who need a fallback when the primary model is unavailable. It is less obviously aimed at a solo developer calling one model, where the gateway adds a process to run and a config format to learn for very little gain.

## How routing, retries and guardrails actually work

The mechanism is a config object attached to the client. Instead of branching in application code, you describe what should happen: how many times to retry, which guardrails to apply, and (in the configurations the README links to) how to fall back or load balance across targets. The gateway receives an OpenAI-shaped request, applies the config, calls the provider, and returns a response in the same shape.

Guardrails run on the output. In the README's example, an output guardrail checks whether the response contains a word from a list and denies the response if it does, with the retry config causing the request to be attempted again rather than returned to the caller. That is a loop, not a filter: a denied response is retried, so a model that keeps producing the blocked word will burn the configured number of attempts before the call fails. Worth knowing before you set attempts to a high number.

The runtime is deliberately portable. The package.json shows three execution paths: a Node server (start:node running build/start-server.js), a Cloudflare Workers build via wrangler, and the default development script which runs under workerd. The Dockerfile is a two-stage node:20-alpine build that compiles with rollup, reinstalls production dependencies, and exposes port 8787. A single gateway deployment can therefore live on a VM, in a container platform, or on Workers without changing the request format your application uses.

## Installing the gateway and making a first guarded call

The README's quickstart runs the gateway straight from npm. Node.js and npm are the only stated prerequisites, and the service listens on port 8787.

```bash
# Run the gateway locally (needs Node.js and npm)
npx @portkey-ai/gateway
```

After that command, the README says the Gateway is running on http://localhost:8787/v1 and the Gateway Console on http://localhost:8787/public/. The console is where local request logs are collected, which is the fastest way to confirm traffic is actually passing through rather than going direct.

If you prefer a container, the repository ships a compose file that pulls the published image and maps the same port.

```yaml
version: '3'
services:
  web:
    ports:
      - "8787:8787"
    image: "portkeyai/gateway:latest"
    restart: always
```

The client side is a Python package. The README example sets the provider and passes the provider key as Authorization, then calls chat.completions.create with a model name.

```python
# pip install -qU portkey-ai

from portkey_ai import Portkey

client = Portkey(
    provider="openai", # or 'anthropic', 'bedrock', 'groq', etc
    Authorization="sk-***" # the provider API key
)

client.chat.completions.create(
    messages=[{"role": "user", "content": "What's the weather like?"}],
    model="gpt-4o-mini"
)
```

The part worth copying is the config, because it shows the shape of the whole system. Here the client is rebuilt with a config that retries five times and denies any output containing a specific word.

```python
config = {
  "retry": {"attempts": 5},

  "output_guardrails": [{
    "default.contains": {"operator": "none", "words": ["Apple"]},
    "deny": True
  }]
}

# Attach the config to the client
client = client.with_options(config=config)
```

The README states that this call would always respond with "Bat" because the guardrail denies replies containing "Apple" and the retry config retries five times before giving up. If you run it, expect latency to rise with the number of retries, and expect a hard failure when every attempt is denied.

## Where the gateway gets in the way

The config is a JSON-shaped object passed at call time, which means routing policy lives in application code after all, just in a different form. If you want to change a fallback order across a fleet of services, you are changing and redeploying each caller unless you route config through something else. The README does not describe a shared config store for the self-hosted gateway.

Guardrails as implemented in the example are word-list checks, not semantic moderation. A contains check on a list of words is cheap and predictable, and also trivially bypassed by spelling variation or by a model that phrases the same idea differently. Treat it as a tripwire for known-bad strings, not as a safety boundary.

Retries change cost and latency in ways that are easy to forget. Five attempts against a paid model means up to five billable calls for one user request when the guardrail or the provider keeps failing, and the caller waits for all of them. Any retry setting needs to be chosen with the failure mode in mind, not just the success path.

Finally, the gateway is another hop. The README claims sub-millisecond latency and a 122kb footprint, which describes the gateway's own overhead rather than end-to-end response time; the provider call still dominates. If your application is a single service calling a single provider with no fallback requirement, the gateway is a component you now have to deploy, monitor and upgrade for no routing benefit.

## Alternatives and the difference in approach

The most direct alternative is to skip the gateway and use each provider's OpenAI-compatible endpoint with a thin client-side wrapper. OpenAI's own SDKs are listed among the supported libraries, so an application can point an OpenAI client at a different base URL and keep most of its code. The difference is where policy lives: with a wrapper you write retries, fallbacks and filtering yourself, in your language, with your own tests. You avoid running a service and you avoid a config format, but you also own the failure handling.

A second path is to use a hosted gateway rather than self-hosting. The README links a Hosted Gateway alongside the open-source repository and recommends Portkey Cloud for deployment. That removes the operational work and the upgrade cadence, at the cost of sending prompts and provider credentials through someone else's infrastructure and depending on their availability. For teams with data-residency constraints, the self-hosted path is the reason to use this project at all.

The third option is a language-native router inside your existing framework. The README lists LangChain, LlamaIndex, Autogen and CrewAI integrations, which means you can keep the gateway and still use those frameworks; the choice is not strictly either/or. What the gateway adds over a framework-level router is that the policy applies to every caller regardless of language, including services that are not Python or JavaScript.

## Maintenance, licensing and upgrade cost

The repository is not archived, but the last push was on 2026-05-25, which is roughly four months before the date of writing. The most recent release listed is v1.15.2 on 2026-01-12, preceded by v1.15.1 on 2025-12-24 and v1.15.0 on 2025-12-23. The npm package version in package.json matches 1.15.2. On that evidence the project is not being actively developed right now, and the README's own banner supports a reading of transition rather than steady state: it announces Gateway 2.0 as a pre-release, says the core enterprise gateway is "merging into open-source with our 2.0 release", and points readers at a 2.0.0 branch.

That has a practical consequence. Anything you build against the 1.x config format may face a migration when 2.0 lands, and the README does not document a migration path, a compatibility guarantee or a rollback procedure. Pinning the npm version rather than tracking main is the low-risk posture, and testing against the 2.0.0 branch separately is how you find out early whether your configs survive.

The licence is MIT, which is permissive and places few obligations on how you redistribute or modify the code. Two things to keep in mind that are not legal advice: the repository contains no separate licence file for the plugins directory, so if you ship a third-party plugin, check what applies to it; and the project is developed by a company that sells a hosted and enterprise product, so the open-source gateway and the commercial offering should be evaluated as separate dependencies with separate lifecycles.

## Conclusion

Adopt it if you are already calling several providers directly and want retries, fallbacks and output filtering in one place without writing that layer yourself, and if you are willing to pin the npm package rather than track main. Do not adopt it if you need a gateway that is being actively developed right now: the last push to main was on 2026-05-25, the newest release is v1.15.2 from 2026-01-12, and the README points at a pre-release 2.0 branch instead. Verify three things before you commit: that the provider you need appears in the supported list, that a Docker image tagged the way you deploy it actually exists, and how you will supply provider keys in your environment, because the quickstart passes an Authorization value inline.

## FAQ

### How do I install the Portkey AI Gateway?

The README's quickstart runs it directly with npx @portkey-ai/gateway, which needs Node.js and npm and serves on port 8787. The repository also ships a Dockerfile and a docker-compose.yaml that pulls the portkeyai/gateway:latest image and maps port 8787.

### How do I use the Portkey AI Gateway?

Install the portkey-ai Python package, create a Portkey client with a provider name and the provider API key as Authorization, then call chat.completions.create with a model. Routing rules such as retries and output guardrails are attached with client.with_options(config=config).

### What ports does the Portkey AI Gateway use?

The README states the gateway runs on http://localhost:8787/v1 and the Gateway Console on http://localhost:8787/public/. The Dockerfile exposes 8787 and the compose file maps 8787:8787.

### Is the Portkey AI Gateway free to use?

The repository is licensed under MIT, which permits modification and redistribution. The README also links a Hosted Gateway and an Enterprise offering, which are separate from the open-source code.

### Can the Portkey AI Gateway run on Cloudflare Workers?

Yes. The package.json includes a dev:workerd script running wrangler dev on src/index.ts and a deploy script using wrangler deploy, and the README lists Cloudflare among the deployment guides.

## Sources

- [License: MIT](https://github.com/Portkey-AI/gateway/blob/main/LICENSE)
- [Portkey-AI/gateway on GitHub](https://github.com/Portkey-AI/gateway)
- [Project website](https://portkey.ai/features/ai-gateway)
- [README](https://github.com/Portkey-AI/gateway/blob/main/README.md)
- [Releases](https://github.com/Portkey-AI/gateway/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/portkey-ai-gateway
