All concepts
Concept

What is LLM gateway?

An LLM gateway (also called an LLM proxy or model router) is a service that sits between your application and one or more model providers, giving you a single endpoint for routing, credentials, retries and usage tracking. Instead of calling OpenAI, Anthropic or a self-hosted model directly, your code calls the gateway.

Published September 28, 2026

How an LLM gateway works

The mechanism is a reverse proxy with provider awareness. Your application sends a request in one wire format, usually the OpenAI chat completions shape, to the gateway's base URL. The gateway authenticates the caller, looks up which upstream model the request should reach, rewrites or forwards the payload, attaches the provider credential, and returns the provider's response in the same format. From the application's point of view there is one API; the provider differences are absorbed at the hop in between.

Around that core forwarding path, gateways add a small set of recurring features. Credential management keeps provider keys out of application code and lets one key pool serve many callers. Routing decides which upstream handles a request, by model name, by weight, by priority or by fallback order when the first choice fails. Retries and timeouts cover provider errors and slow responses. Usage accounting records tokens, cost and latency per request, which is what makes chargeback or billing possible. Guardrails inspect prompts or outputs before they leave or reach the caller. Caching can return a stored answer for a repeated prompt.

The important constraint is that a gateway is on the critical path of every request. It adds a network hop and a process that must stay available. If it holds conversation state, streaming responses become harder to proxy correctly, because tokens arrive incrementally and the gateway must pass them through without buffering the whole answer. Config reload behaviour also matters: some gateways read routes from an external store and hot-load changes, others require a restart. Those are architectural choices, not marketing points, and they decide how the gateway fits your deployment.

When you need one, and when you do not

The clearest case is provider sprawl. Once an application calls more than one vendor, or the same vendor through several accounts, the code accumulates per-provider branches for authentication, error handling and response parsing. A gateway collapses that into one client and one configuration file. The related case is operational control: you want per-team keys, spend limits, request logs or a fallback when a provider has an outage, and you do not want to rebuild those in every service.

A second case is multi-tenant or resale access. If you meter model usage for other people, you need accounting and often billing in the same layer that routes the traffic, because splitting them creates reconciliation work.

The negative case is just as concrete. If an application calls exactly one provider and will keep doing so, a gateway is an extra service to run, monitor and upgrade for no routing benefit. Our analysis of LiteLLM puts it plainly: it is worth adopting when provider sprawl, not model quality, is your bottleneck, and it is the wrong tool when you only ever call one vendor. A local development setup that talks to one model endpoint has the same problem. And if your only goal is a nicer chat interface, a gateway is the wrong layer; it routes API traffic, it does not replace an application.

A middle case is a single provider with several API keys or accounts. That still justifies a small proxy, because key rotation and failover are gateway functions even when the provider count is one. GPT-Load is built for exactly this shape: API-key channels and subscription accounts behind one base URL, with scheduling, failover and usage accounting in a bundled UI, according to its project description and our analysis.

Common pitfalls and limits

Latency and availability are the first limit. Every request pays the gateway hop, and the gateway becomes a component whose failure takes down all model access at once. High availability therefore has to be solved at the gateway layer too, which usually means more than one instance and a shared configuration source.

Feature claims need reading closely. Bifrost's description advertises cluster mode, guardrails and low overhead, but our analysis notes that several headline features sit behind enterprise deployments, so the open-source path does not include everything the headline implies. Portkey's gateway is described as an MIT-licensed TypeScript service with retries, fallbacks, load balancing and output guardrails, and our analysis adds that the maintenance picture is not as current as the README implies. That is a reason to check the repository's last push date rather than the feature list.

Upgrade paths are another trap. GPT-Load's 2.0 line is still at release candidate, and per our analysis it cannot read 1.x data, so the upgrade is not a drop-in for existing installations. Kong's split between a Postgres-backed deployment and a DB-less container changes what you can configure at runtime; the DB-less mode is simpler to run but not equivalent. APISIX stores routes in etcd and hot-loads plugins without restarts, which suits dynamic configuration but is a poor fit if you want a single binary with no external store, as our analysis states.

Finally, guardrails and caching are not free correctness. A guardrail that inspects output adds latency and can block legitimate answers; a cache keyed only on prompt text can return stale or wrong answers when model versions change. Neither is a substitute for application-level validation.

How it shows up in open-source projects

The projects below all describe themselves as gateways or proxies, but they occupy different layers. BerriAI/litellm ships as a Python SDK and as a self-hosted proxy on port 4000, per our analysis. Its description lists cost tracking, guardrails, load balancing and logging across providers including Bedrock, Azure, OpenAI, Anthropic, VertexAI, vLLM and Nvidia NIM.

Kong/kong is a Lua-based API, LLM and MCP gateway that runs from a docker-compose stack, a DB-less container or Kubernetes, and it has an official Kubernetes Ingress Controller. Its breadth means configuration is a real surface area, and our analysis frames the Postgres versus DB-less split and plugin cost as the decisions that matter. apache/apisix is a cloud-native API gateway that stores routes in etcd and hot-loads plugins without restarts; it suits teams that need dynamic configuration and is a poor fit if you want a single binary with no external store. TykTechnologies/tyk is a Go reverse proxy fronting REST, GraphQL, gRPC, TCP and MCP traffic with authentication, rate limiting and analytics; the Docker path is the documented quickest start, with a Redis dependency and a licence file that is not a plain identifier as the trade-offs. higress-group/higress is built on Istio and Envoy and extended through Wasm plugins, targeting LLM traffic and MCP server hosting, and it installs either as a single Docker container or through Helm.

Portkey-AI/gateway is an MIT-licensed TypeScript service that puts retries, fallbacks, load balancing and output guardrails in front of any LLM provider; our analysis singles out the routing configuration as the interesting part. maximhq/bifrost is a Go gateway that puts OpenAI, Anthropic, Bedrock, Vertex and others behind one OpenAI-compatible endpoint with fallbacks and load balancing, and it starts in one command, though several headline features sit behind enterprise deployments. tbphp/gpt-load is a self-hosted gateway for multi-channel, multi-credential setups covering API keys and subscription accounts, with scheduling, failover, request logs and usage.

Two projects place the gateway inside a larger product. coaidev/coai is a Go backend plus React frontend that combines a multi-provider LLM gateway, a channel routing layer and credit or subscription billing behind one admin panel, aimed at operators who want to resell or meter model access. InsForge/InsForge bundles Postgres, auth, storage, edge functions and an LLM gateway behind an MCP server and a CLI, so a coding agent can configure the backend itself; our analysis notes that the repository's documentation has gaps, so the gateway is one component of a platform rather than a standalone choice.

Choosing between them

The first question is deployment shape. If you already run Kubernetes and want an ingress controller plus AI routing, Kong, APISIX and Higress are in that family, with different configuration stores: Kong uses Postgres or a DB-less file, APISIX uses etcd, Higress uses Istio and Envoy with Wasm extensions. If you want a process you start next to your app, LiteLLM's port 4000 proxy, Portkey's TypeScript service and Bifrost's one-command start are the closer fit. GPT-Load is aimed at the multi-credential case rather than general routing.

The second question is scope. A gateway that only routes is easier to reason about than one that also bills, hosts MCP servers or provisions a backend. CoAI and InsForge are the wide end of that spectrum. The narrower tools are not worse; they simply leave billing and application scaffolding to other systems.

The third question is maintenance evidence. Read the last push date and the archived flag, not the feature list. Our analysis of Portkey's gateway notes that the maintenance picture is not as current as the README implies, and GPT-Load's 2.0 line is still at release candidate with no path to read 1.x data. Those are the facts that decide whether an upgrade is routine or a migration.

In practice

An LLM gateway is a reverse proxy for model traffic: one endpoint, provider credentials held centrally, routing and fallback rules, and usage accounting. It earns its place when you call more than one provider or need per-caller control, and it is dead weight when you call one provider from one service. To go further, read the README and configuration reference of LiteLLM, Portkey's gateway and Bifrost for the routing model, then check each repository's last push date and archived flag before committing.

BerriAI/litellmSelf-hosted AI gateway with a Rust core and Python SDK that calls 100+ LLM providers in OpenAI format, adding cost tracking, guardrails, and load balancing.59,624 stars · PythonKong/kongThe API and AI Gateway. Kong runs natively on Kubernetes thanks to its official Kubernetes Ingress Controller.44,219 stars · Luaapache/apisixThe Cloud-Native API Gateway and AI Gateway17,168 stars · LuaInsForge/InsForgeThe all-in-one, open-source backend platform for agentic coding. InsForge gives your coding agent database, auth, storage, compute, hosting, and AI gateway to ship full-stack apps end-to-end.13,022 stars · TypeScriptPortkey-AI/gatewayA blazing fast AI Gateway with integrated guardrails. Route to 1,600+ LLMs, 50+ AI Guardrails with 1 fast & friendly API.13,022 stars · TypeScriptTykTechnologies/tykOpen Source API and AI Gateway supporting REST, GraphQL, TCP, gRPC and MCP (Model Context Protocol)10,837 stars · Gohigress-group/higress🤖 AI Gateway | AI Native API Gateway9,474 stars · Gocoaidev/coai🚀 Next Gen Multi-tenant AI One-Stop Solution. Builtin Admin & Billing System. Enterprise-Grade Unified LLM Gateway Support for 200+ Models And 35+ Providers, Load Balacing w/ Priority-base Routing, Cost Management, Chat Share, Cloud Sync, Credit/Subscription Billing, All File Parsing, Web Search, Built-in Model Cache.9,307 stars · TypeScriptmaximhq/bifrostFastest enterprise AI gateway (50x faster than LiteLLM) with adaptive load balancer, cluster mode, guardrails, 1000+ models support & <100 µs overhead at 5k RPS.8,140 stars · Gotbphp/gpt-loadSelf-hosted AI gateway for multi-channel, multi-credential setups — API keys and subscription accounts, scheduling, failover, request logs and usage. 自托管 AI 网关:多渠道多凭据统一接入,含密钥与订阅账号、调度容错、日志与用量。7,009 stars · GoBlockRunAI/ClawRouterThe agent-native LLM router for autonomous agents. Every frontier model behind one wallet, <1ms local routing, USDC payments on Base & Solana via x402.6,614 stars · TypeScriptweave-os/routerModel router for agentic systems. Routes every prompt to the right model in <50ms. Cut costs 40-70% with just an endpoint change.5,416 stars · Go

Sources

  1. BerriAI/litellm repository
  2. Kong/kong repository
  3. apache/apisix repository
  4. Portkey-AI/gateway repository
  5. InsForge/InsForge repository