CrabTrap: an LLM-as-a-judge forward proxy for AI agent traffic
An LLM-as-a-judge HTTP proxy to secure agents in production
At a glance
- What is it?
- Brex's CrabTrap intercepts outbound HTTP and HTTPS from AI agents, applies static URL rules and an LLM policy judge, and writes every decision to PostgreSQL. It is a forward proxy with a narrow job, and the README is unusually explicit about what it will not do.
- Who is it for?
- Adopt CrabTrap if you already run agents with broad outbound credentials and want a per-agent, natural-language policy layer with an audit trail; the docker compose quickstart and the QUICKSTART.md walkthrough are the places to begin. Do not adopt it if you need inbound filtering, response redaction, WebSocket frame inspection, or a human approval queue, because the README states it does none of those.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 11 days ago.
- What is it written in?
- Mainly Go, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 2, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem CrabTrap takes on: agents with credentials and no outbound policy
An agent that can call Slack, Gmail and GitHub holds credentials for all three. When the agent's reasoning goes wrong, nothing between it and those APIs asks whether the call should happen. CrabTrap is aimed at exactly that gap. It is an HTTP/HTTPS forward proxy that sits between agents and external APIs, evaluating every outbound request before it reaches the internet. The README frames the audience narrowly: teams running AI agents that call external services and want guardrails. It is not a general egress filter for human traffic, and it is not positioned as one. The unit of policy is the agent, not the network segment, which is why the admin API creates users such as [email protected] and issues each one a gateway_auth_token. That token is the proxy credential the agent presents, so policy follows the agent identity rather than the source IP. The trust boundary is the proxy itself, and the README says so directly: the proxy sees all request content in cleartext, including Authorization and Cookie headers, and does not redact it. Anyone who wants a data-loss-prevention layer in the same box is looking at the wrong project.
Two-tier evaluation: static rules first, LLM judge second
The request path has five documented stages. The agent points HTTP_PROXY and HTTPS_PROXY at CrabTrap. CrabTrap terminates TLS using a per-host certificate generated from its own CA and decrypts the request. The decrypted request is matched against static URL pattern rules using prefix, exact or glob matching, with optional HTTP method filters. If a rule matches, the decision is immediate and no LLM call happens; deny rules always outrank allow rules. Only when no static rule matches does the LLM judge evaluate the request against the agent's natural-language security policy. Allowed requests are forwarded, denied requests receive a 403 with the reason, and every request, decision and response is recorded in PostgreSQL. That ordering is the design's main cost control. Deterministic rules absorb the predictable traffic, and the model is reserved for the ambiguous remainder, which matters because every unmatched request is a paid inference call on the critical path. The judge's own robustness has two documented measures: request payloads are JSON-encoded and policy content is JSON-escaped before reaching the model, and a circuit breaker trips after five consecutive LLM failures and reopens after a 10 second cooldown. When the judge is unavailable, the configured fallback is deny by default, with passthrough available as an explicit choice. Deny-by-default is the safer setting and also the one that turns an LLM outage into an agent outage; the README presents both without recommending one. The certificate cache defaults to 10,000 entries, and per-IP rate limiting uses a token bucket at 50 requests per second with a burst of 100.
Installing CrabTrap with docker compose and making a first proxied request
CrabTrap ships as a container image, quay.io/brexhq/crabtrap:latest, and the repository's docker-compose.yml runs it next to postgres:17-alpine. The compose file maps 8080 and 8081 to the host, sets DATABASE_URL to postgres://crabtrap:secret@postgres:5432/crabtrap, mounts a named certs volume at /app/certs, and waits on a pg_isready healthcheck before starting the gateway. Bring both services up, then copy the generated CA certificate out of the container:
docker compose up -d
docker compose cp crabtrap:/app/certs/ca.crt ./ca.crtThe CA is what makes the interception work. The agent must trust it, or every HTTPS request through the proxy fails certificate validation. Next, create an admin user and capture the web token that the command prints on its last line:
admin_token=$(docker compose exec -it crabtrap ./gateway create-admin-user test-admin \
| tail -n1 | cut -d" " -f2)With that token you can create a non-admin agent identity through the admin API on port 8081 and pull its gateway_auth_token out of the response. The README gives this example:
token=$(curl -X POST http://localhost:8081/admin/users \
-H "Content-Type: application/json" \
-H "Authorization: Bearer ${admin_token}" \
-d '{"id": "[email protected]", "is_admin": false}' \
| jq -r '.channels[] | select(.channel_type == "gateway_auth") | .gateway_auth_token')That token is the username in the proxy URL, and the password is empty. Point curl at the proxy on 8080 and pass the CA explicitly:
curl -x http://${token}:@localhost:8080 \
--cacert ca.crt https://httpbin.org/getA successful run returns the httpbin JSON body, and the request appears in the audit trail. The admin UI is at localhost:8081, where the same admin_token logs you in. Configuration lives in config/gateway.yaml.example, whose sections cover proxy, tls, approval, llm_judge, database, audit, alerting and log_level. The approval section is where you set the mode to llm or passthrough and the judge timeout, which defaults to 30 seconds. The audit destination defaults to stderr and can be set to stdout or a file path. If you would rather build from source, the Makefile's build target compiles the web UI first, copies web/dist into cmd/gateway/web/dist for embedding, and then builds the gateway binary with version, commit and build date injected through ldflags.
The boundaries the README draws around CrabTrap
The project lists its own non-goals, and they are worth reading before anything else. CrabTrap is not a WAF or an inbound firewall; it is a forward proxy for agent-originated traffic and does not inspect requests arriving at your services. It does not redact sensitive data, so headers such as Authorization and Cookie pass through in cleartext. It provides no human-in-the-loop approval: there is no approval queue, no Slack prompt, and no escalation path, and decisions are made automatically by static rules and the LLM judge. It does not filter API responses, which are streamed back to the agent unexamined. It does not inspect WebSocket frames; only the upgrade request is evaluated, and frames pass through uninspected once the connection is upgraded. There is a second-order consequence worth naming. Because responses are never examined, a tool result that carries a prompt injection into the agent's context is outside CrabTrap's view. The project's injection defense is one-directional: it protects the judge's own prompt from the request payload, not the agent from the response body. Teams that need response-side filtering are looking at a different class of product. The other operational sharp edge is the deny-by-default fallback. With the circuit breaker tripping after five consecutive LLM failures, a provider outage stops agent traffic entirely until the cooldown expires and the judge recovers, unless the operator has deliberately switched the fallback to passthrough. That is a deliberate trade of availability for safety, and it should be a decision, not a default nobody noticed.
Policy builder and eval: turning observed traffic into rules
Two features distinguish CrabTrap from a plain filtering proxy. The policy builder is an agentic loop with tools that analyzes observed traffic and drafts security policies automatically, which addresses the cold-start problem: an operator staring at an empty rule list has little idea which URL patterns their agents actually hit. The eval system replays historical audit log entries against a policy to measure accuracy, so a drafted or edited policy can be scored against real recorded traffic before it goes live. Together they form a loop: observe, draft, replay, then enable. The eval system depends on the audit trail being complete and on PostgreSQL retaining the entries you want to replay, so retention policy is effectively part of the eval workflow. Denial alerting notifies bot managers when a new URL pattern is denied, deduplicated with a configurable cooldown, and docs/alerting.md covers the notification channels and the sender interface. The admin UI ties these together with an audit trail viewer, a policy editor, eval results and agent management. None of this is a substitute for a policy that a human has reviewed; the builder drafts, and the eval measures, but the README does not claim either one guarantees correctness.
CrabTrap compared with a gateway that speaks MCP or an egress firewall
The nearest alternative in spirit is an MCP gateway or broker, which mediates agent tool calls at the protocol layer: the agent asks the gateway for a tool, and the gateway decides whether to expose it. That approach gives you typed tool schemas and per-tool authorization, but it only covers traffic that goes through the MCP server. An agent that shells out to curl, or a tool that holds its own HTTP client, bypasses it entirely. CrabTrap sits lower in the stack and catches anything that honors HTTP_PROXY and HTTPS_PROXY, which is a broader net but a blunter one: it sees URLs and bodies, not tool semantics. A traditional egress firewall or secure web gateway is the other comparison. Those are built for human and machine traffic at large, with mature rule languages, and some now offer TLS inspection. They generally lack a per-agent natural-language policy and an LLM judge, and they are not designed around an agent identity token. CrabTrap's advantage is that the policy reads like a sentence about what an agent may do, and the eval system scores that sentence against recorded traffic. Its disadvantage is that the judge is a probabilistic component in the request path, with a circuit breaker and a fallback mode that a firewall does not need.
Maintenance, licensing and what upgrading costs you
CrabTrap is MIT licensed, which permits commercial use, modification and redistribution provided the copyright notice and permission notice are retained; the LICENSE file in the repository root is the authority, and this is not legal advice. The repository is not archived, and the last push was on 2026-07-27, which is recent enough that the project is under current development rather than dormant. Releases have been frequent and small: v0.0.1 on 2026-04-17, v0.0.2 on 2026-05-22 and v0.0.3 on 2026-06-25. All three are 0.0.x, which is a fair signal that the configuration schema and the admin API are still settling. That matters for upgrade cost. The admin API is what your provisioning scripts call to create users and read tokens, and the YAML config keys are what your deployment sets, so both are surfaces where a breaking change is cheap for the maintainers and expensive for you. Practical mitigations are visible in the repository: pin the image tag rather than tracking quay.io/brexhq/crabtrap:latest, keep your gateway.yaml under version control, and diff it against config/gateway.yaml.example after each release. The Makefile also lets you override VERSION, COMMIT and DATE at build time, which is useful if you build your own binary and want the running gateway to report which revision it is. The PostgreSQL dependency is a second ongoing cost: the audit table grows with every proxied request and every recorded response, so retention and vacuum planning belong in the deployment, not in a later cleanup.
Editorial conclusion
Adopt CrabTrap if you already run agents with broad outbound credentials and want a per-agent, natural-language policy layer with an audit trail; the docker compose quickstart and the QUICKSTART.md walkthrough are the places to begin. Do not adopt it if you need inbound filtering, response redaction, WebSocket frame inspection, or a human approval queue, because the README states it does none of those. Before trusting it in production, verify your own deployment against the documented failure modes: the LLM judge's fallback mode, the circuit breaker threshold, and the fact that the proxy holds Authorization and Cookie headers in cleartext.
Frequently asked questions
What is CrabTrap from Brex?
It is an HTTP/HTTPS forward proxy that sits between AI agents and external APIs, evaluating each outbound request against static URL rules and an LLM policy judge, then forwarding or blocking it and logging the decision to PostgreSQL.
Can CrabTrap block requests to internal or private networks?
Yes. The README lists SSRF protection that blocks private networks including RFC 1918 ranges, loopback, link-local, Carrier-Grade NAT, and IPv6 ULA, NAT64 and 6to4 addresses, with DNS-rebinding prevention.
What happens when the LLM judge is unavailable?
A circuit breaker trips after five consecutive LLM failures and reopens after a 10 second cooldown. The fallback mode is configurable as deny, which is the default, or passthrough.
Does CrabTrap inspect API responses or WebSocket traffic?
No. Only outbound requests are evaluated; responses from upstream APIs are streamed back to the agent unexamined, and only the WebSocket upgrade request is evaluated, with frames passing through uninspected after the upgrade.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/brexhq-crabtrap)