Model or dataset
ThinkWatchProject/ThinkWatch avatar
ThinkWatchProject/ThinkWatch

ThinkWatch: an AI bastion host for API and MCP traffic

Enterprise AI bastion host for secure AI API and MCP access, with unified proxying, RBAC, audit logs, rate limiting, and cost tracking across OpenAI, Anthropic, Gemini, and self-hosted LLMs.

814 stars22 forksRustNOASSERTION

At a glance

What is it?
ThinkWatch puts every model call and MCP tool invocation behind one gateway with RBAC, audit logs and cost tracking. Here is what the repository actually documents, and where its design choices bite.
Who is it for?
Adopt ThinkWatch if you already run shared AI credentials across several teams and need per-user attribution plus MCP identity propagation, and you can operate PostgreSQL, Redis and ClickHouse. Do not adopt it if you want a single static API key in front of one provider, or if your MCP upstreams issue only shared service tokens you cannot replace.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository last received commits 6 days ago.
What is it written in?
Mainly Rust, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem ThinkWatch targets: shared keys and unattributable spend

The README opens with a list of failures it intends to fix: API keys hardcoded in .env files and passed around in chat, no record of which model a team used, every developer holding direct access to every model and MCP tool, and monthly bills nobody can attribute. Those are operational problems, not model problems, and they appear once more than one team shares the same provider accounts.

The framing the project chooses is an SSH bastion. The README states that just as an SSH secure gateway is the single path through which server access must flow, ThinkWatch is the single path through which AI access must flow. That analogy sets the deployment shape: clients point at the gateway, the gateway holds the provider credentials, and the console at :3001 is where administrators issue and revoke access. It is aimed at platform or security engineers in an organization large enough to have several AI clients (Claude Code, Cursor, CI pipelines) and at least one compliance question to answer.

Gateway and console: two ports, one control plane

The architecture diagram in the README shows two listeners. The gateway on :3000 accepts AI API and MCP traffic from clients such as Claude Code, Cursor, custom agents and CI/CD pipelines, and forwards it to OpenAI, Anthropic, Google Gemini, Azure OpenAI or AWS Bedrock. The console on :3001 serves the management UI and the admin API, and it is where an operator browser connects.

The gateway is multi-format. According to the README it natively serves OpenAI Chat Completions at /v1/chat/completions, Anthropic Messages at /v1/messages and OpenAI Responses at /v1/responses on the same port, and performs automatic format conversion for Anthropic, Gemini, Azure OpenAI and the Bedrock Converse API behind that unified interface. Streaming responses are passed through as SSE with token counting done in real time.

Routing is partly automatic and partly manual, and the split matters. The README says active providers are loaded from the database at startup and registered in the model router, with default prefixes gpt-, o1-, o3- and o4- going to OpenAI, claude- to Anthropic and gemini- to Google. Azure and Bedrock require explicit model registration. If you run mostly Azure, expect more configuration than the prefix table suggests.

Access is issued as virtual API keys with the tw- prefix. A single tw- token can work on both the AI gateway and the MCP gateway, gated by a per-key surfaces allowlist. Keys have a lifecycle: automatic rotation with grace periods, a per-key inactivity timeout, expiry warnings and background policy enforcement. Quotas are composable. The README describes sliding windows at 1m, 5m, 1h, 5h, 1d and 1w, token budgets on daily, weekly and monthly periods, and scoping per user, per API key, per provider or per MCP server. Per-model weighting lets one model's tokens count more than another's against the same quota through input_multiplier and output_multiplier. Upstream failures are handled by a three-state circuit breaker (Closed, Open, HalfOpen) and retries with exponential backoff and jitter.

MCP identity propagation is the differentiating mechanism

Most of the MCP section describes one design decision and its consequences. The README states plainly that the upstream server sees the real end user, not a shared service account, and contrasts this with gateways that pin one admin token to the server config so the upstream audit log records every action as the same account.

Concretely, every MCP request carries the calling user's own OAuth token or personal access token for services such as GitHub, Notion, Linear, Slack, Atlassian, Feishu, GitLab, Cloudflare, Google and Discord. Those credentials are stored AES-256-GCM encrypted in a table named mcp_user_credentials. A user can bind several accounts to the same server (work and personal GitHub, for example), label them, and mark one as default. An API-key-to-account override pins different tw- keys to different upstream accounts on the same server, so a Cursor key can act as a personal GitHub identity while a CI key acts as a service bot.

Onboarding is protocol-driven. The README describes a one-paste flow: paste an MCP URL, and the probe performs a JSON-RPC initialize to trigger WWW-Authenticate, follows the resource_metadata hint, walks RFC 9728 to RFC 8414 to RFC 7591, fetches authorization server metadata at the path-aware well-known location, and runs Dynamic Client Registration when the upstream advertises it. When DCR is unavailable, the UI presents three steps: copy the callback URL, register the app upstream, paste the Client ID back. The gateway also detects token_endpoint_auth_methods_supported of ["none"], propagates is_public_client, and hides the Client Secret field for issuers such as Feishu that do not use one. For upstreams that only accept static tokens, a vault stores PATs and integration tokens with the same per-user surface, and tokens are verified at paste time.

This is a genuine architectural commitment, and it raises the bar for what your upstreams must support. An MCP server that only accepts a shared token cannot participate in per-user identity, however well the rest of the gateway works.

Installing ThinkWatch and making a first request

The repository does not ship a one-line install for the server. The Cargo workspace has six members (crates/server, crates/gateway, crates/mcp-gateway, crates/auth, crates/common, crates/test-support), the Makefile drives a local development stack, and the deploy directory holds Docker Compose and Helm assets. The .env.example is explicit that you should not write .env by hand: it instructs you to run the secret generator instead, which fills secrets with openssl rand and selects hostnames for the chosen topology.

Start by generating a development environment file. The template states that the --dev mode writes .env with localhost hostnames, while --prod writes .env.production with container names and also emits a ClickHouse user configuration file.

bash
bash deploy/generate-secrets.sh --dev

Next bring up the infrastructure. The Makefile's infra target reads the generated .env and starts PostgreSQL, Redis, ClickHouse, Zitadel and RustFS through the development Compose file. The --remove-orphans flag is there because rustfs_init is intentionally a one-shot container.

bash
docker compose -f deploy/docker-compose.dev.yml --env-file .env up -d --remove-orphans

With infrastructure running, `make dev` starts the backend and frontend detached and returns immediately. The Makefile prints the ports it bound: the gateway on :3000, the console on :3001, and the frontend on :5173. Process IDs and logs land under .dev-run/, and the infra containers are deliberately not touched by dev-stop.

bash
make dev
make dev-status
make dev-logs

Your first real request is a chat completion sent to the gateway instead of the provider. The README lists /v1/chat/completions as a natively served endpoint and states that Cursor, Continue, Cline, Claude Code and the official SDKs can treat the gateway as a drop-in replacement. Point the client's base URL at the gateway and authenticate with a tw- virtual key issued from the console. Because provider credentials live on the gateway, the client never holds the upstream key.

Operational weight and where ThinkWatch is the wrong tool

The dependency list is the first thing to weigh. The README's badges and the development Compose file cover PostgreSQL, Redis, ClickHouse, Zitadel and RustFS, and the .env.example exposes DB_MAX_CONNECTIONS, DATABASE_URL, REDIS_URL, JWT_SECRET and ENCRYPTION_KEY as required configuration. That is a real platform to run, not a sidecar. A team that wants a thin proxy in front of one provider will spend more time on this stack than the proxy saves.

The build profile is also tuned for a production server binary rather than fast iteration. The workspace Cargo.toml sets lto = true, codegen-units = 1 and strip = "symbols", with a comment that this trims roughly 15 to 20 percent off the stripped binary at the cost of a longer link step (about 30 seconds on Apple silicon, several minutes on x86 CI). Contributors should expect slow release builds, and the file notes that operators load symbols from a separate debuginfo package when they need them, so debugging a stripped production binary is a deliberate extra step.

The toolchain floor is rust-version 1.85, pinned to the first stable rustc shipping edition 2024. The comment explains the reason: without it, contributors on older toolchains get cryptic let-chain and edition errors instead of a clear version message. The actual toolchain is tracked in rust-toolchain.toml, currently stable. Anyone building from source on a distribution toolchain older than 1.85 will fail fast, which is the intent.

Where ThinkWatch is the wrong tool: single-provider setups, individual developers, and any environment where you cannot run the database and cache tier. The MCP side has its own boundary. Per-user identity only helps if the upstream issues per-user credentials; if your MCP servers authenticate with one organization token and cannot be changed, the gateway's main differentiator does not apply to you. Finally, the README does not document rollback behaviour for a bad key rotation, so treat rotation grace periods as the safety mechanism and test them before you rely on them.

Compared with a plain reverse proxy or an SDK-level wrapper

The obvious alternative is a plain reverse proxy such as nginx or Envoy in front of the provider APIs. That approach can terminate TLS, hold upstream credentials and forward requests, and it costs almost nothing to operate if you already run one. What it does not do is understand the payload. A generic proxy cannot count tokens per model, cannot apply input_multiplier weighting, cannot convert an Anthropic Messages request into a Bedrock Converse call, and cannot attach a per-user OAuth token to an MCP request. Those capabilities require parsing the protocol, which is what the Rust crates here exist to do.

The second alternative is wrapping the provider SDK inside your own application code. That keeps deployment simple and gives you application-level logging, but every client must be modified, and the wrapper is per-language. ThinkWatch's bet is the opposite: clients keep speaking their native API and the gateway does the work. The README's claim that the gateway is a drop-in replacement for Cursor, Continue, Cline, Claude Code and the official SDKs only holds if that bet is correct, so it is the first thing to test with your own client.

The third alternative is using the provider's own organization and project features. Those give you spend limits and usage reports inside one vendor. They stop at the vendor boundary, which is exactly where a multi-provider organization needs a single view. ThinkWatch's cost tracking is per-model with team attribution, and the README presents that as the answer to bills nobody can explain.

Licence, maintenance and upgrade cost

The workspace Cargo.toml declares license = "BUSL-1.1" for the package, while the repository metadata reports the licence as NOASSERTION and the top level contains LICENSE, LICENSING.md and LICENSING.zh-CN.md. The Business Source License is not an open source licence in the OSI sense, and its terms change over time, so read LICENSING.md and the LICENSE file before you plan a deployment. Nothing here is legal advice; the point is that the licence text, not the badge, governs what you may do.

On maintenance: the repository is not archived, and the last push was on 2026-08-01. The most recent releases are v1.0.1 and v1.0.0, both dated 2026-05-27, while the workspace version in Cargo.toml is 1.0.2, which suggests unreleased work on main. The repository also carries a CHANGELOG.md, a cliff.toml for changelog generation and a renovate.json, so dependency updates and release notes appear to be automated rather than manual. There is no published upgrade procedure in the README; database migrations are handled through sqlx's migrate feature in the workspace dependencies, so upgrading across versions means applying migrations to PostgreSQL and rebuilding the binary. Budget for a maintenance window, not a hot swap.

Editorial conclusion

Adopt ThinkWatch if you already run shared AI credentials across several teams and need per-user attribution plus MCP identity propagation, and you can operate PostgreSQL, Redis and ClickHouse. Do not adopt it if you want a single static API key in front of one provider, or if your MCP upstreams issue only shared service tokens you cannot replace. Verify first that your MCP servers support RFC 9728 discovery or accept static tokens, because per-user identity is the whole point of the MCP gateway and the README does not document a fallback for upstreams that offer neither.

Frequently asked questions

What is ThinkWatch used for?

It is a gateway that sits between AI clients and model providers, proxying OpenAI, Anthropic, Gemini, Azure OpenAI and Bedrock traffic plus MCP tool calls. It adds virtual API keys, rate limits, budgets, audit logging and per-model cost tracking in one control plane.

How do I install ThinkWatch?

The repository does not document a one-line install. The documented path is to run bash deploy/generate-secrets.sh --dev to produce .env, start the development infrastructure with the Compose file, then run make dev to launch the gateway on :3000 and the console on :3001.

Does ThinkWatch work with Cursor and Claude Code?

The README states that the gateway natively serves the OpenAI Chat Completions, Anthropic Messages and OpenAI Responses endpoints on a single port and works as a drop-in replacement for Cursor, Continue, Cline, Claude Code and the official SDKs. Verify this with your own client before rolling it out.

What does ThinkWatch require to run?

The development stack starts PostgreSQL, Redis, ClickHouse, Zitadel and RustFS, and .env.example lists DATABASE_URL, REDIS_URL, JWT_SECRET and ENCRYPTION_KEY among the required variables. The workspace also pins a minimum Rust version of 1.85 for building from source.

Official sources

  1. Issues
  2. Project website
  3. README
  4. Releases
  5. ThinkWatchProject/ThinkWatch on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/thinkwatchproject-thinkwatch.svg)](https://hysenlabs.com/projects/thinkwatchproject-thinkwatch)