# gemini-web2api-go: the Gemini web app behind an OpenAI-shaped endpoint

> This is a protocol reverse proxy, not a wrapper. It speaks the browser protocol of the public Gemini web app and exposes it as an OpenAI-compatible API, which is why it needs a TLS client that impersonates a browser and why its dependency list is mostly that library plus three database drivers.

**zexadev/gemini-web2api-go** — 把 Google Gemini 网页反代成 OpenAI 兼容 API · Reverse Google Gemini's web protocol into an OpenAI-compatible API. Single binary, Chrome 146 fingerprint, SQLite, built-in admin dashboard.

- Repository: https://github.com/zexadev/gemini-web2api-go
- Stars: 478 · Forks: 93
- Language: Go
- License: MIT
- Published: 2026-09-20 · Updated: 2026-09-20 · Language: en
- Canonical page: https://hysenlabs.com/projects/zexadev-gemini-web2api-go

## It is a protocol reverse proxy, and the project says so plainly

The architecture is drawn as a diagram in the documentation and the arrows matter. A client, whether an SDK or one of several tools named as examples, sends an OpenAI-shaped request to a local path. The proxy then talks to the public web app using the protocol that web app's browser uses.

The project is explicit about what this is not. It is not a wrapper around the official generative language endpoint, and the distinction is not cosmetic. The official endpoint is a documented service with keys and a quota. The web app is a different service with a different access model, and a proxy to it inherits the access model of the web app.

The endpoint surface is small. There is a chat completions path, a models list, and a responses path. There is also an asynchronous video path shaped like another vendor's video API, where a request creates a job, a second request polls it, and a third downloads the file, with a note that this one requires a logged-in paid-tier account. The port is set at startup and the whole thing is one process, so the container route is a single run with three settings:

```bash
docker run -d --name gemini-web2api \
  -p 127.0.0.1:8083:8083 \
  -v "$PWD/data:/data" \
  -e ADMIN_TOKEN=your-admin-token \
  ghcr.io/zexadev/gemini-web2api-go:latest
```

Streaming behaviour is described with the same precision. Ordinary chat is genuinely incremental, forwarding each upstream frame as it arrives. Two cases are not: requests carrying tool definitions, and the responses path, both of which collect first and send afterwards. If your client depends on streaming to render, that difference is the one that will surprise you.

## The browser impersonation is the load-bearing dependency

The module file is short, and it explains the design better than the feature list does.

The direct dependencies are a TLS client library, an HTTP transport built for it, a tokenizer, a UUID helper, and three database drivers. Almost everything else in the list is an indirect dependency of the first two, and reading the indirect list tells you what the TLS client drags in: a User-Agent impersonation package, a QUIC variant of it, a WebSocket package, a SOCKS proxy library, Brotli compression, and an elliptic curve field arithmetic package.

The documentation says the impersonation simulates the fingerprint of a recent Chrome release rather than using a library's default handshake. That is the mechanism by which the proxy is not distinguishable from a browser at the transport layer, and everything else in the project is arranged around it.

The rate limiting follows from the same model. Limiting is per egress IP address across three dimensions at once, concurrency, requests per minute, and requests per hour. The startup banner shows the default values. And the proxy pool feature is described as making every proxy its own independent rate-limit slot, so adding an exit address adds capacity rather than sharing a limit.

The cookie pool sits alongside it. Multiple accounts rotate oldest-used-first, each account is pinned to its own exit address, and there is automatic renewal plus a keepalive on a timer. The panel exposes a per-account check and marks expired ones for re-import, which is the operational reality of holding a set of live login states.

## Request content is not stored, and the panel is built around that

The storage claim is stated as a rule rather than a feature: prompt and reply content never enter the database, and only metadata is kept, namely length, duration, model, and status.

That constraint shapes the whole admin panel, which is a useful way to see what the author thought a self-hosted proxy operator needs. The overview has rolling key figures, a dual-axis trend of request volume against median latency, grouped statistics by model and by proxy, a view of how much of the per-address rate limit has been consumed, and a one-click connectivity diagnostic.

The request records view is a metadata list with filtering by status and by model and pagination. It is a log you can search without having decided in advance that you want to keep everyone's prompts.

Retention is also two-tier. Detailed request rows are kept for thirty days and aggregate statistics are kept indefinitely, with a retention setting for the detail window. If you point this at a shared gateway, that is the line between what a colleague can see and what they cannot.

The credentials are managed in the panel and stored in the database, and the configuration split is stated in the container file. Tuning parameters, the rate limit allowances, retries, timeouts, default model, fingerprint selection, retention days, and feature switches all live in the panel and take effect on save. The environment variables are described as optional locks, for deployments where you do not want runtime changes to be possible at all.

## A named volume, not a bind mount, and the error message explains why

The container file contains one of the more useful comments in the project, and it is about a bug that will otherwise cost you an afternoon.

The volume section uses a named volume rather than mounting a host directory, and the comment explains that the image runs as a non-root user with a specific numeric identity. A bind mount would take ownership from the host directory, which does not exist before the first run, which Docker then creates owned by root, which the container then cannot write to. The symptom is a database open error with a numeric code.

A named volume inherits ownership from the image, so the container can write. And if you insist on a bind mount, the comment gives the exact command to fix the ownership first.

The port mapping is equally deliberate. It binds to the loopback address on the host rather than to all interfaces, with a comment showing how to expose it to other containers on the same network without publishing it. The admin token has a placeholder default, and the comment on that line says an empty value leaves the panel unauthenticated, which is only acceptable while bound to loopback.

Two more details worth copying. The health check runs the binary's own version flag rather than an HTTP probe, so it verifies the process can start rather than that it can serve. And the default command line is supplied in the image, with a warning that overriding it replaces the whole thing rather than adding to it, including the database path.

## The build stage stays on the host architecture on purpose

The container build has a comment that explains a decision most projects get wrong by accident.

The builder stage is pinned to the build machine's own platform rather than the target platform. The reason given is that on a multi-architecture build, a builder that follows the target platform compiles the arm64 image under emulation, and since compiling is CPU intensive, that is an order of magnitude slower. Pinning the builder to the host architecture and setting the target architecture as a build argument means the cross-compiler does the work natively.

The prepare stage is pinned the same way for a different reason. Its whole job is creating two directories and changing ownership to a numeric user and group. That is architecture independent, so there is no reason to pull a target-architecture base image for it and pay for another emulation round.

The runtime stage is a distroless static image, described in the comment as a couple of megabytes with no shell and no package manager, which is a reasonable minimum attack surface for a service that holds login states. It runs as the non-root user, and the directories were pre-created in the prepare stage precisely because distroless has no directory-creation command.

The build itself disables cgo, with the reason that the SQLite driver used is pure Go. That is what makes a static binary possible at all, and it is why the same code cross-compiles for six platforms from one machine. Symbol stripping is enabled for a stated size saving of about five megabytes, with path trimming alongside it.

## The search tool is a second interface on the same port

The same process also mounts a protocol server on a separate path, and it exposes exactly one tool.

The tool takes a query and returns an answer that the web app compiled with its own web search, followed by a list of source links. It works anonymously, so a cookie is not required. The documentation is clear that this is the only tool, and that the tool for fetching a specific URL was not built.

Because it is on the same port and in the same process, there is no second deployment. The client configuration is a URL and an authorization header carrying the same API key as the chat interface. The transport is a streaming HTTP one, so a remote client connects by URL and inherits the account pool, the proxy pool, and the rate limiting.

That last point is the interesting design choice. A search tool that runs through the same credential pool as the chat path means the rate limits and the account rotation apply to searching as well, and it means there is one thing to monitor rather than two.

The health endpoint is worth noting separately because it is unauthenticated. It returns a status, a version, and the model list, which makes it usable as a liveness probe by anything that expects a plain HTTP check.

## One client cannot connect, and the documentation names it

The client integration section is mostly a single table: a base URL, an API key from the panel, and a model name. Several clients are named as compatible because they speak the OpenAI shape.

The exceptions are the interesting part, and there are two.

One is a command-line agent that uses the responses path rather than chat completions, and the documentation confirms that path is implemented while repeating that it is not incremental streaming. So the client works, with the caveat from the streaming section.

The other is named specifically and ruled out. A different command-line client expects the vendor's native request path, a versioned model-and-method shape, and this project exposes only the OpenAI-shaped surface with no native path implemented. The documentation says so directly rather than leaving you to discover the 404.

For deployments that fan out through a gateway, there are instructions for creating a channel: pick the OpenAI type, point the base URL at the proxy, and list the models you want. There is a note about what to use when the gateway is itself in a container, which is the case that catches people out because the address changes.

The last item in the section is a warning rather than a feature. The project reverse-proxies a consumer web service, and the terms governing that service are not the terms of an API agreement. Nothing in the documentation resolves that, which is the first thing to settle before you depend on it.

## Conclusion

This suits someone who already has a Google account and wants a free tier behind an OpenAI-shaped endpoint for local tooling, and who is willing to run a proxy that depends on an undocumented web protocol. It is a poor fit for anything you cannot fix yourself when Google changes that protocol, and a poor fit for production traffic, since a reverse proxy of a consumer web app has no availability guarantee. Before you run it, check three things: that you understand what you are depending on, because this is not Google's documented interface and the terms of that web service are separate from any API agreement; whether you need the model that requires a cookie, since the faster ones work anonymously and the pro one does not; and that the admin token is set, because the container file shows the default value and an empty token means the panel has no authentication. The newest release is v4.20.1 from 2026-09-15 and the last push was 2026-09-29.

## FAQ

### What is gemini-web2api-go and is it an official Google API?

It is a self-hosted proxy that turns the public Gemini web app into an OpenAI-compatible chat completions endpoint. The documentation is explicit that it is not a wrapper around the official generative language endpoint, and that it reverse-proxies the browser protocol instead, which is why it needs no Google API key and no paid quota.

### Does gemini-web2api-go store my prompts and replies?

No. The stated rule is that prompt and reply content never enter the database and only metadata is kept, namely length, duration, model, and status. Detailed request rows are kept for thirty days and aggregate statistics indefinitely, and the admin panel's request view is built around metadata only.

### What happens if I use a bind mount for the gemini-web2api-go data directory?

The container cannot write to it and fails opening the database. The image runs as a non-root user, and a bind mount takes ownership from the host directory, which Docker creates owned by root when it does not already exist. The project uses a named volume instead, and gives the ownership command if you want a bind mount.

### Can I use the Gemini CLI with gemini-web2api-go?

No. That client expects the vendor's native request path with a versioned model-and-method shape, and this project exposes only the OpenAI-shaped surface with no native path implemented. The documentation names the client specifically rather than leaving you to find the failure yourself.

### What does gemini-web2api-go need to serve its own web search tool?

Nothing beyond running the backend. The protocol server is mounted in the same process on the same port and exposes one tool that returns an answer plus source links, and it works anonymously without a cookie. Remote clients connect by URL over a streaming HTTP transport and reuse the same account pool and rate limits.

## Sources

- [Issues](https://github.com/zexadev/gemini-web2api-go/issues)
- [License: MIT](https://github.com/zexadev/gemini-web2api-go/blob/main/LICENSE)
- [README](https://github.com/zexadev/gemini-web2api-go/blob/main/README.md)
- [Releases](https://github.com/zexadev/gemini-web2api-go/releases)
- [zexadev/gemini-web2api-go on GitHub](https://github.com/zexadev/gemini-web2api-go)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/zexadev-gemini-web2api-go
