# AIStudioToAPI: putting a browser automation shim in front of Google AI Studio

> An Express service that logs into Google AI Studio with a real browser, then exposes the resulting Gemini models through OpenAI, Gemini and Anthropic shaped HTTP endpoints, complete with a web console and VNC account management.

**iBUHub/AIStudioToAPI** — A wrapper that exposes Google AI Studio Build as OpenAI, Gemini, and Anthropic compatible APIs.（一个将 Google AI Studio Build 封装为兼容 OpenAI / Gemini / Anthropic 风格 API 的工具）

- Repository: https://github.com/iBUHub/AIStudioToAPI
- Stars: 1,705 · Forks: 288
- Language: JavaScript
- License: NOASSERTION
- Published: 2026-10-06 · Updated: 2026-10-06 · Language: en
- Canonical page: https://hysenlabs.com/projects/ibuhub-aistudiotoapi

## What the service actually does under the hood

The description is unusual enough to be worth unpacking. AIStudioToAPI does not call a Google API. It drives a browser, signs into the AI Studio Build web application as a person would, and translates your HTTP request into interactions with that page. The service then translates the results back out into whatever shape the calling client expects.

That design is visible in the dependency list of the root package.json. Playwright is pinned at 1.59.1 as a runtime dependency, not a development one, and the Dockerfile installs a long list of X11 and browser libraries alongside xvfb and x11vnc so a browser can run headless but still be watched. The browser itself is Camoufox, a privacy-focused Firefox fork, which the setup script downloads on first run.

The practical consequence is that request latency, model availability and feature support depend on a web application you do not control. When Google changes the AI Studio interface, this project has to change with it. The repository was last pushed on 2026-09-19 and the most recent release, v1.3.7, shipped on 2026-09-03, so the adaptation work is still happening.

## Getting credentials in is the hard part of the setup

There are two installation paths and they are not equivalent. Running directly on Windows, macOS or Linux means cloning the repository and executing the auth setup script:

```bash
git clone https://github.com/iBUHub/AIStudioToAPI.git
cd AIStudioToAPI
```

```bash
npm run setup-auth
```

That script downloads the Camoufox browser, opens it at AI Studio, and waits while you complete the Google login by hand. The credentials it saves land in a directory named configs/auth as files called auth-N.json, with N counting from zero.

The Docker path deliberately skips that step, because the container ships a graphical environment. You start the container and then use the web console: press the add account button, land on a VNC page showing the live browser, log into Google, press save, and the account file appears. The console can also download an existing auth file from one container and upload it into another.

One asymmetry is stated plainly and it matters for planning. Adding accounts through VNC is not supported when you run directly on your own machine; the script is the only route there, and VNC login works only inside the Docker container. Credentials also cannot be injected through environment variables at the time of writing.

## One server, three request shapes

The API surface is the part worth mapping, because the whole point is that a client written against OpenAI can talk to it unchanged. On port 7860 the OpenAI compatible group handles GET /v1/models to list models, POST /v1/chat/completions for chat and image generation, POST /v1/embeddings, and POST /v1/responses, which is the newer Responses API shape. A separate endpoint, /v1/responses/input_tokens, counts input tokens for that API.

The Gemini native group is the closest to the underlying reality. It serves GET /v1beta/models and the standard generateContent, streamGenerateContent, embedContent and batchEmbedContents routes, plus a predict endpoint for Imagen image generation.

The Anthropic compatible group is the smallest: GET /v1/models and POST /v1/messages, with /v1/messages/count_tokens for token counting.

Two details recur across all three groups. Streaming comes in real and fake variants, which is useful when a client insists on a stream that the underlying page interaction cannot genuinely provide. And both the OpenAI chat endpoint and the Gemini generateContent endpoint handle images and speech as well as text, which means text-to-speech and image generation are reachable through the same port.

## The Docker command that starts the whole thing

The documented container invocation is short enough to read as a unit:

```bash
docker run -d \
  --name aistudio-to-api \
  -p 7860:7860 \
  -v /path/to/auth:/app/configs/auth \
  -v /path/to/data:/app/data \
  -e API_KEYS=your-api-key-1,your-api-key-2 \
  -e TZ=Asia/Shanghai \
  --restart unless-stopped \
  ghcr.io/ibuhub/aistudio-to-api:latest
```

The two volume mounts are the operational core. The first persists the Google auth files so a container restart does not force you to log in again, and the second persists request statistics, which are written as JSON lines to a usage-stats.jsonl file inside the data directory. API_KEYS is a comma-separated list used to authenticate incoming client requests, and it is separate from your Google credentials.

The environment file template is more revealing about the security posture than the README is. There are optional WEB_CONSOLE_USERNAME and WEB_CONSOLE_PASSWORD settings for the console, and if neither is set the system falls back to using API_KEYS to log into the console itself. Secure cookies default to off so that a plain HTTP setup works, login rate limiting is on by default at five failed attempts inside a fifteen minute window, and there is an INITIAL_AUTH_INDEX setting for choosing which stored account to start with.

One deprecation notice is worth noting because it shows the project's direction. The WebSocket port environment variable is no longer supported; the port is fixed at 9998 and cannot be customised.

## Where the output formats and the code generators live

This is not a proxy to a paid API, and the pricing question has a specific answer. The service itself is free software, but it consumes Google AI Studio sessions, so what it costs depends on your account rather than on a price list here. Model additions show up in the release notes as small commits: v1.3.7 added a Gemini 3.8 Flash model, v1.3.6 added Gemini 3.7 Flash and a Robotics-ER 2 preview, and v1.3.5 added 3.5 Flash Lite and 3.6 Flash.

The repository tree shows a normal Node service layout with src/, ui/, scripts/, configs/ and docs/, plus a vite.config.js and eslint, stylelint and prettier configuration. The ui directory is a Vite frontend, which is why the start script has a prestart step that builds it, and why there is a separate quick-start script for when the assets are already built and you only want a fast restart.

For a graphical front end, the README points at AMC WebUI, a local-first workflow UI that already knows how to treat this service as a third-party Gemini compatible backend. You point it at the /v1beta path, enter a key matching API_KEYS, and it will talk to the container. There are also deployment notes for reverse proxying through Nginx, with the full configuration in a separate document under docs.

Two hosting platforms get explicit retirement notices. Claw Cloud Run stopped its product and related services as of 2026-05-11, and Zeabur stopped new projects on shared clusters as of 2026-03-15, though services already running there are unaffected. Both tutorials remain in the repository as historical reference.

## What to weigh before you deploy this

The honest framing is that this is a browser session wearing an API costume. That buys you multi-account support and access to whatever the web UI exposes, including image and speech models. It costs you the guarantees a real API gives: stable request semantics, predictable latency, and a contract Google will not change under you.

It also moves your Google authentication into a file on a mounted volume. That is the mechanism the entire service depends on, and it is worth being clear-eyed about. Whoever can read the auth directory can act as you in AI Studio. The project's own documentation does not attempt to hide this, and there is no environment variable route that would let you keep credentials out of the container.

For local use the shape is small: Express on 7860, a Vite-built console, a browser running under Xvfb, and a WebSocket on 9998 for internal browser communication. The Node image is node:24-slim, and the Dockerfile installs VNC and browser dependencies in a single stage with the UI built in.

If your requirement is an OpenAI-shaped client talking to a Gemini model, this works and the documentation is thorough about the endpoints. If your requirement is a production dependency with a support contract, the browser underneath is the part to think about.

## Conclusion

This project solves a specific problem that official Gemini endpoints do not: it uses a web session rather than an API key, and multi-account support is built in from the start. If you need that, the Docker route is the documented one, because VNC account login only works inside the container. If you do not need it, you are adding a browser, a login flow and a translation layer between your client and Google, and an official key would be simpler and more predictable. Before deploying, read the account management section and decide whether storing Google auth files on a mounted volume is something you want in your threat model, since that is the mechanism the whole design rests on.

## FAQ

### Which API formats can AIStudioToAPI serve at the same time?

Three. It exposes OpenAI compatible endpoints such as /v1/chat/completions and /v1/embeddings, Gemini native endpoints under /v1beta, and an Anthropic compatible pair at /v1/models and /v1/messages. All of them run on port 7860 from the same process.

### How do I add a Google account when running outside Docker?

Run the npm run setup-auth script after cloning. It downloads the Camoufox browser, opens AI Studio, and walks you through the Google login interactively, saving the result as auth-N.json files under configs/auth. The VNC account flow documented in the README only works inside the container.

### Does AIStudioToAPI need an API key from me?

It does not use a Google API key. It authenticates by driving a logged-in browser session, which is why there is an account setup step at all. The API_KEYS setting is unrelated to Google: those are the keys your own clients must present to this service.

### Can I point AMC WebUI at this service?

Yes. The README describes AMC WebUI as an existing front end that already supports this project as a third-party Gemini compatible backend. Configure it with the /v1beta base address and an API key that matches one of the values in API_KEYS.

### What license does the project use?

A LICENSE file is present at the root of the repository, but the licence identifier is not stated in the project's own description, so it is worth reading that file directly rather than assuming a standard permissive licence.

## Sources

- [iBUHub/AIStudioToAPI on GitHub](https://github.com/iBUHub/AIStudioToAPI)
- [Issues](https://github.com/iBUHub/AIStudioToAPI/issues)
- [README](https://github.com/iBUHub/AIStudioToAPI/blob/main/README.md)
- [Releases](https://github.com/iBUHub/AIStudioToAPI/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/ibuhub-aistudiotoapi
