gemini-web2api: Using Gemini's Web Interface as an OpenAI-Compatible Local API
Convert Google Gemini web into OpenAI-compatible API. Zero auth, cross-platform, single file.
At a glance
- What is it?
- gemini-web2api is a single Python file that wraps Google Gemini's web interface and serves it as an OpenAI-compatible API at localhost:8081, enabling any OpenAI SDK client or tool to switch to Gemini without an API key or billing setup.
- Who is it for?
- gemini-web2api is a good fit for developers who want to use OpenAI-compatible clients or tools against Gemini's web models without subscribing to the official API. The setup is minimal: a single Python file and one optional dependency.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 50 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 3, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The Problem gemini-web2api Solves
Google's Gemini API requires creating a project, enabling billing, and obtaining an API key. For developers who want to experiment with Gemini models in tools that expect an OpenAI endpoint, or who want to avoid API costs for personal use, the official route involves either paying or being limited to the free tier's rate limits.
gemini-web2api takes a different approach. It intercepts the same web interface that the Gemini web app uses and re-exposes the responses through a local HTTP server that speaks the OpenAI Chat Completions protocol. Any client that works against /v1/chat/completions or /v1/models can point its base URL at localhost:8081 and treat the server as a drop-in replacement for an OpenAI endpoint.
The project targets developers working with desktop AI tools like Cherry Studio or ChatBox, CLI tools that accept an OpenAI base URL, and code that uses the OpenAI Python SDK, all without changing the client code beyond the base URL and an optional API key.
How the Web Interface Interception Works
The server sends requests to the same endpoints that the Gemini web application uses in a browser, constructing the necessary headers and form fields including the XSRF token and session cookies when authentication is needed. For anonymous use, Gemini's web interface permits access to the Flash models without any login, which is what the default configuration uses.
When a client sends a request to /v1/chat/completions, the server maps the model name in the request to a Gemini web model identifier, formats the messages into the web API's expected structure, and returns either a complete JSON response or a server-sent events stream depending on whether the client requested streaming. Tool calls are supported in OpenAI Function Calling format and translated to the web interface's equivalent mechanism.
The Gemini CLI native protocol is also supported through /v1beta/models endpoints, which allows the google/gemini-cli tool to use gemini-web2api as a backend by setting the GOOGLE_GEMINI_BASE_URL environment variable. The Codex CLI Responses API is available at /v1/responses for OpenAI Codex compatibility.
Installing and Running the Server
The only required dependency is httpx, which enables streaming support. Install it and start the server:
pip install httpx
python gemini_web2api.pyThe server starts at http://localhost:8081/v1. To test it with curl:
curl http://localhost:8081/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer sk-your-key" \
-d '{"model":"gemini-3.5-flash","messages":[{"role":"user","content":"Hello!"}]}'For a Docker deployment, build the image and mount a config file:
cp config.example.json config.json
docker build -t gemini-web2api .
docker run -d --name gemini-web2api -p 8081:8081 -v ./config.json:/app/config.json gemini-web2apiOr use Docker Compose:
cp config.example.json config.json
docker compose up -dThe README notes that Docker's default bridge network mode can cause empty responses because Gemini's upstream may reject requests that appear to come from a non-browser context. Switching to host networking resolves this on Linux. Create config.json in the working directory with your preferred settings; by default the server binds on 0.0.0.0 port 8081 with no authentication.
Model Selection and Thinking Depth Control
The server exposes several model names that map to different Gemini web configurations. gemini-3.6-flash and gemini-3.5-flash both route to the all-around Flash model with an output limit of approximately 12,000 characters. gemini-3.5-flash-thinking activates extended thinking mode and produces up to approximately 20,000 characters. gemini-flash-lite is the fastest and lightest option at approximately 10,000 characters output.
For any model that supports thinking, you can control the reasoning depth by appending @think=N to the model name in your request, where 0 is the deepest (default) and 4 is the shallowest. For example:
gemini-3.5-flash-thinking@think=2gemini-3.1-pro requires a Gemini Advanced paid subscription cookie to actually route to the Pro model. Without a valid cookie, requests to gemini-3.1-pro fall back silently to Flash. A free Google account cookie will authenticate the session but will not unlock Pro routing.
Cookie Setup, XSRF Tokens, and What Can Break
Anonymous access works for Flash models. Pro model access requires extracting session cookies from a Gemini Advanced account in Chrome's DevTools and saving them in a file. The required cookie names are SID, HSID, SSID, APISID, SAPISID, and __Secure-1PSID. Pass the cookie file at startup:
python gemini_web2api.py --cookie-file cookie.txtAuthenticated requests may also require the page XSRF token, which appears as SNlM0e in the Gemini web page source. If authenticated requests return HTTP 400 with an XSRF error, refreshing the Gemini web page gives a new token to update in config.json. Google rotates these values, so long-lived deployments need periodic token refreshes.
If the account URL contains an index path such as /u/1/, the auth_user field in config.json must match that index. The README warns that failing to match this value causes request failures.
The project has no GitHub releases. It tracks Gemini's web interface format, which Google can change without notice. The gemini_bl field in config.json holds a build label string that may need updating when Google deploys a new web app version.
Compared with HanaokaYuzu/Gemini-API and the Official SDK
HanaokaYuzu/Gemini-API is a Python library that also provides unofficial programmatic access to Gemini, focusing on a Python API interface rather than an HTTP server. It is better suited to Python scripts that want to call Gemini programmatically without running a local server, while gemini-web2api is better suited to tools that speak HTTP and expect an OpenAI-compatible endpoint.
The official Google Generative AI Python SDK targets the documented Gemini API with proper authentication tokens, predictable rate limits, and guaranteed API stability across model versions. gemini-web2api works without API keys and billing, but its behavior depends on the web interface remaining stable, and it cannot guarantee that model outputs will be the same as what the official API produces for the same input. Google can change the web format at any time.
For production applications, teams should use the official API. For personal tools, local experiments, and integrations where a paid API subscription is not available, gemini-web2api provides a working alternative as long as the operator is prepared for occasional breakage when Google updates its web app.
Editorial conclusion
gemini-web2api is a good fit for developers who want to use OpenAI-compatible clients or tools against Gemini's web models without subscribing to the official API. The setup is minimal: a single Python file and one optional dependency. The main constraint is that Pro model access requires a Gemini Advanced subscription and manual cookie extraction, and the XSRF token may need periodic refreshing when using authenticated routes. If you need a production-grade integration that is stable against Google's web changes, the official Gemini API is the safer path; this tool is better suited to personal use, local experiments, and tool integrations where API billing is not an option.
Frequently asked questions
Does gemini-web2api work without any Google account or API key?
Anonymous access works for all Flash models. No Google account or API key is required for basic use. The Pro model silently falls back to Flash without a Gemini Advanced account cookie.
How do you use gemini-web2api with the Gemini CLI?
Set GEMINI_API_KEY to any non-empty string and GOOGLE_GEMINI_BASE_URL to http://localhost:8081, then run the gemini command. The server handles the native Gemini protocol endpoints at /v1beta/models for both listing models and generating content.
What is the output length limit for gemini-web2api models?
The README lists approximate limits per model: gemini-3.5-flash-thinking produces up to around 20,000 characters, gemini-3.5-flash-thinking-lite around 15,000, most Flash models around 12,000, and gemini-flash-lite around 10,000. These figures come from the README's model table and may change when Google updates the web interface.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/sophomoresty-gemini-web2api)