Model or dataset
foxhui/WebAI2API avatar
foxhui/WebAI2API

WebAI2API ships a placeholder auth token and an endpoint that exports cookies

WebAI2API: 基于 Camoufox 的网页 AI 转 API 工具,支持 LMArena/Gemini等,多窗口并发与账号隔离。 | Web AI to OpenAI API via Camoufox. Supports LMArena/Gemini and more, multi-window concurrency & account isolation.

1,372 stars358 forksJavaScriptMIT

At a glance

What is it?
A browser-driving proxy that turns consumer web AI sites into an OpenAI shaped API. The documentation here is the Chinese README, and it is candid about what the tool does to stay undetected.
Who is it for?
Read the terms of service of every site you intend to proxy before anything else, because the project itself puts that first in its own disclaimer and its stated purpose is to look like a person rather than like a client. If the answer is no, stop here.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 86 days ago.
What is it written in?
Mainly JavaScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 2, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The example config ships sk-change-me-to-your-secure-key

On first run the program copies config.example.yaml to data/config.yaml, and the server section of that example reads:

yaml
server:
  # 监听端口
  port: 3000
  # 鉴权 API Token (可使用 npm run genkey 生成)
  # 该配置会对 API 接口和 WebUI 生效
  auth: sk-change-me-to-your-secure-key

One token guards both the API and the web management console, so anyone holding it has the API and the panel. A placeholder like that is normal in an example file, and the fix is documented in the same file, since `npm run genkey` generates a key. What the file does not do is refuse to start on the default.

Two other details sit in the same block. Configuration changes only take effect after a restart, so a rotated key means a restart before it counts. And the documentation notes that a remote user reaches the console by replacing localhost with the server address, which means the default posture is reachable from the network.

/v1/cookies exports live session credentials

Three endpoints are documented. POST /v1/chat/completions carries the traffic, GET /v1/models lists what is configured. The third is GET /v1/cookies, which exists so other tools can reuse the login this project already maintains:

bash
curl "http://localhost:3000/v1/cookies?name=browser_default&domain=lmarena.ai" \
  -H "Authorization: Bearer YOUR_API_KEY"

Those cookies are real sessions on the target site, not scoped tokens for this API. Automatic session renewal is a listed feature, which means the exported value is current rather than stale.

The endpoint takes an optional instance name defaulting to `default` and an optional domain filter, so a caller can pull cookies for one browser profile and one site. That is a useful feature for a person operating their own accounts and an attractive target for anyone who obtains the API token. It has no separate permission level in the configuration shown.

The console is plain text and the image turns on VNC by default

The deployment section carries a security warning with three parts. The Docker image enables a virtual display with Xvfb and a VNC service by default. That VNC server can be reached from the console's virtual display panel. And the console's transport is not encrypted, so a public environment should use an SSH tunnel or HTTPS.

The image exposes two ports, 3000 and 5900, where 5900 is the VNC port. The compose file publishes only 3000, and so does the documented docker run command, so the default compose path keeps VNC inside the container while still leaving it enabled.

Both files also set a two gigabyte shared memory limit, and the compose file adds init: true with restart unless-stopped. The recommended remote access path is a tunnel:

bash
ssh -L 3000:127.0.0.1:3000 root@服务器IP

The dependency list is the anti-detection toolchain

The stated main feature is human-like interaction: simulated typing and mouse trajectories, with feature disguise for getting past automation detection. The dependencies make that concrete. camoufox-js drives the browser, fingerprint-generator produces browser fingerprints, ghost-cursor-playwright-port generates mouse movement, and proxy-chain, https-proxy-agent and socks-proxy-agent carry the traffic.

Two more install-time behaviours are worth knowing. The repository has a patches/ directory, which for a pnpm project means edited third-party packages, and the package manifest runs a postinstall script.

The mode documentation is unusually blunt about the consequence. Headed mode is the default, headless mode saves resources but may be detected by the sites being proxied, and the guidance is to keep headed mode, or a virtual display, running long term to reduce the chance of a risk trigger. A login mode flag temporarily forces headed operation and disables automation, which is what an interactive sign in needs.

Streaming is the only path that queues without limit

The concurrency behaviour is the part that will shape your client code. Requests are executed as real browser operations, so timing varies. When the backlog of queued tasks exceeds the configured number, non-streaming requests are rejected outright rather than queued.

Streaming mode sends heartbeat packets and, in the project's words, can queue indefinitely to avoid a timeout. Two keepalive styles are selectable in configuration. Comment mode is the default and sends an SSE `:keepalive` comment, which is the standard mechanism and the most compatible. Content mode sends empty content data packets, reserved for clients that will only reset a timeout when they receive JSON.

So the same request, differing only in the stream flag, changes from fail fast to wait as long as you like. The image API has its own limits: up to ten pictures, PNG, JPEG, GIF or WebP, required as Base64 data URLs, and the server converts everything to JPG on the way in.

Node 20 is documented but not declared anywhere in the manifest

The environment section asks for Node.js v20.0.0 or later with ABI 115 or newer. The package manifest declares no engines field at all, so the floor exists as prose and not as something npm enforces.

The container runs a different version again. The Dockerfile is built on node:22-bookworm and sets PLAYWRIGHT_SKIP_BROWSER_DOWNLOAD before running the init script, which is what fetches the browser. So the documented floor is 20, the image is 22, and the manifest is silent.

The manifest also describes the package as an automated image generation tool built on Playwright and Camoufox, which is narrower than the README, where the same tool covers text, image and video across eleven listed sites. It declares version 3.0.0 with no published release attached, and a CHANGELOG.md that holds the history instead.

Four symbols divide supported from impossible, and one table misuses them

The support table has three capability columns across eleven sites, and the legend defines four marks. A tick means supported. A cross means not supported now but possible later. A barred circle means the site itself does not offer the capability. A droplet marks results that carry a watermark which cannot be removed.

Two sites are marked with the droplet: Google Gemini for images and video, and Sora for video. Gemini Enterprise Business is the only entry with a tick in all three columns, and Sora is the only other one with video at all.

The two negative marks are not used consistently. Google Flow has a barred circle for text and a cross for images, which are defined as different things, so the table is not telling a reader whether Flow might add images later. The last row of the table is a placeholder reading that the list continues. The documentation points at GET /v1/models as the authoritative source of what is actually configured, which is the right instinct for a table this ambiguous.

Editorial conclusion

Read the terms of service of every site you intend to proxy before anything else, because the project itself puts that first in its own disclaimer and its stated purpose is to look like a person rather than like a client. If the answer is no, stop here. If the answer is yes, the engineering is worth reading: a browser-driven OpenAI compatible surface, per instance proxying and account isolation, and a table that honestly marks which sites cannot do what. Two things to fix before it faces a network you do not control. The example configuration ships a placeholder token that covers both the API and the web console, and the web console transmits in plain text, so bind it to localhost, tunnel it, or put TLS in front of it. Then decide whether you want /v1/cookies switched off, because that endpoint hands out live session credentials.

Frequently asked questions

What does WebAI2API do?

It drives a real browser through Camoufox and Playwright to talk to consumer web AI sites such as LMArena, Gemini, zAI, ChatGPT and DeepSeek, then exposes those conversations through an OpenAI shaped API. Accounts are isolated per browser instance, and each instance can use its own proxy.

How do I secure the WebAI2API web console?

Change the auth token in data/config.yaml, which can be generated with npm run genkey, and remember that the token covers both the API and the console. The console's transport is not encrypted, so bind it to localhost, tunnel it over SSH, or put HTTPS in front of it before exposing it.

Does WebAI2API work without a display server?

It can, but not in plain headless mode. The documentation states headless mode may be detected by the sites being proxied and recommends headed mode or a virtual display such as Xvfb for long term running. Both the docker run command and the compose file set a two gigabyte shared memory limit for this reason.

Why do WebAI2API requests get rejected?

Non-streaming requests are rejected outright once the number of queued tasks exceeds the configured threshold, because each request runs as real browser work with variable timing. Streaming mode sends heartbeat packets and can queue indefinitely, using SSE keepalive comments by default or empty data packets in content mode.

Which sites does WebAI2API support?

The table covers LMArena, Gemini Enterprise Business, Nano Banana Free, zAI, Google Gemini, ZenMux, ChatGPT, DeepSeek, Sora, Google Flow and Doubao, across text, image and video. Gemini Enterprise Business is marked supported for all three, and Gemini and Sora results carry a watermark that cannot be removed. GET /v1/models is the authoritative list for a given configuration.

Official sources

  1. foxhui/WebAI2API on GitHub
  2. Issues
  3. License: MIT
  4. Project website
  5. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/foxhui-webai2api.svg)](https://hysenlabs.com/projects/foxhui-webai2api)