gemini-skill: Automate Gemini's Web Interface via CDP as an MCP Server
gemini drawing MCP & skill through browser, can be used in openclaw or any agent that supports MCP. Gemini画图 MCP和sill,支持龙虾或任何agent使用٩(๑>◡<๑)۶
At a glance
- What is it?
- gemini-skill is a Node.js project that drives the Gemini web application through the Chrome DevTools Protocol and exposes each operation, image generation, text chat, image upload, and session management, as a standard MCP tool callable from any MCP-compatible agent. It works by automating a real browser session rather than calling Gemini's official API.
- Who is it for?
- gemini-skill suits developers who want to add Gemini image generation to an MCP agent workflow without an official Gemini API key, and who are willing to maintain a Chrome session tied to a Google account. It is a poor choice for production services that require stable, versioned API contracts, since the entire mechanism depends on the Gemini web app's HTML selectors and UI structure, which the project has already needed to update twice in 2026 for new Gemini UI releases.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 12 days ago.
- What is it written in?
- Mainly JavaScript, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 28, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What gemini-skill Does and Who It Is For
gemini-skill targets developers building agentic workflows that need Gemini's image generation capabilities but do not have or want to use Google's official Gemini API. Instead of calling a REST endpoint, it opens a Chrome browser, logs into the Gemini web app under a real Google account, and issues commands through the Chrome DevTools Protocol. The result is passed back to the calling agent as an MCP tool response.
The primary users are developers working with MCP-compatible agents such as Claude Code, OpenClaw, or any other client that can connect to an MCP server over stdio JSON-RPC. The project also includes a SKILL.md file in the repository, which is a skill definition for agents that consume skill formats, such as Claude Code's skill system.
Daemon Architecture and the CDP Automation Stack
The README describes a layered architecture with distinct responsibilities. At the top, an MCP client (the calling agent) communicates with mcp-server.js over stdio using JSON-RPC. mcp-server.js handles the MCP protocol and routes each tool call to gemini-ops.js, which implements the actual Gemini operations such as sending a prompt, waiting for image generation to finish, or extracting the output image. gemini-ops.js connects to the browser through browser.js, which manages the CDP connection. A separate daemon process in the daemon/ directory handles the browser process lifecycle.
The README states the daemon keeps the browser running after an MCP tool call completes, releasing it only after 30 minutes of inactivity. A fresh MCP tool call will start the daemon automatically if it is not running. This on-demand startup means the first call in a cold session takes longer than subsequent ones.
The stealth layer is explicit: the project uses puppeteer-extra-plugin-stealth to hide the webdriver flag and simulate a real browser fingerprint. The README notes that reusing OpenClaw's browser session (by setting BROWSER_DEBUG_PORT=18800) loses this protection, and recommends keeping the default port of 40821 to let gemini-skill manage its own browser instance.
Installing gemini-skill and Configuring the Environment
Prerequisites are Node.js 18 or later and a Chrome, Edge, or Chromium browser installed on the system. A Google account must be logged in to the browser before the first use.
Clone the repository and install dependencies:
git clone https://github.com/WJZ-P/gemini-skill.git
cd gemini-skill
npm installCopy the .env.example file to .env or .env.development (the latter is git-ignored and suited for local secrets) and set the relevant variables. The key browser variables are:
BROWSER_DEBUG_PORT=40821
BROWSER_HEADLESS=false
BROWSER_PROTOCOL_TIMEOUT=60000
OUTPUT_DIR=./gemini-imageThe README recommends running with BROWSER_HEADLESS=false on first use so the browser window is visible for the Google account login. Once the session is persisted in the browser user data directory, headless mode can be enabled for subsequent runs.
The MCP server is started with:
npm run mcpwhich executes node src/mcp-server.js. This command starts the server over stdio, ready to accept JSON-RPC calls from any MCP-compatible client.
Image Generation, Watermark Removal, and Session Management Tools
The README lists the MCP tool set as: AI image generation (send a text prompt and receive the generated image with automatic watermark removal), multi-turn text conversation, image upload for reference-based generation, image extraction from the active session (supporting base64 and full-size CDP download), session management (new session, temporary session, model switching, navigation to history sessions), and, as of v1.1.1, thinking depth control via gemini_set_thinking_depth with standard and extended options.
The automatic watermark removal is a specific processing step applied to downloaded images. The README describes it as removing the Gemini watermark from downloaded output. The mechanism is not documented in the README beyond naming it as a feature.
The model options in v1.1.1 are pro, flash, and flash-lite, replacing the older pro, quick, and think identifiers. The README notes that the UI selectors for model switching were updated to use label text matching rather than fixed selectors, which makes the switch more tolerant of Gemini's internationalized interface.
Atlas Cloud Provider as a Secondary MCP Path
The repository also bundles a minimal Atlas Cloud provider that adds an OpenAI-compatible API path alongside the browser automation main path. The .env.example shows three relevant variables: ATLAS_BASE_URL (default https://api.atlascloud.ai/v1), ATLAS_API_KEY (empty, to be set locally), and ATLAS_MODEL (default openai/gpt-4o-mini).
The README describes the Atlas Cloud integration as a non-invasive addition: it does not alter the Gemini browser automation path, adds an OpenAI-compatible endpoint, and supports model enumeration, standard chat, and streaming chat. The project positions this as a sample for using gemini-skill as a combined Gemini automation tool and a model-as-a-service provider entry point.
This secondary path depends on an Atlas Cloud API key that the user must obtain separately. The README does not document what happens when the key is absent; from the .env.example the ATLAS_API_KEY defaults to empty.
Fragility of UI Automation and Where This Approach Fails
The fundamental risk of driving a web interface through CDP is that every UI change in Gemini can break the automation. The README records two updates in May 2026 alone (v1.1.0 and v1.1.1), each required to adapt to a new Gemini UI release. Selectors like gem-icon-button[arialabel="upload and tools"] and download-generated-image-button are specific to the DOM structure of a particular Gemini release; Google can rename or restructure these at any time without notice.
The official Gemini API from Google takes the opposite approach: it is a versioned REST API with documented endpoints, structured request and response schemas, and explicit deprecation policies. A production service that calls the official Gemini API does not depend on the browser UI staying stable. gemini-skill is better suited to personal automation workflows, proof-of-concept agents, and cases where using the web app is a deliberate choice (for example, to use a free-tier account without API billing), rather than for services where uptime and reliability matter.
There is also a login dependency that the Gemini API does not have: the browser must hold an active Google account session. If the session expires or is invalidated, automation stops until it is renewed manually.
Maintenance Status and License
The last push to the repository was on 2026-08-01, less than two months before 2026-09-28. The project reached version 1.1.1, and the README documents two UI compatibility updates shipped in May 2026, which shows the project is tracking Gemini UI changes actively. There are no GitHub releases; versioning is tracked in the README's update section.
The package.json declares the license as ISC. The repository also contains a LICENSE file. ISC is a permissive license functionally equivalent to MIT, allowing use, modification, and redistribution with attribution. There is no contributor license agreement documented in the repository.
Editorial conclusion
gemini-skill suits developers who want to add Gemini image generation to an MCP agent workflow without an official Gemini API key, and who are willing to maintain a Chrome session tied to a Google account. It is a poor choice for production services that require stable, versioned API contracts, since the entire mechanism depends on the Gemini web app's HTML selectors and UI structure, which the project has already needed to update twice in 2026 for new Gemini UI releases. Before adopting it, verify that the target environment can run a Chrome instance and has a Google account available for the persistent browser session.
Frequently asked questions
How does gemini-skill compare to a Claude Code skill?
gemini-skill is an MCP server that controls the Gemini web app via CDP and exposes image generation and chat as MCP tools. The repository also includes a SKILL.md file, which is a skill definition that tells an agent such as Claude Code how to invoke those MCP tools. The two are complementary: gemini-skill is the backend that drives Gemini, while SKILL.md is the agent-side description of how to use it.
Does gemini-skill require a paid Gemini subscription?
The README requires only a Google account logged into the browser. It does not document any paid tier requirement. The tool drives the Gemini web interface through browser automation, so it uses whatever account tier is logged in to the browser session.
Can gemini-skill run in headless mode on a server?
Yes. The BROWSER_HEADLESS environment variable can be set to true for headless operation. The README recommends running with headless false on first use so the browser window is visible for the Google account login step. Once the session is stored in the browser user data directory, headless mode is usable for subsequent runs.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/wjz-p-gemini-skill)