# gemma4-browser-extension: On-Device AI Agent in Chrome Using Transformers.js and Gemma 4

> gemma4-browser-extension is an open-source Chrome extension by Nico Martin that runs an AI agent entirely on-device via WebGPU and Transformers.js, using the ONNX-format Gemma 4 model from HuggingFace. All inference happens locally: no data is sent to external servers.

**nico-martin/gemma4-browser-extension** — On-device AI agent Chrome extension powered by Transformers.js and Gemma 4

- Repository: https://github.com/nico-martin/gemma4-browser-extension
- Stars: 1,160 · Forks: 193
- Language: TypeScript
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/nico-martin-gemma4-browser-extension

## What the Extension Does and What It Cannot Do

The extension provides an AI agent accessible from a Chrome side panel. The agent understands natural language and uses a set of tools to interact with the browser. The tools fall into three categories: tab management, webpage interaction, and browsing history search.

For tab management, the agent can list open tabs with their titles and URLs, switch to a specific tab, open a new URL in the foreground or background, and close specific tabs. For webpage interaction, the agent can extract structured content from the current page and return the most relevant sections based on a natural language query, and it can highlight and scroll to specific elements on the page. For history search, the agent can search browsing history using semantic natural language queries rather than exact keyword matching.

All of this runs locally. The README states plainly: all processing happens on the device and no data is sent to external servers. The intelligence comes from the Gemma 4 ONNX model and the all-MiniLM-L6-v2 embedding model, both of which are downloaded from HuggingFace on first use and cached locally.

What the extension cannot do: it has no persistent memory across browser sessions beyond the IndexedDB history store, it cannot fill forms or click buttons, and its tab management capabilities are limited to the operations listed above.

## The Three-Component Architecture

The extension is built around three components that each handle a specific role. Understanding why the architecture is split this way matters for developers who want to adapt the code.

The background service worker is the AI engine. It loads the Transformers.js models once when the extension starts and keeps them in memory for the lifetime of the service worker. The README explains the reasoning: loading multi-gigabyte models repeatedly would be impractical, and the background context can handle computationally intensive ML inference without blocking user interactions. The background script processes all inference requests, executes the tools the agent calls, and handles feature extraction.

The side panel is the user interface. It is a React application that displays the chat interface and maintains conversation history across tab switches. Unlike a popup, the side panel stays open while the user browses, preserving context. It communicates with the background script via chrome.runtime.sendMessage and chrome.runtime.onMessage.addListener, so inference requests are non-blocking.

Content scripts run in the context of individual web pages. They have access to the DOM, which neither the background script nor the side panel can reach directly. When the agent calls ask_website or highlight_website_element, the side panel sends a message to the background script, which dispatches a tool call to the content script for the active tab. The content script extracts headings, paragraphs, and lists from the page, generates embeddings using all-MiniLM-L6-v2, and returns the most relevant sections.

## Installing from Source via Developer Mode

The extension is installed as an unpacked extension in Chrome's developer mode. There are no prebuilt releases.

First, clone the repository and install dependencies:

```bash
git clone <repository-url>
cd tfjs-agentgemma-extension
pnpm install
```

Then build the extension:

```bash
pnpm run build
```

For development with automatic rebuilding on changes:

```bash
pnpm run dev
```

After the build completes, load the extension in Chrome:

1. Open chrome://extensions/
2. Enable Developer mode using the toggle in the top right
3. Click Load unpacked
4. Select the dist folder produced by the build

The extension requires Chrome 113 or newer with WebGPU support. On first use, the models download automatically from HuggingFace. This is a one-time download; subsequent launches reuse the cached models.

## The RAG Implementation for Webpage Interaction

The ask_website tool uses Retrieval-Augmented Generation to answer questions about the current page. The implementation has two parts that the README describes in the architecture documentation.

When ask_website is called, the content script extracts structured content from the current page: headings, paragraphs, and lists. This structured extraction is more reliable than a full DOM dump because it skips navigation elements, ads, and boilerplate that would add noise to the embedding search.

The extracted content is then sent to the background script, where the all-MiniLM-L6-v2 embedding model generates vector embeddings for each chunk. The agent's query is also embedded, and cosine similarity between the query embedding and the content embeddings determines which sections are most relevant. Those sections are returned to the agent as context for its response.

The highlight_website_element tool complements ask_website. Once the agent identifies a relevant section, it can instruct the content script to scroll to and visually highlight that section on the page. This closes the loop between finding relevant content and directing the user's attention to it.

The history search tool uses the same embedding approach. The extension stores vector embeddings for page titles, descriptions, and URLs in IndexedDB as the user browses. The find_history tool searches those stored embeddings with a natural language query, with optional time-based filtering.

## Hardware Requirements and Performance Constraints

WebGPU is the non-negotiable requirement. The extension uses WebGPU for all model inference, which means it runs on any GPU that supports WebGPU: Metal on macOS, Vulkan or DX12 on Windows, and Vulkan on Linux. Chrome 113 and newer expose WebGPU as an API. Older Chrome versions and browsers without WebGPU support cannot run the extension.

The Gemma 4 model used is onnx-community/gemma-4-E2B-it-ONNX from HuggingFace. The E2B in the name indicates a 2 billion effective parameter count. This is a small model by current standards, which makes it practical for browser inference. The ONNX format enables efficient execution through the WebGPU backend of Transformers.js without requiring model conversion by the user.

The first launch requires downloading the model, which is multi-gigabyte. The README does not specify the exact size, but ONNX models at this scale are typically in the 1 to 4 gigabyte range. Chrome's storage permissions requested by the extension include storage, used to save settings and model cache.

Inference is slower than a cloud API call. The background context runs inference locally on the user's GPU, and the latency for a typical query depends on the GPU's capabilities and the length of the context. The README does not publish benchmark numbers.

## What This Extension Is Not

The extension is a demonstration and a starting point, not a production-grade assistant. The README title calls it a Transformers.js Gemma 4 Browser Assistant and describes it as demonstrating an effective architecture for integrating Transformers.js into browser extensions. The code is structured to show the pattern clearly rather than to handle edge cases or scale to enterprise use.

There is no server, no authentication, and no synchronization across devices. The history vector database lives in IndexedDB in the user's browser. Clearing browser data clears it. Opening the extension in a different browser or on a different machine starts with an empty history database.

The extension permissions include host_permissions for all URLs, which is necessary for the content script to access page content for RAG. Users who are concerned about the breadth of that permission should read the extension manifest before loading it. The README documents all permissions and explains why each one is needed.

A comparison point is browser-integrated AI features like Google's Gemini in Chrome or Microsoft Copilot in Edge. Those are cloud-backed and send the current page or query to a server. This extension sends nothing off-device, which is the core privacy trade-off in favor of local inference.

## Maintenance and License

The repository is published under the MIT license and the last push was on 2026-08-13. It is not archived, though there are no GitHub releases. The package.json lists version 0.2.1.

The extension uses @huggingface/transformers version 4.2.0 and React 19. The build tooling is Vite 7 with the vite-plugin-web-extension plugin. These are relatively recent versions, which means developers building on this code are working with current toolchain choices rather than legacy dependencies.

The MIT license permits unrestricted use including in commercial products. The HuggingFace models have their own licenses: the Gemma 4 ONNX model is from Google's Gemma family, which has usage terms available on the model card at huggingface.co/onnx-community/gemma-4-E2B-it-ONNX. Teams building commercial products on this extension should review those model terms separately from the extension's own MIT license.

## Conclusion

This extension is a practical starting point for developers who want to build a privacy-preserving browser assistant that keeps all inference local. The architecture is documented in detail and demonstrates patterns that are reusable in other Transformers.js extension projects. The one hard prerequisite is Chrome 113 or newer with WebGPU support; the extension will not work without it. On first use, the models download automatically and are cached, but the Gemma 4 ONNX model is multi-gigabyte, so the first launch on a slow connection takes noticeable time. The repository has no releases, so there is no versioned download: installation requires cloning and loading the unpacked dist folder via Chrome's developer mode.

## FAQ

### Does the Gemma 4 browser extension work offline after first setup?

Once the models have been downloaded and cached on first use, inference runs entirely on-device via WebGPU. The tab management, page RAG, and history search tools all use the cached models and local browser APIs, so no internet connection is needed for those features after the initial model download.

### What browser does the Gemma 4 extension support?

The README specifies Chrome with WebGPU support, requiring Chrome 113 or newer. WebGPU support in other Chromium-based browsers is not mentioned in the README.

### How does the extension search browsing history without sending data to a server?

The extension stores vector embeddings for page titles, descriptions, and URLs in IndexedDB as the user browses. The find_history tool embeds the natural language query locally using all-MiniLM-L6-v2 and performs cosine similarity search against those stored embeddings, with optional time-based filtering. Everything runs in the browser process.

## Sources

- [Issues](https://github.com/nico-martin/gemma4-browser-extension/issues)
- [License: MIT](https://github.com/nico-martin/gemma4-browser-extension/blob/main/LICENSE)
- [nico-martin/gemma4-browser-extension on GitHub](https://github.com/nico-martin/gemma4-browser-extension)
- [README](https://github.com/nico-martin/gemma4-browser-extension/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/nico-martin-gemma4-browser-extension
