Model or dataset
kessler/gemma-gem avatar
kessler/gemma-gem

Gemma Gem: a Chrome extension that runs Gemma 4 on-device through WebGPU

Gemma Gem runs Google's Gemma 4 model entirely on-device via WebGPU — no API keys, no cloud, no data leaving your machine.

970 stars105 forksTypeScriptApache-2.0

At a glance

What is it?
Gemma Gem is a TypeScript Chrome extension that hosts Google's Gemma 4 model in an offscreen document and drives an agent loop over the current page. The useful part is the tool split; the hard part is the hardware.
Who is it for?
Adopt Gemma Gem if you want to see a browser agent run against a real page without sending page content to a hosted API, and you have a Chrome or Edge build with WebGPU and at least 4 GB of GPU-addressable memory for E2B. Do not adopt it if your target machines are integrated-GPU laptops with 8 GB of RAM, if you need Firefox support (pnpm build:firefox exists in package.json but the README documents Chrome only), or if you need a published release to pin.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 125 days ago.
What is it written in?
Mainly TypeScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 26, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What Gemma Gem does that a hosted chat sidebar cannot

The pitch is narrow and concrete: a model that reads and acts on the page you are already looking at, with no API key and no request leaving the machine. The README states it can "read pages, click buttons, fill forms, run JavaScript, and answer questions about any site you visit." That combination, local inference plus DOM actuation, is the reason to care. A cloud assistant can read a page you paste into it; it cannot click a selector for you without a round trip.

The audience is developers and technically comfortable Chrome users who want to experiment with browser agents on their own hardware. It is not a product for someone who wants a chat box that always works. The README's own hardware table is marked "Estimated minimal requirements (not benchmarked on real devices)," which is an honest caveat and also a warning: nobody has published numbers for this build.

Three contexts, one agent loop: how Gemma Gem is wired

The architecture is split across three extension contexts, and the split matters because each one has different privileges.

The offscreen document hosts the model through @huggingface/transformers and WebGPU, and runs the agent loop. The service worker routes messages between the content script and the offscreen document, and handles two tools itself: take_screenshot and run_javascript. The content script injects the gem icon and the shadow DOM chat overlay, and executes the DOM tools: read_page_content, click_element, type_text, scroll_page.

That placement is deliberate. Screenshots and arbitrary JS execution need broader reach than the page's own content script, so they run from the service worker. Selecting and typing into elements is page-scoped work, so it stays in the content script. The README lists run_javascript as executing "in the page context with full DOM access," which is the sharpest edge in the design and the one to think about before installing.

Models are q4f16 quantized ONNX builds from onnx-community, with 128K context. The README notes that at long contexts the KV cache adds 10-20% memory overhead on top of model weights, so the 128K ceiling is not free.

Installing Gemma Gem and getting a first answer out of it

Setup is a two-command build followed by a manual extension load. The README gives pnpm install and pnpm build, then instructs you to load the extension from .output/chrome-mv3-dev/ in chrome://extensions with developer mode on.

bash
pnpm install
pnpm build

The build script is wxt build --mode development, so this produces the development build with logging and source maps. For a quieter bundle, package.json defines pnpm build:prod, which runs wxt build with logging silenced and minified output.

bash
pnpm build:prod

After loading the unpacked extension, open any page and click the gem icon in the bottom-right corner. The README says to wait for the model to load, with progress shown on the icon and in the chat. The first run downloads the model; the README puts the cache at roughly 500 MB for E2B and 1.5 GB for E4B, and says it is cached after that first run. Then ask a question about the page, or request an action.

Model choice lives behind the gear icon in the chat header, and the README states the selection persists across sessions. The same panel holds the thinking toggle, a max-iterations cap on tool call loops, shortcut rebinding, a clear-context button, and a per-hostname disable. Defaults for the shortcuts are Alt+G to toggle and Escape to close, and both are rebindable by clicking the field and pressing the combination you want.

The hardware gate is the real adoption cost

WebGPU is not optional here, and the README is specific about it: Chrome 113+ or Edge 113+, plus the shader-f16 GPU feature. If your GPU does not expose shader-f16, the model does not run, and no amount of RAM compensates.

The memory table is where most laptops fall out. E2B needs an estimated 4 GB of GPU VRAM or shared memory and 6-8 GB of system RAM; E4B needs 6 GB and 8-16 GB. The README lists integrated GPUs (Intel Xe Arc, Apple M1+, Qualcomm Adreno) as workable "with sufficient shared memory," and describes performance as slow on integrated GPUs, normal on mid-range discrete GPUs, fast on high-end GPUs. Mobile is listed as slow outright. Since these figures are explicitly unbenchmarked, treat the table as a floor rather than a specification, and expect the first model load to be the moment you find out.

One more practical gap: the repository has no retrieved releases. Version 0.3.0 in package.json is the only version signal, so anyone wanting a pinned artifact has to build from source.

Where Gemma Gem is the wrong tool

The per-hostname disable setting exists for a reason. A model that can read page content, type into inputs and execute JavaScript in the page context is a poor fit for banking, healthcare or internal admin surfaces, and the README does not document a confirmation step before run_javascript fires. There is no rollback described either: the README does not document undoing an action the agent took on a page.

The tool loop is capped by a max-iterations setting, which is a guard against runaway tool calls rather than a correctness guarantee. Nothing in the README describes validation of the selectors the model produces, so a hallucinated CSS selector is a click that lands nowhere or somewhere unintended.

Context handling is also page-scoped and manual. Clearing context resets conversation history "for the current page," which means the mental model is one conversation per page, not one assistant with memory. If you want an agent that remembers what you did yesterday across sites, this is not it.

Finally, the README documents Chrome and Edge only, even though package.json carries dev:firefox, build:firefox and zip:firefox scripts. Those scripts are not documented as supported, and the requirements section names Chrome with WebGPU support.

How it compares to a browser automation framework

The closest thing to a real alternative is Playwright or Puppeteer driving a local model, or a browser extension that calls a hosted API. The difference is where the model and the decision loop live.

In Playwright, your script decides what to click and the model, if any, sits outside the browser. Gemma Gem inverts that: the model sits inside the browser in an offscreen document, and the decision loop is the agent loop running there. You do not write selectors; the model proposes them from the page text it read through read_page_content. That is more flexible on unfamiliar pages and far less predictable on pages where a wrong click has consequences. Playwright also gives you a test runner, retries and trace files; the README describes none of that here.

Against a hosted-API extension, the trade is capability for privacy. A hosted model is larger and faster on the same laptop, and it works on machines without shader-f16. Gemma Gem's answer is that page content never leaves the machine, which is the whole point of running q4f16 Gemma 4 E2B or E4B locally.

Licence, maintenance and what an upgrade costs

The repository is Apache-2.0, stated both in the licence file and in package.json. That is a permissive licence with an explicit patent grant, and it places no copyleft obligation on your own code. It says nothing about the model weights, which come from separate onnx-community repositories on Hugging Face and carry their own terms; check those before shipping anything. This is not legal advice.

The last push to the default branch was on 2026-05-29. The repository is not archived. There are no retrieved releases, so there is no changelog-driven upgrade path to follow beyond CHANGELOG.md in the repository root.

Upgrade cost is dominated by model size rather than code. Switching from E2B to E4B in settings moves the download from roughly 500 MB to 1.5 GB and raises the estimated GPU memory requirement from 4 GB to 6 GB, with system RAM going from 6-8 GB to 8-16 GB. On a machine that only just ran E2B, that switch is the upgrade that fails. Dependency-wise the surface is small: wxt, @huggingface/transformers, @kessler/gemma-agent and marked, with TypeScript as a dev dependency.

Editorial conclusion

Adopt Gemma Gem if you want to see a browser agent run against a real page without sending page content to a hosted API, and you have a Chrome or Edge build with WebGPU and at least 4 GB of GPU-addressable memory for E2B. Do not adopt it if your target machines are integrated-GPU laptops with 8 GB of RAM, if you need Firefox support (pnpm build:firefox exists in package.json but the README documents Chrome only), or if you need a published release to pin. Verify two things before rollout: that your hardware exposes the shader-f16 GPU feature the README lists as required, and that run_javascript in the page context is acceptable under your own security policy, since the README documents full DOM access with no sandbox described.

Frequently asked questions

Is Gemma Gem the same as Gemini?

No. Gemma Gem is a Chrome extension that runs Google's Gemma 4 model locally through WebGPU. The README describes it as an on-device assistant with no API keys and no cloud, which is a different thing from a hosted service.

What exactly is Gemma Gem?

It is a browser AI agent built on Gemma 4 that reads pages, clicks buttons, fills forms, runs JavaScript and answers questions about the site you are on. The README states the model runs entirely on-device via WebGPU.

How do I install the Gemma Gem Chrome extension?

Run pnpm install and pnpm build, then load the unpacked extension from .output/chrome-mv3-dev/ in chrome://extensions with developer mode enabled. The README requires Chrome with WebGPU support and notes the model is cached after the first run.

What hardware does Gemma Gem need to run Gemma 4?

The README's estimated table lists 4 GB of GPU VRAM or shared memory and 6-8 GB of system RAM for E2B, and 6 GB with 8-16 GB of RAM for E4B, on Chrome 113+ or Edge 113+ with the shader-f16 GPU feature. It labels these estimates as not benchmarked on real devices.

Which tools can Gemma Gem run on a page?

read_page_content, click_element, type_text and scroll_page run in the content script, while take_screenshot and run_javascript run in the service worker. The README says run_javascript executes in the page context with full DOM access.

Official sources

  1. Issues
  2. kessler/gemma-gem on GitHub
  3. License: Apache-2.0
  4. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/kessler-gemma-gem.svg)](https://hysenlabs.com/projects/kessler-gemma-gem)