Library / SDK
oxylabs/google-ai-mode-scraper avatar
oxylabs/google-ai-mode-scraper

oxylabs/google-ai-mode-scraper: a hosted wrapper for Google AI Mode answers

Scrape Google AI Mode responses without blocks on a large scale.

3,625 stars32 forksJavaLicense varies

At a glance

What is it?
The repository is not a scraper you run. It is a thin set of request samples for Oxylabs' Web Scraper API, which returns Google AI Mode response_text, links and citations as JSON. Here is what that means for adoption.
Who is it for?
Adopt it if you need Google AI Mode answers as structured JSON at volume and you would rather pay Oxylabs than run headless browsers and proxies yourself; the sample code is short enough to port to any HTTP client. Do not adopt it if you need something you can run offline, audit end to end, or use without an Oxylabs account, because the repository contains no scraper.
Can I use it commercially?
Not without permission. GitHub finds no licence file in the repository, and without a licence all rights are reserved by default: you may read the code but not reuse it. Check the README, or ask the authors, before using it.
Is it still maintained?
Yes. The repository last received commits 40 days ago.
What is it written in?
Mainly Java, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 28, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What oxylabs/google-ai-mode-scraper actually solves

Google AI Mode answers are generated per query and per location. The result page is assembled with a headless browser, and the answer block, the citation list and the right-hand link box are not exposed through a documented endpoint you can call directly. Anyone who wants those answers as data has to solve the same three problems first: getting a request through without being blocked, rendering the page so the answer exists in the DOM, and turning the rendered page into fields a program can store.

The repository addresses the third problem and outsources the first two. It is a set of request samples, in Python and other languages under Code examples/, against a hosted endpoint. The README describes the product as built on the Web Scraper API, which handles proxies, headless browsers and automated request systems on Oxylabs' side. The audience is therefore narrow and specific: teams doing SEO or generative engine optimisation work, people assembling training or evaluation datasets from AI answers, and data engineers who need a repeatable feed of AI Mode output without maintaining a rendering farm.

It is not a library you import and not a crawler you schedule. There is no parser in the repository, no proxy pool, no browser automation. If you were hoping to clone something and point it at Google, the README's first instruction is to register on the Oxylabs dashboard for a trial, which tells you where the work happens.

The request path: one POST, one parsed object

The mechanism is a single HTTP POST to https://realtime.oxylabs.io/v1/queries with basic authentication and a JSON body. Three parameters are marked mandatory in the README table: source, which must be google_ai_mode; query, the prompt, which cannot exceed 400 characters; and render, which must be set to html for this source. Everything else is optional.

parse is the switch that matters most. It defaults to false, so a request without it returns the rendered page rather than fields. Set parse to true and the response carries a content object with prompt, response_text, citations, links and parse_status_code. citations is the list shown in the main answer block, each entry a URL plus the text that was quoted; links is the separate box on the right side of the page. The output sample also shows markdown_text as a field, and the README points to a separate page on Markdown output for integration with AI tooling.

The envelope around that content is operational: job_id, created_at, updated_at, page, status_code, and the final Google URL including udm=50 and a uule parameter that encodes the location. geo_location is the input that drives uule, and the README links to the localisation documentation rather than listing accepted values. callback_url turns the call asynchronous by pushing results to your own endpoint, which is the only path the README documents for avoiding a blocking wait on long jobs.

One design consequence is worth naming. Because render is forced to html and parse is a separate flag, the same request can return either raw HTML or structured fields. That is convenient for debugging a query whose parse_status_code looks wrong, but it means two code paths in your client for one logical operation.

Installing the Python sample and reading your first response

There is nothing to install from this repository. The README's Python sample needs only the requests library, which you install into your environment first.

bash
pip install requests

The script below is the README's request sample with the credentials left as placeholders. Replace USERNAME and PASSWORD with the Web Scraper API credentials from your Oxylabs dashboard, then run it.

python
import json
import requests

payload = {
    'source': 'google_ai_mode',
    'query': 'best health trackers under $200',
    'render': 'html',
    'parse': True,
    'geo_location': 'United States'
}

response = requests.post(
    'https://realtime.oxylabs.io/v1/queries',
    auth=('USERNAME', 'PASSWORD'),
    json=payload,
)

print(response.json())

The call posts to the realtime endpoint and prints the JSON. On success you should see a results array whose single element contains content.response_text, content.citations and content.links, plus job_id and a status_code of 200. The README's own sample output shows exactly that shape for the query "best health trackers under $200", including a parse_status_code of 12000 next to the parsed content.

To keep the result, the README appends a write step that dumps the same object to response.json with indent=2. That file is what you would diff against the repository's output-sample.json when checking whether a field you depend on has moved. Note that the README does not document retry behaviour, rate limits or how a failed parse surfaces beyond parse_status_code, so your first real task after this run is deciding what your client does when that code is not the value you expect.

The 400-character prompt limit and the fields that move

The hard constraint in the parameter table is the query length: 400 characters maximum. For short consumer-style prompts that is ample. For a detailed instruction with context, role framing and output formatting, it is not, and the README offers no documented workaround such as prompt chaining or session continuation. If your use case is long-form prompting, this source will not carry it.

The second limitation is field variability. The README states plainly that depending on the search query, both the number of items and the fields included can change. That is a real operational risk for anyone writing a fixed schema. A query that returns citations may return none for a different prompt, and the parser's behaviour on those queries is not described. The safest posture is to treat citations and links as optional arrays and to keep the raw JSON alongside whatever you extract, so a schema change does not cost you the source data.

Third, the repository is silent on several things a production user would ask about: how long a job can take before a timeout, whether callbacks retry on failure, and what the scraper status codes beyond the ones linked in the README mean for billing. The README links to the developers.oxylabs.io status code page rather than reproducing the list, so that page, not this repository, is where you resolve those questions.

Finally, the licence is not stated anywhere in the repository metadata or the README. That is unusual for a public sample repository and it matters if you intend to vendor the sample files into your own codebase.

Firecrawl and Tavily solve a different half of the problem

The repository topics list firecrawl-alternative and tavily-alternative, which frames the comparison. The difference is in what each tool treats as the unit of work.

Firecrawl is aimed at crawling and converting arbitrary sites to Markdown or structured output. Its input is a URL or a domain, and its output is page content. Tavily is a search API aimed at returning ranked results and extracted content for a query, typically for retrieval in a language model pipeline. Both put the retrieval decision in your hands: you choose the sources.

This project's input is a prompt, not a URL, and its output is Google's own synthesised answer with the citations Google chose. You are not deciding which pages matter; you are capturing a third party's selection and phrasing. For generative engine optimisation work that is the point, because you want to see what the answer engine says and which sources it credits. For building a retrieval corpus it is the wrong shape entirely, since you inherit Google's citation bias and get no control over coverage.

There is also an architectural difference. Firecrawl and Tavily can be self-hosted or called directly depending on plan, while this repository has no local execution path at all. Every request goes through Oxylabs' infrastructure, which is what removes the proxy and rendering burden and also what makes an Oxylabs account mandatory.

Maintenance, licence and what you are actually depending on

The repository is not archived, and the last push was on 2026-08-21. That is recent enough that the samples track the current API shape, but the repository has no releases, so there is no versioned artefact to pin. Your dependency is not the Git repository at all. It is the realtime.oxylabs.io endpoint and the google_ai_mode source behind it. Upgrading means reading the developers.oxylabs.io documentation for changes to parameters, output fields or status codes, not pulling a new tag.

That reframes the cost calculation. You carry no browser, no proxy pool and no rendering infrastructure, which is the saving. You do carry an account, credentials in your configuration, and a per-request cost that the README does not state. The README mentions a free trial through the dashboard but gives no pricing, so budget planning has to happen on Oxylabs' site rather than in this repository.

The licence question is unresolved. The repository metadata shows no licence, and the README does not mention one. Without a stated licence, the default position for most jurisdictions is that no rights are granted beyond what the platform's terms allow, so copying the sample files into a commercial codebase is a question for Oxylabs, not an assumption. Nothing here is legal advice; the practical step is to ask Oxylabs directly before vendoring the code.

Editorial conclusion

Adopt it if you need Google AI Mode answers as structured JSON at volume and you would rather pay Oxylabs than run headless browsers and proxies yourself; the sample code is short enough to port to any HTTP client. Do not adopt it if you need something you can run offline, audit end to end, or use without an Oxylabs account, because the repository contains no scraper. Before committing, verify three things: that your Oxylabs plan covers the google_ai_mode source, that 400 characters is enough for your prompts, and that the parser fields you plan to store (response_text, citations, links) are stable for your queries by comparing runs against output-sample.json. The licence is not stated in the repository, so treat redistribution of the sample files as unresolved until Oxylabs confirms it in writing.

Frequently asked questions

Is AI scraping illegal?

The repository does not address legality. It documents a hosted API that sends prompts to Google AI Mode and returns parsed JSON, and it points to Oxylabs' terms and documentation for how the service may be used. Questions about the legality of scraping in your jurisdiction are outside what the README covers.

Is Google scraping legal?

The README does not discuss the legality of scraping Google. It describes the google_ai_mode source, the required render value of html, and the parsed output fields, and directs readers to developers.oxylabs.io for the rest.

Can I remove AI Mode from Google?

The README does not cover Google's interface settings or disabling AI Mode in a browser. This project only sends prompts to AI Mode through Oxylabs' API and returns the answers as JSON.

Are there any AI scrapers?

This repository is one: it targets Google AI Mode through the web scraper API, with source set to google_ai_mode and parse set to true for structured output. The README also lists firecrawl-alternative and tavily-alternative among the repository topics for adjacent tools.

Official sources

  1. Issues
  2. oxylabs/google-ai-mode-scraper on GitHub
  3. Project website
  4. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/oxylabs-google-ai-mode-scraper.svg)](https://hysenlabs.com/projects/oxylabs-google-ai-mode-scraper)