# Oxylabs Google News Scraper: Free Python Tool and API

> Oxylabs Google News Scraper is a free Python command-line tool that extracts articles from Google News by topic ID and writes them to a CSV file. The repository also documents the paid Oxylabs SERP API for structured, at-scale news data collection.

**oxylabs/google-news-scraper** — Use Google News Scraper API to obtain the latest global news for your project, including a wide range of sources, headlines, URLs, and publication dates from the Google News platform.

- Repository: https://github.com/oxylabs/google-news-scraper
- Website: https://oxylabs.io/products/scraper-api/serp/google?utm_source=877&utm_medium=affiliate&groupid=877&utm_content=google-news-scraper-github&transaction_id=102c8d36f7f0d0e5797b8f26152160
- Stars: 3,382 · Forks: 25
- Language: Python
- License: not declared
- Published: 2026-09-23 · Updated: 2026-09-23 · Language: en
- Canonical page: https://hysenlabs.com/projects/oxylabs-google-news-scraper

## What the Free Google News Scraper Retrieves

The free component of this repository targets Google News topic pages. Google News organises articles into named topics visible at the top of its homepage, such as Business, Technology, or Sports. Each topic page has a URL that includes a base64-encoded topic ID after the `/topics/` path segment. The tool takes that ID as input and returns the articles listed on that page at the time of the run.

The output file is named `articles.csv` and appears in the directory where you run the tool. The README does not specify the exact column names in the CSV, but the tool is described as retrieving sources, headlines, URLs, and dates. This positions it as useful for news monitoring, media research, or building a dataset of recent coverage on a specific topic. It does not support arbitrary keyword search across Google News; it is limited to the predefined topic IDs that Google exposes in the topic page URLs.

## Installing with Poetry

The project requires Python 3.11. The `pyproject.toml` pins this with `python = "^3.11"`. Installation uses the same Make-based approach as the Oxylabs Amazon Scraper from the same author:

```bash
make install
```

The Makefile runs `pip install poetry==1.8.2` and then `poetry install`. Pinning Poetry to version 1.8.2 rather than the latest release keeps the dependency lock file consistent across machines. The installed dependencies include `selenium` and `webdriver-manager` for browser automation, `httpx` for HTTP requests, `bs4` (BeautifulSoup4) for HTML parsing, `pydantic` and `pydantic-settings` for data validation and configuration, `pandas` for CSV output, and `click` for command-line argument handling.

The presence of both Selenium and httpx suggests the tool uses a browser for initial page loading, then httpx or BeautifulSoup for parsing the returned content. Either way, a Google Chrome installation on the host machine is required, since `webdriver-manager` only manages the ChromeDriver binary and not Chrome itself.

## Finding a Topic ID and Running the Scraper

The workflow for a first scrape requires two steps. First, open Google News in a browser and click on a topic from the header navigation. The URL in the address bar will contain a path like `/topics/` followed by a long base64 string. That string is the topic ID. For the Business topic, the README shows an example ID:

```bash
make scrape TOPIC_ID=CAAqJggKIiBDQkFTRWdvSUwyMHZNRGx6TVdZU0FtVnVHZ0pWVXlnQVAB
```

Note that the Makefile defines an internal variable named `TOPIC` (not `TOPIC_ID`), so there is a discrepancy between the README's example and the Makefile. The README documentation instructs you to use `TOPIC_ID`, but the underlying Make target checks for `$(TOPIC)`. If you run the README example exactly and it fails with the error message about a missing topic ID, try passing `TOPIC` instead of `TOPIC_ID`. The Makefile internally calls `poetry run python -m google_news_scraper --topic=$(TOPIC)`.

After a successful run, `articles.csv` appears in the current directory. The README notes that the tool shows progress in the terminal while running.

## Architecture: Selenium, httpx, and BeautifulSoup Together

The dependency stack in `pyproject.toml` includes both Selenium-based browser automation and the `httpx` HTTP client alongside `bs4` for HTML parsing. This combination is common when a page requires JavaScript execution for its initial load (handled by Selenium) but the actual article data can be extracted with simpler HTML parsing (handled by BeautifulSoup4) once the rendered content is available.

The `webdriver-manager` package downloads a ChromeDriver binary that matches the version of Chrome installed on the machine, eliminating the manual ChromeDriver installation step. The `pydantic` library validates the data extracted from each article before it is written to the CSV, which catches structural issues at parse time rather than silently writing malformed rows. This is the same validation approach used in the Oxylabs Amazon Scraper repository, and both projects share the same author.

Because Google News is a JavaScript-heavy single-page application, a purely static HTTP approach would return an empty page skeleton. The Selenium layer loads the topic page in Chrome, waits for the articles to render, and then makes the rendered DOM available for extraction.

## The Paid Oxylabs Google News API

The second half of the repository covers the paid Oxylabs SERP API, specifically its Google News mode. The API uses the `google_search` source with a `udm` context parameter set to `12`, which instructs Google to return news results. The README gives a Python example:

```python
payload = {
    'source': 'google_search',
    'query': 'adidas',
    'parse': True,
    'context': [
        {'key': 'udm', 'value': '12'},
    ],
}
```

Requests go to `https://realtime.oxylabs.io/v1/queries` with HTTP Basic Auth using an Oxylabs username and password. The API returns data in JSON, HTML, or Markdown format. When `parse` is `True`, the response is structured JSON. The README states the API delivers a list of sources, titles, URLs, and dates. Unlike the free tool, the API handles proxy rotation and JavaScript rendering on Oxylabs infrastructure, so no local browser is required. A free trial offers up to 2,000 results before any payment is needed.

## Where the Free Tool Falls Short

The free scraper is limited to topic pages as defined by Google News's own topic navigation. You cannot pass an arbitrary search query; you can only scrape a topic that Google has already classified and published at a `/topics/` URL. This is a narrower interface than keyword search, and it means you cannot use this tool to monitor coverage of a specific company or event unless Google has created a matching topic page for it.

There is also no built-in rate limiting, retry logic, or proxy rotation. Google will eventually detect repeated automated requests from the same IP address and return a CAPTCHA or an empty response. When that happens, the tool either fails or writes an incomplete CSV with no error indication. For reliable, recurring collection, the paid API is the only option this repository documents.

A comparable alternative is the `pygooglenews` Python library, which wraps Google News RSS feeds rather than scraping the web interface. RSS feeds do not require a browser, making `pygooglenews` significantly lighter to run. The trade-off is that RSS feeds contain fewer articles per topic and update less frequently than the full Google News page. For production news monitoring with structured data requirements, neither the free tool nor `pygooglenews` matches what the paid Oxylabs API offers in terms of reliability and result volume.

## Conclusion

Developers looking for a quick, locally-run starting point for Google News data will find this tool workable for single-topic exports. The dependency on Selenium and a Chrome browser means it is not portable to serverless environments, and there is no built-in mechanism for handling Google's rate limiting. Teams that need reliable, ongoing collection at scale should evaluate the paid Oxylabs API, which requires a free trial account before any requests are possible.

## FAQ

### Is Google scraping legal?

Scraping publicly visible Google pages sits in a legal grey area that depends on jurisdiction, the terms of service you agreed to, and how you use the data. Google's terms of service restrict automated access. The repository does not provide legal guidance on this question.

### Does Google allow data scraping?

Google's terms of service prohibit scraping its services without explicit permission. The company does offer official APIs for some data, including the Google News API, but the free scraper in this repository does not use those official channels.

### Is web scraping illegal?

Web scraping is not universally illegal, but its legality depends on the jurisdiction, the website's terms of service, what data is being collected, and how it is used. The `hiQ Labs v. LinkedIn` case in the United States addressed some of these questions, but the legal situation varies widely by country and continues to evolve.

## Sources

- [Issues](https://github.com/oxylabs/google-news-scraper/issues)
- [oxylabs/google-news-scraper on GitHub](https://github.com/oxylabs/google-news-scraper)
- [Project website](https://oxylabs.io/products/scraper-api/serp/google?utm_source=877&utm_medium=affiliate&groupid=877&utm_content=google-news-scraper-github&transaction_id=102c8d36f7f0d0e5797b8f26152160)
- [README](https://github.com/oxylabs/google-news-scraper/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/oxylabs-google-news-scraper
