Oxylabs Amazon Scraper: Free Python Tool and Commercial API
Free Trial Amazon Scraper API for extracting search, product, offer listing, reviews, question and answers, best sellers and sellers data.
At a glance
- What is it?
- Oxylabs Amazon Scraper is an open-source Python tool that extracts product listings from Amazon department pages into a local CSV file. It also documents how to reach the paid Oxylabs Scraper API when department-page output is not enough.
- Who is it for?
- Developers who need a quick way to collect product titles, ASINs, URLs, and prices from a single Amazon department page will find this tool usable after running `make install`. It does not reach individual product pages, review sections, Q&A threads, or best-seller lists.
- Can I use it commercially?
- Not without permission. GitHub finds no licence file in the repository, and without a licence all rights are reserved by default: you may read the code but not reuse it. Check the README, or ask the authors, before using it.
- Is it still maintained?
- Yes. The repository last received commits 103 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What the Free Scraper Collects from Amazon Department Pages
The free component of this repository is a command-line tool that scrapes a single Amazon department page and writes the results to a file named `amazon_products.csv`. The CSV contains five columns: `title` (the product name as it appears in the listing), `url` (the full URL to the product's Amazon page), `asin_code` (the product's unique ASIN identifier), `image_url` (the URL of the main product image), and `price` (the listed price, which is left empty if the product is out of stock). When a listed product is out of stock, the tool notifies you in the terminal rather than silently omitting the row. This design means the output file will include rows with an empty price field, so downstream code must account for that.
The scope is intentionally narrow. The tool works with department browse pages, not search results pages, individual product detail pages, or category-filtered views that require JavaScript interaction beyond what the tool performs. If you need data from a specific product ASIN, or if you want the full review text, the Q&A section, or offers from third-party sellers, the free tool provides none of those. The README is explicit: for those cases, use the paid API described in the second half of the repository.
Installing the Tool with Poetry and Make
The project requires Python 3.11. The `pyproject.toml` declares this constraint directly: `python = "^3.11"`. The install process uses Make to standardise the setup steps. Running the install command from the repository root handles both the package manager installation and the dependency installation in one step:
make installInternally, the Makefile runs `pip install poetry==1.8.2` followed by `poetry install`. This pins Poetry to version 1.8.2 rather than letting pip pull the latest release, which helps keep the dependency resolution consistent. After the install completes, the relevant packages include `selenium` (browser automation), `webdriver-manager` (automatic ChromeDriver version management), `selenium-wire` (a Selenium extension that allows traffic interception), `pydantic` and `pydantic-settings` (data validation and settings management), `pandas` (for writing the CSV output), and `click` (CLI argument parsing). Because Selenium drives a real browser, you also need Google Chrome installed on the machine where you run the tool. The `webdriver-manager` library will download a matching ChromeDriver automatically on first run, but Chrome itself must be present.
Running Your First Amazon Department Scrape
The scraping workflow has two steps. First, open Amazon in a browser, navigate to a department page such as Computers & Accessories, and copy the full URL from the address bar. Then run the scrape command with that URL as the argument:
make scrape URL="https://www.amazon.com/s?i=specialty-aps&bbn=16225009011&rh=n%3A%2116225009011%2Cn%3A541966&ref=nav_em__nav_desktop_sa_intl_computers_and_accessories_0_2_5_6"The README emphasises quoting the URL with double quotes. Amazon department URLs typically contain query parameters with ampersands and percent-encoded characters, and an unquoted URL will be split by the shell at each ampersand, causing the Make command to fail with a parsing error. After the tool finishes, `amazon_products.csv` appears in the current directory with one row per listed product. Rows for out-of-stock products are included, with the `price` column left blank and a terminal notification printed during the run.
The Make target itself calls `poetry run python -m amazon_scraper --url="$(URL)"`, so the actual entry point is the `amazon_scraper` package inside `src/`. If you prefer to skip Make and invoke Poetry directly, that command works the same way once dependencies are installed.
How Selenium and selenium-wire Drive the Browser
The scraper does not send plain HTTP requests to Amazon. It launches a Chrome browser through Selenium, navigates to the department URL, and reads the rendered DOM. This matters because Amazon's department pages load product listings dynamically, and a static HTTP request would return a page skeleton without the product data.
The `webdriver-manager` package handles the browser driver automatically. On the first run it checks the installed version of Chrome and downloads a matching `chromedriver` binary to a local cache directory. This removes the manual step of downloading and version-matching ChromeDriver, which is a common source of setup failures in Selenium-based tools.
`selenium-wire` extends the standard Selenium `webdriver.Chrome` with the ability to intercept HTTP requests and responses. The tool uses this to inspect the network traffic the browser generates while loading the department page, which allows it to capture structured data that appears in API responses before it is rendered to the DOM. `pydantic` validates the extracted data against a schema, so type errors surface immediately rather than being written silently into the CSV. The `blinker` dependency is pinned to `<1.8.0` in `pyproject.toml`, which is a constraint imposed by `selenium-wire`'s compatibility requirements.
This approach means the machine running the scraper must have Chrome installed, must have network access to Amazon, and will make real browser connections that Amazon can observe and rate-limit.
The Paid Oxylabs API: Seven Amazon Data Sources
The second half of the repository is documentation and code examples for the Oxylabs Web Scraper API, which is a commercial service requiring a separate account and a paid plan or free trial. The API supports seven source values for Amazon:
- `amazon`: accepts any Amazon URL and returns whatever page type that URL represents - `amazon_search`: returns search results for a given search term - `amazon_product`: returns the product detail page for a given ASIN - `amazon_pricing`: returns all offer listings for a given ASIN - `amazon_questions`: returns the Q&A page for a given ASIN - `amazon_bestsellers`: returns best-seller rankings for a taxonomy node - `amazon_sellers`: returns information for a given seller
All sources except `amazon` return structured data. The `amazon` source depends on the URL: if it maps to a page type that Oxylabs can parse, structured data is available when you set the `parse` parameter to `true`. The API handles proxy rotation, IP management, and JavaScript rendering on its infrastructure. You send a POST request to `https://realtime.oxylabs.io/v1/queries` with your credentials and the payload, and receive the data in JSON, HTML, or Markdown format.
Limitations of the Free Tool and When the API Becomes Necessary
The free scraper has three constraints that matter in practice. First, it only reads department browse pages. If you need to build a dataset of individual product specifications, review summaries, or seller information, the tool provides no path to that data. Second, because it runs Chrome locally, it is not deployable to serverless environments or containers without a headless Chrome setup. Third, there is no built-in proxy rotation or session management. Amazon will eventually block repeated requests from the same IP address, and when that happens, the scraper fails silently or returns an incomplete page without any retry logic.
Apify offers a cloud-based Amazon scraper actor that runs on Apify's infrastructure without requiring a local Chrome installation or IP management. The trade-off is that Apify charges per compute unit used, whereas this repository's free tool has no per-request cost once running locally. Apify's actor also covers more Amazon page types than the free tool, which is closer in scope to what the paid Oxylabs API provides. The choice between them depends on whether you need local execution, cloud execution, or the specific structured outputs of the Oxylabs API sources.
Editorial conclusion
Developers who need a quick way to collect product titles, ASINs, URLs, and prices from a single Amazon department page will find this tool usable after running `make install`. It does not reach individual product pages, review sections, Q&A threads, or best-seller lists. For those use cases, the paid Oxylabs Scraper API must be evaluated on its own terms, beginning with a trial account at oxylabs.io.
Frequently asked questions
Is scraping Amazon illegal?
Scraping publicly visible Amazon pages sits in a legal grey area that varies by jurisdiction and by the terms of service you agreed to. Amazon's terms of service prohibit automated access, but the legality under general law depends on what data you collect and how you use it. The README for this tool does not provide legal guidance.
Is it possible to scrape Amazon?
Yes, it is technically possible. This repository demonstrates one approach using a Selenium-driven Chrome browser to load department pages and extract product listings. Amazon actively detects and blocks automated traffic, so scrapers must account for rate limiting and IP blocks.
What is the Amazon Scraper tool in this repository?
It is a free, open-source Python command-line tool that opens an Amazon department page in a headless Chrome browser, extracts the listed products, and writes title, URL, ASIN code, image URL, and price into a CSV file named `amazon_products.csv`.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/oxylabs-amazon-scraper)