oxylabs/how-to-scrape-amazon-product-data: a Python tutorial repo for titles, prices and ratings
The process of extracting product data from Amazon using Python, including titles, ratings, prices, images, and descriptions.
At a glance
- What is it?
- This repository is a step-by-step Python guide to pulling product name, rating, price, images and description from Amazon, plus a pointer to Oxylabs' paid Scraper API. It teaches requests and BeautifulSoup, and the README is upfront that Amazon blocks the first request you send.
- Who is it for?
- Adopt this repository if you want a readable walkthrough of Amazon product page structure and a working requests plus BeautifulSoup starting point for a small, low-volume job, and if you accept that the tutorial stops before proxies, retries or browser rendering.
- Can I use it commercially?
- Not without permission. GitHub finds no licence file in the repository, and without a licence all rights are reserved by default: you may read the code but not reuse it. Check the README, or ask the authors, before using it.
- Is it still maintained?
- Yes. The repository last received commits 114 days ago.
- What is it written in?
- GitHub does not report a main language for this repository.
Answers come from the project's GitHub data, last synced on September 28, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What the Oxylabs Amazon product data guide actually covers
The repository is a tutorial, not a library. There is no package to import and no CLI to run. The README walks through a single script that fetches one Amazon product detail page and extracts five fields: product name, product rating, product price, product images and product description. The intended reader is someone who already writes some Python and wants to see where those values sit in Amazon's HTML. The guide also covers a category or search results page, which the README describes as the page that shows search results and carries the product URL you need before you can reach the detail page. The final step writes the collected rows to a CSV file with pandas. Everything is done with requests, beautifulsoup4, lxml and pandas, all installed from PyPI. There is no database, no scheduler and no queue. If you want a maintained extraction service rather than a lesson, the README points at Oxylabs' own Scraper API at the end, which is a different product with a different cost model.
Why the first request fails and what the header dictionary changes
The most useful part of the guide is the failure it documents before any parsing happens. A plain requests.get against a product URL returns a body containing the line about contacting [email protected], and the README states the status code may be 503 rather than 200. The explanation given is that Amazon can tell the request did not come from a browser. The fix shown is a headers dictionary with a user-agent and an accept-language value, passed to the optional headers parameter of get. The README notes that a user-agent alone is sometimes enough and sometimes not, and offers a longer dictionary with Accept-Encoding, Accept and a Referer pointing at a Google search URL. It also suggests rotating user-agent strings and retrying. That is honest about the mechanism, and it is also where the tutorial's scope ends. There is no proxy pool, no backoff schedule, no cookie handling and no captcha discussion. The README says that if you need JavaScript rendering you will have to use tools like Playwright or Selenium, and leaves it there. Treat the header trick as a way to see the HTML once, not as an access strategy that holds up under volume.
Installing the packages and scraping one product page
The README's setup section creates a virtual environment and installs four packages. The commands below are the macOS and Linux versions exactly as the README gives them; the README also lists Windows variants that use python instead of python3 and a different activation path.
python3 -m venv .env
source .env/bin/activate
python3 -m pip install requests beautifulsoup4 lxml pandasAfter activation, the guide has you create a file named amazon.py and send a request to a Bose QuietComfort 45 product URL, printing the response text. The point of that first run is to see the block message, not the product HTML. Then you add the headers dictionary and pass it to get.
custom_headers = {
'user-agent': 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/108.0.0.0 Safari/537.36',
'accept-language': 'en-GB,en;q=0.9',
}
response = requests.get(url, headers=custom_headers)Run it with `python3 amazon.py`. What you should see is the product page HTML rather than the automated access notice, and a status code of 200 rather than 503. The README warns this outcome is not guaranteed and that the longer header set may be needed instead. From there the guide locates each field by inspecting the page in a browser, right-clicking the element and reading the highlighted markup, which is how it arrives at the span tag for the product title.
Category pages, the ASIN route, and where the tutorial stops being enough
The README splits Amazon pages into two kinds and is clear about the consequence: a category page gives you the title, image, rating, price and the product URL, while the description only exists on the detail page. That means a complete record requires two requests per product, one to discover the URL and one to open it. The guide's later sections cover scraping products from search results, extracting product details, and scraping products by ASIN. The ASIN route is the more stable of the three, because it addresses a product directly instead of depending on a search results layout. The limitation is structural rather than a bug. Selectors derived from a live Amazon page are tied to markup that Amazon changes, and the README gives no selector versioning, no fixture HTML and no test suite, so nothing tells you when a selector has gone stale. There is also no rate limiting in the code as shown, which matters because the guide itself demonstrates that Amazon responds to request patterns. For a few dozen products this is fine. For a catalogue, the missing pieces are the ones the README delegates to Playwright, Selenium or the paid API.
The paid Scraper API alternative, and how it differs from the tutorial
The README's final section, titled as an easier solution to extract Amazon data, is Oxylabs' own Scraper API. The difference in approach is not cosmetic. The tutorial runs on your machine, from your IP address, with headers you maintain, and returns raw HTML that you parse with BeautifulSoup. The API is a hosted endpoint: you send a request to Oxylabs and receive structured product data, with the proxy rotation and rendering handled on their side. That removes the header rotation and the 503 problem from your code, and it removes your ability to inspect the HTML when a field looks wrong. It also introduces a recurring cost and a dependency on a vendor's extraction schema, so a field Amazon renames is the vendor's problem rather than a selector you patch. If you are comparing hosted options, the related search terms point at Apify as another name people look at; the same trade applies, in that you are buying managed extraction instead of writing it. The README does not benchmark the API against the script, and it does not state a price.
Maintenance, licence and what the repository does not tell you
The repository was last pushed on 2026-06-08, which is more than three months before today, and it is not archived. There are no releases, so there is no version to pin and no changelog to read. The top level holds README.md and an images directory, and no licence file appears there, so the terms under which you may reuse the code are not stated in the repository itself; if you plan to copy the script into something you ship, that is worth resolving before you do, and it is a question for the maintainers rather than a legal conclusion you can draw from the file listing. Upgrade cost is close to zero in the sense that there is no dependency graph to manage, and high in the sense that every Amazon markup change is yours to fix. The README also links to a longer blog version of the guide, which is where any added detail lives. Note the affiliate parameters on the promotional image at the top of the README; the guide is published by a vendor that sells the alternative it recommends.
Editorial conclusion
Adopt this repository if you want a readable walkthrough of Amazon product page structure and a working requests plus BeautifulSoup starting point for a small, low-volume job, and if you accept that the tutorial stops before proxies, retries or browser rendering. Do not adopt it if you need a maintained production pipeline, because the README documents one request path and then hands the harder cases to a paid API, and the repository has no releases and no licence file in its top level. Before you build on it, verify two things yourself: the CSS selectors the guide relies on against a live Amazon page today, since Amazon markup changes, and the terms that apply to the data you intend to collect.
Frequently asked questions
Is it legal to scrape data from Amazon?
The repository does not answer this. The README documents that Amazon blocks automated requests and returns an error, and it points to a contact address for automated access discussions, but it makes no statement about legality or terms of use.
How can I scrape product data from Amazon with Python?
The README installs requests, beautifulsoup4, lxml and pandas, sends a GET request with a user-agent and accept-language header to a product URL, then locates the title, rating, price, image and description in the returned HTML and exports the rows to CSV with pandas.
Is web scraping legal or illegal?
The repository does not address this question. It covers request mechanics, the 503 response Amazon returns to non-browser requests, and how to send browser-like headers instead.
Is AI scraping illegal?
The repository does not address this question. Its scope is a Python walkthrough of extracting product name, rating, price, images and description from an Amazon product page with requests and BeautifulSoup.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/oxylabs-how-to-scrape-amazon-product-data)