AutoScraper: teaching a Python scraper by example instead of by selector
A Smart, Automatic, Fast and Lightweight Web Scraper for Python
At a glance
- What is it?
- AutoScraper learns extraction rules from a list of sample values you paste in, then reuses them on other pages. It is a small, MIT-licensed Python library, last pushed on 2026-07-29, and it fits one narrow job well.
- Who is it for?
- Adopt AutoScraper when the page you are scraping has a stable HTML shape and you want extraction rules without writing CSS or XPath by hand. Do not adopt it for pages that render their content with JavaScript, because the library works from the URL or the HTML you hand it, and the README's build signature takes a url or an html string, not a browser.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 63 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What AutoScraper removes from a scraping job
The usual Python scraping loop is: open the page, inspect the DOM, write a CSS or XPath selector, run it, watch it return the wrong nodes, edit the selector. AutoScraper replaces the selector-writing step. You give it a URL plus a list of values you already know are on that page, and it works out the rules itself. The README states the project "gets a url or the html content of a web page and a list of sample data which we want to scrape from that page" and that this sample data "can be text, url or any html tag value of that page."
The audience is narrow and specific. It suits someone who needs a handful of fields from a family of pages that share a template, and who would rather not maintain selectors when the markup shifts slightly. It does not suit someone who needs to click through a login, wait for a chart to render, or paginate an infinite scroll. The library is a parsing and rule-learning layer, not a browser.
How the rule learning works in practice
The core object is AutoScraper. You call build(url, wanted_list) and it fetches the page, locates each value you named, and derives a set of matching rules from the surrounding structure. Those rules are the learned model. You then call get_result_similar on a different URL to pull values that look like the samples, or get_result_exact to pull the same positions in the same order as your wanted_list.
That split matters. get_result_similar is the fuzzy mode: it returns everything the learned rules consider a sibling of your examples. get_result_exact is the positional mode, and the README is explicit that it "will retrieve the data as the same exact order in the wanted list." If you want a title, a star count and an issues URL from a GitHub repository page, you pass all three in one wanted_list and use the exact mode. If you want every related question title on a Stack Overflow page, you pass one title and use the similar mode.
The dependency list is short: requests, bs4 and lxml, per setup.py. There is no headless browser, no Selenium, no Playwright. That is the whole architecture, and it is why the package installs quickly and why it cannot see anything JavaScript produced after load.
Installing AutoScraper and running a first build
The README gives three install routes. The PyPI route is the one most people want, and it is a single command:
pip install autoscraperAfter that, the README's first example builds a scraper against a Stack Overflow question page and prints the related titles it found:
from autoscraper import AutoScraper
url = 'https://stackoverflow.com/questions/2081586/web-scraping-with-python'
wanted_list = ["What are metaclasses in Python?"]
scraper = AutoScraper()
result = scraper.build(url, wanted_list)
print(result)What you should see is a Python list of question titles, with your sample among them. The README's own output shows nine titles, and the sample string appears third in that list. If you get an empty list, the most likely cause is that the value you put in wanted_list is not present in the HTML that requests received.
Once a build works, the README shows how to persist it so you do not refetch and relearn on every run:
scraper.save('yahoo-finance')
scraper.load('yahoo-finance')The README does not document the file format that save writes, nor how many files it produces, so treat the saved artifact as opaque and version it alongside your code.
Passing proxies and headers without forking the library
AutoScraper forwards request arguments rather than hiding them. The README shows a proxies dictionary passed through build via request_args:
proxies = {
"http": 'http://127.0.0.1:8001',
"https": 'https://127.0.0.1:8001',
}
result = scraper.build(url, wanted_list, request_args=dict(proxies=proxies))Because request_args is handed to requests, the same channel carries custom headers, which is the practical answer to sites that reject the default user agent. The README does not state that headers are supported explicitly; it says you "can also pass any custom requests module parameter," and shows proxies as the example. Read that as a general escape hatch, and test the specific parameter you need.
There is a second input path worth knowing about. build accepts html=html_content instead of a url, per the README's parenthetical in the Yahoo Finance example. That is the hook for anyone who wants to fetch with their own client and only use AutoScraper for the parsing step.
Where AutoScraper breaks down
The clearest limitation is JavaScript. The README's Yahoo Finance example carries a warning that "you should update the wanted_list if you want to copy this code, as the content of the page dynamically changes." That is a polite way of saying the sample values are transient. A price that renders client-side will not be in the HTML that requests returns, and build will find nothing to learn from. For that class of page you need a rendering layer in front, and AutoScraper's own documentation does not provide one.
The second limitation is model fragility. The learned rules are derived from one page's structure. If the target site changes its markup, the model does not raise a clear error, it returns fewer or different values. The README does not describe any validation step, confidence score, or drift detection. You are responsible for checking output length and shape yourself.
The third is scope. This is a library, not a service. There is no scheduler, no queue, no retry policy, no rate limiter, and no storage layer in the README. Everything around the extraction is yours to build. The repository does include a tests directory, but the README does not describe what those tests cover, so do not read their existence as a coverage guarantee.
AutoScraper against a selector-based parser
The natural alternative is writing selectors directly against BeautifulSoup or lxml, both of which are already AutoScraper's dependencies. The difference is where the effort sits. With BeautifulSoup you decide the rule and the machine applies it; the rule is explicit in your source, reviewable, and fails loudly when soup.select returns an empty list. With AutoScraper the machine proposes the rule and you accept it; the rule is implicit in a saved model, and the failure mode is a shorter list rather than an exception.
That trade favors AutoScraper when the page structure is awkward or when you are scraping many similar pages whose exact class names you do not want to hardcode. It favors hand-written selectors when the field is business-critical, when you need to know exactly why a value was extracted, or when the page changes often enough that you want a loud failure. A middle path exists and the README supports it: fetch with your own client, pass html= into build, and keep the network layer under your own control while letting AutoScraper handle extraction.
Licence, upkeep and what the release history says
The project is MIT licensed, per both the LICENSE file and the license field in setup.py, which also declares Development Status :: 4 - Beta and python_requires >=3.6. MIT is permissive: you can use the library in closed-source products, and the only real obligation is keeping the copyright notice. This is a description of the licence terms, not legal advice; check the LICENSE file yourself before shipping.
Upgrade cost is the part to weigh. The last push to the repository was on 2026-07-29, so the codebase is not dormant, but the most recent release in the list is v1.1.14 from 2022-07-17, with v1.1.12 and v1.1.11 before it in 2021. That gap between commits and tagged releases means anyone pinning to a version number is pinning to code from 2022. The README's own install instructions include a git route, pip install git+https://github.com/alirezamika/autoscraper.git, which is how you would get post-release changes. The README does not document a changelog, a deprecation policy, or what changed between 1.1.11 and 1.1.14, so budget time for reading the diff if you take the git route.
Editorial conclusion
Adopt AutoScraper when the page you are scraping has a stable HTML shape and you want extraction rules without writing CSS or XPath by hand. Do not adopt it for pages that render their content with JavaScript, because the library works from the URL or the HTML you hand it, and the README's build signature takes a url or an html string, not a browser. Before relying on it, verify three things: that your wanted_list values actually appear in the fetched HTML, that the built model survives a second URL, and that scraper.save writes the files you expect. The last push was on 2026-07-29, so the repository is not abandoned, but the most recent release listed is v1.1.14 from 2022-07-17, and the README does not document a deprecation or migration path for the 1.x model files.
Frequently asked questions
Is web scraping legal or illegal?
The README and the rest of the repository material say nothing about legality, terms of service, or robots.txt. The project ships under the MIT licence, which covers the code, not the sites you point it at.
Do hackers use web scraping?
Nothing in the README, setup.py or the repository layout addresses this. AutoScraper is a general-purpose extraction library and the documentation describes only benign usage such as reading question titles, stock prices and repository metadata.
Can web scraping be detected?
AutoScraper does not address detection at all; the README's only related mechanism is passing custom requests parameters such as proxies and headers through request_args. Anything beyond that is outside what the documentation covers.
What are the risks of web scraping?
The README does not discuss risks. The closest thing to a documented hazard is its warning that the Yahoo Finance example's wanted_list must be updated because the page content changes dynamically, which is a correctness risk rather than a legal or operational one.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/alirezamika-autoscraper)