Model or dataset
itsOwen/CyberScraper-2077 avatar
itsOwen/CyberScraper-2077

CyberScraper 2077: an LLM-driven scraper with a Streamlit front end

A Powerful web scraper powered by LLM | OpenAI, Gemini & Ollama

3,273 stars360 forksPythonMIT

At a glance

What is it?
CyberScraper 2077 wraps Patchright, Tor and a choice of OpenAI, Gemini, LiteLLM or Ollama models behind a Streamlit UI. It is MIT-licensed and installs from requirements.txt, but the README leaves several operational questions open.
Who is it for?
Adopt CyberScraper 2077 if you want a Streamlit front end over an LLM extraction loop and you are comfortable running Patchright and, if needed, Tor yourself. Skip it if you need a headless library with a stable API, or if you cannot send page content to a hosted model.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 4 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What CyberScraper 2077 is for

The project targets people who want to pull structured data off web pages without writing per-site selectors. The README frames it around OpenAI, Gemini and local LLM models, with a Streamlit GUI as the main surface. The intended audience is closer to an analyst who can run a Python environment than to a backend engineer wiring a scraper into a pipeline. The repository layout supports that reading: main.py sits at the top level, app/ and src/ hold the application and logic, and tests/ exists alongside a Dockerfile and a requirements.txt that pins every dependency.

The extraction loop is model-driven. Instead of a CSS or XPath rule per field, the page content is handed to a language model, which decides what to return. That is the whole value proposition, and it is also where the cost sits. The README itself concedes the point in the Ollama section, stating that OpenAI and Gemini are recommended because those models follow instructions well, and that open-source models may need prompt tuning and extra filters. Treat that sentence as the project's own statement about where reliability comes from.

How the extraction pipeline is put together

The dependency list tells most of the story. patchright==1.57.2 is pinned with the comment that it is an undetected Playwright fork, and beautifulsoup4 plus lxml handle parsing. On top of that sit langchain, langchain-openai and langchain-google-genai, which is how prompts and model calls are routed. aiohttp, aiohttp-socks and async-lru cover async fetching and caching, and PySocks provides the Tor path. Output writers are listed too: openpyxl, xlsxwriter and Markdown, matching the JSON, CSV, HTML, SQL and Excel export claim.

So the data flow is: fetch a page through Patchright or an async HTTP client, parse it, hand the relevant text to a model through LangChain, and write the model's structured answer into one of the export formats. Caching is described in the README as content-based and query-based, using an LRU cache and a custom dictionary, which matters because every uncached page view is a paid API call. Tor support is real in the Dockerfile, not just a badge: the image installs tor and tor-geoipdb, appends SocksPort 9050 and ControlPort 9051 to /etc/tor/torrc, and sets CookieAuthentication 1.

Installing CyberScraper 2077 and running a first scrape

The README requires Python 3.10 or higher. Clone the repository, create a virtual environment, install the pinned dependencies, then install the browser that Patchright drives. Note that the README's step 4 says playwright install while the Dockerfile runs patchright install chromium; the pinned package is patchright, so the Dockerfile is the more consistent reference for what actually gets installed.

bash
git clone https://github.com/itsOwen/CyberScraper-2077.git
cd CyberScraper-2077
virtualenv venv
source venv/bin/activate
pip install -r requirements.txt
playwright install

Keys go into the environment before launch. The README shows both variables together, and the app reads them at startup.

bash
export OPENAI_API_KEY="your-api-key-here"
export GOOGLE_API_KEY="your-api-key-here"

If you would rather not use a hosted provider, the README documents a LiteLLM proxy, which lets you point the app at any model the proxy exposes. The configured models show up in the sidebar prefixed with litellm:, and the base URL defaults to http://localhost:4000/v1.

bash
export LITELLM_API_KEY="your-proxy-key"
export LITELLM_BASE_URL="http://localhost:4000/v1"
export LITELLM_MODELS="my-gpt-model,my-claude-model"

The Docker route avoids the local Python setup. The README builds the image with a tag and runs it with both keys passed through, publishing the Streamlit port.

bash
docker build -t cyberscraper-2077 .
docker run -p 8501:8501 -e OPENAI_API_KEY="your-actual-api-key" -e GOOGLE_API_KEY="your-actual-api-key" cyberscraper-2077

After that, the app is reachable on port 8501. Paste a URL, choose a model in the sidebar, and start the extraction; the result is what you then export through the format buttons. The README does not document what a failed extraction looks like in the UI, so expect to read the terminal output when a model returns something unusable.

Where CyberScraper 2077 breaks down

The captcha bypass is the clearest limitation, because the README states it directly: appending -captcha to the URL only works natively and does not work on Docker. If you deploy the container, that feature is gone. The same README describes the current browser feature, which uses your local browser instance to get past bot detection, with the instruction to use it only when necessary. That is a manual escape hatch, not something you can schedule.

Model reliability is the second failure mode, and the project names it. Open-source models through Ollama may need prompt tuning and additional filters, and the README warns that generation speed depends on your hardware. A scraper whose output quality varies with the model you picked is hard to put behind an automated job. There is also an obvious data boundary: unless you use Ollama or a self-hosted LiteLLM proxy, page content leaves your machine for a third-party API, which rules the tool out for internal or regulated pages.

Finally, the page navigation feature is labelled BETA in the README. Multi-page traversal is the part of scraping where selector-free extraction is least predictable, so treat that mode as experimental rather than as a supported workflow.

CyberScraper 2077 compared with ScrapeGraphAI

The related searches for this project keep pulling in ScrapeGraphAI, so the comparison is worth making concrete. Both are AI web scraping projects in Python, and both put a language model between the page and the extracted fields. The difference is in the delivery: CyberScraper 2077 ships a Streamlit application as its primary interface, so the expected user opens a browser, pastes a URL and reads a table. ScrapeGraphAI is oriented around being called from Python code, which suits embedding extraction into a script or a service.

That distinction drives practical choices. With CyberScraper 2077 you get a GUI, export buttons for JSON, CSV, HTML, SQL and Excel, Google Sheets upload, and a Tor route configured in the Docker image. With a library-first tool you get an importable function and whatever deployment you build around it. If your team already runs a scheduler and wants extraction as one step in a pipeline, the Streamlit surface is overhead. If you want a person to paste links and download a spreadsheet, the GUI is the point.

A second difference is provider plumbing. CyberScraper 2077 supports OpenAI, Gemini, LiteLLM and Ollama, and the LiteLLM proxy path is documented with three environment variables. That is a flexible arrangement, but it also means configuration lives in environment variables and the sidebar rather than in a config file you can commit.

Maintenance, licence and the cost of keeping it running

The repository is not archived, and the last push was on 2026-09-09. The licence is MIT, which permits commercial use and modification provided the copyright notice and permission notice are retained; that is a statement about the licence text, not legal advice, and you should read LICENSE in the repository if the distinction matters to you.

The upgrade cost is dominated by the pinned dependency set. requirements.txt fixes exact versions of streamlit, langchain, langchain-openai, langchain-google-genai, openai, patchright and roughly twenty others. LangChain in particular moves quickly, and the project pins langchain==1.2.6 with langchain-community==0.4.1 and langchain-text-splitters==1.1.0. Bumping any one of those means re-testing the extraction path, because prompt construction runs through that stack. The Patchright pin adds a second constraint: an undetected browser fork has to track upstream Chromium to keep working against sites that block automation, and the Dockerfile installs patchright install chromium at build time, so a rebuild can pull a different browser than the one you tested against.

The ongoing cost that is easy to miss is per-page model spend. Caching reduces repeat calls, but a new URL is a new call. Budgeting for that is part of adopting the tool, not an afterthought.

Editorial conclusion

Adopt CyberScraper 2077 if you want a Streamlit front end over an LLM extraction loop and you are comfortable running Patchright and, if needed, Tor yourself. Skip it if you need a headless library with a stable API, or if you cannot send page content to a hosted model. Before committing, verify the model you intend to use actually returns the structured output you need, and check how the Docker image reaches an Ollama instance on the host.

Frequently asked questions

Can web scraping be detected?

Yes, and CyberScraper 2077 treats detection as a real problem rather than a solved one. The README documents a stealth mode, an undetected Playwright fork in the dependency list, and a current browser option that uses your local browser instance, which it says helps bypass bot detections. It also notes that the captcha bypass only works natively and not on Docker.

What is the best AI scraping tool?

This cannot be ranked from the project's own documentation. What can be said is that CyberScraper 2077 uses OpenAI, Gemini, LiteLLM or Ollama for extraction, and its README recommends OpenAI and Gemini over open-source models because they follow instructions better.

Is web scraping difficult to learn?

CyberScraper 2077 lowers the barrier for the extraction step by replacing per-site selectors with a language model, and it puts a Streamlit GUI in front so the workflow is paste a URL and read a table. The setup still expects Python 3.10 or higher, a virtual environment, pip install -r requirements.txt and a browser install.

Official sources

  1. Issues
  2. itsOwen/CyberScraper-2077 on GitHub
  3. License: MIT
  4. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/itsowen-cyberscraper-2077.svg)](https://hysenlabs.com/projects/itsowen-cyberscraper-2077)