Model or dataset
ScrapingBee/chatgpt-scraper-api avatar
ScrapingBee/chatgpt-scraper-api

chatgpt-scraper-api is one README file and four parameters

Collect structured responses from a ChatGPT scraper by sending a prompt with valid ChatGPT scraping API credentials. Enable live search, inject HTML context, and automate intelligent scraper ChatGPT workflows in just a few parameters.

1,098 stars1 forksUnknownLicense varies

At a glance

What is it?
Documentation for a hosted endpoint that takes a prompt and three options and returns structured JSON, published as a single file in a repository with no code, no licence and no tests. The parameter list explains how to inject a target page into the model's context without saying which parameter names that page, and the JSON shape you get back rests on a sentence in your own prompt.
Who is it for?
chatgpt-scraper-api fits someone who wants to try a model-driven extraction endpoint quickly and is comfortable treating the response as untrusted until they validate it, because the whole interface is four parameters and a bearer token. It does not fit a team that needs a versioned client library, a documented error contract or a licence to copy code from, since none of those are in the repository.
Can I use it commercially?
Not without permission. GitHub finds no licence file in the repository, and without a licence all rights are reserved by default: you may read the code but not reuse it. Check the README, or ask the authors, before using it.
Is it still maintained?
Yes. The repository last received commits 82 days ago.
What is it written in?
GitHub does not report a main language for this repository.

Answers come from the project's GitHub data, last synced on October 4, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The repository contains exactly one file

The top level of the project is a single entry: `README.md`. There is no source directory, no package manifest, no test suite and no licence file, and the licence field recorded for the repository is empty. What the file is, then, is a documentation post for a hosted product, and it is written like one. The same keyword appears in several grammatical forms across the page, describing a ChatGPT scraper, a scraper ChatGPT integration and chatgpt scraping workflows, and the opening line calls the project production-ready while the closing line calls it a solid technical foundation. The product itself lives elsewhere, at a documented endpoint, and the project's own homepage field points at the vendor's feature page rather than at the code. One practical consequence for anyone reading this as an engineering guide: the code samples in the file carry no stated licence, since there is no licence to read, and there is no version history to pin, because the repository publishes no releases.

Four parameters, and none of them names the page

The request flow is drawn as a single line, from client to API to optional web search to HTML retrieval to AI processing to a JSON response, and the behaviour is controlled by four parameters. `prompt` says what to analyse or extract. `search`, when true, makes the API run a live web search before generating the response. `country_code` sets the geolocation of those results. `add_html` injects raw HTML from the target page into the model's context. Three of those four are fully specified, and the fourth is where the documentation stops. Nothing on the page takes a URL. The examples pass a prompt such as extracting a product name, a price and an availability, and switch `add_html` on, but no parameter identifies which page to fetch, so either the address is embedded in the prompt text or the retrieval step works on whatever the search returned. For a reader writing the request, that is the first thing to test, because it decides whether you can point the endpoint at a specific URL at all.

With both retrieval switches off, nothing is scraped

The page describes the `search` option's negative case explicitly: when it is false, the model answers using its existing knowledge and the context provided to it. Combined with `add_html` left off, which is what the first code example does, the call is a language model request with no retrieval in it at all. Nothing is fetched, nothing is injected, and the answer comes from the weights. That is worth knowing for two reasons. The same endpoint and the same credentials serve both a scraper and a chat completion, so a misconfigured parameter does not fail loudly, it returns a plausible answer about a product it never looked at. And a monitoring job built on this that silently loses its HTML injection will produce confident, well-formed, wrong data rather than an error. The use-case list on the page leans on exactly the mode that needs the switches on, and the examples split between the two.

The Node path hides the endpoint, the Python path shows it

Two integrations are documented and they are not symmetrical. The Node route goes through the vendor's own package:

bash
npm install scrapingbee
javascript
const { ScrapingBeeClient } = require('scrapingbee');
const client = new ScrapingBeeClient('YOUR_API_KEY');

Note what that package is called and what it is not. The repository is chatgpt-scraper-api, and the module you install is `scrapingbee`, the general client for the vendor's other products, so the ChatGPT method is one method on a broader client rather than a library dedicated to this feature. It also means the Node example never shows a host, a route or a content type, because the SDK holds those. The Python example skips the SDK entirely and posts to the endpoint directly, which is the only place the actual URL appears on the page:

python
url = "https://app.scrapingbee.com/api/v1/chatgpt"

with the key as a bearer token and a JSON content type. If you are writing in a third language, the Python example is the complete spec.

Return JSON is a request in your own prompt, not a declared schema

The extraction examples put the output contract in the prompt text:

code
Extract:
- product_name
- price
- availability
Return JSON.

and the example response shows what comes back:

json
{
  "product_name": "Wireless Headphones",
  "price": 129.99,
  "availability": "In Stock"
}

The two blocks line up, which makes the flow look deterministic, and it is the one place in the file where an assumption is doing quiet work. Nothing on the page states that the endpoint is called with a structured-output or JSON-schema mode, and nothing describes a field type, a required key, a null convention or what happens when a field is absent from the page. The example values are also the canonical illustration rather than a captured response, so they are not evidence about formatting. The engineering consequence is straightforward: parse the response as untrusted, validate the keys you asked for, and decide what a missing field means, because the model is following an instruction rather than satisfying a contract.

Four status codes and one sentence of production advice

The error section is short and useful only as a starting list. Four responses are named: 401 for an invalid API key, 403 forbidden, 429 when a rate limit is exceeded, and 500 for a server error. The advice that follows is to implement retry logic and monitor usage limits in production. What is absent is the part that makes retry logic safe to write. There is no error body shape, so there is nothing to log or parse to tell a bad key from an exhausted quota, no guidance on which codes are worth retrying and which are terminal, no timeout, and no numbers behind the 429, so nothing to back off proportionally. The geotargeting option is documented with the same pattern, with `country_code` set to a two-letter code and used for region-specific pricing, local search analysis and country-based competitor monitoring, which tells you the option exists without telling you how many regions are supported. Pricing, credits and quotas are absent from the page entirely, so the cost of a search-enabled call with HTML injection is a question for the vendor's own pages rather than for this repository.

Editorial conclusion

chatgpt-scraper-api fits someone who wants to try a model-driven extraction endpoint quickly and is comfortable treating the response as untrusted until they validate it, because the whole interface is four parameters and a bearer token. It does not fit a team that needs a versioned client library, a documented error contract or a licence to copy code from, since none of those are in the repository. Four things to check before you build on it: how you identify the page, since the option that injects its HTML does not take a URL and the page must arrive some other way; what a request costs, because the page never mentions pricing or credits; what happens with both search and HTML injection off, which is a model call with no retrieval at all; and how you validate the output, because the structured response is produced by a model obeying an instruction to return JSON rather than by a schema you declared. The vendor sells the service and the vendor's product page is the project's homepage, so terms and pricing live there rather than in the repository. The last commit is dated July 15, 2026 and there are no releases.

Frequently asked questions

What is chatgpt-scraper-api?

A documentation repository for a hosted endpoint that takes a prompt plus three options and returns structured JSON, intended for model-driven web extraction. The repository itself contains only a README file, with no code, no tests and no stated licence.

How do I call the chatgpt-scraper-api?

In Node, install the `scrapingbee` package, create a `ScrapingBeeClient` with your key and call `client.chatGPT` with a prompt and a params object. The Python example skips the SDK and posts JSON to `https://app.scrapingbee.com/api/v1/chatgpt` with a bearer token.

What do the chatgpt-scraper-api parameters do?

`prompt` defines what to extract, `search` triggers a live web search before the answer is generated, `country_code` sets the geolocation of those results, and `add_html` injects the target page's raw HTML into the model's context. No parameter carries the target URL.

Is chatgpt-scraper-api free?

The page says nothing about pricing, credits or the cost of a request. The only rate-related information is a 429 response for exceeding limits, so the cost model has to be checked on the vendor's own product pages.

What errors does the chatgpt-scraper-api return?

Four are listed: 401 for an invalid API key, 403 forbidden, 429 for a rate limit exceeded and 500 for a server error. The only production guidance is to implement retry logic and monitor usage limits, with no error body shape, timeout or retry policy given.

Official sources

  1. Issues
  2. Project website
  3. README
  4. ScrapingBee/chatgpt-scraper-api on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/scrapingbee-chatgpt-scraper-api.svg)](https://hysenlabs.com/projects/scrapingbee-chatgpt-scraper-api)