Model or dataset
oxylabs/perplexity-scraper avatar
oxylabs/perplexity-scraper

oxylabs/perplexity-scraper: a hosted Perplexity API for structured answers

Perplexity Scraper Track brand mentions, analyze rankings, and gain competitor intelligence from Perplexity. Get started in minutes.

3,029 stars7 forksJavaLicense varies

At a glance

What is it?
The repository is a thin client for Oxylabs' Web Scraper API, not a self-hosted scraper. It returns Perplexity answers as parsed JSON, Markdown, HTML or PNG, and every request depends on an Oxylabs account.
Who is it for?
Adopt it if you already pay for Oxylabs' Web Scraper API and want Perplexity answers plus structured sources, related queries and tab metadata without running your own browser fleet. Do not adopt it as a free or self-hosted scraper: the repository contains examples, not a scraping engine, and the README points to the Oxylabs dashboard for a trial.
Can I use it commercially?
Not without permission. GitHub finds no licence file in the repository, and without a licence all rights are reserved by default: you may read the code but not reuse it. Check the README, or ask the authors, before using it.
Is it still maintained?
Yes. The repository last received commits 14 days ago.
What is it written in?
Mainly Java, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What oxylabs/perplexity-scraper actually is

The name suggests a scraper you clone and run. The repository is closer to a set of request recipes. Its top level holds a README, a handful of PNG assets, one full output sample, and a directory called Code examples. The primary language is Java, and the README's own sample is Python. There is no build file, no CLI entry point, no Docker image described, and no release has been published. What you get is documentation of an HTTP contract: post a payload to the Oxylabs Web Scraper API, get back the rendered Perplexity page in whichever format you asked for.

That distinction matters for who this is for. If your job is to track how a brand appears inside Perplexity answers, or to collect the sources Perplexity cites for a topic, the repository tells you the exact field names to expect. If your job is to run scraping infrastructure on your own machines, this repository does not help. The README states that the API handles rendering, parsing and delivery, and that you do not manage proxies or browsers. Those responsibilities sit with Oxylabs, and so does the bill.

The request and response mechanism

Every call is a POST to a single endpoint with a JSON payload. Two parameters are mandatory: source, which selects the Perplexity scraper, and prompt, the question sent to Perplexity, capped at 8000 characters. Everything else is optional. parse controls whether you receive parsed data, and it defaults to false, which is the first trap: send a payload without it and you get an HTML document rather than the JSON structure the README spends most of its length describing.

The remaining optional parameters shape where and how the request runs. geo_location localizes the prompt by country, callback_url turns the call into a push model where results land on your own endpoint, and browser_instructions passes custom instructions when JavaScript is rendered. Authentication is HTTP basic auth against your Oxylabs credentials.

The parsed output is the interesting part. A result carries prompt_query, the model name such as turbo, an answer_results array that the README describes as a Markdown JSON tree, and answer_results_md holding the same answer as a Markdown string. Alongside the answer sit additional_results, which groups UI elements like sources_results, images_results and shopping_results, plus top_sources, top_images, inline_products, related_queries and displayed_tabs. Sources are objects with url and title. The README is explicit that the number of items and fields can vary with the prompt, so any parser you write against this shape needs to tolerate missing keys.

Installing it and sending a first Perplexity prompt

There is nothing to install in the usual sense. The README directs you to register on the Oxylabs dashboard for a trial, then call the API with your credentials. The Python sample below is the one the README gives, with the payload structure unchanged. Replace USERNAME and PASSWORD with your Oxylabs credentials.

python
import requests
from pprint import pprint

payload = {
    'source': 'perplexity',
    'prompt': 'top 3 smartphones in 2025, compare pricing across US marketplaces',
    'geo_location': 'United States',
    'parse': True,
    'callback_url': 'https://your-server.com/oxylabs-callback'
}

response = requests.post(
    'https://data.oxylabs.io/v1/queries',
    auth=('USERNAME', 'PASSWORD'),
    json=payload
)

pprint(response.json())

Run it and you should see a JSON body whose results array contains a job_id, a status_code of 200, the Perplexity search url, and a content object. Inside content you get prompt_query, model, answer_results, answer_results_md, additional_results with sources_results, and related_queries. The repository's output-perplexity-scraper.json file shows the full shape if you want to inspect it before writing a parser.

The callback_url in that sample points at a placeholder host. Leave it out if you want a synchronous pull response, or set it to a real endpoint you control. The README links to a separate page on push and pull integration for the callback contract, and does not describe retry behaviour for failed deliveries.

Where the hosted model constrains you

The clearest limitation is that nothing here works without an Oxylabs subscription and network access to data.oxylabs.io. There is no offline mode, no local rendering, and no fallback. If your environment blocks outbound calls to that host, the repository is inert.

The second constraint is the prompt ceiling. The README sets a maximum of 8000 characters for prompt, and does not document what happens when you exceed it. Long research questions with pasted context will need trimming before they reach the API.

The third is output variability. Because the number of items and fields depends on the prompt, a schema you validate strictly against one response may reject the next. Fields like inline_products only appear when a product-related query triggers shopping results, and top_images only when the Images tab is populated. Code that assumes a fixed set of keys will break on ordinary prompts.

Finally, parse defaults to false. A team that omits it gets HTML and may conclude the structured output does not exist, when the README documents the opposite. This is a documentation default that costs debugging time.

Compared with running your own headless browser

The obvious alternative is a self-hosted headless browser stack, for example Playwright or Selenium driving a real browser against perplexity.ai, with a parser built on top. The difference is where the work sits. A self-hosted stack gives you control over request pacing, session handling and the exact browser version, and it costs only compute. It also puts the anti-bot problem, the proxy pool and the rendering failures on your team, which the README names as the things the Oxylabs API removes.

The trade is inverted for data shape. A self-hosted scraper returns whatever the page renders, and you write the extraction logic. The Oxylabs API returns named fields such as answer_results_md, sources_results and related_queries already separated, which is the part that is genuinely tedious to maintain by hand. If your need is one-off and small, a browser script is cheaper. If your need is a recurring feed of Perplexity answers with their cited sources, the hosted contract is the shorter path, provided you accept the dependency.

Maintenance, licensing and what the repository leaves open

The repository is not archived and the last push was on 2026-09-16, so it is current. That recency reflects example upkeep rather than a versioned library: no releases have been published, so there is no changelog to track and no pinned version to upgrade. Your upgrade surface is the API itself. When Oxylabs changes parameter names or the parsed output shape, the README and the code examples are what move, and your integration follows. Pin nothing, because there is nothing to pin.

The licence is not stated anywhere in the repository, which is a real gap for a repository you might copy code from. Treat the examples as documentation of an API contract rather than as licensed source until you confirm terms with Oxylabs. That is a factual observation about what the repository declares, not a legal opinion.

The README is also silent on rate limits, quota accounting, error codes beyond the status_code field in a result, and the retry semantics of callback_url deliveries. Those are the questions to put to Oxylabs support before a production integration, because the repository cannot answer them.

Editorial conclusion

Adopt it if you already pay for Oxylabs' Web Scraper API and want Perplexity answers plus structured sources, related queries and tab metadata without running your own browser fleet. Do not adopt it as a free or self-hosted scraper: the repository contains examples, not a scraping engine, and the README points to the Oxylabs dashboard for a trial. Before committing, verify your prompt fits the 8000-character limit, confirm whether you need parse set to true or the Markdown field, and check the licence situation, which the repository does not state.

Frequently asked questions

Is AI scraping illegal?

The repository does not address legality. It documents an HTTP API contract and leaves compliance questions to the user.

Is web scraping legal or illegal?

The README says nothing about the legal status of scraping. It describes the Perplexity scraper as a way to collect AI-generated responses and structured metadata through the Oxylabs Web Scraper API.

Which AI is best for scraping data?

The repository only covers Perplexity. Its scraper sends a prompt to Perplexity and returns the answer as HTML, parsed JSON, Markdown or a PNG, with no comparison to other models.

Is scraping with BeautifulSoup legal?

The repository does not discuss BeautifulSoup or any other parsing library. Its examples call the Oxylabs Web Scraper API over HTTP.

Official sources

  1. Issues
  2. oxylabs/perplexity-scraper on GitHub
  3. Project website
  4. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/oxylabs-perplexity-scraper.svg)](https://hysenlabs.com/projects/oxylabs-perplexity-scraper)