Open-source project
oxylabs/web-scraper-api avatar
oxylabs/web-scraper-api

Oxylabs Web Scraper API: One Request for Proxies, CAPTCHA, Rendering, and Parsing

All-in-one web scraping API for real-time, large-scale data extraction – proxies, CAPTCHA management, JS rendering, and parsing in a single request returning HTML or structured JSON.

494 stars1 forksUnknownLicense varies

At a glance

What is it?
Oxylabs Web Scraper API is a commercial service that consolidates proxy rotation, CAPTCHA handling, JavaScript rendering, and structured data parsing behind a single HTTP endpoint. It serves engineering teams building high-volume, production-grade data extraction pipelines who want to avoid managing separate infrastructure for each stage of the scraping stack.
Who is it for?
Engineering teams that need reliable, at-scale data extraction from public websites without operating their own proxy infrastructure, CAPTCHA solver, headless browser fleet, and parser pipeline will find the consolidation in Web Scraper API useful.
Can I use it commercially?
Not without permission. GitHub finds no licence file in the repository, and without a licence all rights are reserved by default: you may read the code but not reuse it. Check the README, or ask the authors, before using it.
Is it still maintained?
Yes. The repository last received commits 14 days ago.
What is it written in?
GitHub does not report a main language for this repository.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What Web Scraper API Consolidates and Who Uses It

A typical self-managed scraping stack requires at minimum: a proxy provider for IP rotation, a service or logic for CAPTCHA solving, a headless browser or rendering service for JavaScript-heavy pages, and a parser to turn raw HTML into structured data. Each component has its own failure modes, rate limits, and maintenance surface.

Web Scraper API replaces all four with a single endpoint. Send a request specifying a target URL or query. The API handles proxy rotation using Oxylabs' proxy network, manages access interruptions and CAPTCHA challenges automatically, optionally renders JavaScript, and optionally parses the response into structured JSON. The caller receives either raw HTML or ready-to-use JSON.

The README describes the product as built to meet enterprise standards, including SOC 2 Type II compliance. The stated use cases are search engine data, e-commerce pricing and product data, travel platform data, real estate listings, and generic web pages. The infrastructure is maintained by Oxylabs' in-house engineering team.

Dedicated Sources and Parsers for Major Targets

Beyond the universal source that works with any public URL, Web Scraper API ships with dedicated sources tuned for specific platforms. Dedicated sources return structured JSON without requiring the caller to write or maintain parsing logic.

Search engine sources cover Google Search, Google Ads, Google Trends, Google Lens, Google Scholar, and Bing. AI platform sources include ChatGPT, Perplexity, and Google AI Mode. E-commerce sources include Amazon, Walmart, eBay, Etsy, Best Buy, Target, Costco, AliExpress, Alibaba, Lazada, Flipkart, and MercadoLibre, among others. The full list is at developers.oxylabs.io/api-targets.

For dedicated sources, the caller can pass either a full URL or parameterized inputs: a search query, a product ID, a video ID, and similar. Setting parse: true in the request body returns structured JSON. Without it, the API returns raw HTML.

The universal source works for any public website. When using it, the caller is responsible for any further parsing or structuring of the HTML response, unless a Custom Parser is configured.

JavaScript Rendering and Browser Instructions

For dynamic pages that require JavaScript to load content, the API's Custom Browser feature handles rendering. Setting render: html in the request body returns the rendered HTML. Setting render: png returns a Base64-encoded screenshot.

For pages that require user-like interaction before the content is accessible, the browser_instructions parameter accepts a list of interaction steps. The README shows an example that types a query into an eBay search field, clicks submit, and waits for results:

json
{
  "source": "universal",
  "url": "https://www.ebay.com/",
  "render": "html",
  "browser_instructions": [
    { "type": "input", "value": "pizza boxes", "selector": { "type": "xpath", "value": "//input[@class='gh-tb ui-autocomplete-input']" } },
    { "type": "click", "selector": { "type": "xpath", "value": "//input[@type='submit']" } },
    { "type": "wait", "wait_time_s": 5 }
  ]
}

Browser instructions support input, click, and wait step types (among others, as the full list is at developers.oxylabs.io). The Oxylabs dashboard provides a Playground for generating and testing browser instructions without manual coding.

Geolocation, Localization, and Custom Parsing

The geo_location parameter selects the proxy server's exit location, so the response reflects what a user in that country, state, or city would see. For SERP sources, this adjusts the search results page. For e-commerce sources, specific geo values are required by certain marketplaces; the per-target documentation covers the valid values.

Additional localization parameters include domain (the top-level domain to query), locale (the interface language), and results_language (the language of the returned results). These allow region-specific extraction without managing multiple proxy locations manually.

Custom Parser is a free feature that lets callers define their own field extraction logic on top of a raw scraping result. The parsed output contains only the fields specified, in a structured format. Parser Presets save and reuse parsing configurations across jobs. Both are configurable through the API or the Oxylabs dashboard.

Scheduler, OxyCopilot, and Output Delivery

The Scheduler feature runs batch or recurring scraping jobs without the caller needing to manage timing logic. A job can be set up to run once or on a schedule, and results are delivered to a configured storage endpoint.

OxyCopilot is an AI assistant built into the Oxylabs dashboard. The README describes it as helping with setup. It does not provide a separate API; it is a configuration aid within the dashboard interface.

Output delivery supports two modes. Pull mode means the caller polls the API to collect results after the job completes. Push mode delivers results to a storage destination the caller provides. The choice depends on whether the workload is synchronous (one request, one immediate response) or asynchronous (large batch, collect later).

Rate limits and batch request caps are documented at the API level but not specified with numeric values in this repository's README. The full rate limit documentation and output delivery options are at developers.oxylabs.io.

Where Web Scraper API Is the Wrong Tool

Web Scraper API is a commercial service with a usage-based pricing model. The repository contains documentation, code examples, and integration guides, not source code for the service itself. Teams with open-source requirements, a need to run the scraper on-premises, or a tight per-request budget that makes a managed service impractical should evaluate other options.

For straightforward scraping of static pages that do not require CAPTCHA bypass or JavaScript rendering, a Python-based open-source stack (Scrapy, httpx, or requests) combined with a residential proxy provider chosen separately is often more cost-effective.

The dedicated source parsers cover a specific set of targets. Websites not in that list fall back to the universal source, which returns raw HTML. If the target changes its HTML structure, any custom parsing logic breaks without notice. Dedicated source parsers abstract that maintenance, but they only exist for the listed platforms.

The README also includes a comparison of Web Scraper API against Oxylabs' own Web Unblocker and Headless Browser products, noting that each fits different use cases: Web Scraper API for full data extraction with parsing, Web Unblocker for bypass without extraction, and Headless Browser for complex multi-step browser automation.

How Web Scraper API Differs from Running Playwright Directly

Playwright is an open-source browser automation library that runs a real Chromium, Firefox, or WebKit instance. A team running Playwright can render any JavaScript-heavy page, interact with it programmatically, and extract whatever data is needed. It is free to run and highly configurable.

The trade-off is infrastructure. Each Playwright instance consumes significant CPU and memory. At scale, maintaining a fleet of headless browser workers, handling their crashes and memory leaks, rotating IPs to avoid blocks, and solving CAPTCHAs requires substantial engineering investment.

Web Scraper API offloads that infrastructure. The browser_instructions parameter provides a subset of Playwright's interaction capabilities sufficient for most data extraction scenarios, without the caller needing to manage the browser fleet, the proxy pool, or the CAPTCHA solver. For teams whose core product is not scraping infrastructure but who need scraped data as an input, that trade-off often makes sense.

Editorial conclusion

Engineering teams that need reliable, at-scale data extraction from public websites without operating their own proxy infrastructure, CAPTCHA solver, headless browser fleet, and parser pipeline will find the consolidation in Web Scraper API useful. The right reason to look elsewhere is open-source requirements or tight per-request budget constraints: this is a commercial service with a usage-based pricing model, and the repository contains documentation rather than source code. Teams already using a well-maintained open-source scraping stack and needing only occasional CAPTCHA bypass may find the full API more than they need. Check the dedicated sources list at developers.oxylabs.io/api-targets before committing, since parser accuracy varies per target and new targets are added over time.

Frequently asked questions

What is the Oxylabs Web Scraper API?

Web Scraper API is an all-in-one web data collection service from Oxylabs. It combines proxy rotation, CAPTCHA handling, JavaScript rendering, and response parsing in a single HTTP endpoint, returning either raw HTML or structured JSON from any public website.

What does web scraping to API mean?

In the context of Web Scraper API, it means the scraping result is delivered as an API response: the caller sends one HTTP request with a target URL, and the API handles all scraping steps and returns the extracted data. The caller receives a structured JSON payload without managing proxies, rendering, or parsing separately.

Does Oxylabs Web Scraper API require separate proxy setup?

No. Proxy rotation is included in the service. The README describes it as an all-in-one solution where the built-in proxy rotator uses Oxylabs' proxy network and handles IP interruptions automatically as part of the request lifecycle.

Official sources

  1. Issues
  2. oxylabs/web-scraper-api on GitHub
  3. Project website
  4. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/oxylabs-web-scraper-api.svg)](https://hysenlabs.com/projects/oxylabs-web-scraper-api)