# oxylabs/ai-crawler-py: Natural Language Web Extraction Without Custom Selectors

> AI-Crawler is an experimental Python SDK from Oxylabs AI Studio that crawls websites and extracts data using a natural language prompt, returning results as structured JSON or Markdown. The repository is archived and depends on a commercial API key, making it a client for a hosted service rather than a self-contained open-source tool.

**oxylabs/ai-crawler-py** — Crawl a website starting from a URL, find relevant pages, and extract data – all guided by your natural language prompt.

- Repository: https://github.com/oxylabs/ai-crawler-py
- Website: https://aistudio.oxylabs.io/apps/crawl
- Stars: 3,263 · Forks: 12
- Language: Unknown
- License: not declared
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/oxylabs-ai-crawler-py

## Natural Language Prompts Replace CSS Selectors

Traditional web scrapers like Scrapy or BeautifulSoup require you to write CSS or XPath selectors that target specific HTML elements. When a site redesigns its markup, every selector breaks and the scraper stops working until someone manually updates it. AI-Crawler takes a different approach: you supply a natural language prompt describing what data you want, and a crawl agent uses that description both to select which pages to visit and to extract the matching content from each one.

The tool is positioned for developers and data scientists who want structured output from public websites without building per-site scraper logic. The README describes the goal as letting users focus on analysis rather than maintaining custom scrapers. You provide a starting URL, write a plain-English description of the content you are after, choose an output format, and optionally supply a schema to shape the JSON output.

The key trade-off is control. Traditional scrapers put you in charge of exactly which URLs get visited and what gets extracted, at the cost of writing and maintaining selectors. AI-Crawler delegates those decisions to a server-side AI model, eliminating selector maintenance but introducing unpredictability: the same prompt against the same site may produce different results on different runs, and there is no documented way to inspect or override the URL selection logic.

## How the Crawl Pipeline Works

The README describes a four-step process. First, provide a starting URL. Second, write a natural language prompt describing the content you want. Third, choose an output format: JSON or Markdown. Fourth, if you chose JSON, supply a schema that tells the parser how to structure the extracted data.

The AiCrawler class in the oxylabs_ai_studio Python package handles both crawl orchestration and data extraction. The generate_schema() method accepts a plain-English description of the fields you need and returns an OpenAPI schema object, which you then pass to crawl() via the schema parameter. The crawler accepts a render_javascript flag (default False) for pages that require JavaScript execution, and a geo_location parameter in ISO2 format for requests through Oxylabs proxies.

The return_sources_limit parameter caps the number of source pages returned, with a default of 25. All AI reasoning about which pages are relevant runs server-side at Oxylabs. The Python SDK is a thin client that sends your prompt and parameters to the Oxylabs AI Studio API and returns the results. This means the quality and consistency of results depends entirely on the hosted service, not on anything you can inspect or tune locally.

## Installing the SDK and Running a First Crawl

The SDK requires Python 3.10 or later and an Oxylabs AI Studio API key. The README directs users to sign up at aistudio.oxylabs.io for a free trial with 1,000 credits. Install the package with pip:

```bash
pip install oxylabs-ai-studio
```

The README gives this example, which initializes the crawler, generates a schema from a plain-English description of the fields to extract, and runs a crawl against Oxylabs' sandbox product listing:

```python
from oxylabs_ai_studio.apps.ai_crawler import AiCrawler
import json

crawler = AiCrawler(api_key="your_api_key")

schema = crawler.generate_schema(prompt="want to parse name, platform, price")
print(f"Generated schema: {schema}")

url = "https://sandbox.oxylabs.io/products"
result = crawler.crawl(
    url=url,
    user_prompt="Find all Halo games for Xbox",
    output_format="json",
    schema=schema,
    render_javascript=False,
)
```

The README shows the response as a list of objects, each containing a data key with the extracted fields and a src key with the source URL. For this example, the output includes products like "Halo: Reach" and "Halo 3" paired with their source URLs:

```json
[
  {
    "data": {
      "items": [
        {"name": "Halo: Reach", "platform": "Xbox platform", "price": 84.99}
      ]
    },
    "src": "https://sandbox.oxylabs.io/products/141"
  }
]
```

Passing output_format="markdown" returns page content as Markdown rather than parsed JSON. The schema parameter is described as mandatory for JSON output; without it, structured field extraction is undefined.

## Request Parameters and the Return Sources Limit

The crawl() method accepts seven parameters, two of which are mandatory. The url parameter is the starting URL the crawler explores, and user_prompt is the natural language instruction that guides both page selection and content extraction. The output_format parameter accepts "json" or "markdown" and defaults to "markdown" when omitted.

The schema parameter is an OpenAPI schema object. For JSON output you either write one yourself or use generate_schema() to produce it from a description. The render_javascript flag defaults to False; setting it to True enables JavaScript rendering for single-page applications and dynamically loaded content. The geo_location parameter accepts a two-letter ISO country code and routes requests through Oxylabs proxies in that geography.

The return_sources_limit parameter caps how many source pages the crawl returns, defaulting to 25. Raising that limit or the request rate requires a higher-tier plan. There is no parameter documented in the README for controlling crawl depth, following external links, or setting a timeout. The README also does not document how credits are consumed per request or whether a failed extraction still costs credits.

## Where AI-Crawler Falls Short

The tool is unsuitable for any use case that requires deterministic, reproducible extraction. Because the URL selection and content parsing happen inside AI models running server-side, there is no guarantee that repeated runs of the same prompt against the same site will return the same pages or the same fields. This makes it difficult to build reliable data pipelines on top of AI-Crawler.

The credits model creates a hard ceiling on volume. The free trial provides 1,000 credits, and the entry-level plan gives 3,000 credits per month at $12. For projects that need to crawl thousands of pages daily, credits run out quickly. The README does not document whether credits are consumed per page visited, per request, or per API call.

The tool cannot be self-hosted. Every request goes through the Oxylabs API. Organizations with data-residency requirements or strict policies about sending web content to third-party services cannot use it. There is also no documented error recovery or retry logic for failed extractions, no way to inspect which URLs the crawler chose, and no mechanism to exclude specific URL patterns from the crawl.

## Archived Repository, Missing License, and API Dependency

The repository was archived on GitHub. The last push was on 2026-08-21. Archiving disables new issues and pull requests, so there is no official channel for reporting problems with the SDK. For a project that depends on a live API, this is particularly concerning: if the API endpoint changes or authentication requirements shift, the archived SDK cannot be patched.

The repository does not specify a license. GitHub's terms of service allow public repositories to be viewed and forked, but the absence of an explicit license means no permission is granted to use, modify, or redistribute the code commercially. Engineers who need to fork or bundle the SDK should contact Oxylabs before doing so.

As a practical comparison: Scrapy, the widely used open-source Python scraping framework, runs entirely on your own infrastructure with no API dependency or usage credits. It requires writing CSS or XPath selectors per site, but those selectors are deterministic, testable, and free to run at any volume. The choice between the two comes down to whether you prefer the maintenance burden of selectors or the unpredictability and cost of AI-assisted extraction.

## Conclusion

Engineers who need a quick, one-off extraction from a public website and already use Oxylabs services will find this SDK the fastest path to structured output: install one package, write a natural language prompt, and receive JSON without writing any selectors. It is the wrong tool for production pipelines that need deterministic results, self-hosted infrastructure, or high-volume extraction within a tight credits budget. The archived repository means no future bug fixes or API compatibility updates. Before adopting it, verify that the Oxylabs AI Studio API still accepts new registrations and confirm how credits are counted per crawl, since the archived SDK cannot be patched with clarifications.

## FAQ

### What does an AI crawler do?

An AI crawler uses AI models to identify and visit pages relevant to a natural language prompt, then extracts content from those pages without requiring manually written CSS or XPath selectors. Oxylabs AI-Crawler returns results as structured JSON using a schema you provide, or as Markdown.

### What Python version does oxylabs/ai-crawler-py require?

The README specifies Python 3.10 or later. The SDK is installed as the oxylabs-ai-studio package via pip.

### Does AI-Crawler require a paid Oxylabs account?

Yes. The tool requires an Oxylabs AI Studio API key to operate. A free trial provides 1,000 credits; the entry-level paid plan starts at $12 per month for 3,000 credits at a rate of 1 request per second. The README does not document a permanent free tier.

## Sources

- [Issues](https://github.com/oxylabs/ai-crawler-py/issues)
- [oxylabs/ai-crawler-py on GitHub](https://github.com/oxylabs/ai-crawler-py)
- [Project website](https://aistudio.oxylabs.io/apps/crawl)
- [README](https://github.com/oxylabs/ai-crawler-py/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/oxylabs-ai-crawler-py
