Model or dataset
oxylabs/oxylabs-ai-studio-py avatar
oxylabs/oxylabs-ai-studio-py

Oxylabs AI Studio Python SDK: prompt-driven scraping, crawling and browsing

Structured data gathering from any website using AI-powered scraper, crawler, and browser automation. Scraping and crawling with natural language prompts. Equip your LLM agents with fresh data. AI Studio python SDK for intelligent web data gathering.

3,395 stars35 forksPythonMIT

At a glance

What is it?
The oxylabs-ai-studio package wraps Oxylabs AI Studio endpoints (AI-Scraper, AI-Crawler, Browser Agent, Search, Map) behind five Python classes. It is a thin client for a paid hosted service, so the real decision is about the API key, not the code.
Who is it for?
Adopt it if you already pay for Oxylabs AI Studio or want prompt-driven extraction without maintaining your own headless browser fleet, and if a Python 3.10+ service that sends page content to a third party is acceptable. Do not adopt it if you need offline extraction, a self-hosted pipeline, or deterministic selectors that never change between runs.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
No. The owners have archived the repository on GitHub, so it is read-only and no longer receives changes.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What oxylabs-ai-studio actually solves

The package answers a narrow question: how do you get structured records out of a page whose markup you do not control and do not want to study? Instead of writing CSS or XPath selectors, you describe the fields in a prompt and let the hosted service decide where they live. The README's scraper example asks for developer, platform, type, price, game title, genre and description, then calls generate_schema on that prompt before scraping anything.

That inversion is the whole product. Selector-based scrapers break when a site ships a redesign; a prompt-driven extractor is supposed to survive it because the instruction is about meaning, not position. The cost is that you no longer know exactly which element produced a value, and the SDK does not give you a selector to inspect afterwards.

The intended audience is Python teams that already treat scraping as a data-acquisition step inside a larger pipeline: RAG ingestion, price monitoring, agent grounding. The repository topics name those use cases directly (ai-grounding, ai-training, ai-search). If you are scraping one stable page twice a month, this is more machinery than the job needs.

Five entry points, one hosted backend

The SDK exposes five apps under src/oxylabs_ai_studio, each a class constructed with api_key: AiCrawler, AiScraper, BrowserAgent, AiSearch and AiMap. There is no local extraction engine. Every call is an HTTP request to Oxylabs AI Studio, which is why httpx and tenacity sit in the dependency list: tenacity is there for retries against a remote service, not for parsing.

The classes differ by how much freedom the remote agent gets. AiScraper.scrape takes one URL and returns it in the requested format. AiCrawler.crawl starts at a URL and follows links, with return_sources_limit capping how many sources come back (default 25). BrowserAgent.run drives a page with a natural-language instruction, and the README's example tells it to use the search bar to find a game and read its price. AiSearch.search queries a search engine and optionally returns markdown content, with a maximum limit of 50. AiMap.map returns candidate URLs from a site, filtered by search_keywords and a user_prompt.

Two design details are worth noticing. First, output_format accepts json, markdown, csv, screenshot and toon, and json, csv and toon all require a schema. Second, render_javascript defaults to False everywhere, and on the scraper it can also be set to "auto" so the service decides whether rendering is needed. That default matters: a JavaScript-heavy page scraped with the default will return whatever the server sends before scripts run.

Install and a first scrape with a generated schema

The package requires Python 3.10 or above and an API key. Installation is a single pip command, and the distribution name differs from the import name: you install oxylabs-ai-studio and import from oxylabs_ai_studio.

bash
pip install oxylabs-ai-studio

After that, the fastest way to see the mechanism is the scraper with a generated schema. The README's example uses the sandbox URL https://sandbox.oxylabs.io/products/3, which is a safe target to run against while you are still deciding whether the service fits.

python
from oxylabs_ai_studio.apps.ai_scraper import AiScraper

scraper = AiScraper(api_key="<API_KEY>")

schema = scraper.generate_schema(
    prompt="want to parse developer, platform, type, price game title, genre (array) and description"
)
print(f"Generated schema: {schema}")

result = scraper.scrape(
    url="https://sandbox.oxylabs.io/products/3",
    output_format="json",
    schema=schema,
    render_javascript=False,
    optimize_content=True,
)
print(result)

What you should see: generate_schema returns a JSON schema derived from your prompt, and scrape returns a result object whose .data holds records matching that schema. If your schema and your output_format disagree, expect an error rather than a silent fallback, because json, csv and toon all require the schema argument.

For crawling, the same pattern applies with a different class and a user_prompt instead of a schema:

python
from oxylabs_ai_studio.apps.ai_crawler import AiCrawler

crawler = AiCrawler(api_key="<API_KEY>")
result = crawler.crawl(
    url="https://oxylabs.io",
    user_prompt="Find all pages with proxy products pricing",
    output_format="markdown",
    render_javascript=False,
    return_sources_limit=3,
    geo_location="United States",
)
for item in result.data:
    print(item)

Setting return_sources_limit to 3 keeps the first run cheap and makes the output easy to read before you widen it to the default 25. The repository also ships runnable files under examples/, including scrape_generated_schema.py, crawl_pydantic_schema.py and search_instant.py, which are the quickest way to see the intended call shapes without writing them yourself.

Search has two endpoints and the SDK picks for you

AiSearch.search is the one place where the SDK makes a routing decision on your behalf. According to the README, when limit is 10 or less and return_content is False, the call automatically goes to the instant endpoint at /search/instant, which returns results immediately without polling. Anything else goes through the polling path.

python
from oxylabs_ai_studio.apps.ai_search import AiSearch

search = AiSearch(api_key="<API_KEY>")
result = search.instant_search(query="lasagna recipe", limit=10)
print(result.data)

instant_search is capped at 10 results and accepts query, limit and geo_location. The full search accepts up to 50 results, a render_javascript flag and a return_content flag that defaults to True, meaning you get markdown page content back unless you turn it off. Geo targeting differs between the two: the full search documents ISO 2-letter codes, country names and coordinate formats, while instant_search points at Google Ads GeoTargets for canonical location names. If you pass a location string that works in one and not the other, that mismatch is the likely cause.

Where the abstraction leaks

The honest limitation is that this SDK cannot do anything the hosted service cannot do, and it cannot tell you when the service got something wrong. A prompt-driven extractor returns a plausible value even when the page did not contain one. There is no selector in the response to audit, and the README does not document a confidence score or a validation step. If your pipeline needs to prove that a price came from a specific element, you are building that check yourself, probably by comparing the extracted value against a second source.

The second leak is JavaScript. render_javascript defaults to False on the scraper, the crawler and the search. The scraper can be set to "auto", but the crawler and search only take a boolean, so you either always render or never render. Browser instructions, which let you click, type and wait before capture, require render_javascript=True on the scraper. A site that loads its prices after a scroll will return empty fields under the defaults, and the failure looks like a bad schema rather than a rendering problem.

Third, this is a metered service. max_credits exists on the crawler as a ceiling, and the README mentions credits but does not publish a rate table, so you cannot estimate cost from the repository alone. Crawling with the default return_sources_limit of 25 and rendering enabled is the expensive combination, and nothing in the SDK warns you before the call.

Finally, there is no rollback story. The README does not document version pinning, a changelog policy, or what happens to your calls when the remote API changes. The published release is v0.2.19 from 2025-11-20, while pyproject.toml in the repository declares version 0.2.22, so the source tree and the released artifact were not in step at the time of writing.

How it compares to a self-hosted crawler

Firecrawl is the comparison the repository's own topics invite, and the difference is architectural rather than cosmetic. A self-hosted crawler such as Firecrawl runs the fetch, the JavaScript rendering and the markdown conversion inside your own infrastructure, so page content never leaves your network and your costs are compute rather than credits. You own the failure modes: browser versions, memory, retry storms.

Oxylabs AI Studio inverts that. You send URLs and prompts to a remote service that runs its own rendering and extraction stack, and you get structured data back. The SDK is roughly a typed wrapper over those endpoints, which is why it fits in a few hundred lines and why the dependency list is five packages long. If your constraint is data residency or per-page cost predictability, the self-hosted route wins. If your constraint is engineering time and you need geo-targeted rendering across many countries without running proxies, the hosted route is the shorter path.

A narrower alternative is to skip AI extraction entirely and keep selector-based scraping with a plain HTTP client plus a parser. That is more brittle across redesigns but completely deterministic and free of per-call billing. The prompt-driven approach is worth its cost only when the pages you target change often enough that selector maintenance is the bigger expense.

Licence, maintenance and upgrade cost

The repository is MIT licensed, which covers the SDK source: you can read it, fork it, and vendor it. It does not cover the service. The API key is the real licence boundary here, and the MIT grant says nothing about how many requests you may make, what you may do with the returned data, or what happens if Oxylabs changes the endpoint contract. Those terms live in the AI Studio agreement, not in LICENSE.

The repository is not archived, and the last push was on 2026-08-21, so work has landed within the last month. That is the only maintenance signal available; the README does not publish a support policy or a deprecation window. The released version is v0.2.19 from 2025-11-20, and pyproject.toml declares 0.2.22, so pinning to a tag and reading the diff before upgrading is the cautious move. Because the SDK is a thin HTTP client, most upgrades will be parameter additions rather than breaking rewrites, but a remote-side change to an endpoint can break you without any version bump at all. Tooling is standard: hatchling builds the wheel, mypy runs in strict mode, and ruff is configured with a broad rule set including bandit checks.

Editorial conclusion

Adopt it if you already pay for Oxylabs AI Studio or want prompt-driven extraction without maintaining your own headless browser fleet, and if a Python 3.10+ service that sends page content to a third party is acceptable. Do not adopt it if you need offline extraction, a self-hosted pipeline, or deterministic selectors that never change between runs. Before committing, verify your AI Studio API key actually authenticates, check the credit cost of the crawl and browser agent calls you plan to run, and read the current parameter list for AiMap.map, which the README truncates mid-example.

Frequently asked questions

How much does Oxylabs AI Studio cost per month?

The README and pyproject.toml do not publish a price or a credit rate table. The crawler exposes a max_credits parameter as a ceiling per call, which implies usage is metered in credits, but the cost of a credit is not stated in the repository.

How do I install and start using the Oxylabs AI Studio Python SDK?

Install with pip install oxylabs-ai-studio on Python 3.10 or above, then construct a class such as AiScraper with your API key and call scrape on a URL. The repository ships runnable examples under examples/, including scrape_generated_schema.py and crawl_markdown.py.

Where does the Oxylabs AI Studio Python SDK get its API key?

The README requires an API key but does not document where to obtain one; every example passes it as the api_key argument to the class constructor. The homepage listed for the project is aistudio.oxylabs.io.

Official sources

  1. License: MIT
  2. oxylabs/oxylabs-ai-studio-py on GitHub
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/oxylabs-oxylabs-ai-studio-py.svg)](https://hysenlabs.com/projects/oxylabs-oxylabs-ai-studio-py)