# Browser Agent from Oxylabs AI Studio: natural language browser automation in Python

> Browser Agent is a hosted browser automation tool driven by prompts or step lists, packaged as the oxylabs-ai-studio Python client. It removes selector maintenance, but it depends on an API key, a paid credit balance and a live service.

**oxylabs/browser-agent-py** — AI Browser Agent is an advanced Browser AI tool developed by Oxylabs AI Studio that automates real user browsing tasks using natural language instructions.

- Repository: https://github.com/oxylabs/browser-agent-py
- Website: https://aistudio.oxylabs.io/apps/browser_agent?utm_source=877&utm_medium=affiliate&utm_campaign=ai_studio&utm_content=browser-agent-py&groupid=877&transaction_id=102f49063ab94276ae8f116d224b67
- Stars: 1,563 · Forks: 2
- Language: Unknown
- License: not declared
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/oxylabs-browser-agent-py

## What Browser Agent solves, and who ends up using it

Static scraping scripts break when a site changes a class name or reshuffles its markup. Browser Agent's answer is to stop describing the page and start describing the task. The README frames the contrast directly: traditional frameworks such as Puppeteer or Selenium "rely on writing selectors and scripts for every action," while the agent takes instructions in plain language and performs clicks, typing, navigation and scrolling itself.

The audience is therefore not the engineer who enjoys writing a resilient crawler. It is the engineer who has a one-off or low-volume flow that is awkward to script: a checkout simulation that has to add an item, apply a coupon and confirm, a travel search that filters results before extracting prices, a job board where postings sit behind a search box. The README lists exactly these as practical use cases. It is also a reasonable fit for teams that already pull data through Oxylabs and want the same account to cover interactive pages.

The cost of that convenience is that the browsing logic is no longer yours. You describe intent, the service decides the steps. That is a real trade, not a free upgrade.

## How the hosted agent turns a prompt into structured output

The client is thin. You construct a BrowserAgent with an API key and call run() with a starting URL, a user prompt and an output format. The README's parameter table lists url and user_prompt as mandatory, output_format as optional and defaulting to markdown, schema as mandatory when JSON is requested, and geo_location as an optional ISO2 proxy location. That last parameter is the tell: the browsing happens on Oxylabs infrastructure, not on your machine, and the request can be routed through a proxy in a chosen country.

Two input styles are documented. You can hand over a natural language prompt, or you can supply a structured step list of actions named click, type, navigate, wait and extract. The second style matters for repeatability: a step list is closer to a script than a sentence, and it narrows what the agent is free to improvise.

Output is one of four formats. json returns data shaped by an OpenAPI schema, markdown returns readable text, html returns the raw page, and screenshot returns a PNG. The schema can be written by hand or produced by generate_schema(), which takes a plain description of the fields you want, such as "game name, platform, review stars and price", and returns a schema object you pass back into run(). For JSON runs this is not optional; the README marks the schema as mandatory for that format.

## Installing oxylabs-ai-studio and running a first task

The README requires Python 3.10 or above and an API key, and points to a free trial with 1000 credits for registration. Installation is a single pip command:

```bash
pip install oxylabs-ai-studio
```

With the package installed, the first real task combines schema generation with a browsing run. The README's example uses the Oxylabs sandbox store, asks the agent to search for a game and find its price, and prints the parsed data:

```python
from oxylabs_ai_studio.apps.browser_agent import BrowserAgent

browser_agent = BrowserAgent(api_key="<API_KEY>")

schema = browser_agent.generate_schema(
    prompt="game name, platform, review stars and price"
)

result = browser_agent.run(
    url="https://sandbox.oxylabs.io/",
    user_prompt="Find if there is game 'super mario odyssey' in the store. If there is, find the price. Use search bar to find the game.",
    output_format="json",
    schema=schema,
)
print(result.data)
```

What you should see is a result whose data field carries the parsed payload. The README's sample output shows a games array with game_name, platform, review_stars and price, and notes that review_stars came back null for that run. That null is worth noticing: a schema guarantees shape, not that every field will be populated.

Screenshots follow the same call with a different format. The returned payload is base64, so the README decodes it before writing the file:

```python
import base64
from oxylabs_ai_studio.apps.browser_agent import BrowserAgent

browser_agent = BrowserAgent(api_key="<API_KEY>")

result = browser_agent.run(
    url="https://sandbox.oxylabs.io/",
    user_prompt="Go to the website and take a screenshot of the home page",
    output_format="screenshot",
)

with open("screenshot.png", "wb") as f:
    f.write(base64.b64decode(result.data.content["data"]))
```

The repository itself contains only a README, a banner image and a sample screenshot, so the client code above is the whole of the documented surface. A JavaScript SDK exists in a separate repository, oxylabs-oxylabs-ai-studio-js, if Python is not your runtime.

## Where Browser Agent is the wrong tool

The README is candid about the first limit: sites with advanced bot detection "may require advanced setup." That is a soft way of saying the hosted browser can be blocked, and the documentation does not describe what that setup consists of. If your target site actively fights automation, you are buying an unknown amount of work.

The second limit is structural. Every run is a network call to Oxylabs AI Studio with an API key. There is no offline mode, no local browser, and no described way to inspect or replay the agent's step-by-step decisions after the fact. For a pipeline that must run in an air-gapped environment, or one where you have to show exactly which request produced which value, that is disqualifying.

Third, the README states that JSON output requires a schema. There is no documented mode where you pass a prompt and receive arbitrary structured JSON without first defining or generating the shape. If your extraction target is genuinely unknown until the page loads, you are working against the design rather than with it.

Finally, cost. The README mentions a free trial of 1000 credits but does not publish a per-run credit price for Browser Agent, so anyone planning volume has to check the pricing separately. The README also cuts off mid-sentence in its own FAQ answer to whether the tool is free, which leaves that question open in the repository itself.

## Browser Agent against Playwright and Puppeteer

The obvious alternative is a local automation library: Playwright or Puppeteer, both of which drive a browser you install and control. The difference in approach is where the intelligence sits. With Playwright you write the selectors, the waits and the click order, and you can run the whole thing on your own machine, in CI, or inside a container with no outbound dependency on a vendor. Browser Agent inverts that: you write a sentence, the service writes the steps, and the browser lives on Oxylabs infrastructure.

That inversion changes failure behaviour. A Playwright script fails at a known line with a known selector. A prompt-driven agent can take a different route on two consecutive runs and still return a well-formed JSON object, which is harder to test. The README's structured step list narrows this gap by fixing the action sequence, but the interpretation of each step is still the agent's.

The practical split is volume and volatility. High-volume, stable targets are cheaper and more predictable as a maintained Playwright script. Low-volume, shifting targets where writing selectors costs more than the run itself are where the hosted agent earns its place. Note also that Browser Agent is not a scraping framework replacement for plain pages: for static HTML, an HTTP client plus a parser is still the smaller tool.

## Maintenance, upgrades and licence questions

The repository was last pushed on 2026-04-02, roughly five months before this writing. It is not archived, but a single push date with no retrieved releases does not support a claim about how actively it is developed. Treat the client as a stable wrapper around a service whose real behaviour is versioned on the server side, and pin the oxylabs-ai-studio version in your requirements file so a pip upgrade does not silently change the client under you.

The upgrade risk is asymmetric. A server-side change to how the agent interprets prompts can alter your results without any change to the package you installed, and the README does not document a rollback path or a way to pin agent behaviour. If reproducibility matters, capture the schema you used and the raw output alongside each run so a shift is visible.

On licensing: the repository does not state a licence, and the package is distributed through PyPI. That means you cannot confirm from this material what terms govern the client code, and the service itself is a commercial product with a credit balance. Anyone planning to redistribute the client, or to embed it in a product, should confirm the licence and the service terms directly rather than assuming either. This is a gap in the repository, not a legal conclusion.

## Conclusion

Adopt Browser Agent when the browsing flow is described more easily in words than in selectors, and when you can absorb a per-run credit cost and an outbound dependency on Oxylabs AI Studio. Do not adopt it if you need offline execution, an auditable open source agent loop, or a self-hosted browser you control. Before committing, install oxylabs-ai-studio on Python 3.10 or above, run the sandbox example against https://sandbox.oxylabs.io/, and confirm your credit balance and the schema your JSON output actually needs.

## FAQ

### What is Browser Agent from Oxylabs used for?

It automates multi-step browsing tasks described in natural language or as a step list, performing clicks, typing, navigation, scrolling and screenshots, then returning data as JSON, Markdown, HTML or a PNG. The README lists e-commerce checkout simulation, travel search automation, job search scraping and event discovery as use cases.

### Is Browser Agent safe to use?

The README says Browser Agent works on most public websites but that you must make sure your use case complies with the target site's Terms of Service and applicable laws. It also notes that sites with advanced bot detection may require advanced setup, and does not describe what that setup involves.

### What can browser agents do?

According to the README, Browser Agent executes clicks, inputs, navigation and scrolling, runs multi-step flows, interacts with JavaScript-rendered pages, and extracts structured JSON once the browsing sequence completes. It can also return Markdown, raw HTML or a PNG screenshot of the browser content.

## Sources

- [Issues](https://github.com/oxylabs/browser-agent-py/issues)
- [oxylabs/browser-agent-py on GitHub](https://github.com/oxylabs/browser-agent-py)
- [Project website](https://aistudio.oxylabs.io/apps/browser_agent?utm_source=877&utm_medium=affiliate&utm_campaign=ai_studio&utm_content=browser-agent-py&groupid=877&transaction_id=102f49063ab94276ae8f116d224b67)
- [README](https://github.com/oxylabs/browser-agent-py/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/oxylabs-browser-agent-py
