Model or dataset
LoseNine/ruyipage avatar
LoseNine/ruyipage

RuyiPage wraps a patched Firefox in WebDriver BiDi and ships the fingerprint file with it

RuyiPage is a next-generation Python web automation framework built on the WebDriver BiDi protocol, designed to bypass bot detection. It self-debugs with AI, uses trace logs to analyze any network flow, and uses a fingerprinted Firefox browser to pass site detection checks.

1,826 stars220 forksPythonBSD-3-Clause

At a glance

What is it?
A Python automation framework for data capture and anti-bot analysis, paired with a custom Firefox build that rewrites fingerprints in native getters and can trace every JS call the page makes.
Who is it for?
RuyiPage is two products sharing one package name, and a reader should be clear about which half they need. If you want Playwright-style Python control over Firefox, the `ruyipage` package alone is enough and the paired runtime is optional, since `browser_path` accepts any Firefox you already have.
Can I use it commercially?
Yes. BSD-3-Clause is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 4 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 9, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Two layers: a BiDi library and a Firefox that was edited underneath it

The package is described as a next-generation Python web automation framework built on the WebDriver BiDi protocol, and the README is quick to add that it can intercept arbitrary request and response packets. What makes it more than a driver wrapper is that it ships its own Firefox engine for Windows and Linux, described as battle-tested against detection.

That second layer matters because most of the features the README advertises are implemented in the browser rather than in Python. Getting a closed shadow root node, having `isTrusted` come back true on constructed events, keeping `navigator.webdriver` false, and rewriting a canvas hash are all things a Python library cannot fake convincingly from outside the page. They have to exist in the engine.

The claim about constructed events is specific enough to be worth parsing. Alongside native actions such as click, input, and hover, the library supports attaching `ruyi: true` to constructed `Event`, `InputEvent`, `MouseEvent`, and `KeyboardEvent` objects so the resulting `isTrusted` value resembles real interaction. That is a browser-side flag, not a Python property.

The protocol choice also explains some API shape. BiDi is the second WebDriver standard, and it exposes capabilities WebDriver Classic never had, which is why a page debugger with breakpoints, conditional breakpoints, log breakpoints, stepping, a call stack, scope inspection, source reading, pause on exception, and property watchpoints fits in as `page.debugger` rather than being bolted on.

There is also an interoperability path the README calls out: support for taking over automation of fingerprint browsers such as ADS, plus HTTP and SOCKS5 proxies with password support and a separate password proxy per tab. The per-tab detail is the kind of granularity that only matters when several identities are running concurrently.

The fingerprint is a text file, not a set of Python kwargs

The paired engine takes its entire external identity from a plain text file passed as `--fpfile=`. One file decides the user agent, language, timezone, screen, CPU core count, touch capability, canvas and audio perturbation, WebGL graphics parameters, a font whitelist, WebRTC, the speech voice list, geolocation, and HTTP or SOCKS5 proxy credentials.

The implementation detail is what makes this approach defensible rather than a wrapper trick. The README states that every value is rewritten inside native C++ getters, that no JavaScript function is wrapped, so `toString()` still reports `[native code]`, and that the prototype chain shape is unchanged. A page that inspects a getter and finds a JavaScript wrapper has found the automation; a page that finds a native getter cannot tell the difference. `navigator.webdriver` is stated to be permanently false.

For most cases the README says you do not need to write that file by hand. `opts.smart_fingerprint()` probes the exit IP, matches language and timezone to it, and picks one of 22 real device hardware profiles to generate a set automatically. When you do need to pin a specific device model or adjust a single field, the documentation to read is `fingerprint/fpfile-fingerprint.md`, which per the README covers every key, its aliases, the APIs it affects, and how to verify the result.

The framework is also explicit that this layer does not depend on BiDi, and that page JavaScript cannot see it. That separation is the cleanest statement of intent in the whole README: the fingerprint lives below the automation protocol, so removing the automation layer would not remove the fingerprint.

MOZ_DOM_TRACE turns the browser into a log of what the page probed

The second engine capability is a tracing facility driven entirely by environment variables set before launch, and it is aimed at a different task than automation. It is for working out what a page does at runtime, which is the question you ask when an anti-bot script behaves differently than you expect.

When the variables are set, the engine writes what happens in the page runtime to a log: every JavaScript function call with its arguments, return value, closure variables, and call tree, plus DOM property reads and writes, cookie and storage operations, exceptions, WASM module and cross-boundary calls, HTTP messages, and WebSocket frames. It can also dump every host interface this build exposes, which gives you a direct way to compare what the page looked for against what this environment actually provides.

There is a second group of `MOZ_DOM_API_*` variables that go further and fix the return values of `Date.now`, `performance.now`, `Math.random`, and `crypto.getRandomValues` directly. For a framework whose selling point is fingerprint consistency, being able to pin the clock and the randomness is a natural extension rather than a surprise.

powershell
$env:MOZ_DOM_TRACE = "1"
$env:MOZ_DOM_TRACE_FILE = "D:\\trace\\run.jsonl"
$env:MOZ_DOM_JSCALL_TRACE = "1"
& firefox.exe --new-instance -no-remote -profile "D:\\profile"

The documentation burden is handled with three files. `trace/trace-cheatsheet.md` is a lookup of 63 switches with their values, defaults, and conditions under which each takes effect. `trace/README.md` covers output formats, how to follow a causal chain, and troubleshooting. And `trace/tools/jscall_decode.py` is the decoder for the binary JSCall output, which is the part you will need if you leave the JSONL mode behind.

Installing the paired runtime, or pointing at the Firefox you already have

The install story is deliberately close to Playwright's, and the README frames it that way. The base package is one command:

bash
pip install ruyiPage --upgrade

Async support is an optional extra rather than a separate distribution, which keeps the default install small.

bash
pip install ruyiPage[async] --upgrade

That extra pulls in `greenlet` and `websockets`, and the README states the synchronous API is unaffected by it. The dependency list in `pyproject.toml` confirms the split: the base requires only `websocket-client>=1.0`, and the `async` extra adds `greenlet>=3.2` and `websockets>=13.0`.

The runtime install is a separate step, run after the package is installed.

bash
python -m ruyipage install

The README is candid about two properties of that command. It downloads the recommended Firefox runtime from the project's GitHub releases into the user cache directory, and it does not verify a hash by default, which the README justifies by pointing out it lets the project update release assets directly. If you prefer to check the artifact yourself, there is a dry-run that prints the download and install plan without fetching anything, plus a `--from-file` path for offline installation of an archive you downloaded yourself, and a `--force` flag to re-download.

The step is genuinely optional. The README says that if you already have Firefox installed, or you are using a portable or fingerprint build, you can skip it and pass the path explicitly through `browser_path`, and that an explicit `browser_path` always wins over the downloaded runtime. `python -m ruyipage path` prints the installed Firefox path and `python -m ruyipage doctor` reports installation status, so the two are easy to check before you debug anything else.

Sync and async differ by more than syntax

The minimal synchronous entry point is four statements, and it is worth quoting in full because it establishes the shape of everything above it.

python
from ruyipage import FirefoxPage

page = FirefoxPage()
page.get("https://www.example.com")
print(page.title)
page.quit()

The async version lives in a separate `ruyipage.aio` module and uses `launch()` rather than constructing `FirefoxPage` directly. The README states that method names are identical between the two and that only `async`/`await` is added, but that properties become methods: `page.title` turns into `await page.get_title()`. Element interaction is the same story, with `el.click_self()` and `el.input("hello async")` becoming awaited calls.

That is an honest design rather than a claim: an attribute that is a plain property in one and a coroutine in the other cannot be made identical, and the README says so plainly instead of implying a drop-in swap. Two worked examples ship at the repository root, `quickstart_bing_search_async.py` and `quickstart_cloudflare_async.py`, alongside their synchronous counterparts.

The example directory is the best map of the feature surface, because the 24 files are numbered in the order a reader would meet the concepts: navigation, element finding, interaction, wait conditions, action chains, screenshots, JavaScript evaluation, cookies, tabs, scrolling, network interception, console listening, iframes, shadow DOM, window management, emulation, web extension, and downloads. That ordering is itself a statement about what this framework considers foundational.

The working tree is three versions ahead of the last release

`pyproject.toml` declares version 1.2.72. The newest release on the repository is v1.2.69, dated 2026-09-08, with v1.2.66 and v1.2.58 before it. The last recorded push is 2026-09-16, later than the newest tag. So the tree and the last paired runtime are not the same thing, and since the runtime is downloaded from releases, a PyPI install of 1.2.72 would fetch a browser build from 1.2.69.

The releases themselves track the browser, not the library, which explains the pattern. Both v1.2.69 and v1.2.66 exist to move the bundled engine to Firefox 155.0, with v1.2.66 being the move from an earlier 155.0a1 build to the release build, and each one names the platform assets, for example `firefox-155.0.en-US.win64-20260907.zip` for Windows x64 and a `.tar.xz` for Linux x86_64.

The one release with real detail is v1.2.58, and it is instructive because it describes a Firefox 155 compatibility problem rather than a feature. Normal `launch()` sessions now start on a script-accessible `about:blank` so that page properties are readable immediately without a remote-allow-system-access flag, while the native landing page for private browsing is preserved when `private=True`. Its verification section lists 338 passing tests in the fast suite, 15 in an async smoke run, and 9 in a release gate.

The test configuration in `pyproject.toml` explains how that suite is organized. Markers separate core launch behavior (`smoke`), module-level regressions (`feature`), multi-module workflows (`integration`), a pre-release minimum set (`release`), tests that need a real Firefox or BiDi session (`browser`), async tests, and fast tests that need none of those. That split is why a library whose headline feature depends on a custom browser build can still run a large suite quickly.

Metadata is consistent otherwise. The license is BSD-3-Clause in both the repository field and `pyproject.toml`, Python 3.9 and newer is required with classifiers up to 3.12, and the project is marked Development Status 4, Beta. The README links an English translation at `README_EN.md`, which is present in the tree.

Editorial conclusion

RuyiPage is two products sharing one package name, and a reader should be clear about which half they need. If you want Playwright-style Python control over Firefox, the `ruyipage` package alone is enough and the paired runtime is optional, since `browser_path` accepts any Firefox you already have. If you need the fingerprint file to survive a detection check, the patched core is the product, and `opts.smart_fingerprint()` is the starting point rather than hand-written values. The trace facility under the `MOZ_DOM_*` variables is arguably the most broadly useful part, because it answers what a page actually probed regardless of whether you intend to defeat anything. One thing to settle first: `pyproject.toml` reads version 1.2.72 while the newest published release is v1.2.69, so the working tree is ahead of the last tagged runtime. Install from PyPI for the paired Firefox, or point `browser_path` at your own build.

Frequently asked questions

Is RuyiPage based on Selenium or Playwright?

Neither. It is built on the WebDriver BiDi protocol, the second WebDriver standard, rather than on the WebDriver Classic protocol that Selenium uses. It also ships its own patched Firefox runtime, which is optional: you can pass `browser_path` and use any Firefox you already have.

Do I need the bundled Firefox to use RuyiPage?

No. `python -m ruyipage install` downloads the paired runtime from GitHub releases, but the README says you can skip that step if you already have Firefox or use a portable or fingerprint build, and an explicit `browser_path` always takes precedence over the downloaded runtime. The fingerprint features are the part that genuinely need the patched core.

How does RuyiPage set its browser fingerprint?

Through a plain text file passed as `--fpfile=`, which controls user agent, language, timezone, screen, CPU cores, touch, canvas and audio perturbation, WebGL parameters, a font whitelist, WebRTC, speech voices, geolocation, and proxy credentials. Values are rewritten inside native C++ getters rather than by wrapping JavaScript, so `toString()` still reads `[native code]`.

How do I get trace logs from RuyiPage's Firefox build?

Set the `MOZ_DOM_*` environment variables before launch. `MOZ_DOM_TRACE` turns tracing on and `MOZ_DOM_TRACE_FILE` chooses the JSONL output path, while `MOZ_DOM_JSCALL_TRACE` adds JavaScript call capture with arguments, return values, closure variables, and the call tree. The switch reference is `trace/trace-cheatsheet.md`.

Official sources

  1. Issues
  2. License: BSD-3-Clause
  3. LoseNine/ruyipage on GitHub
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/losenine-ruyipage.svg)](https://hysenlabs.com/projects/losenine-ruyipage)