Open-source project
g1879/DrissionPage avatar
g1879/DrissionPage

DrissionPage: driving a real browser and raw HTTP from one Python object

Python based web automation tool. Powerful and elegant.

12,497 stars1,148 forksPythonNOASSERTION

At a glance

What is it?
A Python automation library that treats a browser session as something you can also send packets through, without webdriver in the middle.
Who is it for?
DrissionPage is worth a look when your automation problem is that a page needs a real browser to render but your data volume is too high for one request per page, since that is exactly the seam this library is built around. The dependency list is small and readable, `lxml` does the parsing, and there is no driver to version-match against your browser build.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository last received commits 23 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 7, 2026, and from our analysis. They are not legal advice.

Editorial analysis

One object for the browser and the wire

Most Python browser automation stacks split into two camps. One drives a browser through a standardized protocol and sees the page as a tree of elements. The other sends HTTP requests directly and parses the response as text. DrissionPage's premise is that these are not competing approaches but two views of the same session, and that a useful tool should let you move between them without starting over.

The README states this directly: it can control a browser, it can send and receive packets, and it can combine the two, taking the convenience of browser automation together with the efficiency of requests. That is the whole idea in one line, and it is why the dependency list includes both `requests` and `lxml` rather than just a driver client.

The size of the project is unusual for this category. 12,469 stars and 1,149 forks puts it in the same visibility range as the major automation frameworks, with 259 open issues and a last push on 2026-09-14. The default branch is `master`, the language is Python, and the repository tree is small: a `DrissionPage/` package directory, a `tests/` directory, a `requirements.txt` and a `MANIFEST.in`.

No webdriver, and what that buys

The feature list leads with what the author calls a fully self developed core, and the comparison drawn is against Selenium. The specific claims are the interesting part, because each one is a consequence of dropping the driver protocol:

The stated advantages over Selenium are that the library is not built on webdriver, that you do not download a different driver for each browser version, and that it runs faster. Each of those is a real operational saving. Driver management is the part of Selenium that generates the most recurring pain, and the mismatch between an installed browser and its expected driver is a failure mode that arrives on someone else's release schedule.

The remaining claims are about how the page is presented to your code. You can find elements across iframe boundaries without switching into and out of them, and an iframe is treated as an ordinary element rather than a context you have to enter. Several tabs can be operated at once with no switching, which changes how you write concurrent work against one browser. A shadow root in a non open state can be handled, and you can screenshot a whole page including content outside the viewport.

A short dependency list that explains the design

The requirements file is short enough to read in one glance, and each entry hints at a capability:

text
requests
lxml
cssselect
DrissionGet>=1.2.1
DrissionRecord
websocket-client
click
tldextract>=3.4.4
psutil
ftfy

`lxml` and `cssselect` together account for the parsing claim in the README, which says the built-in lxml engine raises parsing speed by several orders of magnitude. That is a plausible claim for HTML parsing and it explains why the library can hand you both an element API and raw lxml nodes.

The two packages prefixed `Drission` are the author's own, and they split the job: `DrissionGet` is the HTTP side, which is how a single session can speak to a site over the wire, and `DrissionRecord` is the traffic recording and playback side. `websocket-client` is there because browser sessions speak DevTools over a websocket, `psutil` for process management of the browser itself, `click` for the command line entry point, `tldextract` for correct cookie domain handling, and `ftfy` for repairing text encoding damage. Nothing in this list is a browser binary or a driver.

Design choices aimed at unstable pages

The feature section spends most of its length on the ergonomics, and the recurring theme is waiting. The README lists automatic waiting and retry everywhere as the feature that makes an unstable network controllable and programs stable. This is a deliberate contrast with frameworks that leave waiting to the caller, and it is the single most consequential design decision in the library.

The rest of the list is a set of small conveniences that add up. There is a minimal locator syntax intended to make finding elements easier, a download manager so you get reliable file handling while a browser is in use, the ability to reuse an already open browser across runs so you are not launching from scratch every time, and ini file configuration loaded automatically to avoid a pile of settings. Page object mode is offered as a wrapper suitable for testing and extension. The README closes this section by saying there are more details it will not list and inviting people to find them by using it, which is a common way for a project with this many small affordances to avoid writing a list nobody reads.

Licensing and the terms of use

This is the part of the repository worth reading carefully, and it is unusual. The license field is unasserted, and the README carries an explicit terms of use section. It allows anyone to use or distribute the source for personal identity, limited to study and lawful non profit purposes. Individuals or organisations without authorisation from the copyright holder may not use the project in source or binary form for commercial purposes.

The conditions attach automatic revocation of that authorisation if any term is breached. They prohibit applying the library to projects that violate local law or ethics, projects that may harm others, and any attacking or harassing behaviour. One term is technical rather than moral: users must respect the Robots protocol and may not collect data that the law or a site's Robots file disallows.

The disclaimer is standard in structure and worth noting anyway. All conduct is the user's own responsibility, disputes and consequences arising from use are unrelated to the copyright holder, and the holder accepts no liability for losses from any defect in the library. For an individual developer or a research project this is workable. For a company, the commercial restriction means the terms have to be settled with the author before the library goes into a product, and that conversation is not something the repository answers.

Release cadence and where the documentation lives

The release list is short. v4.1.0.17 was published on 2025-03-21, v4.1.0.13 on 2024-12-07, and v4.0.4.23 on 2024-05-28. The v4.1.0.17 notes are detailed enough to be worth reading as a description of what the project considers worth changing: elements gained a `child_count` attribute, every `Settings` property gained a setter, English error and prompt text was added, `LocatorError` and `UnknownError` exceptions were introduced, and `ShadowRoot` became able to return text or numeric results from xpath. Smaller entries rename `WrongURLError` to `IncorrectURLError`, change `suffixes_list_path` to `suffixes_list`, and change the parameter name of `ChromiumElement.attr()` from `attr` to `name`.

That level of detail is a good sign for a library, and it is also a reminder of the cost. A rename that reads as cleanup in the notes is a breaking change for anyone who used the old name, which is why the four-digit patch versions here encode real API movement.

Documentation is external. The README is a summary in Chinese that points to DrissionPage.cn as the official site, and the repository is mirrored across gitee, github and gitcode. The README is dominated by a sponsor block for proxy providers, which is a practical signal about who the audience is: people doing collection and cross border data work, where proxy configuration is a first class requirement rather than an afterthought.

Editorial conclusion

DrissionPage is worth a look when your automation problem is that a page needs a real browser to render but your data volume is too high for one request per page, since that is exactly the seam this library is built around. The dependency list is small and readable, `lxml` does the parsing, and there is no driver to version-match against your browser build. Two things deserve a decision before you commit: the terms of use are personal and non commercial, so a company evaluating it needs the author's permission, and the substantive documentation lives on an external site rather than in the repository. For personal and research work the fastest route is `pip install DrissionPage`, then the online documentation on session modes.

Frequently asked questions

What is DrissionPage Python used for?

DrissionPage is a Python library for automating web pages that can both control a browser and send HTTP packets, combining the two in a single session. It targets collection, scraping and browser automation tasks where rendering in a real browser and raw request efficiency both matter.

How does DrissionPage compare with Selenium?

The README's own comparison lists what it claims over Selenium: no webdriver dependency, no separate driver download per browser version, faster execution, element lookup across iframes without switching context, simultaneous operation on multiple tabs without switching, whole page screenshots including content outside the viewport, and handling of non open shadow roots.

Do I need to install a browser driver to use DrissionPage?

No. The library is built without webdriver, which is the source of the README's claim that you do not need to download a different driver for each browser version. Its dependencies are requests, lxml, cssselect, websocket-client, click, tldextract, psutil and ftfy, plus the author's own DrissionGet and DrissionRecord packages.

Can I use DrissionPage in a commercial project?

Not without permission. The README's terms of use allow personal use and distribution limited to study and lawful non profit purposes, and state that individuals or organisations without authorisation from the copyright holder may not use the project in source or binary form for commercial purposes. Breach of any listed condition revokes the authorisation automatically.

Which version of DrissionPage is current?

The latest tagged release in the repository is v4.1.0.17, published on 2025-03-21, after v4.1.0.13 in December 2024 and v4.0.4.23 in May 2024. The last push to the repository was on 2026-09-14, so development is active even though releases are not frequent.

Official sources

  1. g1879/DrissionPage on GitHub
  2. Issues
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/g1879-drissionpage.svg)](https://hysenlabs.com/projects/g1879-drissionpage)