Library / SDK
gawel/pyquery avatar
gawel/pyquery

pyquery: jQuery selectors over lxml, twenty years on

A jquery-like library for python

2,378 stars188 forksPythonNOASSERTION

At a glance

What is it?
A Python library that loads XML or HTML from a string, file, lxml tree or URL and then lets you select with the selectors you already know.
Who is it for?
pyquery sits in a gap that is easy to describe: BeautifulSoup gives you traversal without a selector language, and lxml gives you XPath without a friendly API, while pyquery gives you CSS selectors and jQuery method names over lxml. That is still worth having, especially for people moving from front end work.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository last received commits 72 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 7, 2026, and from our analysis. They are not legal advice.

Editorial analysis

A stated scope and an unapologetic origin

The README is short and unusually direct about what the library is and is not. It says pyquery allows you to make jQuery queries on XML documents, that the API is as much as possible similar to jQuery, and that it uses lxml for fast XML and HTML manipulation.

The origin story is in the second paragraph, and it is worth reading because it explains the design. The author says this is not, or at least not yet, a library to produce or interact with JavaScript code, and gives the reason plainly: he liked the jQuery API, missed it in Python, and told himself to build jQuery in Python. That is the whole design brief.

It explains decisions that would otherwise look arbitrary. The constructor takes several input types rather than one. The selector syntax includes jQuery extensions that are not valid CSS. And the traversal methods are named the way a front end developer would expect rather than the way a tree library would.

The scale is modest and steady. 2,377 stars, 188 forks, 61 open issues, last pushed on 2026-07-27, default branch `master`. The README describes the project as actively developed on a Git repository, and mentions a policy of giving push access to anyone who wants it and then reviewing what they do. The site is pyquery.rtfd.org, and the repository topics are `css`, `jquery`, `lxml`, `python` and `python3`.

Five ways to load a document

The quick start covers the constructor's input types in five lines, which is the fastest way to understand the API:

python
>>> from pyquery import PyQuery as pq
>>> from lxml import etree
>>> d = pq(\"<html></html>\")
>>> d = pq(etree.fromstring(\"<html></html>\"))
>>> d = pq(filename=path_to_html_file)

Five inputs are supported: a string, an lxml document, a file via the `filename` keyword, a URL via the `url` keyword, and a URL with a custom `opener` callback if you need to control fetching. The `opener` form takes a lambda, which means the library does not impose an HTTP client on you.

That last point matters more than it looks. pyquery declares only two hard dependencies, `lxml>=2.1` and `cssselect>=1.5.0`, and neither of those fetches anything. If you want HTTP, you pass your own opener, which means your existing session, proxy configuration, retry policy and headers all keep working. There is a `test` extra that pulls in `requests`, `webob`, `webtest`, `pytest` and `pytest-cov`, so the fetching code paths are exercised in the test suite rather than shipped as a dependency.

Reading and writing inner HTML in four lines

Once loaded, the object behaves like the `$` in jQuery, which the README says directly. Selecting by id returns a list whose repr is the element summary, and the `html` and `text` methods read and write content:

python
>>> p = d("#hello")
>>> print(p.html())
>>> p.html("you know <a href='http://python.org/'>Python</a> rocks")
>>> print(p.text())
you know Python rocks

Two things are worth noticing in that sequence. The setter form of `html()` replaces the inner HTML and returns the selection, so calls chain the way jQuery calls do. And `text()` strips the markup, returning `you know Python rocks` rather than the raw string with the anchor in it. The repr of a selection, `[<p#hello.hello>]`, shows the tag, the id and the classes, which is the jQuery convention rather than the lxml one.

The selector is a CSS selector passed to cssselect, which is why that library is a dependency rather than a bundled parser. Anything cssselect understands, pyquery understands, and anything it does not, the jQuery pseudo classes below cover.

Pseudo classes that CSS never standardised

The last section of the README documents a feature that has no equivalent in plain CSS, and it is the part most likely to be useful to someone arriving from front end work. The library supports pseudo classes that are available in jQuery but are not standard in CSS: `:first`, `:last`, `:even`, `:odd`, `:eq`, `:lt`, `:gt`, `:checked`, `:selected` and `:file`.

Position filters are the reason people reach for a jQuery-shaped library in the first place. In CSS you would use `:nth-child`, which counts every sibling, whereas `:eq(2)` selects the third match of your selector specifically. If you arrived from JavaScript and wrote `:eq` by habit, cssselect alone would reject it:

python
>>> d('p:first')
[<p#hello.hello>]

The `:file`, `:checked` and `:selected` entries are form state filters, useful when you are scraping or driving a page where the value of an input matters more than its markup.

One inconsistency to be aware of: the README's documented import path is `from pyquery import PyQuery`, while the project's own long description in `setup.py` refers to it as the `PyQuery` class while GitHub and PyPI name the distribution `pyquery`. The `as pq` alias is the documented idiom and the one the documentation site uses.

Who maintains it, and under what terms

The packaging metadata answers the maintenance question directly. The author is Olivier Lauzanne, with a copyright header dating the project to 2008, and the maintainer is Gael Pasgrimaud. The version string in `setup.py` is 2.1.1.dev0, a development marker that tells you the package is tracked on a rolling basis rather than cut at release boundaries.

The classifiers list Python 3.11, 3.12 and 3.13, and the development status is marked as 5, Production/Stable. Notably the classifiers do not list anything older than 3.11, which is a useful signal about which interpreters are actually supported now even though the library itself is old enough to predate all of them.

The license is BSD, stated in `setup.py` and in `LICENSE.txt`. Worth noting that the repository's own license metadata field is unasserted, so the authoritative statement is the one in the source and in the file.

There is another detail in `setup.py` that says something about how releases work. The long description is assembled by reading `README.rst` and `CHANGES.rst` from disk and substituting them into a template, so the news section of the package page on PyPI is the changelog file verbatim. That is an old school but reliable approach, and it means the changelog is the release history. There are no GitHub releases on this repository at all, so `CHANGES.rst` is where you look for what changed.

Repository layout and a Mercurial ghost

The tree is small and reveals the project's age and its tooling habits. There is a `pyquery/` package directory, `tests/`, and a `docs/` directory that is the source for the ReadTheDocs site. Test configuration is split across `conftest.py`, `pytest.ini` and `tox.ini`, with the build badge in the README pointing at a workflow that runs tox.

Two entries stand out. `README_fixt.py` is a fixture module, and the README references it in a commented out doctest line that calls a `getfixture` helper with `readme_fixt`. In other words, the README's own examples are partly wired into the test suite through that fixture, which is why the doctest examples work despite depending on variables like `your_url` and `path_to_html_file` that the reader is expected to supply.

The other is `.hgignore`. Mercurial is long gone from mainstream Python tooling, and its presence in a repository that also has `.github/` is a fossil that dates the project precisely. `.hgignore` was not removed during a later migration, which suggests the repository was moved rather than rebuilt.

The practical consequence for a reader is that documentation is the thin part. The README gives you the constructor, two content methods and the pseudo class list. Everything else, including traversal, attributes, and the API surface around form elements, is in the documentation site at pyquery.rtfd.org and in the docstrings.

Editorial conclusion

pyquery sits in a gap that is easy to describe: BeautifulSoup gives you traversal without a selector language, and lxml gives you XPath without a friendly API, while pyquery gives you CSS selectors and jQuery method names over lxml. That is still worth having, especially for people moving from front end work. The version is 2.1.1.dev0, the maintainer is Gael Pasgrimaud, the original author is Olivier Lauzanne, and the last push was on 2026-07-27. There are no GitHub releases at all, so track the version through the changelog rather than through tags. Start with `PyQuery` from a string, and use the `:first`, `:eq` and `:lt` pseudo classes that CSS does not define but jQuery does.

Frequently asked questions

What is pyquery and what is it used for?

pyquery is a Python library for making jQuery style queries against XML and HTML documents. It uses lxml for parsing and cssselect for selectors, so you select elements with CSS selectors and then read or replace their content with jQuery method names. It is for parsing documents, not for producing or interacting with JavaScript code.

How do I load a document in pyquery?

The PyQuery constructor accepts five forms: an XML string, an lxml document, a file with the filename keyword, a URL with the url keyword, or a URL with a custom opener callback. The opener form lets you supply your own fetching behaviour, so pyquery does not force an HTTP client on you.

Does pyquery support jQuery pseudo classes such as :eq and :first?

Yes. The README lists the pseudo classes available in jQuery that are not standard in CSS: :first, :last, :even, :odd, :eq, :lt, :gt, :checked, :selected and :file. These are handled by the library on top of cssselect, which otherwise only understands standard CSS selectors.

What are pyquery's dependencies?

Only two are required: lxml>=2.1 and cssselect>=1.5.0. Testing adds a separate extra with requests, webob, webtest, pytest and pytest-cov. Because no HTTP client is a hard dependency, you can pass your own opener to the constructor and keep your existing session, proxy and retry behaviour.

How is pyquery maintained and what version is current?

Olivier Lauzanne is the original author and Gael Pasgrimaud is the maintainer, with the project dating to 2008. The version in setup.py is 2.1.1.dev0 and the classifiers list Python 3.11 through 3.13. There are no GitHub releases, so CHANGES.rst is the release history, and the last push was on 2026-07-27.

Official sources

  1. gawel/pyquery on GitHub
  2. Issues
  3. Project website
  4. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/gawel-pyquery.svg)](https://hysenlabs.com/projects/gawel-pyquery)