Framework
jae-jae/QueryList avatar
jae-jae/QueryList

QueryList: a jQuery-style PHP scraper built on phpQuery

:spider: The progressive PHP crawler framework! 优雅的渐进式PHP采集框架。

2,689 stars424 forksPHPLicense varies

At a glance

What is it?
QueryList wraps phpQuery in a chainable API for CSS3 selectors, list extraction and Guzzle-backed HTTP requests. It is aimed at PHP developers who already think in jQuery selectors, and it needs PHP 8.1 or newer.
Who is it for?
Adopt QueryList when your scraping logic is selector-driven and your stack is already PHP 8.1 or newer, because the rules array plus query() pipeline maps directly onto CSS selectors and the HTTP layer accepts Guzzle options such as proxy and timeout. Do not adopt it for JavaScript-rendered pages unless you are prepared to add the headless-browser plugin, and do not expect the core package to handle concurrency.
Can I use it commercially?
Not without permission. GitHub finds no licence file in the repository, and without a licence all rights are reserved by default: you may read the code but not reuse it. Check the README, or ask the authors, before using it.
Is it still maintained?
Yes. The repository last received commits 16 days ago.
What is it written in?
Mainly PHP, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 24, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem QueryList solves for PHP developers

Most PHP scraping starts the same way: fetch a page with curl or Guzzle, then reach for DOMDocument or a regex to pull fields out. DOMDocument's XPath is verbose, and regexes break the moment the markup changes. QueryList's answer is to keep the fetching and the extraction in one chainable object and to expose the extraction through CSS3 selectors, which the README describes as having "the same CSS3 DOM selector as jQuery".

The intended user is a PHP developer who already knows jQuery syntax and wants that knowledge to transfer to server-side scraping. The README's feature list is written for that person: DOM traversal and manipulation, a generic list-crawling mode, content filtering through selectors, and a plugin system for the things the core does not do. It is not a distributed crawling platform, and nothing in the README suggests scheduling, queueing or persistence. If you need a fleet of workers sharing a frontier, this is the wrong layer.

How the selector chain and the rules pipeline actually work

The core is built on phpQuery, and the README states this directly: QueryList is "based on phpQuery". That dependency explains the API shape. A QueryList instance holds a loaded document, find() narrows to a node set, and terminal accessors such as text(), attrs() and htmls() pull values out. Because the manipulation methods return the same kind of object, you can keep chaining: the README shows append(), removeClass(), replaceWith(), clone() and appendTo() used in a single expression.

List extraction is the second path, and it is the one most scraping jobs want. You pass an associative array of field names to rules(), where each value is a two-element array of selector and accessor, then call query() and getData(). The README's Google example uses 'title'=>array('h3','text') and 'link'=>array('h3>a','href'), and the printed result is a numbered array of associative rows. That two-element convention is the whole rule format: selector first, then the method or attribute name.

Network access goes through Guzzle. The README labels the section "HTTP Client (GuzzleHttp)" and demonstrates passing a third argument of options that includes proxy, timeout and headers. Cookies are set the same way, as a Cookie header string copied from a browser. There is also a dedicated encoding() method that takes an output charset and an optional input charset, with automatic detection when the input is omitted.

Installing QueryList and scraping a first list

The README lists one requirement, PHP >= 8.1, and one installation method, Composer. Run this from your project root:

bash
composer require jaeger/querylist

Once the autoloader is in place, the smallest useful script is a single expression that fetches a page and reads an attribute. This is the README's GitHub image example:

php
QueryList::get('https://github.com')->find('img')->attrs('src');

For a real extraction job, define rules and call query()->getData(). The README's Google search example is the template for this:

php
$data = QueryList::get('https://www.google.co.jp/search?q=QueryList')
    ->rules([
        'title'=>array('h3','text'),
        'link'=>array('h3>a','href')
    ])
    ->query()->getData();

print_r($data->all());

The result is an array of rows, each with a title and a link key. If the site serves a legacy charset, insert encoding() before find(), as the README does with QueryList::get('https://top.etao.com')->encoding('UTF-8','GB2312'). When the request needs a proxy or custom headers, pass a third options array with keys such as 'proxy', 'timeout' and 'headers', exactly as the README's httpbin example shows.

Where QueryList stops: JavaScript pages, concurrency and licence status

The README is candid that multithreaded crawling and JavaScript-rendered pages are plugin territory, not core behaviour. The feature list puts "Multithreaded crawl" and "Crawl JavaScript dynamic rendering page (PhantomJS/headless WebKit)" under the sentence "Through plug-ins you can easily implement things like". So a single QueryList::get() call is one synchronous request. If your target renders its content client-side, the selector chain will see an empty document unless you install and run one of those plugins, and the README does not document how to configure them.

A second gap is the licence. The repository metadata does not state one, and the README does not name one either. That matters more for a library you will ship than for a personal script, because without a stated licence the default copyright position applies. The README does not document any upgrade or migration procedure between major versions, so the release history is the only signal about change velocity.

Finally, the README's own examples point at Google and GitHub. Both sites actively block automated requests, and the README says nothing about rate limiting, retries or respect for robots.txt. The library will happily send the request; whether the response is useful is your problem.

QueryList compared with a standalone Guzzle plus DOMDocument script

The honest alternative for a PHP developer is to skip the wrapper and write Guzzle plus DOMDocument or Symfony DomCrawler yourself. The difference in approach is where the abstraction sits. QueryList bundles the HTTP client and the DOM into one object and gives the DOM a jQuery-shaped surface, so a selector you would write in a browser mostly works unchanged. A hand-rolled Guzzle script keeps the two concerns separate: you get the raw response, then hand the body to whatever parser you prefer.

That separation is an advantage when you need a parser other than CSS selectors, when you want to swap the HTTP client, or when you want to reuse the same fetch layer across non-HTML formats. It is a disadvantage when the job is exactly "grab these fields from this page", because you end up writing the same fetch-plus-parse glue each time. QueryList's rules array is a compact way to express that glue once.

The other realistic alternative is a full crawling framework in another language, which brings scheduling and distributed queues that QueryList does not attempt. Choose that when the work is a crawl across a large site rather than a scrape of a handful of pages.

Maintenance, releases and what to check before adopting

The repository is not archived, and the last push was on 2026-09-14. The most recent tagged release is V4.4.7 from 2026-02-02, preceded by V4.4.6 on 2024-12-13 and V4.4.5 on 2024-07-16. Those gaps are worth reading carefully: roughly eighteen months separated V4.4.5 and V4.4.6, and about fourteen months separated V4.4.6 and V4.4.7. Commits landing after a tag do not necessarily mean a new tag is imminent.

For upgrade cost, the README does not describe a migration path between major versions, so pinning the version in composer.json is the only defence the documentation supports. The PHP >= 8.1 requirement is the other constraint to plan around: if your application still runs on an older PHP, you cannot install this without upgrading the runtime first.

The licence question is the one to settle before you depend on it. The repository metadata and the README do not state a licence identifier, so there is nothing here to tell you whether redistribution in a closed product is permitted. That is a question for your own legal review, not something the project's documentation answers.

Editorial conclusion

Adopt QueryList when your scraping logic is selector-driven and your stack is already PHP 8.1 or newer, because the rules array plus query() pipeline maps directly onto CSS selectors and the HTTP layer accepts Guzzle options such as proxy and timeout. Do not adopt it for JavaScript-rendered pages unless you are prepared to add the headless-browser plugin, and do not expect the core package to handle concurrency. Before committing, confirm the licence terms yourself, since the repository metadata does not state a licence, and check that your target pages are reachable from the machine that will run the crawl.

Frequently asked questions

What is QueryList in PHP?

QueryList is a PHP web scraper and crawler framework based on phpQuery, which gives it a CSS3 selector and DOM manipulation API modelled on jQuery. It also provides a rules-based list extraction mode and a Guzzle-backed HTTP client.

What is a query list in QueryList's sense?

In QueryList, the list-crawling mode is expressed through rules(): you map field names to a selector and an accessor, call query(), then getData() to receive an array of rows. The README's example maps 'title' to array('h3','text') and 'link' to array('h3>a','href').

How do I install QueryList?

The README gives one method, Composer, with the command composer require jaeger/querylist. The only stated requirement is PHP 8.1 or newer.

Does QueryList handle JavaScript-rendered pages?

Not in the core package. The README lists crawling JavaScript dynamic rendering pages as something you implement through plugins, naming PhantomJS and headless WebKit, and it does not document the plugin configuration.

What licence does QueryList use?

The repository metadata does not state a licence, and the README does not name one. That leaves the terms undetermined from the project's own documentation.

Official sources

  1. Issues
  2. jae-jae/QueryList on GitHub
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/jae-jae-querylist.svg)](https://hysenlabs.com/projects/jae-jae-querylist)