# AutoCrawler: three entry points, one undocumented flag, and a v2.0.0 with no release

> A Selenium based image crawler for Google and Naver with multiprocess downloads and a keyword file at the repository root. Its own argument help contradicts its own full-resolution section, and the macOS fix it recommends turns off a system security check.

**YoongiKim/AutoCrawler** — Google, Naver multiprocess image web crawler (Selenium)

- Repository: https://github.com/YoongiKim/AutoCrawler
- Stars: 1,694 · Forks: 424
- Language: Python
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/yoongikim-autocrawler

## Three entry points, and a flag in the usage line with no entry

The how-to-use section is five steps: install Chrome, run the editable install, put keywords in a file, run the program, collect output. The run step names three different ways to start the same thing, the console command, the module form, and `python main.py`, the last of which is labelled as being kept for backwards compatibility.

```bash
pip install -e .
autocrawler [--skip] [--threads 4] [--google] [--naver] [--full] [--face] [--no_gui auto] [--limit 0]
```

The keyword list is a plain keywords.txt file that you edit in the repository, one keyword per line, and output lands in a download directory. The usage line above lists a `--limit` flag with a default of 0, but the argument reference that follows never describes it. The same block also ends mid-word on the no-GUI entry, whose description stops at the letters Accele. So the one number a user would most want to set, a cap on how much to download, is present in the summary and absent from the explanation.

## The full resolution section still documents the pre-2.0 syntax

There is a version note and then a contradiction. The note says that as of v2.0 the boolean flags take a flag and no-flag pair rather than a flag with true or false, and that no_gui is the exception. The argument reference agrees, listing full and no-full with the default set to no-full. Then a later section, headed full resolution mode, tells you to get full size JPG, GIF and PNG files by specifying `--full true`. That is the old calling convention, in a package whose own note says it no longer accepts it, so the one section that explains a performance trade-off is written against syntax the parser has moved past. The trade-off itself is stated clearly elsewhere: full resolution instead of thumbnails is marked slow. Which flag combination actually produces full size images has to be worked out from the argument reference rather than from the section that explains the feature.

## Five boolean pairs and one flag that still takes a value

The argument reference is explicit about defaults, which is the part worth reading carefully. Skip is on by default and means a keyword is skipped when its download directory already exists, which the text ties to re-downloading, so a second run over the same keywords is a no-op unless the flag is turned off. Both search engines are on by default, so a plain run queries Google and Naver at once. Full resolution is off by default because it is slow. Face search is off. Only no_gui is left as a three-value flag, auto, true or false, for headless mode, and the note confirms it was deliberately not converted to a pair. Threads defaults to 4 and is counted in worker processes rather than threads, which is a distinction that matters on a machine where the default process count already competes with Chrome for memory.

## Imbalance detection compares each keyword against half the average

Crawling different keywords through the same interface produces wildly different result counts, and the project has a small feature for noticing when that happens. When the crawl finishes, it compares the number of files per download directory against the average and reports any directory holding under 50 percent of it. The threshold is fixed at half rather than configurable, and the suggested remedy is to delete those directories and run again. That advice is sound for a search engine that rate-limits or truncates results, and it also means a keyword that legitimately returns fewer images will be reported every time and deleted each time. Combined with the default skip behaviour, a removed directory is also the only way to make the crawler revisit a keyword, since skip is keyed on the directory existing.

## The CAPTCHA path hands a real window to a person before workers start

Google's bot detection is described as IP reputation based rather than per-browser, so even testing bursts can get an address flagged for a while, and a fresh browser with valid cookies can still be asked again. What the tool does about it is worth being precise about, because it is not a bypass. Before any worker process starts, the main process opens one visible Chrome window against Google using a persistent profile directory at ./chrome-profile. If that window shows a challenge, the terminal prints the URL and asks the operator to solve it in the window and press Enter. The detector polls for a few seconds before giving up, because the widget can take a moment to render, which is why a challenge sometimes seems to be skipped. After Enter, the cookies are saved into that profile, and each crawl task gets its own temporary copy of it, since Chrome will not let two processes share one profile directory.

## The macOS fix includes turning off a system security check

A whole section of the page is about chromedriver binaries that hang or become unkillable on macOS, which the page attributes to Gatekeeper rejecting ad-hoc signed binaries without a Team Identifier. The diagnosis command is short.

```bash
spctl -a -vv "$(which chromedriver)"
```

If the answer is rejected, the page explains that on macOS 15 and later the per-file exception command no longer works, returning This operation is no longer supported, and that the remaining override is `sudo spctl --global-disable`. That is a machine wide switch which exposes an allow apps downloaded from anywhere toggle in System Settings. The page does offer a lighter path first: running chromedriver once directly in an interactive terminal, not through a script, has been enough to get it trusted. One more detail explains recurring failures: the driver builder prefers a chromedriver already on PATH over downloading one through webdriver-manager, and each distinct binary needs its own approval, so a copy sitting under the webdriver-manager cache directory is a different file from the one on PATH.

## Xvfb with access control disabled is how the server recipe works

Running the crawler on a remote machine takes three commands, one of which is a security decision. The recipe installs a virtual frame buffer and screen, starts a screen session, and launches the crawler against display 99, with the frame buffer backgrounded and the display set inline.

```bash
screen -S s1
Xvfb :99 -ac & DISPLAY=:99 autocrawler
```

The -ac flag disables access control on that X display, which is what lets a non-interactive session draw into it, and the result is a display anyone who can reach the machine can connect to. The same section is candid about the maintenance model. Selectors live in three files under the collectors directory, two named for the engines and one for shared helpers, and the page says you may need to fix the XPath expressions yourself because the sites change consistently. It gives a browser devtools workflow for finding them, points at an XPath reference on w3schools, notes you can test expressions with the find function in devtools, and records that as of 2026-08 all four collector functions were re-verified against the live DOM with several selector and timing bugs fixed.

## A 2.0.0 package with no releases and a docs directory the page never mentions

The packaging file is where the durable facts are. The project is named autocrawler at version 2.0.0, requires Python 3.11 or newer, and depends on three packages: requests at 2.31 or newer, selenium at 4.20 or newer, and webdriver-manager at 4.0 or newer, with pytest as the only development extra. The source lives under src, which is why an editable install is the documented first step, and the console entry point points at a cli main function inside the package, which is the origin of the autocrawler command. The license is declared as a file reference to Apache-2.0. The repository itself has no releases at all, so a version number in the manifest is the only version number, and the last push to the default branch, master, was on 2026-08-30. A docs directory and a tests directory both exist at the root, and neither is mentioned anywhere in the instructions.

## Conclusion

AutoCrawler suits someone who needs a folder per keyword of images from Google or Naver image search and is willing to maintain the selectors when the sites change, which the project openly expects. Three things to know before a long run. The `--full` example in the readme uses pre-2.0 syntax and will not work as written, use `--full` on its own. On macOS the documented workaround for an untrusted chromedriver includes `sudo spctl --global-disable`, which is a machine wide weakening of a security check, so try running chromedriver once in an interactive terminal first. And the crawl is bound to your IP reputation on Google, so a burst of testing is enough to trigger a CAPTCHA that only a person can clear.

## FAQ

### How do I install and run AutoCrawler?

Install Chrome, run pip install -e . to install the autocrawler package and its dependencies, write your search keywords into keywords.txt, then run autocrawler, python -m autocrawler, or python main.py. Files are written into the download directory.

### What does the --skip flag do in AutoCrawler?

It skips a keyword when its download directory already exists, and it is on by default. The readme ties it to re-downloading, so a second run over the same keywords does nothing unless you turn the flag off.

### How does AutoCrawler deal with a Google CAPTCHA?

Before any worker processes start, it opens one visible Chrome window against Google with a persistent profile. If a challenge appears you solve it in that window and press Enter, the cookies are saved into ./chrome-profile, and each crawl task then gets its own copy of that profile.

### What does AutoCrawler require in terms of Python and dependencies?

Python 3.11 or newer, with requests 2.31 or newer, selenium 4.20 or newer and webdriver-manager 4.0 or newer as runtime dependencies. pytest is the only development extra.

### How do I fix AutoCrawler when Google or Naver change their pages?

You edit the XPath selectors yourself in the collector files under src/autocrawler/collectors, using Chrome developer tools to capture the target and test expressions. As of 2026-08 all four collector functions had been re-verified against the live DOM, and a NOTE comment at the top of each records what changed.

## Sources

- [Issues](https://github.com/YoongiKim/AutoCrawler/issues)
- [License: Apache-2.0](https://github.com/YoongiKim/AutoCrawler/blob/master/LICENSE)
- [README](https://github.com/YoongiKim/AutoCrawler/blob/master/README.md)
- [YoongiKim/AutoCrawler on GitHub](https://github.com/YoongiKim/AutoCrawler)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/yoongikim-autocrawler
